An infrared and visible light image fusion method, device and electronic device
By extracting multi-level features of infrared and visible light images and fusion using modal-guided cross-attention blocks, the problems of data loss and large amount of calculation during image fusion in the prior art are solved, and efficient and accurate image fusion effect is achieved.
Patent Information
- Application Number
- CN202410570755.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-05-09
AI Technical Summary
The prior art has problems of data loss and large calculations when fusing infrared and visible light images, and has failed to fully utilize the potential of Transformer.
A method of fusion between infrared and visible light images is proposed. By acquiring registered infrared images and visible light images, feature information at each level is extracted, and the shallow and deep features are directly fused through modal-guided cross-attention blocks to reduce information loss and reduce calculation amount.
It realizes efficient and accurate image fusion, retains infrared texture details information and visible background, reduces the amount of calculation, and avoids information loss during feature extraction.
Smart Images

Figure CN118587105B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computational photography science and technology, and particularly relates to an infrared and visible light image fusion method, device, electronic device, and storage medium. Background Art
[0002] Limited by the theoretical and technological development of hardware devices, images captured by a single sensor or a single shooting setup cannot effectively and comprehensively depict the imaging scene. We hope to organically combine multiple images captured under different sensors or different shooting setups, while extracting and retaining their respective useful information, suppressing or removing redundant and noise information, so as to obtain a new image that is more suitable for human or machine analysis and understanding, and is convenient for subsequent use in high-level vision tasks. Therefore, the image fusion technology was born. Image fusion technology includes infrared and visible light image fusion, medical image fusion, multi-exposure image fusion, multi-focus image fusion, and multi-spectral and panchromatic image fusion. Among them, infrared and visible light image fusion technology has been introduced into various fields and has become one of the most widely used image fusion technologies at present.
[0003] The existing method for obtaining a fused image based on deep learning is a simple stacking of Transformers and convolutional blocks. After generating the fused features, the missing feature information in the source images is no longer integrated again, resulting in information loss. Secondly, the existing Transformer-based fusion methods do not fully utilize the potential of Transformers, and at the same time increase the computational amount when introducing Transformers.
[0004] In summary, how to design an efficient and accurate infrared and visible light image fusion method is an urgent problem to be solved at present. Summary of the Invention
[0005] The present application aims to solve at least one of the technical problems in the related art to some extent.
[0006] To this end, the first object of the present application is to propose an infrared and visible light image fusion method to solve problems such as data loss and large computational amount during image fusion by existing technical means.
[0007] The second object of the present application is to propose a device.
[0008] The third object of the present application is to propose an electronic device.
[0009] The fourth object of the present application is to propose a computer-readable storage medium.
[0010] To achieve the above object, the first aspect embodiment of the present application proposes an infrared and visible light image fusion method, including:
[0011] Obtain the registered infrared image and visible light image;
[0012] Extract the feature information of the infrared image and the visible light image to obtain the feature information of each level;
[0013] Perform preliminary fusion on the features of the deepest layer in the feature information of each level to obtain preliminary fusion features;
[0014] Fuse the preliminary fusion features with the shallow features in the feature information of each level to obtain the final fusion features;
[0015] Perform image reconstruction based on the final fusion features to obtain the final fusion image.
[0016] Preferably, the extracting the feature information of the infrared image and the visible light image to obtain the feature information of each level includes:
[0017] Extract shallow features from the infrared image and the visible light image by using two shared convolutional layers;
[0018] Use an efficient long-range attention block to perform deep feature extraction on the shallow features, divide the extracted deep features into multiple groups, calculate cross-attention for each group using different windows, splice and integrate the cross-attention, and then perform normalization processing to obtain the feature information of each level.
[0019] Preferably, the feature extraction calculation formula is:
[0020]
[0021] Where, respectively represent the feature information of each level extracted from the infrared image and the visible light image, I ir is the registered infrared image, I vi is the registered visible light image.
[0022] Preferably, the performing preliminary fusion on the features of the deepest layer in the feature information of each level includes:
[0023] Splice the features of the deepest layer in the feature information of each level, and use two convolutional layers to generate a preliminary fusion feature, where the fusion process formula is:
[0024]
[0025] Where, is the preliminary fusion feature, C(·) is the channel dimension splicing operation, is the feature of the deepest layer.
[0026] Preferably, the step of fusing the preliminary fusion feature with the shallow features in the hierarchical feature information to obtain the final fusion feature includes:
[0027] Using the cross-attention mechanism to directly integrate the preliminary fusion feature and the shallow features in the hierarchical feature information into the fusion feature. First, local feature extraction is performed, and then the grouped multi-size window attention calculation strategy is used to perform cross-attention calculation to obtain the final fusion feature.
[0028] Preferably, the cross-attention calculation formula is:
[0029]
[0030] where MGCAB i is the cross-attention block guided by the i-th modality, is the input feature.
[0031] Preferably, the step of performing image reconstruction based on the final fusion feature to obtain the final fusion image includes:
[0032] Using six efficient cross-attention blocks to restore the shallow feature information. The two efficient cross-attention blocks in the middle position use the moving window strategy to calculate the window attention, and LeakyReLU is used as the activation function to generate the final fusion image.
[0033] To achieve the above object, an infrared and visible light image fusion device is proposed in the second aspect embodiment of the present application, including:
[0034] An image acquisition module for acquiring registered infrared images and visible light images;
[0035] A feature extraction module for extracting the feature information of the infrared image and the visible light image to obtain hierarchical feature information;
[0036] A preliminary fusion module for preliminarily fusing the deepest features in the hierarchical feature information to obtain a preliminary fusion feature;
[0037] A final fusion module for fusing the preliminary fusion feature with the shallow features in the hierarchical feature information to obtain a final fusion feature;
[0038] An image reconstruction module for performing image reconstruction based on the final fusion feature to obtain a final fusion image.
[0039] To achieve the above object, an electronic device is proposed in the third aspect embodiment of the present application, including: a processor, and a memory communicatively connected to the processor;
[0040] The memory stores computer-executable instructions;
[0041] The processor executes the computer-executable instructions stored in the memory to implement the method described in any one of the above.
[0042] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a computer-readable storage medium, including computer-executable instructions stored in the computer-readable storage medium, and the computer-executable instructions are used to implement the method described in any one of the above when executed by a processor.
[0043] An infrared and visible light image fusion method provided by the present application utilizes the features of infrared and visible light images, generates fusion features through multiple convolutional layers, directly fuses the extracted shallow and deep features into fusion features through a modality-guided cross-attention block, enhances the fusion features, and at the same time avoids information loss caused by the continuous deepening of the feature extraction process. By introducing a shared attention, shifted convolution, and grouped multi-size window attention calculation mechanism, a good fusion effect is achieved while reducing the computational complexity.
[0044] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings
[0045] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0046] Figure 1 is a flowchart of the first specific embodiment of an infrared and visible light image fusion method provided by the present invention;
[0047] Figure 2 is a general framework flowchart of a modality-guided Transformer;
[0048] Figure 3 is a network structure diagram of an infrared and visible light image fusion method based on deep learning;
[0049] Figure 4 is a structural block diagram of an infrared and visible light image fusion device provided by an embodiment of the present invention. Detailed Embodiments
[0050] The core of the present invention is to provide an infrared and visible light image fusion method, device, electronic device, and storage medium, which guide the process of information integration through deep interaction between different modality features and reduce information loss during the fusion process.
[0051] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0052] Please refer to Figure 1 , Figure 1 which is a flowchart of the first specific embodiment of an infrared and visible light image fusion method provided by the present invention; the specific operation steps are as follows:
[0053] Step S101: Obtain registered infrared images and visible light images;
[0054] Step S102: Extract the feature information of the infrared image and the visible light image to obtain feature information at each level;
[0055] Extract shallow features from the infrared image and the visible light image by using two shared convolutional layers;
[0056] Use an efficient long-range attention block to perform deep feature extraction on the shallow features, divide the extracted deep features into multiple groups, calculate cross-attention for each group using different windows, splice and integrate the cross-attention, and then perform normalization processing to obtain feature information at each level.
[0057] The feature extraction calculation formula is:
[0058]
[0059] where, respectively represent the feature information at each level extracted from the infrared image and the visible light image, I ir is the registered infrared image, I vi is the registered visible light image.
[0060] Step S103: Perform preliminary fusion on the deepest layer features in the feature information at each level to obtain preliminary fusion features;
[0061] Splice the deepest layer features in the feature information at each level, and use two convolutional layers to generate a preliminary fusion feature, where the fusion process formula is:
[0062]
[0063] where, is the preliminary fusion feature, C(·) is the channel dimension splicing operation, is the deepest layer feature.
[0064] Step S104: Fuse the preliminary fusion feature with the shallow features in the hierarchical feature information to obtain the final fusion feature;
[0065] Utilize the cross-attention mechanism to directly integrate the preliminary fusion feature and the shallow features in the hierarchical feature information into the fusion feature. First, perform local feature extraction, and then adopt the grouped multi-size window attention calculation strategy to perform cross-attention calculation to obtain the final fusion feature.
[0066] The cross-attention calculation formula is:
[0067]
[0068] where MGCAB i is the cross-attention block guided by the i-th modality, is the input feature.
[0069] Step S105: Based on the final fusion feature, perform image reconstruction to obtain the final fusion image.
[0070] Utilize six efficient cross-attention blocks to restore the shallow feature information. The two efficient cross-attention blocks in the middle position adopt the moving window strategy to calculate the window attention, and use LeakyReLU as the activation function to generate the final fusion image.
[0071] This embodiment provides an infrared and visible light image fusion method. By using the features of infrared and visible light images obtained by the feature extraction module, preliminary fusion features are generated through multiple convolutional layers. The extracted shallow and deep features are directly integrated into the fusion feature through the modality-guided cross-attention block, which enhances the fusion feature while avoiding information loss caused by the continuous deepening of the feature extraction process. The fusion image output by the modality-guided Transformer can retain the visible light background and infrared key targets while fully preserving the infrared texture detail information. By introducing the shared attention, shifted convolution, and grouped multi-size window attention calculation mechanisms, a good fusion effect is achieved while reducing the computational complexity.
[0072] Based on the above embodiment, this embodiment describes the infrared and visible light image fusion method. As Figure 2 shown, specifically as follows:
[0073] The overall framework of the modality-guided Transformer proposed in this embodiment. This network mainly consists of three parts: a feature extraction module, a modality-guided fusion module, and an image reconstruction module;
[0074] The feature extraction module performs shallow local feature extraction and deep global feature extraction through convolutional layers and efficient long-range attention blocks respectively. In order to retain both the rich detailed texture information in the shallow features and the key target information in the deep features, the feature extraction module will retain the feature maps output by the efficient long-range attention blocks at different depths and use them jointly as the input to the feature fusion module.
[0075] The modality-guided feature fusion module integrates features at different levels obtained from the feature extraction module into the fused features. This process uses a carefully designed modality-guided cross-attention block, which consists of an efficient long-range attention block and an efficient long-range cross-attention block. The modality-guided cross-attention block integrates the information of infrared features and visible light features into the fused features by calculating the cross-attention between infrared features and visible light features and the fused features respectively under the guidance of modality information.
[0076] Finally, in the image reconstruction module, a series of efficient long-range attention blocks are used for deep feature recovery, and then convolutional layers are used for the reconstruction of the fused image.
[0077] Specifically as follows:
[0078] The input of a network for an infrared and visible light image fusion method is a registered infrared image I ir ∈R H×W×1 and a visible light image I vi ∈R H×W×3 , and the output is a fused image I f ∈R H×W×3 generated by the feature extraction module, the modality-guided fusion module, and the image reconstruction module. The specific processing process includes the following steps:
[0079] First, through the feature extraction module E F shallow features with richer detailed information and deep features with more prominent key target information are obtained from infrared and visible light images. The feature extraction module includes two convolutional layers and twelve efficient long-range attention blocks for feature extraction. Specifically, first, two shared convolutional layers are used to extract shallow features from infrared and visible light images. The convolutional kernel sizes of these two convolutional layers are set to 3×3, the stride is 1, and the output channels are set to 30 and 60 respectively. The extracted shallow features are input into the efficient long-range attention blocks, and six efficient long-range attention blocks are used to perform deep feature extraction on the infrared image and the visible light image respectively. As Figure 3As shown in the figure, the efficient long-range attention block first performs local feature extraction, which includes two shifted convolutions and a ReLU activation layer. The features strengthened by local feature extraction are evenly divided into Z groups, and cross attention is calculated on each group with different window sizes. The cross attention calculated by these groups is then concatenated and integrated through 1×1 convolution, and batch normalization is applied for processing. The six efficient long-range attention blocks are further divided into three groups, where the first efficient long-range attention block in each group shares the calculated attention score with the subsequent efficient long-range attention blocks. When calculating attention, the second group uses a cyclic window shift strategy to expand the perception range of the feature extraction module and deeply mine the long-range dependencies between pixels. The features output by the second efficient long-range attention block in each group are retained as the input of the modality-guided fusion module.
[0080] The extracted features at each level are input into a carefully designed modality-guided fusion module, which includes two convolutional layers and three modality-guided cross-attention blocks. Each modality-guided cross-attention block contains six efficient attention blocks and two efficient cross-attention blocks. In the modality-guided fusion module, the deepest features are first and Splicing is used as input and a preliminary fusion feature is generated through two convolutional layers Conv In order to eliminate the information loss caused by the deepening network structure in the feature extraction stage, this embodiment proposes a modality-guided cross-attention block, which uses the cross-attention mechanism to directly integrate the extracted shallow and deep features into the fused features. The input of the efficient cross-attention block is infrared features, visible light features, and fused features. After receiving the input, similar to the efficient attention block, the efficient cross-attention block first performs local feature extraction, and then performs cross-attention calculation, and also adopts the grouped multi-size window attention calculation strategy. The second efficient cross-attention block obtains the attention score from the previous efficient cross-attention block without calculating it again. The one in the middle of the three modality-guided cross-attention blocks adopts a moving window strategy when calculating the window attention. For the i-th modality-guided cross-attention block MGCAB i (i={1,2,3}), given input features It can be output It is expressed as:
[0081]
[0082] Finally, the fusion module integrates the infrared and visible light features to generate fused features as the input of the image reconstruction module. Output of the last modality-guided cross-attention block The final fusion feature F is the output of the fusion module f, and is input into the subsequent image reconstruction module.
[0083] The image reconstruction module consists of six efficient cross-attention blocks and three convolutional layers. The six efficient cross-attention blocks help to recover the fused shallow feature information, enhance the deep feature information, and ultimately improve the quality of the fused image. The two efficient cross-attention blocks in the middle position adopt a moving window strategy when calculating the window attention. The three 3×3 convolutional layers generate the fused image. Their strides are set to 1, and the output channels are set to 30, 15, and 1 in sequence. Except for the last layer, all convolutional layers use LeakyReLU as the activation function. The fused features integrated by the modality-guided fusion module are input into the image reconstruction module RI to generate the final fused image, and this process can be expressed as:
[0084] I f = R I (F f )
[0085] An infrared and visible light image fusion method provided by an embodiment of the present invention uses an infrared and visible light image fusion framework combining a convolutional neural network and a Transformer to fuse infrared images and visible light images. It uses convolutional layers to extract shallow features, reduce the modality differences between infrared and visible light images, and then uses the Transformer structural components to extract multi-level image features, retaining rich detailed information in the shallow features and prominent significant target information in the deep features. Through the deep interaction between different modality features, it guides the process of information integration, reduces information loss during the fusion process, and directly integrates the extracted shallow and deep features into the fused features through the modality-guided cross-attention blocks, enhancing the fused features while avoiding information loss caused by the continuous deepening of the feature extraction process.
[0086] Please refer to Figure 4 , Figure 4 which is the structural block diagram of an infrared and visible light image fusion device provided by an embodiment of the present invention; the specific device may include:
[0087] An image acquisition module 100 that acquires registered infrared images and visible light images;
[0088] A feature extraction module 200 that extracts the feature information of the infrared image and the visible light image to obtain multi-level feature information;
[0089] A preliminary fusion module 300 that preliminarily fuses the deepest features in the multi-level feature information to obtain preliminary fusion features;
[0090] A final fusion module 400 that fuses the preliminary fusion features with the shallow features in the multi-level feature information to obtain final fusion features;
[0091] The image reconstruction module 500 performs image reconstruction based on the final fused features to obtain a final fused image.
[0092] An infrared and visible light image fusion device according to this embodiment is used to implement the foregoing infrared and visible light image fusion method. Therefore, the specific implementation manners in an infrared and visible light image fusion device can be seen in the embodiment part of the foregoing infrared and visible light image fusion method. For example, the image acquisition module 100, the feature extraction module 200, the preliminary fusion module 300, the final fusion module 400, and the image reconstruction module 500 are respectively used to implement steps S101, S102, S103, S104, and S105 in the foregoing infrared and visible light image fusion method. Therefore, its specific implementation manners can refer to the descriptions of the corresponding various part embodiments and will not be elaborated herein.
[0093] To implement the above embodiments, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0094] To implement the above embodiments, the present application also proposes a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are used to implement the method provided in the foregoing embodiments when executed by a processor.
[0095] To implement the above embodiments, the present application also proposes a computer program product including a computer program, and the computer program implements the method provided in the foregoing embodiments when executed by a processor.
[0096] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0097] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and signing an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to safeguard and protect access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0098] This application is expected to provide an implementation scheme for users to selectively prevent the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.
[0099] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0100] Furthermore, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0101] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0102] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0103] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0104] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0105] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0106] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for fusing infrared and visible light images, characterized in that: include: Acquire registered infrared images and visible light images; Extracting feature information of the infrared image and the visible light image to obtain feature information of each level; Preliminarily fusing the deepest features of the feature information at each level to obtain preliminary fused features; Fusing the preliminary fusion features with the shallow features in the feature information of each level to obtain the final fusion features; Perform image reconstruction based on the final fusion feature to obtain a final fusion image; The extracting feature information of the infrared image and the visible light image to obtain feature information at each level includes: Extracting shallow features from the infrared image and the visible light image using two shared convolutional layers; The shallow features are used to extract deep features using efficient long-range attention blocks, and the extracted deep features are divided into multiple groups. Each group uses a different window to calculate cross attention. The cross attentions are spliced and integrated and then normalized to obtain feature information at each level.
2. The infrared and visible light image fusion method according to claim 1, characterized in that: The feature extraction calculation formula is: in, Represent the features of each level extracted from infrared images and visible light images, respectively. ir is the registered infrared image, I vi is the registered visible light image.
3. The infrared and visible light image fusion method according to claim 1, characterized in that: The preliminary fusion of the deepest features of the feature information at each level includes: The deepest features of each level of feature information are concatenated, and a preliminary fusion feature is generated using two convolutional layers, where the fusion process formula is: in, is the initial fusion feature, C(·) is the channel dimension concatenation operation, The deepest feature.
4. The infrared and visible light image fusion method according to claim 1, characterized in that: The fusing of the preliminary fusion features with the shallow features in the feature information of each level to obtain the final fusion features comprises: The preliminary fusion features and the shallow features in the feature information of each level are directly integrated into the fusion features by using the cross-attention mechanism. First, local features are extracted, and then a grouped multi-size window attention calculation strategy is used to perform cross-attention calculation to obtain the final fusion features.
5. The infrared and visible light image fusion method according to claim 4, characterized in that: The cross attention calculation formula is: Among them, MGCAB i is the cross attention block guided by the i-th modality, is the input feature.
6. The infrared and visible light image fusion method according to claim 1, characterized in that: The performing image reconstruction based on the final fusion feature to obtain the final fusion image comprises: Six efficient cross-attention blocks are used to restore shallow feature information. The two efficient cross-attention blocks in the middle position use a moving window strategy to calculate the window attention, and LeakyReLU is used as the activation function to generate the final fused image.
7. An infrared and visible light image fusion device, characterized in that: include: An image acquisition module, which acquires the registered infrared image and visible light image; A feature extraction module extracts feature information of the infrared image and the visible light image to obtain feature information at each level, wherein shallow features are extracted from the infrared image and the visible light image using two shared convolutional layers, deep features are extracted from the shallow features using an efficient remote attention block, the extracted deep features are divided into multiple groups, each group uses a different window to calculate cross attention, the cross attentions are spliced and integrated, and then normalized to obtain feature information at each level; A preliminary fusion module, which performs preliminary fusion of the deepest features of the feature information at each level to obtain preliminary fusion features; The final fusion module fuses the preliminary fusion features with the shallow features in the feature information of each level to obtain the final fusion features; The image reconstruction module performs image reconstruction based on the final fusion feature to obtain a final fusion image.
8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Infrared image and visible light image fusion method, system, device and terminal
CN114140366A
Cited By
Infrared visible light image fusion method based on multi-branch extraction and cross-modal propagation
CN122368710A