Infrared and visible light image fusion method based on frequency semantic compensation cooperation
By adopting a frequency semantic compensation collaboration method in the fusion of infrared and visible light images, combining the semantic and frequency information of the image, the problems of poor image fusion effect and loss of feature information in the prior art are solved, and high-quality fusion images rich in features are achieved.
Patent Information
- Application Number
- CN202510007749.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
Smart Images

Figure CN119941525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of infrared and visible light image fusion, and in particular to an infrared and visible light image fusion method based on frequency semantic compensation collaboration. Background Art
[0002] The infrared and visible light image fusion technology is mainly used to solve the problem of limited information expression of single-modal images, while improving the richness and availability of image information. Due to the limitations of imaging equipment, infrared images can capture the thermal radiation emitted by objects and effectively highlight targets, but lack texture details. In contrast, visible light images contain rich texture details, but it is difficult to distinguish targets under extreme conditions. Therefore, fusing infrared images with visible light images with complementary features can obtain an image rich in information, providing more accurate data for subsequent advanced computer vision tasks such as military surveillance, target tracking, and prominent target detection.
[0003] Existing image fusion methods are mainly divided into traditional image fusion methods and image fusion methods based on deep learning. Although traditional methods have achieved good fusion performance to a certain extent, they still have some limitations and challenges, such as manual feature extraction leading to complex calculations and high requirements for image registration of different modalities. With the great achievements of deep learning in the field of computer science, it has been widely used in image fusion to obtain a more complete image representation because it can accurately capture the characteristics of multimodal data and the complementarity between modalities. Although the existing deep learning fusion methods have achieved relatively good results, there are still some problems that need to be solved. First, in the process of image feature extraction and fusion, the huge network model may lead to the loss of basic information of the source image; second, it is difficult for existing methods to effectively combine the semantic information and frequency information of the image; finally, in multimodal fusion, the high and low frequency information of the image is easily lost, resulting in incomplete fused images or biased towards one modality. Summary of the invention
[0004] In view of this, the present invention provides an infrared and visible light image fusion method based on frequency semantic compensation collaboration, which can effectively combine the two types of information to generate a fused image rich in frequency and semantic features, and solve the problems of initial feature information loss during the image fusion process and excessive bias of the fused image towards a certain mode.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] The present invention provides an infrared and visible light image fusion method based on frequency semantic compensation collaboration, comprising the following steps:
[0007] Step 1: Obtain the data set of the fusion network and divide the data set into a training set and a test set;
[0008] Step 2: construct an infrared and visible light image fusion network for feature extraction and image restoration, the fusion network includes: a hierarchical semantic feature extraction unit, a frequency and semantic feature cross compensation module, and an image upsampling reconstruction module;
[0009] Step 3: Preprocess the images in the training set and normalize them;
[0010] Step 4: Train the fusion network under the guidance of the loss function to obtain a trained fusion network model;
[0011] Step 5: Use the test set data of the fusion network to test the trained network model obtained in step 4 to obtain a fused image.
[0012] Preferably, in step 1, the training set of the fusion network is from the RoadScene dataset, M 3 FD dataset and MSRS dataset, specifically, all 221 image pairs of RoadScene dataset, M 3 The FD dataset contains 200 random image pairs and the MSRS dataset contains 1,179 random image pairs, which are composed of 1,600 image pairs from different times and locations. The test set mainly consists of 25 image pairs from the TNO dataset, 50 image pairs from the MSRS dataset, and the M 3 The FD dataset consists of 100 image pairs.
[0013] Preferably, in step 2, the hierarchical semantic feature extraction unit is composed of two branches, each branch contains three semantic coherence enhancement modules with consistent structures, and each hierarchical semantic feature extraction unit integrates a semantic space attention module to extract the semantic feature information of the infrared image and the visible light image respectively, and connects the semantic features of the two modalities by channel to generate a multimodal semantic feature layer L msf .
[0014] Preferably, in step 2, the frequency and semantic feature cross-compensation module adopts a cross-compensation mechanism, which aims to carefully preserve the semantic information in the input features and accurately extract and fuse the frequency information; first, the multimodal semantic feature layer L msf Through a semantic space attention module and a convolution block, a fused semantic feature layer L is generated. fsf , the convolution block consists of a 5×5 convolution layer and a LeakyReLU activation layer; generating a fusion semantic feature layer L fsfAfter that, it passes through three semantic space attention modules and convolution blocks in turn to protect the semantic information and prevent it from being lost during the processing. The convolution block consists of a 3×3 convolution layer and a LeakyReLU activation layer. Through the convolution operation of different channels, three groups of Sobel operators and two-dimensional wavelet transform work together to obtain the fusion semantic feature layer L rich in semantic information. fsf Extract the frequency information; extract the three sets of low-frequency information L i , i∈1,2,3 is compensated to the fusion semantic feature layer L in the form of residual fsf , generate the low-frequency semantic fusion feature layer L lfsff , three groups of Sobel high frequency information SH i , i∈1, 2, 3 are passed to the image upsampling and reconstruction module for subsequent image reconstruction.
[0015] Preferably, in step 2, the image upsampling reconstruction module first extracts the Sobel high-frequency information SH from three groups of different channel numbers. i , i∈1,2,3, processed by a 2×2 upsampling transposed convolution layer to ensure fusion with the low-frequency semantic feature layer L lfsff Secondly, the feature information of the two parts is fused by element-by-element addition, and then processed through five convolution blocks to obtain the final fused image; among these five convolution blocks, the first four are composed of a 3×3 convolution layer and a LeakyReLU activation layer, and the last convolution block is composed of a 3×3 convolution layer and a Tanh activation layer.
[0016] Preferably, the semantic space attention module is composed of three feature streams, the upper feature streams are processed by Sobel gradient operator and accompanied by a 1×1 convolution layer; the middle feature stream contains the initial input features; the lower feature stream is the normalized calculation of the deviation between the initial input features and their mean; by adding the elements of these three feature streams one by one, and processing them through another 1×1 convolution layer, and finally adding them to the initial image features, accurate semantic space information F is extracted from the initial input features. si .
[0017] Preferably, the six semantic coherence enhancement modules of the dual branches are internally integrated with a semantic space attention module, and firstly pass through a convolution block and a semantic space attention module to extract accurate semantic space information F si, the convolution block consists of a 3×3 convolution layer and a LeakyReLU activation layer; then through the three-branch feature flow, the upper branch consists of two 5×5 convolution layers, a LeakyReLU activation layer and a Sigmoid activation layer, the lower branch consists of two 3×3 convolution layers, a LeakyReLU activation layer and a Sigmoid activation layer, and the middle branch keeps the semantic space information F si The upper and lower branches are respectively multiplied element-by-element with the middle branch and then added to the middle branch; finally, they are added to the semantic space information F after two 3×3 convolutional layers and two LeakyReLU activation layers. si The rich semantic features F are obtained by adding the residuals sf .
[0018] Preferably, in step 3, the 1600 pairs of infrared images and visible light images registered at different times and locations are segmented into 128*128 patch images, and the obtained patch images are used to construct a training set for the fusion network. During the network training process, a Batch Normalization layer is added before the activation function layer of each module to accelerate the training process of the fusion network and improve the performance of the model.
[0019] Preferably, in step 4, the fusion network is trained under the guidance of the loss function to obtain a trained fusion network model. The loss function The calculation formula is shown in formula (1):
[0020]
[0021] In formula (1), is the semantic-based pixel loss, is the texture loss of details, α and β are hyper parameters;
[0022] In formula (1), the semantic-based pixel loss The calculation formula is shown in formula (2):
[0023]
[0024]
[0025] In formula (2), ||||1 represents the L1 norm, F is the fused image, VI is the visible light image, and IR is the infrared image;
[0026] In formula (3), It represents the pixel value threshold result of the N image, which is 1 when it is greater than or equal to 0.5, otherwise it is 0. N(x, y) represents the pixel value of the original image at position (x, y);
[0027] In formula (1), the detail-based texture loss The calculation formula is shown in formula (4):
[0028]
[0029]
[0030] In formula (4), H and W represent the height and width of the fused image respectively. It means that the Sobel operator calculates the gradient information of N images, and ||||2 means the L2 norm;
[0031] In formula (5), G x and G y They are the horizontal and vertical convolution kernels of the Sobel operator, respectively. The horizontal and vertical convolution kernels of the Sobel operator are usually defined as:
[0032]
[0033] * represents the convolution operation, and N is the original image.
[0034] Preferably, in step 5, the test set of the fusion network mainly consists of the TNO dataset, the MSRS dataset and the M 3 The FD data set is composed of the test set, and the data in the test set is input into the trained fusion network model obtained in step 4; during the test process, the tone characteristics of the fused image are consistent with the input visible light image, that is, when the input image is grayscale, the output fused image also appears in grayscale mode; and when the input image is color level, the output fused image also shows the corresponding color characteristics; this operation ensures the consistency of the fused image with the original visible light image in the color dimension.
[0035] Compared with the prior art, the infrared and visible light image fusion method based on frequency semantic compensation cooperation provided by the present invention has the following beneficial effects:
[0036] 1. The present invention aims to solve the problems that semantic information and frequency information are not fully utilized in the image fusion process, resulting in poor fusion image effect and that initial features are easily lost and difficult to retain when the features are processed using the network. Compared with traditional methods, the present invention has better performance in processing image fusion tasks in various environments or situations.
[0037] 2. When processing information of two different natures, semantics and frequency, the present invention adopts a hierarchical semantic feature extraction unit and a semantic space attention module therein to extract semantic information, uses a frequency and semantic feature cross-compensation module to extract frequency information, and retains the semantic information, fully considering the characteristics of the two types of information. Finally, an upsampling reconstruction module is used to obtain a fused image with good visual perception and rich texture details. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0039] Figure 1 is a flow chart of the steps of the present invention;
[0040] Figure 2 It is the overall model diagram of the present invention;
[0041] Figure 3 Schematic diagram of the structure of the semantic coherence enhancement module and the semantic space attention module of the present invention;
[0042] Figure 4 The calculation process of Sobel and two-dimensional convolution wavelet transform in the frequency and semantic feature cross compensation module of the present invention;
[0043] Figure 5 A comparison diagram of a set of infrared images, grayscale visible light images and fused images in a low-resolution scene of the present invention;
[0044] Figure 6 This is a comparison diagram of a set of infrared images, color visible light images and fused images in a high-resolution scene of the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0048] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.
[0049] like Figure 1 FIG. 1 shows the steps of a first embodiment of a method for fusion of infrared and visible light images based on frequency semantic compensation cooperation provided by the present invention. In this embodiment, the following steps are included:
[0050] Step 1: Get the dataset of the fusion network and divide the dataset into training set and test set:
[0051] The training set is selected from the public image fusion dataset RoadScene dataset, M 3 FD dataset and MSRS dataset, specifically, all 221 image pairs of RoadScene dataset, M 3 The FD dataset contains 200 random image pairs and the MSRS dataset contains 1,179 random image pairs, which are composed of 1,600 image pairs from different times and locations. The test set mainly consists of 25 image pairs from the TNO dataset, 50 image pairs from the MSRS dataset, and the M 3 The FD dataset consists of 100 image pairs.
[0052] Step 2: Construct an infrared and visible light image fusion network for feature extraction and image restoration. The fusion network includes: a hierarchical semantic feature extraction unit, a frequency and semantic feature cross compensation module, and an image upsampling reconstruction module. The overall network model of this embodiment is shown in FIG. Figure 2 shown.
[0053] The hierarchical semantic feature extraction unit consists of two branches, each of which contains three semantic coherence enhancement modules with the same structure. Each semantic coherence enhancement module integrates a semantic space attention module to extract the semantic feature information of infrared images and visible light images respectively, and connects the semantic features of the two modalities by channel to generate a multimodal semantic feature layer L msfThe semantic spatial attention module consists of three feature streams. The upper feature stream is processed by the Sobel gradient operator and accompanied by a 1×1 convolution layer. The middle feature stream contains the initial input features. The lower feature stream is the normalized calculation of the deviation between the input image features and their mean. The elements of these three feature streams are added one by one, processed by another 1×1 convolution layer, and finally added to the initial input features to extract accurate semantic spatial information F from the initial input features. si In this embodiment, the semantic coherence enhancement module of the semantic space attention module is integrated as follows: Figure 3 shown.
[0054] The frequency and semantic feature cross compensation module first converts the multimodal semantic feature layer L msf Through a semantic spatial attention module and a convolution block, the convolution block consists of a 5×5 convolution layer and a LeakyReLU activation layer to generate a fused semantic feature layer L fsf ; After the fusion semantic feature layer L fsf After that, it passes through three semantic space attention modules and convolution blocks in turn to protect the semantic information and prevent it from being lost during the processing. The convolution block consists of a 3×3 convolution layer and a LeakyReLU activation layer. Through the convolution operation of different channels, three groups of Sobel operators and two-dimensional wavelet transform work together to obtain the fusion semantic feature layer L rich in semantic information. fsf Extract the frequency information; extract the three sets of low-frequency information L i , i∈1,2,3 is compensated to the fusion semantic feature layer L in the form of residual fsf , generate the low-frequency semantic fusion feature layer L lfsff , three groups of Sobel high frequency information SH i , i∈1, 2, 3 are passed to the image upsampling reconstruction module for subsequent image reconstruction. The calculation process of Sobel and two-dimensional convolution wavelet transform in this embodiment is as follows: Figure 4 shown.
[0055] The image upsampling and reconstruction module first extracts the Sobel high-frequency information SH from three groups of different channel numbers. i , i∈1,2,3, processed by a 2×2 upsampling transposed convolution layer to ensure fusion with the low-frequency semantic feature layer L lfsff The size of the two parts is kept consistent. Secondly, the feature information of the two parts is fused by element-by-element addition, and then processed through five convolution blocks to obtain the final fused image. Among these five convolution blocks, the first four are composed of a 3×3 convolution layer and a LeakyReLU activation layer, while the last convolution block is composed of a 3×3 convolution layer and a Tanh activation layer.
[0056] Step 3: Preprocess the images in the training set and normalize them;
[0057] The 1,600 pairs of infrared images and visible light images registered at different times and locations were segmented into 128*128 patch images, and the obtained patch images were used to construct the training set of the fusion network. During the network training process, a batch normalization layer was added before the activation function layer of each module to accelerate the training process of the fusion network and improve the performance of the model.
[0058] Step 4: Train the fusion network under the guidance of the loss function to obtain a trained fusion network model; loss function The calculation formula is shown in formula (6):
[0059]
[0060] In formula (6), is the semantic-based pixel loss, is the texture loss of details, α and β are hyper parameters;
[0061] In formula (6), the semantic-based pixel loss The calculation formula is shown in formula (7):
[0062]
[0063]
[0064] In formula (7), ||||1 represents the L1 norm, F is the fused image, VI is the visible light image, and IR is the infrared image;
[0065] In formula (8), It represents the pixel value threshold result of the N image, which is 1 when it is greater than or equal to 0.5, otherwise it is 0. N(x, y) represents the pixel value of the original image at position (x, y);
[0066] In formula (6), the detail-based texture loss The calculation formula is shown in formula (9):
[0067]
[0068]
[0069] In formula (9), H and W represent the height and width of the fused image respectively. It means that the Sobel operator calculates the gradient information of N images, and ||||2 means the L2 norm;
[0070] In formula (10), Gx and G y They are the horizontal and vertical convolution kernels of the Sobel operator, respectively. The horizontal and vertical convolution kernels of the Sobel operator are usually defined as:
[0071]
[0072] * represents the convolution operation, and N is the original image.
[0073] The learning rate is set to 0.0005, the Adam optimizer is used to optimize the network, the epoch is set to 100, the batch size is set to 16, and the parameters α and β in the loss function are set to 1 and 4 respectively. This embodiment is implemented based on the pytorch framework, and all experiments are performed on NVIDIA RTX A4000 GPU and Intel XeonW-2145CPU.
[0074] Step 5: Use the test set data of the fusion network to test the trained network model obtained in step 4 to obtain a fused image. During the test, the hue characteristics of the fused image are consistent with the input visible light image, that is, when the input image is grayscale, the output fused image also appears in grayscale mode; and when the input image is color, the output fused image also exhibits corresponding color characteristics; this operation ensures that the fused image is consistent with the original visible light image in the color dimension.
[0075] In order to verify the fusion effect of the fused image obtained in step 5, this embodiment specially selects two groups of fused images from the test set for display. The two groups of fused images are respectively as follows: Figure 5 and Figure 6 As shown. Figure 5 and Figure 6 It can be seen that:
[0076] 1) The semantically salient objects of each group of fused images are completely preserved, and the visual effect is significantly improved. The semantically salient objects mentioned above can be seen from the enlarged part on the left side of the figure;
[0077] 2) Each group of fused images contains more texture detail information, and some detail features can be clearly observed in a variety of environments. The above texture detail information can be seen from the enlarged part on the right side of the figure.
[0078] In addition, this application also uses the TNO data set in the test set of the fusion network to test the DenseFuse fusion method, FusionGAN fusion method, U2Fusion fusion method, MUFusion fusion method, AEFusion (average and nuclear) fusion method, IJMLC fusion method, FAFusion fusion method, CrossFuse fusion method, MDA fusion method and DSFusion fusion method. The test results are shown in Table 1.
[0079] In Table 1, Ours refers to the image fusion method described in this embodiment, AG is the average gradient, EN is the information entropy, FMI_pixel is the pixel feature mutual information, FMI_dct is the discrete cosine feature mutual information, FMI_w is the wavelet feature mutual information, Q AB / F is the gradient-based fusion performance, SD is the standard deviation, and SF is the spatial frequency. The bold indicates the best performance among all the compared methods, and the underlined indicates the second best performance.
[0080] Table 1: Comparison results of different fusion methods on quantitative indicators
[0081]
[0082] It can be seen from Table 1 that:
[0083] 1) The image fusion method described in this embodiment can obtain the highest AG value, FMI_w value, SD value and SF value, which shows that the fused image obtained by the fusion method described in this example can obtain richer texture detail information, edge information, contrast information and wavelet-based feature information;
[0084] 2) The image fusion method described in this embodiment can obtain the second highest FMI_pixel value, FMI_dct value and Q AB / F This shows that the fused image obtained by the fusion method described in this example can obtain relatively rich feature mutual information, retain more effective content at the edge, and better meet human visual requirements;
[0085] 3) The EN value obtained by the image fusion method described in this embodiment is lower than that of MUFusion and DSFusion in the prior art. This is mainly because this example emphasizes the retention of information from the two methods of semantics and frequency, which reduces the degree of retention of grayscale information of the fused image and the source image.
Claims
1. A method for fusion of infrared and visible light images based on frequency semantic compensation collaboration, characterized in that: The steps include: Step 1: Obtain the data set of the fusion network and divide the data set into a training set and a test set; Step 2: construct an infrared and visible light image fusion network for feature extraction and image restoration, the fusion network includes: a hierarchical semantic feature extraction unit, a frequency and semantic feature cross compensation module, and an image upsampling reconstruction module; Step 3: Preprocess the images in the training set and normalize them during the training process; Step 4: Train the fusion network under the guidance of the loss function to obtain a trained fusion network model; Step 5: Use the test set data of the fusion network to test the trained network model obtained in step 4 to obtain a fused image.
2. The infrared and visible light image fusion method based on frequency semantic compensation collaboration according to claim 1 is characterized in that: In step 2, the hierarchical semantic feature extraction unit consists of two branches, each of which contains three semantic coherence enhancement modules with the same structure. Each semantic coherence enhancement module integrates a semantic space attention module to extract the semantic feature information of infrared images and visible light images respectively, and connect the semantic features of the two modalities by channel to generate a multimodal semantic feature layer L msf .
3. The infrared and visible light image fusion method based on frequency semantic compensation collaboration according to claim 1 is characterized in that: In step 2, the frequency and semantic feature cross-compensation module adopts a cross-compensation mechanism, which aims to carefully preserve the semantic information in the input features and accurately extract and fuse the frequency information; first, the multimodal semantic feature layer L msf Through a semantic space attention module and a convolution block, a fused semantic feature layer L is generated. fsf , the convolution block consists of a 5×5 convolution layer and a LeakyReLU activation layer; generating a fusion semantic feature layer L fsf After that, it passes through three semantic space attention modules and convolution blocks in turn to protect the semantic information and prevent it from being lost during the processing. The convolution block consists of a 3×3 convolution layer and a LeakyReLU activation layer. Through the convolution operation of different channels, three groups of Sobel operators and two-dimensional wavelet transform work together to obtain the fusion semantic feature layer L rich in semantic information. fsf Extract the frequency information; extract the three sets of low-frequency information L i , i∈1,2,3 is compensated to the fusion semantic feature layer L in the form of residual fsf , generate the low-frequency semantic fusion feature layer L lfsff , three groups of Sobel high frequency information SH i , i∈1, 2, 3 are passed to the image upsampling and reconstruction module for subsequent image reconstruction.
4. The infrared and visible light image fusion method based on frequency semantic compensation collaboration according to claim 1 is characterized in that: In step 2, the image upsampling reconstruction module first extracts the Sobel high-frequency information SH from three groups of different channel numbers. i , i∈1,2,3, processed by a 2×2 upsampling transposed convolution layer to ensure fusion with the low-frequency semantic feature layer L lfsff The sizes of the two parts are kept consistent; secondly, the feature information of the two parts is fused by element-by-element addition, and then processed through five convolution blocks to obtain the final fused image; among the five convolution blocks, the first four are composed of a 3×3 convolution layer and a LeakyReLU activation layer, and the last convolution block is composed of a 3×3 convolution layer and a Tanh activation layer.
5. The infrared and visible light image fusion method based on frequency semantic compensation collaboration according to claim 1 is characterized in that: In step 3, the registered infrared images and visible light images at different times and locations are segmented into 128*128 patch images. The obtained patch images are used to construct the training set of the fusion network, and the image feature data are normalized during the training process.
6. The infrared and visible light image fusion method based on frequency semantic compensation collaboration according to claim 1 is characterized in that: In step 4, the fusion network is trained under the guidance of the loss function to obtain a trained fusion network model. The loss function The calculation formula is shown in formula (1): In formula (1), is the semantic-based pixel loss, is the texture loss of details, α and β are hyper parameters; In formula (1), the semantic-based pixel loss The calculation formula is shown in formula (2): In formula (2), ∥∥1 represents the L1 norm, F is the fused image, VI is the visible light image, and IR is the infrared image; In formula (3), It represents the pixel value threshold result of the N image. When it is greater than or equal to 0.5, the value is 1, otherwise it is 0. N(x, y) represents the pixel value of the original image at position (x, y); In formula (1), the detail-based texture loss The calculation formula is shown in formula (4): In formula (4), H and W represent the height and width of the fused image respectively. It means that the Sobel operator calculates the gradient information of N images, and ∥∥2 means the L2 norm; In formula (5), G x and G y They are the horizontal and vertical convolution kernels of the Sobel operator, respectively. The horizontal and vertical convolution kernels of the Sobel operator are usually defined as: * represents the convolution operation, and N is the original image.
7. The infrared and visible light image fusion method based on frequency semantic compensation cooperation according to claim 1 is characterized in that: In step 5, the data in the test set is input into the trained fusion network model obtained in step 4. During the test, the tone characteristics of the fused image are consistent with the input visible light image, that is, when the input image is grayscale, the output fused image also appears in grayscale mode; and when the input image is color, the output fused image also exhibits corresponding color characteristics. This operation ensures that the fused image is consistent with the original visible light image in the color dimension.
8. The infrared and visible light image fusion method based on frequency semantic compensation collaboration according to claim 2 is characterized in that: The semantic space attention module consists of three feature streams, and the upper feature stream is processed by the Sobel gradient operator and accompanied by a 1×1 convolution layer; The intermediate feature stream contains the initial input features; The lower feature stream is the normalized calculation of the deviation between the initial input feature and its mean; by adding the elements of these three feature streams one by one, and then processing them through another 1×1 convolution layer, and finally adding them to the initial input feature, the accurate semantic space information F is extracted from the initial input feature. si .
9. The infrared and visible light image fusion method based on frequency semantic compensation cooperation according to claim 2 is characterized in that: The six semantic coherence enhancement modules of the dual branches are integrated with the semantic space attention module. First, they pass through a convolution block and the semantic space attention module to extract accurate semantic space information F si , the convolution block consists of a 3×3 convolution layer and a LeakyReLU activation layer; then through the three-branch feature flow, the upper branch consists of two 5×5 convolution layers, a LeakyReLU activation layer and a Sigmoid activation layer, the lower branch consists of two 3×3 convolution layers, a LeakyReLU activation layer and a Sigmoid activation layer, and the middle branch keeps the semantic space information F si The upper and lower branches are respectively multiplied element-by-element with the middle branch and then added to the middle branch; finally, they are added to the semantic space information F after two 3×3 convolutional layers and two LeakyReLU activation layers. si The rich semantic features F are obtained by adding the residuals sf .