Visible light and SAR image fusion method and device based on master-slave convolution and ViT
Through the master-slave convolution and ViT method, combined with Restormer and Biformer modules, feature extraction and fusion are performed, the problem of SAR image noise influence is solved, and high-quality visible light and SAR image fusion results are generated, overcoming the shortcomings of traditional methods and existing deep learning methods.
Patent Information
- Application Number
- CN202510184709.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-17
AI Technical Summary
The existing image fusion method is difficult to effectively process the spot noise of SAR images, resulting in black artifacts in the fusion image. The traditional methods and existing deep learning methods have limited capabilities when extracting visible light and SAR image features, and cannot meet the complete retention and fusion quality of information.
Using the method based on master-slave convolution and ViT, the feature extraction and fusion process are optimized through feature rough extraction, low-frequency feature extraction, master-slave pattern high-frequency feature extraction, feature matching fusion and feature reconstruction, combined with the Restormer and Biformer modules, the master-slave convolution attention module and feature matching fusion module are used to optimize the feature extraction and fusion process.
It effectively overcomes the influence of SAR image noise, improves the generation quality of the fusion image, reduces black artifacts, retains rich texture information and details, and generates a higher quality fusion image.
Smart Images

Figure CN120163737A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a visible light and SAR image fusion method and device based on master-slave convolution and ViT. Background Art
[0002] With the development of remote sensing technology and computer vision, the fusion of visible light and SAR images has important application values in many fields, such as target recognition, topographic mapping, disaster monitoring, etc. Visible light images can provide rich color and texture information, while SAR images have the advantages of penetrating clouds and working all day long. The fusion of the two can achieve information complementarity. However, due to the large difference in the imaging principles of the two types of images, their fusion faces many challenges. Especially different from multi-modal fusion methods such as infrared visible light and medical image fusion, the speckle noise of SAR images has a huge impact on the extraction and reconstruction of image features, and the commonly used image fusion schemes cannot meet the requirements of SAR and visible light fusion.
[0003] Image fusion techniques can be roughly divided into two categories. One category is to use traditional methods, and the other category is to use deep learning methods. Traditional methods mainly include methods based on multi-scale transformation, saliency, subspace, and sparse representation. Traditional methods usually preprocess visible light and SAR images first, then extract features using wavelet transform, principal component analysis, etc., and finally fuse them according to rules such as weighted average. Deep learning methods construct neural networks, such as directly learning fusion weights with shallow neural networks, or constructing deep convolutional neural networks (CNNs). Deep learning methods generally adopt the autoencoder architecture to implement. The autoencoder architecture extracts features by setting convolutional layers of different scales in the encoder, then designs fusion rules for features of different modalities to obtain fused features, and then obtains the fused image through the decoder and optimizes the model with a loss function including content loss and structural loss.
[0004] The feature extraction of traditional methods is insufficient. When extracting features of visible light and SAR images, it is difficult to capture the rich multi-scale and multi-modal features of the images. Visible light images contain a large amount of color and texture details, and SAR images have unique radar echo features that can reflect the terrain and ground object structures. However, the expression ability of traditional methods for these complex features is limited, resulting in the loss of some detail information in the fused image. Therefore, deep learning methods are mainly used for image fusion now.
[0005] Most of the existing image fusion methods based on deep learning are for fusing infrared and visible light images, medical images and other fields. The data they mainly face is images with less noise or mainly additive noise. The data quality in these directions is relatively high and has little impact on the fusion results. However, the speckle noise in SAR images is caused by the interference of coherent waves and generally has the characteristics of multiplicative noise. Its statistical characteristics are relatively complex and difficult to eliminate. Since convolutional neural networks mainly focus on capturing local features, their ability to identify speckle noise spread throughout the image is limited. They cannot capture noise information, resulting in confusion between noise and image content during reconstruction and generating artifacts. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a visible light and SAR image fusion method and device based on master-slave convolution and ViT.
[0007] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0008] The present invention provides a visible light and SAR image fusion method based on master-slave convolution and ViT, including:
[0009] Obtain at least one pair of images to be fused, where each pair of images to be fused includes a visible light image and a SAR image;
[0010] For each pair of images to be fused, perform feature rough extraction and low-frequency feature extraction on the pair of images to be fused successively to obtain the rough extraction features and low-frequency features of the visible light image and the rough extraction features and low-frequency features of the SAR image;
[0011] Perform high-frequency feature extraction on the rough extraction features of the visible light image and the rough extraction features of the SAR image respectively to obtain the high-frequency features with a master-slave pattern of the visible light image and the high-frequency features with a master-slave pattern of the SAR image; wherein, the high-frequency features with a master-slave pattern refer to features in which the ratio of the included slave features to the master features is a preset ratio;
[0012] Perform feature matching and fusion on the high-frequency features of the visible light image and the high-frequency features of the SAR image, and perform feature matching and fusion on the low-frequency features of the visible light image and the low-frequency features of the SAR image respectively to obtain the fused high-frequency features and the fused low-frequency features;
[0013] Perform feature reconstruction according to the fused high-frequency features and the fused low-frequency features to obtain the fused image of the pair of images to be fused.
[0014] The present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0015] The memory is used for storing a computer program;
[0016] When the processor is used to execute the program stored in the memory, the steps of the above-mentioned visible light and SAR image fusion method based on master-slave convolution and ViT are realized.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0018] The method provided by the present invention can overcome the problem of black artifacts generated in the fusion process caused by noise interference when the traditional convolutional neural network processes SAR images by performing feature rough extraction, low-frequency feature extraction, high-frequency feature extraction with a master-slave pattern, matching fusion of high and low-frequency features, and feature reconstruction of the fused high and low-frequency features, and improves the generation quality of the fused image.
[0019] The following will further elaborate on the present invention in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of a visible light and SAR image fusion method based on master-slave convolution and ViT provided by an embodiment of the present invention;
[0021] Figure 2 is another flowchart of a visible light and SAR image fusion method based on master-slave convolution and ViT provided by an embodiment of the present invention;
[0022] Figure 3 is a structural and working flowchart of a high-frequency feature extraction module provided by an embodiment of the present invention;
[0023] Figure 4 is a structural and working flowchart of a trained feature matching and fusion module provided by an embodiment of the present invention;
[0024] Figure 5 are two working flowcharts of a model in the first training stage provided by an embodiment of the present invention;
[0025] Figure 6 is a comparison chart of the fusion effects of the present invention and multiple different existing algorithms. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0027] Figure 1 and Figure 2 are both the schematic flowcharts of a visible light and SAR image fusion method based on master-slave convolution and ViT provided by the embodiments of the present invention. Combining Figure 1 and Figure 2 as shown, the method includes:
[0028] S101. Obtain at least one pair of images to be fused, where each pair of images to be fused includes a visible light image V and a SAR image S.
[0029] S102. For each pair of images to be fused, perform feature rough extraction and low-frequency feature extraction on this pair of images to be fused successively, and obtain the rough extraction feature V′ and low-frequency feature B of the visible light image V and the rough extraction feature S′ and low-frequency feature B of the SAR image S .
[0030] S103. Perform high-frequency feature extraction on the rough extraction feature V′ of the visible light image and the rough extraction feature S′ of the SAR image respectively, and obtain the high-frequency feature D with a master-slave pattern of the visible light image V and the high-frequency feature D with a master-slave pattern of the SAR image S ; where the high-frequency feature with a master-slave pattern refers to a feature in which the ratio of the slave feature to the master feature included is a preset ratio.
[0031] S104. Perform feature matching and fusion on the high-frequency feature D of the visible light image V and the high-frequency feature D of the SAR image S , and perform feature matching and fusion on the low-frequency feature B of the visible light image V and the low-frequency feature B of the SAR image S respectively, and obtain the fused high-frequency feature D F and the fused low-frequency feature B F .
[0032] S105. Perform feature reconstruction according to the fused high-frequency feature D F and the fused low-frequency feature B F to obtain the fused image F of this pair of images to be fused.
[0033] In the present invention, the above S102 can be implemented through S1021 to S1022:
[0034] S1021. Input the visible light image and the SAR image in this pair of images to be fused into the trained feature rough extraction module respectively, and the trained feature rough extraction module outputs the rough extraction feature V′ of the visible light image and the rough extraction feature S′ of the SAR image respectively.
[0035] Specifically, the feature rough extraction module is an image restoration module (i.e., the Restormer module),
[0036] V' = R(V), S' = R(S). This module has strong generalization ability and robustness, can effectively restore image quality under different image degradation conditions, and can capture the overall information of the image.
[0037] S1022. Input the rough extraction features V' of the visible light image and the rough extraction features S' of the SAR image into the trained low-frequency feature extraction module respectively. The trained low-frequency feature extraction module outputs the low-frequency feature B of the visible light image V and the low-frequency feature B of the SAR image S .
[0038] Specifically, the low-frequency feature extraction module is the Biformer module, B V = B(V'), B s = B(S'). The Biformer module can effectively capture the local details and overall features of the objects in the image through its local-global information fusion mechanism, extract the low-frequency background information in the image, and obtain the low-frequency feature B of the visible light image v and the low-frequency feature B of the SAR image s .
[0039] In the present invention, the above S103 can be implemented through S1031 to S1032:
[0040] S1031. Input the rough extraction features V' of the visible light image into the trained high-frequency feature extraction module. The trained high-frequency feature extraction module outputs the high-frequency feature D with a master-slave pattern of the visible light image V .
[0041] S1032. Input the rough extraction features S' of the SAR image into the trained high-frequency feature extraction module. The trained high-frequency feature extraction module outputs the high-frequency feature D with a master-slave pattern of the SAR image S .
[0042] In the present invention, the high-frequency feature extraction part adopts the master-slave convolutional attention module (i.e., high-frequency feature extraction) proposed by the present invention. This module is composed of a main convolutional block and a slave convolutional block. Different from other existing solutions, the structures of these two convolutional blocks are exactly the same, both consisting of 3*3 convolution + BN layer (batch normalization) + Tanh layer, and their master-slave status is adjusted according to the number of branches. In the present invention, the structure of high-frequency feature extraction is three branches and two convolutions, that is, the first branch uses convolutional module 1, and the second and third branches both adopt the same convolutional module 2. Exemplarily, such as Figure 3As shown, the trained high-frequency feature extraction module successively includes: three convolutional branches, a dimension concatenation module, and a spatial attention module. Each convolutional branch includes a convolutional block and a channel attention module cascaded after the convolutional block. Moreover, the convolutional block in the first convolutional branch is different from the convolutional block in the second convolutional branch, and the convolutional block in the second convolutional branch is the same as the convolutional block in the third convolutional branch. Specifically, the convolutional block in the first convolutional branch is CNN1, and the convolutional blocks in the second and third convolutional branches are both CNN2. The structures of the convolutional blocks in each convolutional branch are the same. Specifically, each convolutional block in each convolutional branch successively includes: 3*3 convolution, BN layer, and Tanh layer; and the trained parameters of CNN1 are different from the trained parameters of CNN2. Based on this, V' or S' is respectively input into the three convolutional branches. After being processed by the convolutional block in each convolutional branch, the convolved features are obtained. The convolved features enter the channel attention module. After the channel attention mechanism is performed by the channel attention module, features of size C*H*W are obtained. Then, a dimension concatenation module is used to concatenate these three features after the channel attention mechanism along the channel dimension to obtain a feature of size 3C*H*W. The feature of size 3C*H*W is then sent to the spatial attention module for spatial attention processing to obtain D V or D S . Here, the specific situation of the feature obtained after concatenation is that the ratio of the feature obtained through convolutional block 1 to the feature obtained through convolutional block 2 is 1:2, forming a master-slave pattern. In this way, when performing spatial attention subsequently, each pixel region of the obtained feature is composed of the features of one convolutional block 1 and two convolutional blocks 2, and then the number of channels is compressed through one layer of convolution to obtain the high-frequency feature D v and D s , D V =D(V'), D S =D(S').
[0043] In the present invention, the above S104 can be implemented through S1041 to S1044:
[0044] S1041. Input the high-frequency feature D V of the visible light image and the high-frequency feature D S of the SAR image into the trained feature matching and fusion module.
[0045] S1042. The trained feature matching and fusion module uses the trained parameter matrix θ to perform an affine transformation on the high-frequency feature D S of the SAR image, and fuses the high-frequency feature AT(D S ) after the affine transformation of the SAR image with the high-frequency feature D V of the visible light image to obtain the fused high-frequency feature DF , then D F = D V + AT(D S ), where AT represents a transformation directly using θ. It should be noted that this transformation is the same as the affine transformation.
[0046] S1043. Input the low-frequency feature B of the visible light image V and the low-frequency feature B of the SAR image S into the trained feature matching and fusion module.
[0047] S1044. The trained feature matching and fusion module uses the trained parameter matrix θ to perform an affine transformation on the low-frequency feature B of the SAR image S , and fuses the low-frequency feature AT(B S ) after the affine transformation of the SAR image with the low-frequency feature B of the visible light image V to obtain the fused low-frequency feature B F , then B F = B V + AT(B S ).
[0048] Figure 4 is a structural schematic diagram and a working flowchart of the trained feature matching and fusion module of the present invention. As Figure 4 shown, the trained feature matching and fusion module includes: a parameter matrix generation module and a module for performing affine transformation operations. The input of the parameter matrix generation module is the high-frequency feature of the SAR image, and the output is the trained parameter matrix θ, which has 6 parameters. The parameter matrix generation module includes: two cascaded sub-modules; the first sub-module successively includes: a Cov2d layer, a max pooling layer, a ReLU layer, a Cov2d layer, a max pooling layer, and a ReLU layer; the second sub-module successively includes: a fully connected layer, a ReLU layer, and a fully connected layer. The parameter matrix θ is controlled by a matching loss function, which will be described in the following model training process.
[0049] In the present invention, the trained feature rough extraction module is used for feature rough extraction, the trained low-frequency feature extraction module is used for low-frequency feature extraction, the trained high-frequency feature extraction module is used for high-frequency feature extraction, the trained feature matching and fusion module is used for feature matching and fusion, and the trained feature reconstruction module is used for feature reconstruction. The trained feature rough extraction module, the trained low-frequency feature extraction module, the trained high-frequency feature extraction module, the trained feature matching and fusion module, and the trained feature reconstruction module constitute the trained image fusion model. The above method provided by the present invention can be implemented by the trained image fusion model. Exemplarily, the feature reconstruction module also uses the Restormer module.
[0050] In the present invention, the trained image fusion model is obtained by the following method:
[0051] S1. Obtain a training set; wherein, the training set contains multiple pairs of sample images to be fused.
[0052] Specifically, first obtain multiple pairs of visible light and SAR images as original samples, and then perform preprocessing on them. Specifically: screen the SAR and visible light images, remove the mismatched image pairs, then randomly select a preset number (for example, 20 pairs) as the test set, and then crop the remaining SAR and visible light images into image blocks of a preset size (for example, size 128×128) to expand the training set, and perform normalization processing on the obtained image blocks for use as the training set.
[0053] S2. Construct an initial model for the first training stage; wherein, the initial model for the first training stage includes: an initial feature rough extraction module, an initial low-frequency feature extraction module, an initial high-frequency feature extraction module, and an initial feature reconstruction module.
[0054] It should be noted that the initial module in the present invention refers to a module whose parameters to be trained are initial parameters. In the present invention, the initial parameters of each initial module are determined according to actual needs. It should be noted that in the initial high-frequency feature extraction module, the initial parameters of the convolution blocks in the first convolution branch are different from those of the convolution blocks in the second convolution branch, while the initial parameters of the convolution blocks in the second convolution branch are the same as those of the convolution blocks in the third convolution branch.
[0055] S3. Use the training set to perform multiple iterative trainings on the initial model of the first training stage to obtain the model of the first training stage.
[0056] Exemplarily, Figure 5 are two working flowcharts of the model of the first training stage, as Figure 5 shown. Each time a visible light image or a SAR image is input into the model of the first training stage, correspondingly, the model can reconstruct a visible light image or a SAR image.
[0057] In the present invention, in the first training stage, the feature matching fusion module is not used, and only the feature extraction framework of the model is trained. The loss functions used include: mean square error loss function L MSE , structural similarity loss function L SSIM and decomposition loss function L decomp .
[0058] For an image I, the calculation formulas of each loss function are as follows:
[0059]
[0060] Among them, I i is the i-th pixel value of the image I, is the i-th pixel value of the reconstructed image I, and N is the total number of pixels in the image.
[0061]
[0062] Among them, the SSIM formula is:
[0063]
[0064] Among them, μ I is the average value of the pixel values of I, is the average value of the pixel values of, is the variance of the pixel values of I, is the variance of the pixel values of, is the covariance of the pixel values of I and . c1 = (k1L) 2 , c2 = (k2L) 2 are constants used to maintain stability. L is the dynamic range of pixel values. k1 = 0.01, k2 = 0.03.
[0065]
[0066] Among them, CC(,) represents calculating the correlation coefficient between the two, and ∈ is used to prevent the denominator from being 0.
[0067] In the present invention, the total loss function in the first training stage is:
[0068] L stage1 = L MSE (V) + L MSE (S) + L SSIM (V) + L SSIM (S) + L decomp .
[0069] S4. Based on the model in the first training stage, construct the initial model in the second training stage; among them, the initial model in the second training stage includes: the feature rough extraction module obtained in the first training stage, the low-frequency feature extraction module obtained in the first training stage, the high-frequency feature extraction module obtained in the first training stage, the initial feature matching and fusion module, and the feature reconstruction module obtained in the first training stage.
[0070] S5. Use the training set to perform multiple iterative trainings on the initial model in the second training stage to obtain a trained image fusion model.
[0071] In the present invention, the loss function used in the second training stage includes: L int , the gradient loss function L grad , L decomp and the matching loss function L θ .
[0072] The calculation formulas of each loss function are as follows:
[0073]
[0074] Among them, H and W respectively represent the height and width of the image, V and S respectively represent the input visible light image and SAR image, F represents the fused image finally output by the model, max(V, S) represents taking the maximum value of each pixel position in V and S, and F - max(V, S) represents the difference between the pixel value at each pixel position in F and the maximum value of the corresponding pixel position in V and S.
[0075]
[0076] This loss function is a gradient loss function, which is used to measure the difference between the generated image and the reference image in terms of gradient features such as edges. represents taking the gradient of the pixel value at each pixel position in F, and Similarly, ||.||1 represents taking the L1 norm.
[0077]
[0078] The matching loss function is used to constrain the affine transformation, θ i represents the i-th parameter in the generated parameter matrix, I′ is the identity matrix, and I′ i is the i-th parameter in the identity matrix.
[0079] In the present invention, the total loss function in the second training stage is:
[0080] L stage2 = L int + L grad + L decomp + L θ .
[0081] The present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0082] The memory is used to store a computer program;
[0083] A processor, when executing a program stored in a memory, implements the steps of the above-mentioned visible light and SAR image fusion method based on master-slave convolution and ViT.
[0084] The present invention comprehensively uses different Vision Transformer modules, effectively ensuring the model's grasp of the global information of the image, and making distinctions in image feature restoration and extraction of low-frequency information of the image, improving the model's understanding and processing ability of the image. Moreover, it overcomes the problem that Vision Transformer ignores local feature details, resulting in reduced distinguishability between the background and the foreground. The present invention proposes a master-slave convolution attention module, which is different from the previous ideas of using convolution. Instead of increasing the distinguishability by stacking convolution layers or using different convolution modules, it uses convolution blocks with the same structure, controls the master-slave pattern of convolution by controlling the number of branches, and forms a three-branch two-convolution form. After each branch passes through convolution, channel attention is used for processing, and then the feature maps of the 3 branches are concatenated by channel to form a pattern with a ratio of slave feature to master feature of 1:2. Then, it is sent to spatial attention to re-integrate the features according to the master-slave ratio of 1:2, effectively improving the image feature extraction ability. The present invention proposes a feature matching and fusion module, which proposes a new feature fusion method, registers the image at the feature layer rather than the element layer, effectively reducing the complex structure of the model and improving the quality of the fused image.
[0085] The superiority of the present invention is illustrated by some comparison diagrams below. Figure 6 It is a comparison diagram of the fusion effects of the present invention and multiple different existing algorithms. Figure 6 Each picture in it is an overhead view of a block image. Among them, (a) is a visible light image, (b) is a SAR image, (c) is a fused image generated by IFCNN, (d) is a fused image generated by the DenseFuse network, (e) is a fused image generated by the DIDFuse network, (f) is a fused image generated by NestFuse, (g) is a fused image generated by PMGI, (h) is a fused image generated by SDNet, and (i) is a fused image generated by the network model proposed by the present invention. It can be clearly seen that there are no suspicious targets on the rooftops of the buildings framed by the red boxes in both the visible light and SAR images, but there are obvious black artifact squares in (c) to (g), and this is effectively solved in (i).
[0086] In terms of objective evaluation indicators, 20 pairs of SAR and visible light images are randomly selected, and after being processed by the network, the index table 1 is obtained. Among them, the ones in bold black are the best, and the ones with * are the second best. Obviously, the present invention maintains a leading level relative to other algorithms. This shows that the fusion framework and the rules of the fusion algorithm proposed by the present invention can better solve the fusion problem.
[0087] Table 1
[0088] EN SF SCD SSIM AG PSNR IFCNN 6.74 23.05 1.29 0.84 9.43 11.79 DenseFuse 6.58 18.25 1.32 1.09* 7.63 18.57 DIDFuse 7.4 34.01* 1.86 1.05 13.68* 15.41 NestFuse 6.67 19.48 1.42 1.1 8.16 18.21* PMGI 6.65 18.17 1.36 1.07 7.55 17.69 SDNet 6.78 24.33 1.38 1.07 9.78 16.52 The present invention 7.12* 35.16 1.53* 1.08 14.83 16.32
[0089] In summary, the algorithm proposed by the present invention can make up for the large differences between the two source images, while relatively completely retaining rich texture information and detail information, and solving the fusion artifact problem generated by the CNN network, and can generate higher-quality fused images.
[0090] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.
[0091] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0092] In the specification, the word "including" does not exclude other components or steps, and "a" or "one" does not exclude the case of a plurality. Certain measures are recited in mutually different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0093] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A visible light and SAR image fusion method based on master-slave convolution and ViT, characterized in that: include: Acquire at least one pair of fused images, wherein each pair of fused images includes a visible light image and a SAR image; For each image to be fused, performing rough feature extraction and low-frequency feature extraction on the image to be fused in sequence, to obtain rough extracted features and low-frequency features of the visible light image and rough extracted features and low-frequency features of the SAR image; The coarsely extracted features of the visible light image and the coarsely extracted features of the SAR image are respectively subjected to high-frequency feature extraction to obtain high-frequency features of the visible light image with a master-slave pattern and high-frequency features of the SAR image with a master-slave pattern; wherein the high-frequency features with a master-slave pattern refer to features in which the ratio of slave features to master features is a preset ratio; Performing feature matching and fusion on the high-frequency features of the visible light image and the high-frequency features of the SAR image, and performing feature matching and fusion on the low-frequency features of the visible light image and the low-frequency features of the SAR image, respectively obtaining fused high-frequency features and fused low-frequency features; Feature reconstruction is performed according to the fused high-frequency features and the fused low-frequency features to obtain a fused image of the pair of images to be fused.
2. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 1 is characterized in that: The method of extracting high-frequency features from the coarsely extracted features of the visible light image and the coarsely extracted features of the SAR image respectively to obtain high-frequency features of the visible light image with a master-slave pattern and high-frequency features of the SAR image with a master-slave pattern includes: Inputting the coarsely extracted features of the visible light image into a trained high-frequency feature extraction module, the trained high-frequency feature extraction module outputs high-frequency features of the visible light image with a master-slave pattern; wherein the ratio of slave features to master features in the high-frequency features is 1:2; Inputting the coarsely extracted features of the SAR image into the trained high-frequency feature extraction module, and the trained high-frequency feature extraction module outputs the high-frequency features of the SAR image with a master-slave pattern; Among them, the trained high-frequency feature extraction module includes three convolution branches, a dimensional splicing module and a spatial attention module in sequence; each convolution branch includes a convolution block and a channel attention module cascaded after the convolution block, and the convolution block in the first convolution branch is different from the convolution block in the second convolution branch, and the convolution block in the second convolution branch is the same as the convolution block in the third convolution branch.
3. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 2 is characterized in that: The convolution blocks in each convolution branch include: 3*3 convolution, BN layer and Tanh layer in sequence; the trained parameters of the convolution block in the first convolution branch are different from the trained parameters of the convolution block in the second convolution branch, and the trained parameters of the convolution block in the second convolution branch are the same as the trained parameters of the convolution block in the third convolution branch.
4. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 1, characterized in that: The performing feature matching and fusing of the high-frequency features of the visible light image and the high-frequency features of the SAR image, and the performing feature matching and fusing of the low-frequency features of the visible light image and the low-frequency features of the SAR image, respectively obtaining fused high-frequency features and fused low-frequency features, comprises: Inputting the high-frequency features of the visible light image and the high-frequency features of the SAR image into a trained feature matching and fusion module; The trained feature matching and fusion module uses the trained parameter matrix to perform affine transformation on the high-frequency features of the SAR image, and performs feature fusion on the high-frequency features of the SAR image after affine transformation with the high-frequency features of the visible light image to obtain the fused high-frequency features; Inputting the low-frequency features of the visible light image and the low-frequency features of the SAR image into the trained feature matching and fusion module; The trained feature matching and fusion module uses the trained parameter matrix to perform affine transformation on the low-frequency features of the SAR image, and performs feature fusion on the affine transformed low-frequency features of the SAR image and the low-frequency features of the visible light image to obtain the fused low-frequency features.
5. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 4 is characterized in that: The trained feature matching fusion module includes: a parameter matrix generation module and a module for performing affine transformation operations; The input of the parameter matrix generation module is the high-frequency features of the SAR image, and the output is the trained parameter matrix; The parameter matrix generation module includes: two cascaded sub-modules; the first sub-module includes: Cov2d layer, maximum pooling layer, ReLU layer, Cov2d layer, maximum pooling layer and ReLU layer in sequence; the second sub-module includes: fully connected layer, ReLU layer and fully connected layer in sequence.
6. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 1, characterized in that: The step of successively performing rough feature extraction and low-frequency feature extraction on the pair of images to be fused to obtain low-frequency features of the visible light image and low-frequency features of the SAR image includes: Inputting the visible light image and the SAR image in the pair of images to be fused into the trained rough feature extraction modules respectively, and the trained rough feature extraction modules output the rough extracted features of the visible light image and the rough extracted features of the SAR image respectively; The coarsely extracted features of the visible light image and the coarsely extracted features of the SAR image are respectively input into the trained low-frequency feature extraction modules, and the trained low-frequency feature extraction modules respectively output the low-frequency features of the visible light image and the low-frequency features of the SAR image.
7. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 6, characterized in that: The trained rough feature extraction module is a trained Restormer module.
8. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 6, characterized in that: The trained low-frequency feature extraction module is a trained Biformer module.
9. The visible light and SAR image fusion method based on master-slave convolution and ViT according to claim 1, characterized in that: The trained feature rough extraction module is used for feature rough extraction, the trained low-frequency feature extraction module is used for low-frequency feature extraction, the trained high-frequency feature extraction module is used for high-frequency feature extraction, the trained feature matching and fusion module is used for feature matching and fusion, and the trained feature reconstruction module is used for feature reconstruction; The trained feature rough extraction module, the trained low-frequency feature extraction module, the trained high-frequency feature extraction module, the trained feature matching fusion module and the trained feature reconstruction module constitute a trained image fusion model; the trained image fusion model is trained by the following method: Acquire a training set; wherein the training set contains multiple sample images to be fused; Constructing an initial model of the first training stage; wherein the initial model of the first training stage includes: an initial feature rough extraction module, an initial low-frequency feature extraction module, an initial high-frequency feature extraction module and an initial feature reconstruction module; Using the training set to perform multiple iterative training on the initial model of the first training stage to obtain a model of the first training stage; Based on the model of the first training stage, construct an initial model of the second training stage; wherein the initial model of the second training stage includes: a feature rough extraction module obtained in the first training stage, a low-frequency feature extraction module obtained in the first training stage, a high-frequency feature extraction module obtained in the first training stage, an initial feature matching and fusion module, and a feature reconstruction module obtained in the first training stage; The training set is used to perform multiple iterative training on the initial model of the second training stage to obtain the trained image fusion model.
10. An electronic device comprising a processor, a communication interface, a memory and a communication bus, characterized in that: The processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 1-9 when executing the program stored in the memory.