Fusion Image Model Training Method, Generation Method and Device for Multi-Source Remote Sensing Data

By building a fusion image model for multi-source remote sensing data, using a dual-stream network and a pseudo-symmetric network combined with a core weight adaptive optimization module and attention mechanism, the characteristics of multi-source remote sensing data are extracted and fused, and the problem that the existing technology cannot differentiate the extraction and fusion features is achieved, and high-quality fusion image generation is achieved.

CN116310634BActive Publication Date: 2025-06-03BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310141473.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-07
Publication Date
2025-06-03
Estimated Expiration
2043-02-07

AI Technical Summary

Technical Problem

The prior art cannot differentiate the features of the synthetic aperture radar image and visible light image, so that the feature fusion cannot be achieved and the fusion image is generated.

Method used

The fusion image model training method for multi-source remote sensing data is adopted, and the characteristics of multi-source remote sensing data are extracted and fused to generate fusion images by building a dual-stream network and a pseudo-symmetric network, combining the core weight adaptive optimization module and attention mechanism.

Benefits of technology

The information complementarity between multi-source remote sensing data is realized, and the amount and readability of the fused image are improved. The generated fused image not only retains the texture information in the synthetic aperture radar image, but also retains the visible light spectrum information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310634B_ABST
    Figure CN116310634B_ABST
Patent Text Reader

Abstract

The present invention provides a method, a generation method and a device for training a fusion image model for multi-source remote sensing data, including: obtaining a plurality of synthetic aperture radar images and visible light image pairs; obtaining an initial fusion image model and a discriminator; the initial fusion image model includes a feature extraction module, a feature fusion module and a fusion image generation module; the feature extraction module is a dual-stream network, each branch adopts a preset neural network with the same structure, each preset neural network is composed of a plurality of convolutional blocks, the parameters of the front-set number of convolutional blocks on the two branches are kept symmetric, and a kernel weight adaptive optimization module is arranged between each convolutional block; the fusion image generation module is constructed by a generative adversarial network composed of a generator and a discriminator, and a positive and negative sample training model is constructed based on the fusion image, the synthetic aperture radar image and the visible light image to obtain a fusion image model for generating a fusion image of the synthetic aperture radar image and the visible light image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for training and generating a fusion image model for multi-source remote sensing data. Background Art

[0002] With the development of remote sensing technology, different sensors such as optical, thermal infrared, and microwave can be deployed as information acquisition payloads on remote platforms such as satellites and airplanes, providing multi-source remote sensing data with multi-spectral, multi-sensor, and multi-resolution to ground researchers. Due to different imaging principles and technical conditions, the remote sensing data obtained by a single remote sensor cannot comprehensively reflect the characteristics of the target object. If multiple data with different characteristics are combined to complement each other, their respective advantages can be exerted, their respective deficiencies can be made up, the ground target can be reflected more comprehensively, and stronger information interpretation ability and more reliable analysis results can be provided. The multi-source remote sensing image fusion technology combines and matches multi-source information, and intelligently synthesizes multi-source remote sensing image data or interpretation products of the same area to generate more accurate, complete, and reliable fusion data than a single information source.

[0003] Taking the fusion of two multi-source remote sensing data, synthetic aperture radar image (SAR image) and visible light image, as an example, the first prior art solution is to adopt a basic data fusion solution. At the feature level, by using the coupled non-negative matrix factorization unmixing method, the visible light image and the SAR image are mixed at the high-dimensional feature level and then reduced in dimension again; or the fusion problem is redefined as a recombined estimation of the core tensor and the dictionary based on coupled sparse tensor decomposition. At the image level, in a more naive pixel weighted average mode, the texture information is calculated by superimposing on the split intensity-color-saturation channels. Although this solution takes into account the large differences between the SAR image and the visible light image in data format and data representation, however, based on traditional fusion methods, the texture and structure information of the fusion result is often very blurred in visual performance, and the generation speed is slow and requires repeated iteration; at the same time, due to the accumulation of error gradients, the intrinsic speckle noise of the SAR image will inevitably cause the neural network to degenerate, affecting the network performance; there is also no effective association between the features of the SAR image and the visible light image. The second prior art solution is to use a deep learning method to process remote sensing images and generate fusion images based on a generative adversarial network. However, the existing deep neural network-based solutions use a completely symmetric Siamese network-like symmetric network for feature extraction, lacking consideration of image differences, and ultimately unable to achieve effective feature fusion. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method, a generation method, and a device for training a fusion image model for multi-source remote sensing data, so as to eliminate or improve one or more defects existing in the prior art, and solve the problem that the prior art cannot differentially extract the features of synthetic aperture radar images and visible light images, and thus cannot achieve feature fusion to generate a fusion image.

[0005] On the one hand, the present invention provides a method for training a fusion image model for multi-source remote sensing data, characterized in that the method includes the following steps:

[0006] Obtain a multi-source remote sensing data set, where the multi-source remote sensing data set contains a plurality of data strips, and each data strip includes a synthetic aperture radar image and a visible light image of the same area at the same time generated based on a preset registration algorithm; the number of channels of the synthetic aperture radar image is replicated to be equal to the number of channels of the visible light image;

[0007] Obtain an initial fusion image model and a discriminator; the initial fusion image model includes a feature extraction module, a feature fusion module, and a fusion image generation module; the feature extraction module is a dual-stream network, and each branch uses a preset neural network with the same structure. The preset neural network is composed of a plurality of convolutional blocks, and the parameters of the first set number of convolutional blocks on the two branches are kept symmetric. Each convolutional block is provided with a kernel weight adaptive optimization module; wherein, the weight adaptive optimization module performs global average pooling operation on each input synthetic aperture radar image or each visible light image, calculates the weight coefficient of the corresponding image based on a plurality of preset image feature evaluation units, and uses the weight coefficient to weight the convolutional kernel weight corresponding to the image feature evaluation unit to update the weight of the convolutional kernel in the corresponding convolutional block; the fusion image generation module is used as a generator to construct a generative adversarial network with the discriminator;

[0008] Input the synthetic aperture radar images of a preset batch number of data strips into the first branch of the dual-stream network to extract the synthetic aperture radar image feature map, and input the visible light images into the second branch of the dual-stream network to extract the visible light image feature map;

[0009] Input the synthetic aperture radar image feature map and the visible light image feature map into the feature fusion module. Based on a preset attention mechanism, extract their respective features by exchanging the query matrices of the synthetic aperture radar image feature map and the visible light image feature map, and use the feature layer to superimpose the features extracted from the synthetic aperture radar image feature map and the visible light image feature map to output a fusion feature vector;

[0010] Input the fusion feature vector into the fusion image generation module to generate a fusion image;

[0011] The synthetic aperture radar images in each data bar and the corresponding generated fused images are stitched into non-real images, which are marked as negative samples; the synthetic aperture radar images and the visible light images are stitched into real images, which are marked as positive samples; the non-real images and the real images are input into the discriminator for training; the non-real images are input into the trained discriminator, and the discriminator parameters are fixed, and the generator is trained according to the discrimination results of the discriminator; the loss functions of the generator and the discriminator are respectively constructed, and the generator and the discriminator that meet the preset performance are trained to finally obtain a fused image model.

[0012] In some embodiments, the two-stream network constructs a pseudo-symmetric network based on the Siamese network; each branch uses a preset neural network with the same structure, and the preset neural network is composed of 5 convolutional blocks; the parameters of the first 3 convolutional blocks on the two branches are the same, and the parameters of the last 2 convolutional blocks are different.

[0013] In some embodiments, a multi-receptive field feature extraction module is further provided between each convolutional block of the preset neural network; the image feature map processed by the current convolutional block is input into the multi-receptive field feature extraction module, and the features of the image feature map with a preset number of different degrees of receptive fields from local to global are extracted and superimposed by a preset number of parallel convolutional layers, and then input into the next convolutional block; among them, the calculation formula of the receptive field can be expressed as:

[0014]

[0015] Among them, RF l represents the receptive field size of the l-th layer; K’ l represents the convolutional kernel size of the l-th layer; S i represents the stride of the i-th layer.

[0016] In some embodiments, based on a preset attention mechanism, by exchanging the query matrices of the synthetic aperture radar image feature map and the visible light image feature map, their respective features are extracted, and it further includes:

[0017] Using a preset convolutional layer to compress the number of channels of the synthetic aperture radar image feature map or the visible light image feature map, and performing linear mapping to obtain the first key-value matrix, the first value matrix, and the first query matrix of the synthetic aperture radar image feature map, as well as the second key-value matrix, the second value matrix, and the second query matrix of the visible light image feature map; using the first query matrix for the feature processing of the visible light image feature map; using the second query matrix for the feature processing of the synthetic aperture radar image feature map;

[0018] In the feature processing of the synthetic aperture radar image feature map, perform matrix dot multiplication on the subsequent features obtained based on the first key value matrix and the first value matrix, calculate the first vector value through the Softmax function, perform dot product fusion on the first vector value and the subsequent features obtained based on the second query matrix, and perform a convolution operation; in the feature processing of the visible light image feature map, perform matrix dot multiplication on the subsequent features obtained based on the second key value matrix and the second value matrix, calculate the second vector value through the Softmax function, perform dot product fusion on the second vector value and the subsequent features obtained based on the first query matrix, and perform a convolution operation; the corresponding calculation process can be expressed as:

[0019] Attention(K,V,Q)=f(Softmax(f T (K)·f(V))·f T (Q other ));

[0020] Where K, V, and Q represent learnable key value matrix, value matrix, and query matrix; (·) T represents transpose.

[0021] In some embodiments, use a feature layer to superimpose the features extracted from the synthetic aperture radar image feature map and the visible light image feature map, output a fused feature vector, and the calculation formula is:

[0022] F fusion =concat(g(Attention S )+g(Attention H ));

[0023] Where g represents a convolution operation; Attention S represents the feature after fusion of the synthetic aperture radar image feature map; Attention H represents the feature after fusion of the visible light image feature map.

[0024] In some embodiments, the fusion image generation module is used as a generator to construct a generative adversarial network with the discriminator; the generative adversarial network adopts a conditional generative adversarial network, takes the synthetic aperture radar image as a conditional input, and optimizes the generator to generate a fusion image; the discriminator consists of 6 convolutional layers, and each convolutional layer has a batch normalization layer and a rectified linear unit activation layer; the output result of the discriminator generates a discriminant result of the authenticity of the fusion image through a preset Sigmoid layer.

[0025] In some embodiments, the loss functions of the generator and the discriminator are respectively constructed, and further include:

[0026] The calculation formula of the loss function of the discriminator is as follows:

[0027]

[0028] where E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I F represents the fused image; I OPT represents the visible light image; D(I SAR , I F ) represents the discrimination result obtained by inputting the non-real image (I SAR , I F ) into the discriminator; D(I SAR , I OPT ) represents the discrimination result obtained by inputting the real image (I SAR , I OPT ) into the discriminator.

[0029] In some embodiments, constructing the loss functions of the generator and the discriminator respectively further includes:

[0030] Constructing the joint loss of the generator based on the generative adversarial loss, the mean absolute error loss, and the texture loss;

[0031] The calculation formula of the generative adversarial loss is as follows:

[0032]

[0033] where E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I OPT represents the visible light image; D(·) represents the discrimination result obtained by the discriminator; G(·) represents the fused image generated by the generator.

[0034] The calculation formula of the mean absolute error loss is as follows:

[0035]

[0036] where E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I OPT represents the visible light image; I F represents the fused image.

[0037] The calculation formula of the texture loss is as follows:

[0038]

[0039] where E[·] represents taking the average expectation; ISAR represents the synthetic aperture radar image; I OPT represents the visible light image; I F represents the fused image; represents the gray-level co-occurrence matrix of the brightness component of the fused image; represents the gray-level co-occurrence matrix of the brightness component of the synthetic aperture radar image.

[0040] The generation adversarial loss, the mean absolute error loss, and the texture loss are weighted to obtain the loss function of the generator, and the calculation formula is:

[0041]

[0042] wherein, L GAN (G) represents the generation adversarial loss; represents the mean absolute error loss; λ is a hyperparameter; L GLCM (G) represents the texture loss; μ represents the weight coefficient of the texture loss.

[0043] On the other hand, the present invention provides a method for generating a fused image for multi-source remote sensing data, characterized in that the method includes the following steps:

[0044] Obtain a synthetic aperture radar image and a visible light image in the same area at the same time generated based on a preset registration algorithm to be fused; the number of channels of the synthetic aperture radar image is replicated to be equal to the number of channels of the visible light image;

[0045] Input the synthetic aperture radar image and the visible light image into the fused image model obtained by the method for training a fused image model for multi-source remote sensing data described in any one of the above, so as to obtain the fused image of the synthetic aperture radar image and the visible light image.

[0046] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods mentioned above are implemented.

[0047] The beneficial effects of the present invention are at least:

[0048] The present invention provides a method, a generation method and a device for training a fusion image model for multi-source remote sensing data. By constructing a fusion image model, features of multi-source remote sensing data with differences are extracted and fused to generate a fusion image, making up for the limitations of a single sensor itself, realizing information complementarity between multi-source remote sensing data, and improving the information content and readability of the fusion image. The fusion image model includes a feature extraction module, a feature fusion module and a fusion image generation module. The feature extraction module is a dual-stream network, and each branch adopts a preset deep convolutional neural network with the same structure. Based on the anti-noise interference and translational invariance of the deep convolutional neural network, reliable fusion image generation is realized. A pseudo-symmetric network is constructed, and a kernel weight adaptive optimization module is introduced into the pseudo-symmetric network, forcing the feature extraction network with shared parameters to adaptively optimize features according to the characteristics of the input image while learning the consistency of multi-source images. The feature fusion module performs information feature fusion of multi-source images based on the attention mechanism, and performs weighted superposition after calculating the feature correlation, better emphasizing the complementary information in the multi-source images and outputting fusion features with more complete feature semantic expressions. The fusion image generation module constructs a conditional generative adversarial network with a discriminator for supervised training. In addition to constructing a conventional loss, texture information of the image is mined through a gray-level co-occurrence matrix to construct a texture loss, and supervised training is performed based on both color and texture, so that the finally generated fusion image retains both the texture information in the synthetic aperture radar image and the visible light spectrum information. The data information of the fusion image generated based on the fusion image model can provide higher-quality and more information-rich basic data for subsequent remote sensing data interpretation tasks.

[0049] Additional advantages, objects, and features of the present invention will be partly described in the following description, and will partly become apparent to those of ordinary skill in the art after studying the following. Or they can be learned through the practice of the present invention. The objects and other advantages of the present invention can be realized and obtained by the structure specifically pointed out in the specification and the drawings.

[0050] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:

[0052] Figure 1 It is a schematic diagram of the steps of a method for training a fusion image model for multi-source remote sensing data in an embodiment of the present invention.

[0053] Figure 2It is a schematic structural diagram of a fusion image model for multi-source remote sensing data in an embodiment of the present invention.

[0054] Figure 3 It is a schematic structural diagram of a two-stream network with independent parameters, a Siamese network with consistent parameters, and a pseudo-symmetric network in an embodiment of the present invention.

[0055] Figure 4 It is a schematic structural diagram of a single-branch structure of a two-stream network in an embodiment of the present invention.

[0056] Figure 5 It is a schematic structural diagram of a dynamic convolution weight optimization module in an embodiment of the present invention.

[0057] Figure 6 It is a flowchart of an adaptive convolution kernel update algorithm of a kernel weight adaptive optimization module in an embodiment of the present invention.

[0058] Figure 7 It is a flowchart of an algorithm of a feature fusion module based on an attention mechanism in an embodiment of the present invention.

[0059] Figure 8 It is a schematic structural diagram of a conditional generative adversarial network constructed by a fusion image generation module and a discriminator in an embodiment of the present invention. Detailed implementation manners

[0060] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the implementation manners and the accompanying drawings. Here, the illustrative implementation manners and descriptions of the present invention are used to explain the present invention, but do not limit the present invention.

[0061] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0062] It should be emphasized that the term "including / containing" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0063] Here, it should also be noted that if not otherwise specified, the term "connection" in this article can refer not only to direct connection but also to indirect connection with an intermediate.

[0064] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0065] It should be emphasized here that the step marks mentioned below do not limit the sequence of each step. Instead, it should be understood that the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0066] To solve the problem that the prior art cannot differentially extract the features of synthetic aperture radar images and visible light images, and thus cannot achieve feature fusion to generate fused images, the present invention provides a method for training a fused image model for multi-source remote sensing data, as Figure 1 shown. The method includes the following steps S101 to S103:

[0067] Step S101: Obtain a multi-source remote sensing data set, which contains multiple data strips. Each data strip includes a synthetic aperture radar image and a visible light image of the same area at the same time generated based on a preset registration algorithm. Among them, the number of channels of the synthetic aperture radar image is replicated to be equal to the number of channels of the visible light image.

[0068] Step S102: Obtain an initial fused image model and a discriminator. The initial fused image model includes a feature extraction module, a feature fusion module, and a fused image generation module. Among them, the feature extraction module is a dual-stream network. Each branch uses a preset neural network with the same structure. Each preset neural network is composed of multiple convolutional blocks. The parameters of the first set number of convolutional blocks on the two branches are kept symmetric. Each convolutional block is provided with a kernel weight adaptive optimization module. The weight adaptive optimization module performs global average pooling operation on each input synthetic aperture radar image or each visible light image, calculates the weight coefficient of the corresponding image based on multiple preset image feature evaluation units, and weights the convolutional kernel weight corresponding to the image feature evaluation unit with the weight coefficient to update the weight of the convolutional kernel in the corresponding convolutional block. The fused image generation module, as a generator, and the discriminator construct a generative adversarial network.

[0069] Step S103: Stitch the synthetic aperture radar images in each data strip with the corresponding generated fused images into non-real images, marked as negative samples; stitch the synthetic aperture radar images and visible light images into real images, marked as positive samples; input the non-real images and real images into the discriminator for training; input the non-real images into the trained discriminator, fix the parameters of the discriminator, and train the generator according to the discrimination result of the discriminator; respectively construct the loss functions of the generator and the discriminator, train to obtain a generator and a discriminator that meet the preset performance, and finally obtain a fused image model.

[0070] In step S101, a multi-source remote sensing image pair with spatial consistency and temporal consistency generated based on a preset registration algorithm is obtained, and a multi-source remote sensing data set is constructed. Among them, in the present invention, two types of multi-source remote sensing data, synthetic aperture radar (SAR) images and visible light images, are taken as examples to design a fusion image model for synthetic aperture radar images and visible light images to generate a fusion image of synthetic aperture radar images and visible light images.

[0071] Synthetic aperture radar images are two-dimensional images generated by synthetic aperture radar sensors actively emitting microwaves and receiving echoes. The high-resolution coherent imaging system for receiving echoes can collect data all day and in all climates. Visible light images receive the spectrum in the frequency band of 0.38 - 0.76 microns, and can show the color information of ground objects through color quantization. However, it is vulnerable to natural climate. Clouds, occlusion, etc. will hinder the transmission of the visible light band spectrum. Therefore, such images cannot provide all-day and all-weather remote sensing images.

[0072] Considering that the initial synthetic aperture radar image is single-channel, the channels of the initial synthetic aperture radar image are copied three times. After copying, the number of channels of the synthetic aperture radar image is the same as that of the visible light image to ensure the consistent symmetry of the subsequent two-stream network for feature extraction. Exemplarily, the size of the synthetic aperture radar image and the visible light image is (256, 256, 3).

[0073] In step S102, an initial fusion image model and a discriminator are constructed, as Figure 2 shown, which is the structural framework diagram of the whole model. In the initial fusion image model, it can be divided into a feature extraction module, a feature fusion module, and a fusion image generation module according to functions.

[0074] In the feature extraction module, a two-stream network is adopted. Each branch adopts a preset neural network with the same structure. Each preset neural network is composed of multiple convolutional blocks. The parameters of the front set number of convolutional blocks on the two branches are symmetric. Each convolutional block is provided with a kernel weight adaptive optimization module.

[0075] Conventional two-stream networks can be divided into two categories. One is a network with completely independent feature parameters and a symmetric network structure, as Figure 3 (a) shown. Each branch of this network accurately learns the features of the corresponding input image. However, since the two-branch networks are completely independent, the number of parameters is often large, and the feature coupling in the feature fusion stage is low. As Figure 3(b) shows a Siamese neural network, and its sub-networks have exactly the same structure, parameters, and weights. Considering the above two classifications, in order to better maintain the characteristics of multi-source inputs and at the same time achieve the acquisition of same-sex features, in the present invention, a two-stream network constructs a pseudo-symmetric network based on the Siamese network, as shown in Figure 3 (c). The multi-source data features are acquired with a consistent network structure and partial parameter sharing as the symmetric result, maximizing the flexibility of feature extraction. Among them, the parameters in the shallow feature extraction part are set the same, and the parameters in the deep feature extraction part are set differently, enabling each branch of the two-stream network to have the ability to adapt to features and also pay attention to different detailed features. Ultimately, the two-stream feature extraction process can obtain a higher image data correlation while maintaining the features of the original image.

[0076] In some embodiments, in the two-stream network, each branch adopts a preset neural network with the same structure. Each preset neural network is composed of 5 convolutional blocks. The parameters of the first 3 convolutional blocks on the two branches are set the same, and the parameters of the last 2 convolutional blocks are set differently.

[0077] As shown in Figure 4 , it is the structure diagram of a single branch of the two-stream network, including 5 convolutional blocks G-C1, G-C2, G-C3, G-C4, G-C5 for feature extraction, and 1 convolutional kernel G-concat for adjusting the number of channels. Each convolutional kernel includes a convolutional layer, a batch normalization (BN) layer, and a rectified linear unit (ReLU) activation layer. Considering that it is computationally expensive to keep the size of the input image during convolution, each convolutional block downsamples the input image through a pooling operation. It should be emphasized that excessive downsampling may cause the feature map to lack accurate global information. Therefore, in the present invention, the sampling step between convolutional kernels is 2. At the same time, to optimize the feature extraction process, each convolutional layer is subjected to batch normalization and ReLU processing.

[0078] The feature extraction module is further optimized below. Considering that if more global context information is to be captured from the input image, a larger receptive field is required. For a standard Convolutional Neural Network (CNN), the traditional method of expanding the receptive field is to use a larger convolutional kernel size and stack more convolutional layers, but this operation may lead to an exponential expansion of training parameters, making the network difficult to train. Another method of expanding the receptive field is to stack more pooling layers to expand the receptive field by reducing the dimension of the feature map and maintaining significant features. In the present invention, a multi-receptive field feature extraction module that can be applied to multiple basic feature extraction networks is designed between each convolutional block to enable the network to have multi-level feature extraction capabilities. Among them, the basic feature extraction networks include ResNet, VggNet, AlexNet, etc.

[0079] The structural diagram of the multi-receptive field feature extraction module is as shown Figure 4 in the dotted area, including three parallel convolutional layers. From top to bottom, each convolutional layer is correspondingly configured with 1×1, 3×3, and 3×3 convolutional kernels with a dilation rate of 2. Specifically, the image feature map obtained by processing the current convolutional block is input into the multi-receptive field feature extraction module. Three parallel convolutional layers are used to extract the features of the image feature map with three different degrees of receptive fields from local to global. The output feature maps of the three parallel convolutional layers are superimposed, and the 1×1 convolutional kernel is used to fuse the features between channels and adjust the number of channels of the feature map to make its output consistent with the number of channels at the input, and then input into the next convolutional block.

[0080] In some embodiments, the calculation formula of the receptive field is as shown in formula (1):

[0081]

[0082] In formula (1), RF l represents the receptive field size of layer l; K’ l represents the convolutional kernel size of layer l; S i represents the stride of the i-th layer.

[0083] Among them, the calculation formula of K’ l is as shown in formula (2):

[0084] K’ l = K l +(K l -1)(r l -1); (2)

[0085] In formula (2), K l represents the size of the original convolutional kernel; r l represents the dilation rate of the dilated convolution of layer l.

[0086] In some embodiments, the multi-receptive field feature extraction module can be set between each convolutional block, or can be set between some convolutional blocks. Exemplarily, the multi-receptive field feature extraction module is only set between the convolutional blocks G-C3 and G-C4 and between G-C4 and G-C5 for high-level feature extraction.

[0087] In the present invention, the multi-receptive field feature extraction module is designed based on the parallel structure of dilated convolution. While adjusting the network features and kernel functions based on convolution, combined with the idea of multi-scale convolution, five convolutional layers with a stride of 1 and three convolutional layers with a stride of 2 are alternately used to reduce the compression degree of the feature map, and while considering the computational cost, better retain the feature integrity of the input image.

[0088] Furthermore, considering that a basic assumption of convolutional neural networks is that all samples should use the same convolutional parameters, but the input of multi-source remote sensing image data fusion itself has differences, it is difficult to adapt to multi-source data with completely different feature representation forms through unified weight parameters. Therefore, if we want to achieve the cognition of multi-source remote sensing data, we cannot only consider improving the capacity of the model, increasing the parameters, depth, and number of channels of the model, because this will lead to an increase in the computational complexity of the model and an increase in the deployment difficulty. In the present invention, a mode of adaptively optimizing the convolutional kernel weight based on the original features of the image is designed, and a kernel weight adaptive optimization module is constructed in each convolutional block, and the convolutional kernel parameters are calculated through the input image to break the traditional static convolutional characteristics.

[0089] First, it is elaborated from a theoretical level. To improve the adaptability of convolutional operations for sample feature extraction, the present invention considers calculating the correlation between as many convolutional kernels and the image content as possible, and assigns different weights to convolutional kernels with different image contents. By transforming the image into a knowledge vector composed of multiple expert knowledge α 1 ,α 2 ,...,α n , where n is the self-set number of experts, which can be modified according to actual needs. The knowledge vector can represent the sample attributes, and the convolutional layer kernel functions are weighted by the knowledge vector to achieve the adaptability of convolution to samples. To more effectively improve the model capacity, the number of experts can be increased during the network design process, which can more effectively improve the network's cognition of samples than increasing the convolutional kernel size.

[0090] However, the computational cost of weighting samples to each convolutional kernel in the network is huge. As shown in Figure 5 (a), the repetitive weighted convolution process of the convolutional kernel and the original image content vector will slow down the model calculation process. Therefore, on this basis, the present invention considers combining the expert knowledge only once to ensure high-efficiency inference while improving the model capacity, as shown in Figure 5 (b). Convolutional operations have linear characteristics. In fact, the previous linear serial combination σ(α 1 (W 1 *x)+…+α n (W n *x)) can be transformed into the parallel section of the subsequent convolutional module σ((α 1 W 1 +…+α n W n )*x). Therefore, after the convolution transformation, n convolutional operations can be simplified to 1 convolutional operation, so as to realize that the expert knowledge is combined only once, and high-efficiency inference can be maintained while improving the model capacity.

[0091] As shown in Figure 6As shown, it is a schematic flow diagram of the adaptive convolution kernel weight update algorithm. Among them, N represents the batch size batch_size, H and h represent the widths of the corresponding images, W and w represent the lengths of the corresponding images, C and c represent the number of channels of the corresponding images. The adaptive convolution kernel weight update algorithm includes the following steps:

[0092] Perform global average pooling operation on the input N synthetic aperture radar images or visible light images of the preset batch quantity, and output (N, C).

[0093] Through the fully connected layer, calculate the weight coefficients of all image feature evaluation units for different input images, and output (N, mum_experts), where num_experts is the number of image feature evaluation units, which is equivalent to the number of experts described above. Then, normalize it to (0, 1) through the ReLU layer and output (N, num_experts) to obtain the scores of each image feature evaluation unit. For ease of understanding, the image feature evaluation unit can be regarded as a "unit" convolution kernel, similar to a unit vector used to construct an adaptive convolution kernel adapted to the full sample.

[0094] Multiply the obtained weight coefficients of each input image by the "unit" convolution kernel weights related to num_experts image feature evaluation units through matrix multiplication to perform corresponding updates on each "unit" convolution kernel, and output (N, h×w×c in ×c out ).

[0095] Perform segmentation in the batch quantity dimension, obtain each image in each batch, and perform convolution kernel weight weighting on each image to achieve the adaptability of the convolution kernel, and output *1, h×w×c in ×c out ).

[0096] Perform feature extraction and stacking on the original input images in the batch quantity dimension, and output the feature map (N, H, W, C) of the input images, and the dimension size of the output result is the same as that of the original input images.

[0097] Based on the above designs of the multi-receptive field feature extraction module and the kernel weight adaptive optimization module, the feature extraction module can complete feature mapping at different levels and realize an adaptive basic feature extraction network for multi-scale and multi-source images.

[0098] In some embodiments, the network parameter configuration in the feature extraction module is shown in Table 1:

[0099] Table 1

[0100]

[0101]

[0102] Among them, stride represents the step size, and @3*3 and @1*1 both represent the corresponding convolutional kernel sizes.

[0103] Based on the above-set feature fusion module, the synthetic aperture radar images of a preset batch quantity of data strips are input into the first branch of the dual-stream network to extract the synthetic aperture radar image feature map, and the visible light image is input into the second branch of the dual-stream network to extract the synthetic aperture radar image feature map.

[0104] In the feature extraction module designed above, the pseudo-symmetric dual-stream network can adapt to the features of images from different data sources, enabling the neural network with a pseudo-symmetric network structure to have the ability to extract cross-source features while forcing the extracted features to have consistency in the feature space. However, the feature fusion of multi-source remote sensing images has not been actually realized. Therefore, the present invention introduces an attention mechanism, injects the vector of the synthetic aperture radar image as an attention vector into the visible light image feature and fuses it with the visible light image feature, and injects the vector of the visible light image as an attention vector into the synthetic aperture radar image feature and fuses it with the synthetic aperture radar image feature, thereby completing the fusion of dual-stream features.

[0105] In the prior art, non-local attention can extract the information of non-local receptive fields, perform convolutional integration with self-features, and achieve the direct fusion of global information. The interaction between any two positions can directly capture long-range dependencies through calculation without considering the pixel position distance. This operation generates self-attention on the feature map in a global manner. Therefore, inspired by the non-local attention neural network, the present invention considers calculating the response of a certain position as the weighted sum of the features of all positions in the input feature map, that is, the feature response between corresponding pixels of multi-source images.

[0106] As Figure 7 shown, the feature fusion module based on the preset attention mechanism realizes the feature fusion of the synthetic aperture radar image feature map and the visible light image feature map, specifically including the following steps:

[0107] Use a preset convolutional layer to compress the number of channels of the synthetic aperture radar image feature map and the visible light image feature map. Exemplarily, the preset convolutional layer structure is 1×1×1; perform a linear mapping on the compressed synthetic aperture radar image feature map and visible light image feature map to obtain the first key matrix K S 、the first value matrix V S and the first query matrix Q H of the synthetic aperture radar image feature map, as well as the second key matrix K H 、the second value matrix V H and the second query matrix Q S; The first query matrix Q obtained from the synthetic aperture radar image feature map H is used for feature processing of the visible light image feature map, and the second query matrix Q obtained from the visible light image feature map S is used for feature processing of the synthetic aperture radar image feature map.

[0108] In the feature processing of the synthetic aperture radar image feature map, the subsequent features obtained based on the first key matrix K S and the first value matrix V S are subjected to matrix dot multiplication, and the first vector value in the range of [0, 1] is calculated through the Softmax function. The first vector value is dot product fused with the subsequent features obtained based on the second query matrix Q H and a convolution operation is performed. In the feature processing of the visible light image feature map, the subsequent features obtained based on the second key matrix K H and the second value matrix V H are subjected to matrix dot multiplication, and the second vector value in the range of [0, 1] is calculated through the Softmax function. The second vector value is dot product fused with the subsequent features obtained based on the first query matrix Q S and a convolution operation is performed.

[0109] Among them, the calculation process of feature fusion is shown in formula (3):

[0110] Attention(K, V, Q) = f(Softmax(f T (K) · f(V)) · f T (Q other )); (3)

[0111] Among them, K, V, Q represent learnable key matrix, value matrix and query matrix; (·) T represents transpose.

[0112] The input synthetic aperture radar image feature map is dot multiplied with the corresponding fused features again to obtain the relationship between each pixel of the synthetic aperture radar image and the pixels of the visible light image; the input visible light image feature map is dot multiplied with the corresponding fused features again to obtain the relationship between each pixel of the visible light image and the pixels of the synthetic aperture radar image.

[0113] Finally, the fused features corresponding to the synthetic aperture radar image feature map and the visible light image feature map are superimposed through the feature layer, and the fused feature vector is output. The calculation formula is shown in formula (4):

[0114] F fusion = concat(g(Attention S ) + g(AttentionH )); (4)

[0115] where g represents the convolution operation; Attention S represents the feature after the fusion of synthetic aperture radar image feature maps; Attention H represents the feature after the fusion of visible light image feature maps.

[0116] Based on the above design, in the feature fusion module, the synthetic aperture radar image feature map and the visible light image feature map are processed by the symmetric attention mechanism, and the global correlation is obtained by using matrix multiplication and performing the Softmax operation on the result, obtaining the attention-weighted feature map between the two images, endowing pixel-level cross-attention between heterogeneous data, realizing the construction of the multi-source image feature circulation channel, and finally inputting the fusion feature vector generated by high-dimensional fusion through the feature layer superposition method into the fusion image generation module to generate the corresponding fusion image.

[0117] The fusion image generation module generates a color image with the material information of the synthetic aperture radar image and including the color information of the visible light image according to the fusion feature vector, and restores it to a fusion image with the same size as the input image through decoding.

[0118] The premise of image fusion is to accurately extract the features of multi-source images through effective learning, synthesize these multi-source information, and improve the interpretability of the fusion image. As Figure 8 shown, in order to effectively fuse synthetic aperture radar images and visible light images, the present invention is based on the Conditional Generative Adversarial Networks (cGAN), takes the synthetic aperture radar image as the conditional input, and constructs a generative adversarial network with the fusion image generation module as the generator and discriminator.

[0119] In step S103, the positive and negative sample pairs are constructed based on the fusion image, the synthetic aperture radar image, and the visible light image to train the generator and discriminator. The discriminator designed by the present invention can not only be used to measure the similarity between the generated fusion image and the input visible light image, but also be used to measure the similarity between the mosaic of the synthetic aperture radar image and the visible light image and the mosaic of the synthetic aperture radar image and the fusion image, so as to objectively reflect the quality of the fusion image and promote the continuous optimization of the generator. At the same time, following the idea of cGAN, taking the synthetic aperture radar image as the conditional input, guiding the generator to generate high-quality fusion images.

[0120] In some embodiments, the discriminator consists of 6 convolutional layers, and each convolutional layer performs batch normalization and LeakyReLU processing, which can avoid network instability caused by overfitting and gradient sparsity during the training process. The output result of the discriminator generates the discriminant result of the authenticity of the fused image through a preset Sigmoid layer. The generator follows the transposed convolution operation, corresponding to the feature extraction module, and consists of 5 transposed convolutional layers. Similarly, each transposed convolutional layer performs normalization and Leaky ReLU processing.

[0121] The network parameter configurations of the generator and the discriminator are shown in Tables 2 and 3 respectively:

[0122] Table 2

[0123]

[0124] Table 3

[0125]

[0126]

[0127] Based on the above designs of each module and the parameter settings of each structure, the initial fused image model is trained to obtain a fused image model that can be used to generate fused images of synthetic aperture radar images and visible light images. Among them, the training process of the fused image generation module in the initial fused image model is relatively more important. In the present invention, based on the conditional generative adversarial network, the fused image generation module is used as the generator to construct a generative adversarial network with the discriminator, and is trained in an alternating training manner, specifically including the following steps:

[0128] First, train the discriminator, which can be denoted as D. Input the fused image generated by the generator into the discriminator, and the fused image can be denoted as I F , and the length, width, and number of channels of the fused image are (W, H, C) respectively. Input the fused image I F and stack it with the channels of the synthetic aperture radar image, and the synthetic aperture radar image can be denoted as I SAR , thus, obtaining the spliced image (I SAR , I F ). Since the fused image I F is an image generated by the generator and belongs to a non-real image, therefore, the true value of (I SAR , I F ) is 0, that is, a negative sample, and input the spliced image (I SAR , I F ) into the discriminator. Stack the channels of the synthetic aperture radar image I SAR with the channels of the visible light image, and the visible light image can be denoted as I OPT , thus, obtaining the spliced image (ISAR , I OPT ), the synthetic aperture radar image I SAR and the visible light image I OPT are both real data in the multi-source remote sensing dataset and belong to real images. Therefore, (I SAR , I OPT ) corresponds to a true value of 1, that is, a positive sample. The spliced image (I SAR , I OPT ) is input into the discriminator. The discriminator is trained using the above two types of data with the aim of enabling the discriminator to learn to distinguish real images from non-real images composed of generated images.

[0129] Then, the generator is trained. The generator can be denoted as G. The non-real image (I SAR ) spliced from the synthetic aperture radar image I F and the fused image I SAR , I F ) is input into the discriminator, and the discriminator parameters are kept fixed while the generator is trained. Among them, since the generator is required to generate positive samples close to real images, the true value of the input negative sample (I SAR , I F ) is modified to 1 to ensure that the discriminator does not update its parameters. During the training of the generator, only by continuously adjusting the parameters to make the generated fused image I F get closer and closer to the real image can the loss value obtained by the discriminator be continuously reduced, so as to achieve the purpose of training and optimizing the generator.

[0130] During the training process, the loss function is used to measure the difference between the forward prediction result and the target result. Iteratively optimizing the network parameters is a process of continuously decreasing the loss function. During the network training process, it is forced to make the difference between the fused image generated by the generator and the real image smaller and smaller, and to gradually enhance the discriminative ability of the discriminator for real images and the generated fused images.

[0131] In some embodiments, the loss function of the discriminator adopts the cross-entropy between the discriminator output result and the true value, as shown in formula (5):

[0132]

[0133] where, E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I F represents the fused image; I OPT represents the visible light image; D(I SAR , I F ) represents the discriminative result obtained by inputting the non-real image (I SAR , I F ) into the discriminator; D(ISAR , I OPT ) represents the discrimination result obtained by inputting the real image (I SAR , I OPT ) into the discriminator.

[0134] In some embodiments, a joint loss of the generator is constructed based on a generative adversarial loss, a mean absolute error loss, and a texture loss.

[0135] The generative adversarial loss is the cross-entropy between the result obtained by using the output image of the generator as the input to the discriminator and the ground truth, and is used to measure the quality of the result generated by the generator, that is, the probability of being determined as a real image. The calculation formula is shown in formula (6):

[0136]

[0137] where E[·] represents taking the average expectation; I SAR represents a synthetic aperture radar image; I OPT represents a visible light image; D(·) represents the discrimination result obtained by the discriminator; G(·) represents the fused image generated by the generator.

[0138] The mean absolute error loss is an l1 loss function, which is used to force the fused image to be closer to the visible light image in terms of color information. The calculation formula is shown in formula (7):

[0139]

[0140] where E[·] represents taking the average expectation; I SAR represents a synthetic aperture radar image; I OPT represents a visible light image; I F represents the fused image.

[0141] Considering that the lightness component of a color image only reflects the degree of light and dark contrast of the color and contains less texture information, in order to make the texture features of the fused image I F closer to those of the synthetic aperture radar image I SAR , the present invention introduces a texture loss function based on a gray-level co-occurrence matrix (GLCM). The L1 norm of the gray-level co-occurrence matrices of the fused image I F and the synthetic aperture radar image I SAR is used as a measure of texture similarity.

[0142] The hue F , saturation , and value component of the fused image I are extracted by using the HSV color model transformation, and the hue SAR , saturation of the synthetic aperture radar image I and the brightness component Calculate the fused image I F brightness component and synthetic aperture radar image I SAR gray-level co-occurrence matrix of the brightness component, as shown in formula (8):

[0143]

[0144] wherein, represents the gray-level co-occurrence matrix of the brightness component of the fused image I F gray-level co-occurrence matrix of the brightness component; represents the gray-level co-occurrence matrix of the brightness component of the synthetic aperture radar image I SAR gray-level co-occurrence matrix of the brightness component.

[0145] Use the L1 norm to constrain the similarity of texture features, and the calculation formula is as shown in formula (9):

[0146]

[0147] wherein, E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I OPT represents the visible light image; I F represents the fused image; represents the gray-level co-occurrence matrix of the brightness component of the fused image; represents the gray matrix of the brightness component of the synthetic aperture radar image.

[0148] Weight the generative adversarial loss, mean absolute error loss and texture loss to obtain the combined loss of the generator, and the calculation formula is as shown in formula (10):

[0149]

[0150] wherein, L GAN (G) represents the generative adversarial loss; represents the mean absolute error loss; λ is a hyperparameter, preferably, λ = 100; L GLCM (G) represents the texture loss; μ represents the weight coefficient of the texture loss, and the theoretical value of μ = 100.

[0151] By alternately training the generator and the discriminator, based on deep neural network parameter optimization methods such as Adam or SDG (stochastic gradient descent), optimize the network model parameters, obtain the fusion image generation network parameters adapted to synthetic aperture radar images and visible light images, and obtain the final fusion image model. Experimental data proves that the fusion image obtained by training in the present invention is also adapted to the generation of fusion images of other multi-source remote sensing data.

[0152] The present invention also provides a method for generating a fused image for multi-source remote sensing data, the method comprising the following steps S201 to S202:

[0153] Step S201: Obtain a synthetic aperture radar image and a visible light image in the same area at the same time generated based on a preset registration algorithm to be fused. Among them, the number of channels of the synthetic aperture radar image is replicated to be equal to the number of channels of the visible light image.

[0154] Input the synthetic aperture radar image and the visible light image into the fused image model obtained by the method for training a fused image model for multi-source remote sensing data described above to obtain a fused image of the synthetic aperture radar image and the visible light image.

[0155] Proved by experimental data, the fused image trained by the present invention is also applicable to the generation of fused images of other multi-source remote sensing data. Therefore, the method for generating a fused image for multi-source remote sensing data provided by the present invention is also applicable to the generation of fused images of other multi-source remote sensing data.

[0156] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method for training a fused image model for multi-source remote sensing data and the method for generating a fused image for multi-source remote sensing data.

[0157] Corresponding to the above method, the present invention also provides a device, which includes a computer device. The computer device includes a processor and a memory. A computer instruction is stored in the memory, and the processor is used to execute the computer instruction stored in the memory. When the computer instruction is executed by the processor, the device implements the steps of the method described above.

[0158] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the aforementioned method for deploying an edge computing server. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0159] In summary, the present invention provides a method, a generation method and a device for training a fusion image model for multi-source remote sensing data. By constructing a fusion image model, features of multi-source remote sensing data with differences are extracted and fused to generate a fusion image, which makes up for the limitations of a single sensor itself, realizes information complementarity between multi-source remote sensing data, and improves the information content and readability of the fusion image. The fusion image model includes a feature extraction module, a feature fusion module and a fusion image generation module. The feature extraction module is a dual-stream network, and each branch adopts a preset deep convolutional neural network with the same structure. Based on the anti-noise interference and translational invariance of the deep convolutional neural network, reliable fusion image generation is realized. A pseudo-symmetric network is constructed, and a kernel weight adaptive optimization module is introduced into the pseudo-symmetric network, forcing the feature extraction network with shared parameters to adaptively optimize features according to the characteristics of the input image while learning the consistency of multi-source images. The feature fusion module performs information feature fusion of multi-source images based on the attention mechanism, and performs weighted superposition after calculating the feature correlation, better emphasizing the complementary information in the multi-source images and outputting fusion features with more complete feature semantic expressions. The fusion image generation module constructs a conditional generative adversarial network as a generator and a discriminator, and performs supervised training. In addition to constructing a conventional loss, texture information of the image is mined through a gray-level co-occurrence matrix to construct a texture loss, and supervised training is performed based on both color and texture, so that the finally generated fusion image retains both the texture information in the synthetic aperture radar image and the visible light spectrum information. The data information of the fusion image generated based on the fusion image model can provide higher-quality and more information-rich basic data for subsequent remote sensing data interpretation tasks.

[0160] Those of ordinary skill in the art should understand that the various exemplary components, systems and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0161] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0162] In the present invention, features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0163] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and variations can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for training a fusion image model for multi-source remote sensing data, characterized in that, the method comprises the following steps: Obtain a multi-source remote sensing data set, where the multi-source remote sensing data set contains multiple data strips, and each data strip includes a synthetic aperture radar image and a visible light image of the same area at the same time generated based on a preset registration algorithm; the number of channels of the synthetic aperture radar image is replicated to be equal to the number of channels of the visible light image; Obtain an initial fusion image model and a discriminator; the initial fusion image model includes a feature extraction module, a feature fusion module, and a fusion image generation module; the feature extraction module is a dual-stream network, and each branch uses a preset neural network with the same structure. The preset neural network is composed of multiple convolutional blocks, and the parameters of the first set number of convolutional blocks on the two branches are kept symmetric. Each convolutional block is provided with a kernel weight adaptive optimization module; wherein, the weight adaptive optimization module performs global average pooling operation on the input synthetic aperture radar images or visible light images, calculates the weight coefficients of the corresponding images based on multiple preset image feature evaluation units, and uses the weight coefficients to weight the convolutional kernel weights corresponding to the image feature evaluation units to update the convolutional kernel weights in the corresponding convolutional blocks; the fusion image generation module, as a generator, constructs a generative adversarial network with the discriminator; Input the synthetic aperture radar images of a preset batch number of data strips into the first branch of the dual-stream network to extract synthetic aperture radar image feature maps, and input the visible light images into the second branch of the dual-stream network to extract visible light image feature maps; Input the synthetic aperture radar image feature maps and the visible light image feature maps into the feature fusion module. Based on a preset attention mechanism, extract their respective features by exchanging the query matrices of the synthetic aperture radar image feature maps and the visible light image feature maps, and use the feature layer to superimpose the features extracted from the synthetic aperture radar image feature maps and the visible light image feature maps to output a fusion feature vector; Input the fusion feature vector into the fusion image generation module to generate a fusion image; Stitch the synthetic aperture radar images in each data strip with the corresponding generated fusion images into non-real images, marked as negative samples; stitch the synthetic aperture radar images and the visible light images into real images, marked as positive samples; input the non-real images and the real images into the discriminator for training; input the non-real images into the trained discriminator, fix the discriminator parameters, and train the generator according to the discrimination results of the discriminator; respectively construct the loss functions of the generator and the discriminator, and train to obtain a generator and a discriminator that meet the preset performance, so as to finally obtain a fusion image model.

2. The method for training a fusion image model for multi-source remote sensing data according to claim 1, characterized in that, The two-stream network constructs a pseudo-symmetric network based on the Siamese network; each branch uses a preset neural network with the same structure, and the preset neural network is composed of 5 convolutional blocks; the parameter settings of the first 3 convolutional blocks on the two branches are the same, and the parameter settings of the last 2 convolutional blocks are different.

3. The method for training a fusion image model for multi-source remote sensing data according to claim 1, wherein, a multi-receptive field feature extraction module is further provided between each convolutional block of the preset neural network; the image feature map processed by the current convolutional block is input into the multi-receptive field feature extraction module, and a preset number of parallel convolutional layers are used to extract and stack the features of the preset number of different degrees of receptive fields of the image feature map from local to global, and then input into the next convolutional block; among them, the calculation formula of the receptive field can be expressed as: Among them, RF l represents the receptive field size of layer l; K' l represents the convolutional kernel size of layer l; S i represents the stride of the i-th layer.

4. The method for training a fusion image model for multi-source remote sensing data according to claim 1, wherein, based on a preset attention mechanism, by exchanging the query matrices of the synthetic aperture radar image feature map and the visible light image feature map respectively to extract their respective features, it further includes: using a preset convolutional layer to compress the number of channels of the synthetic aperture radar image feature map or the visible light image feature map, and performing a linear mapping to obtain the first key-value matrix, the first value matrix and the first query matrix of the synthetic aperture radar image feature map, and the second key-value matrix, the second value matrix and the second query matrix of the visible light image feature map; using the first query matrix for the feature processing of the visible light image feature map; using the second query matrix for the feature processing of the synthetic aperture radar image feature map; in the feature processing of the synthetic aperture radar image feature map, perform matrix dot multiplication on the subsequent features obtained based on the first key-value matrix and the first value matrix, calculate the first vector value through the Softmax function, perform dot product fusion on the first vector value and the subsequent features obtained based on the second query matrix, and perform a convolutional operation; in the feature processing of the visible light image feature map, perform matrix dot multiplication on the subsequent features obtained based on the second key-value matrix and the second value matrix, calculate the second vector value through the Softmax function, perform dot product fusion on the second vector value and the subsequent features obtained based on the first query matrix, and perform a convolutional operation; the corresponding calculation process can be expressed as: Attention(K, V, Q) = f(Softmax(f T (K) · f(V)) · f T (Q other )); Among them, K, V, and Q represent learnable key-value matrix, value matrix, and query matrix; (·) T represents transpose.

5. The method for training a fusion image model for multi-source remote sensing data according to claim 4, wherein, using a feature layer to stack the features extracted from the synthetic aperture radar image feature map and the visible light image feature map, and output a fusion feature vector, and the calculation formula is: F fusion = concat(g(Attention S ) + g(Attention H )); Among them, g represents the convolution operation; Attention S represents the feature after the fusion of the synthetic aperture radar image feature maps; Attention H represents the feature after the fusion of the visible light image feature maps.

6. The method for training a fusion image model for multi-source remote sensing data according to claim 1, wherein, The fused image generation module serves as a generator to construct a generative adversarial network with the discriminator; the generative adversarial network adopts a conditional generative adversarial network, takes the synthetic aperture radar image as a conditional input, and optimizes the generator to generate a fused image; the discriminator consists of 6 convolutional layers, and each convolutional layer has a batch normalization layer and a rectified linear unit activation layer; The output result of the discriminator generates a discrimination result on the authenticity of the fused image through a preset Sigmoid layer.

7. The method for training a fused image model for multi-source remote sensing data according to claim 1, wherein, Constructing the loss functions of the generator and the discriminator respectively further includes: The calculation formula of the loss function of the discriminator is: Among them, E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I F represents the fused image; I OPT represents the visible light image; D(I SAR , I F ) represents the discrimination result obtained by inputting the non-real image (I SAR , I F ) into the discriminator; D(I SAR , I OPT ) represents the discrimination result obtained by inputting the real image (I SAR , I OPT ) into the discriminator.

8. The method for training a fused image model for multi-source remote sensing data according to claim 1, wherein, Constructing the loss functions of the generator and the discriminator respectively further includes: Constructing a joint loss of the generator based on generative adversarial loss, mean absolute error loss, and texture loss; The calculation formula of the generative adversarial loss is: Among them, E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I OPT represents the visible light image; D(·) represents the discrimination result obtained by the discriminator; G(·) represents the fused image generated by the generator; The calculation formula of the mean absolute error loss is: where E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I OPT represents the visible light image; I F represents the fused image; The calculation formula of the texture loss is: where, E[·] represents taking the average expectation; I SAR represents the synthetic aperture radar image; I OPT represents the visible light image; I F represents the fused image; represents the gray-level co-occurrence matrix of the lightness component of the fused image; represents the gray-level co-occurrence matrix of the lightness component of the synthetic aperture radar image; Weight the generative adversarial loss, the mean absolute error loss, and the texture loss to obtain the loss function of the generator, and the calculation formula is: Among them, L GAN (G) represents the generative adversarial loss; represents the mean absolute error loss; λ is a hyperparameter; L GLCM (G) represents the texture loss; μ represents the weight coefficient of the texture loss.

9. A method for generating a fused image for multi-source remote sensing data, wherein, The method includes the following steps: Obtain a synthetic aperture radar image and a visible light image in the same area at the same time generated based on a preset registration algorithm to be fused; the number of channels of the synthetic aperture radar image is replicated to be equal to the number of channels of the visible light image; Input the synthetic aperture radar image and the visible light image into the fused image model obtained by the method for training a fused image model for multi-source remote sensing data according to any one of claims 1 to 8 to obtain a fused image of the synthetic aperture radar image and the visible light image.

10. A computer-readable storage medium, on which a computer program is stored, wherein, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Remote sensing image change detection method and system based on adversarial double-auto-encoder network

    CN115063359A

  • Infrared and visible light image fusion method based on multi-discriminator generative adversarial network

    CN115601282A