A multi-source remote sensing image processing method and system
Through the multi-source remote sensing image processing method based on the Resnet-34 network, using technical means such as multi-modal residual attention fusion blocks, the problem of low feature extraction and classification efficiency of high-score images and Lidar data in multi-source remote sensing data is solved, and more efficient feature extraction and classification effects are achieved.
Patent Information
- Application Number
- CN202310516741.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-05-09
AI Technical Summary
The prior art cannot efficiently realize the feature extraction and classification of high-score image data and Lidar data in multi-source remote sensing data.
Using a multi-source remote sensing image processing method based on Resnet-34 network, the spatial structure and elevation features of Lidar and high-score image data are extracted and fused to achieve efficient feature extraction and classification of multi-source data through multi-modal residual attention fusion block, improved SE fusion module and multi-layer feature global context information learning module.
The feature extraction and classification efficiency of Lidar data and high-score image data is improved, the performance of the network in certain tasks is improved, and the classification accuracy is improved while ensuring the computational volume.
Smart Images

Figure CN116452819B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image data processing, and in particular to a method and system for extracting and classifying spatial features of multi-source remote sensing images. Background Art
[0002] High Resolution Remote Sensing can provide detailed images with meter-level or even sub-meter-level spatial resolution, which can clearly express the spatial structure and surface texture characteristics of ground objects. Advanced LiDAR systems can excellently supplement the elevation information of high-resolution (HR) images. As the research frontier in the field of remote sensing, high-resolution image classification and LiDAR point cloud data classification play an important role in a wide range of applications such as urban planning, environmental monitoring, agricultural and forestry management, and land inventory. Semantic segmentation technology assigns a label to each pixel in an image and has been applied to precision agriculture, urban planning, and natural resource inventory.
[0003] With the rapid development of remote sensing technology, it is more convenient to obtain multi-source remote sensing data. The fusion classification of high-resolution data and Lidar data is a feasible research direction. Multi-source remote sensing data fusion can be divided into three levels: pixel level, feature level and decision level. Feature-level fusion refers to the feature extraction of remote sensing image data before fusion to generate feature vectors, and then feature fusion to combine the features meaningfully, which has high credibility and accuracy.
[0004] At present, research on the collaborative application of Lidar data and high-resolution images focuses on feature extraction and classification of objects, mainly on texture, structure, color, intensity and height information. It is impossible to efficiently realize the feature extraction and classification of high-resolution image data and Lidar data in multi-source remote sensing data. Summary of the invention
[0005] In view of the problems existing in the above-mentioned fields, the present invention proposes a multi-source remote sensing image processing method and system, which can solve the technical problem of not being able to efficiently realize feature extraction and classification of high-resolution data and Lidar data in multi-source remote sensing data.
[0006] In order to solve the above technical problems, the present invention discloses a multi-source remote sensing image processing method, comprising the following steps:
[0007] Obtain multiple data enhancement sample matrices of high-resolution image data and Lidar image data to be fused;
[0008] According to the acquired multiple data enhancement sample matrices, the spatial structure of Lidar and high-resolution image data is extracted based on the Resnet-34 network; the spatial information and elevation features of Lidar image data are extracted using Lidar data information, and the spatial features of high-resolution image data are extracted using high-resolution image data information;
[0009] Constructing a multimodal residual attention fusion block, the multimodal residual attention fusion block includes a fusion module and a deep separable residual module; inputting spatial features and elevation features into the fusion module, using the fusion module to add and fuse the channel weights of the elevation features and spatial features obtained, and inputting them into the deep separable residual module for the next step of feature extraction to obtain multimodal feature extraction data;
[0010] An improved SE fusion module is proposed. The improved SE fusion module includes two fully connected layers. The fully connected layers are used to compress and restore the channel dimensions of different stages of the acquired multimodal feature extraction data, learn the mutual correlation between different channels, and achieve feature reuse at different stages by fusing the feature information between channels of multimodal feature extraction data at different stages of the network, so as to obtain cross-layer fused output data.
[0011] Constructing a multi-layer feature global context information learning module FFD, wherein the multi-layer feature global context information learning module FFD includes a common convolution layer and a Fourier frequency domain transform layer; adding and fusing the obtained cross-layer fusion output data with the sampling data of the adjacent previous stage, inputting the data into the common convolution layer for spatial feature extraction, and then using the Fourier frequency domain transform layer to convert the image into the frequency domain to obtain the global context information in the frequency domain;
[0012] The acquired global context information in the frequency domain is sent to the classifier for classification to obtain the classification result.
[0013] Preferably, the step of obtaining a plurality of data enhancement sample matrices of the high-resolution image data and Lidar image data to be fused comprises the following steps:
[0014] Obtain the original image matrix of high-resolution image data and Lidar image data;
[0015] According to the original image data of the original image matrix, data enhancement samples are obtained by randomly flipping in the vertical direction and the horizontal direction, respectively, wherein the flip angles in the horizontal direction and the vertical direction are both in the range of [0°, 270°].
[0016] Preferably, the multimodal residual attention fusion block, i.e., the RAF module, is constructed to fuse the lateral features of the RGB encoder to the corresponding lateral features of the DSM encoder, and multi-level feature fusion is used for both RGB and DSM encoder features, and then cross-modal fusion is performed. Let D mrepresents the output of cross-modal fusion, then the output feature calculation formula of cross-modal fusion is:
[0017]
[0018] in, represents the side output encoder features of the m-th encoder of the RGB branch, represents the side output encoder feature of the m-th encoder of the DSM branch, m represents the m-th feature encoder, and RAF represents the multimodal residual attention fusion block.
[0019] Preferably, the learning of the correlation between different channels includes learning the correlation between the fifth-stage multimodal feature extraction data and the previous four-stage multimodal feature extraction data channels, fusing the feature information between the multimodal feature data channels at different stages of the network, generating high-resolution semantic instructions by fusing the high-level features of the fifth layer and the cross-modal fusion features of the first to fourth layers obtained from the multimodal residual attention fusion block, and sequentially fusing the cross-modal fusion features of the fifth layer into the cross-modal fusion features of the first to fourth layers to achieve feature reuse of the fifth layer. The formula for SE fusion of multimodal feature extraction data at different stages is:
[0020] F=SE(Concat[D m ,D 5 ]),m=1,......,4
[0021] Among them, F represents the output of cross-layer fusion, Concat[·] represents the concatenation operation, and D m represents the output of cross-modal fusion, D 5 Represents the output of the fifth cross-modal fusion.
[0022] Preferably, the construction of the multi-layer feature global context information learning module FFD comprises the following steps:
[0023] The context information of the two adjacent stages, the output data of the cross-layer fusion and the upsampled data of the adjacent previous stage, is sent to the 3×3 convolution for shallow feature extraction, and then sent to the 2D FFT, Hadamard product and 2D inverse FFT to combine the spatial information and the frequency domain information to extract the context information of the adjacent stages;
[0024] The context information of the extracted adjacent stages is sent to the 3×3 convolution for secondary information extraction to obtain the global context information in the frequency domain;
[0025] A 1×1 convolution is used to adjust the dimension of the channel.
[0026] Preferably, a multi-source remote sensing image processing system is also included, comprising:
[0027] The data acquisition module is used to acquire multiple data enhancement sample matrices of the high-resolution image data and Lidar image data to be fused; extract the spatial structure of the Lidar data based on the Resnet-34 network according to the acquired multiple data enhancement sample matrices; extract the spatial information and elevation features of the Lidar image data using the Lidar data information, and extract the spatial features of the high-resolution image data using the high-resolution image data information;
[0028] The multimodal residual attention fusion block is used to input the spatial features and elevation features into the feature fusion module, and use the fusion module to add and fuse the channel weights of the obtained elevation features and spatial features, and input them into the deep separable residual module for the next step of feature extraction to obtain multimodal feature extraction data;
[0029] The improved SE fusion module includes two fully connected layers, which compress and restore the channel dimensions of different stages of the acquired multimodal feature extraction data through the fully connected layers, learn the correlation between different channels, and achieve feature reuse at different stages by fusing the feature information between channels of multimodal feature extraction data at different stages of the network to obtain cross-layer fused output data;
[0030] The multi-layer feature context information learning module FFD includes a common convolution layer and a Fourier frequency domain transform layer, which is used to add and fuse the obtained cross-layer fusion output data with the sampling data of the adjacent previous stage, input it into the common convolution layer for spatial feature extraction, and then use the Fourier frequency domain transform layer to convert the image into the frequency domain to obtain the global context information in the frequency domain;
[0031] The classification module sends the acquired global context information in the frequency domain into the classifier to obtain the classification result.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The present invention uses the Resnet-34 network as an encoder, which can realize the extraction of the spatial structure of Lidar data information and high-resolution image data information; the constructed multimodal residual attention fusion block uses a fusion module to add and fuse the channel weights of the obtained elevation features and spatial features to obtain multimodal feature extraction data, providing data support for the information fusion of the improved SE fusion module; through the constructed global context information fusion module FFD, frequency domain information is obtained, which can improve the performance of CNN in certain tasks. By comparing with the classical method, it is concluded that the method proposed by the present invention improves the network classification accuracy while ensuring the amount of calculation, and improves the feature extraction and classification efficiency of Lidar data and high-resolution image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the process of the present invention;
[0035] Figure 2 It is the overall network framework diagram of the present invention;
[0036] Figure 3 Schematic diagram of the multimodal residual attention fusion structure of the present invention;
[0037] Figure 4 It is a schematic diagram of the improved SE module fusion structure of the present invention;
[0038] Figure 5 It is a schematic diagram of the global context information fusion structure of the present invention. DETAILED DESCRIPTION
[0039] The following will be combined with the attached embodiment of the present invention Figure 1-5 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms described in the present invention are only used to describe specific implementation methods and are not used to limit the present invention.
[0040] In complex scenes, feature fusion based on simple splicing often fails to achieve better classification performance because it ignores the complementarity between multi-source data, and this method will further increase the dimension of the features, resulting in increased computational costs. At the same time, the features extracted in the deep layer during CNN network feature extraction, that is, the high-level features, have higher semantic information, but lower spatial resolution. The shallower features, that is, the low-level features, have higher spatial information and contain lower semantic information of building features. Due to the up-sampling and down-sampling process in the encoder-decoder network, the information of these features has been lost. How to reduce the loss of feature information and further fuse deep features with shallow features is still a problem that needs further research. Since the interdependence of global context is crucial to the representation of multi-source data, the existing multi-source remote sensing data fusion methods pay less attention to context dependency and have limitations in multi-scale feature extraction of remote sensing images.
[0041] Example
[0042] like Figure 1As shown, the present invention proposes a multi-source remote sensing image processing method, which extracts the spatial top features of high-resolution image data and the spatial information and elevation features of Lidar respectively through a Resnet-34-based extraction module, and fuses the weights obtained from the Lidar image data and the weights obtained from the high-resolution image data through a multimodal residual attention fusion module to achieve preliminary multi-source data fusion; the fused features are input into a deep separable convolution residual block to extract fused features in the Lidar image data; the SE fusion module is used to fuse the features of the last stage with the features of the previous stages to achieve feature reuse; the global context frequency domain change module is used to establish a connection between the global spatial information and the frequency domain information, and the global context information of the frequency domain is learned; the final fused features are sent to the classifier through the classification module to obtain the classification results.
[0043] The specific steps of a multi-source remote sensing image processing method proposed by the present invention are as follows:
[0044] Step S1: Acquire multiple data enhancement sample matrices of high-resolution image data and Lidar image data to be fused;
[0045] Step S2: using the Resnet-34 network to extract the basic spatial structure of the Lidar and high-resolution image data; after inputting multiple data enhancement sample matrices of the high-resolution image data and the Lidar image data into the Resnet-34 network feature extraction module respectively, the spatial information and elevation features of the Lidar image data are extracted using the Lidar data information, and the spatial features of the high-resolution image data are extracted using the spatial features of the high-resolution image;
[0046] Step S3: construct a multimodal residual attention fusion block, which includes a fusion module and a deep separable residual module, input the spatial features and elevation features into the fusion module, use the fusion module to add and fuse the channel weights of the obtained elevation features and spatial features, and input them into the deep separable residual module for the next step of feature extraction, so as to better fuse the elevation information of Lidar and the spatial information of high-resolution, so as to make full use of each modal feature from the multimodal data and obtain multimodal feature extraction data;
[0047] Step S4: Propose an improved SE fusion module and construct feature fusion modules at different stages, which mainly fuse the channel information of feature extraction data at different stages. It includes two FC fully connected layers and a rectified linear unit (ReLU) layer. The first fully connected layer FC mainly compresses the channel information of the fifth stage feature extraction data obtained above and the channel information of the feature extraction data obtained from the first to fourth stages in the channel dimension, and activates it through the rectifier layer. The second fully connected layer FC is used to restore the number of compressed channels and perform downstream tasks. The fully connected layer compresses and restores the channel dimension of the obtained multimodal feature extraction data at different stages, learns the mutual correlation between different channels, and realizes feature reuse at different stages by fusing the feature information between channels of multimodal feature extraction data at different stages of the network, and obtains cross-layer fused output data.
[0048] Step S5: Construct a multi-layer feature global context information learning module FFD, which includes a common convolution layer and a Fourier frequency domain transform layer. The obtained cross-layer fusion output data is added and fused with the sampling data of the adjacent previous stage, and input into the common convolution layer for spatial feature extraction. Then, the Fourier frequency domain transform layer is used to convert the image to the frequency domain, obtain the global context information in the frequency domain, and change the spatial information to the frequency domain.
[0049] Step S6: Send the finally acquired global context information in the frequency domain to the classifier to obtain a classification result.
[0050] In step S1, multiple enhanced data of high-resolution image data and Lidar image data to be fused are obtained, including:
[0051] Obtain the original image matrix of high-resolution image data and Lidar image data;
[0052] According to the original image data in the original image matrix, data enhancement samples are obtained by randomly flipping in the vertical and horizontal directions, respectively, where the flip angles in the horizontal and vertical directions are both in the range of [0°, 270°].
[0053] In step S2, a dual-branch encoder-decoder backbone network, such as Figure 2As shown. The main features include: The encoder part of each branch of Lidar is mainly based on the ResNet-34 network. It consists of a series of residual blocks, whose inputs are added to the outputs through skip connections. The RGB and DSM streams have the same network settings, except that the first convolution block on the DSM branch has only one channel because the DSM input is displayed as one channel. The first encoder block consists of a 7×7 convolution layer with a stride of 2, a batch normalization (BN) layer, and a rectified linear unit (ReLU) layer. The second encoder is a 3×3 maximum pooling layer with a stride of 2 and the first residual layer of the Resnet-34 network, which enables our network to flexibly handle different input branches, and then the remaining three residual layers are added in sequence. Except for the first residual layer, the other three residual layers have a residual unit, which downsamples the feature map and increases the feature channel by 2 times. Therefore, in the encoder network, the final output resolution is 1 / 32 of the original input data. The structure of the decoder part is symmetrical to the encoder part mentioned above. Five deconvolution layers with 2× upsampling residual units are sequentially added at the end of the encoder part so that the size of the final output is equal to the input.
[0054] The present invention uses an FFD Fourier frequency domain transform layer to add the features of these encoders to the features of the decoder, so that the decoder network can obtain a more refined feature map, and the resulting network has a typical FCNs output.
[0055] In step S3, the multimodal residual attention fusion block, such as Figure 3 As shown in the figure, it consists of a data fusion part and a feature extraction part. The data fusion module contains different convolution and sigmoid layers to predict the importance weights. The feature extraction part uses a depth-separable convolution residual block to reduce the amount of calculation while ensuring the extraction accuracy of the network and preventing the gradient disappearance phenomenon.
[0056] Since the features generated by the DSM encoder have little edge information, this paper designs the RAF module to fuse the horizontal features of the RGB encoder into the corresponding horizontal features of the DSM encoder. m represents the output of cross-modal fusion, then the output features of cross-modal fusion can be mathematically expressed as:
[0057]
[0058] Among them, RAF represents the feature fusion module, represents the side output encoder features of the m-th encoder of the RGB branch, represents the side output encoder feature of the m-th encoder of the DSM branch, and m represents the m-th feature encoder.
[0059] In practical applications, a multi-level feature fusion method is used to fuse deep (fifth layer) features and shallow features (first to fourth layers). By fusing the high-level features of the deeper layer (fifth layer) and the shallower layer (first to fourth layers) features obtained from the encoder part of the model, high-resolution semantic instructions are generated. The multi-level feature fusion formula of SE fusion features is:
[0060] F=SE(Concat[D m ,D 5 ]),m=1,......,4 (2)
[0061] Among them, F represents the output of cross-layer fusion, Concat[·] represents the concatenation operation, and D m represents the output of cross-modal fusion, D 5 Represents the output of the fifth cross-modal fusion.
[0062] Data fusion includes two inputs, namely Lidar feature X l and high-scoring feature X h , use two 1×1 convolutions to adjust the number of channels for residual connection, and convert the Lidar feature X l , high-scoring feature X h After splicing, 7×7 convolution is used for shallow feature extraction, and then two 1×1 convolutions are used for residual connection. The Sigmoid activation function is used to obtain the attention weight. The Lidar feature X l , High Score X h Multiply and concatenate with (1-weight) and weight to get the fused feature map. The process can be expressed as:
[0063] F = Concat[X h ,X l ] (3)
[0064] w=δ(f4f3((f1(F)+f2(F)+f1(F))) (4)
[0065] X fusion =Concat[w×X h ,(1-w)×X l ] (5)
[0066] Among them, X fusion represents the information after the fusion of high-resolution and Lidar information, X h represents high-scoring features, X l represents Lidar features, δ(·) represents the Sigmoid activation function; f1, f2, f3, and f4 represent 1×1, 7×7, 5×5, and 3×3 convolution operations; w represents the obtained weight matrix.
[0067] The depthwise separable convolution module consists of a BN layer, a ReLU layer, a depthwise separable convolution, and a residual block.
[0068] like Figure 4 As shown, in step S4, the channel dimension is compressed and restored through the fully connected layer to learn the mutual correlation between different channels; the improved SE fusion module is mainly used to squeeze and excite the inter-channel information of the different stages of the fusion features obtained above to perform inter-channel fusion to learn the complementary information between channels. Its input is two information, one of which comes from the last stage X5 of feature extraction, and the other comes from the first four stages of feature extraction. i , these two pieces of information are sent to the FC fully connected layer for channel squeezing, the ReLU activation layer, and another FC fully connected layer for channel restoration. After restoring its features, the residual connection is performed with the original unsqueezed features. The process can be expressed as:
[0069] X e =Concat[X i ,X5] i=1,......,4 (6)
[0070] X t =FC(RELU(FC(X e )))+X e (7)
[0071] Among them, Concat[·] represents the concatenation operation, X e Represents the cascade feature output of the two stage features, X t Represents the output features of the improved SE module.
[0072] like Figure 5 As shown, in step S5, the discrete Fourier transform layer is used to convert the image to the frequency domain, and the frequency domain information is used to improve the performance of CNN in certain tasks. The context information of the adjacent stage is sent to the 3×3 convolution for shallow feature extraction, and then sent to the 2D FFT, Hadamard product and 2D inverse FFT to better combine the spatial information and frequency domain information, extract the context information, and send the result to the 3×3 convolution for further information extraction, and use 1×1 convolution to adjust the dimension of the channel.
[0073] The present invention also provides a multi-source remote sensing image processing system, comprising:
[0074] The data acquisition module is used to acquire multiple data enhancement sample matrices of the high-resolution image data and Lidar image data to be fused; extract the spatial structure of the Lidar data based on the Resnet-34 network according to the acquired multiple data enhancement sample matrices; extract the spatial information and elevation features of the Lidar image data using the Lidar data information, and extract the spatial features of the high-resolution image data using the high-resolution image data information;
[0075] The multimodal residual attention fusion block is used to input the spatial features and elevation features into the feature fusion module, and use the fusion module to add and fuse the channel weights of the obtained elevation features and spatial features, and input them into the deep separable residual module for the next step of feature extraction to obtain multimodal feature extraction data;
[0076] The improved SE module fusion module includes two fully connected layers and one nonlinear activation layer. The multi-modal feature extraction data obtained is layered and fused through the fully connected layer. The extracted features of the last layer (the fifth layer) are fused with the features of the bottom layer (the first to fourth layers). While retaining the spatial information, the channel dimension is compressed and restored through the fully connected layer to learn the correlation between different channels. The feature information between channels of the multi-modal feature extraction data at different stages of the fusion network is better fused to achieve the reuse of the fifth layer's features and obtain the output data of cross-layer fusion.
[0077] The multi-layer feature context information learning module FFD includes a common convolution layer and a Fourier frequency domain transform layer, which is used to add and fuse the obtained cross-layer fusion output data with the sampling data of the adjacent previous stage, input it into the common convolution layer for spatial feature extraction, and then use the Fourier frequency domain transform layer to convert the image into the frequency domain to obtain the global context information in the frequency domain;
[0078] The classification module sends the acquired global context information in the frequency domain into the classifier to obtain the classification result.
[0079] Specifically, this embodiment uses the internationally recognized standard public data set Vaihingen, which contains 33 remote sensing images of different sizes, each of which is extracted from a larger top-level orthophoto image, and the spatial resolution of the top-level image and DSM is 9 cm. The remote sensing image format is an 8-bit TIFF file, consisting of three bands: near infrared, red, and green. DSM is a single-band TIFF file, and the grayscale level (corresponding to the DSM height) is a 32-bit floating point value encoding with a total of 6 classes (impervious surfaces, buildings, low vegetation, trees, cars, and clutter). Another data set uses the internationally recognized standard public data set Postdam, which contains 38 images of the same size, and the spatial resolution of the top-level image and DSM is 5 cm. Similar to the Vaihingen area, this data set is also composed of three-band remote sensing TIFF files and a single-band DSM. The area coverage size of each remote sensing image is the same. Its classification is consistent with Vaihingen.
[0080] In order to highlight the advantages of the present invention, the present invention uses three methods for comparative experiments, namely Fusenet, HAfusenet and CMGFNet.
[0081] As shown in Tables 1 and 2, the average values of the three evaluation indicators of the multi-source data fusion method are listed respectively. As can be seen from Tables 1 and 2, the Vaihingen dataset and the Postdam dataset achieve OA of 94.7 and 97.3 respectively. For the F1 and FIOU indicators, the method of the present invention is higher than other comparison algorithms.
[0082] For the Vaihingen dataset, the OA of the proposed method is improved by 4.4%, 2.3% and 1.7% compared with Fusenet, HAfusenet and CMGFNet, respectively. For the Postdam dataset, the OA of the proposed method is improved by 2.8%, 1.5% and 1.2% compared with Fusenet, HAfusenet and CMGFNet, respectively. In particular, in the comparative experiment, Fusenet and HAfusenet are official open source codes, while CMGFNet is not an official code, but an existing code adapted from the structure on github.
[0083] Table 1 Comparison of Vaihingen data classification accuracy (%)
[0084] Fusenet HAfusenet CMGFNet our OA=90.3 OA=92.4 OA=93.0 OA=94.7 F1=94.2 F1=95.1 F1=95.7 F1=96.2 FIOU = 89.2 FIOU=86.4 FIOU = 87.6 FIOU = 90.1 Number of parameters: 207793164 Number of parameters: 88977554 Number of parameters: 43684274 Number of parameters: 120407229 Running time: 3603 Running time: 6577 Running time: 4078 Running time: 4790
[0085] Table 2 Comparison of Postdam data classification accuracy (%)
[0086]
[0087]
[0088] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
[0089] In addition, unless otherwise specified, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which the invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.
Claims
1. A multi-source remote sensing image processing method, characterized in that: The steps include: Obtain multiple data enhancement sample matrices of high-resolution image data and Lidar image data to be fused; According to the acquired multiple data enhancement sample matrices, the spatial structure of Lidar and high-resolution image data is extracted based on the Resnet-34 network; Use Lidar data information to extract the spatial information and elevation features of Lidar image data, and use high-resolution image data information to extract the spatial features of high-resolution image data; Constructing a multimodal residual attention fusion block, the multimodal residual attention fusion block includes a fusion module and a deep separable residual module; inputting spatial features and elevation features into the fusion module, using the fusion module to add and fuse the channel weights of the elevation features and spatial features obtained, and inputting them into the deep separable residual module for the next step of feature extraction to obtain multimodal feature extraction data; An improved SE fusion module is proposed. The improved SE fusion module includes two fully connected layers. The fully connected layers are used to compress and restore the channel dimensions of different stages of the acquired multimodal feature extraction data, learn the mutual correlation between different channels, and achieve feature reuse at different stages by fusing the feature information between channels of multimodal feature extraction data at different stages of the network, so as to obtain cross-layer fused output data. Constructing a multi-layer feature global context information learning module FFD, wherein the multi-layer feature global context information learning module FFD includes a common convolution layer and a Fourier frequency domain transform layer; adding and fusing the obtained cross-layer fusion output data with the sampling data of the adjacent previous stage, inputting the data into the common convolution layer for spatial feature extraction, and then using the Fourier frequency domain transform layer to convert the image into the frequency domain to obtain the global context information in the frequency domain; The acquired global context information in the frequency domain is sent to the classifier for classification to obtain the classification result.
2. A multi-source remote sensing image processing method according to claim 1, characterized in that: The step of obtaining a plurality of data enhancement sample matrices of the high-resolution image data and the Lidar image data to be fused comprises the following steps: Obtain the original image matrix of high-resolution image data and Lidar image data; According to the original image data of the original image matrix, data enhancement samples are obtained by randomly flipping in the vertical direction and the horizontal direction, respectively, wherein the flip angles in the horizontal direction and the vertical direction are both in the range of [0°, 270°].
3. A multi-source remote sensing image processing method according to claim 1, characterized in that: The multimodal residual attention fusion block, i.e., the RAF module, is constructed to fuse the lateral features of the RGB encoder to the corresponding lateral features of the DSM encoder, and to perform cross-modal fusion of the RGB and DSM encoder features. Let D m represents the output of cross-modal fusion, then the output feature calculation formula of cross-modal fusion is: in, represents the side output encoder features of the m-th encoder of the RGB branch, represents the side output encoder feature of the m-th encoder of the DSM branch, m represents the m-th feature encoder, and RAF represents the multimodal residual attention fusion block.
4. A multi-source remote sensing image processing method according to claim 1, characterized in that: The learning of the correlation between different channels includes learning the correlation between the fifth-stage multimodal feature extraction data and the previous four-stage multimodal feature extraction data channels, fusing the feature information between the multimodal feature data channels at different stages of the network, generating high-resolution semantic instructions by fusing the high-level features of the fifth layer and the cross-modal fusion features of the first to fourth layers obtained from the multimodal residual attention fusion block, and sequentially fusing the cross-modal fusion features of the fifth layer into the cross-modal fusion features of the first to fourth layers to achieve the reuse of the fifth layer features. The formula for SE fusing the multimodal feature extraction data at different stages is: F=SE(Concat[D m ,D 5 ]),m=1,......,4 Among them, F represents the output of cross-layer fusion, Concat[·] represents the concatenation operation, and D m represents the output of cross-modal fusion, D 5 Represents the output of the fifth cross-modal fusion.
5. The multi-source remote sensing image processing method according to claim 1, characterized in that: The construction of the multi-layer feature global context information learning module FFD includes the following steps: The context information of the two adjacent stages, the output data of the cross-layer fusion and the upsampled data of the adjacent previous stage, is sent to the 3×3 convolution for shallow feature extraction, and then sent to the 2D FFT, Hadamard product and 2D inverse FFT to combine the spatial information and the frequency domain information to extract the context information of the adjacent stages; The context information of the extracted adjacent stages is sent to the 3×3 convolution for secondary information extraction to obtain the global context information in the frequency domain; A 1×1 convolution is used to adjust the dimension of the channel.
6. A multi-source remote sensing image processing system, characterized in that: include: A data acquisition module, used to acquire multiple data enhancement sample matrices of high-resolution image data and Lidar image data to be fused; According to the acquired multiple data enhancement sample matrices, the spatial structure of Lidar data is extracted based on the Resnet-34 network; Use Lidar data information to extract the spatial information and elevation features of Lidar image data, and use high-resolution image data information to extract the spatial features of high-resolution image data; The multimodal residual attention fusion block is used to input the spatial features and elevation features into the feature fusion module, and use the fusion module to add and fuse the channel weights of the obtained elevation features and spatial features, and input them into the deep separable residual module for the next step of feature extraction to obtain multimodal feature extraction data; The improved SE fusion module includes two fully connected layers, which compress and restore the channel dimensions of different stages of the acquired multimodal feature extraction data through the fully connected layers, learn the correlation between different channels, and achieve feature reuse at different stages by fusing the feature information between channels of multimodal feature extraction data at different stages of the network to obtain cross-layer fused output data; The multi-layer feature context information learning module FFD includes a common convolution layer and a Fourier frequency domain transform layer, which is used to add and fuse the obtained cross-layer fusion output data with the sampling data of the adjacent previous stage, input it into the common convolution layer for spatial feature extraction, and then use the Fourier frequency domain transform layer to convert the image into the frequency domain to obtain the global context information in the frequency domain; The classification module sends the acquired global context information in the frequency domain into the classifier to obtain the classification result.
Citation Information
Patent Citations
Remote sensing scene classification method based on Gabor filter and covariance pooling
CN114005046A
Ground feature classification method based on attention twin network and multi-modal fusion features
CN115512162A