Coastal zone feature classification method and system based on spatio-temporal fusion of sequence remote sensing images
By constructing a spatiotemporal fusion coastal land feature classification model and integrating high- and low-resolution image features, the constraint problem between the temporal resolution and spatial resolution of satellite remote sensing images is solved, and high-precision long-term time series monitoring of coastal land features is achieved.
Patent Information
- Application Number
- CN202510006544.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In existing technologies, the temporal resolution and spatial resolution of satellite remote sensing images constrain each other, resulting in low spatial resolution of long-term remote sensing images and temporal discontinuity of high spatial resolution images, making it difficult to achieve high-precision long-term monitoring of coastal zones.
By constructing a spatiotemporal fusion coastal land feature classification model, using a high-frequency-low-frequency feature encoder, a multi-resolution decomposition module, a dense convolutional Transformer module and a convolutional Transformer decoder, we fuse high- and low-resolution image features, reconstruct long-term high-spatial resolution reconstructed images, and obtain a coastal land feature classification map.
It improves the classification accuracy of coastal landforms, realizes long-term and high-spatial-resolution monitoring of the coastal zone, reduces the deviation caused by the dual-source satellite remote sensing imaging mechanism, and improves the accuracy of monitoring data.
Smart Images

Figure CN119919729B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a sequence remote sensing image coastal zone ground object classification method and system based on space-time fusion. BACKGROUND
[0002] The coastal zone is an important transition area between land and sea, and the coastal zone ecosystem has a very important value to China's economy and social life. However, due to human activities and natural factors, the coastal zone ecosystem is facing severe challenges. How to dynamically and accurately monitor the changes of the coastal zone ecosystem has a very important strategic significance for the sustainable development of the ocean. Satellite remote sensing technology is an important means to obtain large-scale and long-time coastal zone data. Due to budget and technical limitations, the time resolution and spatial resolution of remote sensing images obtained by a single sensor are mutually restricted. The spatial resolution of long-time remote sensing images is low, and the time interval of high spatial resolution images is long and discontinuous. However, the monitoring of the coastal zone requires long-time and high spatial resolution remote sensing images. SUMMARY
[0003] The present application aims to provide a sequence remote sensing image coastal zone ground object classification method and system based on space-time fusion to obtain long-time and high spatial resolution monitoring data of the coastal zone and improve the classification accuracy of coastal zone ground objects.
[0004] In order to achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0005] In a first aspect, the present application provides a sequence remote sensing image coastal zone ground object classification method based on space-time fusion, which comprises the following steps:
[0006] Obtaining sample images, the sample images including high spatial resolution images at prior time and low spatial resolution images at predicted time;
[0007] Constructing a space-time fusion coastal zone ground object classification model, training the coastal zone ground object classification model through the sample images until the composite loss of the coastal zone ground object classification model is lower than the set loss threshold, and obtaining a trained coastal zone ground object classification model;
[0008] Obtaining remote sensing images to be classified, and classifying the coastal zone ground objects in the low spatial resolution images at the predicted time in the remote sensing images to be classified through the trained coastal zone ground object classification model; wherein the remote sensing images to be classified include high spatial resolution images at prior time and low spatial resolution images at predicted time.
[0009] Preferably, the constructing a space-time fusion coastal zone ground object classification model, training the coastal zone ground object classification model through the sample images, comprises:
[0010] constructing a coastal zone feature classification model of spatio-temporal fusion; wherein the coastal zone feature classification model comprises a high-low frequency feature encoder, a high-low frequency fusion module, an inverse transform, a convolutional Transformer decoder, and a classifier; the high-low frequency feature encoder comprises a composite low-resolution image generation module, a multi-resolution decomposition module, and a dense convolutional Transformer module; the high-low frequency fusion module comprises a high-frequency fusion module and a low-frequency fusion module;
[0011] obtaining a high spatial resolution image and a low spatial resolution image in a sample image, performing down-resolution processing on the high spatial resolution image through the composite low-resolution image generation module to obtain a degraded image of the same scale as the low spatial resolution image, and fusing the low spatial resolution image and the degraded image to obtain a composite low-resolution image;
[0012] decomposing the high spatial resolution image and the composite low-resolution image through the multi-resolution decomposition module, respectively, extracting a high-resolution high-frequency image and a high-resolution low-frequency image from the high spatial resolution image, and extracting a low-resolution high-frequency image and a low-resolution low-frequency image from the composite low-resolution image;
[0013] extracting hierarchical features of the high-resolution high-frequency image, the high-resolution low-frequency image, the low-resolution high-frequency image, and the low-resolution low-frequency image through the dense convolutional Transformer module, respectively, and obtaining frequency domain local-global features through pixel convolution fusion; wherein the frequency domain local-global features comprise high-resolution high-frequency features, high-resolution low-frequency features, low-resolution high-frequency features, and low-resolution low-frequency features;
[0014] fusing the extracted frequency domain local-global features through the high-frequency fusion module and the low-frequency fusion module to obtain high-frequency fusion features and low-frequency fusion features; and obtaining spatio-temporal fusion features through inverse transform of the low-frequency fusion features and the high-frequency fusion features;
[0015] reconstructing the spatio-temporal fusion features through the convolutional Transformer decoder to obtain a long-time-series high spatial resolution reconstructed image, and transmitting the reconstructed image to the classifier to obtain a long-time-series coastal zone feature classification map.
[0016] Preferably, the high-resolution high-frequency image and the high-resolution low-frequency image are extracted from the high spatial resolution image, and the low-resolution high-frequency image and the low-resolution low-frequency image are extracted from the composite low-resolution image, comprising:
[0017] decomposing the high spatial resolution image through a first NSPF to obtain a first high-resolution low-frequency image and a first high-resolution high-frequency image;
[0018] The first high-score low-frequency image is further decomposed by a second NSPF to obtain a second high-score low-frequency image and a second high-score high-frequency image, and the second high-score low-frequency image is decomposed by a third NSPF to obtain a third high-score low-frequency image and a third high-score high-frequency image; the third high-score low-frequency image is the high-score low-frequency image;
[0019] The first high-resolution high-frequency map, the second high-resolution high-frequency map and the third high-resolution high-frequency map are respectively filtered by the first NSDF, the second NSD F and the third NSDF to obtain the first high-resolution high-frequency sub-band map, the second high-resolution high-frequency sub-band map and the third high-resolution high-frequency sub-band map, and the first high-resolution high-frequency sub-band map, the second high-resolution high-frequency sub-band map and the third high-resolution high-frequency sub-band map constitute the high-resolution high-frequency map.
[0020] Preferably, extracting the low-resolution high-frequency image and the low-resolution low-frequency image from the composite low-resolution image comprises:
[0021] Decompose the composite low-resolution image by the first NSPF to obtain the first low-resolution low-frequency image and the first low-resolution high-frequency image;
[0022] The first low-score low-frequency image is further decomposed by a second NSPF to obtain a second low-score low-frequency image and a second low-score high-frequency image, and the second low-score low-frequency image is decomposed by a third NSPF to obtain a third low-score low-frequency image and a third low-score high-frequency image; the third low-score low-frequency image is the low-score low-frequency image;
[0023] The first low-resolution high-frequency map, the second low-resolution high-frequency map and the third low-resolution high-frequency map are filtered by the first NSDF, the second NSD F and the third NSDF respectively to obtain the first low-resolution high-frequency sub-band map, the second low-resolution high-frequency sub-band map and the third low-resolution high-frequency sub-band map. The first low-resolution high-frequency sub-band map, the second low-resolution high-frequency sub-band map and the third low-resolution high-frequency sub-band map constitute the low-resolution high-frequency map.
[0024] Preferably, the NSPF includes a low-pass filter and a high-pass filter, the low-pass filter includes a low-pass decomposition filter and a low-pass reconstruction filter, and the high-pass filter includes a high-pass decomposition filter and a high-pass reconstruction filter; the NSPF satisfies the identity L f (I)L c (I)+G f (I)G c (I) = 1;
[0025] The NSDF includes a sector filter and a checkerboard filter, the sector filter includes a sector decomposition filter and a sector reconstruction filter, the checkerboard filter includes a checkerboard decomposition filter and a checkerboard reconstruction filter; the NSDF satisfies the identity S f (I)S c (I)+C f (I)C c (I) = 1;
[0026] where I is an input image, L f (·) is a low-pass decomposition filter function, L c (·) is a low-pass reconstruction filter function, G f (·) is a high-pass decomposition filter function, G c (·) is a high-pass reconstruction filter function; S f (I) is a sector decomposition filter function, S c (I) is a sector reconstruction filter function, C f (I) is a chessboard decomposition filter function, C c (I) is a chessboard reconstruction filter function.
[0027] Preferably, the hierarchical features and the frequency domain local-global features of the high-resolution high-frequency image, the high-resolution low-frequency image, the low-resolution high-frequency image and the low-resolution low-frequency image are extracted by the dense convolution Transformer module respectively, and are calculated by the following formula:
[0028] TF1 = f CT (TFI)
[0029] TF2 = f CT (f concat (TFI, TF1)
[0030] TF3 = f CT (f concat (TFI, TF1, TF2)
[0031] TF4 = f CT (f concat (TFI, TF1, TF2, TF3)
[0032] TFO = W 1 (f concat (TFI, TF1, TF2, TF3, TF4)
[0033] where TF1 and TFO are the input image and the output frequency domain local-global features of the dense convolution Transformer module respectively, the input image includes the high-resolution high-frequency image, the high-resolution low-frequency image, the low-resolution high-frequency image and the low-resolution low-frequency image, and the frequency domain local-global features include the high-resolution high-frequency feature, the high-resolution low-frequency feature, the low-resolution high-frequency feature and the low-resolution low-frequency feature; TF i is the hierarchical feature output by the i-th convolution Transformer module, i = 1, 2, 3, 4, f CT (·) is a function of the convolution Transformer module, f concat (·) is a function of the splicing operation, W 1is a 1x1 pixel convolutional fusion dense feature and adjusts the data dimension;
[0034] The convolutional Transformer module is composed of layer normalization, multi-convolution head transposed attention, local enhancement feedforward network, and residual connection, and the mathematical expression is:
[0035] T=f CT (x);
[0036]
[0037] Wherein, x represents the input of the convolutional Transformer module, X represents the normalized image of x, W q,1 (·), W k ,1 (·) and W v,1 (·) represent 1x1 pixel convolutional projection to generate Q, K and V respectively, and represent 3x3 deep convolutional projection to generate Q, K and V respectively, Q, K and V represent the query, key and value of the multi-convolution head transposed attention respectively, R(·) is a function of deformation operation, β is a learnable scaling parameter, A is the generated attention map, F1 is the enhanced feature after self-attention weight, F is the corrected feature after multi-convolution head transposed attention, T1 is the normalized image of F, T2 is the feature after local enhancement of T1, W 1 (·) is a 1x1 pixel convolution, GELU(·) is a Gaussian error linear unit function, is a 3x3 deep convolution, and T represents the output of the convolutional Transformer module.
[0038] Preferably, the high-frequency fusion module and the low-frequency fusion module respectively fuse the extracted frequency domain local-global features to obtain high-frequency fusion features and low-frequency fusion features, each of which includes three cross-convolution attention modules, the cross-convolution attention module is composed of layer normalization, multi-cross-convolution head transposed attention, local enhancement feedforward network, and residual connection, and is realized by the following formula:
[0039]
[0040]
[0041] Wherein, x f1 represents high-frequency features or low-frequency features, x f2 represents low-frequency features or high-frequency features, X f1 and X f2 represent the normalized features of x f1 and x f2 , Wqf,1 (·), W kf,1 (·) and W vf,1 (·) respectively represent the projection generation Q f , K f and V f 1x1 pixel convolution, and respectively represent the projection generation Q f , K f and V f 3x3 deep convolution, Q f , K f and V f represent the query, key and value of the multi-cross convolution head transpose attention, CA is the generated cross attention map, CF1 is the cross enhanced feature after the cross attention weight, CF is the cross fusion feature after the multi-cross convolution head transpose attention, FT1 is the image after the normalization of CF, FT2 is the feature after the enhanced local of FT1, FT represents the high-frequency fusion feature or low-frequency fusion feature output by the fusion module.
[0042] Preferably, the reconstruction image of long time sequence and high spatial resolution is obtained by reconstructing the spatio-temporal fusion feature through the convolutional Transformer decoder, comprising:
[0043] The decoding and restoration of the spatio-temporal fusion feature is performed by 4 convolutional Transformer modules in the decoder, to recover the image change spectral information and detail information;
[0044] The high spatial resolution image is reconstructed by 2 3x3 convolutions and ReLU in the decoder and the last 3x3 convolution;
[0045] The reconstruction image at the prediction moment is generated by the following formula:
[0046]
[0047] Wherein, G p represents the reconstruction image at the prediction moment p, represents 3x3 convolution followed by ReLU activation function, and STFF is the spatio-temporal fusion feature.
[0048] Preferably, the composite loss of the coastal ground object classification model is calculated by the following method:
[0049] The reconstruction image and the corresponding reference image are obtained;
[0050] determine a composite loss of the coastal zone feature classification model based on the reconstructed image and the corresponding reference image; wherein the composite loss comprises a pixel-level loss and a feature-level loss, the pixel-level loss comprises a spectral consistency loss, a structure consistency loss and an image content loss;
[0051] The expression of the feature loss L F is:
[0052]
[0053] wherein N is a batch size, and respectively represent the mean value of the reconstructed image G p and the reference image T p in the feature map of the VGG-19 pre-training model;
[0054] The expression of the spectral consistency loss L pe is:
[0055]
[0056] The expression of the structure consistency loss L MS_SSIM is:
[0057]
[0058] wherein MS_SSIM(G p ,T p ) is a multi-scale structure similarity index; l H , c h and s h are the H and h scale representations of luminance l, contrast c and structure s, and respectively represent the mean value of the reconstructed image G p and the reference image T p , represents the covariance between the reconstructed image G p and the reference image T p , and respectively represent the variance of the reconstructed image G p and the reference image T p , b1 and b2 are constants to avoid the denominator being 0; SSIM(G p ,T p ) is a structure similarity index;
[0059] The expression of the image content loss is:
[0060]
[0061] The composite loss function L Z The expression is:
[0062]
[0063] Wherein, gamma 1, lambda 2, xi 3 and Is a weight coefficient.
[0064] In a second aspect, the embodiments of the present application provide a sequence remote sensing image coastal zone ground object classification system based on space-time fusion, the system comprises:
[0065] At least one processor;
[0066] At least one memory for storing at least one program;
[0067] When the at least one program is executed by the at least one processor, the at least one processor implements the sequence remote sensing image coastal zone ground object classification method based on space-time fusion as described in any one of the above.
[0068] The beneficial effects of the present application are: by reducing the resolution of high spatial resolution images, obtaining degraded images of the same scale as low spatial resolution images, and generating composite low-resolution images with high-resolution image physical properties by fusing with low spatial resolution images, reducing the deviation caused by the imaging mechanism of dual-source satellite remote sensing images; using a multi-resolution decomposition module to decompose the high spatial resolution image and the composite low-resolution image to obtain high-resolution high-frequency images and high-resolution low-frequency images, low-resolution high-frequency images and low-resolution low-frequency images; extracting the frequency domain local-global features of the above high-resolution high-frequency images, high-resolution low-frequency images, low-resolution high-frequency images and low-resolution low-frequency images through the dense convolution Transformer module; and fusing the extracted high-resolution high-frequency features and low-resolution high-frequency features to obtain high-frequency fusion features, and fusing the extracted high-resolution low-frequency features and low-resolution low-frequency features to obtain low-frequency fusion features; then obtaining space-time fusion features by inverse transforming the low-frequency fusion features and the high-frequency fusion features; reconstructing the space-time fusion features through the decoder composed of the convolution Transformer module to obtain long-time sequence high spatial resolution reconstruction images; passing it to the classifier to obtain long-time sequence coastal zone ground object classification map. Through the present application, long-time sequence monitoring data of the coastal zone can be obtained, and the classification accuracy of the coastal zone ground object can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0070] Figure 1 1 is a flow chart of a method for classifying coastal features using sequence remote sensing images based on spatiotemporal fusion in an embodiment of the present invention;
[0071] Figure 2 1 is an overall structural diagram of the spatiotemporal fusion coastal zone feature classification model according to an embodiment of the present invention;
[0072] Figure 3 is a non-subsampled contourlet transform decomposition diagram of a high spatial resolution image in an embodiment of the present invention;
[0073] Figure 4 is a non-subsampled contourlet transform decomposition diagram of a composite low-resolution image in an embodiment of the present invention;
[0074] Figure 5 yes Figure 2 The structure diagram of the dense convolution Transformer module;
[0075] Figure 6 yes Figure 5 The structure diagram of the convolutional Transformer module;
[0076] Figure 7 yes Figure 6 The structure diagram of the multi-convolutional head transposed attention module;
[0077] Figure 8 yes Figure 6 Structural diagram of the local enhanced feedforward network;
[0078] Figure 9 yes Figure 2 Structural diagram of the mid-high frequency fusion module / low frequency fusion module;
[0079] Figure 10 yes Figure 2 The structure diagram of the convolutional Transformer decoder;
[0080] Figure 11 It is a structural diagram of a coastal land feature classification system based on spatiotemporal fusion of sequence remote sensing images in an embodiment of the present invention. DETAILED DESCRIPTION
[0081] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects of the present invention so as to fully understand the purpose, scheme and effect of the present invention. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other unless there is any conflict.
[0082] The time resolution and the spatial resolution of remote sensing images obtained by a single sensor are mutually restricted, the spatial resolution of long-time sequence remote sensing images is low, and the time of high spatial resolution images is discontinuous, and the monitoring of the coastal zone requires long-time sequence high spatial resolution remote sensing images, therefore, a sequence remote sensing image coastal zone feature classification method and system based on spatio-temporal fusion are provided to obtain long-time sequence monitoring data of the coastal zone and improve the classification accuracy of the features of the coastal zone. In the embodiment, first, in order to reduce the deviation caused by the imaging mechanism of the dual-source satellite remote sensing image, the high spatial resolution image is processed to reduce the resolution to obtain a degraded image of the same size as the low spatial resolution image, and the low spatial resolution image and the degraded image are fused to generate a composite low-resolution image with high-resolution image physical characteristics. Second, the high spatial resolution image and the composite low-resolution image are decomposed by a multi-resolution decomposition module to obtain high-resolution high-frequency images HHI and high-resolution low-frequency images HLI, low-resolution high-frequency images LHI and low-resolution low-frequency images LLI. Third, a dense convolutional Transformer module is designed to extract the frequency domain local-global features of the high-resolution high-frequency images HHI, the high-resolution low-frequency images HLI, the low-resolution high-frequency images LHI and the low-resolution low-frequency images LLI; and the extracted high-resolution high-frequency features HHI and the low-resolution high-frequency features LHF are fused to obtain high-frequency fusion features FHF, and the extracted high-resolution low-frequency features HLF and the low-resolution low-frequency features LLI are fused to obtain low-frequency fusion features FLF; and the low-frequency fusion features FLF and the high-frequency fusion features FHF are inversely transformed to obtain spatio-temporal fusion features STFF. Fourth, a decoder composed of a convolutional Transformer module is designed to reconstruct the spatio-temporal fusion features to obtain long-time sequence high spatial resolution remote sensing images; and the long-time sequence high spatial resolution remote sensing images are transmitted to a classifier to obtain long-time sequence coastal zone feature classification images. Fifth, a composite loss function composed of a content loss, a structure loss, a spectral difference loss and a feature loss is constructed to optimize the model to obtain the optimal result. The spatial resolution and the time resolution of the remote sensing image are improved by spatio-temporal fusion to improve the classification accuracy and realize long-time sequence accurate monitoring of the features of the coastal zone. The experimental results on the public CIA and LGC data sets show that the fusion performance and the classification ability of the application are better in subjective evaluation and objective evaluation.
[0083] Referring to Figure 1 The application provides a sequence remote sensing image coastal zone feature classification method based on spatio-temporal fusion, which comprises the following steps:
[0084] S100, acquiring sample images, the sample images comprising high spatial resolution images at a prior time and low spatial resolution images at a predicted time;
[0085] S200, construct a coastal zone ground object classification model based on spatio-temporal fusion, train the coastal zone ground object classification model through the sample image until the compound loss of the coastal zone ground object classification model is lower than a set loss threshold, and obtain a trained coastal zone ground object classification model;
[0086] Specifically, in the process of training the coastal zone ground object classification model, the compound loss of the coastal zone ground object classification model is determined, and when the compound loss of the coastal zone ground object classification model is lower than a set loss threshold, the training of the coastal zone ground object classification model is stopped, and a trained coastal zone ground object classification model is obtained.
[0087] S300, obtain a remote sensing image to be classified, and classify coastal zone ground objects in a low spatial resolution image at a predicted time in the remote sensing image to be classified through the trained coastal zone ground object classification model; wherein the remote sensing image to be classified comprises a high spatial resolution image at a prior time and a low spatial resolution image at a predicted time.
[0088] As an improvement of the above embodiment, in S200, the construction of the coastal zone ground object classification model based on spatio-temporal fusion, the training of the coastal zone ground object classification model through the sample image comprises:
[0089] S210, construct a coastal zone ground object classification model based on spatio-temporal fusion; wherein the coastal zone ground object classification model comprises a high-frequency-low-frequency feature encoder, a high-frequency-low-frequency fusion module, an inverse transform, a convolutional Transformer decoder, and a classifier; the high-frequency-low-frequency feature encoder comprises a composite low-resolution image generation module, a multi-resolution decomposition module, and a dense convolutional Transformer module; and the high-frequency-low-frequency fusion module comprises a high-frequency fusion module and a low-frequency fusion module.
[0090] It should be noted that long-time sequence remote sensing images have low spatial resolution, and have the characteristics of foreign objects with the same spectrum and the same objects with different spectra, resulting in low ground object classification accuracy; and high spatial resolution images have a large time interval, which is not conducive to long-time sequence monitoring of ground objects; in view of the above problems, the present application proposes a sequence remote sensing image coastal zone ground object classification method based on spatio-temporal fusion to obtain coastal zone long-time sequence monitoring data and improve the classification accuracy of coastal zone ground objects, and the overall structure is as shown in Figure 2 . Figure 2 In the formula, G t represents a high spatial resolution image at any time t, DG t is a degraded image of the high spatial resolution image G t , L p is a low spatial resolution remote sensing image at a predicted time p, G p is a high spatial resolution reconstructed image at a predicted time p. The degraded image DG tTo first compare the high spatial resolution image G according to the resolution comparison between the high spatial resolution and low spatial resolution images t Perform a resolution reduction operation and then perform an upsampling operation to the same resolution as the high spatial resolution image G t Images of the same size are obtained.
[0091] S220, obtaining a high spatial resolution image and a low spatial resolution image from the sample image, performing resolution reduction processing on the high spatial resolution image using the composite low-resolution image generation module to obtain a degraded image of the same scale as the low spatial resolution image, and fusing the low spatial resolution image with the degraded image to obtain a composite low-resolution image; wherein the resolution of the high spatial resolution image is greater than that of the low spatial resolution image;
[0092] Specifically, the high spatial resolution image G t Can be decomposed into low spatial resolution images L p Low-frequency images with consistent spatial resolution and high-resolution images G t High-frequency detail images with consistent spatial resolution. t High-frequency detail image and low-frequency image L p The spectral information of the prediction time p high spatial resolution image G p Considering the difference between dual-source satellite images caused by the imaging mechanisms of different sensors, in order to reduce the physical deviation between low spatial resolution images and high spatial resolution images, the information of high spatial resolution images is introduced into low spatial resolution images. t Perform the resolution reduction operation and then obtain the image G with high spatial resolution. t Same size, get downgraded image DG t , which has the physical properties of a high spatial resolution image. Then, the degraded image DG t With low spatial resolution image L p A composite low-resolution image is obtained after fusion consisting of splicing, 1×1 convolution, and ReLU function.
[0093] S230, decomposing the high spatial resolution image and the composite low-resolution image by the multi-resolution decomposition module, extracting a high-resolution high-frequency map and a high-resolution low-frequency map from the high spatial resolution image, and extracting a low-resolution high-frequency map and a low-resolution low-frequency map from the composite low-resolution image;
[0094] Specifically, a high-resolution high-frequency image and a high-resolution low-frequency image are extracted from the high spatial resolution image, and a low-resolution high-frequency image and a low-resolution low-frequency image are extracted from the composite low-resolution image.
[0095] As an improvement of the above embodiment, in S230, the extracting the high-frequency image and the low-frequency image from the high spatial resolution image comprises:
[0096] S2311, the high spatial resolution image G t obtained by the first NSPF decomposition, and the first high-frequency image HFHI obtained by the first NSPF decomposition;
[0097] S2312, the first high-frequency image HFLI is further decomposed by the second NSPF to obtain the second high-frequency image HSLI and the second high-frequency image HSHI, and the second high-frequency image HSLI is further decomposed by the third NSPF to obtain the third high-frequency image HTLI and the third high-frequency image HTHI; the third high-frequency image HTLI is the high-frequency image HLI;
[0098] S2313, the first high-frequency image HFHI, the second high-frequency image HSHI and the third high-frequency image HTHI are filtered by the first NSDF, the second NSDF and the third NSDF respectively to obtain the first high-frequency sub-band image, the second high-frequency sub-band image and the third high-frequency sub-band image, and the first high-frequency sub-band image, the second high-frequency sub-band image and the third high-frequency sub-band image constitute the high-frequency image HHI.
[0099] Specifically, the high spatial resolution image G t is decomposed into a multi-scale and multi-directional high-frequency image and a low-frequency image, i.e. the high-frequency image HHI and the low-frequency image HLI, by a multi-resolution decomposition module, and the multi-resolution decomposition module adopts NSCT (non-subsampled contourlet transform). The NSCT includes NSPF (non-subsampled pyramid filter) and NSDF (non-subsampled directional filter), and in the present application, three decomposition operations are performed to form a three-level pyramid structure, as shown in Figure 3 the high spatial resolution image G tThe first high-frequency image HFLI is decomposed by a second NSPF to obtain a second high-frequency image HSLI and a second high-frequency image HSHI. The second high-frequency image HSLI is decomposed by a third NSPF to obtain a third high-frequency image HTLI (i.e. the high-frequency image HLI) and a third high-frequency image HTHI. The first high-frequency image HFHI, the second high-frequency image HSHI and the third high-frequency image HTHI are filtered by a first NSDF, a second NSDF and a third NSDF respectively to obtain a first high-frequency sub-band image, a second high-frequency sub-band image and a third high-frequency sub-band image, which constitute the high-frequency image HHI.
[0100] As an improvement of the above embodiment, in S230, the extracting of the low-frequency image and the high-frequency image from the composite low-resolution image comprises:
[0101] S2321, the composite low-resolution image is decomposed by a first NSPF to obtain a first low-frequency image LFLI and a first high-frequency image LFHI;
[0102] S2322, the first low-frequency image LFLI is decomposed by a second NSPF to obtain a second low-frequency image LSLI and a second high-frequency image LSHI. The second low-frequency image LSLI is decomposed by a third NSPF to obtain a third low-frequency image LTLI and a third high-frequency image LTHI. The third low-frequency image LTLI is the low-frequency image LLI;
[0103] S2323, the first high-frequency image LFHI, the second high-frequency image LSHI and the third high-frequency image LTHI are filtered by a first NSDF, a second NSDF and a third NSDF respectively to obtain a first low-frequency sub-band image, a second low-frequency sub-band image and a third low-frequency sub-band image, which constitute the low-frequency image LHI.
[0104] Specifically, the composite low-resolution image is decomposed by a non-subsampled contourlet transform into a multi-scale and multi-directional high-frequency image and a low-frequency image, i.e. the low-frequency image LLI and the low-frequency image LHI. As shown in FIG. 2, the low-frequency image LLI is decomposed by a first NSPF to obtain a first high-frequency image HFLI. The first high-frequency image HFLI is decomposed by a second NSPF to obtain a second high-frequency image HSLI and a second high-frequency image HSHI. The second high-frequency image HSLI is decomposed by a third NSPF to obtain a third high-frequency image HTLI (i.e. the high-frequency image HLI) and a third high-frequency image HTHI. The first high-frequency image HFHI, the second high-frequency image HSHI and the third high-frequency image HTHI are filtered by a first NSDF, a second NSDF and a third NSDF respectively to obtain a first high-frequency sub-band image, a second high-frequency sub-band image and a third high-frequency sub-band image, which constitute the high-frequency image HHI. Figure 4As shown, the composite low-resolution image is decomposed by the first NSPF to obtain a first low-resolution low-frequency image LFLI and a first low-resolution high-frequency image LFHI, the first low-resolution low-frequency image LFLI is further decomposed by the second NSPF to obtain a second low-resolution low-frequency image LSLI and a second low-resolution high-frequency image LSHI, then, the second low-resolution low-frequency image LSLI is decomposed by the third NSPF to obtain a third low-resolution low-frequency image LTLI (i.e. the above-mentioned low-resolution low-frequency image LLI) and a third low-resolution high-frequency image LTHI. The first low-resolution high-frequency image LFHI, the second low-resolution high-frequency image LSHI and the third low-resolution high-frequency image LTHI are filtered by the first NSDF, the second NSDF and the third NSDF respectively to obtain a first low-resolution high-frequency sub-band image, a second low-resolution high-frequency sub-band image and a third low-resolution high-frequency sub-band image, and the three low-resolution high-frequency sub-band images constitute the low-resolution high-frequency image LHI.
[0105] As an improvement of the above-mentioned embodiment, the NSPF comprises a low-pass filter and a high-pass filter, the low-pass filter comprises a low-pass decomposition filter and a low-pass reconstruction filter, and the expression is [L f (I), L c (I)], the high-pass filter comprises a high-pass decomposition filter and a high-pass reconstruction filter, and the expression is [G f (I), G c (I)], and the NSPF satisfies the identity (1).
[0106] The NSDF comprises a sector filter and a checkerboard filter, the sector filter comprises a sector decomposition filter and a sector reconstruction filter, and the expression is [S f (I), S c (I)], the checkerboard filter comprises a checkerboard decomposition filter and a checkerboard reconstruction filter, and the expression is [C f (I), C c (I)], and the NSDF satisfies the identity (2).
[0107] L f (I) L c (I) + G f (I) G c (I) = 1 (1)
[0108] S f (I) S c (I) + C f (I) C c (I) = 1 (2)
[0109] Wherein, I is an input image, L f (·) is a low-pass decomposition filter function, L c (·) is a low-pass reconstruction filter function, G f (·) is a high-pass decomposition filter function, G c(·) is a high-pass reconstruction filter function, S f (I) is a fan decomposition filter function, S c (I) is a fan reconstruction filter function, C f (I) is a chessboard decomposition filter function, C c (I) is a chessboard reconstruction filter function.
[0110] S240, respectively extracting the high-resolution high-frequency image, the high-resolution low-frequency image, the low-resolution high-frequency image and the low-resolution low-frequency image by the dense convolution Transformer module, and obtaining the frequency domain local-global feature through pixel convolution fusion; wherein the frequency domain local-global feature includes high-resolution high-frequency feature, high-resolution low-frequency feature, low-resolution high-frequency feature and low-resolution low-frequency feature;
[0111] Specifically, the obtained high-resolution high-frequency image HHI and high-resolution low-frequency image HLI, low-resolution high-frequency image LHI and low-resolution low-frequency image LLI are passed into the dense convolution Transformer module to extract multi-level dense local-global features, and high-resolution high-frequency feature HHI, high-resolution low-frequency feature HLF, low-resolution high-frequency feature LHF and low-resolution low-frequency feature LLF are obtained through convolution fusion.
[0112] It should be noted that the time and computational complexity of ordinary visual Transformer increases with the square of the input spatial resolution, which is not suitable for large spatial resolution images, and is not conducive to the extraction of local features, and the long-distance dependence ability obtained by convolution is weak, so the present application uses convolution Transformer to extract the features of large-scale remote sensing images. As shown in Figure 5 The dense convolution Transformer module includes 4 convolution Transformer modules and 1 convolution fusion layer, the convolution Transformer module combines the advantages of convolution in obtaining local features and Transformer in obtaining long-distance dependence, and can obtain local-global features. The dense convolution Transformer module uses dense connection to form the continuous memory of the hierarchical local-global features extracted by the convolution Transformer module, and the implementation process is as formula (3).
[0113]
[0114] Wherein, TFI and TFO are the input image and the output frequency domain local-global feature of the dense convolution Transformer module, the input image includes high-resolution high-frequency image, high-resolution low-frequency image, low-resolution high-frequency image and low-resolution low-frequency image, and the frequency domain local-global feature includes high-resolution high-frequency feature, high-resolution low-frequency feature, low-resolution high-frequency feature and low-resolution low-frequency feature; TF iis the hierarchical feature output by the i-th convolutional Transformer module, i = 1, 2, 3, 4, f CT (·) is a function of the convolutional Transformer module, f concat (·) is a function of the concatenation operation, W 1 (·) is a 1x1 pixel convolution that fuses dense features and adjusts the data dimension.
[0115] The structure of the convolutional Transformer module is shown in Figure 6 , Figure 7 and Figure 8 , which includes layer normalization, multi-convolution head transpose attention, local enhancement feedforward network, and residual connection, as shown in equation (4). First, the input is layer normalized to obtain layer normalized features, and then 3 1x1 pixel convolutions are used to aggregate cross-channel context information and 3x3 deep convolutions are used to encode channel spatial context information to generate Q, K, and V. Then, Q and K are deformed, and the deformed Q and K matrices are multiplied and Softmaxed to obtain attention A. A is multiplied with the deformed V, deformed, and 1x1 pixel convolved to add to the input to obtain corrected features. The theoretical expression of the above process is shown in equation (5). The corrected features are then layer normalized to generate normalized corrected features, which are projected to increase the feature dimension by 1x1 pixel convolution, and local information is obtained by 3x3 deep convolution. Then, the Gaussian error linear unit is activated to aggregate cross-channel local context information and adjust the number of channels to extract useful local information. The theoretical expression of the above process is shown in equation (6). In this technical field, Q, K, and V represent the query, key, and value of the multi-convolution head transpose attention, respectively.
[0116] T = f CT (x) (4)
[0117]
[0118]
[0119] where x represents the input image of the convolutional Transformer module, X represents the normalized image of x, W q,1 (·), W k,1 (·), and W v,1 (·) represent 1x1 pixel convolutions used to project to generate Q, K, and V, respectively. and respectively represent 3x3 depth convolution for projecting to generate Q, K and V, R(·) is a function of deformation operation, β is a learnable scaling parameter, A is the generated attention map, F1 is the enhanced feature after self-attention weight, F is the corrected feature after multi-convolution head transpose attention, T1 is the image after normalization of F, T2 is the feature after enhancement of local T1, W 1 (·) is 1x1 pixel convolution, GELU(·) is Gaussian error linear unit function, is 3x3 depth convolution, T represents the output of the convolutional Transformer module.
[0120] S250, the extracted frequency domain local-global features are fused by the high-frequency fusion module and the low-frequency fusion module respectively to obtain high-frequency fusion features and low-frequency fusion features; the low-frequency fusion features and the high-frequency fusion features are obtained by inverse transformation to obtain space-time fusion features;
[0121] Specifically, the generated high-frequency features HHF and low-frequency features LHF are transmitted to the high-frequency fusion module to realize high-frequency component feature fusion to obtain high-frequency fusion features FHF, the high-frequency features HLF and the low-frequency features LLF are transmitted to the low-frequency fusion module to obtain low-frequency fusion features FLF, and then the NSCT inverse transformation is performed to obtain space-time fusion features STFF. The high-frequency fusion module and the low-frequency fusion module have the same structure, as shown in Figure 9 , the high-frequency fusion module and the low-frequency fusion module are both composed of 3 cross convolution attention modules. The cross convolution attention module is composed of layer normalization, multi-cross convolution head transpose attention, local enhancement feedforward network and residual connection. The multi-cross convolution head transpose attention realizes the interactive fusion of information with different attributes, and aggregates the complementary information of the two. It is generated by different input projections Q f , K f and V f . For the high-frequency information branch, the high-frequency features HHF are projected by 1x1 pixel convolution and 3x3 depth convolution to generate K f , V f , and the low-frequency features LH F are projected by 1x1 pixel convolution and 3x3 depth convolution to generate Q f . For the low-frequency information branch, the low-frequency features LLF are projected by 1x1 pixel convolution and 3x3 depth convolution to generate K f , V f , and the low-frequency features HLF are projected by 1x1 pixel convolution and 3x3 depth convolution to generate Q f . Then Q f , K f , V f are deformed, and the deformed Q f , K fThe matrix is multiplied and a Softmax operation is performed to obtain a cross-attention graph CA, and CA is combined with the deformed V f After the multiplication operation, the deformation operation, and the 1x1 pixel convolution, the cross-fusion feature CF is obtained by adding the input. The theoretical expression of the above process is as shown in equation (7).
[0122] After the cross-fusion feature is normalized, the normalized cross-fusion feature is generated. Then, the feature dimension is increased by 1x1 pixel convolution projection, the local information is obtained by 3x3 deep convolution, and then the non-linear Gaussian error linear unit is activated. The cross-channel local context information is aggregated and the channel number is adjusted by 1x1 pixel convolution, so as to extract useful local information. The theoretical expression of the specific implementation of the above process is as shown in equation (8).
[0123]
[0124] wherein x f1 and x f2 represent the input of the high-frequency fusion module or the low-frequency fusion module. Specifically, x f1 represents the high-resolution high-frequency feature or the low-resolution low-frequency feature, x f2 represents the high-resolution low-frequency feature or the low-resolution high-frequency feature, X f1 and X f2 represent the normalized features of x f1 and x f2 , W qf,1 (·), W kf,1 (·) and W vf,1 (·) represent 1x1 pixel convolution for projecting to generate Q f , K f and V f , and represent 3x3 deep convolution for projecting to generate Q f , K f and V f , CA is the generated cross-attention graph, CF1 is the cross-enhanced feature after the cross-attention weight, CF is the cross-fusion feature after the multi-cross convolution head transpose attention, FT1 is the image after the normalization of CF, FT2 is the feature after the enhanced local of FT1, and FT represents the output of the fusion module. Specifically, FT represents the high-frequency fusion feature or the low-frequency fusion feature output by the fusion module.
[0125] After the high-frequency fusion module / low-frequency fusion module composed of three cross-convolution attention modules, the high-frequency fusion feature FH F and the low-frequency fusion feature FLF are obtained, and then the spatio-temporal fusion feature STFF is obtained after the NSCT inverse transform.
[0126] S260, reconstructing the spatio-temporal fusion features through the convolutional Transformer decoder to obtain a long-time high-spatial-resolution reconstructed image, and transmitting the reconstructed image to a classifier to obtain a long-time coastal zone feature classification map.
[0127] Specifically, the spatio-temporal fusion features STFF are transmitted to a decoder composed of a convolutional Transformer module to decode the fusion features and reconstruct a high-spatial-resolution image by convolution, and the structure of the decoder is as shown in Figure 10 The four convolutional Transformer modules in the decoder are used for decoding and restoring the spatio-temporal fusion features, to recover the spectral information and detail information of the image changes; the two 3x3 convolutions and ReLU in the decoder and the last 3x3 convolution are used for reconstructing the high-spatial-resolution image, and the implementation process is expressed as formula (9).
[0128]
[0129] wherein G p represents a predicted high-spatial-resolution reconstructed image at time p generated by the coastal zone feature classification model, represents a 3x3 convolution followed by a ReLU activation function.
[0130] The reconstructed image G p generated by the fusion is classified into different types of coverage features after the classifier, and support vector machine classifiers (SVM) and neural network classifiers (NN) can be used as the classifier to classify and process the features. The classification results are compared with the classification results of the fusion results of the high-spatial-resolution images obtained by the existing spatio-temporal fusion methods.
[0131] To better optimize the spatio-temporal fusion coastal zone feature classification model, in the training stage of the spatio-temporal fusion coastal zone feature classification model, the present application optimizes the spatio-temporal fusion coastal zone feature classification model by using a composite loss function, which includes a pixel-level loss and a feature-level loss. The pixel-level loss includes a spectral consistency loss, a structure consistency loss and an image content loss, and the feature-level loss is obtained on a VGG-19 pre-trained model and is taken from the feature loss between the reconstructed image G p and a reference image T p , and the expression of the feature loss L F is formula (10).
[0132]
[0133] wherein N is a batch size, and respectively represent the reconstructed image G p and the reference image T pThe feature maps of the VGG-19 pre-trained model.
[0134] The spectral consistency loss is represented by a spectral vector angle function, which is the consistency of the spectral vector direction of the reconstructed image G p and the reference image T p . The expression of the spectral consistency loss L pe is shown in equation (11).
[0135]
[0136] The structural consistency loss is used to evaluate the difference in details and structure between the reconstructed image G p and the reference image T p . The structural consistency loss is evaluated by a multi-scale structural similarity index, and the expression of the structural consistency loss L MS-SSIM is shown in equation (12). The expression of the structural similarity index SSIM(G p , T p ) is shown in equation (13).
[0137]
[0138]
[0139] where MS_SSIM(G p , T p ) is a multi-scale structural similarity index, l H , c h and s h are the H and h scale representations of the luminance l, contrast c and structure s in equation (13), and represent the mean values of the reconstructed image G p and the reference image T p , respectively, represents the covariance between the reconstructed image G p and the reference image T p , and represent the variances of the reconstructed image G p and the reference image T p , respectively, and b1 and b2 are constants to avoid the denominator being zero; SSIM(G p , T p ) is a structural similarity index.
[0140] The image content loss is expressed by the mean square error of the reconstructed image G p and the reference image T p , and the expression of the mean square error loss L MSE is shown in equation (14).
[0141]
[0142] The composite loss L of the model Z is expressed as formula (15), γ1, λ2, ξ3 and are weight coefficients.
[0143]
[0144] To verify the effectiveness of the method provided by the present application, the following comparative experiments are carried out:
[0145] Experimental setup: the model implementation framework based on deep learning is PyTorch, the STARFM and FSDAF methods are implemented in IDL, and the workstation device used includes 1 NVIDIA Tesla A100 PCIe GPU with 40 GiB of video memory and Intel(R) Xeon(R) Gold 6326 CPU @ 2.90 GHz. The image block size for training the network is 256x256xC, the batch size is 8, and the training round epoch is 200. The hyperparameter settings in the loss function are γ1=1, λ2=1, ξ2=0.5 and The model is optimized using Adam, and the initial learning rate is 2x10 -4 The deep learning model-based experiment is implemented on a GPU, and the traditional model experiment is implemented on a CPU.
[0146] Comparative experiment: comparative experiments are carried out on the public CIA and LGC data sets with mainstream spatio-temporal fusion methods, wherein the traditional methods in the comparative methods include STARFM and FSDAF, and the deep learning methods include EDCSTFN and GAN-STFM. The experimental results of the model are subjected to subjective visual evaluation and objective index evaluation, and the spatio-temporal fusion results are classified, and the classification results are subjected to index evaluation. Table 1 shows the objective indexes of the fusion results. Table 2 shows the index evaluation results of the classification results. It can be seen that the method proposed in the present application has better ability in terms of spectral change information and structure detail preservation, and the objective index evaluation of the present application in table 1 is also better. The classification index of the present application in table 2 is also better. The above proves that the high spatial resolution image quality obtained by the method of the present application is good, and good results are obtained in feature classification.
[0147] Table 1: objective index evaluation of the fusion results of the comparative methods on the CIA data
[0148]
[0149] Table 2: objective index evaluation of the classification of the fusion results of the comparative methods
[0150]
[0151] Conclusion: The application proposes a kind of coastal zone feature classification method based on spatio-temporal fusion sequence remote sensing image to obtain long-time sequence monitoring data of coastal zone and improve the classification accuracy of coastal zone features. First, generate the degraded image of high spatial resolution image, and then fuse it with low spatial resolution image to obtain composite low-resolution image;Then, the multi-scale decomposition method NSCT is used to extract the high-frequency information and low-frequency information of high spatial resolution image and composite low-resolution image;Then, the high-frequency-low-frequency feature encoder and high-frequency / low-frequency fusion module extract and fuse the features of the above high-frequency information and low-frequency information, and obtain the spatio-temporal fusion feature through inverse transformation;Finally, the convolutional Transformer and convolutional module decode the spatio-temporal fusion feature to obtain long-time sequence high spatial resolution reconstruction image, and use different classifiers to classify the features. The application improves the spatial resolution and time resolution of remote sensing image through spatio-temporal fusion, thereby improving the classification accuracy and realizing long-time sequence precision monitoring of coastal zone features. The experimental results on CIA and LGC data sets show that the application has better fusion ability and classification results in subjective evaluation and objective evaluation.
[0152] Compared with the prior art, the technology of the application includes the following improvements:
[0153] In order to reduce the deviation caused by the imaging mechanism of dual-source satellite remote sensing image, the high spatial resolution image is processed to obtain a degraded image with the same size as the low spatial resolution image, and the degraded image is fused with the low spatial resolution image to generate a composite low-resolution image with high-resolution physical characteristics.
[0154] The multi-resolution decomposition module is used to decompose the high spatial resolution image and the composite low-resolution image to obtain high-resolution high-frequency image HHI and high-resolution low-frequency image HLI, low-resolution high-frequency image LHI and low-resolution low-frequency image LLI.
[0155] A dense convolutional Transformer module is designed to extract the frequency domain local-global features of the above HHI, HLI, LHI and LLI, and fuse the high-resolution high-frequency features HHI and low-resolution high-frequency features LHF to obtain high-frequency fusion features FHF, and fuse the high-resolution low-frequency features HLF and low-resolution low-frequency features LHF to obtain low-frequency fusion features FLF, and then inverse transform the low-frequency fusion features FLF and high-frequency fusion features FHF to obtain spatio-temporal fusion features STFF.
[0156] A decoder composed of convolutional Transformer module is designed to reconstruct the spatio-temporal fusion features to obtain long-time sequence high spatial resolution image, which is transmitted to the classifier to obtain long-time sequence coastal zone feature classification map.
[0157] A composite loss function optimization model composed of content loss, structure loss, spectral difference loss and feature loss is constructed to obtain an optimal result.
[0158] Experiments are conducted on the CIA and LGC data sets, and results show that the method proposed in the application improves the spatial resolution and learns more rich spatiotemporal change information to improve the classification accuracy.
[0159] Corresponding to the method of Figure 1 , with reference to Figure 11 , an embodiment of the application provides a sequence remote sensing image coastal zone feature classification system based on spatiotemporal fusion, comprising:
[0160] at least one processor;
[0161] at least one memory for storing at least one program;
[0162] When the at least one program is executed by the at least one processor, the at least one processor implements the method described above.
[0163] It can be seen that the contents in the above method embodiments are applicable to the system embodiments, the system embodiments specifically implement the functions same as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.
[0164] Those of ordinary skill in the art can understand that all or some of the systems in the above disclosed method can be implemented as software, firmware, hardware and appropriate combinations thereof. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, as known to those of ordinary skill in the art, communication media generally includes computer readable instructions, data structures, program modules or other data in modulated data signals such as carrier waves or other transmission mechanisms, and can include any information delivery medium.
[0165] The above is a specific description of the preferred embodiments of the present disclosure, but the present disclosure is not limited to the above-described embodiments. Those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present disclosure, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present disclosure.
Claims
1. A method for classifying coastal features based on spatio-temporal fusion of sequential remote sensing images, characterized in that, The method comprises the following steps: acquiring sample images, the sample images comprising high spatial resolution images at a prior time and low spatial resolution images at a prediction time; constructing a spatio-temporal fusion coastal zone feature classification model, training the coastal zone feature classification model through the sample images until the composite loss of the coastal zone feature classification model is lower than a set loss threshold, and obtaining a trained coastal zone feature classification model; acquiring remote sensing images to be classified, and classifying coastal zone features in low spatial resolution images at a prediction time in the remote sensing images to be classified through the trained coastal zone feature classification model; wherein the remote sensing images to be classified comprise high spatial resolution images at a prior time and low spatial resolution images at a prediction time; the constructing a spatio-temporal fusion coastal zone feature classification model, training the coastal zone feature classification model through the sample images comprises: constructing a spatio-temporal fusion coastal zone feature classification model; wherein the coastal zone feature classification model comprises a high-frequency-low-frequency feature encoder, a high-frequency-low-frequency fusion module, an inverse transform, a convolutional Transformer decoder, and a classifier; the high-frequency-low-frequency feature encoder comprises a composite low-resolution image generation module, a multi-resolution decomposition module, and a dense convolutional Transformer module; the high-frequency-low-frequency fusion module comprises a high-frequency fusion module and a low-frequency fusion module; acquiring high spatial resolution images and low spatial resolution images in the sample images, performing resolution reduction processing on the high spatial resolution images through the composite low-resolution image generation module to obtain downgraded images of the same scale as the low spatial resolution images, fusing the low spatial resolution images and the downgraded images to obtain composite low-resolution images; decomposing the high spatial resolution images and the composite low-resolution images through the multi-resolution decomposition module to extract a high-resolution high-frequency image and a high-resolution low-frequency image from the high spatial resolution images, and to extract a low-resolution high-frequency image and a low-resolution low-frequency image from the composite low-resolution images; extracting hierarchical features of the high-resolution high-frequency image, the high-resolution low-frequency image, the low-resolution high-frequency image, and the low-resolution low-frequency image through the dense convolutional Transformer module respectively, and obtaining frequency domain local-global features through pixel convolution fusion; wherein the frequency domain local-global features comprise high-resolution high-frequency features, high-resolution low-frequency features, low-resolution high-frequency features, and low-resolution low-frequency features; fusing the extracted frequency domain local-global features through the high-frequency fusion module and the low-frequency fusion module to obtain high-frequency fusion features and low-frequency fusion features; and obtaining spatio-temporal fusion features through inverse transform of the low-frequency fusion features and the high-frequency fusion features; reconstructing the spatio-temporal fusion features through the convolutional Transformer decoder to obtain a long-time-series high spatial resolution reconstructed image, and transmitting the reconstructed image to the classifier to obtain a long-time-series coastal zone feature classification map.
2. The method of claim 1, wherein, the extracting a high-resolution high-frequency image and a high-resolution low-frequency image from the high spatial resolution images, and extracting a low-resolution high-frequency image and a low-resolution low-frequency image from the composite low-resolution images comprises: The high spatial resolution image is decomposed by a first NSPF to obtain a first high-frequency low-frequency image and a first high-frequency high-frequency image; The first high-frequency low-frequency image is decomposed by a second NSPF to obtain a second high-frequency low-frequency image and a second high-frequency high-frequency image, and the second high-frequency low-frequency image is decomposed by a third NSPF to obtain a third high-frequency low-frequency image and a third high-frequency high-frequency image; the third high-frequency low-frequency image is the high-frequency low-frequency image; The first high-frequency high-frequency image, the second high-frequency high-frequency image, and the third high-frequency high-frequency image are filtered by a first NSDF, a second NSDF, and a third NSDF respectively to obtain a first high-frequency high-frequency sub-band image, a second high-frequency high-frequency sub-band image, and a third high-frequency high-frequency sub-band image, and the first high-frequency high-frequency sub-band image, the second high-frequency high-frequency sub-band image, and the third high-frequency high-frequency sub-band image constitute the high-frequency high-frequency image.
3. The method of claim 1, wherein, The low-frequency high-frequency image and the low-frequency low-frequency image are extracted from the composite low-resolution image, including: The composite low-resolution image is decomposed by a first NSPF to obtain a first low-frequency low-frequency image and a first low-frequency high-frequency image; The first low-frequency low-frequency image is decomposed by a second NSPF to obtain a second low-frequency low-frequency image and a second low-frequency high-frequency image, and the second low-frequency low-frequency image is decomposed by a third NSPF to obtain a third low-frequency low-frequency image and a third low-frequency high-frequency image; the third low-frequency low-frequency image is the low-frequency low-frequency image; The first low-frequency high-frequency image, the second low-frequency high-frequency image, and the third low-frequency high-frequency image are filtered by a first NSDF, a second NSDF, and a third NSDF respectively to obtain a first low-frequency high-frequency sub-band image, a second low-frequency high-frequency sub-band image, and a third low-frequency high-frequency sub-band image, and the first low-frequency high-frequency sub-band image, the second low-frequency high-frequency sub-band image, and the third low-frequency high-frequency sub-band image constitute the low-frequency high-frequency image.
4. The method according to claim 2 or 3, characterized in that, The NSPF includes a low-pass filter and a high-pass filter, the low-pass filter including a low-pass decomposition filter and a low-pass reconstruction filter, the high-pass filter including a high-pass decomposition filter and a high-pass reconstruction filter; the NSPF satisfies the identity L f (I) L c (I) + G f (I) G c (I) = 1; The NSDF includes a sector filter and a checkerboard filter, the sector filter including a sector decomposition filter and a sector reconstruction filter, the checkerboard filter including a checkerboard decomposition filter and a checkerboard reconstruction filter; the NSDF satisfies the identity S f (I)S c (I)+C f (I)C c (I) = 1; where I is the input image, L f (·) is a low-pass decomposition filter function, L c (·) is a low-pass reconstruction filter function, G f (·) is a high-pass decomposition filter function, G c (·) is a high-pass reconstruction filter function; S f (I) is a sector decomposition filter function, S c (I) is a sector reconstruction filter function, C f (I) is a checkerboard decomposition filter function, C c (I) is a checkerboard reconstruction filter function.
5. The method of claim 1, wherein, The hierarchical features and the frequency domain local-global features of the high-frequency high-frequency image, the high-frequency low-frequency image, the low-frequency high-frequency image, and the low-frequency low-frequency image are extracted by the dense convolution Transformer module, and are calculated by the following formula: Among them, TFI and TFO are the frequency domain local-global features of the input image and output of the dense convolution Transformer module respectively. The input image includes high-resolution high-frequency images, high-resolution low-frequency images, low-resolution high-frequency images and low-resolution low-frequency images, and the frequency domain local-global features include high-resolution high-frequency features, high-resolution low-frequency features, low-resolution high-frequency features and low-resolution low-frequency features; TF i is the hierarchical feature output by the i-th convolutional Transformer module, i = 1, 2, 3, 4, f CT (·) is the function of the convolutional Transformer module, f concat (·) is the function of the splicing operation; The convolution Transformer module is composed of layer normalization, multi-convolution head transpose attention, local enhancement feedforward network, and residual connection, and is mathematically expressed as: T = f CT (x); where x represents the input of the convolutional Transformer module, X represents the normalized image of x, W q,1 (·), W k,1 (·) and W v,1 (·) represent 1x1 pixel convolution for generating Q, K and V respectively, and 3x3 deep convolution for generating Q, K and V respectively, Q, K and V represent the query, key and value of the multi-convolution head transpose attention respectively, R(·) is a function of the deformation operation, β is a learnable scaling parameter, A is the generated self-attention graph, F1 is the enhanced feature after the self-attention weight, F is the corrected feature after the multi-convolution head transpose attention, T1 is the normalized image of F, T2 is the feature of T1 after the enhanced local, W 1 (·) is 1x1 pixel convolution, GELU(·) is Gaussian error linear unit function, 3x3 deep convolution, and T represents the output of the convolutional Transformer module.
6. The method of claim 5, wherein, The high-frequency fusion module and the low-frequency fusion module fuse the extracted frequency domain local-global features to obtain high-frequency fusion features and low-frequency fusion features, including three cross-convolution attention modules, the cross-convolution attention module is composed of layer normalization, multi-cross-convolution head transpose attention, local enhancement feedforward network, and residual connection, and is realized by the following formula: wherein x f1 represents high-score high-frequency features or low-score low-frequency features, x f2 represents high-score low-frequency features or low-score high-frequency features, X f1 and X f2 represent normalized features of x f1 and x f2 , W qf,1 (·), W kf,1 (·) and W vf,1 (·) represent 1×1 pixel convolution for generating Q f , K f and V f , respectively, and represent 3×3 deep convolution for generating Q f , K f and V f , respectively, Q f , K f and V f represent queries, keys and values of the multi-cross convolution head transpose attention, CA is a generated cross attention map, CF1 is a cross enhanced feature after cross attention weight, CF is a cross fusion feature after multi-cross convolution head transpose attention, FT1 is an image after normalization of CF, FT2 is a feature after enhanced local of FT1, and FT represents a high-frequency fusion feature or a low-frequency fusion feature output by the fusion module.
7. The method of claim 6, wherein, The long-time high spatial resolution reconstructed image is obtained by reconstructing the spatio-temporal fusion features by the convolution Transformer decoder, including: The spatio-temporal fusion features are decoded and restored by the four convolution Transformer modules in the decoder to recover the image change spectral information and detail information; The high spatial resolution image is reconstructed by the two 3*3 convolutions with ReLU and the last 3*3 convolution in the decoder; The reconstructed image at the prediction time is generated by the following formula: where G p denotes the reconstructed image at time p, denotes a 3x3 convolution followed by a ReLU activation function, and STFF is the spatio-temporal fusion feature.
8. The method of claim 1, wherein, The composite loss of the coastal feature classification model is calculated by the following method: The reconstructed image and the corresponding reference image are obtained; determine a composite loss of the coastal zone feature classification model based on the reconstructed image and a corresponding reference image; wherein the composite loss comprises a pixel-level loss and a feature-level loss, and the pixel-level loss comprises a spectral consistency loss, a structure consistency loss and an image content loss; The feature loss L F The expression is: where N is the batch size, and respectively represent the reconstructed image G p and the reference image T p at the feature map of the VGG-19 pre-trained model; The spectral consistency loss L pe The expression is: The structural consistency loss L MS_SSIM The expression is: where MS_SSIM(G p ,T p ) is a multi-scale structural similarity index; l H is an H-scale representation of luminance l, c h is an h-scale representation of contrast c, s h is an h-scale representation of structure s, and represent the mean of reconstructed image G p and reference image T p , respectively, represents the covariance between reconstructed image G p and reference image T p , and represent the variance of reconstructed image G p and reference image T p , respectively, b1 and b2 are constants to avoid the denominator being zero; SSIM(G p ,T p ) is a structural similarity index. an expression of the image content loss is: The composite loss function L Z The expression is: wherein γ1, λ2, ξ3and are weight coefficients.
9. A spatio-temporal fusion-based sequence remote sensing image coastal zone feature classification system, characterized in that, The system comprises: at least one processor; at least one memory for storing at least one program; when the at least one program is executed by the at least one processor, the at least one processor implements the spatio-temporal fusion-based sequence remote sensing image coastal zone feature classification method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Reference data non-sensitive remote sensing image space-time fusion model construction method
CN112529828A
Multi-source remote sensing image fusion method and system
CN118968245A