Water quality remote sensing image spatio-temporal data prediction method based on residual dense block improved VPTR
By constructing a video prediction model based on residual dense blocks, the problem of difficult to capture space-time relationships and low accuracy in traditional water quality remote sensing image processing is solved, and high-precision water quality remote sensing image prediction is achieved.
Patent Information
- Application Number
- CN202510438447.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional water quality remote sensing image processing methods are difficult to effectively capture the spatial and temporal relationship between images, especially in long-term prediction, which will lead to the loss of details and timing information. The existing video prediction models have low accuracy and noise problems when processing remote sensing image data, which cannot meet the needs of water quality remote sensing image prediction.
Using a video prediction model based on residual dense blocks, combined with a residual dense block encoder, an improved video prediction Transformer model and a PatchGAN discriminator, the image prediction accuracy is improved by constructing a chlorophyll a inversion model, data preprocessing, feature extraction and spatiotemporal feature encoding.
It enhances the feature extraction ability of water quality remote sensing images, improves the model's expression ability and prediction accuracy, can better capture spatial and temporal information and details changes, and improves the accuracy and stability of remote sensing image prediction.
Smart Images

Figure CN120339865A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of artificial intelligence technology and environmental science, relates to remote sensing monitoring technology, and particularly relates to a spatio - temporal prediction method for a video prediction model of water quality remote sensing image data. Background Art
[0002] Water quality monitoring and ecological health assessment have become important topics in environmental protection and water resource management.
[0003] Traditional water quality monitoring methods mainly rely on manual field sampling and laboratory analysis. Although they can provide accurate water quality data, this method can only reflect the water quality status at a certain moment and specific location, and it is difficult to capture the dynamic changes of water body quality. Therefore, it is impossible to comprehensively understand the spatio - temporal variation characteristics of water bodies. With the rapid development of remote sensing technology, large - scale water quality monitoring based on remote sensing images has become an effective supplementary method. Remote sensing images can provide real - time and dynamic water quality information of large - area water bodies, greatly improving the monitoring efficiency and being able to give early warnings of water quality change trends.
[0004] Water quality remote sensing images usually have complex spatio - temporal variation characteristics. Especially in the analysis of dynamic water body pollution monitoring, temperature changes, and biological distribution, the changes in water quality often have strong temporal dependence. Traditional image processing methods are difficult to fully consider the spatio - temporal relationship between images when predicting water quality changes, especially in long - term prediction tasks, which often lead to the loss of details and temporal information.
[0005] To solve this problem, the video prediction model, as an important research direction in the field of computer vision, has been gradually introduced into the prediction task of water quality remote sensing images. The video prediction model can infer the changes in future frame images by capturing the spatio-temporal information in the image sequence. Among them, the autoencoder model, as an effective unsupervised learning method, plays an important role in extracting the latent features and temporal dependencies of images. The autoencoder can compress the high-dimensional image into a low-dimensional latent space representation and then reconstruct the original image through the decoder, which can effectively extract spatio-temporal information and provide support for the prediction of future images. However, when dealing with complex spatio-temporal dependencies, especially in the long-term prediction of water quality remote sensing images, the traditional autoencoder model often ignores the detailed changes and subtle temporal changes in the images, resulting in a decrease in prediction accuracy. Therefore, to address the above problems, the present invention proposes an improved autoencoder-based video prediction method. By introducing a residual dense block structure, the ability to extract various features in water quality remote sensing images can be enhanced, the flow of information in the network can be strengthened, the problem of gradient disappearance can be effectively alleviated, the training of deep networks can be promoted, and the expression ability and performance of the model can be improved. At the same time, the SMU activation function is used to replace the traditional ReLU activation function, which can keep the neurons active in the deep neural network and promote more effective gradient propagation, thereby further improving the performance of the water quality remote sensing image prediction model. In addition, the contrastive loss function is used to strengthen the similarity learning of images in the latent space, which can more accurately capture spatio-temporal information and detailed changes, thereby improving the prediction accuracy of water quality remote sensing images. The combination of these technologies provides a more accurate remote sensing image processing method for water quality monitoring and assessment. Summary of the Invention
[0006] The video prediction model can effectively capture and utilize the spatio-temporal information of remote sensing images, enhance the prediction ability of future image changes, and significantly improve the accuracy and stability of remote sensing image prediction by introducing the technology of the video prediction model. However, there are currently two problems in remote sensing image prediction based on the video prediction model: First, there are problems such as distortion and noise in the original remote sensing image data, resulting in the inability to directly apply the original remote sensing image data to the prediction of water quality; Second, the existing video prediction models have low accuracy when processing remote sensing image data and cannot effectively capture the key feature information in remote sensing images, requiring efficient feature extraction and denoising technologies. Therefore, the traditional video prediction model cannot fully meet the needs of water quality remote sensing image prediction. The purpose of the present invention is to provide a method for inverting and predicting water quality remote sensing images based on a video prediction model with a residual dense block, providing a new idea for the expansion and prediction of remote sensing images in the cross-disciplinary field of artificial intelligence technology and environmental science.
[0007] A method for retrieving and predicting water quality remote sensing images provided by the present invention preprocesses the original remote sensing images through the following step one to obtain a dataset of chlorophyll a retrieval images, and constructs a video prediction model based on a residual dense block through the following steps two and five, including an encoder module based on a residual dense block, an improved VPTR module, a decoder module based on an encoder of a residual dense block, and a generative discriminant module, for water quality prediction. The video prediction model based on a residual dense block is expressed as Residual in Residual Dense Block and Generative Adversarial Networks of Video Prediction Transformer, abbreviated as RRDB-GAN-VPTR.
[0008] The method of the present invention first obtains the remote sensing image time series of the target water area, and then executes the following five steps:
[0009] Step one, construct a chlorophyll a retrieval module for the target water area. For the remote sensing images provided by the China Resources Satellite Data Center, data preprocessing is performed on the original remote sensing images by means of radiometric calibration, atmospheric correction, and orthorectification; water body extraction is performed on the processed remote sensing images; for chlorophyll a, the pixel values of sensitive bands are selected and the measured chlorophyll a data of the target water area are used to construct a chlorophyll a retrieval model, and batch processing is performed to obtain the time series data of chlorophyll a retrieval images;
[0010] Step two, construct an encoder module based on a residual dense block. The encoder part performs convolutional calculations on the remote sensing images of the target water area at multiple scales and combines residual connections to extract the deep features of the remote sensing images of the target water area; the shallow image features obtained through convolution; the dense residual block can abstract continuous features into more dimensions, fully learn the local features of the image, and improve the generalization ability of the model to construct a residual dense module (Residual Dense Block, RDB) and downsample to extract the deep features of the image. The residual dense module consists of four convolutional layers, an SMU activation function, and a dense feature convolutional layer with a convolution kernel size of 1×1; the RDB modules are connected through dense connections. After fully capturing the internal features of the module, the deep features of the image are obtained through the downsampling layer, and finally the residual connection is used to output the image features generated by the encoder, and the image feature sequence generated by the encoder is made to match the input size of the VPTR module;
[0011] Step 3: Construct a feature encoder module and a decoder module based on the improved VPTR. In the improved VPTR model, to enhance the mutual information between image feature sequences, the model introduces the SupCon contrast loss function. This loss function maximizes the distance between similar features and minimizes the distance between different features through contrastive learning of image feature sequences. In this way, the model can more effectively learn image features with high similarity between different time steps, thereby improving the accuracy and robustness of future image prediction. Input the target water area image feature sequence generated by the encoder part in Step 2 into the transformer encoder module of the improved VPTR. The transformer encoder module combines a local spatial multi-head self-attention module and a temporal multi-head attention module to perform spatio-temporal feature learning and encoding on the image feature sequence. The transformer decoder module has one more temporal multi-head attention layer and one output conversion layer than the encoder module, and finally obtains the image feature sequence at the future moment;
[0012] Step 4: Construct a decoder module based on a residual dense block encoder. The decoder part uses an upsampling layer and the SMU activation function. Send the image feature sequence at the future moment obtained by the improved VPTR model into the decoder module of the encoder, so that the spatial dimension of the feature map gradually returns to the size of the input image, and give the remote sensing image of the target water area at the future moment. Let Steps 2, 3, and 4 constitute the generator module;
[0013] Step 5: Construct a discriminator module based on PatchGAN. Use convolutional layers, batch normalization layers, and ReLU activation functions to divide the input image into multiple small blocks, process each small block independently, and give a fine-grained remote sensing image of the target water area at the future moment.
[0014] Compared with the prior art, the advantages of the present invention are as follows:
[0015] (1) The present invention constructs a new chlorophyll inversion model by establishing a statistical relationship between remote sensing data and ground-measured water quality data.
[0016] (2) The present invention constructs an encoder module of a residual dense block. By using dense connections and residual connections, the feature extraction function is strengthened, feature information is fully utilized, and the SMU activation function is used to avoid gradient disappearance.
[0017] (3) The present invention constructs an improved video prediction model based on a residual dense block, and introduces the supcon contrast loss function to improve the mutual information between the image feature sequence at the future moment predicted by the video prediction module and the image feature sequence generated by the residual dense block encoder, further improving the prediction accuracy of the model.
[0018] (4) By constructing a decoder module based on the residual dense block autoencoder module, the present invention effectively combines the image features generated by the encoder, can better process time series information, and outputs high-quality water area remote sensing images.
[0019] (5) By constructing a discriminator module based on PatchGAN and combining modules such as feature extraction and spatio-temporal feature encoding in the previous steps, the present invention forms a complete generation and discrimination structure, ensuring the high-quality generation of future moment images. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a schematic flow chart of the method for retrieving and predicting water quality remote sensing images of the present invention;
[0021] Figure 2 is a schematic diagram of the water quality remote sensing image prediction structure based on RRDB-GAN-VPTR of the present invention;
[0022] Figure 3 is a schematic diagram of the encoder structure based on the residual dense block of the RRDB-GAN-VPTR module of the present invention;
[0023] Figure 4 is a schematic diagram of the PatchGAN discriminator structure of the RRDB-GAN-VPTR model of the present invention;
[0024] Figure 5 is a schematic diagram of the decoder structure of the RRDB-GAN-VPTR model of the present invention;
[0025] Figure 6 is a sample of the original remote sensing image data in the Beijing area of the present invention;
[0026] Figure 7 is a comparison diagram of the preprocessing of remote sensing image data of the present invention;
[0027] Figure 8 is a sample of the retrieved image data of chlorophyll a in Guanting Reservoir of the present invention;
[0028] Figure 9 is the complete time series image data of chlorophyll a in Guanting Reservoir of the present invention;
[0029] Figure 10 is a comparison diagram of the results of the ablation experiment of the retrieved image prediction of chlorophyll a in Taihu Lake of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0031] The present invention proposes a method for retrieving and predicting water quality remote sensing images, which consists of two parts: batch retrieval of remote sensing images by an inversion model and prediction of water quality remote sensing images by an improved VPTR model based on a residual dense block. Among them, the process of constructing the chlorophyll-a inversion image based on the remote sensing image and the measured chlorophyll-a data of the target water area is step one, which solves the problem that the remote sensing image of the target water area cannot be directly used for water quality prediction; the process of constructing the improved VPTR model based on the residual dense block is steps two, three, four, and five, which solves the problem of low accuracy of traditional video prediction methods.
[0032] The flow of the water quality remote sensing image retrieval and prediction method according to the embodiment of the present invention is as Figure 1 shown; the improved VPTR prediction structure based on the residual dense block according to the embodiment of the present invention is as Figure 2 shown. First, obtain the time series of remote sensing images of the target water area, and then the method of the present invention will be specifically described through the following five steps.
[0033] Step one: Construct an inversion model to realize batch chlorophyll-a inversion images. In the embodiment of the present invention, the water quality remote sensing images all refer to chlorophyll-a inversion images.
[0034] The original remote sensing images are affected by the external environment and have distortion problems. In order to eliminate the influence, data preprocessing needs to be performed on the remote sensing images. The data preprocessing adopts the methods of radiometric calibration, atmospheric correction, and orthorectification. The present invention uses the radiometric calibration in ENVI 5.3, the FLASSH atmospheric correction module, and the orthorectification workflow tool to realize data preprocessing. The spectral response function is cited from the China Resource Satellite Application Center. In the orthorectification process, the global DEM data with a resolution of 30 seconds (about 900m) provided by ENVI itself is used, and cubic convolution interpolation is adopted. Since the area covered by the remote sensing satellite image is much larger than the target area, the preprocessed remote sensing image is cropped for the target area, which facilitates the processing of remote sensing satellite image data and can reduce the workload of the computer, thereby shortening the data processing time. The water body is extracted using the Normalized Difference Water Index (NDWI). NDWI takes advantage of the fact that the reflection ability difference between the green light band and the near-infrared band of the water body is the largest, and further highlights the reflection characteristics of the water body through normalization calculation. Using the GF-1 spectral band, NDWI is formalized in the form of a formula as:
[0035]
[0036] In the embodiment of the present invention, B2 is the reflectance of the green light band of the GF-1 data, and B4 is the reflectance of the near-infrared band of the GF-1 data.
[0037] Different water quality parameters have different sensitivities to different bands. Therefore, it is necessary to select an appropriate band combination as the input factor of the model. In the present invention, a chlorophyll a inversion model is constructed based on the pixel values of the sensitive bands selected for chlorophyll a and the measured chlorophyll a data of the target water area, and is formalized in the formula as:
[0038] Y = 553.9e-2.908x (2)
[0039] where Y is the inversion concentration value of chlorophyll a, x is the pixel value of the B1 / B2 band, and e is the natural constant.
[0040] Step 2: Construct an encoder module based on a residual dense block to generate a target water area image feature sequence.
[0041] The specific structure diagram of the encoder module based on the residual dense block is as shown in Figure 3 shown. Assume I t-9 , I t-8 ,... I t and F t-9 , F t-8 , …, F t are the input image and the output reconstructed image feature sequence respectively, and I t is the image at time t in the time series data. First, it enters the encoder and uses a convolutional layer to extract the preliminary shallow features F t 1 of the image. The formula is as follows:
[0042] F t 1 = H COM (I t ) (3)
[0043] where H COM represents the composite function operation of convolution and the SMU activation function.
[0044] The dense residual block can abstract continuous features into more dimensions, fully learn the local features of the image, and improve the generalization ability of the model. F t 1 is used as the input of the residual dense module. The residual dense block module of the present invention is composed of four convolutional layers, the SMU activation function, and a dense feature convolutional layer with a convolution kernel size of 1×1. The specific structure is as shown in Figure 2 shown. The number of convolution kernels in the convolutional layer is 128, and the size of the 4-layer convolution kernels is 3×3. The SMU activation function can effectively maintain the active state of neurons and can maintain effective gradient transmission in the deep network. The dense feature convolutional layer is mainly used for multi-channel feature fusion.
[0045] In the dense residual network, first, the feature information is adaptively fused in each layer, and then they are passed to the next layer for feature fusion. It can input the shallow features F extracted by the encoder t 1 into each layer of the dense residual network. Then, the output F of the c-th convolutional layer t r,c can be expressed by the following formula:
[0046]
[0047] where τ represents the SMU activation function, r represents the r-th RDB module, ω t r,c represents the weight of the c-th convolutional layer of the r-th RDB module at time t, and [F t 1 , F t r,1 , F t r,2 , …, F t r,c-1 refers to the concatenation of the shallow feature F at time t t 1 and the feature maps generated by convolutional layers 1, …, (c - 1) in the r-th RDB module at time t.
[0048] The present invention adopts a method of local feature fusion to adaptively fuse the features from all Conv layers in the previous RDB module and the current RDB module. The 1×1 convolutional layer in the residual dense module of the present invention can adaptively control the output information, that is, the local feature fusion operation, and perform local feature fusion on the input of the previous residual dense module and the outputs of all convolutional layers in the current residual dense module, and then output F t r,LFF , and the following is the formula:
[0049]
[0050] where represents the 1×1 Conv layer operation in the r-th residual dense module, LFF represents the local feature fusion operation, and F t r,LF represents the output after local feature fusion of the r-th RDB module at time t.
[0051] To further improve the information flow of the convolutional layer, residual learning is introduced in the RDB module, and the following is the formula:
[0052] F t RDB,r = F t r,LFF + Ft 1 (6)
[0053] Among them, F t RDB,r represents the output of the r-th RDB module at time t.
[0054] Each residual dense block is connected by dense connections, and after fully capturing the internal features of the module, F is output t RDB , and the following is the formulation:
[0055]
[0056] Among them, H DFF (·) represents the calculation of global feature fusion information in all residual dense modules, represents the residual coefficient, and F t RDB represents the output after passing through r RDB modules at time t.
[0057] Taking F t RDB as the input of the downsampling layer, the deep image feature F t 2 is extracted, and the following is the formulation:
[0058] F t 2 = H conv (F t RDB ) (8)
[0059] Among them, H conv (·) is the convolution calculation of the downsampling layer.
[0060] Finally, for the output F t 2 of the entire residual dense module and the downsampling layer, a residual connection is adopted to obtain the image feature F t generated by the encoder, and the following is the formulation:
[0061] F t = F t 1 + F t 2 (9)
[0062] Step 3: Construct an improved contrast loss function based on VPTR to predict the image feature sequence at future times.
[0063] The image feature sequence F t-9 , F t-8 , …, F tAs the encoder input of the VPTR model, the VPTR model was proposed by Ye et al.
[0064] After the generated image feature sequence is input into the VPTR model, since the transformer decoder part is prone to generating duplicate feature sequences, to solve this problem, the present invention improves the contrast loss function of VPTR. In order to enable the model to better capture the mutual information between the predicted value and the true value at a certain time, a contrast loss based on supcon is adopted, and the specific formulation is as follows:
[0065]
[0066] Where A represents all samples corresponding to the feature vector v, and P represents the positive sample v corresponding to the feature vector v + , and v + represents the spatial true feature vector corresponding to all feature vectors v within a certain time period, and v - represents all samples except the positive sample v + . S(·,·) represents the cosine similarity function.
[0067] Therefore, the improved VPTR contrast loss function is:
[0068]
[0069] Where X∈[t + 1,t + 10] is the future moment image feature sequence predicted by the VPTR model, F X ,X∈[t + 1,t + 10] is the true future moment image feature sequence generated based on the residual dense block encoder, v s + represents the feature vector at the spatial position s of F X , represents the set of feature vectors at all other spatial positions of F X , represents the feature vector at the spatial position s, represents the set of feature vectors at all other spatial positions, S l represents the total number of all spatial positions in the feature map as H×W, and sg is the stop gradient operation.
[0070] Step Four: Construct a decoder module based on the residual density block encoder to realize the reconstruction of the feature sequence into an image.
[0071] The decoder module based on the residual density block encoder constructed by the present invention is composed of an upsampling layer, an SMU activation function, and a convolutional layer, so as to reconstruct the image. The specific structure is asFigure 4 As shown in the figure. The feature sequence at time t+1 in the predicted future moment image feature sequence obtained in Step 3 is used as the input of the decoder, and then is amplified through two transposed convolution operations to obtain The specific formula is:
[0072]
[0073] where H Norm (·) is the layer normalization function, W1 and W2 are the weights of the upsampling layer, is the transposed convolution operation.
[0074] Finally, the image is output through the convolutional layer F R The specific formula is: The specific formula is:
[0075]
[0076] where δ is the Tanh activation function, and H co (·) is the convolution operation.
[0077] Step 5: Construct a discriminator module based on PatchGAN to realize the output of the predicted future water quality image.
[0078] The PatchGAN discriminator module constructed in the present invention is a complete convolutional topology network, which is composed of a convolutional layer, a normalization layer, and a LeakyReLU activation function layer. The specific structure is as Figure 5 shown. Its core idea is to divide the input image into multiple small blocks and independently distinguish each small block. Due to the independent processing of each small block, this discriminative network enables the model to pay more attention to local structures and details rather than global consistency. The output is an N×N matrix.
[0079] The receptive field size of each convolutional layer of the discriminator is shown in Table 1, where I is the size of the input image, and k i (i = 1, 2, 3, 4, 5) is the convolutional kernel size, Padding i is the padding amount, Stride i is the stride size, and n i is the size of the output matrix.
[0080] Table 1 Receptive field size of PatchGAN
[0081]
[0082] Example 1:
[0083] The dataset used in this invention is the remote sensing images provided by the China Resource Satellite Application Center. Considering that chlorophyll a is used as the main indicator for water ecological prediction modeling, therefore, this invention uses the chlorophyll a inversion image as the dataset for modeling. However, the original remote sensing image data of the target water area is affected by the external environment and has problems such as distortion, and cannot be directly used for modeling prediction. Therefore, it is necessary to perform data preprocessing on it and use the chlorophyll a inversion model for batch inversion. The experimental dataset in this invention is shown in Table 2.
[0084] Table 2 Experimental Dataset of This Invention
[0085]
[0086] Among them, a sample of GF-1 WFV remote sensing image data of a reservoir in North China, China, is as Figure 6 shown. The method of this invention is processed as follows in Steps 1 to 5.
[0087] Step 1: Construct an inversion model to achieve batch chlorophyll a inversion images.
[0088] Radiometric calibration is the process of converting the digital value (DN) recorded by the sensor into the radiance value. Radiometric calibration can be divided into three categories according to the calibration location, namely laboratory calibration, on-board or on-satellite calibration, and field calibration. This invention uses Radiometric Calibration in the ENVI 5.3 software toolbox to process the remote sensing images. The comparison of the images before and after radiometric correction is as Figure 7 shown.
[0089] Atmospheric correction is the process of eliminating the distortion caused by atmospheric scattering, absorption, and reflection in the remote sensing image and converting the radiance or apparent reflectance into the true surface reflectance. Generally speaking, atmospheric correction can be divided into two types: statistical type and physical type. The physical type is to establish a causal relationship according to the physical laws of the remote sensing system to obtain a physical model. It is an abstraction of reality and is relatively complex, such as the classic atmospheric radiation transfer model. This invention only involves physical atmospheric correction. For specific parameter settings, the GF-1 model is selected, the resolution is selected as 16m, the average elevation parameter needs to use GMTED2010 built in ENVI to calculate the average elevation of the specific image, the atmospheric model is selected according to the longitude, latitude, and acquisition time of the data, and the specific spectral response function of GF-1 WFV is provided by the China Resource Satellite Application Center. The comparison of the images before and after atmospheric correction is as Figure 7 shown.
[0090] Remote sensing images are bound to be affected by the external environment during the shooting process. Among these factors, changes in scale, changes in the attitude / orientation of the sensor, and system errors generated by the sensor will cause the pixel geographic location coordinates of the objects in the remote sensing satellite image to be displaced from the actual situation. In order to eliminate the image deformation caused by the terrain or camera orientation and correct the plane coordinates of the image, the image needs to be orthorectified. The present invention mainly utilizes the orthorectification process tool provided by the ENVI software to perform batch radiation calibration on the GF-1WFV data, using the global DEM data with a resolution of 30 seconds (about 900m) provided by ENVI, and adopts cubic convolution interpolation. The image contrast after orthorectification is as follows: Figure 7 shown.
[0091] Because the area covered by remote sensing satellite images is much larger than the target area, in order to facilitate the processing of remote sensing satellite image data and reduce the workload of computers, thereby shortening the data processing time, the remote sensing image is cropped, and then the normalized difference water index is used to extract the water body in the target area to obtain the water body image of Guanting Reservoir.
[0092] Since the composition and state of water bodies will affect the spectral characteristics of water bodies, it is possible to reduce redundant information and extract target information by combining bands. Different water quality parameters have different sensitivities to different bands, so it is necessary to select appropriate band combinations as input factors of the model. The present invention selects B2 and B4 bands as sensitive bands to establish a chlorophyll a inversion model, and uses ArcGIS software to unify the color scale of the chlorophyll a inversion image. The chlorophyll a inversion image data sample of Guanting Reservoir is as follows: Figure 8 shown.
[0093] Step 2: Construct an encoder module based on residual dense blocks to generate a feature sequence of the target water area image.
[0094] In network training, in order to reduce the difficulty of training and perfectly reconstruct all images of the entire training set, the present invention adopts segmented training. In the first part, the VPTR module is excluded, and the encoder and decoder modules based on the residual dense block are trained separately. In the second part of training, only the VPTR module parameters are updated to reduce the complexity of the model as much as possible. The segmented training can simplify the training process of the model.
[0095] In the encoder module based on residual dense blocks, the initial learning rate of this experiment is set to 0.02, the batch size is set to 64, and the overall training cycle is set to 3000 cycles. Secondly, the Adam optimizer is chosen, and its adaptive learning rate adjustment feature helps the model to optimize the loss function more effectively during training.
[0096] The dataset obtained by the encoder module based on the residual dense block is input with data samples such as Figure 9 as shown. Random horizontal and vertical flips are adopted as data augmentation to generate an image feature sequence. Among them, the network layer parameters of the encoder module based on the residual dense block are shown in Table 3.
[0097] Table 3 Parameter settings of the encoder based on the residual dense block
[0098] Network layer Number of input channels Number of output channels Activation function Convolutional layer 3 64 SMU Residual dense block layer 64 64 SMU Downsampling layer 64 528 None
[0099] During the training process of this network, in each batch of embodiments of the present invention, 20-time Guanting Reservoir chlorophyll a inversion images are used as input. These images contain rich spatio-temporal information, which helps the model capture the dynamic changes of the Guanting Reservoir chlorophyll a inversion images. Through learning these historical images, the model can generate an image feature sequence with obvious features. Each layer of convolution in the residual dense block is connected to all previous convolutional layers, so that the feature information extracted by each layer can be fully utilized. At the same time, the characteristics of dense connection can effectively improve the problems of network gradient dispersion and gradient explosion, and also play a role in reducing overfitting. The present invention effectively utilizes the outputs of each residual dense block through dense connection of multiple residual dense block structures, so that the module can extract features more fully and effectively in the deep model.
[0100] Step 3: Construct an improved contrast loss function based on VPTR to predict the image feature sequence at future times.
[0101] The Guanting Reservoir chlorophyll a inversion image feature sequence F t-9 , F t-8 , …, F t obtained in Step 2 is fed into the VPTR model. In this experiment, the initial learning rate of the VPTR model is set to 1e-4, the number of input channels is set to 528, and the number of training epochs is set to 6000 epochs. The convolutional kernel size of the local spatial multi-head self-attention is 4×4. The number of encoder layers is 4, and the number of decoder layers is 8. The Adam optimizer is selected to optimize all transformers, and its characteristic of adaptive learning rate adjustment helps the model optimize the loss function more effectively during the training process.
[0102] Step 4: Construct a decoder module based on the residual density block to reconstruct the feature sequence into an image.
[0103] The predicted future-time image feature sequence obtained in Step 3 is used as the input of the decoder. The specific network layer parameters of the decoder based on the residual density block constructed by the present invention are shown in Table 4.
[0104] Parameter settings of the chlorophyll a inversion image feature decoder for Guanting Reservoir
[0105] Network layer Parameter Activation function ReflectionPad2d intputsize = 3 None Conv2D_1 Filters = 128, Kernelsize = 3 SMU Conv2D_2 Filters = 64, Kernelsize = 3 SMU ReflectionPad2d intputsize = 3 None Conv2D_3 Filters = 3, Kernelsize = 7 SMU
[0106] Step 5: Construct a discriminator module based on PatchGAN to realize the output for predicting the water quality image at a future moment.
[0107] The PatchGAN discriminator module constructed in the present invention consists of a convolutional layer, a normalization layer, and a LeakyReLU activation function layer. The specific network parameters are shown in Table 5 in detail.
[0108] Table 5 Parameter settings of PatchGAN discriminator
[0109] Network layer Number of convolutional kernels Size of convolutional kernel Stride Padding Activation function Conv2D-1 64 4×4 2 1 LeakyReLU Conv2D-2 128 4×4 2 1 LeakyReLU Conv2D-3 256 4×4 2 1 LeakyReLU Conv2D-4 512 4×4 1 1 LeakyReLU Output layer 1 4×4 1 1 None
[0110] During the training process of this network, in each batch of embodiments of the present invention, the chlorophyll a inversion images of Guanting Reservoir at the previous 10 moments are used as inputs. These images contain rich spatio-temporal information, which helps the model capture the dynamic changes of the chlorophyll a inversion images of Guanting Reservoir. By learning these historical images, the model can predict the chlorophyll a inversion image of Guanting Reservoir at the 11th moment, realizing an effective prediction of the chlorophyll a inversion image of Guanting Reservoir at a future moment. The structural characteristics of the residual dense block and dense connection enable it to pay attention to the local details of the image when processing remote sensing images, improving the prediction accuracy and robustness.
[0111] Verified by experiments, when comparing the three models of VPTR, GAN-VPTR, and RRDB-GAN-VPTR in predicting the chlorophyll a inversion images of Guanting Reservoir at the next 10 moments, the specific comparison results are as Figure 10 shown. Table 6 shows the average indicators and improvement effects of each ablation experiment part. Introducing the residual dense block improves the prediction accuracy of the model.
[0112] Table 6 Evaluation indicators of the ablation experiment of the RRDB-GAN-VPTR model
[0113] Model SSIM PSNR LPIPS Optimization ratio VPTR 0.71732 23.9544 44.6995 0% GAN-VPTR 0.74115 25.7152 43.6842 4.31% RRDB-GAN-VPTR 0.76113 26.9355 42.2153 3.60%
[0114] Through Figure 10 Comparing the prediction results of each model and analyzing the results of the three evaluation indicators of each model in Table 6, the following conclusions can be obtained:
[0115] Compared with the two models of GAN-VPTR and VPTR, the RRDB-GAN-VPTR model proposed in the present invention is the model with the best effect in predicting the chlorophyll a inversion images of Guanting Reservoir. The two comparison models do not add dense residual blocks for modeling and prediction.
[0116] First, comparing the three evaluation metrics of the VPTR model and the GAN-VPTR model, the SSIM, PSNR, and LPIPS values of the VPTR model are 0.71732, 23.9544, and 44.6995 respectively, and the SSIM, PSNR, and COSIN values of the GAN-VPTR model are 0.74115, 25.7152, and 43.6842 respectively. Compared with the VPTR model, the SSIM metric and the PSNR metric are increased by 0.02383 and 1.7608 respectively, while the LPIPS is decreased by 1.0153, and the average model optimization percentage is 4.31%.
[0117] Second, comparing the three evaluation metrics of the GAN-VPTR model and the RRDB-GAN-VPTR model, the SSIM, PSNR, and COSIN values of the RRDB-GAN-VPTR model are 0.76113, 26.9355, and 42.2153 respectively. Compared with the VPTR model, the SSIM metric and the PSNR metric are increased by 0.01998 and 1.2203 respectively, while the LPIPS is decreased by 1.4689, and the average model optimization percentage is 3.60%.
[0118] As can be seen from Table 6, the method of the present invention can relatively accurately predict the chlorophyll a inversion image of Guanting Reservoir. By comparing the prediction results obtained by the RRDB-GAN-VPTR model with the prediction results without using the dense residual block, it can be known that the use of the dense residual block proposed by the present invention for feature extraction of the image can make the most of the feature information of the image, avoid the loss of feature information, and improve the prediction accuracy.
[0119] Except for the technical features described in the specification, they are all well-known technologies to those skilled in the art. The present invention omits the description of well-known components and well-known technologies to avoid redundancy and unnecessarily limit the present invention. The described embodiments in the above examples do not represent all embodiments consistent with the present application. Based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A remote sensing water quality prediction method for an improved VPTR model based on residual dense blocks, characterized in that The method first obtains the original remote sensing image time series of the target area, and then performs the following steps: Step 1: construct a chlorophyll inversion module for the target waters, which includes: data preprocessing, water body extraction, construction of a chlorophyll inversion model, and batch processing to obtain chlorophyll remote sensing image time series data; Step 2, construct an encoder module based on residual dense blocks, perform multi-scale convolution calculations on the target water area remote sensing image and combine residual connections to extract deep target water area remote sensing image features; The residual dense block contains four convolutional layers and SMU activation functions as well as a dense feature convolution layer with a 1*1 convolution kernel size. It captures the internal features of the module through dense connections, obtains the deep features of the image through the downsampling layer, and uses residual connections to output the image feature sequence generated by the encoder. Step 3: Construct a feature encoder module and a decoder module for an improved contrastive loss function based on the VPTR model to predict the image feature sequence F t for prediction; The target water area image feature sequence generated by the encoder part in step 2 is input into the transformer encoder module of the improved VPTR. The encoder combines the local spatial multi-head self-attention module and the temporal multi-head attention module to perform spatiotemporal feature learning encoding. The decoder adds a temporal multi-head attention layer and an output conversion layer compared to the encoder to obtain the image feature sequence at future moments. Step 4, construct a decoder module based on the residual dense block encoder; Using the upsampling layer and SMU activation function, the image feature sequence at the future moment obtained by the improved VPTR model is sent to the decoder module based on the residual dense block encoder to output the remote sensing image of the target water area at the future moment. Step 5: Build a discriminator module based on patchGAN; Through convolutional layers, batch normalization layers and ReLU activation functions, fine-grained analysis of remote sensing images of target waters at future moments is provided.
2. The method according to claim 1, wherein The step 1 comprises the following steps: (1.1) For all original remote sensing images, the radiometric calibration, FLASSH atmospheric correction module and orthorectification process tools in ENVI were used for data preprocessing. The atmospheric correction module selected the GF-1 data, and the spectral response function was quoted from the China Resources Satellite Application Center. In the orthorectification process, the global DEM data with a resolution of 30 seconds (about 900m) provided by ENVI was used, and the cubic convolution interpolation method was adopted. (1.2) The water body is extracted using the normalized difference water index NDWI, where the NDWI calculation formula is: Where B2 is the reflectivity of the green band of the GF-1 data, and B4 is the reflectivity of the near-infrared band of the GF-1 data. (1.3) Construct a chlorophyll inversion model to associate the selected sensitive band pixel values with the measured chlorophyll data to achieve quantitative inversion of chlorophyll. The specific construction formula of the model is: Y=553.9e-2.908x Where Y is the chlorophyll inversion concentration value, and x is the pixel value of the B1 / B2 band.
3. The method according to claim 1, characterized in that, The step 2 comprises the following steps: Use an encoder to process the input image I t Perform preliminary feature extraction through a composite operation of a convolutional layer and the SMU activation function as follows: Among them, H COM represents the composite function operation of convolution and the SMU activation function. The shallow features extracted by the encoder are input into each layer of the dense residual network, and the output of the c-th convolutional layer in the d-th RDB module is as follows: where τ represents the SMU activation function, ω r,c represents the weight of the c-th convolutional layer, refers to concatenating the feature map of the (d - 1)-th RDB module and the feature maps generated by convolutional layers 1, ..., (c - 1) in the d-th RDB module. The input of the residual dense module and the outputs of all convolutional layers in the current residual dense module are fused locally and then output The operation is as follows: Among them, represents the 1×1 Conv layer operation in the d-th residual dense module. The residual learning mechanism is used to further improve the information flow of the convolutional layer. Specifically, the input of the residual dense module in the d-th RDB module is output after being fused with the local features A residual network is adopted to form the final output of the fused features The operation is as follows: 。 4. The method according to claim 3, wherein In the step 2, performing dense connection and downsampling of the residual dense block includes: Each residual dense block is connected by a dense connection, and outputs after fully capturing the internal features of the module The operations are as follows: Among them, H DFF (·) represents the calculation of global feature fusion information in all residual dense modules, represents the residual coefficient. Take as the input of the downsampling layer, and further extract the deep features of the image by convolution operation The operation is as follows: Among them, H conv (·) is the convolution calculation of the downsampling layer. Finally, the input to the entire residual dense module and the output of the downsampling layer are connected by a residual connection to obtain the image feature F generated by the encoder t , which is formulated as follows: 。 5. The method according to claim 4, characterized in that, In the step 3, the contrast loss of supcon used is as follows: Where A represents all samples corresponding to the eigenvector v, and P represents the positive sample v corresponding to the eigenvector v + , and v + represents the spatial true eigenvector corresponding to all eigenvectors v within a certain time period, and v - represents all samples except the positive sample v + . S(·,·) represents the cosine similarity function. Therefore, the improved VPTR contrast loss function is calculated as follows: Among them is the future moment image feature sequence predicted by the VPTR model, F X , X ∈ [t + 1, t + 10] is the real future moment image feature sequence generated by the residual dense block encoder, v s + represents F X the feature vector at the spatial position s, represents F X the set of feature vectors at all other spatial positions, represents the feature vector at the spatial position s, represents the set of feature vectors at all other spatial positions, S l represents that the total number of all spatial positions in the feature map is H × W, and sg is the stop gradient operation.
6. The method according to claim 5, characterized in that, In the step 4, the decoder module based on the residual density block includes: Input the future moment image feature sequence predicted by the VPTR model in step three into the decoder for image reconstruction operations; during the decoding process, first perform spatial magnification on the feature sequence through two consecutive transposed convolution operations, and the specific formula is expressed as: Among them, H Norm (·) is the layer normalization function, W1 and W2 are the weights of the upsampling layer, is the transposed convolution operation. Finally, apply a convolutional layer and a Tanh activation function to complete the output of the image, and the operation is as follows: where δ is the Tanh activation function, and H co (·) is the convolution calculation.
7. The method according to claim 1, wherein In the said step five, the discriminator module based on PatchGAN includes: Discriminate between the predicted image obtained using the first four steps and the real image. The module consists of a convolutional layer, a normalization layer, and a LeakyReLU activation function layer. The output of the discriminator network is an N×N matrix, which is used to represent the authenticity score of each small block in the image, so as to judge the authenticity of the overall image. Specifically, the receptive field operation of the discriminator network is as follows: where I is the size of the input image, and k i (i = 1, 2, 3, 4, 5) is the size of the convolutional kernel, Padding i is the padding amount, Stride i is the stride size, and n i is the size of the output matrix.