Remote Sensing Monitoring Method and System for Rivers and Lakes Based on Dual-Branch Interaction and Semantic Alignment
Feature extraction and semantic alignment across the cross-resolution dual-branch interactive network are solved, and the cross-resolution semantic alignment problem of high and low-resolution remote sensing images is realized, efficient and low-label-dependent river and lake remote sensing monitoring is achieved, detection accuracy and computing efficiency are improved, and it is suitable for real-time monitoring of edge devices.
Patent Information
- Application Number
- CN202510553645.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing technology is difficult to effectively solve the problem of cross-resolution semantic alignment of high and low resolution remote sensing images, resulting in limited automation and large-scale applications of river and lake monitoring, and the cost of manual labeling is high and the computational complexity is high, making it difficult to meet the real-time monitoring needs.
A river and lake remote sensing monitoring method based on dual-branch interaction and semantic alignment is adopted, and feature extraction and semantic alignment are performed through cross-resolution dual-branch interaction networks, combining multi-scale recursive super-resolution module, hierarchical multi-scale Transformer module and self-supervised pseudo-label generation strategy to reduce the need for manual labeling and improve detection accuracy and computing efficiency.
Significantly improve detection accuracy, reduce manual labeling requirements by 70%, and support cross-scene detection with resolution differences of 4 to 8 times. It is suitable for edge devices and meets real-time monitoring needs.
Smart Images

Figure CN120071165B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of remote sensing image processing and artificial intelligence, and particularly to a remote sensing monitoring method and system for rivers and lakes based on double-branch interaction and semantic alignment. Background Art
[0002] Monitoring of rivers and lakes is the core task of river and lake ecological governance and long-term supervision, which is directly related to the stability of the water ecosystem. Traditional monitoring methods highly rely on manual interpretation of high-resolution remote sensing images, and professional personnel are required to visually identify and label targets such as buildings and sand mining sites in the images. However, the efficiency of manual interpretation is low, and the processing time for a single image can reach several hours. Moreover, affected by subjective judgment, it is difficult to ensure the annotation consistency, which severely restricts the feasibility of large-scale monitoring.
[0003] In recent years, change detection methods based on deep learning have made remarkable progress in single-resolution scenarios. For example, models such as U-Net and Siamese networks achieve automatic feature extraction and classification through end-to-end training. However, these methods are difficult to cope with the resolution difference problem of multi-source remote sensing images. For example, when high-resolution satellite images (such as the Gaofen-7 with a resolution of 0.5 meters) and low-resolution emergency images (such as the Jilin-1 with a resolution of 2 meters) are used in combination, directly upsampling the low-resolution images will result in texture blurring and detail loss. Taking a certain river and lake area as an example, after the low-resolution images are upsampled by bicubic interpolation, the contour blurring or positioning error will lead to an increase in the false detection and missed detection rates.
[0004] In addition, existing supervised learning methods require fully annotated data of two periods to train the change detection model, while the annotation cost of emergency images is high and the timeliness is poor in actual applications. Although existing semi-supervised methods (such as pseudo-label generation based on generative adversarial networks) can reduce the annotation requirements, they have insufficient adaptability to cross-resolution features, and the generalization ability of the model decreases by about 30% when training with high- and low-resolution images mixed.
[0005] The problem of insufficient cross-resolution semantic alignment is also prominent. After traditional methods improve the resolution of low-resolution images through super-resolution reconstruction, a single-branch network is used for detection. However, due to the simple feature interaction mechanism, it is difficult to fuse multi-scale context information. For example, there are sawtooth effects at the boundaries of small targets (such as ponds) reconstructed in super-resolution images, resulting in a decrease in segmentation accuracy. Although the Transformer-based model enhances semantic alignment through the global self-attention mechanism, its number of parameters is as high as hundreds of millions, making it difficult to be deployed on edge computing devices and unable to meet the real-time monitoring requirements. And due to the lack of a collaborative interaction mechanism in traditional double-branch networks, the feature redundancy between the high- and low-resolution branches increases, the computational complexity rises, and the actual application effect is limited.
[0006] The above problems jointly restrict the automated and large-scale application of river and lake monitoring. There is an urgent need for a change detection algorithm that takes into account cross-resolution semantic alignment, low annotation dependence, and efficient computation to improve the detection accuracy and practicality in complex scenarios. Summary of the Invention
[0007] To address the deficiencies of the prior art, the present invention provides a method and system for remote sensing monitoring of rivers and lakes based on dual-branch interaction and semantic alignment;
[0008] On the one hand, a method for remote sensing monitoring of rivers and lakes based on dual-branch interaction and semantic alignment is provided, including:
[0009] Obtain high-resolution images and low-resolution images of the area to be monitored; the high-resolution images are collected at a first time point, and the low-resolution images are collected at a second time point; the first time point is earlier than the second time point;
[0010] Input the high-resolution images and low-resolution images of the area to be monitored into a trained cross-resolution dual-branch interaction network to obtain the change results of the two images; among them, the trained cross-resolution dual-branch interaction network is used to extract features from the high-resolution images to obtain high-resolution features; regard the high-resolution features as the first features; extract features from the low-resolution images to obtain low-resolution features; further extract features from the low-resolution features to achieve semantic alignment with the first features to obtain second features; perform cross-resolution semantic fusion on the first features and the second features to obtain fusion features; perform change prediction on the fusion features to obtain the change prediction results of the two images.
[0011] On the other hand, a system for remote sensing monitoring of rivers and lakes based on dual-branch interaction and semantic alignment is provided, including:
[0012] An acquisition module configured to: obtain high-resolution images and low-resolution images of the area to be monitored; the high-resolution images are collected at a first time point, and the low-resolution images are collected at a second time point; the first time point is earlier than the second time point;
[0013] A prediction module, configured to: obtain a high-resolution image and a low-resolution image of an area to be monitored, input them into a trained cross-resolution double-branch interaction network, and obtain the change result of the two images; wherein, the trained cross-resolution double-branch interaction network is used to extract features from the high-resolution image to obtain high-resolution features; regard the high-resolution features as the first features; extract features from the low-resolution image to obtain low-resolution features; further extract features from the low-resolution features to achieve semantic alignment with the first features to obtain second features; perform cross-resolution semantic fusion on the first features and the second features to obtain fusion features; perform change prediction on the fusion features to obtain the change prediction result of the two images.
[0014] The above technical solution has the following advantages or beneficial effects: The detection accuracy is significantly improved; the manual annotation requirement is reduced by 70% through the self-supervised pseudo-label generation strategy, greatly reducing the annotation cost; the algorithm supports cross-scene detection with a resolution difference of 4 to 8 times, and has strong cross-resolution robustness; the calculation efficiency is high; it is also applicable to edge device deployment, meeting the practical requirements of real-time monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not constitute an improper limitation of the invention.
[0016] Figure 1 It is a flowchart of the method for Embodiment 1. DETAILED DESCRIPTION
[0017] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs.
[0018] Embodiment 1
[0019] This embodiment provides a remote sensing monitoring method for rivers and lakes based on double-branch interaction and semantic alignment;
[0020] As Figure 1 shown, the remote sensing monitoring method for rivers and lakes based on double-branch interaction and semantic alignment includes:
[0021] S101: Obtain a high-resolution image and a low-resolution image of the area to be monitored; the high-resolution image is collected at a first time point, and the low-resolution image is collected at a second time point; the first time point is earlier than the second time point;
[0022] S102: Input the high-resolution image and the low-resolution image of the area to be monitored into the trained cross-resolution dual-branch interaction network to obtain the change results of the two images;
[0023] Among them, the trained cross-resolution dual-branch interaction network is used for:
[0024] Extract features from the high-resolution image to obtain high-resolution features; regard the high-resolution features as the first features;
[0025] Extract features from the low-resolution image to obtain low-resolution features; further extract features from the low-resolution features to achieve semantic alignment with the first features and obtain the second features;
[0026] Perform cross-resolution semantic fusion on the first features and the second features to obtain fused features;
[0027] Perform change prediction on the fused features to obtain the change prediction results of the two images.
[0028] Furthermore, obtain the high-resolution image and the low-resolution image of the area to be monitored;
[0029] Among them, the high-resolution image means that the ground resolution of the remote sensing image is gradually improved from 10m, 5m, 2m, 1m, or even 0.6m.
[0030] Among them, the low-resolution image means that the ground resolution of the remote sensing image is greater than 10m.
[0031] Furthermore, in S102: Input the high-resolution image and the low-resolution image of the area to be monitored into the trained cross-resolution dual-branch interaction network to obtain the change results of the two images. Among them, the trained cross-resolution dual-branch interaction network includes: a first branch and a second branch;
[0032] The first branch is provided with a first multi-scale recursive super-resolution module. The input value of the first multi-scale recursive super-resolution module is the high-resolution image, and the output value of the first multi-scale recursive super-resolution module is the first feature;
[0033] The second branch is provided with a second multi-scale recursive super-resolution module and a hierarchical multi-scale Transformer module connected in sequence. The input value of the second multi-scale recursive super-resolution module is the low-resolution image, the output value of the second multi-scale recursive super-resolution module is the low-resolution feature, the input value of the hierarchical multi-scale Transformer module is the low-resolution feature, and the output value of the hierarchical multi-scale Transformer module is the second feature;
[0034] The first feature and the second feature are both input into the cross-resolution semantic fusion module. The cross-resolution semantic fusion module outputs a fused feature, and the fused feature is input into a change predictor, which outputs the change result of the two images.
[0035] Further, the extracting features from the high-resolution image to obtain high-resolution features includes:
[0036] Using a first multi-scale recursive super-resolution module to extract features from the high-resolution image to obtain high-resolution features.
[0037] Further, the first multi-scale recursive super-resolution module includes: a ResNet-50 network.
[0038] Further, the first multi-scale recursive super-resolution module includes:
[0039]
[0040] Among them, represents the high-resolution feature; represents the high-resolution image; represents the ResNet50 network.
[0041] Further, the extracting features from the low-resolution image to obtain low-resolution features includes:
[0042] Using a second multi-scale recursive super-resolution module to extract features from the low-resolution image to obtain low-resolution features.
[0043] Further, the second multi-scale recursive super-resolution module includes: a first convolutional layer, a first super-resolution module, a second super-resolution module, a third super-resolution module, and a sub-pixel convolutional module connected in sequence.
[0044] Further, the second multi-scale recursive super-resolution module includes:
[0045]
[0046] Among them, represents a 5×5 convolution, represents a sub-pixel convolution, represents the first super-resolution module, represents the low-resolution feature, represents the low-resolution image, represents the second super-resolution module, represents the third super-resolution module.
[0047] It should be understood that the low-resolution image is input into the second multi-scale recursive super-resolution module, and the high-frequency details are gradually restored through the recursive residual structure and the dynamic upsampling technique. The second multi-scale recursive super-resolution module uses a 5×5 convolution to extract initial features, and then enhances the texture reconstruction ability through a three-level super-resolution module (each group includes a 3×3 depthwise separable convolution, a dynamic upsampling layer, and a channel attention mechanism). Finally, a sub-pixel convolution is used to generate a super-resolution image. In the dual-branch feature extraction stage, a ResNet-50 backbone network with shared weights is used to process the original high-resolution image and the reconstructed super-resolution image respectively to extract multi-scale features.
[0048] Further, the internal structures of the first super-resolution module, the second super-resolution module, and the third super-resolution module are the same. The first super-resolution module includes:
[0049] A separable convolution layer, an upsampling layer, a channel attention mechanism layer, and a first adder connected in sequence;
[0050] The input end of the separable convolution layer is the input end of the first super-resolution module;
[0051] The output end of the first adder is the output end of the first super-resolution module;
[0052] The input end of the separable convolution layer is also connected to the input end of the first adder.
[0053] It should be understood that separable convolution includes spatially separable convolutions and depthwise separable convolution.
[0054] Further, the internal working processes of the first super-resolution module, the second super-resolution module, and the third super-resolution module are the same. The first super-resolution module includes:
[0055]
[0056] Among them, represents a 3×3 convolution, represents upsampling, represents channel attention, represents the input value of the first super-resolution module.
[0057] Further, the further feature extraction of the low-resolution features to achieve semantic alignment with the first feature and obtain the second feature includes:
[0058] The hierarchical multi-scale Transformer module is adopted to further extract features from the low-resolution features to achieve semantic alignment with the first feature and obtain the second feature.
[0059] Furthermore, the hierarchical multi-scale Transformer module includes:
[0060] The third branch, the fourth branch, and the fifth branch in parallel; the input ends of the third branch, the fourth branch, and the fifth branch are all used to input low-resolution features;
[0061] The structures of the third branch, the fourth branch, and the fifth branch are the same;
[0062] The third branch includes: a feature splitting unit, a self-attention mechanism layer, a global sparse sampling unit, and a multiplier connected in sequence; the input end of the multiplier is also connected to the output end of the self-attention mechanism layer; the input end of the feature splitting unit is the input end of the third branch, and the output end of the multiplier is the output end of the third branch;
[0063] The output ends of the third branch, the fourth branch, and the fifth branch are all connected to the cross-scale skip connection layer, and the output end of the cross-scale skip connection layer is the output end of the hierarchical multi-scale Transformer module.
[0064] Furthermore, the formula expression of the hierarchical multi-scale Transformer module includes:
[0065]
[0066] Where represents the feature splitting unit, represents the self-attention mechanism layer, represents the global sparse sampling unit, represents the cross-scale skip connection layer.
[0067] Furthermore, the feature splitting unit splits the feature into a series of several feature patches (patches) by means of a sliding window.
[0068] It should be understood that the self-attention mechanism layer means that after the feature is linearly mapped respectively to obtain Q, K, and V, and then the attention matrix is calculated through the following formula:
[0069]
[0070] Where is the dimension of the input feature.
[0071] Further, the global sparse sampling unit refers to sampling 10% of the eigenvalue from the output value of the self-attention mechanism layer to form a new feature matrix, and using the new feature matrix as the global sparse attention matrix.
[0072] Further, the cross-scale skip connection layer refers to concatenating the output values of the third branch, the fourth branch, and the fifth branch along the channel dimension.
[0073] It should be understood that since the high-resolution branch retains the fine spatial information of the original image, only the low-resolution branch is sent into the hierarchical multi-scale Transformer module to enhance semantic consistency. The hierarchical multi-scale Transformer module divides the feature map into 8×8 windows to calculate the local window self-attention to reduce the computational complexity. At the same time, it randomly samples 10% of the tokens across windows to construct the global sparse attention matrix, and fuses the shallow details and deep semantic features through the cross-scale skip connection. The formula is as follows:
[0074]
[0075] Among them, represents concatenating the features along the channel dimension, represents the output value of the third branch, represents the output value of the fourth branch, represents the output value of the fifth branch.
[0076] Further, the cross-resolution semantic fusion of the first feature and the second feature to obtain the fusion feature includes:
[0077] Using the cross-resolution semantic fusion module to perform cross-resolution semantic fusion on the first feature and the second feature to obtain the fusion feature.
[0078] Further, the cross-resolution semantic fusion module includes:
[0079] The parallel sixth branch and seventh branch; the input end of the sixth branch is used to input the first feature; the input end of the seventh branch is used to input the second feature;
[0080] The sixth branch includes: a global average pooling layer, a linear mapping layer, an activation function layer, a second adder, and a spatial attention layer connected in sequence; the input end of the second adder is also connected to the input end of the global average pooling layer; the input end of the global average pooling layer is the input end of the sixth branch;
[0081] The seventh branch includes: a third adder; the input end of the third adder is connected to the output end of the spatial attention layer, and the output end of the third adder is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is the output end of the cross-resolution semantic fusion module; the input end of the third adder is the input end of the seventh branch.
[0082] Further, the cross-resolution semantic fusion module includes:
[0083]
[0084] Among them, represents the global average pooling layer, represents the linear mapping layer, represents the activation function layer, represents the spatial attention layer, represents the convolutional layer, represents the fused feature.
[0085] It should be understood that in the feature fusion stage, the cross-resolution semantic fusion module dynamically integrates the dual-branch features through channel recalibration and spatial attention. Channel recalibration performs global average pooling on the high-resolution feature map, generates channel weights through a fully connected layer and Sigmoid activation, and screens key channels (such as features related to sand and gravel pits). Spatial attention uses hybrid dilated convolution (dilation rates = 2, 4, 6) to construct multi-scale receptive fields and focuses on regions sensitive to changes (such as newly added illegal construction areas). Finally, residual connection adds the calibrated dual-branch feature maps element by element and integrates them through 1×1 convolution to retain detailed information.
[0086] Further, the step of performing change prediction on the fused feature to obtain the change prediction results of the two images includes:
[0087] Using a change predictor to perform change prediction on the fused feature to obtain the change prediction results of the two images.
[0088] Further, the change predictor includes:
[0089]
[0090] Among them, is the input feature of the change predictor, is a 3×3 convolutional layer, is the batch normalization layer, is the ReLU activation enhancement function;
[0091] Subsequently, the feature map is upsampled by bilinear interpolation to restore the spatial resolution, and 3×3 convolution is used to further refine the edge features, and finally through 1×1 convolution Compress the number of channels to 1, and generate a change probability map through the Sigmoid activation function , and the formula is as follows:
[0092] .
[0093] Furthermore, the trained cross-resolution dual-branch interaction network also includes, during the training process: adding a semantic predictor to the cross-resolution dual-branch interaction network; the training steps of the cross-resolution dual-branch interaction network include:
[0094] Construct a training set, which includes high-resolution images and low-resolution images. The time points for collecting the high-resolution images and low-resolution images are different, and the time point for collecting the high-resolution images is earlier than that for collecting the low-resolution images; the high-resolution images have no labels, and the low-resolution images are provided with ground truth labels, and the ground truth labels are used to annotate different regions in the images; different regions include, but are not limited to: grassland, quarry, pond;
[0095] Input the training set into the cross-resolution dual-branch interaction network. The high-resolution images are processed through the first branch to obtain first features; the low-resolution images are processed through the second branch to obtain second features;
[0096] Input the second features and the ground truth labels of the low-resolution images into the semantic predictor, train the semantic predictor, and obtain the trained semantic predictor;
[0097] The trained semantic predictor receives the first features and outputs pseudo-labels of the high-resolution images according to the first features; calculate the difference between the pseudo-labels of the high-resolution images and the ground truth labels of the low-resolution images to obtain a first difference, perform morphological opening on the first difference, and perform dynamic confidence filtering on the result of the opening operation to obtain self-supervised change detection labels;
[0098] Calculate the total loss function according to the self-supervised change detection labels and the predicted values of the change predictor. When the value of the total loss function no longer decreases, stop training to obtain the trained cross-resolution dual-branch interaction network.
[0099] Specifically, the semantic predictor only exists in the cross-resolution dual-branch interaction network during the training process. After the training is completed, during the actual use stage of the network, the semantic predictor is not required.
[0100] Furthermore, the semantic predictor includes:
[0101]
[0102] Among them, is the input feature of the semantic predictor; is an intermediate variable;
[0103] Subsequently, the feature map is upsampled to restore the spatial resolution through bilinear interpolation, the edge features are further refined using 3×3 convolution, and finally the number of channels is compressed to the corresponding number of classes through 1×1 convolution, and a semantic probability map is generated through the Sigmoid activation function , and the formula is as follows:
[0104] .
[0105] It should be understood that to reduce the dependence on manual annotation, a self-supervised pseudo-label generation strategy is adopted. During the training phase, the features of the second epoch and the ground truth labels are fed into the semantic predictor to calculate the semantic loss, and the semantic predictor is trained. Then, the features of the first epoch are fed into the semantic predictor to obtain a prediction map, which is used as the pseudo-label of the first epoch. The difference between the pseudo-label of the first epoch and the ground truth label of the second epoch is calculated, combined with morphological opening operation (kernel size 3×3) to eliminate isolated noise points, generating weak labels for low-score images, and through dynamic confidence filtering, only pixels with a prediction probability higher than 0.85 are retained as pseudo-labels to avoid false supervision in low-confidence regions. The above process can be formulated as:
[0106]
[0107]
[0108] where, represents the semantic predictor, is the pseudo-label of the first epoch, is the ground truth label of the second epoch, represents the opening operation, represents the dynamic confidence filtering.
[0109] Furthermore, the morphological opening operation on the first difference includes: image erosion and dilation. After the image is eroded, noise is removed, but the image is also compressed; then the eroded image is dilated, which can remove noise and retain the original image.
[0110] Furthermore, the dynamic confidence filtering on the result of the opening operation includes: setting the initial threshold to 0.6; filtering out regions with a confidence level lower than the set threshold and retaining regions with a confidence level higher than the set threshold.
[0111] Furthermore, the total loss function includes:
[0112]
[0113]
[0114]
[0115] Among them, represents the total number of pixels, represents the total number of categories, represents the cumulative category loss weight, and respectively represent the true value and predicted value of the th sample with respect to category ; represents the predicted change region, represents the true change region.
[0116] It should be understood that during the training process, a mixed loss function is introduced, combining weighted cross-entropy loss and Dice loss. Among them, the weighted cross-entropy loss is used to calculate the semantic loss, and the Dice loss is used to calculate the change loss. The model parameters are optimized by the AdamW optimizer and the cosine annealing learning rate scheduler, and the initial learning rate is set to , the batch size is set to 8, and the training period is 200.
[0117] Furthermore, the construction of the training set includes:
[0118] Crop the images around rivers and lakes from remote sensing images of two periods to form remote sensing images of rivers and lakes at the same location in different periods, and visually interpret one of the remote sensing images to form label data. Assume that the first period does not contain label data and the second period contains label data. Scale the cropped remote sensing images of rivers and lakes and their labels to obtain 256×256 tiles; randomly select some tiles and flip them randomly up, down, left, and right; calculate the mean and standard deviation of the RGB three channels, and standardize the tiles. Finally, the size of the obtained training tiles is 256×256×3 (length × width × number of channels).
[0119] In the change map generation stage, the images of the two periods are fed into the model. The semantic segmenter outputs five types of semantic segmentation maps (background, house, shed, pond, sand mining site), and the two-period difference map obtains the preliminary change region by pixel-by-pixel subtraction. For post-processing, the Otsu algorithm is used for adaptive threshold segmentation (default 0.6), and closing operation (kernel size 5×5) is used to fill holes, and connected component analysis is used to filter out noise with an area less than 50 pixels.
[0120] The present invention proposes a cross - resolution dual - branch interaction network. Its core innovations include a multi - scale recursive super - resolution module, a hierarchical multi - scale Transformer, a self - supervised pseudo - label generation strategy, and a cross - resolution semantic fusion module. Among them, the multi - scale recursive super - resolution module reconstructs the high - frequency details of low - resolution images through a recursive residual structure and dynamic upsampling technology, reducing information loss; the hierarchical multi - scale Transformer fuses local window attention and global sparse sampling mechanisms to enhance cross - resolution semantic alignment ability; the self - supervised pseudo - label generation strategy uses the prediction results of the high - resolution branch to generate pseudo - labels for low - resolution images, significantly reducing the need for manual annotation; the cross - resolution semantic fusion module dynamically fuses the features of the two branches through channel recalibration and spatial attention to improve the discriminability of changed regions. This algorithm supports cross - scene detection with a resolution difference of 4 to 8 times under the condition of only requiring single - epoch labeled data.
[0121] Embodiment 2
[0122] This embodiment provides a remote sensing monitoring system for rivers and lakes based on dual - branch interaction and semantic alignment, including:
[0123] An acquisition module, which is configured to: acquire a high - resolution image and a low - resolution image of the area to be monitored; the high - resolution image is acquired at a first time point, and the low - resolution image is acquired at a second time point; the first time point is earlier than the second time point;
[0124] A prediction module, which is configured to: acquire the high - resolution image and the low - resolution image of the area to be monitored, input them into the trained cross - resolution dual - branch interaction network, and obtain the change results of the two images; among them, the trained cross - resolution dual - branch interaction network is used to extract features from the high - resolution image to obtain high - resolution features; regard the high - resolution features as the first features; extract features from the low - resolution image to obtain low - resolution features; further extract features from the low - resolution features to achieve semantic alignment with the first features to obtain second features; perform cross - resolution semantic fusion on the first features and the second features to obtain fused features; perform change prediction on the fused features to obtain the change prediction results of the two images.
[0125] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A remote sensing monitoring method for rivers and lakes based on dual-branch interaction and semantic alignment, characterized in that Including: Obtain the high-resolution image and the low-resolution image of the area to be monitored; The high-resolution image is collected at the first time point, and the low-resolution image is collected at the second time point; The first time point is earlier than the second time point; Input the high-resolution image and the low-resolution image of the area to be monitored into the trained cross-resolution double-branch interaction network to obtain the change result of the two images; Among them, the trained cross-resolution double-branch interaction network is used to extract features from the high-resolution image to obtain high-resolution features. The cross-resolution double-branch interaction network includes: a first branch and a second branch; The first branch is provided with a first multi-scale recursive super-resolution module. The input value of the first multi-scale recursive super-resolution module is the high-resolution image, and the output value of the first multi-scale recursive super-resolution module is the first feature; The second branch is provided with a second multi-scale recursive super-resolution module and a hierarchical multi-scale Transformer module connected in sequence. The input value of the second multi-scale recursive super-resolution module is the low-resolution image, the output value of the second multi-scale recursive super-resolution module is the low-resolution feature, the input value of the hierarchical multi-scale Transformer module is the low-resolution feature, and the output value of the hierarchical multi-scale Transformer module is the second feature; Both the first feature and the second feature are input into the cross-resolution semantic fusion module. The cross-resolution semantic fusion module outputs the fusion feature, and the fusion feature is input into the change predictor. The change predictor outputs the change result of the two images; regard the high-resolution feature as the first feature; extract features from the low-resolution image to obtain the low-resolution feature; further extract features from the low-resolution feature to achieve semantic alignment with the first feature to obtain the second feature; perform cross-resolution semantic fusion on the first feature and the second feature to obtain the fusion feature; perform change prediction on the fusion feature to obtain the change prediction result of the two images.
2. The remote sensing monitoring method for rivers and lakes based on dual-branch interaction and semantic alignment according to claim 1, wherein The extracting features from the low-resolution image to obtain the low-resolution feature includes: Adopt the second multi-scale recursive super-resolution module to extract features from the low-resolution image to obtain the low-resolution feature; The second multi-scale recursive super-resolution module includes: a first convolutional layer, a first super-resolution module, a second super-resolution module, a third super-resolution module, and a sub-pixel convolutional module connected in sequence; The second multi-scale recursive super-resolution module includes: ; Among them, represents a 5×5 convolution, represents sub-pixel convolution, represents the first super-resolution module, represents low-resolution features, represents a low-resolution image, represents the second super-resolution module, represents the third super-resolution module.
3. The remote sensing monitoring method for rivers and lakes based on dual-branch interaction and semantic alignment according to claim 2, characterized in that, The internal structures of the first super-resolution module, the second super-resolution module, and the third super-resolution module are the same. The first super-resolution module includes: A depthwise separable convolutional layer, an upsampling layer, a channel attention mechanism layer, and a first adder connected in sequence; The input end of the depthwise separable convolutional layer is the input end of the first super-resolution module; The output end of the first adder is the output end of the first super-resolution module; The input end of the depthwise separable convolutional layer is also connected to the input end of the first adder; The internal working processes of the first super-resolution module, the second super-resolution module, and the third super-resolution module are the same. The first super-resolution module includes: ; Among them, represents a 3×3 convolution, represents upsampling, represents channel attention, represents the input value of the first super-resolution module.
4. The remote sensing monitoring method for rivers and lakes based on dual-branch interaction and semantic alignment according to claim 1, characterized in that, Performing further feature extraction on the low-resolution features to achieve semantic alignment with the first feature and obtain the second feature, including: Using a hierarchical multi-scale Transformer module to perform further feature extraction on the low-resolution features to achieve semantic alignment with the first feature and obtain the second feature; The hierarchical multi-scale Transformer module includes: The third branch, the fourth branch, and the fifth branch in parallel; the input ends of the third branch, the fourth branch, and the fifth branch are all used to input the low-resolution features; The structures of the third branch, the fourth branch, and the fifth branch are the same; The third branch includes: a feature splitting unit, a self-attention mechanism layer, a global sparse sampling unit, and a multiplier connected in sequence; the input end of the multiplier is also connected to the output end of the self-attention mechanism layer; the input end of the feature splitting unit is the input end of the third branch, and the output end of the multiplier is the output end of the third branch; The output ends of the third branch, the fourth branch, and the fifth branch are all connected to the cross-scale skip connection layer, and the output end of the cross-scale skip connection layer is the output end of the hierarchical multi-scale Transformer module; The formula expression of the hierarchical multi-scale Transformer module includes: ; Among them, represents a feature splitting unit, represents a self-attention mechanism layer, represents a global sparse sampling unit, represents a cross-scale skip connection layer.
5. The remote sensing monitoring method for rivers and lakes based on dual-branch interaction and semantic alignment according to claim 1, characterized in that, Performing cross-resolution semantic fusion on the first feature and the second feature to obtain the fused feature, including: Using a cross-resolution semantic fusion module to perform cross-resolution semantic fusion on the first feature and the second feature to obtain the fused feature; The cross-resolution semantic fusion module includes: The sixth branch and the seventh branch in parallel; the input end of the sixth branch is used to input the first feature; the input end of the seventh branch is used to input the second feature; The sixth branch includes: a global average pooling layer, a linear mapping layer, an activation function layer, a second adder, and a spatial attention layer connected in sequence; the input end of the second adder is also connected to the input end of the global average pooling layer; the input end of the global average pooling layer is the input end of the sixth branch; The seventh branch includes: a third adder; the input end of the third adder is connected to the output end of the spatial attention layer, and the output end of the third adder is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is the output end of the cross-resolution semantic fusion module; the input end of the third adder is the input end of the seventh branch; The cross-resolution semantic fusion module includes: ; Among them, represents the global average pooling layer, represents the linear mapping layer, represents the activation function layer, represents the spatial attention layer, represents the convolutional layer, represents the fused feature.
6. The remote sensing monitoring method for rivers and lakes based on dual-branch interaction and semantic alignment according to claim 1, characterized in that Performing change prediction on the fused feature to obtain the change prediction results of the two images, including: Using a change predictor to perform change prediction on the fused feature to obtain the change prediction results of the two images; The change predictor includes: ; Among them, is the input feature of the change predictor, is a 3×3 convolutional layer, is a batch normalization layer, is the ReLU activation enhancement function; Subsequently, the feature map is upsampled by bilinear interpolation to restore the spatial resolution, and 3×3 convolution is used to further refine the edge features. Finally, 1×1 convolution is used to compress the number of channels to 1, and a change probability map is generated through the Sigmoid activation function , and the formula is as follows: 。 7. The method for remote sensing monitoring of rivers and lakes based on dual-branch interaction and semantic alignment according to claim 1, characterized in that, A trained cross-resolution double-branch interaction network, and during the training process, it also includes: adding a semantic predictor to the cross-resolution double-branch interaction network; the training steps of the cross-resolution double-branch interaction network include: Construct a training set, where the training set includes high-resolution images and low-resolution images. The time points at which the high-resolution images and low-resolution images are collected are different, and the time point for collecting high-resolution images is earlier than the time point for collecting low-resolution images; the high-resolution images have no labels, and the low-resolution images are provided with ground truth labels, and the ground truth labels are used to annotate different regions in the images; Input the training set into the cross-resolution dual-branch interaction network. The high-resolution images are processed through the first branch to obtain first features; the low-resolution images are processed through the second branch to obtain second features; Input the second features and the ground truth labels of the low-resolution images into the semantic predictor, and train the semantic predictor to obtain the trained semantic predictor; The trained semantic predictor receives the first features and outputs pseudo-labels of the high-resolution images according to the first features; calculate the difference between the pseudo-labels of the high-resolution images and the ground truth labels of the low-resolution images to obtain a first difference, perform morphological opening on the first difference, and perform dynamic confidence filtering on the result of the opening operation to obtain self-supervised change detection labels; Calculate the total loss function according to the self-supervised change detection labels and the predicted values of the change predictor. When the value of the total loss function no longer decreases, stop training to obtain the trained cross-resolution dual-branch interaction network.
8. The method for remote sensing monitoring of rivers and lakes based on dual-branch interaction and semantic alignment according to claim 7, characterized in that, The semantic predictor includes: ; Among them, is the input feature of the semantic predictor; is an intermediate variable; Subsequently, bilinear interpolation is used to upsample the feature map to restore the spatial resolution, and 3×3 convolution is utilized to further refine the edge features. Finally, 1×1 convolution is employed to compress the number of channels to the corresponding number of categories, and a semantic probability map is generated through the Sigmoid activation function , and the formula is as follows: 。 9. The remote sensing monitoring system for rivers and lakes based on dual-branch interaction and semantic alignment is characterized in that, including: An acquisition module, which is configured to: acquire high-resolution images and low-resolution images of the area to be monitored; The high-resolution images are collected at a first time point, and the low-resolution images are collected at a second time point; the first time point is earlier than the second time point; A prediction module, which is configured to: acquire the high-resolution images and low-resolution images of the area to be monitored, input them into the trained cross-resolution dual-branch interaction network, and obtain the change results of the two images; among them, the trained cross-resolution dual-branch interaction network is used to extract features from the high-resolution images to obtain high-resolution features. The cross-resolution dual-branch interaction network includes: a first branch and a second branch; The first branch is provided with a first multi-scale recursive super-resolution module. The input value of the first multi-scale recursive super-resolution module is the high-resolution image, and the output value of the first multi-scale recursive super-resolution module is the first feature; The second branch is provided with a second multi-scale recursive super-resolution module and a hierarchical multi-scale Transformer module connected in sequence. The input value of the second multi-scale recursive super-resolution module is the low-resolution image, the output value of the second multi-scale recursive super-resolution module is the low-resolution feature, the input value of the hierarchical multi-scale Transformer module is the low-resolution feature, and the output value of the hierarchical multi-scale Transformer module is the second feature; The first feature and the second feature are both input into the cross-resolution semantic fusion module. The cross-resolution semantic fusion module outputs a fused feature, and the fused feature is input into a change predictor, which outputs the change result of the two images. The high-resolution feature is regarded as the first feature. Feature extraction is performed on the low-resolution image to obtain a low-resolution feature. Further feature extraction is performed on the low-resolution feature to achieve semantic alignment with the first feature, resulting in the second feature. The first feature and the second feature are subjected to cross-resolution semantic fusion to obtain a fused feature. The fused feature is used for change prediction to obtain the change prediction result of the two images.
Citation Information
Patent Citations
Transform galaxy image super-resolution method based on space and channel characteristics
CN118350994A
Remote-sensing image change detection method based on siamese network
WO2024217541A1