Rice drawing network based on space-time attention of double branches
The dual-branch network with TSA and CDC blocks addresses the issue of incomplete spatio-temporal feature fusion in deep learning models, enhancing rice field mapping accuracy by capturing detailed spatial and temporal features in SAR imagery.
Patent Information
- Application Number
- CN202510335557.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-15
AI Technical Summary
The existing rice mapping model has shortcomings in the fusion of space-time features, resulting in the loss of time information and weakening of rice features, ignoring the inherent consistency between agricultural plots and boundaries, affecting the accuracy of boundary extraction.
Using a spatiotemporal attention network based on dual branches, combined with central differential convolution and adaptive weighting mechanism, the spatiotemporal characteristics of remote sensing images are fully extracted through TSA Block and CDC Conv Block, and the model's ability to express boundary details and complex spatial features is enhanced.
The model's identification accuracy of rice planting areas is significantly improved, the accuracy of rice mapping and temporal correlation are improved, and the ability to express complex spatial characteristics is enhanced.
Smart Images

Figure CN120318558A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a rice mapping network based on dual-branch spatio-temporal attention. Background Art
[0002] The purpose of the present invention is to solve the technical problems in the existing rice mapping in the U-shaped network and spatio-temporal attention network, including insufficient extraction of rice edge information and combination of time and space attention, resulting in loss of time information. Thus, a rice mapping network based on central difference convolution and spatio-temporal attention is proposed.
[0003] In terms of models, with the rapid development of artificial intelligence, deep learning has become one of the most cutting-edge technologies in the field of crop remote sensing mapping. DL crop mapping models have shifted from the early single spatial information extraction to the joint extraction of spatio-temporal information. Compared with traditional methods, these models use structures such as LSTM, RNN, and TCN to capture the dependencies of time series, and can better fit the process of crop phenological growth, but there are certain limitations in dealing with complex spatio-temporal relationships. Therefore, recently, spatio-temporal attention mechanisms have received increasing attention in crop mapping research based on remote sensing time series data. Spatio-temporal attention mechanisms can capture the temporal and spatial dependencies in data, which allows the model to dynamically focus on input features at different time steps and spatial steps, facilitating the understanding of the changes in crop growth over time and the mutual influence between different regions. [Granot] Based on the spatio-temporal attention mechanism, pixel-level crop recognition in "Panoptic Segmentation of Satellite Image Time Series with Convolutional Temporal Attention Networks" was achieved using time series data, and the accuracy was significantly improved compared to traditional models. In addition, the combination of pixel set encoder and lightweight time attention (PSE+LTAE) can obtain rich spatial range and time information of crop fields for classifying remote sensing time series images. However, although temporal modules such as Transformer perform well in capturing the temporal information of single features, they tend to ignore the effective fusion of spatial information and time information when dealing with time series data. This limitation may lead to insufficient understanding of complex spatio-temporal relationships, thereby affecting the comprehensive analysis of the crop growth process.
[0004] The incomplete spatio-temporal feature fusion strategy adopted by the current deep learning crop distribution extraction model fails to fully utilize the time-series SAR images, thus weakening the rice features, reducing the spatio-temporal correlation between different polarization modes, and ignoring the internal consistency between agricultural plots and boundaries, which easily leads to the inconsistency between the extracted boundaries and the actual situation; therefore, it is necessary to propose a rice mapping network based on dual-branch spatio-temporal attention to solve the above problems. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a rice mapping network based on dual-branch spatio-temporal attention. By combining the static and dynamic attention mechanisms through the TSA Block, the spatio-temporal features in the remote sensing ground object images are fully extracted. Through the CDC Conv Block, the central difference convolution and ordinary convolution are combined with the adaptive weighting mechanism. The central difference convolution can better capture the detail changes of the image. Combined with the adaptive weighting, it can dynamically adjust the convolution information and effectively improve the model's expression ability for boundary details and complex spatial features.
[0006] To achieve the above technical effects, the technical solution adopted by the present invention is: A rice mapping network based on the spatio-temporal sequence of Sentinel-1 radar images, including a dual-branch network SimTA, a spatio-temporal sequence attention module TSABlock, and a spatial information extraction module CDCConv based on central difference convolution; the two input ends of the dual-branch network respectively receive the VV and VH polarization data of the Sentinel-1 radar images, the output of the dual-branch network is connected to the CDCConv module, after two layers of CDCConv processing, channel superposition is performed, and the superimposed features are input into the TSABlock module, and finally the rice mapping result is output.
[0007] Preferably, the dual-branch network includes an encoding-decoding structure, an edge information extraction module, and a spatio-temporal attention module; among them, both the encoding layer and the decoding layer include multiple layers of structures. The two branches of the first-layer encoder respectively process the VV and VH data, and use CDCConv for feature extraction, and its output is used as the input of the first decoding layer and is transmitted to the second-layer encoder at the same time; the output of the second-layer encoder is connected to the penultimate decoding layer, and feature fusion is performed with the output of the previous decoding layer through a skip connection; the third-layer encoder uses CDCConv, and its output is connected to the fourth-layer encoder and performs skip fusion with the output of the second decoding layer.
[0008] Preferably, the feature extraction operation of the encoding layer includes: The input features sequentially pass through the first Conv layer, the BatchNorm layer, and the second Conv layer to generate the output feature F(1,1). At the same time, the input features pass through the CDCConv layer and the second BatchNorm layer to generate the output feature F(1,2). F(1,1) and F(1,2) are added together with adaptive weights to form the encoded feature.
[0009] Preferably, the feature extraction operation of the encoding layer includes: the encoded feature of the first layer is input into the first Conv layer of this layer in the second encoding layer, the output is input into the BatchNorm layer of the encoding layer of this layer, and the BatchNorm output of the encoding layer is input to the second Conv layer to obtain the output result; The encoded feature of the second layer is input into the third encoding layer, input into the first Conv layer of this layer, the output is input into the BatchNorm layer of the third-layer encoding layer, and the BatchNorm output of the encoding layer is input to the second Conv layer to obtain the output result F(3,1); and the feature is input into the first CDCConv layer of the encoding layer, the output of the CDCConv layer is input into the second BatchNorm of the encoding layer to obtain the output result F(3,2), and F(3,1) and F(3,2) are added using adaptive weights, and the output obtains the encoded feature of the third layer; The encoded feature of the third layer is input into the fourth encoding layer, input into the first Conv layer of this layer, the output is input into the BatchNorm layer of the encoding layer of this layer, and the BatchNorm output of the encoding layer is input to the second Conv layer to obtain the output result.
[0010] Preferably, the output features of the VV and VH encoders of the dual-branch network are fused by channel stacking at each layer, specifically including concatenating the output features of the encoders from the first layer to the fourth layer in the channel dimension.
[0011] Preferably, the operation of the spatio-temporal sequence attention module TSABlock includes: Input the feature fusion result of the fourth layer into the dual-branch network of the first layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the first layer to obtain feature F(1,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(1,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(1,2) to obtain feature F(1,3), and the feature obtained by cross-multiplying F(1,1) and F(1,3) is dot-multiplied with the original feature to obtain the output result of the first layer; The spatio-temporal information extraction feature of the first layer is input into the two-branch network of the second layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain feature F(2,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(2,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(2,2) to obtain feature F(2,3), and the feature obtained by cross-multiplying F(2,1) and F(2,3) is dot-multiplied with the original feature to obtain the output result of the second layer; The spatio-temporal information extraction feature of the second layer is input into the two-branch network of the third layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain feature F(3,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(3,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(3,2) to obtain feature F(3,3), and the feature obtained by cross-multiplying F(3,1) and F(3,3) is dot-multiplied with the original feature to obtain the output result of the third layer; The spatio-temporal information extraction features of the third layer and the spatio-temporal information extraction features of the second layer are input into the double-branch network of the third layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain feature F(4,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(4,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(4,2) to obtain feature F(4,3), and the feature obtained by cross-multiplying F(4,1) and F(4,3) is dot-multiplied with the original feature to obtain the output result of the fourth layer.
[0012] Preferably, the residual structure sequentially includes a BatchNorm layer, a Conv layer, a DWConv layer, and a second Conv layer. Its output is added to the original features and then input into a third Conv layer. After being processed by a BatchNorm layer and a fourth Conv layer, it is added to the residual branch to generate the feature F(x, 3).
[0013] Preferably, the image restoration operation of the decoding layer includes: Input the extraction result of the spatio-temporal sequence attention of the fourth layer into the first decoding layer and into the first Conv layer of this layer. Input this output into the BatchNorm layer of the decoding layer of this layer. The output of the BatchNorm of the decoding layer is input to the second Conv layer to obtain the output result; The decoded feature of the first layer and the encoded feature of the third layer are added and input into the second decoding layer and into the first Conv layer of this layer. Input this output into the BatchNorm layer of the third layer decoding layer. The output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result F(2, 1); and the feature is input to the first CDCConv layer of the decoding layer. The output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(2, 2). F(2, 1) and F(2, 2) are added using adaptive weights, and the output obtains the encoded feature of the first layer; The decoded feature of the second layer and the encoded feature of the second layer are added and input into the third decoding layer, and into the first Conv layer of this layer. Input this output into the BatchNorm layer of the decoding layer of this layer. The output of the BatchNorm of the decoding layer is input to the second Conv layer to obtain the output result; The decoded feature of the third layer and the encoded feature of the first layer are added and input into the second decoding layer and into the first Conv layer of this layer. Input this output into the BatchNorm layer of the third layer decoding layer. The output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result F(4, 1); and the feature is input to the first CDCConv layer of the decoding layer. The output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(4, 2). F(4, 1) and F(4, 2) are added using adaptive weights, and the output obtains the encoded feature of the first layer.
[0014] Preferably, the central difference convolution CDCConv extracts edge features by calculating the difference values between the central pixel and the surrounding pixels, and its formula is: ; Where is the central pixel value, is the neighborhood pixel value, is the weight of the convolution kernel.
[0015] Preferably, the rice mapping result output by the network is classified through the Softmax function to generate a spatial distribution probability map of the rice planting area.
[0016] The beneficial effects of the present invention compared with the prior art are as follows: 1. The incomplete spatio-temporal feature fusion strategy adopted by the prior art deep learning crop distribution extraction model fails to fully utilize the time series SAR images, thus weakening the rice features and reducing the spatio-temporal correlation between different polarization methods. The present invention designs a TSA Block by combining static and dynamic attention mechanisms, aiming to fully extract the spatio-temporal features in the remote sensing ground object images. This module introduces large kernel convolution and ordinary convolution to enhance the ability to extract spatial information within the frame, and at the same time uses SENet to capture the temporal information between frames, significantly improving the model's understanding and classification performance of ground object changes.
[0017] 2. The models of the prior art ignore the internal consistency between agricultural plots and boundaries, which easily leads to the extracted boundaries not conforming to the actual situation. The present invention designs a CDC Conv Block, which combines central difference convolution and ordinary convolution with an adaptive weighting mechanism. Its central difference convolution can better capture the detailed changes in the image, and combined with the adaptive weighting, it can dynamically adjust the convolution information, effectively improving the model's expression ability for boundary details and complex spatial features. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the overall model structure implemented by the present invention; Figure 2 is Figure 1 the schematic diagram of the central difference convolution module in Figure 3 is Figure 1 the schematic diagram of the spatio-temporal attention module in DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Embodiment 1: As Figure 1 shown, a rice mapping network based on Sentinel-1 radar images of spatio-temporal sequences includes a dual-branch network SimTA, a spatio-temporal sequence attention module TSABlock, and a spatial information extraction module CDCConv based on central difference convolution; the two input ends of the dual-branch network respectively receive the VV and VH polarization data of the Sentinel-1 radar images, the output of the dual-branch network is connected to the CDCConv module, after two-layer CDCConv processing, channel superposition is performed, and the superimposed features are input into the TSABlock module, and finally the rice mapping result is output.
[0020] Preferably, the dual-branch network includes an encoder-decoder structure, an edge information extraction module, and a spatio-temporal attention module; among them, both the encoding layer and the decoding layer include multiple layers of structures. The two branches of the first-layer encoder respectively process VV and VH data, and use CDCConv for feature extraction, and its output is used as the input of the first decoding layer and is simultaneously transmitted to the second-layer encoder; the output of the second-layer encoder is connected to the second-to-last decoder layer, and feature fusion is performed with the output of the previous decoder layer through a skip connection; the third-layer encoder uses CDCConv, its output is connected to the fourth-layer encoder, and skip fusion is performed with the output of the second decoder layer.
[0021] As Figure 2 shown, preferably, the feature extraction operation of the encoding layer includes: The input feature passes through the first Conv layer, the BatchNorm layer, and the second Conv layer in sequence to generate the output feature F(1,1). At the same time, the input feature passes through the CDCConv layer and the second BatchNorm layer to generate the output feature F(1,2). F(1,1) and F(1,2) are added together through adaptive weights to form the encoded feature.
[0022] Preferably, the feature extraction operation of the encoding layer includes: the encoded feature of the first layer is input into the first Conv layer of the second encoding layer and input into this layer, and this output is input into the BatchNorm layer of the encoding layer of this layer. The BatchNorm output of the encoding layer is input to the second Conv layer to obtain the output result; The encoded feature of the second layer is input into the third encoding layer and input into the first Conv layer of this layer. This output is input into the BatchNorm layer of the third encoding layer. The BatchNorm output of the encoding layer is input to the second Conv layer to obtain the output result F(3,1); and the feature is input into the first CDCConv layer of the encoding layer. The output of the CDCConv layer is input into the second BatchNorm of the encoding layer to obtain the output result F(3,2). F(3,1) and F(3,2) are added using adaptive weights, and the output obtains the encoded feature of the third layer; The encoded feature of the third layer is input into the fourth encoding layer and input into the first Conv layer of this layer. This output is input into the BatchNorm layer of the encoding layer of this layer. The BatchNorm output of the encoding layer is input to the second Conv layer to obtain the output result.
[0023] Preferably, the output features of the VV and VH encoders of the dual-branch network are fused through channel stacking at each layer, specifically including splicing the output features of the first to fourth layer encoders in the channel dimension.
[0024] As Figure 3As shown, preferably, the operations of the spatio-temporal sequence attention module TSABlock include: Input the feature fusion result of the fourth layer into the dual-branch network of the first layer: The feature is input into the parallel structure layer of AvgPooling and MaxPooling in the first layer. The output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the first layer to obtain the feature F(1,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer. The output of this BatchNorm layer is input into the first Conv layer. The output of the first Conv layer is input into the first DW Conv layer. The first DWConv layer is input into the second Conv layer. The output of the second Conv layer added to the original feature is output as F(1,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer. The output of the second BatchNorm layer is input into the fourth Conv layer. The output of the fourth Conv layer is input into the third BatchNorm layer. The output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(1,2) to obtain the feature F(1,3). The feature obtained by cross-multiplying F(1,1) and F(1,3) is dot-multiplied with the original feature to obtain the output result of the first layer; Input the spatio-temporal information extraction feature of the first layer into the dual-branch network of the second layer: The feature is input into the parallel structure layer of AvgPooling and MaxPooling in the first layer. The output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain the feature F(2,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer. The output of this BatchNorm layer is input into the first Conv layer. The output of the first Conv layer is input into the first DW Conv layer. The first DWConv layer is input into the second Conv layer. The output of the second Conv layer added to the original feature is output as F(2,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer. The output of the second BatchNorm layer is input into the fourth Conv layer. The output of the fourth Conv layer is input into the third BatchNorm layer. The output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(2,2) to obtain the feature F(2,3). The feature obtained by cross-multiplying F(2,1) and F(2,3) is dot-multiplied with the original feature to obtain the output result of the second layer; Input the spatio-temporal information extraction feature of the second layer into the dual-branch network of the third layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain the feature F(3,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(3,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(3,2) to obtain the feature F(3,3), and the feature obtained by cross-multiplying F(3,1) and F(3,3) is dot-multiplied with the original feature to obtain the output result of the third layer; The spatio-temporal information extraction features of the third layer and the spatio-temporal information extraction features of the second layer are input into the double-branch network of the third layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain the feature F(4,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(4,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(4,2) to obtain the feature F(4,3), and the feature obtained by cross-multiplying F(4,1) and F(4,3) is dot-multiplied with the original feature to obtain the output result of the fourth layer.
[0025] Preferably, the residual structure sequentially includes a BatchNorm layer, a Conv layer, a DWConv layer, and a second Conv layer. Its output is added to the original features and then input into a third Conv layer. After being processed by the BatchNorm layer and the fourth Conv layer, it is added to the residual branch to generate the feature F(x,3).
[0026] Preferably, the image restoration operation of the decoding layer includes: Input the extraction result of the spatio-temporal sequence attention of the fourth layer into the first decoding layer, and input it into the first Conv layer of this layer. Input this output into the BatchNorm layer of the decoding layer of this layer, and the output of the BatchNorm of the decoding layer is input to the second Conv layer to obtain the output result; The decoded feature of the first layer and the encoded feature of the third layer are added and input into the second decoding layer. Input it into the first Conv layer of this layer. Input this output into the BatchNorm layer of the third layer decoding layer, and the output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result F(2,1); and the feature is input to the first CDCConv layer of the decoding layer, and the output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(2,2). F(2,1) and F(2,2) are added using adaptive weights, and the output obtains the encoded feature of the first layer; The decoded feature of the second layer and the encoded feature of the second layer are added and input into the third decoding layer. Input it into the first Conv layer of this layer. Input this output into the BatchNorm layer of the decoding layer of this layer, and the output of the BatchNorm of the decoding layer is input to the second Conv layer to obtain the output result; The decoded feature of the third layer and the encoded feature of the first layer are added and input into the second decoding layer. Input it into the first Conv layer of this layer. Input this output into the BatchNorm layer of the third layer decoding layer, and the output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result F(4,1); and the feature is input to the first CDCConv layer of the decoding layer, and the output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(4,2). F(4,1) and F(4,2) are added using adaptive weights, and the output obtains the encoded feature of the first layer.
[0027] Preferably, the central difference convolution CDCConv extracts edge features by calculating the difference values between the central pixel and the surrounding pixels, and its formula is: ; Where is the central pixel value, is the neighborhood pixel value, is the weight of the convolution kernel.
[0028] Preferably, the rice mapping result output by the network is classified through the Softmax function to generate a spatial distribution probability map of the rice planting area.
[0029] Example 2: This example discloses an example of a specific network structure and execution process: (1) Encoder structure: The feature is input to the first Conv layer of the encoding layer, and this output is input to the BatchNorm layer of the first encoding layer. The output of the first BatchNorm of the encoding layer is input to the first Conv layer to obtain the output result F(1,1); and the feature is input to the first CDCConv layer of the encoding layer. The output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(1,2). F(1,1) and F(1,2) are added using adaptive weights, and the output obtains the encoding feature of the first layer.
[0030] The encoding feature of the first layer is input to the second encoding layer and input to the first Conv layer of this layer. This output is input to the BatchNorm layer of the encoding layer of this layer, and the output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result.
[0031] The encoding feature of the second layer is input to the third encoding layer and input to the first Conv layer of this layer. This output is input to the BatchNorm layer of the third encoding layer, and the output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result F(3,1); and the feature is input to the first CDCConv layer of the encoding layer. The output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(3,2). F(3,1) and F(3,2) are added using adaptive weights, and the output obtains the encoding feature of the third layer.
[0032] The encoding feature of the third layer is input to the fourth encoding layer and input to the first Conv layer of this layer. This output is input to the BatchNorm layer of the encoding layer of this layer, and the output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result.
[0033] (2) Feature fusion: The output features of the encoder of VV in the first layer and the encoder of VH are superimposed on the channels. The output features of the encoder of VV and the encoder of VH in the second layer are superimposed on the channels. The output features of the encoder of VV and the encoder of VH in the third layer are superimposed on the channels. The output features of the encoder of VV and the encoder of VH in the fourth layer are superimposed on the channels.
[0034] (3)Spatio-temporal attention: The feature fusion result of the fourth layer is input into the double-branch network of the first layer. The feature is input into the parallel structure layer of AvgPooling and MaxPooling in the first layer. The output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the first layer to obtain the feature F(1,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer. The output of this BatchNorm layer is input into the first Conv layer. The output of the first Conv layer is input into the first DW Conv layer. The first DWConv layer is input into the second Conv layer. The output of the second Conv layer added to the original feature is output as F(1,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer. The output of the second BatchNorm layer is input into the fourth Conv layer. The output of the fourth Conv layer is input into the third BatchNorm layer. The output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(1,2) to obtain the feature F(1,3). The feature obtained by cross-multiplying F(1,1) and F(1,3) is dot-multiplied with the original feature to obtain the output result of the first layer.
[0035] The spatio-temporal information extraction feature of the first layer is input into the double-branch network of the second layer. The feature is input into the parallel structure layer of AvgPooling and MaxPooling in the first layer. The output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain the feature F(2,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer. The output of this BatchNorm layer is input into the first Conv layer. The output of the first Conv layer is input into the first DW Conv layer. The first DWConv layer is input into the second Conv layer. The output of the second Conv layer added to the original feature is output as F(2,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer. The output of the second BatchNorm layer is input into the fourth Conv layer. The output of the fourth Conv layer is input into the third BatchNorm layer. The output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(2,2) to obtain the feature F(2,3). The feature obtained by cross-multiplying F(2,1) and F(2,3) is dot-multiplied with the original feature to obtain the output result of the second layer.
[0036] The spatio-temporal information extraction features of the second layer are input into the dual-branch network of the third layer. The features are input into the parallel structure layer of AvgPooling and MaxPooling in the first layer, and the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain the feature F(3,1); the features are input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, and the output of the second Conv layer added to the original features is output F(3,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, and the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature added to F(3,2) to obtain the feature F(3,3). The feature obtained by cross-multiplying F(3,1) and F(3,3) is dot-multiplied with the original features to obtain the output result of the third layer.
[0037] The spatio-temporal information extraction features of the third layer and the spatio-temporal information extraction features of the second layer are input into the dual-branch network of the third layer. The features are input into the parallel structure layer of AvgPooling and MaxPooling in the first layer, and the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain the feature F(4,1); the features are input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, and the output of the second Conv layer added to the original features is output F(4,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, and the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature added to F(4,2) to obtain the feature F(4,3). The feature obtained by cross-multiplying F(4,1) and F(4,3) is dot-multiplied with the original features to obtain the output result of the fourth layer.
[0038] (4) Image decoding part: Input the extraction result of the spatio-temporal sequence attention of the fourth layer into the first decoding layer, input it into the first Conv layer of this layer, input this output into the BatchNorm layer of the decoding layer of this layer, and input the output of the BatchNorm layer of the decoding layer into the second Conv layer to obtain the output result.
[0039] Add the decoding features of the first layer and the encoding features of the third layer and input them into the second decoding layer. Input them into the first Conv layer of this layer, input this output into the BatchNorm layer of the third layer decoding layer, and input the output of the BatchNorm layer of the encoding layer into the second Conv layer to obtain the output result F(2,1); and input the features into the first CDCConv layer of the decoding layer, input the output of the CDCConv layer into the second BatchNorm of the encoding layer to obtain the output result F(2,2), and add F(2,1) and F(2,2) using adaptive weights to output the encoding features of the first layer.
[0040] Add the decoding features of the second layer and the encoding features of the second layer and input them into the third decoding layer. Input them into the first Conv layer of this layer, input this output into the BatchNorm layer of the decoding layer of this layer, and input the output of the BatchNorm layer of the decoding layer into the second Conv layer to obtain the output result.
[0041] Add the decoding features of the third layer and the encoding features of the first layer and input them into the second decoding layer. Input them into the first Conv layer of this layer, input this output into the BatchNorm layer of the third layer decoding layer, and input the output of the BatchNorm layer of the encoding layer into the second Conv layer to obtain the output result F(4,1); and input the features into the first CDCConv layer of the decoding layer, input the output of the CDCConv layer into the second BatchNorm of the encoding layer to obtain the output result F(4,2), and add F(4,1) and F(4,2) using adaptive weights to output the encoding features of the first layer.
[0042] Example 3: This example demonstrates the preliminary preparations and construction results in the specific network construction process: 1. Parameter setting: All experiments in the study were conducted on an NVIDIA RTX 3090 GPU, and the code was implemented using PyTorch. The proposed network was optimized using the Adam optimizer, and the momentum decay indicator was set to 0.0002 and 2 for the learning rate and batch size, respectively. The initial learning rate was set to 0.001, and MultiStepLR was used to dynamically adjust between the two; in addition, this embodiment synthesized Sentinel images and CDL of Arkansas, USA in 2017, 2018 and 2019 as real data sets to evaluate the mapping model; 2017 and 2018 were used as training sets and validation sets with a ratio of 8:2, and 2019 was used as a test set to verify temporal generalization. The current excellent rice mapping methods of SAR images are added for comparison with the method of the present invention, and these methods are as follows: Unet, TFBS, ConvLSTM, UTAE, SIMVP. 2. The experimental results are shown in Table 1 below:
[0043] Table 1: Comparison of experimental results The results are shown in Table 1. The F1 values of the overall recognition are all above 89%, and the IoU values are all above 80%. Compared with the general situation, the performance decrease is smaller. Among VV, VH and their fusion methods, SimTA has the highest classification accuracy, among which the F1 value of rice reaches 90.9%, which is 1.8%, 0.9%, 1.3%, 1.5% and 0.4% higher than Unet, TFBS, ConvLST, SimVP and UTAE respectively. This shows that SimTA has more obvious advantages than other models in cross-year applications. Among them, in feature-level fusion, the results of SimTA are close to those of U-TAE. The use of jump connections in the semantic layer of the decoder can transmit semantic information, but retain shallow information to promote more accurate classification. However, in other fusion strategies, the multi-scale fusion strategy of U-TAE did not show its advantages. In other fusion strategies, only a certain level is combined, so more emphasis is placed on the effect of the model on the two factors. In the field of SAR mapping, due to the poor imaging quality, it is necessary to extract as much edge and texture information as possible in the shallow layer. SimTA significantly enhances the feature extraction capability and suppresses the influence of noise by combining the adaptive weights of central difference convolution and ordinary convolution. It takes into account the modeling of global and local information at the same time, and greatly reduces computational redundancy while ensuring high performance. It is especially suitable for task scenarios such as SAR image mapping with complex noise and sparse information.
[0044] 3. Ablation experiment: In order to study the effectiveness of the two modules and evaluate their impact on model performance, the details are shown in Table 2: Table 2: Comparison of ablation experiment results;
[0045] As shown in Table 2, an ablation experiment was conducted on the basic model SimTA; by gradually adding CDConvBlock and TSABlock to the basic model (Baseline), the performance changes under each configuration were systematically analyzed. The experimental results show that the overall accuracy of the basic model is 0.886, and the mean intersection over union is 0.773. After adding feature-level fusion, OA reaches 0.900 and mIoU reaches 0.810. After adding CDC, the overall accuracy is improved to 0.903 and mIoU increases to 0.816, indicating that CDConvBlock significantly enhances the performance. Further, when adding TA, the performance of the model is further improved, with OA reaching 0.907 and mIoU being 0.826, showing that TA is also effective.
[0046] When CDConvBlock and TSABlock are used simultaneously, the model performs best, with the overall accuracy reaching 0.911 and mIoU being improved to 0.831, showing the synergistic effect of these two modules. In addition, the ablation experiment also shows that in the IoU metrics of different classes, the introduction of CDConvBlock and TSABlock can effectively improve the classification effect, especially the significant improvement in RiceIoU and OtherPaddyIoU, further verifying the effectiveness and complementarity of the modules. Combined with the innovative code implementation, these experimental results provide a solid empirical basis for the performance improvement of the SimTA model, highlighting the importance of module design in deep learning models.
Claims
1. A rice mapping network based on dual-branch spatio-temporal attention, characterized in that, It includes a dual-branch network SimTA, a spatio-temporal sequence attention module TSABlock, and a spatial information extraction module CDCConv based on central difference convolution. The two input ends of the dual-branch network respectively receive the VV and VH polarization data of Sentinel-1 radar images. The output of the dual-branch network is connected to the CDCConv module. After two layers of CDCConv processing, channel stacking is performed, and the stacked features are input into the TSABlock module, and finally the rice mapping result is output.
2. The rice mapping network based on dual-branch spatio-temporal attention according to claim 1, characterized in that, The dual-branch network includes an encoder-decoder structure, an edge information extraction module, and a spatio-temporal attention module. Among them, both the encoding layer and the decoding layer contain multiple-layer structures. The two branches of the first-layer encoder respectively process the VV and VH data, and CDCConv is used for feature extraction. Its output is used as the input of the first decoding layer and is simultaneously transmitted to the second-layer encoder. The output of the second-layer encoder is connected to the penultimate-layer decoder, and feature fusion is performed with the output of the previous decoding layer through a skip connection. The third-layer encoder uses CDCConv, and its output is connected to the fourth-layer encoder and performs skip fusion with the output of the second-layer decoder.
3. The rice mapping network based on dual-branch spatio-temporal attention according to claim 2, characterized in that, The feature extraction operation of the encoding layer includes: The input feature passes through the first Conv layer, the BatchNorm layer, and the second Conv layer in sequence to generate the output feature F(1,1). At the same time, the input feature passes through the CDCConv layer and the second BatchNorm layer to generate the output feature F(1,2). F(1,1) and F(1,2) are added together through adaptive weights to form the encoded feature.
4. The rice mapping network based on dual-branch spatio-temporal attention according to claim 3, characterized in that, The feature extraction operation of the encoding layer includes: The encoded feature of the first layer is input into the second encoding layer and then into the first Conv layer of this layer. The output of this layer is input into the BatchNorm layer of the encoding layer of this layer. The output of the BatchNorm layer of the encoding layer is input into the second Conv layer to obtain the output result. The encoded feature of the second layer is input into the third encoding layer and then into the first Conv layer of this layer. The output of this layer is input into the BatchNorm layer of the third-layer encoding layer. The output of the BatchNorm layer of the encoding layer is input into the second Conv layer to obtain the output result F(3,1). And the feature is input into the first CDCConv layer of the encoding layer. The output of the CDCConv layer is input into the second BatchNorm layer of the encoding layer to obtain the output result F(3,2). F(3,1) and F(3,2) are added together using adaptive weights, and the output obtains the encoded feature of the third layer. The encoded feature of the third layer is input into the fourth encoding layer and then into the first Conv layer of this layer. The output of this layer is input into the BatchNorm layer of the encoding layer of this layer. The output of the BatchNorm layer of the encoding layer is input into the second Conv layer to obtain the output result.
5. The rice mapping network based on dual-branch spatio-temporal attention according to claim 2, wherein The output features of the VV and VH encoders of the dual-branch network are fused through channel stacking, specifically including concatenating the output features of the first to fourth layer encoders in the channel dimension.
6. A rice mapping network based on dual-branch spatio-temporal attention according to claim 2, characterized in that The operations of the spatio-temporal sequence attention module TSABlock include: Input the feature fusion result of the fourth layer into the double-branch network of the first layer: The features are input into the parallel structure layer of AvgPooling and MaxPooling in the first layer. The outputs of the parallel structure layer of AvgPooling and MaxPooling in the first layer are input into the MLP layer of the first layer to obtain the feature F(1,1). The features are input into the residual structure of the first layer. It is first input into the first BatchNorm layer. The output of this BatchNorm layer is input into the first Conv layer. The output of the first Conv layer is input into the first DW Conv layer. The first DWConv layer is input into the second Conv layer. The output of the second Conv layer added to the original features is output as F(1,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer. The output of the second BatchNorm layer is input into the fourth Conv layer. The output of the fourth Conv layer is input into the third BatchNorm layer. The output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature added to F(1,2) to obtain the feature F(1,3). The feature obtained by cross-multiplying F(1,1) and F(1,3) is dot-multiplied with the original features to obtain the output result of the first layer; Input the spatio-temporal information extraction features of the first layer into the double-branch network of the second layer: The features are input into the parallel structure layer of AvgPooling and MaxPooling in the first layer. The outputs of the parallel structure layer of AvgPooling and MaxPooling in the first layer are input into the MLP layer of the second layer to obtain the feature F(2,1). The features are input into the residual structure of the first layer. It is first input into the first BatchNorm layer. The output of this BatchNorm layer is input into the first Conv layer. The output of the first Conv layer is input into the first DW Conv layer. The first DWConv layer is input into the second Conv layer. The output of the second Conv layer added to the original features is output as F(2,2) and input into the third Conv layer. The output of the third Conv layer is input into the second BatchNorm layer. The output of the second BatchNorm layer is input into the fourth Conv layer. The output of the fourth Conv layer is input into the third BatchNorm layer. The output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature added to F(2,2) to obtain the feature F(2,3). The feature obtained by cross-multiplying F(2,1) and F(2,3) is dot-multiplied with the original features to obtain the output result of the second layer; Input the spatio-temporal information extraction features of the second layer into the double-branch network of the third layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain feature F(3,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(3,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(3,2) to obtain feature F(3,3), and the feature obtained by cross-multiplying F(3,1) and F(3,3) is dot-multiplied with the original feature to obtain the output result of the third layer; The spatio-temporal information extraction features of the third layer and the spatio-temporal information extraction features of the second layer are input into the double-branch network of the third layer: The parallel structure layer of AvgPooling and MaxPooling in the first layer of feature input, the output of the parallel structure layer of AvgPooling and MaxPooling in the first layer is input into the MLP layer of the second layer to obtain feature F(4,1); the feature is input into the residual structure of the first layer, which is first input into the first BatchNorm layer, the output of this BatchNorm layer is input into the first Conv layer, the output of the first Conv layer is input into the first DW Conv layer, the first DWConv layer is input into the second Conv layer, the output of the second Conv layer is added to the original feature, and the output F(4,2) is input into the third Conv layer, the output of the third Conv layer is input into the second BatchNorm layer, the output of the second BatchNorm layer is input into the fourth Conv layer, the output of the fourth Conv layer is input into the third BatchNorm layer, the output of the third BatchNorm layer is input into the fourth Conv layer to obtain the output feature, which is added to F(4,2) to obtain feature F(4,3), and the feature obtained by cross-multiplying F(4,1) and F(4,3) is dot-multiplied with the original feature to obtain the output result of the fourth layer.
7. The rice mapping network based on dual-branch spatio-temporal attention according to claim 6, characterized in that, The residual structure sequentially includes a BatchNorm layer, a Conv layer, a DWConv layer, and a second Conv layer. Its output is added to the original feature and then input into the third Conv layer. After being processed by the BatchNorm layer and the fourth Conv layer, it is added to the residual branch to generate feature F(x,3).
8. A rice mapping network based on dual-branch spatio-temporal attention according to claim 2, characterized in that, The image restoration operation of the decoding layer includes: The extraction result of the spatio-temporal sequence attention of the fourth layer is input into the first decoding layer, and input into the first Conv layer of this layer. The output is input into the BatchNorm layer of the decoding layer of this layer, and the output of the BatchNorm of the decoding layer is input to the second Conv layer to obtain the output result; The decoding feature of the first layer and the encoding feature of the third layer are added together and input into the second decoding layer, and input into the first Conv layer of this layer. The output is input into the BatchNorm layer of the third layer decoding layer, and the output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result F(2,1); and the feature is input to the first CDCConv layer of the decoding layer, and the output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(2,2). F(2,1) and F(2,2) are added using the adaptive weight, and the output obtains the encoding feature of the first layer; The decoding feature of the second layer and the encoding feature of the second layer are added together and input into the third decoding layer, and input into the first Conv layer of this layer. The output is input into the BatchNorm layer of the decoding layer of this layer, and the output of the BatchNorm of the decoding layer is input to the second Conv layer to obtain the output result; The decoding feature of the third layer and the encoding feature of the first layer are added together and input into the second decoding layer, and input into the first Conv layer of this layer. The output is input into the BatchNorm layer of the third layer decoding layer, and the output of the BatchNorm of the encoding layer is input to the second Conv layer to obtain the output result F(4,1); and the feature is input to the first CDCConv layer of the decoding layer, and the output of the CDCConv layer is input to the second BatchNorm of the encoding layer to obtain the output result F(4,2). F(4,1) and F(4,2) are added using the adaptive weight, and the output obtains the encoding feature of the first layer.
9. The rice mapping network based on dual-branch spatio-temporal attention according to claim 1, characterized in that, The central difference convolution CDCConv extracts edge features by calculating the difference values between the central pixel and the surrounding pixels, and its formula is: ; Among them is the central pixel value, is the neighborhood pixel value, is the convolution kernel weight.
10. A rice mapping network based on dual-branch spatio-temporal attention according to claim 1, characterized in that The rice mapping result output by the network is classified through the Softmax function to generate a spatial distribution probability map of the rice planting area.