Remote sensing crop recognition method and system based on multi-modal time attention, and medium
By constructing a remote sensing crop identification method based on multimodal temporal attention, and utilizing an improved Attention UNet network architecture and a lightweight temporal attention encoder, the problems of shallow fusion layer and weak temporal coordination in the collaborative identification of optical and SAR images are solved, achieving high-precision crop identification and monitoring.
Patent Information
- Application Number
- CN202611114162.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies for crop identification using optical and SAR images suffer from drawbacks such as shallow fusion levels, weak temporal coordination, insufficient feature utilization, and lack of physical mechanisms, making it difficult to achieve high-precision crop identification.
A remote sensing crop identification method based on multimodal temporal attention is constructed. By using an improved Attention UNet network architecture, spatial features of optical and SAR images are extracted respectively. A multimodal temporal matching mechanism is introduced and combined with a lightweight temporal attention encoder to dynamically align and fuse multimodal features, and deeply explore the phenological evolution patterns of crops.
It achieves high-precision and rapid crop identification, makes full use of the complementary features of optical and SAR images, maintains high robustness under complex climatic conditions, and is suitable for large-scale agricultural remote sensing monitoring.
Smart Images

Figure CN122637232A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image processing technology, specifically to a remote sensing crop identification method, system, and medium based on multimodal temporal attention. Background Technology
[0002] Refined crop identification and automated mapping play a crucial role in agricultural production and management, and are an urgent need arising from the rapid development of information technology. With the advancement of remote sensing technology, remote sensing data has evolved into a multi-source, multi-modal, and multi-temporal data format. Therefore, multi-modal and multi-temporal imagery data of the same land area can be acquired as needed, making it possible to improve crop identification accuracy by leveraging the complementarity of multi-modal data. In the current technological context, large-scale, fine-scale crop identification using machine learning or deep learning techniques has made significant progress. Existing technical solutions can be categorized from two perspectives: data sources and technical solutions. From the perspective of data sources, they can be divided into: crop identification methods based on single-period high-resolution optical remote sensing imagery, methods based on multi-period optical remote sensing imagery, methods based on multi-period synthetic aperture radar (SAR) imagery, and methods based on the synergy of optical and SAR imagery. From the perspective of technical solutions, they can be divided into: crop identification methods based on machine learning models and crop identification methods based on deep learning models.
[0003] Crop identification based on single-period high-resolution optical remote sensing images and crop identification based on multi-period optical images are widely used for fine crop identification due to the rich spectral information and high readability of optical images.
[0004] In existing technologies, although there are schemes using optical and SAR fusion methods for fine identification, they often suffer from shortcomings such as shallow fusion levels, weak temporal coordination, insufficient feature utilization, and lack of physical mechanisms. Patent CN 115130547B identifies target crops by calculating the distance matching similarity between the temporal polarization features of crops and standard crops. However, due to speckle noise in SAR image data pairs and issues such as overlay, perspective contraction, and shadowing in side-looking radar imaging, the accuracy of crop identification models is affected, and relying solely on single-polarization SAR image data is insufficient to meet the needs of precision agriculture. Accurate identification of remotely sensed crops depends on stable and reliable data support. However, due to the large differences in crop planting distribution and significant differences in crop growth phenology, relying solely on single-phase images is insufficient to accurately identify target crops at different growth stages and in different regions. Furthermore, for crop identification technologies using multiple temporal images, spaceborne optical sensors are easily affected by adverse weather conditions such as clouds and rain, leading to data loss and an inability to fully capture crop growth changes during critical growth periods. This poses a significant challenge to the robustness of models in dealing with missing temporal data and their ability to generalize to different regions. Synthetic Aperture Radar (SAR), as an active microwave remote sensing technology, has all-weather, all-time data acquisition capabilities, is sensitive to surface vegetation structure, and can penetrate clouds to obtain surface information, effectively compensating for the shortcomings of optical imagery.
[0005] Patent CN 113420838 B uses a multi-scale feature extraction network and a spatial / spectral attention mechanism to fuse enhanced features extracted from optical and SAR images, achieving crop classification using collaborative optical and SAR image data. However, this scheme uses a weighted fusion method to fuse features extracted from optical and SAR image data. This direct fusion method ignores the physical meaning and differences between the two different modalities of data, leading to feature space redundancy or even conflicts. When there are significant differences in the temporal sequence of the two modalities of data, the fused SAR image data may also have a negative impact on the accurate identification of crops.
[0006] Patent CN 116664959 B utilizes fused optical and SAR time-series image data to identify crops through a pre-trained random forest crop classification model. This scheme employs a stacking and stitching method for fusing optical and SAR time-series data, without deep, complementary cross-modal interaction between the optical and SAR image data at the feature level. Summary of the Invention
[0007] To address the shortcomings of existing methods that use optical and SAR fusion for precise crop identification, such as shallow fusion levels, weak temporal coordination, insufficient feature utilization, and lack of physical mechanisms, this invention provides a remote sensing crop identification method, system, and medium based on multimodal temporal attention. An improved Attention UNet network architecture is constructed to build a multimodal temporal attention fusion identification model, extracting optical and SAR spatial features from optical and SAR temporal datasets respectively. A multimodal temporal matching mechanism is introduced to dynamically align and fuse multimodal spatial features based on observation dates. This is further combined with a lightweight temporal attention encoder to capture temporal dependencies, deeply mining crop phenological evolution patterns, and ultimately achieving high-precision, rapid, end-to-end crop identification.
[0008] The above-mentioned technical objective of the present invention is achieved through the following technical solution:
[0009] This solution provides a remote sensing crop identification method based on multimodal temporal attention, the method including:
[0010] Collect remote sensing image data at multiple different time points during the crop growth cycle, including SAR image data and optical image data;
[0011] The remote sensing image data is preprocessed to construct an optical time-series dataset and a SAR time-series dataset;
[0012] A multimodal temporal attention fusion recognition model is constructed to identify crops by combining the optical temporal dataset and the SAR temporal dataset. Multi-period optical spatial features and multi-period SAR spatial features are extracted from the optical and SAR temporal datasets, respectively. First, the optical and SAR spatial features are temporally aligned, then arranged chronologically. Multi-period optical and SAR spatial features are fused using the same date fusion and different date splicing methods to construct temporal spatial fusion features. The temporal spatial fusion features are positionally encoded, and a multi-head attention mechanism is introduced to extract attention-weighted temporal aggregate features based on the temporal spatial fusion features. The crop identification result is obtained by decoding the optical spatial features, SAR spatial features, and temporal aggregate features.
[0013] A further optimization scheme is that the SAR image data includes: HH polarization band data transmitted horizontally and received horizontally, HV polarization band data transmitted horizontally and received vertically, VH polarization band data transmitted vertically and received horizontally, and VV polarization band data transmitted vertically and received vertically.
[0014] A further optimization scheme is that the construction methods for the optical time-series dataset and the SAR time-series dataset include:
[0015] The optical image data is first radiometrically calibrated and then atmospherically corrected, and then normalized.
[0016] The SAR image data is first radiometrically calibrated and then converted into backscattering coefficients by decibel conversion. The SAR image data after decibel conversion is then normalized. Finally, the normalized optical image data is embedded into the normalized SAR image data according to the Julian coding method.
[0017] Selected according to the crop growth cycle An optical time-series dataset was constructed from optical image data at optical time steps, and selected... A SAR time series dataset is constructed from SAR image data at each SAR time step.
[0018] A further optimized scheme is that the extraction methods for the optical spatial features and SAR spatial features include:
[0019] Two independent, temporally shared, improved Attention UNet encoders are used as spatial encoders to progressively extract local spatial features from optical temporal datasets and SAR temporal datasets, respectively.
[0020] An attention gating mechanism is introduced to dynamically weight the low-level features output by each spatial encoder: the low-level features are mapped to the same number of intermediate channels as the high-level features output by the decoder, and the high-level features are upsampled to match the size of the low-level features. Then, optical spatial attention masks and SAR spatial attention masks are generated.
[0021] By splicing high-level features with weighted low-level features, and then using continuous convolution to reduce dimensionality, we obtain optical spatial features and SAR spatial features with unified dimensions.
[0022] A further optimized approach is to include methods for obtaining spatial fusion features, including:
[0023] Acquire Julian days from optical image data and SAR image data, map the Julian days to feature dimensions, and inject spatial features;
[0024] Calculate the Julian diurnal difference between the optical time step and the SAR time step. When the difference between Julian and Japanese If the time interval is less than 1 day, it is considered to be consistent in time sequence; otherwise, it is considered to be inconsistent in time sequence.
[0025] The spatial fusion feature is obtained by weighted fusion of temporally consistent optical spatial features and SAR spatial features.
[0026] A further optimized solution is that the method for obtaining the temporal aggregation features includes:
[0027] The spatial fusion features are transformed into temporal fusion features, and positional encoding is performed based on a sine-cosine encoding method to inject temporal positional information into the temporal fusion features;
[0028] The temporal fusion features after injecting temporal location information are mapped to queries, keys, and values, and the temporal attention weights are calculated by scaling dot product attention.
[0029] Based on the multilayer perceptron, the temporal fusion features are restored to obtain temporal aggregate features, which fuse key phenological information from all time steps.
[0030] A further optimization is that the temporal aggregation features are acquired based on a lightweight temporal attention encoder.
[0031] A further optimization scheme is that the multimodal temporal attention fusion recognition model is a dual-branch AttentionUNet network architecture. The AttentionUNet network architecture introduces a temporal parameter sharing mechanism, which reuses the same set of AttentionUNet parameters for all time steps of the same modality.
[0032] This solution also provides a remote sensing crop identification system based on multimodal temporal attention, used to implement the aforementioned remote sensing crop identification method based on multimodal temporal attention. The system includes:
[0033] The acquisition module is used to acquire remote sensing image data at multiple different time points during the crop growth cycle. The remote sensing image data includes SAR image data and optical image data.
[0034] The preprocessing module is used to preprocess the remote sensing image data to construct an optical time-series dataset and a SAR time-series dataset.
[0035] The identification module is used to construct a multimodal temporal attention fusion identification model, which combines the optical temporal dataset and the SAR temporal dataset for crop identification: Multi-period optical spatial features and multi-period SAR spatial features are extracted from the optical and SAR temporal datasets respectively; the optical and SAR spatial features are first temporally aligned, then arranged in chronological order, and multi-period optical and SAR spatial features are fused according to the method of fusion for the same date and splicing for different dates to construct a temporal spatial fusion feature; the temporal spatial fusion feature is positionally encoded, and a multi-head attention mechanism is introduced to extract attention-weighted temporal aggregated features based on the temporal spatial fusion feature; the crop identification result is obtained by decoding based on the optical spatial features, SAR spatial features, and temporal aggregated features.
[0036] This solution also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, can implement the remote sensing crop identification method based on multimodal temporal attention as described above.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. This invention provides a remote sensing crop identification method, system, and medium based on multimodal temporal attention; it constructs a multimodal temporal attention fusion identification model based on an improved Attention UNet network architecture, extracting optical spatial features and SAR spatial features from optical temporal datasets and SAR temporal datasets respectively; it introduces a multimodal temporal matching mechanism, dynamically aligns and fuses multimodal spatial features according to the observation date, and then combines a lightweight temporal attention encoder to capture temporal dependencies, deeply mining the phenological evolution patterns of crops, and finally achieving end-to-end high-precision and rapid crop identification; it fully utilizes the feature complementarity of optical image data and SAR image data, effectively fusing the spectral-texture of optical image data and the texture, structure, scattering, and polarization features of SAR image data, while capturing the temporal dependencies of crop growth cycles, maintaining high robustness under complex climatic conditions, and is suitable for large-scale agricultural remote sensing monitoring scenarios.
[0039] 2. This invention provides a remote sensing crop identification method, system, and medium based on multimodal temporal attention; using time series as a benchmark, it achieves multimodal feature fusion for the same date and splicing of multimodal features for different dates through dynamic date matching, avoiding the loss of temporal information caused by interpolation, preserving the temporal independence of different modal data features, and effectively improving the accuracy of crop identification.
[0040] 3. This invention provides a remote sensing crop identification method, system, and medium based on multimodal temporal attention. By introducing a multi-head attention mechanism, the importance of time steps in temporal data is quantified, and sine-cosine position coding avoids temporal confusion, thus improving the robustness of the multimodal temporal attention fusion identification model to random missing time-series data under complex weather conditions. Furthermore, the temporal attention mechanism enables the multimodal temporal attention fusion identification model to quantify the importance of different time steps, capture the global key phenological features of crop growth, and effectively improve the accuracy of the model in crop identification.
[0041] 4. This invention provides a remote sensing crop identification method, system, and medium based on multimodal temporal attention; it adopts a network architecture of dual-branch spatial coding, multimodal temporal fusion, temporal attention aggregation, and decoding classification. The decoder uses skip connections to reuse the intermediate features (such as texture, boundary, etc.) of the dual-branch spatial encoder and fuses them with the temporal aggregation features, solving the problem of crop boundary ambiguity caused by temporal aggregation, and realizing effective complementarity between spatial features and temporal features. Attached Figure Description
[0042] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:
[0043] Figure 1 This is a schematic diagram of the remote sensing crop identification method based on multimodal temporal attention;
[0044] Figure 2 This is a schematic diagram of the network architecture of a multimodal temporal attention fusion recognition model;
[0045] Figure 3 A schematic diagram of the network architecture for obtaining temporal aggregation features;
[0046] Figure 4 This is a schematic diagram of the structure of a remote sensing crop identification system based on multimodal temporal attention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0048] To address the shortcomings of existing schemes using optical and SAR fusion methods for fine identification, such as shallow fusion levels, weak temporal coordination, insufficient feature utilization, and lack of physical mechanisms, this solution provides the following embodiments:
[0049] Example 1
[0050] This embodiment provides a remote sensing crop identification method based on multimodal temporal attention, such as... Figure 1 As shown, the methods include:
[0051] Step 1: Collect remote sensing image data at multiple different time points during the crop growth cycle. The remote sensing image data includes SAR image data and optical image data.
[0052] In this embodiment, the optical image data selected are the B2, B3, B4, B8 and B11 channels of the Sentinel-2 satellite, which can effectively reflect the spectral characteristics of crops such as chlorophyll content and vegetation coverage.
[0053] In this embodiment, the SAR image data selected is HH polarization band data transmitted horizontally and received horizontally, HV polarization band data transmitted horizontally and received vertically, VH polarization band data transmitted vertically and received horizontally, and VV polarization band data transmitted vertically and received vertically. This data can penetrate clouds and rain, providing scattering characteristics of crop structures. Due to differences in the original resolution of the optical image data, the B11 channel is resampled to 10m using bilinear interpolation to ensure spatial resolution consistency with other optical bands and SAR image data. The specific resampling process is as follows:
[0054] ;
[0055] in, Represents the raw pixel values of optical image data. Indicates the target pixel coordinates. Indicates the interpolation window index; and These represent the floor values of x and y, respectively. This represents the pixel value that was resampled.
[0056] Step 2: Preprocess the remote sensing image data to construct optical time-series datasets and SAR time-series datasets;
[0057] Specifically, this step includes the following methods:
[0058] S21, the optical image data is first radiometrically calibrated and then atmospherically corrected. Since there is a difference in dimensions between the optical reflectance image data collected by Sentinel-2 satellite and the SAR backscattering coefficient collected by Sentinel-1 satellite, this embodiment performs radiometric calibration on the optical image data, specifically converting the original brightness DN value of the pixel into reflectance and performing atmospheric correction to eliminate the influence of aerosols.
[0059] S22, the SAR image data is first radiometrically calibrated and then the scattering coefficient is converted. The SAR image data after the scattering coefficient conversion is normalized. Finally, the acquisition time sequence of the optical image data is embedded into the normalized SAR image data. In this embodiment, the SAR data acquired by Sentinel-1 satellite is radiometrically calibrated, then converted into backscattering coefficients by decibel conversion, and normalized to the range of 0~1 using the maximum and minimum values.
[0060] S23, selected according to the crop growth cycle An optical time-series dataset is constructed from optical image data at optical time steps. Select A SAR time-series dataset is constructed from SAR image data at each SAR time step. ;in, B represents the set of real number matrices; B represents the batch size. and These represent the time series lengths of optical image data and SAR image data, respectively. and H represents the number of data channels for optical image data and SAR image data, respectively. H represents the image height and W represents the image width.
[0061] Step 3, construct as follows Figure 2 The multimodal temporal attention fusion recognition model shown combines optical temporal datasets and SAR temporal datasets for crop identification: Multi-period optical spatial features and multi-period SAR spatial features are extracted from the optical and SAR temporal datasets, respectively; the optical and SAR spatial features are first temporally aligned, then arranged chronologically, and fused using the same-date fusion and different-date splicing methods to construct a temporal spatial fusion feature; the temporal spatial fusion feature is positionally encoded, and a multi-head attention mechanism is introduced to extract attention-weighted temporal aggregation features based on the temporal spatial fusion feature; the crop identification result is obtained by decoding the optical spatial features, SAR spatial features, and temporal aggregation features.
[0062] Optical spatial features include both optical spatial features and spectral features of optical time-series datasets;
[0063] SAR spatial features include SAR spatial features, texture features, and structural features of SAR time series datasets;
[0064] Specifically, methods for extracting optical spatial features and SAR spatial features include:
[0065] S31 uses two independent temporally shared Attention UNet encoders as spatial encoders to progressively extract local spatial features from optical temporal datasets and SAR temporal datasets, respectively.
[0066] S32 introduces an attention gating mechanism to dynamically weight the low-level features output by each spatial encoder: the low-level features and the high-level features output by the decoder are mapped to the same number of intermediate channels, and the high-level features are upsampled to match the size of the low-level features. Then, optical spatial attention masks and SAR spatial attention masks are generated by ReLU activation function and 1×1 convolution, respectively.
[0067] S33, by splicing high-level features with weighted low-level features, and then reducing the dimensionality through continuous convolution, obtains optical spatial features and SAR spatial features with unified dimensions.
[0068] Spatial feature extraction is the technical foundation of crop identification. Simultaneously capturing the spectral features of optical image data and the texture-structure features of SAR image data, and through effective feature fusion, helps improve the accuracy of crop identification. This scheme proposes to use a temporally shared Attention UNet encoder as a spatial encoder, which strengthens crop planting boundary information through an attention gating mechanism, while reusing parameters to reduce computational load; taking multimodal temporal data (optical temporal dataset and SAR temporal dataset) as input, each is processed separately by an independent temporally shared Attention UNet encoder branch.
[0069] like Figure 2 As shown in this embodiment, the multimodal temporal attention fusion recognition model is a dual-branch AttentionUNet network architecture. The dual-branch AttentionUNet network architecture introduces a temporal parameter sharing mechanism, which reuses the same set of AttentionUNet parameters for all time steps of the same modality. This is implemented by a lightweight temporal attention encoder L-TAE.
[0070] Taking the encoder branch corresponding to the optical temporal dataset as an example, the temporal shared attention UNet encoder, as a spatial encoder, contains one consecutive convolutional block (DoubleConv), four downsampling blocks (Down), and four attention gates (AttentionGate). The consecutive convolutional block consists of two consecutive 3×3 convolutional layers, a ReLU function, and max pooling, progressively extracting features and reducing spatial resolution to extract local spatial features of the crop. The mathematical form of the consecutive convolutional block is:
[0071] ; ;
[0072] in, Indicates the input image; This represents the features extracted after the first 3×3 convolution in a consecutive convolutional block; () represents a two-dimensional convolution operation; () indicates the number of groups, which is suitable for small-batch training scenarios of remote sensing data and avoids unstable normalization effects; ReLU() indicates the activation function; k indicates the kernel size and p indicates the amount of padding. This represents the extracted output features.
[0073] The decoder's decoding process is mainly completed by four downsampling blocks. Through multiple convolutional pooling operations, the dimensionality of the features is reduced and the receptive field is expanded, achieving multi-scale progressive feature extraction.
[0074] ;
[0075] Where s represents the pooling step size, This represents the feature map extracted by the previous layer of the network; Indicates max pooling; Represents consecutive convolutional blocks in the UNet network; This represents the multi-scale progressive features extracted from downsampling blocks after continuous convolution.
[0076] In this embodiment, the Attention UNet network architecture adds an attention gating mechanism to the low-level features output by the dynamically weighted encoder, reducing redundant low-level features extracted by skip connections; it suppresses irrelevant regions of the input image while highlighting salient features in specific local regions. The calculation process is as follows: A 1×1 convolution maps the high-level features G output by the decoder to the low-level features X output by the encoder to the same number of intermediate channels, and upsamples the high-level features to match the size of the low-level features; its mathematical expression is:
[0077] ; ;
[0078] Subsequently, a spatial attention mask is generated by using the ReLU activation function and a 1×1 convolution. The calculation method is as follows:
[0079] ;
[0080] ;
[0081] ;
[0082] in, Represents high-level features; k represents the kernel size; Indicates the number of channels after mapping; Indicates the number of intermediate channels; Indicates low-level features; () denotes a bilinear interpolation function; Indicates a spatial attention mask; Indicates the shape and size of low-level features; Indicates batch normalization; This represents the low-level features output by the decoder; ) represents a 1×1 convolution; This represents the high-level feature that matches the low-level feature size after upsampling; G represents the high-level feature output by the decoder. This scheme uses a spatial attention masking mechanism to highlight the spectral and textural features of crops while suppressing background noise from soil and water. Subsequently, by concatenating high-level features with weighted low-level features, semantic information and spatial details are preserved simultaneously. After dimensionality reduction processing, a spatial feature map with uniform dimensionality is formed, ultimately outputting the optical spatial features. SAR spatial characteristics .
[0083] Methods for obtaining spatial fusion features include:
[0084] G31 acquires the Julian day of optical image data and SAR image data, maps the Julian day to a feature dimension, and injects spatial features;
[0085] G32, calculates the Julian diurnal difference between the optical time step and the SAR time step. When the difference between Julian and Japanese If the time interval is less than 1 day (i.e., the optical image data and SAR image data were observed on the same day), it is determined that the time sequence is consistent; otherwise, it is determined that the time sequence is inconsistent.
[0086] G33 obtains spatial fusion features by weighted fusion of temporally consistent optical spatial features and SAR spatial features.
[0087] The decoder consists of multiple upsampling blocks, each upsampling block (Up) comprising bilinear interpolation, attention-weighted feature concatenation, and a continuous convolutional block (DoubleConv); its mathematical expression is as follows:
[0088] ;
[0089] ;
[0090] ;
[0091] Where G represents high-level spatial features; This represents the spatial characteristics after upsampling; scale represents the channel compression factor. () represents the channel dimension concatenation function; This indicates the number of feature channels before upsampling; the concatenated dual-channel features can be compressed into single-channel features using consecutive convolutional blocks (DoubleConv). This represents the stitching result of optical branch and SAR branch features; Indicates the dimension of the splicing channel.
[0092] To adapt to multi-temporal scenarios, this embodiment introduces a temporal parameter sharing mechanism in the Attention UNet network architecture: the same set of Attention UNet parameters is reused for all time steps of the same modality. Through the merging and splitting of batch and temporal dimensions, its mathematical expression is as follows:
[0093] ;
[0094] in, Indicates the initial timing input; The network structure represents the Attention UNet encoding structure; B represents the batch size; T represents the time series month length; C represents the number of image channels; H represents the image height; W represents the image width; batch processing of multi-time series data is achieved through dimensional changes; and view() represents the tensor operation function.
[0095] Multimodal data fusion often faces spatiotemporal alignment issues. Traditional methods unify the temporal sequence through interpolation, but this approach loses the unique temporal information of different modalities. This solution uses the Julian Day (DOY) for temporal alignment and feature fusion, dynamically aligning features from different modalities according to time order. This enables adaptive feature fusion processing, combining features from the same date with features from different dates, and using optical temporal spatial features. SAR temporal and spatial characteristics Julian's Day and optical imaging data Julian Day and SAR image data As input, the temporal encoder is mapped to temporal features consistent with the spatial feature dimension; the Julian Day DOY is converted to a normalized value and mapped to the feature dimension, and spatial features are injected to distinguish temporal discriminativeness. Its mathematical expression is as follows:
[0096] ;
[0097] ;
[0098] DOY stands for Julian Japan; This refers to the unified Julian Japan; () represents the dimension expansion function; Linear represents a linear layer; The number of channels represents the temporal characteristics, consistent with the number of channels representing the spatial characteristics.
[0099] Calculate the DOY difference between the optical and SAR time steps, when the difference is... If the observation period is less than 1 day (i.e., the optical and SAR image data were observed on the same day), it is considered a time-consistent feature; otherwise, it is considered a time-inconsistent feature.
[0100] Features: For time-consistent optical and SAR features, a weighted fusion method is used to fuse spatial features.
[0101] The optical and SAR spatial features extracted by the spatial encoder are embedded in the temporal dimension and then concatenated along the channel dimension. The dimensionality is reduced to 64 channels through 1×1 convolution, achieving nonlinear fusion of multimodal features. Its mathematical expression is as follows:
[0102] ;
[0103] in, This represents the features after nonlinear fusion; idx represents the temporal index after fusion; Cat represents the concatenation of features along the channel dimension; Conv() represents 1×1 convolution; and GroupNorm represents group normalization. This represents the optical temporal feature at temporal coding position i; This represents the SAR temporal feature at temporal coding position j.
[0104] During feature fusion, the independence of optical features and SAR data features is preserved, and the calculation formula is as follows:
[0105] ;
[0106] Finally, arrange all according to DOY time sequence. Generate fusion features with a unified temporal dimension. ;in, N represents the number of fusionable time-series logs; This represents the features obtained after nonlinear fusion with a convolution kernel of type k.
[0107] like Figure 3 As shown, the methods for obtaining temporal aggregation features include:
[0108] T31 transforms spatial fusion features into temporal fusion features and performs position encoding based on a sine-cosine encoding method, injecting temporal position information into the temporal fusion features;
[0109] T32 maps the temporal fusion features after injecting temporal location information to queries, keys, and values, and calculates temporal attention weights by scaling dot product attention.
[0110] T33, based on the multilayer perceptron, restores the dimensions of the temporal fusion features to obtain temporal aggregate features, and fuses key phenological information from all time steps.
[0111] Crop growth exhibits significant phenological differences, and traditional temporal fusion strategies can mask the temporal contribution of key phenological stages. The temporal aggregation features in this scheme are acquired based on a lightweight temporal attention encoder, such as... Figure 3As shown, a lightweight temporal attention encoder is used to capture the dynamic growth dependence of crops by quantifying the relative importance of different time steps through multi-head attention. Its input data is the spatial fusion feature obtained by temporal fusion of optical spatial features and SAR spatial features. The specific implementation process is as follows:
[0112] Spatial integration features The dimension is from B× ×C×H×W is converted to B×H×W× ×C treats each spatial location pixel as an independent temporal sample.
[0113] In temporal feature extraction, positional encoding is performed using a sine-cosine coding method to represent temporal features. The mathematical expression for injecting timing location information is as follows:
[0114] ;
[0115] ;
[0116] ;
[0117] Where, pos {0,…, -1} represents the position index; d {0,…,C / 2-1} represents the channel pair index; Represents the position encoding matrix; This represents the temporal characteristics after the injection of location information.
[0118] Will The mapping is represented as query (Q), key (K), and value (V). Temporal attention weights are calculated using scaled dot product attention, as shown in the following formula:
[0119] ;
[0120] ;
[0121] ;
[0122] In the formula, Q(m) represents the query matrix at the m-th time step; K(n) represents the key matrix at the n-th time step; and V represents the value matrix. The number of attention heads is indicated; in this embodiment, attention head 1, attention head 2, and attention head 3 are included. This represents a scaling factor to prevent the normalized Softmax gradient from vanishing due to excessively large attention scores. Indicates the time step index; This represents the attention score of the m-th time step to the n-th time step; This represents the normalized attention weights; Temporal features representing attention-weighted processing;
[0123] Finally, the dimensions are reduced to aggregated global temporal features of size B×C×H×W using a multilayer perceptron (MLP). It integrates key phenological information from all time steps.
[0124] This approach utilizes the decoder portion of the Attention UNet network architecture for decoding. The decoder comprises upsampling blocks and output convolutional blocks, with its input being the temporal aggregated features from the temporal attention encoder. The intermediate features of the dual-branch spatial encoder are used to gradually recover the spatial dimensions through upsampling, ultimately outputting pixel-level crop classification results.
[0125] It consists of four upsampling blocks, each containing bilinear interpolation upsampling, an attention gating mechanism (AttentionGate), and a continuous convolutional block (DoubleConv) for progressively aggregating temporal features. Perform sampling.
[0126] Taking the fourth upsampling block as an example, its calculation formula is as follows:
[0127] ;
[0128] , ;
[0129] ;
[0130] In the formula, This indicates the output feature of the third upsampled block; Interpolate() represents bilinear interpolation, and scale=2 represents the upsampling factor; , representing the attention-weighted low-level feature at t=1; This represents the weighted intermediate features computed through attention gating; AttentionGate() represents the attention gating, and the weighted features are... Highlight local details of crops; Cat indicates channel-dimensional splicing; DoubleConv() indicates reducing the spliced features to 64 dimensions and outputting the features. .
[0131] The features output by the decoder are mapped to the corresponding crop category using a 1×1 convolution. The calculation formula is as follows:
[0132] ;
[0133] ;
[0134] in, The predicted value representing the crop category is processed by the Softmax function to obtain the pixel-level category probability; This represents the features output after the last upsampling block. The number of input feature channels, The number of output channels is the number of crop categories. This represents the probability that the b-th sample and the (h, w)-th pixel belong to the c-th crop category; exp() represents the exponential function; the denominator is the sum of the exponents of the predicted values for all crop categories.
[0135] This scheme uses time series data as a benchmark and achieves multimodal feature fusion for the same date through dynamic date matching. The fusion scheme, which combines multimodal features from different dates, avoids the loss of temporal information caused by interpolation, preserves the temporal independence of different modal data features, and effectively improves crop identification accuracy. By introducing a multi-head attention mechanism to quantify the importance of time steps in the time series data, and using sine-cosine positional encoding to avoid temporal confusion, the robustness of the multimodal temporal attention fusion identification model to random missing time series data under complex weather conditions is improved. Furthermore, the temporal attention mechanism enables the multimodal temporal attention fusion identification model to quantify the importance of different time steps, capture global key phenological features of crop growth, and effectively improve the model's accuracy in crop identification. A network architecture of dual-branch spatial coding, multimodal temporal fusion, temporal attention aggregation, and decoding classification is adopted. The decoder reuses intermediate features (such as texture and boundaries) from the dual-branch spatial encoder through skip connections to fuse with temporal aggregated features, solving the problem of crop boundary ambiguity caused by temporal aggregation and achieving effective complementarity between spatial and temporal features.
[0136] This scheme employs a dual-branch Attention UNet spatial encoding, a lightweight temporal attention encoder (L-TAE), and a spatiotemporal collaborative decoding architecture for both encoding and decoding. In the encoding phase, the dual-branch spatial encoder is configured independently for optical and SAR image data. Each encoder's attention gating projects high-level semantic features and low-level spatial features to an intermediate dimension via 1×1 convolutions. After bilinear interpolation alignment, spatial codes are generated, and then low-level features are weighted to highlight crop planting boundaries. Each upsampling block reuses the encoder's attention-weighted intermediate features, concatenates them with temporally aggregated features, and then refines them, improving the model's ability to identify crops planted in fragmented field plots.
[0137] Example 2
[0138] This embodiment provides a remote sensing crop identification system based on multimodal temporal attention, such as... Figure 4 As shown, the system for implementing the remote sensing crop identification method based on multimodal temporal attention described in Example 1 includes:
[0139] The acquisition module is used to acquire remote sensing image data at multiple different time points during the crop growth cycle. The remote sensing image data includes SAR image data and optical image data.
[0140] The preprocessing module is used to preprocess the remote sensing image data to construct an optical time-series dataset and a SAR time-series dataset.
[0141] The identification module is used to construct a multimodal temporal attention fusion identification model, which combines the optical temporal dataset and the SAR temporal dataset for crop identification: Multi-period optical spatial features and multi-period SAR spatial features are extracted from the optical and SAR temporal datasets respectively; the optical and SAR spatial features are first temporally aligned, then arranged in chronological order, and multi-period optical and SAR spatial features are fused according to the method of fusion for the same date and splicing for different dates to construct a temporal spatial fusion feature; the temporal spatial fusion feature is positionally encoded, and a multi-head attention mechanism is introduced to extract attention-weighted temporal aggregated features based on the temporal spatial fusion feature; the crop identification result is obtained by decoding based on the optical spatial features, SAR spatial features, and temporal aggregated features.
[0142] Example 3
[0143] An embodiment provides a computer-readable medium having a computer program stored thereon. This computer program, when executed by a processor, can implement the remote sensing crop identification method based on multimodal temporal attention as described in Embodiment 1; specifically, it implements the following steps:
[0144] Step 1: Collect remote sensing image data at multiple different time points during the crop growth cycle. The remote sensing image data includes SAR image data and optical image data.
[0145] Step 2: Preprocess the remote sensing image data to construct an optical time-series dataset and a SAR time-series dataset;
[0146] Step 3: Construct a multimodal temporal attention fusion recognition model, combining the optical temporal dataset and the SAR temporal dataset for crop identification: Extract multi-period optical spatial features and multi-period SAR spatial features from the optical and SAR temporal datasets respectively; First, align the optical and SAR spatial features temporally, then arrange them in chronological order, fusing the multi-period optical and SAR spatial features according to the same date and different date combinations to construct a temporal spatial fusion feature; Encode the temporal spatial fusion feature by position, and introduce a multi-head attention mechanism to extract attention-weighted temporal aggregation features based on the temporal spatial fusion feature; Decode the optical spatial features, SAR spatial features, and temporal aggregation features to obtain the crop identification result.
[0147] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A remote sensing crop identification method based on multimodal temporal attention, characterized in that the method... include: Collect remote sensing image data at multiple different time points during the crop growth cycle, including SAR image data and optical image data; The remote sensing image data is preprocessed to construct an optical time-series dataset and a SAR time-series dataset; A multimodal temporal attention fusion recognition model is constructed, and crop recognition is performed by combining the optical temporal dataset and the SAR temporal dataset: multi-period optical spatial features and multi-period SAR spatial features are extracted based on the optical temporal dataset and the SAR temporal dataset, respectively. First, the optical spatial features and SAR spatial features are time-series aligned, then arranged in chronological order, and multiple periods of optical spatial features and multiple periods of SAR spatial features are fused according to the method of fusion of the same date and splicing of different dates to construct a time-series spatial fusion feature. The temporal-spatial fusion features are positionally encoded, and a multi-head attention mechanism is introduced to extract attention-weighted temporal aggregation features based on the temporal-spatial fusion features; Crop identification results are obtained by decoding based on optical spatial features, SAR spatial features, and temporal aggregation features.
2. The remote sensing crop identification method based on multimodal temporal attention according to claim 1, characterized in that, The SAR image data includes: HH polarization band data transmitted horizontally and received horizontally, HV polarization band data transmitted horizontally and received vertically, VH polarization band data transmitted vertically and received horizontally, and VV polarization band data transmitted vertically and received vertically.
3. The remote sensing crop identification method based on multimodal temporal attention according to claim 1, characterized in that, The methods for constructing the optical time series dataset and the SAR time series dataset include: The optical image data is first radiometrically calibrated and then atmospherically corrected, and then normalized. The SAR image data is first radiometrically calibrated and then converted into backscattering coefficients by decibel conversion. The SAR image data after decibel conversion is then normalized. Finally, the normalized optical image data is embedded into the normalized SAR image data according to the Julian coding method. Selected according to the crop growth cycle An optical time-series dataset was constructed from optical image data at optical time steps, and selected... A SAR time series dataset is constructed from SAR image data at each SAR time step.
4. The remote sensing crop identification method based on multimodal temporal attention according to claim 3, characterized in that, The methods for extracting optical spatial features and SAR spatial features include: Two independent, temporally shared, improved Attention UNet encoders are used as spatial encoders to progressively extract local spatial features from optical temporal datasets and SAR temporal datasets, respectively. An attention gating mechanism is introduced to dynamically weight the low-level features output by each spatial encoder: the low-level features are mapped to the same number of intermediate channels as the high-level features output by the decoder, and the high-level features are upsampled to match the size of the low-level features. Then, optical spatial attention masks and SAR spatial attention masks are generated. By splicing high-level features with weighted low-level features, and then using continuous convolution to reduce dimensionality, we obtain optical spatial features and SAR spatial features with unified dimensions.
5. The remote sensing crop identification method based on multimodal temporal attention according to claim 4, characterized in that, Methods for obtaining spatial fusion features include: Acquire Julian days from optical image data and SAR image data, map the Julian days to feature dimensions, and inject spatial features; Calculate the Julian diurnal difference between the optical time step and the SAR time step. When the difference between Julian and Japanese If the time interval is less than 1 day, it is considered to be consistent in time sequence; otherwise, it is considered to be inconsistent in time sequence. The spatial fusion feature is obtained by weighted fusion of temporally consistent optical spatial features and SAR spatial features.
6. The remote sensing crop identification method based on multimodal temporal attention according to claim 1, characterized in that, The method for obtaining the temporal aggregation features includes: The spatial fusion features are transformed into temporal fusion features, and positional encoding is performed based on a sine-cosine encoding method to inject temporal positional information into the temporal fusion features; The temporal fusion features after injecting temporal location information are mapped to queries, keys, and values, and the temporal attention weights are calculated by scaling dot product attention. Based on the multilayer perceptron, the temporal fusion features are restored to obtain temporal aggregate features, which fuse key phenological information from all time steps.
7. The remote sensing crop identification method based on multimodal temporal attention according to claim 6, characterized in that, The temporal aggregation features are obtained based on a lightweight temporal attention encoder.
8. The remote sensing crop identification method based on multimodal temporal attention according to claim 1, characterized in that, The multimodal temporal attention fusion recognition model is a dual-branch Attention UNet network architecture. The Attention UNet network architecture introduces a temporal parameter sharing mechanism, which reuses the same set of Attention UNet parameters for all time steps of the same modality.
9. A remote sensing crop identification system based on multimodal temporal attention, characterized in that, The system is used to implement the remote sensing crop identification method based on multimodal temporal attention as described in any one of claims 1-8, the system comprising: The acquisition module is used to acquire remote sensing image data at multiple different time points during the crop growth cycle. The remote sensing image data includes SAR image data and optical image data. The preprocessing module is used to preprocess the remote sensing image data to construct an optical time-series dataset and a SAR time-series dataset. The identification module is used to construct a multimodal temporal attention fusion identification model, which combines the optical temporal dataset and the SAR temporal dataset for crop identification: Multi-period optical spatial features and multi-period SAR spatial features are extracted from the optical and SAR temporal datasets respectively; the optical and SAR spatial features are first temporally aligned, then arranged in chronological order, and multi-period optical and SAR spatial features are fused according to the method of fusion for the same date and splicing for different dates to construct a temporal spatial fusion feature; the temporal spatial fusion feature is positionally encoded, and a multi-head attention mechanism is introduced to extract attention-weighted temporal aggregated features based on the temporal spatial fusion feature; the crop identification result is obtained by decoding based on the optical spatial features, SAR spatial features, and temporal aggregated features.
10. A computer-readable medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, can implement the remote sensing crop identification method based on multimodal temporal attention as described in any one of claims 1-8.
Citation Information
Patent Citations
Polarimetric SAR crop classification method, system, equipment and medium based on multi-feature joint time series matching
CN115130547B