This invention relates to a
remote sensing interpretation method and
system that integrates multi-source spatiotemporal spectral features and a visual model. First, optical imagery,
harmonic data, and
synthetic aperture radar (SAR) data of the target area are acquired and preprocessed. Next, spatiotemporal features of the optical imagery and
harmonic data are extracted separately and fused using a cross-attention mechanism to generate fused features. Then, the fused features and SAR data are subjected to image
serialization, temporal encoding, and
spatial encoding processing, and spatiotemporal features are extracted using a self-attention mechanism. Finally, the spatiotemporal features are decoded and the multi-source feature
convolution results are fused to generate pixel-level
land cover classification results. This invention introduces
harmonic data to capture the temporal variation trend of ground features, utilizes the visual
Transformer self-attention mechanism to capture spatiotemporal interdependencies, and combines multi-
source data to capture multi-dimensional features of ground features, avoiding the limitations of a single
data source. Furthermore, the fusion of multi-source temporal features allows the model to adapt to different geographical environments, improving generalization performance.