A marine ecological dynamic prediction method and system based on multi-modal remote sensing and spatiotemporal graph

By deeply fusing multimodal remote sensing features and using a spatiotemporal ecological process inference network, the problems of multimodal data fusion and spatiotemporal modeling in marine ecological remote sensing prediction are solved, achieving high-precision, interpretable future trajectory prediction and causal attribution, and supporting smart ocean management.

CN121708503BActive Publication Date: 2026-04-17SHANDONG MARINE RESOURCE AND ENVIRONMENT RESEARCH INSTITUTE (SHANDONG MARINE ENVIRONMENTAL MONITORING CENTER SHANDONG AQUATIC PRODUCTS QUALITY INSPECTION CENTER)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG MARINE RESOURCE AND ENVIRONMENT RESEARCH INSTITUTE (SHANDONG MARINE ENVIRONMENTAL MONITORING CENTER SHANDONG AQUATIC PRODUCTS QUALITY INSPECTION CENTER)
Filing Date
2026-02-11
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing marine ecological remote sensing prediction technologies cannot capture the spatial heterogeneity between modes when fusing multimodal data, lack accuracy in modeling spatiotemporal dynamic processes, and lack interpretability in prediction results, making it difficult to achieve high-precision, interpretable, and intelligent monitoring.

Method used

A marine ecological dynamic prediction method based on multimodal remote sensing and spatiotemporal mapping is adopted. By deeply fusing multimodal remote sensing features, using a spatiotemporal ecological process inference network (STEP-Net) that couples spatial diffusion and temporal inertia, and using deep Taylor decomposition for causal attribution, a high-precision and interpretable future trajectory prediction is generated.

Benefits of technology

It improves the accuracy and interpretability of marine ecological prediction, achieves high-precision future trajectory prediction and causal attribution explanation, and supports the precise management of smart oceans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708503B_ABST
    Figure CN121708503B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of marine ecological monitoring, in particular to a marine ecological dynamic prediction method and system based on multi-modal remote sensing and a space-time graph. The method comprises the following steps: performing multi-modal remote sensing feature depth fusion according to obtained remote sensing image data, wherein the multi-modal remote sensing feature depth fusion comprises multi-modal feature map extraction and space embedding, local cross-modal feature interaction and alignment, and global context aggregation and node embedding generation; constructing a space-time ecological process reasoning network STEP-Net based on coupled space diffusion and time inertia, wherein the space-time ecological process reasoning network STEP-Net comprises multi-order space diffusion field modeling and ecological state transition unit construction; and utilizing deep space-time representation generated by the STEP-Net to perform future-oriented and uncertain trajectory prediction and cause and effect attribution based on deep Taylor decomposition. The application constructs a brand-new technical framework integrating fine fusion, mechanism reasoning and interpretable prediction around the precise prediction and scientific early warning demand of marine ecological system dynamic evolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marine ecological monitoring technology, and in particular to a method and system for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping. Background Technology

[0002] Remote sensing technology has become a mainstream method for large-scale, long-term marine ecological monitoring, enabling the acquisition of key ecological factors such as seawater temperature, chlorophyll concentration, and sea surface roughness at a macroscopic scale. However, in response to the need for dynamic evolution prediction and early warning of complex marine ecological events such as red tides and green tides, existing technologies still face systemic bottlenecks in areas such as the refinement and adaptability of multimodal data fusion, the mechanistic modeling of spatiotemporal dynamic processes, and the quantification of uncertainty and causal attribution explanation of prediction results. These limitations make it difficult to support the construction of a high-precision, interpretable, and decision-making intelligent marine ecological monitoring system.

[0003] First, at the level of multimodal remote sensing data fusion, existing technologies are far from achieving a level of refinement and adaptability. Although the industry has recognized the necessity of fusing multi-source data such as optical, synthetic aperture radar (SAR), and sea surface temperature (SST) data, mainstream fusion methods—such as simple feature vector concatenation or weighted summation based on global attention—have a fundamental flaw: they assume that the correlations between different modes are homogeneous across the entire observation space. However, in real marine environments, such correlations exhibit high spatial heterogeneity. For example, in a localized area of ​​a closed bay, a small increase in water temperature may be the dominant factor driving a surge in chlorophyll concentration; while in another open area influenced by ocean currents, changes in sea surface roughness may be more strongly associated with biological aggregation. Existing methods cannot capture these dynamically changing intermodal cooperative patterns at different geographical locations and scales, resulting in inaccurate feature representations after fusion, information redundancy, and even conflicts, thus limiting the performance ceiling of downstream prediction models.

[0004] Secondly, at the level of spatiotemporal dynamic process modeling, existing models provide a superficial simulation of ecosystem evolution. Traditional image segmentation networks (such as U-Net) completely ignore the time dimension, while time series models (such as LSTM) isolate spatial interactions. Even the more advanced graph neural network architectures combined with recurrent units (such as GCN+GRU) are mechanistically agnostic. Standard GRU or LSTM units cannot distinguish, from a physical or ecological perspective, the source of the information they receive: whether it is a diffusion effect from the spatial neighborhood, the temporal inertia of their own state, or an external shock from newly observed data. This leads to deviations between the model's simulation of complex processes such as algal bloom diffusion, drift, and decay and the actual physical processes, making it difficult to further improve prediction accuracy.

[0005] Finally, regarding the usability and credibility of the prediction results, existing technologies have two major shortcomings. First, most prediction models only provide deterministic predictions for a single future point in time (e.g., "an algal bloom will occur in 24 hours"), and this "black and white" result cannot meet the needs of modern risk management. Decision-makers need to understand the uncertainty of event development, that is, the multiple possible evolutionary paths and their corresponding probabilities, in order to conduct tiered responses and optimize resource allocation. Second, most deep learning models are "black boxes," and their prediction results lack interpretability. Even if the model accurately warns of risks, it cannot answer the crucial question of how it was achieved. The lack of a causal attribution mechanism that can accurately trace and quantify the contribution of each input factor makes it difficult for decision-makers to fully trust the warning results, and it cannot provide effective guidance for subsequent scientific research and precise governance.

[0006] Therefore, there is an urgent need to develop a novel marine ecological prediction method that can achieve spatially adaptive multimodal fine fusion, incorporate a spatiotemporal evolution inference engine that conforms to ecological dynamics, and provide probabilistic future trajectory prediction and high-precision causal attribution explanation, so as to comprehensively break through the current technological bottlenecks. Summary of the Invention

[0007] To address the aforementioned issues in the four core dimensions of current marine ecological remote sensing prediction—data fusion, process modeling, prediction paradigm, and result reliability—this invention provides a method and system for dynamic prediction of marine ecology based on multimodal remote sensing and spatiotemporal mapping.

[0008] In a first aspect, the present invention provides a method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping, which adopts the following technical solution:

[0009] A method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping includes:

[0010] Acquire remote sensing image data;

[0011] Deep fusion of multimodal remote sensing features is performed based on the acquired remote sensing image data, including multimodal feature map extraction and spatial embedding, local cross-modal feature interaction and alignment, and global context aggregation and node embedding generation;

[0012] A spatiotemporal ecological process inference network STEP-Net based on coupled spatial diffusion and temporal inertia is constructed, including multi-order spatial diffusion field modeling and ecological state transition unit construction.

[0013] We utilize the deep spatiotemporal representations generated by STEP-Net to perform future-oriented, uncertain trajectory prediction and causal attribution based on deep Taylor decomposition.

[0014] Furthermore, the multimodal feature map extraction and spatial embedding includes using a convolutional neural network to map remote sensing image data into feature maps that preserve spatial structure. , represented as:

[0015] Where mod is the modal identifier. The original input image for the modality; This represents the feature extraction network for the corresponding modality; The output modal feature map is then processed by gridding, and each feature block is flattened to form a one-dimensional vector sequence.

[0016] Where P is the total number of pixel blocks after the spatial grid is divided; Indicates a splitting operation; Indicates the flattening operation; The original feature vector after flattening the p-th pixel block; p is the position index of the pixel block. Finally, Transformer position encoding is used to generate position encoded values ​​for each position index p and dimension index i:

[0017] ,

[0018] Where i is the dimension index of the feature vector; This represents the dimension of the uniform feature vector processed internally by the model. This represents the position encoding vector at position p; the flattened feature block vector is then passed through a linear projection layer for dimensionality reduction, and the corresponding position encoding vector is added to obtain the final feature block vector carrying position information. :

[0019] ,

[0020] Wherein, Linear(.): a linear projection layer; Let be the feature vector of the p-th pixel block carrying location information.

[0021] Furthermore, the local cross-modal feature interaction and alignment includes employing a multi-head cross-attention mechanism to process the input feature block vector. Generate the required Q, K, V vectors using a linear projection matrix:

[0022] Where i is the index of the attention head; Nh is the total number of attention heads; The feature vector of the p-th pixel block is the input. This indicates that the i-th attention head is used to generate a learnable linear projection matrix of query, key, and value. Let represent the query, key, and value vector generated by the i-th head. Then, for each head, the attention output is calculated independently using the scaled dot product attention formula:

[0023] ,in, The key vector dimension for each attention head; Let T be the output vector of the i-th attention head, and let T be the transpose of all... After all individual calculations are completed, the outputs are stitched together and then processed through a linear projection layer. The results are then fused to obtain the final output of the multi-head attention layer:

[0024] ,

[0025] Concat(.) is a quantity concatenation operation; This is the final output linear projection layer of the head attention module; The final aggregated output for multi-head attention;

[0026] To integrate the original information and stabilize training, the output of the multi-head attention is added to the original query vector via a residual connection, and then layer normalization is performed.

[0027] ,in, The intermediate feature vectors are then processed through residual connections and layer normalization. Finally, a nonlinear transformation is performed via a feedforward neural network, followed by further residual connections and layer normalization to complete the final update of the optical modal feature blocks, resulting in a composite of local information from other modalities. :

[0028] Where FFN(.) is a feedforward neural network, To update the obtained optical modal feature blocks.

[0029] Furthermore, the global context aggregation and node embedding generation includes inputting a locally aligned sequence of feature blocks into a multi-layer global self-attention Transformer encoder, enabling each pixel block to interact with other pixel blocks using depth information, thereby constructing a global context-aware representation. , represented as:

[0030] ,in, The input is a sequence of locally aligned feature blocks; Standard Transformer encoder for global context modeling; This is the final representation of the p-th pixel block, which includes global context information. To generate the embedding vector representing the entire region, an attention-based pooling mechanism is used, and a single-layer MLP activated by tanh is used to compute an unnormalized importance score for each pixel block. :

[0031] ,

[0032] in, Let be the unnormalized importance score of the p-th pixel block; These are the learnable parameters for the attention pooling network; The activation function is used; these scores are converted into attention weights that sum to 1 using the Softmax function. :

[0033] , where j is the index used for summation over all pixel blocks; Let be the attention weight of the p-th pixel block in the weighted pooling; based on the learned weights, perform a weighted summation on the final representations of all pixel blocks to obtain the final multimodal ecosystem embedding vector belonging to the n-th graph node at time t. :

[0034] Where n is the index of the graph node, This represents the final multimodal ecological embedding vector belonging to the nth graph node at time t.

[0035] Furthermore, the multi-order spatial diffusion field modeling includes constructing a multi-order spatial diffusion field for aggregating multi-order neighborhood information by stacking multi-layer graph convolutional networks (GCNs). First, the output of the HCAT module... Left multiplication by normalized adjacency matrix Preliminary neighborhood aggregation features were obtained. : ,in, This is the normalized geographical adjacency matrix; The node feature matrix is ​​input from the HCAT module. This is a first-order neighborhood aggregation feature matrix; then, linear transformations and nonlinear activations are applied to the aggregated features to generate a first-order diffusion field. :

[0036] Where ReLU is the activation function; It is the learnable weight matrix of the first layer GCN; This represents the first-order spatial diffusion field; finally, using the first-order diffusion field as input, second-order neighborhood information aggregation is performed first: ,in, This is the normalized geographical adjacency matrix; The neighborhood aggregation feature matrix is ​​a second-order matrix; then, a linear transformation and nonlinear activation are performed to generate a second-order diffusion field. :

[0037] Where ReLU is the activation function; It is the learnable weight matrix of the second-layer GCN; Represents a second-order spatial diffusion field; to dynamically integrate influences from different spatial distances, a set of learnable weights is used. By weighted and fused diffusion fields of different orders, the final integrated spatial diffusion field is obtained. :

[0038] ,

[0039] in, For learnable fusion weight parameters ; This is the final spatial diffusion field matrix that integrates the effects of multiple neighborhoods.

[0040] Furthermore, the construction of the ecological state transition unit includes using the three basic gating mechanisms calculated by ESU to independently adjust the intensity of external forcing, spatial diffusion, and temporal inertia, respectively generating a candidate state component dominated by external forcing. A candidate state component dominated by "spatial diffusion" :

[0041] ,

[0042] in, These are candidate components driven by external forcing and candidate components driven by spatial diffusion, respectively. This represents the tanh activation function; This represents all learnable weight matrices and bias vectors within the ESU unit; to dynamically balance the influence of driving forces, a novel source-dominated gate is introduced using the ESU. Based on current new observations and spatial diffusion trends, the sources upon which state changes depend are dynamically determined, expressed as:

[0043] ,in, This represents all learnable weight matrices and bias vectors within the ESU unit, where W-type matrices apply to the current input and U-type matrices apply to the loop state. This represents the Sigmoid activation function; The gating vector represents the dominant gate from which the source originates;

[0044] Finally, by utilizing all the gating mechanisms, the candidate components are weighted and fused to form a final candidate state that is more physically interpretable. :

[0045] ,

[0046] in, These are, respectively, candidate components driven by external forcing, candidate components driven by spatial diffusion, and the final fusion candidate state; Represents element-wise multiplication; utilizes time inertia gates. The historical state of the previous moment and the currently generated candidate state By performing a weighted combination, the final state update at the current moment is completed, resulting in... :

[0047] .

[0048] Furthermore, the use of the deep spatiotemporal representation generated by STEP-Net for future-oriented, uncertain trajectory prediction includes employing an attention-based autoregressive decoder to predict the k-th future step from complex historical information. The autoregressive decoder dynamically extracts the most relevant context, first based on the state of the previous time step. , with each node in the historical state sequence Unnormalized alignment scores are calculated using a small alignment network. :

[0049] Where T represents the current time; k represents the index of the predicted future time step; This represents the state vector of the nth node. This represents the hidden state of the internal loop unit of the decoder at time T +k - 1; Learnable parameters for the attention alignment network; This represents the alignment score between the current state of the decoder and the state of historical node n.

[0050] Furthermore, the method of using the deep spatiotemporal representation generated by STEP-Net for future-oriented, uncertain trajectory prediction also includes normalizing the alignment score using the Softmax function to obtain attention weights. The portion of the historical state sequence to focus on in the current prediction step is determined using attention weights, and is expressed as:

[0051] ,in, Attention weights represent the importance of the historical state of the nth node when predicting the kth step. Based on the calculated attention weights, the historical state sequence is weighted and summed to generate a context vector containing the most critical historical information. :

[0052] ,in, This generates a context vector; finally, the context vector is concatenated with the current state of the decoder, and then passed through an output linear layer and a softmax function to generate the probability distribution of the current prediction step across all C ecological categories. :

[0053] ,

[0054] in, This is the final output linear layer of the decoder; This represents the probability distribution matrix for the final output at the k-th future step.

[0055] Furthermore, the causal attribution based on deep Taylor decomposition includes hierarchical correlation propagation (LRP) based on deep Taylor decomposition. Rules define the contribution of neurons. The main part and negative part :

[0056] ,

[0057] in, Index representing a neuron; Indicates the first The activation value of layer neuron j; Indicates the connection of the first Layer neuron j and the first The weights of layer neuron k; This represents the contribution of neuron j to neuron k; The function indicates that only the positive values ​​are retained; This indicates that only the negative part should be retained; according to The rules will be the first Layer correlation score Backpropagation is performed based on positive and negative contributions, respectively, up to the [number]th [unit]. Neurons in the layer; first calculate the forward propagation component. and negative propagation component :

[0058] ,

[0059] in, This is the hierarchical index of the neural network; Index representing a neuron; yes The correlation score of layer neurons; Represents the positive and negative correlation components that propagate from neuron k to neuron j; It is a hyperparameter of the LRP rule; This is a very small stable term; finally, the total correlation score of neuron j is obtained by summing all the positive and negative correlation components from the upper node k. :

[0060] By applying the correlation propagation rule layer by layer from the final output layer of the model, the predicted correlation of the top layer is ultimately distributed to the neurons of the bottom layer without loss.

[0061] Secondly, a marine ecological dynamic prediction system based on multimodal remote sensing and spatiotemporal mapping includes:

[0062] The data acquisition module is configured to acquire remote sensing image data;

[0063] The feature fusion module is configured to perform deep fusion of multimodal remote sensing features based on the acquired remote sensing image data, including multimodal feature map extraction and spatial embedding, local cross-modal feature interaction and alignment, and global context aggregation and node embedding generation.

[0064] The model building module is configured to construct the STEP-Net spatiotemporal ecological process inference network based on coupled spatial diffusion and temporal inertia, including multi-order spatial diffusion field modeling and ecological state transition unit construction.

[0065] The prediction module is configured to perform future-oriented, uncertain trajectory prediction and causal attribution based on deep Taylor decomposition using the deep spatiotemporal representation generated by STEP-Net.

[0066] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping.

[0067] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping.

[0068] In summary, the present invention has the following beneficial technical effects:

[0069] Compared with the multiple technical bottlenecks of existing marine ecological remote sensing prediction methods, such as coarse multimodal information fusion, lack of spatiotemporal evolution modeling mechanism, single prediction paradigm and poor interpretability of results, this invention focuses on the needs of accurate prediction and scientific early warning of "dynamic evolution of marine ecosystems" and constructs a brand-new technical framework that integrates refined fusion, mechanistic reasoning and interpretable prediction.

[0070] First, by employing a multimodal remote sensing feature deep fusion module (HCAT) based on hierarchical attention transformers, this invention addresses the challenge of traditional fusion methods failing to handle the "spatial heterogeneity" of intermodal correlations. Its "local alignment-global aggregation" mechanism adaptively captures and models the inherent collaborative patterns of optical, SAR, and SST data at different spatial scales, generating a unified ecological state representation far more accurate and robust than simple stitching or global weighting. Second, the spatiotemporal ecological process inference network (STEP-Net) coupled with spatial diffusion and temporal inertia designed in this invention, particularly its core Ecological State Transition Unit (ESTU), represents a significant innovation in traditional spatiotemporal modeling. By establishing independent gating mechanisms for the three core driving forces—spatial diffusion, temporal inertia, and external forcing—ESTU can simulate state evolution in a more physical and ecologically sound manner, accurately depicting the diffusion, drift, and formation / dissipation processes of phenomena such as algal blooms, thus improving the accuracy of dynamic prediction. Finally, the probabilistic decoding and causal attribution module (PTD-CA) for future trajectories in this invention achieves a dual breakthrough in both the practicality and reliability of predictions. On the one hand, it upgrades the prediction paradigm from "deterministic single-point prediction" to "probabilistic trajectory prediction," providing crucial quantitative evidence of uncertainty for risk assessment and tiered early warning. On the other hand, its attribution mechanism based on deep Taylor decomposition can accurately trace the source, visually explaining the key driving factors behind the early warning in the form of a "heat map," completely breaking the "black box" attribute of the model and greatly enhancing decision-makers' trust in the early warning results.

[0071] In typical nearshore eutrophication and algal bloom dynamic prediction tasks, the method of this invention improves the accuracy (ACC) of ecological state prediction for the next 24 hours to 93.2%; for critical severe algal bloom events, the false negative rate (MR) is significantly reduced from over 10% in traditional spatiotemporal models to 5.8%; at the same time, the false positive rate (FAR) is also controlled at an excellent level of 4.6%; and the consistency assessment (Att-Consist) between its causal attribution results and the judgment of oceanographic experts is as high as 92.3%. This invention combines high precision, high reliability, and strong interpretability, providing strong technical support for realizing precise and scientific management of the smart ocean. Attached Figure Description

[0072] Figure 1 This is a schematic diagram of a marine ecological dynamic prediction method based on multimodal remote sensing and spatiotemporal mapping according to Embodiment 1 of the present invention;

[0073] Figure 2 This is a schematic diagram comparing the positive evaluation indicators of each model in Embodiment 1 of the present invention;

[0074] Figure 3 This is a schematic diagram comparing the negative evaluation indicators of each model in Embodiment 1 of the present invention;

[0075] Figure 4 This is a schematic diagram of the performance trade-offs of each model in Embodiment 1 of the present invention;

[0076] Figure 5 This is a radar schematic diagram illustrating the overall performance of each model in Embodiment 1 of the present invention;

[0077] Figure 6 This is a thermodynamic diagram illustrating the performance of each model in Embodiment 1 of the present invention. Detailed Implementation

[0078] The present invention will be further described in detail below with reference to the accompanying drawings.

[0079] Example 1

[0080] Reference Figure 1 This embodiment of a method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping includes:

[0081] This invention proposes a method for predicting the dynamics of marine ecosystems based on multimodal remote sensing and spatiotemporal mapping networks. The core idea of ​​this technical solution is not only to integrate the surface information of different remote sensing data, but also to deeply explore their inherent physical and ecological coupling relationships, and to explicitly model and reason about the spatiotemporal evolution of the ecosystem within a unified framework. The overall technical path is precisely designed into three progressively advancing core modules: (1) a multimodal remote sensing feature deep fusion module based on hierarchical attention transformers (HCAT); (2) a spatiotemporal ecological process reasoning network coupling spatial diffusion and temporal inertia (STEP-Net); and (3) a probabilistic decoding and causal attribution module for future trajectories (PTD-CA).

[0082] (1) Multimodal remote sensing feature deep fusion module based on hierarchical attention transformer (HCAT)

[0083] This module aims to address the information loss and conflict issues during the fusion of multi-source heterogeneous remote sensing data, particularly the core challenge of "spatial heterogeneity," which existing methods cannot capture intermodal correlations. To this end, this module designs a "micro-to-macro" fusion paradigm. Its execution process includes three closely linked sub-steps: feature map extraction and spatial embedding, local cross-modal feature interaction and alignment, and global context aggregation and node embedding generation. Ultimately, it generates a highly condensed, representative multimodal ecological embedding vector for each spatial node in the downstream spatiotemporal network.

[0084] 1) Multimodal feature map extraction and spatial embedding

[0085] Because the original remote sensing images differ greatly in physical meaning and data structure, direct fusion without processing will inevitably lead to a decline in model performance. Therefore, the primary task of this step is to transform the images of different modalities from pixel space to a unified deep feature space containing rich semantic information, and to inject spatial location information into these features. This is a prerequisite for the effective operation of the subsequent attention mechanism.

[0086] First, in order to extract a unified representation containing high-level semantic information from raw remote sensing images with varying structures, this invention employs independent convolutional neural networks (CNNs) as encoders for each modality, processing the input images... After depthwise convolution, nonlinear activation, and pooling operations, it is mapped to a feature map that preserves the spatial structure. :

[0087] ,

[0088] Here, mod is a modal identifier, taking one of {opt, sar, sst}. mod=opt represents the optical remote sensing mode, mainly derived from multispectral or hyperspectral sensors, with features including water color, chlorophyll concentration inversion, etc.; mod=sar represents the synthetic aperture radar mode, with features including sea surface roughness, wind and wave distribution, sea surface floating object information, etc.; mod=sst represents the sea surface temperature mode, with features derived from infrared remote sensing or microwave radiometers, reflecting the impact of large-scale sea temperature changes on ecological processes. The original input image for modality mod is a three-dimensional tensor; This represents the feature extraction network for the corresponding modality; This is the output modal feature map.

[0089] To process spatial information in a sequential manner, we perform grid-based segmentation of the extracted feature maps and flatten each feature block to form a one-dimensional vector sequence:

[0090] ,

[0091] Where P is the total number of pixel blocks after the spatial grid is divided; Indicates a splitting operation; Indicates the flattening operation; The original feature vector after flattening the p-th pixel block; p is the position index of the pixel block, with a value range of [1, P].

[0092] To enable the model to perceive the original spatial location of each feature block, we employ the classic Transformer sine / cosine positional encoding method. First, we generate positional encoding values ​​for each position index p and dimension index i:

[0093] ,

[0094] Where i is the dimension index of the feature vector; This represents the dimension of the uniform feature vector processed internally by the model. This represents the position encoding vector at the p-th position.

[0095] The flattened feature block vector is then passed through a linear projection layer for dimensionality reduction, and the corresponding positional encoding vector is added to obtain the final feature vector carrying positional information. :

[0096] ,

[0097] Wherein, Linear(.): a linear projection layer used for dimensional transformation; Let be the feature vector of the p-th pixel block that ultimately carries location information, and let be the dimension.

[0098] 2) Local cross-modal feature interaction and alignment

[0099] After obtaining the unified feature block sequence across modalities, the next key challenge is how to deeply interact these features at the microscale to capture their synergistic effects across the same geographic location. The core of this step is to employ a multi-head cross-attention mechanism, forcing the model to independently learn the nonlinear dependencies between optical, SAR, and SST features from multiple "viewpoints" (i.e., multiple "heads") within each pixel block.

[0100] This process uses one modality (e.g., optics) as the "query," and combinations of other modalities as the "key" and "value." For the i-th attention head, firstly, the input feature block vector needs to be... (mod is the modality identifier, taking the value of one of {opt, sar, sst}) Generate the required Q, K, V vectors for this head using the linear projection matrix unique to this head:

[0101] ,

[0102] Where i is the index of the attention head, and its value ranges from [1, ..., ... ]; Nh represents the total number of attention heads; Let mol∊{opt,sar,sst} be the feature vector of the p-th pixel block. This indicates that the i-th attention head is used to generate a learnable linear projection matrix of query, key, and value. This represents the query, key, and value vector generated from the i-th header.

[0103] Subsequently, each head independently calculates its attention output using the scaled dot product attention formula:

[0104] ,

[0105] in, The key vector dimension for each attention head is... ; Let T be the output vector of the i-th attention head, and T be the transpose.

[0106] In all After all the calculations have been completed, their outputs are concatenated and passed through a final linear projection layer. The results are then fused to obtain the final output of the multi-head attention layer:

[0107] ,

[0108] Concat(.) is a quantity concatenation operation; The final output linear projection layer of the head attention module has the following weight matrix: ; This is the final aggregated output of multi-head attention.

[0109] To integrate the original information and stabilize training, the output of the multi-head attention is added to the original query vector via a residual connection, and then layer normalization is performed:

[0110] ,

[0111] in, Residual connections and intermediate feature vectors after layer normalization.

[0112] Finally, a nonlinear transformation is performed using a feedforward neural network (FFN), followed by residual connections and layer normalization to complete the final update of the optical modal feature blocks, resulting in a product that incorporates local information from other modalities. :

[0113] ,

[0114] Here, FFN(.) is a feedforward neural network, consisting of two linear layers and a nonlinear activation function. The final output is an updated optical modal feature block obtained after a round of deep cross-modal interaction and alignment.

[0115] 3) Global context aggregation and node embedding generation

[0116] Local alignment addresses the issue of feature interactions at the microscopic level, but marine ecological events are macroscopic phenomena whose evolution depends on interactions across a large spatial scale. Therefore, this step aims to capture the long-range dependencies between any two spatial locations (pixel blocks) and effectively aggregate global information into a final embedding vector that can represent each spatial node (e.g., primitive).

[0117] To this end, we input the locally aligned feature block sequence output from the previous step into a multi-layer global self-attention Transformer encoder, enabling each pixel block to interact with all other pixel blocks on depth information, thereby constructing a global context-aware representation. :

[0118] ,

[0119] in, The input is a sequence of locally aligned feature blocks; The standard Transformer encoder used for global context modeling has an internal structure similar to (1) of (2), but here it is self-attention; This is the final representation of the p-th pixel block, which includes global context information.

[0120] After obtaining the pixel block sequence rich in global context, we employ an attention-based pooling mechanism to generate an embedding vector representing the entire region (i.e., a graph node). This mechanism first computes an unnormalized importance score for each pixel block using a small neural network (a single-layer MLP with tanh activation). :

[0121] ,

[0122] in, Let be the unnormalized importance score of the p-th pixel block; Let be the learnable parameters of the attention pooling network, namely the weight matrix, bias vector, and weight vector, respectively. This is the activation function.

[0123] These scores are converted into attention weights that sum to 1 using the Softmax function. :

[0124] ,

[0125] Where j is the index used for summation over all pixel blocks; represents the attention weight of the p-th pixel block in weighted pooling.

[0126] Based on the learned weights, the final representations of all pixel blocks are weighted and summed to obtain the final multimodal ecosystem embedding vector belonging to the nth graph node at time t. :

[0127] ,

[0128] Where n is the index of the graph node, This represents the final multimodal ecological embedding vector belonging to the nth graph node at time t.

[0129] (2) Spatiotemporal ecological process reasoning network coupled with spatial diffusion and temporal inertia (STEP-Net)

[0130] This module aims to simulate and extrapolate the spatiotemporal evolution of ecosystems, with its core function being to address the shortcomings of existing models in that the "mechanism is unknown" regarding the intrinsic driving forces of evolution. This module characterizes spatial interactions through multi-order spatial diffusion field modeling and innovatively employs the Ecological State Transition Unit (ESTU) to precisely simulate temporal state transitions, thereby enabling in-depth reasoning about ecological processes.

[0131] 1) Modeling of multi-order spatial diffusion fields

[0132] The spatial impact of marine ecosystems is not limited to directly adjacent areas, but also has transmission effects over long distances. In order to comprehensively capture this multi-layered spatial dependence, this step constructs a "multi-level spatial diffusion field" that can aggregate multi-level neighborhood information by stacking multi-layer graph convolutional networks (GCNs).

[0133] First, perform first-order neighborhood information aggregation, that is, aggregate the output of the HCAT module. Left multiplication by normalized adjacency matrix Preliminary neighborhood aggregation features were obtained. :

[0134] ,

[0135] in, This is the normalized geographical adjacency matrix; The node feature matrix is ​​input from the HCAT module. It is a first-order neighborhood aggregation feature matrix.

[0136] Linear transformation and nonlinear activation are applied to the aggregated features to generate a first-order diffusion field. :

[0137] ,

[0138] Where ReLU is the activation function; It is the learnable weight matrix of the first layer GCN; It represents a first-order spatial diffusion field.

[0139] To capture the effects at greater distances, we take the first-order diffusion field as input and repeat the above process, first performing second-order neighborhood information aggregation:

[0140] ,

[0141] in, This is the normalized geographical adjacency matrix; It is a second-order neighborhood aggregation feature matrix.

[0142] Then, linear transformation and nonlinear activation are performed to generate a second-order diffusion field. :

[0143] ,

[0144] Where ReLU is the activation function; It is the learnable weight matrix of the second-layer GCN; This represents a second-order spatial diffusion field.

[0145] To dynamically integrate the influences from different spatial distances, we use a set of learnable weights. By weighted and fused diffusion fields of different orders, the final integrated spatial diffusion field is obtained. :

[0146] ,

[0147] in, Let be the learnable fusion weight parameters, be a scalar, and ; This is the final spatial diffusion field matrix that integrates the influence of multiple neighborhoods.

[0148] 2) Ecological State Transition Unit (ESTU)

[0149] This step is the core innovation of the present invention, which aims to design a brand-new cyclic unit whose state update mechanism is more in line with the laws of ecological dynamics and can explicitly decouple and simulate the three core driving forces behind state evolution: external forcing, spatial diffusion, and temporal inertia.

[0150] First, ESU calculates three fundamental gating mechanisms, each used to independently regulate external forcing (from new observations). Spatial diffusion (from the diffusion field) ) and time inertia (from historical states) The strength of )

[0151] ,

[0152] in, These are gating vectors with values ​​between [0,1], representing external forced gate, spatial diffusion gate, and time inertial gate, respectively. This represents the Sigmoid activation function; This represents all learnable weight matrices and bias vectors within the ESU unit, where W-type matrices apply to the current input and U-type matrices apply to the loop state.

[0153] Each candidate state component dominated by "external forcing" is generated. A candidate state component dominated by "spatial diffusion" :

[0154] ,

[0155] in, These are candidate components driven by external forcing and candidate components driven by spatial diffusion, respectively. This represents the tanh activation function; This represents all learnable weight matrices and bias vectors within the ESU unit.

[0156] To dynamically balance the influence of these two driving forces, ESU introduces a completely new "source-dominated gate". Based on current new observations and spatial diffusion trends, this phylogenetic tree dynamically determines which source the state changes should rely more on:

[0157] ,

[0158] in, This represents all learnable weight matrices and bias vectors within the ESU unit, where W-type matrices apply to the current input and U-type matrices apply to the loop state. This represents the Sigmoid activation function; The gating vector represents the dominant gate from which the source originates.

[0159] By utilizing all gating mechanisms, the two candidate components are weighted and fused to form the final, more physically interpretable candidate state. :

[0160] ,

[0161] in, These are, respectively, candidate components driven by external forcing, candidate components driven by spatial diffusion, and the final fusion candidate state; This indicates element-wise multiplication.

[0162] Finally, using time inertia gates The historical state of the previous moment and the currently generated candidate state By performing a weighted combination, the final state update at the current moment is completed, resulting in... :

[0163] ,

[0164] (3) Probabilistic Decoding and Causal Attribution Module for Future Trajectories (PTD-CA)

[0165] The responsibility of this module is to use the deep spatiotemporal representation generated by STEP-Net to perform future-oriented, uncertain trajectory predictions and to provide high-precision, reliable causal explanations for the prediction results.

[0166] 1) Probabilistic trajectory decoding based on attention mechanism

[0167] In order to generate the ecological state evolution trajectory for multiple future time steps and quantify its uncertainty, this step employs an attention-based autoregressive decoder.

[0168] When predicting the k-th future step, in order to draw upon complex historical information... The decoder dynamically extracts the most relevant context, first based on the state of the previous time step. , with each node in the historical state sequence The unnormalized alignment score is calculated using a small alignment network. :

[0169] ,

[0170] Where T represents the current time; k represents the index of the predicted future time step; This represents the state vector of the nth node. This represents the hidden state of the internal loop unit of the decoder at time T + k - 1; Learnable parameters (weight vector and weight matrix) for the attention alignment network. This represents the alignment score between the current state of the decoder and the state of historical node n.

[0171] The alignment scores are normalized using the Softmax function to obtain the attention weights. This weight determines which parts of the historical state sequence should be "focused" on in the current prediction step:

[0172] ,

[0173] in, Attention weights represent the importance of the historical state of the nth node when predicting the kth step.

[0174] Based on the calculated attention weights, the historical state sequence is weighted and summed to generate a context vector containing the most critical historical information. :

[0175] ,

[0176] in, This is the generated context vector.

[0177] The context vector is concatenated with the current state of the decoder, and then passed through an output linear layer and a softmax function to generate the probability distribution of the current prediction step across all C ecological categories. :

[0178] ,

[0179] in, This is the final output linear layer of the decoder; This represents the probability distribution matrix of the final output at the k-th future step.

[0180] 2) Causal Attribution Based on Deep Taylor Decomposition

[0181] To open the "black box" of the model and trace a high-risk prediction back to its initial input features, this step employs the Hierarchical Relevance Propagation (LRP) algorithm based on deep Taylor decomposition. rule.

[0182] Define the contribution of a neuron The main part and negative part This represents both promoting and inhibiting effects:

[0183] ,

[0184] in, Index representing a neuron; Indicates the first The activation value of layer neuron j; Indicates the connection of the first Layer neuron j and the first The weights of layer neuron k; This represents the contribution of neuron j to neuron k; The function means to keep only the positive values ​​and set all negative values ​​to 0; This means that only negative values ​​are retained, and all positive values ​​are changed to 0.

[0185] according to The rules will be the first Layer correlation score Backpropagation is performed based on positive and negative contributions, respectively, up to the [number]th [unit]. Neurons in the layer. First, calculate the forward propagation component. and negative propagation component :

[0186] ,

[0187] in, This is the hierarchical index of the neural network; Index representing a neuron; yes The correlation score of layer neurons; Represents the positive and negative correlation components that propagate from neuron k to neuron j; These are hyperparameters of the LRP rule, used to control the weighting of positive and negative contributions. They are typically set to... ; This is a very small, stable term used to prevent the denominator from being zero.

[0188] Finally, the total relevance score of neuron j is obtained by summing all the positive and negative relevance components from the upper node k. :

[0189] ,

[0190] By applying this correlation propagation rule backward layer by layer, starting from the final output layer of the model, the predictive correlation of the top layer can be completely and losslessly distributed to the neurons at the bottom layer, which represent each pixel of the original input image.

[0191] Experimental verification

[0192] To systematically verify the performance advantages of the method of this invention in real-world marine ecological prediction tasks, we constructed a multimodal, long-term remote sensing dataset based on the Bohai Sea in China. The dataset covers observational data from major eutrophication-prone areas of the Bohai Sea from 2018 to 2023. The data consists of three types of synchronous or quasi-synchronous remote sensing data sources: ① Optical mode: Sentinel-2 L2A multispectral imagery, used to extract proxy variables such as water color index and chlorophyll a concentration. ② SAR mode: Sentinel-1 GRD data, used to extract physical information such as sea surface wind field and roughness. ③ SST mode: MODIS daily sea surface temperature product. The data samples are arranged in a 7-day observation sequence to predict the ecological status for the next 1-3 days. We invited marine ecology experts to annotate the data and defined four ecological states: 0 - normal water body; 1 - slightly eutrophic; 2 - moderate algal bloom (visible but scattered algal bloom patches); 3 - severe algal bloom (large-scale, high-density algal bloom coverage). A total of 1,200 valid time-series samples were constructed, of which 900 were used for model training and 300 were used for performance testing.

[0193] Comparison Methods: To comprehensively evaluate the performance of the method in this invention and highlight its innovation in fusion, spatiotemporal modeling, and overall architecture, we set up the following four representative comparison methods: ① U-Net (Optical Modality Only): A classic image segmentation network that uses only optical images to perform pixel-level prediction of algal bloom regions at the next time step, representing a single-modality, non-temporally modeled baseline method. ② ConvLSTM (Concatenation Fusion): A mainstream spatiotemporal sequence prediction model where we simply concatenate the three-modality data at the input end before processing. This method can capture spatiotemporal information but lacks explicit spatial topology modeling, and the fusion method is simple. ③ GCN-GRU (Concatenation Fusion): A relatively advanced graph spatiotemporal network baseline that constructs a regional graph structure and models spatial and temporal dependencies through GCN and standard GRU respectively. This method also uses simple concatenation fusion, and its recurrent units do not possess the mechanistic design of this invention. ④ ST-Transformer (Attention Fusion): An advanced spatiotemporal prediction model based on Transformer that uses an attention mechanism to handle spatiotemporal dependencies. We implemented a version that uses simple global attention weighting for multimodal fusion to compare with the refined hierarchical fusion performance of the HCAT module in this invention. ⑤ The method of this invention: namely, the end-to-end prediction model proposed in this paper that fully integrates HCAT, STEP-Net (with built-in ESU), and PTD-CA modules.

[0194] All methods were evaluated under the same training and test sets. Evaluation metrics included: ① Prediction Accuracy (ACC): The overall accuracy of predicting the ecological state category of all nodes for the next 24 hours. ② F1 Score: A comprehensive evaluation of the model's precision and recall under imbalanced categories, with particular focus on its ability to identify algal blooms (a minority category). ③ False Negative Rate (MR): For the "severe algal bloom" category, the proportion of times the model incorrectly predicts it as another category; a key indicator of the reliability of the early warning system. ④ False Alarm Rate (FAR): The proportion of times the model incorrectly predicts "normal water bodies" as any anomalous category (levels 1, 2, and 3). ⑤ Attribution Consistency (Att-Consist): A blind review by three domain experts of the attribution results for 100 early warning events, evaluating the consistency between the dominant factors output by the model and the experts' judgments based on oceanographic knowledge.

[0195] Table 1. Comparison of data from different methods under five major indicators.

[0196] Method Name Fusion method Spatiotemporal modeling ACC F1-Score MR (Severe Algal Bloom) FAR (Frequency Availability) Att-Consist U-Net single mode Spatial CNN 74.8% 0.69 26.5% 16.3% N / A ConvLSTM Simple splicing CNN+LSTM 82.1% 0.78 17.2% 11.8% N / A GCN-GRU Simple splicing GCN+GRU 87.5% 0.86 12.4% 9.5% 65.8% ST-Transformer Global attention Transformer 89.2% 0.88 9.8% 8.1% 71.3% Method of the present invention Hierarchical Attention (HCAT) GCN+ESTU(STEP-Net) 93.2% 0.92 5.8% 4.6% 92.3%

[0197] Since U-Net is a pure image segmentation prediction network, its design goal is to output a category label for each pixel (e.g., whether it is algal bloom). The model itself does not have built-in modules for generating structured causal explanations or analyzing the contribution of different ecological factors. Although ConvLSTM is a spatiotemporal prediction model, its core architecture relies on predicting future states through convolutional recurrent units. It also lacks a specially designed attribution mechanism that can trace back and quantify the importance of each input feature (such as optical, SAR, SST). Therefore, neither of these methods can produce an interpretable attribution report as defined in this invention. This metric is not suitable for evaluating them, hence the "N / A" status on the "Attribution Consistency (Att-Consist)" metric.

[0198] The experimental results are shown in Table 1. Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6As shown, the systematic advantages of the method of this invention are clearly demonstrated. U-Net, as a single-modal static model, has the lowest performance across all metrics, with a false negative rate of 26.5%, which is unacceptable in practical early warning applications. This highlights the necessity of multimodal fusion and temporal modeling. Although ConvLSTM introduces temporal and multimodal information, its performance improvement is limited due to the lack of explicit modeling of spatial topological relationships and its coarse fusion method. GCN-GRU and ST-Transformer, as more advanced baselines, show significant performance improvements, proving the effectiveness of spatial relationships and attention mechanisms. However, GCN-GRU's simple splicing fusion prevents it from fully utilizing intermodal collaborative information; while ST-Transformer employs attention, its global fusion method is less effective than the refined hierarchical fusion of the HCAT module in this invention, and their shared standard cyclic units or temporal mechanisms are less accurate than the ESU units in STEP-Net in simulating ecological processes. Therefore, there is still a significant gap in the crucial false negative and false positive rates.

[0199] The five indicators are evaluated using different methods and cannot be directly displayed on the same radar chart. This invention retains the positive indicators unchanged. For the negative indicators, this invention applies a "1 - value" reverse processing, uniformly converting all indicators into a "higher value, better" format for display on the same chart. To prevent normalized values ​​from being 0 (or very close to 0), when multiple such indicators appear, the method shrinks back to the center point in multiple directions on the radar chart, thus presenting a "straight line" or "point". To avoid this problem, this invention introduces a "biased offset ε" mechanism, forcibly compressing the normalized values ​​to the range [ε, 1 - ε], ensuring that all methods have at least some visual display space on the chart.

[0200] In comparison, the method of this invention achieved the best performance across all five evaluation metrics. Its prediction accuracy (93.2%) and F1 score (0.92) were significantly superior. Crucially, the false negative rate (MR) for severe algal blooms was reduced to only 5.8%, and the false positive rate (FAR) was lowered to 4.6%, demonstrating that the early warning system of this invention possesses both high sensitivity and high reliability. Furthermore, the attribution consistency of up to 92.3% fully demonstrates the powerful capabilities and scientific value of the explanatory module of this invention, far exceeding other models with explanatory potential.

[0201] In summary, the experimental results are highly consistent with the theoretical analysis, fully demonstrating the advanced nature, effectiveness, and practicality of the method of this invention in the task of predicting marine ecological dynamics.

[0202] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping.

[0203] A terminal device includes a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping.

[0204] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping, characterized in that, include: Acquire remote sensing image data; The acquired remote sensing image data is used to perform deep fusion of multimodal remote sensing features (HCAT), which includes multimodal feature map extraction and spatial embedding, local cross-modal feature interaction and alignment, and global context aggregation and node embedding generation. A spatiotemporal ecological process inference network STEP-Net based on coupled spatial diffusion and temporal inertia is constructed, including multi-order spatial diffusion field modeling and ecological state transition unit (ESTU) construction. Using the deep spatiotemporal representations generated by STEP-Net, we can perform future-oriented, uncertain trajectory prediction and causal attribution based on deep Taylor decomposition. The multi-order spatial diffusion field modeling includes constructing a multi-order spatial diffusion field for aggregating multi-order neighborhood information by stacking multi-layer graph convolutional networks (GCNs). First, the output of the HCAT module... Left multiplication by normalized adjacency matrix Preliminary neighborhood aggregation features were obtained. : ,in, This is the normalized geographical adjacency matrix; The node feature matrix is ​​input from the HCAT module. This is a first-order neighborhood aggregation feature matrix; then, linear transformations and nonlinear activations are applied to the aggregated features to generate a first-order diffusion field. : Where ReLU is the activation function; It is the learnable weight matrix of the first layer GCN; This represents the first-order spatial diffusion field; finally, using the first-order diffusion field as input, second-order neighborhood information aggregation is performed first: ,in, This is the normalized geographical adjacency matrix; The neighborhood aggregation feature matrix is ​​a second-order matrix; then, a linear transformation and nonlinear activation are performed to generate a second-order diffusion field. : Where ReLU is the activation function; It is the learnable weight matrix of the second-layer GCN; Represents a second-order spatial diffusion field; to dynamically integrate influences from different spatial distances, a set of learnable weights is used. By weighted and fused diffusion fields of different orders, the final integrated spatial diffusion field is obtained. : , in, For learnable fusion weight parameters ; This is the final spatial diffusion field matrix that integrates the effects of multiple neighborhoods; The construction of the ecological state transition unit involves using the three basic gating mechanisms calculated by ESU to independently adjust the intensity of external forcing, spatial diffusion, and temporal inertia, thereby generating a candidate state component dominated by external forcing for each. And a candidate state component dominated by "spatial diffusion". : , in, These are candidate components driven by external forcing and candidate components driven by spatial diffusion, respectively. This represents the tanh activation function; This represents all learnable weight matrices and bias vectors within the ESU unit; to dynamically balance the influence of driving forces, a novel source-dominated gate is introduced using the ESU. Based on current new observations and spatial diffusion trends, the sources upon which state changes depend are dynamically determined, expressed as: ,in, This represents all learnable weight matrices and bias vectors within the ESU unit, where W-type matrices apply to the current input and U-type matrices apply to the loop state. This represents the Sigmoid activation function; The gating vector represents the dominant gate from which the source originates; Finally, by utilizing all the gating mechanisms, the candidate components are weighted and fused to form a final candidate state that is more physically interpretable. : , in, These are, respectively, candidate components driven by external forcing, candidate components driven by spatial diffusion, and the final fusion candidate state; Represents element-wise multiplication; utilizes time inertia gates. The historical state of the previous moment and the currently generated candidate state By performing a weighted combination, the final state update at the current moment is completed, resulting in... : .

2. The method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping according to claim 1, characterized in that, The multimodal feature map extraction and spatial embedding includes using a convolutional neural network to map remote sensing image data into feature maps that preserve spatial structure. , is represented as: Where mod is the modal identifier. The original input image for the modality; This represents the feature extraction network for the corresponding modality; The output modal feature map is defined as follows: opt represents the optical remote sensing mode; sar represents the synthetic aperture radar mode; and sst represents the sea surface temperature mode. The extracted feature map is then divided into grids, and each feature block is flattened to form a one-dimensional vector sequence. Where P is the total number of pixel blocks after the spatial grid is divided; Indicates a splitting operation; Indicates the flattening operation; This represents the original feature vector after the p-th pixel block is flattened; p is the position index of the pixel block. Finally, Transformer position encoding is used to generate position encoded values ​​for each position index p and dimension index i: , Where i is the dimension index of the feature vector; This represents the dimension of the uniform feature vector processed internally by the model. This represents the position encoding vector at position p; the flattened feature block vector is then passed through a linear projection layer for dimensionality reduction, and the corresponding position encoding vector is added to obtain the final feature block vector carrying position information. : , Wherein, Linear(·): a linear projection layer; Let be the feature vector of the p-th pixel block carrying location information.

3. The method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping according to claim 2, characterized in that, The local cross-modal feature interaction and alignment includes employing a multi-head cross-attention mechanism to process the input feature block vector. Generate the required Q, K, V vectors using a linear projection matrix: Where i is the index of the attention head; Nh is the total number of attention heads; The feature vector of the p-th pixel block is the input. This indicates that the i-th attention head is used to generate a learnable linear projection matrix of query, key, and value. Let represent the query, key, and value vector generated by the i-th head. Then, for each head, the attention output is calculated independently using the scaled dot product attention formula: ,in, The key vector dimension for each attention head; Let T be the output vector of the i-th attention head, and T be the transpose of all... After all individual calculations are completed, the outputs are stitched together and then processed through a linear projection layer. The results are then fused to obtain the final output of the multi-head attention layer: , Concat(·) is a vector concatenation operation; This is the final output linear projection layer of the head attention module; The final aggregated output for multi-head attention; To integrate the original information and stabilize training, the output of the multi-head attention is added to the original query vector via a residual connection, and then layer normalization is performed. in, The intermediate feature vectors are then processed through residual connections and layer normalization. Finally, a nonlinear transformation is performed via a feedforward neural network, followed by further residual connections and layer normalization to complete the final update of the optical modal feature blocks, resulting in a composite of local information from other modalities. : Where FFN(·) is a feedforward neural network, To update the obtained optical modal feature blocks.

4. The method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping according to claim 3, characterized in that, The global context aggregation and node embedding generation includes inputting a locally aligned sequence of feature blocks into a multi-layer global self-attention Transformer encoder, enabling each pixel block to interact with other pixel blocks on depth information, thereby constructing a global context-aware representation. , is represented as: ,in, The input is a sequence of locally aligned feature blocks; Standard Transformer encoder for global context modeling; This is the final representation of the p-th pixel block, which includes global context information. To generate the embedding vector representing the entire region, an attention-based pooling mechanism is used, and a single-layer MLP activated by tanh is used to compute an unnormalized importance score for each pixel block. : , in, Let be the unnormalized importance score of the p-th pixel block; These are the learnable parameters for the attention pooling network; The activation function is used; these scores are converted into attention weights that sum to 1 using the Softmax function. : , where j is the index used for summation over all pixel blocks; Let be the attention weight of the p-th pixel block in the weighted pooling; based on the learned weights, perform a weighted summation on the final representations of all pixel blocks to obtain the final multimodal ecosystem embedding vector belonging to the n-th graph node at time t. : Where n is the index of the graph node, This represents the final multimodal ecological embedding vector belonging to the nth graph node at time t.

5. The method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping according to claim 4, characterized in that, The deep spatiotemporal representation generated by STEP-Net is used for future-oriented, uncertain trajectory prediction. This includes employing an attention-based autoregressive decoder to predict the k-th future step from complex historical information. The autoregressive decoder dynamically extracts the most relevant context, first based on the state of the previous time step. , with each node in the historical state sequence Unnormalized alignment scores are calculated using a small alignment network. : Where T represents the current time; k represents the index of the predicted future time step; This represents the state vector of the nth node. This represents the hidden state of the internal loop unit of the decoder at time T + k -1; Learnable parameters for the attention alignment network; This represents the alignment score between the current state of the decoder and the state of historical node n.

6. The method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping according to claim 5, characterized in that, The method of using the deep spatiotemporal representation generated by STEP-Net for future-oriented, uncertain trajectory prediction also includes normalizing the alignment score using the Softmax function to obtain attention weights. The portion of the historical state sequence to focus on in the current prediction step is determined using attention weights, and is expressed as: ,in, Attention weights represent the importance of the historical state of the nth node when predicting the kth step. Based on the calculated attention weights, the historical state sequence is weighted and summed to generate a context vector containing the most critical historical information. : ,in, This generates a context vector; finally, the context vector is concatenated with the current state of the decoder, and then passed through an output linear layer and a softmax function to generate the probability distribution of the current prediction step across all C ecological categories. : , in, This is the final output linear layer of the decoder; This represents the probability distribution matrix for the final output at the k-th future step.

7. The method for predicting marine ecological dynamics based on multimodal remote sensing and spatiotemporal mapping according to claim 6, characterized in that, The causal attribution based on deep Taylor decomposition includes hierarchical correlation propagation (LRP) based on deep Taylor decomposition. Rules define the contribution of neurons. The main part and negative part : , in, Index representing a neuron; Indicates the first The activation value of layer neuron j; Indicates the connection of the first Layer neuron j and the first The weights of layer neuron k; This represents the contribution of neuron j to neuron k; The function indicates that only the positive values ​​are retained; This indicates that only the negative part should be retained; according to The rules will be the first Layer correlation score Backpropagation is performed based on positive and negative contributions, respectively, up to the [number]th [unit]. Neurons in the layer; first calculate the forward propagation component. and negative propagation component : , in, This is the hierarchical index of the neural network; Index representing a neuron; yes Correlation score of layer neurons; Represents the positive and negative correlation components that propagate from neuron k to neuron j; These are hyperparameters of the LRP rules; This is a very small stable term; finally, the total correlation score of neuron j is obtained by summing all the positive and negative correlation components from the upper node k. : By applying the correlation propagation rule layer by layer from the final output layer of the model, the predicted correlation of the top layer is ultimately distributed to the neurons of the bottom layer without loss.

8. A marine ecological dynamic prediction system based on multimodal remote sensing and spatiotemporal mapping, executing the marine ecological dynamic prediction method based on multimodal remote sensing and spatiotemporal mapping as described in claim 1, characterized in that, include: The data acquisition module is configured to acquire remote sensing image data; The feature fusion module is configured to perform deep fusion of multimodal remote sensing features based on the acquired remote sensing image data, including multimodal feature map extraction and spatial embedding, local cross-modal feature interaction and alignment, and global context aggregation and node embedding generation. The model building module is configured to construct the STEP-Net spatiotemporal ecological process inference network based on coupled spatial diffusion and temporal inertia, including multi-order spatial diffusion field modeling and ecological state transition unit construction. The prediction module is configured to perform future-oriented, uncertain trajectory prediction and causal attribution based on deep Taylor decomposition using the deep spatiotemporal representation generated by STEP-Net.

Citation Information

Patent Citations

  • Thermal runaway prevention and heat dissipation optimization method and system based on capacitor module

    CN120180071A

  • User behavior prediction system and method based on multi-modal data fusion

    CN120832498A