Ocean red tide anomaly detection method and system fusing multi-source remote sensing and graph neural network

By integrating multi-source remote sensing with graph neural networks, the problems of insufficient multi-source data fusion and spatiotemporal dynamic feature extraction in red tide detection were solved, high-precision red tide anomaly detection and trend analysis were achieved, and the level of intelligent response to red tide disasters was improved.

CN120656076AActive Publication Date: 2025-09-16SHANDONG MARINE RESOURCE AND ENVIRONMENT RESEARCH INSTITUTE (SHANDONG MARINE ENVIRONMENTAL MONITORING CENTER SHANDONG AQUATIC PRODUCTS QUALITY INSPECTION CENTER)

Patent Information

Application Number
CN202510824421.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-16
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing red tide detection methods have weak multi-source data fusion capabilities and insufficient spatiotemporal dynamic feature extraction, making it difficult to achieve high-precision red tide warnings. They also lack the ability to identify atypical samples and sudden abnormal events, resulting in prediction lags and misjudgments.

Method used

By integrating multi-source remote sensing with graph neural networks, remote sensing image data, drone image data, and monitoring point data are acquired to perform data preprocessing, feature extraction, and fusion. A spatiotemporal graph structure is constructed, and a cross-modal comparative self-supervised learning mechanism is used for consistent representation learning. Heterogeneous feature fusion and joint representation are then performed to ultimately perform anomaly detection and early warning.

Benefits of technology

It has significantly improved the meticulousness and global perception capabilities of red tide characteristic modeling, enhanced the understanding of the nonlinear coupling relationship between red tide inducing factors, achieved high-precision red tide event detection and trend analysis, and improved the intelligence and foresight of red tide disaster response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656076A_ABST
    Figure CN120656076A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of red tide anomaly detection, in particular to an ocean red tide anomaly detection method and system fusing multi-source remote sensing and a graph neural network. The method comprises the following steps: acquiring remote sensing image data, unmanned aerial vehicle image data and monitoring data of a monitoring point; performing data preprocessing on the acquired remote sensing image data and unmanned aerial vehicle image data; performing feature extraction and feature fusion on the remote sensing image and the unmanned aerial vehicle image to obtain remote sensing feature data; constructing a space-time diagram structure based on the monitoring data of the monitoring points to obtain diagram structure data; based on a cross-modal comparison self-supervised learning mechanism, carrying out consistency representation learning on a remote sensing feature mode and a graph structure feature mode; by introducing multi-source heterogeneous data and fusing a graph neural network modeling means, the limitation of a single data driving method in the aspects of coarse red tide recognition granularity, low space-time precision and the like is effectively broken through, and the meticulous property and global perception ability of red tide feature modeling are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of red tide anomaly detection, and in particular to a method and system for detecting marine red tide anomaly by integrating multi-source remote sensing and graph neural networks. Background Art

[0002] With global climate change and increasing human activity in coastal areas, the marine ecosystem faces increasingly severe challenges. The frequent occurrence of harmful algal blooms, such as red tides, has attracted widespread attention. Red tides, a typical sudden, regional, and rapidly evolving marine disaster, often cause reduced marine fishery production, damage coastal tourism, and disrupt aquatic ecosystems. In severe cases, they can even endanger human health and the safety of coastal infrastructure. To achieve early detection, dynamic tracking, and accurate early warning of marine red tides, it is urgent to build an intelligent anomaly detection technology system with high temporal and spatial resolution and strong environmental adaptability.

[0003] Existing red tide monitoring and detection methods fall into three main categories: First, traditional monitoring methods based on water quality factors rely on physical and chemical parameters such as nutrients, chlorophyll, and temperature collected at fixed monitoring stations. These methods use empirical formulas or threshold rules to determine red tide risk. While these methods have a certain scientific basis, their spatial coverage is limited, making them inadequate for dynamic monitoring of wide-area oceans. Second, image recognition methods based on remote sensing images utilize visible or infrared bands to invert ocean color anomalies for macroscopic identification. These methods offer advantages over wide areas and high frequencies, but suffer from cloud cover, insufficient spatial resolution, and high target heterogeneity. Third, data-driven machine learning methods rely on historical images or monitoring data to train classifiers or predictive models. While these methods can improve detection efficiency, they often rely on a large number of labeled samples and lack the ability to model spatiotemporal dynamic features, making it difficult to capture the complex causal relationships and cross-modal feature interactions in red tide evolution. In practical applications, these methods generally face challenges such as prediction lag, regional misjudgments, and difficulty integrating across scales, making them ineffective in supporting intelligent monitoring and scientific decision-making for the marine ecosystem.

[0004] Current red tide detection methods mainly rely on single remote sensing image processing, threshold rule-based warning models, or static classification models trained on historical samples. They have the following prominent defects: (1) weak multi-source data fusion capabilities, making it difficult to effectively integrate remote sensing images, drone images, and red tide prediction factors such as nutrients and temperature, resulting in one-sided information utilization, limited monitoring sensitivity, and limited spatial coverage; (2) lack of modeling mechanism for spatiotemporal evolution laws. Most existing methods are static analysis, which makes it difficult to capture the occurrence and development trends of red tides, resulting in delayed predictions and insufficient accuracy; (3) the model's ability to identify atypical samples and sudden abnormal events is insufficient. When there are insufficient samples or obvious scene changes, it is easy to make misjudgments or omissions, making it difficult to support real-time and stable risk warning needs. Therefore, it is urgent to propose an intelligent red tide detection system that integrates multi-source remote sensing images and red tide factor data, and has spatiotemporal modeling and graph structure reasoning capabilities. By constructing a spatiotemporal heterogeneous graph and introducing graph neural networks for high-order information extraction, accurate detection, trend analysis, and explainable warning of red tide events can be achieved, thereby improving the system's comprehensive perception and intelligent response capabilities. Summary of the Invention

[0005] In order to solve the problems of difficulty in multi-source data fusion, insufficient spatiotemporal dynamic feature extraction capabilities, and lack of cross-modal consistency modeling in the process of marine red tide anomaly detection, the present invention provides a marine red tide anomaly detection method and system that integrates multi-source remote sensing and graph neural networks.

[0006] In the first aspect, the present invention provides a method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks, which adopts the following technical solutions: A method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks, including: Acquire remote sensing image data, drone image data, and monitoring data from monitoring points; Perform data preprocessing on the acquired remote sensing image data and UAV image data; Perform feature extraction and feature fusion on remote sensing images and UAV images to obtain remote sensing feature data; Construct a spatiotemporal graph structure based on the monitoring data of the monitoring points to obtain graph structure data; Based on the cross-modal contrast self-supervised learning mechanism, the consistent representation learning of remote sensing feature modalities and graph structure feature modalities is carried out; Heterogeneous feature fusion and joint representation based on consistency learning; Anomaly detection and early warning are performed based on the fusion characterization results.

[0007] Furthermore, the acquired remote sensing image data and drone image data are preprocessed, including firstly performing physical consistency and geometric consistency processing on the remote sensing image and the drone image, and uniformly standardizing all factor data; then performing spatiotemporal alignment and missing data completion, including time alignment of the image data and the factor data according to the timestamp; for the missing monitoring point data caused by the buoy being offline, spatial weighted interpolation is used to complete the data; a time series-based interpolation and image completion method is introduced, and an image reconstruction method guided by optical flow is used to realize content prediction and completion of the frame obscured by fog. Assume that the image frame of the drone under continuous time acquisition is I t-1 with I t+1 , the goal is to reconstruct the missing intermediate image frame I t , first calculate the forward optical flow F between the two frames t-1 →t+1 and reverse optical flow F t+1 →t-1, the reconstruction value of a pixel position in the intermediate frame is estimated by using the bidirectional optical flow and the time weight α, which is expressed as: , Among them, (x, y) represents a pixel position in the image, (u, v) represents the motion vector of the pixel from time point a to b, and F a→b Represents the optical flow field from frame a to frame b, u1 and v1 represent the optical flow field from I t-1 to I t+1 The optical flow of u2 and v2 from I t+1 to I t-1 Optical flow.

[0008] Furthermore, the feature extraction and feature fusion of the remote sensing image and the UAV image include setting the UAV image sequence of length T as { X u} u=1 T , each frame image X u To represent multispectral image data at time t, a lightweight convolutional module is first used to extract local perceptual features. Subsequently, the self-attention mechanism in the Transformer architecture is introduced to model the dynamic dependencies between image sequences, thereby obtaining the behavioral patterns of key areas in the temporal dimension and ultimately obtaining a sequence-level dynamic embedding representation. In the spatial high-resolution feature extraction branch, a single remote sensing image is input, and multi-level spatial semantic features are extracted through a pre-trained multi-scale residual network ResNet-50 + ASPP. At the same time, a spatial attention mechanism is introduced to recalibrate the feature distribution to highlight potential abnormal areas in the image.

[0009] Furthermore, the spatiotemporal graph structure is constructed based on the monitoring data of the monitoring points, including constructing a dynamic adjacency graph based on a sliding time window, and at each time step t, constructing an adjacency matrix based on the historical sequence {t-w+1,…,t} with a window length of w using feature similarity A t , capturing the dynamic connection strength between nodes at the current moment; in view of the strong volatility of marine environmental variables at different time scales, a sliding window decomposition mechanism is introduced to divide the long-term time series data into multiple sliding short sequences to extract short-term dynamic features and avoid long-term stable trends covering up sudden anomalies. At each time step t, the node feature tensor under the current sliding window is constructed S t ∈ R w×N×F , and input the graph convolution operation; finally, to adapt to the corresponding differences of nodes, a node-level personality parameter vector is constructed θ i , introduces a personalized weight control mechanism, expressed as: , in, represents the final graph convolution output of node i at time t, represents the trainable personality weight of node i, reflecting its preference for factor response, stands for element-wise multiplication, represents the input features of node j at time t, Representative Node i In time t The set of adjacent nodes.

[0010] Furthermore, the consistent representation learning of remote sensing feature modalities and graph structure feature modalities based on cross-modal contrastive self-supervised learning mechanism includes projecting the two modalities into a shared semantic space to achieve mapping consistency of the two modalities in the global semantic space, and using the InfoNCE loss based on contrastive learning to construct a global modality alignment target, which is expressed as: , Among them, sim(a,b) represents cosine similarity, is the temperature coefficient, which is used to adjust the smoothness of the distribution. Represents a set of negative sample graph structures containing different time steps, Represents the feature vector generated by the spatiotemporal graph neural network, Represents the feature vector produced by the dual-branch encoder.

[0011] Furthermore, the self-supervised learning mechanism based on cross-modal contrast is used to learn the consistency representation of remote sensing feature modalities and graph structure feature modalities, and also includes the introduction of a local semantic alignment mechanism to enhance the semantic alignment capability between modalities at a fine-grained level. Specifically, the remote sensing image is divided into K fixed spatial window areas, and each area is extracted through a convolutional encoder to extract a local embedded representation. At the same time, each node in the graph structure is embedded as , the goal is to enable each local area of ​​the image to find the semantically closest site node in the graph structure and build a cross-modal local alignment pair. The alignment loss is expressed as: , in, v represents a node in the graph structure, and sim(a,b) represents the similarity function.

[0012] Furthermore, the heterogeneous feature fusion and joint representation based on consistency learning includes introducing a modal attention mechanism to adaptively adjust the fusion weights of different modal features according to task relevance, wherein the image modal feature is set to h I ∈R d , the graph structure modal feature is h G ∈R d , the fusion weight is calculated using the shared attention network, which is expressed as: , Among them, W is the learnable attention parameter, h I represents the image modality feature, h G Represents the graph structure modal features. The final fusion representation is: .in, and Represents the fusion weight.

[0013] Furthermore, the heterogeneous feature fusion and joint representation based on consistency learning also includes using a semantic channel attention mechanism to enhance the semantic dimension highly related to red tide in the fused feature, and performing weighted adjustment on the fused joint representation in the channel dimension. h fused , the Squeeze-and-Excitation mechanism is used to construct the channel attention weight, and the semantic dimension is added to the fusion feature to obtain the feature h recon , the features h recon Input into the Transformer-based spatiotemporal modeling module to construct a multi-time multi-modal joint representation sequence. Represents the fused feature sequence of the past T time steps, which is input to the temporal encoder (Temporal Transformer) as: , in, It represents the fused dynamic joint spatiotemporal feature sequence, which serves as the input of the downstream red tide warning and spatial reasoning module.

[0014] Furthermore, the anomaly detection and early warning based on the fusion characterization result includes adopting an unsupervised anomaly detection method based on an autoencoder, setting the joint feature of the input ∈R d , the autoencoder consists of an encoder (•) and decoder Ψ(•), which models the normal state by minimizing the reconstruction error and performs anomaly detection based on the set threshold to judge the reconstruction error; the time series prediction model is trained using the historical fusion feature sequence to predict the feature expression at the next moment or several moments in the future, and further predict the probability of red tide occurrence.

[0015] The second aspect is a marine red tide anomaly detection system that integrates multi-source remote sensing and graph neural networks, including: The data acquisition module is configured to acquire remote sensing image data, drone image data, and monitoring data of monitoring points; The preprocessing module is configured to perform data preprocessing on the acquired remote sensing image data and the UAV image data; The remote sensing feature module is configured to perform feature extraction and feature fusion on the remote sensing image and the UAV image to obtain remote sensing feature data; A graph structure module is configured to construct a spatiotemporal graph structure based on the monitoring data of the monitoring points to obtain graph structure data; The consistency module is configured to learn the consistent representation of remote sensing feature modalities and graph structure feature modalities based on a cross-modal comparative self-supervised learning mechanism; The joint module is configured to perform heterogeneous feature fusion and joint representation based on consistency learning; The early warning module is configured to perform anomaly detection and early warning based on the fusion characterization results.

[0016] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device to describe a method for detecting marine red tide anomalies that integrates multi-source remote sensing and graph neural networks.

[0017] In a fourth aspect, the present invention provides a terminal device comprising a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; and the computer-readable storage medium is used to store multiple instructions, wherein the instructions are suitable for being loaded and executed by the processor to implement the method for detecting marine red tide anomalies that integrates multi-source remote sensing and graph neural networks.

[0018] In summary, the present invention has the following beneficial technical effects: Compared with the existing technology, the marine red tide spatiotemporal modeling and early warning system of the present invention, which constructs a graph neural network based on remote sensing images, drone images and red tide predictor factor data, has the following beneficial effects: by introducing multi-source heterogeneous data (including remote sensing images, drone monitoring images and various environmental factors) and integrating graph neural network modeling methods, it effectively breaks through the limitations of the single data-driven method in terms of coarse granularity and low spatiotemporal accuracy in red tide identification, and significantly improves the meticulousness and global perception ability of red tide feature modeling.

[0019] By constructing a cross-modal comparative self-supervised learning mechanism, unified semantic alignment between image modality and graph structure modality is achieved, enhancing the model's ability to understand the potential correlation patterns of multimodal data; combining the attention mechanism and the joint representation module, the nonlinear coupling relationship between red tide inducing factors is further explored, improving the interpretability and reliability of the red tide triggering mechanism modeling; on this basis, the introduction of anomaly detection and time series prediction modules can perform high-precision dynamic deduction of the occurrence and evolution trends of marine red tides, and realize visual early warning, effectively improving the intelligence and foresight of red tide disaster response. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of a method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks according to Example 1 of the present invention; Figure 2 1 is a schematic diagram comparing the red tide anomaly detection models ACC and F1 according to Example 1 of the present invention; Figure 3 1 is a schematic diagram comparing radars of different models for red tide anomaly detection according to Example 1 of the present invention; Figure 4 1 is a thermal diagram of the spatial distribution of red tide anomaly detection according to Example 1 of the present invention; Figure 5 It is a schematic diagram of the time series prediction curve of red tide anomaly detection concentration in Example 1 of the present invention. DETAILED DESCRIPTION

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] Example 1 Reference Figure 1In this embodiment, a method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks includes: A method and system for detecting marine red tide anomalies that integrates multi-source remote sensing imagery, drone imagery, and red tide prediction factors. The overall system design consists of six core modules: an image and monitoring data preprocessing module, which standardizes and aligns remote sensing imagery, drone imagery, and factor data; a dual-branch feature encoding module, which extracts dynamic and spatial features from time series images and high-resolution images, respectively; a spatiotemporal graph structure modeling module, which uses the monitoring area as a graph node to construct a spatiotemporal graph structure that integrates factors such as water temperature, salinity, and pH; a cross-modal comparative self-supervised learning module, which learns consistent representations of image features and graph structure features; an attention fusion and joint representation module, which weightedly fuses multimodal features to enhance the expressive power of red tide-induced information; and an anomaly detection and prediction warning module, which identifies and predicts red tide anomalies based on fused features and outputs visual warnings.

[0023] S1. Image and monitoring data preprocessing module, In the task of detecting marine red tide anomalies, the raw data comes from a variety of sources, mainly including satellite remote sensing images with high spatial resolution but slow temporal updates, drone patrol images with high temporal resolution but limited coverage, and red tide predictor data from monitoring buoys or sensors, such as water temperature (Temperature), salinity (Salinity), pH (pH) and chlorophyll concentration (Chlorophyll-a Concentration). The heterogeneity, inconsistent scales and potential missing data of these data seriously restrict the accuracy of subsequent feature extraction and modeling. Therefore, this module aims to standardize, spatially and temporally align and structure all data to form a standardized data set that can be used for graph modeling and deep learning input. The specific processing flow includes the following three parts: 1) Preprocessing of remote sensing images and drone images, addressing the physical and geometric consistency issues of remote sensing images and drone images. Remote sensing images are often affected by the atmosphere, clouds, solar altitude, etc., and require radiation correction to restore the true reflectivity of the ground objects. Convert the original digital value (DN value) of the remote sensing image to the radiation brightness received by the sensor. .

[0024] , Among them, DN represents the digital value recorded by the sensor, G represents the gain of the sensor calibration, BRepresents the sensor calibration offset. Furthermore, atmospheric correction aims to convert the reflectance of the surface-atmosphere system (i.e., the value seen in remote sensing images) into true surface reflectance, eliminating interference from atmospheric aerosols, molecular scattering, water vapor, and other factors. This improves the spectral consistency of images acquired at different times.

[0025] , in, represents the surface reflectivity, Represents the reflectivity of the top of the atmosphere obtained from remote sensing images. represents the path reflectivity, and Represents the transmittance on the solar incident path and the sensor observation path, S represents the atmospheric scattering term.

[0026] Due to the low sensor height and variable viewing angle, drone images have significant geometric distortion. Therefore, geometric distortion correction and orthorectification are required to ensure that the image can be aligned with the remote sensing image in a unified coordinate system. In terms of spatial alignment, to achieve high-precision registration between remote sensing and drone images, a feature point matching algorithm is used to extract and align key areas between images so that images of different modalities have a consistent spatial reference system. Feature points are extracted from remote sensing images and drone images respectively, and the scale-invariant feature transform (SIFT) algorithm is used to extract local image feature descriptors that are scale- and rotation-invariant. Let image X be the image of the drone image. r represents remote sensing images, X u ={X u1 , X u2 , ..., X uT} represents the drone image, and the extracted feature point sets are: , in, f ri , f uj Represents the first i , j feature points, each of which contains attributes such as position, scale, direction, and a descriptor vector d i , using Euclidean distance to calculate the similarity of the two feature vectors, obtaining an initial set of matching pairs. Due to issues such as perspective differences, occlusions, and repeated textures during feature matching, matching pairs may contain a large number of incorrect matching points. To address this, the RANSAC (Random Sample Consensus) algorithm is introduced to perform robustness screening and geometric transformation model estimation on the initial matching pairs, evaluating the matching consistency of each point: , Among them, x i Represents the feature points in the remote sensing image, x i ′ represents the corresponding point in the UAV image, H Represents the homography transformation matrix to be estimated, mapping the remote sensing image coordinates to the drone image, and ||·|| represents the Euclidean distance. Let the remote sensing image be X r , the drone time series image is X u ={X u1 , X u2 , ..., X uT The image is then resampled using bilinear interpolation to unify its spatial resolution, ensuring consistent input dimensions for downstream models. Finally, the image is normalized (0-1 normalization) and enhanced (contrast stretching) to improve feature representation.

[0027] 2) Standardize the time series of red tide factor data. Monitoring factor data is typically collected periodically by buoys, ocean sensors, and other equipment. With high temporal granularity and diverse dimensions, it serves as a crucial driver for red tide anomaly detection. These factors include water temperature (T), salinity (S), pH, and chlorophyll concentration (Chl). These factors vary widely in their ranges and trends. Directly inputting them into the model can easily bias the training process toward features with larger numerical scales. Therefore, all factor data must first be standardized.

[0028] Let the factor vector of the i-th monitoring point at time t be F(x i t ) = [T i t , S i t , pH i t , Chl i t ], using the Z-score standardization method, it is converted to a distribution with a mean of 0 and a variance of 1 using the following formula: , in, and are the mean and standard deviation of the i-th factor in all monitoring data, ensuring that the weight of each factor is at the same order of magnitude during the training phase to avoid gradient update offset.

[0029] In addition, in order to enhance the stationarity of the time series, differential operation can be introduced to weaken the interference of periodic fluctuations on the modeling results.

[0030] , Among them, x t represents the observation value of the original time series at time t, x t-1 Represents the observation value at the previous moment, y t Represents the differenced sequence value.

[0031] 3) Spatiotemporal alignment and missing data completion. Due to the complex observation environment and poor equipment stability, multi-source data often suffer from time inconsistencies and spatial omissions. To ensure the spatiotemporal alignment of the inputs of the graph neural network and the time series model, the following three steps are required: Time synchronization: align image data and factor data based on timestamps. For example, if remote sensing images are collected daily and factor data are recorded hourly, you can set a time window to extract the factor value closest to the image time or perform linear interpolation: , Spatial interpolation completion: For missing monitoring point data due to reasons such as buoy offline, spatial weighted interpolation is used to complete the data. i The missing value is , by its neighborhood The effective data is completed by distance weighting: , Image occlusion processing: In case of image information loss due to fog occlusion, the integrity of image data can be enhanced by using an interpolation algorithm based on image sequence (such as image completion based on Optical Flow). This paper introduces an interpolation and image completion method based on time series, and uses the image reconstruction technology guided by optical flow (Optical Flow-based FrameInterpolation) to achieve content prediction and completion of fog-occluded frames. Assume that the image frames collected by the drone in continuous time are I t-1 with I t+1 , the goal is to reconstruct the missing intermediate image frame I t First, calculate the forward optical flow F between the two frames t-1 →t+1 and reverse optical flow F t+1 →t-1, defined as follows: , Among them, (x, y) represents a pixel position in the image, (u, v) represents the motion vector (optical flow) of the pixel from time point a to b, and F a→b Represents the optical flow field from frame a to frame b. Through the bidirectional optical flow and the time weight α, the reconstruction value of a pixel position in the intermediate frame can be estimated: , Among them, u1 and v1 represent the t-1 to I t+1The optical flow of u2 and v2 from I t+1 to I t-1 Optical flow.

[0032] S2. Dual-branch feature encoding and spatiotemporal fusion module, In the task of detecting marine red tide anomalies, different image modalities contain different dimensions of information: remote sensing images have the advantage of high spatial resolution, capturing large-scale spatial structure and regional anomalies; while drone imagery has the advantage of high temporal resolution, reflecting dynamic changes within local areas in a timely manner. To fully exploit the complementary information contained in these two types of images, the system designed a dual-branch feature encoding and spatiotemporal fusion module, constructing a time series feature branch and a spatial texture feature branch respectively, and achieving spatiotemporal semantic fusion through a cross-scale attention mechanism.

[0033] 1) In the time series feature extraction branch, the drone image sequence with a length of T is set to { X u} u=1 T , each frame image X u Represents the multispectral image data at time t. In order to preserve the long-range dependencies and regional level change trajectories within the sequence, a lightweight convolution module is first used to extract local perception features: , in, f t ∈R d For the t Feature representation of frame images, θ is a network parameter. The self-attention mechanism in the Transformer architecture is then introduced to model the dynamic dependencies between image sequences to obtain the behavior pattern of key areas in the time dimension. Its attention weight is defined as follows: , Among them, Q,K,V∈R T×dk Represent query, key, and value matrices respectively, d k is the feature dimension, and Softmax ensures the normalization of attention weights. Finally, the sequence-level dynamic embedding representation H={ h t} t=1 T , as an encoding of the dynamic behavior of the region.

[0034] 2) In the spatial high-resolution feature extraction branch, the input is a single remote sensing image, and multi-level spatial semantic features are extracted through a pre-trained multi-scale residual network (ResNet-50 + ASPP): , in, g ∈R d is the spatial structure feature, The ASPP module is used to expand the receptive field to perceive regional context. In order to highlight potential abnormal areas in the image (such as the edges of red tide clusters or mutation areas), the spatial attention mechanism is introduced to recalibrate the feature distribution: , in, W 1, W 2 is the weight matrix of the attention module, b 1, b 2 is the bias term, σ is the Sigmoid function, and ⊙ represents element-wise multiplication to achieve explicit enhancement of salient areas.

[0035] 3) In the spatiotemporal fusion stage, in order to achieve cross-modal semantic alignment, the dynamic feature h t With static features g ′ is fused in a unified embedding space and input into a fully connected network after vector concatenation to achieve nonlinear remapping: , Among them, z t Represents the fused multimodal spatiotemporal representation. In order to further enhance the discriminative ability of fused features in prediction tasks, a gating mechanism or cross-attention module can be introduced to achieve semantic enhancement: , Among them, Gate represents the gating unit, which is used to adaptively adjust the importance weight of each modal feature in different scenarios.

[0036] S3. Spatiotemporal graph structure modeling module, During the spatiotemporal evolution of marine red tides, various monitoring data (such as water temperature, salinity, and pH) are typically collected continuously by multiple automated monitoring stations located in different sea areas. These monitoring stations are topologically connected in geographic space, and the collected data exhibit distinct time series characteristics. Therefore, to accurately characterize the dependency structure of various factors during spatial propagation and temporal evolution, constructing a spatiotemporal graph structure model is crucial. The core goal of this module is to use monitoring stations as graph nodes, integrate hydrological environmental factors to construct a spatial topological graph, and then combine the graph neural network (GNN) with the temporal modeling module to achieve high-dimensional dynamic modeling of the marine environment.

[0037] 1) Dynamic graph modeling, building a dynamic adjacency graph based on a sliding time window. To address the issue of monitoring the evolution of influence relationships between nodes over time, this module uses a sliding time window to construct a dynamic graph structure. Specifically, at each time step t, based on the historical sequence {t-w+1,…,t} with a window length of w, an adjacency matrix is ​​constructed using feature similarity. A t , capturing the dynamic connection strength between nodes at the current moment. The edge weight calculation formula is: , in, represents the feature matrix of node i over the past w time steps, where F is the factor dimension. ||·|| represents the L2 norm, which measures the Euclidean distance between node feature sequences. Representative Node i and j At the moment t The dynamic similarity edge weight of N Represents the total number of nodes in the graph. Dynamic adjacency matrix A t It reflects the time-varying structural dependencies between nodes and provides a structural basis for subsequent graph convolution operations.

[0038] 2) Sliding decomposition mechanism to model short-term time series dynamic changes. Considering the strong volatility of marine environmental variables at different time scales, this module introduces a sliding window decomposition mechanism to divide long-term time series data into multiple sliding short sequences to extract short-term dynamic features and avoid long-term stable trends masking sudden anomalies. At each time step t, the node feature tensor under the current sliding window is constructed. S t ∈ R w×N×F , and input the graph convolution operation: , Among them, H t ∈R N×d Represents the encoding output of the node at time t, and d is the output dimension. D t represent A t The degree matrix of A t ( i , j ) is the weighted sum of . W t The learnable weight matrix of the graph convolution layer, represents a nonlinear activation function (such as ReLU), S tThe feature sequence of the nodes within the sliding window is used as the input signal. The sliding mechanism can enhance the model's ability to perceive sudden changes (such as rapid increases in ocean factors) and seasonal changes, improving the sensitivity and robustness of the model.

[0039] 3) Personalized graph convolution modeling to adapt to node response differences. Different monitoring stations have significant differences in their response mechanisms to red tide factors due to their locations (such as nearshore, deep sea, estuary) or sensor deployment conditions. To model this heterogeneity, a node-level personalized parameter vector is designed. θ i , introducing a personalized weight control mechanism: , in, represents the final graph convolution output of node i at time t, represents the trainable personality weight of node i, reflecting its preference for factor response, stands for element-wise multiplication, represents the input features of node j at time t, Representative Node i In time t The set of adjacent nodes.

[0040] S4. Cross-modal contrastive self-supervised learning module, In the spatiotemporal modeling of red tides based on multi-source information fusion, remote sensing images and graph-structured data constructed based on ocean monitoring stations exhibit significant differences in data modality, perception mode, temporal resolution, and spatial density. This modal heterogeneity not only increases the difficulty of cross-modal learning, but also leads to problems of information redundancy and inconsistent representation. Traditional fusion strategies rely on supervised learning methods, which are limited by the scarcity of red tide samples and the difficulty of manual labeling, making it difficult to obtain sufficient training signals. Therefore, this module introduces a cross-modal comparative self-supervised learning mechanism to drive the learning of consistent representations between remote sensing image modalities and monitoring graph structure modalities in an unsupervised manner, thereby improving the robustness and generalization capabilities of downstream spatiotemporal prediction and reasoning.

[0041] 1) Global modal alignment. In red tide early warning systems, remote sensing image data typically contains large-scale marine environmental features, while graph-structured data, composed of multiple monitoring stations, reflects local numerical indicators such as water temperature, salinity, pH, and chlorophyll concentration. To achieve consistent mapping of the two modalities in the global semantic space, the two modalities are projected into a shared semantic space, and the global modal alignment objective is constructed using the InfoNCE loss based on contrastive learning: , Among them, sim(a,b) represents cosine similarity, is the temperature coefficient, which is used to adjust the smoothness of the distribution. Represents a set of negative sample graph structures containing different time steps, Represents the feature vector generated by the spatiotemporal graph neural network, Represents the feature vector produced by the dual-branch encoder.

[0042] 2) Red tide formation has significant spatial heterogeneity and often occurs in certain sea areas or near local sites. Therefore, relying solely on global embedding may mask the semantic response of key local areas. This module further introduces a local semantic alignment mechanism to enhance the semantic alignment capability between modalities at a fine-grained level. Specifically, the remote sensing image is divided into K fixed spatial window regions (e.g., 16×16 patches), and each region is extracted through a convolutional encoder to represent the local embedding. At the same time, each node in the graph structure is embedded as The goal is to enable each local area of ​​the image to find a semantically closest site node in the graph structure and construct a cross-modal local alignment pair. The alignment loss is as follows: , in, v represents a node in the graph structure, and sim(a,b) represents the similarity function.

[0043] This loss encourages the local spatial structure of the image to obtain the most similar semantic mapping in the graph node space, so that the local features remain consistent and enhance the model's perception and generalization capabilities of local red tide change areas.

[0044] 3) Region-aware negative sampling. In cross-modal contrastive learning, the quality of negative samples directly determines the effectiveness of the learning signal. If the negative samples are too similar to the positive samples, it will lead to blurred model learning objectives, slow convergence, and even gradient vanishing problems. To this end, a region-aware negative sampling strategy is designed. When constructing image negative samples and graph structure negative samples, the joint constraints of spatial distance and graph topological distance are considered to eliminate pseudo-negative samples with similar semantics. Assume that the positive samples come from image region p + , the corresponding graph structure node is v + , then the candidate set of negative sample areas is {p -}, the effective negative sample set is defined as: , Among them, ||p - -p + ||2 represents the Euclidean distance in the image pixel space, GraphDist(v + ,v -) represents the shortest path graph distance between sites, and δ and η represent the minimum spatial distance and minimum structural distance thresholds. Combining global modality alignment, local semantic alignment, and region-aware negative sampling, the final cross-modal contrastive self-supervision joint training objective is as follows: , Among them, λ1 and λ2 are weighted coefficients used to balance the representation alignment losses of different scales.

[0045] S5. Attention Fusion and Joint Representation Module, The formation and evolution of red tides are often driven by a combination of environmental factors, including sea surface temperature, salinity, nutrient concentration, light, wind, and hydrodynamic changes. Remote sensing imagery and monitoring site map data capture the morphology and triggering mechanisms of red tides from different perspectives. The previous module extracts deep features from each modality and learns cross-modal consistency. However, effectively integrating these heterogeneous features and performing joint representation remains crucial for red tide prediction accuracy and model generalization.

[0046] 1) Modal attention weighted fusion. Due to the dynamic differences in the ability of remote sensing images and monitoring graph structures to express red tide information, image information may be more sensitive at certain moments (such as obvious red tide-stained water bodies), while at other moments, graph structure data (such as changes in nutrients) may reflect the potential trend of red tide in advance. Therefore, the modal attention mechanism is first introduced to adaptively adjust the fusion weights of different modal features based on task relevance. Let the image modal feature be h I ∈R d , the graph structure modal feature is h G ∈R d , use the shared attention network to calculate the fusion weights: , Among them, W is the learnable attention parameter, h I represents the image modality feature, h G Represents the graph structure modal features. The final fusion representation is: , in, and Represents the fusion weight. This mechanism dynamically allocates the contribution ratio of modal information to the final fusion feature based on the expression strength of modal information at different time points, effectively improving the flexibility and discriminability of red tide cause representation.

[0047] 2) To further enhance the semantic dimensions in the fused features that are highly relevant to red tides (such as sudden rises in water temperature, high chlorophyll value areas, and abnormal reflection areas), a semantic channel attention mechanism is used to weight the fused joint representation in the channel dimension. This mechanism guides the model to focus on key semantic channels by learning the importance weights between feature channels. The Squeeze-and-Excitation (SE) mechanism is used to construct channel attention weights, and the semantic dimension is added to the fused features to obtain h recon : , in, W 1. W 2 is the weight of dimensionality reduction and dimensionality increase, represents the Sigmoid function, Represents element-by-element multiplication. Through this mechanism, the model can automatically suppress redundant channels irrelevant to red tide (such as cloud interference and shore background) and enhance the response strength of key channels, thereby improving overall semantic consistency and predictive discriminability.

[0048] 3) Construction of spatiotemporal joint representation: Red tide is a complex spatiotemporal evolution phenomenon, and its spatial expansion and temporal evolution process are continuous and coupled. It is difficult to form a complete prediction perspective by relying on a single frame or local information. Therefore, the fusion feature h recon It is further input into the Transformer-based spatiotemporal modeling module to construct a multi-time and multi-modal joint representation sequence. Represents the fused feature sequence of the past T time steps, which is input to the temporal encoder (TemporalTransformer): , in, This represents the fused dynamic joint spatiotemporal feature sequence, which serves as the input to the downstream red tide warning and spatial reasoning module. This joint feature not only integrates the two modalities of remote sensing and monitoring map structure, but also explicitly models their coupling relationship on the time axis, helping to form a complete evolutionary chain representation.

[0049] S6. Anomaly detection and prediction warning module, In the marine red tide monitoring system, to achieve real-time identification and trend prediction of abnormal red tide events, this module uses the multimodal joint features obtained by fusing the previous modules as input. Through deep learning and statistical anomaly detection algorithms, it builds a comprehensive early warning platform that integrates anomaly detection, trend prediction, and visual early warning. This module can not only capture sudden anomalies in the spatiotemporal distribution of red tides, but also predict the trend of possible red tide events in the future, providing a scientific basis and intuitive display for prevention and control decisions. It specifically includes the following three parts: 1) Anomaly detection, in fusion features Based on the anomaly detection module, the purpose is to identify whether there is a red tide anomaly at the current moment. To this end, an unsupervised anomaly detection method based on an autoencoder is adopted. Assume that the joint feature of the input ∈R d , the autoencoder consists of an encoder (•) and decoder Ψ(•), which models the normal state by minimizing the reconstruction error: , in, is the reconstructed feature. The reconstruction error is given in the form of mean square error (MSE): , The assumption of anomaly detection is that under normal conditions, the model can reconstruct the input features well, but when anomalies occur, the reconstruction error will increase significantly. Let the threshold ϵ be the preset upper limit of the reconstruction error. If , This method uses unsupervised learning autoencoders to capture the inherent patterns of the data, enabling the system to accurately identify abnormal phenomena even in the absence of labeled samples.

[0050] 2) Trend prediction: To achieve dynamic prediction of the evolution trend of red tide, this module further constructs a neural network model based on time series prediction, such as the long short-term memory network (LSTM) time series model. On this basis, the time series prediction model is trained using the historical fusion feature sequence to predict the feature expression at the next moment or several moments in the future, and further predict the probability of red tide occurrence. Assuming the prediction model is F pred (•), the fusion feature of the predicted future time T+1 is expressed as: , Next, the fused feature map output by the prediction model is converted into the probability of red tide anomaly: , Where σ(•) is a sigmoid function that maps the output to the interval [0, 1], representing the probability of a red tide event. When the probability exceeds the preset probability threshold γ, the occurrence of a red tide anomaly is predicted and future trend forecast information is output.

[0051] After anomaly detection and trend prediction are complete, the system needs to present the detection results to decision makers and users in an intuitive manner. The early warning visualization module visualizes information such as anomaly distribution, future change trends, and predicted probabilities in the form of charts and time series curves.

[0052] Experimental verification: To validate the effectiveness of this study's proposed spatiotemporal modeling and early warning system for marine red tides, constructed using remote sensing imagery, drone imagery, and red tide predictor data, a field simulation and multi-source data-driven experimental platform was constructed in a typical coastal area prone to red tides. The platform collected multimodal data, including remote sensing imagery, low-altitude high-definition drone imagery, and temperature, salinity, pH, and chlorophyll concentration from monitoring buoys. The total sample size reached 11,000 sets, covering the complete red tide evolution cycle (occurrence, development, and decline) and demonstrating rich spatial and temporal variations. To ensure the experiment's closeness to actual early warning needs, complex environmental variables such as typhoon interference, cloud cover, degraded remote sensing image quality, and missing data from some monitoring nodes were introduced during the design process to comprehensively test the model's robustness and generalization capabilities under uncertain scenarios.

[0053] The comparison method selected the current representative red tide detection and prediction models, including ConvLSTM based on the fusion of convolution and temporal modeling, ASTGCN combined with graph attention mechanism, Multi-Modal Fusion Net (MMFN) with multimodal input, and the proposed Cross-Modal Matching Transformer (CMMT). All models were compared under a unified data set partitioning (training set: validation set: test set = 6:2:2), consistent optimization strategy and number of training rounds to ensure evaluation fairness. The performance evaluation indicators include six indicators: accuracy (ACC), F1-Score, spatial positioning error, warning lead time, false alarm rate and model inference delay, and a comprehensive evaluation was conducted from three dimensions: detection accuracy, timeliness and deployment feasibility. The experimental results are as follows Figure 2 、 Figure 3 As shown in Table 1, the proposed method outperforms the comparison model in all indicators, verifying its significant advantages in multimodal modeling and red tide prediction tasks.

[0054] Table 1 Data comparison of different methods under six indicators Model Name ACC F1 Early warning lead time Positioning error False positive rate Inference latency Transformer 84.7% 82.3% 1.6 days 12.4km 9.8% 2.5min ASTGCN 87.1% 85.0% 2.1 days 10.7km 8.1% 3.2min MMFN 85.5% 83.9% 1.8 days 11.5km 8.9% 4.0min CMMT 88.4% 86.8% 2.4 days 9.2km 7.2% 4.6min ConvLSTM 82.2% 80.7% 1.5 days 13.1km 10.4% 3.0min Method of the present invention 91.3% 89.6% 3 days 6.3km 5.1% 3.8min from Figure 2 、 Figure 3 As can be seen from Table 1, traditional methods such as the Transformer, ASTGCN, MMFN, CMMT, and ConvLSTM all demonstrate some predictive ability in the red tide anomaly detection task, but still exhibit significant deficiencies in key performance dimensions. The Transformer has some advantages in global modeling, capable of modeling long-term dependencies in time series. However, its limited ability to model regional spatial details leads to large localization errors, with a false alarm rate as high as 9.8%. ASTGCN, by introducing a graph-structured modeling approach, effectively integrates temporal and spatial factors, outperforming the Transformer in terms of warning lead time and accuracy. However, its static graph structure struggles to adapt to the rapidly evolving dynamics of red tides. MMFN and CMMT improve predictive accuracy through multimodal fusion, achieving F1 scores of 83.9% and 86.8% respectively. However, their fusion approach is relatively shallow, lacking cross-modal alignment and semantic enhancement, resulting in a still-high false alarm rate and significantly increased inference latency. ConvLSTM has certain advantages in time series modeling, but due to the lack of explicit spatial structure expression, its prediction accuracy is low, as shown by the lowest ACC and F1 values ​​(82.2% and 80.7%).

[0055] In contrast, the method of the present invention is based on key technologies such as dual-branch feature encoding, cross-modal contrastive learning self-supervision, attention fusion and graph structure modeling. It fully integrates multi-source information such as remote sensing images, drone images and red tide inducing factors, and achieves a higher level of semantic consistency and structural sensitivity modeling capabilities. In the experiment, this method outperformed other comparison methods in six evaluation indicators: the accuracy rate reached 91.3%, the F1 value reached 89.6%, and it was significantly ahead in warning lead time (3.0 days) and spatial positioning accuracy (error 6.3km). The false alarm rate was reduced to 5.1%, and the inference delay was controlled within 3.8 minutes. The above results fully verified the practicality and advancement of this method in red tide anomaly identification and trend prediction, and it has good engineering deployment prospects and emergency response value.

[0056] In order to verify the model's ability to perceive spatial anomaly distribution, the system constructed a spatial heat map of the probability of red tide occurrence. The results are as follows: Figure 4 As shown. The different shades of color in the figure correspond to the probability of red tide occurrence in different monitoring areas. The darker the color, the higher the red tide risk in the area. It can be seen that the abnormal distribution predicted by the method of the present invention has obvious spatial clustering. The high-risk areas are concentrated in the coastal sea areas that are greatly affected by human activities. The spatial boundaries are clear and the change trend is consistent with historical monitoring data. This result shows that this method not only has a high prediction accuracy, but also can effectively locate the red tide abnormal area on a spatial scale, providing a refined reference basis for subsequent early warning and response measures.

[0057] In order to verify the model's ability to predict the dynamic trend of red tide concentration, the system conducted a comparative analysis of the red tide concentration prediction results of different models at typical monitoring sites. The results are as follows: Figure 5 The figure shows the time-varying trends (t, t+1, t+2, etc.) of the actual red tide concentration curve and the predicted curves of each comparison model. It can be seen that the concentration trajectory predicted by the proposed method closely matches the actual observed values, accurately capturing the rise and fall of red tides, with the best fitting results at mutation points and inflection points. Experimental results demonstrate that the proposed method offers higher prediction accuracy and response sensitivity in time-series modeling of red tide concentrations, facilitating continuous dynamic monitoring of red tide development and trend warnings.

[0058] Example 2 This embodiment provides a marine red tide anomaly detection system that integrates multi-source remote sensing and graph neural networks, including: The data acquisition module is configured as follows: A computer-readable storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device, for a method for detecting marine red tide anomalies that integrates multi-source remote sensing and graph neural networks.

[0059] A terminal device includes a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor to implement a method for detecting marine red tide anomalies that integrates multi-source remote sensing and graph neural networks.

[0060] The above are all preferred embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks, characterized in that: include: Acquire remote sensing image data, drone image data, and monitoring data from monitoring points; Perform data preprocessing on the acquired remote sensing image data and UAV image data; Perform feature extraction and feature fusion on remote sensing images and UAV images to obtain remote sensing feature data; Construct a spatiotemporal graph structure based on the monitoring data of the monitoring points to obtain graph structure data; Based on the cross-modal contrast self-supervised learning mechanism, the consistent representation learning of remote sensing feature modalities and graph structure feature modalities is carried out; Heterogeneous feature fusion and joint representation based on consistency learning; Anomaly detection and early warning are performed based on the fusion characterization results.

2. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural network according to claim 1 is characterized in that: The data preprocessing of the acquired remote sensing image data and UAV image data includes firstly performing physical consistency and geometric consistency processing on the remote sensing image and the UAV image, and performing unified standardization processing on all factor data; Then perform spatiotemporal alignment and missing data completion, which includes temporal alignment of image data and factor data according to timestamps; For the missing data of monitoring points caused by buoy offline, spatial weighted interpolation is used to complete the data. The interpolation and image completion method based on time series is introduced, and the image reconstruction method guided by optical flow is used to realize the content prediction and completion of the frame blocked by fog. Assume that the image frame collected by the drone in continuous time is I t-1 with I t+1 , the goal is to reconstruct the missing intermediate image frame I t , first calculate the forward optical flow F between the two frames t-1 →t+1 and reverse optical flow F t+1 →t-1, the reconstruction value of a pixel position in the intermediate frame is estimated by using the bidirectional optical flow and the time weight α, which is expressed as: , Among them, (x, y) represents a pixel position in the image, (u, v) represents the motion vector of the pixel from time point a to b, and F a→b Represents the optical flow field from frame a to frame b, u1 and v1 represent the optical flow field from I t-1 to I t+1 The optical flow of u2 and v2 from I t+1 to I t-1 Optical flow.

3. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural network according to claim 2 is characterized in that: The feature extraction and feature fusion of remote sensing images and UAV images are performed, including in the time series feature extraction branch, setting the UAV image sequence of length T as { X u } u=1 T , each frame image X u To represent multispectral image data at time t, we first use a lightweight convolutional module to extract local perceptual features. We then introduce the self-attention mechanism in the Transformer architecture to model the dynamic dependencies between image sequences, thereby capturing the temporal behavior of key regions and ultimately obtaining a sequence-level dynamic embedding representation. In the spatial high-resolution feature extraction branch, the input is a single remote sensing image. Multi-level spatial semantic features are extracted through the pre-trained multi-scale residual network ResNet-50 + ASPP. At the same time, in order to highlight potential abnormal areas in the image, a spatial attention mechanism is introduced to recalibrate the feature distribution.

4. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural network according to claim 3 is characterized in that: The spatiotemporal graph structure is constructed based on the monitoring data of the monitoring points, including constructing a dynamic adjacency graph based on a sliding time window. At each time step t, an adjacency matrix is ​​constructed based on the historical sequence {t-w+1,…,t} with a window length of w using feature similarity. A t , capturing the dynamic connection strength between nodes at the current moment; in view of the strong volatility of marine environmental variables at different time scales, a sliding window decomposition mechanism is introduced to divide the long-term time series data into multiple sliding short sequences to extract short-term dynamic features and avoid long-term stable trends covering up sudden anomalies. At each time step t, the node feature tensor under the current sliding window is constructed S t ∈ R w×N×F , and input the graph convolution operation; finally, to adapt to the corresponding differences of nodes, a node-level personality parameter vector is constructed θ i , introduces a personalized weight control mechanism, expressed as: , in, represents the final graph convolution output of node i at time t, represents the trainable personality weight of node i, reflecting its preference for factor response, stands for element-wise multiplication, represents the input features of node j at time t, Representative Node i In time t The set of adjacent nodes.

5. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural network according to claim 4 is characterized in that: The described cross-modal contrastive self-supervised learning mechanism is used to learn the consistent representation of remote sensing feature modalities and graph structure feature modalities. This includes projecting the two modalities into a shared semantic space to achieve mapping consistency between the two modalities in the global semantic space, and using the InfoNCE loss based on contrastive learning to construct a global modality alignment target, which is expressed as: , Among them, sim(a,b) represents cosine similarity, is the temperature coefficient, which is used to adjust the smoothness of the distribution. Represents a set of negative sample graph structures containing different time steps, Represents the feature vector generated by the spatiotemporal graph neural network, Represents the feature vector produced by the dual-branch encoder.

6. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks according to claim 5 is characterized in that: The described cross-modal contrast self-supervised learning mechanism is based on which the remote sensing feature modality and the graph structure feature modality are represented consistently. It also includes introducing a local semantic alignment mechanism to enhance the semantic alignment capability between modalities at a fine-grained level. Specifically, the remote sensing image is divided into K fixed spatial window regions, and each region is extracted with a convolutional encoder to represent the local embedding. At the same time, each node in the graph structure is embedded as , the goal is to enable each local area of ​​the image to find the semantically closest site node in the graph structure and build a cross-modal local alignment pair. The alignment loss is expressed as: , in, v represents a node in the graph structure, and sim(a,b) represents the similarity function.

7. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks according to claim 6 is characterized in that: The heterogeneous feature fusion and joint representation based on consistency learning includes introducing a modal attention mechanism to adaptively adjust the fusion weights of different modal features according to task relevance. I ∈R d , the graph structure modal feature is h G ∈R d , the fusion weight is calculated using the shared attention network, which is expressed as: , Among them, W is the learnable attention parameter, h I represents the image modality feature, h G Represents the graph structure modal features, and the final fusion representation is: , in, and Represents the fusion weight.

8. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural network according to claim 7 is characterized in that: The heterogeneous feature fusion and joint representation based on consistency learning also includes using a semantic channel attention mechanism to enhance the semantic dimension highly related to red tide in the fused feature, and performing weighted adjustment on the fused joint representation in the channel dimension. h fused , the Squeeze-and-Excitation mechanism is used to construct the channel attention weight, and the semantic dimension is added to the fusion feature to obtain the feature h recon ,Will h recon Input into the Transformer-based spatiotemporal modeling module to construct a multi-time multi-modal joint representation sequence. Represents the fused feature sequence of the past T time steps, which is input to the temporal encoder (Temporal Transformer) as: , in, It represents the fused dynamic joint spatiotemporal feature sequence, which serves as the input of the downstream red tide warning and spatial reasoning module.

9. The method for detecting marine red tide anomalies by integrating multi-source remote sensing and graph neural networks according to claim 8 is characterized in that: The abnormality detection and warning based on the fusion characterization results includes adopting an unsupervised abnormality detection method based on an autoencoder, assuming the joint feature of the input ∈R d , the autoencoder consists of an encoder (•) and decoder Ψ(•), which models the normal state by minimizing the reconstruction error and performs anomaly detection based on the set threshold to judge the reconstruction error; the time series prediction model is trained using the historical fusion feature sequence to predict the feature expression at the next moment or several moments in the future, and further predict the probability of red tide occurrence.

10. A marine red tide anomaly detection system integrating multi-source remote sensing and graph neural network, characterized in that: include: The data acquisition module is configured to acquire remote sensing image data, drone image data, and monitoring data of monitoring points; The preprocessing module is configured to perform data preprocessing on the acquired remote sensing image data and the UAV image data; The remote sensing feature module is configured to perform feature extraction and feature fusion on the remote sensing image and the UAV image to obtain remote sensing feature data; The graph structure module is configured to construct a spatiotemporal graph structure based on the monitoring data of the monitoring points to obtain graph structure data; The consistency module is configured to learn the consistent representation of remote sensing feature modalities and graph structure feature modalities based on a cross-modal comparative self-supervised learning mechanism; The joint module is configured to perform heterogeneous feature fusion and joint representation based on consistency learning; The early warning module is configured to perform anomaly detection and early warning based on the fusion characterization results.

Citation Information

Patent Citations

  • Deep learning model and land cover classification method and device

    CN116704363A

  • Marine ecology-oriented time-space diagram neural network anomaly detection method and system

    CN119312267A

  • Red tide anomaly prediction method and system based on dual-channel space-time diagram neural network

    CN119339245A

  • Construction method of unmanned aerial vehicle remote sensing image classification model

    CN119399655A

  • Hyperspectral remote sensing image red tide detection method and system

    CN120071139A

Cited By

  • Physical prior and spatio-temporal evolution fused remote sensing image ocean green tide monitoring method and system

    CN120833561A

  • A remote sensing image marine green tide monitoring method and system fusing physical priori and space-time evolution

    CN120833561B

  • Multi-source heterogeneous data processing method and system for complex river section surveying and mapping

    CN120873980A

  • Method for detecting strong wind weather

    CN120932054A

  • Intelligent quantitative method and system for automatic tea drink production

    CN121209437A