Red tide anomaly detection method and system based on improved multi-mode Transform
By using an improved multimodal Transformer for cross-modal feature learning and knowledge distillation, the problem of insufficient cross-modal modeling capability in existing red tide detection methods is solved, enabling accurate detection and real-time early warning of red tide anomalies, which is suitable for resource-constrained marine edge equipment.
Patent Information
- Application Number
- CN202511075687.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-01
AI Technical Summary
Existing red tide detection methods suffer from insufficient cross-modal modeling capabilities, limited utilization of semantic information, and ambiguous location of abnormal areas. They also struggle to achieve the collaborative utilization of semantic information such as remote sensing images and red tide text descriptions. Furthermore, traditional methods are computationally expensive in resource-constrained marine edge devices, making it difficult to meet the requirements of real-time performance and accuracy.
An improved multimodal Transformer is adopted, which performs cross-modal feature learning through a hierarchical Transformer with a multimodal capsule mechanism. It combines a semantic path-guided attention mechanism for image-semantic feature alignment and uses multimodal knowledge distillation for red tide anomaly prediction. A hierarchical Transformer architecture with cross-modal specific-shared structure and multimodal capsule mechanism is constructed to achieve hierarchical understanding and dynamic aggregation of cross-modal features, and the model complexity is reduced through knowledge distillation.
It achieves accurate perception of complex red tide scenes from multiple angles, improves the semantic consistency and complementarity between image spatial structure and text semantic tags, has stable multimodal representation capabilities, and supports real-time red tide anomaly detection and early warning on resource-constrained devices.
Smart Images

Figure CN120913074A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of red tide anomaly detection, and in particular to a red tide anomaly detection method and system based on an improved multi-modal Transformer. BACKGROUND
[0002] Red tide not only affects fishery resources and aquaculture output, but also causes coastal tourism losses, aggravates water eutrophication, and even causes degradation of nearshore water ecological functions. Red tide events have the characteristics of complex spatio-temporal evolution, multiple sources of causes, short monitoring response period, etc. Therefore, it is urgent to build an intelligent detection and early warning technology system for red tide anomalies to support early perception, accurate identification and dynamic tracking of abnormal marine phenomena.
[0003] Existing red tide detection methods can be roughly divided into three categories: (1) physical and chemical monitoring methods relying on fixed points or manual sampling, mainly collecting parameters such as nutrient salts, chlorophyll, water temperature, pH value, etc., and determining the risk of red tide occurrence by setting thresholds or establishing empirical models, but this method has limited spatial coverage and is difficult to support large-scale dynamic monitoring; (2) visual recognition methods based on remote sensing images, which use medium and high resolution satellite or unmanned aerial vehicle images to observe changes in sea surface color, algal density distribution, etc., and have wide-area observation and frequent update capabilities, but are easily affected by weather and have insufficient recognition accuracy for weak signal or boundary ambiguous red tide; (3) model methods based on machine learning, which use a large number of historical image data or monitoring records to train deep neural networks for discrimination or prediction, and have high automation level, but the effect is easily degraded in the case of insufficient samples and non-uniform data modalities, especially in effectively fusing complex visual information in remote sensing images and semantic factors in monitoring text.
[0004] In addition, although current red tide detection methods have initially introduced image recognition and machine learning techniques, there are still the following outstanding problems: (1) limited modal fusion capability, existing methods mostly rely on single image or numerical factors, making it difficult to cooperatively utilize remote sensing images and semantic information such as red tide text description, resulting in fragmented multi-source information and difficulty in fully restoring red tide evolution characteristics; (2) insufficient accuracy of abnormal area positioning, traditional image threshold or regional mean-based segmentation methods cannot accurately extract weak signal red tide boundaries, especially in complex backgrounds such as cloud cover and uneven lighting, recognition robustness is poor; (3) traditional multi-modal fusion methods have redundant model parameters and high computational cost, making it difficult to deploy in resource-constrained marine edge devices, and ignoring the preservation of cross-modal semantic transfer capability, which easily causes feature loss and expression deviation, making it difficult to meet the dual requirements of real-time and accuracy in the context of red tide anomaly detection.
[0005] Therefore, in view of the limitations of the above-mentioned method in cross-modal modeling capability, limited utilization of semantic information, and fuzzy positioning of abnormal areas, there is an urgent need to propose a red tide anomaly detection method and system based on an improved multi-modal Transformer. SUMMARY
[0006] To solve the above-mentioned problems, the present application provides a red tide anomaly detection method and system based on an improved multi-modal Transformer.
[0007] In a first aspect, the present application provides a red tide anomaly detection method based on an improved multi-modal Transformer, which adopts the following technical solution: A red tide anomaly detection method based on an improved multi-modal Transformer, comprising: acquiring remote sensing images and text data; performing data preprocessing on the acquired remote sensing images and text data, respectively; performing visual positioning and text selection based on the preprocessed data; performing cross-modal feature learning based on a multi-modal capsule mechanism hierarchical Transformer; performing image-semantic feature alignment optimization based on a semantic path guided attention mechanism; performing multi-modal knowledge distillation on the optimized features; performing marine red tide anomaly prediction based on a student model after knowledge distillation based on fused multi-modal information.
[0008] Further, the data preprocessing on the acquired remote sensing images and text data, respectively, includes, for the remote sensing images, removing high-frequency noise components based on wavelet transform DWT, extracting image key points and their descriptions using scale-invariant feature transform SIFT, and constructing an image feature set; for red tide text data, eliminating noise information in unstructured text through text cleaning, extracting semantic relationships between recognized entities, using a structured relationship classifier based on a multi-head attention mechanism to obtain structured triples; and performing timestamp unification and standardization, image feature completion, and pixel-level interpolation recovery based on adjacent pixels in the image spatial local area to estimate the content of the missing area, represented as: , wherein, represents the pixel value of the (x, y) position in the estimated image, i.e., the interpolated and completed pixel point, represents the pixel value of the (x+i, y+j) position of the known adjacent pixel, Omega represents the interpolation window range, wij represents the interpolation weight coefficient, and Z represents the weight normalization factor.
[0009] Further, the visual positioning and text selection based on the pre-processed data include constructing a cross-modal attention map based on a remote sensing image, extracting deep visual features through an image encoder ViT, extracting global text semantics through a language encoder BERT, and constructing an attention map A using a cross-modal attention mechanism, each value A i,j indicating the correlation degree of the position (i, j) in the image with the red tide text semantics, and then using a U-Net decoding network to restore the attention map A to a spatial mask of the original image size, and normalizing through a Sigmoid function to obtain a preliminary abnormal probability mask: , wherein σ represents a Sigmoid activation function, and each value M0(i, j) of the output represents the probability that the pixel (i, j) in the image is an abnormal area, and finally introducing a mask optimizer constructed based on a SAM framework to generate a fine boundary mask using the original image detail information, the language prompt vector and the initial mask.
[0010] Further, the hierarchical Transformer based on the multi-modal capsule mechanism performs cross-modal feature learning, including using a hierarchical Transformer architecture that integrates a modal-specific-shared structure and a multi-modal capsule mechanism to extract modal difference features and shared high-order semantics in stages, and realizing hierarchical understanding and dynamic aggregation across modalities, wherein for image data, a local attention map is constructed based on the mask Mfinal to obtain image modality embedding; for text data, a pre-trained text encoder is used to generate a word vector sequence; a modal-specific-shared structure is constructed by dividing the input features into modal-specific features and shared features through the introduction of a gating mechanism; and a multi-modal capsule mechanism based on dynamic routing is introduced, and an initial low-order capsule vector U={u1,...,u n is obtained based on a trainable projection matrix W ij , and a prediction vector of a low-order capsule i for a high-order capsule j is obtained ij , and a high-order capsule vector is dynamically weighted and aggregated, and the vector module length is compressed using a squash function to generate a probabilistic semantic entity representation v j : , wherein s j represents the vector before aggregation, and the final output is a shared semantic fusion feature , a model-specific feature , and and a high-order capsule output V={v1,...,v kThe cascading composition forms a multi-level semantic representation: .
[0011] Further, the image-semantic feature alignment optimization based on the semantic path guided attention mechanism includes obtaining a word embedding sequence h T ={e1,...,e k} from the input text after the encoder, constructing a semantic path set P by means of event extraction, method recognition, spatial relationship and causal expression in the text: , wherein e i represents the embedding vector of the i-th word, each path p i represents a series of entities connected by semantic relationships r j (i) , and a weighted aggregation is used to construct a path representation vector for the node embedding in each path: , wherein, is the attention weight of the j-th node in the path, and PE(j) is the position encoding item, and finally a path embedding set is obtained. represents the representation vector of the p k th path.
[0012] Further, the image-semantic feature alignment optimization based on the semantic path guided attention mechanism further includes introducing the path embedding into the fusion semantic space output by the previous module for attention adjustment, so that the semantic path serves as a reasoning clue to guide the model to focus on the logical key points, and the multi-modal hierarchical modeling and the capsule aggregation output the fusion feature , which represents the feature of each visual / text fusion unit, and the guided attention weight is obtained by calculating the matching degree between each feature and all semantic paths: , wherein W n is a learnable projection matrix, and β i,j represents the degree to which the fusion feature is affected by the path semantic, and based on the attention distribution, a guided fusion representation Z fused is constructed: .
[0013] Further, the multi-modal knowledge distillation of the optimized feature includes fusing the feature vector Z fused of the image feature and the semantic feature with the detailed text description T LLMThe input teacher model is based on a large-scale Transformer structure to model the input end-to-end, and the encoder part is stacked by multi-head self-attention and feedforward network, and the output of the l-th layer is represented as: , where FFN represents a feedforward network, MHAtt represents a multi-head self-attention mechanism, H (0) represents the initial feature, Z fused represents the feature vector fused by the image feature and the semantic feature, T LLM represents the detailed text description generated by the prompt word driving module, BERT represents the encoding process, and the teacher model outputs a high-order semantic vector h teacher .
[0014] Further, the multi-modal knowledge distillation of the optimized feature further includes simplifying the student model structure to a small number of attention layers and a small feedforward network during the distillation process based on a multi-element distillation loss, inputting the same feature, and outputting a predicted feature h student , wherein the distillation loss includes soft target distillation, feature alignment distillation, and inter-layer attention alignment, and is represented as: where h teacher represents the predicted feature generated by the teacher model, h student represents the predicted feature generated by the student model, KL represents the KL divergence, σ represents the softmax, τ is the temperature coefficient, A(l) is the attention weight matrix of the l-th layer, and the comprehensive total loss is: , wherein, , and represent weight coefficients.
[0015] Further, the student model based on the knowledge distillation of the fused multi-modal information is used for marine red tide anomaly prediction, which includes performing anomaly scoring based on the fused feature sequence output by the student model, and for the feature z s (t) output by the student model at time t, an autoencoder structure is used for unsupervised modeling, and the reconstruction output of the autoencoder is defined as , and the reconstruction error is used as the basis for anomaly scoring, and the anomaly score is represented as , wherein A (t) is the anomaly score at the current time, is the feature vector reconstructed by the autoencoder, and the fused feature output by the student model is compared with the initial anomaly text description The joint input large model generates an updated abnormal text description, and the large language model G θ(⋅) completes semantic understanding and abnormal language generation of the input, and finally obtains an enhanced red tide abnormal warning text , wherein, represents a context-aware, feature-driven abnormal description text, represents an initial abnormal text description, represents the fusion features output by the student model, and Prompt represents the prompt text.
[0016] In a second aspect, a red tide anomaly detection system based on an improved multi-modal Transformer includes: A data acquisition module configured to acquire remote sensing images and text data; A preprocessing module configured to perform data preprocessing on the acquired remote sensing images and text data, respectively; A selection module configured to perform visual positioning and text selection based on the preprocessed data; An alignment module configured to perform cross-modal feature learning and feature alignment based on image-text feature encoding; An optimization module configured to perform image-semantic feature alignment optimization based on a prompt word driven generation mechanism; A distillation module configured to perform multi-modal knowledge distillation on the optimized features; A prediction module configured to perform marine red tide anomaly prediction based on a student model after knowledge distillation of fused multi-modal information.
[0017] In a third aspect, the present application provides a computer-readable storage medium having a plurality of instructions stored therein, the instructions being adapted to be loaded and executed by a processor of a terminal device to implement the red tide anomaly detection method based on the improved multi-modal Transformer.
[0018] In a fourth aspect, the present application provides a terminal device including a processor and a computer-readable storage medium, the processor being configured to implement the instructions, and the computer-readable storage medium being configured to store a plurality of instructions, the instructions being adapted to be loaded and executed by the processor to implement the red tide anomaly detection method based on the improved multi-modal Transformer.
[0019] In summary, the present application has the following beneficial technical effects: Compared with existing technologies, the multimodal large-scale red tide anomaly detection method proposed in this invention, which integrates semantic and visual features, has the following significant advantages: First, the system integrates multi-source heterogeneous data such as remote sensing images and red tide historical records, and combines image and text preprocessing, visual positioning, keyword extraction and other modules to achieve multi-angle accurate perception of complex red tide scenes, breaking through the bottleneck of traditional methods that are difficult to accurately identify red tide areas under the condition of single data and limited information dimensions.
[0020] Secondly, a dual-channel image-text feature encoding structure is constructed, and a shared attention mechanism and cross-modal feature alignment strategy are introduced, which significantly improves the semantic consistency and complementarity between image spatial structure and text semantic labels, enabling the system to still have stable multimodal representation capabilities under weak supervision.
[0021] Third, the system enhances the model's understanding of red tide semantic concepts and its ability to express language through a prompt word-driven generation module. Furthermore, by combining the teacher-student model architecture, it completes lightweight distillation of knowledge from large models, effectively balancing prediction performance and edge deployment requirements.
[0022] Finally, based on the fusion features, an anomaly detection and trend prediction model is constructed, which, together with the visualization early warning output module, can not only realize dynamic monitoring and real-time early warning of red tide anomalies, but also has good generalization ability and spatiotemporal adaptability, providing key support for intelligent monitoring and emergency response in complex marine environments. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of a red tide anomaly detection method based on an improved multimodal Transformer according to Embodiment 1 of the present invention.
[0024] Figure 2 A comparison diagram of ACC and F1 models in Embodiment 1 of the present invention.
[0025] Figure 3 Accuracy variation graphs for different Dropout ratios in Embodiment 1 of the present invention. Detailed Implementation
[0026] The present invention will be further described in detail below with reference to the accompanying drawings.
[0027] Example 1 Reference Figure 1 This embodiment of a red tide anomaly detection method based on an improved multimodal Transformer includes: Acquire remote sensing images and text data; Data preprocessing is performed on the acquired remote sensing images and text data respectively; Visual localization and text selection based on preprocessed data; Cross-modal feature learning and feature alignment based on image-text feature coding; Image-semantic feature alignment optimization based on prompt word-driven generation mechanism; Multi-modal knowledge distillation on optimized features; Ocean red tide anomaly prediction based on student model after knowledge distillation based on fusion of multi-modal information.
[0028] Specifically: S1. Image and text data preprocessing module, In the anomaly detection task of marine red tide, remote sensing images and red tide related text records jointly constitute the dual source of information input. Remote sensing images can provide visual information in a wide range and multiple time phases, while text data such as monitoring logs, forecast records, and expert analysis reports contain rich semantic and experiential knowledge. Due to the modal, structural, and quality differences between images and text data, scientific preprocessing methods must be used for standardization and structuralization to ensure the accuracy and consistency of subsequent model training and feature fusion. This module mainly includes the following three steps: 1) Remote sensing image preprocessing. Remote sensing images, as the core information source for red tide detection, have the advantages of high resolution and wide spatio-temporal coverage, but also face noise and distortion introduced by factors such as atmospheric disturbance, cloud cover, and sensor drift. The method based on wavelet transform (DWT) is used to remove high-frequency noise components. Wavelet transform is a tool that simultaneously has time and frequency local analysis capabilities, suitable for detail extraction and noise removal in remote sensing images. Let the input image be I(x,y), and perform two-dimensional discrete wavelet transform (DWT) on it to obtain four subbands: , Among them, LL represents the low-frequency approximation subband, which retains the main structural information of the image; LH, HL, and HH represent high-frequency detail subbands, which contain edge and noise information. To remove noise, set the threshold λ and perform soft threshold compression on the high-frequency coefficients: , Among them, represents the high-frequency subband coefficient, represents the compressed coefficient, and sign(⋅) represents the sign function. The compressed coefficients are reconstructed into an image through inverse wavelet transform (IDWT) to obtain the denoised image I denoise . Due to factors such as sensor perspective changes and terrain differences, remote sensing images often have geometric shifts. To achieve image alignment, scale-invariant feature transform (SIFT) is used to extract image key points and their descriptors, and an image feature set is constructed: , Among them, (xi y i ) is the i-th key point position, d i is its descriptor vector. Since remote sensing images are often collected from different times, sensors or angles in the actual acquisition process, the images may have spatial position offset. By registration, two images can be "aligned" to make a one-to-one correspondence between pixels. After extracting the feature point set from two images I1, I2, the Euclidean distance is used for matching to obtain the initial corresponding point set. To eliminate matching errors and outliers, the random sample consensus algorithm (RANSAC) is used to obtain the perspective transformation matrix H, and the optimization objective is: , where p i is the matching point in image I1, p i ' is the corresponding point in I2, H represents the transformation matrix, and finally the registered image I aligned is obtained. Different remote sensing images have large differences in pixel value dynamic range due to different imaging sensors and lighting conditions. To enhance the adaptability of the model, the minimum-maximum normalization is used to unify the pixel value to the interval [0, 1]: , where I aligned represents the pixel value of the registered image, I min and I max represent the minimum and maximum values of the image pixel value. Normalization operation can improve the convergence speed and stability of the feature extraction stage, and ensure the consistency of different batches of images in the model input space.
[0029] 2) Red tide text data cleaning and structuring: Red tide related text data mainly includes: marine monitoring reports, research paper abstracts, news reports, etc. Due to the diversity of these texts, different formats, and a lot of noise (such as redundant descriptions, inconsistent terminology, and chaotic space-time elements), they need to be cleaned and structured to improve the efficiency of subsequent model processing and semantic understanding. The goal of text cleaning is to eliminate noise information in unstructured text, including redundant punctuation, format symbols, redundant descriptions, etc., and to standardize terminology. Named entity recognition is used to identify key entities related to red tide from text, such as: time (Time), location (Location), parameter (Parameter), value (Value), and event type (Event). A BERT-based sequence labeling model is used, with BIO encoding for annotation training. Let the sentence be: , where w i is the i-th word, and the corresponding context representation h i is obtained after BERT encoding, and the optimal annotation path is obtained by CRF layer decoding: , where, represents the transition probability function, y represents the label sequence, including B-LOC (location start), I-LOC (location middle), O (other), etc. The relationship extraction target is to extract the semantic relationship between the identified entities, and the structured relationship classifier based on the multi-head attention mechanism is used. The model represents two entities as: , where, represents the context vector splicing of the first entity, which is input into the relationship classifier: , Finally, the structured triple is obtained. In order to facilitate subsequent multi-modal alignment with image data, the text is structured into a unified JSON structure format.
[0030] 3) Time alignment and missing data completion. In order to realize the effective fusion of remote sensing images and text information in multi-modal red tide detection, the time alignment problem between different data sources needs to be solved first, and the detection stability and robustness under the condition of sample missing need to be improved through the completion strategy. Two parts of timestamp unification and standardization and image feature completion need to be handled: timestamp unification and standardization, let the remote sensing image sequence be: , The text data sequence is: , where, t img , t txt represent the timestamps of images and texts. By constructing a time window Δt, image-text data with a time difference less than the window threshold are paired: , Image feature completion. When the remote sensing image is missing (such as being blocked by clouds), pixel-level interpolation is performed based on the adjacent pixels in the local region of the image space to estimate the content of the missing region: , where, represents the estimated pixel value of the (x, y) position in the image, i.e. the interpolated and completed pixel point, represents the pixel value of the (x+i, y+j) position of the known adjacent pixel, Ω represents the interpolation window range, w ij represents the interpolation weight coefficient, and Z represents the weight normalization factor.
[0031] S2. Visual positioning and text selection module, In the task of marine red tide anomaly detection, remote sensing images contain rich information of ocean surface spectrum, which can capture red tide precursors such as seawater discoloration, phytoplankton aggregation, and water turbidity. However, due to the wide coverage of images, the target area has different scales and fuzzy boundaries, and red tide often occurs in a regional and sparse distribution. If the entire image is used directly for anomaly detection, it is easy to be disturbed by irrelevant backgrounds such as coastlines, fishing boats, and clouds. Therefore, it is urgent to design a precise and interpretable anomaly region visual positioning mechanism to focus attention on spatial regions with potential red tide characteristics, thereby improving detection efficiency and accuracy. This module proposes a visual positioning and mask prediction method based on multi-modal pre-trained models, constructs a mask prediction head (Mask Head) to realize pixel-level segmentation of suspected red tide areas in remote sensing images, and introduces a keyword selector to filter out semantic labels in the text sequence that are significantly related to red tide, enhancing the model's semantic perception and interpretability.
[0032] 1) Construct a cross-modal attention map to find potential abnormal regions. The remote sensing image I extracts its deep visual features through the image encoder (ViT), and the output is an embedding tensor: , where h, w are the spatial dimensions of the down-sampled feature map, and d is the number of feature channels. The monitoring text sequence T = {w1, w2,..., w n is input into the language encoder (BERT) to extract global text semantic embeddings: , To establish the association between image spatial regions and text semantics, a cross-modal attention mechanism is used to construct an attention map A, whose each value A i,j represents the relevance of position (i, j) in the image to red tide text semantics, and the calculation method is: , where W q ,W k are learnable linear mapping matrices, and softmax ensures that the sum of attention weights is 1, is the visual feature of the (i, j) point in the image. The attention map A can be regarded as a low-resolution "abnormal heat map".
[0033] 2) Coarse-grained mask generation to locate suspected areas. The resolution of the attention map A is usually low and cannot be directly used for fine-grained region labeling. A lightweight U-Net decoder is designed to restore it to a spatial mask of the original image size. The decoder gradually restores the spatial resolution through multiple layers of upsampling (deconvolution) and outputs the mask logits: , Further normalized by the Sigmoid function, the preliminary anomaly probability mask is obtained: , where σ represents the Sigmoid activation function, and each value M0(i,j) of the output represents the probability that the pixel (i,j) in the image is an abnormal area.
[0034] 3) Combine semantics and image details to refine the boundary mask. The preliminary mask M0 is usually fuzzy and not accurate enough, so a mask optimizer based on the SAM framework is introduced to generate a high-quality fine boundary mask using the original image detail information, language prompt vector and initial mask. The input image is encoded using a Vision Transformer to output an image feature tensor E I The prompt encoder receives the user-provided prompt information and converts it into a prompt vector E P The image feature E I , the prompt embedding E P , and the coarse mask M0 are input together for fusion modeling, and a high-resolution mask logits is output: , where E I represents the image feature, E P represents the prompt vector feature, and M0 represents the coarse mask.
[0035] 4) Keyword selector, identify significant semantic labels of red tide in text. Monitoring text often appears semantic signals such as "sea water turns red", "algal bloom", "high density of phytoplankton", etc., which are important semantic basis for judging the occurrence of red tide. Therefore, a lightweight keyword selector is designed to judge the target noun of each token in the input text. Let the hidden state sequence output by the encoder after processing the text sequence be: , where h represents the text encoding feature. A linear classifier with a sigmoid activation function is used to output the probability that each token is a target keyword: , where W s represents the weight and b s represents the bias. Given a threshold τ, the final selected keyword set is: , S3. Hierarchical multi-modal feature modeling module, In the task of red tide anomaly detection, the spatial texture information of image modality and the event semantics of text modality have significant differences and potential complementarity. To accurately model the deep correlation between the two and improve the structured expression ability of abnormal patterns, this module designs a hierarchical Transformer architecture that integrates modality-specific and shared structures and multi-modal capsule mechanisms to extract modality difference features and shared high-order semantics in stages, achieving cross-modal hierarchical understanding and dynamic aggregation.
[0036] 1) Based on the mask M final Construct a local attention map to obtain the image modality embedding. Let the original remote sensing image be I, and the high-resolution mask M final , construct a local attention map based on the salient region enhancement strategy: , where ⊙ represents pixel-wise multiplication, Blur(⋅) represents image Gaussian blur operation, and λ∈[0,1] controls the background degradation degree. This processing emphasizes the model's ability to focus on suspected abnormal areas and weakens the disturbance of non-target areas. The processed image is input into the visual encoder CLIP-ViT with frozen parameters to extract high-order visual embeddings: , Project to cross-modal shared space: , where W I represents the weight matrix, and b I represents the bias vector.
[0037] 2) Let the original text input be X T ={w1,...,w n}, and the high-confidence semantic labels extracted by the keyword selector be K select ={k1,...,k m}. Use the pre-trained text encoder (CLIP-TextEncoder) to generate a sequence of word vectors: , For the keyword set K, construct a focused representation in the word vector space through an attention mechanism: , where is the word embedding, and α j is the semantic importance weight of each keyword. Then use linear projection to unify the dimensions: , where W T represents the weight matrix, and b T represents the bias vector.
[0038] 3) Modality-Specific-Shared Structure Construction. A gating mechanism is introduced to divide the input features into modality-specific features and shared features. Taking image modality as an example, the gating process is as follows: , Among them, g I Represents the image feature Z I Gated networks, and Represents learnable weights and biases, σ represents the Sigmoid function, and ⊙ represents element-wise multiplication. , These represent the specific and shared features of the image modality, respectively. Similarly, the specific and shared features of the text modality are obtained. and .Will and Feed the data into a collaborative Transformer structure for shared semantic modeling: , 4) Multimodal capsule mechanism for dynamic aggregation of higher-order semantic structures. To further model the higher-order semantic structures and spatial aggregation patterns between text and image modalities, a multimodal capsule mechanism based on dynamic routing is introduced. Perform a linear mapping to obtain the initial low-order capsule vector U={u1,...,u n Based on the trainable projection matrix W ij Obtain the prediction vector of lower-order capsule i for higher-order capsule j. : , By routing coefficient c ij For higher-order capsule vectors Dynamic weighted aggregation is performed, and the vector magnitude is compressed using the squash function to generate a probabilistic semantic entity representation v. j : , Among them, s j The vector before aggregation is represented by the squash function, which preserves the direction and only compresses the magnitude to (0,1), making the vector length interpretable as the "probability of feature entity existence". The final output is composed of shared semantic fusion features. Model-specific characteristics and With the higher-order capsule output V={v1,...,v k Cascaded structures form multi-level semantic representations: .
[0039] S4. Semantic path guided attention module, To improve the understanding depth and detection accuracy of complex abnormal patterns, a semantic path guided attention mechanism is proposed, which aims to mine the spatial orientation information and causal chain structure in the red tide description text, construct structured semantic path embedding, and guide the model to realize more human cognitive abnormal detection logic modeling in the multi-modal semantic fusion space, providing more explainable and task-oriented feature basis for subsequent knowledge compression and lightweight deployment.
[0040] 1) The word embedding sequence h T = {e1,...,e k} obtained after encoding the input text is used to identify the spatial relationships and causal expressions in the text by event extraction, dependency syntax analysis or knowledge template method, and construct a semantic path set P: , where e i represents the embedding vector of the i-th word, and each path p i represents a series of entities or events connected by semantic relationships r j (i) (e.g. "temperature rise → water eutrophication → red tide outbreak"), which embodies the reasoning chain of red tide occurrence.
[0041] 2) To embed the representation space of these paths into the model, the weighted aggregation method is used to construct the path representation vector for the node embedding in each path: , where is the attention weight of the j-th node in the path (reflecting its semantic importance in the path), and PE(j) is the position encoding item, which is used to preserve the order information. Finally, the path embedding set , where represents the representation vector of the p k th path.
[0042] 3) The path embedding is introduced into the fusion semantic space output by the previous module to adjust the attention, so that the semantic path serves as a "reasoning clue" to guide the model to focus on the logical key points. The fusion feature of multi-modal hierarchical modeling and capsule aggregation output is , which represents the feature of each visual / text fusion unit. We calculate the matching degree between each feature and all semantic paths to obtain the guided attention weight: , where W nis a learnable projection matrix, β i,j denotes the fused feature pathway influenced by the semantics. Based on this attention distribution, we construct the guided fused representation Z fused : , S5. Multi-modal knowledge distillation and lightweight deployment module, To realize the rapid deployment and efficient operation of the red tide anomaly detection model on the device, this module is based on image-semantic joint representation, aiming to significantly compress the parameter size and reduce the computational complexity while maintaining the ability of the model in image anomaly perception and semantic understanding, to realize the availability and real-time performance of the model in the actual environment.
[0043] 1) Teacher model design and training. The feature vector Z fused is input into the teacher model, which adopts a large-scale Transformer structure to model the input end-to-end. The encoder part is stacked by multi-head self-attention and feed-forward network, and the output of the l-th layer can be represented as: , where FFN represents the feed-forward network, MHAtt represents the multi-head self-attention mechanism, H (0) represents the initial feature, and Z fused represents the feature vector fused by image features and semantic features. In the last layer, the teacher model outputs a high-order semantic vector h teacher .
[0044] 2) Student model distillation and lightweight deployment. To transfer the above high-dimensional ability to a lightweight model, we design a multi-element distillation loss. In the distillation process, the student model structure is simplified to a small number of attention layers and a small feed-forward network, and the same feature is input to output the predicted feature h student . The distillation loss consists of three parts: soft target distillation, feature alignment distillation, and inter-layer attention alignment, which are: where h teacher represents the predicted feature generated by the teacher model, h student represents the predicted feature generated by the student model, KL represents the KL divergence, σ represents the softmax, τ is the temperature coefficient, and A(l) is the attention weight matrix of the l-th layer. The comprehensive total loss is: , where , and represent the weight coefficients.
[0045] S6. Red tide anomaly detection and intelligent early warning module, This module aims to realize accurate identification of marine red tide anomaly, time series risk evolution modeling and multi-dimensional early warning output based on the student model output of multi-modal information fusion. In the context of complex marine ecological environment and multi-source heterogeneous red tide triggering factors, single modal information is easily disturbed by local observation errors or abnormal disturbances, making it difficult to support stable and interpretable early warning judgments. Therefore, the system relies on the semantic fusion feature vector sequence generated by the student model to build an end-to-end anomaly detection and intelligent early warning mechanism.
[0046] 1) Abnormal score generation and judgment mechanism. To accurately identify whether there is a red tide anomaly at the current time, the module first generates an abnormal score based on the fusion feature sequence output by the student model. Let the feature output by the student model at time t be z s (t) , and use the autoencoder structure for unsupervised modeling. The autoencoder consists of an encoder function ϕ(·) and a decoder function ψ(·), which aims to learn a low-dimensional compression and restoration mapping on the normal data distribution. Define the reconstruction output as: , and take the reconstruction error as the basis for abnormal score, the abnormal score can be further defined as: , where A (t) is the abnormal score at the current time, is the feature vector after reconstruction of the autoencoder. When A(t) exceeds the preset threshold , i.e. A(t)> , it is determined that there is a potential red tide anomaly event at the current time.
[0047] 2) Abnormal type interpretation and multi-dimensional early warning output. While implementing anomaly detection, the system further analyzes the semantic structure and time series evolution characteristics of the anomaly. Based on the fusion feature z s (t) output by the student model, a feature-semantic mapping mechanism is established combining the key factor labels extracted from the red tide monitoring text (such as "water temperature rise", "nutrient salt enrichment", "frontal stability", etc.). The contribution of each type of semantic factor to the abnormal score is calculated through attention weighting, realizing the explanatory modeling of the possible causes of the abnormal event. Define the candidate semantic factor set as {l1, l2,..., l K}, and the attention score of the fusion feature in the factor embedding space is: , where e lk represents the embedding vector of semantic factor l k , Wa are learnable weight parameters. Attention score a k (t) can be used to explain the risk contribution of each semantic factor in the current anomaly, thereby assisting in analyzing the cause of the anomaly, the evolution path, and the possible impact range.
[0048] 3) Structured early warning information release and visualization. The system organizes information such as anomaly score A(t), anomaly type explanation weight {a k (t)}, corresponding spatio-temporal positioning, risk level, historical evolution trend, etc. to generate structured intelligent early warning output. The output includes anomaly detection results (whether abnormal, abnormal intensity), key cause explanation (high-weight semantic factors and their contribution), risk level classification (dynamic classification according to score threshold). The system supports visual output and supports information visualization in the form of charts, time series curves, etc.
[0049] Experimental verification: To verify the effectiveness of the proposed method and system for red tide anomaly detection based on an improved multi-modal Transformer, this paper constructs a real scene experimental platform covering multi-modal information in a typical nearshore red tide frequent area in China. The data sources include medium and high resolution remote sensing images, historical red tide observation records, monitoring buoy measured factors (temperature, salinity, pH, chlorophyll concentration, etc.), and manually labeled anomaly reports, forming a total of multiple groups of samples, covering the whole process of red tide evolution stage (initial appearance - expansion - outbreak - subsidence). To enhance the adaptability and challenge of the experiment to the real scene, complex situations such as remote sensing image cloud cover, observation record text redundancy, and missing monitoring indicators are simulated during data construction, fully testing the robustness and generalization ability of the system under uncertain conditions.
[0050] The comparative experiment selects the current mainstream multi-modal anomaly detection and marine intelligent perception model, including cross-modal matching Transformer (CMMT), ConvLSTM using a space-time modeling structure, ASTGCN combining graph neural network and attention mechanism, BLIP-2 based on graph-text alignment structure, CLIP multi-modal fusion representation model, and the method proposed in this paper. Training and testing are carried out under unified data division (training set: validation set: test set = 6:2:2), consistent optimizer and learning rate strategy to ensure the fairness of model comparison. The performance evaluation indicators cover accuracy, F1 value, abnormal area spatial positioning error, early warning lead time, false alarm rate and model inference latency, which comprehensively evaluate the performance of each method from the three dimensions of detection accuracy, response timeliness and deployment efficiency. Experimental results show that the proposed method performs best in all evaluation indicators, fully verifying the significant advantages of the system in multi-modal data processing and red tide detection tasks, and has good application prospect and promotion value.
[0051] Table 1 Comparison of data of different methods in six indicators Model name ACC F1 Spatial positioning error Early warning lead time False alarm rate Inference delay CMMT 86.7% 84.2% 5.3km 1.8 days 11.5% 480s ConvLSTM 82. 4% 79.6% 6.8km 1.2 days 15.8% 210s ASTGCN 84.1% 81.3% 6.1km 1.4 days 13.2% 320s BLIP-2 88.3% 85.7% 4.9km 2.0 days 10.3% 920s CLIP 85.6% 83.1% 5.6km 1.7 days 12.1% 690s The method of the present invention 91.4% 89.2% 4.3km 2.4 days 8.4% 180s As can be seen from Table 1 and Figure 2 , the mainstream methods such as CMMT, ConvLSTM, ASTGCN, BLIP-2 and CLIP all show certain performance basis in red tide anomaly detection tasks, but there are still obvious short boards in key indicator dimensions. ConvLSTM has strong memory ability in time series modeling, but lacks spatial dependence structure modeling, resulting in low detection accuracy, with accuracy and F1 value of 82.4% and 79.6%, respectively, the lowest among all models, and spatial positioning error of 6.8 km. ASTGCN models the correlation between time and space through graph attention mechanism, and has better early warning lead time (1.4 days) than ConvLSTM, but due to the static nature of the graph structure, it is difficult to adapt to complex and variable red tide propagation paths, and the false alarm rate remains at 13.2%. BLIP-2 and CLIP introduce a graph-text joint learning mechanism to enhance the model's understanding of remote sensing images and text information, achieving accuracy of 88.3% and 85.6%, respectively, and good performance in spatial positioning error control (both less than 5.6 km), but both have high inference delay (> 600 s), which makes it difficult to meet the real-time warning requirements. CMMT adopts a cross-modal matching mechanism, which outperforms the above models in F1 value (84.2%) and early warning capability (1.8 days), but still has the limitations of high false alarm rate (11.5%) and unstable fusion features.
[0052] In contrast, the method proposed in the present application realizes significant improvement in accuracy (91.4%), F1 value (89.2%), spatial positioning error (4.3 km) and early warning lead time (2.4 days) by introducing visual positioning and mask prediction, multi-modal dual-channel encoding feature mechanism and distillation optimization deployment strategy, and the inference delay is controlled within 180s, which shows excellent detection accuracy, response speed and deployment feasibility, and verifies the strong robustness and practical value of the method in complex multi-modal environment.
[0053] In order to verify the robustness performance of the model under input disturbance conditions, the system artificially sets different proportions of data dropout in the experiment to simulate the actual scene of observation data loss or sensor abnormality, and the results are shown in Figure 3 As shown in the figure, the change trend of the accuracy (ACC) of each model is shown when the Dropout ratio is gradually increased from 0% to 40%. The method of the present application (red solid line) maintains high accuracy under each level of disturbance, and its ACC only decreases slightly from 91.4% to 86%, with an overall fluctuation of less than 6%, showing significant robustness. In contrast, CLIP and BLIP-2 are more sensitive to disturbance, and their accuracy decreases to 78% and 80%, respectively, while ConvLSTM and ASTGCN degrade more severely under high Dropout conditions, with an accuracy of about 73%. The experimental results show that the method has strong fault tolerance and can effectively resist performance degradation caused by input abnormalities, and is suitable for uncertain early warning tasks in complex marine monitoring scenarios.
[0054] A computer-readable storage medium, wherein a plurality of instructions are stored, the instructions being adapted to be loaded and executed by a processor of a terminal device to implement the red tide anomaly detection method based on the improved multi-modal Transformer.
[0055] A terminal device, comprising a processor and a computer-readable storage medium, the processor being configured to implement instructions; and the computer-readable storage medium being configured to store a plurality of instructions, the instructions being adapted to be loaded and executed by the processor to implement the red tide anomaly detection method based on the improved multi-modal Transformer.
[0056] The above are preferred embodiments of the present application, and are not intended to limit the protection scope of the present application, therefore: any equivalent changes made in the structure, shape, principle of the present application should be covered within the protection scope of the present application.
Claims
1. A red tide anomaly detection method based on an improved multi-modal Transformer, characterized in that, The method comprises the following steps: acquiring remote sensing images and text data; respectively pre-processing the acquired remote sensing images and text data; based on the pre-processed data, visual positioning and text selection are performed; cross-modal feature learning is performed based on a multi-modal capsule mechanism hierarchical Transformer; image-semantic feature alignment optimization is performed based on a semantic path guided attention mechanism; multi-modal knowledge distillation is performed on the optimized features; ocean red tide anomaly prediction is performed based on a student model after knowledge distillation of fused multi-modal information.
2. The method according to claim 1, wherein, The pre-processing of the acquired remote sensing images and text data respectively comprises the following steps: for the remote sensing images, high-frequency noise components are removed based on wavelet transform DWT, image key points and their descriptions are extracted using scale invariant feature transform SIFT, and an image feature set is constructed; for the red tide text data, noise information in unstructured text is eliminated through text cleaning, semantic relationships are extracted between the recognized entities, and structured triples are obtained using a structured relationship classifier based on a multi-head attention mechanism; timestamp unification and standardization, image feature completion, and pixel-level interpolation recovery based on the local region of the image space are performed, and the content of the missing area is estimated, expressed as: , wherein, represents the pixel value of the (x, y) position in the estimated image, i.e. the pixel to be interpolated, represents the pixel value of the known neighboring pixel of the (x+i, y+j) position, Ω represents the interpolation window range, wij represents the interpolation weight coefficient, and Z represents the weight normalization factor.
3. The method according to claim 2, wherein, The visual positioning and text selection based on the preprocessed data include constructing a cross-modal attention map based on a remote sensing image, extracting deep visual features through an image encoder ViT, extracting global text semantics through a language encoder BERT, constructing an attention map A using a cross-modal attention mechanism, each value A i,j represents the correlation degree of the position (i, j) in the image with the red tide text semantics, and the attention map A is restored to a spatial mask of the original image size using a U-Net decoding network, and normalized through a Sigmoid function to obtain a preliminary abnormal probability mask: , wherein σ represents a Sigmoid activation function, and each value M0(i,j) of the output represents the probability that the pixel (i,j) in the image is an abnormal area, and finally a mask optimizer based on the SAM framework is introduced to generate a fine boundary mask using the original image detail information, the language prompt vector, and the initial mask.
4. The red tide anomaly detection method based on the improved multi-modal Transformer according to claim 3, characterized in that, The hierarchical Transformer based on the multi-modal capsule mechanism performs cross-modal feature learning, including using a hierarchical Transformer architecture that fuses a modal-specific-shared structure and a multi-modal capsule mechanism to extract modal difference features and shared high-order semantics in stages, to achieve cross-modal hierarchical understanding and dynamic aggregation, wherein, for image data, a local attention map is constructed based on a mask Mfinal to obtain image modal embedding; for text data, a pre-trained text encoder is used to generate a word vector sequence; a modal-specific-shared structure is constructed by introducing a gating mechanism to divide input features into modal-specific features and shared features; a multi-modal capsule mechanism based on dynamic routing is also introduced, an initial low-order capsule vector U={u1,...,u n} is obtained based on linear mapping, a prediction vector of a low-order capsule i for a high-order capsule j is obtained based on a trainable projection matrix W ij , a high-order capsule vector is dynamically weighted and aggregated by a routing coefficient c ij , and a squash function is used to compress the vector module length to generate a probabilistic semantic entity representation v j : , where s j denotes the vector before aggregation, and the final output is fused by shared semantic features , model-specific features and with high-order capsule outputs V={v1,...,v k} to form multi-level semantic representations: .
5. The method according to claim 4, wherein, The image-semantic feature alignment optimization based on the semantic path guided attention mechanism includes a word embedding sequence h T ={e1,...,e k} obtained after the input text passes through an encoder, and a semantic path set P is constructed by means of event extraction, method recognition, spatial relationship and causal expression existing in the text. , where e i denotes the embedding vector of the ith word, and p i denotes the path representation vector for path p j (i) connected entities, and a weighted aggregation of the node embeddings in each path is used to construct the path representation vector: , wherein, is the attention weight of the jth node in the path, PE(j) is the positional encoding term, and the final path embedding set is obtained wherein, represents the representation vector of the pth k path.
6. The method according to claim 5, wherein, The image-semantic feature alignment optimization based on the semantic path guided attention mechanism further comprises introducing a path embedding into a fusion semantic space of an output of the previous module for attention adjustment, so that the semantic path guides the model to focus on logical key points, and the fusion feature output by the multi-modal hierarchical modeling and capsule aggregation is , represents the feature of each visual / text fusion unit, and the guided attention weight is obtained by calculating the matching degree between each feature and all semantic paths: , where W n is a learnable projection matrix, β i,j denotes the fused feature The degree of influence of the path semantics, based on the attention distribution, constructs the guided fused representation Z fused : .
7. The method according to claim 6, wherein, The optimized features are subjected to multi-modal knowledge distillation, including a feature vector Z that fuses image features and semantic features fused a detailed text description T generated by the prompt word driving module LLM An input teacher model models the input end to end based on a large-scale Transformer structure, and an encoder part is stacked by multi-head self-attention and a feedforward network. The output of the i-th layer is represented as: , Wherein, FFN represents a feedforward network, MHAtt represents a multi-head self-attention mechanism, H (0) represents an initial feature, Z fused represents a feature vector fused by image features and semantic features, T LLM represents a detailed text description generated by the prompt word driving module, BERT represents an encoding process, and a high-order semantic vector h teacher is output by the teacher model at the last layer.
8. The method according to claim 7, wherein, The multi-modal knowledge distillation on the optimized features further comprises, based on a multi-element distillation loss, simplifying a student model structure into a small number of attention layers and a small feedforward network in a distillation process, inputting the same features, and outputting predicted features h student wherein the distillation loss comprises soft target distillation, feature alignment distillation, and inter-layer attention alignment, and is expressed as: Among them, h teacher h represents the predicted features generated by the teacher model. student The predicted features generated by the student model represent the KL divergence, σ represents softmax, τ is the temperature coefficient, A(l) is the attention weight matrix of the l-th layer, and the total loss is: , wherein , and represent weight coefficients.
9. The method according to claim 8, wherein, The student model based on the knowledge distillation of the fusion multi-modal information is used for marine red tide anomaly prediction, including performing anomaly scoring based on a fusion feature sequence output by the student model, and the feature output by the student model at time t is z s (t) An auto-encoder structure is used for unsupervised modeling, and the reconstruction output of the auto-encoder is defined as The reconstruction error is used as the basis of the anomaly score, and the anomaly score is represented as Wherein, A (t) is the anomaly score at the current time, is the feature vector after reconstruction of the auto-encoder, the fusion feature output by the student model and the initial anomaly text description generated by the prompt word driving module are jointly input into a large model to generate an updated anomaly text description, through a large language model G θ(⋅) , semantic understanding and anomaly language generation are completed, and an enhanced red tide anomaly warning text is finally obtained , wherein, represents a context-aware, feature-driven anomaly description text, represents an initial anomaly text description, represents a fused feature output by the student model, and Prompt represents a prompt text.
10. A red tide anomaly detection system based on an improved multi-modal Transformer, characterized in that, The method comprises the following steps: a data acquisition module configured to acquire remote sensing images and text data; a pre-processing module configured to respectively pre-process the acquired remote sensing images and text data; a selection module configured to perform visual positioning and text selection based on the pre-processed data; an alignment module configured to perform cross-modal feature learning and feature alignment based on image-text feature encoding; an optimization module configured to perform image-semantic feature alignment optimization based on a prompt word driven generation mechanism; a distillation module configured to perform multi-modal knowledge distillation on the optimized features; a prediction module configured to perform ocean red tide anomaly prediction based on a student model after knowledge distillation of fused multi-modal information.
Citation Information
Patent Citations
Multi-mode-based adaptive remote sensing image target detection method and system
CN116994108A
Safety propaganda and education training knowledge graph and data management method and system based on AI
CN118585658A
Position and semantic optimization method and system for remote sensing visual question and answer
CN118862002A
Grounded visual question answering method based on daynamic two-level visual information fusion
US20250140124A1
Text-guided multi-modal relationship extraction method and apparatus
WO2025130069A1
Cited By
Algal bloom risk remote sensing intelligent identification method and system
CN121121530A
Multi-mode and reinforcement learning intelligent evaluation method and system
CN121210929A
Foreign matter intelligent identification and dynamic monitoring method based on remote sensing data fusion
CN121259611A
Method and system for detecting abnormal deformation of water conservancy project
CN121482044A
A method and system for detecting abnormal deformation of a hydraulic engineering
CN121482044B