Cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning

Through the method of multimodal data fusion and deep learning, combined with multispectral cameras and infrared thermal imaging technology, the spatial characteristics and time series characteristics of cyanobacterial blooms are extracted, which solves the accuracy and complexity problems of cyanobacterial bloom monitoring in low visibility environments and realizes efficient cyanobacterial bloom prediction and early warning.

CN120411797BActive Publication Date: 2025-10-14ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510924750.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-14
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

The existing cyanobacterial bloom monitoring technology suffers from degraded imaging quality in low-visibility environments, making it difficult to effectively capture the spatial distribution, morphological evolution, and temporal diffusion trends of cyanobacterial blooms. In addition, multispectral data processing is complex, and high data integrity and quality requirements are required, resulting in a decrease in prediction accuracy.

Method used

Multimodal data fusion and deep learning methods are adopted. Image data is generated through multispectral cameras and infrared thermal imaging technology. Convolutional neural networks and long short-term memory networks are combined to extract spatial and temporal features. Single-class support vector machines and attention mechanisms are used to identify and predict outliers, and a bidirectional spatiotemporal prediction network is constructed.

Benefits of technology

It can effectively suppress scattering interference under extreme weather conditions, improve feature extraction rate and prediction accuracy, break through the limitations of traditional methods of temporal and spatial feature separation, and achieve high-reliability early warning of cyanobacteria blooms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411797B_ABST
    Figure CN120411797B_ABST
Patent Text Reader

Abstract

The embodiment of the application is suitable for the field of computer vision environment monitoring, and provides a cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning. In the monitoring method, after the initial data obtained is preprocessed, a variation point detection is introduced to reduce human factor interference, reduce calculation amount, and enhance model adaptability; a convolutional neural network (CNN) is used to extract multi-layer features, and the features are input into a long short-term memory (LSTM) network for time series modeling, so as to improve model performance; on this basis, the features of the CNN and the LSTM are fused, so as to further improve model accuracy; finally, the variation point detection is introduced again, rich feature information and multi-modal complementarity are used to enhance image key information, and detection precision and model robustness are improved. The monitoring method of the embodiment of the application breaks through the limitation of traditional method space-time feature segmentation, significantly improves prediction continuity and accuracy, and provides high reliability support for cyanobacterial bloom early warning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present application belongs to the technical field of environmental monitoring of computer vision, and particularly relates to a cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning. BACKGROUND

[0002] In recent years, the frequency and intensity of cyanobacterial blooms have significantly increased, which has become a major environmental problem threatening the safety of aquatic ecosystems, the quality of drinking water sources, and human health. Therefore, timely and accurate monitoring and identification of cyanobacterial blooms have become an urgent and critical technical problem in the current research field for preventing and controlling cyanotoxin poisoning events.

[0003] The existing methods for improving the accuracy of cyanobacterial target detection and prediction accuracy can be divided into two categories: deep learning-based image processing methods (DL-IPM) and deep learning-based multi-spectral data methods (DL-MDM). The former mainly relies on RGB image data and uses convolutional neural networks (CNN) and attention mechanisms to directly extract features from images and perform classification or detection. The latter combines multi-spectral data (such as visible light, near-infrared, etc.) and uses deep learning models (such as Transformer) to process multi-band information, which can better capture the spectral characteristics of cyanobacteria.

[0004] Specifically, DL-IPM mainly relies on visible light image data, so in low-visibility environments such as rain, snow, and haze, light scattering and spectral distortion can cause imaging quality degradation, affecting the extraction and identification of cyanobacterial features. In addition, DL-IPM usually processes spatial features and time series independently, making it difficult to effectively capture the dynamic coupling relationship between the spatial distribution and morphological evolution of cyanobacterial blooms and the temporal diffusion trend, thereby limiting the accuracy of the prediction model.

[0005] In related technologies, although DL-MDM utilizes the rich spectral information of multi-spectral data, the data dimension is high and the processing is complex, requiring more complex preprocessing and feature extraction methods. At the same time, although multi-spectral data can alleviate the impact of low-visibility environments to some extent, spectral distortion and scattering can still cause imaging quality degradation in extreme weather conditions. In addition, DL-MDM has high requirements for data integrity and quality, and data missing or uneven collection intervals can lead to decreased prediction accuracy of the model, making it difficult to adaptively correct. SUMMARY

[0006] The purpose of the embodiment of the present application is to provide a cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning to at least solve the technical problem of poor model prediction accuracy in related cyanobacterial bloom monitoring technologies.

[0007] To achieve the above-mentioned purpose, the embodiment of the present application provides the following technical solutions.

[0008] According to one embodiment of the present application, a cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning is provided;

[0009] comprising the following steps:

[0010] Step S1, using a multi-spectral camera to collect the reflected light and scattered light of the cyanobacterial distribution area, combining with infrared thermal imaging technology for multi-spectral fusion to generate image data;

[0011] Step S2, taking the image data as the input of the improved FAST algorithm to obtain candidate mutation points, screening the candidate mutation points based on a dynamic threshold and removing isolated noise points through DBSCAN clustering to output key regions with dense spatial distribution;

[0012] Step S3, based on the screened key regions, fusing visible light texture features and infrared thermal imaging temperature gradients to generate a joint feature map and screening out abnormal regions;

[0013] Step S4, using a convolutional neural network to extract spatial features and combining a long short-term memory network to model time series and output time sequence features, combining a space-time attention mechanism to dynamically calculate weights, and weighting the spatial features and time sequence features based on the calculated weights;

[0014] Step S5, introducing a one-class support vector machine to identify abnormal points of the fused features, and constructing a bidirectional space-time prediction network based on an attention mechanism, in which the prediction network, through a multi-head self-attention mechanism, extracts global space-time dependency, decodes features in forward / reverse time sequence, generates cyanobacterial density distribution prediction values, and obtains a prediction result;

[0015] Step S6, combining historical data and prediction results of the cyanobacterial distribution area to generate a space-time evolution map of cyanobacterial diffusion trend for visual analysis.

[0016] Further, in step S4, the step of using a convolutional neural network to extract spatial features comprises:

[0017] constructing a double-branch convolutional neural network for extracting multi-modal features of cyanobacteria based on an HSV color space;

[0018] by constructing a three-level feature pyramid of low-level features, medium-level features and high-level features, the spatial information of different scales is fused; wherein, in the low-level feature layer, the basic visual information including hue boundary and local texture in the image is captured; in the medium-level feature layer, through 3x3 transpose convolution upsampling, algal spot morphology and saturation gradient are extracted to balance detail and semantic information; in the high-level feature layer, after global average pooling of the medium-level features, the channels are compressed to generate a compact global feature vector, enhancing the generalization ability of the model to complex scenes;

[0019] After processing the data through the three-level pyramid, the data is fused across the hierarchy, and the high-level features are upsampled to 64x64 and spliced with the medium-level features and the low-level features to form 64x64x864, and then compressed to 64x64x512 through 1x1 convolution, which not only eliminates redundant information but also greatly reduces the amount of calculation.

[0020] Further, in the constructed double-branch convolutional neural network, the hue branch layer adopts 5x5 hollow convolution to expand the receptive field and capture the hue boundary features, and outputs 128x128x32 feature maps after 3x3 pooling layer; the saturation-lightness joint branch layer extracts texture features through 3x3 convolution and introduces an SE attention module to dynamically enhance key regions, wherein the weight matrix 、 outputs 128x128x64 feature maps; the features of the two branches are spliced into 128x128x96, compressed to 128x128x128 through 1x1 convolution, and the cross-modal interaction is enhanced.

[0021] Further, in step S4, the step of combining the long short-term memory network to model the time sequence and output the time sequence features includes:

[0022] The 64x64x512 spatiotemporal feature map sequence is unfolded along the spatial dimension into 64x64 independent time sequences, which are input into a bidirectional LSTM network: the forward LSTM layer processes the sequence in chronological order to generate forward hidden states ; the backward LSTM layer processes the sequence in reverse chronological order to generate backward hidden states ;

[0023] Finally, the forward and backward hidden states of each spatial position are spliced along the channel dimension to output a 64x64x1024 spatiotemporal coupling feature matrix, which completely preserves the spatiotemporal evolution law of cyanobacterial diffusion.

[0024] Further, in step S4, the step of combining the long short-term memory network to model the time sequence and output the time sequence features includes:

[0025] The 64x64x512 spatial features output by the CNN are expanded to 1024 channels through 1x1 convolution, and the dimension is aligned with the 64x64x1024 time sequence features of the LSTM;

[0026] The joint features are extracted through 3x3 convolution and ReLU activation in 512 channels, and then compressed through 1x1 convolution to generate an attention map. After Softmax normalization, the fusion weight a of each position is obtained, a∈[0,1], and the weight a is expressed as:

[0027] ;

[0028] The CNN feature and the LSTM feature are fused by weighting with weight a and weight 1-a, and the fused feature is weighted is represented as:

[0029] ;

[0030] In the formula, represents the spatial feature map output by the CNN, with a dimension of 64x64x512; represents the time sequence feature map output by the LSTM, with a dimension of 64x64x1024; represents the joint feature map of 64x64x1536 generated by splicing the CNN and LSTM features along the channel dimension; represents a 3x3 convolution kernel with an output channel number of 512, used for extracting the joint spatio-temporal feature; represents a 1x1 convolution kernel with an output channel number of 1, used for compressing the feature to generate an attention weight map; represents a normalization operation, which maps the weight to the interval [0, 1] and ensures that the sum of the weights is equal to 1, a represents a dynamic weight map with a dimension of 64x64x1, and represents an element-wise multiplication, which is used to weight the CNN and LSTM features at each position.

[0031] Further, in step S5, the step of introducing the one-class support vector machine to identify the abnormal points of the fused features includes:

[0032] The output of the CNN-LSTM spatio-temporal feature fusion as a variation point;

[0033] The abnormal detection algorithm based on the one-class support vector machine regards each position of the 64x64 grid as an independent sample, and a total of 4096 feature vectors are generated, each with a dimension of 1024;

[0034] The principal component analysis (PCA) is used to reduce the feature vectors to 32 dimensions, and the decision function is used for variation point detection. When f(x)=-1, it is marked as a variation point;

[0035] The abnormal probability heat map output based on the variation point detection , a candidate region coordinate set satisfying is extracted ;

[0036] wherein the constructed abnormal detection algorithm is represented as: ;

[0037] The decision function is represented as: f(x)=sgn(w·Φ(x)-ρ);

[0038] In the formula, (x+1, y) represents the coordinates of the pixel adjacent to the right of the current pixel (x, y), (x-1, y) represents the coordinates of the pixel adjacent to the left of the current pixel (x, y), w is a weight vector of the hyperplane, is a mapping of the data point x in the high-dimensional space, and is a bias term; sgn() represents a sign function.

[0039] Further, in step S5, a global space-time dependency is extracted by a multi-head self-attention mechanism, and is represented as:

[0040]

[0041] In the formula, Q, K, and V are query matrix, key matrix, and value matrix respectively; d k represents a feature dimension;

[0042] The generated cyanobacterial density distribution prediction value is represented as

[0043] Further, in step S6, a cyanobacterial bloom outbreak threshold is defined , and is represented as: , wherein is a historical cyanobacterial density mean, is a standard deviation; when the prediction value , the grid is marked as a high-risk area, an early warning heat map is output, and the risk level is marked.

[0044] Compared with the prior art, the cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning has the following beneficial effects:

[0045] First, the present application realizes a breakthrough improvement in environmental adaptability and feature extraction efficiency through the design of multi-modal data fusion and deep learning architecture; for the spectral distortion problem of traditional methods in low-visibility scenes, through the multispectral fusion technology of infrared thermal imaging and visible light data, combined with radiation correction and atmospheric correction algorithm, the scattering interference under extreme weather such as rain, snow and haze is effectively suppressed; the brightness and chroma features are independently processed by using the HSV color space, and the complementary information fusion is realized by dynamic weight distribution, which significantly improves the effective feature extraction rate and reduces the feature loss rate;

[0046] ​​Secondly, the application solves the dual bottleneck of efficiency and accuracy of traditional models through a dual-variant point detection mechanism and a space-time feature fusion architecture; an improved corner detection and texture analysis algorithm is introduced to quickly locate the key area and eliminate noise interference, thereby greatly reducing the redundant calculation amount; a CNN-LSTM cascade architecture is combined to model the time series dynamics through a bidirectional LSTM and perform attention weighted fusion with the CNN spatial features, thereby realizing holographic capture of the spatiotemporal evolution law of cyanobacterial diffusion; this design breaks through the limitations of the traditional method of splitting space-time features and significantly improves the prediction continuity and accuracy, thereby providing high reliability support for cyanobacterial bloom early warning. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some embodiments of the application.

[0048] Figure 1 A system architecture diagram of the cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning of the application;

[0049] Figure 2 An implementation flowchart of the cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning of the application;

[0050] Figure 3 A sub-flowchart of the cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning of the application; DETAILED DESCRIPTION

[0051] In order to make the objects, technical solutions and advantages of the application more clear, the following will further describe the application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and not to limit the application.

[0052] The specific implementation of the application will be described in detail in combination with specific embodiments.

[0053] Please refer to Figure 1 and Figure 2 In an embodiment of the application, a cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning is provided, which comprises the following steps:

[0054] Step S1, using a multi-spectral camera to collect the reflected light and scattered light of the cyanobacterial distribution area, combining with infrared thermal imaging technology to perform multi-spectral fusion, and generating image data;

[0055] In step S1 of the present application, the satellite and the unmanned aerial vehicle carry multi-spectral cameras to collect the reflected light and scattered light of the cyanobacteria distribution area, and combine the infrared thermal imaging technology to perform multi-spectral fusion to generate high-quality image data; then, after the preprocessing steps of radiation correction, atmospheric correction and data enhancement, the image quality and the generalization ability of the model are significantly improved;

[0056] In the present application, a high-precision multi-spectral camera is carried on a satellite or unmanned aerial vehicle platform to obtain reflected light and scattered light data of the cyanobacteria distribution area from different angles and time periods to generate remote sensing data containing image and spectral information and tags Infrared thermal imaging technology is combined to perform multi-spectral fusion to generate multi-spectral images containing temperature information and retaining clear contours and tags ;

[0057] Further, the initial data obtained is preprocessed, and the preprocessing process includes data enhancement, data standardization, data size unification, image correction and color space conversion to improve the image quality and the generalization ability of the model;

[0058] Specifically, in the data acquisition step, a high-precision multi-spectral camera is carried on a satellite or unmanned aerial vehicle platform to obtain reflected light and scattered light data of the cyanobacteria distribution area from different angles and time periods to generate remote sensing data containing image and spectral information. These data are transmitted to the ground receiving station through signal modulation, and after demodulation and decompression processing, the preliminary optical image and tags is obtained. At the same time, a non-cooled infrared thermal imager carried by the satellite or unmanned aerial vehicle synchronously collects infrared thermal imaging data of the cyanobacteria bloom area.

[0059] The device is based on the principle of thermal radiation, which generates a gray-scale image reflecting the temperature distribution by detecting the 8-14 mu wavelength infrared radiation emitted from the surface of the water body and tags ;

[0060] Then image fusion is performed, and the specific process of image fusion includes:

[0061] First, the SIFT feature point matching is performed to spatially align the visible light image and the infrared thermal image to ensure that the pixel points at the same geographical coordinates are one-to-one corresponding, and the SIFT feature point matching is represented as:

[0062]

[0063]

[0064] wherein G(x,y,σ) is a Gaussian kernel, is the original image, and σ is the scale parameter;

[0065] For the obtained optical image, the gray level co-occurrence matrix (GLCM) is calculated for each color channel (R, G, B), and typical texture features are extracted to represent the morphological profile of cyanobacteria. The process of extracting typical texture features is represented as:

[0066]

[0067] wherein, represents the joint probability of gray levels i and j at direction θ and distance d, and the square term (i-j) 2 represents the amplified gray level difference, highlighting the texture changes in the cyanobacteria aggregation area;

[0068] For the obtained gray image, temperature gradient features are extracted, and the Otsu algorithm is used to locate the abnormal high temperature area CTAZ, reflecting the surface temperature changes caused by cyanobacteria aggregation. The Otsu algorithm is represented as:

[0069]

[0070] wherein, represents the maximum inter-class variance, respectively represent the pixel proportion of the foreground (class C0) and the background (class C1) after threshold t segmentation, 、 respectively represent the average gray values of the two types of pixels;

[0071] Then, the dynamic weight distribution formula is used to distribute the weight of the two images, and the pixel-level fusion formula is used to fuse the two images together to generate a multispectral image and label , represented as:

[0072]

[0073]

[0074]

[0075] wherein, represents the spectral response intensity of the visible light sensor at wavelength λ, represents the spectral response intensity of the infrared sensor at wavelength λ; represents the fusion weight of the visible light image, represents the fusion weight of the infrared image, represents the pixel value of the visible light image at coordinates (x, y), represents the pixel value of the infrared thermal imaging image at coordinates (x, y), represents the pixel value of the visible light image at coordinates (x, y).

[0076] Further, after obtaining the fusion image, the obtained fusion image is subjected to radiation correction, the sensor self noise, environmental interference and radiation distortion in the imaging process are eliminated, and the real radiation value of the blue algae is restored;

[0077] The step of performing radiation correction on the fusion image is represented as:

[0078]

[0079] Wherein, represents the radiation brightness of the wave band λ, DN represents the original numerical value recorded by the sensor, represents the gain coefficient, represents the bias value;

[0080] After radiation correction, atmospheric correction processing is performed, specifically, the influence of atmospheric scattering and absorption is eliminated by dark object subtraction (DOS), and the real reflectivity of the ground surface is inverted, represented as:

[0081]

[0082] Wherein, represents the ground reflectivity, represents the radiation brightness of the top layer of the atmosphere, represents the atmospheric path radiation, represents the solar irradiance, represents the solar zenith angle;

[0083] Further, in the step of pre-processing the obtained initial data in the present application:

[0084] Data enhancement: the data set is expanded and the generalization ability of the model is enhanced by means of inversion, rotation, random cropping and the like;

[0085] Data standardization: for the obtained RGB image, the mean and standard deviation of R, G and B channels are calculated, and the mean and variance method is used for standardization, which effectively avoids the problem of gradient explosion or disappearance of data gradient, and the standardization processing is represented as:

[0086]

[0087] Wherein, is the standardized pixel value, x is the original pixel value, μ is the mean of the channel, and σ is the standard deviation of the channel;

[0088] Data size unification: the data set is uniformly adjusted to the size of 256*256, which is convenient for subsequent operation of the data set;

[0089] Image correction: divide the image into multiple regions, calculate the corresponding gamma value according to the brightness of each region, calculate the gamma value of each region to correct, enhance the contrast of image details, weaken the dependence of the model on light, and the expression of calculating the gamma value of each region to correct is:

[0090]

[0091] Wherein, Indicates the pixel value of the transformed image, Indicates the pixel value of the original image, and gamma indicates the calculated gamma value.

[0092] Color space conversion: convert the RGB image into the HSV color space, so that the color feature and the brightness feature can be processed independently, facilitate the extraction and analysis of color information under different light conditions, and thus enhance the robustness of the model.

[0093] Therefore, for data acquisition and data processing thereof, the present application adopts a mode of remote sensing combined with infrared remote sensing and HSV color space, effectively solves the limitation of traditional methods on light and shooting angle under adverse weather such as rain, snow and haze, and improves the robustness and adaptability of the model.

[0094] Please continue to refer to Figure 1 And Figure 2 The monitoring method provided by the embodiments of the present application further comprises:

[0095] Step S2, input the image data as the input of the improved FAST algorithm to obtain candidate mutation points, filter the candidate mutation points based on a dynamic threshold and remove isolated noise points through DBSCAN clustering to output a key region with a dense spatial distribution;

[0096] Step S2 is the first mutation point detection, through the mutation point detection technology, the key region is quickly located and the calculation amount is reduced, and the adaptability of the model is enhanced;

[0097] Specifically, in step S2, the improved FAST corner detection algorithm is used to extract local texture features in combination with GLCM: the data is input into the improved FAST algorithm to obtain candidate mutation points V, and then the contrast of the region is quantified to distinguish the light and dark, the entropy value represents the texture complexity, and the correlation describes the structural regularity, and the texture features Fi={Contrasti, Entropyi, Correlationi}, Contrasti represents the contrast feature, Entropyi represents the entropy value feature, and Correlationi represents the correlation feature.

[0098] Wherein:

[0099] The contrast is represented as:

[0100]

[0101] The entropy value is represented as:

[0102]

[0103] The correlation is represented as:

[0104]

[0105] In the formula, represents the joint probability of gray levels i and j under the direction theta and the distance d; the square term represents amplifying the gray difference and highlighting the texture change of the cyanobacteria aggregation area; respectively represent the standard deviation of the x-axis and the y-axis; , respectively represent the mean value of gray levels i and j under the direction theta and the distance d;

[0106] Subsequently, V is screened based on the dynamic threshold tau, and isolated noise points are removed through DBSCAN clustering, and the spatially dense distribution of key regions is output .

[0107] Step S3, based on the screened key region, fusion visible light texture features and infrared thermal imaging temperature gradient to generate joint feature map, screening out abnormal area;

[0108] Specifically, in step S3 of the present application, based on the key region screened in step S1 , fusion visible light texture features and infrared thermal imaging temperature gradient to generate joint feature map M(x,y), screening out abnormal area;

[0109] In the formula,

[0110]

[0111] In the formula, represents the gradient of the temperature field in the horizontal direction, which is used to quantify the local temperature change rate; T represents the temperature; (x+1,y) and (x-1,y) represent the coordinates of the right and left adjacent pixels of the current pixel (x,y).

[0112] Step S4, using convolutional neural network to extract spatial features, and combining long short-term memory network to model time sequence and output time sequence features, combining spatial-temporal attention mechanism to dynamically calculate weight, based on the calculated weight to weight fusion spatial features and time sequence features;

[0113] In step S4 of the present application, spatial features are extracted based on the abnormal area obtained in step S3, i.e., excluding the part of the abnormal area, by using a convolutional neural network;

[0114] Step S5, introduce one-class support vector machine to identify abnormal points of the fusion features, and construct a bidirectional spatio-temporal prediction network based on attention mechanism, in the constructed prediction network, the global spatio-temporal dependence is extracted through the multi-head self-attention mechanism, the features are decoded in the forward / inverse time sequence, the cyanobacteria density distribution prediction value is generated, and the prediction result is obtained;

[0115] Step S6, combining the historical data and the prediction result of the cyanobacteria distribution area, a spatio-temporal evolution graph of cyanobacteria diffusion trend for visual analysis is generated.

[0116] Please refer to Figure 3 In an implementation manner of the present application, the step of extracting spatial features by using a convolutional neural network in step S4 specifically includes:

[0117] Step S41, constructing a double-branch convolutional neural network for extracting multi-modal features of cyanobacteria based on HSV color space;

[0118] Specifically, when extracting spatial features, the input is an HSV image of 256x256x3, and the parameters are The H (hue), S (saturation), and V (lightness) channels are normalized to [0, 1] respectively, wherein the normalization processing is represented as:

[0119]

[0120] In the formula, Hnorm represents the normalized hue feature, Snorm represents the normalized saturation feature, Vnorm represents the normalized lightness feature;

[0121] Further, in the constructed double-branch convolutional neural network, the hue branch layer (H channel) adopts 5x5 dilated convolution (dilated rate=2) to expand the receptive field and capture the hue boundary feature, and outputs a 128x128x32 feature map after a 3x3 pooling layer (step length=2); the saturation-lightness joint branch layer (S+V channel) extracts texture features through 3x3 convolution, and introduces an SE attention module to dynamically enhance key areas, wherein the weight matrix 、 R represents that the element data in the matrix is a real number, and a 128x128x64 feature map is output; the features of the two branches are spliced into 128x128x96, compressed to 128x128x128 through 1x1 convolution, and the cross-modal interaction is enhanced;

[0122] wherein, in the step of introducing the SE attention module to dynamically enhance the key region, is represented as:

[0123]

[0124] wherein, W1 and W2 represent weight matrices; S c represents the attention weight map processed by the SE attention module, used to dynamically enhance the features of the key region; F SV represents the feature map of the saturation-lightness joint branch layer;

[0125] Step S42, by constructing a three-level feature pyramid of low-level features, medium-level features and high-level features, the spatial information of different scales is fused; wherein, in the low-level feature layer, the basic visual information including hue boundary and local texture in the image is captured; in the medium-level feature layer, the algal spot morphology and saturation gradient are extracted by 3×3 transpose convolution upsampling, and the detail and semantic information are balanced; in the high-level feature layer, the global average pooling is performed on the medium-level features, and the channel is compressed to generate a compact global feature vector, thereby enhancing the generalization ability of the model to complex scenes;

[0126] Specifically, in the low-level feature layer (128×128×128), the basic visual information including hue boundary and local texture in the image is captured, represented as:

[0127]

[0128] wherein, represents the weight matrix for extracting low-level features;

[0129] Specifically, in the medium-level feature layer (64×64×256), the algal spot morphology and saturation gradient are extracted by 3×3 transpose convolution (stride=2) upsampling, represented as:

[0130]

[0131] Specifically, in the high-level feature layer (32×32×512), the channel is compressed after global average pooling (GAP) of the medium-level features, represented as:

[0132] Step S43, after processing the data by the three-level pyramid, the data is fused across levels, the high-level features are upsampled to 64×64, spliced with the medium-level features and the low-level features to form 64×64×864, and then compressed to 64×64×512 data by 1×1 convolution, thereby eliminating redundant information and greatly reducing the calculation amount.

[0133] Specifically, the data fusion across levels is represented as:

[0134] ​

[0135] wherein, is the fusion feature, is the weight matrix, BN represents batch normalization, accelerates training and prevents gradient disappearance; ReLU is an activation function that enhances non-linear expression capability; F mid is the low-level feature input, is a 3x3 transpose convolution with a stride of 2, upsampling the resolution from 64x64 to 128x128; represents global average pooling on the middle-level feature (64x64x256), with an output of 1x1x256; represents a 1x1 convolution kernel that expands the number of channels from 256 to 512 to enhance feature expression capability; represents upsampling the high-level feature 32x32x512 to 64x64x512; Concat represents three groups of features concatenated along the channel dimension.

[0136] Further, in step S4, the step of modeling the time series with the combination of long short-term memory networks to output time sequence features includes: expanding the 64x64x512 spatio-temporal feature map sequence along the spatial dimension into 64x64 independent time series, and inputting the sequence into a bidirectional LSTM network: the forward LSTM layer processes the sequence in chronological order, generating forward hidden state ; the backward LSTM layer processes the sequence in reverse chronological order, generating backward hidden state ;

[0137] Finally, the forward and backward hidden states of each spatial position are concatenated along the channel dimension to output a 64x64x1024 spatio-temporal coupling feature matrix, fully preserving the spatio-temporal evolution law of cyanobacterial diffusion.

[0138] Further, in step S4, the step of weighting and fusing spatial features and time sequence features based on calculation includes:

[0139] The 64x64x512 spatial features output by the CNN are expanded to 1024 channels through a 1x1 convolution, aligning with the 64x64x1024 time sequence feature dimension of the LSTM;

[0140] Combining the spatial-temporal attention mechanism to dynamically calculate the weight : through a 3x3 convolution of 512 channels and ReLU activation to extract joint features, and then through a 1x1 convolution to compress the channels to generate an attention map, after Softmax normalization, the fusion weight of each position is obtained α∈[0,1], the weight α is represented as:

[0141]

[0142] The CNN feature and the LSTM feature are fused by weighting with the weight a and the weight 1-a, and the fused feature is is represented as:

[0143] ;

[0144] In the formula, represents a spatial feature map output by the CNN, and the dimension is 64x64x512; represents a time sequence feature map output by the LSTM, and the dimension is 64x64x1024; represents that the CNN feature and the LSTM feature are spliced along the channel dimension to generate a joint feature map of 64x64x1536; represents a 3x3 convolution kernel, the output channel number is 512, and the joint space-time feature is extracted; represents a 1x1 convolution kernel, the output channel number is 1, the feature is compressed to generate an attention weight map; Softmax represents a normalization operation, the weight is mapped to the interval [0, 1], and the sum of the weights is ensured to be equal to 1, a represents a dynamic weight map, the dimension is 64x64x1, and represents an element-by-element multiplication, which is used to weight the CNN and LSTM features at each position.

[0145] Further, in step S5, the step of introducing the one-class support vector machine to identify the abnormal points of the fused feature includes:

[0146] The output of the fused CNN-LSTM space-time feature As a variant point, the spatial form and the time dynamics of the integrated multi-modal data are integrated to form a joint feature map containing space-time coupling information. In addition, the fused feature also retains the temperature gradient information of the infrared thermal imaging and the texture features of the visible light image, providing a data basis for multi-modal complementary analysis;

[0147] The one-class support vector machine (One-Class SVM) based anomaly detection algorithm regards each position of the 64x64 grid as an independent sample, and a total of 4096 feature vectors are generated, each feature vector has a dimension of 1024;

[0148] The principal component analysis PCA is used to reduce the feature vector to 32 dimensions, and the decision function is used for variant point detection. When f(x)=-1, the variant point is marked;

[0149] Anomaly probability heat map output based on variant point detection The candidate region coordinate set satisfying is extracted The candidate region coordinate set C is used to locate the high-risk area, further analyze and process, and better understand the diffusion trend;

[0150] Wherein, the constructed anomaly detection algorithm is represented as:

[0151] The decision function is represented as: ;

[0152] In the formula, The coordinates of the candidate region are represented as: The gradient of the temperature field in the horizontal direction is represented as T, and T represents the temperature; (x+1, y) represents the coordinates of the right adjacent pixel of the current pixel (x, y), (x-1, y) represents the coordinates of the left adjacent pixel of the current pixel (x, y), and w is the weight vector of the hyperplane, is the mapping of the data point x in the high-dimensional space, and rho is a bias term, and sgn() represents a sign function.

[0153] In steps S2 and S4 of the present application, the variation point detection is adopted, and a double variation point detection mechanism is adopted, so as to reduce the calculation amount, reduce the calculation cost, reduce the interference of non-natural factors, remove the variation points, and improve the accuracy of the model.

[0154] Further, in step S5, based on the space-time coupling feature matrix , a bidirectional space-time prediction network based on the attention mechanism is constructed; the space-time feature sequence of 30 consecutive days is input in time sequence , and a cascaded Transformer-LSTM architecture is used for prediction;

[0155] Further, in step S5, the global space-time dependency is extracted through the multi-head self-attention mechanism, and is represented as:

[0156] ;

[0157] In the formula, Q, K, and V are query matrix, key matrix, and value matrix, respectively; represents the feature dimension; and T represents the transpose of the matrix;

[0158] The generated cyanobacterial density distribution prediction value is represented as .

[0159] In step S6, based on the historical data statistics and ecological carrying capacity evaluation, the cyanobacterial bloom outbreak threshold is defined, and is represented as: , wherein, is the historical average cyanobacterial density, is the standard deviation; when the prediction value , the grid is marked as a high-risk area, an early warning heat map is output, and the risk level is marked.

[0160] Further, the prediction module in step S5 of the embodiment of the present application introduces a comparison mechanism, and the image obtained by the time difference is continuously corrected to make accurate prediction, reducing the data requirement of the traditional prediction method.

[0161] While the embodiments of the application have been disclosed in connection with the specification and examples herein, it should be understood that it can be employed in various other ways. Additional modifications can be made by those skilled in the art without departing from the scope of the present application. Therefore, the present application is not limited to the specific details and examples shown and described herein.

Claims

1. A cyanobacteria bloom monitoring method based on multimodal data fusion and deep learning, characterized by: The following steps are involved: Step S1: using a multispectral camera to collect reflected light and scattered light from the cyanobacteria distribution area, and combining it with infrared thermal imaging technology to perform multispectral fusion to generate image data; Step S2: Using the image data as input for the improved FAST algorithm to obtain candidate variation points, screening the candidate variation points based on a dynamic threshold and removing isolated noise points through DBSCAN clustering, and outputting key areas with dense spatial distribution; Step S3: Based on the selected key areas, the visible light texture features and the infrared thermal imaging temperature gradient are integrated to generate a joint feature map to screen out abnormal areas; Step S4: Using a convolutional neural network to extract spatial features based on the abnormal area, and combining it with a long short-term memory network to model the time series and output temporal features, dynamically calculating weights using a spatial-temporal attention mechanism, and weightedly fusing spatial and temporal features based on the calculated weights; Step S5: Introduce a single-class support vector machine to identify outliers in the fused features and construct a bidirectional spatiotemporal prediction network based on the attention mechanism. In the constructed prediction network, the global spatiotemporal dependency is extracted through the multi-head self-attention mechanism, and the features are decoded in forward / reverse time order to generate a predicted value of the cyanobacteria density distribution and obtain a prediction result. Step S6: Combining historical data of the cyanobacteria distribution area with the prediction results, a spatiotemporal evolution diagram of the cyanobacteria diffusion trend for visualization analysis is generated.

2. The cyanobacteria bloom monitoring method based on multimodal data fusion and deep learning according to claim 1 is characterized in that: In step S4, the step of extracting spatial features using a convolutional neural network includes: A two-branch convolutional neural network was constructed based on the HSV color space to extract multimodal features of cyanobacteria; By constructing a three-level feature pyramid of low-level features, mid-level features, and high-level features, spatial information of different scales is integrated; In the low-level feature layer, basic visual information including hue boundaries and local textures in the image is captured; In the mid-level feature layer, 3×3 transposed convolution upsampling is used to extract algae spot morphology and saturation gradient, balancing details and semantic information. In the high-level feature layer, the intermediate features are globally averaged and pooled, and the channels are compressed to generate a global feature vector. After processing the data through the three-level pyramid, the data is fused across levels, the high-level features are upsampled to 64×64, and concatenated with the mid-level features and low-level features to form 64×64×864, and then compressed to 64×64×512 through 1×1 convolution.

3. The cyanobacteria bloom monitoring method based on multimodal data fusion and deep learning according to claim 2 is characterized in that: In the constructed dual-branch convolutional neural network, the hue branch layer uses a 5×5 dilated convolution to expand the receptive field and capture the hue boundary features. After a 3×3 pooling layer, it outputs a 128×128×32 feature map. The saturation-lightness joint branch layer extracts texture features through a 3×3 convolution, and introduces the SE attention module to dynamically enhance the key areas, which can be expressed as: ; Among them, W1 and W2 represent weight matrices; S c represents the attention weight map after processing by the SE attention module, which is used to dynamically enhance the features of the key areas; F SV Feature map representing the saturation-lightness joint branch layer; weight matrix 、 , output a 128×128×64 feature map, R indicates that the element data in the matrix is ​​a real number; The two branch features are concatenated into 128×128×96 and compressed to 128×128×128 through 1×1 convolution.

4. The method for monitoring cyanobacteria blooms based on multimodal data fusion and deep learning according to claim 3, characterized in that: In step S4, the steps of modeling the time series and outputting the time series features by combining the long short-term memory network include: The 64×64×512 spatiotemporal feature map sequence is expanded into 64×64 independent time series along the spatial dimension and input into the bidirectional LSTM network: the forward LSTM layer processes the sequence in time order and generates the forward hidden state. ; The backward LSTM layer processes the sequence in reverse time order to generate the backward hidden state ; The forward and backward hidden states of each spatial position are concatenated along the channel dimension to output a 64×64×1024 spatiotemporal coupling feature matrix.

5. The method for monitoring cyanobacteria blooms based on multimodal data fusion and deep learning according to claim 4, characterized in that: In step S4, the step of weightedly fusing spatial features and temporal features based on the calculated weights includes: The 64×64×512 spatial features output by CNN are expanded to 1024 channels through 1×1 convolution, which aligns with the 64×64×1024 temporal feature dimensions of LSTM. The joint features are extracted through 512-channel 3×3 convolution and ReLU activation, and the attention map is generated by 1×1 convolution compression channel. After Softmax normalization, the fusion weight α∈[0,1] of each position is obtained. The weight α is expressed as: ; CNN features and LSTM features are weighted and fused according to weight α and weight 1-α. The weighted fused features Expressed as: ; Where, F cnn Represents the spatial feature map of CNN output, with a dimension of 64×64×512; Represents the temporal feature map of LSTM output, with a dimension of 64×64×1024; Indicates concatenating CNN and LSTM features along the channel dimension to generate a 64×64×1536 joint feature map; Represents a 3×3 convolution kernel with 512 output channels, used to extract spatiotemporal joint features; Represents a 1×1 convolution kernel with an output channel of 1, compressing features to generate an attention weight map; Softmax represents a normalization operation, mapping weights to the interval [0, 1] to ensure that the sum of weights is equal to 1; α represents a dynamic weight map with a dimension of 64×64×1; ⊙ represents element-by-element multiplication, which is used to weight the CNN and LSTM features at each position.

6. The method for monitoring cyanobacteria blooms based on multimodal data fusion and deep learning according to claim 5, characterized in that: In step S5, a single-class support vector machine is introduced to identify outliers based on the fused features, including: The output after fusion of CNN-LSTM spatiotemporal features Data as mutation points; An anomaly detection algorithm based on a single-class support vector machine is constructed. Each position of the 64×64 grid is regarded as an independent sample, generating a total of 4096 feature vectors, each with a dimension of 1024; Principal component analysis (PCA) is used to reduce the eigenvector to 32 dimensions, and a decision function is used to detect mutation points. When f(x) = -1, it is marked as a mutation point. Abnormal probability heat map based on the output of mutation point detection , extract satisfaction The candidate region coordinate set ; Among them, the constructed anomaly detection algorithm is expressed as: ; The decision function is expressed as: f(x)=sgn(w·Φ(x)-ρ); Where, Represents the coordinates of the candidate region; represents the horizontal gradient of the temperature field, T represents the temperature; (x+1,y) represents the coordinates of the adjacent pixel to the right of the current pixel (x,y), (x-1,y) represents the coordinates of the adjacent pixel to the left of the current pixel (x,y), w is the weight vector of the hyperplane, Φ(x) is the mapping of the data point x in the high-dimensional space, ρ is a bias term; sgn() represents the sign function.

7. The method for monitoring cyanobacteria blooms based on multimodal data fusion and deep learning according to claim 6, characterized in that: In step S5, the global spatiotemporal dependency is extracted through the multi-head self-attention mechanism, which is expressed as: ; Where Q is the query matrix, K is the key matrix, and V is the value matrix; k Represents feature dimension; T represents the transpose of the matrix; The generated predicted value of cyanobacteria density distribution is expressed as .

8. The method for monitoring cyanobacteria blooms based on multimodal data fusion and deep learning according to claim 7, characterized in that: In step S6, the cyanobacteria bloom threshold is defined , expressed as: ; in, is the historical average density of cyanobacteria, is the standard deviation; when the predicted value When the grid is marked as a high-risk area, the warning heat map is output , and mark the risk level.

Citation Information

Patent Citations

  • Cyanobacterial bloom cross-water-area prediction method and device based on transfer learning

    CN115330053A

  • Automatic detection of sea floating objects from satellite imagery

    US20240013531A1