A method, medium and system for identifying and warning of green tides in coastal high-point monitoring

Through deep learning technology, the green tide semantic segmentation model is constructed, combined with the algorithm model of PSPNet and ShuffleNetV2, and the ECA attention mechanism and multi-scale feature fusion strategy are introduced, which solves the problem of inefficient green tide monitoring in the existing technology and achieves efficient and accurate green tide recognition and early warning.

CN119540781BActive Publication Date: 2025-06-20BEIHAI FORECASTING CENT OF STATE OCEANIC ADMINISTRATION ((QINGDAO MARINE FORECASTING STATION OF STATE OCEANIC ADMINISTRATION) (QINGDAO MARINE ENVIRONMENT MONITORING CENT OF STATE OCEANIC ADMINISTRATION))
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510088231.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-20
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The existing visual interpretation and manual threshold setting methods are inefficient in green tide monitoring, with great artificial influence, making it difficult to achieve real-time and efficient monitoring.

Method used

Deep learning technology is used to build a green tide semantic segmentation model. By constructing a high-quality green tide image dataset, a deep learning algorithm model combining PSPNet and ShuffleNetV2 is designed, and an ECA attention mechanism and multi-scale feature fusion strategy are introduced to achieve accurate identification and early warning of green tides.

Benefits of technology

It significantly improves the accuracy and efficiency of green tide recognition, can better capture the characteristics of green tides in complex sea surface environments, reduce the instability of single-frame recognition, provide more stable and reliable monitoring results, and establish a complete early warning mechanism to promptly detect the development trend of green tides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540781B_ABST
    Figure CN119540781B_ABST
Patent Text Reader

Abstract

The present invention provides a method, medium and system for monitoring and identifying and warning of green tides at high points along the coast, belonging to the technical field of green tide identification and warning, including: collecting a large number of green tide images taken by high-point monitoring, performing JSON data annotation on the green tide contour, converting the JSON annotation into a PNG mask image, and dividing it into a training set, a validation set and a test set. Designing a deep learning green tide semantic segmentation algorithm, selecting PSPNet and replacing the backbone network with ShuffleNetV2, adding an ECA attention mechanism and multi-scale feature fusion, and jointly using FocalLoss and BCELoss to solve the problem of class imbalance. Training the model using the training data set and updating the weights through error backpropagation. Evaluating the performance using the validation set, adjusting the hyperparameters to optimize the model, and selecting the optimal model. Inputting the image to be identified into the model to obtain the green tide segmentation mask, and judging whether it exceeds the threshold or shows abnormal changes according to the segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of green tide identification and early warning, and more specifically, relates to a method, medium, and system for identifying and warning green tides through high-point monitoring along the coast. Background Art

[0002] The green tide phenomenon in coastal areas is an abnormal phenomenon in the marine environment, usually caused by the massive reproduction of specific algae. These algae usually release a large amount of harmful biotoxins, which have a negative impact on the balance of the marine ecosystem and biodiversity. In addition, the green tide may also cause excessive oxygen consumption, resulting in hypoxia in the water area, affecting the survival and reproduction of fishery resources, and having a long-term adverse impact on the marine ecosystem.

[0003] In the past few decades, due to the impact of global climate change and human activities, the frequent occurrence and expansion of the green tide have become a global problem. Traditional monitoring methods are often limited by aspects such as labor costs, time delays, and monitoring ranges. Therefore, a more intelligent, efficient, and real-time monitoring method is needed to detect and handle potential green tide problems in a timely manner.

[0004] In terms of green tide information extraction, currently, most methods use visual interpretation and manual threshold setting. This method relies on expert knowledge and is easy to operate, but has disadvantages such as low efficiency and large human influence. Especially during green tide emergencies, these drawbacks are more obvious. Summary of the Invention

[0005] In view of this, the present invention provides a method, medium, and system for identifying and warning green tides through high-point monitoring along the coast, which can solve the technical problem of low efficiency existing in the existing visual interpretation and manual threshold setting methods.

[0006] The present invention is implemented as follows:

[0007] The first aspect of the present invention provides a method for identifying and warning green tides through high-point monitoring along the coast, which includes the following steps:

[0008] S10. Construct a green tide semantic segmentation image dataset: Collect a large number of green tide images taken by high-point monitoring, and perform JSON data annotation on the green tide contours in the green tide images to obtain the annotated image data;

[0009] S20. Data processing: Convert the JSON format data annotation into a PNG mask image, and divide it into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10% to obtain a green tide semantic segmentation image training dataset;

[0010] S30. Design the prototype of the deep learning green tide semantic segmentation algorithm model: Select the PSPNet semantic segmentation algorithm, replace the backbone network with ShuffleNetV2, and combine the ECA attention mechanism equations with the multi-scale feature fusion strategy, and jointly use the FocalLoss and BCELoss loss functions to solve the class imbalance problem and improve the generalization ability of the model;

[0011] S40. Train the green tide semantic segmentation model: Use the green tide semantic segmentation image training dataset to train the prototype of the deep learning green tide semantic segmentation algorithm model, and update the weights of the prototype of the deep learning green tide semantic segmentation algorithm model through error backpropagation;

[0012] S50. Model evaluation and optimization: Use the validation set to evaluate the performance of the trained prototype of the deep learning green tide semantic segmentation algorithm model, adjust the hyperparameters to optimize the model performance, and select the optimal model as the green tide semantic segmentation model;

[0013] S60. Obtain the image to be recognized: Obtain the real-time image data from the coastal high-point monitoring system as the image to be recognized;

[0014] S70. Green tide recognition: Preprocess the image to be recognized, input it into the trained green tide semantic segmentation model, and obtain the segmentation mask of the green tide;

[0015] S80. Early warning judgment: According to the segmentation mask, judge whether the green tide area exceeds the preset threshold or shows abnormal changes. If the conditions are met, trigger the early warning system and notify the relevant management personnel.

[0016] Among them, the ECA attention mechanism equations include an adaptive convolution kernel size calculation equation, a global average pooling equation, a one-dimensional convolution operation equation, a Sigmoid activation function equation, and an attention weighting equation.

[0017] 1. Adaptive convolution kernel size calculation equation: Used to calculate the convolution kernel size adapted to the number of input channels.

[0018] 2. Global average pooling equation: Compresses the input feature map into a one-dimensional vector.

[0019] 3. One-dimensional convolution operation equation: Performs a one-dimensional convolution operation using the calculated adaptive convolution kernel size.

[0020] 4. Sigmoid activation function equation: Maps the output of the one-dimensional convolution to between 0 and 1 to obtain the attention weight.

[0021] 5. Attention weighting equation: Applies the obtained attention weight to the original input feature map.

[0022] These five equations form the complete ECA attention mechanism, which can effectively capture the dependencies between channels and improve the feature extraction ability of the model.

[0023] Furthermore, the equation for calculating the adaptive convolution kernel size is specifically expressed as:

[0024] ;

[0025] In the formula: is the number of channels of the input feature map; is an adjustable scaling factor, optimized within the range of [0.1, 10] through grid search; is the bias term, optimized within the range of [-5, 5] through random search; is to round down to the nearest odd number.

[0026] Among them, the global average pooling equation is specifically expressed as:

[0027] ;

[0028] In the formula: is the global average pooling result of the th channel; , are the height and width of the input feature map; is the value of the input feature map at the th channel and position .

[0029] Among them, the one-dimensional convolution operation equation is specifically expressed as:

[0030] ;

[0031] In the formula: is the one-dimensional convolution output of the th channel; is the adaptively calculated convolution kernel size; is the learnable one-dimensional convolution kernel weight, optimized through backpropagation; is the globally average pooled result after circular padding.

[0032] Among them, the Sigmoid activation function equation is specifically expressed as:

[0033] ;

[0034] In the formula: is the attention weight of the th channel; is the steepness parameter of the Sigmoid function, optimized through gradient descent; It is the central offset of the Sigmoid function and is optimized by gradient descent.

[0035] Among them, the attention weighting equation is specifically expressed as:

[0036] ;

[0037] In the formula: is the value of the output feature map at the th channel and position ; is the value of the input feature map at the th channel and position ; is the attention weight of the th channel; is a small constant to prevent division by zero error, and the default setting is ; , is a learnable spatial attention parameter and is optimized by backpropagation; is the hyperbolic tangent function used to introduce spatial attention.

[0038] Among them, the formula for combining the ECA attention mechanism equations with the multi-scale feature fusion strategy is expressed as follows:

[0039] Step 1, multi-scale feature extraction, which is specifically expressed as:

[0040] ;

[0041] In the formula: : The feature map of the th layer; : The feature extraction function of the th layer; : The input image; : The total number of layers;

[0042] Step 2, ECA attention mechanism, applied to each scale, which is specifically expressed as:

[0043] ;

[0044] In the formula: represents the ECA attention mechanism equations;

[0045] Step 3, multi-scale feature fusion:

[0046] ;

[0047] In the formula: : After being processed by the ECA attention mechanism, the Feature map of the layer; : An upsampling function that upsamples the feature map to the same size as the highest-resolution feature map; : The fusion weight of the layer features, learned to satisfy ;

[0048] Step 4, generation of the final feature map, specifically expressed as:

[0049] ;

[0050] where is a learnable transformation function for adjusting the number of channels of the fused features.

[0051] In the formula: : The number of channels of the layer feature map; , : The height and width of the layer feature map; , : The calculation parameter of the ECA adaptive convolution kernel size of the layer; : The one-dimensional convolution kernel weight of the layer; , : The Sigmoid function parameter of the layer; : The small constant of the layer to prevent division by zero error; , : The spatial attention parameter of the layer.

[0052] Specifically, the step S10 specifically includes:

[0053] Step 101, use a high-point monitoring system to collect a large number of green tide images under different times and weather conditions, ensuring that the image resolution is not lower than 1920×1080 pixels;

[0054] Step 102, preprocess the collected images, including denoising, color equalization, and geometric correction;

[0055] Step 103, use a professional image annotation tool to accurately contour-label the green tide area in the green tide image to generate a JSON format annotation file;

[0056] Step 104, have multiple experienced experts participate in the annotation process to ensure the accuracy and consistency of the annotation;

[0057] Step 105: Pair and save the original images and the corresponding JSON annotation files to form a complete annotation dataset;

[0058] Step 106: Collect no less than 10,000 valid green tide images, ensure that the proportion distribution of the green tide coverage area in the images is uniform, and there are appropriate sample distributions from 1% to 90%.

[0059] Among them, the step S20 specifically includes:

[0060] Step 201: Develop a data conversion script to convert the JSON format annotation file into a binary PNG mask image, with the green tide area marked as 255 and the non-green tide area marked as 0;

[0061] Step 202: Use the OpenCV library to read the original image and the corresponding mask image, and adjust them to a unified size of 512×512 pixels;

[0062] Step 203: Use the stratified random sampling method to divide the dataset into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10%;

[0063] Step 204: Use the green tide coverage rate as the stratification basis to divide the dataset into three levels: low coverage rate, medium coverage rate, and high coverage rate, and then perform random sampling within each level;

[0064] Step 205: Save the divided dataset in a format suitable for reading by the deep learning framework, such as TFRecord or HDF5 file;

[0065] Step 206: Compress and store the image data, using a lossless compression algorithm such as PNG compression to reduce the storage space and reading time while ensuring the image quality.

[0066] Among them, the step S30 specifically includes:

[0067] Step 301: Select PSPNet as the basic semantic segmentation algorithm and replace the backbone network of PSPNet with the lightweight ShuffleNetV2;

[0068] Step 302: Integrate the ECA mechanism into each stage of ShuffleNetV2 to enhance the model's ability to model the inter-channel dependence relationship;

[0069] Step 303: Design a multi-scale feature fusion strategy, extract feature maps from different stages of ShuffleNetV2, adjust the number of channels through 1×1 convolution, use bilinear interpolation to upsample all feature maps to the same spatial resolution, and obtain the fused feature map through weighted summation;

[0070] Step 304: Design a loss function, and combine FocalLoss and BCELoss to solve the class imbalance problem and improve the generalization ability of the model;

[0071] Step 305: Linearly combine the two loss functions, and determine the initial value of the balance factor through cross-validation.

[0072] Among them, the step S40 specifically includes:

[0073] Step 401: Initialize the ShuffleNetV2 backbone network with the weights pre-trained on the ImageNet dataset, and use the He initialization method for the newly added layers;

[0074] Step 402: Set the training hyperparameters, use the Adam optimizer, set the initial learning rate to 0.001, the weight decay coefficient to 0.0001, and use the cosine annealing learning rate scheduling strategy;

[0075] Step 403: Implement data augmentation strategies, including random horizontal flipping, random vertical flipping, random rotation, random scaling, and random adjustment of brightness, contrast, and saturation;

[0076] Step 404: Start the iterative training process. In each epoch, input the augmented training data into the model, perform forward propagation calculation to obtain the prediction results, calculate the loss function value, use the automatic differentiation technique to calculate the gradient, and update the model parameters through the backpropagation algorithm;

[0077] Step 405: At the end of each epoch, evaluate the model performance on the validation set, and record key metrics such as accuracy and IoU;

[0078] Step 406: Implement an early stopping strategy. When the performance on the validation set has not improved for 5 consecutive epochs, end the training in advance;

[0079] Step 407: Adopt the gradient clipping technique to limit the L2 norm of the gradient within 5 to prevent the gradient explosion problem.

[0080] Among them, the step S50 specifically includes:

[0081] Step 501: Use the validation set to conduct a detailed evaluation of the trained model, calculate multiple evaluation metrics, including pixel-level accuracy, mean intersection over union, F1 score, precision, and recall;

[0082] Step 502: Conduct error analysis, classify and count the mispredicted samples in the validation set, and identify the main error types of the model;

[0083] Step 503: Analyze the attention distribution of the model through visualization technology to understand whether the image regions focused by the model are reasonable;

[0084] Step 504: Based on the error analysis results, adjust the model hyperparameters and structure, such as increasing the sampling weight of small targets, adding more fine-grained features in the feature pyramid, increasing the regularization strength, and adjusting the parameters of the ECA module;

[0085] Step 505: Implement 5-fold cross-validation, calculate the evaluation metrics for each fold, and take the average as the final performance metric;

[0086] Step 506: Conduct ensemble learning on the best-performing model configurations, adopt model averaging and test-time augmentation techniques, select 3-5 of the best-performing model configurations in cross-validation, and take the average of the prediction results for each test sample.

[0087] Among them, the specific steps of step S60 include:

[0088] Step 601: Select a suitable high-point location in the coastal area and install a high-resolution camera device with a resolution not lower than 4K and a field of view angle not less than 90 degrees;

[0089] Step 602: According to the seasonal characteristics of the occurrence of green tides, adjust the image acquisition frequency, set a higher acquisition frequency during the high-incidence period of green tides, such as once every 5 minutes, and reduce the frequency to once every 30 minutes in other periods;

[0090] Step 603: Perform real-time preprocessing on the collected original images, including denoising, color correction, and geometric correction. Use a Gaussian filter to remove image noise, perform automatic white balance and contrast enhancement, and apply a lens distortion correction algorithm;

[0091] Step 604: Use a high-speed network to transmit the preprocessed images to the central processing server in real time and compress the images using a lossless compression algorithm;

[0092] Step 605: On the server side, store the received images in a high-performance distributed file system;

[0093] Step 606: Develop an automated quality detection algorithm to evaluate each received image, including clarity evaluation, exposure check, and cloud and fog occlusion detection;

[0094] Step 607: For images that fail the quality detection, the system automatically adjusts the camera parameters and re-collects them. At the same time, establish an artificial review mechanism to regularly spot-check the automatically collected images.

[0095] Among them, the specific steps of step S70 include:

[0096] Step 701: Adjust the acquired image to the input size of 512×512 pixels required by the model, scale it using the bicubic interpolation algorithm, and then perform image normalization;

[0097] Step 702: Input the preprocessed image into the selected optimal green tide semantic segmentation model, and use inference acceleration frameworks such as TensorRT or ONNXRuntime to optimize the model inference speed;

[0098] Step 703: During the inference process, enable half-precision calculation. For large-size images, adopt the sliding window technique to divide the image into multiple overlapping small blocks, perform predictions separately, and then splice and restore the results to the original size;

[0099] Step 704: Apply threshold segmentation to the probability map output by the model, and use morphological operations to refine the segmentation result to eliminate small noise regions and fill small holes;

[0100] Step 705: Apply conditional random field as a post-processing step to optimize the accuracy of the segmentation boundary;

[0101] Step 706: Perform time series analysis on the recognition results at multiple consecutive time points, and use the Kalman filter to smooth the green tide boundary to reduce the instability of single-frame recognition;

[0102] Step 707: Overlay the recognition result on the original image to generate a vectorized green tide contour, and save the recognition result in the GeoTIFF format, including geographic coordinate information.

[0103] Among them, the specific steps of step S80 include:

[0104] Step 801: Set warning thresholds at multiple levels, including mild warning, moderate warning, and severe warning, and these thresholds are dynamically adjusted according to the characteristics of different sea areas and seasonal changes;

[0105] Step 802: Use time series analysis techniques, such as autoregressive integrated moving average model, to model the change trend of the green tide area and predict the short-term development trend of the green tide;

[0106] Step 803: Use spatial autocorrelation analysis methods, such as Moran's index, to evaluate the spatial aggregation degree of the green tide;

[0107] Step 804: Establish a multi-factor evaluation model based on fuzzy logic. After standardizing each index, the comprehensive risk level is obtained through fuzzy rule inference;

[0108] Step 805: Use the local outlier factor algorithm to detect abnormal patterns in the development of the green tide, set the local outlier factor threshold, and situations exceeding this threshold will trigger special attention;

[0109] Step 806: When the comprehensive evaluation result exceeds the preset threshold or an anomaly is detected, automatically trigger the early warning system and send early warning messages to relevant management personnel through multiple channels;

[0110] Step 807: Push the early warning information to the visualization platform, use geographic information system technology to visually display the distribution of green tides and the early warning area on the electronic map, establish an early warning response mechanism, and start corresponding emergency plans according to different early warning levels.

[0111] The second aspect of the present invention provides a computer-readable storage medium, wherein program instructions are stored in the computer-readable storage medium, and when the program instructions run, they are used to execute the above-mentioned method for monitoring and identifying green tides at coastal high points and giving early warnings.

[0112] The third aspect of the present invention provides a system for monitoring and identifying green tides at coastal high points and giving early warnings, which includes the above-mentioned computer-readable storage medium.

[0113] Compared with the prior art, the beneficial effects of the method, medium, and system for monitoring and identifying green tides at coastal high points provided by the present invention are as follows:

[0114] First of all, by constructing a green tide semantic segmentation image dataset with strong pertinence, this method significantly improves the effect of model training. By collecting a large number of green tide images taken by high-point monitoring and performing precise JSON data annotation, the quality and diversity of the dataset are ensured. This high-quality dataset provides a solid foundation for subsequent deep learning model training, enabling the model to better adapt to complex sea surface environments and different forms of green tides.

[0115] Secondly, the deep learning green tide semantic segmentation algorithm model designed by this method significantly improves the accuracy and efficiency of green tide identification. By combining PSPNet and ShuffleNetV2, and introducing the ECA attention mechanism and multi-scale feature fusion strategy, this method effectively improves the recognition ability of small-area green tides and green tide edge areas while ensuring the lightweight of the model. Especially the introduction of the ECA attention mechanism enables the model to better capture the dependencies between channels, improving the efficiency and accuracy of feature extraction.

[0116] Furthermore, this method makes full use of spatio-temporal information, significantly improving the stability and reliability of the monitoring results. By performing time series analysis on the recognition results of multiple consecutive time points and using the Kalman filter to smooth the green tide boundary, the instability of single-frame recognition is effectively reduced, and the continuity and accuracy of green tide monitoring are improved. This method enables the system to better cope with interference factors such as sea surface fluctuations and light changes, providing more stable and reliable monitoring results.

[0117] Finally, this method establishes a complete early warning mechanism, which can timely detect the development trend of green tides and issue early warnings. By setting multi-level early warning thresholds and combining time series analysis and spatial autocorrelation analysis, this method can comprehensively evaluate the development situation of green tides. In particular, the introduction of a multi-factor evaluation model based on fuzzy logic and the local outlier factor algorithm enables the system to more sensitively capture the abnormal patterns in the development of green tides and trigger early warnings in a timely manner. This early warning mechanism greatly improves the foresight and practicality of green tide monitoring and provides strong support for the prevention and control of green tides in coastal areas.

[0118] In summary, the method of the present invention solves the technical problem of low efficiency existing in the existing visual interpretation and manually setting thresholds. BRIEF DESCRIPTION OF THE DRAWINGS

[0119] Figure 1 is a flowchart of the method provided by the present invention;

[0120] Figure 2 is the flowchart of the ECA attention mechanism in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0121] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0122] As Figure 1 shown, it is a flowchart of a method for identifying and early warning green tides by coastal high-point monitoring provided by the present invention. This method includes the following steps:

[0123] S10. Construct a green tide semantic segmentation image dataset: Collect a large number of green tide images taken by high-point monitoring, and perform JSON data annotation on the green tide contours in the green tide images to obtain the annotated image data;

[0124] S20. Data processing: Convert the JSON format data annotation into a PNG mask image, and divide it into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10% to obtain a green tide semantic segmentation image training dataset;

[0125] S30. Design a prototype of a deep learning green tide semantic segmentation algorithm model: Select the PSPNet semantic segmentation algorithm, replace the backbone network with ShuffleNetV2, and combine the improved ECA attention mechanism equations with the multi-scale feature fusion strategy, and jointly use the FocalLoss and BCELoss loss functions to solve the class imbalance problem and improve the generalization ability of the model;

[0126] S40. Train the green tide semantic segmentation model: Use the green tide semantic segmentation image training dataset to train the prototype of the deep learning green tide semantic segmentation algorithm model, and update each weight of the prototype of the deep learning green tide semantic segmentation algorithm model through error backpropagation.

[0127] S50. Model evaluation and optimization: Use the validation set to evaluate the performance of the trained prototype of the deep learning green tide semantic segmentation algorithm model, adjust the hyperparameters to optimize the model performance, and select the optimal model as the green tide semantic segmentation model.

[0128] S60. Obtain the image to be recognized: Obtain real-time image data from the coastal high-point monitoring system as the image to be recognized.

[0129] S70. Green tide recognition: Preprocess the image to be recognized, input it into the trained green tide semantic segmentation model, and obtain the segmentation mask of the green tide.

[0130] S80. Early warning judgment: According to the segmentation mask, judge whether the green tide area exceeds the preset threshold or shows abnormal changes. If the conditions are met, trigger the early warning system and notify the relevant management personnel.

[0131] The specific implementation manners of the above steps are described in detail below:

[0132] The specific implementation manner of step S10 is: Construct a green tide semantic segmentation image dataset. The purpose of this step is to create a high-quality and diverse green tide image dataset to provide a basis for subsequent deep learning model training. The specific implementation process includes: First, use high-point monitoring systems (such as drones, high-tower cameras, etc.) to collect a large number of green tide images at different times and under different weather conditions to ensure the diversity and representativeness of the data. During the collection process, it should be noted that the resolution of the images is not less than 1920×1080 pixels to ensure sufficient detailed information. Second, preprocess the collected images, including operations such as denoising, color equalization, and geometric correction, to improve the image quality. Then, use professional image annotation tools (such as LabelMe or CVAT) to accurately contour-label the green tide areas in the green tide images to generate annotation files in JSON format. During the annotation process, multiple experienced experts should participate together to ensure the accuracy and consistency of the annotation. Finally, pair and save the original images and the corresponding JSON annotation files to form a complete annotated dataset. To ensure the scale and quality of the dataset, it is recommended to collect no less than 10,000 effective green tide images and ensure that the proportion distribution of the green tide coverage area in the images is uniform, with appropriate sample distributions from 1% to 90%.

[0133] The specific implementation of step S20 is: data processing. The purpose of this step is to convert the labeled data into a format suitable for deep learning model training and reasonably divide the dataset. The specific implementation process includes: First, develop a data conversion script to convert the JSON-formatted annotation file into a binary PNG mask image. During the conversion process, the green tide area is marked as 255 (white), and the non-green tide area is marked as 0 (black). Second, use the OpenCV library in Python to read the original image and the corresponding mask image, and adjust them to a unified size (such as 512×512 pixels) to meet the input requirements of the deep learning model. Then, use the stratified random sampling method to divide the dataset into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10%. Stratified sampling can ensure that the distribution of the green tide coverage area in each subset is similar to that of the overall dataset, improving the representativeness of the samples. During the division process, use the green tide coverage rate as the stratification basis, divide the dataset into three levels: low coverage rate (0-30%), medium coverage rate (30-70%), and high coverage rate (70-100%), and then perform random sampling within each level. Finally, save the divided dataset in a format suitable for deep learning frameworks (such as TensorFlow or PyTorch) to read, such as TFRecord or HDF5 files. To improve the data loading efficiency, the image data can be compressed and stored, using a lossless compression algorithm such as PNG compression to reduce the storage space and reading time while ensuring the image quality.

[0134] The specific implementation of step S30 is as follows: Design the prototype of the deep learning green tide semantic segmentation algorithm model. The purpose of this step is to construct an efficient and accurate green tide semantic segmentation model architecture. The specific implementation process includes: First, select PSPNet (Pyramid Scene Parsing Network) as the basic semantic segmentation algorithm because PSPNet can effectively capture context information at different scales and is suitable for processing complex green tide scenarios. Second, replace the backbone network of PSPNet with the lightweight ShuffleNetV2 to improve the computational efficiency and real-time performance of the model. ShuffleNetV2 significantly reduces the amount of computation and the number of parameters while maintaining good performance through channel splitting and shuffling operations. Then, integrate the ECA (Efficient Channel Attention) mechanism into each stage of ShuffleNetV2 to enhance the model's ability to model the interdependencies between channels. The ECA mechanism achieves efficient channel attention calculation by adaptively determining the size of the one-dimensional convolutional kernel. Next, design a multi-scale feature fusion strategy to weight-fuse feature maps at different levels to comprehensively utilize the detailed information of the low-level and the semantic information of the high-level. Specifically, extract feature maps from different stages of ShuffleNetV2, adjust the number of channels through 1×1 convolution, then use bilinear interpolation to upsample all feature maps to the same spatial resolution, and finally obtain the fused feature map through weighted summation. The weight coefficients are learned through backpropagation. Finally, design a loss function that combines FocalLoss and BCELoss (Binary Cross Entropy Loss) to address the class imbalance problem and improve the generalization ability of the model. The expression of FocalLoss is: , where is the probability predicted by the model, is the class weight, is the focusing parameter (usually set to 2). The expression of BCELoss is: BCE ( y , y ^ )=-[ y log ( y ^ )+(1- y ) log ( 1- y ^ )] , where is the true label, is the predicted value of the model. Linearly combine the two loss functions: , where is the balance factor, determined by cross-validation, and the initial value can be set to 0.5.

[0135] The following is the complex formula representation of the improved ECA attention mechanism equations, including detailed parameter, variable explanations, and calculation steps:

[0136] 1. The equation for calculating the adaptive convolution kernel size is specifically expressed as:

[0137] ;

[0138] In the formula: is the number of channels of the input feature map; is the adjustable scaling factor, optimized within the range of [0.1, 10] through grid search; is the bias term, optimized within the range of [-5, 5] through random search; is the floor function to the nearest odd number.

[0139] 2. The global average pooling equation is specifically expressed as:

[0140] ;

[0141] In the formula: is the global average pooling result of the th channel; , are the height and width of the input feature map; is the value of the input feature map at the th channel and position .

[0142] 3. The one-dimensional convolution operation equation is specifically expressed as:

[0143] ;

[0144] In the formula: is the one-dimensional convolution output of the th channel; is the adaptively calculated convolution kernel size; is the learnable one-dimensional convolution kernel weight, optimized through backpropagation; is the globally averaged pooled result after circular padding.

[0145] 4. The Sigmoid activation function equation is specifically expressed as:

[0146] ;

[0147] In the formula: is the attention weight of the th channel; is the steepness parameter of the Sigmoid function, optimized by gradient descent; is the central offset of the Sigmoid function, optimized by gradient descent.

[0148] 5. The attention-weighted equation is specifically expressed as:

[0149] ;

[0150] In the formula: is the value of the output feature map at the th channel and position ; is the value of the input feature map at the th channel and position ; is the attention weight of the th channel; is a small constant to prevent division-by-zero errors, usually set to ; , is a learnable spatial attention parameter, optimized by backpropagation; is the hyperbolic tangent function, used to introduce spatial attention.

[0151] The following is a description of the above calculation steps:

[0152] 1. Calculate using the adaptive convolution kernel size calculation equation;

[0153] 2. Apply global average pooling to the input feature map to obtain ;

[0154] 3. Use the one-dimensional convolution operation equation to perform convolution on to obtain ;

[0155] 4. Input into the Sigmoid activation function to obtain the channel attention weight ;

[0156] 5. Use the attention-weighted equation to apply to the input feature map to obtain the output feature map ;

[0157] This ECA attention mechanism equation system introduces additional parameters and a spatial attention mechanism to improve the model's expressive ability and flexibility. By optimizing these parameters, it can better adapt to different data distributions and task requirements.

[0158] The formula combining the ECA attention mechanism equation set with the multi-scale feature fusion strategy is as follows:

[0159] 1. Multi-scale feature extraction, specifically expressed as:

[0160] ;

[0161] In the formula: : The feature map of the th layer; : The feature extraction function (such as convolution, pooling, etc.) of the th layer; : The input image; : The total number of layers.

[0162] 2. ECA attention mechanism (applied to each scale), specifically expressed as:

[0163] ;

[0164] Among them, the function includes the following steps:

[0165] a) Adaptive convolution kernel size calculation:

[0166] ;

[0167] b) Global average pooling:

[0168] ;

[0169] c) One-dimensional convolution operation:

[0170] ;

[0171] d) Sigmoid activation:

[0172] ;

[0173] e) Attention weighting:

[0174] ;

[0175] 3. Multi-scale feature fusion:

[0176] ;

[0177] In the formula: : The feature map of the th layer after being processed by the ECA attention mechanism; : The upsampling function that upsamples the feature map to the same size as the highest-resolution feature map; : The The fusion weights of the layer features are obtained through learning and satisfy .

[0178] 4. Generation of the final feature map, which is specifically expressed as:

[0179] :

[0180] where is a learnable transformation function, such as a 1x1 convolution, used to adjust the number of channels of the fused features.

[0181] In the formula: : The number of channels of the -th layer feature map; , : The height and width of the -th layer feature map; , : The calculation parameters of the ECA adaptive convolution kernel size of the -th layer; : The one-dimensional convolution kernel weights of the -th layer; , : The Sigmoid function parameters of the -th layer; : The small constant of the -th layer to prevent division by zero errors; , : The spatial attention parameters of the -th layer.

[0182] The following is a description of the above calculation steps:

[0183] 1. Perform multi-scale feature extraction on the input image to obtain ;

[0184] 2. Apply the ECA attention mechanism to each scale of the feature map to obtain ;

[0185] 3. Perform weighted fusion on the processed multi-scale features to obtain ;

[0186] 4. Generate the final feature map through the final transformation function ;

[0187] This combined method applies the ECA attention mechanism to the feature maps at each scale, and then integrates the multi-scale information through weighted fusion. This method can effectively capture features at different scales, while using the ECA attention mechanism to enhance important channel information, improving the accuracy and robustness of the model for green tide recognition.

[0188] The specific implementation of step S40 is as follows: Train the green tide semantic segmentation model. The purpose of this step is to obtain a high-performance green tide semantic segmentation model through iterative optimization using the prepared dataset and the designed model architecture. The specific implementation process includes: First, initialize the model parameters. For the ShuffleNetV2 backbone network, use the weights pre-trained on the ImageNet dataset for initialization to accelerate convergence and improve model performance. For the newly added layers (such as the ECA module, multi-scale feature fusion layer, etc.), adopt the He initialization method, that is, randomly sample the initial weights from a truncated normal distribution with a mean of 0 and a standard deviation of , where is the number of input units of this layer. Second, set the training hyperparameters. Adopt the Adam optimizer, set the initial learning rate to 0.001, and the weight decay coefficient to 0.0001. Use the cosine annealing learning rate scheduling strategy to dynamically adjust the learning rate during training. The batch size is set according to the GPU memory capacity, and usually 16 or 32 can be selected. Then, implement the data augmentation strategy, including random horizontal flipping (probability 0.5), random vertical flipping (probability 0.5), random rotation (-15° to 15°), random scaling (0.8 to 1.2 times), and random adjustment of brightness, contrast, and saturation (within the range of ±20%). These augmentation strategies can improve the generalization ability of the model and its adaptability to various environmental conditions. Next, start the iterative training process. In each epoch, input the augmented training data into the model, perform forward propagation calculation to obtain the prediction results, and then calculate the loss function value. Use automatic differentiation technology to calculate the gradients and update the model parameters through the backpropagation algorithm. At the end of each epoch, evaluate the model performance on the validation set and record key metrics such as accuracy, IoU (Intersection over Union), etc. Implement the early stopping strategy. When the performance on the validation set has not improved for 5 consecutive epochs, end the training early. Finally, adopt the gradient clipping technique to limit the L2 norm of the gradients within 5 to prevent the gradient explosion problem. The entire training process continues until the preset maximum number of epochs (such as 100) is reached or the early stopping condition is triggered. During the training process, save the model checkpoint every 5 epochs for subsequent analysis and selection of the best model.

[0189] The specific implementation of step S50 is: model evaluation and optimization. The purpose of this step is to comprehensively evaluate the performance of the trained model and optimize the model by adjusting hyperparameters and model structure. The specific implementation process includes: First, use the validation set to conduct a detailed evaluation of the trained model. Calculate multiple evaluation metrics, including PixelAccuracy, Mean Intersection over Union (MIoU), F1-score, Precision, and Recall. For the green tide recognition task, special attention is paid to the IoU metric, and its calculation formula is: , where TP is the true positive, FP is the false positive, and FN is the false negative. Set the IoU threshold to 0.5 and calculate the Average Precision (AP) at different IoU thresholds. Second, conduct error analysis. Classify and count the mispredicted samples in the validation set to identify the main error types of the model, such as over-segmentation, under-segmentation, inaccurate boundaries, etc. Analyze the attention distribution of the model through visualization techniques (such as GradCAM) to understand whether the image regions focused on by the model are reasonable. Then, based on the error analysis results, adjust the model hyperparameters and structure. For example, if it is found that the model has poor performance in recognizing small-area green tides, the sampling weight of small targets can be increased or more fine-grained features can be added to the feature pyramid. If the model shows overfitting, the regularization strength can be increased, such as increasing the weight decay coefficient or introducing a dropout layer. Adjust the parameters of the ECA module, such as changing the and values in the calculation of the adaptive convolution kernel size to optimize the effect of the channel attention mechanism. For the PSPNet part, different pyramid pooling scales can be tried, such as {1, 2, 3, 6} or {1, 3, 5, 7}, to capture richer multi-scale information. Next, implement cross-validation. After merging the training set and the validation set, conduct 5-fold cross-validation to more comprehensively evaluate the stability and generalization ability of the model performance. Calculate the evaluation metrics for each fold and take the average as the final performance metric. Set the model selection criteria, such as MIoU not less than 0.85 and F1-score not less than 0.9. Finally, perform ensemble learning on the best-performing model configurations. Adopt techniques such as Model Averaging and Test Time Augmentation (TTA). Specifically, select 3-5 of the best-performing model configurations in cross-validation and take the average of the prediction results for each test sample. In the test stage, perform multiple transformations (such as flipping, scaling) on each input image, and comprehensively combine the prediction results after transformation to obtain the final prediction. This ensemble strategy can further improve the stability and accuracy of the model.

[0190] The specific implementation of step S60 is as follows: Obtain the image to be recognized. The purpose of this step is to obtain high-quality image data in real time from the coastal high-point monitoring system, providing input for subsequent green tide recognition. The specific implementation process includes: First, configure the high-point monitoring system. Select a suitable high-point location (such as a lighthouse, high-rise building, or hill) in the coastal area and install a high-resolution imaging device. The resolution of the imaging device should be no less than 4K (3840×2160 pixels), and the field of view angle should be no less than 90 degrees to ensure sufficient coverage of the sea area. Second, set the image acquisition parameters. Adjust the image acquisition frequency according to the seasonal characteristics of green tide occurrence. During the high-incidence period of green tide (usually from May to August), set a higher acquisition frequency, such as once every 5 minutes; in other periods, the frequency can be reduced to once every 30 minutes to balance the data volume and system load. Metadata such as timestamp, GPS coordinates, and camera orientation should be recorded synchronously during image acquisition for subsequent analysis. Then, implement image preprocessing. Perform real-time preprocessing on the acquired original images, including denoising, color correction, and geometric correction. Use a Gaussian filter to remove image noise, and the filter kernel size can be set to 3×3 or 5×5, dynamically adjusted according to the image quality. Perform automatic white balance and contrast enhancement to adapt to different lighting conditions. If the camera has distortion, apply a lens distortion correction algorithm, such as the Brown-Conrady model, and the correction parameters are obtained through offline calibration. Next, realize data transmission and storage. Use a high-speed network (such as 5G or dedicated line) to transmit the preprocessed images to the central processing server in real time. Compress the images using a lossless compression algorithm (such as JPEG2000) to reduce the transmission bandwidth requirement while ensuring image quality. On the server side, store the received images in a high-performance distributed file system, such as the Hadoop Distributed File System (HDFS), to support large-scale parallel processing. Finally, optionally, establish an image quality control mechanism. Evaluate each received image. Quality detection includes clarity evaluation (calculating the image gradient variance using the Laplacian operator, with the threshold set to 100), exposure check (calculating the image histogram to ensure that the proportion of pixel values within the range of 10-245 is not less than 95%), and cloud occlusion detection (using a threshold segmentation method based on the HSV color space, with the cloud coverage threshold set to 30%). For images that do not pass the quality detection, the system automatically adjusts the camera parameters (such as exposure time, ISO sensitivity, etc.) and re-acquires. At the same time, establish an artificial review mechanism to regularly spot-check the automatically acquired images to ensure system stability and data quality.

[0191] The specific implementation of step S70 is: green tide identification. The purpose of this step is to use the trained deep learning model to accurately segment the green tide area from the real-time acquired images. The specific implementation process includes: First, image preprocessing. Adjust the image obtained in step S60 to the input size required by the model (such as 512×512 pixels), and use the bicubic interpolation algorithm for scaling to maintain the smoothness of the image. Then perform image normalization, scale the pixel values to the range [-1,1], and the calculation formula is: , where and are the mean and standard deviation of the training set respectively. Second, model inference. Input the preprocessed image into the optimal green tide semantic segmentation model selected in step S50. Use inference acceleration frameworks such as TensorRT or ONNXRuntime to optimize the model inference speed. During the inference process, enable half-precision (FP16) calculation to improve the calculation efficiency while ensuring the accuracy. For large-size images, adopt the sliding window technique, divide the image into multiple overlapping small blocks (such as 256×256 pixels, with an overlap rate of 50%), perform predictions separately, and then splice and restore the results to the original size. Then, post-processing optimization. Apply threshold segmentation to the probability map output by the model, and the threshold is determined through ROC curve analysis, and the initial value can be set to 0.5. Use morphological operations (such as opening and closing operations) to refine the segmentation results, eliminate small noise regions and fill small holes. The kernel size of the opening and closing operations can be set to 3×3 or 5×5, and dynamically adjusted according to the typical size of the green tide patches. Apply conditional random field (CRF) as a post-processing step to optimize the accuracy of the segmentation boundary. The energy function of CRF is defined as: , where is the unary potential term, is the pairwise potential term. The unary potential is based on the output probability of the model, and the pairwise potential uses a Gaussian kernel to model the similarity between pixels. Then, multi-temporal fusion. Perform time series analysis on the recognition results at multiple consecutive time points (such as within the most recent 1 hour), and use the Kalman filter to smooth the green tide boundary to reduce the instability of single-frame recognition. The state transition equation of the Kalman filter is: , and the observation equation is: , where is the state transition matrix, is the observation matrix, and They are the process noise and the observation noise respectively. Finally, the results are visualized and stored. The recognition results are superimposed on the original image, and different colors or transparencies are used to represent the probability intensity of the green tide area. A vectorized green tide contour is generated to facilitate subsequent area calculation and change analysis. The recognition results are saved in the GeoTIFF format, including geographical coordinate information, which is convenient for integration with other Geographic Information System (GIS) data. At the same time, key parameters (such as the green tide area, the central coordinates, the maximum extension direction, etc.) are extracted and stored in a time series database (such as InfluxDB) to support subsequent trend analysis and early warning judgment.

[0192] The specific implementation of step S80 is: early warning judgment. The purpose of this step is to promptly detect abnormal situations and trigger the early warning mechanism based on the green tide recognition results. The specific implementation process includes: First, set the early warning thresholds. According to historical data and expert experience, multiple levels of early warning thresholds are set. For example, the proportion of the green tide area in the monitoring area can be set as the main indicator, and the initial thresholds can be set as: mild warning 5%, moderate warning 10%, and severe warning 20%. These thresholds should be dynamically adjusted according to the characteristics of different sea areas and seasonal changes. Second, spatio-temporal analysis. Time series analysis techniques, such as the Autoregressive Integrated Moving Average model (ARIMA), are used to model the change trend of the green tide area. The expression of the ARIMA model is: , where is the order of differencing, is the number of autoregressive terms, is the number of moving average terms. By fitting the ARIMA model, the development trend of the green tide in the short term (such as the next 24 hours) can be predicted. At the same time, spatial autocorrelation analysis methods, such as Moran's I, are used to evaluate the spatial aggregation degree of the green tide. The calculation formula of Moran's I is: , where is the number of spatial units, is the spatial weight matrix, and It is the attribute value at the corresponding position. A Moran's Index greater than 0 indicates a positive correlation, and less than 0 indicates a negative correlation. Then, a multi-factor comprehensive evaluation is carried out. In addition to the green tide area, other factors such as the growth rate, spatial distribution, and marine meteorological conditions also need to be considered. A multi-factor evaluation model based on fuzzy logic is established. After standardizing each index, the comprehensive risk level is obtained through fuzzy rule reasoning. The general form of the fuzzy rule is: IF (Condition 1) AND (Condition 2) OR (Condition 3) THEN (Conclusion). For example, "IF (The area growth rate is high) AND (The spatial aggregation degree is high) OR (The predicted area exceeds the threshold after 24 hours) THEN (Trigger a moderate warning)". Next, anomaly detection. The Local Outlier Factor (LOF) algorithm is used to detect abnormal patterns in the development of the green tide. The LOF algorithm calculates the local density deviation of each data point relative to its k-nearest neighbors. Points with LOF values much greater than 1 are regarded as anomalies. The calculation of LOF is based on the concepts of k-distance and reachability distance. The specific calculation steps include: determining the k-distance neighborhood, calculating the reachability distance, calculating the local reachability density, and calculating the LOF value. Set the LOF threshold (such as 2.0), and situations exceeding this threshold will trigger special attention. Finally, warning trigger and notification. When the comprehensive evaluation result exceeds the preset threshold or an anomaly is detected, the warning system is automatically triggered. The warning information is sent to relevant management personnel through multiple channels, including text messages, emails, mobile application push, etc. The warning information should include: warning level, green tide area and location, development trend prediction, recommended measures to be taken, etc. At the same time, the warning information is pushed to the visualization platform, and Geographic Information System (GIS) technology is used to visually display the green tide distribution and warning area on the electronic map. An early warning response mechanism is established, and corresponding emergency plans are activated according to different warning levels, such as increasing the monitoring frequency, deploying prevention and control equipment, and organizing personnel for on-site disposal.

[0193] Through the implementation of the above detailed steps, the present invention provides a comprehensive and efficient method for identifying and warning green tides in coastal high points. This method combines advanced deep learning technology, image processing technology, and spatio-temporal data analysis technology to achieve accurate identification, dynamic monitoring, and timely warning of green tides. By constructing a high-quality dataset, designing an optimized deep learning model, implementing real-time image acquisition and processing, and conducting multi-dimensional warning analysis, this method can effectively support the green tide monitoring and management work in coastal areas and provide important decision-making support for preventing and controlling green tide disasters.

[0194] Specifically, the principle of the present invention is:

[0195] First of all, a high-quality dataset is the key foundation for the performance of deep learning models. This method constructs a highly targeted green tide semantic segmentation image dataset, providing rich and diverse samples for model training. Such a dataset not only includes green tide images under different times and weather conditions, but also provides detailed green tide contour information through accurate JSON data annotation. Such a dataset can help the model learn the characteristics of green tides in various complex environments, thereby improving the generalization ability and recognition accuracy of the model.

[0196] Secondly, the deep learning green tide semantic segmentation algorithm model designed by this method integrates a variety of advanced technologies to achieve efficient and accurate green tide recognition. The choice of PSPNet as the basic semantic segmentation algorithm is because it can effectively capture context information at different scales and is suitable for processing complex sea surface scenes. Replacing the backbone network with ShuffleNetV2 greatly reduces the computational complexity and makes the model more suitable for real-time processing. The introduction of the ECA attention mechanism enables the model to adaptively adjust the importance of different channels, effectively improving the efficiency and accuracy of feature extraction. The multi-scale feature fusion strategy can comprehensively utilize feature information at different levels, which is beneficial to improving the model's recognition ability for small-area green tides and green tide edges.

[0197] Furthermore, this method makes full use of spatio-temporal information to improve the stability and reliability of monitoring results. Through time series analysis of the recognition results at multiple consecutive time points, the system can capture the dynamic characteristics of green tide changes, rather than relying solely on single-frame images. The application of the Kalman filter is based on the modeling of the movement law of green tides. Through the process of prediction and correction, it effectively smooths the recognition results and reduces the influence of random noise. The comprehensive utilization of such spatio-temporal information conforms to the objective law of green tide evolution, so it can provide more stable and reliable monitoring results.

[0198] Finally, the early warning mechanism established by this method is based on multi-dimensional analysis of the development trend of green tides. By setting multi-level early warning thresholds, the system can trigger different levels of early warnings according to the severity of green tides. Time series analysis techniques (such as the ARIMA model) can capture the trend and periodic characteristics of green tide area changes, providing a basis for short-term prediction. Spatial autocorrelation analysis (such as the Moran index) can evaluate the spatial aggregation degree of green tides, helping to identify the diffusion pattern of green tides. The multi-factor evaluation model based on fuzzy logic comprehensively considers various indicators and can more comprehensively evaluate the green tide risk. The introduction of the local outlier factor algorithm improves the system's sensitivity to abnormal green tide development patterns, which is conducive to early detection of potential large-scale green tide disasters.

[0199] In summary, the technical solution of the present invention forms a complete green tide identification and early warning system through means such as data-driven, algorithm optimization, spatio-temporal information fusion, and multi-dimensional analysis. This method meets the actual needs of green tide monitoring and the objective laws of green tide evolution, so it can effectively solve the problems existing in the prior art and provide efficient, accurate, and real-time green tide monitoring and early warning services.

[0200] To better understand and implement the present invention, a specific Embodiment 1 of the present invention is provided below. Embodiment 1 completes the construction and training of the green tide semantic segmentation model: The specific implementation manner of the construction of the green tide semantic segmentation model in Embodiment 1 is as follows.

[0201] S1. The construction method of the green tide semantic segmentation model includes: constructing a green tide semantic segmentation image dataset; training the green tide semantic segmentation model based on the green tide semantic segmentation dataset.

[0202] S1.1 Constructing the green tide semantic segmentation image dataset

[0203] The method for constructing the green tide semantic segmentation image dataset includes: collecting green tide images taken by high-point monitoring and performing data annotation on the green tide contour to obtain the annotated image data; dividing the annotated image data into a training set, a validation set, and a test set according to a preset ratio to obtain the green tide semantic segmentation image dataset.

[0204] Perform manual data annotation for green tide semantic segmentation through the labelme software. The generated json annotation file contains categories and the position information of each point. Match the annotation information of each image from the original json annotation file of the dataset, convert it into a mask image, and finally divide all the image data into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10% to finally form the green tide segmentation image dataset.

[0205] S1.2 Training the green tide semantic segmentation model

[0206] (1) Design a deep learning green tide semantic segmentation algorithm model.

[0207] Improve the PSPNet algorithm for green tide semantic segmentation. First, some improvements are made to the PSPNet algorithm. These include improvements to the backbone network, adding an attention mechanism module to achieve better recognition effects, and improvements to the loss function. Then, the green tide semantic segmentation model is trained and validated. The improved PSPNet algorithm forms a new algorithm, ECA-ShuffledPSPNet.

[0208] The green tide images in the input high - point monitoring are used to extract features through the ShuffleNetV2 backbone network. The obtained feature maps are subjected to sub - region average pooling. The feature pyramid based on the ECA attention mechanism is divided into four levels. The first level is the coarsest level that performs global average pooling on each feature map, and the obtained feature vectors are processed through the ECA attention mechanism. The second level divides the feature map into 2×2 sub - regions, then performs average pooling on each sub - region, and the obtained feature vectors are processed through the ECA attention mechanism. The third level divides the feature map into 3×3 sub - regions, then performs average pooling on each sub - region, and the obtained feature vectors are processed through the ECA attention mechanism. The fourth level is the finest level that divides the feature map into 6×6 sub - regions, then performs pooling on each sub - region, and the obtained feature vectors are processed through the ECA attention mechanism. These pyramid features processed through the ECA attention mechanism are directly upsampled to the same size as the input features, and then merged with the input features for a concat operation to obtain global features. Finally, a final prediction map is generated through a convolutional layer.

[0209] 1) Replace the PSPNet backbone network with ShuffleNetV2

[0210] In the application scenario of green tide semantic segmentation, an efficient, stable, and accurate model is required to process a large amount of monitoring video or image data. Since this scenario requires real - time processing of input data, the model must have high - speed inference. At the same time, green tide semantic segmentation requires the model to accurately identify and segment the green tide area, so the model needs to have high classification accuracy and semantic segmentation precision. Considering the limited hardware resources in the actual application scenario, the model should minimize the computational amount and memory occupancy as much as possible to improve resource utilization.

[0211] Based on the above requirements, it is a suitable choice to replace the PSPNet backbone network with ShuffleNetV2. As a lightweight convolutional neural network, ShuffleNetV2 has high computational efficiency and low resource occupancy, which can meet the real - time requirements. At the same time, it also has good feature extraction ability and model stability, which can ensure high classification accuracy and semantic segmentation precision. Therefore, replacing the PSPNet backbone network with ShuffleNetV2 can better meet the requirements of the green tide semantic segmentation application scenario.

[0212] In summary, replacing the PSPNet backbone network with ShuffleNetV2 can bring benefits and significance such as improved computational efficiency, enhanced model stability, and maintained model performance. This replacement helps to optimize the model's performance, making the model more efficient, stable, and reliable in practical applications.

[0213] 2) Combine the ECA attention mechanism with the multi-scale feature fusion strategy.

[0214] The ECA attention mechanism is combined with the multi-scale feature fusion strategy. By applying the ECA attention mechanism to feature maps of different scales in the PSPNet feature pyramid, the model can more effectively fuse and utilize feature information from coarse to fine. ECA combined with multi-scale fusion can help the model better handle details and global information in complex scenarios.

[0215] The ECA attention mechanism is a lightweight channel attention mechanism. It learns channel attention through a 1D convolutional layer and reduces computational complexity. The ECA attention mechanism avoids dimensionality reduction and instead uses 1D convolution to achieve local cross-channel interaction, thereby extracting dependencies between channels.

[0216] The implementation process of the ECA attention mechanism is shown in Figure 2:

[0217] First, perform global average pooling on the input feature map. This step compresses the original two-dimensional feature map [h, w, c] into a one-dimensional vector [1, 1, c], where h and w represent the height and width of the feature map respectively, and c represents the number of channels.

[0218] Next, calculate the adaptive one-dimensional convolutional kernel size k based on the number of channels of the feature map. This k will be used for the subsequent attention weight calculation, and the specific calculation method is as follows:

[0219] , where C is the number of input channels, b = 1, = 2;

[0220] Then, use k for the one-dimensional convolution operation to obtain the weights for each channel. The Sigmoid activation function is used to map the weights between (0 - 1). This weight reflects the importance of each channel of the feature map.

[0221] Finally, multiply the obtained normalized weights with the original input feature map channel by channel to generate a weighted feature map. This step multiplies each channel of the feature map by the corresponding weight, thereby emphasizing important channels and suppressing unimportant channels.

[0222] Through the above steps, the ECA attention mechanism realizes the attention weighting process of the input feature map, improving the performance and generalization ability of the model.

[0223] The computational complexity of the ECA attention mechanism is relatively low and it is easy to optimize. The attention calculation of ECA is a natural regularization method that does not require additional regularization techniques and is not prone to overfitting.

[0224] 3) Combining the Focal Loss and BCELoss functions to improve the model performance

[0225] Since the green tide data is less than the background data, the number of category samples is unbalanced, which usually affects the effect of the classification algorithm. To solve the problem of class imbalance and improve the generalization ability of the model, the present invention combines the Focal Loss and BCELoss.

[0226] Focal Loss is a loss function for dealing with the problem of class imbalance. It can pay more attention to the samples that are difficult to classify, reduce the weight of the samples that are easy to classify, and thus improve the attention degree to small targets.

[0227] The formula of Focalloss:

[0228] ;

[0229] Where:

[0230] ;

[0231] reflects the degree of closeness to the ground truth, that is, the category y. The larger it is, the closer it is to the category y, that is, the more accurate the classification. >0 is an adjustable factor.

[0232] BCELoss is a loss function designed specifically for binary classification problems, which can directly measure the prediction accuracy and robustness of the model. BCELoss pays attention to the samples that are difficult to classify by adjusting the weights, thereby reducing the weights of the samples that are easy to classify, and helping to improve the generalization ability of the model. By minimizing the cross-entropy loss, the model can gradually adjust the weights and biases to make the prediction results as close as possible to the true results.

[0233] The formula of BCELoss:

[0234] L bce = 1 N ∑ i - [ y i log ( p i )+(1 - y i log ( 1 - p i ))] ;

[0235] In the formula: yi represents the true value of sample i; represents the probability that sample i is predicted as the positive class; N is the total number of image pixels.

[0236] The disadvantage of BCELoss is that it is sensitive to class imbalance, but it can improve the generalization ability of the model. FocalLoss is specifically designed to handle class imbalance problems. Therefore, combining FocalLoss with BCELoss can consider both class imbalance and prediction accuracy.

[0237] The formula for the final loss function:

[0238] ;

[0239] Combining FocalLoss with BCELoss can give full play to their respective advantages, improve the generalization ability of the model while ensuring the accuracy of the model.

[0240] (2) Training of the green tide semantic segmentation model

[0241] The training process of the green tide segmentation model is as follows: First, prepare an image dataset containing green tides, input the images into the improved semantic segmentation network for forward propagation to obtain the predicted labels for each pixel. Use the combined loss function to calculate the difference between the predicted labels and the actual labels. According to the calculated loss, update the weight parameters of the network through the backpropagation algorithm. Error backpropagation calculates the gradients and propagates the gradients layer by layer in reverse, thereby adjusting the weight parameters to reduce the loss. Use the Adam optimizer to update the weight parameters in the network according to the gradients. After each training epoch, use the test dataset to evaluate the performance of the model. The evaluation metrics include accuracy, precision, recall, etc. By comparing the evaluation results of different training epochs, the improvement or decline trend of the model performance can be observed. According to the model evaluation results, adjust the hyperparameters to optimize the model performance. Through iterative training, continuously optimize the weights and hyperparameters of the model to improve the accuracy and stability of semantic segmentation.

[0242] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention. Different from Example 1, in this Example 2, an improved ECA attention mechanism is used: This Example 2 takes a coastal area of a certain city as the background and details the implementation process of the coastal high-point monitoring and green tide identification and early warning method of the present invention. As an important coastal city in China, this city faces serious green tide threats every summer, which has had a major impact on the local tourism, fishery and marine ecosystem. In order to better monitor and early warn of the occurrence and development of green tides, the local marine environment monitoring department decided to adopt the method of the present invention to establish a set of efficient and accurate green tide identification and early warning systems.

[0243] Step S10: Construct a green tide semantic segmentation image dataset

[0244] First, 10 high - point monitoring stations were selected along the coast of the city, including different types of areas such as beaches, ports, and islands. A 4K ultra - high - definition camera (with a resolution of 3840×2160 pixels) was installed at each monitoring station, with a field of view angle of 120 degrees, covering a sea area with a radius of about 5 kilometers.

[0245] During the high - incidence period of green tides from May 1st to August 31st, 2023, images were collected at each monitoring station every 10 minutes, and a total of 123,760 original images were collected. These images covered different weather conditions (sunny, cloudy, overcast), different time periods (daytime, dusk, night), and different sea conditions (calm, light waves, large waves).

[0246] Next, 5 experienced marine ecological experts used the LabelMe tool to annotate these images. During the annotation process, the experts carefully outlined the green - tide contours in each image and saved the annotation results in JSON format. To ensure the annotation quality, each image was independently annotated by at least two experts, and then the final annotation results were determined through cross - validation.

[0247] After screening and quality control, 50,000 high - quality annotated images were finally selected as the dataset. Among these images, the distribution of the green - tide coverage area is as follows:

[0248] 1 - 10% coverage rate: 15,000 images;

[0249] 11 - 30% coverage rate: 20,000 images;

[0250] 31 - 60% coverage rate: 10,000 images;

[0251] 61 - 90% coverage rate: 5,000 images;

[0252] Step S20: Data processing;

[0253] A Python script was developed to convert the JSON - formatted annotation files into binary PNG mask images. After conversion, the green - tide area was marked as 255 (white), and the non - green - tide area was marked as 0 (black).

[0254] The OpenCV library was used to uniformly resize the original images and the corresponding mask images to a size of 512×512 pixels. To maintain the image quality, the bicubic interpolation algorithm was used for scaling.

[0255] Then, using the stratified random sampling method, the dataset is divided into a training set (40,000 images), a validation set (5,000 images), and a test set (5,000 images) in the ratio of 80%, 10%, and 10%. Random sampling is performed within each coverage interval to ensure that the distribution of the green tide coverage area in each subset is similar to that of the overall dataset.

[0256] Finally, the divided dataset is saved in the TFRecord format to improve the data reading efficiency. At the same time, the PNG compression algorithm is used to losslessly compress the images, reducing the dataset size from the original 256GB to 150GB.

[0257] Step S30: Design the prototype of the deep learning green tide semantic segmentation algorithm model

[0258] Based on the PSPNet architecture, the backbone network is replaced with ShuffleNetV2. The configuration of ShuffleNetV2 is as follows:

[0259] Stage 1: The number of output channels is 24, repeated 2 times;

[0260] Stage 2: The number of output channels is 116, repeated 4 times;

[0261] Stage 3: The number of output channels is 232, repeated 8 times;

[0262] Stage 4: The number of output channels is 464, repeated 4 times;

[0263] An ECA module is added after each ShuffleNetV2 unit, and the calculation parameters for the adaptive convolution kernel size are set to γ = 2 and β = 1.

[0264] The multi-scale feature fusion strategy extracts feature maps from the 4 stages of ShuffleNetV2, uniformly adjusts the number of channels to 128 through 1×1 convolution, and then upsamples all feature maps to 1 / 4 of the input resolution using bilinear interpolation. Finally, the fused feature map is obtained through weighted summation, and the initial weights are set to [0.1, 0.2, 0.3, 0.4].

[0265] The loss function adopts a linear combination of FocalLoss and BCELoss, and the weight ratio is initially set to 1:1. The parameters of FocalLoss are set to α = 0.25 and γ = 2.

[0266] Step S40: Train the green tide semantic segmentation model

[0267] The model is implemented using the PyTorch deep learning framework. The training is carried out on 4 NVIDIA Tesla V100 GPUs, and the batchsize is set to 64.

[0268] The Adam optimizer is adopted, with the initial learning rate set to 0.001 and the weight decay coefficient set to 0.0001. The cosine annealing learning rate scheduling strategy is used, and the total number of training epochs is set to 100.

[0269] The data augmentation strategy includes:

[0270] - Random horizontal flipping (probability 0.5);

[0271] - Random vertical flipping (probability 0.5);

[0272] - Random rotation (-15° to 15°);

[0273] - Random scaling (0.8 to 1.2 times);

[0274] - Random brightness adjustment (±20%);

[0275] - Random contrast adjustment (±20%);

[0276] - Random saturation adjustment (±20%);

[0277] During the training process, the model performance is evaluated on the validation set after each epoch. The main evaluation metrics are the mean intersection over union (MIoU) and pixel-level accuracy.

[0278] The early stopping strategy is implemented. When the MIoU on the validation set does not improve for 5 consecutive epochs, the training is terminated early. At the same time, the gradient clipping technique is adopted to limit the L2 norm of the gradient within 5.

[0279] Step S50: Model evaluation and optimization

[0280] A detailed evaluation is carried out on the validation set, and the following performance metrics are obtained as shown in Table 1:

[0281] Table 1 Performance metric table

[0282] index value Mean Intersection over Union (MIoU) 0.892 Pixel-level accuracy 0.967 F1 score 0.934 Precision 0.941 Recall 0.927

[0283] Through error analysis, it is found that the model has a certain problem of missed detection when dealing with small-area green tides (coverage rate < 5%). Therefore, the sampling weight of small targets is increased, and finer-grained features are added to the feature pyramid.

[0284] The parameters of the ECA module are adjusted, with γ adjusted to 1.5 and β adjusted to 0.8 to improve the sensitivity of the model to channel attention.

[0285] 5-fold cross-validation is performed, and the average MIoU reaches 0.901 with a standard deviation of 0.008, indicating the stable performance of the model.

[0286] Finally, select the top 3 model configurations with the best performance in cross-validation for integration, and adopt model averaging and test-time augmentation techniques. Test-time augmentation includes horizontal flipping, vertical flipping, and 90-degree rotation, and the average of the 8 transformation results of each test sample is taken.

[0287] After optimization, the final performance on the test set is shown in Table 2 below:

[0288] Table 2 Optimized Performance Index Table

[0289] index value Mean Intersection over Union (MIoU) 0.915 Pixel-level accuracy 0.978 F1 score 0.952 Precision 0.958 Recall 0.946

[0290] Step S60: Obtain the image to be recognized

[0291] Twenty new high-point monitoring stations were selected along the coast of the city, and an 8K ultra-high-definition camera (resolution 7680×4320 pixels) was installed at each station, with a field of view of 135 degrees, covering a sea area with a radius of about 8 kilometers.

[0292] The parameter settings of the camera equipment are as follows:

[0293] - High-incidence period of green tides (May - August): Images are collected every 3 minutes;

[0294] - Other periods: Images are collected every 15 minutes;

[0295] The image preprocessing process includes:

[0296] 1. Use a 5×5 Gaussian filter to remove image noise;

[0297] 2. Automatic white balance to ensure color restoration;

[0298] 3. Contrast Limited Adaptive Histogram Equalization (CLAHE) to improve image details;

[0299] 4. Use the Brown-Conrady model for lens distortion correction;

[0300] Adopt the JPEG2000 compression algorithm to compress the original image (about 100MB) to about 20MB, with a compression ratio of 5:1.

[0301] Use a 5G private network to transmit the compressed image to the central processing server in real time. The server adopts a distributed storage system and uses the Hadoop Distributed File System (HDFS) to store image data.

[0302] The image quality control algorithm is set as follows:

[0303] - Sharpness evaluation: Calculate the variance of the image gradient using the Laplacian operator, and set the threshold to 120;

[0304] - Exposure check: Ensure that the proportion of pixel values within the range of 15 - 240 is not less than 97%;

[0305] - Cloud and fog occlusion detection: Use the threshold segmentation method in the HSV color space, and set the cloud and fog coverage threshold to 25%.

[0306] Step S70: Green tide recognition

[0307] Image preprocessing: Use the bicubic interpolation algorithm to scale the image to 512×512 pixels, and then perform normalization to scale the pixel values to the range of [-1, 1].

[0308] Model inference: Use the TensorRT inference acceleration framework to run the optimized model on the NVIDIA Tesla T4 GPU. Enable FP16 half-precision calculation, and the average processing time for a single image is 15 milliseconds.

[0309] For 8K raw images, adopt the sliding window technique to divide the image into 25 overlapping small blocks of 512×512 pixels (overlap rate 50%), perform predictions separately, and then splice and restore the results to the original size.

[0310] Post-processing optimization:

[0311] 1. Apply threshold segmentation with the threshold set to 0.5;

[0312] 2. Use a 5×5 structuring element for morphological opening and closing operations to eliminate small noise regions and fill small holes;

[0313] 3. Apply the conditional random field (CRF) to optimize the segmentation boundary, and set the σ of the Gaussian kernel to 3;

[0314] Time series analysis: Use the Kalman filter to smooth the recognition results of 10 consecutive frames (30 minutes). The parameter settings of the Kalman filter are as follows:

[0315] - Process noise covariance Q = 0.01 * np.eye(4);

[0316] - Observation noise covariance R = 0.1 * np.eye(2);

[0317] - Initial state covariance P = 1000 * np.eye(4);

[0318] Step S80: Early warning judgment

[0319] Set the early warning threshold:

[0320] - Mild early warning: The green tide area accounts for 5% of the monitoring area;

[0321] - Moderate warning: The area of the green tide accounts for 10% of the monitored area;

[0322] - Severe warning: The area of the green tide accounts for 20% of the monitored area;

[0323] Time series analysis: Use the ARIMA(2,1,2) model to model the green tide area data for the past 7 days and predict the development trend of the green tide in the next 24 hours.

[0324] Spatial autocorrelation analysis: Use the Moran index to evaluate the spatial aggregation degree of the green tide, and set the spatial weight matrix as inverse distance weighting.

[0325] Multi-factor evaluation model: Establish an evaluation model based on fuzzy logic. The input variables include the green tide area, growth rate, spatial aggregation degree, and meteorological conditions (temperature, wind direction, wind speed). Example of fuzzy rules:

[0326] - IF (the area IS large) AND (the growth rate IS fast) AND (the spatial aggregation degree IS high) THEN (the risk IS very high);

[0327] - IF (the area IS medium) AND (the growth rate IS slow) AND (the meteorological conditions IS unfavorable) THEN (the risk IS medium);

[0328] Anomaly detection: Use the Local Outlier Factor (LOF) algorithm, set k = 5, and the LOF threshold is 2.0.

[0329] Warning trigger and notification:

[0330] - When the comprehensive risk assessment result exceeds 0.7 or the LOF value exceeds 2.0, trigger the warning system;

[0331] - Send warning messages to relevant management personnel through SMS, email, and mobile app push;

[0332] - The warning message includes: warning level, green tide area and location, 24-hour development trend prediction, and recommended measures;

[0333] - Use ArcGIS to build a visualization platform to display the green tide distribution and warning area in real time on an electronic map.

[0334] System operation results:

[0335] This system was put into use during the high-incidence period of the green tide from May 1st to August 31st, 2024. During this period, a total of 3,942,000 images were processed, and 387 green tide events were identified, among which:

[0336] - Mild warning: 276 times;

[0337] - Moderate warning: 89 times;

[0338] - Severe warning: 22 times;

[0339] The average processing time of the system is 0.8 seconds per image, including the entire process of image preprocessing, model inference, post-processing, and early warning judgment. The average transmission delay of early warning information is 30 seconds.

[0340] By comparing with traditional satellite remote sensing and on-site marine investigation methods, this system shows significant advantages in the early identification of green tides. Among the 22 severe warning events, 19 times the warning was issued 24 - 48 hours earlier than the traditional method, winning valuable response time for the emergency management department.

[0341] The accuracy of the system has also been verified. Through random spot checks and on-site verification, the average error of the estimated green tide area by the system is within ±7%, and the average deviation of the position accuracy is less than 100 meters.

[0342] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention.

Claims

1. A method for identifying and warning green tides at coastal high points, characterized in that: The following steps are involved: S10, collecting a large number of green tide images taken by high-point monitoring, and annotating the green tide contours in the green tide images with JSON data to obtain annotated image data; S20, converting the data annotations in JSON format into PNG mask images, and dividing them into a training set, a validation set, and a test set to obtain a green tide semantic segmentation image training data set; S30, select the PSPNet semantic segmentation algorithm, replace the backbone network with ShuffleNetV2, and combine the ECA attention mechanism equation group with the multi-scale feature fusion strategy, and combine the Focal Loss and BCELoss loss functions to solve the category imbalance problem and improve the generalization ability of the model; S40, using the green tide semantic segmentation image training data set to train a deep learning green tide semantic segmentation algorithm model prototype, and updating each weight of the deep learning green tide semantic segmentation algorithm model prototype by error back propagation; S50, using the validation set to evaluate the performance of the trained deep learning green tide semantic segmentation algorithm model prototype, adjusting the hyperparameters to optimize the model performance, and selecting the optimal model as the green tide semantic segmentation model; S60, acquiring real-time image data from a coastal high point monitoring system as an image to be identified; S70, preprocessing the image to be identified, inputting it into the trained green tide semantic segmentation model, and obtaining the segmentation mask of the green tide; S80, judging whether the green tide area exceeds a preset threshold or has abnormal changes according to the segmentation mask, and if the conditions are met, triggering an early warning system and notifying relevant management personnel; Among them, the formula combining the ECA attention mechanism equation group with the multi-scale feature fusion strategy is expressed as follows: Step 1: Multi-scale feature extraction, specifically expressed as: ; Where: : No. Feature map of the layer; : No. Feature extraction function of the layer; : Input image; : Total number of layers; Step 2: ECA attention mechanism is applied to each scale, specifically expressed as: ; Where: represents the ECA attention mechanism equation group; Step 3: Multi-scale feature fusion: ; Where: : The first Layer feature map; : Upsampling function, which upsamples the feature map to the same size as the highest resolution feature map; : No. The fusion weight of the layer features is obtained through learning and satisfies ; Step 4: The final feature map is generated, which is specifically expressed as: ; in It is a learnable transformation function used to adjust the number of channels of the fused features.

2. A coastal high point monitoring green tide identification and early warning method according to claim 1, characterized in that: The ECA attention mechanism equation group includes an adaptive convolution kernel size calculation equation, a global average pooling equation, a one-dimensional convolution operation equation, a Sigmoid activation function equation, and an attention weighted equation.

3. A coastal high point monitoring green tide identification and early warning method according to claim 2, characterized in that: The adaptive convolution kernel size calculation equation is specifically expressed as: ; Where: is the number of channels of the input feature map; is an adjustable scaling factor, optimized in the range of [0.1, 10] by grid search; is the bias term, which is optimized in the range of [-5, 5] by random search; Round down to the nearest odd number.

4. A coastal high point monitoring green tide identification and early warning method according to claim 3, characterized in that: The global average pooling equation is specifically expressed as: ; Where: For the Global average pooling result of channels; , is the height and width of the input feature map; The input feature map is Channels, positions The value of .

5. A coastal high point monitoring green tide identification and early warning method according to claim 4, characterized in that: The one-dimensional convolution operation equation is specifically expressed as: ; Where: For the One-dimensional convolution output of channels; The convolution kernel size is calculated adaptively; The one-dimensional convolution kernel weights can be learned and optimized through back-propagation; It is the global average pooling result after loop padding.

6. A coastal high point monitoring green tide identification and early warning method according to claim 5, characterized in that: The Sigmoid activation function equation is specifically expressed as: ; Where: For the The attention weight of each channel; is the steepness parameter of the Sigmoid function, optimized by gradient descent; is the center shift of the Sigmoid function, optimized by gradient descent.

7. A coastal high point monitoring green tide identification and early warning method according to claim 6, characterized in that: The attention weighted equation is specifically expressed as: ; Where: The output feature map is Channels, positions The value of The input feature map is Channels, positions The value of For the The attention weight of each channel; is a small constant to prevent division by zero errors. The default setting is ; , are learnable spatial attention parameters optimized via back-propagation; is the hyperbolic tangent function, which is used to introduce spatial attention.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, which, when executed, are used to execute the coastal high point monitoring green tide identification and early warning method described in any one of claims 1-7.

9. A coastal high point monitoring green tide identification and early warning system, characterized in that: Contains the computer-readable storage medium of claim 8.

Citation Information

Patent Citations

  • Remote sensing image green tide information extraction method based on deep learning and super-resolution

    CN112966580A

  • Red tide detection method and system, medium, computer equipment and terminal

    CN117911885A