Shape structure reinforced lightweight plankton classification method based on multi-scale gating attention

By introducing multi-scale gated attention modules and wavelet transformations into the plankton classification algorithm, combined with Class-Balanced Focal Loss, the shortcomings of existing algorithms in terms of lightweight and generalization capabilities are solved, and efficient and accurate plankton image classification is achieved.

CN120148028APending Publication Date: 2025-06-13ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510147192.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing plankton classification algorithms have problems such as insufficient lightweighting, low accuracy, inadequate network structure, and weak generalization capabilities, and are difficult to effectively implement in scenarios where computing resources are limited.

Method used

A lightweight plankton image classification method based on GhostNet is proposed. By introducing multi-scale gated attention module (MGLR) and wavelet transform, the overall structure perception ability of the network is enhanced, and the problem of uneven sample distribution is optimized through Class-Balanced Focal Loss.

Benefits of technology

It realizes more accurate plankton image classification, improves the model's sensitivity to details and shape profiles, significantly improves classification accuracy and stability, and is suitable for resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148028A_ABST
    Figure CN120148028A_ABST
Patent Text Reader

Abstract

The invention discloses a shape structure enhanced lightweight plankton classification method based on multi-scale gating attention, and the method comprises the steps: obtaining plankton image data, and constructing a data set; preprocessing the image; designing and constructing a lightweight neural network architecture based on the GhostNet; an MGLR (Multi-Scale Gated Long-Range) module is put forward, a gating mechanism and large-kernel multi-scale convolution are combined, weights of different scale features are dynamically adjusted, and the perception ability of the model to global information is enhanced; wavelet transform is introduced to refine a convolution structure, and image signals are decomposed into different frequency components, so that low-frequency information and high-frequency information are extracted respectively; introducing a category balance factor to process a category imbalance problem; and outputting a classification result by training the optimization model. The classification precision and the model efficiency can be remarkably improved, and the performance is excellent especially under the low resolution or complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of plankton classification, and particularly relates to a plankton classification method based on deep learning. Background Art

[0002] Plankton are numerous in nature and play an important role in the ecosystem. They are the basis of many food chains and support rich biodiversity. They are not only an important food source for other aquatic organisms but also play a key role in climate regulation and oxygen production. Therefore, monitoring the changes in plankton is crucial for understanding the response mechanism of the ecosystem. With the rapid development of the Video Plankton Recorder (VPR), the image acquisition speed has been significantly improved, resulting in a large number of plankton images being collected. However, these images are often mixed with dust, debris, and artifacts, making it difficult for even experienced experts to avoid the inherent uncertainty of classifying blurred images when annotating data. The number of plankton images is huge, reaching the order of hundreds of millions. Traditional manual classification methods are not only time-consuming and laborious but also inefficient. How to efficiently process such a large dataset has become an important challenge. In recent years, with the rapid development of computer vision technology, convolutional neural networks (CNNs) have been widely used in the automatic recognition of plankton images.

[0003] At first, Al-Barazanchi et al. proposed using a 10-layer CNN as a feature extractor to replace the traditional feature engineering that relies on expert knowledge. This method effectively overcomes the limitations of traditional technologies in terms of feature quality, implementation complexity, and time cost, providing a new idea for the automatic classification of plankton images. At the same time, random forest and SVM are used for classification to improve the performance and efficiency of the model. With the development of deep learning, deep networks perform better in feature information extraction. Three years later, they adopted a deeper neural network on this basis to further improve the classification accuracy of the dataset. The introduction of "residual connection" effectively solves the problem of gradient disappearance or explosion that occurs when the network is too deep, thus promoting the development of the network towards deeper levels. Xiu et al. used a 50-layer ResNet network to process a plankton 121 classification task and achieved an optimal accuracy of 73.1%. Similarly, Cui et al. also used ResNet50 for classification on a private dataset and improved the traditional full classification head based on this using multi-class SVM, achieving a precision and recall rate of over 94%. In addition to using network models in the general field, Dai et al. proposed a network specifically for plankton classification called ZooplanktoNet based on AlexNet and VGGNet. Their 11-layer model achieved the best result in this task, with an accuracy of 93.7%, which is 2.6% higher than the second-best solution.

[0004] Compared with terrestrial organism images, plankton images usually have lower resolution, blurred details, and most mainstream plankton datasets are grayscale images, which means they contain less information compared to normal images. Therefore, extracting effective features has become the core challenge in classification. To obtain more information, inspired by GoogleNet, Py et al. designed a multi-scale initial layer to extract features in parallel by converting the input image into three different sizes, thus fusing features of different scales and enhancing the feature representation ability. Since color information is often missing in plankton classification, shape and texture features become the main basis. Dai et al. first proposed using texture features to describe species differences and constructed a hybrid network composed of three parallel CNNs, combining the Scharr operator and Canny edge detection to extract shape and texture features, which improved the accuracy by 1% compared to the baseline, but the number of parameters increased significantly. In contrast, Cui et al. used Gaussian filtering and logarithmic image enhancement to extract shape features and texture information. They concatenated the processed image with the original image to form a three-channel image similar to RGB as the initial feature input into the Alexnet-based model. This method only requires training one model. Although the improvement is less than 1%, it significantly reduces the required computational resources. Different from image enhancement, Cheng et al. used polar coordinate representation to solve the classification problem caused by angular differences, while Ellen et al. enhanced the classification performance of the network by introducing geometric shapes and context metadata.

[0005] Data imbalance is another major challenge in plankton classification. The uneven distribution of classes causes the classifier to be biased towards strong classes, affecting the accuracy of the model. To address the long-tail problem in the dataset, Lee et al. generated a balanced subset by random sampling and fine-tuned AlexNet with this subset, increasing the F1 score from 17% to 33%. Wang et al. proposed a data threshold method. By preferentially training small classes and freezing their weights, when training the VGG16 model, the F1 score reached 0.54. In addition, generative adversarial networks (GANs) are often used to handle imbalanced datasets. Wang et al. generated difficult samples through an architecture similar to CGAN to improve the classification accuracy; Li et al. used CycleGAN to generate fake images to mitigate the negative impact of rare samples. Although these methods can enhance the robustness and accuracy of the model, static strategies may cause "data drift", that is, the distribution of training data does not match the data distribution in actual applications, resulting in a decline in model performance and inability to meet actual needs.

[0006] In recent years, ensemble learning has developed rapidly in this field. Lumini et al. tested the fusion strategies of different CNNs and integrated 11 independent classifiers using the sequential floating forward selection (SFFS) feature selection method. After multiple rounds of training, their best model achieved an F1 score of up to 0.953 on the public dataset, exceeding all previous models. Kerr et al. fine-tuned multiple CNNs, froze all the trained weights, only kept the last softmax layer, connected all the outputs and fused them with an independently trained external multi-layer perceptron (MLP), achieving an F1 score of 96% on the 104-classification task. Similarly, Kyathanahally et al. aggregated 6 existing CNNs to improve accuracy and robustness, while Maracani et al. combined CNNs and pre-trained Vision Transformer models, promoting the progress of plankton classification technology and becoming the state-of-the-art technology in the current field. Different from traditional CNNs, Luo et al. achieved spatial sparse CNNs through a pruning process, thus reducing the connections. This method can learn more quickly when processing plankton images, reduce the risk of overfitting, and reduce the computational cost. However, due to the highly variable texture and lighting conditions of plankton images, the model may not perform well on some types of plankton images. In terms of lightweight, Guo et al. were the first to use the lightweight ShuffleNetV2 to classify plankton. As a lightweight network, ShuffleNetV2 can be accelerated by FPGA in other fields and applied to in-situ embedded devices.

[0007] The traditional model paradigm in the field of plankton research has been difficult to meet the needs of modern technological development. With the continuous evolution of deep learning, the limitations of traditional deep networks in terms of performance and efficiency have become increasingly prominent. Continuing to rely on these relatively old models not only makes it difficult to fully utilize computing resources but also restricts the further improvement of performance. At the same time, how to reduce the demand for computing resources while maintaining high performance has become a challenge that urgently needs to be solved in this field. Although ensemble learning can significantly improve model performance, it also consumes a large amount of computing resources and fails to achieve efficient operation on the basis of lightweight. This makes it difficult to effectively implement ensemble learning methods in scenarios with limited computing resources, restricting their wide application in practical applications. Summary of the Invention

[0008] The object of the present invention is to propose an image classification algorithm based on a deep learning network for the problems of the existing plankton classification algorithms, such as lack of light weight, low accuracy, insufficiently advanced network structure, and weak generalization ability. This algorithm uses GhostNet as the benchmark network for feature extraction, adds the proposed attention module (MGLR) to the network to enhance the ability to capture global information. And wavelet transform is introduced to refine the convolutional structure, decomposing the image signal into different frequency components, so as to extract low-frequency and high-frequency information respectively. The method combining MGLR and wavelet transform strengthens the overall structure perception ability of the network. Finally, by introducing a class balance factor, the situation of uneven sample distribution is optimized, the recognition ability of the model for difficult samples is improved, the long-tail problem in the dataset is effectively addressed, and the purpose of efficient classification is achieved.

[0009] The present invention provides a method for strengthening the shape structure and lightening the weight of plankton classification based on multi-scale gated attention, including the following steps:

[0010] Step 1, obtaining plankton images and constructing a dataset; each image in the dataset includes an image label, and the image label is the plankton species type;

[0011] Step 2: Constructing a plankton classification model, denoted as the MWG-GhostNet model in the present invention. The plankton classification model uses GhostNet as the backbone network. The structure of the plankton classification model, from input to output, includes: an input layer Input, several bottleneck layers MWC-bneck, a global average pooling layer, a dimensionality increasing layer, and a classification layer;

[0012] The input of the input layer Input is the image obtained in Step 1;

[0013] Each bottleneck layer MWC-bneck includes: Ghost module 1 in the GhostNet network, Ghost module 2 in the GhostNet network, the MGLR attention module, and a wavelet convolution module; Ghost module 1 increases the dimension of the input features of the bottleneck layer MWC-bneck, MGLR extracts global information from the input features of the bottleneck layer MWC-bneck and performs pointwise multiplication with the output of Ghost module 1, and inputs it into the wavelet convolution module. The wavelet convolution module is used to respond to the low-frequency information in the reweighted feature map and strengthen the contour structure information; Ghost module 2 reduces the dimension of the output of the wavelet convolution module;

[0014] The feature map output by the global average pooling layer is unfolded into a one-dimensional vector, dimensionally increased through the dimensionality increasing layer, and then mapped to the dimension of the number of categories using the classification layer for the classification task;

[0015] Step 3: Use the plankton dataset to train the plankton classification model and obtain the trained plankton classification model;

[0016] Step 4: Model classification prediction

[0017] Obtain the plankton image to be classified, input it into the trained plankton classification model, and obtain the classification result.

[0018] Preferably, in Step 3, the Class-Balanced Focal Loss function is used as an optimizer to train the plankton classification model by the method of backpropagation gradient update.

[0019] Preferably, in Step 2, the number of bottleneck layers MWC-bneck is more than 16.

[0020] Preferably, in Step 2, the dimensionality increase layer includes a 1x1 convolution.

[0021] Preferably, in Step 2, the classification layer includes a Softmax classifier.

[0022] Preferably, the classification method is applied to underwater automatic monitoring of plankton.

[0023] Preferably, in Step 2, the MGLR attention module performs horizontal aggregation and vertical aggregation according to the following formula;

[0024]

[0025] The input feature map of the MGLR attention module is x, F H and F W respectively represent the convolution kernel operations in the horizontal and vertical directions, σ represents the Sigmoid activation function for gating operations; x h,w represents each pixel in x. In horizontal aggregation, each pixel in the feature map is weighted and summed with the features of other pixels in the same row of this pixel. After the result passes through gating activation, it is multiplied point by point with the original pixel to obtain the horizontally aggregated feature map α′ hw ; The vertical aggregation is based on the horizontal aggregation, aggregates along the w axis, and multiplies through gating operations to finally obtain the aggregated feature map α hw ;

[0026] During the horizontal aggregation and vertical aggregation processes, convolution kernels of several different sizes are used for feature extraction in the horizontal and vertical directions.

[0027] Preferably, the several different-sized convolution kernels are specifically the following convolution kernels: 1x5, 1x7, 1x9.

[0028] The present invention proposes a lightweight plankton image classification method (MWC-GhostNet) based on GhostNet. This method mainly consists of three parts: the Multi-scale Gating (MGLR) module: aggregates long-range dependencies in the horizontal and vertical directions through multi-scale asymmetric convolutional kernels, dynamically learns the weights of features at each scale using a gating mechanism, and fuses information from different spatial scales, thereby enhancing the model's ability to capture image details and global information. Introduce wavelet transform in the expansion layer to capture various information in the feature map through different frequencies, and improve the model's ability to extract the shape contours of plankton. Effectively address the problem of dataset imbalance through Class-Balanced Focal Loss, and help the model improve its performance in the classification task of difficult samples. Especially in the case of class imbalance, the classification accuracy is enhanced.

[0029] The beneficial effects of the present invention are:

[0030] The Multi-scale Gating Long-range (MGLR) module proposed by the present invention can effectively aggregate global and local information in the image, improve the model's sensitivity to details and shape contours, and thus achieve more accurate plankton image classification. Through the combination of wavelet transform and the MGLR module, the model shows stronger structure perception ability when extracting the shape and texture features of plankton, and effectively addresses the problems of complex background and insufficient details.

[0031] The present invention introduces Class-Balanced Focal Loss, optimizes the processing of difficult samples for the class imbalance problem in the dataset, and effectively improves the classification accuracy of the model on the long-tail dataset. Combining multiple optimization strategies and modules improves the classification accuracy and stability of the model. Especially in the plankton classification task, it has stronger adaptability and robustness compared with traditional methods.

[0032] The present invention can not only significantly improve the accuracy of plankton image classification, but also achieve low computational resource consumption in practical applications, and is suitable for various resource-constrained devices. The classification accuracy of this method on two datasets is as high as 78.43% and 98.76% respectively, and the F1 scores reach 0.6702 and 0.9618. The model has excellent prediction accuracy, good convergence and stability, and can be directly used for automatic monitoring of plankton points. Brief Description of the Drawings

[0033] Figure 1 It is the structure diagram of the bottleneck layer of the MWC-GhostNet network constructed by the present invention;

[0034] Figure 2It is the model diagram of the Multi-Scale Gated Long-Range Attention (MGLR) attention structure constructed by the present invention;

[0035] Figure 3 It is the module weight diagram of MGLR attention in different bottleneck layers of the MWC-GhostNet network constructed by the present invention;

[0036] Figure 4 It is the single-channel example diagram of wavelet transform in the expansion layer of the MWC-GhostNet network constructed by the present invention;

[0037] Figure 5 It is the overall structure flow chart of the MWC-GhostNet network constructed by the present invention;

[0038] Figure 6 It is the specific sample diagram of the Kaggle121 and PMID2019 datasets used by the present invention

[0039] Figure 7 It is the sample quantity distribution diagram of the Kaggle121 and PMID2019 datasets used by the present invention

[0040] Figure 8 It is the comparison diagram of confusion matrices of the MWC-GhostNet network constructed by the present invention on 2 datasets

[0041] Figure 9 It is the class activation comparison diagram of the present invention on the Kaggle121 dataset

[0042] Figure 10 It is the class activation comparison diagram of the present invention on the PMID2019 dataset

[0043] Figure 11 It is the comparison curve diagram of loss functions of the present invention on 2 datasets Detailed implementation manners

[0044] The present invention will be further described below with reference to the accompanying drawings.

[0045] The present invention includes the following steps:

[0046] Step 1: Obtain the plankton image dataset and perform preprocessing:

[0047] A total of 2 datasets are used in the experimental process of the present invention, and the specific sample images are as shown in the appendix Figure 6As shown, the Kaggle121 dataset is sourced from the ISIIS imaging system and consists of grayscale images taken in the Florida Straits, USA, from May to June 2014. This dataset was first released in the national data science competition in 2015. In this study, we selected 30,336 images with labels as the original annotated version, which altogether contain 121 categories. According to the method in the literature, we divided these images into 24,342 training images and 5,994 test images for subsequent experiments.

[0048] The PMID2019 dataset is a microscopic image dataset for intelligent marine agriculture plankton detection. The specific sample images are as attached Figure 6 As shown, it is divided into 24 categories, and each image contains multiple plankton objects. In this study, we cropped the images with annotation boxes and classified them, removed some incorrect samples, and finally selected 20 categories, a total of 10,131 images, which were divided into a training set and a test set at an 8:2 ratio.

[0049] Attached Figure 7 shows the sample distribution of each category in Kaggle121 and PMID2019 and the imbalance of the datasets. The abscissa represents different species categories, and the ordinate represents the sample quantity for the species. We performed data augmentation on the datasets through horizontal and vertical flipping, random rotation, horizontal or vertical translation, and random cropping. The training set is used to train the network model, and the test set is used to test the classification accuracy of the model.

[0050] Step 2: Construct the MGLR-WTConv-CBFocal-GhostNet model:

[0051] We proposed MWC-GhostNet, a lightweight CNN model designed specifically for plankton image classification. This model constructs the network by stacking multiple MWC-bnecks (as shown in the attachment Figure 1 ), where MGLR and WTConv are two core components.

[0052] In the Ghost Module, only 3x3 small kernel convolutions are adopted. Although this design is computationally efficient and suitable for environments with limited computing resources, small kernel convolutions can only capture information within a local window and lack attention to global information, which becomes an important bottleneck in network performance. In the plankton recognition task, shape and texture are key information. Without capturing global information, it is easy to misjudge species with similar textures, thus affecting the classification performance of the model. To enhance the model's ability to capture key features, we propose a multi-scale gated Multi-Scale Gated Long-Range Attention (MGLR) module, aiming to improve the ability to capture global information. MGLR also includes two parts: horizontal aggregation and vertical aggregation, and a gating mechanism is introduced on this basis. By dynamically controlling the transmission of information flow, MGLR can adaptively select the most relevant features, avoid interference from redundant information and noise, and thus effectively improve the quality of feature representation. The mathematical description of this method can be expressed as:

[0053]

[0054] Assume the original input feature is x, where F H and F W represent the convolution kernel operations in the horizontal and vertical directions respectively, and σ represents the Sigmoid activation function for gating operations. x h,w represents each pixel in the input feature map, and α′ hw represents the feature map after horizontal aggregation. In horizontal aggregation, the features at each position in the feature map are weighted and summed with the features at other positions in the same row of this position. After passing through the gating activation, the result is multiplied pointwise with the original pixels to obtain the horizontally aggregated feature map α′ hw . Similarly, vertical aggregation is based on horizontal aggregation, aggregates along the w axis, and multiplies through gating operations to finally obtain the aggregated feature map α hw . During the horizontal and vertical aggregation processes, multiple convolution kernels of different sizes (1x5, 1x7, 1x9) are used for feature extraction in the horizontal and vertical directions. This can effectively capture information at different scales and improve the model's perception ability of details and global shapes in plankton images. Att Figure 3 shows the weight heat maps generated by the MGLR attention mechanism in different bottleneck layers. As the network depth increases and the feature map size gradually shrinks, it gradually focuses on the overall information, and the output of this module is used to re-weight the extended features.

[0055] The MGLR module combines a gating mechanism and large-kernel multi-scale convolutions, which can dynamically adjust the weights of features at each scale, enhance the representation ability of important features, and suppress irrelevant information. This design effectively improves the model's ability to capture long-range dependencies, ensuring stronger global perception while maintaining lightness and efficiency. Through multi-scale feature fusion and dynamic weighting, MGLR performs excellently in plankton image classification. At the same time, depthwise separable convolutions are adopted to reduce the computational overhead and improve the computational efficiency of the model. This design optimizes the performance significantly without sacrificing classification accuracy while ensuring high efficiency.

[0056] Many CNN models tend to focus more on color and have a lower preference for shape. However, color information is often lacking in plankton images, which makes many efficient networks perform poorly in plankton image classification tasks. Therefore, we introduce wavelet transform to refine the convolutional structure to enhance the shape contour information in plankton images. Equation 3 describes a complete first-level wavelet convolution process. Assume the input tensor is x, where conv represents the depth convolution module. The wavelet transform (WT) is used to process the input to generate a wavelet component representation (W) containing four different frequency combinations. Next, these different frequency components are subjected to small-kernel depth convolutions, and then the features are reconstructed through the inverse wavelet transform (IWT). Finally, the output of the IWT is added to the convolution result of the original input tensor to obtain the final output.

[0057] F(x) = Conv(x) + IWT(Conv(W, WT(x))) (3)

[0058] The wavelet transform decomposes the image into low-frequency and high-frequency components, extracting the overall structure and shape contour of the image respectively. Att Figure 4 shows an example of a single-channel second-level wavelet convolution. Through the separable convolution processing of frequency components, wavelet convolution can extract richer feature information, highlighting the detailed features, especially the contours and edges of plankton. Using multi-frequency components can capture more information and improve the recognition ability of shape and edge features. In addition, wavelet convolution improves the performance while maintaining the computational efficiency by expanding the receptive field and increasing a very small number of model parameters.

[0059] The specific model structure diagram is as shown in Att Figure 5The pre - processed enhanced image shown above first completes the down - sampling and up - dimensionality of the feature map through the ConvStem module, which not only reduces the number of parameters but also extracts key low - level features. The difference between the grayscale image and the color image lies only in the different initial input channel numbers. Subsequently, the feature map passes through 16 MWG - bneck modules in sequence, capturing diverse shape features layer by layer. The MGLR module is embedded in the expansion layer of each bottleneck, and the down - sampling process of the feature map is implemented using WTConv. Next, the feature map is up - dimensionalized through a 1×1 convolution and compressed along the channel direction through global average pooling to retain the final channel descriptor. Finally, dropout is added to the up - dimensionalized fully - connected layer to prevent overfitting, and classification is completed through a Softmax classifier to output the final result. The relevant parameters of the MWC - GhostNet model are shown in Table 1. Extension represents the number of channels after the MWG - bneck passes through the expansion layer, and k represents the convolution kernel size of WTConv when stride = 2.

[0060] Table 1. Model Parameter Table

[0061]

[0062] Step 3: Use the plankton dataset to train the MWG - GhostNet model. To address the long - tail problem in plankton classification, we use the Class - Balanced Focal Loss (CB Focal Loss) function to improve the classification performance of the model. Previous studies usually dealt with the long - tail problem by resampling or inverse weighting according to the sample frequency, but these methods were not very effective in practical applications. Using weights can compensate for the minority classes, enabling the model to pay more attention to these less - represented classes during training and improving the classification performance. However, correctly adjusting α is usually time - consuming and depends on the researcher's experience. CB Focal Loss proposes a class balance, calculating the weight of each class according to Equation 4, where n i represents the actual number of samples of a certain class, and β is a balance coefficient used to control the balance degree of class weights. Under the same conditions, when the actual number of samples increases, the effective number of samples will decrease, thereby reducing the obtained weight w i .

[0063]

[0064] CB Focal Loss(pt)= - w′i(1 - p t ) γ log(p t ) (6)

[0065] As shown in Equation 6, replacing the α parameter in Focal Loss with this type of balance can effectively solve the problem that some categories have too strong or too weak influence due to improper weight setting. In actual experiments, we ensure a more balanced influence among categories by normalizing the weights (see Equation 5). During the training process, CB Focal Loss is used to replace the conventional Cross-Entropy Loss (CE Loss), enabling the model to appropriately focus on each category during training, especially assigning higher weights to those difficult samples that are easily overlooked. This helps to strengthen the learning of minority classes and effectively avoid overfitting or underfitting problems caused by class imbalance.

[0066] Step 4: Model classification prediction:

[0067] First, obtain the test set images. These images have not been seen by the model during the training process, so they can provide an objective evaluation criterion to help understand the performance of the model in actual applications. Second, the MWC-GhostNet model loads the trained weights and passes these images into the trained MWC-GhostNet model for prediction to output the final result.

[0068] Then, conduct a comparative analysis of the algorithm performance.

[0069] 1. Analysis of classification accuracy results

[0070] To study the image classification effect of the MWC-GhostNet model, a comparative experiment is adopted. The comparison methods include the established plankton classification methods in the field, ResNet50+SVM, ZOOPLANKTON, InceptionNet, ShuffleNet, and the recent efficient lightweight networks MobileNetV2, MobileNetV3, EfficientNet b1, GhostNet, GhostnetV2. The evaluation metrics include the number of model parameters, FLOPs, Top-1 accuracy, Top-5 accuracy, and F1 score.

[0071] Table 2. Comparative experiments of different models on the Kaggle121 dataset

[0072]

[0073]

[0074] On the Kaggle121 dataset, we conducted experiments according to the existing configurations. Each model was trained for 100 epochs in the same experimental environment. Except for the proposed model, the cross-entropy loss function was used by default for the other models. The experimental results are shown in Table 2. Among them, Xiu et al. used VGG19 and ResNet32 networks for classification tasks on the Kaggle121 dataset, and we followed their experimental results as the optimal benchmark on this dataset. MWC-GhostNet achieved the best performance in terms of Top-1 accuracy, Top-5 accuracy, and F1 score, which were 78.43%, 96.95%, and 0.6702 respectively. Although traditional VGG, ResNet, ZOOPLANKTON, and InceptionNet have higher numbers of parameters and FLOPs, they did not show any advantages in performance. Compared with other lightweight convolutional neural network (CNN) models, MWC-GhostNet showed significant advantages in FLOPs. Although the number of parameters of MWC-GhostNet is slightly higher than that of MobileNetV2, its accuracy is significantly better than that of MobileNetV2, with the Top-1 accuracy increased by 4.12%. Compared with GhostNet and GhostNetV2, MWC-GhostNet improved the F1 score by 5.83% and 3.39% respectively, although the numbers of parameters and computational complexity are similar. MobileNetV3, EfficientNet b1, and GhostNet have improved accuracy compared to MobileNetV2, possibly due to the use of the SE module. In contrast, EfficientNet b1 has a larger number of parameters, but its performance on F1 is 8.43% lower than that of MWC-GhostNet. This performance gap can be attributed to the fact that the MGLR module in MWC-GhostNet shows stronger performance than the SE module in the plankton recognition task and can capture key global information.

[0075] Merge means that we merged the Kaggle121 dataset according to the biological sub-classifications using the method of Orenstein et al., and finally formed 37 classes. The results of MWC-GhostNet are significantly better than those of the pre-trained AlexNet, with the Top-1 accuracy increased by 6.39%, while the number of parameters is only 6.63% of that of AlexNet. The experiment shows that when the details of image features are less, using a model with a larger number of parameters for classification may not necessarily achieve better results.

[0076] Table 3. Comparative experimental results on the PMID2019 dataset

[0077]

[0078]

[0079] Table 3 shows the experimental results on the PMID2019 dataset. Each model was trained for 30 epochs in the same environment. MWC-GhostNet achieved the best performance in terms of Top-1 accuracy, Top-5 accuracy, and F1-score, which were 98.81%, 99.85%, and 0.9618 respectively, significantly leading other models, indicating its high accuracy in class prediction. Traditional models still did not show advantages in various metrics. EfficientNet b1 achieved sub-optimal results on this dataset, benefiting from its advanced architecture and larger number of parameters. Among lightweight networks, MobileNetV2 and GhostNet showed poor performance due to the lack of fine-grained structures for image feature extraction.

[0080] The experimental results show that among the above models, MWC-GhostNet demonstrated state-of-the-art performance in the plankton classification task, showing high practical application potential with its excellent accuracy and low computational overhead.

[0081] 2. Model Visualization Analysis

[0082] The original GhostNet was used as the baseline network and compared with our improved model. Attached Figure 8 shows the confusion matrices of our proposed model and the baseline model on two datasets. The rows of the matrix represent the predicted classes, and the columns represent the true labels. We used TPR (True Positive Rate) as the evaluation metric to intuitively show the recall rate of the model for each class. The green area represents the proportion of correct predictions, and the red area represents the proportion of incorrect predictions. The input was fed into the trained MWC-GhostNet model for prediction to output the final results.

[0083] To more intuitively show the improvement effect of MWC-GhostNet, we extracted the deep features of the layer before the global pooling layer. The features of each channel were superimposed to generate a class activation map (CAM) to show the regions of interest of the model. The results are shown in Attached Figure 9 and Figure 10As shown in the figure. The first row in the figure is the original image, the second row is the CAM generated by the baseline network, and the third row is the CAM generated by MWC-GhostNet. On the Kaggle121 dataset, our model shows more concentrated attention areas on plankton images, especially in the central area and complex structures, significantly outperforming the baseline model. In (a) to (c), MWC-GhostNet focuses on the highly activated areas in the central region of the organism, while the attention of the baseline model is more dispersed; in (c) to (f), the model has more precise attention to the contour of the long-strip plankton and can clearly capture the structural details of the organism, indicating that our improvement can enhance the shape recognition ability.

[0084] Similarly, as shown in the appendix Figure 10 As shown, on the PMID2019 dataset, compared with the baseline model, our model has higher accuracy and stronger background suppression ability in plankton recognition. For example, in (a) and (b), the activation areas of the baseline model are more dispersed, while our model focuses more on the main part of the plankton, showing a more precise localization effect. In addition, in (c), (e), and (f), there are still some activations in the heatmap of the baseline model on the background, which may lead to misjudgment, while our model significantly reduces the activation of the irrelevant background, showing a better background suppression effect. Especially in Figure (d), the plankton presents a multi-segment structure, and our model can more completely focus on each segment structure, demonstrating superior structure perception ability.

[0085] Overall, our model performs excellently in the focusing accuracy of the target area, which is mainly due to the advantages of the MGLR module in global information extraction and processing. By combining the MGLR module and wavelet transform, our model significantly enhances the ability to perceive structural information. The MGLR module not only improves the multi-scale attention ability of features but also shows excellent effects in feature selection and background suppression, ensuring that plankton can be accurately separated from the complex background, thus improving the adaptability and classification accuracy of the model in practical applications.

[0086] 3. Analysis of the Influence of Loss Function Parameters

[0087] To verify the influence of the loss function on the prediction results, the effects of three loss functions, namely CE Loss, Focal Loss, and CB FocalLoss, on the F1 score were compared. To optimize the performance, we conducted hyperparameter searches for γ and β within the set range, where the value of γ is {1, 2}, the value of β is {0.9, 0.99, 0.999, 0.9999}, and the α of Focal Loss remains the default value to evaluate the influence of γ and β on the model performance. The appendix Figure 11 shows the search curves on the two datasets.

[0088] As shown in Figure 11 (a), on Kaggle121, when the Focal Loss lacks the adjustment of α, the performance is even worse than that of the CE Loss, indicating that the weights are very crucial for the regulation of the imbalanced dataset. When using the CB Focal Loss with a smaller β, the F1 result exceeds the other two loss functions. Similarly, PMID2019 (as shown in Figure 11 (b)) achieves the maximum F1 score at β = 0.99. In addition, the same situation exists in both datasets. As β increases, the F1 score decreases instead. This is mainly because the increase in β leads to a decrease in the class weights, reducing the influence of this class on the loss calculation and causing the model to not fully converge within a limited number of iterations. The experimental results show that the CB Focal Loss can significantly improve the F1 score, indicating that the model has effectively improved in identifying difficult samples in the dataset. In practical applications, the convergence speed directly affects resource consumption. Therefore, appropriate hyperparameters should be selected according to the task requirements to achieve the best balance between performance and resource consumption.

[0089] 4. Analysis of the Results of the Algorithm Ablation Experiment

[0090] To further verify the contribution of the proposed module to the model performance, a systematic ablation experiment was conducted on Kaggle121. Analyze its impact on the classification effect to clarify the role of each module in the overall performance improvement. According to the results in Table 6, introducing the multi-level WTConv can effectively improve the model performance on the basis of the original DWConv. On the Kaggle121 dataset, the number of parameters of the DWConv is 4.05M, and the FLOPs is 0.392G; after introducing the WTConv, the number of parameters of 3 levels is 4.31M, and the FLOPs is 0.479G. The number of parameters and FLOPs only increase by 0.26M and 0.087G respectively. Although the accuracy (0.6478) of the 3-level WTConv has decreased, the accuracy of the 2-level WTConv reaches 0.6639, and the recall rate is 0.6365, showing a significant balance between classification accuracy and comprehensiveness. The 2-level wavelet transform performs best among multiple indicators, demonstrating the effectiveness of refining the convolutional hierarchy as an innovative design. Generally speaking, further increasing the number of wavelet transform layers will increase the computational cost and the number of parameters of the model, but the performance improvement is limited and may lead to overfitting or introduce noise.

[0091] Table 4 Comparison Results of Multi-Level Wavelet Transform

[0092]

[0093] Table 5 shows the ablation experiment results of this study with common efficient attention mechanism modules (SE, CBAM, CA, DFC, EMA) and its own module. Except that Baseline added the SE module to some layers according to the original network architecture, the remaining experiments were similar to MGLR, embedding the attention module into each layer of the network to replace the SE module. Although the SE module has the smallest FLOPs, it has a large number of parameters and the worst accuracy performance. CBAM has improved in precision, but has a low recall rate, resulting in poor overall balance, and its accuracy is similar to that of the SE module. While maintaining a low number of parameters and FLOPs, CA and DFC have improved in various indicators. EMA achieved the best performance in accuracy and F1 score through cross-space multi-scale fusion, but its computational cost is as high as 1.94G, which may lead to slow inference speed in the case of limited hardware facilities. MGLR shows the sub-optimal solution among all modules, and its recall rate and precision reach 0.6808 and 0.6274 respectively, which are the best values among all modules. Although its accuracy and F1 are slightly lower than those of EMA, its number of parameters and FLOPs are significantly reduced, making it suitable for application in the embedded device environment.

[0094] Table 5. Ablation Experiment Table

[0095]

[0096]

[0097] The experimental results show that after adding WTConv and CB Focal Loss on the basis of MGLR, our model shows significant performance improvement in the plankton image classification task. Without pre-training, the accuracy, precision, recall rate and F1 score of our model reach 78.43%, 0.6953, 0.6639 and 0.6702 respectively, which are improved by 1.81%, 4.57%, 5.26% and 5.83% compared with Baseline.

[0098] In practical applications, transfer learning, as a common technical means, can effectively accelerate the training process and improve the performance of the model on the target task. We used a model pre-trained on a large-scale dataset (denoted by "#") for training. After transfer learning, the accuracy, precision, recall, and F1-score of the model reached 79.03%, 0.7046, 0.6684, and 0.6718 respectively, showing further improvement compared to the results of the non-pre-trained model. It is worth noting that our model has a similar number of parameters to the Baseline, and the computational cost (FLOPs) only increased by 0.19G, indicating that the model can maintain a low computational overhead while improving performance. This makes our method not only achieve a significant improvement in classification accuracy but also possess good computational efficiency, making it suitable for embedded devices and hardware environments with limited resources in practical applications.

Claims

1. A shape structure enhanced lightweight plankton classification method based on multi-scale gated attention, characterized in that: The following steps are involved: Step 1, obtaining a plankton image and constructing a data set; each image in the data set includes an image label, and the image label is a plankton species type; Step 2: Construct a plankton classification model. The plankton classification model uses GhostNet as the backbone network. The structure of the plankton classification model includes, from input to output: an input layer, several bottleneck layers MWC-bneck, a global average pooling layer, a dimensionality increase layer, and a classification layer; The input of the input layer Input is the image obtained in step 1; Each bottleneck layer MWC-bneck includes: Ghost module 1 in the GhostNet network, Ghost module 2 in the GhostNet network, MGLR attention module, and wavelet convolution module; the Ghost module 1 increases the dimension of the input features of the bottleneck layer MWC-bneck, and MGLR extracts global information from the input features of the bottleneck layer MWC-bneck and multiplies them point by point with the output of Ghost module 1, and inputs them into the wavelet convolution module, which is used to respond to low-frequency information in the reweighted feature map and strengthen the contour structure information; the Ghost module 2 reduces the dimension of the output of the wavelet convolution module; The feature map output by the global average pooling layer is expanded into a one-dimensional vector and then upgraded through the dimension-enhancing layer. Then, the classification layer is used to map it to the dimension of the number of categories for classification tasks. Step 3: Use the plankton dataset to train the plankton classification model, and use the trained plankton classification model; Step 4: Model classification prediction Obtain the plankton image to be classified, pass it into the trained plankton classification model, and obtain the classification result.

2. The shape structure enhanced lightweight plankton classification method based on multi-scale gated attention as claimed in claim 1, characterized in that: In step 3, the Class-Balanced Focal Loss loss function is used as the optimizer to train the plankton classification model by back-propagation gradient update.

3. The shape structure enhanced lightweight plankton classification method based on multi-scale gated attention as claimed in claim 1, characterized in that: In step 2, the number of the bottleneck layer MWC-bneck is greater than 16.

4. The shape structure enhanced lightweight plankton classification method based on multi-scale gated attention as claimed in claim 1, characterized in that: In step 2, the dimension-raising layer includes a 1x1 convolution.

5. The shape structure enhanced lightweight plankton classification method based on multi-scale gated attention as claimed in claim 1, characterized in that: In step 2, the classification layer includes a Softmax classifier.

6. The shape structure enhanced lightweight plankton classification method based on multi-scale gated attention as claimed in claim 1, characterized in that: The classification method is applied to underwater automated monitoring of plankton.

7. The shape structure enhanced lightweight plankton classification method based on multi-scale gated attention as claimed in claim 1, characterized in that: In step 2, the MGLR attention module performs horizontal aggregation and vertical aggregation according to the following formula; The input feature map of the MGLR attention module is x,F H and F W Respectively represent the convolution kernel operations in the horizontal and vertical directions, σ represents the Sigmoid activation function, which is used for gating operations; x h,w Represents each pixel in x. In horizontal aggregation, each pixel in the feature map is weighted summed with the features of other pixels in the same row of the pixel. After gated activation, the result is multiplied point by point with the original pixel to obtain the horizontally aggregated feature map α′ hw ; The vertical aggregation is based on the horizontal aggregation, and the aggregation is performed along the w axis, and multiplied by the gate operation to finally obtain the aggregated feature map α hw ; In the process of horizontal aggregation and vertical aggregation, several convolution kernels of different sizes are used to extract features in the horizontal and vertical directions.

8. The shape structure enhanced lightweight plankton classification method based on multi-scale gated attention as claimed in claim 7, characterized in that: In step 2, the convolution kernels of different sizes are specifically the following convolution kernels: 1x5, 1x7, and 1x9.