A plankton fine-grained classification method and system based on a progressive down-sampling three-stage architecture
Patent Information
- Application Number
- CN202610947075.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-29
AI Technical Summary
[0008]本发明的目的在于针对现有深度学习模型在浮游生物细粒度分类任务中存在的因过度下采样导致微观形态特征丢失、特征提取模块难以兼顾全局与局部细节、以及缺乏通道特征动态门控约束等技术问题,提出一种基于渐进式下采样三阶段架构的浮游生物细粒度分类方法及系统
[0024]1.本发明在特征提取模块中仅执行两次下采样操作,使特征提取模块输出的特征图空间分辨率维持在原始输入尺寸的1/16(例如输入为224×224像素时,最终特征图14×14像素),这种设计在物理结构上最大程度地保留了浮游生物的微观高频形态信息;通顺,本发明的下采样模块基于倒残差瓶颈结构,并集成通道注意力机制以对特征响应进行重校准,从而缓解空间下采样造成的细粒度信息丢失。
Smart Images

Figure CN122454570B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of planktonic classification technology, specifically relating to a fine-grained planktonic classification method and system based on a progressive downsampling three-stage architecture. Background Technology
[0002] Plankton are primary producers and key components of marine and inland water ecosystems. Their population succession, abundance changes, and community structure are core indicators reflecting water quality, assessing fishery resources, and providing early warning of ecological and environmental health (such as red tides and algal blooms). Therefore, achieving in-situ, efficient, and automated classification and identification of plankton is of great practical significance for responding to national ecological protection strategies and improving the level of aquatic biodiversity monitoring.
[0003] Traditional planktonic identification primarily relies on manual microscopic examination. This method is not only time-consuming and labor-intensive, but also highly dependent on the experience of taxonomists, making it subjective and unsuitable for the demands of modern, high-frequency, and large-scale operational automated monitoring. With the rapid development of deep learning and machine vision technologies, image classification methods based on convolutional neural networks (CNNs) and vision transformers have been increasingly applied to the automated identification of planktonic organisms, improving classification efficiency to some extent. However, existing deep learning models still face significant technical bottlenecks and limitations when applied to complex, fine-grained planktonic classification tasks, specifically in the following aspects:
[0004] 1. Oversampling in mainstream networks leads to loss of fine-grained features: Currently, mainstream visual backbone networks (such as ConvNeXt and Swin Transformer) generally adopt a four-stage hierarchical architecture. To control computational complexity and expand the receptive field, these networks perform spatial downsampling during feature extraction, resulting in the spatial resolution of the final feature map being drastically compressed to 1 / 32 or even lower than the original input image. Since images of planktonic organisms typically exhibit extremely high inter-class similarity (different species have very similar morphologies) and intra-class polymorphism (different poses and features of the same planktonic organism can have multiple forms), their key classification criteria often rely on extremely subtle high-frequency local structures such as tentacles, cilia, caudal forks, and textures. The 1 / 32 spatial compression of existing networks causes these fine-grained features, occupying very few pixels, to be smoothly or irreversibly lost in the deep network, directly limiting the upper limit of fine-grained classification accuracy.
[0005] 2. Existing feature extraction modules struggle to balance global dependencies with local details: Most feature extraction modules within mainstream visual models employ standard self-attention or basic spatial reduction attention mechanisms, often overemphasizing the establishment of global macroscopic dependencies while severely neglecting strong structural relationships between pixels in local neighborhoods. During feature dimensionality reduction, this easily leads to blurring, breakage, or distortion of microscopic local details such as planktonic edge cilia and transparent films, resulting in insufficiently refined spatial features that cannot effectively support demanding fine-grained classification tasks.
[0006] 3. Traditional feedforward networks lack dynamic gating constraints on channel features: Multilayer perceptrons (MLPs) or feedforward neural networks (FFNs), widely used in traditional deep learning feature extraction modules, typically employ static nonlinear activation functions, transforming the feature streams of all channels indiscriminately. In fine-grained planktonic identification, the core clues for distinguishing highly similar species are often hidden in a few specific channels. Existing modules, lacking dynamic gating and filtering mechanisms, cannot adaptively suppress background noise and amplify key identification features, greatly limiting the model's ability to represent high-dimensional features.
[0007] In summary, there is an urgent need for a planktonic classification method that can retain high spatial resolution in the macroscopic physical architecture and possess extremely strong feature extraction capabilities in the microscopic feature extraction module (taking into account both local spatial detail enhancement and channel dynamic gating) in order to completely break through the accuracy bottleneck of existing automated monitoring systems. Summary of the Invention
[0008] The purpose of this invention is to address the technical problems of existing deep learning models in the fine-grained classification of plankton, such as loss of microscopic morphological features due to oversampling, difficulty in the feature extraction module to take into account both global and local details, and lack of dynamic gating constraints for channel features. This invention proposes a fine-grained plankton classification method and system based on a progressive downsampling three-stage architecture. This method can achieve extremely strong feature extraction and representation while maintaining high spatial resolution, effectively addressing the challenge of classifying highly similar fine-grained species, and achieving efficient and accurate classification.
[0009] In a first aspect, the present invention provides a planktonic fine-grained classification method based on a progressive downsampling three-stage architecture, the method comprising:
[0010] Images of the plankton to be tested are acquired and input into a plankton classification model. The plankton classification model includes a progressive feature mapping module, a feature extraction module, and a classification head connected in sequence. The progressive feature mapping module is used to downsample the plankton images. The feature extraction module is used to process the downsampled feature map and input the processing result into the classification head to obtain the types of plankton in the plankton images.
[0011] The feature extraction module includes three feature extraction sub-blocks; a downsampling module is connected in series between adjacent feature extraction sub-blocks; in the downsampling module, the input feature map is processed by a series of pointwise convolutional layers and depthwise separable convolutions to obtain an intermediate feature map; the intermediate feature map is processed by a compression-excitation attention module, and the processing result is fused with the intermediate feature map before a convolution operation is performed to obtain the output feature map of the downsampling module.
[0012] Preferably, the feature extraction sub-block includes a local feature enhancement module and a global feature extraction module connected in series; the local feature enhancement module is used to enhance the local spatial micro-details of the input feature map; the global feature extraction module is used to take into account global long-range dependency modeling.
[0013] Preferably, in the local feature enhancement module, local features are extracted from the input feature map through a series of normalization layers and depthwise separable convolutional layers, and the processing result is fused with the input feature map to obtain a first fused feature map; the first fused feature map is processed using a series of normalization layers and a feedforward network, and the processing result is fused with the first fused feature map to obtain the local enhanced feature map output by the local feature enhancement module.
[0014] Preferably, in the global feature extraction module, the input feature map is processed by a series of normalization layers and a global mixer, and the processing result is fused with the input feature map to obtain a global feature map; the global feature map is processed by a series of normalization layers and a gated feature reconstruction module, and the processing result is fused with the global feature map to obtain the output feature map of the global feature extraction module.
[0015] Preferably, in the global mixer, the input feature map is processed by a series of deep convolutional layers and normalization layers. The processing result is then projected into a value vector and a key vector, and an attention score is calculated by combining the result with the query vector projected from the input feature map to obtain the attention output. The attention output is then processed by a deep convolutional layer, and the processing result is fused with the attention output before being projected to obtain the output feature map of the global mixer.
[0016] Preferably, the gated feature reconstruction module includes a value branch and a gate branch; the value branch processes the input feature map through a concatenated convolutional layer and an activation function to obtain the output result of the value branch; the gate branch processes the input feature map through a convolutional layer to obtain the output result of the gate branch; the output results of the value branch and the gate branch are fused and further processed through a convolutional layer to obtain the output feature map of the gated feature reconstruction module.
[0017] Preferably, the progressive feature mapping module includes multiple convolutional layers, and the stride of a single convolutional layer is no greater than 2.
[0018] Preferably, the plankton classification model is trained using a dataset containing images of different plankton. During training, the cross-entropy loss function is used to calculate the error between the probability distribution predicted by the model and the real species labels, and to guide the parameter update of the plankton classification model.
[0019] As a preferred approach, the planktonic images in the dataset are preprocessed before model training. The preprocessing method is as follows: after unifying the size of all planktonic images, the planktonic images are processed by random rotation, translation, cropping, and contrast adjustment.
[0020] Secondly, this invention provides a plankton fine-grained classification system based on a progressive downsampling three-stage architecture, which is used to execute the aforementioned plankton fine-grained classification method. The plankton fine-grained classification system includes an image acquisition module, a preprocessing module, a plankton classification module, a training module, and a classification output module. The image acquisition module is used to acquire plankton images. The preprocessing module is used to perform scale transformation and data augmentation on the plankton images. The plankton classification module is used to classify the plankton types in the input plankton images. The training module is used to train the plankton classification model in the plankton classification module. The classification output module is used to display the plankton classification results obtained by the plankton classification module in real time.
[0021] Thirdly, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the memory stores the computer program; and the processor executes the aforementioned plankton fine-grained classification method.
[0022] Fourthly, the present invention provides a readable storage medium storing a computer program; when executed by a processor, the computer program is used to implement the aforementioned plankton fine-grained classification method.
[0023] The beneficial effects of this invention are:
[0024] 1. This invention performs only two downsampling operations in the feature extraction module, maintaining the spatial resolution of the feature map output by the feature extraction module at 1 / 16 of the original input size (e.g., when the input is 224×224 pixels, the final feature map is 14×14 pixels). This design preserves the microscopic high-frequency morphological information of plankton to the greatest extent in terms of physical structure. Furthermore, the downsampling module of this invention is based on an inverted residual bottleneck structure and integrates a channel attention mechanism to recalibrate the feature response, thereby mitigating the loss of fine-grained information caused by spatial downsampling.
[0025] 2. This invention innovatively proposes a progressive downsampling three-stage architecture, which effectively alleviates the irreversible loss of microscopic high-frequency morphological features such as tentacles and cilia of plankton, significantly improves the model's sensitivity to details and shape contours, and provides underlying visual cues for the accurate differentiation of fine-grained similar species.
[0026] 3. This invention introduces a global mixer in the feature extraction module, which significantly reduces the computational complexity of processing high-resolution feature maps. On the other hand, through the deep convolutional branches concatenated after attention operations, it effectively compensates for the loss of high-frequency edges and strong correlations of local pixels during spatial reduction, achieving accurate fusion of global macro-dependencies and local micro-details. At the same time, this invention also employs a gated feature reconstruction module in the feature extraction module, enabling the model to obtain nonlinear constraint capabilities similar to a "dynamic gate," which can adaptively suppress noise channels caused by complex water backgrounds, while multiplying the specific channel features containing key identification information, greatly improving the expression accuracy of high-dimensional features.
[0027] 4. This invention not only significantly improves the accuracy of fine-grained classification of plankton images, achieving a Top-1 accuracy of up to 97.94% on datasets such as HZWPlankton8, but also demonstrates excellent prediction accuracy, good convergence, and stability. It meets the reliability and timeliness requirements of underwater high-throughput in-situ imaging instruments for data processing, strongly supporting the practical application of automated in-situ plankton monitoring systems. Attached Figure Description
[0028] Figure 1 This image displays 20 classes of copepod samples from the ZooplanktonBench50 dataset.
[0029] Figure 2 This is a schematic diagram of the overall architecture of the planktonic classification model in this invention.
[0030] Figure 3 This is a schematic diagram of the downsampling module in this invention.
[0031] Figure 4 This is a schematic diagram of the feature extraction sub-block in this invention.
[0032] Figure 5 This is a schematic diagram illustrating the convergence analysis of training loss in this invention.
[0033] Figure 6 This is a comparison chart of class activations for 20 representative copepod classes on the ZooplanktonBench50 dataset, based on the present invention.
[0034] Figure 7The image shows a comparison of class activations of the baseline model EfficientNet on the ZoolanktonBench50 dataset for 20 representative copepod classes. Detailed Implementation
[0035] The present invention will be further described below with reference to the accompanying drawings.
[0036] A fine-grained plankton classification method based on a progressive downsampling three-stage architecture is proposed. The plankton fine-grained classification system includes an image acquisition module, a preprocessing module, a plankton classification module, a training module, and a classification output module. The image acquisition module acquires plankton images. The preprocessing module performs scaling and data augmentation on the plankton images. The plankton classification module classifies the plankton types from the input plankton images. The training module trains the plankton classification model in the plankton classification module. The classification output module displays the plankton classification results acquired by the plankton classification module in real time.
[0037] This method for fine-grained classification of plankton includes the following steps:
[0038] Step 1: Obtain the plankton image dataset and perform preprocessing:
[0039] This invention uses four representative plankton datasets, as shown in Table 1, which cover various application scenarios such as in-situ monitoring and microscopic imaging, to comprehensively verify the generalization performance of the model.
[0040] Table 1 Overview of the experimental zooplankton image dataset
[0041]
[0042] The ZooplanktonBench50 dataset contains 11,677 images across 50 classes of marine zooplankton. The dataset originates from the In-Situ Imaging System for Zooplankton (ISIIS), deployed on July 24, 2011, in the northern Gulf of Mexico. The 20 copepod classes included in this dataset exhibit extremely high inter-class similarity. Differences between these classes are often only evident in high-frequency microscopic details such as tentacle length and tail fork shape. Figure 1 As shown, due to interference from illumination changes and semi-transparent occlusion in in-situ imaging, extremely high demands are placed on the resolution preservation ability of the visual backbone network in deep features.
[0043] The ZooCAMNet76 dataset contains 17,506 images representing 76 classes of plankton. The dataset originated from a long-term monitoring project in the Bay of Biscay from May 2016 to November 2020, collected using a continuous walking fish egg sampler (CUFES) with a mesh size of 315µm. The original database contained 1.28 million images (93 classes). After denoising and filtering for non-living and plant samples, 17,506 images belonging to the 76 classes of zooplankton were retained. The dataset has a highly complex internal hierarchy, encompassing 22 classes of copepods, 9 classes of hydroids, and 7 classes of molluscs.
[0044] The ZooScanNet101 dataset contains 24,199 images representing 101 classes of plankton. It is the largest plankton dataset in terms of temporal span and global geographic coverage, with original imagery spanning nearly 25 years from 1995 to 2019. The images were acquired by the ZooScan system through fishing in the world's oceans and laboratory scanning. A 24,199 experimental subset of 101 classes was constructed from the original 1.45 million scans. This dataset exhibits extremely high biodiversity, comprehensively covering 20 copepods, 21 hydroids, and 17 molluscs, among others. This high morphological complexity and spatiotemporal distribution make it an ideal benchmark for evaluating algorithms handling large-scale, fine-grained classification problems.
[0045] The HZWPlankton8 dataset is an internal dataset built by our research team, with images collected from the in-situ monitoring system in Hangzhou Bay. This dataset contains 3674 high-resolution images of eight common nearshore plankton species (such as the delicate water flea, the obese arrow worm, and the Shengshan cup jellyfish).
[0046] All datasets were divided into training, validation, and test sets in a 6:2:2 ratio to increase sample diversity and alleviate model overfitting. To address the common background noise interference in plankton images, preprocessing and data augmentation were performed on the plankton images in the dataset: First, all original images were uniformly scaled to 224×224 pixels; second, the RandAugment enhancement strategy was introduced, using random rotation, translation, cropping, and contrast adjustment to improve model robustness.
[0047] Step 2: Construct a planktonic classification model
[0048] In plankton identification tasks, shape, edge texture, and microscopic appendages (such as tentacles and tail forks) are key information for distinguishing highly similar species. However, mainstream visual models (such as MetaFormer and Swin Transformer) generally adopt a four-stage architecture, ultimately downsampling the feature map to 1 / 32 of its original size. This deep spatial compression leads to the smooth or irreversible loss of high-frequency fine-grained features that occupy very few pixels. To address this, this invention proposes a progressive downsampling three-stage architecture. The first downsampling module of this architecture uses a progressive feature mapping module, and the latter two downsampling modules consist of MBConv downsampling modules. Unlike traditional models, this invention stops the downsampling operation after the third feature extraction stage of the plankton classification model, maintaining the spatial resolution of the feature map output by the backbone network at 1 / 16 of the original input size (e.g., when the input is 224×224 pixels, the final feature map is 14×14 pixels). This design preserves the microscopic high-frequency morphological information of plankton to the greatest extent in terms of physical structure.
[0049] like Figure 2 As shown, the planktonic classification model includes a progressive feature mapping module, a feature extraction module, and a classification head. To further reduce information loss in the initial stage, this invention abandons the traditional single-layer large-stride convolutional kernel (such as...). (with a step size of 4), instead adopting three cascaded units. Convolutional layers form a progressive feature mapping module. In this module, the stride of the first two convolutional layers is set to 2, and the stride of the last convolutional layer is set to 1. This design enables the feature map to be downsampled to 1 / 4 while expanding the receptive field more smoothly, capturing richer low-level morphological details.
[0050] The feature extraction module comprises three sub-blocks. Downsampling modules are cascaded between adjacent feature extraction sub-blocks. This design is based on an inverse residual bottleneck structure and integrates a channel attention mechanism to recalibrate the feature response, thereby mitigating the loss of fine-grained information caused by spatial downsampling. Given input features... .
[0051] like Figure 3 As shown, in the downsampling module, the input features first pass through a... Pointwise convolution expands the channels to obtain (Expanding ratio e = 4), then with a step size of 2. Depthwise separable convolution is used to perform spatial downsampling to obtain intermediate feature maps. After adaptive recalibration of the channel responses by the compression-stimulation attention module (SE), the data is fused with the intermediate feature maps, and finally... Convolution reduces its channel number from Projecting to the next stage Its complete computation flow can be represented as:
[0052]
[0053]
[0054] like Figure 4 As shown, unlike traditional single-space or single-channel mixers, the feature extraction sub-block employs an interleaved "convolution-attention" hybrid design to construct a cascaded residual structure. The feature extraction sub-block includes a series of local feature enhancement modules and a global feature extraction module, aiming to balance global long-range dependency modeling with local spatial micro-detail enhancement.
[0055] In the local feature enhancement module, through Depthily separable convolutional (DWConv) layers extract local features from the input feature map to capture spatial relationships between adjacent pixels (such as high-frequency local features like the edges and textures of plankton). Group Normalization (GN) and a feedforward network (FFN) are combined to lay the foundation for fine-grained features. Given input features... The mathematical expression for its local enhancement process is:
[0056]
[0057]
[0058] The global feature extraction module focuses on semantic aggregation of the global context. It sequentially establishes long-range dependencies across spatial boundaries using a global mixer, and then maps and refines channel dimensions using a multilayer perceptron. Within the global feature extraction module, concatenated group normalization and a global mixer are employed to enhance local features. The process is performed, and the results are compared with local enhancement features. Perform residual fusion to obtain global features. This enables global feature interaction across the spatial dimension. After establishing global dependencies in the spatial dimension, to avoid homogenizing all features and suppress water background noise, a series of group normalization and gated feature reconstruction modules are used to process the global features. Process the data and combine the results with global features. Residual fusion is performed to obtain the output feature map of the global feature extraction module. The above process is expressed by the following formula:
[0059]
[0060]
[0061] The global mixer is internally implemented using a Locally Enhanced Spatial Reduction Attention (LeSRA) mechanism. For clarity in subsequent formulas, let the normalized input features be... This mechanism maintains high spatial resolution of the query vector Query(Q) while simultaneously applying depthwise convolutions with stride R to the input features. Spatial dimensionality reduction is performed to generate low-resolution key vectors Key(K) and value vectors Value(V).
[0062]
[0063]
[0064]
[0065] Wherein, projection matrix , , , The semantic similarity between high-resolution local details and low-resolution global structure was calculated. The dimension representing the key vector prevents the Softmax function from entering the saturation region where the gradient is minimal, thus ensuring the numerical stability of training. Global feature interaction is achieved through scaling the dot product attention to obtain the attention output. .
[0066]
[0067] To compensate for the loss of smoothness in local details due to space reduction, this unit strictly connects a [missing information - likely a parameter] at the attention output end. Depthwise convolution is used as a local enhancement residual branch to re-extract high-frequency information. Finally... The output Y is obtained by fusing and projecting two parts:
[0068]
[0069]
[0070] In the gated feature reconstruction module, let the normalized features be... This module contains two parallel paths: the value branch uses a non-monotonic SiLU activation function; the gate branch is entirely data-driven and does not use an activation function. Both are... Convolution implementation:
[0071]
[0072]
[0073] Through the element-wise Hadamard product of the two branches, the model adaptively amplifies foreground features and attenuates background interference:
[0074]
[0075] Finally passed The final feature is obtained by substituting the projection convolution output into the outer residual connection.
[0076]
[0077] To balance computational overhead, the intermediate hidden layer dimension of the SwiGLU MLP is... Asymmetric parametric alignment constraints are used: ,in Indicates the number of input channels. This indicates the normal channel spread ratio of the MLP. This indicates a round-down operation.
[0078] Step 3: Model Training
[0079] like Figure 5 As shown, a preprocessed plankton dataset is used to iteratively train a plankton classification model. This embodiment is implemented based on the PyTorch deep learning framework, and the core configuration of the training process is as follows: First, the cross-entropy loss function is used to calculate the error between the probability distribution predicted by the model and the real species labels. Second, the AdamW optimizer is used, along with a cosine annealing strategy to dynamically schedule the learning rate, and the model parameters are updated through the backpropagation algorithm. During training, a learning rate preheating and gradient pruning mechanism is set to prevent gradient explosion during deep network training from scratch, ensuring the numerical stability of the training process. In addition, to further enhance the robustness of the model and alleviate overfitting, hybrid data augmentation strategies such as RandAugment, MixUp, and CutMix are introduced during training to expand sample diversity. After multiple iterations until the model's training loss converges, the optimized plankton classification model is finally obtained.
[0080] Step 4: Comparative Analysis of Model Classification Prediction and Algorithm Performance
[0081] First, test set images are acquired. These images are not seen by the model during training, thus providing an objective evaluation metric to help understand the model's true generalization performance in real-world applications. Second, this invention (PlanktonNet model) loads the trained weights and feeds the test set images into the already trained PlanktonNet model for prediction, outputting the final classification result. Then, algorithm performance comparison and analysis are performed. To comprehensively evaluate the performance of the PlanktonNet model, this embodiment conducts an in-depth comparison from two dimensions: classification accuracy and model visualization.
[0082] 1. Classification accuracy results analysis
[0083] To investigate the image classification performance of the PlanktonNet model, a systematic comparative experiment was conducted on four datasets, and the results are shown in Table 2. The benchmark methods used for comparison covered the three most representative backbone networks in the current field of computer vision (e.g., ConvNeXt V2, EfficientNetV2); (2) hybrid architectures (e.g., CAFormer, SwiftFormer); and (3) visual Transformers, such as SHViT. Evaluation metrics included Top-1 accuracy and F1 score.
[0084] Table 2 Performance metrics on different datasets
[0085]
[0086] Comparative experiments on the ZoonucktonBench50 dataset show that existing strong baseline models (such as ConvNeXt V2 and CAFormer) achieve Top-1 accuracies of 87.49% and 87.24%, respectively. Analysis reveals that these models perform well in general object recognition tasks, but exhibit performance bottlenecks when processing aquatic plankton. This is because the general downsampling paradigm leads to excessive compression of feature map resolution, resulting in the loss of high-frequency microstructural information such as copepod tentacles and setae due to reduced spatial resolution. The PlanktonNet proposed in this invention employs a three-stage architecture to maintain the resolution of deep feature maps at 1 / 16 of the original input, avoiding excessive spatial resolution reduction. Experimental data (Table 2) shows that this invention achieves an accuracy of 91.45% under the same conditions, an improvement of approximately 4% compared to the comparative models. This result demonstrates the significant technical advantage of maintaining high spatial resolution in deep networks for resolving fine biological morphological features.
[0087] To further verify the universality of this invention, tests were conducted on several publicly available datasets, including ZooCAMNet76, ZooScanNet101, and HZWPlankton8. Experimental results confirm that this invention possesses excellent cross-domain generalization ability. On the HZWPlankton8 dataset, which exhibits typical class imbalance, this invention achieved a Top-1 accuracy of 97.94% and an F1 score of 96.92%, both superior to ConvNeXt V2-B and EfficientViT-L2. When processing long-tailed distributed data, this invention did not overfit to the majority class. This demonstrates that even for specific species with scarce samples, this invention can maintain extremely high discrimination accuracy and robustness.
[0088] exist Figure 5 The figure illustrates the loss convergence trajectory during training to assess the model's convergence stability. It compares the performance of PlanktonNet (green curve) with two representative baseline models: ConvNeXt V2-B (orange curve) and MambaOut-Base (blue curve). The solid line depicts the average loss across three independent repeated experiments, while the shaded area reflects the range of experimental results. An embedded plot on the right further shows a magnified view of the later stages of training (Epochs 250-300). The observations demonstrate that PlanktonNet not only achieves the fastest convergence speed and the lowest final loss but also exhibits high training stability: its fluctuations between different experiments are minimal (very narrow shaded band), a stark contrast to the wider variance band of ConvNeXt.
[0089] 2. Model Visualization Analysis
[0090] Obtain the confusion matrices of this invention and the baseline model on four datasets. Rows in the matrix represent predicted classes, and columns represent true labels. The True Positive Rate (TPR) is used as the evaluation metric to visually represent the model's recall rate for each class. The pre-trained PlanktonNet model is then used to predict the final output. The matrix shows a diagonal advantage, with the vast majority of samples clustered along the main diagonal, confirming that PlanktonNet maintains high classification robustness across a wide distribution spanning 50 to 101 classes.
[0091] To more intuitively demonstrate the improved effects of the PlanktonNet three-stage architecture and feature extraction module, EfficientNet was used as the baseline model, and a Class Activation Map (CAM) comparison analysis was performed between it and the model of this invention. By extracting deep network features and overlaying the features of each channel, a CAM heatmap was generated to show the high-frequency attention regions of the model in the image, such as... Figure 6 ,7 As shown.
[0092] The highlighted areas of the heatmap accurately cover high-frequency morphological features that are crucial in taxonomy, such as the antennae of copepods and the tail forks. This ability to capture fine structures is thanks to the innovative topological design of the feature extraction module in PlanktonNet, which achieves a deep fusion of local fine-grained feature extraction and global context modeling, proving that the model has indeed learned biologically significant discriminative features.
[0093] Meanwhile, PlanktonNet demonstrated good signal-to-noise ratio control when processing in-situ images containing a large number of suspended particles. Comparative experiments showed that when the baseline model was often misled by bubbles or non-biological debris in the background, this invention could effectively and strictly constrain the activation range within the biological ontological contour (as shown in the attached figure). Figure 6 (As shown). This anti-interference capability is attributed to the dynamic gating mechanism of the SwiGLU MLP unit: by learning an adaptive suppression threshold on the feature channels, the model uses the gating mechanism to dynamically suppress non-informative background noise in the feature map, thereby maintaining the purity of feature extraction in the low-contrast deep-sea imaging environment.
[0094] 3. Analysis of Algorithm Ablation Experiment Results
[0095] To further verify the contributions of the proposed progressive three-stage architecture and the core components in the feature extraction module to the overall model performance, this embodiment conducted systematic ablation experiments on the ZoonucktonBench50 dataset. By progressively replacing or removing specific network components, the impact of the downsampling module, the selection of the feature extraction module, the receptive field size in the local feature enhancement stage, and the fine-grained design of the feature extraction module on the final classification results (including throughput, computational cost GFLOPs, and Top-1 accuracy) was analyzed, thereby clarifying the substantial role of each innovative design in improving overall performance.
[0096] To quantitatively evaluate the effectiveness of each component of PlanktonNet, all comparison models were trained from scratch with consistent hyperparameters. The experiments first verified the design rationale of the downsampling module. Conventional visual models often use ordinary convolutions with a stride of 2 for downsampling, but in plankton recognition tasks, this rapid spatial compression leads to information loss. Experimental data shows that when the inverse residual-based downsampling module MBConv was replaced with simple stride convolution downsampling, the Top-1 accuracy plummeted from the baseline of 91.16% to 87.95%. This is mainly because key discrimination criteria for plankton, such as tentacles and tail forks, are high-frequency, subtle features. Direct, rapid convolution destroys their pixel-level morphological integrity. Introducing the MBConv structure with SE attention, utilizing a channel recalibration mechanism, effectively preserves these high-frequency gradient information that are easily lost at low resolution.
[0097] Based on the established downsampling module, a step-by-step decoupling analysis was performed on the core computational units of the feature extraction module. Removing the depthwise separable convolution responsible for local feature extraction caused the accuracy to drop to 89.47%, confirming that introducing spatial inductive bias before global interaction is a prerequisite for effectively capturing plankton texture. Further analysis showed that while removing the feedforward network (FFN) alone reduced the computational cost to 8.93 GFLOPs, the accuracy also dropped to 90.88%, attributed to the model losing its nonlinear mapping capability in the channel dimension. More importantly, when both FFN and DWConv were removed simultaneously (i.e., only the global mixer was retained), the accuracy dropped to a minimum of 89.36%. This result indicates that a simple global attention mechanism lacks the inductive bias required to capture local details and cannot independently complete high-precision, fine-grained recognition tasks. In the comparison of channel mixers, the accuracy using the standard MLP was only 88.91%, while the SwiGLU MLP used in PlanktonNet improved it to 91.16%. This 2.25% gain is attributed to the data-driven gate branch in the SwiGLU MLP acting as a dynamic filter, effectively suppressing the response of background noise such as marine snow and bubbles during the feature selection stage using the Hadamard product, thereby significantly enhancing the feature expression of translucent biological subjects under low-contrast conditions. Furthermore, in the module comparison within the global mixer, Spatial Reducing Attention (SRA) demonstrated strong feature extraction capabilities, significantly outperforming pure convolution (90.71%) and pure pooling (90.15%) with an accuracy of 91.16%. This result confirms the effectiveness of global modeling in handling morphological variations in plankton. However, the LeSRA model achieves a further breakthrough on this benchmark, pushing the accuracy to 91.33%.
[0098] Finally, the impact of hyperparameter configuration on model numerical stability was explored at the micro-design level. Experimental data showed that as the kernel size increased from 3×3 to 7×7, the model accuracy steadily improved from 89.30% to 91.16%, indicating that the expanded local receptive field could better integrate the complete organ structure of plankton. However, when further increased to 9×9, the accuracy dropped back to 90.54%. This non-monotonic change revealed the limitations of receptive field expansion: in the case of relatively simple backgrounds in plankton images, excessively large kernels mixed in redundant pure background noise, thereby reducing the effective signal-to-noise ratio of the feature maps. Meanwhile, regarding the selection of normalization layers, given that high-resolution feature maps forced a batch size limit of 26, GroupNorm, with its independence from batch dimensions, showed better statistical stability than Batch Norm (90.99%) and Layer Norm (90.54%). Considering computational efficiency, the channel expansion factor was... Setting it to 8 / 3 is more efficient than high-consumption systems. The configuration (14.26 GFLOPs) achieves lossless accuracy preservation while consuming only 11.59 GFLOPs, establishing the optimal trade-off between parameter efficiency and feature representation capability.
Claims
1. A fine-grained classification method for planktonic organisms based on a progressive downsampling three-stage architecture, characterized in that: The method includes: Images of the plankton to be tested are acquired and input into a plankton classification model. The plankton classification model includes a progressive feature mapping module, a feature extraction module, and a classification head connected in sequence. The progressive feature mapping module is used to downsample the plankton images. The feature extraction module is used to process the downsampled feature map and input the processing result into the classification head to obtain the types of plankton in the plankton images. The feature extraction module includes multi-layer feature extraction sub-blocks and two-layer downsampling modules. The downsampling module is used to perform downsampling operations on the output feature maps of some feature extraction sub-blocks. In the downsampling module, the input feature map is processed by a series of pointwise convolutional layers and depthwise separable convolutions to obtain an intermediate feature map. The intermediate feature map is processed by a compression-excitation attention module, and the processing result is fused with the intermediate feature map before performing a convolution operation to obtain the output feature map of the downsampling module. The feature extraction sub-block includes a cascaded local feature enhancement module and a global feature extraction module; the local feature enhancement module is used to enhance the local spatial micro-details of the input feature map; the global feature extraction module is used to take into account global long-range dependency modeling. In the local feature enhancement module, local features are extracted from the input feature map through a series of normalization layers and depthwise separable convolutional layers, and the processing result is fused with the input feature map to obtain a first fused feature map; the first fused feature map is processed using a series of normalization layers and a feedforward network, and the processing result is fused with the first fused feature map to obtain the local enhanced feature map output by the local feature enhancement module; In the global feature extraction module, the input feature map is processed by a series of normalization layers and a global mixer, and the processing result is fused with the input feature map to obtain the global feature map. The global feature map is then processed by a series of normalization layers and a gated feature reconstruction module, and the processing result is fused with the global feature map to obtain the output feature map of the global feature extraction module.
2. The planktonic fine-grained classification method based on a progressive downsampling three-stage architecture according to claim 1, characterized in that: In the global mixer, the input feature map is processed by a series of deep convolutional layers and normalization layers. The processed result is then projected into a value vector and a key vector. An attention score is calculated by comparing the result with the query vector projected from the input feature map, and the attention output is obtained. The attention output is processed using deep convolutional layers, and the processed result is fused with the attention output before being projected to obtain the output feature map of the global mixer.
3. The planktonic fine-grained classification method based on a progressive downsampling three-stage architecture according to claim 1, characterized in that: The gated feature reconstruction module includes a value branch and a gate branch. The value branch processes the input feature map through a concatenated convolutional layer and an activation function to obtain the output result of the value branch. The gate branch processes the input feature map through a convolutional layer to obtain the output result of the gate branch. The output results of the value branch and the gate branch are fused and further processed through a convolutional layer to obtain the output feature map of the gated feature reconstruction module.
4. The planktonic fine-grained classification method based on a progressive downsampling three-stage architecture according to claim 1, characterized in that: The progressive feature mapping module includes multiple convolutional layers, and the stride of a single convolutional layer is no greater than 2.
5. The planktonic fine-grained classification method based on a progressive downsampling three-stage architecture according to claim 1, characterized in that: The plankton classification model is trained using a dataset containing images of different plankton. During training, the cross-entropy loss function is used to calculate the error between the probability distribution predicted by the model and the real species labels, and to guide the parameter updates of the plankton classification model.
6. The planktonic fine-grained classification method based on a progressive downsampling three-stage architecture according to claim 5, characterized in that: Before model training, the planktonic images in the dataset are preprocessed. The preprocessing method is as follows: after unifying the size of all planktonic images, the planktonic images are processed by random rotation, translation, cropping and contrast adjustment.
7. A fine-grained planktonic classification system based on a progressive downsampling three-stage architecture, characterized in that: This system is used to execute a plankton fine-grained classification method based on a progressive downsampling three-stage architecture as described in claim 1. The plankton fine-grained classification system includes an image acquisition module, a plankton classification module, and a classification output module. The image acquisition module is used to acquire images of the plankton to be tested. The plankton classification module is used to classify the plankton types in the tested plankton images. The classification output module is used to display the classification results of the plankton in the tested plankton images in real time.
Citation Information
Patent Citations
A marine survey image enhancement system
CA3036667A1
Image classification model construction method, image classification method, image classification device, image classification equipment and storage medium
CN118675000A