Phytoplankton microscopic image classification model construction method

The ResNeXtViT model enhances floating plant classification by integrating ResNeXt and Transformer branches with adaptive convolution and attention mechanisms, addressing image quality issues and subtle species differences to improve classification accuracy.

CN120318818APending Publication Date: 2025-07-15Hefei Comprehensive Science Center Environmental Research Institute
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510460104.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing phytoplankton classification model is insufficient in the case of unstable microscopic image quality, confusion of phytoplankton categories and large morphological changes, making it difficult to effectively distinguish subtle differences.

Method used

The ResNeXtViT model is adopted, combined with the ResNeXt branch and the Transformer branch, and the feature extraction capability is enhanced through the wavelet convolution attention module and the adaptive convolution module, the feature fusion module is used to integrate local details and global structures, and the model performance is optimized through cross entropy, central features and key point feature loss functions.

Benefits of technology

It significantly improves the adaptability and classification accuracy of the model under complex imaging conditions, can extract target information more fine-grained, and improves the ability to identify phytoplankton morphology and subtle differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318818A_ABST
    Figure CN120318818A_ABST
Patent Text Reader

Abstract

The invention discloses a phytoplankton microscopic image classification model construction method, and belongs to the field of intelligent detection. For the problems of unstable microscopic image quality and small difference between phytoplankton classes, a double-branch ResNeXtViT model is provided: a ResNeXt branch enhances the texture feature extraction capability through a wavelet convolution attention module, and improves the adaptability to irregular shapes in combination with an adaptive convolution module; a Transform branch is used for capturing global context information; the feature fusion module integrates local details and a global structure. A combined data enhancement strategy including noise addition and color migration is adopted, and a joint loss function is designed to optimize intra-class compactness and local discrimination capability. Experiments show that the accuracy of the method on a Chaohu lake phytoplankton data set reaches 94.11%, which is superior to that of ResNet, ViT and other models, and the method can be deployed to an intelligent analyzer to realize rapid water quality monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent detection of phytoplankton, and particularly relates to a method for constructing a classification model of phytoplankton microscopic images. Background Art

[0002] Phytoplankton are the primary producers in the aquatic ecosystem, fixing carbon dioxide through photosynthesis and providing a basic energy source for aquatic organisms. At the same time, the diversity monitoring of phytoplankton is an important part of the evaluation of aquatic ecological quality. The accurate identification and classification of phytoplankton not only help to deeply understand the ecological functions and community dynamics, but also provide a scientific basis for water quality monitoring, ecological health assessment and water resource management.

[0003] The identification methods of phytoplankton mainly include traditional microscope analysis and modern intelligent technologies. Traditional methods rely on morphological characteristics such as cell morphology, size and pigment type for classification, which require professional operation, are time-consuming and easily affected by subjectivity. Modern intelligent technologies use deep learning algorithms to automatically extract the characteristics of phytoplankton and achieve rapid and accurate classification. However, the current models are still insufficient in the classification of phytoplankton. On the one hand, the clarity and resolution of microscopic images are easily affected by light sources, focusing, electronic noise and optical noise. On the other hand, phytoplankton belonging to the same genus are similar in texture details and cell structure, especially multicellular phytoplankton only have minor differences; at the same time, the morphological differences of the same species of phytoplankton are relatively large due to seasonal and water environmental influences. The classification of phytoplankton is more complex than traditional fine-grained classification tasks (such as cat and dog classification), and the category confusion of phytoplankton based on deep learning is more serious. Further improving the accuracy of phytoplankton classification remains a challenge. Summary of the Invention

[0004] In view of the problem of category confusion of phytoplankton in the background art, the present invention proposes a method for constructing a classification model of phytoplankton microscopic images.

[0005] The technical solution of the present invention is as follows:

[0006] A method for constructing a classification model of phytoplankton microscopic images includes the following steps:

[0007] Step 1: Collect water samples and fix phytoplankton with Lugol's reagent, and obtain original microscopic images by using microscopic imaging technology;

[0008] Step 2: Clean the data and construct a training set and a test set, and perform combined data augmentation on the training set;

[0009] Step 3: Construct a double-branch classification model ResNeXtViT, including:

[0010] ResNeXt branch: the first three residual modules are followed by a wavelet convolution attention module, and the fourth residual module is followed by an adaptive convolution module;

[0011] Transformer branch: extracts global features through image segmentation and self-attention mechanism;

[0012] Feature fusion module: align and fuse dual-branch features;

[0013] Joint loss function: includes cross entropy loss, center feature loss and key point feature loss;

[0014] Step 4: Train the model and select the optimal model;

[0015] Step 5: Deploy the model to the phytoplankton smart analyzer.

[0016] Further optionally, the use of Lugo reagent in step 1 can prevent sample degradation and achieve long-term preservation, while fixing the cell structure and staining the cells, making the microscopic imaging clearer. Due to differences in human operation, different amounts of Lugo reagent used will lead to large differences in cell staining depth and color, causing the same type of phytoplankton to present different colors.

[0017] Further optionally, in the step 2 of constructing the data set, the original microscopic images are first cleaned to remove incomplete and blurred images, modify erroneous labels, and perform combined data enhancement on the training set to increase the diversity of training samples. The main enhancement methods include:

[0018] Randomly add Gaussian and salt and pepper noise; randomly perform horizontal, vertical and biaxial flipping within the range of ±10 degrees; randomly block with a black rectangle of 5% of the image area; randomly scale within the range of ±10% of the original image; randomly select 2 methods from the above methods for combination;

[0019] Use HSV enhancement to adjust the image hue, saturation, and brightness within the range of ±0.2, ±0.4, and ±0.3 respectively;

[0020] Color migration is used to migrate the colors of certain types of phytoplankton using the Gaussian distribution method.

[0021] Further optionally, the wavelet convolution attention module in step 3 is intended to expand the receptive field and improve the expression ability of edge and texture information. It uses wavelet decomposition combined with the attention mechanism to separate features of different scales in the frequency domain, enhance the receptive field, and improve the modeling ability of fine-grained information. The core process is as follows:

[0022] S1. The input feature map X is decomposed into low-frequency components and high-frequency components by applying the Haar wavelet transform. The low-frequency components retain global information, and the high-frequency components capture local details and edge textures.

[0023] S2. The low-frequency components are directly restored to the spatial size of the input feature map X through the inverse wavelet transform to reconstruct the overall structural information and provide the global background and main contours in the feature map.

[0024] S3. The high-frequency information undergoes feature extraction through 1×1 and 3×3 double-branch convolutions to enhance local details with multi-scale features, and the channel attention mechanism is combined to enhance the expression of key high-frequency information and suppress unimportant high-frequency components.

[0025] S4. The input feature map X passes through a 3×3 standard convolution to obtain a feature map to further extract local features and enhance the non-linear representation ability. Then, the features after the inverse transformation of the low-frequency components, and the features after the convolution of the high-frequency components are weighted and fused to obtain the final output feature map :

[0026]

[0027] Among them, , and are all learnable parameters, and the weights of the feature map , the features after the inverse wavelet transform of the low-frequency components, and the features after the convolution operation of the high-frequency components are calculated respectively.

[0028] Further optionally, the adaptive convolution module in step three depends on the AKConv convolution and the channel attention mechanism, combines adaptive feature extraction and channel-level dynamic weight assignment to enhance the model's sensitivity to the target shape, scale, and key features, thereby improving the classification performance. The core process is as follows:

[0029] S1. The input feature map X generates an offset through a 3×3 convolution to adjust the sampling position of the input feature map, enabling the convolution kernel to adaptively capture important information.

[0030] S2. Calculate the standard grid sampling points, and combine the offset to obtain the final dynamic sampling coordinates. Through bilinear interpolation, the features are resampled from the input feature map by weighted summation of the four nearest grid points.

[0031] S3. Convolve these adjusted features with N×1, and add BN and SiLU activation to enhance the feature extraction ability.

[0032] S4. In combination with the channel attention mechanism, the output features of AKConv are used to generate channel weights through global average pooling and 1×1 convolution, enabling the network to focus on key feature channels, thereby further enhancing the feature expression ability.

[0033] Further optionally, for the Transformer branch in step three, the ViT model is used to extract global features. First, the input image is divided into image patches of a fixed size. After each patch is flattened, it is converted into a vector of a fixed dimension through a linear mapping layer and then added with positional encoding to retain spatial position information. Then, this sequence passes through the Transformer encoder, and each layer contains multi-head self-attention and a feed-forward network, which are used to capture global dependencies and extract features. Finally, ViT obtains the global feature representation of the image through pooling.

[0034] Further optionally, the branch feature fusion module in step three adjusts the tokens extracted by ViT at the intermediate stage to align them with the ResNeXt feature map in the spatial dimension, ensuring the consistency of feature fusion. Then, the global features of ViT and the local features of ResNeXt are fused through channel concatenation, and the self-attention mechanism is used to optimize feature interaction. Finally, the fused features are processed using 1×1 convolution, batch normalization, and activation functions to restore to the same dimension as the ResNeXt features. At the final stage, the features after the final pooling of ResNeXt are concatenated with the features corresponding to the classification token ([CLS] token) extracted by ViT and fed into the classifier for class prediction.

[0035] Further optionally, the loss function in step three mainly consists of three parts, and its formula is as follows:

[0036]

[0037] Among them, represents the classification loss function of ResNeXtViT, , and are the cross-entropy loss, the center feature loss, and the key point feature loss respectively, , and are the weights of each loss respectively;

[0038] The cross-entropy loss function measures the difference between the predicted distribution and the true distribution, and evaluates the accuracy of the model by calculating the logarithmic loss of the predicted probability corresponding to the true class. In PyTorch, CrossEntropyLoss directly accepts logits as input and automatically applies softmax internally and calculates the loss.

[0039]

[0040] Among them, represents the number of training samples, represents the th training sample, represents the class of the th sample, represents the predicted class probability distribution of the th sample;

[0041] The central feature loss enhances the feature compactness of samples of the same class, especially the within-class compactness of the minority classes, by minimizing the distance between the features of each sample and the class center, thereby improving the classification effect on the minority classes. At the same time, L2 normalization is introduced in the central feature loss to reduce the influence of scale differences, making the model training more stable; its formula is as follows:

[0042]

[0043] Among them, represents the number of training samples, represents the th training sample, represents the class label of the th sample, represents the th sample's feature vector, which is after L2 normalization, represents the central feature vector corresponding to class ;

[0044] The key-point feature loss promotes the accurate classification of the model on local features by minimizing the gap between the predicted class distribution of each key point and the true class. Compared with the traditional whole-image classification method, it can extract the local information of the target in a finer granularity, enhancing the discriminative ability of the model, especially in the case where the target has subtle differences; its formula is as follows:

[0045]

[0046] Among them, represents the number of training samples, represents the number of classes, represents the number of key points set in each sample, represents the th training sample, represents the th class, represents the th key point, represents each key point The score for the category represents the true label of the th image in the category .

[0047] An electronic device includes a processor and a memory. The memory stores program instructions that, when executed by the processor, implement the above method.

[0048] A computer-readable storage medium stores a computer program that, when executed, implements the above method.

[0049] Advantageous effects:

[0050] The present invention discloses a method for constructing a classification model for phytoplankton microscopic images. Aiming at the problem of unstable image quality caused by light changes, focus deviation, and noise interference during microscopic imaging, a set of data enhancement strategies are formulated to improve the adaptability of the model to different imaging conditions. The ResNeXtViT model proposed in the present invention adopts a dual-branch structure, including a ResNeXt branch for capturing detailed information and a Transformer branch for capturing global information. In the ResNeXt branch, the wavelet convolutional attention module uses wavelet decomposition and attention mechanism to expand the receptive field and enhance the expression ability of edge and texture information, improving the modeling ability for fine-grained information. The adaptive convolutional module relies on AKConv convolution and channel attention mechanism, and through adaptive feature extraction and channel-level dynamic weight allocation, enhances the sensitivity of the model to the shape, scale, and key features of phytoplankton. The feature fusion module uses the self-attention mechanism to optimize feature interaction in the intermediate stage between the ResNeXt branch and the Transformer branch. The loss function improves the classification effect for minority classes, can extract the local information of the target more finely, and enhances the discriminative ability of the model. The present invention discloses a classification model for phytoplankton microscopic images, which effectively alleviates the problems of unstable microscopic image quality, difficult to distinguish subtle differences between phytoplankton categories, and large morphological changes of different species, and significantly improves the adaptability and classification accuracy of the model under complex imaging conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic flowchart of a method for constructing a classification model for phytoplankton microscopic images.

[0052] Figure 2 is a structural diagram of the ResNeXtViT model.

[0053] Figure 3 is a structural diagram of the wavelet convolutional attention module.

[0054] Figure 4It is a structural diagram of an adaptive convolution module.

[0055] Figure 5 It is a structural diagram of a feature fusion module.

[0056] Figure 6 It is a visualization result diagram of phytoplankton image features. Detailed implementation manners

[0057] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. However, the following embodiments are only for explaining the present invention, and the protection scope of the present invention should include all the contents of the claims. Moreover, through the description of the following embodiments, those skilled in the art can fully implement all the contents of the claims of the present invention.

[0058] Embodiment

[0059] The following details illustrate the implementation manner of the present invention, and the overall process is as Figure 1 shown.

[0060] Step 1, sample collection, fixation and data preprocessing. The phytoplankton studied in the present invention is mainly in the waters of Chaohu Lake. Water samples are collected at representative locations such as Gushan, Guishan, Huanglu and the center of the eastern half of the lake, and Lugol's reagent is used to fix the cell structure and stain the cells. Then, a high-resolution original microscopic image is obtained using a 20x Nikon microscope (numerical aperture 0.45).

[0061] Step 2, the collected original microscopic images are first cleaned, incomplete or blurred images are removed, and incorrect labels are corrected to construct a high-quality initial data set. And the data is randomly divided into a training set and a test set according to the ratio of 8:2. Subsequently, data augmentation is performed on the training set to expand data diversity and simulate the possible illumination changes, focusing deviations and noise interferences in the microscopic imaging process, so as to improve the adaptability and generalization ability of the model to different imaging conditions. The data augmentation includes: noise addition, image flipping, random occlusion, image scaling, HSV enhancement and color transfer. Among them, noise addition includes Gaussian and salt-and-pepper noise; image flipping performs horizontal, vertical and biaxial flipping in the range of ±10 degrees; random occlusion is performed with a black rectangular frame covering 5% of the image area; random scaling randomly scales the image within the range of ±10%; HSV enhancement adjusts the hue, saturation and brightness within the ranges of ±0.2, ±0.4 and ±0.3 respectively; color transfer performs color transformation based on the Gaussian distribution for certain specific categories of phytoplankton.

[0062] Step 3, build a phytoplankton microscopic image classification model ResNeXtViT through the PyTorch architecture and Python language, and its network structure is as Figure 2As shown. The model adopts a dual-branch structure, including a ResNeXt branch for extracting local features and a Transformer branch for capturing global features. Among them, the ResNeXt branch enhances the ability to extract detailed features through a wavelet convolution attention module and an adaptive convolution module, improving the sensitivity to the morphological features of phytoplankton. At different stages, the two branches integrate information through a feature fusion module, effectively integrating local details and global structures, enhancing the ability to identify minute morphological differences in phytoplankton, and improving classification accuracy.

[0063] The backbone network of the ResNeXt branch uses the resnext50 structure. The initial stage includes a 7×7 convolution and a max-pooling layer to extract basic features and reduce the computational amount. Subsequently, the network successively passes through Layer1, Layer2, Layer3, and Layer4 for deep feature extraction. Among them, a wavelet convolution attention module is introduced after Layer1, Layer2, and Layer3 to enhance the expression ability of edge and texture features, while an adaptive convolution module is added after Layer4 to improve the perception ability of the morphology and scale changes of phytoplankton. Finally, the features are globally average-pooled to map the high-dimensional features to a fixed-length vector representation, generating a global feature vector with a shape of [batch, 2048], providing a compact and discriminative expression for subsequent classification and feature fusion.

[0064] The wavelet convolution attention module aims to increase the receptive field and enhance the expression ability of edge and texture information. The specific structure is as Figure 3 shown. First, the input feature map X is decomposed into a low-frequency component and a high-frequency component by applying the Haar wavelet transform. The low-frequency component is directly restored to the spatial size of the input feature map X through the inverse wavelet transform. The high-frequency information undergoes feature extraction through a 1×1 and 3×3 dual-branch convolution to enhance the local details with multi-scale features, and combines a channel attention mechanism to enhance the expression of key high-frequency information and suppress unimportant high-frequency components. The input feature map X passes through a 3×3 standard convolution to obtain the feature map , and then , the feature after the inverse transformation of the low-frequency component, and the feature after the convolution of the high-frequency component are weighted and fused to obtain the final output feature map , and the specific weighting method is as follows:

[0065]

[0066] Among them, , and are all learnable parameters, calculating the weights of the feature map , the feature after the inverse wavelet transform of the low-frequency component, and the feature after the convolution operation of the high-frequency component respectively;

[0067] The adaptive convolution module enhances the model's sensitivity to object shape, scale, and key features. The specific structure is as Figure 4 shown. First, offsets are generated through 3×3 convolution, dynamic sampling coordinates are calculated by combining with standard grid sampling points, and features are resampled from the input feature map using bilinear interpolation. Then, the adjusted features are convolved with N×1, combined with BN and SiLU activation to enhance the feature extraction ability. Finally, channel weights are generated through global average pooling and 1×1 convolution.

[0068] The backbone network of the Transformer branch uses the vit-base structure to perform global feature extraction with the standard Transformer structure. The input image is first divided into fixed-size image patches through a linear projection layer. Each image module is 16×16 in size and flattened into a 768-dimensional vector, generating a total of 196 patches. Subsequently, it is processed by 12 layers of Transformer encoders, each layer containing multi-head self-attention and a feed-forward network for modeling global information and long-range dependencies. Finally, a global feature vector with a feature dimension of [batch, 768] corresponding to the classification token ([CLS] token) is extracted, providing rich global context information for subsequent feature fusion and classification.

[0069] As Figure 5 shown, the feature fusion module integrates the features of the two branches at different stages, mainly fusing Layer 3 and Layer 4 of resnext50 with the 8th and 12th layer encoders of vit-base. Layer 3 is responsible for intermediate features, focusing on local structures and fine-grained features. The corresponding 8th layer encoder has both local and global receptive fields, and complementing it can enhance the ability to extract subtle features. Layer 4 is responsible for high-level features, focusing on global semantic information. Combining with the 12th layer encoder with the most global features improves the overall morphological understanding and class discrimination ability, thus taking into account both local details and global information and improving classification accuracy.

[0070] The final features of the two branches are concatenated in the channel dimension to [batch, 2816], and then sent to the classifier to calculate the class probabilities through softmax. The loss function of the model ResNeXtViT consists of three parts, and its formula is as follows:

[0071]

[0072] Among them, represents the classification loss function of ResNeXtViT, , and are the cross-entropy loss, center feature loss, and key point feature loss respectively, , and are the weights of each loss respectively.

[0073] Step 4: Train on the Ubuntu 22.04 platform. The device is equipped with an NVIDIA GeForce RTX 3090 (24GB video memory) and an Intel(R) Xeon(R) Silver 4210 CPU (2.20GHz). The specific training parameters are shown in Table 1.

[0074] Table 1 Detailed Information of Training Parameters

[0075] Warm-up training can prevent gradient oscillation caused by too large an initial learning rate, improve training stability, and is especially suitable for large-scale training. Mixed-precision training (FP16) reduces video memory occupancy and accelerates calculations, while combining with the AMP mechanism to maintain numerical stability. The cosine annealing learning rate scheduler simulates the learning process, converges quickly in the initial stage, and adjusts slowly in the later stage to avoid falling into local optima, ultimately improving the model accuracy. The combination of the three makes the training more efficient and stable, and improves the final classification performance. The loss function used in model training consists of three parts, and its formula is as follows:

[0076]

[0077] Among them, represents the classification loss function of ResNeXtViT, , and are the cross-entropy loss, central feature loss, and key point feature loss respectively, , and are the weights of each loss respectively;

[0078] The cross-entropy loss function measures the difference between the predicted distribution and the true distribution, and evaluates the accuracy of the model by calculating the logarithmic loss of the predicted probability corresponding to the true category. In PyTorch, CrossEntropyLoss directly accepts logits as input and automatically applies softmax internally and calculates the loss;

[0079]

[0080] Among them, represents the number of training samples, represents the th training sample, represents the th sample's category, represents the The predicted class probability distribution of a single sample;

[0081] The central feature loss enhances the feature compactness of samples of the same class (especially the within-class compactness of the minority class) by minimizing the distance between the features of each sample and the class center, thereby improving the classification effect on the minority class. At the same time, L2 normalization is introduced in the central feature loss to reduce the influence of scale differences, making the model training more stable; its formula is as follows:

[0082]

[0083] Among them, represents the number of training samples, represents the th training sample, represents the class label of the th sample, represents the feature vector of the th sample, which is after L2 normalization, represents the central feature vector corresponding to class

[0084] Step five, to verify the effectiveness of the proposed method, it is compared with classical models such as ResNet, ConvNeXt, ViT, and Swin-Transformer, and accuracy (ACC), macro-average (MAC), and F1-score are used as evaluation metrics. Accuracy measures the overall classification accuracy of the model, macro-average ensures the same contribution of each class, which is suitable for the case of class imbalance, while F1-score takes into account both Precision and Recall, providing a more comprehensive evaluation of classification performance. The self-made Chaohu phytoplankton dataset is used for training, and the specific results are shown in Table 2.

[0085] Table 2 Test results of the Chaohu phytoplankton dataset

[0086] As can be seen from Table 2, the model proposed in the present invention is superior to other models in terms of accuracy, macro-average, and F1-score. Compared with classical models such as ResNet, ConvNeXt, ViT, and Swin-Transformer, the proposed method effectively enhances the sensitivity to the shape, scale, and key features of phytoplankton through the wavelet convolutional attention module, adaptive convolutional module, and dual-branch feature fusion, thereby improving the overall classification performance.

[0087] Step 6. To evaluate the contributions of various innovative strategies in the microscopic image classification task of phytoplankton, this article gradually added key modules such as the wavelet convolutional attention module, the adaptive convolutional module, and the dual-branch feature fusion through ablation experiments, and analyzed their effects on the model performance. The detailed results of the ablation experiments are shown in Table 3.

[0088] Table 3 Results of Ablation Experiments

[0089] As can be seen from Table 3, the results of the ablation experiments show that each innovative strategy effectively improves the classification performance of phytoplankton microscopic images. The wavelet convolutional attention module enhances the extraction of edge and texture features, the adaptive convolutional module improves the adaptability to irregular shapes, the dual-branch structure combines local details and global information to enhance the feature expression ability, and the joint loss function optimizes the classification accuracy and robustness. The results of the ablation experiments verify the effectiveness of the proposed method.

[0090] Step 7. To further verify the performance of the model proposed in the present invention, we conducted a visualization analysis of the feature maps for some phytoplankton. Feature maps are the key information extracted by the neural network at different levels and are used to characterize the significant features of the input image. To observe whether the model pays attention to the target area, we used Grad-CAM to visualize the features of the final layer of ResNeXtViT. The visualization results of the feature maps are as Figure 6 shown. The attention areas of the model mainly focus on the targets. Especially for multi-cellular phytoplankton with a large aspect ratio, the model successfully captures its fine structural features. At the same time, for phytoplankton with variable morphologies, the model can adapt to their different morphologies and ignore background information during the classification process, further verifying the effectiveness of the proposed method in the phytoplankton classification task.

[0091] Step 8. Export the best model in onnx format, deploy it on a phytoplankton intelligent analyzer, and load the model through onnxruntime for inference to achieve fast and accurate classification of phytoplankton.

[0092] From the results of the implementation cases, it can be seen that a method for constructing a microscopic image classification model of phytoplankton proposed in the present invention effectively improves the classification accuracy of phytoplankton microscopic images. This method improves the generalization ability of the model through data augmentation strategies, reduces the impact of focus deviation and noise interference in the microscopic imaging process on the classification performance. At the same time, it enhances the extraction ability of details such as edges and textures, and improves the adaptability of the model to phytoplankton with different morphologies. Through the feature fusion of the dual-branch, local detail information and global structural information are fully combined, enabling the model to more accurately capture the features of phytoplankton with variable morphologies.

[0093] The foregoing are only specific embodiments of the present application to enable persons skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for constructing a microscopic image classification model of phytoplankton, characterized in that The method comprises the following steps: Step 1: Collect water samples and fix phytoplankton using Lugo reagent, and obtain original microscopic images using microscopic imaging technology; Step 2: Clean the data and construct the training set and test set, and perform combined data enhancement on the training set; Step 3: Build a two-branch classification model ResNeXtViT, including: ResNeXt branch: the first three residual modules are followed by a wavelet convolution attention module, and the fourth residual module is followed by an adaptive convolution module; Transformer branch: extracts global features through image segmentation and self-attention mechanism; Feature fusion module: align and fuse dual-branch features; Joint loss function: includes cross entropy loss, center feature loss and key point feature loss; Step 4: Train the model and select the optimal model; Step 5: Deploy the model to the phytoplankton smart analyzer.

2. The method for constructing a microscopic image classification model of phytoplankton according to claim 1, characterized in that, The use of Lugo reagent in step 1 can prevent sample degradation and achieve long-term preservation, while fixing the cell structure and staining the cells to make microscopic imaging clearer. Due to differences in human operation, different amounts of Lugo reagent used will lead to large differences in cell staining depth and color, causing the same type of phytoplankton to present different colors.

3. The method for constructing a phytoplankton microscopic image classification model according to claim 1, characterized in that, In the process of constructing the data set in step 2, the original microscopic images are first cleaned to remove incomplete and blurred images, modify wrong labels, and perform combined data enhancement on the training set to increase the diversity of training samples. The main enhancement methods include: randomly adding Gaussian and salt and pepper noise; randomly flipping horizontally, vertically and biaxially within the range of ±10 degrees; randomly blocking with a black rectangular frame of 5% of the image area; randomly scaling within the range of ±10% of the original image; randomly selecting 2 methods of the above methods for combination; using HSV enhancement to adjust the image hue (Hue), saturation (Saturation), and brightness (Value), respectively, within the range of ±0.2, ±0.4 and ±0.3; using color migration, the color of certain types of phytoplankton is migrated using the Gaussian distribution method.

4. The method for constructing a microscopic image classification model of phytoplankton according to claim 1, wherein The wavelet convolution attention module in step 3 is designed to expand the receptive field and improve the expression ability of edge and texture information. It uses wavelet decomposition combined with attention mechanism to separate features of different scales in the frequency domain, enhance the receptive field, and improve the modeling ability of fine-grained information. The process is as follows: S1, the input feature map X is decomposed into low-frequency components and high-frequency components by applying Haar wavelet transform, where the low-frequency components retain global information and the high-frequency components capture local details and edge textures; S2, the low-frequency components are directly restored to the spatial size of the input feature map X through inverse wavelet transform to reconstruct the overall structural information and provide the global background and main contours in the feature map; S3, high-frequency information is extracted through 1×1 and 3×3 dual-branch convolution to enhance local details with multi-scale features, and combined with the channel attention mechanism to enhance the expression of key high-frequency information and suppress unimportant high-frequency components; S4. The input feature map X undergoes a 3×3 standard convolution to obtain a feature map , to further extract local features and enhance the non-linear representation ability, and then , the features after the inverse transformation of the low-frequency components, and the features after the convolution of the high-frequency components are weighted and fused to obtain the final output feature map : Among them, , and are all learnable parameters, and the weights of the feature map , the feature after the inverse wavelet transform of the low-frequency component, and the feature after the convolution operation of the high-frequency component are calculated respectively.

5. The method for constructing a microscopic image classification model of phytoplankton according to claim 1, characterized in that, In step 3, the adaptive convolution module relies on AKConv convolution and channel attention mechanism, combines adaptive feature extraction and channel-level dynamic weight allocation to enhance the model's sensitivity to target shape, scale, and key features, thereby improving the classification performance. The process is as follows: S1. The input feature map X generates an offset through a 3×3 convolution to adjust the sampling position of the input feature map, enabling the convolution kernel to adaptively capture important information; S2. Calculate the standard grid sampling points, and combine the offset to obtain the final dynamic sampling coordinates. Through bilinear interpolation, re-sample the features from the input feature map by weighted summation of the four nearest grid points; S3. Convolve these adjusted features with N×1, and add BN and SiLU activation to enhance the feature extraction ability; S4. Combine the channel attention mechanism, generate channel weights by global average pooling and 1×1 convolution of the output features of AKConv, enabling the network to focus on key feature channels, thereby further improving the feature expression ability.

6. The method for constructing a microscopic image classification model of phytoplankton according to claim 1, characterized in that, In the Transformer branch of step 3, the ViT model is used to extract global features. First, the input image is divided into image patches of a fixed size. After each patch is flattened, it is converted into a vector of a fixed dimension through a linear mapping layer, and then added with position encoding to retain spatial position information. Then, this sequence passes through the Transformer encoder, each layer containing multi-head self-attention and a feed-forward network, used to capture global dependencies and extract features. Finally, ViT obtains the global feature representation of the image through pooling.

7. A method for constructing a microscopic image classification model of phytoplankton as claimed in claim 1, wherein The branch feature fusion module in step 3 adjusts the tokens extracted by ViT in the intermediate stage to align them with the ResNeXt feature map in the spatial dimension, ensuring the consistency of feature fusion. Then, the global features of ViT and the local features of ResNeXt are fused through channel concatenation, and the self-attention mechanism is used to optimize feature interaction. Finally, the fused features are processed using 1×1 convolution, batch normalization, and activation functions to restore to the same dimension as the ResNeXt features; In the final stage, the features after ResNeXt pooling are concatenated with the features corresponding to the classification token ([CLS] token) extracted by ViT and sent to the classifier for class prediction.

8. The method for constructing a microscopic image classification model of phytoplankton according to claim 1, characterized in that, The loss function in step 3 mainly consists of three parts, and its formula is as follows: Among them, represents the classification loss function of ResNeXtViT, , and are the cross-entropy loss, the central feature loss, and the key-point feature loss respectively, , and are the weights of each loss respectively; The cross-entropy loss function measures the difference between the predicted distribution and the true distribution, and evaluates the accuracy of the model by calculating the logarithmic loss of the predicted probability corresponding to the true class. In PyTorch, CrossEntropyLoss directly accepts logits as input and will automatically apply softmax and calculate the loss internally. Its formula is as follows: Among them, represents the number of training samples, represents the th training sample, represents the category of the th sample, represents the predicted category probability distribution of the th sample; The central feature loss enhances the feature compactness of samples of the same class, especially the within-class compactness of the minority classes, by minimizing the distance between the features of each sample and the class center, thereby improving the classification effect on the minority classes. At the same time, L2 normalization is introduced in the central feature loss to reduce the influence of scale differences, making the model training more stable. The formula is as follows: Among them, represents the number of training samples, represents the th training sample, represents the class label of the th sample, represents the feature vector of the th sample, which is after L2 normalization, represents the central feature vector corresponding to the class ; The key-point feature loss promotes the accurate classification of the model on local features by minimizing the gap between the predicted class distribution of each key point and the true class. Compared with the traditional whole-image classification method, it can extract the local information of the target in a finer granularity, enhancing the discriminative ability of the model, especially in the case where the targets have subtle differences. The formula is as follows: Among them, represents the number of training samples, represents the number of classes, represents the number of key points set in each sample, represents the th training sample, represents the th class, represents the th key point, represents the score of each key point for class respectively, represents the th image's true label for class respectively.

9. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores program instructions. When the program instructions are executed by the processor, the method according to any one of claims 1-8 is implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored, and when the program is executed, the method according to any one of claims 1-8 is implemented.

Citation Information

Cited By

  • Fine-grained image classification method based on deep wavelet feature fusion

    CN120673186A

  • Fine-grained image classification method based on deep wavelet feature fusion

    CN120673186B

  • Phytoplankton chromatography sequence identification method and phytoplankton chromatography sequence model building method

    CN121838155A

  • Optical imaging method and system for intelligent plankton analysis

    CN122409436A

  • Plankton fine-grained classification method and system based on progressive down-sampling three-stage architecture

    CN122454570A