Marine plankton image classification method
By introducing deep learning algorithms and traditional morphological features into the image classification of marine plankton, an automated image classification model was established, which solved the problem of insufficient classification accuracy and robustness in the existing technology, and achieved efficient and automated image classification.
Patent Information
- Application Number
- CN202510290404.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art relies on manually set geometric morphological features in marine plankton image classification, resulting in the accuracy and robustness of classification in the face of complex and variable plankton species, requiring a lot of manual correction.
Introduce deep learning algorithms and traditional morphological features to establish a classification model based on geometric morphological features, deep learning features and fusion features to realize the automation of image classification.
It greatly improves the efficiency of marine plankton image classification, improves the accuracy and robustness of classification, and reduces the dependence on artificial correction.
Smart Images

Figure CN120164030A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological image classification processing, and particularly to a method for classifying marine plankton images. Background Art
[0002] Marine plankton diversity is the cornerstone for maintaining the stability of the marine ecosystem and provides many indispensable ecosystem service functions. The planktonic food web is a key link in the material and energy cycle of the entire marine ecosystem and plays an important role in the global carbon cycle and climate regulation. With human activities and climate change, the sharp increase in the abundance of specific plankton in coastal waters can lead to the occurrence of ecological disasters such as red tides, bringing huge ecological risks and affecting the development of fisheries and tourism. Therefore, the species and abundance of plankton are important indicators for evaluating the health of the marine ecosystem, and the world is continuously strengthening the monitoring and research of plankton communities to better understand and protect the marine ecological environment.
[0003] FlowCAM is a device that combines microscopic imaging technology and flow cytometry technology, capable of collecting images of particles in liquid samples and extracting multi-dimensional geometric morphological features, thereby realizing image recognition and ecological statistics. It is a commonly used device for monitoring and researching the community structure of marine plankton. However, the image classification algorithm built into the VisualSpreadsheet software of FlowCAM mainly relies on manually set geometric morphological features. When facing complex and variable plankton species, the accuracy and robustness of classification are limited, and a large amount of manual correction is still required to obtain more accurate results.
[0004] In view of this, it is necessary to provide a new technical solution to solve the above problems. Summary of the Invention
[0005] To solve the above technical problems, the present application provides a method for classifying marine plankton images. By introducing a deep learning algorithm and combining it with traditional morphological features, the automation of image classification is achieved, and the classification efficiency is greatly improved compared with traditional methods.
[0006] A method for classifying marine plankton images includes:
[0007] Using a flow particle imaging analyzer to collect plankton images in the seawater of the sea area to be studied;
[0008] Manually specifying the categories of some plankton images in the collected sea area to be studied to establish an original classification dataset of marine plankton images;
[0009] Converting the original classification dataset of marine plankton images into a standard classification dataset of marine plankton images;
[0010] Build and train a classification model based on geometric morphological features, a classification model based on deep learning features, and a classification model based on the fusion of geometric morphological features and deep learning features, and screen out the optimal classification model for marine plankton images;
[0011] Use the selected optimal model to identify the plankton images of the collected samples to be detected, and group, count, and store the identified plankton images by category.
[0012] Preferably, manually assign category designations to the plankton images collected in the research sea area, and establish an original classification dataset for marine plankton images, including:
[0013] Use the FlowCAM imaging analysis system to collect plankton images to be classified in the samples;
[0014] Use Visual Spreadsheet software to manually assign category designations to the collected plankton particle images;
[0015] Save the plankton particle images of the same category in the form of a puzzle as a TIF format, and save the meta-information including position and geometric morphological parameters of each particle image in the puzzle as an FLB format file, so as to establish an original classification dataset for marine plankton images under VisualSpreadsheet software.
[0016] Preferably, the conversion of the original classification dataset of marine plankton images into a standard classification dataset for marine plankton images includes:
[0017] Traverse the FLB files of various types of plankton, and parse them into CSV spreadsheets respectively;
[0018] Merge all the obtained CSV spreadsheets, and delete duplicate rows based on the file name field of the image, only keeping the first record of each image, to obtain a merged spreadsheet;
[0019] Group the merged spreadsheet based on the file name field of the image, and each group corresponds to the same plankton category;
[0020] Randomly divide the data in each group into a training set, a validation set, and a test set;
[0021] Traverse each row of data in the merged spreadsheet, split the plankton image puzzle into several single particle images, and establish a standard dataset.
[0022] Preferably, the traversal of each row of data in the merged spreadsheet, splitting the plankton image puzzle into several single particle images, and establishing a standard dataset includes:
[0023] Read the fields representing the jigsaw puzzle file name, the category to which the image belongs, the image file name, the X coordinate of the upper left corner of the particle image, the Y coordinate of the upper left corner of the particle image, the width of the particle image, the height of the particle image, and the dataset to which it belongs in each line;
[0024] Use an image processing tool to read the image using the jigsaw puzzle file name, and crop out each particle image from the jigsaw puzzle image using the upper left corner coordinates and width and height information of the particle image, and save the cropped particle images to the corresponding folders according to the training set, validation set, test set, and image category.
[0025] Preferably, establish and train a classification model based on geometric morphological features, a classification model based on deep learning features, and a classification model based on the fusion of geometric morphological features and deep learning features, and screen out the optimal classification model for marine plankton images, including:
[0026] Based on the geometric morphological features of each marine plankton particle recorded in the merged spreadsheet and the training set, validation set, or test set to which it belongs, use several machine classification models in the PyCaret machine learning library, and adopt classical classification algorithms for training and parameter tuning to establish a classification model based only on geometric morphological features;
[0027] Based on the standard classification dataset, use the PyTorch computing package, and adopt the method of transfer learning to train several existing deep neural network models, adjust the model parameters on the training set, and determine the optimal classification model based on geometric morphological features on the validation set;
[0028] Take the output before the fully connected layer of the deep neural network as the deep features, use the machine learning model in the PyCaret computing package as the classification model to train and tune the parameters of the classification model for the deep features output before the fully connected layer of the deep neural network, and establish a classification model based only on deep features;
[0029] Based on the classification dataset StandardDataset, use the PyTorch computing package, and adopt the method of transfer learning to train several existing deep neural network models, adjust the model parameters on the training set, and determine the optimal classification model based on deep features on the validation set;
[0030] Concatenate the deep feature vectors extracted from before the fully connected layer of the deep neural network and the traditional morphological feature vectors exported from the VisualSpreadsheet software to form a new fused feature vector;
[0031] Train a classification model and optimize its parameters for the fused feature vectors on the PyCaret computing package to establish a classification model using the features after fusing geometric morphological features and depth features;
[0032] Calculate the evaluation metrics for the prediction results and the true results of the above model on the test set through the F1 score in the Python scikit-learn library, and select the optimal classification model for marine plankton images.
[0033] Preferably, in the process of calculating the evaluation metrics for the prediction results and the true results of the above model on the test set through the F1 score in the Python scikit-learn library and selecting the optimal classification model for marine plankton images, the expression of the F1 score is:
[0034]
[0035] wherein,
[0036]
[0037] In the formula, F1Score is the F1 score; Precision is the precision rate; Recall is the recall rate; TP represents true positive; FP represents false positive; FN represents false negative.
[0038] Preferably, the method of using the optimal classification model for marine plankton images to identify the plankton images of the collected samples to be detected, grouping, counting and storing the identified plankton images by category includes:
[0039] Use the FlowCAM imaging analysis system to collect the plankton images to be classified in the sample, and obtain the corresponding LST file and the TIF file corresponding to the LST file;
[0040] Parse the LST file into a CSV spreadsheet;
[0041] Obtain all the data in the CSV spreadsheet row by row, and rely on the file names of the plankton puzzles of various categories in the CSV spreadsheet to split the puzzle images in the TIF file into single-particle images, and save the position information and size information of each single-particle image;
[0042] Traverse all the image files of the sample to be classified, read the file name without the extension as the image number; use the marine plankton image classification model to predict and identify the image, and obtain the predicted category label corresponding to each image;
[0043] Save the image number and the predicted category label of each sample to a CSV spreadsheet file, and convert the CSV spreadsheet file to the CLA file format.
[0044] Preferably, it further includes: reading the CLA format file containing the predicted class labels using VisualSpreadsheet software and performing manual correction.
[0045] Compared with the prior art, the present application has at least the following beneficial effects:
[0046] 1. By introducing a deep learning algorithm and combining it with traditional morphological features, the present invention realizes the automation of image classification, and significantly improves the classification efficiency compared with traditional methods.
[0047] 2. The present invention provides a complete process from the puzzle image to the standardized data set, ensuring the standardization and consistency of data preparation, and laying a standardized data foundation for model training and prediction.
[0048] 3. The present invention supports multiple deep learning frameworks and models, and can select the best solution by comparing the performance of different models, showing excellent scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Some specific embodiments of the present invention will be described in detail hereinafter with reference to the accompanying drawings in an exemplary but not restrictive manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale.
[0050] In the figures:
[0051] Figure 1 is a schematic diagram of the overall process of the present invention.
[0052] Figure 2 is a visualization diagram of the model classification result of the embodiment of the present invention in VisualSpreadsheet software. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0054] The present invention extracts the implicit depth features of marine plankton images by introducing a deep neural network, and fuses them with the displayed geometric morphological features extracted by VisualSpreadsheet software, so as to establish a more accurate recognition method for marine plankton images, improve the image recognition accuracy and efficiency of FlowCAM equipment, and support the research and health guarantee of offshore plankton ecology.
[0055] As Figure 1 shown, a method for classifying marine plankton images includes the following steps:
[0056] S1. Use a flow cytometry particle imaging analyzer to collect plankton images in the seawater of the sea area to be studied;
[0057] S2. Designate categories for the collected plankton images in the study sea area, and establish an original classification dataset of marine plankton images;
[0058] S1. Convert the original classification dataset of marine plankton images into a standard classification dataset of marine plankton images;
[0059] S1. Establish and train a classification model based on geometric morphological features and a classification model based on deep learning features, and select the optimal model;
[0060] S1. Use the selected optimal model to identify the plankton images of the samples to be detected collected, and group, count and store the identified plankton images according to categories.
[0061] In this embodiment, a method for classifying marine plankton images includes the following process:
[0062] I. Establishment of the original classification dataset of marine plankton images
[0063] Use a flow cytometry particle imaging analyzer (FlowCAM) to collect plankton images in the seawater of the sea area to be studied, establish an original classification dataset RawDataset in FLB format in VisualSpreadsheet, and then convert it into a standard classification dataset StandardDataset of marine plankton images. In practical applications, this method is relatively convenient to operate.
[0064] The method for establishing a classification dataset in FLB format is as follows:
[0065] Use Visual Spreadsheet software to designate categories for the plankton images in the study sea area collected by FlowCAM, and establish an original classification dataset RawDataset of marine plankton images under this software;
[0066] Among them, images of marine plankton of the same category are saved in TIF format as a jigsaw puzzle, and meta-information such as the position of each particle image in the jigsaw puzzle and geometric shape parameters is saved as an FLB format file.
[0067] II. Establishment of a standard marine plankton image classification dataset
[0068] The original classification dataset RawDataset in FLB format can only be recognized and read by VisualSpreadsheet software and cannot be used in general image reading algorithms and deep learning model training. Therefore, it needs to be converted.
[0069] (1) Parse the FLB file into a spreadsheet file
[0070] Traverse the FLB files of each plankton category and parse them into CSV spreadsheet files respectively, as shown in Table 1. The steps are as follows:
[0071] (a) Use Python or other text reading tools to extract the text from the 2nd to 66th lines of the FLB file, all separated by "|"; retain the first string to obtain 65 strings, denoted as Str65.
[0072] (b) Use Python Pandas or other table parsing tools, specify the delimiter as "|", and the data type as string, to read the strings from the 67th line to the last line of the FLB file. The data read should contain 65 columns, and the field names of each column are specified as Str65.
[0073] (c) Insert the corresponding plankton category of the FLB file into the spreadsheet as a new column "species_name", such as the species_name column in Table 1.
[0074] Table 1 Column names and examples of the generated CSV table
[0075]
[0076]
[0077] (2) Merge the spreadsheets
[0078] Merge all the csv files obtained in the previous step, and delete duplicate rows based on the image_id field, only retaining the first record of each image, to obtain the merged spreadsheet, denoted as DF.
[0079] (3) Divide the dataset
[0080] Group the spreadsheet DF based on the "species_name" field, with each group corresponding to the same plankton category; randomly divide the data in each group into a training set, a validation set, and a test set; and add a new column "set" to the spreadsheet DF to mark the type of training set, validation set, and test set to which each data belongs, where the training set, validation set, and test set are marked as 'train', 'val', and 'test' respectively.
[0081] (4) Check the file integrity
[0082] Obtain the unique values of the "collage_file" field in the spreadsheet DF, which are the file names of the plankton jigsaw puzzles for each category, and check whether the corresponding jigsaw files are complete. If any are missing, exit; if complete, continue.
[0083] It should be noted that checking the integrity of the jigsaw files is a non-essential step for checking whether the corresponding jigsaw files are complete. Omitting this step can also achieve the image classification of marine plankton, but if there are missing cases, it will affect the statistics of marine plankton.
[0084] (5) Establish a standard dataset
[0085] Traverse each row of data in the spreadsheet DF, split the marine plankton image jigsaw into several single-particle images, and establish the standard dataset StandardDataset. The specific operation steps are as follows:
[0086] (a) Read the "collage_file", "species_name", "image_id", "image_x", "image_y", "image_w", "image_h", and "set" fields in each row. The above fields correspond to the jigsaw file name, the category to which the image belongs, the unique identifier of the image, the X and Y coordinates of the upper left corner of the particle image, the width and height of the particle image, and the dataset to which it belongs.
[0087] (b) Use the opencv image processing tool in the Python environment. Read the image using the jigsaw file name in the above information, crop each particle image from the jigsaw image using the upper left corner coordinates and width and height information of the particle image, and save the cropped particle images to the corresponding folders according to the dataset and image category.
[0088] Among them, the image file name is the unique identifier of the image, and the image format is a non-compressed or lossless compression format such as BMP or PNG; the directory structure is that the dataset folder contains different image category folders, and each image category folder includes the corresponding cropped particle images. This dataset is the standard dataset StandardDataset that can be used for training the deep neural network model.
[0089] III. Establishment of Marine Plankton Image Classification Model
[0090] (1) Establishment of Classification Model Based on Geometric Morphological Features
[0091] Based on the geometric morphological features (GeometricFeatures) of each marine plankton particle recorded in the spreadsheet DF and the affiliated training set, validation set and test set, use the PyCaret machine learning library or other machine learning toolkits, and adopt classical classification algorithms for training and parameter tuning to establish a classification model only based on geometric morphological features.
[0092] Among them, the training set and the validation set are combined for model training and parameter tuning, and the model performance is evaluated on the test set to determine the optimal classification model based on geometric morphological features.
[0093] Among them, 18 classification algorithms in the PyCaret machine learning library are used for training and parameter tuning, including logistic regression, K-nearest neighbor, naive Bayes, decision tree, support vector machine, radial basis function support vector machine, Gaussian process classifier, multi-layer perceptron, ridge regression, random forest, quadratic discriminant analysis, adaptive boosting, gradient boosting classifier, linear discriminant analysis, extreme forest, LightGBM, CatBoost and dummy model; use the compare_models function in the PyCaret package for the above model training, and use tune_model for model parameter tuning. During the model training process, preprocessing data of normalization (normalize = True) and principal component analysis for dimensionality reduction (PCA, pca = True) are enabled, and grid method is used for hyperparameter tuning.
[0094] (2) Establishment of Classification Model Based on Deep Learning Features
[0095] (a) Based on the classification dataset StandardDataset, use PyTorch or other deep learning frameworks, and adopt the method of transfer learning to train ResNet, EfficientNet v2, Swin Transformer v2 or other deep neural network models. Adjust the model parameters on the training set and determine the optimal model on the validation set.
[0096] (b) Use the output before the fully connected layer of the deep neural network as the deep feature DeepFeatures, and use the classification model in the PyCaret machine learning library to train and tune the parameters of the classification model for this deep feature to establish a classification model only based on deep features.
[0097] Among them, the training set and the validation set are combined for model training and parameter tuning, and the model performance is evaluated on the test set to determine the optimal classification model based on geometric morphological features. This process is similar to the model training, parameter tuning, and performance evaluation processes of the classification model based on geometric morphological features, and will not be elaborated here.
[0098] (3) Establishment of a classification model based on fused features
[0099] (a) Extract deep features before the fully connected layer of the deep neural network, denoted as Fd with a length of m, and export traditional morphological features GeometricFeatures from FlowCAMVisualSpreadsheet, denoted as Ft with a length of n.
[0100] (b) Concatenate the two one-dimensional feature vectors Fd and Ft to form a new feature vector Fn, expressed as: Fn = [Fd, Ft]; the length of the resulting vector Fn is m + n.
[0101] (c) Use the classification models in the PyCaret machine learning library to train and tune the parameters of the classification model for the newly fused features. Combine the training set and the validation set for model training and tuning, evaluate the model performance on the test set, and determine the optimal classification model based on the fused features. This process is similar to the model training, parameter tuning, and performance evaluation processes of the classification model based on geometric morphological features, and will not be elaborated here.
[0102] (4) Model performance comparison and selection
[0103] Use the F1 Score, Accuracy, Precision, and Recall calculation functions in the Python scikit-learn library to calculate the relevant evaluation metrics for the prediction results and the true results of the above models on the test set, and select the model with the highest F1 score as the optimal model, with the other metrics as references.
[0104]
[0105] Among them: TP represents True Positives, TN represents True Negatives, FP represents False Positives, and FN represents False Negatives.
[0106] In this embodiment, it can be seen from the comparison of the marine plankton image classification models that the method combining the deep features extracted by the EfficientNet-v2l neural network and the geometric morphological features, and then combined with the LDA classifier has the highest accuracy and F1 score, as shown in Table 2. Therefore, this classification model is used as the optimal classification model for marine plankton images to classify marine plankton images.
[0107] Table 2 Comparison of Marine Plankton Image Classification Models
[0108]
[0109]
[0110] III. Inference Application of Marine Plankton Image Classification Model
[0111] After a new marine plankton sample is collected using FlowCAM, the obtained LST metadata and corresponding mosaics cannot be recognized by common image reading modules such as opencv and need to be converted to a common image format.
[0112] In addition, after the prediction result is obtained, it is converted into the cla format recognizable by VisualSpreadsheet software, which can facilitate personnel familiar with this software to perform manual correction. Although the results obtained using deep learning methods also require manual correction, the image recognition accuracy has been greatly improved, and the need for manual work has been significantly reduced. The specific steps are as follows:
[0113] (1) Parse the LST file into a spreadsheet file
[0114] After collecting marine plankton images using FlowCAM and VisualSpreadsheet software, they are stored in different folders according to samples. The folders should include LST files and corresponding TIFF files. Parse the LST file into a CSV spreadsheet file, and the steps are as follows:
[0115] (a) Use Python or other text reading tools to extract the text from the 2nd to 66th lines of the LST file, all separated by "|"; retain the first string to obtain 65 strings, denoted as Str65.
[0116] (b) Use Python Pandas or other table parsing tools, specify the delimiter as "|", and the data type as string, to read the strings from the 67th line to the last line of the LST file. The read data should contain 65 columns, and the field names of each column are specified as Str65.
[0117] (2) Check the integrity of the mosaic file
[0118] Traverse all the above spreadsheets to obtain the unique values of the collage_file field. This value is the file name of the plankton puzzle for each category. Check whether the corresponding puzzle files are complete. If any are missing, exit; if complete, continue.
[0119] (3) Crop into particle images
[0120] Obtain the data of all the above spreadsheets row by row, and split the puzzle images into single-particle images. The specific operation steps are as follows:
[0121] (a) Read the collage_file, id, image_x, image_y, image_w, and image_h fields in each row. The above fields correspond to the puzzle file name, image number, the X and Y coordinates of the upper left corner of the particle image, the width, and the height of the particle image respectively.
[0122] (b) Use the opencv image processing tool in the Python environment. Read the image using the puzzle file name in the above information, and crop each particle image from the puzzle image using the upper left corner coordinates and width and height information of the particle image, and save it in the same folder. The folder name is the sample number. Among them, the image file name is the image number, and the image format is BMP or PNG and other uncompressed or lossless compression formats.
[0123] (4) Image category prediction
[0124] Traverse all the image files of each sample, and read the file name without the extension as the image number; use the optimal classification model for marine plankton images to predict and identify the images, and obtain the predicted category labels corresponding to each image; save the image IDs and predicted category labels of each sample as a CSV file.
[0125] (5) Generate a CLA file
[0126] Convert the CSV file to the CLA file format. The steps are as follows:
[0127] (a) Obtain the folder name where the CSV file is located as the sample number, obtain the total number of categories of the predicted categories in the CSV file, and initialize the list in the following format: "V3", "1", sample number, total number of categories. Among them, the first two strings are fixed version identifiers.
[0128] (b) Group the data in the CSV file according to the predicted categories, obtain all the image IDs and the total number of images in each group, and generate a string list for each group in the following format: category name, "OR", "0", number of images, all image IDs.
[0129] (c) Connect the elements in the list into a string using the line break character \n to form the content of the CLA file.
[0130] IV. Manual Correction of the Classification Results of Marine Plankton Images
[0131] Use the Visual Spreadsheet software to view the generated CLA file and perform manual correction on the specifications of all image categories.
[0132] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0133] It should be noted that the terms "first", "second", etc. in the description, claims and drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein.
[0134] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for classifying marine plankton images, characterized in that: include: Use a flow particle imaging analyzer to collect images of plankton in the seawater of the sea area to be studied; Manually assign categories to the collected plankton images in the research area and establish an original classification dataset of marine plankton images; Convert the original classification dataset of marine plankton images into a standard classification dataset of marine plankton images; Establish and train classification models based on geometric morphological features, classification models based on deep learning features, and classification models based on the fusion of geometric morphological features and deep learning features to screen out the optimal classification model for marine plankton images; The optimal classification model of marine plankton images is used to identify the collected plankton images of the samples to be tested, and the identified plankton images are grouped, counted and stored according to categories.
2. The marine plankton image classification method according to claim 1, characterized in that: The artificial classification of the collected plankton images in the research sea area is performed to establish an original classification data set of marine plankton images, including: The FlowCAM imaging analysis system was used to collect images of the plankton to be classified in the samples; Visual Spreadsheet software was used to assign categories to the collected plankton particle images; The images of marine plankton particles of the same category were saved in TIF format in the form of puzzles, and the meta-information of each particle image in the puzzle, including position and geometric morphological parameters, was saved in FLB format files, thereby establishing the original classification data set of marine plankton images under VisualSpreadsheet software.
3. The marine plankton image classification method according to claim 2, characterized in that: The method of converting the original classification dataset of the marine plankton image into a standard classification dataset of the marine plankton image comprises: Iterate through the FLB files of each type of plankton and parse them into CSV spreadsheets; Merge all the obtained CSV spreadsheets, delete duplicate rows based on the image file name field, and only keep the first record of each image to obtain the merged spreadsheet; The merged spreadsheet was grouped based on the image's file name field, with each group corresponding to the same plankton category; The data in each group is randomly divided into training set, validation set and test set; Traverse each row of data in the merged spreadsheet, split the marine plankton image mosaic into several single particle images, and establish a standard data set.
4. The marine plankton image classification method according to claim 3, characterized in that: The method traverses each row of data in the merged spreadsheet, divides the marine plankton image mosaic into a number of single particle images, and establishes a standard data set, including: Read the fields in each line representing the puzzle file name, image category, image file name, X coordinate of the upper left corner of the particle image, Y coordinate of the upper left corner of the particle image, width of the particle image, height of the particle image, and the data set to which it belongs; Use image processing tools to read the image using the puzzle file name, and use the upper left corner coordinates and width and height information of the particle image to crop each particle image from the puzzle image, and save the cropped particle images to the corresponding folder according to the dataset and image category.
5. The marine plankton image classification method according to claim 4, characterized in that: Establish and train classification models based on geometric morphological features, classification models based on deep learning features, and classification models based on the fusion of geometric morphological features and deep learning features, and screen out the optimal classification model for marine plankton images, including: Based on the geometric morphological characteristics of each marine plankton particle recorded in the merged spreadsheet and the training set, validation set or test set to which it belongs, several machine classification models in the PyCaret machine learning library are used, and classical classification algorithms are used for training and parameter tuning to establish a classification model based only on geometric morphological characteristics; Based on the standard classification data set, using the PyTorch computing package and the transfer learning method, several existing deep neural network models are trained, the model parameters are adjusted on the training set, and the optimal classification model based on geometric morphological features is determined on the validation set; The output before the fully connected layer of the deep neural network is used as the deep feature, and the machine learning model in the PyCaret computing package is used as the classification model to train the classification model and optimize the parameters of the deep features of the output before the fully connected layer of the deep neural network, and establish a classification model based only on deep features; Based on the classification dataset StandardDataset, using the PyTorch computing package and the transfer learning method, several existing deep neural network models are trained, the model parameters are adjusted on the training set, and the optimal classification model based on deep features is determined on the validation set; The deep feature vector extracted from the fully connected layer of the deep neural network and the traditional morphological feature vector exported from the VisualSpreadsheet software are spliced to form a new fused feature vector; Training and parameter tuning of the classification model for the fused feature vectors on the PyCaret computing package; The F1 score in the Python scikit-learn library is used to calculate the evaluation index of the predicted results and actual results of the above model on the test set to screen out the optimal classification model for marine plankton images.
6. The marine plankton image classification method according to claim 5, characterized in that: The F1 score in the Python scikit-learn library is used to calculate the evaluation index of the optimal classification model of marine plankton images by comparing the predicted results and the actual results of the above model on the test set. The expression of the F1 score is: in, In the formula, F1 Score is the F1 score; Precision is the precision rate; Recal is the recall rate; TP means true positive; FP means false positive; FN means false negative.
7. The marine plankton image classification method according to claim 6, characterized in that: The method uses the optimal classification model for marine plankton images to identify the collected plankton images of the sample to be detected, and groups, counts and stores the identified plankton images according to categories, including: The FlowCAM imaging analysis system is used to collect the images of the plankton to be classified in the sample, and the corresponding LST files and TIF files corresponding to the LST files are obtained; Parse the LST file into a CSV spreadsheet; All data in the CSV spreadsheet are obtained line by line, and the mosaic images in the TIF file are divided into single particle images according to the file names of the plankton mosaics of each category in the CSV spreadsheet, and the position information and size information of each single particle image are saved; Traverse all image files of the samples to be classified, read the file name without the extension as the image number; use the marine plankton image classification model to predict and identify the image, and obtain the predicted category label corresponding to each image; Save the image number and predicted category label of each sample into a CSV spreadsheet file, and convert the CSV spreadsheet file to the CLA file format.
8. The marine plankton image classification method according to claim 7, characterized in that: Also includes: VisualSpreadsheet software was used to read the CLA format file containing the predicted category labels and perform manual correction.