Algae classification method based on algae photosynthetic pigment characteristic difference

By establishing an algae photosynthetic pigment database and using LNN and CNN algorithms, digital classification of algae is achieved, which solves the problems of low efficiency and high subjectivity of traditional methods, improves the accuracy of algae classification and the compatibility of machine learning, and provides technical support for ecological risk assessment.

CN120673915APending Publication Date: 2025-09-19KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510786525.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional algae classification methods are inefficient and highly subjective, making them difficult to efficiently integrate with machine learning algorithms, which limits the accuracy of predicting specific differences in algal toxicity responses.

Method used

A database of algae photosynthetic pigments was established, integrating the molecular formula, molecular weight, elemental composition and fluorescence spectral characteristics of photosynthetic pigments. A classification model was constructed using LNN and CNN algorithms to achieve digital classification of algae.

Benefits of technology

It realizes the digital and intelligent classification of algae, improves the classification efficiency and accuracy, supports the deep integration of algae toxicity response characteristics and compound toxicity data, and provides technical support for accurate ecological risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673915A_ABST
    Figure CN120673915A_ABST
Patent Text Reader

Abstract

The invention discloses an algae classification method based on algae photosynthetic pigment characteristic difference, which comprises the following steps: S1, obtaining photosynthetic pigment types, molecular formulas and fluorescence spectrograms of different algae, and establishing an algae photosynthetic pigment database; s2, the chemical indexes and the spectral features are fused into multi-dimensional feature vectors, and standardization processing is carried out on the multi-dimensional feature vectors; s3, dividing the data set into a training set and a test set, and training and optimizing a classification model; and S4, AI-based algae classification verification and result presentation. According to the method, the photosynthetic pigment characteristics of the algae are converted into digital indexes, so that a quantitative and efficient new method is provided for algae classification, a new way is opened up for application of machine learning in algae toxicity prediction, the prediction and evaluation ability of the algae in an aquatic ecosystem affected by toxic and harmful substances is improved, and the method is suitable for popularization and application. And a technical support is provided for environmental protection and ecological risk early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital classification of algae, and relates to a classification method for algae based on the differences in the characteristics of algae photosynthetic pigments. Specifically, the present invention is a classification system based on the characteristics of algae photosynthetic pigments, integrating chemical parameters such as pigment molecular weight and elemental composition, and fluorescence spectral characteristics, combined with LNN and CNN algorithms. Background Art

[0002] As a key component of aquatic ecosystems, algae classification is crucial for ecological research, water quality monitoring, and biodiversity conservation. Traditional classification methods rely primarily on morphological observations and physiological and biochemical experiments, which are limited by low efficiency and subjectivity. Different algal phyla differ in the composition of their photosynthetic pigments. These pigments not only play a central role in photosynthesis, but their unique molecular formulas and molecular weights also serve as potential distinguishing features between algae.

[0003] Machine learning technology is increasingly demonstrating significant potential in biological classification. It can automatically construct classification models based on input feature data, enabling rapid and accurate classification of biological samples. Machine learning has also achieved breakthroughs in compound toxicity prediction, enabling mathematical modeling of complex toxicity mechanisms through the construction of quantitative structure-activity relationship (QSAR) models. However, its application remains limited by the unstructured nature of biological taxonomic data. Currently, because the algal classification system has not yet been digitized, traditional biological classification features are difficult to effectively integrate with machine learning algorithms. This results in the models being unable to quantify and analyze the specific differences in toxicity responses across phyla, such as green algae and diatoms, limiting prediction accuracy. Therefore, this paper proposes an algal classification method based on the differences in the properties of algal photosynthetic pigments. This method uses the differences in the photosynthetic pigments of different algae, combined with chemical parameters such as molecular formula and molecular weight, and spectral characteristics, to digitally classify algae, ultimately constructing an algal classification system compatible with machine learning algorithms. Summary of the Invention

[0004] This invention aims to propose an algae classification method based on the differences in algal photosynthetic pigment characteristics. First, a systematic database of algal photosynthetic pigments is established, integrating physicochemical parameters such as molecular formula, molecular weight, and elemental composition, as well as spectral characteristics such as fluorescence spectra. Feature encoding is then used to convert pigment composition differences into a quantifiable digital matrix. Finally, a classification model compatible with machine learning algorithms is constructed. This method transcends the qualitative description limitations of traditional classification and provides a new paradigm for intelligent algae identification. In the future, it will also enable the deep integration of algal toxicity response characteristics with compound toxicity data, providing technical support for precise ecological risk assessment.

[0005] The technical solution of the present invention is: a method for classifying algae based on the differences in algae photosynthetic pigment characteristics, comprising: S1: Obtain the types, molecular formulas and fluorescence spectra of photosynthetic pigments of different algae and establish an algae photosynthetic pigment database; S2: Fusion of chemical indicators and spectral features into multidimensional feature vectors and standardization; S3: Divide the dataset into training and test sets, and use the liquid neural network algorithm (LNN) and convolutional neural network (CNN) to train and optimize the classification model; S4: AI-based algae classification verification and result presentation.

[0006] Furthermore, the step S1 specifically includes: We collected photosynthetic pigment types, molecular formulas, molecular weights, elemental compositions, and fluorescence spectra of major algae, including those from the Chlorophyta, Cyanobacteria, Bacillariophyta, Dinophyta, Phaeophyta, Rhodophyta, Euglenophyta, Charophyta, Cryptophyta, Chrysophyta, Glaucophyta, and Xanthophyta, from databases such as PubChem, KEGG, and Web of Science. This included core pigments such as chlorophyll, phycobilins, carotenoids, and xanthophylls. By integrating these chemical indicators and fluorescence spectra, we established a scalable algal photosynthetic pigment database that supports dynamic updates and is compatible with multi-source data.

[0007] Furthermore, the step S2 specifically includes: The theoretical molecular weight and elemental composition ratio of the pigment were calculated based on the molecular formula to form a basic chemical index vector. The quantitative central features of the fluorescence spectrum were extracted, including the main peak position, secondary peak position, secondary peak offset, spectral symmetry index, half-peak width, and the ratio of the main peak to secondary peak areas. The chemical indicators and spectral features were fused into a multidimensional feature vector, and the Z-score normalization method was used to normalize the data to eliminate dimensional differences.

[0008] Preferably, the multidimensional feature vector includes a chemical index vector of the theoretical molecular weight and elemental composition ratio (C:O:H) calculated based on the molecular formula, and a spectral feature vector of the main peak position, secondary peak offset, spectral symmetry index, half-peak width, and main peak to secondary peak area ratio of the fluorescence spectrum.

[0009] Furthermore, the step S3 specifically includes: The dataset was split into training and test sets at a 7:3 ratio. Model training was performed using the Liquid Neural Network (LNN) algorithm and the Convolutional Neural Network (CNN). During training, algorithm hyperparameters, such as the learning rate and number of iterations, were adjusted to optimize model performance. The trained model was validated using the test set, and the model with the highest classification accuracy was selected as the algae classification model.

[0010] Furthermore, the step S4 specifically includes: The AI ​​is fed with machine learning data analysis results, fluorescence spectra, and their spectral characteristics. By learning the types of photosynthetic pigments present and their spectral peaks, the AI ​​determines the types, phyla, and corresponding digital indicators of photosynthetic pigments in unknown algae samples. The results can be visualized, including classification confidence, spectral fitting curves, and similarity comparisons with the database.

[0011] The technical solution of the present invention also includes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the algae classification method based on the differences in algae photosynthetic pigment characteristics of the present invention.

[0012] The technical solution of the present invention also includes a computer device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it can process dynamic spectral time series data through LNN, extract spectral image features through CNN, and jointly output classification results.

[0013] Compared with the existing technology, it has the following beneficial effects: 1) Break through the limitations of qualitative descriptions in traditional classification methods, convert qualitative descriptions into quantitative indicators such as molecular formulas and spectral characteristics, realize digital and intelligent classification of algae, and improve classification efficiency and accuracy.

[0014] 2) It is compatible with machine learning and helps to achieve a deep integration of algal toxicity response characteristics and compound toxicity data, better understand the role of algae in the ecosystem and its response to environmental changes, and provide strong technical support for accurate ecological risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Flowchart of the establishment method of the present invention.

[0016] Figure 2 It is a specific framework flow chart of the present invention. DETAILED DESCRIPTION

[0017] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below with reference to the accompanying drawings and embodiments.

[0018] By converting the characteristics of algae photosynthetic pigments into digital indicators, this invention not only provides a new quantitative and efficient method for algae classification, but also opens up new avenues for the application of machine learning in algae toxicity prediction. It helps to improve the ability to predict and evaluate the effects of toxic and harmful substances on algae in aquatic ecosystems, and provides technical support for environmental protection and ecological risk warning.

[0019] Example 1 The algae classification method based on the differences in algae photosynthetic pigment characteristics of this embodiment includes the following steps: S1: Obtain the types, molecular formulas and fluorescence spectra of photosynthetic pigments of different algae and establish an algae photosynthetic pigment database; S2: Fusion of chemical indicators and spectral features into multidimensional feature vectors and standardization; S3: Divide the dataset into training and test sets, and use the liquid neural network algorithm (LNN) and convolutional neural network (CNN) to train and optimize the classification model; S4: AI-based algae classification verification and result presentation.

[0020] like Figure 2 Specifically, data on the types, molecular formulas, molecular weights, elemental compositions, and fluorescence spectra of photosynthetic pigments of major algae, including Chlorophyta, Cyanobacteria, Bacillariophyta, Dinophyta, Phaeophyta, Rhodophyta, Euglenophyta, Charophyta, Cryptophyta, Chrysophyta, Glaucophyta, and Xanthophyta, were collected from databases such as PubChem, KEGG, and Web of Science. This included core pigments such as chlorophyll, phycobilins, carotenoids, and xanthophylls. By integrating these chemical indicators and fluorescence spectra, a scalable algal photosynthetic pigment database was established, supporting dynamic updates and compatibility with multi-source data.

[0021] The theoretical molecular weight and elemental composition ratio of the pigment were calculated based on the molecular formula to form a basic chemical index vector. The quantitative central features of the fluorescence spectrum were extracted, including the main peak position, secondary peak position, secondary peak offset, spectral symmetry index, half-peak width, and the ratio of the main peak to secondary peak areas. The chemical indicators and spectral features were fused into a multidimensional feature vector, and the Z-score normalization method was used to normalize the data to eliminate dimensional differences.

[0022] The dataset was split into training and test sets at a ratio of 7:3, and model training was performed using the Liquid Neural Network (LNN) and Convolutional Neural Network (CNN) algorithms. During training, the LNN algorithm leveraged its advantages in processing dynamic systems and time-varying information, leveraging its ability to simulate dynamic processes to process fluorescence spectra of algae at different growth stages. Simultaneously, the powerful image feature extraction capabilities of the Convolutional Neural Network (CNN) were leveraged to perform image-level feature extraction and enhance spectral pattern recognition. Model performance was optimized by adjusting algorithm hyperparameters, such as the learning rate and number of iterations. The trained models were validated on the test set, and the model with the highest classification accuracy was selected as the algae classification model.

[0023] The machine learning model's analysis results, along with the algae sample's fluorescence spectra and their characteristic spectral peaks, are integrated and fed into the AI. Using this learned feature knowledge, the AI ​​accurately identifies the types of photosynthetic pigments in unknown algae samples, determines the algae's phylum based on the pigment's composition, and outputs corresponding digital indicators, providing a quantitative basis for algae classification.

[0024] There are four key technical points in the identification of this example: 1) integration and construction of the algae photosynthetic pigment database; 2) data standardization processing; 3) training and optimization of the classification model; 4) AI-based algae classification verification and result presentation.

[0025] 1. Integration and construction of algae photosynthetic pigment database, including: Through literature research, we obtained chemical parameters such as pigment composition, molecular formula, molecular weight, and elemental composition for different algae, as well as spectral characteristics such as fluorescence spectrum peaks, spectral symmetry index, and peak area. We then set up a reasonable data table structure based on the data characteristics. For example, we set up a "pigment chemical property table" to record chemical indicators such as pigment molecular formula, molecular weight, and elemental composition, as shown in Table 1.

[0026] Table 1 Chemical properties of different algae pigments 2. Data standardization, including: Based on the molecular formula of the pigment, the theoretical molecular weight and element molar ratio are calculated using chemometric methods to form a basic chemical index vector V chem =[ M w, C:H:O]. For example, the main pigments in Phaeophyta are chlorophyll a, chlorophyll c1, chlorophyll c2, β-carotene, fucoxanthin, and diatomaxanthin, and their molecular formulas are C 55 H 72 N4O5Mg、C 35 H 30 O5N4Mg、C 35 H 28 O5N4Mg、C 40 H 56 、C 42 H 58 O6、C 40 H 54O2, the molecular weights are 893.49, 610.94, 608.93, 536.87, 658.91, 566.86, the total molecular weight is 3876, the C, H, O element composition ratios are 55:72:5, 35:30:5, 35:28:5, 40:56:0, 42:58:6, 40:54:2, and the total element composition ratio is C:H:O=247:298:23, then its chemical index vector can be written as V chem =[3876, 247:298:23].

[0027] Extract the main peak position λmain, secondary peak position λsecondary, secondary peak offset Δλsecondary, spectrum symmetry index SI, half-peak width FWHM, main peak to secondary peak area ratio A from the fluorescence spectrum. PC:PE , forming the spectral feature vector V spec=[λmain, λsecondary, Δλsecondary, SI, FWHM, A PC:PE For example, the main peak of Peridinium is at 680nm, the secondary peak is at 450nm, the main peak area is 3500, the secondary peak area is 1800, the spectrum symmetry index is 0.45, and the half-peak width is 40. V spec= [680, - 230, 0.45, 40, 3500:1800].

[0028] The chemical index vector and the spectral feature vector are sequentially spliced ​​to form a multidimensional feature vector Vtotal =[ V chem, V spec].

[0029] The Z-score standardization method is used to normalize the multidimensional feature vector. The calculation formula is: Among them, x is the standardized data, x i is the original data, μ is the mean value of the original data, and σ is the standard deviation of the original data.

[0030] 3. Classification model training and optimization, including: The processed dataset was divided into a training set and a test set in a ratio of 7:3. Stratified sampling was used to ensure that the ratio of each algae sample in the training and test sets was consistent with that in the original dataset, avoiding the impact of sample imbalance on model training results.

[0031] Build a liquid neural network (LNN) model and set the number of neurons in the input layer, the number of hidden layers and neurons, and the number of neurons in the output layer. Specifically: The number of neurons in the input layer is consistent with the dimension of the multidimensional feature vector. Assuming that the feature vector used for algae classification contains 10 dimensions, the number of neurons in the input layer is set to 10.

[0032] The number of hidden layers is generally set to 3-5 layers, and the number of neurons is determined according to an empirical formula, that is, the average number of input and output layers. Assuming that the number of neurons in the input layer is 10 and the number of neurons in the output layer is 12, the number of neurons in each hidden layer can be considered to be set to (10+12) / 2=11.

[0033] The output layer has the same number of neurons as the number of algae categories. For example, if we want to classify 12 different algae, the output layer will have 12 neurons, one for each algae category. During training, these neurons use activation functions (such as the softmax function) to convert the input from the previous layer into a probability distribution, representing the probability that the input sample belongs to each algae category.

[0034] Build a convolutional neural network (CNN) model, including multiple convolutional layers, pooling layers, and fully connected layers.

[0035] Convolutional layer: Convolution kernels of different sizes (e.g., 3×3, 5×5) are set. The kernels are slid across the input data to extract local features of the algae image. The number of kernels is also adjusted to control the richness of the extracted features.

[0036] Pooling layer: Uses maximum pooling or average pooling to downsample the feature maps output by the convolutional layer, reducing feature dimensions and computational complexity while retaining important feature information. Optimize the pooling layer parameters to determine the appropriate pooling window size and stride.

[0037] Fully connected layer: Adjust the number of neurons in the fully connected layer to integrate and classify the output features of the pooling layer.

[0038] The model was trained using the Adam optimizer, which adaptively adjusts the learning rate to improve training efficiency and model convergence speed. The cross-entropy loss function was selected as the loss function to measure the difference between the model's predicted probability distribution and the true category distribution. This guides the update direction of the model parameters and makes the model's predictions more accurate.

[0039] During the training process of the above model, hyperparameters such as learning rate (range: 0.001-0.1) and number of iterations (range: 100-1000) need to be adjusted.

[0040] In this example, after training, both the LNN and CNN models were validated using the test set. Model performance was comprehensively evaluated by calculating metrics such as classification accuracy, precision, recall, and F1 score on the test set. The test results of the two models were compared, and the model with the highest classification accuracy and excellent performance on other performance metrics was selected as the final algae classification model.

[0041] The calculation formula for accuracy is: The calculation formula for precision is: The calculation formula of recall rate Recall is: The calculation formula of F1 value is: Among them, TP (True Positives): true positive examples, which are predicted to be positive examples and are actually positive examples; FP (False Positives): false positive examples, which are predicted to be positive examples but are actually negative examples; FN (false Negatives): false negative examples, which are predicted to be negative examples but are actually positive examples; TN (True Negatives): true negative examples, which are predicted to be negative examples and are actually negative examples.

[0042] 4. AI-based algae classification verification and results presentation, including: The machine learning data analysis results, fluorescence spectra, and their spectral characteristics are fed into the AI ​​system. Using the learned feature knowledge, the AI ​​identifies the type and spectral peaks of the photosynthetic pigments in the unknown algae sample. This information is then compared with the data in the algae photosynthetic pigment database to determine the unknown algae sample's phylum and corresponding digital indicators. Simultaneously, a classification confidence chart, spectral fitting curves, and a similarity comparison report with the database are generated.

[0043] The above describes the basic principles and main features of the present invention and the advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from all points of view, the embodiments should be regarded as illustrative and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and range of equivalents of the claims be included in the present invention. Any reference signs in the claims should not be construed as limiting the claim to which they relate.

[0044] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A method for classifying algae based on differences in algae photosynthetic pigment characteristics, characterized in that: The following steps are involved: S1: Obtain the types, molecular formulas and fluorescence spectra of photosynthetic pigments of different algae and establish an algae photosynthetic pigment database; S2: Fusion of chemical indicators and spectral features into multidimensional feature vectors and standardization; S3: Divide the dataset into training and test sets to train and optimize the classification model; S4: AI-based algae classification verification and result presentation.

2. The algae classification method based on the difference in algae photosynthetic pigment characteristics according to claim 1, characterized in that: The step S1 specifically includes: collecting the photosynthetic pigment type, molecular formula, molecular weight, elemental composition and fluorescence spectrum data of algae from a database, wherein the algae include any one or more of green algae, cyanobacteria, diatoms, dinoflagellates, brown algae, red algae, euglenophyceae, charophyceae, cryptophyceae, chrysophyceae, glaucophyceae, and xanthophyceae, but are not limited to the above algae; the pigments include any one or more of chlorophyll, phycobilin, carotenoid, and xanthophyceae, but are not limited to the above pigments; integrating the above chemical indicators and fluorescence spectrum data to establish an expandable algae photosynthetic pigment database that supports dynamic updates and compatibility with multi-source data.

3. The algae classification method based on the difference in algae photosynthetic pigment characteristics according to claim 1, characterized in that: The step S2 specifically includes: calculating the theoretical molecular weight and elemental composition ratio of the pigment based on the molecular formula to form a basic chemical index vector; extracting the quantitative central features of the fluorescence spectrum, including the main peak position, secondary peak position, secondary peak offset, spectral symmetry index, half-peak width, and the main peak to secondary peak area ratio; fusing the chemical indicators and spectral features into a multidimensional feature vector, and using the Z-score normalization method to normalize the data to eliminate dimensional differences.

4. The algae classification method based on the difference in algae photosynthetic pigment characteristics according to claim 3, characterized in that: The multidimensional feature vector includes a chemical index vector of a theoretical molecular weight and an elemental composition ratio (C:O:H) calculated based on a molecular formula, and a spectral feature vector of a main peak position, a secondary peak offset, a spectral symmetry index, a half-peak width, and an area ratio of a main peak to a secondary peak of a fluorescence spectrum.

5. The algae classification method based on the difference in algae photosynthetic pigment characteristics according to claim 1, characterized in that: The step S3 specifically includes: dividing the data set into a training set and a test set in a ratio of 7:3, using a liquid neural network algorithm (LNN) and a convolutional neural network (CNN) to train the model; during the training process, optimizing the model performance by adjusting the algorithm's hyperparameters; verifying the trained model with the test set, and ultimately determining the model with the highest classification accuracy as the algae classification model.

6. The algae classification method based on the difference in algae photosynthetic pigment characteristics according to claim 1, characterized in that: The step S4 specifically includes: feeding the AI ​​with the data analysis results of machine learning, the fluorescence spectrum and its spectral characteristics; the AI ​​determines the type of photosynthetic pigments in the unknown algae sample, the category to which it belongs, and the corresponding digital indicators by learning the types of photosynthetic pigments contained and the spectral characteristic peaks; the results support visual output, including classification confidence, spectral fitting curve and similarity comparison with the database.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it can process dynamic spectral time series data through LNN, extract spectral image features through CNN, and jointly output classification results.