Blue-green algae type distinguishing method based on three-dimensional fluorescence spectrum

By preprocessing and feature extraction of three-dimensional fluorescence spectral data, combined with LightGBM multi-classifier, the problem of difficulty in distinguishing cyanobacter species in the prior art is solved, and a high-accurate cyanobacter species discrimination is achieved.

CN120198704APending Publication Date: 2025-06-24ZHEJIANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510095508.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish different types of cyanobacteria, especially when the fluorescence signal is close and the background fluorescence signal is complex, resulting in frequent misjudgment during the classification process.

Method used

The cyanobacter species discrimination method based on three-dimensional fluorescence spectrum was used to pre-process the data through the Delaunay triangle interpolation method, the blank background solvent subtraction method and the Savitzky-Golay polynomial surface smoothing method, and the spectral features were extracted in combination with the parallel factor analysis method and the peak ratio method, and finally the LightGBM multi-classifier was used for classification.

Benefits of technology

Accurate judgment of different cyanobacter species is achieved, classification accuracy is improved, misjudgment rate is reduced, and the performance is good especially in real freshwater environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198704A_ABST
    Figure CN120198704A_ABST
Patent Text Reader

Abstract

The invention discloses a blue-green algae type discrimination method based on a three-dimensional fluorescence spectrum. The method comprises the following steps: S0, obtaining a to-be-detected sample; the method comprises the following steps: S1, acquiring original three-dimensional fluorescence spectrum data; s2, preprocessing the original three-dimensional fluorescence spectrum data; rayleigh scattering and Raman scattering in the original three-dimensional fluorescence spectrum are removed, and smooth denoising is carried out on the original three-dimensional fluorescence spectrum; s3, spectral features are extracted based on a parallel factor-peak ratio method, and spectral feature vectors are obtained; and S4, carrying out missing value supplement and data standardization processing on the obtained spectral feature vectors, inputting the processed feature vectors into a trained LightGBM multi-classifier, outputting probabilities by a classification model, and selecting a category with the maximum probability. According to the method, the distribution mechanism of the blue-green algae fluorescent pigment is combined, and the fluorescent component characteristics and the relative content characteristics of the fluorescent components of the blue-green algae are comprehensively considered, so that the three-dimensional fluorescence spectrum information of the blue-green algae sample can be more comprehensively, completely and meticulously represented, and the variety discrimination of different blue-green algae with smaller difference is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional fluorescence cyanobacteria detection, and specifically relates to a method for discriminating cyanobacteria species based on three-dimensional fluorescence spectroscopy. Background Art

[0002] The massive reproduction of cyanobacteria in fresh water is not only one of the important inducements of water eutrophication, but also accompanied by the release of cyanobacteria secondary metabolites such as microcystins and neurotoxins, resulting in liver damage and nerve poisoning caused by human intake, posing a serious threat to the sustainable development of the ecological environment and human health and safety. Therefore, it is very important to identify the species of cyanobacteria. On the one hand, the threat and attention of toxic and harmful cyanobacteria are high; on the other hand, there are differences in emergency treatment means for different species of cyanobacteria. However, there are some key problems in the research on the identification of cyanobacteria species based on three-dimensional fluorescence technology.

[0003] First of all, the spectral differences of different species of cyanobacteria are small. Since different species of cyanobacteria under the same genus of cyanobacteria have similar types of fluorescent pigments. Therefore, their fluorescence signals are relatively close, the number and position of fluorescence peaks are similar, and multiple groups of fluorescence peaks overlap, making it difficult to effectively distinguish them through conventional feature extraction and classification recognition algorithms. At the same time, real fresh water environments usually contain various fluorescent substances, and these substances will generate complex background fluorescence signals. When the cyanobacteria concentration is low, the characteristic fluorescence peak information is easily masked by the fresh water background fluorescence signal due to insufficient fluorescence intensity, resulting in misjudgment during the classification process. Summary of the Invention

[0004] The purpose of the present invention is to overcome the problems existing in the current cyanobacteria classification and recognition methods, and provide a method for discriminating cyanobacteria species based on three-dimensional fluorescence spectroscopy.

[0005] In order to achieve the above purpose, the present invention is realized through the following technical solutions: A method for discriminating cyanobacteria species by three-dimensional fluorescence spectroscopy, comprising the following steps: S0. Obtain a sample to be detected; Collect an algal solution sample, filter it through a vacuum filtration device and a 0.45 μm water-based filter membrane to obtain a sample to be measured; S1. Obtain the original three-dimensional fluorescence spectral data; Take 2 mL of the sample to be measured and place it in a quartz cuvette of a three-dimensional fluorescence spectrometer to measure the three-dimensional fluorescence spectrum of the solution; S2. Pretreat the original three-dimensional fluorescence spectral data; Remove Rayleigh scattering in the original three-dimensional fluorescence spectrum using Delaunay triangle interpolation method, remove Raman scattering in the original three-dimensional fluorescence spectrum using blank background solvent subtraction method, and smooth and denoise the original three-dimensional fluorescence spectrum using Savitzky-Golay polynomial surface smoothing method; S3. Extract spectral features based on parallel factor-peak ratio method Extract the fluorescence component features of the three-dimensional fluorescence spectrum preprocessed in step S2 based on parallel factor analysis method, extract the fluorescence peak ratio features of the three-dimensional fluorescence spectrum preprocessed in step S2 based on peak ratio method, and combine the fluorescence component features and fluorescence peak ratio features to obtain a spectral feature vector; S4. Perform classification of LightGBM multi-classifier Perform missing value supplementation and data standardization processing on the obtained spectral feature vector, input the processed feature vector into the trained LightGBM multi-classifier, the classification model outputs the probability of the current sample to be tested corresponding to each existing category, and select the category with the highest probability as the cyanobacteria category to which the current sample to be tested belongs.

[0006] The described S3. includes the following steps: S3.1 Comprehensively determine the number of factors F of the parallel factor analysis method based on kernel consistency diagnostic method and residual analysis method; S3.2 Decompose the three-dimensional fluorescence spectrum preprocessed in S2 based on parallel factor analysis method to obtain a concentration score matrix A, an excitation loading matrix B, and an emission loading matrix C, where the excitation loading matrix and the emission loading matrix are used to characterize the fluorescence component features of the three-dimensional fluorescence spectrum; S3.3 Find the emission wavelength ex F and the excitation wavelength em F corresponding to the peak points of each factor matrix B and matrix C, so as to determine that the wavelengths of the F peak points corresponding to the F factors are S3.4 Search for the f (f = 1,..., F) -th peak point in the three-dimensional fluorescence spectrum S3.5 Search for the strongest fluorescence peak point in a small range near and update it as the new peak point S3.6 Take the N×N region centered on as the characteristic peak region of the f -th factor in the three-dimensional fluorescence spectrum of this sample; S3.7 Repeat steps S3.4 - S3.6 to obtain F characteristic peak regions, that is, F N×N matrices, and these matrices are the fluorescence peak regions corresponding to the three-dimensional fluorescence spectrum; S3.8 Calculate the ratio of each pair of the F peak regions. Here, the matrix ratio is matrix point division, that is, the corresponding elements in the matrix are divided by each other, so as to obtain an N×N ratio matrix, which is the fluorescence pigment ratio feature of the sample obtained by the peak ratio algorithm; S3.9 Straighten the fluorescence component features of the three-dimensional fluorescence spectrum extracted in S3.2 and the fluorescence peak ratio features of the three-dimensional fluorescence spectrum extracted in S3.8 by rows, and combine them end to end to obtain the feature vector of the three-dimensional fluorescence spectrum of the sample to be measured.

[0007] The described S4. includes the following steps: S4.1 Process the missing data in the feature vector obtained in S3 by using the median filling method, that is, calculate the median of each column and fill the missing value with this value. In addition, standardize the data to ensure the comparability of the combined features; S4.2 Calculate the weighted sum of the prediction results of all decision trees in the LightGBM multi-classifier; S4.3 Convert the output of S4.2 into a probability distribution through the Softmax function, and select the category with the highest probability as the category to which the current cyanobacteria sample belongs.

[0008] The training process of the LightGBM multi-classifier in the described S4. includes the following steps: S5.1 Obtain training samples, including different types of cyanobacteria; S5.2 Obtain the original three-dimensional fluorescence spectrum data of the training samples; S5.3 Preprocess the original three-dimensional fluorescence spectrum data of the training samples; S5.4 Extract the spectral features of the preprocessed three-dimensional fluorescence spectrum of the training samples based on the parallel factor-peak ratio method; S5.5 Perform missing value supplementation and data standardization processing on the obtained spectral feature vectors, input the processed feature vectors into the LightGBM multi-classifier for training, set hyperparameters, including learning rate, tree depth, number of leaf nodes, etc., and construct multiple decision trees to ensure that the model can fully learn the spectral features of different types of cyanobacteria, and obtain a trained LightGBM multi-classifier.

[0009] Compared with the existing methods, the present invention has the following beneficial effects: The method of the present invention combines the distribution mechanism of cyanobacteria fluorescence pigments, comprehensively considers the fluorescence component features and the relative content features of fluorescence components of cyanobacteria, can more comprehensively, completely and meticulously characterize the three-dimensional fluorescence spectrum information of cyanobacteria samples, and realize the discrimination of different cyanobacteria species with small differences. Description of the Drawings

[0010] Figure 1This is the flow chart of the method for discriminating cyanobacteria species based on three-dimensional fluorescence spectroscopy of the present invention.

[0011] Figure 2 This is the schematic diagram of extracting features by the peak ratio algorithm.

[0012] Figure 3 This is the schematic diagram of the three-dimensional fluorescence spectra of Oscillatoria, Nostoc, Aphanizomenon flos-aquae, Chroococcus, Microcystis aeruginosa, Phormidium, and Anabaena.

[0013] Figure 4A 、 Figure 4B This is the analysis result of the parallel factor analysis method. The left part is the emission loading matrix, the middle part is the excitation loading matrix, and the right part is the component fluorescence spectrum.

[0014] Figure 5 This is the characteristic peak region obtained by searching. The upper figure is the characteristic peak region of Chroococcus, and the lower figure is the characteristic peak region of Microcystis aeruginosa.

[0016] Figure 6 This is the ratio feature extracted by the peak ratio algorithm. The upper figure is the characteristic peak ratio matrix of Chroococcus, and the lower figure is the characteristic peak ratio matrix of Microcystis aeruginosa.

[0017] Figure 7 This is the confusion matrix of the cyanobacteria classification results of different algorithms in the laboratory scenario.

[0018] Figure 8 This is the confusion matrix of the cyanobacteria classification results of the algorithm in the real fresh water environment. Detailed implementation manners

[0019] The following further elaborates the present invention in conjunction with the accompanying drawings and embodiments.

[0020] As Figure 1 shown, a method for discriminating cyanobacteria species based on three-dimensional fluorescence spectroscopy includes the following steps: S0. Obtain the sample to be detected; Select a suitable sampling tool, such as a water pump, a deep water sampler or an automatic water sampler, etc., to collect the algal liquid sample, and filter the suspended substances and fine particles in the water through a vacuum filtration device and a 0.45 μm water system filter membrane to obtain the sample to be detected; S1. Obtain the original three-dimensional fluorescence spectrum data; Take 2 mL of the sample to be detected and place it in a quartz cuvette of a three-dimensional fluorescence spectrometer to measure the three-dimensional fluorescence spectrum of the solution; S2. Pretreat the original three-dimensional fluorescence spectrum data; Remove Rayleigh scattering in the original three-dimensional fluorescence spectrum using Delaunay triangle interpolation method, remove Raman scattering in the original three-dimensional fluorescence spectrum using blank background solvent subtraction method, and smooth and denoise the original three-dimensional fluorescence spectrum using Savitzky-Golay polynomial surface smoothing method; S3. Extract spectral features based on parallel factor-peak ratio method Extract the fluorescence component features of the three-dimensional fluorescence spectrum preprocessed in step S2 based on parallel factor analysis method, extract the fluorescence peak ratio features of the three-dimensional fluorescence spectrum preprocessed in step S2 based on peak ratio method, and combine the fluorescence component features and fluorescence peak ratio features to obtain a spectral feature vector; S4. Perform classification using LightGBM multi-classifier Perform missing value supplementation and data standardization processing on the obtained spectral feature vector, input the processed feature vector into the trained LightGBM multi-classifier, the classification model outputs the probability of the current sample to be tested corresponding to each existing category, and select the category with the highest probability as the cyanobacteria category to which the current sample to be tested belongs.

[0021] As Figure 2 shown, the said S3. includes the following steps: S3.1 Comprehensively determine the number of factors F of the parallel factor analysis method based on kernel consistency diagnostic method and residual analysis method; S3.2 Decompose the three-dimensional fluorescence spectrum preprocessed in S2 based on parallel factor analysis method to obtain a concentration score matrix A, an excitation loading matrix B, and an emission loading matrix C, where the excitation loading matrix and the emission loading matrix are used to characterize the fluorescence component features of the three-dimensional fluorescence spectrum; S3.3 Find the emission wavelength ex F and excitation wavelength em F corresponding to the peak points of each factor matrix B and matrix C, so as to determine that the wavelengths of the F peak points corresponding to the F factors are S3.4 Search for the f (f = 1,..., F) -th peak point in the three-dimensional fluorescence spectrum S3.5 Search for the strongest fluorescence peak in a small range near and update it as the new peak point S3.6 Take the N×N area centered on as the characteristic peak area of the f -th factor in the three-dimensional fluorescence spectrum of this sample; S3.7 Repeat steps S3.4~S3.6 to obtain F characteristic peak areas, that is, F N×N matrices, and these matrices are the fluorescence peak areas corresponding to the three-dimensional fluorescence spectrum; S3.8 Calculate the ratio of each pair of the F peak regions. The matrix ratio here is matrix point division, that is, the corresponding elements in the matrix are divided by each other, so as to obtain an N×N ratio matrix, which is the fluorescence pigment ratio feature of the sample obtained by the peak ratio algorithm; S3.9 Straighten the fluorescence component features of the three-dimensional fluorescence spectrum extracted in S3.2 and the fluorescence peak ratio features of the three-dimensional fluorescence spectrum extracted in S3.8 row by row, and combine them end to end to obtain the feature vector of the three-dimensional fluorescence spectrum of the sample to be measured.

[0022] The described S4. includes the following steps: S4.1 Use the median filling method to process the missing data in the feature vector obtained in S3, that is, calculate the median of each column and fill the missing value with this value. In addition, standardize the data to ensure the comparability of the combined features; S4.2 Calculate the weighted sum of the prediction results of all decision trees in the LightGBM multi-classifier; S4.3 Convert the output of S4.2 into a probability distribution through the Softmax function, and select the category with the highest probability as the category to which the current cyanobacteria sample belongs.

[0023] The training process of the LightGBM multi-classifier in the described S4. includes the following steps: S5.1 Obtain training samples, including different types of cyanobacteria; S5.2 Obtain the original three-dimensional fluorescence spectrum data of the training samples; S5.3 Preprocess the original three-dimensional fluorescence spectrum data of the training samples; S5.4 Extract the spectral features of the preprocessed three-dimensional fluorescence spectrum of the training samples based on the parallel factor-peak ratio method; S5.5 Perform missing value supplementation and data standardization processing on the obtained spectral feature vector, input the processed feature vector into the LightGBM multi-classifier for training, set hyperparameters, including learning rate, tree depth, number of leaf nodes, etc., and construct multiple decision trees to ensure that the model can fully learn the spectral features of different types of cyanobacteria, and obtain a trained LightGBM multi-classifier. Example

[0024] Use 7 kinds of cyanobacteria solutions cultured in the laboratory and 3 kinds of cyanobacteria solutions in the real fresh water environment as samples for experimental verification. The specific experimental steps are as follows:

[0025] (1) Cyanobacteria culture. Purchase cyanobacteria strains. All experimental algae are cultured according to the standard of "Chemicals - Algal Growth Inhibition Test" (GB / T 21805 - 2008). After one week, take the cyanobacteria solution for subculture again. The ratio of algal liquid to culture medium is 1:15, and the culture period is 14 days.

[0026] (2) Dataset division. For 7 species of cyanobacteria with a 14 - day culture period, 5 concentration gradients of samples are prepared for each species of cyanobacteria every day, and finally a total of 490 samples are obtained (laboratory scenario). Take 20% of the samples as the test set, and the remaining 80% of the samples as the training set. For 3 species of cyanobacteria, 4 different freshwater waters, and 6 algal density ratios, a total of 72 samples are finally obtained (real freshwater scenario). Take 25% of the samples as the test set, and the remaining 75% of the samples as the training set.

[0027] (3) Detection. The fluorescence measurement instrument uses Aqualog from Horiba as the detection device. The excitation wavelength setting range is 240nm - 800nm, the wavelength interval is 5nm, the emission wavelength setting range is 243.544nm - 823.84nm, the wavelength interval is about 2.33nm, the CCD gain is Medium, the integration time is 0.1s, and the sensitivity S / N is 20000:1. The sample is measured in a quartz detection cuvette with an optical path of 1cm. The dimension of the single fluorescence spectral matrix EEM obtained during the measurement is 250*113. Through the detection software of the Aqualog instrument, the emission part is automatically interpolated to 250nm - 800nm with an interval of 2.5nm, that is, the size of the single three - dimensional fluorescence spectral map becomes 221*113.

[0028] (4) Raman scattering removal. Use the background subtraction method to perform Raman scattering removal on the three - dimensional fluorescence spectral data of all cyanobacteria samples.

[0029] (5) Rayleigh scattering removal. Select the region where the emission wavelength is equal to or twice the excitation wavelength plus or minus 20nm as the Rayleigh scattering band of the sample, and use the Delaunay triangular interpolation method to perform Rayleigh scattering removal on the three - dimensional fluorescence spectral data of all cyanobacteria samples.

[0030] (6) Smoothing and denoising. Use the Savitzky - Golay polynomial surface smoothing method to smooth and denoise the three - dimensional fluorescence spectrum after subtracting Raman scattering and Rayleigh scattering, and obtain the pre - processed three - dimensional fluorescence spectrum.

[0031] Figure 3 Schematic diagrams of the three - dimensional fluorescence spectra of Oscillatoria, Nostoc, Aphanizomenon flos - aquae, Chroococcus, Microcystis aeruginosa, Phormidium, and Anabaena.

[0032] Figure 4A 、 Figure 4BThe analysis results of PARAFAC. The left figure is the emission loading matrix, the middle figure is the excitation loading matrix, and the right figure is the component fluorescence spectrum. Here, the combination of kernel consistency diagnosis method and residual analysis method is used to determine that the number of factors of PARAFAC is 8. Then, based on the number of 8 factors, PARAFAC is carried out to obtain the emission loading matrix and excitation loading matrix of each factor.

[0033] Figure 5 The characteristic peak regions obtained by searching. The upper figure is the characteristic peak region of Chroococcus, and the lower figure is the characteristic peak region of Microcystis aeruginosa. Based on the PARAFAC results, the peak points of the excitation loading matrix and the emission loading matrix corresponding to each factor can be obtained. The combination of these peak points is the peak point coordinates of the corresponding component of this factor in the three-dimensional fluorescence spectrum. Search for the real peak points in a small range near the peak point coordinates in the three-dimensional fluorescence spectrum of the sample, update them as new peak points, and take the 7×7 region around the peak points to obtain the characteristic peak regions.

[0034] Figure 6 The ratio features extracted by the peak ratio algorithm. The upper figure is the characteristic peak ratio matrix of Chroococcus, and the lower figure is the characteristic peak ratio matrix of Microcystis aeruginosa. That is Figure Five The obtained 8 peak ratio regions are divided pairwise to obtain 28 7×7 ratio matrices.

[0035] Figure 7 The confusion matrix of the cyanobacteria classification results in the laboratory scenario of different feature extraction algorithms. When the classifier is selected as LightGBM, only comparing the feature extraction algorithms, it can be seen that the proposed PARAFAC-peak ratio algorithm adds the feature of fluorescence peak ratio on the basis of the trilinear decomposition algorithm, which can more comprehensively describe the fluorescence differences between different species of cyanobacteria. The final classification accuracy is 93.9%, the weighted average of precision is 94.1%, the weighted average of recall is 93.9%, and the weighted average of F1 value is 93.%, showing better performance compared with the other four algorithms. Comparing the classification algorithms, it can be seen that the LightGBM classifier shows better classification accuracy under different feature extraction algorithms, which mainly benefits from the improvement of LightGBM in data processing and feature selection. First, LightGBM adopts a more efficient way to find split points, enabling the model to more accurately capture the effective information of the input fluorescence features. In addition, LightGBM also performs excellently in dealing with high-dimensional data and sparse features, and can maintain a high classification accuracy in complex scenarios.

[0036] Figure 8It is the confusion matrix of the classification results of cyanobacteria in the real fresh - water environment. It can be seen that the overall classification accuracy is 94.4%, the weighted average of precision is 95.2%, the weighted average of recall is 94.4%, and the weighted average of F1 - value is 94.4%. The classification effect is good, and only one Anabaena sample is misjudged as Chroococcus. The sample data here includes four fresh - water environments with different water qualities. Thus, it can be shown that when the model is applied to the identification of cyanobacteria in the real fresh - water environment, the model still has a good classification effect, and for fresh - water bodies with different water qualities, the classification and identification effects are all excellent.

[0037] Therefore, the fluorescence spectral feature extraction method proposed in the present invention combines the parallel factor and the peak - ratio algorithm. The parallel factor is used to extract the dimension - reduced information characterizing the fluorescence components in the cyanobacteria spectrum, and the peak - ratio is used to extract the information on the relative content of different fluorescence components in the cyanobacteria spectrum. The combination of the two can clearly and comprehensively describe the fluorescence specificity of different cyanobacteria. It provides efficient and reliable technical support for the monitoring and management of cyanobacterial blooms, promotes the development of environmental monitoring technology towards the direction of intelligence and refinement to a certain extent, and also provides an important scientific basis for water - body ecological protection.

[0038] The technical features of the above - described embodiments can be further combined. For the sake of brevity of description, not all possible combinations of the technical features in the above - described embodiments are described. However, as long as the combination of these technical features does not conflict, it should be considered as the scope recorded in this specification.

[0039] The above - described embodiments only express several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A method for distinguishing cyanobacteria species based on three-dimensional fluorescence spectroscopy, characterized in that: The following steps are involved: S0. Obtain the sample to be tested; Collect algae liquid samples, filter them through a vacuum filtration device and a 0.45 μm water filter membrane to obtain samples to be tested; S1. Obtaining original three-dimensional fluorescence spectrum data; Take 2 mL of the sample to be tested and place it in the quartz cuvette of the three-dimensional fluorescence spectrometer to measure the three-dimensional fluorescence spectrum of the solution; S2. Preprocessing the original three-dimensional fluorescence spectrum data; The Delaunay triangle interpolation method was used to remove Rayleigh scattering from the original three-dimensional fluorescence spectrum, the blank background solvent subtraction method was used to remove Raman scattering from the original three-dimensional fluorescence spectrum, and the Savitzky-Golay polynomial surface smoothing method was used to smooth and denoise the original three-dimensional fluorescence spectrum. S3. Extraction of spectral features based on parallel factor-peak ratio method Extracting the fluorescence component features of the three-dimensional fluorescence spectrum preprocessed in step S2 based on the parallel factor analysis method, extracting the fluorescence peak ratio features of the three-dimensional fluorescence spectrum preprocessed in step S2 based on the peak ratio method, combining the fluorescence component features and the fluorescence peak ratio features to obtain a spectral feature vector; S4. Classification with LightGBM multi-classifier The obtained spectral feature vector is supplemented with missing values ​​and data is normalized. The processed feature vector is input into the trained LightGBM multi-classifier. The classification model outputs the probability of the current sample to be tested corresponding to each existing category, and the category with the largest probability is selected as the cyanobacteria category to which the current sample to be tested belongs.

2. The method according to claim 1, characterized in that In the step S2., the three-dimensional fluorescence spectrum after removing Rayleigh scattering and Raman scattering is subjected to Savitzky-Golay smoothing, a polynomial is used for fitting, the fitting coefficient is obtained by the least squares method, and then the smoothing value of each point in the smoothing window is calculated to obtain the smoothed three-dimensional fluorescence spectrum data.

3. The method according to claim 1, characterized in that The S3. comprises the following steps: S3.1 Comprehensive determination of the number of factors in the parallel factor analysis method based on the kernel consistency diagnostic method and residual analysis method F ; S3.2 decomposes the three-dimensional fluorescence spectrum preprocessed by S2 based on parallel factor analysis to obtain the concentration score matrix A , excitation load matrix B and the launch load matrix C , where the excitation load matrix and the emission load matrix are used to characterize the fluorescence component characteristics of the three-dimensional fluorescence spectrum; S3.3 Find each factor matrix B and matrix C The emission wavelength corresponding to the peak point ex F and excitation wavelength em F , thereby determining F The corresponding factor F The wavelength of the peak point is ; S3.4 Searching for the first f ( f = 1, …, F ) peak points ; S3.5 Search for the strongest point of the fluorescence peak in a small area nearby and update it as the new peak point ; S3.6 As the center, take N × N The area is taken as the first f The characteristic peak area of ​​each factor; S3.7 Repeat steps S3.4 to S3.6 to obtain F characteristic peak area, that is, F indivual N × N Matrices, these matrices are the fluorescence peak areas corresponding to the three-dimensional fluorescence spectra; S3.8 pairs F The peak regions are compared with each other. The matrix ratio here is the matrix point division, that is, the corresponding elements in the matrix are divided, so as to obtain indivual N × N The ratio matrix is ​​the fluorescent pigment ratio feature obtained by the peak ratio algorithm for this sample; S3.9 straightens the fluorescence component features of the three-dimensional fluorescence spectrum extracted by S3.2 and the fluorescence peak ratio features of the three-dimensional fluorescence spectrum extracted by S3.8 by rows, connects them end to end and combines them to obtain the feature vector of the three-dimensional fluorescence spectrum of the sample to be tested.

4. The method according to claim 1, characterized in that: The S4. comprises the following steps: S4.1 uses the median filling method to process the missing data in the feature vector obtained in S3, that is, calculate the median of each column and fill the missing value with this value; in addition, the data is standardized to ensure the comparability of the combined features; S4.2 calculates the weighted sum of all decision tree prediction results in the LightGBM multi-classifier; S4.3 converts the output of S4.2 into a probability distribution through the Softmax function, and selects the category with the largest probability as the category to which the current cyanobacteria sample belongs.

5. The method according to claim 1, characterized in that The training process of the LightGBM multi-classifier in S4. includes the following steps: S5.1 obtain training samples, including different types of cyanobacteria; S5.2 obtains original three-dimensional fluorescence spectrum data of training samples; S5.3 preprocessing the original three-dimensional fluorescence spectrum data of the training samples; S5.4 extracts the spectral features of the three-dimensional fluorescence spectra of the training samples after preprocessing based on the parallel factor-peak ratio method; S5.5 fills in missing values ​​and performs data standardization on the obtained spectral feature vector, inputs the processed feature vector into the LightGBM multi-classifier for training, sets hyperparameters, including learning rate, tree depth, number of leaf nodes, and constructs multiple decision trees to ensure that the model can fully learn the spectral characteristics of different categories of cyanobacteria and obtain a trained LightGBM multi-classifier.

Citation Information

Cited By

  • Vertical blue-green algae population identification method and system based on multi-modal sensing and intelligent reconstruction

    CN122413335A

  • A differentiated antibiotic fluorescence fingerprint integrated classification and identification method based on a deep residual attention network

    CN122758168A