Water environment phytoplankton dynamic monitoring system based on machine learning
Through the water environment phytoplankton dynamic monitoring system based on machine learning, the problem of low accuracy in traditional monitoring methods has been solved, efficient monitoring of the dynamic changes of phytoplankton has been achieved, and the monitoring accuracy and adaptability have been improved.
Patent Information
- Application Number
- CN202510882682.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-10
AI Technical Summary
The accuracy of monitoring phytoplankton through water quality sensors in existing technologies is low, and they cannot adapt to the dynamic changes of phytoplankton in the water environment. They require manual sampling, which leads to a decrease in monitoring accuracy.
A dynamic monitoring system for phytoplankton in water environments based on machine learning is adopted, including a preprocessing module, a machine learning model building module, a real-time update module and a system deployment module. The monitoring accuracy is improved through data collection, multimodal data fusion, data cleaning, feature extraction, machine learning model building and real-time updating.
The accuracy of phytoplankton monitoring has been significantly improved, making the monitoring data more in line with the actual situation and adapting to the dynamic changes of the water environment.
Smart Images

Figure CN120766014A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field, and in particular to a dynamic monitoring system for phytoplankton in an aquatic environment based on machine learning. Background Art
[0002] In the aquatic environment, phytoplankton plays a vital role in maintaining the ecological balance of the water body and biodiversity. However, phytoplankton is prone to over-reproduction, causing algal blooms, which seriously threaten water resources and fisheries, and even have adverse effects on human health. In order to prevent the above situation, it is necessary to monitor the phytoplankton in the aquatic environment. The current traditional method is to set up a variety of water quality sensors in the water environment to conduct simple data monitoring, and at the same time conduct regular manual sampling to observe the growth status of phytoplankton, thereby achieving the monitoring purpose.
[0003] However, in the aforementioned existing technologies, although water quality sensors can play a certain monitoring role, most of the data are complex and useless, and the accuracy is low. Therefore, manual sampling is required. Phytoplankton has dynamic characteristics in the water environment, and various parameters will change in real time. Manual sampling and setting sensors alone cannot adapt to dynamic situations, resulting in a significant decrease in the final monitoring accuracy. Summary of the Invention
[0004] The purpose of the present invention is to provide a dynamic monitoring system for phytoplankton in water environments based on machine learning, so as to solve the problem that although water quality sensors can play a certain monitoring role in the prior art, most of the data are complex and useless, and the accuracy is low, so manual sampling is required. Phytoplankton has dynamic characteristics in the water environment, and various parameters will change in real time. Manual sampling and setting of sensors alone cannot adapt to dynamic situations, resulting in a significant decrease in the final monitoring accuracy.
[0005] To achieve the above objectives, the present invention provides a dynamic monitoring system for phytoplankton in a water environment based on machine learning, comprising a preprocessing module, a machine learning model construction module, a real-time update module, and a system deployment module, wherein the preprocessing module, the machine learning model construction module, the real-time update module, and the system deployment module are connected in sequence;
[0006] The preprocessing module is used to collect and preprocess phytoplankton data;
[0007] The machine learning model building module is used to build a monitoring model based on preprocessed data;
[0008] The real-time update module is used to update the model in real time;
[0009] The system deployment module is used to deploy the monitoring system.
[0010] The preprocessing module includes a data acquisition unit, a multimodal data fusion unit, a data cleaning unit and a feature extraction unit, and the data acquisition unit, the multimodal data fusion unit, the data cleaning unit and the feature extraction unit are connected in sequence;
[0011] The data acquisition unit is used to set a camera to collect images, set a flow cytometer to collect data on phytoplankton in the water environment, and rely on multiple water quality sensors to obtain water environment data, thereby obtaining optical images, flow cytometric data, and environmental data respectively;
[0012] The multimodal data fusion unit is used to fuse the optical image, flow cytometry data and environmental data;
[0013] The data cleaning unit is used to clean the optical image, flow cytometry data and environmental data;
[0014] The feature extraction unit is used to extract features from optical images, flow cytometric data and environmental data.
[0015] Wherein, the multimodal data fusion unit includes a data synchronization subunit, a data calibration subunit and a data fusion unit, and the data synchronization subunit, the data calibration subunit and the data fusion unit are connected in sequence;
[0016] The data synchronization subunit is used to align the optical image, flow cytometry data and environmental data in time using a timestamp synchronization technology to eliminate time errors;
[0017] The data calibration subunit is used to regularly calibrate the hardware equipment for data acquisition;
[0018] The data fusion unit is used to fuse the optical image, flow cytometric data and environmental data by using a weighted average method.
[0019] Wherein, the data cleaning unit includes an image data cleaning subunit, a flow cytometry data cleaning subunit and an environmental parameter data cleaning subunit, and the image data cleaning subunit, the flow cytometry data cleaning subunit and the environmental parameter data cleaning subunit are connected in sequence;
[0020] The image data cleaning subunit is used to remove noise in the optical image using a median filter image processing technique, then enhance the image contrast using histogram equalization and contrast stretching, and finally remove interference objects in the image using morphological operations;
[0021] The flow cytometry data cleaning subunit is used to clean the data output by the flow cytometer to remove abnormal values and invalid numbers;
[0022] The environmental parameter data cleaning subunit is used to filter and calibrate the water quality sensor data to eliminate noise and drift.
[0023] Wherein, the feature extraction unit includes an image data extraction subunit, a flow cytometry data extraction subunit and an environmental parameter data extraction subunit, and the image data extraction subunit, the flow cytometry data extraction subunit and the environmental parameter data extraction subunit are connected in sequence;
[0024] The image data extraction subunit is used to extract the morphological characteristics, texture characteristics and color characteristics of phytoplankton from the cleaned image;
[0025] The flow cytometry data extraction subunit is used to extract the fluorescence intensity, forward scattered light intensity and side scattered light intensity of the flow cytometer data;
[0026] The environmental parameter data extraction subunit is used to extract statistics from the environmental data, where the statistics include mean, maximum, minimum and change rate.
[0027] Wherein, the machine learning model construction module includes a model selection unit and a model training unit, and the model selection unit and the model training unit are connected in sequence;
[0028] The model selection unit is used to select multiple machine learning models to predict phytoplankton and fuse the output results;
[0029] The model training unit is used to train the model.
[0030] The model selection unit includes a convolutional neural network subunit, a long short-term memory network subunit and a model fusion subunit, and the convolutional neural network subunit, the long short-term memory network subunit and the model fusion subunit are connected in sequence;
[0031] The convolutional neural network subunit is used to automatically learn features in the image using a convolutional neural network to accurately identify the types of phytoplankton;
[0032] The long short-term memory network subunit is used to use the long short-term memory network to capture long-term dependencies in time series data and accurately predict the dynamic changes of phytoplankton;
[0033] The model fusion subunit is used to fuse the outputs of the convolutional neural network and the long short-term memory network to obtain a fusion model, combining the image classification results and the biomass prediction results.
[0034] The model training unit includes a data set division subunit, a model training subunit, a model tuning subunit and a model evaluation subunit, and the data set division subunit, the model training subunit, the model tuning subunit and the model evaluation subunit are connected in sequence;
[0035] The data set division subunit is used to divide the extracted features into a training set, a validation set and a test set, the training set is used for model training, the validation set is used for model tuning, and the test set is used for model performance evaluation;
[0036] The model training subunit is used to train the convolutional neural network and the long short-term memory network using the training set, adjust the model parameters, and optimize the model performance;
[0037] The model tuning subunit is used to tune the convolutional neural network and the long short-term memory network using the validation set, and adjust the parameters of the learning rate, batch size and regularization term;
[0038] The model evaluation subunit is used to evaluate the tuned model using the test set, calculate the accuracy, recall, F1 value, root mean square error and mean absolute error, and evaluate the performance of the model.
[0039] The real-time update module includes an online learning unit, an incremental learning unit, and a learning result adjustment unit, and the online learning unit, the incremental learning unit, and the learning result adjustment unit are connected in sequence;
[0040] The online learning unit is used to use an online learning algorithm to update the parameters of the fusion model when new data arrives;
[0041] The incremental learning unit is used to design an incremental learning strategy to regularly integrate new data into the fusion model;
[0042] The learning result adjustment unit is used to evaluate the model performance after the model is updated, and perform version rollback if it is found that the model performance has degraded.
[0043] The system deployment module includes a hardware deployment unit, a software integration unit and a cloud platform support unit, and the hardware deployment unit, the software integration unit and the cloud platform support unit are connected in sequence;
[0044] The hardware deployment unit is used to deploy hardware equipment used for monitoring phytoplankton in the water environment;
[0045] The software integration unit is used to deploy software modules used for monitoring phytoplankton in water environments;
[0046] The cloud platform support unit is used to set up an upper cloud platform, store monitoring data and provide computing resources.
[0047] The present invention provides a water environment phytoplankton dynamic monitoring system based on machine learning, wherein the preprocessing module is used to collect and preprocess phytoplankton data; the machine learning model construction module is used to construct a monitoring model based on the preprocessed data; the real-time update module is used to update the model in real time; and the system deployment module is used to deploy the monitoring system. Thus, by collecting water environment data and phytoplankton data and constructing a prediction model based on machine learning, and relying on real-time updating of the model, the dynamic situation of phytoplankton can be accurately monitored, the monitoring accuracy can be significantly improved, and the monitoring data can be more in line with the actual situation. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.
[0049] Figure 1 It is a schematic diagram of the water environment phytoplankton dynamic monitoring system based on machine learning of the present invention.
[0050] Figure 2 It is a schematic diagram of the preprocessing module of the present invention.
[0051] Figure 3 It is a schematic diagram of the multimodal data fusion unit of the present invention.
[0052] Figure 4 It is a principle diagram of the data cleaning unit of the present invention.
[0053] Figure 5 It is a schematic diagram of the feature extraction unit of the present invention.
[0054] Figure 6 It is a schematic diagram of the machine learning model building module of the present invention.
[0055] Figure 7 It is a schematic diagram of the model selection unit of the present invention.
[0056] Figure 8 It is a schematic diagram of the model training unit of the present invention.
[0057] Figure 9 It is a principle diagram of the real-time update module of the present invention.
[0058] Figure 10 It is a schematic diagram of the system deployment module of the present invention.
[0059] 1-preprocessing module, 101-data acquisition unit, 102-multimodal data fusion unit, 1021-data synchronization subunit, 1022-data calibration subunit, 1023-data fusion unit, 103-data cleaning unit, 1031-image data cleaning subunit, 1032-flow cytometry data cleaning subunit, 1033-environmental parameter data cleaning subunit, 104-feature extraction unit, 1041-image data extraction subunit, 1042-flow cytometry data extraction subunit, 1043-environmental parameter data extraction subunit, 2-machine learning model construction module, 201-Model selection unit, 2011-Convolutional neural network subunit, 2012-Long short-term memory network subunit, 2013-Model fusion subunit, 202-Model training unit, 2021-Dataset partitioning subunit, 2022-Model training subunit, 2023-Model tuning subunit, 2024-Model evaluation subunit, 3-Real-time update module, 301-Online learning unit, 302-Incremental learning unit, 303-Learning result adjustment unit, 4-System deployment module, 401-Hardware deployment unit, 402-Software integration unit, 403-Cloud platform support unit. DETAILED DESCRIPTION
[0060] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but should not be understood as limiting the present invention.
[0061] See also Figures 1 to 10 The present invention provides a system for monitoring phytoplankton in aquatic environments based on machine learning, which specifically includes:
[0062] The preprocessing module 1 is used to collect and preprocess phytoplankton data;
[0063] Specifically include:
[0064] The multimodal data fusion unit 1023102 is used to fuse the optical image, flow cytometry data and environmental data;
[0065] Specifically include:
[0066] The data synchronization subunit 1021 is used to align the optical image, flow cytometry data and environmental data in time using a timestamp synchronization technology to eliminate time errors;
[0067] The data synchronization subunit 1021 ensures that data from different data sources remain consistent in the time dimension to facilitate subsequent fusion and analysis.
[0068] The data calibration subunit 1022 is used to regularly calibrate the hardware equipment used for data acquisition;
[0069] The data calibration subunit 1022 ensures the accuracy and stability of the data acquisition hardware equipment.
[0070] The data fusion unit 1023 is used to fuse the optical image, flow cytometry data and environmental data using a weighted average method.
[0071] Since optical images, flow cytometry data, and environmental data may come from different sensors and data sources, their dimensions, ranges, and distribution characteristics may vary. The purpose of data normalization is to eliminate these differences so that data from different data sources can be compared and fused on the same scale.
[0072] The weighted average method integrates information from different data sources, leveraging the strengths of each and compensating for the shortcomings of a single source. By adjusting weights, the contribution of different data sources to the fusion process can be controlled. For example, if a data source is of higher quality or contains more useful information, it can be assigned a higher weight. Conversely, if a data source contains more noise or uncertainty, it can be assigned a lower weight. In phytoplankton monitoring, optical images provide information on phytoplankton morphology and distribution, flow cytometric data provides information on cell size and fluorescence characteristics, and environmental data provides information on the physical and chemical properties of the water. Fusion of these data using the weighted average method can yield a more comprehensive and accurate description of phytoplankton status. Therefore, data fusion unit 1023 fuses data from different data sources using a specific algorithm to generate a more comprehensive and accurate description of the water environment and phytoplankton status. This principle is based on the weighted average method, assigning different weights to different data sources based on their reliability and importance, achieving optimized data fusion.
[0073] The data cleaning unit 103 is used to clean the optical image, flow cytometry data and environmental data;
[0074] Specifically include:
[0075] The image data cleaning subunit 1031 is used to remove noise in the optical image using a median filter image processing technique, then enhance the image contrast using histogram equalization and contrast stretching, and finally remove interference in the image using morphological operations;
[0076] Image data cleaning subunit 1031 processes optical image data, improving image quality through denoising, enhancement, and morphological operations. This process utilizes median filtering to remove image noise, histogram equalization and contrast stretching to enhance image contrast, and morphological operations to remove interference, making the image clearer and easier to analyze.
[0077] The flow cytometry data cleaning subunit 1032 is used to clean the data output by the flow cytometer to remove abnormal values and invalid numbers;
[0078] During the data acquisition process, flow cytometers may be affected by factors such as instrument noise, cell aggregation, debris interference, electronic component fluctuations, environmental interference, and operational errors, resulting in outliers and invalid values in the data. Data cleaning can remove this noise and interference, making the data more accurate and reliable, and providing a high-quality data foundation for subsequent analysis. Specifically, the Z-score method identifies outliers by calculating the degree of deviation of the data point from the mean. The Z-score indicates the number of standard deviations between the data point and the mean. Data points with an absolute Z-score value greater than 3 are generally considered outliers. These outliers are then deleted.
[0079] The environmental parameter data cleaning subunit 1033 is used to filter and calibrate the water quality sensor data to eliminate noise and drift.
[0080] Environmental parameter data cleaning subunit 1033 specifically processes water quality sensor data, eliminating noise and drift through filtering and calibration. This process uses filtering algorithms to remove high-frequency noise and calibration algorithms to eliminate sensor drift errors, ensuring data stability and accuracy.
[0081] The feature extraction unit 104 is used to extract features from the optical image, flow cytometry data and environmental data.
[0082] Specifically include:
[0083] The image data extraction subunit 1041 is used to extract the morphological characteristics, texture characteristics and color characteristics of phytoplankton from the cleaned image;
[0084] The extracted morphological, texture, and color features can be used as input for machine learning models to classify and identify phytoplankton. By training the classification model, automatic identification of phytoplankton species can be achieved, improving classification accuracy and efficiency. Furthermore, by regularly extracting morphological, texture, and color features of phytoplankton, dynamic changes in phytoplankton communities can be monitored. Changes in these features can reflect changes in phytoplankton species composition, abundance, and physiological state, providing timely and accurate information for water environment management.
[0085] The flow cytometry data extraction subunit 1042 is used to extract the fluorescence intensity, forward scattered light intensity and side scattered light intensity of the flow cytometer data;
[0086] Fluorescence intensity can reflect the metabolic activity or physiological state of cells. For example, highly active cells may produce stronger fluorescence signals. By using different fluorescent dyes, specific components within cells (such as DNA, RNA, proteins, etc.) can be labeled, and the fluorescence intensity can reflect the content or distribution of these components. The intensity of forward scattered light is closely related to the size of the cell and can therefore be used to estimate the size of the cell. This is of great significance for understanding the structure and dynamic changes of phytoplankton communities. The intensity of side scattered light can reflect the complexity of the internal structure of the cell, such as the size and shape of the cell nucleus and the amount and distribution of particulate matter within the cell.
[0087] The environmental parameter data extraction subunit 1043 is used to extract statistics from the environmental data. The statistics include mean, maximum, minimum and change rate.
[0088] Values can be used to assess the overall environmental quality of a water body. For example, by calculating the mean values of environmental parameters such as water temperature, pH, dissolved oxygen, and nutrient concentration over a period of time, we can understand the average environmental conditions of the water body and provide basic information for phytoplankton growth. Maximum values can be used to detect abnormal events in environmental parameters. For example, a sudden increase in the maximum value of water temperature or nutrient concentration may indicate abnormal events such as pollutant discharge or eutrophication, which may have a significant impact on the growth and distribution of phytoplankton. Minimum values can be used to identify environmental limiting factors affecting phytoplankton growth. For example, when the minimum value of dissolved oxygen concentration falls below a certain threshold, it may indicate hypoxia in the water body, which may pose a threat to the survival and reproduction of phytoplankton. Rates of change can be used to track the dynamic changes in environmental parameters. For example, by calculating the rate of change of environmental parameters such as water temperature, pH, or nutrient concentration, we can understand the changing trends of these parameters over time and provide dynamic information for phytoplankton monitoring.
[0089] The machine learning model building module 2 is used to build a monitoring model based on preprocessed data;
[0090] Specifically include:
[0091] The model selection unit 201 is used to select multiple machine learning models to predict phytoplankton and fuse the output results;
[0092] Specifically include:
[0093] The convolutional neural network subunit 2011 is used to automatically learn features in the image using a convolutional neural network to accurately identify the types of phytoplankton;
[0094] The Convolutional Neural Network Subunit 2011 uses convolutional neural networks (CNNs) to perform feature learning and classification recognition on optical images;
[0095] First, the convolution layer's learnable convolution kernel (filter) slides over the input image, performing a convolution operation to generate a feature map. Each convolution kernel can detect specific local patterns in the image, such as edges and textures. In phytoplankton species recognition, the convolution layer can automatically learn low-level features in the image, such as edges and contours. These features are the basis for distinguishing different types of phytoplankton.
[0096] Then, a pooling layer is used to reduce the dimensionality of the feature map, reducing the amount of computation while retaining important features. Common pooling operations include max pooling and average pooling. In phytoplankton species identification, the pooling layer can reduce the spatial size of the feature map and reduce the complexity of the model while retaining feature information useful for species identification.
[0097] Finally, the fully connected layer integrates the features extracted by the previous layers and outputs the classification result. Each neuron in the fully connected layer is connected to all neurons in the previous layer, undergoing linear transformations through weights and biases. Nonlinearity is then introduced through activation functions (such as ReLU and Sigmoid). In phytoplankton species recognition, the fully connected layer integrates the features extracted by the convolutional and pooling layers, and maps these features to different phytoplankton species using the weights and biases learned during training, enabling accurate species identification.
[0098] The long short-term memory network subunit 2012 is used to use the long short-term memory network to capture long-term dependencies in time series data and accurately predict the dynamic changes of phytoplankton;
[0099] The Long Short-Term Memory Network subunit 2012 utilizes a long short-term memory network (LSTM) to model and predict time series data. LSTM controls the flow of information through three gating mechanisms: input gate, forget gate, and output gate. The input gate determines how much of the current input information can enter the cell state; the forget gate determines which information in the cell state is forgotten; and the output gate determines which information in the cell state is output to the next time step. These gating mechanisms enable LSTM to selectively retain or forget historical information, thereby capturing long-term dependencies in time series data. In predicting phytoplankton dynamics, LSTM can remember past environmental parameters and phytoplankton population changes, providing important evidence for future predictions.
[0100] At the same time, it relies on the cell state in the long-short-term memory network to transmit long-term dependent information; it is updated through a gating mechanism to retain information useful for future predictions; in the prediction of dynamic changes in phytoplankton, the cell state can preserve the changing trends of historical environmental parameters and phytoplankton numbers, providing LSTM with long-term memory capabilities, enabling it to more accurately predict the dynamic changes of phytoplankton.
[0101] The model fusion subunit 2013 is used to fuse the outputs of the convolutional neural network and the long short-term memory network to produce a fused model that combines the image classification results and biomass prediction results. The model fusion subunit 2013 fuses the outputs of the CNN and LSTM to form a comprehensive fusion model. In feature-level fusion, the output features of the CNN and LSTM are merged at an earlier stage. This means that the image features extracted by the CNN and the time series features extracted by the LSTM are combined to form a richer feature representation. In decision-level fusion, the CNN and LSTM perform independent predictions and then merge their predictions. For example, the CNN may output a classification result for phytoplankton species, while the LSTM may output a biomass prediction. Finally, these two results are combined using a fusion strategy (such as weighted averaging or voting) to produce a comprehensive prediction result. By fusing the outputs of the CNN and LSTM, the model can simultaneously consider spatial information in the image and temporal information in the time series. CNNs have advantages in image feature extraction, while LSTMs excel at processing time series data. By fusing the outputs of the two, their complementary strengths can be fully utilized.
[0102] In decision-layer fusion, this can be achieved through weighted averaging, which assigns different weights to the outputs of CNN and LSTM, and then calculates the weighted average. The weights can be determined based on the performance of the model or domain knowledge.
[0103] In the fusion strategy, serial fusion can be used: the outputs of CNN and LSTM are serially connected in sequence to form a longer feature vector, which is then input into another fully connected layer for classification or regression tasks; or parallel fusion can be used: the outputs of CNN and LSTM are respectively input into two independent fully connected layers, and then the outputs of the two fully connected layers are fused (such as addition, splicing, etc.), and finally input into another fully connected layer for classification or regression tasks.
[0104] The model training unit 202 is used to train the model.
[0105] Specifically include:
[0106] The data set division subunit 2021 is used to divide the extracted features into a training set, a validation set and a test set, the training set is used for model training, the validation set is used for model tuning, and the test set is used for model performance evaluation;
[0107] The model training subunit 2022 is used to train the convolutional neural network and the long short-term memory network using the training set, adjust the model parameters, and optimize the model performance;
[0108] The model tuning subunit 2023 is used to tune the convolutional neural network and the long short-term memory network using the validation set, and adjust the parameters of the learning rate, batch size and regularization term;
[0109] The model evaluation subunit 2024 is used to evaluate the tuned model using the test set, calculate the accuracy, recall, F1 value, root mean square error and mean absolute error, and evaluate the performance of the model.
[0110] The real-time updating module 3 is used to update the model in real time;
[0111] Specifically include:
[0112] The online learning unit 301 is used to use an online learning algorithm to update the parameters of the fusion model when new data arrives;
[0113] The incremental learning unit 302 is used to design an incremental learning strategy to regularly integrate new data into the fusion model;
[0114] Online learning algorithms typically use stochastic gradient descent (SGD) or its variants (such as Adam, RMSProp, etc.) as optimization algorithms. When new data arrives, the algorithm calculates the prediction error under the current model parameters and updates the model parameters based on the error gradient to minimize the prediction error.
[0115] When new data arrives, the online learning algorithm will input the new data into the CNN and LSTM sub-models respectively, calculate the prediction error of the two sub-models, and update the parameters of the two sub-models based on the error gradient. At the same time, if the fusion model uses a structure such as feature layer fusion or parallel fusion, it is also necessary to calculate the prediction error of the fusion model based on the fused features or outputs and update the parameters of the fusion model;
[0116] Therefore, the online learning algorithm can process newly arrived data in real time and update model parameters, so that the model can adapt to environmental changes and maintain accurate predictions of dynamic changes in phytoplankton; it can achieve real-time updates of the model and improve the real-time performance of the model.
[0117] The learning result adjusting unit 303 is configured to evaluate the model performance after the model is updated, and if the model performance is found to be decreased, the version is rolled back.
[0118] Through the version rollback method, the model can be kept in the best performance state all the time, and thus the performance can be steadily enhanced through continuous iteration.
[0119] The system deployment module 4 is configured to deploy the monitoring system.
[0120] Specifically, the method comprises the following steps.
[0121] The hardware deployment unit 401 is configured to deploy the hardware device used for monitoring the phytoplankton in the water environment.
[0122] The software integration unit 402 is configured to deploy the software module used for monitoring the phytoplankton in the water environment.
[0123] The cloud platform supporting unit 403 is configured to set an upper cloud platform, store the monitoring data, and provide the computing resource.
[0124] The above only discloses one or more preferred embodiments of the present application, and cannot limit the scope of the rights of the present application. Those skilled in the art can understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made according to the claims of the present application still belong to the scope covered by the present application.
Claims
1. A dynamic monitoring system for phytoplankton in water environment based on machine learning, characterized in that: It includes a pre-processing module, a machine learning model construction module, a real-time update module and a system deployment module, wherein the pre-processing module, the machine learning model construction module, the real-time update module and the system deployment module are connected in sequence; The preprocessing module is used to collect and preprocess phytoplankton data; The machine learning model building module is used to build a monitoring model based on preprocessed data; The real-time update module is used to update the model in real time; The system deployment module is used to deploy the monitoring system.
2. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 1, characterized in that: The preprocessing module includes a data acquisition unit, a multimodal data fusion unit, a data cleaning unit and a feature extraction unit, wherein the data acquisition unit, the multimodal data fusion unit, the data cleaning unit and the feature extraction unit are connected in sequence; The data acquisition unit is used to set a camera to collect images, set a flow cytometer to collect data on phytoplankton in the water environment, and rely on multiple water quality sensors to obtain water environment data, thereby obtaining optical images, flow cytometric data, and environmental data respectively; The multimodal data fusion unit is used to fuse the optical image, flow cytometry data and environmental data; The data cleaning unit is used to clean the optical image, flow cytometry data and environmental data; The feature extraction unit is used to extract features from optical images, flow cytometric data and environmental data.
3. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 2, characterized in that: The multimodal data fusion unit includes a data synchronization subunit, a data calibration subunit and a data fusion unit, and the data synchronization subunit, the data calibration subunit and the data fusion unit are connected in sequence; The data synchronization subunit is used to align the optical image, flow cytometry data and environmental data in time using a timestamp synchronization technology to eliminate time errors; The data calibration subunit is used to regularly calibrate the hardware equipment for data acquisition; The data fusion unit is used to fuse the optical image, flow cytometric data and environmental data by using a weighted average method.
4. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 3, characterized in that: The data cleaning unit includes an image data cleaning subunit, a flow cytometry data cleaning subunit and an environmental parameter data cleaning subunit, wherein the image data cleaning subunit, the flow cytometry data cleaning subunit and the environmental parameter data cleaning subunit are connected in sequence; The image data cleaning subunit is used to remove noise in the optical image using a median filter image processing technique, then enhance the image contrast using histogram equalization and contrast stretching, and finally remove interference objects in the image using morphological operations; The flow cytometry data cleaning subunit is used to clean the data output by the flow cytometer to remove abnormal values and invalid numbers; The environmental parameter data cleaning subunit is used to filter and calibrate the water quality sensor data to eliminate noise and drift.
5. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 4, characterized in that: The feature extraction unit includes an image data extraction subunit, a flow cytometry data extraction subunit and an environmental parameter data extraction subunit, and the image data extraction subunit, the flow cytometry data extraction subunit and the environmental parameter data extraction subunit are connected in sequence; The image data extraction subunit is used to extract the morphological characteristics, texture characteristics and color characteristics of phytoplankton from the cleaned image; The flow cytometry data extraction subunit is used to extract the fluorescence intensity, forward scattered light intensity and side scattered light intensity of the flow cytometer data; The environmental parameter data extraction subunit is used to extract statistics from the environmental data, where the statistics include mean, maximum, minimum and change rate.
6. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 5, characterized in that: The machine learning model construction module includes a model selection unit and a model training unit, and the model selection unit and the model training unit are connected in sequence; The model selection unit is used to select multiple machine learning models to predict phytoplankton and fuse the output results; The model training unit is used to train the model.
7. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 6, characterized in that: The model selection unit includes a convolutional neural network subunit, a long short-term memory network subunit and a model fusion subunit, and the convolutional neural network subunit, the long short-term memory network subunit and the model fusion subunit are connected in sequence; The convolutional neural network subunit is used to automatically learn features in the image using a convolutional neural network to accurately identify the types of phytoplankton; The long short-term memory network subunit is used to use the long short-term memory network to capture long-term dependencies in time series data and accurately predict the dynamic changes of phytoplankton; The model fusion subunit is used to fuse the outputs of the convolutional neural network and the long short-term memory network to obtain a fusion model, combining the image classification results and the biomass prediction results.
8. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 7, characterized in that: The model training unit includes a data set division subunit, a model training subunit, a model tuning subunit and a model evaluation subunit, wherein the data set division subunit, the model training subunit, the model tuning subunit and the model evaluation subunit are connected in sequence; The data set division subunit is used to divide the extracted features into a training set, a validation set and a test set, the training set is used for model training, the validation set is used for model tuning, and the test set is used for model performance evaluation; The model training subunit is used to train the convolutional neural network and the long short-term memory network using the training set, adjust the model parameters, and optimize the model performance; The model tuning subunit is used to tune the convolutional neural network and the long short-term memory network using the validation set, and adjust the parameters of the learning rate, batch size and regularization term; The model evaluation subunit is used to evaluate the tuned model using the test set, calculate the accuracy, recall, F1 value, root mean square error and mean absolute error, and evaluate the performance of the model.
9. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 8, characterized in that: The real-time update module includes an online learning unit, an incremental learning unit and a learning result adjustment unit, wherein the online learning unit, the incremental learning unit and the learning result adjustment unit are connected in sequence; The online learning unit is used to use an online learning algorithm to update the parameters of the fusion model when new data arrives; The incremental learning unit is used to design an incremental learning strategy to regularly integrate new data into the fusion model; The learning result adjustment unit is used to evaluate the model performance after the model is updated, and perform version rollback if it is found that the model performance has degraded.
10. The system for dynamic monitoring of phytoplankton in water environment based on machine learning according to claim 9, characterized in that: The system deployment module includes a hardware deployment unit, a software integration unit and a cloud platform support unit, wherein the hardware deployment unit, the software integration unit and the cloud platform support unit are connected in sequence; The hardware deployment unit is used to deploy hardware equipment used for monitoring phytoplankton in the water environment; The software integration unit is used to deploy software modules used for monitoring phytoplankton in water environments; The cloud platform support unit is used to set up an upper cloud platform, store monitoring data and provide computing resources.