Water quality detection method and system based on combination of machine learning and spectrophotometry
The water sample absorbance is measured by spectrophotometer and the data set is mapped using kernel function technology, and the training sample set is dynamically adjusted in combination with active learning strategies. The problem of insufficient extraction of deep-level features and model generalization ability in the existing technology is solved, and high-precision and stable water quality detection prediction is achieved.
Patent Information
- Application Number
- CN202411936779.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The lack of effective means in the prior art to extract deep-level features, resulting in insufficient prediction accuracy of water quality detection. The sample set used by most machine learning models during training is static, and the active learning strategy can not fully utilize the active learning strategy to dynamically adjust the training sample set, limiting the generalization ability and prediction accuracy of the model.
The absorbance of the water sample is measured at multiple preset wavelengths by a spectrophotometer to form a multi-dimensional absorbance data set, and the kernel function technology is used to map it to the high-dimensional space to obtain a data set with enhanced nonlinear characteristics. The known water samples and their corresponding pollutant concentrations are extracted from historical water quality data, a training sample set is generated, and a machine learning model is used to predict pollutant concentrations on the data set, and the training sample set is dynamically adjusted using an active learning strategy.
The prediction accuracy of water quality detection and the generalization ability of the model are improved, more accurate and stable prediction of pollutant concentrations are achieved, and adaptability to future unknown samples is enhanced.
Smart Images

Figure CN120028268A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of water quality detection, and in particular to a water quality detection method based on machine learning combined with spectrophotometry. Background Art
[0002] With the rapid development of industrialization and urbanization, water pollution has become increasingly serious, posing a huge threat to human health and the ecological environment. Traditional water quality testing methods mainly rely on laboratory analysis, such as chemical titration, chromatography, etc. Although these methods are accurate, they are time-consuming, costly, and difficult to achieve large-scale real-time monitoring. In recent years, spectrophotometry technology has been widely used in water quality testing due to its rapid, simple and non-destructive characteristics. However, the complexity and nonlinear characteristics of spectrophotometric data pose challenges to the accurate prediction of pollutant concentrations;
[0003] In the existing technology, although the spectrophotometer can efficiently obtain the absorbance data of water samples, when processing multidimensional absorbance data sets, traditional methods often lack effective means to extract deep-level features, resulting in insufficient prediction accuracy. In addition, the sample sets used in the training process of most machine learning models are usually static, and they fail to make full use of active learning strategies to dynamically adjust the training sample sets, which limits the generalization ability and prediction accuracy of the model. Summary of the invention
[0004] The embodiment of the present invention provides a water quality detection and system based on machine learning combined with spectrophotometry, which is used to solve the problem that there is a lack of effective means to extract deep-level features in the prior art, resulting in insufficient prediction accuracy and the inability to fully utilize active learning strategies to dynamically adjust training sample sets, which limits the generalization ability of the model and the prediction accuracy.
[0005] In the first aspect, an embodiment of the present invention provides a water quality detection method based on machine learning combined with spectrophotometry, including: measuring the absorbance of a water sample at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set; using the multidimensional absorbance data set to map to a high-dimensional space by kernel function technology to obtain a data set with enhanced nonlinear characteristics; extracting known water samples and pollutant concentrations corresponding to the known water samples from historical water quality data to generate a training sample set; applying a machine learning model to predict the pollutant concentration of the data set with enhanced nonlinear characteristics, and dynamically adjusting the training sample set using an active learning strategy to obtain a pollutant concentration prediction value; using the pollutant concentration prediction value, combined with water quality safety standards, to evaluate the safety level of the water sample, and output a water quality report containing pollutant types, concentration levels, and compliance with safe drinking standards.
[0006] In a second aspect, the present application provides a water quality detection system based on machine learning combined with spectrophotometry, including:
[0007] A measurement module, used for measuring the absorbance of the water sample at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set;
[0008] A mapping module, used to utilize the multidimensional absorbance data set to map to a high-dimensional space through a kernel function technique to obtain a data set with enhanced nonlinear characteristics;
[0009] An extraction module is used to extract known water samples and pollutant concentrations corresponding to the known water samples from historical water quality data to generate a training sample set;
[0010] A prediction module, used to apply a machine learning model to predict the pollutant concentration of the data set with enhanced nonlinear characteristics, and to dynamically adjust the training sample set using an active learning strategy to obtain a predicted value of the pollutant concentration;
[0011] The evaluation module is used to use the predicted value of the pollutant concentration in combination with the water quality safety standard to evaluate the safety level of the water sample and output a water quality report containing the pollutant type, concentration level and compliance with the safe drinking standard.
[0012] In a third aspect, an embodiment of the present invention provides a computing device, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute a water quality detection method based on machine learning combined with spectrophotometry as described in any one of the first aspects.
[0013] In a fourth aspect, an embodiment of the present invention provides a computer storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement a water quality detection method based on machine learning combined with spectrophotometry as described in any one of the first aspects.
[0014] In an embodiment of the present invention, the absorbance of a water sample is measured at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set;
[0015] Utilizing the multidimensional absorbance data set, mapping it to a high-dimensional space through a kernel function technique, thereby obtaining a data set with enhanced nonlinear characteristics;
[0016] Applying a machine learning model to predict pollutant concentrations on the data set with enhanced nonlinear characteristics, and using an active learning strategy to dynamically adjust the training sample set to generate predicted pollutant concentrations;
[0017] Using the predicted value of the pollutant concentration and combining it with the water quality safety standard, the water sample is evaluated for its safety level, and a water quality report containing the pollutant type, concentration level, and compliance with the safe drinking standard is output. The technical solution provided by the present invention measures the absorbance of the water sample at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set, which provides a rich information basis and helps to more accurately identify and quantify pollutants. The multidimensional absorbance data set is mapped to a high-dimensional space using kernel function technology, which enhances the nonlinear feature expression ability of the data set. This processing method enables complex and nonlinear relationships to be better presented, improves the prediction accuracy of subsequent machine learning models, extracts known water samples and their corresponding pollutant concentrations from historical water quality data, generates an initial training sample set, and uses an active learning strategy to dynamically adjust the training sample set in the subsequent process. It not only makes full use of existing data resources, but also ensures the generalization ability and real-time adaptability of the model by continuously introducing new samples to optimize the model, and realizes an integrated process from data collection to result interpretation, providing a scientific basis and technical support for water quality management decisions;
[0018] Furthermore, based on the deviation between the initial pollutant concentration estimate and the actual measured value, the uncertainty area of the model prediction is determined. This process can accurately locate the weak links in the model prediction and provide a clear direction for targeted optimization.
[0019] By using the uncertainty region, the best water samples are selected to be added to the training sample set, thereby optimizing the machine learning model. The introduction of active learning strategy enables the model to obtain more valuable information in key areas, effectively improving the prediction accuracy and stability of the model;
[0020] The pollutant concentration of the multidimensional absorbance data set is predicted again through the optimized machine learning model (i.e., the enhanced machine learning model), and finally a more accurate pollutant concentration prediction value is obtained. This method not only improves the accuracy of the prediction, but also enhances the model's generalization ability for future unknown samples.
[0021] These and other aspects of the present invention will become more apparent from the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] Figure 1A flow chart of a water quality detection method based on machine learning combined with spectrophotometry provided in an embodiment of the present invention;
[0024] Figure 2 A schematic diagram of the structure of a water quality detection system based on machine learning combined with spectrophotometry provided in an embodiment of the present invention;
[0025] Figure 3 A schematic diagram of the structure of a computing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0027] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0029] In the prior art, although spectrophotometers can efficiently obtain absorbance data of water samples, traditional methods often lack effective means to extract deep features when processing multidimensional absorbance data sets, resulting in insufficient prediction accuracy. In addition, most machine learning models use static sample sets during training, and fail to make full use of active learning strategies to dynamically adjust training sample sets, which limits the generalization ability and prediction accuracy of the model. Based on this, the present invention provides a water quality detection method based on machine learning combined with spectrophotometry, such as Figure 1 ,include:
[0030] Step 101: measuring the absorbance of a water sample at multiple preset wavelengths using a spectrophotometer to form a multidimensional absorbance data set;
[0031] In this step, a spectrophotometer refers to an instrument used to measure the degree of absorption of light of different wavelengths by a substance. It can work at multiple preset wavelengths to provide information about various components in a water sample. Absorbance refers to the degree to which light is absorbed when passing through a water sample. It is an important parameter for measuring the concentration of dissolved or suspended substances in water. A multidimensional absorbance data set refers to a collection of absorbance values at multiple wavelengths, which can reflect the presence of different pollutants in a water sample and their relative concentrations. The absorbance of a water sample is measured by a spectrophotometer over a range of selected wavelengths, with one absorbance value for each wavelength. All measured absorbance values are collected to form a multidimensional data set that contains the response characteristics of the water sample at different wavelengths.
[0032] Step 102: using the multidimensional absorbance data set, mapping it to a high-dimensional space through a kernel function technique to obtain a data set with enhanced nonlinear features;
[0033] In this step, kernel function technology refers to a mathematical tool used to convert raw data from a low-dimensional space to a higher-dimensional space, making the originally complex nonlinear relationship easier to handle; high-dimensional space refers to a space with a higher dimension than the space where the original data is located, in which the relationship between data points can be more clearly expressed; a data set with enhanced nonlinear features means that after kernel function mapping, the nonlinear features that were originally difficult to capture in the data set are better expressed, which helps to improve the accuracy of subsequent analysis; using kernel function technology, such as radial basis function (RBF) kernel or polynomial kernel, the multidimensional absorbance data set is mapped to a higher-dimensional space, in which the complex nonlinear relationship between data points becomes more obvious, thereby forming a data set with enhanced nonlinear features.
[0034] Step 103: extracting known water samples and pollutant concentrations corresponding to the known water samples from historical water quality data to generate a training sample set;
[0035] In this step, historical water quality data refers to the water quality detection records accumulated in the past, including water sample information collected at different times and locations and their corresponding pollutant concentrations; known water samples refer to those water samples whose pollutant concentrations have been accurately measured and confirmed; training sample set refers to the data set used to train machine learning models, usually containing input features (such as absorbance values) and target outputs (such as pollutant concentrations); select qualified historical water quality data from the database, especially those verified water samples and their corresponding pollutant concentrations, and organize these data into a structured table or file format to form a training sample set to prepare for subsequent machine learning model training.
[0036] Step 104: applying a machine learning model to predict pollutant concentrations on the data set with enhanced nonlinear characteristics, and using an active learning strategy to dynamically adjust the training sample set to obtain a predicted value of pollutant concentrations;
[0037] In this step, the machine learning model refers to a mathematical model constructed by an algorithm, which can predict the output result (such as pollutant concentration) based on the input data (such as absorbance value). Commonly used models include support vector machines, random forests, etc.; active learning strategy refers to an iterative machine learning method. The model will select the most valuable new samples to add to the training set based on the uncertainty of the current prediction results to optimize the model performance; the pollutant concentration prediction value refers to the concentration of a specific pollutant in the water sample predicted by the machine learning model based on the input data; first, the pollutant concentration is predicted for the data set with enhanced nonlinear features using the preliminarily constructed machine learning model, and then the deviation between the predicted results and the actual measured values is evaluated to identify areas with large prediction errors. Then, according to the active learning strategy, representative new water samples are selected for actual measurement, and these new samples and their measurement results are added to the training sample set. Repeat this process until the prediction accuracy of the model reaches the expected standard, and finally the pollutant concentration prediction value is obtained.
[0038] Step 105: Using the predicted value of the pollutant concentration and combining it with the water quality safety standard, the water sample is evaluated for safety level, and a water quality report including the pollutant type, concentration level and compliance with the safe drinking standard is output;
[0039] In this step, water quality safety standards refer to a series of regulations formulated by national or international organizations on the maximum allowable concentrations of pollutants in drinking water and other water bodies for other purposes; safety level assessment refers to determining the safety level of water samples, such as safe, warning, dangerous, etc., based on the predicted pollutant concentrations and corresponding safety standards; water quality report refers to the final output result document, which contains information such as pollutant type, concentration level, and whether it meets safe drinking standards; the safety level of each pollutant is evaluated based on the pollutant concentration prediction value provided by the machine learning model and the applicable water quality safety standards. The overall safety status of the water sample is determined by comprehensively considering the evaluation results of all pollutants. Finally, a detailed water quality report is generated, which not only lists the specific concentrations of various pollutants, but also clearly indicates whether the water sample meets the safe drinking standards, providing a basis for relevant decision-making;
[0040] In a typical water quality testing scenario, imagine an environmental monitoring station dedicated to improving water quality testing capabilities to ensure that local residents have access to safe and reliable drinking water. The monitoring station decides to use an innovative approach - water quality testing technology based on machine learning combined with spectrophotometry, to achieve this goal;
[0041] First, technicians use a high-precision spectrophotometer to measure the absorbance of the collected water samples at multiple preset wavelengths. These wavelengths are carefully selected based on the characteristics of common pollutants to ensure that as much information as possible can be captured. Through a series of measurements, each water sample generates a multidimensional absorbance data set. This data set contains absorbance values at different wavelengths, providing a rich information basis for subsequent analysis;
[0042] Next, in order to better process these complex data, the kernel function technology is used to map the multidimensional absorbance data set to a higher-dimensional space. This process enhances the ability to express nonlinear features in the data set, making the originally difficult-to-capture relationships more obvious and easier to handle. After this transformation, the resulting data set not only retains all the characteristics of the original data, but also reveals more patterns hidden behind the data, providing strong support for subsequent prediction models.
[0043] Then, in order to train the machine learning model, known water samples and their corresponding pollutant concentrations were extracted from the historical water quality database to construct the initial training sample set. These historical data came from water quality monitoring records accumulated over the past years, covering a variety of water bodies and pollutants. In this way, it was ensured that the training sample set was both representative and diverse enough, thus providing a solid foundation for the machine learning model.
[0044] With the training sample set, we began to apply the machine learning model to predict pollutant concentrations on the data set with enhanced nonlinear features. In the initial training stage, the machine learning model gradually adjusted its own parameters to minimize the prediction error by learning from the training sample set. However, we did not stop there, but introduced an active learning strategy to dynamically optimize the training sample set. Specifically, every time the model makes a prediction, it evaluates the deviation between the predicted result and the actual measurement value, and identifies areas with large prediction errors. For these areas with high uncertainty, new water samples are selected for actual measurement, and these new samples and their measurement results are added to the training sample set to further optimize the model performance. By repeating this process, we finally obtained a highly accurate and generalizable enhanced machine learning model.
[0045] Finally, the enhanced machine learning model was used to predict the pollutant concentrations of the new multidimensional absorbance dataset, and detailed pollutant concentration prediction values were obtained. Subsequently, these predictions were compared with national and international water quality safety standards, and a safety level assessment was conducted on each water sample. Based on the assessment results, a corresponding safety level (such as safe, warning, or dangerous) was assigned to each water sample, and a complete water quality report was generated. This report not only lists the specific concentrations of various pollutants, but also clearly indicates whether the water sample meets the safe drinking standards, providing important decision-making basis for relevant departments and the public;
[0046] Through the above steps, rapid, efficient and accurate water quality testing was successfully achieved, significantly improving the capacity and level of water quality management. This method not only saves time and cost, but also makes important contributions to safeguarding public health and environmental protection.
[0047] Based on this, the present invention provides a specific embodiment, wherein the step 104 applies a machine learning model to predict the pollutant concentration of the data set with enhanced nonlinear characteristics, and uses an active learning strategy to dynamically adjust the training sample set to obtain a predicted value of the pollutant concentration, specifically including the following steps:
[0048] Step 201: Preliminarily train the machine learning model according to the data set with enhanced nonlinear features to obtain an initial machine learning model;
[0049] In this step, the data set with enhanced nonlinear features refers to the data set with enhanced nonlinear feature expression ability after being mapped to high-dimensional space by kernel function technology; the machine learning model refers to the mathematical model constructed by the algorithm, which can predict the output result (such as pollutant concentration) based on the input data (such as absorbance value); preliminary training refers to the first training of the model using part of the data to preliminarily adjust the model parameters to ensure that it has a certain predictive ability; the initial machine learning model refers to the model that has basic predictive ability but has not been fully optimized after preliminary training; the machine learning model is preliminarily trained using the data set with enhanced nonlinear features. This process involves selecting appropriate algorithms and parameter settings, and gradually adjusting the model parameters through iterative optimization. Through preliminary training, the machine learning model can preliminarily capture the patterns and relationships in the data to form an initial machine learning model.
[0050] Step 202: using the initial machine learning model to predict the pollutant concentration of the multidimensional absorbance data set to obtain an initial pollutant concentration estimate;
[0051] In this step, the initial machine learning model refers to a model that has basic prediction capabilities but has not been fully optimized after preliminary training; the multidimensional absorbance data set refers to a collection of absorbance values at multiple wavelengths, which can reflect the presence and relative concentrations of different pollutants in water samples; pollutant concentration prediction refers to the use of machine learning models to predict the concentration of specific pollutants in water samples based on input data (such as absorbance values); the initial pollutant concentration estimate refers to the preliminary prediction result given by the initial machine learning model, which may contain certain errors; the initial machine learning model is applied to the multidimensional absorbance data set to predict pollutant concentrations. During this process, the model evaluates new data points based on previously learned patterns and outputs initial pollutant concentration estimates for each water sample. These estimates provide a reference for subsequent optimization.
[0052] Step 203: Determine the uncertainty region of the model prediction based on the deviation between the initial pollutant concentration estimate and the actual measured value;
[0053] In this step, the initial pollutant concentration estimate refers to the preliminary prediction result given by the initial machine learning model, which may contain certain errors; the actual measured value refers to the actual pollutant concentration value measured by laboratory or other precise methods; the deviation refers to the difference between the predicted value and the actual value, which is used to measure the accuracy of the model prediction; the uncertainty area refers to the area with large prediction errors, indicating that the model's prediction in these areas is not accurate enough; compare the deviation between the initial pollutant concentration estimate and the actual measured value, and identify the areas with large prediction errors. These areas are marked as uncertainty areas, which means that there is a large uncertainty in the model's prediction here, which helps to locate key areas that need improvement
[0054] Step 204: using the uncertainty region, adopting an active learning strategy to select the best water sample to add to the training sample set to optimize the machine learning model, and obtaining a reinforced machine learning model;
[0055] In this step, the uncertainty area refers to the area with large prediction errors, indicating that the model's predictions in these areas are not accurate enough; the active learning strategy refers to an iterative machine learning method, in which the model selects the most valuable new samples to add to the training set based on the uncertainty of the current prediction results to optimize the model performance; the best water samples refer to those new samples that can best reveal the uncertainty of the model's predictions, which help improve the generalization ability of the model; the enhanced machine learning model refers to a model whose prediction accuracy and stability are significantly improved after optimization; using the information of the uncertainty area, an active learning strategy is adopted to select the best water samples for actual measurement, and these new samples and their measurement results are added to the training sample set. In this way, valuable data is continuously introduced, the training sample set is dynamically adjusted, and finally the machine learning model is optimized to form a more accurate and stable enhanced machine learning model.
[0056] Step 205: using the enhanced machine learning model to predict the pollutant concentration of the multidimensional absorbance data set again to obtain a predicted pollutant concentration value;
[0057] In this step, the enhanced machine learning model refers to a model that has been optimized and has significantly improved prediction accuracy and stability; the multidimensional absorbance data set refers to a collection of absorbance values at multiple wavelengths, which can reflect the presence and relative concentrations of different pollutants in water samples; the pollutant concentration prediction value refers to the final prediction result given by the enhanced machine learning model, which is usually more accurate and reliable; the enhanced machine learning model is used to predict the pollutant concentration of the multidimensional absorbance data set again. This prediction is based on the optimized model and can provide more accurate and reliable pollutant concentration prediction values. These prediction values not only improve the accuracy of detection, but also enhance the generalization ability for future unknown samples.
[0058] Based on this, the present invention provides a specific embodiment, wherein step 201, preliminarily training a machine learning model according to a data set with enhanced nonlinear features to obtain an initial machine learning model, specifically comprises the following steps:
[0059] Step 301: constructing a basic framework of a machine learning model based on the data set with enhanced nonlinear features and combined with historical water quality data;
[0060] In this step, the dataset with enhanced nonlinear features refers to the dataset with enhanced nonlinear feature expression ability after being mapped to high-dimensional space by kernel function technology; historical water quality data refers to the water quality detection records accumulated in the past, including water sample information collected at different times and locations and their corresponding pollutant concentrations.
[0061] The basic framework of the machine learning model refers to the preliminary mathematical model structure for subsequent training and optimization. The process of building the basic framework of the machine learning model based on the data set with enhanced nonlinear characteristics and combined with historical water quality data involves selecting appropriate algorithms and parameter settings to ensure that the basic framework can adapt to the data characteristics and provide a good starting point for subsequent optimization.
[0062] Step 302: Perform a stability test on the basic framework using a cross-validation method to obtain a test result;
[0063] In this step, the cross-validation method refers to a technology for evaluating model performance. The data set is divided into multiple subsets and used as validation sets in turn to evaluate the stability and generalization ability of the model. The test results refer to various indicators of model performance obtained through cross-validation, such as accuracy, recall rate, etc. The basic framework is tested for stability using the cross-validation method. The data set is divided into multiple subsets and used as training sets and validation sets in turn to evaluate the performance of the model under different data partitions. Finally, a set of test results are obtained, which reflect the stability and generalization ability of the model.
[0064] Step 303: Based on the test results, the hyperparameters in the basic framework are optimized and adjusted to obtain an optimized and adjusted basic framework;
[0065] In this step, hyperparameters refer to parameters in the model that need to be set in advance and are not automatically learned during the training process, such as learning rate, regularization coefficient, etc.; optimization and adjustment refers to adjusting hyperparameters according to test results to improve model performance; based on the test results, the hyperparameters in the basic framework are optimized and adjusted, and the best hyperparameter combination is found through grid search or random search methods to ensure that the model can maintain high prediction accuracy and stability under different data partitions, and finally obtain the optimized and adjusted basic framework.
[0066] Step 304: using the optimized and adjusted basic framework, comprehensively training the data set with enhanced nonlinear features to obtain a preliminarily trained machine learning model;
[0067] In this step, comprehensive training refers to fully training the model using the entire data set to adjust the model's internal parameters to make it better fit the data; the initially trained machine learning model refers to a model that has a certain predictive ability but has not been fully optimized after comprehensive training; using the optimized and adjusted basic framework, the data set with enhanced nonlinear features is comprehensively trained. In this process, the model continuously adjusts its internal parameters to minimize the prediction error. Through comprehensive training, a initially trained machine learning model is obtained, which can accurately predict pollutant concentrations to a certain extent.
[0068] Step 305: using the preliminarily trained machine learning model to predict the multidimensional absorbance data set to obtain a prediction result;
[0069] In this step, the multidimensional absorbance dataset refers to a collection of absorbance values at multiple wavelengths, which can reflect the presence of different pollutants in water samples and their relative concentrations; the prediction result refers to the estimated value of the pollutant concentration in the water sample output by the machine learning model; the multidimensional absorbance dataset is predicted using the preliminarily trained machine learning model, which evaluates new data points based on previously learned patterns and outputs an estimated value of the initial pollutant concentration for each water sample. These prediction results provide an important reference for further optimizing the model.
[0070] Step 306: using the comparison between the prediction result and the actual measurement value, evaluating the initial performance of the preliminarily trained machine learning model, obtaining a performance evaluation result, and adjusting the parameters of the preliminarily trained machine learning model based on the performance evaluation result to construct an initial machine learning model;
[0071] In this step, the prediction result refers to the estimated value of the pollutant concentration in the water sample output by the machine learning model; the actual measurement value refers to the actual pollutant concentration value measured by the laboratory or other precise methods.
[0072] Initial performance evaluation refers to evaluating the performance of the model after initial training by comparing the predicted results with the actual measured values; performance evaluation results refer to various indicators derived from the comparison, such as mean square error (MSE), mean absolute error (MAE) and coefficient of determination (R 2 Score), which is used to measure the prediction accuracy and stability of the model; adjusting model parameters refers to optimizing the internal parameters of the model according to the performance evaluation results to improve the performance of the model; the initial machine learning model refers to the model with good prediction ability and stability after parameter adjustment; using the comparison of prediction results with actual measured values to evaluate the initial performance of the preliminarily trained machine learning model. This process involves calculating the deviation between the predicted value and the true value, and generating detailed performance evaluation indicators such as mean square error, mean absolute error and determination coefficient. Through these indicators, we can fully understand the prediction accuracy and stability of the model;
[0073] Based on the performance evaluation results, adjust the parameters of the initially trained machine learning model. This step may include fine-tuning hyperparameters, introducing regularization terms to reduce the risk of overfitting, etc. This method aims to improve the generalization ability of the model and ensure its prediction accuracy under different water quality conditions;
[0074] Finally, through a series of adjustments and optimizations, an initial machine learning model was built. This model not only performed well on the existing data, but also provided a solid foundation for subsequent further optimization.
[0075] Based on this, the present invention provides a specific embodiment, wherein step 306 uses the comparison between the prediction result and the actual measurement value to evaluate the initial performance of the preliminarily trained machine learning model to obtain a performance evaluation result, and based on the performance evaluation result, adjusts the parameters of the preliminarily trained machine learning model to construct an initial machine learning model, specifically comprising the following steps:
[0076] Step 401: Compare the prediction results with the actual measured values to calculate the prediction error indicators, including mean square error, mean absolute error and R 2 Score, get the performance evaluation index;
[0077] In this step, the prediction result refers to the estimated value of the pollutant concentration in the water sample output by the machine learning model; the actual measurement value refers to the actual pollutant concentration value measured by the laboratory or other precise methods.
[0078] Prediction error indicators refer to various statistics used to measure the prediction accuracy of the model, such as mean square error (MSE), mean absolute error (MAE) and coefficient of determination (R 2 score); performance evaluation indicators refer to specific values derived from the prediction error indicators, which are used to comprehensively evaluate the performance of the model; prediction error indicators including mean square error, mean absolute error and R are calculated by comparing the prediction results with the actual measured values. 2 Scores, these indicators can quantify the accuracy and stability of model predictions. Through this process, detailed performance evaluation indicators are obtained, providing a basis for subsequent optimization.
[0079] Step 402: Based on the performance evaluation index, identify the hyperparameters that have the greatest impact on the performance through sensitivity analysis, and obtain a list of key hyperparameters;
[0080] In this step, sensitivity analysis refers to an analytical method used to determine which hyperparameters have the greatest impact on model performance; hyperparameters refer to parameters in the model that need to be pre-set and are not automatically learned during training, such as learning rate, regularization coefficient, etc.; the key hyperparameter list refers to the set of hyperparameters that have a significant impact on model performance identified after sensitivity analysis; based on performance evaluation indicators, sensitivity analysis is used to identify the hyperparameters that have the greatest impact on performance. This process involves changing the values of each hyperparameter and observing how these changes affect the performance of the model, and finally obtaining a key hyperparameter list that lists those hyperparameters that have a significant impact on model performance;
[0081] In order to identify the hyperparameters that have the greatest impact on model performance, a sensitivity analysis method based on Gaussian Process Regression (GPR) can be used, where the expression of the sensitivity analysis method is as follows:
[0082]
[0083] Among them, S i refers to the sensitivity of the i-th hyperparameter. Sensitivity measures the degree of influence of each hyperparameter on the model performance. By calculating the partial derivative, the impact of each hyperparameter change on the model output can be quantified; f(θ) refers to the performance evaluation index of the machine learning model (such as mean square error, mean absolute error or determination coefficient). The performance evaluation index is derived from the actual performance of the model during training and verification, usually including mean square error (MSE), mean absolute error (MAE) or determination coefficient (R 2 ), these indicators directly reflect the prediction accuracy and stability of the model; θ iRefers to the i-th hyperparameter. Hyperparameters refer to those parameters that need to be set in advance and are not automatically learned during the training process, such as learning rate, regularization coefficient, etc. These parameters directly affect the training process and final performance of the model; α is a weight parameter used to balance the influence of the first-order derivative and the second-order derivative. The weight parameter α is used to balance the influence of the first-order derivative and the second-order derivative. The first-order derivative reflects the direct impact of the hyperparameter on the model performance, while the second-order derivative captures the change in the rate of change of the hyperparameter, which helps to more fully understand the impact of the hyperparameter; It represents the expected value of the hyperparameter distribution p(θ), introduces the expected value calculation in Gaussian process regression (GPR), takes into account the distribution characteristics of the hyperparameters, and makes the sensitivity analysis more robust; Refers to the performance evaluation index about θ i The second-order partial derivative provides information about the rate of change of hyperparameters, helping to identify those hyperparameters that have a significant impact on model performance but have complex nonlinear relationships;
[0084] The overall formula is designed to improve the accuracy and applicability of sensitivity analysis by introducing advanced mathematical tools such as Gaussian process regression and Bayesian optimization. These methods can not only more accurately identify the hyperparameters that have the greatest impact on model performance, but also dynamically adjust the selection of key hyperparameters to adapt to the characteristics and needs of different data sets.
[0085] Step 403: Based on the key hyperparameter list, a Bayesian optimization algorithm is used to perform a global search on the key hyperparameters of the preliminarily trained machine learning model to find the optimal hyperparameter combination and obtain optimization suggestions;
[0086] In this step, the Bayesian optimization algorithm refers to an efficient optimization algorithm that guides the search process by building a probability model to find the optimal solution; the optimal hyperparameter combination refers to the best hyperparameter setting found within a given range to achieve the best model performance; the optimization suggestion refers to the improvement plan proposed based on the results of the Bayesian optimization algorithm; based on the key hyperparameter list, the Bayesian optimization algorithm is used to perform a global search for the key hyperparameters of the initially trained machine learning model, and the optimal hyperparameter combination is finally found by iteratively adjusting the hyperparameters and evaluating the model performance. This process generates optimization suggestions to guide subsequent model adjustments.
[0087] Step 404: adjusting the key hyperparameters of the initially trained machine learning model according to the optimization suggestions, and introducing regularization terms to obtain an optimized machine learning model;
[0088] In this step, the regularization term refers to a technique to prevent overfitting, which limits the model complexity by adding a penalty term; the optimized machine learning model refers to a model with improved performance after hyperparameter adjustment and regularization processing; the key hyperparameters of the initially trained machine learning model are adjusted according to the optimization suggestions, and appropriate regularization terms are introduced to reduce the risk of overfitting. In this way, an optimized machine learning model is obtained, and its predictive and generalization capabilities are significantly improved.
[0089] Step 405: using the optimized machine learning model to predict the multidimensional absorbance data set again to obtain the optimal prediction result;
[0090] In this step, the multidimensional absorbance dataset refers to a collection of absorbance values at multiple wavelengths, which can reflect the presence of different pollutants in water samples and their relative concentrations; the optimal prediction result refers to a more accurate estimate of pollutant concentration output by the optimized model; the multidimensional absorbance dataset is predicted again using the optimized machine learning model, which evaluates new data points based on previously learned patterns and outputs the optimal prediction results for each water sample, which are usually more accurate and reliable.
[0091] Step 406: Verify the performance of the optimized machine learning model by comparing the optimal prediction result with the corresponding actual measurement value, and obtain the final performance verification result;
[0092] In this step, the actual measured value refers to the real pollutant concentration value measured by the laboratory or other precise methods; the final performance verification result refers to the final evaluation index of the model performance obtained by comparing the optimal prediction result with the actual measurement value; the performance of the optimized machine learning model is verified by comparing the optimal prediction result with the corresponding actual measurement value. Through this process, the final performance verification results are obtained. These results reflect the performance of the model after optimization and further confirm the prediction accuracy and stability of the model.
[0093] Step 407: Based on the final performance verification result, adjust the optimized machine learning model parameters to construct a completed initial machine learning model;
[0094] In this step, the final performance verification result refers to the final evaluation index of the model performance obtained by comparing the optimal prediction result with the actual measurement value; adjusting the model parameters refers to further fine-tuning the internal parameters of the model according to the final performance verification result to ensure the best performance; the initial machine learning model refers to the model with high predictive ability and stability that is finally constructed after multiple optimizations and verifications; based on the final performance verification result, the optimized machine learning model parameters are fine-tuned to ensure the accuracy and reliability of the model in predicting all pollutant concentrations, and finally a complete initial machine learning model is constructed. This model not only performs well on existing data, but also has good generalization ability, providing a solid foundation for subsequent applications.
[0095] Based on this, the present invention provides a specific embodiment, in which step 401 uses the comparison between the prediction result and the actual measurement value to calculate the prediction error index, including the mean square error, the mean absolute error and R 2 Score, get the performance evaluation index, specifically including the following steps:
[0096] Step 501: Utilizing the comparison between the prediction result and the actual measurement value, the prediction error index is calculated by statistical method to obtain preliminary performance evaluation data, wherein the error index includes: mean square error, mean absolute error and determination coefficient;
[0097] In this step, statistical methods refer to methods used to quantify the difference between the predicted results and the actual measured values, such as calculating the mean square error (MSE), mean absolute error (MAE), and coefficient of determination (R 2 ); Preliminary performance evaluation data refers to the specific values obtained from the prediction error indicators, which are used to initially understand the performance of the model; by comparing the prediction results with the actual measured values, the prediction error indicators are calculated by statistical methods. These indicators include mean square error, mean absolute error and determination coefficient. Through this process, preliminary performance evaluation data are obtained, which provide a basis for subsequent analysis.
[0098] Step 502: Using the preliminary performance evaluation data, construct a performance evaluation model, and using the performance evaluation model, classify and analyze the prediction errors of different pollutant concentrations, identify the types of pollutants with higher errors, and obtain a list of high-error pollutants;
[0099] In this step, classification analysis refers to analyzing the prediction errors of different pollutant concentrations by category to identify the types of pollutants with larger errors; the high-error pollutant list refers to listing those types of pollutants whose prediction errors are significantly higher than other pollutants; using preliminary performance evaluation data, a performance evaluation model is constructed and used to classify and analyze the prediction errors of different pollutant concentrations. Through this analysis, the types of pollutants with higher errors are identified and a high-error pollutant list is generated.
[0100] Step 503: Analyze the causes of high prediction errors based on the high error pollutant list and the physical and chemical properties of the pollutants to obtain an error cause analysis report;
[0101] In this step, the physicochemical properties refer to the characteristics of the pollutant itself, such as solubility, molecular weight, etc., which may affect the prediction accuracy; the error cause analysis report refers to a document that records in detail the reasons for high prediction errors; based on the list of high error pollutants and combined with the physicochemical properties of the pollutants, the causes of high prediction errors are analyzed, and by comprehensively considering the characteristics of the pollutants and their behavior in the environment, the error cause analysis report is finally obtained.
[0102] Step 504: using the error cause analysis report, propose targeted improvement measures to obtain optimization suggestions, and use the optimization suggestions to perform targeted optimization on the preliminarily trained machine learning model to obtain an optimized machine learning model;
[0103] In this step, improvement measures refer to solutions proposed for the causes of errors, such as increasing the number of samples of specific pollutants or adjusting the model architecture; optimization suggestions refer to clear guidance plans formed based on improvement measures; optimized machine learning model refers to a model with improved performance after targeted optimization; using the error cause analysis report, targeted improvement measures are proposed to obtain optimization suggestions. Based on these suggestions, the initially trained machine learning model is targetedly optimized, and finally an optimized machine learning model is obtained, with improved prediction accuracy and stability.
[0104] Step 505: Utilizing the optimized machine learning model, the prediction error index is calculated again to obtain a target performance evaluation index;
[0105] In this step, the prediction error index refers to the various statistics used to measure the prediction accuracy of the model, such as mean square error, mean absolute error and determination coefficient; the target performance evaluation index refers to the prediction error index recalculated by the optimized model, which reflects the performance of the model after optimization; the prediction error index is calculated again using the optimized machine learning model. Through this process, the target performance evaluation index is obtained. These indicators demonstrate the prediction accuracy and stability of the model after optimization, providing an important reference for subsequent applications.
[0106] Based on this, the present invention provides a specific embodiment, wherein the step 102 uses the multidimensional absorbance data set to map to a high-dimensional space through a kernel function technique to obtain a data set with enhanced nonlinear features, specifically comprising the following steps:
[0107] Step 601: using the multidimensional absorbance data set, selecting a suitable kernel function, mapping the data to a high-dimensional space, and obtaining a preliminary mapping data set;
[0108] In this step, the multidimensional absorbance data set refers to a set of absorbance values at multiple wavelengths, which can reflect the presence and relative concentrations of different pollutants in water samples; the kernel function refers to a mathematical tool used to convert raw data from a low-dimensional space to a higher-dimensional space, making the originally complex nonlinear relationship easier to handle; the high-dimensional space refers to a space with a higher dimension than the space where the original data is located, in which the relationship between data points can be more clearly expressed; the preliminary mapping data set refers to a data set formed after kernel function mapping, which retains all the characteristics of the original data and enhances the expression of nonlinear features; using the multidimensional absorbance data set, select a suitable kernel function, such as radial basis function (RBF) or polynomial kernel, to map the data to a high-dimensional space. Through this process, a preliminary mapping data set is obtained. This data set provides a better basis for subsequent analysis.
[0109] Step 602: using the preliminary mapping data set, screening the features that contribute most to the prediction of pollutant concentration through a feature selection algorithm to obtain a selected feature set;
[0110] In this step, the preliminary mapping data set refers to the data set formed after kernel function mapping, which retains all the characteristics of the original data and enhances the expression ability of nonlinear features; the feature selection algorithm refers to the method used to identify and select the features that are most helpful to the model performance, such as recursive feature elimination (RFE), principal component analysis (PCA), etc.; the selected feature set refers to the feature set that is retained after screening by the feature selection algorithm and contributes most to the prediction of pollutant concentration; using the preliminary mapping data set, the feature selection algorithm is used to screen the features that contribute most to the prediction of pollutant concentration. Through this process, the selected feature set is obtained. The selected feature set not only reduces the data dimension, but also improves the efficiency and accuracy of model training.
[0111] Step 603: using the selected feature set and combining it with domain knowledge to perform feature semantic annotation to obtain an annotated feature set;
[0112] In this step, the selected feature set refers to the set of features that contribute most to the prediction of pollutant concentrations after being screened by the feature selection algorithm; domain knowledge refers to professional knowledge about a specific application field, such as the chemical and physical principles in water quality testing; feature semantic annotation refers to adding a label to each feature that describes its meaning and function to increase the interpretability of the feature; the annotated feature set refers to the set of features that have been semantically annotated, and these features have clear meanings and functions; the selected feature set is used in combination with domain knowledge for feature semantic annotation. Through this process, the annotated feature set is obtained. The annotated feature set not only helps to understand the specific meaning of each feature, but also improves the interpretability of the model results.
[0113] Step 604: construct a feature association network using the annotated feature set, analyze the interactions and dependencies between features, and obtain a feature association network graph;
[0114] In this step, the annotated feature set refers to the feature set that has been processed by semantic annotation, and these features have clear meanings and functions; the feature association network refers to a graphical representation that shows the interactions and dependencies between features; the feature association network diagram refers to a diagram generated by constructing a feature association network, which shows the complex relationships between various features; using the annotated feature set, a feature association network is constructed to analyze the interactions and dependencies between features. Through this process, a feature association network diagram is obtained. The feature association network diagram reveals the complex relationships between various features and provides a basis for further optimizing the model.
[0115] Step 605: Utilizing the feature association network diagram, optimizing the selection of kernel functions and the parameter settings of kernel functions, and obtaining an optimized high-dimensional space data set;
[0116] In this step, the feature association network diagram refers to a diagram generated by constructing a feature association network, which shows the complex relationship between each feature; the selection of the kernel function refers to selecting the most suitable kernel function type according to the feature association network diagram; the parameter setting of the kernel function refers to adjusting the parameters of the kernel function to optimize the model performance; the optimized high-dimensional space data set refers to a data set that is more suitable for machine learning model training after optimization processing; using the feature association network diagram, the selection of the kernel function and the parameter setting of the kernel function are optimized. Through this process, the optimized high-dimensional space data set is obtained. The optimized high-dimensional space data set not only enhances the expression ability of nonlinear features, but also improves the prediction accuracy of the model.
[0117] Step 606: using the optimized high-dimensional space data set, perform data dimensionality reduction processing to obtain a data set with enhanced nonlinear features;
[0118] In this step, the optimized high-dimensional space dataset refers to a dataset that is more suitable for machine learning model training after optimization processing; data dimensionality reduction processing refers to the process of reducing data dimensions while trying to maintain important information. Common methods include principal component analysis (PCA) and linear discriminant analysis (LDA), etc.; a dataset with enhanced nonlinear features refers to a dataset that still retains key nonlinear features after dimensionality reduction processing, which is convenient for subsequent model training and prediction; the optimized high-dimensional space dataset is used to perform data dimensionality reduction processing. Through this process, a dataset with enhanced nonlinear features is obtained. This dataset not only reduces the computational complexity, but also retains the most important nonlinear features, providing high-quality data support for the final model training.
[0119] Based on this, the present invention provides a specific embodiment, wherein the step 105 uses the predicted value of the pollutant concentration in combination with the water quality safety standard to evaluate the safety level of the water sample, and outputs a water quality report including the pollutant type, concentration level, and compliance with the safe drinking standard, specifically including the following steps:
[0120] Step 701: using the predicted pollutant concentration to compare with national and international water quality safety standards to determine the safety level of each pollutant;
[0121] In this step, the predicted pollutant concentration refers to the estimated value of the pollutant concentration in the water sample output by the machine learning model; national and international water quality safety standards refer to the regulations formulated by national or international organizations on the maximum allowable concentrations of pollutants in drinking water and other water bodies for other purposes; the safety level refers to the safety classification of each pollutant based on whether the pollutant concentration exceeds the safety standard; the process of using the predicted pollutant concentration to compare with the national and international water quality safety standards to determine the safety level of each pollutant involves comparing the predicted concentration value with the limit value specified in the standard, providing basic information for subsequent evaluation.
[0122] Step 702: Based on the safety level and in combination with the health impact factors of the pollutants, a comprehensive assessment model is constructed to assess the overall safety of the water sample and obtain a comprehensive assessment result;
[0123] In this step, the health impact factor refers to the potential impact of pollutants on human health, such as carcinogenicity, teratogenicity, etc.; the comprehensive assessment model refers to a mathematical model used to comprehensively consider multiple factors (such as pollutant concentration, health impact factors) to evaluate the overall safety of water samples; the comprehensive assessment result refers to the conclusion about the overall safety of water samples drawn through the comprehensive assessment model; based on the safety level and combined with the health impact factors of pollutants, a comprehensive assessment model is constructed. This model not only considers whether the pollutant concentration exceeds the standard, but also evaluates the potential impact of these pollutants on human health. By evaluating the overall safety of water samples, a comprehensive assessment result is finally obtained.
[0124] Step 703: using the comprehensive evaluation results, classify the water samples into three safety levels: safe, warning, and dangerous, and obtain the water sample safety level;
[0125] In this step, safety level classification refers to the process of dividing water samples into different safety levels based on the comprehensive evaluation results; water sample safety level refers to the final determined water sample safety level, which is divided into three levels: safe, warning and dangerous; using the comprehensive evaluation results, water samples are classified into three levels: safe, warning and dangerous. Through this process, the water sample safety level is obtained, which provides a clear basis for subsequent management and decision-making.
[0126] Step 704: Generate a water quality status map using the water sample safety level, combined with geographic location and time information;
[0127] In this step, the geographic location and time information refer to the collection location of the water sample and its corresponding sampling time; the water quality status map refers to a graphical representation of the water quality safety of different regions, which is convenient for intuitively understanding the water quality conditions in various regions; the water sample safety level is used in combination with the geographic location and time information to generate a water quality status map. The water quality safety level of different regions is marked on the map, showing the water quality conditions that change over time and space, providing an intuitive reference tool for the public and management departments.
[0128] Step 705: using the water quality status map, formulate targeted water quality improvement suggestions and obtain an improvement suggestion report, wherein the water quality improvement suggestions include: pollution source control and water quality purification measures;
[0129] In this step, water quality improvement suggestions refer to solutions proposed for water quality problems in specific areas, such as pollution source control and water purification measures; improvement suggestion report refers to a document that records water quality improvement suggestions in detail to provide guidance for actual operations; using the water quality status map, targeted water quality improvement suggestions are formulated to obtain an improvement suggestion report. The improvement suggestions include pollution source control and water purification measures. The report provides relevant departments with a specific action guide to improve water quality conditions.
[0130] Step 706: using the improvement suggestion report and combining it with the water quality monitoring data, updating the water quality status map and water quality improvement suggestions at a preset time, and obtaining an updated water quality status map and water quality improvement suggestions;
[0131] In this step, water quality monitoring data refers to water quality test data collected regularly, which is used to track changes in water quality; preset time refers to a set time interval, such as monthly or quarterly, which is used to regularly update relevant information; updated water quality status map and water quality improvement suggestions refer to maps and reports that reflect the latest water quality status and improvement suggestions after regular updates; using the improvement suggestion report, combined with the water quality monitoring data, the water quality status map and water quality improvement suggestions are updated at the preset time. Through this process, updated water quality status map and water quality improvement suggestions are obtained to ensure the timeliness and accuracy of the information and support continuous improvement.
[0132] Step 707: Output a complete water quality report using the updated water quality status map and water quality improvement suggestions, wherein the complete water quality report includes: pollutant type, concentration level, compliance with safe drinking standards, water quality safety level, and water quality improvement suggestions;
[0133] In this step, the complete water quality report refers to the final comprehensive water quality analysis document output, which includes pollutant types, concentration levels, compliance with safe drinking standards, water quality safety grades and water quality improvement suggestions; using the updated water quality status map and water quality improvement suggestions, a complete water quality report is output. This report not only includes pollutant types, concentration levels and compliance with safe drinking standards, but also lists in detail the water quality safety grades and water quality improvement suggestions, providing comprehensive information support for relevant decision-making.
[0134] Figure 2 A schematic diagram of the structure of a water quality detection system based on machine learning combined with spectrophotometry is provided for the embodiment of the present application. Figure 2 As shown, the system includes:
[0135] A measuring module 21 is used to measure the absorbance of the water sample at multiple preset wavelengths by a spectrophotometer to form a multi-dimensional absorbance data set;
[0136] A mapping module 22, used to utilize the multidimensional absorbance data set to map to a high-dimensional space through a kernel function technique to obtain a data set with enhanced nonlinear characteristics;
[0137] An extraction module 23 is used to extract known water samples and pollutant concentrations corresponding to the known water samples from historical water quality data to generate a training sample set;
[0138] A prediction module 24, for applying a machine learning model to predict the pollutant concentration of the data set with enhanced nonlinear characteristics, and dynamically adjusting the training sample set using an active learning strategy to obtain a predicted value of the pollutant concentration;
[0139] The evaluation module 25 is used to use the predicted value of the pollutant concentration in combination with the water quality safety standard to evaluate the safety level of the water sample and output a water quality report including the pollutant type, concentration level and compliance with the safe drinking standard.
[0140] Figure 2 The water quality detection system based on machine learning combined with spectrophotometry can perform Figure 1 The implementation principle and technical effect of the xx method described in the embodiment shown will not be repeated. The specific way in which each module and unit performs operations in the water quality detection system based on machine learning combined with spectrophotometry in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.
[0141] Figure 2 The water quality detection system based on machine learning combined with spectrophotometry in the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;
[0142] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .
[0143] The processing component 32 is used for: right 1.
[0144] The processing component 32 includes one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be implemented by one or more application-specific integrated circuits (AICs), digital signal processors (DPs), digital signal processing devices (DPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0145] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0146] Computing devices also include other components, such as input / output interfaces, display components, and communication components.
[0147] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above-mentioned peripheral interface module can be an output device or an input device.
[0148] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0149] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0150] The embodiment of the present invention further provides a computer storage medium storing a computer program, which can achieve the above-mentioned Figure 1 An xx method and system of the illustrated embodiment.
[0151] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0152] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0153] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A water quality detection method based on machine learning combined with spectrophotometry, characterized in that: include: The absorbance of the water sample is measured at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set; Utilizing the multidimensional absorbance data set, mapping it to a high-dimensional space through a kernel function technique, thereby obtaining a data set with enhanced nonlinear characteristics; Extract known water samples and the pollutant concentrations corresponding to the known water samples from historical water quality data to generate a training sample set; Applying a machine learning model to predict pollutant concentrations on the data set with enhanced nonlinear characteristics, and using an active learning strategy to dynamically adjust the training sample set to obtain a predicted value of pollutant concentration; The predicted values of pollutant concentrations are used in combination with water quality safety standards to assess the safety level of water samples, and a water quality report is output that includes pollutant types, concentration levels, and compliance with safe drinking standards.
2. The method according to claim 1, characterized in that The machine learning model is applied to the data set with enhanced nonlinear characteristics to predict the pollutant concentration, and the active learning strategy is used to dynamically adjust the training sample set to obtain the predicted value of the pollutant concentration, including: Preliminarily training the machine learning model based on the data set with enhanced nonlinear features to obtain an initial machine learning model; Using the initial machine learning model to predict pollutant concentrations on a multidimensional absorbance data set to obtain an initial pollutant concentration estimate; Determining a region of uncertainty in model prediction based on deviations between the initial pollutant concentration estimates and actual measured values; Using the uncertainty region, an active learning strategy is adopted to select the best water sample to add to the training sample set to optimize the machine learning model, thereby obtaining a reinforced machine learning model; The enhanced machine learning model is used to predict the pollutant concentration of the multidimensional absorbance data set again to obtain the predicted value of the pollutant concentration.
3. The method according to claim 2, characterized in that The machine learning model is preliminarily trained based on the data set with enhanced nonlinear features to obtain an initial machine learning model, including: Based on the dataset with enhanced nonlinear features and combined with historical water quality data, the basic framework of the machine learning model is constructed; A cross-validation method is used to perform a stability test on the basic framework to obtain a test result; Based on the test results, the hyperparameters in the basic framework are optimized and adjusted to obtain an optimized and adjusted basic framework; Using the optimized and adjusted basic framework, a data set with enhanced nonlinear features is fully trained to obtain a preliminarily trained machine learning model; Using the preliminarily trained machine learning model, predicting the multidimensional absorbance data set to obtain a prediction result; By comparing the predicted results with the actual measured values, the initial performance of the preliminarily trained machine learning model is evaluated to obtain a performance evaluation result. Based on the performance evaluation result, the parameters of the preliminarily trained machine learning model are adjusted to construct an initial machine learning model.
4. The method according to claim 3, characterized in that: By comparing the prediction results with the actual measured values, the initial performance of the preliminarily trained machine learning model is evaluated to obtain a performance evaluation result, and based on the performance evaluation result, the parameters of the preliminarily trained machine learning model are adjusted to construct an initial machine learning model, including: By comparing the predicted results with the actual measured values, the prediction error indicators are calculated, including mean square error, mean absolute error and R 2 Score, get the performance evaluation index; Based on the performance evaluation indicators, identify the hyperparameters that have the greatest impact on performance through sensitivity analysis, and obtain a list of key hyperparameters; According to the key hyperparameter list, a Bayesian optimization algorithm is used to perform a global search on the key hyperparameters of the initially trained machine learning model to find the optimal hyperparameter combination and obtain optimization suggestions; According to the optimization suggestions, key hyperparameters of the initially trained machine learning model are adjusted, and regularization terms are introduced to obtain an optimized machine learning model; Using the optimized machine learning model, the multidimensional absorbance data set is predicted again to obtain the optimal prediction result; By comparing the optimal prediction results with the corresponding actual measured values, the performance of the optimized machine learning model is verified to obtain the final performance verification results; Based on the final performance verification results, the optimized machine learning model parameters are adjusted to construct a completed initial machine learning model.
5. The method according to claim 4, characterized in that By comparing the predicted results with the actual measured values, the prediction error indicators are calculated, including mean square error, mean absolute error and R 2 Score, get performance evaluation indicators, including: By comparing the prediction results with the actual measured values, the prediction error index is calculated by statistical methods to obtain preliminary performance evaluation data, wherein the error index includes: mean square error, mean absolute error and determination coefficient; Using the preliminary performance evaluation data, a performance evaluation model is constructed, and using the performance evaluation model, prediction errors of different pollutant concentrations are classified and analyzed to identify types of pollutants with higher errors, and obtain a list of high-error pollutants; Analyze the causes of high prediction errors based on the high error pollutant list and the physical and chemical properties of the pollutants to obtain an error cause analysis report; Using the error cause analysis report, propose targeted improvement measures to obtain optimization suggestions, and use the optimization suggestions to perform targeted optimization on the initially trained machine learning model to obtain an optimized machine learning model; Using the optimized machine learning model, the prediction error index is calculated again to obtain the target performance evaluation index.
6. The method according to claim 1, characterized in that The multidimensional absorbance data set is used to map to a high-dimensional space through a kernel function technique to obtain a data set with enhanced nonlinear features, including: Using the multidimensional absorbance data set, selecting a suitable kernel function, mapping the data into a high-dimensional space, and obtaining a preliminary mapping data set; Using the preliminary mapping data set, the features that contribute most to the prediction of pollutant concentration are screened by a feature selection algorithm to obtain a selected feature set; Using the selected feature set and combining it with domain knowledge, feature semantic annotation is performed to obtain an annotated feature set; Using the annotated feature set, constructing a feature association network, analyzing the interactions and dependencies between features, and obtaining a feature association network graph; Utilizing the feature association network diagram, optimizing the selection of kernel functions and the parameter settings of the kernel functions, and obtaining an optimized high-dimensional space data set; The optimized high-dimensional space data set is used to perform data dimensionality reduction processing to obtain a data set with enhanced nonlinear features.
7. The method according to claim 1, characterized in that Using the predicted values of pollutant concentrations and in combination with water quality safety standards, the water sample is evaluated for safety level, and a water quality report containing pollutant types, concentration levels, and compliance with safe drinking standards is output, including: Using the predicted pollutant concentrations to compare with national and international water quality safety standards, determine the safety level of each pollutant; Based on the safety level and in combination with the health impact factors of pollutants, a comprehensive assessment model is constructed to assess the overall safety of the water sample and obtain a comprehensive assessment result; Using the comprehensive evaluation results, the water samples are classified into three safety levels: safe, warning, and dangerous, to obtain the water sample safety level; Generate a water quality status map using the water sample safety level combined with geographic location and time information; Utilizing the water quality status map, formulate targeted water quality improvement suggestions and obtain an improvement suggestion report, wherein the water quality improvement suggestions include: pollution source control and water quality purification measures; Using the improvement suggestion report, combined with water quality monitoring data, the water quality status map and water quality improvement suggestions are updated at a preset time to obtain an updated water quality status map and water quality improvement suggestions; The updated water quality status map and water quality improvement suggestions are used to output a complete water quality report, which includes: pollutant types, concentration levels, compliance with safe drinking standards, water quality safety levels, and water quality improvement suggestions.
8. A water quality detection system based on machine learning combined with spectrophotometry, characterized in that: include: A measurement module, used for measuring the absorbance of the water sample at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set; A mapping module, used to utilize the multidimensional absorbance data set to map to a high-dimensional space through a kernel function technique to obtain a data set with enhanced nonlinear characteristics; An extraction module is used to extract known water samples and pollutant concentrations corresponding to the known water samples from historical water quality data to generate a training sample set; A prediction module, used to apply a machine learning model to predict the pollutant concentration of the data set with enhanced nonlinear characteristics, and to dynamically adjust the training sample set using an active learning strategy to obtain a predicted value of the pollutant concentration; The evaluation module is used to use the predicted value of the pollutant concentration in combination with the water quality safety standard to evaluate the safety level of the water sample and output a water quality report containing the pollutant type, concentration level and compliance with the safe drinking standard.
9. A computing device, characterized in that It comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a water quality detection method based on machine learning combined with spectrophotometry as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, a water quality detection method based on machine learning combined with spectrophotometry as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Nonlinear full-spectrum water turbidity quantitative analysis method based on extreme random tree
CN110887798A
River and lake water environment data fusion and sample labeling method and system based on machine learning and computer equipment
CN115146720A
Food safety detection system based on spectral analysis
CN118730953A
Method and device for training crack image detection model based on active learning
CN118918419A
Portable field microbial incubator
CN119040118A
Cited By
Soybean producing area soil micro-ecological risk assessment method and system based on combined pollution
CN120706889A
Sewage treatment water quality intelligent prediction method and system based on discharge characteristics
CN121030241A