A water quality detection method and system based on machine learning combined with spectrophotometry

By measuring the absorbance of water samples with a spectrophotometer and mapping it to a high-dimensional space using kernel function technology, combined with machine learning and active learning strategies, the problem of insufficient prediction accuracy in water quality detection in existing technologies is solved, achieving efficient and accurate water quality detection and assessment.

CN120028268BActive Publication Date: 2026-01-20CSSC HAISHEN MEDICAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411936779.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2026-01-20
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

In existing technologies, spectrophotometry in water quality testing lacks effective means to extract deep features when processing multidimensional absorbance datasets, resulting in insufficient prediction accuracy. Furthermore, machine learning models fail to fully utilize active learning strategies to dynamically adjust the training sample set, limiting the model's generalization ability and prediction accuracy.

Method used

The absorbance of water samples is measured at multiple preset wavelengths using a spectrophotometer to form a multidimensional absorbance dataset. Kernel function technology is used to map the dataset to a high-dimensional space. Known water samples and their pollutant concentrations are extracted to generate a training sample set. Machine learning combined with an active learning strategy is used to dynamically adjust the training sample set to predict pollutant concentrations. Finally, the safety level is assessed in conjunction with water quality safety standards.

Benefits of technology

It improves the prediction accuracy and generalization ability of water quality testing, enabling rapid and accurate water quality testing, generating detailed pollutant concentration and safety level assessment reports, and supporting water quality management decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120028268B_ABST
    Figure CN120028268B_ABST
Patent Text Reader

Abstract

The application provides a water quality detection method based on machine learning combined with spectrophotometry, wherein, in the embodiment of the application, the absorbance of a water sample is measured at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set, the data set is mapped to a high-dimensional space through a kernel function technology to obtain a data set with enhanced nonlinear characteristics, known water samples and the pollutant concentrations corresponding to the known water samples are extracted from historical water quality data to generate a training sample set, a machine learning model is applied to predict the pollutant concentration of the data set with enhanced nonlinear characteristics, an active learning strategy is used to dynamically adjust the training sample set to obtain a pollutant concentration prediction value, the pollutant concentration prediction value is used in combination with water quality safety standards to evaluate the safety level of the water sample, and a water quality report is output, the technical solution provided by the application improves the efficiency and accuracy of water quality detection, and provides strong technical support for environmental protection and public health protection, and the entire process realizes integrated operation from data acquisition to result interpretation, ensuring the scientificity and reliability of water quality management decisions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the application relates to the technical field of water quality detection, and particularly relates to a water quality detection method based on machine learning and combining spectrophotometry. BACKGROUND

[0002] With the rapid development of industrialization and urbanization, water pollution problems are becoming increasingly serious, which poses a great threat to human health and ecological environment. Traditional water quality detection methods mainly rely on laboratory analysis, such as chemical titration, chromatographic analysis, etc. Although these methods are accurate, they are time-consuming, high-cost, and difficult to achieve large-scale real-time monitoring. In recent years, spectrophotometry has been widely used in water quality detection due to its rapid, simple and non-destructive characteristics. However, the complexity and non-linear characteristics of spectrophotometric data pose challenges to the accurate prediction of pollutant concentrations.

[0003] In the prior art, although spectrophotometers can efficiently obtain absorbance data of water samples, traditional methods often lack effective means to extract deep features when processing multi-dimensional absorbance data sets, resulting in insufficient prediction accuracy. In addition, most machine learning models use static sample sets in the training process, which fails to fully utilize the active learning strategy to dynamically adjust the training sample set, limiting the model's generalization ability and prediction accuracy. SUMMARY

[0004] The embodiment of the application provides a water quality detection method based on machine learning and combining spectrophotometry, which solves the problem of insufficient prediction accuracy caused by the lack of effective means to extract deep features in the prior art, and the inability to fully utilize the active learning strategy to dynamically adjust the training sample set, which limits the model's generalization ability and prediction accuracy.

[0005] In the first aspect, the embodiment of the application provides a water quality detection method based on machine learning and combining spectrophotometry, which includes: measuring the absorbance of water samples at multiple preset wavelengths by a spectrophotometer to form a multi-dimensional absorbance data set; using the multi-dimensional absorbance data set, mapping to a high-dimensional space through kernel function technology to obtain an enhanced nonlinear feature data set; extracting known water samples and corresponding pollutant concentrations from historical water quality data to generate a training sample set; applying a machine learning model to predict the pollutant concentration of the enhanced nonlinear feature data set, while dynamically adjusting the training sample set using an active learning strategy to obtain a pollutant concentration prediction value; using the pollutant concentration prediction value, combining water quality safety standards to evaluate the safety level of the water sample, and outputting a water quality report containing pollutant type, concentration level and safety drinking standard compliance.

[0006] In the second aspect, the embodiment of the application provides a water quality detection system based on machine learning and combining spectrophotometry, which includes:

[0007] a measurement module configured to measure absorbance of the water sample at a plurality of preset wavelengths by a spectrophotometer to form a multi-dimensional absorbance dataset;

[0008] a mapping module configured to map the multi-dimensional absorbance dataset to a high-dimensional space by a kernel function technique to obtain a dataset with enhanced nonlinear features;

[0009] an extraction module configured to extract known water samples and pollutant concentrations corresponding to the known water samples from historical water quality data to generate a training sample set;

[0010] a prediction module configured to apply a machine learning model to the dataset with enhanced nonlinear features for pollutant concentration prediction, and dynamically adjust the training sample set by using an active learning strategy to obtain a pollutant concentration prediction value;

[0011] an evaluation module configured to evaluate the safety level of the water sample by using the pollutant concentration prediction value in combination with a water quality safety standard, and output a water quality report containing a pollutant type, a concentration level, and a safety drinking standard compliance.

[0012] In a third aspect, an embodiment of the present application provides a computing device including a processor and a memory, the memory storing a computer program, and the processor being configured to execute the computer program to perform the water quality detection method based on machine learning in combination with spectrophotometry according to any one of the first aspect.

[0013] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing computer program instructions, and the computer program instructions are executed by a processor to implement the water quality detection method based on machine learning in combination with spectrophotometry according to any one of the first aspect.

[0014] In the embodiment of the present application, the absorbance of the water sample is measured at a plurality of preset wavelengths by a spectrophotometer to form a multi-dimensional absorbance dataset;

[0015] The multi-dimensional absorbance dataset is mapped to a high-dimensional space by a kernel function technique to obtain a dataset with enhanced nonlinear features;

[0016] A machine learning model is applied to the dataset with enhanced nonlinear features for pollutant concentration prediction, and a training sample set is dynamically adjusted by using an active learning strategy to generate a predicted pollutant concentration;

[0017] With the pollutant concentration prediction value, combined with the water quality safety standard, the safety grade of the water sample is evaluated, and a water quality report containing the pollutant type, concentration level and safety drinking standard compliance is output, the technical scheme provided by the application measures the absorbance of the water sample at multiple preset wavelengths by the spectrophotometer, forms a multi-dimensional absorbance data set, provides rich information basis, which helps to more accurately identify and quantify the pollutants, and the multi-dimensional absorbance data set is mapped to a high-dimensional space by using the kernel function technology, which enhances the non-linear feature expression ability of the data set. This processing method makes the complex and nonlinear relationship better, improves the prediction accuracy of the subsequent machine learning model, extracts the known water samples and their corresponding pollutant concentrations from the historical water quality data to generate an initial training sample set, and dynamically adjusts the training sample set in the subsequent process by using the active learning strategy, which not only makes full use of the existing data resources, but also optimizes the model by continuously introducing new samples, ensures the generalization ability and real-time adaptability of the model, realizes the integrated process from data acquisition to result interpretation, and provides scientific basis and technical support for water quality management decision;

[0018] Further, based on the deviation between the initial pollutant concentration estimation value and the actual measured value, the uncertainty region of the model prediction is determined, which can accurately locate the weak link in the model prediction and provide a clear direction for targeted optimization;

[0019] The uncertainty region is used to select the best water sample to add to the training sample set, thereby optimizing the machine learning model. The introduction of the active learning strategy enables the model to obtain more valuable information in the key area, effectively improving the prediction accuracy and stability of the model;

[0020] The optimized machine learning model (i.e. the reinforced machine learning model) is used to predict the pollutant concentration of the multi-dimensional absorbance data set again, and more accurate pollutant concentration prediction values are finally obtained. This method not only improves the prediction accuracy, but also enhances the generalization ability of the model to unknown samples in the future.

[0021] These aspects or other aspects of the application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0023] Figure 1A flow chart of a water quality detection method based on machine learning combined with spectrophotometry provided for an embodiment of the present application is shown in the figure.

[0024] Figure 2 A structural schematic diagram of a water quality detection system based on machine learning combined with spectrophotometry provided for an embodiment of the present application is shown in the figure.

[0025] Figure 3 A structural schematic diagram of a computing device provided for an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0026] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings of the embodiments of the present application.

[0027] In some of the descriptions in the specification and claims of the present application and the above-mentioned drawings, a plurality of operations appearing in a specific order are included, but it should be clearly understood that these operations can be executed in parallel or in parallel with the order in which they appear in this text, and the serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the "first", "second", etc. described herein are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence. "First" and "second" are different types.

[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0029] In the prior art, although the spectrophotometer can efficiently obtain the absorbance data of the water sample, when processing the multi-dimensional absorbance data set, the traditional method often lacks effective means to extract deep features, resulting in insufficient prediction accuracy. In addition, the sample set used in the training process of most machine learning models is usually static, and the active learning strategy is not used to dynamically adjust the training sample set, which limits the generalization ability and prediction accuracy of the model. Based on this, the present application provides a water quality detection method based on machine learning combined with spectrophotometry, as shown in Figure 1 , comprising:

[0030] Step 101: Measure the absorbance of the water sample at a plurality of preset wavelengths by a spectrophotometer to form a multi-dimensional absorbance data set.

[0031] In this step, spectrophotometer refers to an instrument used to measure the degree of light absorption by a substance at different wavelengths, capable of operating at multiple pre-set wavelengths to provide information about various components in the water sample; absorbance refers to the degree of light absorption by the water sample, an important parameter for measuring the concentration of dissolved or suspended substances in water; multi-dimensional absorbance dataset refers to a collection of absorbance values at multiple wavelengths, reflecting the presence of different pollutants and their relative concentrations in the water sample; by measuring the absorbance of the water sample at a series of selected wavelength ranges, each wavelength corresponds to an absorbance value. All measured absorbance values are collected to form a multi-dimensional dataset containing the response characteristics of the water sample at different wavelengths.

[0032] Step 102: Using the multi-dimensional absorbance dataset, map it to a high-dimensional space through kernel function technology to obtain a dataset with enhanced non-linear features;

[0033] In this step, kernel function technology refers to a mathematical tool used to convert original data from low-dimensional space to higher-dimensional space, making it easier to handle complex non-linear relationships; high-dimensional space refers to a space with higher dimensions than the original data space, in which the relationship between data points can be more clearly represented; dataset with enhanced non-linear features refers to the non-linear features in the dataset that are difficult to capture after kernel function mapping, which helps improve the accuracy of subsequent analysis; use kernel function technology such as radial basis function (RBF) kernel or polynomial kernel to map the multi-dimensional absorbance dataset to a higher-dimensional space, in which the complex non-linear relationships between data points become more apparent, forming a dataset with enhanced non-linear features.

[0034] Step 103: Extract known water samples and their corresponding pollutant concentrations from historical water quality data to generate a training sample set;

[0035] In this step, historical water quality data refers to past water quality detection records, including water sample information collected at different times and places and their corresponding pollutant concentrations; known water samples refer to those whose pollutant concentrations have been accurately measured and confirmed; training sample set refers to a data collection used to train machine learning models, usually containing input features (such as absorbance values) and target outputs (such as pollutant concentrations); filter historical water quality data from the database that meets the conditions, especially those that have been verified water samples and their corresponding pollutant concentrations, organize these data into structured tables or file formats to form a training sample set, preparing for subsequent machine learning model training.

[0036] Step 104: Apply the machine learning model to the dataset of enhanced nonlinear features for pollutant concentration prediction, while dynamically adjusting the training sample set using an active learning strategy to obtain the pollutant concentration prediction value;

[0037] In this step, the machine learning model refers to a mathematical model constructed by algorithms, which can predict the output results (such as pollutant concentration) based on input data (such as absorbance values). Common models include support vector machines, random forests, etc. The active learning strategy refers to an iterative machine learning method, in which the model selects new samples with the most value to join the training set according to the uncertainty of the current prediction results, to optimize the model performance. The pollutant concentration prediction value refers to the concentration of a specific pollutant in the water sample predicted by the machine learning model based on the input data. First, use the initially constructed machine learning model to predict the pollutant concentration of the enhanced nonlinear feature dataset, then evaluate the deviation between the predicted results and the actual measured values, identify the areas with larger prediction errors, then select representative new water samples for actual measurement according to the active learning strategy, and add these new samples and their measurement results to the training sample set. Repeat this process until the prediction accuracy of the model reaches the expected standard, and finally obtain the pollutant concentration prediction value.

[0038] Step 105: Use the pollutant concentration prediction value to evaluate the safety level of the water sample in combination with water quality safety standards, and output a water quality report containing pollutant type, concentration level, and compliance with safe drinking standards;

[0039] In this step, the water quality safety standards refer to a series of regulations established by the state or international organizations regarding the maximum allowable concentration of pollutants in drinking water and other water bodies for other purposes. The safety level evaluation refers to determining the safety level of the water sample based on the predicted pollutant concentration and the corresponding safety standards, such as safe, warning, dangerous, etc. The water quality report refers to the final output result document, which contains information such as pollutant type, concentration level, and compliance with safe drinking standards. According to the pollutant concentration prediction value provided by the machine learning model, evaluate the safety level of each pollutant by comparing it with the applicable water quality safety standards. Consider all the evaluation results of all pollutants to determine the overall safety status of the water sample. Finally, generate a detailed water quality report that not only lists the specific concentrations of various pollutants, but also clearly indicates whether the water sample meets the safe drinking standards, providing a basis for relevant decisions;

[0040] In a typical water quality detection scenario, imagine an environmental monitoring station dedicated to improving water quality detection capabilities to ensure that local residents can obtain safe and reliable drinking water. The monitoring station decides to adopt an innovative method - water quality detection technology based on machine learning combined with spectrophotometry, to achieve this goal.

[0041] Firstly, the technician uses a high-precision spectrophotometer to measure the absorbance of the collected water samples at multiple pre-set wavelengths These wavelengths are carefully selected based on the characteristics of common pollutants to ensure that as much information as possible is captured Through a series of measurements, a multi-dimensional absorbance data set is generated for each water sample, which contains absorbance values at different wavelengths, providing a rich information base for subsequent analysis;

[0042] Next, in order to better handle these complex data, the multi-dimensional absorbance data set is mapped to a higher-dimensional space using kernel function technology This process enhances the expression of non-linear features in the data set, making previously difficult-to-capture relationships more obvious and easier to handle After this conversion, the resulting data set not only retains all the characteristics of the original data, but also reveals more patterns hidden in the data, providing strong support for subsequent prediction models;

[0043] Then, in order to train the machine learning model, known water samples and their corresponding pollutant concentrations are extracted from the historical water quality database to construct the initial training sample set These historical data come from water quality monitoring records accumulated over the years, covering various types of water bodies and pollutants In this way, the training sample set is both representative and diverse, providing a solid foundation for the machine learning model;

[0044] With the training sample set, the machine learning model is applied to predict the pollutant concentration of the enhanced non-linear feature data set In the initial training phase, the machine learning model learns from the training sample set and gradually adjusts its parameters to minimize prediction error However, it does not stop there, but introduces an active learning strategy to dynamically optimize the training sample set Specifically, whenever the model makes a prediction, it evaluates the deviation between the predicted result and the actual measured value, identifying areas with high prediction error For these areas with high uncertainty, new water samples are selected for actual measurement, and these new samples and their measurement results are added to the training sample set to further optimize model performance Through repeated cycles of this process, a highly accurate and generalizable reinforcement machine learning model is ultimately obtained;

[0045] Finally, the enhanced machine learning model is used to predict the concentration of pollutants for new multi-dimensional absorbance data sets, obtaining detailed predicted values of pollutant concentration. Subsequently, these predicted values are compared with national and international water quality safety standards, and a safety level assessment is made for each water sample. According to the assessment results, each water sample is assigned a corresponding safety level (such as safe, warning or dangerous), and a complete water quality report is generated. This report not only lists the specific concentrations of various pollutants, but also clearly indicates whether the water sample meets the safety drinking water standards, providing important decision-making basis for relevant departments and the public.

[0046] Through the above steps, rapid, efficient and accurate water quality detection is successfully achieved, significantly improving the ability and level of water quality management. This method not only saves time and cost, but also makes an important contribution to public health and environmental protection.

[0047] Based on this, the present application provides a specific embodiment, wherein the step 104 applies a machine learning model to predict the concentration of pollutants based on the enhanced non-linear feature data set, and simultaneously adopts an active learning strategy to dynamically adjust the training sample set, obtaining the predicted value of the concentration of pollutants, specifically including the following steps:

[0048] Step 201: Preliminary training of the machine learning model based on the enhanced non-linear feature data set, obtaining an initial machine learning model;

[0049] In this step, the enhanced non-linear feature data set refers to a data set that has enhanced non-linear feature expression ability after being mapped to a high-dimensional space through kernel function technology; the machine learning model refers to a mathematical model constructed by algorithm, which can predict the output result (such as the concentration of pollutants) according to the input data (such as absorbance value); the preliminary training refers to the first training of the model using part of the data to preliminarily adjust the model parameters, ensuring that it has a certain prediction ability; the initial machine learning model refers to a model that has basic prediction ability but has not been fully optimized after preliminary training; the enhanced non-linear feature data set is used to preliminarily train the machine learning model, which involves selecting appropriate algorithms and parameter settings, and gradually adjusting the model parameters through iterative optimization. Through preliminary training, the machine learning model can preliminarily capture the patterns and relationships in the data, forming an initial machine learning model

[0050] Step 202: Using the initial machine learning model to predict the concentration of pollutants for multi-dimensional absorbance data sets, obtaining initial estimated values of the concentration of pollutants;

[0051] In this step, the initial machine learning model refers to a model that has been preliminarily trained and has basic prediction capabilities but has not been fully optimized. The multi-dimensional absorbance dataset refers to a set of absorbance values at multiple wavelengths that can reflect the presence of different pollutants and their relative concentrations in the water sample. The pollutant concentration prediction refers to predicting the concentration of a specific pollutant in a water sample based on input data such as absorbance values using a machine learning model. The initial pollutant concentration estimate refers to the preliminary prediction results given by the initial machine learning model, which may contain some errors. The initial machine learning model is applied to the multi-dimensional absorbance dataset for pollutant concentration prediction. During this process, the model evaluates new data points based on previously learned patterns and outputs the initial pollutant concentration estimate for each water sample. These estimates provide a reference for subsequent optimization

[0052] Step 203: Determine the uncertainty region of model prediction based on the deviation between the initial pollutant concentration estimate and the actual measured value

[0053] In this step, the initial pollutant concentration estimate refers to the preliminary prediction results given by the initial machine learning model, which may contain some errors. The actual measured value refers to the true pollutant concentration value measured by the laboratory or other accurate methods. The deviation refers to the difference between the predicted value and the actual value, which is used to measure the accuracy of model prediction. The uncertainty region refers to the region with larger prediction errors, indicating that the model's prediction in these regions is not accurate enough. By comparing the deviation between the initial pollutant concentration estimate and the actual measured value, the region with larger prediction errors is identified and marked as the uncertainty region, indicating that the model's prediction in this area has greater uncertainty, which helps to locate the key areas that need to be improved

[0054] Step 204: Use the uncertainty region to select the best water sample to join the training sample set using an active learning strategy to optimize the machine learning model, resulting in a reinforced machine learning model

[0055] In this step, the uncertainty region refers to the region with larger prediction errors, indicating that the model's prediction in these regions is not accurate enough. The active learning strategy refers to an iterative machine learning method in which the model selects the most valuable new samples to join the training set based on the uncertainty of the current prediction results to optimize model performance. The best water sample refers to those new samples that best reveal the uncertainty of the model's prediction, which helps to improve the model's generalization ability. The reinforced machine learning model refers to a model that has been optimized and has significantly improved prediction accuracy and stability. By using the information of the uncertainty region, the active learning strategy is used to select the best water sample for actual measurement, and these new samples and their measurement results are added to the training sample set. In this way, valuable data is continuously introduced, and the training sample set is dynamically adjusted, ultimately optimizing the machine learning model and forming a more accurate and stable reinforced machine learning model

[0056] Step 205: using the reinforced machine learning model to predict the concentration of pollutants again on the multi-dimensional absorbance data set, and obtaining the predicted value of the concentration of pollutants;

[0057] In this step, the reinforced machine learning model refers to the model whose prediction accuracy and stability have been significantly improved after optimization; the multi-dimensional absorbance data set refers to a set composed of absorbance values at multiple wavelengths, which can reflect the existence of different pollutants in the water sample and their relative concentrations; the predicted value of the concentration of pollutants refers to the final prediction result given by the reinforced machine learning model, which is usually more accurate and reliable; using the reinforced machine learning model to predict the concentration of pollutants again on the multi-dimensional absorbance data set, this time the prediction is based on the optimized model, which can provide more accurate and reliable predicted values of the concentration of pollutants, which not only improves the accuracy of detection, but also enhances the generalization ability to unknown samples in the future.

[0058] Based on this, the present application provides a specific embodiment, wherein the step 201, the machine learning model is initially trained according to the data set with enhanced nonlinear features, and an initial machine learning model is obtained, which specifically comprises the following steps:

[0059] Step 301: constructing the basic framework of the machine learning model according to the data set with enhanced nonlinear features and combining historical water quality data;

[0060] In this step, the data set with enhanced nonlinear features refers to a data set that has enhanced nonlinear feature expression ability after being mapped to a high-dimensional space by kernel function technology; the historical water quality data refers to the accumulated water quality detection records, including water sample information collected at different times and places and the corresponding concentration of pollutants

[0061] The basic framework of the machine learning model refers to the initially built mathematical model structure, which is used for subsequent training and optimization; constructing the basic framework of the machine learning model according to the data set with enhanced nonlinear features and combining historical water quality data, this process involves selecting appropriate algorithms and parameter settings to ensure that the basic framework can adapt to the characteristics of the data and provide a good starting point for subsequent optimization.

[0062] Step 302: using cross-validation method to test the stability of the basic framework, and obtaining the test result;

[0063] In this step, the cross-validation method refers to a technique for evaluating model performance by dividing the dataset into multiple subsets and using them as validation sets in turn to evaluate the stability and generalization ability of the model; the test results refer to various indicators of model performance obtained through cross-validation, such as accuracy, recall rate, etc.; the stability of the base framework is tested by cross-validation method by dividing the dataset into multiple subsets, and using them as training set and validation set in turn to evaluate the performance of the model under different data division Finally, a set of test results are obtained, which reflect the stability and generalization ability of the model.

[0064] Step 303: Based on the test results, the hyperparameters in the base framework are optimized and adjusted to obtain an optimized and adjusted base framework;

[0065] In this step, hyperparameters refer to parameters in the model that need to be set in advance and are not automatically learned during the training process, such as learning rate, regularization coefficient, etc.; optimization and adjustment refer to adjusting hyperparameters based on test results to improve model performance; based on test results, hyperparameters in the base framework are optimized and adjusted by methods such as grid search or random search to find the best combination of hyperparameters, ensuring that the model maintains high prediction accuracy and stability under different data division Finally, an optimized and adjusted base framework is obtained.

[0066] Step 304: Use the optimized and adjusted base framework to comprehensively train the data set with enhanced nonlinear features to obtain a preliminary trained machine learning model;

[0067] In this step, comprehensive training refers to fully training the model using the entire dataset to adjust internal parameters to better fit the data; the preliminary trained machine learning model refers to a model that has certain prediction ability but has not been fully optimized after comprehensive training; using the optimized and adjusted base framework, the data set with enhanced nonlinear features is comprehensively trained In this process, the model continuously adjusts internal parameters to minimize prediction errors Through comprehensive training, a preliminary trained machine learning model is obtained, which can accurately predict pollutant concentrations to a certain extent.

[0068] Step 305: Use the preliminary trained machine learning model to predict the multi-dimensional absorbance data set to obtain prediction results;

[0069] In this step, the multi-dimensional absorbance dataset refers to a collection of absorbance values at multiple wavelengths that reflect the presence and relative concentrations of different pollutants in the water sample. The prediction results refer to the estimated values of pollutant concentrations in the water sample output by the machine learning model. Using the preliminary trained machine learning model, the multi-dimensional absorbance dataset is predicted. The model evaluates new data points based on previously learned patterns and outputs initial pollutant concentration estimates for each water sample. These prediction results provide important references for further optimizing the model.

[0070] Step 306: Evaluate the initial performance of the preliminary trained machine learning model using the comparison between the prediction results and actual measurement values, obtaining performance evaluation results. Based on the performance evaluation results, adjust the parameters of the preliminary trained machine learning model to build an initial machine learning model.

[0071] In this step, the prediction results refer to the estimated values of pollutant concentrations in the water sample output by the machine learning model. The actual measurement values refer to the true pollutant concentration values measured by the laboratory or other precise methods

[0072] Initial performance evaluation refers to evaluating the performance of the model after preliminary training by comparing the prediction results with the actual measurement values. Performance evaluation results refer to indicators such as mean square error (MSE), mean absolute error (MAE), and determination coefficient (R² score) derived from the comparison, which are used to measure the prediction accuracy and stability of the model. Adjusting model parameters refers to optimizing internal parameters of the model based on performance evaluation results to improve the performance of the model. The initial machine learning model refers to the model with good prediction ability and stability after parameter adjustment. The process of evaluating the initial performance of the preliminary trained machine learning model using the comparison between the prediction results and actual measurement values involves calculating the deviation between the predicted values and the true values, generating detailed performance evaluation indicators such as mean square error, mean absolute error, and determination coefficient. Through these indicators, the prediction accuracy and stability of the model can be comprehensively understood.

[0073] Based on the performance evaluation results, the parameters of the preliminary trained machine learning model are adjusted. This step may include fine-tuning hyperparameters, introducing regularization terms to reduce the risk of overfitting, and other methods aimed at improving the generalization ability of the model to ensure its prediction accuracy under different water quality conditions.

[0074] Finally, through a series of adjustments and optimizations, an initial machine learning model is built. This model not only performs well on existing data, but also provides a solid foundation for further optimization in the future.

[0075] Based on this, the application provides a specific embodiment, wherein the step 306 comprises the following steps:

[0076] The step 401 comprises the following steps:

[0077] In this step, the prediction result refers to the estimated value of the pollutant concentration in the water sample output by the machine learning model; the actual measured value refers to the true pollutant concentration value measured by a laboratory or other accurate method

[0078] The prediction error index refers to various statistical quantities for measuring the prediction accuracy of the model, such as mean square error (MSE), mean absolute error (MAE), and coefficient of determination (R² score); the performance evaluation index refers to specific numerical values derived from the prediction error index, which are used to comprehensively evaluate the performance of the model; the prediction error index is calculated by comparing the prediction result with the actual measured value, including mean square error, mean absolute error, and R² score; these indicators can quantify the accuracy and stability of the model prediction, and through this process, detailed performance evaluation indicators are obtained, providing a basis for subsequent optimization.

[0079] The step 402 comprises the following steps:

[0080] In this step, the sensitivity analysis refers to an analysis method for determining which hyperparameters have the greatest impact on the performance of the model; the hyperparameters refer to the parameters in the model that need to be pre-set and are not automatically learned during the training process, such as learning rate, regularization coefficient, etc.; the key hyperparameter list refers to the set of hyperparameters that have a significant impact on the performance of the model identified through sensitivity analysis; based on the performance evaluation index, the sensitivity analysis identifies the hyperparameters that have the greatest impact on the performance, which involves changing the values of each hyperparameter and observing how these changes affect the performance of the model, ultimately obtaining a key hyperparameter list that lists the hyperparameters that have a significant impact on the performance of the model;

[0081] In order to identify the hyperparameters that have the greatest impact on the performance of the model, a sensitivity analysis method based on Gaussian Process Regression (GPR) can be used, wherein the expression of the sensitivity analysis method is as follows:

[0082]

[0083] where, is the sensitivity of the i-th hyperparameter, which measures the degree of influence of each hyperparameter on the model performance. By calculating the partial derivative, the impact of the change of each hyperparameter on the model output can be quantified; is the performance evaluation indicator of the machine learning model (such as mean square error, mean absolute error or coefficient of determination), which is derived from the actual performance in the model training and validation process, usually including mean square error (MSE), mean absolute error (MAE) or coefficient of determination (R²), which directly reflects the prediction accuracy and stability of the model; is the i-th hyperparameter, which refers to those parameters that need to be set in advance and are not automatically learned in the training process, such as learning rate, regularization coefficient, etc., which directly affect the training process and final performance of the model; is a weight parameter used to balance the influence of first-order derivative and second-order derivative, and the weight parameter is used to balance the influence of first-order derivative and second-order derivative. The first-order derivative reflects the direct influence of the hyperparameter on the model performance, while the second-order derivative captures the change of the rate of change of the hyperparameter, which helps to understand the influence of the hyperparameter more comprehensively; represents the expected value of the hyperparameter distribution , the expected value calculation in Gaussian Process Regression (GPR) is introduced, which considers the distribution characteristics of the hyperparameters, making the sensitivity analysis more robust; is the second-order partial derivative of the performance evaluation indicator with respect to , the second-order partial derivative provides information about the rate of change of the hyperparameter, helping to identify those hyperparameters that have a significant impact on the model performance but have complex nonlinear relationships; The overall formula is designed to improve the accuracy and applicability of sensitivity analysis by introducing advanced mathematical tools such as Gaussian Process Regression and Bayesian Optimization. These methods not only enable more accurate identification of the hyperparameters that have the greatest impact on model performance, but also dynamically adjust the selection of key hyperparameters to adapt to the characteristics and needs of different data sets.

[0084] Step 403: According to the list of key hyperparameters, use Bayesian optimization algorithm to conduct global search on the key hyperparameters of the preliminary trained machine learning model, find the optimal combination of hyperparameters, and get optimization suggestions;

[0085]

[0086] ​​In this step, Bayesian optimization refers to an efficient optimization algorithm that guides the search process by constructing a probabilistic model to find the optimal solution; the optimal hyperparameter combination refers to the best hyperparameter settings found within a given range, which maximizes model performance; optimization suggestions refer to improvement schemes proposed based on the results of Bayesian optimization; based on the list of key hyperparameters, Bayesian optimization is used to perform a global search on the key hyperparameters of the initially trained machine learning model. By iteratively adjusting the hyperparameters and evaluating model performance, the optimal hyperparameter combination is finally found. This process generates optimization suggestions to guide subsequent model adjustments.

[0087] Step 404: Adjust the key hyperparameters of the initially trained machine learning model according to the optimization suggestions, and introduce a regularization term to obtain the optimized machine learning model.

[0088] In this step, the regularization term refers to a technique to prevent overfitting by adding a penalty term to limit the model complexity; the optimized machine learning model refers to the model whose performance has been improved after hyperparameter tuning and regularization; the key hyperparameters of the initially trained machine learning model are adjusted according to the optimization suggestions, and appropriate regularization terms are introduced to reduce the risk of overfitting. In this way, an optimized machine learning model is obtained, whose predictive ability and generalization ability are significantly improved.

[0089] Step 405: Using the optimized machine learning model, predict the multidimensional absorbance dataset again to obtain the optimal prediction result;

[0090] In this step, the multidimensional absorbance dataset refers to a collection of absorbance values ​​at multiple wavelengths, which can reflect the presence and relative concentration of different pollutants in the water sample; the optimal prediction result refers to the more accurate pollutant concentration estimate output by the optimized model; the optimized machine learning model is used to predict the multidimensional absorbance dataset again, which evaluates the new data points based on the previously learned patterns and outputs the optimal prediction result for each water sample, which is usually more accurate and reliable.

[0091] Step 406: By comparing the optimal prediction result with the corresponding actual measurement value, verify the performance of the optimized machine learning model and obtain the final performance verification result;

[0092] In this step, the actual measurement value refers to the true pollutant concentration value measured by a laboratory or other precise method; the final performance verification result refers to the final evaluation index of the model performance obtained by comparing the optimal prediction result with the actual measurement value; the performance of the optimized machine learning model is verified by comparing the optimal prediction result with the corresponding actual measurement value, and through this process, the final performance verification result is obtained, which reflects the performance of the model after optimization and further confirms the prediction accuracy and stability of the model.

[0093] Step 407: Adjust the parameters of the optimized machine learning model based on the final performance verification result to build a completed initial machine learning model.

[0094] In this step, the final performance verification result refers to the final evaluation index of the model performance obtained by comparing the optimal prediction result with the actual measurement value; adjusting the model parameters refers to further fine-tuning the internal parameters of the model according to the final performance verification result to ensure optimal performance; the initial machine learning model refers to the model with high prediction ability and stability finally built after multiple optimizations and verifications; based on the final performance verification result, the parameters of the optimized machine learning model are fine-tuned to ensure the accuracy and reliability of the model in predicting all pollutant concentrations, and finally a complete initial machine learning model is built, which not only performs well on existing data but also has good generalization ability, providing a solid foundation for subsequent applications.

[0095] Based on this, the present application provides a specific embodiment, wherein the step 401 comprises the following steps:

[0096] Step 501: Calculate the prediction error indicators including mean square error, mean absolute error and R² score by comparing the prediction results with the actual measurement values through statistical methods to obtain preliminary performance evaluation data, wherein the error indicators include: mean square error, mean absolute error and determination coefficient.

[0097] In this step, the statistical method refers to a method for quantifying the difference between the prediction results and the actual measurement values, such as calculating the mean square error (MSE), mean absolute error (MAE) and determination coefficient (R²); the preliminary performance evaluation data refers to specific numerical values obtained from the prediction error indicators, which are used to preliminarily understand the performance of the model; the prediction error indicators including mean square error, mean absolute error and determination coefficient are calculated by comparing the prediction results with the actual measurement values through statistical methods, and through this process, the preliminary performance evaluation data is obtained, which provides a basis for subsequent analysis.

[0098] Step 502: Construct a performance evaluation model using the preliminary performance evaluation data, classify the prediction errors of different pollutant concentrations using the performance evaluation model, identify the pollutant categories with higher errors, and obtain a high-error pollutant list;

[0099] In this step, the classification analysis refers to analyzing the prediction errors of different pollutant concentrations by category to identify the pollutant categories with higher errors; the high-error pollutant list refers to listing those pollutant categories with significantly higher prediction errors than others; using the preliminary performance evaluation data, a performance evaluation model is constructed, which is used to classify the prediction errors of different pollutant concentrations. Through this analysis, the pollutant categories with higher errors are identified, and a high-error pollutant list is generated.

[0100] Step 503: Analyze the causes of high prediction errors according to the high-error pollutant list combined with the physical and chemical properties of pollutants, and obtain an error cause analysis report;

[0101] In this step, the physical and chemical properties refer to the characteristics of the pollutants themselves, such as solubility, molecular weight, etc., which may affect the prediction accuracy; the error cause analysis report refers to a document that records the reasons for high prediction errors in detail; according to the high-error pollutant list, combined with the physical and chemical properties of pollutants, the causes of high prediction errors are analyzed. By comprehensively considering the characteristics of pollutants and their behavior in the environment, the error cause analysis report is ultimately obtained.

[0102] Step 504: Use the error cause analysis report to propose targeted improvement measures and obtain optimization suggestions, and use the optimization suggestions to perform targeted optimization on the initially trained machine learning model to obtain an optimized machine learning model;

[0103] In this step, the improvement measures refer to solutions proposed for error causes, such as increasing the number of samples for specific pollutants or adjusting the model architecture; the optimization suggestions refer to clear guidance schemes formed based on the improvement measures; the optimized machine learning model refers to a model whose performance has been improved after targeted optimization; using the error cause analysis report, targeted improvement measures are proposed to obtain optimization suggestions. According to these suggestions, the initially trained machine learning model is optimized, and finally an optimized machine learning model is obtained, which has improved prediction accuracy and stability.

[0104] Step 505: Use the optimized machine learning model to calculate the prediction error indicators again to obtain the target performance evaluation indicators;

[0105] In this step, the prediction error index refers to various statistics used to measure the model's prediction accuracy, such as mean squared error, mean absolute error, and coefficient of determination; the target performance evaluation index refers to the prediction error index recalculated after the optimization of the model, reflecting the model's performance after optimization; using the optimized machine learning model, the prediction error index is recalculated. Through this process, the target performance evaluation index is obtained. These indexes demonstrate the prediction accuracy and stability of the optimized model, providing an important reference for subsequent applications.

[0106] Based on this, the present invention provides a specific embodiment. Step 102, which utilizes the multidimensional absorbance dataset and maps it to a high-dimensional space using kernel function technology to obtain a dataset with enhanced nonlinear features, specifically includes the following steps:

[0107] Step 601: Using the multidimensional absorbance dataset, select a suitable kernel function to map the data to a high-dimensional space to obtain a preliminary mapped dataset;

[0108] In this step, the multidimensional absorbance dataset refers to a collection of absorbance values ​​at multiple wavelengths, which reflects the presence and relative concentration of different pollutants in the water sample. The kernel function is a mathematical tool used to transform the original data from a low-dimensional space to a higher-dimensional space, making the originally complex nonlinear relationships easier to handle. The high-dimensional space refers to a space with a higher dimension than the original data space, in which the relationships between data points can be more clearly represented. The preliminary mapping dataset refers to the dataset formed after kernel function mapping; this dataset retains all the characteristics of the original data and enhances the expressive power of nonlinear features. Using the multidimensional absorbance dataset, a suitable kernel function, such as the radial basis function (RBF) or a polynomial kernel, is selected to map the data to a high-dimensional space. Through this process, the preliminary mapping dataset is obtained, which provides a better foundation for subsequent analysis.

[0109] Step 602: Using the preliminary mapping dataset, a feature selection algorithm is used to filter the features that contribute most to the prediction of pollutant concentration, resulting in a selected feature set;

[0110] In this step, the preliminary mapping dataset refers to the dataset formed after kernel function mapping, which retains all the characteristics of the original data and enhances the expression ability of nonlinear features; the feature selection algorithm refers to a method for identifying and selecting features that are most helpful to model performance, such as recursive feature elimination (RFE), principal component analysis (PCA), etc.; the selected feature set refers to the feature set that is most contributed to the prediction of pollutant concentration after being screened by the feature selection algorithm; the selected feature set is obtained by screening the features that are most contributed to the prediction of pollutant concentration from the preliminary mapping dataset through the feature selection algorithm. The selected feature set not only reduces the data dimension, but also improves the efficiency and accuracy of model training.

[0111] Step 603: Using the selected feature set, perform feature semantic annotation combined with domain knowledge to obtain the annotated feature set;

[0112] In this step, the selected feature set refers to the feature set that is most contributed to the prediction of pollutant concentration after being screened by the feature selection algorithm; domain knowledge refers to professional knowledge about a specific application field, such as chemical and physical principles in water quality detection; feature semantic annotation refers to adding labels to each feature to describe its meaning and function to increase the interpretability of the feature; the annotated feature set refers to the feature set after semantic annotation processing, which has clear meaning and function; the selected feature set is used to perform feature semantic annotation combined with domain knowledge. Through this process, the annotated feature set is obtained. The annotated feature set not only helps to understand the specific meaning of each feature, but also improves the interpretability of the model results.

[0113] Step 604: Using the annotated feature set, construct a feature correlation network to analyze the interaction and dependency between features, and obtain a feature correlation network graph;

[0114] In this step, the annotated feature set refers to the feature set after semantic annotation processing, which has clear meaning and function; the feature correlation network refers to a graphical representation of the interaction and dependency between features; the feature correlation network graph refers to a chart generated by constructing a feature correlation network, which shows the complex relationships between features; the annotated feature set is used to construct a feature correlation network to analyze the interaction and dependency between features. Through this process, the feature correlation network graph is obtained. The feature correlation network graph reveals the complex relationships between features, providing a basis for further optimizing the model.

[0115] Step 605: Using the feature correlation network graph, optimize the selection of kernel functions and the parameter settings of kernel functions to obtain an optimized high-dimensional space dataset;

[0116] In this step, the feature correlation network graph refers to the graph generated by constructing the feature correlation network, which shows the complex relationship between various features; the selection of the kernel function refers to selecting the most suitable type of kernel function according to the feature correlation network graph; the parameter setting of the kernel function refers to adjusting the parameters of the kernel function to optimize the model performance; the optimized high-dimensional space data set refers to the data set that is more suitable for training the machine learning model after optimization; and the selection of the kernel function and the parameter setting of the kernel function are optimized by using the feature correlation network graph. Through this process, the optimized high-dimensional space data set is obtained. The optimized high-dimensional space data set not only enhances the expression ability of nonlinear features, but also improves the prediction accuracy of the model.

[0117] Step 606: using the optimized high-dimensional space data set, performing data dimension reduction processing to obtain a data set with enhanced nonlinear features;

[0118] In this step, the optimized high-dimensional space data set refers to the data set that is more suitable for training the machine learning model after optimization; the data dimension reduction processing refers to the process of reducing the data dimension while trying to maintain important information, and the commonly used methods include principal component analysis (PCA), linear discriminant analysis (LDA), etc.; the data set with enhanced nonlinear features refers to the data set that still retains key nonlinear features after dimension reduction processing, which is convenient for subsequent model training and prediction; and the data dimension reduction processing is performed by using the optimized high-dimensional space data set. Through this process, the data set with enhanced nonlinear features is obtained. This data set not only reduces the computational complexity, but also retains the most important nonlinear features, providing high-quality data support for the final model training.

[0119] Based on this, the present application provides a specific embodiment, wherein the step 105, using the predicted pollutant concentration, combined with water quality safety standards, to evaluate the safety level of the water sample, and output a water quality report containing pollutant type, concentration level and safety drinking standard compliance, specifically including the following steps:

[0120] Step 701: using the predicted pollutant concentration to determine the safety level of each pollutant by comparing with national and international water quality safety standards;

[0121] In this step, the predicted pollutant concentration refers to the estimated value of the machine learning model output about the concentration of pollutants in the water sample; the national and international water quality safety standards refer to the provisions of the maximum allowable concentration of pollutants in drinking water and other water bodies made by the state or international organizations; the safety level refers to the safety classification of each pollutant according to whether the concentration of pollutants exceeds the safety standard; and the safety level of each pollutant is determined by using the predicted pollutant concentration to compare with the national and international water quality safety standards. This process involves comparing the predicted concentration value with the standard limit value, providing basic information for subsequent evaluation.

[0122] Step 702: Based on the safety level, a comprehensive evaluation model is constructed in combination with the health impact factor of pollutants to evaluate the overall safety of the water sample, and a comprehensive evaluation result is obtained.

[0123] In this step, the health impact factor refers to the potential impact of pollutants on human health, such as carcinogenicity and teratogenicity. The comprehensive evaluation model refers to a mathematical model used to comprehensively consider various factors (such as pollutant concentration and health impact factor) to evaluate the overall safety of the water sample. The comprehensive evaluation result refers to the conclusion about the overall safety of the water sample obtained by the comprehensive evaluation model. Based on the safety level, a comprehensive evaluation model is constructed in combination with the health impact factor of pollutants. This model not only considers whether the concentration of pollutants exceeds the standard, but also evaluates the potential impact of these pollutants on human health. By evaluating the overall safety of the water sample, the comprehensive evaluation result is ultimately obtained.

[0124] Step 703: Using the comprehensive evaluation result, the safety level of the water sample is divided into three levels: safe, warning, and dangerous, and the safety level of the water sample is obtained.

[0125] In this step, the safety level division refers to the process of dividing the water sample into different safety levels according to the comprehensive evaluation result. The safety level of the water sample refers to the final determined safety level of the water sample, which is divided into three levels: safe, warning, and dangerous. Using the comprehensive evaluation result, the safety level of the water sample is divided into three levels: safe, warning, and dangerous. Through this process, the safety level of the water sample is obtained, which provides clear basis for subsequent management and decision-making.

[0126] Step 704: Using the safety level of the water sample, in combination with geographical location and time information, a water quality condition map is generated.

[0127] In this step, the geographical location and time information refer to the collection location of the water sample and its corresponding sampling time. The water quality condition map refers to a graphical representation of the water quality safety in different regions, which facilitates intuitive understanding of the water quality in each region. Using the safety level of the water sample, in combination with geographical location and time information, a water quality condition map is generated. The map marks the water quality safety level of different regions and displays the water quality condition changing with time and space, providing an intuitive reference tool for the public and management departments.

[0128] Step 705: Using the water quality condition map, targeted water quality improvement suggestions are made, and an improvement suggestion report is obtained, wherein the water quality improvement suggestions include pollution source control and water quality purification measures.

[0129] In this step, water quality improvement suggestions refer to solutions proposed for specific regional water quality problems, such as pollution source control and water purification measures; improvement suggestion reports refer to documents that record water quality improvement suggestions in detail, providing guidance for actual operations; using the water quality status map, targeted water quality improvement suggestions are developed to obtain improvement suggestion reports, which include pollution source control and water purification measures, and provide specific action guidelines for relevant departments to improve water quality conditions.

[0130] Step 706: Using the improvement suggestion report, update the water quality status map and water quality improvement suggestions at a preset time, combined with water quality monitoring data, to obtain an updated water quality status map and water quality improvement suggestions;

[0131] In this step, water quality monitoring data refers to regularly collected water quality detection data used to track water quality changes; the preset time refers to a set time interval, such as monthly or quarterly, used to regularly update relevant information; the updated water quality status map and water quality improvement suggestions refer to maps and reports that reflect the latest water quality conditions and improvement suggestions after regular updates; using the improvement suggestion report, combined with water quality monitoring data, update the water quality status map and water quality improvement suggestions at a preset time, through this process, obtain the updated water quality status map and water quality improvement suggestions, ensuring the timeliness and accuracy of the information, supporting continuous improvement.

[0132] Step 707: Using the updated water quality status map and water quality improvement suggestions, output a complete water quality report, which includes: pollutant type, concentration level, safety drinking standard compliance, water quality safety level, and water quality improvement suggestions;

[0133] In this step, the complete water quality report refers to the final output of a comprehensive water quality analysis document, including pollutant type, concentration level, safety drinking standard compliance, water quality safety level, and water quality improvement suggestions; using the updated water quality status map and water quality improvement suggestions, output a complete water quality report, which not only includes pollutant type, concentration level, and safety drinking standard compliance, but also lists water quality safety level and water quality improvement suggestions in detail, providing comprehensive information support for relevant decision-making.

[0134] Figure 2 An embodiment of the present application provides a structure diagram of a water quality detection system based on machine learning combined with spectrophotometry, as shown in Figure 2 The system includes:

[0135] The measurement module 21 is used to measure the absorbance of the water sample at multiple preset wavelengths by a spectrophotometer to form a multi-dimensional absorbance data set;

[0136] The mapping module 22 is configured to map the multi-dimensional absorbance data set to a high-dimensional space by a kernel function technique to obtain an enhanced non-linear feature data set;

[0137] The extraction module 23 is configured to extract known water samples and corresponding pollutant concentrations of the known water samples from historical water quality data to generate a training sample set;

[0138] The prediction module 24 is configured to apply a machine learning model to the enhanced non-linear feature data set for pollutant concentration prediction, and dynamically adjust the training sample set by using an active learning strategy to obtain a pollutant concentration prediction value.

[0139] The evaluation module 25 is configured to evaluate the safety level of the water sample by using the pollutant concentration prediction value and combining water quality safety standards, and output a water quality report containing the pollutant type, concentration level, and safety drinking standard compliance.

[0140] Figure 2 The water quality detection system based on machine learning and combined with spectrophotometry can perform Figure 1 The water quality detection method based on machine learning and combined with spectrophotometry has the same implementation principles and technical effects as the above-mentioned embodiments. The specific operation modes of each module and unit of the water quality detection system based on machine learning and combined with spectrophotometry in the above-mentioned embodiments have been described in detail in the embodiments related to the method, and will not be described in detail here.

[0141] Figure 2 The water quality detection system based on machine learning and combined with spectrophotometry can be implemented as a computing device, such as Figure 3 As shown, the computing device can include a storage component 31 and a processing component 32.

[0142] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.

[0143] The processing component 32 includes one or more processors to execute computer instructions to complete all or part of the steps in the above-mentioned method. Of course, the processing component can also be one or more application-specific integrated circuits (AICs), digital signal processors (DPs), digital signal processing devices (DPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic elements for executing the above-mentioned method.

[0144] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or nonvolatile storage devices, or a combination thereof, such as a static random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk, or a optical disk.

[0145] The computing device further includes other components, such as an input / output interface, a display component, a communication component.

[0146] The input / output interface provides an interface between the processing component and peripheral interface modules, which can be output devices, input devices.

[0147] The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, etc.

[0148] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform, and the computing device can be a cloud server, and the processing component, the storage component, etc. can be a basic server resource rented or purchased from the cloud computing platform.

[0149] The embodiment of the application further provides a computer storage medium storing a computer program, and the computer program can implement the above-mentioned Figure 1 The embodiment shown in the figure provides a water quality detection method and system based on machine learning combined with spectrophotometry.

[0150] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-mentioned system, device and unit can refer to the corresponding process in the foregoing method embodiment, and will not be described here.

[0151] The device embodiment described above is only schematic, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the embodiment scheme. Those skilled in the art can understand and implement without creative labor.

[0152] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0153] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A water quality detection method based on machine learning combined with spectrophotometry, characterized in that, include: The absorbance of water samples was measured at multiple preset wavelengths using a spectrophotometer to form a multidimensional absorbance dataset. Using the aforementioned multidimensional absorbance dataset, a high-dimensional space is mapped through kernel function techniques to obtain a dataset with enhanced nonlinear features; Extract known water samples and their corresponding pollutant concentrations from historical water quality data to generate a training sample set; A machine learning model is applied to predict pollutant concentrations in the dataset with enhanced nonlinear features, while an active learning strategy is used to dynamically adjust the training sample set to obtain predicted pollutant concentrations. Using the predicted pollutant concentration values ​​and in conjunction with water quality safety standards, the safety level of the water sample is assessed, and a water quality report is output that includes the type of pollutant, concentration level, and compliance with safe drinking standards. The application of a machine learning model to predict pollutant concentrations on the dataset with enhanced nonlinear features, and the dynamic adjustment of the training sample set using an active learning strategy to obtain predicted pollutant concentration values, includes: The machine learning model is initially trained based on the dataset with enhanced nonlinear features to obtain the initial machine learning model; The initial machine learning model is then optimized for hyperparameters, wherein a sensitivity analysis method based on Gaussian process regression is used to identify key hyperparameters; the expression for the sensitivity analysis method is as follows: ; in, Refers to the first Sensitivity of each hyperparameter; Refers to the performance evaluation metrics of machine learning models; Refers to the first One hyperparameter; It is a weighting parameter used to balance the effects of the first and second derivatives; Indicates the hyperparameter distribution Expected value; Performance evaluation metrics are related to The second-order partial derivative; Based on the sensitivity A list of key hyperparameters is determined, and the key hyperparameters are optimized using a Bayesian optimization algorithm based on the list of key hyperparameters to obtain an optimized initial machine learning model; The optimized initial machine learning model is used to predict pollutant concentrations in the multidimensional absorbance dataset to obtain initial pollutant concentration estimates. Based on the deviation between the initial pollutant concentration estimate and the actual measurement, the uncertainty region of the model prediction is determined; Using the aforementioned uncertainty region, an active learning strategy is employed to select the optimal water sample to be added to the training sample set to optimize the optimized initial machine learning model, thereby obtaining a reinforced machine learning model. The enhanced machine learning model is used to predict pollutant concentrations again on the multidimensional absorbance dataset to obtain predicted pollutant concentration values.

2. The method according to claim 1, characterized in that, The machine learning model is initially trained based on the dataset with enhanced nonlinear features to obtain an initial machine learning model, including: Based on the dataset with enhanced nonlinear features, combined with historical water quality data, the basic framework for machine learning models is constructed. The basic framework was subjected to stability testing using cross-validation, and the test results were obtained. Based on the test results, the hyperparameters in the basic framework are optimized and adjusted to obtain the optimized basic framework. Using the optimized and adjusted basic framework, a dataset with enhanced nonlinear features is fully trained to obtain a preliminarily trained machine learning model. Using the pre-trained machine learning model, predictions are made on the multidimensional absorbance dataset to obtain prediction results; By comparing the predicted results with the actual measured values, the initial performance of the pre-trained machine learning model is evaluated to obtain the performance evaluation results. Based on the performance evaluation results, the parameters of the pre-trained machine learning model are adjusted to construct the initial machine learning model.

3. The method according to claim 2, characterized in that, By comparing the predicted results with the actual measured values, the initial performance of the pre-trained machine learning model is evaluated to obtain a performance evaluation result. Based on the performance evaluation result, the parameters of the pre-trained machine learning model are adjusted to construct an initial machine learning model, including: By comparing the predicted results with the actual measured values, prediction error indices are calculated, including mean square error, mean absolute error, and R² score, to obtain performance evaluation indices. Based on the performance evaluation metrics, sensitivity analysis is used to identify the hyperparameters that have the greatest impact on performance, resulting in a list of key hyperparameters. Based on the list of key hyperparameters, a Bayesian optimization algorithm is used to perform a global search on the key hyperparameters of the initially trained machine learning model to find the optimal combination of hyperparameters and obtain optimization suggestions. Based on the optimization suggestions, the key hyperparameters of the initially trained machine learning model are adjusted, and a regularization term is introduced to obtain the optimized machine learning model. Using the optimized machine learning model, the multidimensional absorbance dataset is predicted again to obtain the optimal prediction result; By comparing the optimal prediction results with the corresponding actual measured values, the performance of the optimized machine learning model is verified, and the final performance verification results are obtained. Based on the final performance verification results, the parameters of the optimized machine learning model are adjusted, and the initial machine learning model is constructed.

4. The method according to claim 3, characterized in that, By comparing the predicted results with the actual measured values, prediction error indices are calculated, including mean squared error, mean absolute error, and R² score, to obtain performance evaluation indices, including: By comparing the predicted results with the actual measured values, the prediction error index is calculated using statistical methods to obtain preliminary performance evaluation data. The error index includes: mean squared error, mean absolute error, and coefficient of determination. Using the preliminary performance evaluation data, a performance evaluation model is constructed. Using the performance evaluation model, the prediction errors of different pollutant concentrations are classified and analyzed to identify the types of pollutants with high errors, and a list of high-error pollutants is obtained. Based on the list of high-error pollutants and their physicochemical properties, the causes of high prediction errors are analyzed, and an error cause analysis report is obtained. Using the error cause analysis report, targeted improvement measures are proposed to obtain optimization suggestions. Using the optimization suggestions, the initially trained machine learning model is optimized to obtain the optimized machine learning model. Using the optimized machine learning model, the prediction error index is calculated again to obtain the target performance evaluation index.

5. The method according to claim 1, characterized in that, Using the aforementioned multidimensional absorbance dataset, a high-dimensional space is mapped through kernel function techniques to obtain a dataset with enhanced nonlinear features, including: Using the multidimensional absorbance dataset, a kernel function is selected to map the data to a high-dimensional space, resulting in a preliminary mapped dataset; Using the aforementioned preliminary mapping dataset, a feature selection algorithm is used to filter the features that contribute most to the prediction of pollutant concentrations, resulting in a carefully selected feature set. Using the selected feature set and combining domain knowledge, feature semantic annotation is performed to obtain the annotated feature set; Using the labeled feature set, a feature association network is constructed, and the interactions and dependencies between features are analyzed to obtain a feature association network graph; By utilizing the aforementioned feature association network graph, the selection of the kernel function and the parameter settings of the kernel function are optimized to obtain an optimized high-dimensional space dataset; Using the optimized high-dimensional space dataset, dimensionality reduction processing is performed to obtain a dataset with enhanced nonlinear features.

6. The method according to claim 1, characterized in that, Using the predicted pollutant concentrations and in conjunction with water quality safety standards, a safety level assessment is conducted on the water sample, generating a water quality report that includes the pollutant type, concentration level, and compliance with safe drinking water standards, including: By comparing the predicted pollutant concentrations with national and international water quality safety standards, the safety level of each pollutant is determined. Based on the aforementioned safety level and combined with the health impact factors of pollutants, a comprehensive assessment model is constructed to evaluate the overall safety of the water sample and obtain the comprehensive assessment results. Based on the comprehensive assessment results, the water samples are classified into three levels: safe, warning, and dangerous, thus obtaining the water sample safety level. Using the water sample safety level, combined with geographical location and time information, a water quality status map is generated; Using the aforementioned water quality map, targeted water quality improvement recommendations are formulated, resulting in an improvement recommendation report. These recommendations include pollution source control and water purification measures. Using the aforementioned improvement suggestion report, combined with water quality monitoring data, the water quality status map and water quality improvement suggestions are updated at preset times to obtain the updated water quality status map and water quality improvement suggestions; Using the updated water quality map and water quality improvement recommendations, a complete water quality report is generated, which includes: pollutant type, concentration level, compliance with safe drinking standards, water quality safety level, and water quality improvement recommendations.

7. A water quality detection system based on machine learning combined with spectrophotometry, used to execute the water quality detection method based on machine learning combined with spectrophotometry as described in any one of claims 1 to 6, characterized in that, include: The measurement module is used to measure the absorbance of water samples at multiple preset wavelengths using a spectrophotometer, forming a multidimensional absorbance dataset. The mapping module is used to map the multidimensional absorbance dataset to a high-dimensional space using kernel function techniques to obtain a dataset with enhanced nonlinear features; The extraction module is used to extract known water samples and their corresponding pollutant concentrations from historical water quality data to generate a training sample set. The prediction module is used to apply a machine learning model to predict pollutant concentrations on the dataset with enhanced nonlinear features, and simultaneously uses an active learning strategy to dynamically adjust the training sample set to obtain predicted pollutant concentration values. The assessment module is used to assess the safety level of water samples by combining the predicted pollutant concentration values ​​with water quality safety standards, and outputs a water quality report that includes the type of pollutant, concentration level, and compliance with safe drinking standards.

8. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a water quality detection method based on machine learning combined with spectrophotometry as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that, The device contains a computer program that, when executed by a computer, implements a water quality detection method based on machine learning combined with spectrophotometry as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Nonlinear full-spectrum water turbidity quantitative analysis method based on extreme random tree

    CN110887798A

  • River and lake water environment data fusion and sample labeling method and system based on machine learning and computer equipment

    CN115146720A

  • Method and device for training crack image detection model based on active learning

    CN118918419A