Method for measuring ammonia in domestic drinking water by continuous flow injection method
By employing continuous flow injection and data processing techniques, the accuracy and consistency issues of ammonia detection under different water sources and laboratory conditions were resolved, enabling ammonia concentration detection under unified standards and improving the reliability and stability of water quality monitoring.
Patent Information
- Application Number
- CN202511114677.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing detection methods lack accuracy and consistency in detecting ammonia under different water source types and laboratory conditions. There is a lack of systematic method comparison and verification and inter-laboratory data consistency analysis tools, making it difficult to compare water quality monitoring data across regions or laboratories.
The continuous flow injection method was adopted. Outliers and noise were removed through the data preprocessing module, a standardized dataset was established, the correlation between methods was detected by multiple linear regression analysis, a support vector machine model was constructed for nonlinear mapping, a set of calibration parameters was generated, the bias between methods was eliminated, an inter-laboratory data consistency evaluation system was established, and the calibration process was dynamically triggered.
It significantly improves the accuracy and consistency of ammonia detection in multiple water sources, optimizes quality control efficiency, and provides reliable technical support for environmental monitoring.
Smart Images

Figure CN120971680A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of determining ammonia in drinking water, and particularly relates to a method for determining ammonia in drinking water by continuous flow injection. BACKGROUND
[0002] The detection of ammonia in drinking water is a key area for ensuring water quality safety and public health, directly related to whether the drinking water meets health standards and affecting the ecological environment and social well-being. Ammonia, as a common pollutant in water bodies, is derived from agricultural runoff, industrial wastewater, and natural decomposition. Accurate detection of its content is crucial for water quality monitoring and management. However, existing detection methods have significant limitations in practical application. Traditional methods such as spectrophotometry or electrode method often require complex sample pretreatment, time-consuming operation and are easily affected by interfering substances in water samples, resulting in insufficient stability of the results. In addition, different detection methods have different applicability in different types of water sources (such as surface water, groundwater or treated drinking water), and lack of systematic comparison and verification, resulting in insufficient data comparability and difficulty in meeting the precise monitoring needs in multiple scenarios.
[0003] These limitations further highlight the core technical challenges. First, the evaluation system of correlation and accuracy among detection methods is not yet perfect. Due to the different response characteristics of free ammonia, total ammonia and organic nitrogen in different methods, there is a lack of unified standard to evaluate their consistency. For example, some methods may perform well in surface water detection, but have significant errors in groundwater containing complex organic matter, making it difficult to compare water quality monitoring data across regions or laboratories. Second, the comparability of detection results between laboratories is insufficient. The differences in detection equipment, operation process and environmental conditions used by different laboratories result in poor consistency of results of serial detection techniques. For example, in actual monitoring, the same water sample may have significant numerical deviations when detected by continuous flow injection method in different laboratories due to different instrument calibration or operation habits, affecting the accurate judgment of water quality trends.
[0004] Therefore, how to establish a systematic method for comparison and verification and a tool for analyzing the consistency of data between laboratories to ensure the detection accuracy and comparability of continuous flow injection method in different types of water sources and laboratory conditions becomes a key problem of this research. SUMMARY
[0005] The technical problem to be solved by the present application is to provide a method for determining ammonia in drinking water by continuous flow injection method to overcome the deficiencies of the background art.
[0006] The present application adopts the following technical solution to solve the above technical problems: The method for determining ammonia in drinking water by continuous flow injection method comprises: The raw data of ammonia detection of different water source types are acquired, abnormal values and noise interference are removed through a data preprocessing module, a standardized detection data set is established, meanwhile, operation parameters and environmental condition information of each detection method are recorded, clean basic data matrix is obtained for subsequent analysis and processing; According to the basic data matrix, the correlation characteristics between different detection methods are analyzed by using a multivariate linear regression algorithm, the consistency level and deviation source between methods are judged by calculating the linear relationship coefficient and residual distribution mode of the detection results of each method, and the detection method category needing correction is determined; For the identified deviation detection method, a nonlinear mapping relationship model between methods is constructed by using a support vector machine algorithm, the complex method difference characteristics are processed by using a kernel function, a detection result conversion matrix under different water source conditions is established, and a standardized method correction parameter set is obtained; The original detection data are corrected in batches by using the correction parameter set, the systematic deviation between methods is eliminated through matrix operation, ammonia concentration detection results under a unified standard are generated, and the change amplitude and distribution characteristics of the data before and after correction are calculated to judge the effectiveness degree of the correction effect; According to the corrected detection result data, the key factor weights affecting the detection accuracy are analyzed by using a random forest algorithm, the contribution degrees of water source types, interference substance concentrations, operation conditions and other variables to the detection accuracy are identified, and the key monitoring indexes and threshold ranges of quality control are determined; Based on the key monitoring indexes, a data consistency evaluation system between laboratories is established, the standard deviation and coefficient of variation of the detection results of each laboratory are calculated, a consistency scoring matrix is generated, if the score is lower than a preset threshold, a data correction process is triggered, and a final detection report meeting the standard requirements is obtained.
[0007] The technical scheme provided by the embodiment of the application can include the following beneficial effects: The application discloses a comprehensive processing method for ammonia detection data consistency and accuracy of different water source types, solves the systematic deviation and insufficient consistency caused by operation parameters, environmental conditions and interference substances in multi-source and multi-method detection, removes abnormal values and noise through a data preprocessing module, establishes a standardized data set, analyzes the correlation and deviation source between detection methods by using multivariate linear regression, generates a correction parameter set by using a support vector machine to construct a nonlinear mapping model, eliminates the deviation between methods in batches, and generates ammonia concentration results under a unified standard. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 Flow chart of the method for determining ammonia in drinking water by the continuous flow injection method of the present application.
[0009] Figure 2 Schematic diagram of the method for determining ammonia in drinking water by the continuous flow injection method of the present application Figure One .
[0010] Figure 3 Schematic diagram of the method for determining ammonia in drinking water by the continuous flow injection method of the present application Figure Two . DETAILED DESCRIPTION
[0011] In order to enable the personnel in the technical field to better understand the technical solutions in the present specification, the technical solutions in the present specification will be clearly and completely described below in combination with the drawings in the present specification. Obviously, the described embodiments are only some of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by the personnel in the field without making creative efforts should belong to the scope of protection of the present specification.
[0012] As Figures 1-3 , the method for determining ammonia in drinking water by the continuous flow injection method of the present embodiment can specifically include: S101, obtaining ammonia detection raw data of different water source types, removing abnormal values and noise interference through a data preprocessing module, establishing a standardized detection data set, recording operation parameters and environmental condition information of each detection method at the same time, obtaining a clean basic data matrix for subsequent analysis and processing.
[0013] Obtain ammonia detection raw data from multiple water source types, record data points using a sensor acquisition system, and obtain a raw data set. Analyze the raw data set through a preprocessing module, and if the detection data deviates from the preset threshold range, eliminate abnormal values to obtain a preliminary screening data set. Process the preliminary screening data set using a median filter algorithm to eliminate noise interference and obtain a denoising data set. Standardize the denoising data set according to the water source type and environmental conditions, and obtain a standardized data set using a z-score normalization method. Extract operation parameters from the detection method, combine environmental condition information, construct a parameter condition matrix, and obtain a detection context data set. Integrate the standardized data set and the detection context data set by a matrix merging method to generate a clean data matrix. Perform dimensionality reduction processing on the clean data matrix using a principal component analysis algorithm to extract main features and obtain a feature data set for subsequent analysis.
[0014] In one possible implementation, the acquisition of ammonia detection raw data can be performed by a sensor acquisition system, from different water source types such as surface water, groundwater, and industrial wastewater.
[0015] For example, the sampling point of surface water is set in the middle reaches of a river, the sampling point of groundwater is selected in a deep well, and the industrial wastewater is taken from the factory sewage outlet. The sensor records the ammonia concentration every minute, with the unit of mg / L, and continuously collects for 24 hours to obtain a raw data set containing 1440 data points. The ammonia concentration of surface water may fluctuate between 0.1-0.5 mg / L, the ammonia concentration of groundwater may fluctuate between 0.05-0.2 mg / L, and the ammonia concentration of industrial wastewater may be as high as 2-10 mg / L. This multi-source acquisition ensures that the data covers various water quality scenarios, improving the universality of subsequent analysis.
[0016] For example, when the preprocessing module performs outlier rejection on the raw data set, a threshold range of 0-15 mg / L can be set based on the statistical rules of historical water quality data. Assuming that an industrial wastewater data point is recorded as 50 mg / L, which is obviously deviated from the normal range, it may be a sensor failure or a transient pollution, and needs to be rejected. This process generates a preliminary screening data set, which retains about 95% of the valid data points, reduces the risk of misjudgment, and improves the reliability of the data.
[0017] Specifically, the median filter algorithm is used to eliminate noise interference. For example, with a filter window size of 5, for a certain data point, take the two data points before and after it, and calculate the median to replace the original value.
[0018] For example, a certain group of data is [0.3, 0.4, 10.0, 0.5, 0.2], the middle 10.0 is noise, and the median 0.4 is replaced to obtain the denoising data set. This method effectively smooths short-term fluctuations, retains data trends, and is suitable for sensor jitter scenarios in water quality monitoring.
[0019] In one embodiment, z-score normalization is used for standardization processing. Assuming that the mean of the ammonia concentration of surface water is 0.3 mg / L and the standard deviation is 0.1 mg / L, a data point of 0.5 mg / L is calculated as 2.0 after z-score, and the standardized data set is formed after uniform dimension. This method eliminates the magnitude difference of different water sources, facilitates cross-water source comparison, and improves the stability of model training.
[0020] For example, when constructing the parameter condition matrix, the operating parameters such as sensor sensitivity and sampling frequency are extracted, and the environmental conditions such as water temperature and pH value are combined.
[0021] For example, when sampling surface water, the water temperature is 25°C, the pH value is 7.5, and the sensor sensitivity is 0.01 mg / L, forming a matrix row vector. The combined clean data matrix integrates ammonia concentration and environmental information, completely describes the detection context, and facilitates the mining of the correlation between data.
[0022] Specifically, principal component analysis is used for dimensionality reduction. Assuming that the clean data matrix contains 10 variables such as ammonia concentration, water temperature, pH, etc., after analysis, the first 3 principal components explain 90% of the variance, and these features are retained to form a feature dataset. This method reduces computational complexity, highlights key variables, and improves the efficiency and accuracy of subsequent pollution prediction models.
[0023] It should be noted that the above process gradually improves data quality and availability from raw data to feature dataset through multi-level processing, providing reliable support for water quality monitoring and assisting environmental management decisions.
[0024] S102, according to the basic data matrix, the correlation characteristics between different detection methods are analyzed by using multivariate linear regression algorithm, the linear relationship coefficient and residual distribution pattern of each method detection result are calculated, the consistency level and deviation source between methods are judged, and the detection method category that needs to be corrected is determined.
[0025] The ammonia detection results are obtained from the clean data matrix, and the linear relationship coefficients between each detection method are calculated by using multivariate linear regression algorithm to obtain the coefficient matrix. For the coefficient matrix, the significance level of the linear relationship coefficient is analyzed, if the significance level is lower than the preset threshold, it is marked as a low correlation method, and a low correlation method set is obtained. From the low correlation method set, the residual distribution pattern is extracted, and the Kolmogorov-Smirnov test is used to judge whether the residual distribution deviates from the normal distribution to obtain a deviation distribution method list. According to the deviation distribution method list, the deviation characteristics are extracted, and the k-means clustering algorithm is used to group the deviation characteristics to obtain the deviation characteristic classification result. From the deviation characteristic classification result, the detection method corresponding to each group of deviation characteristics is obtained, and it is judged whether the method needs to be corrected. If the deviation characteristics deviate from the preset range, it is determined as a method that needs to be corrected, and a set of methods that need to be corrected is obtained. For the set of methods that need to be corrected, the corresponding operation parameters and environmental conditions are obtained, and the related information is extracted from the detection context dataset to construct a correction parameter matrix. The data of the methods that need to be corrected are adjusted by using the correction parameter matrix, the linear relationship coefficients of the adjusted data are recalculated, and the updated coefficient matrix is obtained.
[0026] For example, after obtaining ammonia detection results from the clean data matrix, a multiple linear regression algorithm can be used to analyze the linear relationship between different detection methods. Multiple linear regression calculates the linear relationship coefficients between methods by fitting the model to form a coefficient matrix. Suppose there are three detection methods: electrochemical method, colorimetric method, and ion selective electrode method, and the clean data matrix contains ammonia concentration data for each method. The electrochemical method data range is 0.2-0.8 mg / L, the colorimetric method is 0.3-1.0 mg / L, and the ion selective electrode method is 0.15-0.7 mg / L. After regression analysis, the coefficient matrix shows that the correlation coefficient between the electrochemical method and the colorimetric method is 0.85, the correlation coefficient between the colorimetric method and the ion selective electrode method is 0.45, and the correlation coefficient between the electrochemical method and the ion selective electrode method is 0.60. This indicates that the electrochemical method and the colorimetric method are highly correlated, while the correlation between other methods is low.
[0027] Specifically, when analyzing the significance level of the coefficient matrix, a significance threshold of 0.05 can be set. The p-value of each correlation coefficient is determined by t-test, and if the p-value is greater than 0.05, it is marked as a low correlation method. Suppose the p-value of the colorimetric method and the ion selective electrode method is 0.08, which exceeds the threshold, and is included in the low correlation method set. This set reflects that the data consistency between methods is poor, which may be caused by sensor characteristics or environmental interference, and needs further analysis.
[0028] In one embodiment, residual distribution patterns are extracted from the low correlation method set, and Kolmogorov-Smirnov test is used to determine whether the residuals conform to the normal distribution. Residuals are the difference between actual values and predicted values of the regression model.
[0029] For example, the residual data range of the colorimetric method and the ion selective electrode method is -0.3 to 0.3 mg / L, and the K-S test shows that the p-value is 0.02, which is less than 0.05, indicating that the residuals deviate from the normal distribution, and are included in the deviated distribution method list. This suggests that there is a systematic deviation between methods, which needs to be corrected to improve data consistency.
[0030] For example, for the deviated distribution method list, extract deviation features such as residual mean and variance, and use k-means clustering algorithm for grouping. Suppose the residual mean is 0.1 mg / L and the variance is 0.05 mg / L, and the clustering is divided into two groups: one group has small deviation, and the mean is close to 0; the other group has large deviation, and the mean deviates from 0.2 mg / L. Through the clustering result, the detection method whose deviation feature deviates from the preset range (for example, the mean exceeds 0.15 mg / L) is determined as the method that needs to be corrected, and is included in the set of methods that need to be corrected, such as the ion selective electrode method.
[0031] Specifically, for the set of methods to be corrected, the operation parameters such as sensor calibration frequency, sampling time, and environmental conditions such as water temperature, pH value are extracted to construct a correction parameter matrix. Assuming that the ion selective electrode method samples at a water temperature of 20°C, a pH of 7.2, and a calibration frequency of once a day, the matrix records this information. Using the correction parameter matrix, the data is adjusted by linear interpolation, for example, the ammonia concentration data of the ion selective electrode method is adjusted by 0.1 mg / L to compensate for system bias.
[0032] In one embodiment, the adjusted data recalculates the linear relationship coefficients and updates the coefficient matrix. After adjustment, the correlation coefficient between the colorimetric method and the ion selective electrode method is improved from 0.45 to 0.75, and the p value is reduced to 0.03, significantly enhancing the significance. This shows that the correction effectively improves the consistency of data between methods, providing a more reliable basis for subsequent water quality analysis.
[0033] S103, for the identified deviation detection method, a nonlinear mapping relationship model between methods is constructed by a support vector machine algorithm, complex method difference features are processed by a kernel function, a detection result conversion matrix under different water source conditions is established, and a standardized method correction parameter set is obtained.
[0034] Using the support vector machine algorithm, feature vectors are extracted from the deviation detection method, a nonlinear mapping model is constructed, and a feature representation of the difference between methods is obtained. According to the feature representation, a radial basis kernel function is applied to calculate the nonlinear relationship between the methods, and a preliminary detection result conversion matrix is generated. Environmental parameters are obtained from the water source condition database, and the preliminary conversion matrix is adjusted to obtain a detection result conversion matrix that adapts to the water source conditions. If the element value of the detection result conversion matrix deviates from the preset threshold, the support vector machine model parameters are adjusted by an iterative optimization algorithm to generate an optimized conversion matrix. The optimized conversion matrix is used to standardize the results of the deviation detection method to obtain corrected detection data. Consistency features are extracted from the corrected detection data, and the cosine similarity is used to calculate the matching degree of the corrected data and the standard data set to judge the correction effect. According to the correction effect, if the consistency features are lower than the preset threshold, additional environmental parameters are extracted from the water source condition database to update the conversion matrix and obtain the final correction parameter set.
[0035] For example, in the business scenario of ammonia detection, when the support vector machine algorithm is used to extract feature vectors, key features such as peak value, fluctuation frequency and background noise level of the detection result can be extracted from the data set of the deviation detection method. The support vector machine maps the feature vectors of different detection methods to a high-dimensional space by constructing a hyperplane to form a nonlinear mapping model.
[0036] Exemplarily, assuming that there are three ammonia detection methods A, B, and C, and the detection results are 10.5 mg / L, 11.2 mg / L, and 9.8 mg / L respectively, the characteristic vector can include the absolute value difference of the results, the time series volatility rate, and the like. The support vector machine uses these characteristics to generate a model capable of distinguishing the differences between methods, and then constructs a preliminary detection result conversion matrix. Such a matrix can map the detection values of different methods to a unified reference framework, which is helpful for subsequent correction.
[0037] In a possible implementation, when the radial basis kernel function is used to calculate the nonlinear relationship between methods, the bandwidth parameter σ can be selected as 0.5 to balance the complexity and generalization ability of the model.
[0038] For example, the difference between the detection results of methods A and B is 0.7 mg / L, and the nonlinear similarity thereof is calculated by the radial basis kernel function to generate an initial element value of the conversion matrix, such as 0.85. Such a matrix reflects the relative consistency between methods, which is convenient for subsequent adjustment. Environmental parameters such as temperature 20°C and pH value 7.2 are extracted from the water source condition database, and the element value is adjusted in combination with the preliminary conversion matrix.
[0039] For example, if high temperature causes the detection value of method A to be too high, the weight thereof can be reduced by adjusting the matrix parameters to generate a conversion matrix adapted to the water source conditions. In this way, environmental changes can be effectively adapted.
[0040] Specifically, if the element value of the conversion matrix deviates from the preset threshold value 0.9, the support vector machine model parameters need to be adjusted by an iterative optimization algorithm.
[0041] For example, the gradient descent method is used to optimize the penalty coefficient C of the support vector machine, and the initial value 1.0 is adjusted to 1.5 to generate an optimized conversion matrix. The optimized matrix can correct the detection value of method A from 10.5 mg / L to 10.2 mg / L, which is closer to the standard value. Such correction improves the reliability of the detection result. The corrected detection data can be further extracted for consistency features such as mean and variance, and the cosine similarity is used to calculate the matching degree with the standard data set.
[0042] For example, the cosine similarity of the corrected data is 0.95, which is higher than the threshold value 0.9, indicating that the correction effect is good.
[0043] For example, if the consistency features are lower than the threshold value, additional parameters such as water turbidity 50 NTU can be extracted from the water source condition database to update the conversion matrix.
[0044] For example, by increasing the turbidity-related weight factor, the matrix element value is adjusted to 0.92, and the final correction parameter set is generated. This set can be used for subsequent parameter configuration of the detection device, ensuring more stable detection results. This multi-level correction method, through the combination of dynamic adjustment of environmental parameters and model optimization, can effectively improve the consistency and accuracy of ammonia detection, providing reliable support for water quality monitoring.
[0045] In S104, the original detection data is batch corrected using the correction parameter set, and the systematic deviation between methods is eliminated through matrix operation to generate ammonia concentration detection results under a unified standard. At the same time, the data change amplitude and distribution characteristics before and after correction are calculated to judge the effectiveness of the correction effect.
[0046] The original detection data is executed by the correction parameter set to perform matrix operation, eliminating systematic deviation and generating standardized ammonia concentration detection results. According to the standardized detection results, the change amplitude of the data before and after correction is calculated, and the statistical analysis method is used to extract the mean and variance of the data to obtain the change amplitude characteristics. If the change amplitude characteristics exceed the preset threshold, the environmental parameters are obtained from the water source environment database, the correction parameter set is adjusted, and the updated correction matrix is generated. The standardized detection results are corrected again through the updated correction matrix to generate optimized ammonia concentration data. The matching degree of the optimized ammonia concentration data and the standard data set is calculated using cosine similarity to obtain a consistency index. If the consistency index is lower than the preset threshold, the statistical characteristics of the optimized ammonia concentration data are extracted, combined with the environmental parameters, and the correction parameter set is updated to generate a final correction matrix. The optimized ammonia concentration data is processed through the final correction matrix to obtain the final standardized detection results.
[0047] For example, in the business scenario of ammonia concentration detection, the original detection data is subjected to matrix operation through the correction parameter set, aiming to eliminate systematic deviation. The original detection data may be affected by detection equipment or environmental factors, such as instrument sensitivity differences or water temperature fluctuations, resulting in inconsistent results. The correction parameter set is usually represented in matrix form and contains weight factors for different detection methods.
[0048] For example, assuming there are three ammonia detection methods A, B, and C, with original detection results of 10.8 mg / L, 11.5 mg / L, and 10.1 mg / L, the correction matrix can be generated through historical data training, containing bias adjustment coefficients for each method. After matrix operation, the standardized results may be adjusted to 10.3 mg / L, 10.4 mg / L, and 10.2 mg / L, significantly reducing the difference between methods. This way ensures that the detection results are more comparable through a unified reference framework.
[0049] In one possible implementation, the amplitude of change of the data before and after correction is calculated, and the mean and variance are extracted as features. The mean of the data before correction is 10.8 mg / L, and the variance is 0.49; the mean after correction is 10.3 mg / L, and the variance is 0.01. The change amplitude feature shows that the reduction in variance indicates an improvement in data consistency. If the change amplitude exceeds a preset threshold, such as a variance greater than 0.1, parameters such as water temperature 22°C and pH value 7.5 are extracted from the water source environment database. These parameters may cause detection bias, for example, high temperature may cause the detection value of method A to be high. By adjusting the weight factor in the correction matrix, such as reducing the weight of method A to 0.1, an updated correction matrix is generated to ensure that the results are closer to the true value.
[0050] For example, during secondary correction, the updated correction matrix further processes the normalized data to generate optimized ammonia concentration data, such as 10.25 mg / L, 10.3 mg / L, and 10.2 mg / L. The cosine similarity is used to evaluate the matching degree of the optimized data and the standard data set. Assuming that the ammonia concentration of the standard data set is 10.2 mg / L, the cosine similarity of the optimized data is 0.96, which is higher than the threshold value 0.9, indicating good consistency. If the consistency index is lower than the threshold value, such as 0.85, the statistical features of the optimized data, such as mean 10.25 mg / L and variance 0.005, are extracted, and the environmental parameters such as turbidity 40 NTU are combined to update the correction parameter set. Turbidity may affect the accuracy of optical detection methods, so the weight factor related to turbidity is increased, and the matrix element value is adjusted to 0.93 to generate the final correction matrix.
[0051] In one possible implementation, the final correction matrix processes the optimized data to obtain the final standardized detection results, such as 10.22 mg / L, 10.23 mg / L, and 10.21 mg / L. Such results are highly consistent and suitable for water quality monitoring equipment configuration to ensure detection stability. Through dynamic adjustment of environmental parameters and multi-level correction of matrix operations, the data reliability is significantly improved, providing accurate support for water quality monitoring. This method forms a rigorous logic chain through feature extraction and consistency evaluation to ensure that the correction process adapts to complex environmental changes.
[0052] S105、According to the corrected detection result data, the key factor weight affecting the detection accuracy is analyzed by the random forest algorithm, the contribution of variables such as water source type, interference substance concentration, and operation condition to the detection accuracy is identified, and the key monitoring indicators and threshold range of quality control are determined.
[0053] The corrected ammonia concentration data is analyzed for feature importance by a random forest algorithm, the contribution weights of water source type, interferent concentration, and operating condition to detection accuracy are calculated, and the variable weight ranking is obtained. According to the variable weight ranking, the variable with the highest weight is extracted from the water source type, the interferent concentration, and the operating condition to generate a list of key influencing factors. Through the list of key influencing factors, combined with the preset quality control standard, the priority monitoring indicators are generated, and the threshold range of each indicator is determined. If the weight value of the key influencing factor is lower than the preset threshold, the environmental parameters are obtained from the water source environment database, the feature input of the random forest model is adjusted, the variable weight is recalculated, and the updated key influencing factors are obtained. The updated key influencing factors are used to generate optimized monitoring indicators and threshold ranges. The support vector machine algorithm is used to classify the optimized monitoring indicators to determine whether the indicators meet the preset threshold range, and the classification result is obtained. According to the classification result, the environmental parameters related to the indicators not meeting the threshold range are extracted from the water source environment database, the threshold range of the monitoring indicators is updated, and the final quality control strategy is obtained.
[0054] For example, in the business scenario of ammonia concentration detection, the corrected ammonia concentration data is analyzed for feature importance by a random forest algorithm, aiming to identify key factors affecting detection accuracy. The random forest algorithm builds multiple decision trees to evaluate the contribution of each feature to the prediction result and generates a feature importance score. Assuming that the analyzed data set contains features such as water source type, interferent concentration, and operating condition, the algorithm calculation result shows that the water source type weight is 0.45, the interferent concentration weight is 0.35, and the operating condition weight is 0.2. This weight ranking indicates that the water source type has the greatest impact on detection accuracy.
[0055] Specifically, according to the weight ranking, the water source type with the highest weight is extracted as the key influencing factor.
[0056] For example, analysis shows that the ammonia concentration detection bias of groundwater is larger than that of surface water, possibly due to the high mineral content in groundwater. By generating a list of key influencing factors, combined with the quality control standard, the priority monitoring indicators are determined, such as setting "groundwater source" in the water source type as the primary indicator, with a threshold range of mineral content below 200 mg / L. If the weight value is lower than the preset threshold of 0.3, for example, the operating condition weight is 0.2, the parameters are extracted from the water source environment database, such as water temperature 25°C and turbidity 30 NTU, the feature input of the random forest model is adjusted, the weight is recalculated, and the updated key influencing factors are obtained, such as water source type 0.48, interferent concentration 0.36, and operating condition 0.16, to generate a new list of key influencing factors.
[0057] In one possible implementation, based on updated key influencing factors, the monitoring indicators are optimized to be groundwater source and turbidity, with threshold ranges of 180-220 mg / L for mineral content and 20-40 NTU for turbidity, respectively. A support vector machine algorithm is then used to classify these indicators and determine whether the thresholds are met.
[0058] For example, in one test, the mineral content was 230 mg / L, exceeding the threshold, and the classification result was unqualified. Relevant parameters were extracted from the database, revealing that high turbidity (40 NTU) might cause optical detection bias. Therefore, the turbidity threshold was updated to 20-35 NTU, forming the final quality control strategy. This strategy ensures more reliable test results by dynamically adjusting the threshold.
[0059] For example, for unqualified classification results, parameters such as turbidity (40 NTU) and pH (7.8) were extracted. Analysis revealed that a high pH value might amplify the influence of interfering substances. When updating monitoring indicators, the pH threshold range was increased to 7.0-7.5 to ensure more accurate classification results. This multi-level analysis and adjustment approach, through feature importance ranking, dynamic threshold optimization, and classification verification, significantly improves the stability of water quality monitoring and supports accurate detection.
[0060] S106. Based on the key monitoring indicators, establish a data consistency evaluation system between laboratories. By calculating the standard deviation and coefficient of variation of the test results of each laboratory, a consistency score matrix is generated. If the score is lower than the preset threshold, the data correction process is triggered to obtain a final test report that meets the standard requirements.
[0061] Based on laboratory test results, the standard deviation and coefficient of variation of each laboratory's data are calculated to generate a consistency score matrix. If any score in the consistency score matrix is lower than a preset threshold, the original data of the relevant test results are retrieved from the laboratory database and processed using a data correction process to obtain a corrected dataset. Based on the corrected dataset, the standard deviation and coefficient of variation are recalculated, and the consistency score matrix is updated to obtain an updated score result. Using the updated score result, the K-means clustering algorithm is used to classify the laboratory test results to determine whether each laboratory's data meets the quality control standards, obtaining a classification result. If the classification result shows that there is laboratory data that does not meet the quality control standards, relevant environmental parameters are extracted from the laboratory database, and the parameter settings of the data correction process are adjusted to obtain an optimized corrected dataset. Based on the optimized corrected dataset, the final consistency score matrix is generated to determine the compliance status of each laboratory's test results. Using the final consistency score matrix, a decision tree algorithm is used to conduct a source analysis on the non-compliant laboratory data to determine the source factors of the abnormal data, obtaining the source analysis result.
[0062] Specifically, in the field of water quality monitoring, consistency analysis of laboratory test results is crucial to ensure data reliability.
[0063] For example, by calculating the standard deviation and coefficient of variation of each laboratory data, the dispersion degree of data can be quantified, so as to evaluate the stability of test results. The standard deviation reflects the fluctuation amplitude of data, and the coefficient of variation eliminates the dimension influence by dividing the standard deviation by the mean value, which is convenient for comparison between different laboratories. Assuming that a laboratory detects the concentration of ammonia nitrogen in water, the mean value of 10 test results is 5 mg / L, the standard deviation is 0.5 mg / L, and the coefficient of variation is 0.1. If the preset coefficient of variation threshold is 0.15, the data of this laboratory preliminarily meets the consistency requirement. Based on this, a consistency score matrix is generated, and the score is based on the weighted calculation of the reciprocal of the coefficient of variation and the deviation of the mean value. Assuming that the scores of laboratories A, B and C are 0.9, 0.7 and 0.85. If the score of laboratory B is lower than the threshold value 0.8, the original data of laboratory B needs to be extracted from the database, such as detection time, instrument calibration state, and data correction process is adopted, such as removing outliers or correcting instrument deviation, and then the standard deviation is recalculated as 0.4 mg / L, the coefficient of variation is 0.08, and the consistency score is updated to 0.88.
[0064] Specifically, the updated score matrix is used for K-means clustering analysis, and the laboratories are divided into two categories of “high consistency” and “low consistency”. Assuming that the clustering result shows that laboratory B still belongs to low consistency, the environmental parameters such as water temperature 30°C and turbidity 50 NTU during detection are extracted, and it is found that high turbidity may interfere with optical detection. Adjust the correction process and add a turbidity compensation model, and the coefficient of variation of the optimized data set is reduced to 0.07, the score is updated to 0.9, and it enters the high consistency category. The final consistency score matrix reflects the compliance status of each laboratory, and laboratory B meets the standard.
[0065] In one possible implementation, decision tree traceability analysis is performed on laboratory data that does not meet the standard.
[0066] For example, the abnormal data of laboratory B may be caused by high turbidity or instrument not calibrated in time. The decision tree analysis shows that when the turbidity exceeds 40 NTU, the detection deviation probability reaches 70%. The traceability result shows that the turbidity needs to be monitored first and the instrument needs to be calibrated regularly to ensure data reliability. This multi-level analysis significantly improves the stability of water quality detection through consistency score, clustering classification and traceability analysis.
[0067] The preferred embodiments of the application disclosed above are only to facilitate the elucidation of the application. The preferred embodiments do not describe all the details of the application and limit the application to the specific embodiments. Obviously, many modifications and variations can be made in light of the teachings above. The description is chosen and described in order to provide the best illustration of the application and its practical application to those skilled in the art and to enable those skilled in the art to best utilize the application. The application is limited only by the claims and their full scope and equivalents.
Claims
1. Method for the determination of ammonia in drinking water by continuous flow injection, characterised in that, The method comprises: Obtain ammonia detection raw data of different water source types, remove outliers and noise interference through a data preprocessing module, establish a standardized detection data set, record the operating parameters and environmental condition information of each detection method, and obtain a clean basic data matrix for subsequent analysis and processing; According to the basic data matrix, the correlation characteristics between different detection methods are analyzed by using a multivariate linear regression algorithm, the linear relationship coefficient and residual distribution pattern of the detection results of each method are calculated, the consistency level and deviation source between methods are judged, and the detection method category that needs to be corrected is determined; For the identified deviation detection method, a nonlinear mapping relationship model between methods is constructed by using a support vector machine algorithm, the kernel function is used to process the complex method difference characteristics, a detection result conversion matrix under different water source conditions is established, and a standardized method correction parameter set is obtained; The correction parameter set is used to correct the original detection data in batches, the systematic deviation between methods is eliminated through matrix operation, the ammonia concentration detection results under the unified standard are generated, the data change amplitude and distribution characteristics before and after correction are calculated, and the effectiveness of the correction effect is judged; According to the corrected detection result data, the key factor weight affecting the detection accuracy is analyzed by using a random forest algorithm, the contribution of water source type, interference substance concentration, operation condition and other variables to the detection accuracy is identified, and the key monitoring index and threshold range of quality control are determined; Based on the key monitoring index, a data consistency evaluation system between laboratories is established, the standard deviation and coefficient of variation of the detection results of each laboratory are calculated, a consistency score matrix is generated, if the score is lower than the preset threshold, the data correction process is triggered, and the final detection report meeting the standard requirements is obtained.
2. The method for determining ammonia in drinking water by continuous flow injection according to claim 1, characterized in that, The method comprises: Obtain ammonia detection raw data of different water source types, remove outliers and noise interference through a data preprocessing module, establish a standardized detection data set, record the operating parameters and environmental condition information of each detection method, and obtain a clean basic data matrix for subsequent analysis and processing, comprising: Obtain ammonia detection raw data from multiple water source types, record data points using a sensor collection system, and obtain an original data set; Through the preprocessing module, the original data set is analyzed, if the detection data deviates from the preset threshold range, the outliers are removed, and a preliminary screening data set is obtained; The preliminary screening data set is processed by using a median filter algorithm to eliminate noise interference, and a denoising data set is obtained; The denoising data set is standardized according to the water source type and environmental conditions, and a z-score normalization method is used to obtain a standardized data set; Extract the operating parameters from the detection method, combine the environmental condition information, construct a parameter condition matrix, and obtain a detection context data set; The standardized data set and the detection context data set are integrated by using a matrix merging method to generate a clean data matrix; The clean data matrix is processed by using a principal component analysis algorithm to reduce the dimension, extract the main features, and obtain a feature data set for subsequent analysis.
3. The method for determination of ammonia in drinking water by continuous flow injection according to claim 1, characterized in that, The method comprises the following steps of: acquiring ammonia detection results from the clean data matrix, calculating linear relationship coefficients between detection methods by using a multivariate linear regression algorithm, and obtaining a coefficient matrix; analyzing a significance level of the linear relationship coefficients for the coefficient matrix, and marking as a low correlation method if the significance level is lower than a preset threshold, to obtain a low correlation method set; extracting a residual distribution pattern from the low correlation method set, and judging whether the residual distribution deviates from a normal distribution by using a Kolmogorov-Smirnov test to obtain a distribution deviation method list; extracting deviation features from the distribution deviation method list, grouping the deviation features by using a k-means clustering algorithm, and obtaining a deviation feature classification result; acquiring detection methods corresponding to each group of deviation features from the deviation feature classification result, judging whether the methods need to be corrected, and determining as a method needing to be corrected if the deviation features deviate from a preset range to obtain a method set needing to be corrected; acquiring operation parameters and environmental conditions corresponding to the method set needing to be corrected, extracting related information from a detection context data set, and constructing a correction parameter matrix; adjusting data of the method set needing to be corrected by using the correction parameter matrix, recalculating linear relationship coefficients of the adjusted data, and obtaining an updated coefficient matrix. The method comprises the following steps of: constructing a nonlinear mapping relationship model between methods by using a support vector machine algorithm, processing complex method difference features by using a kernel function, establishing a detection result conversion matrix under different water source conditions, and obtaining a standardized method correction parameter set, comprising: extracting feature vectors from the deviation detection method by using a support vector machine algorithm, constructing a nonlinear mapping model, and obtaining a feature representation of the method difference; calculating a nonlinear relationship of the method difference by using a radial basis kernel function according to the feature representation, and generating a preliminary detection result conversion matrix; acquiring environmental parameters from a water source condition database, adjusting matrix parameters in combination with the preliminary conversion matrix, and obtaining a detection result conversion matrix adapted to the water source condition; if element values of the detection result conversion matrix deviate from a preset threshold, adjusting support vector machine model parameters by using an iterative optimization algorithm to generate an optimized conversion matrix; standardizing results of the deviation detection method by using the optimized conversion matrix, and obtaining corrected detection data; extracting consistency features from the corrected detection data, calculating a matching degree of the corrected data and a standard data set by using a cosine similarity, and judging a correction effect; if the consistency features are lower than a preset threshold according to the correction effect, extracting additional environmental parameters from the water source condition database, updating the conversion matrix, and obtaining a final correction parameter set. 4. The method for determination of ammonia in drinking water by continuous flow injection according to claim 1, characterized in that, 5. The method for determination of ammonia in drinking water by continuous flow injection according to claim 1, characterized in that, The batch correction processing of the original detection data is performed by using the correction parameter set, the systematic deviation between methods is eliminated through matrix operation, the ammonia concentration detection result under the unified standard is generated, the data change amplitude and distribution characteristics before and after correction are calculated, and the effectiveness degree of the correction effect is judged, including: The matrix operation is performed on the original detection data by using the correction parameter set, the systematic deviation is eliminated, and the standardized ammonia concentration detection result is generated; According to the standardized detection result, the change amplitude of the data before and after correction is calculated, the statistical analysis method is used to extract the mean and variance of the data, and the change amplitude characteristics are obtained; If the change amplitude characteristics exceed the preset threshold, the environmental parameters are obtained from the water source environment database, the correction parameter set is adjusted, and the updated correction matrix is generated; The standardized detection result is corrected again by using the updated correction matrix, and the optimized ammonia concentration data is generated; The matching degree of the optimized ammonia concentration data and the standard data set is calculated by using the cosine similarity, and the consistency index is obtained; If the consistency index is lower than the preset threshold, the statistical characteristics of the optimized ammonia concentration data are extracted, the environmental parameters are combined, the correction parameter set is updated, and the final correction matrix is generated; The final standardized detection result is obtained by processing the optimized ammonia concentration data through the final correction matrix.
6. The method for determination of ammonia in drinking water by continuous flow injection according to claim 1, characterized in that, According to the corrected detection result data, the key factor weight affecting the detection accuracy is analyzed by using the random forest algorithm, the contribution degree of water source type, interference substance concentration and operation condition to the detection accuracy is identified, the key monitoring index and threshold range of quality control are determined, including: The feature importance analysis is performed on the corrected ammonia concentration data by using the random forest algorithm, the contribution weight of water source type, interference substance concentration and operation condition to the detection accuracy is calculated, and the variable weight order is obtained; According to the variable weight order, the variable with the highest weight is extracted from the water source type, the interference substance concentration and the operation condition, and the key influence factor list is generated; Through the key influence factor list, the priority monitoring index is generated by combining the preset quality control standard, and the threshold range of each index is determined; If the weight value of the key influence factor is lower than the preset threshold, the environmental parameters are obtained from the water source environment database, the feature input of the random forest model is adjusted, the variable weight is recalculated, and the updated key influence factor is obtained; The optimized monitoring index and threshold range are generated by using the updated key influence factor; The support vector machine algorithm is used for classification of the optimized monitoring index, and whether each index meets the preset threshold range is judged, and the classification result is obtained; According to the classification result, the environmental parameters related to the index not meeting the threshold range are extracted from the water source environment database, the threshold range of the monitoring index is updated, and the final quality control strategy is obtained.
7. The method for determination of ammonia in drinking water by continuous flow injection according to claim 1, characterized in that, Based on the key monitoring index, a laboratory data consistency evaluation system is established, the standard deviation and coefficient of variation of the detection results of each laboratory are calculated, a consistency score matrix is generated, if the score is lower than the preset threshold, a data correction process is triggered, and a final detection report meeting the standard requirements is obtained, including: The standard deviation and the coefficient of variation of each laboratory data are calculated based on the laboratory test results to generate a consistency score matrix; If any score in the consistency score matrix is lower than a preset threshold, the original data of the relevant test results are obtained from the laboratory database, and the data correction process is used for processing to obtain a corrected data set; The standard deviation and the coefficient of variation are recalculated based on the corrected data set, the consistency score matrix is updated, and an updated score result is obtained; The laboratory test results are classified by using the K-means clustering algorithm based on the updated score result, whether the laboratory data conforms to the quality control standard is judged, and a classification result is obtained; If the classification result shows that there is laboratory data that does not conform to the quality control standard, the relevant environmental parameters are extracted from the laboratory database, the parameter settings of the data correction process are adjusted, and an optimized corrected data set is obtained; A final consistency score matrix is generated based on the optimized corrected data set to determine the conformity state of the laboratory test results; The source factors of the abnormal data are determined by using the decision tree algorithm to perform traceability analysis on the laboratory data that does not conform to the standard based on the final consistency score matrix, and a traceability result is obtained.
Citation Information
Cited By
Crude polysaccharide content determination method and system based on multi-standard comparison
CN122135815A