Online dynamic water fingerprint rapid acquisition and pretreatment device and method
By optimizing the water quality monitoring flow path through graph theory modeling and artificial intelligence technology, the problems of low efficiency, poor accuracy, and poor robustness in existing technologies have been solved, achieving efficient and accurate real-time online water quality monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU EDUCATION UNIV
- Filing Date
- 2026-02-27
- Publication Date
- 2026-04-21
AI Technical Summary
Existing water quality monitoring methods lack system-level optimization, and the integration of artificial intelligence models with online monitoring hardware is insufficient, resulting in low efficiency, poor accuracy, and poor robustness.
Intelligent flow path selection is achieved using graph theory modeling, optical window contamination is predicted by temporal convolutional networks, dilution decision-making is performed by deep neural networks, dilution control is optimized, matrix flip graph search is used to accelerate spectral calculation, and redundant flow paths are configured to improve system fault tolerance.
Monitoring efficiency is improved by 50%, detection accuracy by 70%, calculation speed by 5-10 times, system robustness by 80%, and user experience and operational consistency are greatly improved.
Smart Images

Figure CN121740566B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water quality monitoring technology, specifically to an online dynamic water quality fingerprint rapid acquisition and preprocessing method and device. Background Technology
[0002] The importance of water quality monitoring is increasingly prominent, and it currently relies primarily on laboratory analysis or online optical detection equipment. In recent years, machine learning, particularly deep learning, has made breakthroughs in spectral analysis, time-series prediction, and intelligent control, providing new technological pathways for intelligent water quality monitoring. Deep neural networks possess powerful feature extraction and nonlinear modeling capabilities, enabling them to automatically learn water quality feature representations from raw spectral data; recurrent neural networks and temporal convolutional networks can effectively model time-series data, achieving trend prediction and anomaly detection; reinforcement learning algorithms can learn optimal decision-making strategies through interaction with the environment.
[0003] However, existing research focuses on offline data analysis and lacks system-level solutions that deeply integrate artificial intelligence models with online monitoring hardware. In other words, existing methods usually optimize individual decision-making processes in isolation and lack a globally optimal system-level optimization scheme. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide an online dynamic water quality fingerprint rapid acquisition and preprocessing method. It achieves intelligent flow path selection through graph theory modeling, optimizes dilution control using optimal transport theory, and accelerates spectral calculation using a matrix flip graph search algorithm, thereby improving the efficiency, accuracy and robustness of water quality monitoring.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] A method for rapid online dynamic water quality fingerprint acquisition and preprocessing includes the following steps: intelligently selecting the test flow path from multiple inlet flow paths; determining the contamination level of an optical window based on a light intensity attenuation trend prediction model, and cleaning the optical window when the predicted contamination level exceeds a dynamic threshold, wherein the light intensity attenuation trend prediction model is a temporal convolutional network model; detecting the water sample of the test flow path in the measurement channel corresponding to the cleaned optical window, and inputting the turbidity detection results, colorimetric detection results, and historical dilution effect evaluation data into a dilution decision model, wherein the dilution decision model adopts a deep neural network architecture; diluting the water sample of the test flow path according to the dilution parameters output by the dilution decision model; acquiring the spectral data of the diluted water sample to obtain water quality fingerprint spectral data, and calculating water quality parameters based on the water quality fingerprint spectral data to complete the acquisition and preprocessing of the water quality fingerprint.
[0007] Further, the intelligent selection of the test flow path from multiple inlet flow paths includes: acquiring the physical structure of the entire water quality monitoring system; modeling the entire water quality monitoring system as a first directed weighted graph, where vertices represent key components including sampling points, filters, pumps, mixers, and detectors, edges represent pipelines connecting the key components, and the weight of the edges represents flow time, energy consumption, or failure risk; and outputting a system topology graph model corresponding to the entire water quality monitoring system; based on the system topology graph model, applying a depth-first search algorithm to traverse the system topology graph model from the starting vertex to identify all cut edges, and decomposing the system topology graph model into multiple bilaterally connected components by deleting all identified cut edges, obtaining the bilaterally connected component decomposition result; and identifying the bilaterally connected components... The connection relationships between each bilaterally connected component in the edge-connected component decomposition result are used to determine the potential single-point fault location, and redundant flow paths are configured at the potential single-point fault location. After completing the design of the redundant flow paths, the water sampling problem is formalized into a modified Traveling Salesman Problem (TSP), and a distance matrix is constructed, where the elements of the distance matrix represent the comprehensive cost from one component to another. The cutting plane method is applied to solve the linear programming relaxation problem in the distance matrix to obtain an initial lower bound solution. The branch and bound method is used to gradually reduce the integer gap of the initial lower bound solution until the integer gap is less than a preset threshold to obtain the optimal sampling path. Based on the path switching logic in the optimal sampling path, the opening and closing of the solenoid valves in the multi-flow path structure with redundant design is controlled to realize the flow path selection and determine the flow path to be tested.
[0008] Further, the step of controlling the opening and closing of solenoid valves in a multi-flow-path structure with redundant design to select the flow path and determine the flow path to be tested, based on the path switching logic in the optimal sampling path, includes: monitoring the operating status parameters and performance parameters of each component in the multi-flow-path structure with redundant design; determining whether the operating status parameters deviate from the normal operating parameter threshold range or whether the performance parameters are lower than a preset performance threshold, in order to detect abnormal situations such as pipeline blockage or pump performance degradation, and generating a system abnormal status signal; based on the system abnormal status signal, marking the abnormal components indicated by the system abnormal status signal as unavailable from the system topology model through a dynamic path replanning algorithm, and recalculating the optimal path from the current state to the target state, automatically starting alternative flow paths or recalculating the optimal sampling path to obtain an updated sampling path scheme; based on the updated sampling path scheme, first closing the solenoid valve of the current flow path corresponding to the optimal sampling path scheme, then opening the solenoid valve of the path corresponding to the updated sampling path scheme and controlling the flow rate to be gradually adjusted during the switching process to determine the dynamically optimized flow path to be tested.
[0009] Further, the process of diluting the water sample of the flow path to be tested according to the dilution parameters output by the dilution decision model includes: acquiring the raw water spectral data of the water sample of the flow path to be tested; representing the raw water spectral data as the multidimensional spectral distribution characteristics before dilution; representing the spectral characteristics under ideal measurement conditions as the target spectral distribution characteristics after dilution; defining a cost matrix for transforming the multidimensional spectral distribution characteristics before dilution to the target spectral distribution characteristics after dilution, wherein each element of the cost matrix represents the transmission cost required to transform each spectral feature before dilution into the corresponding spectral feature after dilution; and applying a linear space hypergradient algorithm to iteratively solve the optimal transmission strategy of the cost matrix. The process involves initializing dual variables and calculating hypergradient values in each iteration. The dual variables are then updated according to a step size rule until the gradient norm is less than a preset gradient norm threshold, yielding the optimal transport strategy. A mapping function is established from the spectral distribution before and after dilution to the physical dilution ratio. The optimal transport strategy is substituted into the mapping function to calculate the actual dilution ratio, resulting in the optimal dilution ratio. Based on the optimal dilution ratio and the target total flow rate, the values of the raw water flow rate and the dilution water flow rate are calculated to obtain dilution control parameters. According to the dilution control parameters, the raw water flow rate and the dilution water flow rate are controlled via a proportional valve or peristaltic pump to dilute the water sample in the test flow path, resulting in a diluted water sample.
[0010] Further, the calculation of water quality parameters based on the water quality fingerprint spectral data includes: applying principal component analysis to calculate the covariance matrix and decompose the eigenvalues of the original spectral data matrix corresponding to the water quality fingerprint spectral data, identifying spectral feature vectors, and reconstructing the original spectral data matrix in a structured manner according to the spectral feature vectors to obtain a structured spectral data matrix; based on the structured spectral data matrix, by abstracting the structured spectral data matrix into a second directed weighted graph, wherein the vertices of the second directed weighted graph represent specific elements or blocks in the structured spectral data matrix, and the edges of the second directed weighted graph represent the computational dependencies between the specific elements or blocks, defining a set of matrix flipping operations, and constructing a flipping operation cost model to evaluate the computational cost-benefit ratio of each flip in the set of matrix flipping operations, to obtain a flipped graph model of the spectral data; and designing a heuristic based on the flipped graph model of the spectral data. The search function evaluates the estimated cost of moving from the current state to the target state of the flip graph model of the spectral data. A priority queue is used to manage the flip operation sequence to be explored, and a pruning strategy is developed to eliminate search branches with non-optimal solutions, resulting in an optimized flip operation sequence. Based on the optimized flip operation sequence, key matrix operation patterns in water quality fingerprint analysis are identified, and corresponding flip operation sequence templates are designed for each operation pattern to accelerate the calculation of key matrices in the water quality fingerprint analysis. The accelerated matrix calculation results are used to optimize memory access, reduce cache miss rate, and optimize single instruction multiple data stream instruction set. Multiple data elements are processed in parallel using an embedded processor to obtain real-time spectral analysis data. Based on the spectral characteristics of the real-time spectral analysis data, pollutant concentration, total organic carbon, and chemical oxygen demand are calculated using a preset water quality parameter calculation model to obtain the water quality parameters.
[0011] Furthermore, after calculating the pollutant concentration, total organic carbon, and chemical oxygen demand based on the spectral characteristics of the real-time processed spectral analysis data using a preset water quality parameter calculation model to obtain the water quality parameters, the method further includes: acquiring the real-time processed spectral analysis data, the operating status parameters of the water quality monitoring system, control command execution feedback, and abnormal event alarm information; organizing the real-time processed spectral analysis data, the operating status parameters, the control command execution feedback, and the abnormal event alarm information by timestamp and storing them in an operating record database to obtain an operating record; extracting data within a preset time window from the operating record and performing outlier removal, missing value imputation, and data smoothing preprocessing to obtain preprocessed operating data; and displaying the current values of the water quality parameters in a digital panel format by plotting the real-time processed spectral analysis data as a spectral curve based on the preprocessed operating data.
[0012] Further, based on the system topology graph model, the depth-first search algorithm is applied to traverse the system topology graph model from the starting vertex to identify all cut edges, and the system topology graph model is decomposed into multiple bilaterally connected components by deleting all the identified cut edges to obtain the bilaterally connected component decomposition result. This includes: based on the system topology graph model, by initializing the depth-first search traversal process to record the visit timestamp and minimum reach timestamp of each vertex, recursively traversing all vertices and edges of the system topology graph model starting from the starting vertex to obtain the traversal result; according to the traversal result, by judging the relationship between the minimum reach timestamps of the two endpoints of an edge, when the minimum reach timestamp of one endpoint of an edge is greater than the visit timestamp of the other endpoint, the edge is marked as a cut edge, obtaining a set of all cut edges; based on the set of all cut edges, by deleting all cut edges from the set of all cut edges in the system topology graph model, the system topology graph model is decomposed into multiple connected subgraphs, wherein there are at least two non-repeating paths between any two vertices in each connected subgraph, obtaining the bilaterally connected component decomposition result.
[0013] Furthermore, the application of the linear space hypergradient algorithm to iteratively solve the optimal transport strategy for the cost matrix, initializing the dual variables and calculating the hypergradient value in each iteration, and updating the dual variables according to the step size rule until the gradient norm is less than a preset gradient norm threshold, to obtain the optimal transport strategy, includes: based on the cost matrix, formalizing the optimal transport problem into a linear programming problem, setting constraints as the marginal distribution constraints of the multidimensional spectral distribution characteristics before dilution and the marginal distribution constraints of the target spectral distribution characteristics after dilution, with the objective of minimizing the weighted sum of the cost matrix and the transport plan, to obtain the linear programming model of optimal transport; based on the linear programming model of optimal transport, initializing the dual variables as zero vectors or random vectors, setting initial step size parameters and convergence thresholds. The initial dual variable is obtained by calculating the hypergradient value and determining the adaptive step size in each iteration based on the cost matrix and the current dual variable. The dual variable is then updated based on the adaptive step size and the hypergradient value to obtain the updated dual variable. The gradient norm corresponding to the updated dual variable is calculated, and it is determined whether the gradient norm is less than the gradient norm threshold. When the gradient norm is less than the gradient norm threshold, the iteration is stopped to obtain the converged dual variable. Based on the converged dual variable, the original optimal transmission strategy matrix is calculated from the converged dual variable, where each element of the original optimal transmission strategy matrix represents the optimal transmission amount from the spectral features before dilution to the spectral features after dilution to obtain the optimal transmission strategy.
[0014] This invention also provides an online dynamic water quality fingerprint rapid acquisition and pretreatment device. The device is used to execute the above-mentioned method and includes: a multi-flow-path sampling unit, equipped with a raw water direct monitoring flow path, an automatic dilution monitoring flow path, and a standard sample calibration flow path, each flow path being connected to a main fluid pipe via a solenoid valve; an adaptive dilution unit, connected to the raw water direct monitoring flow path, having a dilution ratio control function based on turbidity or suspended solids concentration; an intelligent cleaning module, including a miniature air compressor, an acid tank, a bactericide tank, and a gas-liquid mixing spray washing structure, used for high-pressure rinsing of the optical window; a spectral detection module, including an ultraviolet-visible spectrometer and an excitation-emission matrix fluorescence spectrometer connected to the main fluid pipe; wherein the spectral detection module adopts a flow cell structure, enabling online measurement of water samples without sampling; and a controller, used to control the flow path switching of the multi-flow-path sampling unit, the dilution ratio of the adaptive dilution unit, the cleaning process of the intelligent cleaning module, the spectral acquisition timing and detection trigger of the spectral detection module, and to achieve data communication.
[0015] The beneficial effects of this invention are as follows:
[0016] Significantly improved monitoring efficiency: By using graph theory-based intelligent flow path selection and optimal sampling path planning, the water sampling problem is formalized into a modified traveling salesman problem and solved using the cutting plane method and branch and bound method. Compared with the traditional polling strategy, the sampling cycle is shortened by 30-50% and the monitoring efficiency is improved by 50%.
[0017] Significantly improved detection accuracy: The optimal transport theory is used to optimize dilution control. The dilution problem is modeled as an optimal transport problem of multidimensional spectral distribution. The optimal dilution ratio is solved by the linear space super gradient algorithm, which reduces the deviation between the diluted spectrum and the ideal measurement conditions by 60%. The relative error of water quality parameter measurement is reduced from ±10% to ±3%, and the detection accuracy is improved by 70%.
[0018] Significantly faster computation speed: The matrix flip graph search algorithm is applied to optimize spectral calculation. By abstracting the spectral data matrix into a flip graph model and using heuristic search to find the optimal flip operation sequence, combined with fast matrix multiplication and SIMD vectorization technology, the speed of spectral calculation is increased by 5-10 times. The calculation time for a single water quality parameter is shortened from 10-30 seconds to 1-3 seconds, realizing true real-time online monitoring.
[0019] The system robustness is significantly enhanced: by identifying system vulnerabilities and configuring redundant flow paths through bilateral connected component decomposition, combined with a dynamic path replanning mechanism, the system's fault tolerance to single-point failures is greatly improved. The system availability rate is increased from 90-95% to over 99.5%, the fault recovery time is shortened from 10-30 minutes to less than 1 minute, and the system downtime is reduced by more than 80%.
[0020] The level of intelligence has been comprehensively improved: the integration of temporal convolutional network models to predict the degree of optical window contamination, deep neural network architecture for dilution decision models, intelligent flow path selection and automatic dilution control reduces manual intervention, improves operational consistency, and enhances user experience through a real-time visual interface. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of an online dynamic water quality fingerprint rapid acquisition and preprocessing method provided by an embodiment of the present invention;
[0023] Figure 2 This is a flowchart of the intelligent flow path selection process of the present invention;
[0024] Figure 3 This is an example of the system topology diagram model of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of an online dynamic water quality fingerprint rapid acquisition and pretreatment device provided in an embodiment of the present invention.
[0026] Figure reference numerals: 11. Raw water direct monitoring flow path; 12. Automatic dilution monitoring flow path; 13. Standard sample calibration flow path; 14. Fluid main pipe; 21. Miniature air compressor; 22. Acid tank; 23. Bactericide tank; 3. Spectroscopic detection module; 4. Controller; 51. Internal circulation pump; 52. Pressure relief valve; 53. Filter; 61. First air valve; 62. Second air valve; 71. First feed pump; 72. Second feed pump; 73. Third feed pump; 81. First solenoid valve; 82. Second solenoid valve; 83. Third solenoid valve; 84. Fourth solenoid valve; 85. Fifth solenoid valve; 86. Sixth solenoid valve; 87. Seventh solenoid valve; 88. Eighth solenoid valve; 9. Turbidity sensor. Detailed Implementation
[0027] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0028] Example 1
[0029] This embodiment provides a method for rapid online dynamic water quality fingerprint acquisition and preprocessing, such as... Figure 1As shown, the test flow path is intelligently selected from multiple inlet flow paths; the degree of contamination of the optical window is determined based on the light intensity attenuation trend prediction model, and the optical window is cleaned when the predicted contamination degree exceeds the dynamic threshold. The light intensity attenuation trend prediction model is a temporal convolutional network model. In the measurement channel corresponding to the cleaned optical window, the water sample of the test flow path is tested. The turbidity detection results, colorimetric detection results, and historical dilution effect evaluation data are input into the dilution decision model, which employs a deep neural network architecture. The water sample of the test flow path is diluted according to the dilution parameters output by the dilution decision model. The diluted water sample is then subjected to spectral acquisition to obtain water quality fingerprint spectral data, and water quality parameters are calculated based on the water quality fingerprint spectral data to complete the water quality fingerprint acquisition and preprocessing.
[0030] The above process specifically includes:
[0031] S1: Intelligently select the flow path to be tested from multiple inlet flow paths.
[0032] In this embodiment, the water quality monitoring system (hereinafter referred to as the system) includes multiple influent flow paths, such as upstream sampling point flow paths, tributary sampling point flow paths, and main stream sampling point flow paths. The flow path selection module intelligently selects the flow path to be monitored from the multiple influent flow paths based on historical water quality data, system operating status, and optimization algorithms.
[0033] Specifically, the flow path selection module first acquires historical water quality data for each inlet flow path, including time-series data of parameters such as pollutant concentration, turbidity, and color. Then, based on graph theory modeling, the entire water quality monitoring system is modeled as a directed weighted graph, where vertices represent key components such as sampling points, filters, pumps, mixers, and detectors, edges represent pipelines connecting the components, and the weight of the edges represents flow time, energy consumption, or failure risk.
[0034] Based on the graph theory model, the flow path selection module applies an optimization algorithm to calculate the optimal sampling path. This optimization algorithm formalizes the water sampling problem as a modified Traveling Salesman Problem (TSP), aiming to minimize the total sampling cost (including time cost, energy cost, and risk cost) while meeting water quality monitoring requirements.
[0035] The flow path selection module controls the opening and closing states of each solenoid valve in the multi-flow path structure based on the calculated optimal sampling path, thus realizing the water sample flow path from the selected test flow path to the detector. For example, in a certain monitoring cycle, if the flow path selection module determines that the mainstream sampling point flow path is the test flow path, it will open the solenoid valve corresponding to the mainstream sampling point and close the solenoid valves corresponding to other sampling points, allowing the mainstream water sample to flow into the subsequent detection channel.
[0036] S2: Based on the light intensity attenuation trend prediction model, determine the degree of contamination of the optical window. When the predicted degree of contamination exceeds the dynamic threshold, clean the optical window.
[0037] In this embodiment, the optical window is a key component of the spectral detection module. When the water sample flows through the optical window, the light beam emitted by the light source passes through the water sample and is received by the detector. During long-term operation, pollutants such as suspended solids and microbial films in the water sample will adhere to the surface of the optical window, causing light intensity attenuation and affecting detection accuracy.
[0038] To promptly detect and clean optical window contamination, this embodiment employs a light intensity attenuation trend prediction model. This model is a Temporal Convolutional Network (TCN) model, which can learn the temporal evolution of light intensity attenuation from historical light intensity data and predict the light intensity value at future times.
[0039] The construction and training process of a temporal convolutional network model includes:
[0040] Temporal convolutional network models consist of multiple causal convolutional layers and residual connection structures. Causal convolutional layers ensure that the model uses only historical data for prediction and does not reveal future information. Residual connection structures alleviate the vanishing gradient problem in deep networks and improve model training efficiency.
[0041] The model input is a historical light intensity time series. ,in, Represents the first within the historical time window The light intensity value collected at all times, For time index, the value is... ; The total length of the historical time window; the model output is the future [number]th [time window]. Predicted light intensity value at time ,in, The predicted time step is defined as the time interval from the end of the historical time window to the predicted time.
[0042] The first sequential convolutional network The layer convolution operation is defined as:
[0043]
[0044] in, Indicates the first Feature map of the layer and The first The kernel weights and biases of the layers. This indicates a causal convolution operation. This is the activation function.
[0045] The residual connection structure adds the input directly to the output:
[0046]
[0047] The model training uses the Mean Squared Error (MSE) loss function:
[0048]
[0049] in, The number of training samples. For the first The true light intensity value of each sample These are the model's predicted values.
[0050] The training data came from the historical operation records of the water quality monitoring system, including light intensity attenuation curves under different pollution levels. The Adam optimizer was used to update the model parameters, with a learning rate of 0.001, a batch size of 32, and 100 training epochs.
[0051] After training, the model is deployed to the controller of the water quality monitoring system. During system operation, the controller collects light intensity data in real time and calls the temporal convolutional network model to predict light intensity at regular intervals (e.g., every 5 minutes).
[0052] When the predicted light intensity value is lower than the dynamic threshold, the optical window is deemed to be contaminated beyond the standard, triggering a cleaning process. The dynamic threshold is adaptively adjusted according to the current water quality conditions; for example, when the water sample has high turbidity, the dynamic threshold is lowered accordingly to enable more frequent cleaning.
[0053] The cleaning process includes the following steps: First, close the solenoid valve of the water sample flow path to stop the water sample flow; then, start the intelligent cleaning module, which generates high-pressure airflow through a micro air compressor to spray acid or bactericide from the spray structure onto the surface of the optical window to remove the attached contaminants; finally, rinse the optical window with pure water, drain the cleaning solution, and restore the water sample flow.
[0054] S3: In the measurement channel corresponding to the cleaned optical window, the water sample of the flow path to be tested is detected, and the turbidity detection result, color detection result, and historical dilution effect evaluation data are input into the dilution decision model.
[0055] In this embodiment, after cleaning is completed, the optical window returns to a clean state, and the light intensity returns to normal. At this time, the solenoid valve of the flow path to be measured is opened, allowing the water sample to flow into the measurement channel.
[0056] The measurement channel is equipped with a turbidity sensor and a colorimetry sensor, used to detect the turbidity and colorimetry of the water sample, respectively. The turbidity sensor, based on the principle of light scattering, measures the intensity of light scattered by suspended particles in the water sample and outputs the turbidity value (unit: NTU). The colorimetry sensor, based on the principle of light absorption, measures the absorbance of the water sample at a specific wavelength (e.g., 436 nm) and outputs the colorimetry value (unit: degrees).
[0057] The controller collects turbidity detection results. Colorimetric test results At the same time, historical dilution effect evaluation data are extracted from the historical database. Historical dilution effect assessment data includes information such as the dilution ratio used in previous monitoring cycles, the spectral quality score after dilution, and the calculation error of water quality parameters.
[0058] The controller combines turbidity detection results, colorimetric detection results, and historical dilution effect evaluation data into an input vector. This information is then input into the dilution decision model.
[0059] The dilution decision model employs a deep neural network architecture, which can learn the complex nonlinear relationship between the dilution ratio and water sample characteristics from the input data and output the optimal dilution parameters.
[0060] Construction and training of dilution decision models:
[0061] The dilution decision model is a multi-layer feedforward neural network, consisting of an input layer, multiple hidden layers, and an output layer. The input layer receives turbidity, chromaticity, and historical dilution effect evaluation data, while the output layer outputs the dilution ratio. and dilution water flow rate .
[0062] The hidden layer uses a fully connected structure, and the calculation formula for each layer is as follows:
[0063]
[0064]
[0065] in, and The first The linear output and activation output of the layer, and For the weight matrix and bias vector, It is an activation function (e.g., ReLU or Sigmoid).
[0066] The output layer uses a linear activation function, and the output dilution parameter is:
[0067]
[0068] Where L is the total number of hidden layers.
[0069] The model was trained using a supervised learning method. The training data consisted of manually labeled dilution experiment data, including optimal dilution ratios and dilution effect scores under different turbidity and chromaticity conditions. The loss function was defined as the mean squared error between the predicted dilution parameters and the optimal dilution parameters.
[0070]
[0071] in, and For the first The optimal dilution ratio and dilution water flow rate for each training sample.
[0072] The model was trained using the backpropagation algorithm and the Adam optimizer, with a learning rate of 0.001, a batch size of 64, and 200 training epochs. An early stopping strategy was employed during training: training was stopped when the validation set loss did not decrease for 10 consecutive epochs to prevent overfitting.
[0073] After training, the model is deployed to the controller for real-time decision dilution.
[0074] S4: Dilute the water sample of the flow path to be tested according to the dilution parameters output by the dilution decision model.
[0075] In this embodiment, the dilution decision model outputs the dilution ratio. and dilution water flow rate Based on these dilution parameters, the controller directs the adaptive dilution unit to dilute the water sample in the test flow path.
[0076] The adaptive dilution unit includes a raw water flow path, a dilution water flow path, proportional valves, and a mixer. The raw water flow path is connected to the flow path to be tested, introducing raw water into the mixer; the dilution water flow path is connected to a pure water source, introducing dilution water into the mixer. The proportional valves control the raw water flow rate. and dilution water flow rate .
[0077] According to the dilution ratio and target total flow (e.g., 100 mL / min), calculate the raw water flow rate and the dilution water flow rate:
[0078]
[0079]
[0080] The controller sends control commands to the proportional valve to adjust the valve opening, ensuring that the raw water flow rate and dilution water flow rate reach the calculated values. The raw water and dilution water are thoroughly mixed in the mixer to form a diluted water sample, which then flows into the spectral detection module.
[0081] For example, when the dilution ratio is 60% and the target total flow rate is 100 mL / min, the original water flow rate is 40 mL / min and the dilution water flow rate is 60 mL / min. The concentration of the mixed water sample is reduced to 40% of the original water concentration, bringing it into the detector's optimal operating range.
[0082] S5: Perform spectral acquisition on the diluted water sample to obtain water quality fingerprint spectral data.
[0083] In this embodiment, the diluted water sample flows into the spectral detection module. The spectral detection module includes an ultraviolet-visible spectrometer and an excitation-emission matrix fluorescence spectrometer, employing a flow cell structure, allowing for online measurement of the water sample without the need for sampling.
[0084] The ultraviolet-visible spectrometer emits a continuous spectrum with wavelengths ranging from 200 to 800 nm. The light beam passes through a water sample in a flow cell and is received by a detector. The detector measures the transmitted light intensity at different wavelengths and calculates the absorbance.
[0085]
[0086] in, wavelength absorbance at that point The intensity of transmitted light. The incident light intensity is [value].
[0087] Ultraviolet-Vis spectrometer output absorbance spectrum It contains absorbance values at 601 wavelength points.
[0088] Excitation-emission matrix fluorescence spectrometers sequentially excite water samples using light sources with different excitation wavelengths, and the detector measures the fluorescence intensity at different emission wavelengths to form a two-dimensional fluorescence spectral matrix. .
[0089] The controller combines the ultraviolet-visible absorbance spectrum and fluorescence spectrum matrix to form water quality fingerprint spectral data. .
[0090] S6: Calculate water quality parameters based on the water fingerprint spectral data to complete the collection and preprocessing of the water fingerprint.
[0091] In this embodiment, the controller calculates water quality parameters such as pollutant concentration, total organic carbon (TOC), and chemical oxygen demand (COD) based on water quality fingerprint spectral data and a preset water quality parameter calculation model.
[0092] Water quality parameter calculation models can employ methods such as multiple linear regression, partial least squares regression, or neural networks. Taking total organic carbon (TOC) calculation as an example, an empirical formula based on 254 nm ultraviolet absorbance is used:
[0093]
[0094] in, These are empirical coefficients, determined through calibration using standard samples. The absorbance is at a wavelength of 254 nm.
[0095] Chemical oxygen demand (COD) is calculated using a multiple linear regression model based on multi-band spectral characteristics:
[0096]
[0097] in, For regression coefficients, For the selected characteristic wavelength, This represents the number of characteristic wavelengths.
[0098] The regression coefficients were obtained by least-squares fitting of the spectral data of the standard samples and the known COD values.
[0099] After the controller calculates the water quality parameters, it stores the results in the database and displays them to the user through a visual interface, thus completing the collection and preprocessing of water quality fingerprints.
[0100] Example 2
[0101] like Figure 2As shown, this embodiment, based on embodiment 1, provides a detailed explanation of step S1, "intelligently selecting the flow path to be tested from multiple inlet flow paths." The physical structure of the entire water quality monitoring system is obtained, and the entire water quality monitoring system is modeled as a first directed weighted graph. Vertices represent key components including sampling points, filters, pumps, mixers, and detectors, and edges represent pipelines connecting these key components. The weight of each edge represents flow time, energy consumption, or failure risk. A system topology graph model corresponding to the entire water quality monitoring system is output. Based on the system topology graph model, a depth-first search algorithm is applied to traverse the system topology graph model from the starting vertex to identify all cut edges. By deleting all identified cut edges, the system topology graph model is decomposed into multiple bilateral connected components, obtaining the bilateral connected component decomposition results. The bilateral connected components are decomposed by identifying each bilateral component in the bilateral connected component decomposition results. The connection relationships between connected components are used to determine potential single-point fault locations, and redundant flow paths are configured at these locations. After completing the design of redundant flow paths, the water sampling problem is formalized as a modified Traveling Salesman Problem (TSP), and a distance matrix is constructed, where the elements of the distance matrix represent the comprehensive cost from one component to another. The cutting plane method is applied to solve the linear programming relaxation problem in the distance matrix to obtain an initial lower bound solution. The branch and bound method is used to gradually reduce the integer gap of the initial lower bound solution until the integer gap is less than a preset threshold, thus obtaining the optimal sampling path. Based on the path switching logic in the optimal sampling path, the opening and closing of the solenoid valves in the multi-flow path structure with redundant design are controlled to achieve flow path selection, thereby determining the flow path to be tested.
[0102] S1.1: Obtain the physical structure of the entire water quality monitoring system, model the entire water quality monitoring system as a first directed weighted graph, and output the system topology graph model.
[0103] In this embodiment, for example, the water quality monitoring system includes multiple physical components, including three sampling points (upstream sampling point, tributary sampling point, and main stream sampling point), three solenoid valves, one filter, one pump, one mixer, one dilution water source, and one detector. These components are connected by pipelines to form a complex fluid network.
[0104] To achieve intelligent flow path selection, a mathematical model of the system's physical structure is first required. This embodiment employs graph theory to model the water quality monitoring system as a first directed weighted graph. .
[0105] Vertex set Includes all key components:
[0106]
[0107] in, These represent upstream sampling points, tributary sampling points, and main stream sampling points, respectively. This represents the corresponding solenoid valve. Represents a filter. Represents a pump. Represents a mixer. Represents dilution water source, Represents the detector, This represents the total number of vertices.
[0108] Edge set Includes piping between connecting components:
[0109]
[0110] For example, This indicates the pipeline from the upstream sampling point to the corresponding solenoid valve. This indicates the piping from the solenoid valve to the filter.
[0111] Edge weight This indicates the flow time, energy consumption, or failure risk. For edges... The weights are defined as follows:
[0112]
[0113] in, For water samples from components Flow to components Time (unit: seconds), Energy consumption during the flow process (unit: joules). For the risk of failure of pipelines or components (dimensionless, range 0-1), This is a weighting coefficient, set according to actual application requirements.
[0114] For example, in a river water quality monitoring station, the pipeline length from the upstream sampling point to the detector is 50 meters, the flow time is 120 seconds, the energy consumption is 50 joules, and the failure risk is 0.1. Then the weight of the corresponding edge is:
[0115]
[0116] (assumption) (These are 1.0, 0.5, and 100 respectively).
[0117] By traversing all components and pipes in the system, a complete system topology model Gsystem is constructed, such as... Figure 3 As shown, the output includes the vertex set, edge set, and weight matrix.
[0118] S1.2: Based on the system topology graph model, the depth-first search algorithm is applied to identify all cut edges, and the system topology graph model is decomposed into multiple bilateral connected components to obtain the bilateral connected component decomposition results.
[0119] In this embodiment, to improve system reliability, it is necessary to identify vulnerable points, i.e., cut edges, in the system topology graph. A cut edge is an edge whose deletion disrupts the connectivity of the graph, preventing some vertices from reaching other vertices. The pipelines or components corresponding to cut edges are single points of failure in the system; once a failure occurs, it will lead to system paralysis.
[0120] This embodiment uses the Depth-First Search (DFS) algorithm to identify all cut edges. The DFS algorithm starts from the initial vertex and recursively traverses all vertices and edges of the graph, recording the visit timestamp and the lowest reachable timestamp for each vertex.
[0121] Access timestamp Represents vertices The time when it is first visited during a DFS traversal. Minimum reachable timestamp. Indicates from vertex Starting from the beginning, find the visit timestamp of the earliest ancestor vertex that can be reached through the edges and back edges of the DFS tree.
[0122] For the edge If the conditions are met:
[0123]
[0124] Then the edge This is a cut edge. The condition indicates that from the vertex... Once you start, you cannot return to the vertex via any other path. or its ancestors, therefore the side It is the only bridge connecting two connected components.
[0125] Applying the Depth-First Search (DFS) algorithm to traverse the system topology graph model Identify all cut edges to obtain the cut edge set. .
[0126] Then, all cut edges are removed from the system topology graph model, decomposing the graph into multiple connected subgraphs. Within each connected subgraph, there are at least two paths between any two vertices with no repeating edges; these are called bilaterally connected components.
[0127] For example, in a water quality monitoring system, two cut edges were identified: the main pipeline connecting the mixer and the detector, and the pipeline connecting the dilution water source. After removing these two cut edges, the system is decomposed into three bilaterally connected components: sampling point-solenoid valve-filter-pump-mixer, dilution water source, and detector.
[0128] Output the decomposition results of bilaterally connected components, including the vertex set and edge set of each bilaterally connected component.
[0129] S1.3: By identifying the connection relationships between each bilateral connected component in the bilateral connected component decomposition result, the potential single point of failure location is determined, and redundant flow paths are configured at the potential single point of failure location.
[0130] In this embodiment, the two-sided connected components are connected by a cut edge. The pipeline or component corresponding to the cut edge is a potential single point of failure location, and redundant flow paths need to be configured to improve the system's fault tolerance. That is, the water sample is controlled to flow through the redundant flow path pre-configured at the potential single point of failure location.
[0131] By analyzing the results of bilateral connected component decomposition, the connectivity relationships between each bilateral connected component are identified. For each cut edge... , where the vertex Belongs to bilateral connected components ,vertex Belongs to bilateral connected components Then cut the edge. It is a connection and The only path.
[0132] Cutting edges Corresponding locations should be designed with backup pipelines or bypasses to create redundant flow paths. The design principle for redundant flow paths is that when the main pipeline fails, water samples can be drawn from the backup pipeline. Flow to To ensure system connectivity.
[0133] For example, for the main pipeline (cut edge) connecting the mixer and the detector, a backup pipeline is designed, connecting another outlet of the mixer to a backup inlet of the detector. When the main pipeline is blocked, the system switches to the backup pipeline by controlling a solenoid valve to continue monitoring.
[0134] For the pipeline connecting to the dilution water source (cut edge), design a backup dilution water source pipeline, connecting to the mixer from another pure water source. When the main dilution water source fails, switch to the backup dilution water source.
[0135] After completing the redundant flow path design, update the system topology model, add the edges corresponding to the spare pipelines, and form a multi-flow path structure with redundant design.
[0136] S1.4: After completing the design of the redundant flow path, the water sampling problem is formalized into a modified Traveling Salesman Problem (TSP). The distance matrix is constructed, and the cutting plane method and branch and bound method are applied to solve it, thus obtaining the optimal sampling path.
[0137] In this embodiment, the water quality monitoring system needs to collect water samples from multiple sampling points sequentially for testing. The choice of sampling order affects monitoring efficiency and cost. This embodiment formalizes the water sampling problem as a modified Traveling Salesman Problem (TSP).
[0138] The traditional TSP problem is defined as: given Given the distances between several cities, find the shortest path that visits each city exactly once and returns to the starting point. The modified TSP problem considers the special constraints of a water quality monitoring system.
[0139] The starting point is fixed at the detector position;
[0140] The order of access should take into account the timeliness and priority of the water samples;
[0141] Path costs include not only distance, but also time, energy consumption, and risks.
[0142] Construct the distance matrix , of which elements Indicates from component To Component Overall cost:
[0143]
[0144] That is, from the components To Component Among all possible paths, the path with the smallest total weight is selected, and its total weight is used as an element of the distance matrix.
[0145] For example, there may be multiple paths from the upstream sampling point to the detector:
[0146] Path 1: Upstream sampling point → Solenoid valve → Filter → Pump → Mixer → Detector;
[0147] Path 2: Upstream sampling point → Solenoid valve → Filter → Pump → Mixer → Backup pipeline → Detector;
[0148] Choose the path with the lowest total weight.
[0149] The modified mathematical model for the TSP problem is as follows:
[0150]
[0151] Constraints:
[0152]
[0153]
[0154]
[0155]
[0156] in, For binary decision variables, when from component Move directly to component hour ,otherwise The third constraint is the sub-loop elimination constraint, which prevents the formation of small loops that do not contain all vertices.
[0157] The TSP problem is an NP-hard problem, and solving it exactly involves high computational complexity. This embodiment uses a combination of the cutting plane method and the branch and bound method to solve it.
[0158] The cutting plane method first solves the linear programming relaxation problem, that is, the integer constraints. relaxation The relaxation problem of linear programming can be solved efficiently, yielding a lower bound solution.
[0159] If the lower bound solution satisfies the integer constraint, it is the optimal solution. Otherwise, identify solutions that violate the sub-loop elimination constraint and add a cutting plane constraint:
[0160]
[0161] in, This is the set of vertices that form a sub-loop.
[0162] Resolve the linear programming problem with the added cutting plane constraint, iterating until no new cutting plane constraint can be added.
[0163] The branch and bound method, based on the cutting plane method, branches the non-integer solutions.
[0164] Choose a non-integer variable Create two subproblems:
[0165] Add constraints to a subproblem Add constraints to another subproblem The subproblem is solved recursively, and the current best integer solution is maintained as an upper bound.
[0166] Prune a branch when its lower bound is greater than or equal to its upper bound. The global optimum is obtained when all branches have been pruned or solved.
[0167] In this embodiment, the integer gap threshold is set to 1%, that is, when the gap is less than 1%, the iteration stops and an approximate optimal solution is obtained.
[0168] For example, in a water quality monitoring station, the modified TSP problem was solved using the cutting plane method and the branch and bound method, and the optimal sampling path was obtained: detector → main stream sampling point → tributary sampling point → upstream sampling point → detector.
[0169] The total cost of this path is 450 (a weighted sum of time, energy consumption, and risk), which is 25% lower than the traditional polling strategy (detector → upstream sampling point → tributary sampling point → mainstream sampling point → detector, total cost 600).
[0170] S1.5: Based on the path switching logic in the optimal sampling path, control the opening and closing of the solenoid valve in the multi-flow path structure with redundant design to achieve flow path selection and determine the flow path to be tested.
[0171] In this embodiment, the controller switches the flow path sequentially according to the optimal sampling path to collect water samples from each sampling point.
[0172] The optimal sampling path includes path switching logic, which specifies the dwell time at each sampling point and the time to switch to the next sampling point. Based on the path switching logic, the controller sends control commands at predetermined times to control the opening and closing states of each solenoid valve in the multi-flow path structure.
[0173] For example, at time The controller opens the solenoid valve corresponding to the main sampling point and closes the solenoid valves corresponding to other sampling points, allowing the main water sample to flow into the detection channel. After the main water sample monitoring is completed, at a specific time... The controller closes the solenoid valve at the main stream sampling point and opens the solenoid valve at the tributary sampling point, switching to tributary water sample monitoring.
[0174] The opening and closing state of the solenoid valve is controlled by digital signals. The controller outputs a high-level signal (e.g., 5V) to open the solenoid valve and a low-level signal (e.g., 0V) to close it. The response time of the solenoid valve is typically 0.1-0.5 seconds. After sending a switching command, the controller waits for the solenoid valve to be fully open or closed before proceeding to the next step.
[0175] During flow path switching, to avoid mixing of water samples from different sampling points, the controller closes the current flow path solenoid valve and waits for the residual water sample in the pipeline to drain (delay 0.5-2 seconds) before opening the next flow path solenoid valve. Simultaneously, the controller gradually adjusts the flow rate, employing a PID control algorithm to smoothly transition the flow rate to the target flow rate, avoiding pressure fluctuations and water sample contamination caused by sudden flow changes.
[0176] Through the above flow path selection process, the controller determines the flow path to be tested in the current monitoring cycle, thus realizing intelligent flow path selection.
[0177] Example 3
[0178] This embodiment, based on Embodiment 2, provides a dynamic path replanning method for detecting anomalies and automatically switching flow paths during system operation. It monitors the operating status parameters and performance parameters of each component in the multi-flow path structure with redundant design; determines whether the operating status parameters deviate from the normal operating parameter threshold range or whether the performance parameters are lower than a preset performance threshold to detect abnormal situations such as pipeline blockage or pump performance degradation, and generates a system abnormal status signal; based on the system abnormal status signal, the abnormal component indicated by the system abnormal status signal is marked as unavailable from the system topology model using a dynamic path replanning algorithm, and the optimal path from the current state to the target state is recalculated, automatically activating alternative flow paths or recalculating the optimal sampling path to obtain an updated sampling path scheme; based on the updated sampling path scheme, the solenoid valve of the current flow path corresponding to the optimal sampling path scheme is first closed, then the solenoid valve of the path corresponding to the updated sampling path scheme is opened, and the flow rate is gradually adjusted during the switching process to determine the dynamically optimized flow path to be tested.
[0179] Step A: Monitor the operating status parameters and performance parameters of each component in the multi-flow path structure with redundant design.
[0180] In this embodiment, the water quality monitoring system is equipped with a variety of sensors to monitor the operating status parameters and performance parameters of each component in real time.
[0181] Operating status parameters include:
[0182] Flow rate of each pipeline measured by the flow sensor (unit: mL / min);
[0183] Pressure in each pipeline measured by the pressure sensor (unit: kPa);
[0184] The solenoid valve status sensor detects the open / closed state (open / closed) of the solenoid valve.
[0185] Pump speed sensor measures pump speed (unit: rpm).
[0186] Performance parameters include:
[0187] The ratio of the pump's actual output pressure to its rated output pressure (pump efficiency);
[0188] The pressure difference of the filter (the difference between the inlet pressure and the outlet pressure, reflecting the degree of filter clogging);
[0189] The light transmittance of the optical window (reflecting the degree of window contamination).
[0190] The controller collects data from each sensor at regular intervals (e.g., 1 second) through the data acquisition module and stores it in the real-time data buffer.
[0191] For example, during a certain monitoring period, the controller collected the following data:
[0192] Mainstream sampling point pipeline flow rate: 95 mL / min;
[0193] Mainstream sampling point pipeline pressure: 150 kPa;
[0194] Solenoid valve status at main sampling points: Open;
[0195] Pump speed: 1500 rpm;
[0196] Pump efficiency: 0.92;
[0197] Filter differential pressure: 20 kPa;
[0198] Optical window transmittance: 0.88;
[0199] The controller continuously monitors these parameters, providing a data foundation for anomaly detection.
[0200] Step B: Determine whether the operating status parameters deviate from the normal operating parameter threshold range or whether the performance parameters are lower than the preset performance threshold, in order to detect abnormal situations such as pipeline blockage or pump performance degradation, and generate a system abnormal status signal.
[0201] In this embodiment, the controller establishes an anomaly detection algorithm and sets the normal operation parameter threshold range and preset performance threshold.
[0202] The threshold range for normal operating parameters is determined based on system design parameters and historical operating data. For example:
[0203] Normal flow rate range: 90-110 mL / min (target flow rate 100 mL / min, allowable deviation ±10%);
[0204] Normal pressure range: 140-160 kPa (target pressure 150 kPa, allowable deviation ±10 kPa);
[0205] Normal pump speed range: 1400-1600 rpm;
[0206] The preset performance threshold is determined based on the component's performance degradation pattern and maintenance strategy. For example:
[0207] Pump efficiency threshold: 0.85 (when the pump efficiency is below 0.85, the pump performance is considered to have deteriorated);
[0208] Filter differential pressure threshold: 30 kPa (when the differential pressure exceeds 30 kPa, the filter is considered clogged);
[0209] Optical window transmittance threshold: 0.80 (when the transmittance is below 0.80, the window is considered contaminated);
[0210] After each data acquisition, the controller determines whether the operating status parameters deviate from the normal range and whether the performance parameters are lower than the preset threshold.
[0211] The judgment logic is as follows:
[0212] If the flow rate is <90 or the flow rate is >110:
[0213] An abnormal signal was generated: "Traffic anomaly";
[0214] If the pressure is <140 or the pressure is >160:
[0215] An abnormal signal was generated: "Pressure Anomaly";
[0216] If the pump efficiency is <0.85:
[0217] An abnormal signal was generated: "Pump performance has decreased";
[0218] If the filter differential pressure is >30:
[0219] An error signal was generated: "Filter clogged";
[0220] If the transmittance of the optical window is <0.80:
[0221] An abnormal signal was generated: "Window contamination";
[0222] For example, during a certain monitoring cycle, the controller detected a sudden drop in the flow rate of the main sampling point pipeline to 60 mL / min, which is lower than the normal range limit of 90 mL / min, and judged it as "abnormal flow". Further analysis revealed that the filter pressure differential increased to 35 kPa, exceeding the threshold of 30 kPa, and was judged as "filter blockage".
[0223] The controller generates system abnormal status signals, including information such as abnormal type, abnormal component, and abnormal time.
[0224] Abnormal signals:
[0225] Type: Filter clogging;
[0226] Component: Filter v7;
[0227] Time: 2023-10-15 14:32:18;
[0228] Details: Filter pressure differential is 35 kPa, exceeding the threshold of 30 kPa;
[0229] Abnormal system status signals trigger the dynamic path replanning process.
[0230] Step C: Based on the system abnormal state signal, the abnormal component indicated by the system abnormal state signal is marked as unavailable from the system topology model by the dynamic path replanning algorithm, and the optimal path from the current state to the target state is recalculated. Alternative flow paths are automatically started or the optimal sampling path is recalculated to obtain the updated sampling path scheme.
[0231] In this embodiment, after receiving a system abnormality signal, the controller immediately starts the dynamic path replanning algorithm.
[0232] The steps of the dynamic path replanning algorithm are as follows:
[0233] From the system topology model In this context, the vertices corresponding to abnormal components are marked as unavailable. For example, filters. If blocked, mark it. Not available.
[0234] Update edge weights: Set the weights of all edges connecting unavailable vertices to infinity. This indicates that these edges are not traversable. For example, edge and The weight is set to .
[0235] Apply shortest path algorithms (such as Dijkstra's algorithm or A*). The algorithm recalculates the optimal path from the current state (current sampling point) to the target state (detector).
[0236] The pseudocode for Dijkstra's algorithm is as follows:
[0237] Initialization: dist[starting point] = 0, dist[other vertices] = ∞;
[0238] The priority queue PQ contains all vertices, sorted by their dist values;
[0239] When PQ is not empty:
[0240] u = the vertex with the smallest dist value in PQ;
[0241] Remove u from PQ;
[0242] For each neighbor v of u:
[0243] If dist[u] + w_uv <dist[v]:
[0244] dist[v] = dist[u] + w_uv;
[0245] Update the position of v in PQ;
[0246] If a new optimal path is found, the alternative flow path will be automatically activated.
[0247] For example, the original path was "mainstream sampling point → solenoid valve → filter → pump → mixer → detector". After the filter became clogged, the new path became "mainstream sampling point → solenoid valve → backup filter → pump → mixer → detector".
[0248] If no new path can be found (all paths are blocked), the optimal sampling path is recalculated, the current sampling point is skipped, and other sampling points are monitored first. For example, if the main sampling point flow path is completely blocked, the sampling path is updated to "detector → tributary sampling point → upstream sampling point → detector", and the main sampling point is temporarily skipped.
[0249] The updated sampling path scheme is output, including new flow path switching logic and solenoid valve control instructions.
[0250] For example, when a filter is clogged, the dynamic path replanning algorithm calculates a new path:
[0251] Updated sampling path scheme:
[0252] Current sampling points: Mainstream sampling points;
[0253] New path:
[0254] Main sampling point → Solenoid valve v5 → Backup filter v12 → Pump v8 → Mixer v9 → Detector v11;
[0255] Solenoid valve control commands:
[0256] Close: v4 (main filter inlet solenoid valve);
[0257] Open: v5 (backup filter inlet solenoid valve);
[0258] Estimated switching time: 5 seconds;
[0259] Step D: Based on the updated sampling path scheme, first close the solenoid valve of the current flow path corresponding to the optimal sampling path scheme, then open the solenoid valve of the path corresponding to the updated sampling path scheme and control the flow rate to be gradually adjusted during the switching process to determine the dynamically optimized flow path to be tested.
[0260] In this embodiment, the controller performs a flow path switching operation based on the updated sampling path scheme.
[0261] The flow path switching employs a smooth transition mechanism to avoid water sample contamination and data interruption.
[0262] First, close the solenoid valve of the current flow path (e.g., the main filter inlet solenoid valve v4) to stop the water sample flow.
[0263] Wait for the residual water sample in the pipeline to drain. The controller delays for 0.5-2 seconds to ensure that the water sample in the main filter pipeline is completely drained, preventing mixing with the water sample in the backup flow path.
[0264] Open the solenoid valve for the new path (e.g., the backup filter inlet solenoid valve v5) to start the backup flow path.
[0265] The flow rate is gradually adjusted to the target flow rate. The controller uses a PID (proportional-integral-derivative) control algorithm, which adjusts the pump speed or proportional valve opening based on real-time feedback from the flow sensor to ensure a smooth flow transition.
[0266] The output of the PID control algorithm is:
[0267]
[0268] in, To control the output (pump speed or valve opening), For flow error, For target traffic, This represents the actual traffic volume. These are the proportional, integral, and differential coefficients, respectively.
[0269] For example, the target flow rate is 100 mL / min, the current actual flow rate is 60 mL / min, and the flow rate error is 40 mL / min. The PID controller calculates the control output and increases the pump speed from 1500 rpm to 1800 rpm, causing the flow rate to gradually increase. After 5 seconds of adjustment, the flow rate stabilizes at 100 mL / min.
[0270] Once the flow rate is confirmed to be stable, the controller outputs the dynamically optimized flow path under test and resumes normal monitoring.
[0271] For example, after the filter becomes clogged and the flow path is switched, the dynamically optimized flow path to be tested is:
[0272] The dynamically optimized flow path to be tested:
[0273] Sampling points: Mainstream sampling points;
[0274] Flow path: Main sampling point → Solenoid valve v5 → Backup filter v12 → Pump v8 → Mixer v9 → Detector v11;
[0275] Flow rate: 100 mL / min;
[0276] Status: Normal monitoring;
[0277] Through dynamic path replanning, the system can switch the flow path within 5 seconds after detecting an anomaly, and the data interruption time is less than 1 minute, ensuring continuous monitoring.
[0278] Example 4
[0279] This embodiment, based on embodiment 1, provides a detailed explanation of step S4, "dilution treatment of the water sample of the flow path to be tested according to the dilution parameters output by the dilution decision model". This step involves acquiring the raw water spectral data of the water sample to be tested, representing it as the multidimensional spectral distribution characteristics before dilution, and simultaneously representing the spectral characteristics under ideal measurement conditions as the target spectral distribution characteristics after dilution. A cost matrix is defined to transform the spectral distribution characteristics from before to after dilution, where each element represents the transmission cost required to convert each spectral feature before dilution to the corresponding spectral feature after dilution. The optimal transmission strategy for the cost matrix is iteratively solved using a linear space hypergradient algorithm. This involves initializing dual variables and calculating the hypergradient value in each iteration, updating the dual variables according to the step size rule until the gradient norm is less than a preset gradient norm threshold, ultimately obtaining the optimal transmission strategy. A mapping function is established from the spectral distribution characteristics before and after dilution to the physical dilution ratio. The optimal transmission strategy is substituted into the mapping function to calculate the actual dilution ratio, obtaining the optimal dilution ratio. Based on the optimal dilution ratio and the target total flow rate, the values of the raw water flow rate and the dilution water flow rate are calculated to obtain the dilution control parameters. Based on the dilution control parameters, the raw water flow rate and the dilution water flow rate are controlled through a proportional valve or peristaltic pump to achieve the dilution treatment of the water sample to be tested, obtaining the diluted water sample.
[0280] Acquisition and feature representation of raw water spectral data:
[0281] In this embodiment, before dilution, the raw water spectral data of the water sample to be tested needs to be collected. The controller sends a command to the spectral detection module, triggering the ultraviolet-visible spectrometer to perform a rapid spectral scan of the raw water. The spectral scan covers a wavelength range of 200 nm to 800 nm, collecting one data point every 1 nm, for a total of 601 absorbance values at different wavelengths. These absorbance values constitute a complete characterization of the raw water spectral data, reflecting the comprehensive optical properties of various pollutants in the water sample.
[0282] To apply optimal transport theory for dilution optimization, the raw water spectral data needs to be converted into a probability distribution. Specifically, each absorbance value of the raw water spectral data is divided by the sum of the absorbance values at all wavelengths to obtain a normalized spectral distribution. The normalized data satisfies the basic property of a probability distribution, that is, the sum of all normalized absorbance values equals 1, and each normalized value is between 0 and 1, which can be regarded as the proportion of that wavelength point in the overall spectral energy distribution. The data processed in this way is called the multidimensional spectral distribution feature before dilution.
[0283] Simultaneously, based on the detector's technical specifications and optimal operating conditions, the diluted target spectral distribution characteristics are defined. The detector's optimal absorbance operating range is typically 0.1 to 1.0 absorbance units. When the absorbance is too low, the signal-to-noise ratio is small, increasing measurement error; when the absorbance is too high, the detector may reach saturation, disrupting the linear response. Therefore, the goal of setting the target spectral distribution characteristics is to scale the entire raw water spectrum to the optimal operating range. By multiplying the raw water spectrum by a target scaling factor, the peak absorbance of the scaled spectrum is positioned at the center of the optimal range, for example, 0.5 absorbance units. The scaled spectral data is also normalized to obtain the diluted target spectral distribution characteristics.
[0284] For example, suppose the peak absorbance of the raw water spectrum is 2.5 absorbance units, which significantly exceeds the detector's optimal operating range of 1.0. To reduce the peak to 0.5, the required target scaling factor is 0.5 divided by 2.5, which equals 0.2. This means the diluted water sample concentration should be 20% of the original water concentration, corresponding to a dilution ratio of 80%. In this way, a preliminary correlation is established between spectral characteristics and dilution ratio.
[0285] Definition and construction of cost matrix:
[0286] In this embodiment, optimal transport theory is used to study how to transform one probability distribution into another with minimal cost. In the water sample dilution optimization problem, the objective is to transform the spectral distribution before dilution into the target spectral distribution after dilution while minimizing the conversion cost. The conversion cost is quantified and defined using a cost matrix.
[0287] Each element of the cost matrix represents the cost of transforming the spectral characteristics of one wavelength point to another. The cost function is defined based on the distance between wavelengths. This embodiment uses Euclidean distance as the cost metric, meaning the distance between two wavelength points is equal to the square of their difference on the wavelength axis. The greater the distance, the higher the transformation cost, because transformation across a large wavelength interval implies a significant change in spectral characteristics.
[0288] The cost matrix is constructed as follows: For any two wavelength points in the spectral data, let their wavelengths be the i-th wavelength and the j-th wavelength, respectively. The element in the i-th row and j-th column of the cost matrix is equal to the square of the absolute value of the difference between the i-th wavelength and the j-th wavelength. The cost matrix has a dimension of 601 rows and 601 columns, corresponding to all possible transformation combinations between the 601 wavelength points. The diagonal elements have a value of zero, indicating that the transformation from the same wavelength point to itself incurs no cost. The values of the off-diagonal elements increase rapidly with the increase of the wavelength interval.
[0289] For example, for the conversion between wavelengths of 200 nm and 200 nm, the cost matrix element value is 200 minus the square of the absolute value of 200, which equals zero. For the conversion between wavelengths of 200 nm and 300 nm, the cost matrix element value is 200 minus the square of the absolute value of 300, which equals 100 squared, or 10000. For the conversion between wavelengths of 200 nm and 800 nm, the cost matrix element value is 200 minus the square of the absolute value of 800, which equals 600 squared, or 360000. It can be seen that as the wavelength interval increases from 100 nm to 600 nm, the conversion cost surges from 10000 to 360000, fully demonstrating the high cost of converting across large wavelength intervals.
[0290] Once the cost matrix is constructed, it provides the core parameters of the objective function for the mathematical modeling of the optimal transport problem.
[0291] Establishing a linear programming model:
[0292] In this embodiment, the optimal transport problem is formalized as a linear programming problem for solution. The goal of the linear programming model is to find a transport plan matrix that describes the probabilistic mass of the transfer from each wavelength point of the pre-dilution spectral distribution to each wavelength point of the post-dilution target spectral distribution. The element in the i-th row and j-th column of the transport plan matrix represents the probabilistic mass proportion of the transfer from the i-th wavelength point of the pre-dilution spectrum to the j-th wavelength point of the post-dilution spectrum.
[0293] The objective function of the linear programming model is to minimize the total transmission cost. The total transmission cost equals the sum of the element-wise products of the cost matrix and the transmission plan matrix. Specifically, the total transmission cost is obtained by multiplying each corresponding element in both the cost matrix and the transmission plan matrix, and then summing all the products. The objective function aims to find the transmission plan matrix that minimizes the total transmission cost.
[0294] The linear programming model needs to satisfy two types of constraints. The first type of constraint is called the row sum constraint, which ensures that the total probabilistic mass transferred from each source wavelength is equal to the probability value of the source distribution at that wavelength. For each row of the transfer plan matrix, the sum of all elements in that row must equal the normalized absorbance value of the spectral distribution before dilution at the corresponding wavelength. This constraint guarantees that all probabilistic masses of the source distribution are transferred out, and no mass is lost.
[0295] The second type of constraint, called the column sum constraint, ensures that the total probabilistic mass transferred to each target wavelength point is equal to the probability value of the target distribution at that wavelength point. For each column of the transfer plan matrix, the sum of all elements in that column must equal the normalized absorbance value of the diluted target spectral distribution at the corresponding wavelength point. This constraint guarantees that all probabilistic masses of the target distribution are transferred from the source distribution, and the transfer process neither introduces new masses nor results in any loss of mass.
[0296] Furthermore, all elements of the transmission plan matrix must be greater than or equal to zero, because the probability mass cannot be negative. This non-negativity constraint ensures the physical rationality of the transmission plan.
[0297] The number of variables in the linear programming model equals the total number of elements in the transport plan matrix, i.e., 601 multiplied by 601 equals 361201 variables. The number of constraints equals the number of rows and constraints plus the number of columns and constraints, i.e., 601 plus 601 equals 1202 constraints. Due to the enormous number of variables and constraints, directly solving this linear programming problem has extremely high computational complexity, requiring the use of efficient optimization algorithms.
[0298] Dual problem transformation and application of the hypergradient algorithm:
[0299] In this embodiment, to reduce computational complexity, the original linear programming problem is transformed into a dual problem for solution. Duality theory is one of the core theories of linear programming, stating that every linear programming problem has a corresponding dual problem, and the optimal solutions of the primal and dual problems are interconnected. The number of variables in the dual problem is equal to the number of constraints in the primal problem, and much smaller than the number of variables in the primal problem; therefore, the computational complexity of solving the dual problem is significantly reduced.
[0300] The variables in the dual problem are called dual variables, consisting of two sets of dual variables, corresponding to the row and column constraints of the primal problem, respectively. The first set of dual variables has 601 wavelength points, each corresponding to a row and constraint. The second set of dual variables also has 601 wavelength points, each corresponding to a column and constraint. The objective function of the dual problem is to maximize a linear expression, which is equal to the inner product of the normalized absorbance values of each wavelength point in the undiluted spectral distribution with the first set of dual variables, plus the inner product of the normalized absorbance values of each wavelength point in the diluted target spectral distribution with the second set of dual variables.
[0301] The constraint of the dual problem requires that the sum of any two dual variables is less than or equal to the corresponding element of the cost matrix. Specifically, for the element in the i-th row and j-th column of the cost matrix, the sum of the i-th element of the first set of dual variables and the j-th element of the second set of dual variables must be less than or equal to the value of that element in the cost matrix. This constraint holds for all elements of the cost matrix, therefore the number of constraints equals the total number of elements in the cost matrix, which is 361,201.
[0302] The hypergradient algorithm in linear space is an iterative optimization algorithm specifically designed to solve large-scale dual problems. The core idea is to iteratively update the dual variable space along the hypergradient direction, gradually approximating the optimal solution. Hypergradient is a generalization of the gradient concept and is applicable to non-smooth optimization problems.
[0303] The algorithm first performs initialization. All elements of the two sets of dual variables are initialized to zero vectors, indicating that no dual constraints are initially applied. An initial step size parameter is set to control the update magnitude of the dual variables in each iteration. A gradient norm threshold is set for convergence criterion; when the gradient norm is less than this threshold, the algorithm is considered to have converged to the vicinity of the optimal solution, and the iteration can be terminated.
[0304] In each iteration, the algorithm first calculates the transport plan matrix based on the current dual variables. The element in the i-th row and j-th column of the transport plan matrix is determined according to the following rule: if the sum of the i-th element of the first set of dual variables and the j-th element of the second set of dual variables is exactly equal to the element in the i-th row and j-th column of the cost matrix, then that element of the transport plan matrix is equal to the normalized value of the i-th wavelength point of the pre-dilution spectral distribution multiplied by the normalized value of the j-th wavelength point of the post-dilution target spectral distribution; otherwise, that element of the transport plan matrix is equal to zero. This rule is called the complementary relaxation condition and is a core property of duality theory.
[0305] The algorithm then calculates the hypergradient. For the i-th element of the first set of dual variables, its hypergradient equals the normalized value of the i-th wavelength point of the undiluted spectral distribution, minus the sum of all elements in the i-th row of the transmission plan matrix. For the j-th element of the second set of dual variables, its hypergradient equals the normalized value of the j-th wavelength point of the diluted target spectral distribution, minus the sum of all elements in the j-th column of the transmission plan matrix. The hypergradient reflects the degree of deviation between the current transmission plan and the constraints.
[0306] Next, the algorithm determines the adaptive step size. The adaptive step size is calculated using the Armijo backtracking search method. This method starts with a relatively large initial step size, calculating the updated dual variable and corresponding objective function value along the hypergradient direction. If the objective function value is a sufficient improvement compared to the current value, satisfying the Armijo criterion, the step size is accepted; otherwise, the step size is halved and the algorithm retryes until a suitable step size is found or the step size is less than a minimum threshold. The Armijo criterion is that the improvement in the objective function is greater than or equal to a certain proportion of the product of the step size and the square of the hypergradient norm; this proportion is typically set to 0.1.
[0307] Based on the determined adaptive step size and hypergradient, the algorithm updates the dual variables. The i-th element of the first set of dual variables is incremented by the adaptive step size multiplied by its hypergradient value; similarly, the j-th element of the second set of dual variables is also incremented by the adaptive step size multiplied by its hypergradient value. The updated dual variables are used in the next iteration.
[0308] The algorithm calculates the gradient norm of the updated dual variable. The gradient norm is equal to the square root of the sum of the squares of all hypergradient values, measuring the overall size of the hypergradient vector. If the gradient norm is less than a preset gradient norm threshold, such as one ten-thousandth, the algorithm is considered to have converged, the iteration stops, and the current dual variable is output as the converged solution; otherwise, the next iteration continues.
[0309] In a water sample dilution optimization example, after 150 iterations, the gradient norm decreased to 5 x 10^-7, satisfying the convergence condition. The transport plan matrix corresponding to the dual variable at convergence is the optimal transport strategy matrix.
[0310] Mapping from optimal transmission strategy to physical dilution ratio:
[0311] In this embodiment, the optimal transmission strategy matrix describes the optimal transformation relationship between the spectral distributions before and after dilution, but the physically defined dilution ratio is not directly given. A mapping function from the optimal transmission strategy to the physical dilution ratio needs to be established.
[0312] This mapping function is based on Beer-Lambert's law. Beer-Lambert's law states that the absorbance of a solution is directly proportional to the concentration of the absorbing substance in the solution, with the proportionality constant being the product of the molar absorptivity and the optical path length. When a solution is diluted, the concentration of the absorbing substance decreases, and the absorbance decreases accordingly. The diluted concentration is equal to the original water concentration multiplied by 1 minus the dilution ratio, and the diluted absorbance is equal to the original water absorbance multiplied by 1 minus the dilution ratio. Therefore, there is a direct relationship between the dilution ratio and the absorbance scaling factor; that is, 1 minus the dilution ratio equals the diluted absorbance divided by the original water absorbance.
[0313] To calculate the dilution ratio from the optimal transmission strategy matrix, we first need to define the concept of spectral energy. Spectral energy is defined as the weighted sum of absorbance values, with the weights being the normalized absorbance values. The spectral energy before dilution is equal to the inner product of the normalized absorbance values at each wavelength point of the spectral distribution before dilution and the original absorbance values. The target spectral energy after dilution is equal to the inner product of the normalized absorbance values at each wavelength point of the target spectral distribution after dilution and the target absorbance value.
[0314] The mapping function is defined as the dilution ratio equal to 1 minus the target spectral energy after dilution divided by the spectral energy before dilution. Substituting the optimal transmission strategy matrix into this mapping function, the actual optimal dilution ratio is calculated.
[0315] In a specific water sample example, the calculated spectral energy of the raw water was 1500 units, while the target spectral energy after dilution was 300 units. According to the mapping function, the optimal dilution ratio is equal to 1 minus 300 divided by 1500, which equals 1 minus 0.2, or 0.8 or 80%. This means that 80% of the volume of dilution water needs to be mixed with 20% of the volume of raw water to achieve the ideal measurement conditions for the spectrum of the diluted water sample.
[0316] Calculation and execution of dilution control parameters:
[0317] In this embodiment, the specific values of the raw water flow rate and the dilution water flow rate are calculated based on the optimal dilution ratio and the target total flow rate to obtain the dilution control parameters. The target total flow rate is determined by the system design parameters and is usually set to 100 ml per minute. This ensures that the detector has a sufficient water sample supply without the flow rate being too high, which would result in the water sample having too short a residence time in the flow cell.
[0318] The raw water flow rate is calculated by multiplying the target total flow rate by 1 and then subtracting the optimal dilution ratio. The dilution water flow rate is calculated by multiplying the target total flow rate by the optimal dilution ratio. The sum of the two equals the target total flow rate, ensuring flow balance.
[0319] In the example of the optimal 80% dilution ratio above, assuming the target total flow rate is 100 ml / min, the raw water flow rate is equal to 100 multiplied by 1 minus 0.8, which is 100 multiplied by 0.2, equaling 20 ml / min. The dilution water flow rate is equal to 100 multiplied by 0.8, equaling 80 ml / min. The dilution control parameters are the raw water flow rate of 20 ml / min and the dilution water flow rate of 80 ml / min.
[0320] The controller, based on dilution control parameters, precisely adjusts the raw water flow rate and the dilution water flow rate by controlling the proportional valve or peristaltic pump in the adaptive dilution unit. The proportional valve is controlled by outputting a voltage signal to control the valve opening; the voltage signal is proportional to the flow rate. The required control voltage is calculated based on the proportional valve's flow-voltage calibration curve. For example, if the proportional valve's flow-voltage coefficient is 0.1 volts per milliliter per minute, then a raw water flow rate of 20 milliliters per minute corresponds to a control voltage of 2.0 volts, and a dilution water flow rate of 80 milliliters per minute corresponds to a control voltage of 8.0 volts.
[0321] The control method for peristaltic pumps is to control the pump speed by outputting a pulse width modulation signal. The pump speed is directly proportional to the flow rate. The required pump speed is calculated based on the peristaltic pump's flow-speed calibration curve. For example, if the peristaltic pump's flow-speed coefficient is 10 revolutions per minute per milliliter per minute (rpm), then a raw water flow rate of 20 ml / min corresponds to a pump speed of 200 rpm, and a dilution water flow rate of 80 ml / min corresponds to a pump speed of 800 rpm.
[0322] The controller sends control commands to adjust the opening of the proportional valve or the speed of the peristaltic pump. Raw water and dilution water enter the mixer from the raw water path and dilution water path respectively, where they are thoroughly mixed. The mixer employs either a static mixing structure or a dynamic stirring structure to ensure uniform mixing of the raw water and dilution water, eliminating concentration gradients or stratification. The mixing time is typically 1 to 3 seconds, sufficient for thorough mixing.
[0323] After mixing, the diluted water sample flows out of the mixer outlet and enters the spectral detection module for spectral acquisition. The controller monitors the spectral characteristics of the diluted water in real time and fine-tunes the dilution ratio through a closed-loop feedback mechanism. If there is a deviation between the diluted spectrum and the target spectrum, the controller calculates the deviation and adjusts the raw water flow rate and the dilution water flow rate according to the proportional-integral control algorithm to further optimize the dilution effect.
[0324] During the dilution of a water sample, the initial dilution ratio was set to 80%. The peak absorbance of the diluted sample was measured to be 0.48 absorbance units, slightly lower than the target value of 0.50. The controller adjusted the dilution ratio to 78%, increasing the proportion of raw water. The peak absorbance of the diluted sample then rose to 0.50, achieving the ideal measurement conditions, and the dilution optimization was completed.
[0325] Example 5
[0326] This embodiment, based on Embodiment 1, provides a detailed explanation of step S6, "Calculating water quality parameters based on the water quality fingerprint spectral data." This step involves applying principal component analysis to calculate the covariance matrix and decompose the eigenvalues of the original spectral data matrix corresponding to the water quality fingerprint spectral data, identifying spectral feature vectors, and then structurally reconstructing the original spectral data matrix based on these feature vectors to obtain a structured spectral data matrix. Based on this structured spectral data matrix, it is abstracted into a second directed weighted graph, where the vertices represent specific elements or blocks in the structured spectral data matrix, and the edges represent computational dependencies between these elements or blocks. A set of matrix flipping operations is defined, and a flipping operation cost model is constructed to evaluate the computational cost-benefit ratio of each flip in the set, resulting in a flipped graph model of the spectral data. Based on this flipped graph model, a heuristic search function is designed to evaluate the transition from the current state to the target state of the flipped graph model. The cost of state estimation is performed, a priority queue is used to manage the flip operation sequence to be explored, and a pruning strategy is developed to eliminate search branches with non-optimal solutions, resulting in an optimized flip operation sequence. Based on the optimized flip operation sequence, the key matrix operation patterns in water quality fingerprint analysis are identified, and corresponding flip operation sequence templates are designed for each operation pattern to accelerate the calculation of key matrices in water quality fingerprint analysis. The accelerated matrix calculation results are used to optimize memory access, reduce cache miss rate, and optimize single instruction multiple data stream instruction set. Multiple data elements are processed in parallel using an embedded processor to obtain real-time spectral analysis data. Based on the spectral characteristics of the real-time spectral analysis data, pollutant concentration, total organic carbon, and chemical oxygen demand are calculated using a preset water quality parameter calculation model to obtain water quality parameters.
[0327] Principal Component Analysis Dimensionality Reduction and Structure Reconstruction:
[0328] In this embodiment, the water fingerprint spectral data contains absorbance measurements at a large number of wavelengths, with a data dimension as high as 601, corresponding to sampling points every 1 nanometer within the wavelength range of 200 to 800 nanometers. Directly performing matrix operations on high-dimensional data leads to extremely high computational complexity, large memory consumption, and long computation time. To reduce computational complexity and extract key spectral features, this embodiment applies principal component analysis (PCA) to perform dimensionality reduction on the original spectral data matrix.
[0329] Principal component analysis (PCA) is a classic method for dimensionality reduction. It projects high-dimensional data into a low-dimensional subspace through linear transformation while preserving the main variability information. The core steps of PCA include covariance matrix calculation, eigenvalue decomposition, and principal component selection.
[0330] First, the covariance matrix of the original spectral data matrix is calculated. The original spectral data matrix has 601 rows multiplied by the number of samples in the column, where the 601 rows correspond to 601 wavelength points, and the columns correspond to multiple spectral samples collected historically. In the case of a single sample, a sample set can be constructed using historically accumulated spectral data. The covariance matrix describes the correlation between different wavelength points; it has 601 rows and 601 columns and is a symmetric matrix.
[0331] The covariance matrix is calculated by first centering the original spectral data matrix, i.e., subtracting the mean of all samples at each wavelength from the absorbance value at that wavelength, resulting in a centered matrix. Then, the centered matrix is multiplied by its transpose and divided by the number of samples minus 1 to obtain the covariance matrix. The element in the i-th row and j-th column of the covariance matrix represents the covariance between the i-th and j-th wavelengths, reflecting the degree of linear correlation between them.
[0332] Next, eigenvalue decomposition is performed on the covariance matrix. Eigenvalue decomposition decomposes the covariance matrix into the product of an eigenvector matrix and an eigenvalue diagonal matrix. Each column of the eigenvector matrix is an eigenvector, corresponding to an eigenvalue of the covariance matrix. The eigenvalue reflects the magnitude of the variance of the data along the direction of the corresponding eigenvector; the larger the eigenvalue, the more significant the variation in the data along that direction.
[0333] Eigenvalue decomposition is implemented using numerical algorithms, such as the QR algorithm or the Jacobi algorithm. After decomposition, 601 eigenvalues and corresponding 601 eigenvectors are obtained. The eigenvalues are arranged in descending order, and the corresponding eigenvectors are called principal components. The first principal component corresponds to the largest eigenvalue, representing the direction of the greatest data variation; the second principal component corresponds to the second largest eigenvalue, representing the direction of the second greatest data variation, and is orthogonal to the first principal component; and so on.
[0334] The principle of principal component selection is to retain a sufficient variance contribution rate, such as 95%. The variance contribution rate is equal to the sum of the first k eigenvalues divided by the sum of all eigenvalues. By calculating the cumulative variance contribution rate, the number of principal components k to be retained is determined. For example, if the cumulative variance contribution rate of the first 50 principal components reaches 95%, then the first 50 principal components are selected, and the remaining 551 principal components are discarded.
[0335] A dimensionality-reduced matrix is constructed based on the selected k principal components. The dimensionality-reduced matrix has a dimension of 601 rows and k columns, with each column being an eigenvector of a selected principal component. Multiplying the original spectral data matrix by the transpose of the dimensionality-reduced matrix yields a dimensionality-reduced spectral data matrix with dimensions of k rows multiplied by the number of samples. In the case of a single sample, the dimensionality-reduced spectral data matrix is a k-row, 1-column vector.
[0336] The dimensionality-reduced spectral data matrix exhibits structured characteristics. First, the dimensionality is significantly reduced, from 601 dimensions to k dimensions, such as 50 dimensions, resulting in a data volume reduction of approximately 91%. Second, the dimensionality-reduced data is more compactly distributed and more correlated in the principal component space, which is beneficial for subsequent matrix operation optimization. Third, the dimensionality-reduced data may exhibit characteristics such as low rank, sparseness, or block structure, which provide possibilities for matrix flipping optimization.
[0337] After the structured reconstruction is completed, the structured spectral data matrix is output as input data for the subsequent matrix flip graph search algorithm.
[0338] Construction of the matrix flip graph model:
[0339] In this embodiment, to further accelerate matrix operations on spectral data, the structured spectral data matrix is abstracted into a second directed weighted graph, referred to as the matrix flip graph model. This model describes the computational dependencies between matrix elements or blocks using graph theory and defines a series of matrix flip operations. By optimizing the sequence of flip operations, the efficiency of matrix operations is improved.
[0340] The vertices of the matrix flip graph model are defined as follows: each vertex represents a specific element or a sub-matrix block in the structured spectral data matrix. For a dimensionality-reduced spectral data vector with dimensions k rows and 1 column, it can be divided into several blocks, each containing several consecutive elements. For example, a 50-row, 1-column vector can be divided into 10 blocks, each containing 5 elements. Each block corresponds to one vertex, for a total of 10 vertices. The block division follows the principle of data locality, grouping adjacent elements into the same block, which is beneficial for caching.
[0341] The edges in the matrix flip graph model are defined as follows: Edges represent computational dependencies between elements or blocks. If the result of computing block j requires data from block i, then there exists a directed edge from vertex i to vertex j. The weight of an edge represents the cost of data transfer or computation, such as memory access time or the number of computational operations. Computational dependencies are determined based on the specific type of matrix operation, such as matrix multiplication, convolution, or correlation calculations.
[0342] Matrix flipping operations are transformations performed on matrices to change the storage order or arrangement of matrix elements, thereby optimizing the performance of subsequent matrix operations. Common matrix flipping operations include row flipping, column flipping, transposition, and block reassembly.
[0343] Row flipping reverses the row order of a matrix, swapping the first and last rows, the second and second-to-last rows, and so on. Column flipping reverses the column order of a matrix. Transpose swaps rows and columns, changing the element in the i-th row and j-th column of the original matrix to the element in the j-th row and i-th column of the transpose. Block rearrangement rearranges the sub-blocks of a matrix, changing their storage order.
[0344] Each flip operation incurs computational costs, including data access, data movement, and register operations. However, flip operations can also bring benefits to subsequent matrix operations, such as improved data locality, better cache hit rates, and the ability to enable vectorized instructions or parallel computing. The cost-benefit ratio of a flip operation is defined as the benefit divided by the cost; the higher the cost-benefit ratio, the more valuable the flip operation.
[0345] The specific method for constructing the cost model for flip operations is as follows: For each flip operation, calculate its cost. The cost mainly includes the number of elements accessed and the number of memory accesses. For example, a row flip operation requires accessing all elements of the matrix, and the cost is proportional to the total number of matrix elements. For a k-row, 1-column vector, the cost of row flipping is k element accesses.
[0346] The benefit is evaluated based on the speedup effect of the transpose operation on subsequent operations. For example, if a vector is transposed into a row vector with 1 row and k columns, subsequent vector dot product operations can utilize contiguous memory access to improve cache hit rate and reduce performance loss caused by cache misses. Assuming the transpose operation speeds up subsequent operations by 10%, the benefit is quantified as 10% of the subsequent operation time.
[0347] The cost-benefit ratio equals the benefit divided by the cost. For example, the cost of a transpose operation is k element accesses, and the benefit is a 10% speedup for subsequent operations. Assuming the subsequent operations take 100 time units, the benefit is 10 time units, and the cost-benefit ratio is 10 divided by k. If k equals 50, the cost-benefit ratio is 0.2. If another flip operation has a cost-benefit ratio of 2.0, then that operation is preferred.
[0348] The output matrix flip graph model includes a vertex set, edge set, edge weight, flip operation set, and cost model, providing complete graph structure and cost information for subsequent heuristic search.
[0349] Heuristic search algorithms optimize the flip operation sequence:
[0350] In this embodiment, the goal is to find a sequence of flip operations such that applying these operations sequentially to the matrix results in the fastest subsequent matrix operations. This problem can be formalized as a state-space search problem, and the A* search algorithm can be used to find the optimal sequence of flip operations.
[0351] A state is defined as the current form of a matrix, i.e., the matrix after a series of flip operations. The initial state is the original structured matrix, and the target state is the matrix form best suited for subsequent operations. An action is defined as an operation in the set of flip operations; after applying an operation from the current state, the user transitions to the new state.
[0352] The core of the A* search algorithm is the cost function, which consists of two parts: the cumulative cost *g* from the initial state to the current state, and the estimated cost *h* from the current state to the target state. The total cost *f* equals *g* plus *h*. The algorithm maintains a priority queue, storing states to be explored in ascending order of total cost *f*. Each time, the state with the smallest *f* value is taken from the priority queue, expanded, a successor state is generated, its cost is calculated, and the successor state is added to the priority queue. This process is repeated until the target state is reached or the priority queue is empty.
[0353] The cumulative cost g is calculated based on the cost of the flip operations. The cumulative cost from the initial state to the current state is equal to the sum of the costs of all flip operations performed along the path. For example, if the current state is obtained through two operations, transpose and reassemble, then g equals the cost of the transpose operation plus the cost of the reassemble operation.
[0354] The estimated cost *h* is a heuristic function that evaluates the estimated cost of moving from the current state to the target state. The design of the heuristic function must satisfy acceptability, meaning the value of *h* should not exceed the true optimal cost, ensuring that the A* algorithm can find the optimal solution. The heuristic function considers the following factors: data locality, cache friendliness, and parallel computing opportunity.
[0355] Data locality measures the contiguousness of matrix elements in memory. Higher contiguousness leads to higher cache hit rates and lower memory access latency. Locality is quantified by calculating the interval between adjacent elements in memory; the smaller the interval, the better the locality.
[0356] Cache friendliness measures how well the matrix access pattern matches the cache line size. Modern processors load data in cache lines, typically 64 bytes each. High cache friendliness means the matrix access pattern fully utilizes the data within a cache line and avoids wasting data within that line.
[0357] Parallel computing measures whether a matrix structure is suitable for Single Instruction Multiple Data Stream (SID) vectorization or multithreaded parallel computation. Vectorization requires data to be aligned to specific boundaries and the computation pattern to be regular. Multithreaded parallelism requires that the task can be decomposed into multiple independent subtasks.
[0358] The heuristic function is defined as a weighted sum of data locality score, cache friendliness score, and parallelism score. The score is normalized to the range of 0 to 1, with higher scores indicating better performance. The heuristic cost h is defined as 1 minus the score, such that states with higher scores have smaller h values, and are closer to the target state. The weighting coefficients are determined based on the importance of each factor in the actual application.
[0359] The specific execution process of the A* search algorithm is as follows: Initialize the priority queue, add the initial state to the queue. The initial state's g value is 0, the h value is calculated according to the heuristic function, and the f value is equal to the h value. Take the state with the smallest f value from the priority queue as the current state. If the current state is the target state, the search ends, and the operation sequence from the initial state to the current state is output.
[0360] Try all possible flip operations on the current state to generate a successor state. For each successor state, calculate a new g value, which is equal to the current state's g value plus the cost of the flip operation. Calculate the h and f values of the successor state. If the successor state has not been visited, or the new g value is less than the recorded g value, add the successor state to the priority queue and update its priority.
[0361] To improve search efficiency, pruning strategies are developed to eliminate search branches that lead to non-optimal solutions. These strategies include using a hash table to record visited states to avoid redundant exploration; maintaining the cost of the current optimal solution as an upper bound, pruning branches if the f-value of a state is greater than or equal to this upper bound; and terminating operation sequences that clearly will not produce the optimal solution early, such as applying the same flip operation twice consecutively, which is equivalent to not performing the operation at all.
[0362] In a spectral data matrix optimization problem, the A* search algorithm, after 80 state expansions, found the optimal sequence of flip operations: transpose, block recombination, and row flipping. This sequence improves the speed of subsequent matrix multiplication operations by 8 times.
[0363] It outputs an optimized sequence of flip operations, providing an operation template for accelerating computation of key matrix operation patterns.
[0364] Accelerated computation of key matrix operation patterns:
[0365] In this embodiment, water fingerprint analysis involves various matrix operations, including convolution for spectral smoothing and feature extraction, correlation calculation for spectral matching and pattern recognition, and matrix multiplication for linear transformation and projection. These key matrix operation patterns are identified, and dedicated flipping operation sequence templates are designed for each pattern to accelerate computation.
[0366] Convolution operations involve sliding windows and element-wise multiplication and addition. To improve cache hit rate, the matrix is stored in row-major order, and the flipping operation sequence is already row-major and does not require additional flipping. The elements of the convolution kernel are multiplied and accumulated one by one with the elements in the matrix window, accessing the matrix row elements sequentially, making full use of the data locality of cache rows.
[0367] Correlation calculations require computing the inner product of two vectors. To leverage single-instruction multiple-data vectorization, vector elements are aligned to 16-byte boundaries, eliminating the need for additional flipping operations after data alignment. Vectorized instructions can compute multiplication and addition of multiple elements at once, significantly improving computation speed.
[0368] Matrix multiplication is the most computationally complex operation, with its complexity proportional to the cube of the matrix's dimension. To accelerate computation, the Strassen algorithm or the Coppersmith-Winograd algorithm is used. The optimal algorithm is selected for submatrices of different sizes and structures.
[0369] Small matrices are computed directly, with no flipping operation sequence. Medium-sized matrices are computed using the Strassen algorithm, where the flipping operation sequence is calculated recursively after partitioning the matrix. The Strassen algorithm divides the matrix into four submatrices and recursively calculates seven submatrix multiplications instead of the eight required by the traditional method, reducing computation by approximately 14%.
[0370] Large matrices are processed using a block-based parallel algorithm, where the operation sequence is reversed by dividing the matrix into blocks, computing them in parallel, and then merging the results. The matrix is divided into multiple sub-blocks, each assigned to a processor core for parallel computation, and the results are then merged. Parallel computation fully utilizes the computing power of multi-core processors, reducing computation time.
[0371] In the calculation of a certain water quality parameter, a matrix multiplication of 50 rows and 50 columns is required. By applying the Strassen algorithm, the matrix is divided into four submatrices of 25 rows and 25 columns each, and the submatrix multiplications are recursively calculated seven times, reducing the computational load by approximately 14%. The matrix calculation time after acceleration is reduced to 86% of the original time.
[0372] Hardware-level memory access and vectorization optimization:
[0373] In this embodiment, based on matrix flipping and fast algorithms, further hardware-level optimizations are implemented, including memory access optimization and single instruction multiple data (SID) instruction set optimization.
[0374] Memory access optimization employs the following strategies: It uses a row-major or column-major sequential access pattern, selecting an appropriate storage order based on the characteristics of matrix operations. For example, in matrix multiplication, if the first matrix is accessed row-wise and the second matrix is accessed column-wise, then the first matrix is stored in row-major order, and the second matrix is stored in column-major order, or transposed and then stored in row-major order, ensuring the continuity of the access pattern.
[0375] Cache prefetching reduces cache misses by prefetching the next data block into the cache while accessing the current data. Embedded processors typically support hardware prefetching mechanisms or can explicitly trigger prefetching via software prefetch instructions. For example, the PLD instruction in ARM processors is used to preload data into the cache.
[0376] Implement data alignment and padding to align data to cache line boundaries, improving memory bandwidth utilization. For data that is less than a full cache line, padding is performed to ensure it occupies a complete cache line. For example, a 50-row, 1-column vector is padded to 64 rows and 1 column, with 14 zero elements added to ensure it occupies a complete cache line, thus avoiding wasted data within the cache line.
[0377] Single-Instruction Multiple-Data (SMD) instruction set optimizations leverage the vector computing capabilities of embedded processors. ARM processors support the NEON instruction set, while Intel processors support SSE or AVX instruction sets, enabling parallel processing of multiple data elements with a single instruction. This implements vectorized matrix multiplication, addition, and element-wise operations, converting scalar operations into vector operations.
[0378] The implementation of vectorized matrix addition is as follows: For vector addition c equals a plus b, use NEON's VADD instruction to calculate the addition of four single-precision floating-point numbers at a time. Iterate through all elements of the vector, processing four elements at a time, until all elements have been calculated. If the vector length is not a multiple of 4, the remaining elements are processed as scalars.
[0379] The vectorized matrix multiplication method is implemented as follows: For each row i of matrix A and each column j of matrix B, process 4 columns at a time. Load the elements of the i-th row of A into a vector register, load the elements of the j-th to j+3-th columns of B into a vector register, calculate the vector multiplication to obtain intermediate results, accumulate the intermediate results into a result vector, and store the result vector in the elements of the i-th row and j-th to j+3-th columns of C. Through vectorization, the speed of matrix multiplication is improved by 4 times.
[0380] Parallel processing leverages the multi-core resources of embedded processors. Matrix operation tasks are distributed across multiple CPU cores for parallel execution. For example, the matrix is divided into blocks, with each core computing one sub-block. Multithreaded programming, employing POSIX threads or the OpenMP parallel programming framework, is used to achieve task parallelism.
[0381] On a quad-core ARM processor, a 50x50 matrix multiplication is decomposed into four 25x50 subtasks, with each core calculating one subtask, reducing the parallel execution time to one-quarter of that of a single core.
[0382] It outputs real-time processed spectral analysis data, including dimensionality-reduced spectral feature vectors and intermediate results after accelerated calculation, providing an efficient data foundation for water quality parameter calculation.
[0383] Calculation and output of water quality parameters:
[0384] In this embodiment, based on real-time processed spectral analysis data, a preset water quality parameter calculation model is applied to calculate various water quality parameters, including total organic carbon, chemical oxygen demand, and pollutant concentration.
[0385] Total organic carbon (TOC) exhibits a linear relationship with UV absorbance at 254 nm. The calculation model states that TOC equals an empirical coefficient multiplied by the absorbance at 254 nm plus a bias term. The empirical coefficient and bias term are determined through calibration using standard samples. For example, at a water quality monitoring station, the calibration coefficient corresponds to 2.5 mg / L TOC per absorbance unit, and the bias term is 0.5 mg / L. When the absorbance at 254 nm is 0.6 absorbance units, the TOC is calculated as 2.5 multiplied by 0.6 plus 0.5, which equals 1.5 plus 0.5, resulting in 2.0 mg / L.
[0386] Chemical oxygen demand (COD) is correlated with multi-band spectral characteristics, and a multiple linear regression model was used. The calculation model states that COD equals the intercept term plus the sum of the products of the absorbance at each characteristic wavelength and the corresponding regression coefficient. The regression coefficients were obtained by least-squares fitting using spectral data of standard samples and known COD values. The fitting objective was to minimize the sum of squares of the differences between the predicted and actual values.
[0387] For example, a water quality monitoring station selected three characteristic wavelengths: 254 nm, 280 nm, and 365 nm. The regression coefficients were 5.0 for the intercept term, 10.0 for 254 nm, 8.0 for 280 nm, and 6.0 for 365 nm. When the absorbance at 254 nm was 0.6, at 280 nm it was 0.5, and at 365 nm it was 0.3, the chemical oxygen demand (COD) was calculated as 5.0 + 10.0 × 0.6 + 8.0 × 0.5 + 6.0 × 0.3, which equals 5.0 + 6.0 + 4.0 + 1.8, resulting in 16.8 mg / L.
[0388] Pollutant concentration calculations are based on absorbance or fluorescence intensity at a specific wavelength. For example, copper ions have a characteristic absorption peak at 600 nm, and their concentration is directly proportional to absorbance. The calculation model is that the copper ion concentration equals the molar absorptivity multiplied by the absorbance at 600 nm. The molar absorptivity is determined through calibration with standard solutions.
[0389] It outputs water quality parameters, including total organic carbon, chemical oxygen demand, and pollutant concentration, and completes the collection and preprocessing of water quality fingerprints to provide accurate data support for water quality monitoring and assessment.
[0390] Example 6
[0391] This embodiment, based on Embodiment 5, provides a system operation visualization method. It acquires real-time processed spectral analysis data, water quality monitoring system operating status parameters, control command execution feedback, and abnormal event alarm information. The data is organized by timestamp and stored in an operation record database to obtain operation records. Data within a preset time window is extracted from the operation records, and outlier removal, missing value imputation, and data smoothing preprocessing are performed to obtain preprocessed operation data. Based on the preprocessed operation data, the real-time processed spectral analysis data is plotted as a spectral curve, and the current values of water quality parameters are displayed in a digital panel, achieving real-time visualization of the system's operating status.
[0392] Data collection and time series organization:
[0393] In this embodiment, the water quality monitoring system generates a large amount of real-time data during operation, requiring systematic collection, organization, and storage. The operational data includes several types: real-time spectral analysis data records the absorbance and fluorescence spectra obtained from each spectral scan, reflecting the optical characteristics of the water quality; operational status parameters record the real-time status of each system component, including flow rate measurements from flow sensors, pressure measurements from pressure sensors, solenoid valve opening / closing status detected by solenoid valve status sensors, and pump speed measurements from pump speed sensors; control command execution feedback records various control commands sent by the controller and their execution status, such as confirmation of solenoid valve opening / closing commands and pump speed adjustment commands; and abnormal event alarm information records abnormal situations detected by the system, such as filter clogging, pump performance degradation, and window contamination.
[0394] The controller acquires the aforementioned data in real time via a data acquisition module. This module connects to the sensors and actuators using standard communication protocols, such as Modbus or CAN bus. The data acquisition frequency is determined based on the data type and monitoring requirements; for example, spectral data is acquired once per measurement cycle, operating status parameters are acquired once per second, and control commands and abnormal events are recorded in real time.
[0395] Each data record is assigned a timestamp, accurate to the millisecond, in the format of year, month, day, hour, minute, second, point, millisecond, for example, October 15, 2023, 14:32:18:523. The timestamp identifies the moment the data was collected and is a key index for organizing time-series data.
[0396] Data is organized by timestamp to form time-series records. Each time-series record uses its timestamp as the primary key, linking various data types to the corresponding moment. For example, a record for a specific moment might include fields such as timestamp, spectral data array, flow rate value, pressure value, solenoid valve status dictionary, pump speed value, control command string, and exception event string.
[0397] Time-series records are stored in a runtime database. This embodiment uses a time-series database, such as InfluxDB or TimescaleDB. Time-series databases are optimized for time-series data, supporting efficient time-range queries, data aggregation, and compressed storage, making them suitable for long-term storage of large amounts of runtime data.
[0398] The database table structure is designed as follows: The table name is system_runtime_log, and the fields include timestamp (primary key), spectrum_data (spectral data stored in JSON format), flow_rate (flow rate floating-point number), pressure (pressure floating-point number), valve_states (solenoid valve status stored in JSON format), pump_speed (pump speed integer), control_command (control command string), and alarm_event (alarm event string). The controller inserts a record into the database at regular intervals, forming a continuous running record.
[0399] Output the execution records and store them in the execution record database for subsequent querying, analysis and visualization.
[0400] Preprocessing and quality control of runtime data:
[0401] In this embodiment, in order to generate high-quality visualization graphics, it is necessary to extract data within a specific time window from the running record database and perform preprocessing, including outlier removal, missing value imputation, and data smoothing.
[0402] The time window selection is determined based on user needs. Common time windows include those displaying short-term trends for the most recent hour, intraday variations for the most recent 24 hours, and cyclical patterns for the most recent 7 days. Data within the time window is retrieved from the runtime database using SQL queries specifying the time range. For example, to query data for the most recent hour, the query condition is that the timestamp is greater than or equal to the current time minus one hour, sorted in ascending order of timestamp. The query results include all records within the time window.
[0403] Outlier removal is used to eliminate abnormal measurements in the operational data, such as abnormal readings caused by sensor malfunctions. This embodiment uses the three-standard-deviation criterion or the interquartile range method to remove outliers.
[0404] The three-standard-deviation criterion involves the following steps: Calculate the mean and standard deviation of the data, and then remove data points that exceed the range of the mean plus or minus three standard deviations. For example, if the mean of flow rate data is 100 ml / min and the standard deviation is 5 ml / min, the normal range is 100 minus 15 to 100 plus 15, which is 85 to 115 ml / min. If a data point has a flow rate of 120 ml / min, exceeding the upper limit of the normal range, it is considered an outlier and removed.
[0405] The steps of the interquartile range (IQR) method are as follows: Calculate the first quartile (Q1) and the third quartile (Q3) of the data. The IQR is equal to Q3 minus Q1. Data points outside the range of Q1 minus 1.5 times IQR to Q3 plus 1.5 times IQR are removed. The IQR method makes fewer assumptions about the data distribution and is suitable for non-normally distributed data.
[0406] Missing value imputation is used to fill in missing records in operational data, such as data loss caused by sensor communication interruptions. This embodiment uses linear interpolation or a time series model to imput missing values.
[0407] The steps of linear interpolation are as follows: For the time t corresponding to the missing value, find the time t-1 corresponding to the previous valid value and the time t-1 corresponding to the next valid value. Calculate the missing value as the average of the two valid values. For example, the flow rate at time 99 is 98 ml / min, the flow rate at time 101 is 102 ml / min, and the flow rate at time 100 is missing. Through linear interpolation, the calculated value is 98 + 102 divided by 2, which equals 100 ml / min.
[0408] Time series model imputation uses autoregressive moving average or exponential smoothing models to predict missing values based on the temporal evolution of historical data. Time series models consider the temporal correlation and trend changes of the data, resulting in more accurate imputation results.
[0409] Data smoothing is used to remove high-frequency noise from the running data, making the visualized curves smoother. This embodiment uses moving average or Savitzky-Golay filtering for data smoothing.
[0410] The steps of a moving average are as follows: For a data point t, calculate the average of w data points before and after it as the smoothed value. The window size w controls the degree of smoothing; the larger w is, the stronger the smoothing effect, but it may lose detailed information. For example, with a window size of 5, when performing an 11-point moving average on traffic data, the smoothed value of each data point is equal to the average of the 11 points before and after it (5 points before and after).
[0411] Savitzky-Golay filtering uses polynomial fitting to smooth data while preserving its trend characteristics. This filtering method achieves smoothing while retaining local features by fitting a low-order polynomial within a moving window and replacing the original data points with the polynomial values.
[0412] The output is preprocessed running data, including outlier removal, missing value imputation, and smoothed time series data, providing a high-quality data foundation for the generation of visualization graphics.
[0413] Spectral curve plotting and digital panel display:
[0414] In this embodiment, based on the preprocessed operational data, a visual graph is generated and displayed to the user through a touchscreen human-computer interaction interface. The visual content includes spectral curves and digital panels.
[0415] Spectral curve plotting converts real-time spectral data into wavelength-absorbance curves. The horizontal axis represents the wavelength range of 200 to 800 nanometers, and the vertical axis represents the absorbance range of 0 to 3 absorbance units. Plotting methods include line graphs and smooth curves. Line graphs connect the absorbance values at each wavelength point to form a broken line. Smooth curves use spline interpolation to generate a smooth curve, resulting in a more aesthetically pleasing visual effect.
[0416] The spectral curves can overlay and display the spectra at multiple points in time, distinguished by different colors, allowing users to observe the trend of spectral changes over time. For example, a blue curve represents the spectrum at the current moment, a green curve represents the spectrum from 1 hour ago, and a red curve represents the spectrum from 24 hours ago. By comparing the spectral curves at different times, users can intuitively understand the temporal evolution of water quality.
[0417] The digital panel displays the current values of water quality parameters in digital form. The digital panel includes the parameter name, value, unit, and status indicator. For example, total organic carbon is displayed as 15.3 mg / L, indicating a normal status, and is indicated in green; chemical oxygen demand is displayed as 42.8 mg / L, indicating a normal status, and is indicated in green; turbidity is displayed as 8.5 NTU, indicating a normal status, and is indicated in green; flow rate is displayed as 98 mL / min, indicating a normal status, and is indicated in green.
[0418] The status indicator is determined based on the comparison between the parameter value and the threshold. If the parameter value is within the normal range, the status is displayed as normal and indicated by green; if the parameter value is close to the threshold, the status is displayed as a warning and indicated by yellow, reminding the user to pay attention; if the parameter value exceeds the threshold, the status is displayed as abnormal and indicated by red, prompting the user to take action.
[0419] The system topology diagram graphically displays the system's topology structure, using different colors to indicate the operational status of each component. Green indicates a component is operating normally, yellow indicates a component's performance is deteriorating and requires attention, and red indicates a component is faulty and needs repair. Users can click on a component in the topology diagram to view its detailed parameters and historical operation records.
[0420] Trend curves plot the changes of water quality parameters over time. The horizontal axis represents time, and the vertical axis represents the parameter value. For example, plotting the total organic carbon trend curve for the most recent 24 hours allows observation of the intraday variation of total organic carbon. Trend curves can be plotted as line graphs or area graphs; area graphs fill the area below the curve for a more intuitive visual effect.
[0421] The touchscreen human-machine interface supports multiple interactive functions. Touch to view detailed information: Clicking on a component or parameter will bring up a detailed information window displaying real-time data, historical trends, set thresholds, and other information for that component or parameter; Drag to zoom the time axis: Drag the time axis on the trend curve to zoom or pan the time range and view historical data for different time periods; Switch display views: Switch between spectral view, parameter view, and topology view to meet different viewing needs; Manual control: Users can manually control flow path switching, set alarm thresholds, start the cleaning process, etc., enabling manual intervention and debugging of the system.
[0422] The system outputs a visual system status, which is displayed in real time through a touchscreen human-computer interaction interface, improving the user experience and allowing users to fully understand the system's operating status and promptly detect and handle abnormal situations.
[0423] Example 7
[0424] This embodiment, based on Embodiment 2, provides a detailed explanation of step S1.2, "Based on the system topology graph model, apply the depth-first search algorithm to identify all cut edges and decompose the system topology graph model into multiple bilaterally connected components." This step initializes the depth-first search traversal process, recording the visit timestamp and minimum reach timestamp of each vertex. Starting from the initial vertex, it recursively traverses all vertices and edges of the system topology graph model to obtain the traversal results. Based on the traversal results, it determines the relationship between the minimum reach timestamps of the two endpoints of an edge. When the minimum reach timestamp of one endpoint of an edge is greater than the visit timestamp of the other endpoint, the edge is marked as a cut edge, resulting in a set of all cut edges. Based on the set of all cut edges, all cut edges are deleted from the system topology graph model, decomposing the system topology graph model into multiple connected subgraphs, where there are at least two non-repeating paths between any two vertices within each connected subgraph, resulting in the bilaterally connected component decomposition result.
[0425] In this embodiment, the depth-first search algorithm starts from the initial vertex of the system's topology graph and recursively visits all reachable vertices and edges. During the traversal, the algorithm records two key timestamps for each vertex: the visit timestamp and the lowest reachable timestamp.
[0426] The visit timestamp indicates the moment when the vertex is first visited during the depth-first search. The lowest reachable timestamp indicates the visit timestamp of the earliest ancestor vertex that can be reached from this vertex through the edges and back edges of the depth-first search tree. These two timestamps are the core basis for identifying cut edges.
[0427] Before the algorithm begins, initialization is performed. All vertices in the system topology graph are marked as unvisited, indicating that these vertices have not yet been explored by the search algorithm. A global time counter is initialized to zero and increments each time a new vertex is visited, assigning a unique visit timestamp.
[0428] For each vertex, prepare three data structure spaces: an access timestamp array to record the time when the vertex is first visited, a lowest reachable timestamp array to record the time when the earliest ancestor that can be reached from the vertex, and a parent node array to record the parent node of the vertex in the depth-first search tree, which is used to determine the type of edge and avoid back edges that backtrack to the parent node.
[0429] The recursive process of depth-first search begins with a selected starting vertex. The starting vertex is typically an important node in the system topology graph, such as a detector node or a master pump node. When a vertex is visited, it is first marked as visited to prevent duplicate visits. Then, a global time counter is incremented, and the current time is assigned to both the vertex's visit timestamp and its lowest reachable timestamp. Initially, the vertex's lowest reachable timestamp equals its visit timestamp, indicating that no earlier reachable ancestor has been found.
[0430] Next, the algorithm iterates through all neighboring vertices of the current vertex. Neighboring vertices are connected to the current vertex via edges in the system's topology graph. For each neighboring vertex, the algorithm needs to determine its visit status and adopt different processing strategies based on different statuses.
[0431] If a neighboring vertex has not yet been visited, it means that the neighboring vertex is a new node in the depth-first search tree. The algorithm sets the current vertex as the parent node of the neighboring vertex and records it in the parent node array. Then, it recursively calls the depth-first search function to visit the neighboring vertex, exploring the subtree rooted at that neighboring vertex.
[0432] After recursion returns, the algorithm needs to update the lowest reachable timestamp of the current vertex. The update rule is to take the smaller value between the lowest reachable timestamp of the current vertex and the lowest reachable timestamps of its neighbors. This update operation reflects the transitivity of lowest reachable timestamps: if an early ancestor can be reached from a neighboring vertex, then that ancestor can also be reached from the current vertex through its neighbors.
[0433] At this point, the algorithm checks whether the edge connecting the current vertex and its neighboring vertices is a cut edge. The criterion for a cut edge is: if the lowest reachable timestamp of a neighboring vertex is strictly greater than the visit timestamp of the current vertex, then the edge is a cut edge. This condition means that starting from a neighboring vertex, it is impossible to reach the current vertex or its ancestor through either a depth-first search tree edge or a back edge. In other words, this edge is the only bridge connecting the subtree containing the neighboring vertex to the rest of the graph; deleting this edge would break the connectivity of the graph. Edges that satisfy the condition are marked as cut edges and added to the cut edge set.
[0434] If a neighboring vertex has already been visited and is not the parent of the current vertex, then a back edge has been found. A back edge is an edge connecting a node in a depth-first search tree to an ancestor node that is not its direct parent. The existence of a back edge means that the current vertex can reach an earlier ancestor vertex through this back edge.
[0435] The algorithm updates the minimum reachable timestamp of the current vertex by taking the smaller value between the current value and the visit timestamps of neighboring vertices. Note that the visit timestamps of neighboring vertices are used here, not the minimum reachable timestamp, because back edges directly connect to ancestor nodes, and the visit timestamps of ancestor nodes are the reachable timestamps. Through this update, the minimum reachable timestamp of the current vertex is reduced, reflecting the ability to reach earlier ancestors via back edges.
[0436] The following explanation uses a water quality monitoring system as an example. This system comprises eleven key components: an upstream sampling point, a tributary sampling point, a main stream sampling point, three solenoid valves, a filter, a pump, a mixer, a dilution water source, and a detector. These components are connected via pipelines to form the system topology.
[0437] The depth-first search starts from the upstream sampling point. The algorithm first visits the upstream sampling point, setting its visit timestamp and lowest reachable timestamp to 1. Then it visits the upstream sampling point's neighbor, the solenoid valve, setting its timestamp to 2 and its parent node to the upstream sampling point. Next, it visits the filter (timestamp 3, parent node: solenoid valve). It continues visiting the pump (timestamp 4, parent node: filter). Then it visits the mixer (timestamp 5, parent node: pump). Starting from the mixer, it visits the detector (timestamp 6, parent node: mixer). Since the detector has no unvisited neighbors, the algorithm recursively returns to the mixer.
[0438] The algorithm then checks the edge from the mixer to the detector. The detector's lowest reachable timestamp is 6, and the mixer's access timestamp is 5. Since 6 is greater than 5, the edge-cutting condition is met, and therefore the edge from the mixer to the detector is marked as a cut edge. This means that if this edge is broken, the detector will lose connection to the rest of the system and will be unable to perform water quality testing.
[0439] The algorithm continues accessing the dilution source from the mixer, setting the timestamp of the dilution source to 7, with the mixer as its parent node. The dilution source has no unvisited neighbors, so the algorithm recursively returns to the mixer. The algorithm checks the edge from the dilution source to the mixer. The lowest reachable timestamp of the dilution source is 7, and the visit timestamp of the mixer is 5. Since 7 is greater than 5, this edge is also marked as a cut edge, indicating that the pipe between the dilution source and the mixer is the only channel for the dilution function.
[0440] The recursion continues and the algorithm checks the edge from the pump to the mixer. The mixer's lowest reachable timestamp is 5, and the pump's access timestamp is 4. Since 5 is greater than 4, this edge is marked as a cut edge. Similarly, the edge from the filter to the pump is also marked as a cut edge.
[0441] After the depth-first search has visited all vertices reachable from the upstream sampling point, the algorithm checks if there are any unvisited vertices. If there are unvisited vertices, the graph is not connected, and the depth-first search needs to be restarted from the other unvisited vertex. In this example, the tributary sampling point and the main stream sampling point also need to be visited; they are connected to the filter via their respective solenoid valves and ultimately flow into the same flow path system.
[0442] After the depth-first search algorithm has traversed all vertices, it outputs the traversal results. The results include the visit timestamp, lowest reachable timestamp, parent node information, and a list of all edges marked as cut edges for each vertex. This information provides a complete data foundation for subsequent cut edge identification and bilateral connected component decomposition.
[0443] In this embodiment, based on the traversal results of the depth-first search, the algorithm systematically examines each edge in the graph to determine whether it is a cut edge. The determination of cut edges is based on the timestamp relationship between the two endpoints of the edge, which is an important method in graph theory for identifying network vulnerabilities.
[0444] For each edge in the depth-first search tree, i.e., the edge connecting a parent node and a child node, the algorithm applies the cut edge criterion. Let the two endpoints of the edge be vertex u and vertex v, where u is the parent node of v. If the lowest reachable timestamp of vertex v is strictly greater than the visit timestamp of vertex u, then the edge is a cut edge.
[0445] The deeper meaning of this condition can be understood from the perspective of graph connectivity. The lowest reachable timestamp of vertex v represents the time when the earliest ancestor can be reached from v. If this time is strictly greater than the visit time of u, it means that all paths from v cannot reach u or u's ancestors, and can only reach vertices visited after u. Therefore, the edge u to v is the only connection between the subtree containing v and the rest of the graph. Deleting this edge will separate the subtree containing v from the rest of the graph, destroying the connectivity of the graph.
[0446] The algorithm iterates through all tree edges recorded during the depth-first search, applying a cut edge criterion to each edge. Edges that meet the criteria are added to the cut edge set. It's important to note that back edges in the graph cannot be cut edges, because their existence itself provides an alternative path; deleting a back edge does not disrupt the graph's connectivity. Therefore, the algorithm only needs to check tree edges.
[0447] In the water quality monitoring system example, the algorithm identified four cut edges, corresponding to pipelines that are critical bottlenecks in the system's flow path. The first cut edge is the pipeline from the filter to the pump. This pipeline is the only channel for water samples to flow from the front-end sampling and filtration module to the back-end pumping and detection module. If this pipeline is blocked or ruptured, the water sample cannot reach the pump, the pump cannot provide power to the system, and the entire monitoring process will be interrupted.
[0448] The second cut-off line is the pipeline from the pump to the mixer. The pump delivers the pressurized water sample to the mixer for dilution. This pipeline is the only path connecting the pumping and dilution functions. If this pipeline malfunctions, such as a pipe rupture causing leakage, or internal scaling increasing flow resistance, the water sample cannot enter the mixer, the dilution function fails, and high-concentration water samples may exceed the detector's measurement range, leading to decreased detection accuracy or detection failure.
[0449] The third pipe is the main pipeline from the mixer to the detector. The diluted water sample flows into the detector through this pipeline for spectral analysis. This pipeline is the only channel for the water sample to reach the detector and carries the final output function of the monitoring system. If this pipeline is blocked, for example due to residual particulate matter deposited in the water sample, or because a loose pipeline joint allows gas to enter and form bubbles, the detector will not receive a stable supply of water sample, and the spectral measurement results will exhibit abnormal fluctuations or be interrupted.
[0450] The fourth section is the pipeline from the dilution water source to the mixer. The dilution water is injected into the mixer through this pipeline and mixed with the raw water according to a set ratio. If this pipeline becomes clogged, for example, due to algae or biofilm formation on the inner wall after long-term use, or due to insufficient pressure in the dilution water supply, the system will be unable to perform dilution processing and will only be able to detect the raw water directly. For high-concentration water samples, this will cause detector saturation and loss of measurement accuracy; for water samples requiring multiple dilutions, it may even damage the detector's optical components.
[0451] The algorithm records the information of these four cut edges into a cut edge set. The information for each cut edge includes detailed parameters such as the vertex identifiers of the two endpoints, the edge weight, the physical component type corresponding to the edge, the pipe length, and the pipe material. This information provides crucial decision-making basis for subsequent redundancy design and system maintenance.
[0452] The output of the cut edge set not only indicates the location of single points of failure in the system, but also provides quantitative indicators for system reliability assessment. The vulnerability of the system can be measured by indicators such as the number of cut edges, the sum of cut edge weights, and the number of nodes affected after a cut edge is broken. The more cut edges there are, the more potential points of failure the system has; the larger the sum of cut edge weights, the longer the total length of the system's critical pipelines, and the higher the risk of failure; the more nodes affected after a cut edge is broken, the wider the impact range of a single point of failure.
[0453] In this embodiment, the algorithm decomposes the system topology graph based on the identified set of cut edges. The decomposition method involves removing all cut edges from the original graph, dividing it into multiple disconnected subgraphs. Each subgraph is called a bilaterally connected component, which possesses important fault-tolerant properties.
[0454] Bilaterally connected components possess the following core property: Within a bilaterally connected component, there exist at least two paths between any two vertices that do not repeat. This means that there are no cut edges within a bilaterally connected component, and deleting any edge within the component will not disrupt its connectivity. Bilaterally connected components exhibit high fault tolerance; even if any edge fails, the system can still maintain connectivity through alternative paths.
[0455] The algorithm first creates a copy of the original system's topology graph to preserve complete information about the original graph for subsequent analysis. Then, it removes all edges from the cut edge set in the copy. The deletion operation is implemented by modifying the graph's data structure. For a graph represented by an adjacency list, deleting an edge u to v means removing vertex v from the adjacency list of vertex u, and simultaneously removing vertex u from the adjacency list of vertex v, maintaining the symmetry of the undirected graph. For a graph represented by an adjacency matrix, deleting an edge u to v means setting the elements in the u-th row and v-th column of the matrix to infinity or a special flag indicating that the edge does not exist.
[0456] After removing all cut edges, the original connected graph is decomposed into multiple disconnected subgraphs. The algorithm uses breadth-first search or depth-first search to identify these connected subgraphs. Starting the search from any unvisited vertex, all reachable vertices are marked as belonging to the same connected component. This process is repeated, starting a new search from another unvisited vertex, until all vertices are assigned to a connected component.
[0457] In the water quality monitoring system example, after removing the four cut edges, the system is decomposed into five bilaterally connected components. The first component contains three sampling points, three solenoid valves, and a filter. Multiple paths exist between these components; for example, the water sample can reach the filter directly via the upstream solenoid valve, or it can detour via a tributary solenoid valve, or it can reach it via the main solenoid valve. This multi-path structure gives this component high fault tolerance; if any one pipeline fails, the water sample can still reach the filter via other paths.
[0458] The second component contains only one vertex: the pump. The pump is connected to the filter and mixer via two cut edges, and is not part of any two-sided connected component, making it a single point of failure. Pump failure will paralyze the entire system because there is no backup pump to provide replacement functionality.
[0459] The third component contains only one vertex: the mixer. The mixer is connected to the pump, dilution water source, and detector via three cut edges, representing a single point of failure. A mixer failure will prevent dilution processing, affecting the detection accuracy of high-concentration water samples.
[0460] The fourth component contains only one vertex: the dilution source. The dilution source is connected to the mixer via a cut edge and is the sole supply source for the dilution function. Failure of the dilution source will cause the dilution function to completely fail.
[0461] The fifth component contains only the vertex of the detector. The detector is connected to the mixer via an edge cut and is the final output of the monitoring system. A detector failure will prevent spectral detection, causing the system to lose its monitoring capability.
[0462] The algorithm outputs the decomposition results of bilaterally connected components. The decomposition results are organized in list form, with each list element corresponding to a bilaterally connected component, containing the vertex set and edge set of that component. For example, the vertex set of the first component includes the upstream sampling point, tributary sampling point, main stream sampling point, three solenoid valves, and a filter, and the edge set contains all the edges connecting these vertices, totaling several pipelines.
[0463] The bilateral connectivity component decomposition results reveal the vulnerability distribution of the system topology. Components containing only a single vertex indicate that the vertex is connected to other parts of the system via cut edges, representing potential single points of failure. In this example, the pump, mixer, dilution source, and detector are all single points of failure, requiring the addition of redundant paths or backup components in the system design to improve the system's fault tolerance.
[0464] The decomposition results also provide priority guidance for system maintenance. When an edge within a bilaterally connected component fails, the system can continue to operate because alternative paths exist within the component, and maintenance can be scheduled during non-urgent periods. However, when a cut edge fails, the system's connectivity will be disrupted, requiring immediate repair or switching to an alternative path; this has the highest maintenance priority.
[0465] The algorithm can also calculate the redundancy of each bilaterally connected component, quantifying the component's fault tolerance. Redundancy is defined as the ratio of the number of edges within a component to the minimum number of edges required to maintain connectivity. For a connected graph with n vertices, the minimum number of edges required to maintain connectivity is n minus 1. The higher the redundancy, the stronger the component's fault tolerance. For components with low redundancy, adding extra connecting edges can be considered to improve system reliability.
[0466] Example 8
[0467] This embodiment, based on embodiment 4, provides a detailed explanation of the step "applying the linear space hypergradient algorithm to iteratively solve the optimal transmission strategy for the cost matrix". This step formalizes the optimal transport problem as a linear programming problem, setting constraints as marginal distribution constraints of the multidimensional spectral distribution characteristics before dilution and marginal distribution constraints of the target spectral distribution characteristics after dilution. The objective is to minimize the weighted sum of the cost matrix and the transport plan, resulting in a linear programming model for optimal transport. Based on this model, the dual variables are initialized as zero or random vectors, with initial step size parameters and a convergence threshold set, resulting in initialized dual variables. Based on these initialized dual variables, the hypergradient value is calculated in each iteration based on the cost matrix and the current dual variables, and an adaptive step size is determined. The dual variables are then updated based on the adaptive step size and the hypergradient value, resulting in updated dual variables. The gradient norm corresponding to the updated dual variables is calculated, and it is determined whether the gradient norm is less than a gradient norm threshold. Iteration stops when the gradient norm is less than the threshold, resulting in converged dual variables. Based on the converged dual variables, the original optimal transport strategy matrix is calculated, where each element represents the optimal transport amount from the pre-dilution spectral characteristics to the post-dilution spectral characteristics, thus obtaining the optimal transport strategy.
[0468] In this embodiment, the optimal transport problem is formalized as a linear programming problem for solution. The goal of the linear programming model is to find a transport plan matrix that describes the probabilistic mass allocation scheme for transport from each wavelength point of the pre-dilution spectral distribution to each wavelength point of the post-dilution target spectral distribution. The design of the transport plan matrix must satisfy two fundamental principles: mass conservation and target matching.
[0469] The dimension of the transmission plan matrix is equal to the square of the number of wavelengths. For spectral data containing 601 wavelengths, the transmission plan matrix has a dimension of 601 rows and 601 columns, totaling 361,201 elements. The element in the i-th row and j-th column of the matrix represents the probability mass ratio of transferring from the i-th wavelength of the undiluted spectrum to the j-th wavelength of the diluted spectrum. This element must take a value between 0 and 1, and the sum of all elements must equal 1, satisfying the basic properties of a probability distribution.
[0470] The objective function of the linear programming model is to minimize the total transmission cost. The total transmission cost is calculated by multiplying corresponding elements of the cost matrix and the transmission plan matrix element-wise and then summing the results. Specifically, multiplication is performed on each pair of corresponding elements in the cost matrix and the transmission plan matrix, and then all multiplications are summed to obtain the total transmission cost. The objective function aims to find the transmission plan matrix that minimizes the total transmission cost.
[0471] Mathematically, the objective function is expressed as the minimum of the sum of the products of the elements in the i-th row and j-th column of the cost matrix for all i from 1 to 601 and the elements in the i-th row and j-th column of the transmission plan matrix. This objective function is a linear function of the elements of the transmission plan matrix, satisfying the linearity requirement of linear programming.
[0472] The linear programming model needs to satisfy two types of constraints to ensure the physical rationality and mathematical correctness of the transmission plan. The first type of constraint is called the row sum constraint or mass outflow constraint, which ensures that the total probabilistic mass transferred from each source wavelength point is equal to the probability value of the source distribution at that wavelength point. For each row of the transmission plan matrix, the sum of all elements in that row must be equal to the normalized absorbance value of the spectral distribution before dilution at the corresponding wavelength point.
[0473] The number of rows and constraints equals the number of wavelength points, totaling 601 constraints. The physical meaning of this constraint is that all probability masses from the source distribution must be transferred to the target distribution without any omissions or losses. If the probability mass of a source wavelength point is 0.005, then the sum of the probability masses transferred from that wavelength point to all target wavelength points must equal 0.005.
[0474] The second type of constraint is called the column sum constraint or mass inflow constraint, which ensures that the total probabilistic mass transferred to each target wavelength point is equal to the probability value of the target distribution at that wavelength point. For each column of the transfer plan matrix, the sum of all elements in that column must be equal to the normalized absorbance value of the diluted target spectral distribution at the corresponding wavelength point.
[0475] The number of columns and constraints is equal to the number of wavelength points, totaling 601 constraints. The physical meaning of this constraint is that all probabilistic masses of the target distribution must be transferred from the source distribution; no mass can be created out of thin air or over-injected. If the probabilistic mass of a target wavelength point is 0.003, then the sum of the probabilistic masses transferred from all source wavelength points to that target wavelength point must equal 0.003.
[0476] Furthermore, all elements of the transmission plan matrix must be greater than or equal to zero, which is a requirement for the non-negativity of the probability mass. The probability mass cannot be negative; this constraint ensures the physical rationality of the transmission plan. The number of non-negativity constraints equals the total number of elements in the transmission plan matrix, totaling 361,201 constraints.
[0477] Combining the objective function and constraints, the linear programming model contains 361,201 decision variables and approximately 1,202 equality constraints plus 361,201 inequality constraints. Due to the sheer number of variables and constraints, directly solving this linear programming problem has extremely high computational complexity, and standard simplex methods or interior-point methods are insufficient to complete the solution in a reasonable time. Therefore, duality theory and efficient iterative algorithms are needed to reduce computational complexity.
[0478] In this embodiment, to reduce computational complexity, the original linear programming problem is transformed into a dual problem for solution. Duality theory is the core theory of linear programming, stating that every linear programming problem has a corresponding dual problem, and the optimal solutions of the original and dual problems satisfy strong duality, meaning that their optimal objective function values are equal. The advantage of the dual problem is that the number of variables is significantly reduced, and the solution efficiency is greatly improved.
[0479] The variables in the dual problem are called dual variables, corresponding to the constraints of the primal problem. Since the primal problem has 601 row and column constraints, the dual problem has two sets of dual variables, each containing 601 variables, for a total of 1202 dual variables. The first set of dual variables corresponds to the row and column constraints, with each variable associated with a source wavelength point; the second set of dual variables corresponds to the column and column constraints, with each variable associated with a target wavelength point.
[0480] The objective function of the dual problem is to maximize a linear expression consisting of two inner products. The first part is the inner product of the normalized absorbance values at each wavelength of the undiluted spectral distribution and the first set of dual variables, representing the quality weights of the source distribution. The second part is the inner product of the normalized absorbance values at each wavelength of the diluted target spectral distribution and the second set of dual variables, representing the quality weights of the target distribution. The dual objective function aims to find the dual variables that maximize the sum of these two inner products.
[0481] The constraint of the dual problem requires that the sum of any two dual variables is less than or equal to the corresponding element of the cost matrix. Specifically, for the element in the i-th row and j-th column of the cost matrix, the sum of the i-th element of the first set of dual variables and the j-th element of the second set of dual variables must be less than or equal to the value of that element in the cost matrix. This constraint holds for all elements of the cost matrix, therefore the number of constraints equals the total number of elements in the cost matrix, 361,201. Although the number of constraints is still large, the number of variables in the dual problem is reduced from 361,201 to 1,202, a reduction of 99.7%, making iterative solutions possible.
[0482] The hypergradient algorithm in linear space is an iterative optimization algorithm specifically designed for solving large-scale dual problems. This algorithm iteratively updates the dual variable space along the hypergradient direction, gradually approximating the optimal solution. Hypergradient is a generalization of the gradient concept to non-smooth optimization problems, applicable when the objective function is not differentiable everywhere.
[0483] The algorithm first initializes the dual variables. There are two initialization strategies: zero initialization and random initialization. Zero initialization initializes all elements of the two sets of dual variables to a zero vector, indicating that no dual constraints are initially imposed. The advantage of zero initialization is its computational simplicity and stable convergence behavior. Random initialization initializes the elements of the dual variables to random values within a certain range, such as uniformly distributed random numbers between -0.1 and +0.1. The advantage of random initialization is that it may escape local optima, thus accelerating the convergence speed.
[0484] This embodiment employs a zero-initialization strategy, initializing all elements of the first and second sets of dual variables to zero. After initialization, the dual variables form two 601-dimensional zero vectors.
[0485] The algorithm sets an initial step size parameter to control the update magnitude of the dual variable in each iteration. The choice of step size parameter requires a trade-off between convergence speed and stability. An excessively large step size may lead to iterative oscillations or divergence, while an excessively small step size will result in slow convergence. This embodiment employs an adaptive step size strategy, setting the initial step size to 1.0, and then dynamically adjusting it according to the Armijo criterion.
[0486] The algorithm sets a gradient norm threshold for convergence determination. When the gradient norm is less than this threshold, the algorithm is considered to have converged to the vicinity of the optimal solution, and the iteration can be terminated. The choice of the gradient norm threshold affects the solution accuracy and computation time. The smaller the threshold, the higher the solution accuracy, but the more iterations are required and the longer the computation time. In this embodiment, the gradient norm threshold is set to one part per million, that is, iteration stops when the gradient norm is less than 0.000001.
[0487] After initialization, the initial dual variables are output, including two sets of zero vectors, providing a starting point for subsequent iterative updates.
[0488] In this embodiment, the core of the linear space hypergradient algorithm is to calculate the hypergradient and determine the adaptive step size in each iteration. The hypergradient reflects the direction and magnitude of the deviation between the current dual variable and the optimal solution, and the adaptive step size controls the distance moved along the hypergradient direction.
[0489] The hypergradient is calculated in each iteration. The dual problem of the optimal transport problem is the Lagrangian function. Defined as:
[0490]
[0491] in, and These are the dual variables corresponding to row and column constraints, respectively. Based on this function, calculate the variables related to... and The hypergradient.
[0492] At the beginning of each iteration, the algorithm first calculates the transport plan matrix based on the current dual variable. The element in the i-th row and j-th column of the transport plan matrix is determined according to the complementary relaxation condition. The complementary relaxation condition is a core property of duality theory, stating that at the optimal solution, if a primal variable is positive, the corresponding dual constraint must be equal; if the dual constraint is strictly inequal, the corresponding primal variable must be zero.
[0493] The specific calculation rules are as follows: Check if the sum of the i-th element of the first set of dual variables and the j-th element of the second set of dual variables is equal to the element in the i-th row and j-th column of the cost matrix. If they are equal, it means the dual constraint holds true, and in this case, the element of the transmission plan matrix is equal to the normalized value of the i-th wavelength point of the pre-dilution spectral distribution multiplied by the normalized value of the j-th wavelength point of the post-dilution target spectral distribution. This product represents the probabilistic quality allocation on the transmission path. If the dual constraint does not hold true, it means the cost of the transmission path is higher than the sum of the dual variables, and the path should not be used for transmission; therefore, the element of the transmission plan matrix is equal to zero.
[0494] By traversing all elements of the cost matrix and calculating the corresponding elements of the transmission plan matrix one by one, the complete transmission plan matrix is obtained. This matrix describes the optimal transmission scheme under the current dual variables.
[0495] Next, the algorithm calculates the hypergradient. For the i-th element of the first set of dual variables, its hypergradient is equal to the normalized value of the i-th wavelength point of the spectral distribution before dilution, minus the sum of all elements in the i-th row of the transfer plan matrix. The sum of the elements in the i-th row represents the total probabilistic mass transferred from the i-th source wavelength point. The sign of the hypergradient reflects the deviation between the mass outflow and the source distribution. If the hypergradient is positive, it indicates that the mass transferred from that source wavelength point is insufficient, and the dual variable needs to be increased to encourage more mass to be transferred from that point; if the hypergradient is negative, it indicates that too much mass is transferred, and the dual variable needs to be reduced.
[0496] For the j-th element of the second set of dual variables, its hypergradient is equal to the normalized value of the j-th wavelength point of the diluted target spectral distribution, minus the sum of all elements in the j-th column of the transfer plan matrix. The sum of the elements in the j-th column represents the total probabilistic mass transferred to the j-th target wavelength point. The sign of the hypergradient reflects the deviation between the mass inflow and the target distribution. If the hypergradient is positive, it indicates that the mass transferred to that target wavelength point is insufficient, and the dual variable needs to be increased to attract more mass transfer; if the hypergradient is negative, it indicates that the transferred mass is excessive, and the dual variable needs to be reduced.
[0497] After the hypergradient is calculated, the algorithm obtains two 601-dimensional hypergradient vectors, which correspond to two sets of dual variables.
[0498] The algorithm then determines the adaptive step size. The adaptive step size is calculated using the Armijo backtracking search method, which progressively reduces the step size until a sufficient descent condition is found. The Armijo criterion is that after moving the step size along the hypergradient direction, the increase in the dual objective function must be greater than or equal to a certain proportion of the product of the step size and the square of the hypergradient norm. This proportion is called the Armijo parameter and is typically set to 0.1.
[0499] The specific search process is as follows: Initialize the trial step size to 1.0. Calculate the new dual variable after moving the trial step size along the hypergradient direction, i.e., the current dual variable plus the trial step size multiplied by the hypergradient. Calculate the new dual objective function value based on the new dual variable. Calculate the actual increase in the dual objective function, which is equal to the new objective function value minus the current objective function value. Calculate the minimum increase required by the Armijo criterion, which is equal to the Armijo parameters multiplied by the trial step size multiplied by the square of the hypergradient norm.
[0500] If the actual increase is greater than or equal to the minimum increase, the attempted step size satisfies the Armijo criterion, and this step size is accepted as the adaptive step size. If the actual increase is less than the minimum increase, the attempted step size is too large. The attempted step size is halved, the new dual variable and the actual increase are recalculated, and the Armijo criterion is checked again. This process is repeated until a step size that satisfies the Armijo criterion is found, or the attempted step size is less than the minimum step size threshold, for example, 0.0001. If the criterion is still not satisfied even after reducing the step size to the minimum threshold, the minimum step size is accepted as the adaptive step size.
[0501] The Armijo backtracking search method can adaptively adjust the step size according to the local curvature of the objective function while ensuring convergence. This avoids oscillations caused by excessively large step sizes and slow convergence caused by excessively small step sizes.
[0502] In a water sample dilution optimization example, at the 50th iteration, the hypergradient norm was 0.05. An attempt was made with a step size of 1.0 to satisfy the Armijo criterion. The increase in the dual objective function was 0.003, which is greater than the minimum increase of 0.1 multiplied by 1.0 multiplied by 0.05 squared, equal to 0.00025. Therefore, a step size of 1.0 was accepted. At the 100th iteration, the hypergradient norm decreased to 0.01. An attempt was made with a step size of 1.0, but this did not satisfy the Armijo criterion. The step size was halved to 0.5, which satisfied the criterion. A step size of 0.5 was then accepted. As the iterations progressed, the hypergradient norm gradually decreased, and the adaptive step size decreased accordingly, indicating that the iterations had entered a fine-tuning phase.
[0503] In this embodiment, the algorithm updates the dual variable based on the determined adaptive step size and the calculated hypergradient. The update rule is to add the adaptive step size multiplied by the hypergradient to the current dual variable to obtain the updated dual variable.
[0504] For the i-th element of the first set of dual variables, the updated value is equal to the current value plus the adaptive step size multiplied by the hypergradient value of that element. The same update rule applies to the j-th element of the second set of dual variables. The update operation is performed on all elements of both sets of dual variables one by one to obtain the complete updated dual variables.
[0505] The updated dual variables reflect the shift along the hypergradient direction. If the hypergradient of a dual variable is positive, the updated variable increases, encouraging the corresponding probability mass flow; if the hypergradient is negative, the updated variable decreases, inhibiting the corresponding probability mass flow. Through this iterative update mechanism, the dual variables are gradually adjusted to near their optimal values.
[0506] After the update, the algorithm calculates the gradient norm of the updated dual variable. The gradient norm is equal to the square root of the sum of the squares of all hypergradient values, measuring the overall magnitude of the hypergradient vector. The smaller the gradient norm, the closer the current dual variable is to the optimal solution, and the closer the hypergradient vector is to zero.
[0507] The gradient norm is calculated as follows: sum the squares of all hypergradient values for the first set of dual variables, add the sum of the squares of all hypergradient values for the second set of dual variables, and then take the square root of the sum. For example, if the first set of dual variables has 601 hypergradients and the second set of dual variables has 601 hypergradients, then the sum of squares contains 1202 squared terms.
[0508] The algorithm determines whether the gradient norm is less than a preset gradient norm threshold. If the gradient norm is less than the threshold, it means the hypergradient vector is small enough, the dual variable is close to the optimal solution, the iteration converges, and the algorithm stops iterating. If the gradient norm is greater than or equal to the threshold, it means the dual variable is still some distance from the optimal solution, and the next round of iteration is required.
[0509] In a water sample dilution optimization example, the algorithm starts iterating from the initial dual variables. After the first iteration, the gradient norm is 5.0, which is much greater than the threshold of 0.000001, so iteration continues. After the 50th iteration, the gradient norm decreases to 0.5, but is still greater than the threshold, so iteration continues. After the 100th iteration, the gradient norm decreases to 0.05, so iteration continues. After the 150th iteration, the gradient norm decreases to 0.0000005, which is less than the threshold of 0.000001, satisfying the convergence condition, and the algorithm stops iterating.
[0510] After the iteration stops, the converged dual variable is output. The converged dual variable is the optimal or near-optimal solution to the dual problem found by the algorithm, corresponding to the optimal propagation strategy of the primal problem.
[0511] In this embodiment, the algorithm calculates the original optimal transmission strategy matrix based on the convergent dual variables. According to the strong duality and complementary relaxation conditions of duality theory, the optimal solution to the primal problem can be uniquely determined from the optimal solution to the dual problem.
[0512] The element in the i-th row and j-th column of the original optimal transmission strategy matrix is calculated according to the following rule: Check whether the sum of the i-th element of the first set of dual variables and the j-th element of the second set of dual variables in the convergent dual variables is equal to the element in the i-th row and j-th column of the cost matrix. If they are equal, it means that the transmission path is used in the optimal solution, and this element of the transmission strategy matrix is equal to the normalized value of the i-th wavelength point of the spectral distribution before dilution multiplied by the normalized value of the j-th wavelength point of the target spectral distribution after dilution. If they are not equal, it means that the transmission path is not used in the optimal solution, and this element of the transmission strategy matrix is equal to zero.
[0513] By traversing all elements of the cost matrix and calculating the corresponding elements of the transmission strategy matrix one by one, the complete original optimal transmission strategy matrix is obtained. Each element of this matrix represents the optimal probabilistic quality for the transfer from a certain wavelength point in the pre-dilution spectrum to a certain wavelength point in the post-dilution spectrum.
[0514] The original optimal transmission strategy matrix satisfies all the constraints of the original problem: the sum of the elements in each row is equal to the normalized value of the spectral distribution before dilution at the corresponding wavelength point, satisfying the mass outflow constraint; the sum of the elements in each column is equal to the normalized value of the target spectral distribution after dilution at the corresponding wavelength point, satisfying the mass inflow constraint; all elements are non-negative, satisfying the non-negativity constraint.
[0515] Meanwhile, the original optimal transmission strategy matrix minimizes the objective function, i.e., the total transmission cost. Multiplying the corresponding elements of this matrix by the cost matrix and summing them yields the total transmission cost that is the smallest among all feasible transmission schemes.
[0516] In a water sample dilution optimization example, the original optimal transport strategy matrix shows that most probabilistic masses are transferred from the absorbance peak wavelength of the pre-dilution spectrum near 300 nm to the target wavelength of the post-dilution spectrum near 300 nm, because the cost at that wavelength is relatively low. A small number of probabilistic masses are transferred from other wavelengths to adjacent wavelengths, because the cost of small-range wavelength shifts is relatively low. Almost no probabilistic masses are transferred across a large wavelength range, because the cost of large-range transfers is very high.
[0517] The original optimal transmission strategy matrix is output, which describes the optimal conversion scheme from the spectral distribution before dilution to the target spectral distribution after dilution, providing a key data foundation for subsequent dilution ratio calculations.
[0518] Example 9
[0519] like Figure 4 As shown, the present invention also provides an online dynamic water quality fingerprint rapid acquisition and preprocessing device for performing the method of any one of the embodiments 1-8 above, including:
[0520] The multi-flow-path sampling unit is equipped with a raw water direct monitoring flow path 11, an automatic dilution monitoring flow path 12, and a standard sample calibration flow path 13. Each flow path is connected to the fluid main pipe 14 through a solenoid valve.
[0521] An adaptive dilution unit, connected to the raw water direct monitoring flow path, has a dilution ratio control function based on turbidity or suspended solids concentration.
[0522] The intelligent cleaning module includes a miniature air compressor 21, an acid tank 22, a bactericide tank 23, and a gas-liquid mixing spray cleaning structure, which is used to perform high-pressure cleaning of the optical window;
[0523] The spectral detection module 3 includes an ultraviolet-visible spectrometer and an excitation-emission matrix fluorescence spectrometer connected to the fluid main pipe 14; wherein the spectral detection module adopts a flow cell structure, which allows water samples to be measured online without sampling.
[0524] Controller 4 is used to control the flow path switching of the multi-flow path sample supply unit, the dilution ratio of the adaptive dilution unit, the cleaning process of the intelligent cleaning module, the spectral acquisition timing and detection trigger of the spectral detection module, and to realize data communication.
[0525] The multi-path sample supply unit is the core of the device's sample supply, integrating three functional flow paths: a raw water direct monitoring flow path 11, an automatic dilution monitoring flow path 12, and a standard sample calibration flow path 13. Each flow path is precisely connected to the fluid main pipe 14 via a solenoid valve, allowing for switching between different flow paths according to monitoring needs and ensuring a stable supply of various samples as required. The adaptive dilution unit is connected to the raw water direct monitoring flow path 11 and has an intelligent function to determine the dilution ratio based on turbidity or suspended solids concentration. The intelligent cleaning module is specifically designed for cleaning the optical window and includes a miniature air compressor 21, an acid tank 22, a bactericide tank 23, and a gas-liquid mixing spray cleaning structure. It achieves efficient cleaning of the optical window through high-pressure rinsing, ensuring the stable operation of the detection components. The gas-liquid mixing spray cleaning structure is the core execution component of the cleaning function, with a built-in nozzle array designed at a specific angle. It mixes the delivered gas and liquid in a preset ratio and then sprays it at high pressure to perform all-round, thorough rinsing of the optical window, ensuring effective cleaning. The spectral detection module 3 is the core detection component for water fingerprint acquisition. It consists of an ultraviolet-visible spectrometer and an excitation-emission matrix fluorescence spectrometer, and adopts a flow cell structure design, enabling direct online measurement of water samples without the need for sampling, significantly improving detection efficiency and real-time performance. The controller 4, as the core control unit of the device, undertakes the coordinated control tasks of the entire system. Specifically, it controls the flow path switching of the multi-flow-path sample supply unit, adjusts the dilution ratio of the adaptive dilution unit, manages the cleaning process of the intelligent cleaning module, coordinates the spectral acquisition timing and detection triggering of the spectral detection module 3, and simultaneously enables data communication between the system and external systems, ensuring fully automated and intelligent operation of the entire device.
[0526] The main body of the device is installed inside the online monitoring system, and it also includes three key components: an internal circulation pump 51, a pressure relief valve 52, and a filter 53. The internal circulation pump 51 is used to maintain the internal fluid circulation of the system and can flush the dead areas of the pipeline to prevent the deposition of particulate matter in the water sample. The filter 53 is installed in the main pipeline of the system and removes large particulate impurities in the water sample through a structure such as a stainless steel filter screen to avoid blockage of subsequent pipelines, valves and detection components. The pressure relief valve 52 is used to regulate the pressure in the pipeline. When the pressure in the pipeline exceeds a preset threshold, it automatically releases pressure to prevent damage to the pipeline, pump body and other core components due to excessive pressure, and ensures the safety and stability of the fluid transmission of the system.
[0527] Based on the above structural design, the device has four core operating modes. The applicable scenarios, flow paths, and control logic of each mode are as follows: The first is the raw water direct monitoring mode (conventional mode). This mode is suitable for conventional monitoring scenarios where the water quality is relatively clear and the turbidity value is within the range of the device. Its flow path is: raw water in → second solenoid valve 82 (open) → turbidity sensor 9 → sixth solenoid valve 86 (open, direct flow) → confluence node → seventh solenoid valve 87 → flow pool of spectral detection module 3 → eighth solenoid valve 88 → discharge. In this mode, the controller controls the first, third, fourth, and fifth solenoid valves to be closed. The raw water flows directly to the confluence node through the bypass pipeline where the sixth solenoid valve 86 is located, thereby bypassing the dilution and mixing structure and quickly reaching the spectral detection module 3, minimizing the sample transmission lag time and ensuring the real-time performance of conventional water quality monitoring.
[0528] The second mode is the adaptive dilution monitoring mode (high turbidity mode). This mode is an automatic switching mode. When the turbidity sensor detects that the turbidity value of the raw water exceeds the set threshold, it automatically switches to this mode. The flow direction is divided into three stages: raw water side, dilution water side, and after mixing. The raw water side path is raw water inlet → second solenoid valve 82 (open) → turbidity sensor 9. At this time, the sixth solenoid valve 86 is closed, forcing water flow into the mixer. The dilution water side path is dilution water inlet → third solenoid valve 83 (open) → fifth solenoid valve 85 (open) → mixer. After the raw water and dilution water are uniformly mixed in the mixer, the mixed liquid is monitored and discharged along the path of confluence node → seventh solenoid valve 87 → flow pool of spectral detection module 3 → eighth solenoid valve 88 → discharge. In this mode, the fifth solenoid valve 85 plays a key role, which can accurately control the timing of dilution water injection into the mixing chamber and ensure the stability and accuracy of the dilution process.
[0529] The third is the standard sample calibration mode (quality control mode), which is mainly used for periodic zero-point verification or span calibration of the device to ensure the accuracy and reliability of monitoring data. Its flow path is: standard sample in → fourth solenoid valve 84 (open) → confluence node → seventh solenoid valve 87 (closed) → flow cell of spectral detection module 3 → eighth solenoid valve 88 → discharge. The core design advantage of this mode is that the standard sample flow path is independent of the raw water pipeline and is directly connected to the front end of the spectral detection module. Structurally, it avoids cross-contamination of samples caused by sharing with the raw water pipeline, thus ensuring the accuracy of calibration results.
[0530] Fourthly, the cleaning and maintenance mode integrates a micro air compressor 21 (compressed air supply), a gas-liquid spray cleaning path, a bactericide tank 23, an acid pickling agent tank 22 (for dissolving inorganic deposits), and dedicated first, second, and third feed pumps 71, 72, and 73 through an intelligent cleaning module. The cleaning execution logic is as follows: the controller triggers the cleaning cycle, or a sudden change in turbidity / absorbance triggers it → the solenoid valve closes the sample passage → the first gas valve 61 and the second gas valve 62 are opened for gas injection → the gas-liquid two-phase flow washes the optical window → acid / bactericide is added as needed for chemical cleaning → the discharge path is opened after cleaning. This module can effectively prevent algae adhesion, particle deposition, organic film contamination, and optical window mirror contamination. During cleaning, the second solenoid valve 82, the third solenoid valve 83, the fourth solenoid valve 84, and the eighth solenoid valve 88 are closed, while the other solenoid valves remain open.
[0531] It should be noted that, in order to achieve the high fault tolerance of the system described in Embodiment 2, the device described in this embodiment has a pre-configured redundant flow path in its physical structure. Although Figure 4 To maintain clarity and simplicity, only a single main passage is shown. However, in the actual device, backup pipelines and matching switching solenoid valves (not shown in the figure, but usually located in the same solenoid valve housing as the ordinary solenoid valve, with an internal switching switch that can immediately switch when the ordinary solenoid valve fails) are connected in parallel between the mixer and the spectral detection module, and between the micro air compressor and each flow path connection point.
[0532] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for rapid online dynamic water quality fingerprint acquisition and preprocessing, characterized in that, Includes the following steps: Intelligent selection of the test flow path from multiple inlet flow paths includes: acquiring the physical structure of the entire water quality monitoring system; modeling the entire water quality monitoring system as a first directed weighted graph, where vertices represent key components including sampling points, filters, pumps, mixers, and detectors, edges represent pipelines connecting the key components, and the weight of the edges represents flow time, energy consumption, or failure risk; and outputting a system topology graph model corresponding to the entire water quality monitoring system; based on the system topology graph model, applying a depth-first search algorithm to traverse the system topology graph model from the starting vertex to identify all cut edges, and decomposing the system topology graph model into multiple bilaterally connected components by deleting all identified cut edges, obtaining the bilaterally connected component decomposition result; and identifying the bilaterally connected components... The connection relationships between the bilaterally connected components in the component decomposition results are used to determine the potential single-point fault locations, and redundant flow paths are configured at these locations. After completing the design of the redundant flow paths, the water sampling problem is formalized into a modified Traveling Salesman Problem (TSP), and a distance matrix is constructed. The elements of the distance matrix represent the comprehensive cost from one component to another. The cutting plane method is applied to solve the linear programming relaxation problem in the distance matrix to obtain an initial lower bound solution. The branch and bound method is used to gradually reduce the integer gap of the initial lower bound solution until the integer gap is less than a preset threshold, thus obtaining the optimal sampling path. Based on the path switching logic in the optimal sampling path, the opening and closing of the solenoid valves in the multi-flow path structure with redundant design are controlled to achieve flow path selection and determine the flow path to be tested. The degree of contamination of the optical window is determined based on the light intensity attenuation trend prediction model. When the predicted degree of contamination exceeds the dynamic threshold, the optical window is cleaned. The light intensity attenuation trend prediction model is a temporal convolutional network model. In the measurement channel corresponding to the cleaned optical window, the water sample of the flow path to be tested is detected, and the turbidity detection result, colorimetric detection result, and historical dilution effect evaluation data are input into the dilution decision model, wherein the dilution decision model adopts a deep neural network architecture; The water sample from the flow path to be tested is diluted according to the dilution parameters output by the dilution decision model. The diluted water sample is subjected to spectral acquisition to obtain water quality fingerprint spectral data, and water quality parameters are calculated based on the water quality fingerprint spectral data to complete the collection and preprocessing of water quality fingerprint.
2. The method according to claim 1, characterized in that, The process of controlling the opening and closing of the solenoid valve in the multi-flow-path structure with redundant design to select the flow path based on the path switching logic in the optimal sampling path, and determining the flow path to be measured, includes: Monitor the operating status parameters and performance parameters of each component in the multi-flow path structure with redundant design; Determine whether the operating status parameters deviate from the normal operating parameter threshold range or whether the performance parameters are lower than the preset performance threshold, so as to detect abnormal situations such as pipeline blockage or pump performance degradation, and generate a system abnormal status signal; Based on the system abnormal state signal, the abnormal component indicated by the system abnormal state signal is marked as unavailable from the system topology model by the dynamic path replanning algorithm, and the optimal path from the current state to the target state is recalculated. Alternative flow paths are automatically started or the optimal sampling path is recalculated to obtain an updated sampling path scheme. Based on the updated sampling path scheme, first close the solenoid valve of the current flow path corresponding to the optimal sampling path scheme, then open the solenoid valve of the path corresponding to the updated sampling path scheme and control the flow rate to be gradually adjusted during the switching process to determine the dynamically optimized flow path to be tested.
3. The method according to claim 1, characterized in that, The process of diluting the water sample of the flow path to be tested according to the dilution parameters output by the dilution decision model includes: The raw water spectral data of the water sample in the flow path to be tested are obtained, and the raw water spectral data are represented as the multidimensional spectral distribution characteristics before dilution. The spectral characteristics under ideal measurement conditions are represented as the target spectral distribution characteristics after dilution. Define a cost matrix for transforming the multidimensional spectral distribution features before dilution to the target spectral distribution features after dilution, wherein each element of the cost matrix represents the transmission cost required to transform each spectral feature before dilution into the corresponding spectral feature after dilution. The optimal transmission strategy of the cost matrix is solved iteratively by applying the linear space hypergradient algorithm. The dual variables are initialized and the hypergradient value is calculated in each iteration. The dual variables are updated according to the step size rule until the gradient norm is less than the preset gradient norm threshold, and the optimal transmission strategy is obtained. A mapping function is established from the spectral distribution before and after dilution to the physical dilution ratio. The optimal transmission strategy is substituted into the mapping function to calculate the actual dilution ratio, and the optimal dilution ratio is obtained. The raw water flow rate and dilution water flow rate are calculated based on the optimal dilution ratio and the target total flow rate to obtain the dilution control parameters; According to the dilution control parameters, the flow rate of the raw water and the flow rate of the dilution water are controlled by a proportional valve or a peristaltic pump to achieve the dilution treatment of the water sample in the test flow path, so as to obtain the diluted water sample.
4. The method according to claim 1, characterized in that, The calculation of water quality parameters based on the water quality fingerprint spectral data includes: Principal component analysis is applied to calculate the covariance matrix and decompose the eigenvalues of the original spectral data matrix corresponding to the water quality fingerprint spectral data, identify the spectral feature vectors, and reconstruct the original spectral data matrix in a structured manner based on the spectral feature vectors to obtain the structured spectral data matrix. Based on the structured spectral data matrix, the structured spectral data matrix is abstracted into a second directed weighted graph, wherein the vertices of the second directed weighted graph represent specific elements or blocks in the structured spectral data matrix, and the edges of the second directed weighted graph represent the computational dependencies between the specific elements or blocks. A set of matrix flipping operations is defined, and a flipping operation cost model is constructed to evaluate the computational cost-benefit ratio of each flip in the set of matrix flipping operations, thus obtaining a flipped graph model of the spectral data. Based on the flip graph model of the spectral data, a heuristic search function is designed to evaluate the estimated cost from the current state of the flip graph model of the spectral data to the target state. A priority queue is used to manage the flip operation sequence to be explored, and a pruning strategy is developed to eliminate search branches with non-optimal solutions, so as to obtain an optimized flip operation sequence. Based on the optimized flipping operation sequence, by identifying the key matrix operation patterns in water quality fingerprint analysis, a corresponding flipping operation sequence template is designed for each operation pattern in the key matrix operation pattern, so as to accelerate the calculation of the key matrix in the water quality fingerprint analysis. The accelerated matrix calculation results are used to optimize memory access, reduce cache miss rate, and optimize single instruction multiple data stream instruction set. The embedded processor is used to process multiple data elements in parallel to obtain real-time spectral analysis data. Based on the spectral characteristics of the real-time processed spectral analysis data, the pollutant concentration, total organic carbon, and chemical oxygen demand are calculated using a preset water quality parameter calculation model to obtain the water quality parameters.
5. The method according to claim 4, characterized in that, After obtaining the water quality parameters by calculating pollutant concentration, total organic carbon, and chemical oxygen demand using a preset water quality parameter calculation model based on the spectral characteristics of the real-time processed spectral analysis data, the method further includes: The real-time processed spectral analysis data, the operating status parameters of the water quality monitoring system, the control command execution feedback and the abnormal event alarm information are acquired, and the real-time processed spectral analysis data, the operating status parameters, the control command execution feedback and the abnormal event alarm information are organized by timestamp and stored in the operation record database to obtain the operation record; Data within a preset time window is extracted from the running records and preprocessed by outlier removal, missing value imputation, and data smoothing to obtain preprocessed running data. Based on the preprocessed operational data, the current values of the water quality parameters are displayed in a digital panel by plotting the real-time processed spectral analysis data into spectral curves.
6. The method according to claim 1, characterized in that, Based on the system topology graph model, a depth-first search algorithm is applied to traverse the system topology graph model from the starting vertex to identify all cut edges. Then, by deleting all identified cut edges, the system topology graph model is decomposed into multiple bilaterally connected components, yielding the bilaterally connected component decomposition results, including: Based on the system topology graph model, the access timestamp and lowest reach timestamp of each vertex are recorded by initializing the depth-first search traversal process. Starting from the starting vertex, all vertices and edges of the system topology graph model are recursively traversed to obtain the traversal result. Based on the traversal results, by judging the relationship between the lowest reachable timestamps of the two endpoints of an edge, when the lowest reachable timestamp of one endpoint of an edge is greater than the access timestamp of the other endpoint, the edge is marked as a cut edge, and the set of all cut edges is obtained. Based on the set of all cut edges, the system topology graph model is decomposed into multiple connected subgraphs by deleting all cut edges from the set of all cut edges in the system topology graph model. In each connected subgraph, there are at least two non-repeating paths between any two vertices, thus obtaining the bilateral connected component decomposition result.
7. The method according to claim 3, characterized in that, The method employs a linear space hypergradient algorithm to iteratively solve for the optimal transport strategy of the cost matrix. It initializes dual variables and calculates hypergradient values in each iteration. The dual variables are updated according to a step size rule until the gradient norm is less than a preset gradient norm threshold, thus obtaining the optimal transport strategy. This includes: Based on the cost matrix, by formalizing the optimal transmission problem into a linear programming problem, and setting the constraints as the marginal distribution constraints of the multidimensional spectral distribution characteristics before dilution and the marginal distribution constraints of the target spectral distribution characteristics after dilution, the linear programming model of optimal transmission is obtained with the goal of minimizing the weighted sum of the cost matrix and the transmission plan. Based on the linear programming model of optimal transmission, the initial dual variables are obtained by initializing the dual variables as zero vectors or random vectors, setting the initial step size parameter and the convergence threshold. Based on the initialized dual variables, the updated dual variables are obtained by calculating the hypergradient value and determining the adaptive step size in each iteration according to the cost matrix and the current dual variables, and updating the dual variables according to the adaptive step size and the hypergradient value. Calculate the gradient norm corresponding to the updated dual variable, determine whether the gradient norm is less than the gradient norm threshold, stop the iteration when the gradient norm is less than the gradient norm threshold, and obtain the converged dual variable; Based on the converged dual variables, the original optimal transmission strategy matrix is calculated from the converged dual variables, wherein each element of the original optimal transmission strategy matrix represents the optimal transmission amount from the spectral features before dilution to the spectral features after dilution, thus obtaining the optimal transmission strategy.
8. An online dynamic water quality fingerprint rapid acquisition and pretreatment device, characterized in that, The apparatus is used to perform the method according to any one of claims 1-7, comprising: The multi-flow-path sampling unit is equipped with a raw water direct monitoring flow path, an automatic dilution monitoring flow path, and a standard sample calibration flow path. Each flow path is connected to the main fluid pipe through a solenoid valve. An adaptive dilution unit, connected to the raw water direct monitoring flow path, has a dilution ratio control function based on turbidity or suspended solids concentration. The intelligent cleaning module includes a miniature air compressor, an acid tank, a bactericide tank, and a gas-liquid mixing spray cleaning structure, which is used to perform high-pressure rinsing of the optical window; The spectral detection module includes an ultraviolet-visible spectrometer and an excitation-emission matrix fluorescence spectrometer connected to the fluid main pipe; wherein the spectral detection module adopts a flow cell structure, enabling online measurement of water samples without the need for sampling. The controller is used to control the flow path switching of the multi-flow-path sample supply unit, the dilution ratio of the adaptive dilution unit, the cleaning process of the intelligent cleaning module, the spectral acquisition timing and detection trigger of the spectral detection module, and to realize data communication.
Citation Information
Patent Citations
Method and device for imaging from spectrum to mass concentration based on physical mechanism deep learning
CN121353613A
Sensor-synchronized spectrally-structured-light imaging
WO2015077493A1