Full-automatic water quality COD (Chemical Oxygen Demand) analysis method and device
Through multi-source signal fusion and manifold learning of absorbance spectra and electrochemical characteristics, combined with hybrid neural network modeling, the efficiency and accuracy problems of existing water quality COD analysis methods in complex water sample scenarios are solved, and high-precision water quality COD analysis is achieved.
Patent Information
- Application Number
- CN202511160316.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing water quality COD analysis methods lack efficiency and accuracy in complex water sample scenarios, and are difficult to meet high-throughput and real-time requirements, mainly due to insufficient ability to process water sample characteristic information.
By obtaining the absorbance spectra and electrochemical characteristics of the target water samples in different bands, feature-level fusion and manifold learning are performed, combined with hybrid neural network modeling, correction and verification processing are performed to generate a fully automatic water quality COD analysis method.
It achieves efficient integration and optimization of complex water samples, improves the accuracy and stability of analysis results, can capture complex nonlinear relationships with high precision, and meets the high-throughput and real-time requirements of modern water quality monitoring.
Smart Images

Figure CN120652073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water quality analysis, and in particular to a fully automatic water quality COD analysis method and device. Background Art
[0002] Chemical oxygen demand (COD) analysis of water quality is used to assess the content of organic pollutants and reducing substances in water bodies. Existing COD analysis methods mainly use chemical titration or laboratory instrument analysis, introducing UV-visible spectroscopy absorbance analysis or electrochemical sensors into COD detection to achieve rapid measurement. These technologies analyze the absorbance or electrochemical signals of water samples in specific bands, and have initially achieved non-chemical COD estimation, reducing the use of some experimental consumables. However, these methods still have certain limitations in practical applications, especially the insufficient analysis efficiency and accuracy in complex water sample scenarios, which makes it difficult to meet the high-throughput and automation requirements of modern water quality monitoring.
[0003] This core problem of existing technologies stems from their inability to process characteristic information of water samples. Existing methods typically rely on a single or limited signal source (such as absorbance or chemical reaction endpoints), failing to fully capture the multidimensional characteristics of the complex organic composition of water samples. For example, in spectral analysis, using only absorbance data from a single band makes it difficult to reflect the synergistic effects of multiple pollutants in a water sample, resulting in unstable signals. This significantly reduces the accuracy of analytical results when faced with water samples with complex compositions and strong background interference (such as industrial wastewater or natural water bodies). This is especially difficult to meet the real-time and high-precision requirements in online monitoring scenarios that require rapid response.
[0004] Therefore, there is an urgent need for a COD analysis method that can efficiently integrate and optimize multi-source signal data and improve analysis performance to overcome the shortcomings of existing technologies in data processing capabilities. Summary of the Invention
[0005] The main purpose of the present invention is to provide a fully automatic water quality COD analysis method and device, aiming to overcome the technical problems that the existing technology cannot fully capture the characteristic information of water samples, resulting in low analysis efficiency and insufficient accuracy.
[0006] In order to solve the above-mentioned problems, the present invention proposes a fully automatic water quality COD analysis method, which comprises: Obtaining the absorbance spectrum and electrochemical characteristics of the target water sample in the first preset band to generate an original signal set; Acquire dynamic features of a fusion liquid of the target water sample and the potassium dichromate solution in a second preset band, and perform feature-level fusion of the dynamic features with the original signal set to generate a reaction feature set; Performing manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set; Perform hybrid neural network modeling processing based on the optimized feature set and output a preliminary COD prediction value; After the preliminary COD prediction value is corrected, the correction result is verified to generate a COD analysis set.
[0007] Furthermore, the step of obtaining the absorbance spectrum and electrochemical characteristics of the target water sample in the first preset band to generate an original signal set includes: Collecting the absorbance spectrum of the target water sample in the 200-235 nm band, performing denoising on the absorbance spectrum to generate a spectral signal matrix; Based on a multi-channel electrochemical sensor array, the redox potential and current response of the target water sample are synchronously collected to obtain an electrochemical signal matrix; Extracting low-frequency components of the spectral signal matrix and reconstructing the signal, performing principal component analysis on the reconstructed signal, and generating a spectral feature matrix; Perform frequency domain feature extraction on the electrochemical signal matrix to obtain the frequency domain feature vector: The spectral feature matrix and the frequency domain feature vector are subjected to Z-score normalization processing, and the spectral feature matrix and the frequency domain feature vector are weightedly spliced and fused according to information entropy to obtain an original signal set.
[0008] Furthermore, the step of obtaining the dynamic characteristics of the fusion liquid of the target water sample and the potassium dichromate solution in the second preset band, and performing feature-level fusion of the dynamic characteristics with the original signal set to generate a reaction feature set includes: The absorbance changes of the fusion liquid of the target water sample and potassium dichromate solution in the 300-450nm band are collected to generate a dynamic spectral dataset containing time and wavelength dimensions; Processing the absorbance sequence of each wavelength in the dynamic spectral data set based on a time series difference algorithm, calculating the change rate of absorbance at adjacent time points, and generating an absorbance change rate matrix; Performing reaction rate feature extraction processing on the absorbance change rate matrix to obtain a reaction rate feature set; collecting an electrochemical signal sequence of the fusion liquid, and performing frequency domain analysis on the electrochemical signal sequence to generate a fusion frequency domain feature set; The reaction rate feature set, the fused frequency domain feature set and the spectral feature vector in the original signal set are fused at the feature level to obtain the reaction feature set.
[0009] Furthermore, the step of performing manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set includes: Constructing a local reconstruction weight matrix according to the k nearest neighbors of each fused feature vector in the reaction feature set, solving the local reconstruction weight matrix globally and mapping it to a three-dimensional manifold space to obtain a manifold feature set; According to a preset neighborhood radius and a preset minimum number of samples, the core points and boundary points of each point in the manifold feature set are calculated, and cluster clusters are generated according to a density reachability relationship to obtain a cluster feature label set; Performing a four-level decomposition on the time series signal of the dynamic reaction metadata in the reaction feature set, extracting low-frequency coefficients and high-frequency coefficients, calculating energy distribution characteristics based on the decomposition coefficients, and generating a time series feature vector set; Calculating the mutual information value between the time series feature vector and the cluster feature label set, and screening to obtain a highly correlated feature subset according to a preset mutual information threshold; The manifold feature set, cluster feature label set, time series feature vector set and high correlation feature subset are weightedly fused by entropy weight method to obtain a comprehensive feature vector set; Regularization and feature optimization processing are performed on the comprehensive feature vector set to obtain an optimized feature set.
[0010] Furthermore, the step of performing hybrid neural network modeling processing according to the optimized feature set and outputting a preliminary COD prediction value includes: Constructing a convolutional neural network model, performing a convolution operation on the optimized feature set based on the convolutional neural network model, and performing dimensionality reduction processing on the convolution output through an average pooling layer to generate a convolutional feature atlas containing multi-scale features; Performing temporal dependency modeling on the convolutional feature atlas to obtain a temporal feature sequence set; Performing attention weighting processing according to the feature importance metadata in the temporal feature sequence set and the optimized feature set to obtain a weighted feature vector set; A two-layer fully connected neural network was constructed, with 256 neurons in the first layer and 1 neuron in the second layer. The weighted feature vector set was input into the fully connected neural network for feature mapping, and the preliminary COD prediction value was obtained as output.
[0011] Furthermore, the step of correcting the preliminary COD prediction value includes: The preliminary COD prediction value was optimized by genetic algorithm hyperparameter optimization to obtain the optimized COD prediction value; Calculating a residual sequence between the optimized COD prediction value and the actual COD value of the target water sample, performing least squares estimation on the residual sequence, and generating a residual correction vector; Constructing a random forest model, predicting the residual correction vector according to the random forest model, and generating an integrated corrected COD value; A consistency check process is performed on the integrated corrected COD value to obtain a correction result set.
[0012] Furthermore, the step of verifying the correction result and generating a COD analysis set includes: The correction result set, optimized feature set and preliminary COD prediction value are subjected to multi-source data fusion processing to obtain a fused feature set; The fused feature set is subjected to particle filtering to obtain a stable COD estimation set, wherein the state vector is defined as the COD value and its rate of change, and the observation vector is the fused feature vector; The stable COD estimation set is segmented, and each segment is processed by discrete cosine transform to obtain a multi-scale feature set; The multi-scale feature set is subjected to random forest regression processing to obtain a COD analysis set.
[0013] The present invention also provides a fully automatic water quality COD analysis device, comprising: an acquisition module, configured to acquire an absorbance spectrum and electrochemical characteristics of a target water sample in a first preset wavelength band and generate an original signal set; a fusion module, configured to obtain dynamic features of a fusion liquid of the target water sample and the potassium dichromate solution in a second preset band, and perform feature-level fusion of the dynamic features with the original signal set to generate a reaction feature set; An optimization module, configured to perform manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set; An output module is used to perform hybrid neural network modeling processing based on the optimized feature set and output a preliminary COD prediction value; The verification module is used to perform correction processing on the preliminary COD prediction value, verify the correction result, and generate a COD analysis set.
[0014] The present invention further provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.
[0015] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above method when executed by a processor.
[0016] Compared with the prior art, this application has the following beneficial effects: This application proposes a fully automatic water quality COD analysis method and device. By acquiring multidimensional signals (such as absorbance spectra and electrochemical characteristics) and performing feature-level fusion, it fully explores the synergistic characteristics of multiple pollutants in water samples and generates a more comprehensive set of reaction features. Through manifold learning and feature optimization, data quality is further improved and background interference and noise effects are reduced. Modeling processing based on hybrid neural networks can accurately capture complex nonlinear relationships and output high-precision preliminary COD prediction values. Finally, through correction and verification processing, the reliability and stability of the analysis results are ensured. This application can efficiently integrate multi-source signal data and perform in-depth optimization processing, overcoming the shortcomings of existing technologies in complex water sample analysis, which have low analysis efficiency and accuracy due to insufficient data processing capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not intended to limit the conditions under which this application can be implemented, and therefore have no substantive technical significance. Any structural modifications, changes in proportional relationships, or adjustments in size, without affecting the efficacy and objectives that can be achieved by this application, should still fall within the scope of the technical contents disclosed in this application.
[0019] Figure 1 This is a schematic diagram of the steps of a fully automatic water quality COD analysis method in one embodiment of the present invention; Figure 2 This is a schematic block diagram of the structure of a fully automatic water quality COD analysis device according to one embodiment of the present invention; Figure 3 is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention; The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0021] Those skilled in the art will appreciate that, unless expressly stated otherwise, the singular forms "a", "an", "above", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when an element is said to be "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any module and all combinations of one or more associated listed items.
[0022] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which this invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as such, will not be interpreted in an idealized or overly formal sense.
[0023] Reference Figure 1 The embodiment of the present invention provides a fully automatic water quality COD analysis method, comprising the following steps: S1: Obtain the absorbance spectrum and electrochemical characteristics of the target water sample in the first preset band to generate an original signal set; In step S1, the absorbance spectrum of the target water sample is collected using an integrated micro-spectrometer to perform high-resolution spectral scanning within a first preset wavelength band (200-235 nm). This wavelength band covers the ultraviolet to visible light range and can effectively capture the characteristic absorption peaks of organic pollutants and reducing substances in the water sample, generating a spectral matrix containing a time series. The matrix records the change in absorbance at each wavelength over time. The electrochemical characteristics of the water sample are synchronously collected using an electrochemical sensor array, mainly including the redox potential (ORP) and current response. These electrochemical signals reflect the electrochemical activity of oxidizable substances in the water sample, such as the current change or potential shift caused by the oxidation reaction of organic matter on the electrode surface. This generates an electrochemical signal matrix, which also contains time series information. The spectral matrix and electrochemical signal matrix were preprocessed. For the spectral matrix, wavelet transform was used for denoising. Specifically, based on the Daubechies 4 wavelet basis, the signal was decomposed into five layers, retaining the low-frequency components to reconstruct the spectral signal. This process effectively filtered out high-frequency noise (such as instrument jitter or ambient light interference) while retaining a smooth signal that reflects the characteristics of organic matter in the water sample. For the electrochemical signal matrix, frequency domain analysis was performed using fast Fourier transform (FFT). Frequency features in the 0.01-50 Hz range were extracted and frequency domain feature vectors were constructed. This frequency range covers the main dynamic responses in the electrochemical reaction. The processed spectral matrix, electrochemical signal matrix, and frequency domain eigenvectors were standardized to eliminate dimensional differences and amplitude deviations between different signal sources. Z-score standardization was performed to achieve a mean of 0 and a standard deviation of 1 for each signal. The standardized spectral eigenvectors, electrochemical eigenvectors, and timestamp information were fused to generate a raw signal set with a data dimension of 10,000 × 256, where 10,000 represents the number of sampling points in the time series and 256 represents the fused feature dimension, encompassing multidimensional spectral and electrochemical information. This raw signal set comprehensively characterizes the spectral and electrochemical properties of the water sample, reflecting both the absorbance peak of benzene series in the UV band and the current response of the oxidation reaction on the electrode surface.
[0024] S2: obtaining dynamic features of the fusion liquid of the target water sample and the potassium dichromate solution in a second preset band, and performing feature-level fusion on the dynamic features and the original signal set to generate a reaction feature set; In step S2, the target water sample collected in step S1 is mixed with potassium dichromate solution in a specific proportion and subjected to a chemical reaction at a constant temperature of 95°C. Potassium dichromate is a strong oxidant and can undergo redox reaction with organic matter in the water sample to generate hexavalent chromium ions and other products. This reaction process is accompanied by changes in absorbance and electrochemical signals. Therefore, a highly sensitive spectrometer and electrochemical sensor are used to collect the absorbance changes and electron transfer rates of the reaction solution in the 300-450 nm band in real time to generate a reaction spectrum matrix and a current time series. The 300-450 nm band covers the characteristic absorption peaks of hexavalent chromium ions and their reaction products, such as Cr 6+ to Cr 3+ The current time series reflects the dynamic characteristics of electron transfer during the reaction process, such as the electron transfer rate on the electrode surface during organic oxidation. After data acquisition, the reaction spectrum matrix is differentially processed, and the reaction rate curve is extracted by calculating the rate of change of absorbance over time. This curve can characterize the oxidation rate of organic matter in the reaction solution. For example, when treating an industrial wastewater sample, the reaction rate curve may show a rapid initial decrease in absorbance, corresponding to the rapid consumption of easily oxidizable organic matter, and a gradual decrease in absorbance in the later stages, reflecting the reaction characteristics of difficult-to-degrade substances. A Hilbert transform is applied to the current time series to extract the instantaneous phase and amplitude, generating a phase eigenvector. This process captures the dynamic oscillation characteristics of the electrochemical signal, extracts low-frequency phase changes in the current signal related to the oxidation rate of organic matter, and generates an eigenvector reflecting the reaction dynamics. Multiscale entropy analysis calculates the complexity characteristics of the reaction spectrum matrix and current time series, generating an entropy eigenvector. This characteristic reflects the nonlinear dynamic characteristics of the reaction system. For example, high-COD water samples have higher entropy values for the reaction spectrum and current signal due to the wide variety of complex organic matter, while low-COD water samples exhibit lower complexity. The introduction of the entropy eigenvector enhances the reaction feature set's ability to characterize the complexity of the water sample. The extracted reaction rate curve, phase eigenvector, and entropy eigenvector are then fused with the spectral eigenvector of the original signal set in step S1 at the feature level. This fusion process aims to integrate multi-source information before and after the reaction to comprehensively characterize the chemical and physical properties of the organic matter in the water sample and form a more comprehensive feature description. To reduce data redundancy and improve computational efficiency, kernel principal component analysis (KPCA) was used to reduce the dimensionality of the fused features, retaining the principal components with a cumulative contribution rate of 90%. The reduced fused feature vectors were generated to form the final reaction feature set with a data dimension of 5000 × 128, where 5000 represents the number of sampling points in the time series and 128 represents the feature dimension after dimensionality reduction, including the fused feature vectors, reaction rate metadata, and complexity features.
[0025] S3: performing manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set; In step S3, the t-distributed stochastic neighbor embedding (t-SNE) algorithm is used to perform manifold learning processing on the fused feature vectors in the reaction feature set. t-SNE is a nonlinear dimensionality reduction method that maps the high-dimensional fused feature vectors (dimension 5000×128) in the reaction feature set to a two-dimensional manifold space. The perplexity is set to 30 to balance the local and global structures. It is iterated 1000 times to ensure convergence, and a two-dimensional manifold feature vector is generated. This process can compress the complex data structure in the reaction feature set that reflects the chemical reaction characteristics of the water samples into a two-dimensional representation that is easier to analyze, highlighting the characteristic differences of water samples with different COD concentrations. After completing manifold learning, the two-dimensional manifold feature vectors were clustered using the spectral clustering algorithm. The number of clusters was set to 5 to capture the potential categories of water sample characteristics, such as high COD, low COD, or reaction patterns of different organic matter types. Spectral clustering constructed a similarity matrix and calculated the eigenvalues and eigenvectors of the Laplace matrix to assign data points to 5 clusters and generate a cluster label vector. This vector assigned a category label to each data point, reflecting the grouping characteristics of the water sample in the reaction feature space. Nonlinear regression analysis was performed on the reaction rate metadata in the reaction feature set, and Gaussian kernel support vector regression (SVR) was used to fit the mapping relationship between reaction rate and COD concentration. The Gaussian kernel function can effectively capture the nonlinear change trend of the reaction rate and generate a regression coefficient vector, which quantifies the contribution of the reaction rate to COD prediction. The complexity features in the reaction feature set were screened using the information gain algorithm. The information gain threshold was set to 0.8, retaining only features with high information content for COD prediction to generate an important feature subset. This process effectively eliminated redundant or low-correlation features. The two-dimensional manifold feature vector, cluster label vector, regression coefficient vector, and important feature subset were weighted and fused using the entropy weight method. The entropy weight method assigns weights based on the information entropy of each feature, ensuring that features with high information content occupy a larger proportion in the fusion. This generates a comprehensive feature vector, which integrates the results of manifold learning, cluster analysis, regression analysis, and feature screening, and comprehensively characterizes the reaction characteristics of the water sample. The comprehensive feature vector was regularized using the L2 norm. The normalization operation controlled the numerical range of the feature vector to prevent deviations caused by feature scale differences in subsequent modeling. The optimized feature set containing the comprehensive feature vector and feature importance metadata was generated.
[0026] S4: performing hybrid neural network modeling processing according to the optimized feature set and outputting a preliminary COD prediction value; In step S4, a hybrid neural network model is constructed, taking as input the comprehensive feature vectors (2000×64) from the optimized feature set. These feature vectors contain the water quality reaction characteristics after manifold learning and feature optimization, such as a comprehensive representation of spectral absorbance and reaction rate. The hybrid model consists of a one-dimensional convolutional neural network and a gated recurrent unit network. The convolutional layer uses 64 convolution kernels of size 5 with a stride of 1. This layer captures local variations in absorbance at specific wavelengths during the reaction. These local features are then subjected to dimensionality reduction using a max-pooling layer of size 2, reducing computational complexity while preserving key feature information, enabling the model to more efficiently process high-dimensional data. The output of the pooling layer is fed into a gated recurrent unit (GRU) layer with 256 units to capture long-term temporal dependencies in the feature vectors. The GRU layer models the temporal variation of the reaction rate and identifies characteristic differences between the initial and stable phases of the reaction, thereby improving prediction accuracy. The output of the GRU layer is compressed into a single node through the fully connected layer to generate a preliminary COD prediction value. This prediction value directly reflects the mapping relationship between the input features and the COD concentration. During the model training process, a five-fold cross-validation method is used to divide the optimized feature set into five equal parts. Four parts are used for training and one part for validation. After training, repeated sampling is performed from the optimized feature set to generate multiple subsets, which are input into the model to calculate the prediction value. The distribution of these prediction values is statistically analyzed to generate a 95% confidence interval. This confidence interval provides a quantitative indicator of the reliability of each prediction value, generating a prediction result set containing a preliminary COD prediction value, a 95% confidence interval, and an attention weight. The data dimension is 2000×3, where 2000 represents the number of sampling points and 3 represents the prediction value, upper and lower bounds of the confidence interval, and attention weight corresponding to each sampling point.
[0027] S5: After correcting the preliminary COD prediction value, verifying the correction result to generate a COD analysis set; In step S5, a multi-source data fusion model is constructed. This model takes as input the comprehensive feature vector from the optimized feature set, the preliminary COD predictions from the prediction result set, and the corrected COD predictions from the correction result set (including the corrected COD value, confidence interval, and residual correction vector). This model integrates these heterogeneous data into a unified feature representation and employs the extended Kalman filter (EKF) algorithm to address the state estimation problem in nonlinear systems. Specifically, the state vector is defined as the COD value and its rate of change, where the COD value reflects the chemical oxygen demand (COD) of water quality, and the rate of change captures its dynamic trend. The observation vector consists of the corrected COD prediction (from the correction result set) and the comprehensive feature vector (from the optimized feature set). Through the iterative update process of the extended Kalman filter, the state estimate is continuously refined based on the observed data to generate a fused COD prediction. During the filtering process, the state transition matrix describes the temporal evolution of the COD value and its rate of change, while the observation matrix maps the state vector to the observation data space. By calculating the Kalman gain, the fusion model balances the weight between the model prediction and the observed data, thereby generating more accurate COD predictions. After generating the fused COD predictions, multi-scale analysis is performed on the fused COD predictions using wavelet transform decomposition. Wavelet transforms decompose the signal into frequency components, revealing the characteristics of the COD predictions at different time scales. In the implementation, the decomposition level is set to 4, and the discrete wavelet transform (e.g., Daubechies wavelet basis) is used to decompose the COD prediction sequence into low-frequency approximation components and high-frequency detail components. The root mean square error (RMSE) is calculated for each scale component to generate a validation metric set. The RMSE is calculated for the deviation of each component from the true COD value or the corrected COD prediction, resulting in a set of validation metrics, such as [0.12, 0.08, 0.05, 0.03], reflecting the prediction accuracy of each scale component. After the validation metrics are generated, the fused COD predictions are corrected using the gradient boosted decision tree (GBDT). This ensemble of multiple weak learners (decision trees) fits complex data relationships, demonstrating strong nonlinear modeling capabilities. In the implementation, the tree depth was set to 5 and the number of trees was set to 100. The gradient boosting algorithm was used to iteratively optimize the COD predictions. During the training process, GBDT takes the fused COD predictions and the validation indicator set as input, learns the residual between the predicted values and the true values, and generates the final COD results through weighted combination.
[0028] In one embodiment, the step of obtaining the absorbance spectrum and electrochemical characteristics of the target water sample in the first preset band to generate an original signal set includes: Collecting the absorbance spectrum of the target water sample in the 200-235 nm band, performing denoising on the absorbance spectrum to generate a spectral signal matrix; Based on a multi-channel electrochemical sensor array, the redox potential and current response of the target water sample are synchronously collected to obtain an electrochemical signal matrix; Extracting low-frequency components of the spectral signal matrix and reconstructing the signal, performing principal component analysis on the reconstructed signal, and generating a spectral feature matrix; Perform frequency domain feature extraction on the electrochemical signal matrix to obtain the frequency domain feature vector: The spectral feature matrix and the frequency domain feature vector are subjected to Z-score normalization processing, and the spectral feature matrix and the frequency domain feature vector are weightedly spliced and fused according to information entropy to obtain an original signal set.
[0029] In the above example, a target water sample is placed in a quartz cuvette and scanned within the 200-235 nm wavelength range using a spectrophotometer. The absorbance data is recorded to generate a continuous absorbance spectrum. De-noising is achieved using algorithms such as wavelet transform or sliding average filtering. For example, discrete wavelet transform (DWT) is used to decompose the signal, filtering out high-frequency noise while retaining low-frequency, significant signals, thereby generating a spectral signal matrix. This matrix forms a two-dimensional data structure with wavelength as the horizontal axis and absorbance as the vertical axis. An electrochemical sensor is constructed that includes multiple electrodes, for example, a three-electrode system consisting of a working electrode (e.g., a glassy carbon electrode), a reference electrode (e.g., an Ag / AgCl electrode), and an auxiliary electrode. The multi-channel design allows for simultaneous acquisition of signals from different electrodes. The target water sample is placed in an electrochemical cell, and by applying a specific potential (e.g., a sweeping potential in cyclic voltammetry), the redox potential and current response are simultaneously measured. The redox potential reflects the chemical activity of oxidizing or reducing species in the water sample, while the current response characterizes the intensity of the redox reaction occurring on the electrode surface. These signals are recorded as time series data by a data acquisition system, forming an electrochemical signal matrix. Each channel corresponds to the signal of an electrode, and each column of the matrix represents the potential or current value at a given time point. Frequency domain analysis of the spectral signal matrix is performed. Fast Fourier transform (FFT) can be used to convert the absorbance signal from the time domain to the frequency domain, extracting low-frequency components to filter out high-frequency noise and transient interference. The low-frequency components contain the main features of the spectral signal and reflect the absorption characteristics of organic matter. After extracting the low-frequency components, the signal is reconstructed using an inverse Fourier transform to obtain a smooth spectral signal. Principal component analysis (PCA) is performed on the reconstructed signal to reduce data dimensionality and extract key features. PCA calculates the covariance matrix of the spectral signal matrix and extracts the principal components (i.e., eigenvectors). The result is a spectral feature matrix, where each row represents a sample and each column represents a principal component score. This significantly reduces data redundancy while preserving the key information of the original spectral signal. Frequency domain feature extraction is performed on the electrochemical signal matrix to generate a frequency domain feature vector. Fast Fourier transform (FFT) can also be used to transform the electrochemical signal matrix into the frequency domain to extract frequency domain features, such as the signal's power spectral density, dominant frequency components, or energy distribution within a specific frequency range. These features reflect the periodicity or dynamic properties of the electrochemical signal and can characterize the intensity and regularity of redox reactions in the water sample. After frequency domain feature extraction, a frequency domain feature vector is generated, in which each element corresponds to a specific frequency domain eigenvalue and is organized in vector form. The mean and standard deviation of the spectral feature matrix and frequency domain feature vector are calculated, and each data point is normalized using the formula (x - μ) / σ, where x is the original data, μ is the mean, and σ is the standard deviation. This eliminates differences in feature dimensions and scales, allowing the spectral and electrochemical features to be integrated into the same dimension. The weight of each feature is calculated based on information entropy.Information entropy measures the uncertainty of each feature to determine its information content. Features with lower information entropy (i.e., higher information content) are assigned higher weights. In practice, the information entropy values of the spectral feature matrix and frequency domain feature vectors are calculated, and then weights are assigned based on the entropy values, for example, using the normalized reciprocal entropy as a weighting factor. Weighted concatenation fusion performs a weighted linear combination or matrix concatenation of the spectral feature matrix and frequency domain feature vectors to generate a comprehensive raw signal set containing the fusion information of spectral and electrochemical features.
[0030] In one embodiment, the step of obtaining dynamic features of the fusion liquid of the target water sample and the potassium dichromate solution in a second preset band, and performing feature-level fusion of the dynamic features with the original signal set to generate a reaction feature set includes: The absorbance changes of the fusion liquid of the target water sample and potassium dichromate solution in the 300-450nm band are collected to generate a dynamic spectral dataset containing time and wavelength dimensions; Processing the absorbance sequence of each wavelength in the dynamic spectral data set based on a time series difference algorithm, calculating the change rate of absorbance at adjacent time points, and generating an absorbance change rate matrix; Performing reaction rate feature extraction processing on the absorbance change rate matrix to obtain a reaction rate feature set; collecting an electrochemical signal sequence of the fusion liquid, and performing frequency domain analysis on the electrochemical signal sequence to generate a fusion frequency domain feature set; The reaction rate feature set, the fused frequency domain feature set and the spectral feature vector in the original signal set are fused at the feature level to obtain the reaction feature set.
[0031] In the above embodiment, the water sample in the 300-450 nm band covers the hexavalent chromium (Cr 6Characteristic absorption peaks of ⁺ (such as the strong absorption of dichromate at approximately 350nm) are suitable for monitoring the dynamics of oxidation reactions in chemical oxygen demand (COD) analysis. Specifically, a target water sample is mixed with a potassium dichromate solution in a specific ratio. A monitoring meter continuously scans the 300-450nm wavelength range, recording absorbance data at different time points during the reaction. Because potassium dichromate oxidation of organic matter is time-dependent, absorbance changes as the reaction progresses. The collected data includes both wavelength (300-450nm) and time (e.g., recording every few seconds), generating a three-dimensional dynamic spectral dataset. This dataset uses wavelength as the horizontal axis, time as the vertical axis, and absorbance as the numerical value, forming a multidimensional data structure encompassing both time and wavelength dimensions. Time series analysis is performed on the absorbance sequence at each wavelength (e.g., 300nm, 301nm, etc.). A first-order difference algorithm is used to calculate the change in absorbance between adjacent time points: ΔA(t) = A(t+1)-A(t), where A(t) is the absorbance at time t and ΔA(t) is the change. To characterize the reaction rate, the rate of change is further calculated: ΔA(t) / Δt, where Δt is the time interval (e.g., 1 second). This rate of change reflects the dynamic rate of the potassium dichromate oxidation reaction at a specific wavelength and is closely related to the oxidation ease of organic matter in the water sample. This process is repeated for all wavelengths to generate an absorbance change rate matrix, where rows correspond to different wavelengths and columns correspond to different time points. The matrix elements are the absorbance change rates for the corresponding wavelengths and time points. Feature engineering is performed on the absorbance change rate matrix to extract key features that can characterize the oxidation reaction rate. For example, statistical analysis can be used to extract the mean, maximum, slope, or integrated area of each wavelength rate of change series. These features reflect the rate characteristics of the reaction at different stages. Using an electrochemical workstation or multi-channel electrochemical sensor array, the electrochemical signals of the fused liquid during potassium dichromate oxidation can be monitored in real time. The reaction rate feature set, the fused frequency domain feature set, and the spectral feature vectors from the original signal set are then fused at the feature level to generate a reaction feature set. Feature-level fusion can be achieved through concatenation, weighted combination, or deep learning methods. For example, the weights of each feature set are calculated based on information entropy, with feature sets with lower entropy (higher information content) receiving higher weights. A fused feature vector is then generated through weighted linear combination. Another approach is to use methods such as principal component analysis or autoencoders to map the three feature sets into a low-dimensional space to generate a comprehensive feature vector. The resulting reaction feature set integrates spectral dynamic features, electrochemical frequency domain features, and original spectral features, enabling a comprehensive characterization of the physicochemical properties of the target water sample during the potassium dichromate oxidation reaction.
[0032] In one embodiment, the step of performing manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set includes: Constructing a local reconstruction weight matrix according to the k nearest neighbors of each fused feature vector in the reaction feature set, solving the local reconstruction weight matrix globally and mapping it to a three-dimensional manifold space to obtain a manifold feature set; According to a preset neighborhood radius and a preset minimum number of samples, the core points and boundary points of each point in the manifold feature set are calculated, and cluster clusters are generated according to a density reachability relationship to obtain a cluster feature label set; Performing a four-level decomposition on the time series signal of the dynamic reaction metadata in the reaction feature set, extracting low-frequency coefficients and high-frequency coefficients, calculating energy distribution characteristics based on the decomposition coefficients, and generating a time series feature vector set; Calculating the mutual information value between the time series feature vector and the cluster feature label set, and screening to obtain a highly correlated feature subset according to a preset mutual information threshold; The manifold feature set, cluster feature label set, time series feature vector set and high correlation feature subset are weightedly fused by entropy weight method to obtain a comprehensive feature vector set; Regularization and feature optimization processing are performed on the comprehensive feature vector set to obtain an optimized feature set.
[0033] In the above embodiment, the k nearest neighbors for each feature vector are found using Euclidean distance or other distance metrics (such as Manhattan distance or cosine similarity). The selection of k nearest neighbors requires a balance between preserving local structure and computational complexity. The value of k can be adjusted empirically based on the size and characteristics of the dataset. After finding the k nearest neighbors, the local reconstruction weight matrix is constructed by minimizing the reconstruction error to calculate the weights between each sample point and its neighbors. This local reconstruction weight matrix is then mapped to a three-dimensional manifold space through global optimization to obtain a manifold feature set. Two parameters are set: the neighborhood radius (ε) and the minimum number of samples (MinPts). The neighborhood radius defines the neighborhood range of a sample point, while the minimum number of samples determines whether a point can be considered a core point. Each point in the manifold feature set is traversed and the number of sample points in its neighborhood is calculated. If the number is greater than or equal to MinPts, the point is considered a core point. If its neighborhood contains core points but the point itself is not a core point, it is considered a boundary point. If it is neither a core point nor a boundary point, it is marked as a noise point. Based on these classifications, clusters are generated using density reachability relationships (i.e., gradually expanding from a core point to its neighborhood). Each cluster represents a collection of samples with similar characteristics within the manifold feature set, and the cluster feature label set is the cluster label assigned to each sample point. The time series signal of the dynamic response metadata is subjected to a four-level wavelet decomposition. The four-level decomposition signal is decomposed into low-frequency approximate components and high-frequency detail components, with each decomposition further decomposing the low-frequency components of the previous level into lower-frequency approximate components and detail components. After the four-level decomposition, the extracted low-frequency coefficients reflect the overall trend and high-order characteristics of the signal, while the high-frequency coefficients capture local variations and detailed features. Based on these decomposition coefficients, energy distribution features are calculated, typically by normalizing the sum of the squares of the decomposition coefficients at each level to obtain the energy contribution at each level. These energy features constitute the time series feature vector set. By calculating the mutual information value between each eigenvector in the time series feature vector set and the cluster feature label set, the joint probability distribution and marginal probability distribution of the eigenvector and label can be estimated. The calculated mutual information value reflects the contribution of each time series feature to the cluster label. A preset mutual information threshold (which can be determined based on experience or statistical analysis) is set, and features with mutual information values above the threshold are screened to form a highly correlated feature subset. The entropy value of each feature set is calculated. Lower entropy values indicate greater information content and should therefore be assigned a higher weight. Through normalization, the entropy values are converted into weight coefficients. The manifold feature set, cluster feature label set, time series feature vector set, and highly correlated feature subset are then linearly weighted and fused according to the weights to generate a comprehensive feature vector set. The comprehensive feature vector set is then normalized (e.g., zero-mean unit variance normalization), and then feature selection or dimensionality reduction algorithms are applied to select the feature subset that contributes most to the target task (e.g., classification or regression), resulting in an optimized feature set.
[0034] In one embodiment, the step of performing hybrid neural network modeling processing based on the optimized feature set and outputting a preliminary COD prediction value includes: Constructing a convolutional neural network model, performing a convolution operation on the optimized feature set based on the convolutional neural network model, and performing dimensionality reduction processing on the convolution output through an average pooling layer to generate a convolutional feature atlas containing multi-scale features; Performing temporal dependency modeling on the convolutional feature atlas to obtain a temporal feature sequence set; Performing attention weighting processing according to the feature importance metadata in the temporal feature sequence set and the optimized feature set to obtain a weighted feature vector set; A two-layer fully connected neural network was constructed, with 256 neurons in the first layer and 1 neuron in the second layer. The weighted feature vector set was input into the fully connected neural network for feature mapping, and the preliminary COD prediction value was obtained as output.
[0035] In the above embodiment, the convolutional neural network structure is designed, including parameters such as the number of convolutional layers, convolution kernel size, and stride. For example, multiple convolutional layers can be used, each with a 3*3 or 5*5 convolution kernel. Local features are extracted from the optimized input feature set using a sliding window approach. The essence of the convolution operation is to extract local patterns or features of the data by performing an inner product operation between the convolution kernel and the input data. An activation function, such as ReLU (Rectified Linear Unit), is added after the convolution layer to zero out negative values, thereby accelerating the convergence of gradient descent. The convolution output is subjected to dimensionality reduction using an average pooling layer. By taking the average value within a specified window, the size of the feature map is reduced, thereby reducing computational complexity and preventing overfitting. In implementation, the pooling window size can be set to 2*2 or 3*3, and sliding with an appropriate stride to generate a convolutional feature map containing multi-scale features. Temporal dependency modeling of the convolutional feature map can be implemented using models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or gated recurrent units (GRUs). Taking LSTM as an example, by introducing memory cells and three gates (input gate, forget gate, and output gate) to control the flow of information, long-term dependencies can be effectively captured. During implementation, the convolutional feature atlas can be expanded along the time dimension to form a sequence input into the LSTM model. Assuming the shape of the convolutional feature atlas is (number of samples, time steps, feature dimensions), the LSTM processes each feature map time-step by time step and outputs a time series feature sequence set. Attention weighting is performed based on the feature importance metadata in the time series feature sequence set and the optimized feature set. By assigning different weights to each element of the input, the more important parts for the task are highlighted. In this embodiment, the feature importance metadata can be a weight calculated by a feature selection algorithm, reflecting the contribution of each feature in the optimized feature set to COD prediction. To implement weighted attention processing, a softmax-based attention mechanism can be used to concatenate or map the temporal feature set with feature importance metadata to generate a comprehensive feature representation. Attention scores are then calculated using a fully connected layer or multi-layer perceptron (MLP). These scores are normalized using a softmax function to obtain weights for each feature. These weights are then multiplied by the corresponding features in the temporal feature set to generate a set of weighted feature vectors. This allows for dynamic adjustment of the model's focus on different features. For example, in COD prediction, certain time points or environmental parameters (such as temperature or pH) may have a greater impact on the results, and the attention mechanism will automatically assign higher weights to these features. A two-layer fully connected neural network is constructed, and the weighted feature vector set is input for feature mapping, outputting a preliminary COD prediction. A fully connected neural network (FCN) is a classic neural network architecture suitable for mapping high-dimensional features into a low-dimensional output space.In this embodiment, 256 neurons are set in the first layer and 1 neuron is set in the second layer. The 256 neurons in the first layer can be regarded as a nonlinear transformation of the weighted feature vector set, capturing the complex relationship between features through a large number of neurons. To implement this layer, matrix multiplication can be used to multiply the input vector by the weight matrix, and a bias term can be added, and then nonlinearity can be introduced through the ReLU activation function. The 1 neuron in the second layer directly outputs a scalar value, that is, the preliminary COD prediction value. During implementation, the second layer can use no activation function to ensure the continuity of the output, which is suitable for regression tasks. To train the entire model, the mean squared error (MSE) can be selected as the loss function, and the Adam optimizer can be used for gradient descent update to obtain the preliminary COD prediction value.
[0036] In one embodiment, the step of correcting the preliminary COD prediction value includes: The preliminary COD prediction value was optimized by genetic algorithm hyperparameter optimization to obtain the optimized COD prediction value; Calculating a residual sequence between the optimized COD prediction value and the actual COD value of the target water sample, performing least squares estimation on the residual sequence, and generating a residual correction vector; Constructing a random forest model, predicting the residual correction vector according to the random forest model, and generating an integrated corrected COD value; A consistency check process is performed on the integrated corrected COD value to obtain a correction result set.
[0037] In the above embodiment, the preliminary COD predictions were subjected to genetic algorithm hyperparameter optimization. A genetic algorithm (GA) is a global optimization algorithm based on the principles of biological evolution. It improves prediction accuracy by optimizing model hyperparameters (such as the learning rate, number of neurons, and convolution kernel size) or post-processing parameters of the predicted values. A GA population is defined as a set of possible hyperparameter combinations, with each combination considered an individual. For example, parameters such as the learning rate and the number of hidden layer neurons can be encoded as binary or real values to form an initial population. The performance of each individual is evaluated using a fitness function, which is typically based on the error between the preliminary COD prediction and the true COD value, such as the mean squared error (MSE). To achieve optimization, the GA iteratively updates the population through selection, crossover, and mutation operations. In practical applications, the population size can be set to 50-100, the number of iterations to 100-200 generations, and the crossover and mutation rates can be adjusted based on computing resources (for example, a crossover rate of 0.8 and a mutation rate of 0.01). After multiple generations of evolution, the genetic algorithm will converge to a set of optimal hyperparameter combinations, and the preliminary COD prediction value will be adjusted based on these parameters to generate an optimized COD prediction value. The residual sequence between the optimized COD prediction value and the true COD value of the target water sample is calculated. The residual sequence refers to the difference sequence between the optimized COD prediction value and the true COD value, which reflects the deviation of the model prediction. During implementation, the true COD value of the target water sample is collected, which can be obtained through laboratory chemical analysis, such as the COD value of the water sample measured using the potassium dichromate method mentioned in the above embodiment. The optimized COD prediction value is subtracted from the true COD value one by one to obtain a residual sequence. The residual sequence is modeled by least squares estimation, with the goal of finding a function (such as a linear function or a polynomial function) to minimize the sum of squares of the residual sequence. During implementation, it can be assumed that the residual sequence satisfies a certain functional form, such as e i = a + b*x i + ε i , where x i is an auxiliary variable related to the residual (possibly from the optimized feature set or timestamp), a and b are parameters to be estimated, and ε iis the random error. Parameters a and b are solved using the least squares interaction multiplication method to obtain a fitted model, which then generates a residual correction vector. This vector can be considered a further correction to the optimized COD prediction, aiming to compensate for the model's systematic bias in certain samples. A random forest model is constructed and used to predict the residual correction vector. The role of the random forest is to model the residual correction vector and further optimize the COD prediction. To implement this step, training data is prepared, with the residual correction vector as the target variable. Input features can include original features from the optimized feature set, the optimized COD prediction value, or other relevant metadata. The random forest model consists of multiple decision trees, each trained using randomly selected samples and features. During implementation, the number of decision trees can be set, and each decision tree is constructed using a random subset of the training data using bootstrap sampling. At each node split, a random subset of features (for example, the square root of the total number of features) is evaluated to reduce correlation between trees. After training, the random forest generates an ensemble corrected COD value by averaging the predictions of all decision trees (for regression tasks). To improve model performance, hyperparameters, such as the maximum depth of the decision tree or the minimum number of leaf node samples, can be adjusted through cross-validation. Consistency testing of the integrated corrected COD values can be performed using a variety of statistical methods, such as calculating the Pearson correlation coefficient between the integrated corrected COD values and the true COD values to assess the degree of consistency between the two. Furthermore, residual analysis can be used to check whether the residuals of the corrected predicted values conform to a normal distribution or exhibit systematic bias. If the data volume is large enough, the distribution of prediction errors before and after correction can be compared using a t-test or F-test to verify whether the correction effect is significant. In practical applications, consistency testing can also be combined with domain knowledge, such as checking whether the corrected COD values are within a reasonable range (e.g., 0-1000 mg / L, depending on the water sample type). If the predicted values for certain samples deviate significantly from the true values, further analysis can be conducted, such as due to data anomalies or model limitations, and these samples can be removed from the corrected result set or marked as data for verification. The corrected result set contains the COD predicted values after all steps.
[0038] In one embodiment, the step of verifying the correction result and generating a COD analysis set includes: The correction result set, optimized feature set and preliminary COD prediction value are subjected to multi-source data fusion processing to obtain a fused feature set; The fused feature set is subjected to particle filtering to obtain a stable COD estimation set, wherein the state vector is defined as the COD value and its rate of change, and the observation vector is the fused feature vector; The stable COD estimation set is segmented, and each segment is processed by discrete cosine transform to obtain a multi-scale feature set; The multi-scale feature set is subjected to random forest regression processing to obtain a COD analysis set.
[0039] In the above embodiment, data fusion technology is used to integrate the three into a fusion feature set, and methods including weighted average fusion, feature splicing or deep learning fusion can be used. For example, the correction result set, the optimized feature set and the preliminary COD prediction value can be directly connected into a high-dimensional vector through feature splicing, or a weighted average method can be used to give different weights according to the confidence or importance of each data and then fuse them. The fusion feature set is processed by particle filtering, which is used to smooth the noise in the fusion feature set and generate a stable COD estimate. When implementing, the state transition model and the observation model are first defined. The state vector is defined as the COD value and its rate of change, that is, {x t = [COD t , ΔCOD t ]}, where COD t Indicates the COD value at time t, ΔCOD t Represents the rate of change of COD value. The state transition model can be modeled as x t = f(x {t-1} , w t ), where f is the state transfer function, w t is the process noise, usually assumed to be Gaussian noise. The observation vector is the fusion feature vector z t , the observation model is z t = h(x t , v t ), where h is the observation function, v t For observation noise, we first initialize a set of particles, each particle represents a possible value of the state vector. For example, we randomly generate N particles {x ti , i=1, 2, ..., N}, and assign an initial weight w to each particle ti =1 / N. Then, the next state of each particle is predicted according to the state transition model, that is, x is calculated by function f ti ; Using the observation vector z t Update the weight of each particle. The weight update formula is w ti ∝ p(z t |x ti ), where p(z t |x ti ) represents the observation likelihood, which can be calculated based on the distribution of observation noise. In this embodiment, the fusion feature vector z tThese observations can be viewed as highly correlated with COD values, and likelihood can be calculated using a regression model or distance metric. After updating the weights, resampling is performed, retaining particles with higher weights and removing those with lower weights to prevent particle degradation. A stable COD estimate set is obtained by weighted averaging the states of all particles. The stable COD estimate set is segmented based on data characteristics or application requirements, such as by fixed time intervals (e.g., hourly) or adaptively segmented based on COD fluctuation trends, using a sliding window or change-point-based segmentation algorithm. After segmentation, a discrete cosine transform is applied to each data segment to decompose the COD estimate into cosine components of different frequencies. Low-frequency components reflect overall trends, while high-frequency components capture local fluctuations. In implementation, a one-dimensional DCT algorithm is applied to the COD value sequence for each data segment to obtain a set of DCT coefficients, which constitute a multi-scale feature set. A random forest regression is performed on the multi-scale feature set, using the multi-scale feature set as input. The target variable is either the true COD value (if used for training) or the aforementioned stable COD estimate (if used for prediction). Specifically, the number of decision trees is set (e.g., 100-500), and each decision tree is constructed by randomly selecting a subset from the training data through bootstrap sampling. When each tree node splits, a random subset of features (e.g., the square root of the total number of features) is randomly selected for evaluation to reduce correlation between trees. After training, the random forest averages the predictions from all decision trees to generate the final COD analysis set. This resulting COD analysis set is a collection of high-precision COD values processed through multi-source fusion, particle filtering, DCT transformation, and random forest regression, making it suitable for practical applications in fully automated water quality analysis.
[0040] Reference Figure 2 , a fully automatic water quality COD analysis device, comprising: An acquisition module 100 is used to acquire an absorbance spectrum and electrochemical characteristics of a target water sample in a first preset wavelength band and generate an original signal set; A fusion module 200 is configured to obtain dynamic features of a fusion liquid of the target water sample and the potassium dichromate solution in a second preset band, and perform feature-level fusion of the dynamic features with the original signal set to generate a reaction feature set; An optimization module 300 is used to perform manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set; An output module 400 is configured to perform hybrid neural network modeling processing based on the optimized feature set and output a preliminary COD prediction value; The verification module 500 is used to perform correction processing on the preliminary COD prediction value, and then perform verification processing on the correction result to generate a COD analysis set.
[0041] Reference Figure 3 In the embodiment of the present application, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as a robot three-dimensional environment database. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a fully automatic water quality COD analysis method is implemented.
[0042] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, a fully automatic water quality COD analysis method is implemented, comprising the steps of: obtaining an absorbance spectrum and electrochemical characteristics of a target water sample in a first preset band to generate an original signal set; obtaining dynamic characteristics of a fusion liquid of the target water sample and a potassium dichromate solution in a second preset band, performing feature-level fusion of the dynamic characteristics with the original signal set to generate a reaction feature set; performing manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set; performing hybrid neural network modeling processing based on the optimized feature set to output a preliminary COD prediction value; performing correction processing on the preliminary COD prediction value, and then verifying the correction result to generate a COD analysis set.
[0043] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0044] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A fully automatic water quality COD analysis method, characterized in that: include: Obtaining the absorbance spectrum and electrochemical characteristics of the target water sample in the first preset band to generate an original signal set; Acquire dynamic features of a fusion liquid of the target water sample and the potassium dichromate solution in a second preset band, and perform feature-level fusion of the dynamic features with the original signal set to generate a reaction feature set; Performing manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set; Perform hybrid neural network modeling processing based on the optimized feature set and output a preliminary COD prediction value; After the preliminary COD prediction value is corrected, the correction result is verified to generate a COD analysis set.
2. A fully automatic water quality COD analysis method according to claim 1, characterized in that, The step of obtaining the absorbance spectrum and electrochemical characteristics of the target water sample in the first preset band to generate an original signal set includes: Collecting the absorbance spectrum of the target water sample in the 200-235 nm band, performing denoising on the absorbance spectrum to generate a spectral signal matrix; Based on a multi-channel electrochemical sensor array, the redox potential and current response of the target water sample are synchronously collected to obtain an electrochemical signal matrix; Extracting low-frequency components of the spectral signal matrix and reconstructing the signal, performing principal component analysis on the reconstructed signal, and generating a spectral feature matrix; Perform frequency domain feature extraction on the electrochemical signal matrix to obtain the frequency domain feature vector: The spectral feature matrix and the frequency domain feature vector are subjected to Z-score normalization processing, and the spectral feature matrix and the frequency domain feature vector are weightedly spliced and fused according to information entropy to obtain an original signal set.
3. A fully automatic water quality COD analysis method according to claim 1, characterized in that, The step of obtaining the dynamic characteristics of the fusion liquid of the target water sample and the potassium dichromate solution in the second preset band, and performing feature-level fusion of the dynamic characteristics with the original signal set to generate a reaction feature set includes: The absorbance changes of the fusion liquid of the target water sample and potassium dichromate solution in the 300-450nm band are collected to generate a dynamic spectral dataset containing time and wavelength dimensions; Processing the absorbance sequence of each wavelength in the dynamic spectral data set based on a time series difference algorithm, calculating the change rate of absorbance at adjacent time points, and generating an absorbance change rate matrix; Performing reaction rate feature extraction processing on the absorbance change rate matrix to obtain a reaction rate feature set; collecting an electrochemical signal sequence of the fusion liquid, and performing frequency domain analysis on the electrochemical signal sequence to generate a fusion frequency domain feature set; The reaction rate feature set, the fused frequency domain feature set and the spectral feature vector in the original signal set are fused at the feature level to obtain the reaction feature set.
4. A fully automatic water quality COD analysis method according to claim 1, characterized in that, The step of performing manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set includes: Constructing a local reconstruction weight matrix according to the k nearest neighbors of each fused feature vector in the reaction feature set, solving the local reconstruction weight matrix globally and mapping it to a three-dimensional manifold space to obtain a manifold feature set; According to a preset neighborhood radius and a preset minimum number of samples, the core points and boundary points of each point in the manifold feature set are calculated, and cluster clusters are generated according to a density reachability relationship to obtain a cluster feature label set; Performing a four-level decomposition on the time series signal of the dynamic reaction metadata in the reaction feature set, extracting low-frequency coefficients and high-frequency coefficients, calculating energy distribution characteristics based on the decomposition coefficients, and generating a time series feature vector set; Calculating the mutual information value between the time series feature vector and the cluster feature label set, and screening to obtain a highly correlated feature subset according to a preset mutual information threshold; The manifold feature set, cluster feature label set, time series feature vector set and high correlation feature subset are weightedly fused by entropy weight method to obtain a comprehensive feature vector set; Regularization and feature optimization processing are performed on the comprehensive feature vector set to obtain an optimized feature set.
5. A fully automatic water quality COD analysis method according to claim 1, characterized in that, The step of performing hybrid neural network modeling processing according to the optimized feature set and outputting a preliminary COD prediction value includes: Constructing a convolutional neural network model, performing a convolution operation on the optimized feature set based on the convolutional neural network model, and performing dimensionality reduction processing on the convolution output through an average pooling layer to generate a convolutional feature atlas containing multi-scale features; Performing temporal dependency modeling on the convolutional feature atlas to obtain a temporal feature sequence set; Performing attention weighting processing according to the feature importance metadata in the temporal feature sequence set and the optimized feature set to obtain a weighted feature vector set; A two-layer fully connected neural network was constructed, with 256 neurons in the first layer and 1 neuron in the second layer. The weighted feature vector set was input into the fully connected neural network for feature mapping, and the preliminary COD prediction value was obtained as output.
6. A fully automatic water quality COD analysis method according to claim 1, characterized in that, The step of correcting the preliminary COD prediction value comprises: The preliminary COD prediction value was optimized by genetic algorithm hyperparameter optimization to obtain the optimized COD prediction value; Calculating a residual sequence between the optimized COD prediction value and the actual COD value of the target water sample, performing least squares estimation on the residual sequence, and generating a residual correction vector; Constructing a random forest model, predicting the residual correction vector according to the random forest model, and generating an integrated corrected COD value; A consistency check process is performed on the integrated corrected COD value to obtain a correction result set.
7. A fully automatic water quality COD analysis method according to claim 1, characterized in that, The step of verifying the correction result and generating a COD analysis set includes: The correction result set, optimized feature set and preliminary COD prediction value are subjected to multi-source data fusion processing to obtain a fused feature set; The fused feature set is subjected to particle filtering to obtain a stable COD estimation set, wherein the state vector is defined as the COD value and its rate of change, and the observation vector is the fused feature vector; The stable COD estimation set is segmented, and each segment is processed by discrete cosine transform to obtain a multi-scale feature set; The multi-scale feature set is subjected to random forest regression processing to obtain a COD analysis set.
8. A fully automatic water quality COD analysis device, characterized in that: include: an acquisition module, configured to acquire an absorbance spectrum and electrochemical characteristics of a target water sample in a first preset wavelength band and generate an original signal set; a fusion module, configured to obtain dynamic features of a fusion liquid of the target water sample and the potassium dichromate solution in a second preset band, and perform feature-level fusion of the dynamic features with the original signal set to generate a reaction feature set; An optimization module, configured to perform manifold learning and feature optimization processing on the reaction feature set to generate an optimized feature set; An output module is used to perform hybrid neural network modeling processing based on the optimized feature set and output a preliminary COD prediction value; The verification module is used to perform correction processing on the preliminary COD prediction value, verify the correction result, and generate a COD analysis set.
9. A computer device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Online monitoring process for components of lead frame electroplating solution
CN121113922A
An on-line monitoring process of plating solution composition for lead frame
CN121113922B
Self-adaptive regulation and control method and device for Fenton oxidation equipment
CN121143234A