Intelligent detection method for biodegradability of printing and dyeing wastewater based on fluorescence component analysis

CN122836009APending Publication Date: 2026-09-29ZHEJIANG PROVINCE JIAHUA PRINTING & DYEING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610707708.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]目前,评估印染废水可生化性的标准方法是五日生化需氧量与化学需氧量的比值法,即BOD5/COD比值法,其中BOD5的测试周期较长,需要五日,无法满足印染废水实时调控的需求,部分印染废水中会存在抑制好氧生物活性的物质,如重金属络合基团中释放的铜离子等等,导致BOD5测试结果失准,使BOD5/COD比值法无法真实反应印染废水的生物处理潜力,在现有技术中,为了改进上述问题,已有研究尝试利用三维荧光光谱对印染废水的性质进行评估,如公开号为CN115165825A的中国专利中通过将三维荧光光谱图划分为数个固定的激发-发射区域,计算各区域积分体积占比,并寻找与BOD5/COD值相关性最高的单一区域进行线性拟合,从而建立预测模型,这样的方式,虽然实现了印染废水的可生化性快速评估,但是存在:依赖于预先设置的固定区域划分,难以捕捉印染废水复杂多变的荧光特征,在印染废水的成分发生变化时,固定区域的代表性下降,导致模型的预测性下降;仅选取了单一关联区域进行建模,而光谱中其他区域同样包含与微生物代谢、有机物腐殖化程度等信息,三维荧光光谱数据与可生化性之间的深层关联未被发掘;预测模型为简单的线性数理统计模型,对非线性关系的刻画能力有限,无法实现与后续处理工艺的智能联动,因而,根据上述缺陷,提出一种新的基于荧光组分解析的印染废水可生化性智能检测方法

Benefits of technology

通过标准化的样品预处理和内滤效应校正,从源头保证数据质量,利用平行因子分析模型解析出具有明确化学意义的独立荧光组分,替代传统粗略的区域积分,从分子层面揭示有机物组成,融合组分贡献度与原始光谱细节并构建高维特征向量,这一技术链条使提取的特征信息更本质、稳定,奠定了高精度基础;将特征向量输入机器学习模型进行训练与预测,能刻画复杂的非线性关系,使模型在独立测试集上的预测精度稳定,对比传统需要5天的BOD5检测时间更短,实现了从即时诊断的目的;实时预测值与工艺调控规则联动,使得方法从单纯检测工具转变为参与生产调控的完整闭环智能系统,可以保证处理工艺的稳定运行与节能降耗。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122836009A_ABST
    Figure CN122836009A_ABST
Patent Text Reader

Abstract

The application discloses a printing and dyeing wastewater biodegradability intelligent detection method based on fluorescence component analysis, and relates to the technical field of wastewater detection.The technical scheme is as follows: the printing and dyeing wastewater biodegradability intelligent detection method based on fluorescence component analysis comprises sample pretreatment, correction of three-dimensional fluorescence spectrum data set, parallel factor analysis model analysis of fluorescence components, construction of high-dimensional characteristic vector, optimization of biodegradability intelligent prediction model, and real-time regulation of biochemical treatment process.Through standardization sample pretreatment and internal filter effect correction, independent fluorescence components with clear chemical significance are analyzed by combining the parallel factor analysis model, the traditional rough regional integral is replaced, the high-dimensional characteristic vector is constructed by fusing component contribution and original spectrum details, the extracted characteristic information is more essential and stable, the characteristic vector is input into a machine learning model for training and prediction, the complex nonlinear relationship can be described, the prediction accuracy of the model on an independent test set is stable, and the detection time is shorter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wastewater detection technology, and more specifically, to an intelligent detection method for the biodegradability of dyeing and printing wastewater based on fluorescence component analysis. Background Technology

[0002] Wastewater generated from fabric dyeing and printing contains a large amount of recalcitrant organic matter such as dyes and auxiliaries. If this type of wastewater is discharged directly, it will have a serious impact on the environment. Therefore, it is necessary to treat the dyeing and printing wastewater through biochemical treatment to meet the discharge standards. Accurate and rapid assessment of the biodegradability of dyeing and printing wastewater is a key prerequisite for optimizing the biochemical treatment process, achieving standard discharge and energy conservation and consumption reduction.

[0003] Currently, the standard method for assessing the biodegradability of dyeing and printing wastewater is the five-day biochemical oxygen demand (BOD5) to chemical oxygen demand (COD) ratio method. However, the BOD5 test cycle is relatively long, requiring five days, which cannot meet the needs of real-time control of dyeing and printing wastewater. Furthermore, some dyeing and printing wastewater contains substances that inhibit aerobic biological activity, such as copper ions released from heavy metal complexes, leading to inaccurate BOD5 test results. Therefore, the BOD5 / COD ratio method cannot truly reflect the biological treatment potential of dyeing and printing wastewater. To address these issues, existing research has attempted to use three-dimensional fluorescence spectroscopy to assess the properties of dyeing and printing wastewater. For example, Chinese patent CN115165825A divides the three-dimensional fluorescence spectrum into several fixed excitation-emission regions, calculates the integral volume ratio of each region, and identifies the relationship between the BOD5 / COD ratio and the actual wastewater treatment potential. While linear fitting of the single region with the highest correlation to establish a predictive model enables rapid assessment of the biodegradability of dyeing and printing wastewater, this approach suffers from several drawbacks: It relies on pre-defined fixed region divisions, making it difficult to capture the complex and variable fluorescence characteristics of dyeing and printing wastewater; the representativeness of the fixed region decreases as the composition of the wastewater changes, leading to a decline in the model's predictive accuracy; it only selects a single correlated region for modeling, while other regions in the spectrum also contain information related to microbial metabolism and the degree of organic matter humification, failing to uncover the deep correlation between three-dimensional fluorescence spectral data and biodegradability; and the predictive model is a simple linear mathematical statistical model with limited ability to characterize nonlinear relationships, making it unable to achieve intelligent linkage with subsequent treatment processes. Therefore, based on these shortcomings, a novel intelligent detection method for the biodegradability of dyeing and printing wastewater based on fluorescence component analysis is proposed. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the present invention aims to provide an intelligent detection method for the biodegradability of dyeing and printing wastewater based on fluorescence component analysis.

[0005] The above-mentioned technical objective of this invention is achieved through the following technical solution: an intelligent detection method for the biodegradability of dyeing and printing wastewater based on fluorescence component analysis, comprising the following steps: S1. Collect comprehensive wastewater samples from different dyeing and printing production stages to form a sample set. Perform standardized pretreatment on the sample set and determine the BOD5 / COD ratio of each sample in the set. Use this as the benchmark true value label for biodegradability assessment and for training the subsequent intelligent prediction model for biodegradability. S2. Under constant temperature and automatic sample injection conditions, a three-dimensional fluorescence spectrometer is used to scan the pretreated sample to obtain the original excitation-emission matrix data. The original excitation-emission matrix data is systematically corrected, and the mathematical correction model of the internal filter effect is applied simultaneously for compensation to obtain the corrected three-dimensional fluorescence spectrum dataset. S3. Import the corrected three-dimensional fluorescence spectrum dataset from step S2 into the parallel factor analysis mathematical model for iterative calculation and signal decoupling. Adaptively determine the number N of independent fluorescent components in the sample through core consistency diagnosis and leave-one-out cross-validation, and extract the excitation-emission load spectrum corresponding to each component and its relative contribution score matrix in the sample. S4. Based on the component contribution scores and the corrected three-dimensional fluorescence spectrum dataset, extract and fuse the standardized contribution scores of each component, the bioavailability index defined by the ratio of the contribution of tryptophan-like and humic acid-like components, and the peak intensity ratio of specific fluorescence regions to construct a high-dimensional feature vector. S5. Using high-dimensional feature vectors as input variables and biodegradability benchmark true value as target variables, divide the training set, validation set and test set, select the ensemble learning model as the basic architecture, and train and fine-tune the ensemble learning model through hyperparameter optimization and validation set early stopping strategy to obtain the final biodegradability intelligent prediction model. S6. For the target dyeing and printing wastewater to be tested, repeat steps S1 and S2 to complete the pretreatment and three-dimensional fluorescence spectrum acquisition and correction. Apply the parallel factor analysis model established in step S3 to analyze the fluorescent components of the target dyeing and printing wastewater. Execute step S4 to construct a high-dimensional feature vector and input it into the fully trained intelligent prediction model for biodegradability in step S5. Output the predicted value of biodegradability of the wastewater in real time. When the predicted value is lower than the preset threshold, automatically trigger or adjust the dosing rate of the external carbon source dosing system to guide the real-time control of the biochemical treatment process.

[0006] The present invention is further configured such that: the goal of gradient dilution in step S1 is to control the ultraviolet absorbance value of the comprehensive wastewater sample at a wavelength of 254 nm to below 0.05, thereby eliminating the influence of the internal filtration effect on the fluorescence intensity measurement.

[0007] The present invention is further configured such that: the specific fluorescent component in step S3 includes an independent fluorescent component having a specific excitation / emission wavelength pair, which is resolved from characteristic wastewater containing reactive dyes, disperse dyes or auxiliaries by parallel factor analysis.

[0008] The present invention is further configured such that: in step S4, the bioavailability index is further used as a priority decision indicator for starting the carbon source addition system in step S6.

[0009] The present invention is further configured such that the ensemble learning model in step S5 is a gradient boosting machine.

[0010] The present invention is further configured such that the biodegradability intelligent prediction model trained in step S5 has the following prediction performance on the independent test set: the coefficient of determination between the predicted value and the measured value by the standard method is greater than or equal to 0.92, and the root mean square error is less than or equal to 0.045.

[0011] The present invention is further configured to: dynamically generate and implement the preset threshold and corresponding carbon source injection acceleration rate adjustment amount in step S6 based on the real-time decision of the preset logic or process control model in the process control rule library.

[0012] An electronic device includes a three-dimensional fluorescence spectroscopy data acquisition module and a data processing module. The three-dimensional fluorescence spectroscopy data acquisition module is used to perform step S2, and the data processing module is used to perform steps S3-S6. The data processing module includes a data memory and a processor, wherein the data memory stores a computer program that can execute the steps on the processor.

[0013] In summary, the present invention has the following beneficial effects: By standardizing sample pretreatment and correcting for internal filtration effects, data quality is ensured from the source. Parallel factor analysis models are used to resolve independent fluorescent components with clear chemical significance, replacing traditional coarse regional integration. This reveals the composition of organic matter at the molecular level. The contribution of components is integrated with the original spectral details to construct a high-dimensional feature vector. This technological chain makes the extracted feature information more essential and stable, laying the foundation for high accuracy. Inputting the feature vector into a machine learning model for training and prediction can characterize complex nonlinear relationships, making the model's prediction accuracy stable on independent test sets. Compared with the traditional 5-day BOD5 detection time, it is shorter, achieving the goal of instant diagnosis. The real-time predicted values ​​are linked with process control rules, transforming the method from a simple detection tool into a complete closed-loop intelligent system participating in production control, which can ensure the stable operation of the processing process and energy saving. Attached Figure Description

[0014] Figure 1 This is a flowchart of the intelligent detection and control method for biodegradability of the present invention; Figure 2This is a block diagram of the electronic device structure of the method of the present invention. Detailed Implementation

[0015] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. Example

[0016] Intelligent detection method for biodegradability of dyeing and printing wastewater based on fluorescence component analysis, such as Figure 1 As shown, it includes the following steps: Step S1, Sample Construction and Standardization Preprocessing: A total of 200 wastewater samples were collected from multiple dyeing and printing enterprises. The sampling must adhere to the principles of representativeness and coverage: the samples must cover the main types of dyeing and printing fibers at the source of the process, including natural fibers, synthetic fibers, and various blended fabrics; within the same dyeing and printing enterprise, wastewater from each stage of the process, including pretreatment, dyeing / printing, and finishing, as well as the final integrated equalization tank, should be collected; similarly, wastewater should be collected at different production cycles, and for intermittent sources, samples should be collected throughout the entire drainage cycle. The aim is to construct a sample set that can fully reflect the diversity of wastewater characteristics in the target area or target process. During collection, clean polyethylene or glass containers should be used. After thorough rinsing at the sampling point, instantaneous water samples or 24-hour mixed water samples should be collected, with a minimum sampling volume of 2L. Samples must be refrigerated during transportation and pretreated and subjected to three-dimensional fluorescence spectroscopy scanning within 48 hours. The preprocessing procedure is as follows: Step 1: After thoroughly shaking the water sample, take 500 mL of each sample from the collection and place it in a centrifuge tube. Centrifuge at 8000 r / min for 15 minutes at 4℃, and then remove the supernatant and bottom precipitate particles. Step 2: Filter the supernatant after centrifugation through a 0.45μm filter membrane to remove colloidal particles and microorganisms, and obtain a clear filtrate; Step 3: Use a UV-Vis spectrophotometer to measure the UV absorbance A of the filtrate obtained in Step 2 at a wavelength of 254 nm. If A > 0.05, use ultrapure water for gradient dilution until A ≤ 0.05 to ensure that the attenuation of fluorescence intensity caused by the internal filtering effect is reduced to a negligible level, thus ensuring the linearity and authenticity of the subsequent fluorescence signal. During the above process, the dilution factor D needs to be recorded simultaneously for the correction of relevant parameters of the original concentration of the fluorescence detection sample during calculation. After completing the above three steps, for each pretreated sample, determine its COD and BOD5 according to the national standard method. The COD value should be determined using the dichromate method. The COD value of the pretreated water sample should be multiplied by the dilution factor D to obtain the original COD value of the water sample. The BOD5 value should be determined using the dilution and inoculation method. After culturing in a constant temperature incubator at 20±1℃ for 5 days, the BOD5 value should be determined. Similarly, the result should be multiplied by the dilution factor D to obtain the original BOD5 value of the water sample. Then, calculate the BOD5 / COD ratio. Use this BOD5 / COD ratio as the true value of biodegradability. This value should be between 0.15 and 0.55. Step S2, Three-dimensional fluorescence spectral acquisition and calibration: Measurement is performed using a fluorescence spectrophotometer equipped with an autosampler. The fluorescence spectrophotometer should be configured as follows: Excitation wavelength Ex: 220nm-450nm, with each excitation band separated by 5nm, for a total of 47 excitation wavelengths; The emission wavelength Em is 250nm-550nm, with each 2nm interval forming a transmission band, and data is collected in a total of 151 emission wavelengths. Acquisition speed: 1200nm / min, photomultiplier tube voltage set to medium response, slit width usually set to 5nm / 5nm (excitation / emission) to obtain sufficient signal intensity and maintain appropriate resolution. Before each scan, baseline calibration with ultrapure water is required to ensure the signal stability of the fluorescence spectrophotometer. The sample pretreated in step S1 is injected into the flow cell. The fluorescence spectrophotometer performs a full-band scan according to the above settings. For each set excitation wavelength, the corresponding complete emission spectrum is recorded. Finally, each sample will generate a raw excitation-emission matrix with a dimension of 47×151. During the above process, an ultrapure water blank sample is scanned simultaneously to obtain its Raman scattering peak intensity. From the raw excitation-emission matrix data of all samples, the corresponding Raman scattering intensity is subtracted according to the wavelength point. For the data of the first-order (Em = Ex) and second-order (Em = 2*Ex) Rayleigh scattering band regions, the data is set to zero or replaced with the average value of the adjacent non-scattering region by interpolation method to eliminate the strong interference signal generated by water molecule Raman scattering and Rayleigh scattering in the optical path in the raw excitation-emission matrix data. Using a UV-Vis spectrophotometer with ultrapure water as a reference, the absorption spectrum of the same pretreated sample was measured in the range of excitation and emission wavelengths that matched the scanning of the fluorescence spectrophotometer. For each data point in the original excitation-emission matrix, its fluorescence intensity attenuation correction factor F due to the internal filtering effect is... c (Ex, Em) are calculated according to the following formula: , In the formula: A(Ex) is the absorbance of the sample at the excitation wavelength Ex, and A(Em) is the absorbance of the sample at the emission wavelength Em; By multiplying the apparent fluorescence intensity of each data point in the original excitation-emission matrix after scattering subtraction by the corresponding correction factor, the corrected fluorescence intensity can be obtained. After performing the above steps on all samples in sequence, the corrected three-dimensional fluorescence spectrum dataset is obtained. Each sample data in this dataset is an excitation-emission matrix that eliminates the main physical interference and can more realistically reflect the fluorescence characteristics of organic matter in water. Step S3, Intelligent analysis of fluorescence components based on parallel factor analysis: a. Format and organize the corrected three-dimensional fluorescence spectrum dataset of all samples obtained in step S2, and import it into the parallel factor analysis mathematical model. This dataset can be represented as a three-dimensional array X with dimensions I×J×K, where I represents the number of samples, J represents the number of emission wavelengths, and K represents the number of excitation wavelengths. Each element X in the array... IJK This represents the corrected fluorescence intensity of the I-th sample at the J-th emission wavelength and the K-th excitation wavelength; b. Parallel factor analysis mathematical model is a trilinear decomposition method based on multivariate curve resolution. The core idea of ​​the parallel factor analysis mathematical model is: assuming that the observed three-dimensional data array X is composed of several fluorescent components with constant spectral shapes, and the concentration of each component in different samples is linearly variable. The model decomposes X into three loading matrices: Emission payload spectral matrix A: dimension J×F, where F is the preset number of components, and each column represents the fluorescence emission spectral profile of a fluorescent component across all emission wavelengths; Excitation load spectral matrix B: dimension K×F, each column represents the fluorescence excitation spectral profile of a fluorescent component across all excitation wavelengths; Score matrix C: dimension I×F, each row corresponds to a sample, and each column represents the relative contribution of the sample to the corresponding fluorescent component; The mathematical expression of the model is: , Where eijk is the residual, the decomposition process is solved by alternating least squares method, aiming to find A, B, C to minimize the sum of squared residuals. In the implementation of this invention, non-negativity constraints are applied to both the load spectrum and the score matrix to ensure that the resolved spectral profile and concentration contribution have physical meaning. c. Determine the number of components F using the following steps: Core consistency diagnosis: Calculate the core consistency of the model under different F values. When the F value is equal to or exceeds the actual chemical group fraction in the data, the core consistency will drop significantly from close to 100%. Select the largest F value that still maintains high core consistency, generally greater than 90%, as the candidate. Leave-one-out cross-validation: The sample set is divided into a training set and a validation set. Models with different F-values ​​are built based on the training set, and the fluorescence data of the validation set is predicted. The sum of squared predicted residuals is compared, and the F-value that minimizes the prediction error is selected. Residual analysis: Examine the residuals e of the model under the selected F-value. IJK Whether the distribution is random and lacks obvious structural features indicates that the model has fully extracted the signal, and the final optimal F is determined through the above steps; d. Result analysis and component confirmation: After the model runs, it outputs the final three matrices A, B, and C. Acquisition and identification of excitation-emission load spectra: A graph is plotted between each column of matrix B and the corresponding column of matrix A to obtain the excitation-emission load spectrum of the fluorescent component. By comparing the excitation-emission load spectrum with the fluorescence spectra of known standard substances, the component is chemically identified. In the application of this invention to dyeing and printing wastewater, at least the following can be stably resolved: Component 1: The maximum excitation / emission peaks are located around 220, 275 / 310 nm, which are identified as tyrosine-like substances and are associated with easily biodegradable protein-like organic matter; Component 2: The maximum excitation / emission peaks are located around 220, 275 / 350 nm, which are identified as tryptophan-like substances and are also associated with easily degradable protein-like organic matter; Component 3: The maximum excitation / emission peaks are located around 260 / 450 nm, indicating that it is a humic acid-like substance, representing humic organic matter that is difficult to biodegrade; Component 4: Characteristic fluorescent components associated with specific dyes or auxiliaries are identified, and the positions of their excitation-emission peaks depend on the specific substance.

[0017] Obtaining the component contribution score matrix: Matrix C is the required component contribution score matrix, where elements C... IF This represents the relative contribution of the I-th sample to the F-th fluorescent component; Step S4, Construction of multi-dimensional fusion feature vector: From the data sources in steps S2 and S3, three types of feature parameters are systematically extracted: a. Component Standardized Contribution: Based on matrix C in step S3, calculate the contribution score C of each row of matrix C for each component. IFPerform standardization, calculate its relative percentage contribution, or perform Z-score standardization: Relative percentage: This yields a set of F-dimensional features with a total sum of 100%; Z-score standardization: , where μ f and σ f These are the mean and standard deviation of the scores of all samples on the f-th component, respectively. This set of F-dimensional features directly quantifies the relative abundance of various fluorescent substances, such as protein-like and humic substances, in the sample, and is the most direct description of the organic composition of the sample. b. Bioavailability Index: Following step S3, the contribution score of the component identified as tryptophan-like amino acids is recorded as C. Trp The contribution score of the component identified as humic acid-like is denoted as C. Humic According to the calculated bioavailability index: Among them, tryptophan-like substances are usually associated with easily biodegradable protein and amino acid organic matter, while humic acid-like substances represent humic substances with complex structures and difficult degradation. The bioavailability index calculated in the above formula reflects the ratio of easily degradable components to difficult-to-degrade components in wastewater. The higher the ratio, the better the biodegradability of the wastewater in theory. c. Peak intensity ratio of specific fluorescence regions: Based on the three-dimensional fluorescence spectral data corrected in step S2, and according to the commonalities of a large number of dyeing and printing wastewater spectra, two to three meaningful excitation-emission region pairs are predefined. Region pair R1: excitation Ex: 270-280nm, emission Em: 330-360nm; Region pair R2: excitation Ex: 250-260nm, emission Em: 420-460nm. For each sample, calculate the volume integral or mean peak intensity of the fluorescence intensity within a specified region using its corrected excitation-emission matrix data, and then calculate the ratio between these regional integral values: This ratio can sensitively reflect the relative intensity changes between different fluorophores, which may be related to specific biochemical processes or pollution sources, and is an effective supplement to the characteristics of component contribution. All the extracted feature parameters are concatenated sequentially to construct a unified high-dimensional feature vector for each sample. The vector dimension is calculated as follows: assuming four components are analyzed and one BAI index and two fluorescence region ratios are calculated, the total dimension of the final feature vector is: 4 (component contribution) + (BAI) + 2 (peak intensity ratio) = 7 dimensions. For the i-th sample, its final feature vector... It can be represented as: ,

[0018] Before inputting feature vectors into the model, the features of the entire training set are usually standardized (e.g., Z-score) to make the mean of each feature 0 and the standard deviation 1, in order to eliminate the difference in units and optimize the training process and performance of the machine learning model. Step S5, Construction and training of a biodegradability intelligent prediction model based on machine learning algorithms: a. Dataset preparation and partitioning The high-dimensional feature vectors constructed for all samples in step S4 are used as the model input variable (X), and paired with the corresponding BOD5 / COD baseline ground truth values ​​measured in step S1 as the target output variable (y), thus forming a complete modeling dataset. To ensure the model has reliable generalization ability, this dataset must be divided into three mutually exclusive subsets: Training set: Used for fitting and learning model parameters; its sample size typically accounts for 60-70% of the total dataset. Validation set: Used to monitor model performance during training, tune hyperparameters, and implement early stopping strategies to suppress overfitting; its sample size is typically 15-20%. Test set: Used for final performance evaluation of the fully trained model, providing an unbiased estimate of its predictive ability on new samples; its sample size is typically 15-20%. b. Model Architecture Selection Dataset partitioning should be performed using a stratified randomized method. Given the significant differences in the sources and processes of dyeing and printing wastewater samples, the partitioning must ensure that the distribution of biodegradability (i.e., B / C value) of samples in the training, validation, and test sets remains basically consistent to avoid introducing systematic bias due to data partitioning; This paper adopts the gradient booster machine from the ensemble learning model as the core modeling architecture. Traditional linear methods, such as linear regression and partial least squares regression, are difficult to fully characterize the complex nonlinear relationship between high-dimensional fluorescence features and biodegradability indicators. On the other hand, single decision tree models are prone to overfitting and insufficient generalization ability. The gradient booster machine can effectively capture the above-mentioned complex nonlinear mapping by building weak learners, usually decision trees, in the sequence and continuously fitting the residuals of the previous stage prediction. In addition, based on the characteristics of tree models, it is not sensitive to the scale of input features. Although standardization can be used to accelerate convergence, it is not necessary. After training, the gradient booster machine can provide a ranking of the contribution of each feature to the prediction results. This can not only verify the effectiveness of the constructed features, such as biogenicity index, but also enhance the interpretability of the biodegradability intelligent prediction model. Therefore, the gradient booster machine achieves a better balance between model bias and variance, and can be used as the basic algorithm architecture of this invention. c. Model training and hyperparameter optimization During the model training process of this invention, the key hyperparameters involved in the gradient boosting machine include: the number of weak learners, the learning rate that controls the contribution weight of each tree, the maximum depth that determines the complexity of a single tree, and the minimum number of samples required for the internal node to be repartitioned. The optimization process is as follows: First, based on the training set, K-fold cross-validation is used in combination with grid search or random search strategies to optimize the system in the predefined hyperparameter space. Then, under each set of parameters, the average performance index of the model on the validation set is calculated through cross-validation, such as negative mean square error or coefficient of determination R², to monitor and evaluate the effect of the parameters. Finally, the hyperparameter combination with the best performance index on the validation set is selected as the final configuration of the biodegradability intelligent prediction model. d. Early stopping strategy and model finalization To suppress overfitting during model training, an early stopping strategy based on the validation set is adopted. Specifically, during the final model training based on optimal hyperparameters, the loss function (e.g., mean squared error) on the independent validation set is monitored in real time. Training automatically terminates when the loss function no longer decreases within a pre-defined number of consecutive iterations. This strategy stops iterations when the model's generalization performance reaches its optimal level, thus avoiding the memorization of data noise due to overtraining and obtaining the final model with the strongest generalization ability. After training is complete, the final model's performance must be independently evaluated and accepted using a test set that was not involved in training or tuning. The evaluation uses the following metrics and methods: Coefficient of determination: The coefficient of determination R² between the model's predicted values ​​and the measured values ​​using standard methods is calculated to measure the model's ability to explain data fluctuations. This invention requires that the fully trained model have an R² of no less than 0.92 on the test set; Root mean square error: This is the root mean square value of the prediction error, which directly reflects the average absolute error level of the predicted value. Performance visualization analysis: Plot a scatter plot of "predicted value - measured value" for the test set samples, and overlay an ideal fitting line for comparison. The closer the scatter points are to the ideal line, the higher the model prediction accuracy. Finally, after the model is built and trained in step S5, an independent test set that has not participated in the training and hyperparameter tuning should be used to conduct the final evaluation and verification of the generalization performance of the final biodegradability intelligent prediction model. The evaluation and verification refer to the indicators and methods mentioned above. Step S6, Intelligent detection and process guidance for the biodegradability of the target wastewater sample: The trained intelligent prediction model for biodegradability is then embedded into the actual wastewater detection system to complete closed-loop control from spectral measurement to process regulation. Its implementation is divided into the following three stages: 1. Standardized online detection and feature extraction of target wastewater For the dyeing and printing wastewater to be tested, the system automatically executes a standardized process that is completely consistent with the model training phase: Pretreatment: Perform high-speed centrifugation, membrane filtration and gradient dilution according to step S1 to ensure that the 254nm UV absorbance is ≤0.05 to eliminate optical interference; Spectral acquisition and correction: Following the method in step S2, three-dimensional fluorescence spectra were acquired using the same instrument parameters, and scattering subtraction and internal filtering effect correction were performed to obtain standardized excitation-emission matrix data; Feature vector construction in real time: The corrected data is input into the established parallel factor analysis model to obtain the contribution score of each fluorescent component; then, according to the rules defined in step S4, the same feature set is extracted and spliced ​​in a predetermined order to form a high-dimensional feature vector that is completely consistent with the structure of the training set. 2. Real-time intelligent prediction of biodegradability The feature vectors generated in real time are input into the trained and optimized biodegradability prediction model, and the model outputs the B / C prediction value instantly. This process can be completed within minutes, achieving an effective replacement for the traditional BOD5 test. 3. Closed-loop process control based on dynamic thresholds The predicted values ​​drive real-time optimization of downstream processes via pre-defined logic: Threshold dynamic determination: The system's preset threshold (e.g., B / C=0.25) and corresponding adjustment strategy are not fixed, but are dynamically calculated and generated by the process control rule base or process control model based on real-time process status, vertical operation data, and effluent quality targets. The process control rule base is essentially a set of "condition-action" rules and mapping logic. Its construction is based on a comprehensive analysis of historical operating data, process mechanism knowledge, and expert experience. The core construction principles include: 1. Threshold Partitioning and Action Mapping: Based on long-term operational data, the predicted biochemical susceptibility value is divided into different control intervals, and specific process adjustment actions are defined for each interval, for example: If the predicted B / C ratio is < 0.25, then the system is considered to have a "severe carbon source shortage," and the action is to increase the carbon source input acceleration rate by 30%. If IF 0.25 ≤ B / C predicted value < 0.35, then the condition is determined to be "slightly insufficient carbon source," and the action is to increase the carbon source input acceleration rate by 15%. If IF 0.35 ≤ B / C predicted value ≤ 0.50, then the carbon source is deemed "suitable". Action: Maintain the current acceleration rate. If the predicted B / C ratio > 0.50, then the system is considered to have "excess carbon source," and the action is to reduce the carbon source input acceleration rate by 20%. 2. Multi-parameter coupled decision-making: Rule-based conditions can be coupled with other real-time process parameters (such as influent flow rate, pH, and ammonia nitrogen concentration) to make decisions more accurate. For example: If the predicted B / C value is < 0.30 AND the influent flow rate is > 80% of the design flow rate, then the action is to add an additional 10% acceleration rate to the baseline adjustment. 3. Time-varying and adaptive adjustment: The thresholds and action quantities in the rule base can be dynamically refreshed through the upper-level management system according to the season, sludge age or target effluent standards, to achieve offline optimization of the strategy; Dynamic generation logic of process control model: When using PID process control model, the setpoint (i.e. B / C value) of PID controller is not constant. It can be output by rule base according to working conditions or calculated in real time by a simple meta-model. For example: setpoint = base value (0.30) + f(influent COD, target ammonia nitrogen removal rate), where f() is a linear correction function fitted according to historical data. The controller calculates the deviation e(t) between the predicted value (process variable PV) and the dynamic setpoint (SP) in real time, and applies the PID formula: ,

[0019] The required carbon source acceleration rate adjustment ∆u(t) to eliminate this deviation is calculated, and finally the specific instruction u(t) to be sent to the actuator is generated. Judgment and Trigger: When the real-time predicted value is lower than the threshold, the system determines that the carbon source is insufficient and automatically triggers the control program; Precise adjustment: Control commands are sent to the carbon source addition system, and its addition rate can be dynamically adjusted in a step-by-step or proportional-integral manner according to the deviation, so as to match the carbon source supply with the actual deficit; Closed-loop control: After the carbon source is added, the system re-executes the detection-prediction-regulation cycle in the next detection cycle, thus forming an adaptive closed-loop control system of "real-time perception-intelligent prediction-precise regulation-feedback verification".

[0020] In summary, standardized sample pretreatment and internal filtration effect correction ensure data quality from the source. Parallel factor analysis models are used to resolve independent fluorescent components with clear chemical significance, replacing traditional coarse regional integration. This reveals the composition of organic matter at the molecular level. By fusing component contributions with original spectral details and constructing high-dimensional feature vectors, this technological chain makes the extracted feature information more essential and stable, laying a foundation for high accuracy. Inputting the feature vectors into machine learning models for training and prediction can characterize complex nonlinear relationships, making the model's prediction accuracy stable on independent test sets. Compared to the traditional 5-day BOD5 detection time, this is shorter, achieving the goal of immediate diagnosis. The real-time predicted values ​​are linked with process control rules, transforming the method from a simple detection tool into a complete closed-loop intelligent system participating in production control, ensuring stable operation of the processing process and energy saving.

[0021] Example 2: An electronic device for realizing an intelligent detection method for the biodegradability of dyeing and printing wastewater based on fluorescence component analysis, such as... Figure 2 As shown, this includes the following two points: 1. Equipment hardware integration and modular composition: The equipment is integrated into an industrial cabinet and mainly consists of the following two core functional modules: Three-dimensional fluorescence spectroscopy online acquisition module: This is a fluorescence spectroscopy analysis unit specifically designed for online water quality monitoring. It integrates an automated sample pretreatment and spectral detection subsystem, specifically including: Automatic sampling and pretreatment unit: Composed of a micro submersible pump, a multi-way valve, a precision injection pump, and an online filtration and centrifugation flow path, responsible for collecting water samples according to a preset program and automatically performing the standardized pretreatment of step S1 of claim 1, including adjusting the ultraviolet absorbance of the sample at 254nm to ≤0.05 in real time through gradient dilution; The spectral detection unit includes a xenon lamp light source, an excitation and emission dual monochromator, a quartz flow cell, a photomultiplier tube detector, and a constant temperature control system. It is responsible for executing step S2 of claim 1, automatically completing the spectral scan, and using a built-in algorithm to perform scattering subtraction and internal filtering effect correction. First control and communication unit: Equipped with an independent microcontroller, used to coordinate the timing control of the above operations, and upload the corrected three-dimensional fluorescence spectrum data through an industrial communication interface; Embedded data processing and control module: As the intelligent computing core of the device, it is implemented using an industrial computer or a high-performance embedded platform, and includes: Non-volatile memory: Solid-state drives store all the software models and configuration files required for operation, including: parallel factor analysis model parameters, feature construction rules, a fully trained biodegradability intelligent prediction model, and process control rule base or process control model parameters. Central Processing Unit: Configured to load and run computer programs in memory, thereby sequentially calling the corresponding models and automatically executing steps S3 to S6 in claim 1 in sequence to complete component analysis, feature construction, biodegradability prediction and regulation decision-making; 2. Equipment automation workflow and software execution logic In automatic operation mode, the device operates according to the following process cycle defined by the stored program: ① Periodic triggering: A new detection and control cycle is initiated by an internal timer or an external signal; ② Spectral data acquisition and uploading: The processor sends a start command to the acquisition module; The microcontroller of the acquisition module sequentially drives the sampling preprocessing and spectral detection units to complete water sample acquisition, standardization preprocessing, spectral scanning and real-time correction; the corrected excitation-emission matrix data is transmitted back to the memory of the data processing module through the communication interface; ③ Data processing and biodegradability prediction: Step S3: The processor calls the parallel factor analysis mathematical model to calculate the contribution score of the current water sample to each fluorescence component; Step S4: Based on the pre-stored rules, the processor calculates features such as standardized contribution and bioavailability index from the scores and spectral data, and fuses them to construct a standardized high-dimensional feature vector; Step S5: Input the feature vector into the prediction model and output the predicted value of biodegradability in real time; ④ Intelligent decision-making and control output: Step S6: The processor compares the predicted value with the preset threshold, or dynamically calculates it through the process control model, to generate specific process control instructions; these instructions are sent to the field actuators, such as the frequency converter of the carbon source dosing pump, through a standard industrial output interface. ⑤ Standby cycle: The device enters standby mode and waits for the next cycle to be triggered.

[0022] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. An intelligent detection method for the biodegradability of dyeing and printing wastewater based on fluorescence component analysis, characterized in that: Includes the following steps: S1. Collect comprehensive wastewater samples from different dyeing and printing production stages to form a sample set. Perform standardized pretreatment on the sample set and determine the BOD5 / COD ratio of each sample in the set. Use this as the benchmark true value label for biodegradability assessment and for training the subsequent intelligent prediction model for biodegradability. S2. Under constant temperature and automatic sample injection conditions, a three-dimensional fluorescence spectrometer is used to scan the pretreated sample to obtain the original excitation-emission matrix data. The original excitation-emission matrix data is systematically corrected, and the mathematical correction model of the internal filter effect is applied simultaneously for compensation to obtain the corrected three-dimensional fluorescence spectrum dataset. S3. Import the corrected three-dimensional fluorescence spectrum dataset from step S2 into the parallel factor analysis mathematical model for iterative calculation and signal decoupling. Adaptively determine the number N of independent fluorescent components in the sample through core consistency diagnosis and leave-one-out cross-validation, and extract the excitation-emission load spectrum corresponding to each component and its relative contribution score matrix in the sample. S4. Based on the component contribution scores and the corrected three-dimensional fluorescence spectrum dataset, extract and fuse the standardized contribution scores of each component, the bioavailability index defined by the ratio of the contribution of tryptophan-like and humic acid-like components, and the peak intensity ratio of specific fluorescence regions to construct a high-dimensional feature vector. S5. Using high-dimensional feature vectors as input variables and biodegradability benchmark true value as target variables, divide the training set, validation set and test set, select the ensemble learning model as the basic architecture, and train and fine-tune the ensemble learning model through hyperparameter optimization and validation set early stopping strategy to obtain the final biodegradability intelligent prediction model. S6. For the target dyeing and printing wastewater to be tested, repeat steps S1 and S2 to complete the pretreatment and three-dimensional fluorescence spectrum acquisition and correction. Apply the parallel factor analysis model established in step S3 to analyze the fluorescent components of the target dyeing and printing wastewater. Execute step S4 to construct a high-dimensional feature vector and input it into the fully trained intelligent prediction model for biodegradability in step S5. Output the predicted value of biodegradability of the wastewater in real time. When the predicted value is lower than the preset threshold, automatically trigger or adjust the dosing rate of the external carbon source dosing system to guide the real-time control of the biochemical treatment process.

2. The intelligent detection method for biodegradability of dyeing and printing wastewater based on fluorescence component analysis according to claim 1, characterized in that: The goal of gradient dilution in step S1 is to control the UV absorbance of the combined wastewater sample at a wavelength of 254 nm to below 0.05, thereby eliminating the influence of the internal filtration effect on the fluorescence intensity measurement.

3. The intelligent detection method for biodegradability of dyeing and printing wastewater based on fluorescence component analysis according to claim 1, characterized in that: The specific fluorescent components in step S3 include independent fluorescent components with specific excitation / emission wavelength pairs that are resolved from characteristic wastewater containing reactive dyes, disperse dyes, or auxiliaries by parallel factor analysis.

4. The intelligent detection method for biodegradability of dyeing and printing wastewater based on fluorescence component analysis according to claim 1, characterized in that: In step S4, the bioavailability index is further used as a priority decision indicator for starting the carbon source addition system in step S6.

5. The intelligent detection method for biodegradability of dyeing and printing wastewater based on fluorescence component analysis according to claim 1, characterized in that: In step S5, the ensemble learning model is a gradient boosting machine.

6. The intelligent detection method for biodegradability of dyeing and printing wastewater based on fluorescence component analysis according to claim 1, characterized in that: The biodegradability intelligent prediction model trained in step S5 has the following prediction performance on the independent test set: the coefficient of determination between the predicted value and the standard method measurement value is greater than or equal to 0.92, and the root mean square error is less than or equal to 0.

045.

7. The intelligent detection method for biodegradability of dyeing and printing wastewater based on fluorescence component analysis according to claim 1, characterized in that: Based on the preset logic or real-time decision-making of the process control model in the process control rule library, the preset threshold and the corresponding carbon source input acceleration rate adjustment in step S6 are dynamically generated and implemented.

8. An electronic device for performing the method according to any one of claims 1 to 7, characterized in that: It includes a three-dimensional fluorescence spectroscopy data acquisition module and a data processing module. The three-dimensional fluorescence spectroscopy data acquisition module is used to execute step S2, and the data processing module is used to execute steps S3-S6. The data processing module includes a data storage device and a processor, wherein the data storage device stores a computer program that can execute the steps on the processor.

Citation Information

Patent Citations

  • Method for evaluating biodegradability of wastewater in printing and dyeing industry based on three-dimensional fluorescence spectrum

    CN115165825A