A Soft Sensing Method, System, Electronic Device and Medium for Sewage Water Quality Index
The initial regression model is constructed through PLS and ELM algorithms, and iteratively updated with RPLS and RELM algorithms, which solves the problems of low measurement accuracy of water quality indicators and model degradation in wastewater treatment, and realizes efficient soft measurement and model maintenance.
Patent Information
- Application Number
- CN202311281633.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-10-07
AI Technical Summary
During the existing sewage treatment process, the measurement accuracy of the water quality indicators is low and the soft measurement model is prone to deterioration, resulting in difficulty in monitoring and maintenance.
The initial regression model is constructed using PLS and ELM algorithms, combined with RPLS and RELM algorithms for iterative updates, and through semi-supervised multi-rate hybrid adaptive model, the recursive partial least squares and recursive limit learning machine is integrated, the labeled data set is expanded, the model prediction performance is improved, and model degradation is overcome.
The measurement accuracy of sewage water quality indicators is improved, the problem of model degradation is solved, and the model maintenance process is simplified.
Smart Images

Figure CN117272241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial process control, and particularly to a soft measurement method, system, electronic device and medium for sewage water quality indexes. Background Art
[0002] The rapid, timely and accurate determination of water quality indexes is a difficult point in ensuring the stable sewage treatment effect. Therefore, soft measurement technology is widely used in the sewage treatment process. However, due to the characteristics of multi-sampling rates of measurement process data and complex and harsh detection environments, the prediction performance of the soft measurement model decreases, ultimately resulting in low accuracy of the measurement results of water quality indexes. And because there is a problem of model degradation in the prediction of the soft measurement model, it is extremely difficult to monitor important quality variables and maintain soft meters in the sewage treatment process. Therefore, there is an urgent need for a sewage water quality index measurement method that can improve measurement accuracy and overcome the problem of model degradation. Summary of the Invention
[0003] The purpose of the present invention is to provide a soft measurement method, system, electronic device and medium for sewage water quality indexes, which can improve measurement accuracy and overcome the problem of degradation of the prediction model.
[0004] To achieve the above purpose, the present invention provides the following solutions:
[0005] A soft measurement method for sewage water quality indexes includes:
[0006] Construct a model using the PLS algorithm based on the first labeled data subset, and construct a model using the ELM algorithm based on the second labeled data subset, respectively obtaining a first initial regression model and a second initial regression model; both the first labeled data subset and the second labeled data subset include multiple sets of water quality index sets; a set of water quality index sets includes an input water quality index set and an output water quality index set;
[0007] Use the first initial regression model to predict the first labeled data subset, and use the second initial regression model to predict the second labeled data subset, respectively obtaining a first labeled prediction error value and a second labeled prediction error value;
[0008] At the current iteration number, input the unlabeled data set at the previous iteration number into the first initial regression model and the second initial regression model respectively, obtaining a first Plus labeled data subset and a second Plus labeled data subset; the unlabeled data set includes multiple input water quality index sets;
[0009] Predict the first Plus - marked data subset using the first initial regression model, and predict the second Plus - marked data subset using the second initial regression model to obtain a first new - marked prediction error value and a second new - marked prediction error value;
[0010] Obtain the confidence levels of each input water quality index set in the unlabeled dataset at the previous iteration number according to the first marked prediction error value and the first new - marked prediction error value;
[0011] Update the first marked data subset, the second marked data subset, and the unlabeled dataset at the previous iteration number according to the target dataset to obtain the first new - marked data subset, the second new - marked data subset, and the unlabeled dataset at the current iteration number; the target dataset is the input water quality index set with the maximum confidence level;
[0012] Update the iteration number and enter the next iteration until the iteration stop condition is reached. Use RPLS to construct a first prediction model based on the first new - marked data subset at the last iteration number, and use RELM to construct a second prediction model based on the second new - marked data subset at the last iteration number;
[0013] Obtain the final prediction model according to the first prediction model and the second prediction model, and the final prediction model is used for soft - sensing of sewage water quality indicators.
[0014] Optionally, input the unlabeled dataset at the previous iteration number into the first initial regression model and the second initial regression model respectively to obtain a first Plus - marked data subset and a second Plus - marked data subset. Specifically, it includes:
[0015] Input the unlabeled dataset at the previous iteration number into the first initial regression model to obtain a first output water quality index set corresponding to the unlabeled dataset at the previous iteration number;
[0016] Input the unlabeled dataset at the previous iteration number into the second initial regression model to obtain a second output water quality index set corresponding to the unlabeled dataset at the previous iteration number;
[0017] Determine the unlabeled dataset at the previous iteration number and the first output water quality index set corresponding to the unlabeled dataset at the previous iteration number as the first Plus - marked data subset;
[0018] Determine the unlabeled dataset at the previous iteration number and the second output water quality index set corresponding to the unlabeled dataset at the previous iteration number as the second Plus - marked data subset.
[0019] Optionally, update the first labeled data subset, the second labeled data subset, and the unlabeled data set at the previous iteration according to the target data set to obtain the first new labeled data subset, the second new labeled data subset, and the unlabeled data set at the current iteration. Specifically, it includes:
[0020] Input the target data set into the first initial regression model to obtain the first output water quality index set corresponding to the target data set;
[0021] Input the target data set into the second initial regression model to obtain the second output water quality index set corresponding to the target data set;
[0022] Add the target data set and the first output water quality index set corresponding to the target data set to the second labeled data subset to obtain the first new labeled data subset at the current iteration;
[0023] Add the target data set and the second output water quality index set corresponding to the target data set to the first labeled data subset to obtain the second new labeled data subset at the current iteration;
[0024] Delete the target data set from the unlabeled data set at the previous iteration to obtain the unlabeled data set at the current iteration.
[0025] Optionally, the final prediction model is: Where Y(x) represents the final prediction model, H′1(x) represents the first prediction model, and H′2(x) represents the second prediction model.
[0026] A soft sensor system for sewage water quality indicators, comprising:
[0027] An initial regression model construction module, configured to construct a model according to the first labeled data subset by using the PLS algorithm, and construct a model according to the second labeled data subset by using the ELM algorithm, respectively obtaining a first initial regression model and a second initial regression model; both the first labeled data subset and the second labeled data subset include multiple groups of water quality index sets; a group of water quality index sets includes an input water quality index set and an output water quality index set;
[0028] A labeled prediction error value calculation module, configured to predict the first labeled data subset by using the first initial regression model, and predict the second labeled data subset by using the second initial regression model, respectively obtaining a first labeled prediction error value and a second labeled prediction error value;
[0029] A labeled data subset expansion module, which is used to input the unlabeled data set in the previous iteration into the first initial regression model and the second initial regression model respectively under the current iteration number to obtain a first Plus labeled data subset and a second Plus labeled data subset; the unlabeled data set includes a plurality of input water quality index sets;
[0030] A new labeled prediction error value calculation module, which is used to predict the first Plus labeled data subset with the first initial regression model and predict the second Plus labeled data subset with the second initial regression model to obtain a first new labeled prediction error value and a second new labeled prediction error value;
[0031] A confidence calculation module, which is used to obtain the confidence of each input water quality index set in the unlabeled data set in the previous iteration according to the first labeled prediction error value and the first new labeled prediction error value;
[0032] An update module, which is used to update the first labeled data subset, the second labeled data subset, and the unlabeled data set in the previous iteration according to the target data set to obtain a first new labeled data subset in the current iteration, a second new labeled data subset in the current iteration, and an unlabeled data set in the current iteration; the target data set is the input water quality index set with the maximum confidence;
[0033] A prediction model construction module, which is used to update the iteration number and enter the next iteration until the iteration stop condition is reached, and construct a first prediction model using RPLS according to the first new labeled data subset in the last iteration number, and construct a second prediction model using RELM according to the second new labeled data subset in the last iteration number;
[0034] A final prediction model determination module, which is used to obtain a final prediction model according to the first prediction model and the second prediction model, and the final prediction model is used for soft measurement of sewage water quality indexes.
[0035] Optionally, the labeled data subset expansion module specifically includes:
[0036] A first output water quality index set determination module, which is used to input the unlabeled data set in the previous iteration into the first initial regression model to obtain a first output water quality index set corresponding to the unlabeled data set in the previous iteration;
[0037] A second output water quality index set determination module, which is used to input the unlabeled data set in the previous iteration into the second initial regression model to obtain a second output water quality index set corresponding to the unlabeled data set in the previous iteration;
[0038] The first Plus-marked data subset determination module is used to determine the unlabeled data set at the previous iteration number and the first output water quality index set corresponding to the unlabeled data set at the previous iteration number as the first Plus-marked data subset;
[0039] The second Plus-marked data subset determination module is used to determine the unlabeled data set at the previous iteration number and the second output water quality index set corresponding to the unlabeled data set at the previous iteration number as the second Plus-marked data subset.
[0040] Optionally, the update module specifically includes
[0041] The PLS model prediction module is used to input the target data set into the first initial regression model to obtain the first output water quality index set corresponding to the target data set;
[0042] The ELM model prediction module is used to input the target data set into the second initial regression model to obtain the second output water quality index set corresponding to the target data set;
[0043] The first new-marked data subset determination module is used to add the target data set and the first output water quality index set corresponding to the target data set to the second marked data subset to obtain the first new-marked data subset at the current iteration number;
[0044] The second new-marked data subset determination module is used to add the target data set and the second output water quality index set corresponding to the target data set to the first marked data subset to obtain the second new-marked data subset at the current iteration number;
[0045] The unlabeled data set determination module at the current iteration number is used to delete the target data set from the unlabeled data set at the previous iteration number to obtain the unlabeled data set at the current iteration number.
[0046] Optionally, the final prediction model is: Where Y(x) represents the final prediction model, H′1(x) represents the first prediction model, and H′2(x) represents the second prediction model.
[0047] An electronic device includes:
[0048] A memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the sewage water quality index soft measurement method according to the above.
[0049] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned soft measurement method for sewage water quality indicators is implemented.
[0050] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0051] The present invention transforms the multi-rate system problem into the problem of imbalance between labeled data and unlabeled data, and integrates two different types of adaptive regression algorithms, RPLS and RELM, to model and train the labeled data, improving the prediction accuracy of the soft measurement model for important indicators in the context of multi-sampling rate data, thereby achieving the monitoring of important variables. At the same time, the model adopts an adaptive method that can effectively overcome the problem of model degradation. Solving the problem of model degradation facilitates maintenance. Description of the Drawings
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0053] Figure 1 It is a flowchart of the soft measurement method for sewage water quality indicators provided by the embodiments of the present invention;
[0054] Figure 2 It is a framework diagram of the soft measurement method for sewage water quality indicators provided by the embodiments of the present invention;
[0055] Figure 3 It is a flowchart of the semi-supervised multi-rate hybrid adaptive model;
[0056] Figure 4 It is a comparison chart of SS regression prediction results;
[0057] Figure 5 It is a comparison chart of SNH regression prediction results;
[0058] Figure 6 It is a comparison chart of SNO regression prediction results;
[0059] Figure 7 It is a comparison chart of COD regression prediction results;
[0060] Figure 8 It is a comparison chart of BOD5 regression prediction results. Detailed Embodiments
[0061] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0063] As Figure 1 and Figure 2 shown, the embodiment of the present invention provides a soft measurement method for sewage water quality indicators, specifically a soft measurement method for sewage water quality indicators based on a semi-supervised multi-rate hybrid adaptive model. This method introduces a semi-supervised regression modeling framework, constructs a semi-supervised multi-rate soft measurement model algorithm, and incorporates an adaptive technology to finally predict important effluent water quality indicators that are not easily measured in the sewage treatment process, improving the model prediction performance and solving the model degradation problem for convenient maintenance. The method includes:
[0064] Using the partial least squares (PLS) algorithm to construct a model based on the first labeled data subset L1, and using the extreme learning machine (ELM) algorithm to construct a model based on the second labeled data subset L2, respectively obtaining the first initial regression model h1 = PLS(L1) and the second initial regression model h2 = ELM(L2); both the first labeled data subset and the second labeled data subset include multiple sets of water quality indicator sets; a set of water quality indicator sets includes an input water quality indicator set and an output water quality indicator set; the input water quality indicator set includes the SS, SNH, and SNO concentrations of the effluent, as well as the important effluent indicators COD and BOD5, and the output water quality indicator set includes the incoming water SS-1, SS-in, SNH-1, SNH-2, SNH-3, SNH-in, SNO-1, SNO-2, SNO-3, SO-1, SO-2, SO-5, Q-tr, Q-in, COD-in.
[0065] Using the first initial regression model h1 to predict the first labeled data subset L1, and using the second initial regression model h2 to predict the second labeled data subset L2, respectively obtaining the first labeled prediction error value R1 and the second labeled prediction error value R2.
[0066] At the current iteration number, the unlabeled dataset at the previous iteration number is respectively input into the first initial regression model and the second initial regression model to obtain a first Plus-labeled data subset and a second Plus-labeled data subset; the unlabeled dataset includes multiple input water quality index sets.
[0067] Use the first initial regression model to predict the first Plus-labeled data subset L1(plus), and use the second initial regression model to predict the second Plus-labeled data subset L2(plus) to obtain a first new label prediction error value R′1 and a second new label prediction error value R′2.
[0068] Obtain the confidence of each input water quality index set in the unlabeled dataset at the previous iteration number according to the first label prediction error value and the first new label prediction error value.
[0069] According to the target dataset x u Update the first labeled data subset, the second labeled data subset, and the unlabeled dataset at the previous iteration number to obtain a first new labeled data subset L1(new) at the current iteration number, a second new labeled data subset L2(new) at the current iteration number, and an unlabeled dataset at the current iteration number; the target dataset is the input water quality index set with the highest confidence.
[0070] Update the iteration number and enter the next iteration. Until the iteration stop condition is reached, use RPLS to construct a first prediction model H1' according to the first new labeled data subset at the last iteration number, and use RELM to construct a second prediction model H'2 according to the second new labeled data subset at the last iteration number. Among them, Recursive Partial Least Squares (RPLS) is based on the original PLS core algorithm. When new data {X1, Y1} arrives, the old data {X, Y} is also used to update the prediction model together. At this time, the data matrices of the two are:
[0071]
[0072] Then multiply the old parameter matrix by the forgetting factor λ (0 < λ ≤ 1) to weaken the influence of the old data on modeling, and then form a new data matrix with the old parameters:
[0073]
[0074] Through the change of the training data X and Y, the T, B, and Q matrices in Equation (4) are also updated. Use RPLS to establish a regression prediction model: H′1 = H′1(x) = RPLS(L1(new)).
[0075] After obtaining the output weight β based on the core algorithm of ELM, the Recursive Extreme Learning Machine (RELM) updates the original data x and y with the new data x new , y new and the forgetting factor λ (0 < λ ≤ 1), that is, x t = [λx, x t , y m = [λy, y t . Finally, the updated x and y are used to learn and train the network structure again to update the weight β. The regression prediction model is established by RELM: H'2 = H'2(x) = RELM(L1(new)).
[0076] The final prediction model is obtained according to the first prediction model and the second prediction model, and the final prediction model is used for soft measurement of sewage water quality indicators.
[0077] In practical applications, a small amount of complete data with the same sampling rate is used as the labeled data L = (X i , Y i ), and the data with inconsistent sampling rates is used as the unlabeled data U(x u ). The labeled data set L is evenly divided into L1 and L2, the unlabeled data set is denoted as U, and the predicted values are calculated using the initial regression models h1 and h2 and
[0078] In practical applications, the unlabeled data sets in the previous iteration are respectively input into the first initial regression model and the second initial regression model to obtain the first Plus labeled data subset and the second Plus labeled data subset, specifically including:
[0079] Input the unlabeled data set in the previous iteration into the first initial regression model to obtain the first output water quality index set corresponding to the unlabeled data set in the previous iteration
[0080] Input the unlabeled data set in the previous iteration into the second initial regression model to obtain the second output water quality index set corresponding to the unlabeled data set in the previous iteration
[0081] Determine the unlabeled data set in the previous iteration and the first output water quality index set corresponding to the unlabeled data set in the previous iteration as the first Plus labeled data subset L1(plus).
[0082] Determine the unlabeled data set at the previous iteration number and the set of second output water quality indicators corresponding to the unlabeled data set at the previous iteration number It is the second Plus-labeled data subset L2(plus).
[0083] In practical applications, the first labeled data subset, the second labeled data subset, and the unlabeled data set at the previous iteration number are updated according to the target data set to obtain the first new labeled data subset, the second new labeled data subset, and the unlabeled data set at the current iteration number. Specifically, it includes
[0084] Input the target first labeled data subset into the first initial regression model to obtain the set of first output water quality indicators corresponding to the target data set
[0085] Input the target second labeled data subset into the second initial regression model to obtain the set of second output water quality indicators corresponding to the target data set
[0086] Add the target data set and the set of first output water quality indicators corresponding to the target data set to the second labeled data subset to obtain the first new labeled data subset at the current iteration number.
[0087] Add the target data set and the set of second output water quality indicators corresponding to the target data set to the first labeled data subset to obtain the second new labeled data subset at the current iteration number.
[0088] Delete the target data set from the unlabeled data set at the previous iteration number to obtain the unlabeled data set at the current iteration number.
[0089] In practical applications, the calculation formulas for R1, R2, R1', and R2' are as follows:
[0090]
[0091]
[0092] In the formula, y i represents x i the true set of output water quality indicators, and h1(x i ) and h2(x i ) represent the set of output water quality indicators predicted by the labeled data set through the regression models PLS and ELM. h1(x u ) and h2(x u ) represent the set of water quality predicted by the unlabeled data set through the regression models PLS and ELM.
[0093] In practical applications, the confidence level ▽ u The specific calculation formula is as follows:
[0094]
[0095] Select the unlabeled data maxx with the largest confidence level u , and the corresponding output regression prediction value As new labeled data, they are respectively put into L1 and L2 to obtain L1(new) and L2(new).
[0096] In practical applications, the final prediction model is: Among them, Y(x) represents the final prediction model, H′1(x) represents the first prediction model, and H′2(x) represents the second prediction model.
[0097] The present invention aims at the performance degradation problem of the soft sensor model and the multi-sampling rate system problem caused by the uneven sampling frequency of the difficult-to-measure variable data in the sewage treatment process. First, the multi-rate system problem is transformed into the problem of unbalanced labeled data and unlabeled data, and then the framework algorithm of semi-supervised co-training is used to mine the useful information hidden in the unlabeled data to expand the labeled data set. In addition, the present invention integrates two different types of adaptive regression algorithms, recursive PLS and recursive ELM, to model and train the labeled data, which can not only enhance the independence of the two groups of regression models, but also improve the diversity of the overall model, so as to solve the modeling problem of weak non-linear data that may be between linear and non-linear. Finally, the online adaptive model updated in real time can effectively overcome problems such as data drift and model degradation and improve the prediction performance of the model.
[0098] The present invention provides a more specific embodiment to introduce the above method in detail as follows:
[0099] This embodiment is a "simulation benchmark model" established by the European Union's Science and Technology Cooperation Organization and the International Water Association based on the No. 1 activated sludge model. The layout of this model is as Figure 3 shown, consisting of a bioreactor (5999m 3 ) mixed by five small units and a secondary sedimentation tank (6000m 3 ). The designed average daily sewage treatment capacity of this equipment is 20000m 3, wastewater containing biodegradable COD of 300 mg / L. The simulation data consists of 14 days of data, including data for sunny days, rainy days, and heavy rain days. There are 1344 sets of data for each weather condition. Fifteen easily measurable variables are selected as input variables from the measurable variables according to the mechanism process flow and expert experience, including influent SS-1, SS-in, SNH-1, SNH-2, SNH-3, SNH-in, SNO-1, SNO-2, SNO-3, SO-1, SO-2, SO-5, Q-tr, Q-in, COD-in. During the experiment, the concentrations of SS, SNH, and SNO in the effluent, as well as the important effluent indicators COD and BOD5, are used as output variables.
[0100] Step 1: Match the multi-sampling rate water quality index data collected by the instrument during the sewage treatment process. The data at the complete sampling rate is used as the labeled data, and the rest is used as the unlabeled data, constituting the labeled data set L and the unlabeled data set U.
[0101] Step 2: Since both the labeled data set and the unlabeled data set have 1344 sets of data, fifteen easily measurable variables are selected as input variables, and five difficult-to-measure effluent water quality indicators are selected as output variables. Therefore, the input data set is The output data set is The labeled data set is divided into L = {(x1, y1), (x2, y2), …, (x 672 , y 672 )}, and the unlabeled data set U = {x 672 ,, x 673 , …, x 1344}.
[0102] Step 3: Divide the labeled data set L = {(x1, y1), (x2, y2), …, (x 672 , y 672 )} into two equal parts: L1 = {(x1, y1), (x2, y2), …, (x 672 , y 672 )}, L2 = {(x 672 , y 672 ), (x 673 , y 673 ), …, (x 1344 , y 1344 )}.
[0103] Step 4: Establish two initial regression models h1 = PLS(L1) and h2 = ELM(L2) using PLS and ELM for L1 and L2 respectively.
[0104] Step 5: Calculate the output variables of the labeled subsets L1 and L2, i.e., the prediction results, through two models h1 = PLS(L1) and h2 = ELM(L2). Compare the two prediction results with the true data results and calculate the prediction error values, denoted as R1 and R2.
[0105] Step 6: Use the initial regression models h1 and h2 to calculate the predicted values for the unlabeled dataset U = {x 672 ,,x 673 ,…,x 1344}, and obtain the first Plus-labeled data subset L and and obtain the first Plus-labeled data subset L and the second Plus-labeled data subset L 1(plus) and the second Plus-labeled data subset L 2(plus) .
[0106] Step 7: Calculate the output variables of the first Plus-labeled data subset L u and the second Plus-labeled data subset L u through the two models h1 = PLS(x 1(plus) ), h2 = ELM(x 1(plus) ), i.e., the prediction results, and calculate the prediction error values, denoted as R′1 and R'2.
[0107] Step 8: Calculate the difference between R and R' as the confidence level and obtain the confidence level of the unlabeled data x u . The specific formula is:
[0108] Step 9: Select the unlabeled data with the maximum confidence level max x u . Calculate the corresponding output regression prediction value of max x u using the PLS and ELM models Put and as new labeled data and cross-put them into L1 and L2 to obtain L 1(new) and L 2(new) , and update the unlabeled dataset U = U - x u .
[0109] Step 10: Repeat the above learning process until the iteration termination condition, i.e., the selection of the specified number of x u is completed. For the finally expanded labeled data subsets L 1(new) and L 2(new) , establish two prediction models H′1(x) and H'2(x) using RPLS and RELM respectively. The final prediction result is determined by the mean of the predicted values of the two models, i.e., Among them, the network parameters of ELM and RELM are specifically set as follows: the number of hidden layer nodes: 45; transfer function: sig; type selection: regression.
[0110] The present invention can effectively predict 5 difficult-to-measure water quality indicators, and the prediction results can be seen Figures 4 to 8 , using two different types of regression models, recursive PLS and recursive ELM, for modeling and prediction, and fully using new data information to update the prediction model, so as to solve the performance degradation problem of the soft sensor model, which provides a strong guarantee for the maintenance of the soft sensor model and is worthy of promotion.
[0111] The embodiment of the present invention also provides a soft sensor system for sewage water quality indicators for the above method, including:
[0112] An initial regression model construction module, configured to construct a model according to the first labeled data subset using the PLS algorithm, and construct a model according to the second labeled data subset using the ELM algorithm, respectively obtaining a first initial regression model and a second initial regression model; both the first labeled data subset and the second labeled data subset include multiple groups of water quality index sets; a group of water quality index sets includes an input water quality index set and an output water quality index set.
[0113] A labeled prediction error value calculation module, configured to predict the first labeled data subset using the first initial regression model, and predict the second labeled data subset using the second initial regression model, respectively obtaining a first labeled prediction error value and a second labeled prediction error value.
[0114] A labeled data subset expansion module, configured to input the unlabeled data set in the previous iteration into the first initial regression model and the second initial regression model respectively at the current iteration number, obtaining a first Plus labeled data subset and a second Plus labeled data subset; the unlabeled data set includes multiple input water quality index sets.
[0115] A new labeled prediction error value calculation module, configured to predict the first Plus labeled data subset using the first initial regression model, and predict the second Plus labeled data subset using the second initial regression model, obtaining a first new labeled prediction error value and a second new labeled prediction error value.
[0116] A confidence level calculation module, configured to obtain the confidence levels of each input water quality index set in the unlabeled data set in the previous iteration according to the first labeled prediction error value and the first new labeled prediction error value.
[0117] An update module, configured to update the first labeled data subset, the second labeled data subset, and the unlabeled data set at the previous iteration number according to the target data set to obtain the first new labeled data subset, the second new labeled data subset, and the unlabeled data set at the current iteration number; the target data set is the set of input water quality indicators with the highest confidence.
[0118] A prediction model construction module, configured to update the iteration number and enter the next iteration until the iteration stop condition is reached, and construct a first prediction model using RPLS based on the first new labeled data subset at the last iteration number, and construct a second prediction model using RELM based on the second new labeled data subset at the last iteration number.
[0119] A final prediction model determination module, configured to obtain a final prediction model according to the first prediction model and the second prediction model, and the final prediction model is used for soft measurement of sewage water quality indicators.
[0120] As an optional implementation manner, the labeled data subset expansion module specifically includes:
[0121] A first output water quality indicator set determination module, configured to input the unlabeled data set at the previous iteration number into the first initial regression model to obtain a first output water quality indicator set corresponding to the unlabeled data set at the previous iteration number.
[0122] A second output water quality indicator set determination module, which inputs the unlabeled data set at the previous iteration number into the second initial regression model to obtain a second output water quality indicator set corresponding to the unlabeled data set at the previous iteration number.
[0123] A first Plus labeled data subset determination module, configured to determine the unlabeled data set at the previous iteration number and the first output water quality indicator set corresponding to the unlabeled data set at the previous iteration number as the first Plus labeled data subset.
[0124] A second Plus labeled data subset determination module, configured to determine the unlabeled data set at the previous iteration number and the second output water quality indicator set corresponding to the unlabeled data set at the previous iteration number as the second Plus labeled data subset.
[0125] As an optional implementation manner, the update module specifically includes:
[0126] A PLS model prediction module, configured to input the target data set into the first initial regression model to obtain a first output water quality indicator set corresponding to the target data set.
[0127] The ELM model prediction module is used to input the target data set into the second initial regression model to obtain the second output water quality index set corresponding to the target data set.
[0128] The first new-labeled data subset determination module is used to add the target data set and the first output water quality index set corresponding to the target data set to the second labeled data subset to obtain the first new-labeled data subset at the current iteration.
[0129] The second new-labeled data subset determination module is used to add the target data set and the second output water quality index set corresponding to the target data set to the first labeled data subset to obtain the second new-labeled data subset at the current iteration.
[0130] The unlabeled data set determination module at the current iteration is used to delete the target data set from the unlabeled data set at the previous iteration to obtain the unlabeled data set at the current iteration.
[0131] As an optional implementation manner, the final prediction model is: Wherein, Y(x) represents the final prediction model, H′1(x) represents the first prediction model, and H′2(x) represents the second prediction model.
[0132] An embodiment of the present invention also provides an electronic device, including:
[0133] A memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the sewage water quality index soft measurement method according to the above method embodiment.
[0134] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the sewage water quality index soft measurement method described in the above method embodiment.
[0135] The beneficial effects of the present invention are as follows:
[0136] The present invention can effectively extract the important features of multi-sampling rate data without losing the correlation between data, establish a soft measurement model algorithm in a multi-sampling rate complex environment, overcome the degradation problem of the soft measurement model under the influence of complex environment, consider the maintenance and adaptive ability of the model when designing the soft measurement model, perform adaptive correction and update on the soft measurement model algorithm, and can effectively overcome problems such as data drift and model degradation and improve the prediction performance of the model.
[0137] The present invention not only uses the framework algorithm of semi-supervised co-training to mine the useful information hidden in incomplete data under a large number of inconsistent sampling rates to solve the multi-rate system problem, but also the co-training algorithm of the present invention will simultaneously use two different types of adaptive regression algorithms to model and train the labeled data, which not only enhances the independence of the two groups of regression models, but also improves the diversity (linear and non-linear) of the overall model, so as to solve the modeling problem of weak non-linear data that may be between linear and non-linear.
[0138] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method part.
[0139] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A soft measurement method for sewage water quality indicators, characterized in that Including: Using the PLS algorithm to construct a model based on the first labeled data subset, and using the ELM algorithm to construct a model based on the second labeled data subset, respectively obtaining a first initial regression model and a second initial regression model; Both the first labeled data subset and the second labeled data subset include multiple sets of water quality index sets; one set of water quality index sets includes an input water quality index set and an output water quality index set; Using the first initial regression model to predict the first labeled data subset, and using the second initial regression model to predict the second labeled data subset, respectively obtaining a first labeled prediction error value and a second labeled prediction error value; At the current iteration number, input the unlabeled data set at the previous iteration number into the first initial regression model and the second initial regression model respectively to obtain a first Plus labeled data subset and a second Plus labeled data subset; The unlabeled data set includes multiple input water quality index sets; Using the first initial regression model to predict the first Plus labeled data subset, and using the second initial regression model to predict the second Plus labeled data subset, obtaining a first new labeled prediction error value and a second new labeled prediction error value; Obtaining the confidence degrees of each input water quality index set in the unlabeled data set at the previous iteration number according to the first labeled prediction error value and the first new labeled prediction error value; Updating the first labeled data subset, the second labeled data subset, and the unlabeled data set at the previous iteration number according to the target data set to obtain a first new labeled data subset at the current iteration number, a second new labeled data subset at the current iteration number, and an unlabeled data set at the current iteration number; the target data set is the input water quality index set with the maximum confidence degree; Updating the iteration number and entering the next iteration until the iteration stop condition is reached. Using RPLS to construct a first prediction model based on the first new labeled data subset at the last iteration number, and using RELM to construct a second prediction model based on the second new labeled data subset at the last iteration number; Obtaining a final prediction model according to the first prediction model and the second prediction model, and the final prediction model is used for soft measurement of sewage water quality indexes.
2. The soft sensor method for sewage water quality index according to claim 1, characterized in that Inputting the unlabeled data set at the previous iteration number into the first initial regression model and the second initial regression model respectively to obtain a first Plus labeled data subset and a second Plus labeled data subset, specifically including: Inputting the unlabeled data set at the previous iteration number into the first initial regression model to obtain a first output water quality index set corresponding to the unlabeled data set at the previous iteration number; Inputting the unlabeled data set at the previous iteration number into the second initial regression model to obtain a second output water quality index set corresponding to the unlabeled data set at the previous iteration number; Determining the unlabeled data set at the previous iteration number and the first output water quality index set corresponding to the unlabeled data set at the previous iteration number as the first Plus labeled data subset; Determine the unlabeled data set at the previous iteration number and the set of second output water quality indicators corresponding to the unlabeled data set at the previous iteration number as the second Plus-labeled data subset.
3. The soft measurement method for sewage water quality indicators according to claim 1, characterized in that, Update the first labeled data subset, the second labeled data subset, and the unlabeled data set at the previous iteration number according to the target data set to obtain the first new labeled data subset, the second new labeled data subset, and the unlabeled data set at the current iteration number. Specifically, it includes: Input the target data set into the first initial regression model to obtain the set of first output water quality indicators corresponding to the target data set; Input the target data set into the second initial regression model to obtain the set of second output water quality indicators corresponding to the target data set; Add the target data set and the set of first output water quality indicators corresponding to the target data set to the second labeled data subset to obtain the first new labeled data subset at the current iteration number; Add the target data set and the set of second output water quality indicators corresponding to the target data set to the first labeled data subset to obtain the second new labeled data subset at the current iteration number; Delete the target data set from the unlabeled data set at the previous iteration number to obtain the unlabeled data set at the current iteration number.
4. The soft sensor method for sewage water quality indexes according to claim 1, characterized in that The final prediction model is as follows: Where Y(x) represents the final prediction model, H′1(x) represents the first prediction model, and H′2(x) represents the second prediction model.
5. A soft sensor system for sewage water quality indicators, characterized in that, It includes: An initial regression model construction module for constructing a model according to the first labeled data subset using the PLS algorithm and constructing a model according to the second labeled data subset using the ELM algorithm to obtain the first initial regression model and the second initial regression model respectively; Both the first labeled data subset and the second labeled data subset include multiple sets of water quality indicator sets; a set of water quality indicator sets includes an input water quality indicator set and an output water quality indicator set; A labeled prediction error value calculation module for predicting the first labeled data subset using the first initial regression model and predicting the second labeled data subset using the second initial regression model to obtain the first labeled prediction error value and the second labeled prediction error value respectively; A labeled data subset expansion module for inputting the unlabeled data set at the previous iteration number into the first initial regression model and the second initial regression model respectively at the current iteration number to obtain the first Plus-labeled data subset and the second Plus-labeled data subset; The unlabeled data set includes multiple input water quality indicator sets; A new labeled prediction error value calculation module for predicting the first Plus-labeled data subset using the first initial regression model and predicting the second Plus-labeled data subset using the second initial regression model to obtain the first new labeled prediction error value and the second new labeled prediction error value; A confidence level calculation module for obtaining the confidence levels of each input water quality indicator set in the unlabeled data set at the previous iteration number according to the first labeled prediction error value and the first new labeled prediction error value; An update module, configured to update the first labeled data subset, the second labeled data subset, and the unlabeled data set at the previous iteration number according to the target data set to obtain the first new labeled data subset, the second new labeled data subset, and the unlabeled data set at the current iteration number; the target data set is the input water quality index set with the maximum confidence. A prediction model construction module, configured to update the iteration number and enter the next iteration until the iteration stop condition is reached, and construct a first prediction model using RPLS based on the first new labeled data subset at the last iteration number, and construct a second prediction model using RELM based on the second new labeled data subset at the last iteration number. A final prediction model determination module, configured to obtain a final prediction model according to the first prediction model and the second prediction model, and the final prediction model is used for soft sensing of sewage water quality indexes.
6. The soft sensor system for sewage water quality indicators according to claim 5, characterized in that The labeled data subset expansion module specifically includes: A first output water quality index set determination module, configured to input the unlabeled data set at the previous iteration number into the first initial regression model to obtain the first output water quality index set corresponding to the unlabeled data set at the previous iteration number. A second output water quality index set determination module, configured to input the unlabeled data set at the previous iteration number into the second initial regression model to obtain the second output water quality index set corresponding to the unlabeled data set at the previous iteration number. A first Plus labeled data subset determination module, configured to determine the unlabeled data set at the previous iteration number and the first output water quality index set corresponding to the unlabeled data set at the previous iteration number as the first Plus labeled data subset. A second Plus labeled data subset determination module, configured to determine the unlabeled data set at the previous iteration number and the second output water quality index set corresponding to the unlabeled data set at the previous iteration number as the second Plus labeled data subset.
7. The soft sensor system for sewage water quality indexes according to claim 5, characterized in that, The update module specifically includes: A PLS model prediction module, configured to input the target data set into the first initial regression model to obtain the first output water quality index set corresponding to the target data set. An ELM model prediction module, configured to input the target data set into the second initial regression model to obtain the second output water quality index set corresponding to the target data set. A first new labeled data subset determination module, configured to add the target data set and the first output water quality index set corresponding to the target data set to the second labeled data subset to obtain the first new labeled data subset at the current iteration number. A second new labeled data subset determination module, configured to add the target data set and the second output water quality index set corresponding to the target data set to the first labeled data subset to obtain the second new labeled data subset at the current iteration number. An unlabeled data set determination module at the current iteration number, configured to delete the target data set from the unlabeled data set at the previous iteration number to obtain the unlabeled data set at the current iteration number.
8. The soft sensor system for sewage water quality indicators according to claim 5, characterized in that, The final prediction model is as follows: Among them, Y(x) represents the final prediction model, H′1(x) represents the first prediction model, and H′2(x) represents the second prediction model.
9. An electronic device, characterized in that, Including: A memory and a processor, the memory is used for storing a computer program, and the processor runs the computer program to enable the electronic device to execute the soft measurement method for sewage water quality indexes according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it implements the soft measurement method for sewage water quality indexes according to any one of claims 1 to 4.
Citation Information
Patent Citations
Sewage monitoring multi-output soft measurement method based on semi-supervised learning
CN112381221A
Industrial process soft measurement modeling method based on semi-supervised ensemble learning
CN112989711A