Sewage treatment system fault diagnosis method based on Borderline SMOTE-LASSONET

The BorderlineSMOTE algorithm is used to balance the sewage treatment data set and combine the LASSONET algorithm to screen key features, which solves the problems of data imbalance and insufficient model interpretability in the existing technology, and improves the accuracy and efficiency of fault diagnosis of sewage treatment system.

CN120067745APending Publication Date: 2025-05-30SHANGHAI INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510039346.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing wastewater treatment system fault diagnosis methods are difficult to effectively deal with highly unbalanced data sets, resulting in the bias of the results toward most categories and lack of methods to improve model interpretability.

Method used

The BorderlineSMOTE algorithm is used to interpolate a few types of samples to generate new samples to achieve the data balance state, and model it in combination with the LASSONET algorithm to filter out the most critical feature set for fault diagnosis.

Benefits of technology

Through data balance and feature screening, the accuracy of fault diagnosis of sewage treatment system and the interpretability of the model are improved, and the detection basis and troubleshooting efficiency of on-site staff are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067745A_ABST
    Figure CN120067745A_ABST
Patent Text Reader

Abstract

The invention relates to a sewage treatment system fault diagnosis method based on Borderline SMOTE-LASSONET. The method comprises the following steps: 1) complementing missing values in sewage data and standardizing the data; 2) expanding minority class data to a balanced state by adopting a Borderline SMOTE algorithm, and generating a new balanced data set; 3) performing LASSONET modeling, and screening out an input feature set and a hyper-parameter corresponding to a state with the highest model accuracy rate; 4) selecting a state with the highest model accuracy rate to verify the test set; and 5) obtaining a final fault diagnosis device. According to the method, unbalanced sewage data are effectively processed, the feature set with the largest contribution to fault diagnosis is screened out, the method has guiding significance on targeted feature detection of field staff, and meanwhile, the fault diagnosis accuracy is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sewage treatment system management, and in particular to a sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET. Background Art

[0002] Sewage treatment plants play a key role in environmental protection and sustainable utilization of water resources. However, sewage treatment is a complex, multi-variable dynamic process with numerous influencing factors, making it difficult to ensure long-term stable operation. When a system failure occurs, quickly and accurately determining the fault location can help on-site staff promptly identify problems and avoid serious consequences. Intelligent algorithms have given new possibilities to the field of sewage treatment system fault diagnosis.

[0003] Chinese Patent CN105740619A discloses a "sewage treatment online fault diagnosis method based on a kernel function weighted extreme learning machine", but it does not consider that the sewage treatment data set is a highly imbalanced set, which easily leads to results being biased towards the majority class. Patent CN109558893A discloses a "fast integrated sewage treatment fault diagnosis method based on a resampling pool". This invention overcomes the problem of highly imbalanced data to a certain extent using the SMOTE algorithm. The SMOTE algorithm randomly interpolates between a minority class sample and its nearest neighbor minority class sample, without considering enhancing the representativeness of minority class samples at the boundary of the majority class. Additionally, in the face of a large number of sewage treatment process variables, improving the interpretability of the model, that is, finding certain features that play an important role in faults, and then enabling on-site staff to focus on such features, has very important practical significance. So far, there are few literatures recorded in the field of sewage treatment system fault diagnosis. Summary of the Invention

[0004] The object of the present invention is to overcome the defects existing in the above-mentioned prior art and provide a sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET. A sewage treatment system fault diagnosis method is proposed. The sewage treatment system fault diagnosis method is based on BorderlineSMOTE-LASSONET. The present invention proposes an algorithm fusion and synergy enhancement method, which organically combines the BorderlineSMOTE algorithm that can solve the data imbalance problem with the LASSONET algorithm with strong fitting ability and interpretability. It can not only overcome the above problems, but also enhance the applicability and stability of the algorithm. The BordelineSMOTE algorithm interpolates new samples for the minority class samples on the boundary of the majority class, so that the number of minority class samples reaches the desired balance state, reduces the degree of sample imbalance, and enhances the representativeness of minority class samples on the decision boundary. The LASSONET algorithm not only inherits the strong fitting ability of the neural network to ensure the classification accuracy, but also takes into account the generalization and interpretability of the LASSO algorithm, and is particularly suitable for the fault diagnosis problem of the sewage treatment system.

[0005] The present invention can be realized through the following technical solutions:

[0006] The technical solutions adopted by the present invention are as follows:

[0007] A sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET includes the following steps:

[0008] 1) Preprocess the sewage treatment data set, complete the missing values in the sewage data and standardize the data;

[0009] 2) Use the BorderlineSMOTE algorithm to expand the minority class data in the standardized sewage treatment data set to a balanced state and generate a new balanced state data set;

[0010] 3) Perform LASSONET modeling, construct a LASSONET model, and screen out the input feature set and hyperparameters corresponding to the state with the highest model accuracy;

[0011] 4) Select the state with the highest model accuracy to verify the test set;

[0012] 5) Obtain the final fault diagnoser: Repeat the step of completing the missing values of the sewage treatment data in step 1), input the test set data to be measured into the trained LASSONET model, and the fault category of the data to be measured can be output.

[0013] Furthermore, it specifically includes the following steps:

[0014] 1) Preprocess the sewage treatment dataset Nold with N features, complete the missing values in the sewage treatment data, generate a new dataset Nnew, and standardize the data;

[0015] 2) Use the BorderlineSMOTE algorithm to expand the minority class sample size until the desired balanced state is reached, obtain the balanced state dataset Nbal, and divide the balanced state dataset Nbal into a training set and a test set;

[0016] 3) Perform LASSONET modeling to construct a LASSONET model;

[0017] 4) Use the training set data to train the LASSONET model to obtain a trained LASSONET model;

[0018] 5) Select the feature set with the highest accuracy in the LASSONET model as the optimal feature set for fault diagnosis;

[0019] 6) Select the model state corresponding to the optimal feature set for fault diagnosis and diagnose the test set;

[0020] 7) Repeat the step of completing the missing values in the sewage treatment data in step 1), input the test set data to be measured into the trained LASSONET model, and the fault category of the data to be measured can be output.

[0021] Furthermore, the missing value completion algorithm used to complete the missing values in the sewage treatment data is the random forest algorithm.

[0022] Furthermore, the model state includes an input feature set and hyperparameters; the hyperparameters include λ, M, and also include Lasso regression parameters, the number of hidden layer neurons in the neural network, the learning rate, the number of iterations, etc. Among them, λ is used to control the sparsity of the features.

[0023] Furthermore, the sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET includes the following steps:

[0024] Step S1: Preprocess the sewage treatment dataset Nold with N features, complete the missing values in the sewage treatment data, generate a new dataset Nnew, and standardize the data.

[0025] Step S2: Use the BorderlineSMOTE algorithm to expand the minority class sample size until the desired balanced state is reached, obtain the balanced state dataset Nbal, and divide the data (balanced state dataset Nbal) into a training set and a test set.

[0026] Step S3: Perform LASSONET modeling and initialize the hyperparameters related to the LASSONET model.

[0027] Step S4: Use the training set data to train the model and generate the training paths of the model, which are used for subsequent model saving and evaluation.

[0028] Step S5: Initialize the number of features, model accuracy, regularization strength parameter, and the list of selected feature names.

[0029] Step S6: Use each state in the model training path of Step S4 for diagnosis, and record the parameters such as the number of features, model accuracy, regularization strength parameter, and the selected feature names.

[0030] Step S7: In Step S6, the feature set with the highest model accuracy is the optimal feature set for fault diagnosis, that is, the feature combination included in the set contributes the most to the fault diagnosis of this sewage treatment system and is key monitored in engineering.

[0031] Step S8: Select the model state (containing corresponding hyperparameters such as λ) corresponding to the optimal feature set for fault diagnosis in the path, diagnose the test set, and calculate and print parameters such as accuracy, precision of each fault state to confirm the final performance of the LASSONET model.

[0032] Step S9: Repeat Step S1 to fill in the missing values of the data to be measured, input the data to be measured into the trained LASSONET model, and the fault category of the data to be measured can be output.

[0033] Furthermore, for the method of filling in the missing values of Nold in Step S1, the random forest algorithm is adopted, and the specific steps are as follows:

[0034] S1.1: Take the feature T with missing values as the label, and form a new feature matrix with the other n - 1 features and the original label.

[0035] S1.2: Take the non - missing part of T as y_train, and the corresponding other n - 1 features and the original label y as x_train; take the missing part of T as y_test, and the corresponding other n - 1 features and the original label y as x_test.

[0036] S1.3: Traverse all features and sort the features with missing values according to the number of missing values.

[0037] S1.4: Start filling from the feature with the fewest missing values, temporarily use 0 to replace the missing values of other features, repeat S1.1 - S1.2, and after each regression prediction, put the predicted value into the original feature matrix, and then continue to fill the next feature.

[0038] The normalization formula described in step S1 is as follows:

[0039]

[0040] where X is the original data, and X scaled is the normalized data, μ is the mean of the feature, and σ is the standard deviation of the feature.

[0041] The BorderlineSMOTE algorithm described in step S2 includes the following steps:

[0042] S2.1 In the dataset Nnew, there are p classes of minority classes J1, J2... Jp. Randomly select a reference sample in the class Jp, calculate the Euclidean distance between the reference sample and all samples in the entire dataset Nnew, and find the n nearest neighbors of the reference sample.

[0043] S2.2 If the number m of the majority samples among these n nearest neighbors satisfies n / 2 < m < n, then we define this sample as a Borderline sample.

[0044] S2.3 For each Borderline sample, randomly select a minority class sample original_sample from the n nearest neighbors of the Borderline sample, generate a new sample using the following formula, and put it into the dataset of this class

[0045] new_sample = original_sample + α × (selected_neighbor - original_sample)(2)

[0046] α is a random number between 0 and 1, used for interpolation between the original sample and the selected neighbor. This parameter controls the position of the new sample between the original sample and the neighbor sample, and selected_neigbbor is the nearest neighbor.

[0047] S2.4 Repeat the above process (steps S2.1 - S2.3) to perform sample expansion on all minority class samples until the number of minority samples reaches the desired balanced state.

[0048] The objective function of the lassonet model described in S3 is:

[0049]

[0050] where Loss(θ, W; X, y) is the cross - entropy loss function, used to measure the model prediction error, and its expression is defined as:

[0051]

[0052] It is the L1 regularization term. By imposing a sparsity constraint on θ, it prompts the model to only select the features that have a significant impact on fault diagnosis.

[0053] In formula (3), θ is the linear coefficient directly associated with the input features and output features, W is the weight matrix of the neural network hidden layer, X is the input matrix, that is, the sewage treatment data set (balanced state data set Nbal), Y is the target variable, that is, various fault state categories of the sewage treatment system. N is the total number of samples, and K is the total number of fault categories. The λ of the regularization term is the regularization strength parameter, which controls the degree of sparsity.

[0054] For the objective function, there are the following constraint conditions:

[0055]

[0056] It represents the j-th column of the weight matrix from the input layer to the first hidden layer, corresponding to the input feature x j The weights of the connections between all neurons in the hidden layer, θ j is the input feature x j The corresponding linear regression weight, M is a positive hyperparameter used to control the strength of regularization. This constraint condition means that if the linear weight of x j is very small or even 0, then the maximum value of its corresponding W j (0) must also be correspondingly small to achieve the purpose of feature selection.

[0057] Compared with the prior art, the present method has the following obvious characteristics and advantages:

[0058] 1. The present invention proposes the BorderlineSMOTE algorithm for the problem of unbalanced sewage treatment process data, interpolates the boundary samples to generate new samples, enriches the number of minority class samples in the data set, makes the number of minority class samples reach the desired balanced state, reduces the degree of sample imbalance, enhances the representativeness of minority class samples on the decision boundary, and further enhances the accuracy of the model.

[0059] 2. The present invention introduces the LASSO jump layer into the feedforward neural network, ranks the importance of different feature combinations through sparse feature selection, so as to screen out the most critical feature set for fault diagnosis. Compared with the traditional method, this not only improves the interpretability of the model, but also provides targeted detection basis for on-site staff, significantly improving the efficiency of fault troubleshooting. By automatically selecting the optimal feature set, it is not necessary to collect and process all the feature data in the sewage treatment process, thus reducing the data collection cost and resource consumption, and optimizing the resource utilization while ensuring the diagnostic effect.

[0060] 3. The fault state of the sewage treatment process is diagnosed by the BorderlineSMOTE-LASSONET algorithm of the present invention, which improves the accuracy of fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is the LASSONET network structure used in the present invention.

[0062] Figure 2 It is the fault diagnosis process of the fault diagnosis method for the sewage treatment system of the present invention.

[0063] Figure 3 It is the confusion matrix of the test results using the algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The description of at least one exemplary embodiment below is actually only illustrative and in no way restrictive of the present invention and its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0065] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0066] Features such as component models, material names, connection structures, control methods, etc. not clearly described in the present invention are regarded as common technical features disclosed in the prior art.

[0067] The design of the present invention relates to a fault diagnosis method for a sewage treatment system based on BorderlineSMOTE-LASSONET, including the steps of: 1) Completing the missing values in the sewage data and normalizing the data; 2) Using the BorderlineSMOTE algorithm to expand the minority class data to a balanced state and generating a new balanced data set; 3) Conducting LASSONET modeling and screening out the input feature set and hyperparameters corresponding to the state with the highest model accuracy; 4) Selecting the state with the highest model accuracy to verify the test set; 5) Obtaining the final fault diagnoser. The present invention effectively processes unbalanced sewage data, screens out the feature set that contributes the most to fault diagnosis, has guiding significance for on-site workers to detect features targeted, and at the same time, the present invention also improves the accuracy of fault diagnosis.

[0068] A fault diagnosis method for a sewage treatment system based on BorderlineSMOTE-LASSONET, the method comprising the following steps:

[0069] Step S1, preprocessing the sewage treatment data set Nold containing N features, completing the missing values of the sewage treatment data, generating a new data set Nnew, and normalizing the data.

[0070] Step S2, using the BorderlineSMOTE algorithm to expand the minority class sample size until reaching the desired balanced state, obtaining a balanced state data set Nbal, and dividing the data (balanced state data set Nbal) into a training set and a test set.

[0071] Step S3, conducting LASSONET modeling and initializing the hyperparameters related to the LASSONET model.

[0072] Step S4, using the training set data to train the model and generating the training paths of the model, which are used for subsequent model saving and evaluation.

[0073] Step S5, initializing the number of features, model accuracy, regularization strength parameter, and the list of selected feature names.

[0074] Step S6, using each state in the model training path of Step S4 for diagnosis, and recording these parameters such as the number of features, model accuracy, regularization strength parameter, and the selected feature names.

[0075] In Step S6, the feature set with the highest model accuracy is the optimal feature set for fault diagnosis, that is, the feature combination included in the set contributes the most to the fault diagnosis of this sewage treatment system and is key monitored in the project.

[0076] Step S8: Select the model state corresponding to the optimal feature set for fault diagnosis in the path (including hyperparameters such as corresponding λ), perform diagnosis on the test set, calculate and print parameters such as accuracy, precision for each fault state, etc. to confirm the final performance of the LASSONET model.

[0077] Step S9: Repeat Step S1 to fill in the missing values of the data to be measured, input the data to be measured into the trained LASSONET model, and the fault category of the data to be measured can be output.

[0078] Furthermore, for the method of filling in the missing values of Nold in Step S1, the random forest algorithm is adopted, and the specific steps are as follows:

[0079] S1.1: Take the feature T with missing values as the label, and form a new feature matrix with the other n - 1 features and the original label.

[0080] S1.2: Take the non - missing part of T as y_train, and the corresponding other n - 1 features and the original label y as x_train; take the missing part of T as y_test, and the corresponding other n - 1 features and the original label y as x_test.

[0081] S1.3: Traverse all features, and sort the features with missing values according to the number of missing values.

[0082] S1.4: Start filling from the feature with the fewest missing values, temporarily replace the missing values of other features with 0, repeat S1.1 - S1.2, and after each regression prediction, put the predicted value into the original feature matrix, and then continue to fill the next feature.

[0083] The standardization formula described in Step S1 is:

[0084]

[0085] where X is the original data, X scaled is the standardized data, μ is the mean of the feature, and σ is the standard deviation of the feature.

[0086] The BorderlineSMOTE algorithm described in Step S2 includes the following steps:

[0087] S2.1: There are p types of minority classes J1, J2... Jp in the dataset Nnew. Randomly select a reference sample in the class Jp, calculate the Euclidean distance between the reference sample and all samples in the entire dataset Nnew, and find the n nearest neighbors of the reference sample.

[0088] S2.2: If the number m of the majority samples among these n nearest neighbors satisfies n / 2 < m < n, then we define this sample as a Borderline sample.

[0089] S2.3 For each Borderline sample, randomly select a minority class sample original_sample from the n nearest neighbors of the Borderline sample, generate a new sample using the following formula, and put it into the dataset of this class

[0090] new_sample = original_sample + α × (selected_neighbor - original_sample)(2)

[0091] α is a random number between 0 and 1, used to interpolate between the original sample and the selected neighbor. This parameter controls the position of the new sample between the original sample and the neighbor sample. selected_neighbor is the nearest neighbor.

[0092] S2.4 Repeat the above process (steps S2.1 - S2.3) to perform sample expansion on all minority class samples until the number of minority samples reaches the desired balanced state.

[0093] For the lassonet model described in S3, its objective function is:[[]]

[0094]

[0095] where Loss(θ, W; X, y) is the cross - entropy loss function, used to measure the model prediction error, and its expression is defined as:[[]]

[0096]

[0097] is the L1 regularization term. By imposing a sparsity constraint on θ, it prompts the model to only select features that have a significant impact on fault diagnosis.

[0098] In formula (3), θ is the linear coefficient directly associated with the input features and output features, W is the weight matrix of the neural network hidden layer, X is the input matrix, that is, the sewage treatment dataset, Y is the target variable, that is, various fault status categories of the sewage treatment system. N is the total number of samples, and K is the total number of fault categories. The λ of the regularization term is the regularization strength parameter, which controls the degree of sparsity.

[0099] For the objective function, there are the following constraint conditions:[[]]

[0100]

[0101] represents the j - th column of the weight matrix from the input layer to the first hidden layer, corresponding to the input feature x jThe weights θ of the connections between all neurons in the hidden layer j is the input feature x j The corresponding linear regression weight, M is a positive hyperparameter used to control the strength of regularization. This constraint means that if the linear weight of x j is very small or even zero, then the corresponding maximum value of W j (0) must also be correspondingly small to achieve the purpose of feature selection.

[0102] Embodiment

[0103] As Figure 2 shown, this embodiment provides a fault diagnosis method for a sewage treatment system based on BorderlineSMOTE-LASSONET. The data in this embodiment comes from the sewage treatment plant data in the University of California database (UCI). There are a total of 527 groups of this data, including 38 feature variables, and a total of 13 categories including the normal state. For the sake of simplicity of description, the state names will be replaced by numbers below, and the distribution is shown in Table 1.

[0104] Table 1 Distribution of the original categories of 527 samples

[0105]

[0106] In view of its practical significance, some states are merged into one category, a total of 4 major categories, as shown in Table 2.

[0107] Table 2 Distribution of the merged categories of 527 groups of data

[0108] Category 1 2 3 4 Quantity 332 116 69 14 Category before merging 1、11 5 9 2、3、4、6、7、8、10、12、13

[0109] Category 1 is the normal state of the sewage treatment system, category 2 is the state with better performance of the sewage treatment system, category 3 is the state where the water treatment system is normal but the influent volume is small, and category 4 is the abnormal state such as the failure of the secondary sedimentation tank (secondary clarifier) and heavy rain weather. Category 1 is regarded as the majority sample, and categories 2, 3, and 4 are all regarded as minority classes.

[0110] Due to the complex situation on the sewage treatment site and the existence of conditions such as damaged detection equipment, there are missing data in this data set. For such a situation, the traditional approach is to delete the samples where the missing values are located, or fill them with the mean value of this feature. If the sample size is large enough, the former can be simply selected, which is obviously not applicable to this data set, and the use of the latter may introduce bias, thus affecting the subsequent fault diagnosis. The present invention uses the random forest method to fill in the missing values, and by combining multiple decision trees, captures the complex non-linear relationships in the data, thereby providing a more accurate estimate of the missing values.

[0111] The specific steps of the BorderlineSMOTE-LASSONET algorithm adopted by the present invention in this embodiment are as follows:

[0112] 1) Complete and standardize the missing values in the original data through the random forest algorithm described in step S1;

[0113] 2) Through the BorderlineSMOTE algorithm described in step S2, expand the sample size of the three types of minority classes until the desired balanced state is reached, obtain the balanced state dataset Nbal, divide the balanced state dataset Nbal into a training set and a test set, and use a 7-3 split. The data comparison before and after expansion is shown in Table 3:

[0114] Table 3 Data comparison before and after expansion

[0115] Category 1 2 3 4 Quantity before expansion 332 116 69 14 Quantity before expansion 332 332 332 332

[0116] 3) As Figure 1 shown, establish a LASSONET model to sparsify the features, predict each state in the training path, screen out the features that have little effect on the diagnostic result based on the overall accuracy rate, sort the prediction process in descending order of the accuracy rate (model accuracy rate), find the feature set with the best classification effect, export the lassonet traversal results, and part of the results are shown in Table 4:

[0117] Table 4 Partial results of lassonet traversal

[0118]

[0119]

[0120] It can be seen from Table 4 that when lambda = 1.43 and the feature variables are sparsified from 38 to 16, the model fault diagnosis accuracy rate is the highest. Therefore, the feature set included in this group is the optimal feature set (optimal feature set for fault diagnosis). Select this model state to verify the test set. Repeat filling the missing values of the data to be measured, input the data to be measured into the trained LASSONET model, and the fault category of the data to be measured can be output.

[0121] In order to more comprehensively evaluate the LASSONET model, draw the confusion matrix of the multi-classification task, and calculate the precision pre and the geometric mean G-mean of each class. As Figure 3 shown, it is the confusion matrix of the simulation results of this embodiment.

[0122] For the N-classification problem, the confusion matrix is an N×N matrix, where each row represents the actual class and each column represents the predicted class. The definition of the confusion matrix is shown in Table 5 below, TP n: Judging a positive sample as a positive class, FP nn : Judging a negative sample as a positive class or a positive class as a negative class, such as TP 2 That is, the actual class is class 2 and the model prediction is also class 2, which means the prediction is correct; TP 21 The actual class is class 2 and the prediction is class 1, TP 12 The actual class is class 1 and the prediction is class 2, prediction error:

[0123] Definition of the confusion matrix in Table 5

[0124] Actual / Predicted Category 1 Category 2 …… Category n Category 1 <![CDATA[TP 1 > <![CDATA[FP 12 > …… <![CDATA[FP 1n > Category 2 <![CDATA[FP 21 > <![CDATA[TP 2 > …… <![CDATA[FP 2n > …… …… …… …… …… Category n <![CDATA[FP n1 > <![CDATA[FP n2 > …… <![CDATA[TP n >

[0125] Accuracy represents the proportion of the number of samples with correct model predictions in the total number of samples. It is an overall measurement index that reflects the performance of the model on the entire dataset. The formula is:

[0126]

[0127] where ConfusionMatrix ij represents the element in the i-th row and j-th column of the confusion matrix.

[0128] Precision represents the proportion of samples actually belonging to a certain class among those classified as that class. It is more suitable for evaluating the classification effect of a single class rather than the overall performance. The formula is:

[0129]

[0130] G-Mean is the geometric mean of the recall rates of all classes and is particularly suitable for the classification evaluation of imbalanced samples. It represents the balance of the classification effects of the model on each class. The higher the value, the better the model performs on all classes.

[0131] Defined as:

[0132]

[0133] where Recall is the recall rate, and its definition is:

[0134]

[0135] Select the random forest model, Adaboost model, and support vector machine model for comparison. The results are shown in Table 6:

[0136] Table 6 Comparison of the random forest model, Adaboost model, support vector machine model with the LASSONET of the present application

[0137] Acc R1pre R2pre R3pre R4pre G-Mean LASSONET 0.951 0.969 0.898 0.967 0.979 0.952 Random Forest 0.925 0.926 0.876 0.906 0.934 0.927 Adaboost 0.779 0.829 0.597 0.822 0.934 0.765 SVM 0.907 0.751 0.976 0.759 0.862 0.905

[0138] In Table 6, ACC is the overall accuracy of the model, R1pre - R4pre are the corresponding precision rates of the four categories respectively, and G - mean is the geometric mean of the model. As can be seen from the table, although the precision rate of the SVM model LASSONET for the second category is higher, in terms of the overall accuracy, G - mean and the precision rates of the other three categories, the performance of the LASSONET model exceeds that of the random forest, Adaboost model and support vector machine model. Thus, it can be proved that the algorithm adopted in the present invention is suitable for the fault diagnosis problem of unbalanced data in the sewage treatment process.

[0139] Comparative example

[0140] To better illustrate the effect of BorderlineSMOTE on the processing of unbalanced data, in this comparative example, the unbalanced data is not processed and the SMOTE algorithm is used to process the unbalanced data respectively, and the remaining steps still adopt the LASSONET modeling scheme of the implementation case. The comparison results are shown in Table 7:

[0141] Table 7 Comparison of the processing of unbalanced data without processing, SMOTE algorithm for processing unbalanced data, and BorderlineSMOTE of the present application for processing unbalanced data

[0142]

[0143] As can be seen from Table 7, the accuracy of not processing the unbalanced data is lower than that of the SMOTE and BorderlineSMOTE algorithms. Only the precision rate of the second category is relatively high. However, the G - Mean value of the model trained with the data processed by the BorderlineSMOTE algorithm is the highest. The results show that the BorderlineSMOTE algorithm has advantages in solving the problem of unbalanced sewage treatment data.

[0144] The above - mentioned embodiments are the implementation manners with better effects of the present invention. However, the implementation manners of the present invention are not limited by the above - mentioned embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.

[0145] The above description of the embodiments is for the convenience of those of ordinary skill in the art to understand and use the invention. Those familiar with the technology can obviously make various modifications to these embodiments easily and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above - mentioned embodiments, and the improvements and modifications made by those skilled in the art without departing from the scope of the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET, characterized in that: It includes the following steps: 1) Preprocess the sewage treatment dataset, complete the missing values in the sewage data and standardize the data; 2) Use the BorderlineSMOTE algorithm to expand the minority class data in the standardized sewage treatment dataset to a balanced state, generating a new balanced dataset; 3) Conduct LASSONET modeling, construct the LASSONET model, and select the input feature set and hyperparameters corresponding to the state with the highest model accuracy; 4) Select the state with the highest model accuracy to verify the test set; 5) Repeat the step of completing the missing values in the sewage treatment data in step 1), input the test set data to be measured into the trained LASSONET model, and the fault category of the data to be measured can be output.

2. According to claim 1, a sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET is characterized in that: In step 1), the missing value completion algorithm for completing the missing values in the sewage treatment data is the random forest algorithm.

3. A sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 2, characterized in that: The specific process of the missing value completion algorithm in step 1) includes the following steps: 1.1) Use the feature T with missing values as the label, and the other n - 1 features and the original label to form a new feature matrix; 1.2) Use the non - missing part of T as y_train, and the corresponding other n - 1 features and the original label y as x_train; use the missing part of T as y_test, and the corresponding other n - 1 features and the original label y as x_test; 1.3) Traverse all features, and sort the features with missing values according to the number of missing values; 1.4) Start filling from the feature with the fewest missing values, temporarily use 0 to replace the missing values of other features, repeat 1.1) - 1.2), and after each regression prediction, put the predicted value into the original feature matrix, and then continue to fill the next feature.

4. The sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 1 is characterized in that: The specific formula for the data standardization method of standardizing the data in step 1) is as follows: Among them, X is the original data, X scaled is the standardized data, μ is the mean of the feature, and σ is the standard deviation of the feature.

5. The sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 1 is characterized in that: The specific process of the BorderlineSMOTE algorithm in step 2) includes the following steps: 2.1) There are p classes of minority classes J1, J2... Jp in the dataset Nnew. Randomly select a reference sample in the class Jp, calculate the Euclidean distance between the reference sample and all samples in the dataset Nnew, and find the n nearest neighbors of the reference sample; 2.2) If the number m of the majority samples among the n nearest neighbors satisfies n / 2 < m < n, then define this sample as a Borderline sample; 2.3) For each Borderline sample, randomly select a minority class sample original_sample from the n nearest neighbors of the Borderline sample, and generate a new sample new_sample using the following formula and put it into the dataset of this class new_sample = original_sample+α×(selected_neighbor - original_sample) (2) where α is a random number between 0 and 1, and selected_neighbor is the nearest neighbor; 2.4) Repeat steps 2.1) to 2.3) to expand the samples of all minority classes until the number of minority samples reaches the desired equilibrium state, and obtain the equilibrium state data set Nbal.

6. A sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 1, characterized in that: The objective function of the LASSONET algorithm in step 3) is: Among them, Loss(θ, W; X, y) is the cross entropy loss function, which is used to measure the model prediction error, and its expression is defined as: is the L1 regularization term, which imposes a sparsity constraint on θ to force the model to select only features that have a significant impact on fault diagnosis; In formula (3), θ is the linear coefficient directly related to the input feature and the output feature, W is the weight matrix of the hidden layer of the neural network, X is the input matrix, i.e., the sewage treatment data set, Y is the target variable, i.e., the various fault status categories of the sewage treatment system, N is the total number of samples, K is the total number of fault categories, and λ of the regularization term is the regularization strength parameter, which controls the degree of sparsity.

7. A sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 6, characterized in that: For the objective function of the LASSONET algorithm, there are the following constraints: W j (0) Represents the jth column of the weight matrix from the input layer to the first hidden layer, which corresponds to the input feature x j The weights of the connections between all neurons in the hidden layer, θ j is the input feature x j The corresponding linear regression weight, M is a positive hyperparameter used to control the strength of regularization.

8. The sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 1 is characterized in that: In step 3), the process of selecting the input feature set and hyperparameters corresponding to the state with the highest model accuracy includes the following steps: 4.1) Initialize the hyperparameters related to the LASSONET model; 4.2) Use the training set data to train the model and generate the training path of the model; 4.3) Initialize the number of features, model accuracy, regularization strength parameter, and a list of selected feature names; 4.4) Use each model state in the training path of the model in step 4.2) for diagnosis, and record the number of features, model accuracy, regularization strength parameter, and selected feature names.

9. The sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 1, characterized in that: Step 4) specifically includes the following process: The feature set with the highest model accuracy is the optimal feature set for fault diagnosis. The model state corresponding to the optimal feature set for fault diagnosis is selected, the test set is diagnosed, and the accuracy and precision of each fault state are calculated and printed to confirm the final performance of the LASSONET model.

10. The sewage treatment system fault diagnosis method based on BorderlineSMOTE-LASSONET according to claim 1, characterized in that: In step 3), the hyperparameters include λ and M.

Citation Information

Patent Citations

  • On-line fault diagnosis method of weighted extreme learning machine sewage treatment on the basis of kernel function

    CN105740619A

  • A rapid integrated sewage treatment fault diagnosis method based on a resampling pool

    CN109558893A