Sewage biological treatment effect prediction method based on machine learning

By identifying abnormal states in the wastewater treatment model and supplementing and optimizing the data accordingly, the problem of insufficient flexibility in responding to anomalies during actual operation was solved, achieving higher prediction accuracy and stability.

CN120823892APending Publication Date: 2025-10-21HUNAN XIANDAO YANGHU RECLAIMED WATER CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510921148.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

In existing technologies, sewage treatment models lack the ability to flexibly respond to and precisely regulate abnormal monitoring and diagnosis during actual operation, resulting in poor prediction accuracy and stability.

Method used

The model is identified by using the fluctuation coefficient and the maximum deviation value. The abnormal state is determined by the load anomaly coefficient and the load characteristic correlation threshold. Load data or characteristic parameter correlation is supplemented, the data resolution of the characteristic interval is adjusted, multi-dimensional optimization is performed, and abnormal hyperparameters are adjusted to improve the adaptability and accuracy of the model.

Benefits of technology

It effectively improves the model's predictive ability and stability, enabling it to more accurately reflect the wastewater treatment process, adapt to complex data changes, and enhance the model's reliability and predictive performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823892A_ABST
    Figure CN120823892A_ABST
Patent Text Reader

Abstract

The invention relates to the field of sewage treatment, in particular to a sewage biological treatment effect prediction method based on machine learning, which comprises the following steps: determining whether a target machine learning model is abnormal or not according to a fluctuation coefficient and a maximum deviation value; if the target machine learning model is abnormal, determining an abnormal state according to a load abnormal coefficient and a load feature correlation threshold; according to the abnormal state, it is determined that load data supplementation is carried out, or feature parameter association supplementation is carried out, or the data resolution of a feature interval is adjusted, or multi-dimensional optimization is carried out; determining an initial interval according to the feature representation degree and the paragraph comparison degree, and determining whether to adjust the initial interval to obtain a feature interval according to an initial interval distribution coefficient and an interval influence degree; and performing multi-dimensional optimization under a preset condition. The accuracy degree of model prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sewage treatment, and in particular to a method for predicting the effect of sewage biological treatment based on machine learning. Background Art

[0002] With the continuous improvement of sewage treatment requirements and the need for refined operation and management of sewage treatment plants, it has become an important research direction to use models to predict the operating status of secondary biochemical treatment processes such as AAO, SBR, and MSBR and optimize the operation strategy. However, in actual applications, due to the complexity of the sewage treatment process and the poor quality and representativeness of training data, there is often a certain error between the model prediction results and the actual measurement values. Therefore, how to improve the accuracy of model prediction is an urgent problem to be solved by technical personnel in this field.

[0003] Chinese patent publication number CN119416066A discloses a sewage treatment effect prediction method and system based on deep learning. The method includes: data collection, data preprocessing, treatment effect prediction model construction, hyperparameter optimization and sewage treatment effect prediction. The original sewage treatment data is obtained through data collection; data preprocessing methods such as data cleaning, data labeling, data encoding, data normalization and data set segmentation are adopted; an improved spatiotemporal attention graph convolution model is used to predict sewage treatment effects, which can effectively extract key time series features in time series data and the complex spatial relationships involved in the sewage treatment process; a hunting optimization algorithm is used to optimize the model hyperparameters, which can conduct targeted exploration of the parameter space when processing sewage treatment data with multiple variables and complex spatiotemporal characteristics. However, the above scheme has the following problems: it does not fully consider the abnormal monitoring and diagnosis and dynamic optimization adjustment of the model in actual operation, which makes the model lack the ability to flexibly respond to abnormal conditions and accurately control them, resulting in poor accuracy and stability of model predictions. Summary of the Invention

[0004] To this end, the present invention provides a method for predicting the effect of sewage biological treatment based on machine learning, which is used to overcome the problem that the existing technology does not fully consider the abnormal monitoring and diagnosis and dynamic optimization adjustment of the model in actual operation, resulting in the model lacking the ability to flexibly respond to abnormal conditions and accurately control, resulting in poor accuracy and stability of model prediction.

[0005] To achieve the above objectives, the present invention provides a method for predicting the effect of sewage biological treatment based on machine learning, comprising:

[0006] Determine whether the target machine learning model has anomalies based on the fluctuation coefficient and maximum deviation value;

[0007] If the target machine learning model has an anomaly, the abnormal state is determined based on the load anomaly coefficient and the load feature association threshold;

[0008] Supplement load data according to abnormal conditions, or supplement characteristic parameters or adjust data resolution of characteristic intervals, or perform multi-dimensional optimization;

[0009] Determine the initial interval based on the feature representativeness and paragraph comparison, and determine whether to adjust the initial interval to obtain the feature interval based on the initial interval distribution coefficient and interval influence;

[0010] Under the preset conditions, multi-dimensional optimization is performed, among which,

[0011] Determine a correlation combination based on an instability parameter correlation degree or a characteristic interval correlation threshold based on an instability parameter characteristic value and an instability parameter proportion, determine an instability correlation combination based on an abnormal reference value, and adjust the data volume of the instability correlation combination based on the abnormal reference value;

[0012] Under the condition of unstable data volume adjustment, abnormal hyperparameters are determined according to the complexity comparison degree, and each abnormal hyperparameter is adjusted according to the deviation coefficient.

[0013] Furthermore, whether the target machine learning model has anomalies is determined based on the fluctuation coefficient and the maximum deviation value, including:

[0014] If the fluctuation coefficient is greater than or equal to the preset fluctuation coefficient or the maximum deviation value is greater than or equal to the preset maximum deviation value, it is determined that the target machine learning model has an anomaly;

[0015] If the fluctuation coefficient is less than the preset fluctuation coefficient and the maximum deviation value is less than the preset maximum deviation value, it is determined that there is no abnormality in the target machine learning model.

[0016] Furthermore, if the abnormal state is a first abnormal state in which the load abnormality coefficient is greater than or equal to a preset load abnormality coefficient, the load data is supplemented.

[0017] Furthermore, if the abnormal state is a second abnormal state in which the load abnormality coefficient is less than a preset load abnormality coefficient and the load characteristic association threshold is less than a preset load characteristic association threshold, characteristic parameter association supplementation or adjustment of the data resolution of the characteristic interval is performed according to the comparison reference value;

[0018] If the comparison reference value is greater than or equal to the preset comparison reference value, the data resolution of the feature interval is reduced;

[0019] If the comparison reference value is less than the preset comparison reference value, the characteristic parameter association supplement is performed.

[0020] Furthermore, the method for confirming the characteristic interval includes:

[0021] Determine the initial interval based on feature representativeness and paragraph comparison;

[0022] If the initial interval distribution coefficient is greater than or equal to the preset initial interval distribution coefficient or the interval influence is less than the preset interval influence, the initial interval is recorded as a characteristic interval;

[0023] If the initial interval distribution coefficient is less than the preset initial interval distribution coefficient and the interval influence is greater than or equal to the preset interval influence, the increased and adjusted initial interval is recorded as the characteristic interval.

[0024] Furthermore, under preset conditions, the associated combination is determined based on the characteristic value of the instability parameter and the proportion of the instability parameter, including:

[0025] If the instability parameter characteristic value is less than the preset instability parameter characteristic value and the instability parameter proportion is less than the preset instability parameter proportion, then the correlation combination is determined according to the instability parameter correlation degree;

[0026] If the instability parameter characteristic value is greater than or equal to the preset instability parameter characteristic value or the instability parameter ratio is greater than or equal to the preset instability parameter ratio, then the associated combination is determined according to the characteristic interval association threshold;

[0027] The preset condition is the third abnormal state or the adjustment abnormal state.

[0028] Furthermore, the association combination whose abnormal reference value is greater than the preset abnormal reference value is recorded as an unstable association combination, and the data volume of the unstable association combination is increased and adjusted based on the abnormal reference value;

[0029] The increase in the data volume of a single unstable association combination is positively correlated with the abnormal reference value corresponding to the unstable association combination.

[0030] Furthermore, under the condition of unstable data volume regulation, abnormal hyperparameters are determined based on the complexity comparison, including:

[0031] If the complexity comparison degree is less than the preset complexity comparison degree, the abnormal hyperparameter is determined based on the regular comparison degree;

[0032] If the complexity comparison degree is greater than or equal to the preset complexity comparison degree, the abnormal hyperparameter is determined based on the floating reference value.

[0033] Furthermore, each abnormal hyperparameter is adjusted according to the deviation coefficient;

[0034] The adjustment amount corresponding to a single abnormal hyperparameter is positively correlated with the deviation coefficient corresponding to the abnormal hyperparameter.

[0035] Furthermore, the average value of the deviation values ​​corresponding to each prediction result of the target machine learning model within a preset time period is compared with the standard deviation to obtain the output reference value, and the regularity comparison degree is determined based on the output reference value and the hyperparameter comparison degree.

[0036] Compared with the prior art, the beneficial effect of the present invention lies in that, in the technical solution of the present invention, the fluctuation coefficient and the maximum deviation value can effectively reflect the large deviation or unstable fluctuation that may occur in the model during the prediction process, ensuring that the prediction results of the model can accurately reflect the actual state of the sewage treatment process, and then when there is an abnormality in the target machine learning model, the abnormal state is determined according to the load abnormality coefficient and the load feature association threshold, and the load abnormality coefficient and the load feature association threshold are used to effectively reflect the degree of data missing and the degree of data characteristics, and then according to the abnormal state, the feature parameter association is adaptively supplemented or the data resolution of the feature interval is improved or multi-dimensional optimization is performed, which can solve the specific problems of the model in a targeted manner and effectively improve the performance and prediction ability of the model. The feature parameter association supplement can enrich the feature information of the data, so that the model can more comprehensively capture the key features in the data; improving the data resolution of the feature interval can more finely reflect the change law of the data, which helps the model to make more accurate predictions; multi-dimensional optimization can improve the model from multiple aspects to make it better adapt to complex sewage treatment data.

[0037] Furthermore, the present invention determines the initial interval through feature representativeness and paragraph comparison, which can automatically identify the most representative and relevant parts of the data, improving the efficiency and objectivity of data processing. The initial interval is adjusted in combination with the initial interval distribution coefficient and interval influence, which can further optimize the selection of feature intervals, ensure that the feature intervals can accurately reflect the main characteristics and changing trends of the data, and provide more valuable information for model prediction.

[0038] Furthermore, the present invention performs multi-dimensional optimization under preset conditions, and can improve the model from multiple aspects at the same time, such as adjusting the amount of data of unstable associated combinations, adjusting hyperparameters, etc., so as to more comprehensively improve the performance and predictive ability of the model, so that it can better adapt to the complex sewage treatment process. Adjusting the amount of data of unstable associated combinations according to abnormal reference values ​​helps to solve the instability problem of the model in a targeted manner, improve the reliability and stability of the model, and judge whether it is necessary to determine abnormal hyperparameters according to regular comparison degrees or floating reference values ​​through complex comparison degrees. It can adapt to the complex changes and uncertainties of sewage treatment data, timely discover and adjust abnormal hyperparameters in the model, and maintain the predictive performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Schematic diagram of the method for predicting the effect of biological sewage treatment based on machine learning of the present invention;

[0040] Figure 2 This is a flow chart of the present invention for determining whether a target machine learning model has an anomaly based on the fluctuation coefficient and the maximum deviation value;

[0041] Figure 3 This is a flow chart of the present invention for determining load data supplementation or characteristic parameter association supplementation or adjusting data resolution or multi-dimensional optimization of characteristic intervals according to abnormal conditions;

[0042] Figure 4 The flowchart of the present invention is to determine abnormal hyperparameters according to the complex comparison degree. DETAILED DESCRIPTION

[0043] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0044] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0045] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0046] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0047] See also Figures 1 to 4 As shown, the present invention provides a method for predicting the effect of sewage biological treatment based on machine learning, comprising:

[0048] Determine whether the target machine learning model has anomalies based on the fluctuation coefficient and maximum deviation value;

[0049] If the target machine learning model has an anomaly, the abnormal state is determined based on the load anomaly coefficient and the load feature association threshold;

[0050] Supplement load data according to abnormal conditions, or supplement characteristic parameters or adjust data resolution of characteristic intervals, or perform multi-dimensional optimization;

[0051] Determine the initial interval based on the feature representativeness and paragraph comparison, and determine whether to adjust the initial interval to obtain the feature interval based on the initial interval distribution coefficient and interval influence;

[0052] Under the preset conditions, multi-dimensional optimization is performed, among which,

[0053] Determine a correlation combination based on an instability parameter correlation degree or a characteristic interval correlation threshold based on an instability parameter characteristic value and an instability parameter proportion, determine an instability correlation combination based on an abnormal reference value, and adjust the data volume of the instability correlation combination based on the abnormal reference value;

[0054] Under the condition of unstable data volume adjustment, abnormal hyperparameters are determined according to the complexity comparison degree, and each abnormal hyperparameter is adjusted according to the deviation coefficient.

[0055] The application scenario of the present invention is to use the target machine learning model to predict the biological treatment effect of the secondary biochemical treatment process of the sewage treatment plant. The secondary biochemical treatment process includes but is not limited to AAO, SBR and MSBR. The present invention specifically takes MSBR as an example. The process of predicting the biological treatment effect of other secondary biochemical treatment processes is the same as that of MSBR, and the details are not repeated here. In the present invention, several historical records are correspondingly set up, and any historical record records the fluctuation coefficient, maximum deviation value, comparison reference value, deviation value, load in the historical process of predicting the biological treatment effect of the MSBR of the sewage treatment plant using the target machine learning model at least once. Abnormal coefficient, load characteristic correlation threshold and comparison reference value, etc., and each historical record corresponds to a qualified mark. The qualified mark records whether the process of using the target machine learning model to predict the biological treatment effect of the MSBR of the sewage treatment plant meets the user's needs. The qualified mark can be recorded manually. It is understandable that the user can determine whether the process of using the target machine learning model to predict the biological treatment effect of the MSBR of the sewage treatment plant meets the needs based on the self-set indicators. The self-set indicators can be but not limited to the number of errors, which will not be elaborated here. The number of errors is the number of times there is an error between the prediction result of the target machine learning model and the actual measurement value;

[0056] The target machine learning model takes the influent water quality indicators and process parameters collected in real time in the MSBR of the sewage treatment plant as input and the predicted effluent water quality indicators as output;

[0057] Influent water quality indicators and effluent water quality indicators include but are not limited to COD, ammonia nitrogen, total nitrogen and total phosphorus. Process parameters include but are not limited to sludge concentration, aeration volume and water temperature. Water quality indicators are measured by a multi-parameter water quality detector, sludge concentration is measured by a sludge concentration sensor, aeration volume is measured by a flow meter, and water temperature is measured by a temperature sensor. This is easy to understand for those skilled in the art and will not be described in detail.

[0058] The present invention is provided with a target coefficient and a related threshold value, and the corresponding relationship between the target coefficient and the related threshold value is expressed by a weight formula, and the weight formula is: target coefficient = weight coefficient × related threshold value. Specifically, the present invention records the increase in the amount of load data, the decrease in the data resolution of the feature interval, the number of load data supplemented by the feature parameters, the value of w, the increase in the amount of data of the unstable associated combination, and the adjustment amount corresponding to the abnormal hyperparameter as the target coefficient, and records the load abnormality coefficient, the comparison reference value, the feature associated reference value, the analysis time, the abnormal reference value, and the deviation coefficient corresponding to the abnormal hyperparameter as the related threshold value. It can be understood that the target coefficients all have corresponding related threshold values, for example, the load The increase in data volume is positively correlated with the load anomaly coefficient, and the positive correlation between the increase in load data volume and the load anomaly coefficient is expressed through a weight formula. The value of the weight coefficient can be determined based on the user's historical experience according to the degree of influence of the load anomaly coefficient on the increase in load data volume, and the value of the weight coefficient can be optimized by combining the historical records of multiple uses of the target machine learning model to predict the biological treatment effect of the MSBR of the sewage treatment plant with the multi-layer perceptron. The value of the weight coefficient optimized by the multi-layer perceptron is easy for technical personnel in this field to understand, and will not be elaborated on. The value principles of the weight coefficients corresponding to other target coefficients and related thresholds are the same, and will not be elaborated on here.

[0059] Specifically, the target machine learning model is determined to have abnormalities based on the fluctuation coefficient and the maximum deviation value, including:

[0060] If the fluctuation coefficient is greater than or equal to the preset fluctuation coefficient or the maximum deviation value is greater than or equal to the preset maximum deviation value, it is determined that the target machine learning model has an anomaly;

[0061] If the fluctuation coefficient is less than the preset fluctuation coefficient and the maximum deviation value is less than the preset maximum deviation value, it is determined that there is no abnormality in the target machine learning model.

[0062] Among them, n predictions are made for the target machine learning model, and the corresponding input conditions such as influent water quality indicators and process parameters in each prediction are kept consistent, that is, for the same parameters, the input data values ​​are the same, and the standard deviation of the deviation values ​​corresponding to each prediction result is recorded as the fluctuation coefficient; the value of n can be determined by the user according to actual needs. The greater the user's demand for improving the accuracy of the fluctuation coefficient determination, the larger the value of n is. A value of n is provided, and n is 20;

[0063] The influent water quality indicators and process parameters collected in real time in the MSBR of the sewage treatment plant are used as model inputs, and the target machine learning model is predicted in real time within a preset time period. The input data for each prediction is based on the influent water quality and process parameters collected in real time. The maximum value of the deviation values ​​corresponding to the prediction results of the target machine learning model within the preset time period is recorded as the maximum deviation value, and the deviation value corresponding to the single prediction result is the maximum value of the sub-deviation values ​​corresponding to each effluent water quality indicator. For a single effluent water quality indicator, the value of the effluent water quality indicator corresponding to the single prediction result of the target machine learning model is recorded as a1, and the value of the effluent water quality indicator actually measured is recorded as a2. The sub-deviation value corresponding to the single effluent water quality indicator = |a1-a2| / (the larger value of a1 and a2). The value of the preset time period can be determined by the user according to actual needs. The greater the user's demand for improving the accuracy of the maximum deviation value determination, the longer the value of the preset time period is. A value of the preset time period is provided, and the preset time period is 48h.

[0064] The values ​​of the preset fluctuation coefficient and the preset maximum deviation value can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of the prediction results, the smaller the values ​​of the preset fluctuation coefficient and the preset maximum deviation value. A method for determining the values ​​of the preset fluctuation coefficient and the preset maximum deviation value is provided to detect and determine the existence of abnormal historical records of the target machine learning model. The average value of the fluctuation coefficient and the average value of the maximum deviation values ​​corresponding to the historical records that can meet the user's needs are recorded as the preset fluctuation coefficient and the preset maximum deviation value, respectively.

[0065] Specifically, if the abnormal state is a first abnormal state in which the load abnormality coefficient is greater than or equal to a preset load abnormality coefficient, the load data is supplemented.

[0066] The abnormal state includes a first abnormal state, a second abnormal state and a third abnormal state. The first abnormal state is that the load abnormality coefficient is greater than or equal to the preset load abnormality coefficient, the second abnormal state is that the load abnormality coefficient is less than the preset load abnormality coefficient and the load characteristic association threshold is less than the preset load characteristic association threshold, and the third abnormal state is that the load abnormality coefficient is less than the preset load abnormality coefficient and the load characteristic association threshold is greater than or equal to the preset load characteristic association threshold;

[0067] The present invention includes several historical operation data, and a single historical operation data includes the numerical values ​​corresponding to each monitoring parameter in the MSBR monitored in real time during the sewage treatment process. The collection position of each monitoring parameter of the single historical operation data can be determined by the user according to actual needs, and there is no specific restriction. The monitoring parameters include water quality indicators and process parameters. The target machine learning model in the present invention corresponds to several training data, and a single training data corresponds to a historical operation data. The single training data is the numerical value corresponding to the monitoring parameter corresponding to each time point in the single historical training data. The number of training data in the present invention is less than the number of historical operation data. The starting time of the single historical operation data is used as the starting point, and an interval point is set every 10 minutes until the sewage treatment is completed. The starting point and each interval point are recorded as time points; the value of each water quality index corresponding to the starting time of the single historical operation data is the influent water quality index, and the value of each water quality index corresponding to the latest time of the single historical operation data is the effluent water quality index;

[0068] For a single inlet water quality indicator in a single training data set, if the comparison reference value corresponding to the inlet water quality indicator is greater than the preset comparison reference value corresponding to the inlet water quality indicator, then the water quality indicator is a load water quality indicator; the comparison reference value corresponding to a single inlet water quality indicator = (the value of the inlet water quality indicator in the single training data - the average value of the values ​​of the inlet water quality indicator corresponding to each historical operation data) / the average value of the values ​​of the inlet water quality indicator corresponding to each historical operation data;

[0069] The value of the preset comparison reference value corresponding to a single inlet water quality index can be determined by the user according to the actual application scenario. The smaller the value of the preset comparison reference value, the greater the user's demand for determining the inlet water quality index as a load index. A value of the preset comparison reference value is provided, and load data in which the inlet water quality index is determined as a load water quality index in the historical records that can meet the user's needs is detected. The minimum value of the comparison reference values ​​corresponding to the inlet water quality index in each load data is recorded as the preset comparison reference value corresponding to the inlet water quality index;

[0070] The training data with loaded water quality indicators are recorded as loaded data;

[0071] Among the prediction results obtained by performing real-time prediction on the target machine learning model within a preset time period, the prediction results with a deviation value greater than the preset deviation value are recorded as abnormal prediction results. The value of the preset deviation value can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of model prediction, the smaller the value of the preset deviation value is. A value of the preset deviation value is provided, and the minimum value of the deviation values ​​corresponding to the abnormal prediction results in the historical records that can meet the user's needs is recorded as the preset deviation value;

[0072] Detect the inlet water quality index corresponding to each abnormal prediction result, and record the abnormal prediction result with load water quality index as load abnormal result;

[0073] Load anomaly coefficient = [number of load anomaly results / (number of prediction results with load water quality indicators + 1)] × [1-(number of load data / total amount of training data)]. It can be understood that the load anomaly coefficient comprehensively reflects the abnormal proportion of the target machine learning model in predicting the effluent water quality related to the load water quality indicators and the impact of the missing load data in the training data on the prediction accuracy, reflecting the reliability and stability of the target machine learning model in predicting the load water quality indicators;

[0074] The load feature association threshold is the average value of the load association degrees corresponding to each load data. For a single load data, the load data is recorded as the target load data, and the load data other than the target load data is recorded as the reference load data. The load association degree corresponding to the target load data is the average value of the sub-load association degrees corresponding to the target load data and each reference load data.

[0075] The sub-load correlation degree corresponding to any two load data is the average value of the sub-characteristic load correlation degrees corresponding to each monitoring parameter;

[0076] For a single monitoring parameter of any two load data, the monitoring parameter is recorded as the target monitoring parameter. The calculation formula of the sub-feature load correlation r corresponding to the target monitoring parameter is:

[0077]

[0078] Where m is the number of time points in the load data with a smaller number of time points; x k and y k are the values ​​of the target monitoring parameters corresponding to the kth time point in the two load data, is x k The average value of the target monitoring parameter corresponding to the first m time points in the corresponding load data, y k The average value of the target monitoring parameter corresponding to the first m time points in the corresponding load data, k = 1, 2, 3, ..., m;

[0079] It can be understood that the load feature association threshold reflects the similarity between the load data in terms of monitoring parameters;

[0080] The values ​​of the preset load anomaly coefficient and the preset load characteristic association threshold can be determined by the user according to the actual application scenario. The smaller the value of the preset load anomaly coefficient, the greater the user's demand for load data supplementation. The larger the value of the preset load characteristic association threshold, the greater the user's demand for characteristic parameter association supplementation or adjustment of data resolution of characteristic interval according to comparison with reference value. A value of a preset load anomaly coefficient and a preset load characteristic association threshold is provided, and the historical records of load data supplementation are detected. The average value of the load anomaly coefficient corresponding to the historical records that can meet the user's needs is recorded as the preset load anomaly coefficient. The historical records of characteristic parameter association supplementation or adjustment of data resolution of characteristic interval according to comparison with reference value are detected, and the average value of the load characteristic association threshold corresponding to the historical records that can meet the user's needs is recorded as the preset load characteristic association threshold.

[0081] The load data supplementation includes: increasing and adjusting the load data volume according to the load anomaly coefficient, wherein the increase value of the load data volume is positively correlated with the load anomaly coefficient; the load data volume is the total amount of load data in the training data.

[0082] Specifically, if the abnormal state is the second abnormal state in which the load abnormality coefficient is less than the preset load abnormality coefficient and the load characteristic association threshold is less than the preset load characteristic association threshold, the characteristic parameter association is supplemented or the data resolution of the characteristic interval is adjusted according to the comparison reference value;

[0083] If the comparison reference value is greater than or equal to the preset comparison reference value, the data resolution of the feature interval is reduced;

[0084] If the comparison reference value is less than the preset comparison reference value, the characteristic parameter association supplement is performed.

[0085] Wherein, the comparison reference value = (load characteristic correlation threshold - historical characteristic correlation threshold) × the average value of the floating deviation values ​​corresponding to each load data;

[0086] The historical feature correlation threshold is the average of the operation correlations of the historical operation data corresponding to each load data. For a single historical operation data, the historical operation data is recorded as the target historical operation data, and the other historical operation data other than the target historical operation data is recorded as the reference historical operation data. The operation correlation corresponding to the target historical operation data is the average of the sub-operation correlations corresponding to the target historical operation data and each reference historical operation data.

[0087] The sub-operation correlation degree corresponding to any two load data is the average value of the sub-operation characteristic correlation degrees corresponding to each monitoring parameter;

[0088] For a single monitoring parameter of any two load data, the monitoring parameter is recorded as the target monitoring parameter. The calculation formula of the sub-operation characteristic correlation degree r0 corresponding to the target monitoring parameter is:

[0089]

[0090] Where m0 is the number of target monitoring parameter values ​​in the smaller amount of historical operation data; x k0 and y k0 are the values ​​of the k0th target monitoring parameters in the two historical operation data, is x k0 The average value of the first m0 target monitoring parameters in the corresponding historical operation data, y k0 The average value of the first m0 target monitoring parameters in the corresponding historical operation data, k0 = 1, 2, 3, ..., m0;

[0091] The floating deviation value corresponding to a single load data is the maximum value among the sub-floating deviation values ​​corresponding to each monitoring parameter in the load data. The sub-floating deviation value corresponding to a single monitoring parameter = the standard deviation of the value corresponding to the monitoring parameter actually monitored in the historical operation data corresponding to the load data - the standard deviation of the value of the monitoring parameter corresponding to each time point in the load data. It can be understood that the comparison reference value reflects the difference in correlation between the load data and the historical operation data, and, combined with the fluctuation characteristics of the monitoring parameters, reflects the degree of fluctuation difference between the historical operation data and the load data of each monitoring parameter in the sewage treatment process.

[0092] The value of the preset comparison reference value can be determined by the user according to the actual application scenario. The larger the value of the preset comparison reference value, the greater the user's demand for feature parameter association supplementation. A method for determining the value of the preset comparison reference value is provided, which detects the historical records of feature parameter association supplementation and records the average value of the comparison reference values ​​corresponding to the historical records that can meet the user's needs as the preset comparison reference value;

[0093] When reducing the data resolution of the feature interval, the data resolution of the feature interval of each training data is reduced, and the reduction value of the data resolution of the feature interval is positively correlated with the comparison reference value;

[0094] The data resolution of the feature interval is the time interval between any two adjacent time points in the feature interval;

[0095] The characteristic parameter association supplementation includes: recording a monitoring parameter whose characteristic association reference value is greater than a preset characteristic association reference value as a characteristic parameter; when supplementing a single characteristic parameter, the number of load data supplemented for the single characteristic parameter is positively correlated with the characteristic association reference value corresponding to the characteristic parameter; when supplementing a single characteristic parameter, selecting historical operation data in ascending order of the association reference value with the characteristic parameter; the association reference value of a single historical operation data with the characteristic parameter is an average value of the sub-operation characteristic association degree of the characteristic parameter corresponding to the historical operation data corresponding to each load data;

[0096] For a single monitoring parameter, the monitoring parameter is recorded as the target parameter. The characteristic correlation reference value corresponding to the target parameter is the average value of the correlation thresholds corresponding to each load data of the target parameter. The correlation threshold corresponding to the target parameter in a single load data is the average value of the sub-characteristic load correlation degrees corresponding to the load data and other load data.

[0097] The value of the preset feature association reference value can be determined by the user according to the actual application scenario. The greater the user's demand for improving the prediction accuracy, the smaller the value of the preset feature association reference value. A method for determining the preset feature association reference value is provided, and the average value of the feature association reference value corresponding to each feature parameter in the historical record that can meet the user's needs is recorded as the preset feature association reference value.

[0098] Specifically, the method for confirming the characteristic interval includes:

[0099] Determine the initial interval based on feature representativeness and paragraph comparison;

[0100] If the initial interval distribution coefficient is greater than or equal to the preset initial interval distribution coefficient or the interval influence is less than the preset interval influence, the initial interval is recorded as a characteristic interval;

[0101] If the initial interval distribution coefficient is less than the preset initial interval distribution coefficient and the interval influence is greater than or equal to the preset interval influence, the increased and adjusted initial interval is recorded as the characteristic interval.

[0102] The sewage treatment time of the historical operation data corresponding to each training data is detected, and the minimum sewage treatment time is recorded as the analysis time. Each data value within the analysis time is selected from the historical operation data corresponding to each training data, and the analysis time is divided into w equal parts to obtain several time intervals. The value of w is positively correlated with the analysis time.

[0103] The feature representativeness corresponding to a single time interval is the average value of the sub-feature representativeness of the time interval in the historical running data corresponding to each training data;

[0104] The representativeness of the sub-feature corresponding to a single time interval in a single historical operation data is the maximum value of the interval reference values ​​corresponding to each monitoring parameter in the historical operation data; the interval reference value corresponding to a single monitoring parameter is the standard deviation of each value corresponding to the monitoring parameter in the single historical operation data in the time interval;

[0105] The paragraph comparison degree corresponding to a single time interval = the feature representativeness corresponding to the time interval - the comparison feature representativeness corresponding to the time interval;

[0106] The representativeness of the comparison feature corresponding to a single time interval is the average value of the representativeness of the sub-comparison features corresponding to the time interval in each training data; the representativeness of the sub-comparison feature corresponding to a single time interval in a single training data is the maximum value of the comparison interval reference value corresponding to each monitoring parameter in the training data; the comparison interval reference value corresponding to a single monitoring parameter is the standard deviation of each value of the monitoring parameter corresponding to each time point in the time interval in the single training data;

[0107] When determining the initial interval based on the feature representativeness and the paragraph comparison degree, the time interval in which the feature representativeness is greater than the preset feature representativeness or the paragraph comparison degree is greater than the preset paragraph comparison degree is recorded as the initial interval;

[0108] The values ​​of the preset feature representativeness and the preset paragraph comparison degree can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of model prediction, the smaller the values ​​of the preset feature representativeness and the preset paragraph comparison degree. A method for determining the values ​​of the preset feature representativeness and the preset paragraph comparison degree is provided. The average value of the feature representativeness and the average value of the paragraph comparison degree corresponding to each initial interval in the historical records that can meet the user's needs are detected, and the values ​​are recorded as the preset feature representativeness and the preset paragraph comparison degree respectively.

[0109] The time between the earliest moment of the earliest initial interval and the latest moment of the latest initial interval is recorded as the analysis time segment. The initial interval distribution coefficient = the length of the analysis time segment / the number of initial intervals;

[0110] The other time intervals outside the initial interval are recorded as non-initial intervals;

[0111] The interval impact is the maximum value of the abnormal impact corresponding to each non-initial interval in the analysis time segment;

[0112] The abnormal impact degree corresponding to a single non-initial interval is the average of the slope differences between the non-initial interval and each initial interval;

[0113] The slope difference between a single non-initial interval and a single initial interval = |slope reference value corresponding to the non-initial interval - slope reference value corresponding to the initial interval|. The slope reference value corresponding to a single time interval is the average of the sub-slope reference values ​​of each monitoring parameter in that time zone. The sub-slope reference value of a single monitoring parameter in that time zone = (the value of the monitoring parameter corresponding to the latest moment in the time interval - the value of the monitoring parameter corresponding to the earliest moment in the time interval) / the duration of the time interval.

[0114] The values ​​of the preset initial interval distribution coefficient and the preset interval influence can be determined by the user according to the actual application scenario. The larger the value of the preset initial interval distribution coefficient and the smaller the value of the preset interval influence, the greater the user's demand for recording the increased adjusted initial interval as a characteristic interval. A value of the preset initial interval distribution coefficient and the preset interval influence is provided, and a historical record of the user recording the increased adjusted initial interval as a characteristic interval is detected. The average value of the initial interval distribution coefficient and the average value of the interval influence corresponding to the historical records that can meet the user's needs are recorded as the preset initial interval distribution coefficient and the preset interval influence, respectively.

[0115] When the initial interval after the increase adjustment is recorded as the characteristic interval, the non-initial intervals in the analysis time segment whose abnormal influence is greater than the preset abnormal influence and each initial interval are recorded as the characteristic intervals;

[0116] The value of the preset abnormality impact can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of the prediction, the smaller the value of the preset abnormality impact. A method for determining the value of the preset abnormality impact is provided, and the average value of the abnormality impact corresponding to each non-initial interval in the feature interval in the historical record that can meet the user's needs is recorded as the preset abnormality impact.

[0117] Specifically, under the preset conditions, the associated combinations are determined based on the characteristic values ​​of the instability parameters and the proportion of the instability parameters, including:

[0118] If the instability parameter characteristic value is less than the preset instability parameter characteristic value and the instability parameter proportion is less than the preset instability parameter proportion, then the correlation combination is determined according to the instability parameter correlation degree;

[0119] If the instability parameter characteristic value is greater than or equal to the preset instability parameter characteristic value or the instability parameter ratio is greater than or equal to the preset instability parameter ratio, then the associated combination is determined according to the characteristic interval association threshold;

[0120] The preset condition is the third abnormal state or the adjustment abnormal state.

[0121] Among them, the abnormal adjustment state is that after the load data is supplemented or the characteristic parameter association is supplemented according to the comparison reference value or the data resolution of the characteristic interval is adjusted, the fluctuation coefficient of the target machine learning model is greater than or equal to the preset fluctuation coefficient or the maximum deviation value is greater than or equal to the preset maximum deviation value;

[0122] For a single water quality indicator, detect the maximum value of the sub-deviation values ​​corresponding to each prediction result of the water quality indicator within the preset time period and record it as the maximum sub-deviation value. If the maximum sub-deviation value is greater than the preset sub-deviation value, mark the water quality indicator as an unstable parameter;

[0123] The value of the preset sub-deviation value can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of the prediction, the smaller the value of the preset sub-deviation value is. A preset sub-deviation value is provided, and the preset sub-deviation value is 10%;

[0124] The instability parameter characteristic value is the average value of the sub-instability characteristic values ​​corresponding to each instability parameter. The sub-instability characteristic value corresponding to a single instability parameter is the standard deviation of the sub-deviation values ​​corresponding to each prediction result of the instability parameter within a preset time period.

[0125] The proportion of unstable parameters = the number of unstable parameters / the total number of monitored water quality indicators;

[0126] The values ​​of the preset instability parameter characteristic value and the preset instability parameter ratio can be determined by the user according to the actual application scenario. The smaller the values ​​of the preset instability parameter characteristic value and the preset instability parameter ratio are, the greater the user's demand for determining the associated combination according to the instability parameter correlation degree is. A value of the preset instability parameter characteristic value and the preset instability parameter ratio is provided, and the historical records of the user determining the associated combination according to the instability parameter correlation degree are detected. The average values ​​of the instability parameter characteristic values ​​and the average values ​​of the instability parameter ratios corresponding to the historical records that can meet the user's needs are recorded as the preset instability parameter characteristic value and the preset instability parameter ratio, respectively.

[0127] Determining an association combination according to the instability parameter correlation degree, including: performing an association analysis on each training data, when performing the association analysis on a single training data, recording the training data as target training data, recording other training data other than the target training data that are not recorded in the association combination as reference training data, recording the target training data and each reference training data whose instability parameter correlation degree with the target training data is greater than or equal to a preset instability parameter correlation degree into an association combination, and continuing to perform the association analysis on the training data that are not recorded in the association combination until all the training data are recorded in the association combination, and then stopping the association analysis;

[0128] Determining an associated combination according to a feature interval association threshold includes: performing a combination analysis on each training data; when performing a combination analysis on a single training data, recording the training data as target data, recording other training data other than the target data that are not recorded in the associated combination as reference data, recording the target data and each reference data whose feature interval association threshold with the target data is greater than or equal to a preset feature interval association threshold into an associated combination, and continuing to perform a combination analysis on the training data that are not recorded in the associated combination until all training data are recorded in the associated combination, then stopping the combination analysis;

[0129] For any two training data, the instability parameter correlation degree is the average value of the sub-correlation degrees corresponding to each instability parameter. The sub-correlation degree corresponding to a single instability parameter = 1-[1 / (|standard deviation of each value corresponding to the instability parameter in one training data - standard deviation of each value corresponding to the instability parameter in the other training data|+1)];

[0130] For any two training data, the feature interval correlation threshold is the average value of the interval correlation of each monitoring parameter in the feature interval of the two training data; the interval correlation corresponding to a single monitoring parameter = 1-[1 / (|standard deviation of each value corresponding to the monitoring parameter in the feature interval of one training data - standard deviation of each value corresponding to the monitoring parameter in the feature interval of the other training data|+1)];

[0131] The values ​​of the preset instability parameter correlation degree and the preset characteristic interval correlation threshold can be determined by the user according to the actual application scenario. The greater the user's demand for improving the prediction accuracy, the larger the values ​​of the preset instability parameter correlation degree and the preset characteristic interval correlation threshold are. A value of the preset instability parameter correlation degree and the preset characteristic interval correlation threshold is provided, and the historical records of determining the correlation combination according to the instability parameter correlation degree are detected, and the average value of the reference instability parameter correlation degree corresponding to each correlation combination in the historical records that can meet the user's needs is recorded as the preset instability parameter correlation degree. The historical records of determining the correlation combination according to the characteristic interval correlation threshold are detected, and the average value of the reference characteristic interval correlation threshold corresponding to each correlation combination in the historical records that can meet the user's needs is recorded as the preset characteristic interval correlation threshold;

[0132] The reference instability parameter correlation degree and the reference feature interval correlation threshold are respectively the instability parameter correlation degree and the feature interval correlation threshold corresponding to any two training data in a single correlation combination.

[0133] Specifically, the association combination whose abnormal reference value is greater than the preset abnormal reference value is recorded as an unstable association combination, and the data volume of the unstable association combination is increased and adjusted based on the abnormal reference value;

[0134] The increase in the data volume of a single unstable association combination is positively correlated with the abnormal reference value corresponding to the unstable association combination.

[0135] Abnormal reference value = quantity abnormality / average value of quantity abnormality corresponding to each unstable association combination in the historical records that can meet user needs + deviation abnormality / average value of deviation abnormality corresponding to each unstable association combination in the historical records that can meet user needs;

[0136] The abnormality of the number of corresponding single association combinations = 1-(the number of training data included in the association combination / the average number of training data corresponding to each association combination);

[0137] The deviation abnormality corresponding to a single association combination = the average deviation reference values ​​corresponding to other association combinations other than the unstable association combination in the historical records that can meet user needs - the deviation reference value corresponding to the association combination. The deviation reference value corresponding to a single association combination is the average fluctuation deviation threshold values ​​corresponding to each monitoring parameter. The fluctuation deviation threshold value corresponding to a single monitoring parameter is the standard deviation of the processing reference values ​​corresponding to the monitoring parameter in each training data of the single association combination. The processing reference value corresponding to the monitoring parameter in a single training data = the value corresponding to the monitoring parameter at the latest time of the training data - the value corresponding to the monitoring parameter at the earliest time of the training data.

[0138] The value of the preset abnormal reference value can be determined by the user according to the actual application scenario. The greater the user's demand for improved prediction accuracy, the smaller the value of the preset abnormal reference value. A preset abnormal reference value is provided, and the average value of the abnormal reference values ​​corresponding to each instability association combination in the historical records that can meet the user's needs is recorded as the preset abnormal reference value;

[0139] When increasing the amount of data for a single unstable association combination, historical operating data with a deviation coefficient greater than a preset deviation coefficient from the association combination is selected for supplementation;

[0140] The deviation coefficient between a single historical operating data and a single associated combination = the average value of the supplementary fluctuation deviation thresholds corresponding to each monitoring parameter - the deviation reference value corresponding to the associated combination. The supplementary fluctuation deviation threshold corresponding to a single monitoring parameter is the standard deviation of the processing reference value corresponding to the monitoring parameter in each analysis data. The analysis data includes the training data in the associated combination and the historical operating data. The processing reference value corresponding to a single monitoring parameter in a single historical operating data = the value of the monitoring parameter corresponding to the latest time of the historical operating data - the value of the monitoring parameter corresponding to the earliest time of the historical operating data.

[0141] The value of the preset deviation coefficient can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of the prediction, the larger the value of the preset deviation coefficient. A value of the preset deviation coefficient is provided, and the average value of the deviation coefficients corresponding to each historical operating data supplemented for each instability-related combination in the historical records that can meet the user's needs is recorded as the preset deviation coefficient.

[0142] Specifically, under the condition of unstable data volume regulation, abnormal hyperparameters are determined based on the complexity comparison, including:

[0143] If the complexity comparison degree is less than the preset complexity comparison degree, the abnormal hyperparameter is determined based on the regular comparison degree;

[0144] If the complexity comparison degree is greater than or equal to the preset complexity comparison degree, the abnormal hyperparameter is determined based on the floating reference value.

[0145] Among them, hyperparameters include but are not limited to learning rate, regularization parameter, batch size, and number of network layers. The data volume adjustment instability condition is that after increasing the data volume of each unstable correlation combination, the fluctuation coefficient of the target machine learning model is greater than or equal to the preset fluctuation coefficient or the maximum deviation value is greater than or equal to the preset maximum deviation value;

[0146] Complexity comparison degree = |data reference value - average of the data reference values ​​corresponding to each historical record that is adjusted for abnormal hyperparameters and can meet user needs|, where the data reference value is the average of the abnormal reference values ​​corresponding to each association combination after increasing the amount of data for the unstable association combination based on the abnormal reference value;

[0147] The value of the preset complexity comparison degree can be determined by the user according to the actual application scenario. The larger the value of the preset complexity comparison degree, the greater the user's demand for determining abnormal hyperparameters based on the regular comparison degree. A method for determining the value of the preset complexity comparison degree is provided. The historical records of determining abnormal hyperparameters based on the regular comparison degree are detected, and the average value of the complexity comparison degree corresponding to the historical records that can meet the user's needs is recorded as the preset complexity comparison degree.

[0148] When determining abnormal hyperparameters based on the regularity comparison degree, the hyperparameter adjusted in the historical record with the largest regularity comparison degree is recorded as the abnormal hyperparameter;

[0149] When determining abnormal hyperparameters based on floating reference values, hyperparameters whose floating reference values ​​are greater than the preset floating reference values ​​are recorded as abnormal hyperparameters;

[0150] The floating reference value corresponding to a single hyperparameter = |value of the hyperparameter / output reference value - average value of the hyperparameter values ​​in the historical records that can meet the user's requirements / average value of the output reference values ​​in the historical records that can meet the user's requirements|;

[0151] The value of the preset floating reference value can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of the prediction, the smaller the value of the preset floating reference value is. A value of the preset floating reference value is provided, and the average value of the floating reference values ​​corresponding to each abnormal hyperparameter in each historical record that can determine the abnormal hyperparameter based on the floating reference value and meet the user's needs is recorded as the preset floating reference value.

[0152] Specifically, each abnormal hyperparameter is adjusted according to the deviation coefficient;

[0153] The adjustment amount corresponding to a single abnormal hyperparameter is positively correlated with the deviation coefficient corresponding to the abnormal hyperparameter.

[0154] For a single hyperparameter, the deviation coefficient = floating reference value / average floating reference value corresponding to the historical records that can meet user needs + regular deviation value / average regular deviation value corresponding to the historical records that can meet user needs;

[0155] Regularity deviation value = the adjustment amount of a single abnormal hyperparameter in the historical record with the highest regularity comparison degree × regularity comparison degree / the average regularity comparison degree corresponding to the historical records that can meet user needs;

[0156] The adjustment coefficient corresponding to a single abnormal hyperparameter = the value of the abnormal hyperparameter / the output reference value - the average value of the values ​​corresponding to the abnormal hyperparameter in the historical records that can meet the user's needs / the average value of the output reference values ​​corresponding to the historical records that can meet the user's needs;

[0157] For a single abnormal hyperparameter,

[0158] If the adjustment coefficient is greater than or equal to 0, the abnormal hyperparameter is adjusted downward according to the deviation coefficient;

[0159] If the adjustment coefficient is less than 0, the abnormal hyperparameter is increased and adjusted according to the deviation coefficient;

[0160] The adjustment amount corresponding to a single abnormal hyperparameter is positively correlated with the deviation coefficient corresponding to the abnormal hyperparameter.

[0161] Specifically, the average value of the deviation values ​​corresponding to each prediction result of the target machine learning model within a preset time period is compared with the standard deviation to obtain the output reference value, and the regularity comparison degree is determined based on the output reference value and the hyperparameter comparison degree.

[0162] Among them, each historical record that adjusts abnormal hyperparameters and can meet user needs is recorded as a reference record. The regularity comparison degree corresponding to a single reference record = |output reference value - output reference value corresponding to the reference record| × hyperparameter comparison degree. Hyperparameter comparison degree = the number of hyperparameters in the reference record that have the same value as the current target machine learning model / the total number of hyperparameters corresponding to the target machine learning model;

[0163] Output reference value = average value of deviation values ​​corresponding to each prediction result of the target machine learning model within the preset time length after increasing the data amount of the unstable associated combination based on the abnormal reference value / standard deviation of deviation values ​​corresponding to each prediction result of the target machine learning model within the preset time length.

[0164] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

[0165] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for predicting the effect of biological sewage treatment based on machine learning, characterized in that: include: Determine whether the target machine learning model has anomalies based on the fluctuation coefficient and maximum deviation value; If the target machine learning model has an anomaly, the abnormal state is determined based on the load anomaly coefficient and the load feature association threshold; Supplement load data according to abnormal conditions, or supplement characteristic parameters or adjust data resolution of characteristic intervals, or perform multi-dimensional optimization; Determine the initial interval based on the feature representativeness and paragraph comparison, and determine whether to adjust the initial interval to obtain the feature interval based on the initial interval distribution coefficient and interval influence; Under the preset conditions, multi-dimensional optimization is performed, among which, Determine a correlation combination based on an instability parameter correlation degree or a characteristic interval correlation threshold based on an instability parameter characteristic value and an instability parameter proportion, determine an instability correlation combination based on an abnormal reference value, and adjust the data volume of the instability correlation combination based on the abnormal reference value; Under the condition of unstable data volume adjustment, abnormal hyperparameters are determined according to the complexity comparison degree, and each abnormal hyperparameter is adjusted according to the deviation coefficient.

2. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 1, characterized in that: Determine whether the target machine learning model has abnormalities based on the fluctuation coefficient and maximum deviation value, including: If the fluctuation coefficient is greater than or equal to the preset fluctuation coefficient or the maximum deviation value is greater than or equal to the preset maximum deviation value, it is determined that the target machine learning model has an anomaly; If the fluctuation coefficient is less than the preset fluctuation coefficient and the maximum deviation value is less than the preset maximum deviation value, it is determined that there is no abnormality in the target machine learning model.

3. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 2, characterized in that: If the abnormal state is the first abnormal state in which the load abnormality coefficient is greater than or equal to the preset load abnormality coefficient, the load data is supplemented.

4. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 3, characterized in that: If the abnormal state is a second abnormal state in which the load abnormality coefficient is less than the preset load abnormality coefficient and the load characteristic association threshold is less than the preset load characteristic association threshold, then the characteristic parameter association is supplemented or the data resolution of the characteristic interval is adjusted according to the comparison reference value; If the comparison reference value is greater than or equal to the preset comparison reference value, the data resolution of the feature interval is reduced; If the comparison reference value is less than the preset comparison reference value, the characteristic parameter association supplement is performed.

5. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 4, characterized in that: Methods for confirming the characteristic interval include: Determine the initial interval based on feature representativeness and paragraph comparison; If the initial interval distribution coefficient is greater than or equal to the preset initial interval distribution coefficient or the interval influence is less than the preset interval influence, the initial interval is recorded as a characteristic interval; If the initial interval distribution coefficient is less than the preset initial interval distribution coefficient and the interval influence is greater than or equal to the preset interval influence, the increased and adjusted initial interval is recorded as the characteristic interval.

6. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 5, characterized in that: Under preset conditions, the associated combinations are determined based on the characteristic values ​​of the instability parameters and the proportion of the instability parameters, including: If the instability parameter characteristic value is less than the preset instability parameter characteristic value and the instability parameter proportion is less than the preset instability parameter proportion, then the correlation combination is determined according to the instability parameter correlation degree; If the instability parameter characteristic value is greater than or equal to the preset instability parameter characteristic value or the instability parameter ratio is greater than or equal to the preset instability parameter ratio, then the associated combination is determined according to the characteristic interval association threshold; The preset condition is the third abnormal state or the adjustment abnormal state.

7. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 6, characterized in that: Recording the association combination whose abnormal reference value is greater than the preset abnormal reference value as an unstable association combination, and increasing the data volume of the unstable association combination based on the abnormal reference value; The increase in the data volume of a single unstable association combination is positively correlated with the abnormal reference value corresponding to the unstable association combination.

8. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 7, characterized in that: Under the condition of unstable data volume regulation, abnormal hyperparameters are determined based on the complexity comparison, including: If the complexity comparison degree is less than the preset complexity comparison degree, the abnormal hyperparameter is determined based on the regular comparison degree; If the complexity comparison degree is greater than or equal to the preset complexity comparison degree, the abnormal hyperparameter is determined based on the floating reference value.

9. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 8, characterized in that: Adjust each abnormal hyperparameter based on the deviation coefficient; The adjustment amount corresponding to a single abnormal hyperparameter is positively correlated with the deviation coefficient corresponding to the abnormal hyperparameter.

10. The method for predicting the effect of biological sewage treatment based on machine learning according to claim 8, characterized in that: The average value of the deviation values ​​corresponding to each prediction result of the target machine learning model within the preset time is compared with the standard deviation to obtain the output reference value, and the regularity comparison degree is determined based on the output reference value and the hyperparameter comparison degree.

Citation Information

Patent Citations

  • Sewage treatment effect prediction method and system based on deep learning

    CN119416066A

  • Dynamic optimization system and method based on sewage treatment

    CN119090101A

  • Method for predicting equipment fault in industrial internet based on AI large model

    CN119357677A

  • Sewage treatment operation support apparatus, sewage treatment operation support system, sewage treatment operation support method, and sewage treatment operation support program

    JP2006021085A