A Machine Learning-Based Intelligent Prediction Method for High-Salinity Wastewater Treatment in Coal Mines and Coal Chemical Industries
By improving the alpha evolution algorithm and optimizing the BP neural network, a machine learning model was constructed, which solved the problem of accurately predicting the treatment effect of high-salt wastewater from coal mines and coal chemical plants. This achieved high-precision prediction of wastewater treatment effect and improved the model's search efficiency and prediction accuracy.
Patent Information
- Application Number
- CN202510810163.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing technologies cannot accurately and timely grasp the treatment effect of high-salt wastewater from coal mines and coal chemical plants. Traditional methods are highly subjective and have large errors, and simple models cannot fully consider the interaction of complex variables, thus limiting the accuracy of prediction.
An improved alpha evolution algorithm was used to optimize the BP neural network. A machine learning model was constructed by improving the alpha operator, boundary constraints, and policy selection. The model was combined with multi-source data for prediction. Flexible boundary processing and hybrid selection strategies were used. A loss function for catalyst activity decay factors was added to optimize the model structure.
It significantly improves the prediction accuracy of high-salinity wastewater treatment effects, meets the industry's demand for high-precision prediction, reduces the risk of getting trapped in local optima, and improves the model's search efficiency and prediction accuracy.
Smart Images

Figure QLYQS_2 
Figure QLYQS_5 
Figure QLYQS_7
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wastewater treatment, and in particular relates to an intelligent prediction method for the treatment effect of high-salt wastewater from coal mines and coal chemical plants based on machine learning. Background Technology
[0002] Large quantities of high-salinity wastewater are generated during coal mining and coal chemical production. This type of wastewater has a complex composition, containing high concentrations of total dissolved solids, chemical oxygen demand (COD), biological oxygen demand (BOD), heavy metal ions, and high levels of suspended solids. Direct discharge without effective treatment can cause serious environmental pollution, threatening ecological balance and human health. Traditional high-salinity wastewater treatment methods rely heavily on experience and conventional monitoring to assess treatment effectiveness, failing to accurately and promptly grasp the various indicators of the treated wastewater. With increasingly stringent environmental standards, precise control over the treatment effectiveness of high-salinity wastewater has become a critical requirement for industry development. However, existing prediction methods have several shortcomings. On the one hand, manual monitoring and experience-based judgment are highly subjective and prone to large errors, making it difficult to meet the requirements for high-precision prediction of wastewater treatment effectiveness. On the other hand, some prediction methods based on simple models cannot fully consider the interactions of numerous complex variables during the treatment process, limiting their accuracy. Summary of the Invention
[0003] To address the technical problems mentioned above, this invention proposes an intelligent prediction method for the treatment effect of high-salt wastewater from coal mines and coal chemical plants based on machine learning.
[0004] To achieve the above objectives, the technical solution adopted by the present invention includes the following steps:
[0005] S1. First, collect multi-source data on the treatment process of high-salt wastewater from coal mines and coal chemical plants, and preprocess the data.
[0006] S2. A machine learning model is constructed by optimizing the BP neural network using an improved alpha evolution algorithm; the improvements to the alpha evolution algorithm include improvements to the alpha operator, boundary constraints, and policy selection.
[0007] The specific improvements to the alpha operator are as follows:
[0008] Step 1: Improve the construction of the evolutionary matrix. First, calculate the fitness value of each individual, then sort the individuals from best to worst according to their fitness values, dividing the population into k levels. The number of individuals in the j-th level is n. j The total number of individuals is N, and the sampling probability of individuals at different levels is... From the j-th layer with probability p j Extract n j ×p j The evolutionary matrix E is composed of ×N individuals;
[0009] Step 2: Improve the calculation of the adaptive basis vector P, as follows: Where ω t For dynamic weights, These are the evolution paths of the previous iteration, and A and B represent matrices.
[0010] Step 3: Improve the random step size. First, improve the decay factor α: Where FEs and MaxFEs are the current function evaluation count and the maximum function evaluation count, respectively, and D t D is an indicator of population diversity, calculated using the average Euclidean distance between individuals. avg Given the historical average of population diversity, the random step size Δr = (ub-lb)·(2R1·R2-R2)·S·α, where ub and lb are the upper and lower bounds of the search space, respectively, S is the scaling factor, and R1 and R2 are real matrices used to generate perturbations.
[0011] S3. Improve the loss function by incorporating catalyst activity attenuation factors;
[0012] S4. Train and validate the constructed model, and output the best model;
[0013] S5. Finally, import the new data into the optimal model to predict the wastewater treatment effect.
[0014] Preferably, the multi-source data consists of input data and predicted output data. The input data includes total dissolved solids (TDS), chemical oxygen demand (COD), biological oxygen demand (BOD), heavy metal ion concentration, suspended solids content, pH value, and treatment process parameters, including reaction temperature, reaction time, reagent content, and wastewater retention time. The predicted output data includes total dissolved solids (TDS), chemical oxygen demand (COD), biological oxygen demand (BOD), heavy metal ion concentration, suspended solids content, and pH value of the effluent.
[0015] Preferably, the number of individuals in the improvement of the alpha operator in step S2 is n. j satisfy And the sampling probability p j satisfy
[0016] Preferably, in step S2, the dynamic weight ω in the calculation of the improved adaptive basis vector P t It gradually decreases as the number of iterations increases, specifically in the form of: Where T is the maximum number of iterations.
[0017] Preferably, the boundary constraint in step S2 is improved by adopting an elastic boundary treatment method: Where e iThis represents the value of an evolved individual in a certain dimension. When an individual goes out of bounds, it is not regenerated completely randomly, but rather generated within a flexible region near the boundary.
[0018] Preferably, the strategy selection in step S2 is improved to a hybrid selection strategy, which combines elite retention and tournament selection. The m individuals with the best fitness in the population are directly retained to the next generation. These individuals are called elite individuals. Then, for the remaining NM individuals, the tournament selection method is used.
[0019] Preferably, the loss function in step S3 is improved as follows: Where N 样 denoted as the total number of samples, and b as the activity of the catalyst.
[0020] Compared with existing technologies, the advantages and positive effects of this invention are as follows: It employs an improved alpha evolution algorithm to optimize the BP neural network, improves the alpha operator to enhance population diversity and accelerate convergence; elastic boundary processing preserves boundary information, and a hybrid selection strategy avoids premature convergence. The loss function incorporates catalyst activity decay factors, making it more closely aligned with actual treatment conditions. These technologies work synergistically, enabling the model to accurately consider complex variables, effectively avoid local optima, significantly improve prediction accuracy, and meet the industry demand for high-precision prediction of high-salinity wastewater treatment effects in coal mines and coal chemical plants. Detailed Implementation
[0021] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described below with reference to embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0022] Numerous specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways than those described herein, and therefore the invention is not limited to the specific embodiments disclosed in the following specification.
[0023] High-salinity wastewater treatment has become a key challenge restricting the sustainable development of the industry. This type of wastewater has an extremely complex composition, containing high concentrations of total dissolved solids, chemical oxygen demand (COD), biological oxygen demand (BOD), heavy metal ions, and a large amount of suspended solids. Improper treatment and direct discharge can cause serious pollution to soil, water bodies, and the atmosphere, endangering ecological balance and human health. Meanwhile, with increasingly stringent environmental standards, traditional methods relying on experience and conventional monitoring to judge wastewater treatment effectiveness are no longer sufficient to meet the industry's urgent need for high-precision prediction of treatment results. This invention proposes an intelligent prediction method for the treatment effect of high-salinity wastewater from coal mines and coal chemical plants based on machine learning.
[0024] First, multi-source data was collected and preprocessed. Various data were collected from the treatment process of high-salt wastewater from coal mines and coal chemical plants, covering influent total dissolved solids (TDS), influent chemical oxygen demand (COD), influent biological oxygen demand (BOD), influent heavy metal ion concentration, influent suspended solids (SSD), and influent pH, as well as treatment process parameters, including reaction temperature, reaction time, reagent concentration, and wastewater retention time. The effluent TDS, effluent COD, effluent BOD, effluent heavy metal ion concentration, effluent SSD, and effluent pH were also collected as predicted output data. These collected data were preprocessed to remove outliers, fill missing values, and standardize the data to ensure comparability between different data types, providing a high-quality data foundation for subsequent model construction.
[0025] Next, considering that the existing alpha evolutionary algorithm blindly explores the search space, resulting in slow convergence, and that in the later stages of evolution, due to the decrease in population diversity, the algorithm is prone to lingering near local optima, further affecting convergence efficiency, this invention improves the alpha operator. First, it improves the construction of the evolutionary matrix by calculating the fitness value of each individual, where the fitness value is the value of the fitness function, which is the reciprocal of the loss function. Then, individuals are sorted from best to worst according to their fitness values, dividing the population into k levels, with n individuals in the j-th level. j The total number of individuals is N, and the sampling probability of individuals at different levels is... From the j-th layer with probability p j Extract n j ×p j An evolutionary matrix E is composed of ×N individuals, where the number of individuals is n. j satisfy And the sampling probability p j satisfy The calculation of the adaptive basis vector P is improved as follows: Where ω t For dynamic weights, Let A and B represent the evolutionary paths of the previous iteration, and let A and B denote matrices. To improve the stochastic step size, first improve the decay factor α: Where FEs and MaxFEs are the current function evaluation count and the maximum function evaluation count, respectively, and D t D is an indicator of population diversity, calculated using the average Euclidean distance between individuals. avgLet Δr be the historical average of population diversity. Then, the random step size Δr = (ub - lb)·(2R1·R2 - R2)·S·α, where ub and lb are the upper and lower bounds of the search space, respectively, S is the scaling factor, and R1 and R2 are real matrices used to generate perturbations. In summary, regarding the improved construction of the evolutionary matrix, stratified sampling based on individual fitness ranking fully utilizes information from individuals at different levels of the population, significantly increasing population diversity. This significantly enhances the algorithm's ability to escape local optima, thereby exploring a broader solution space and increasing the likelihood of finding the global optimum. The improved adaptive basis vector P calculation introduces dynamic weights that change with the number of iterations. This allows the algorithm to actively explore new solution spaces in the early stages of iteration, while focusing on a refined search near better solutions in the later stages. This adaptive adjustment effectively accelerates the convergence speed while ensuring search accuracy and reducing unnecessary computational resource waste. By improving the random step size to correlate it with the number of function evaluations and population diversity, a larger step size can quickly locate possible regions in the early stages of the search. As the algorithm approaches the optimal solution, the step size is reduced to avoid missing the optimal solution. The algorithm is like having intelligent navigation, which can flexibly cope with different search stages and improve its ability to deal with complex optimization problems.
[0026] Furthermore, considering that in existing Alpha Evolutionary algorithms, a common approach to boundary constraints is to directly and completely randomly regenerate individuals when they cross the boundary. While this method is simple, it has significant drawbacks. Completely random generation leads to the loss of valuable information accumulated by individuals near the boundary, causing the algorithm to repeatedly explore already explored areas in subsequent searches, reducing search efficiency and potentially disrupting the overall evolutionary trend of the population, affecting the algorithm's convergence to the optimal solution. Regarding strategy selection, common approaches often employ a single selection strategy, such as simple roulette wheel selection or tournament selection. Roulette wheel selection is easily affected by the distribution of individual fitness values; if the fitness values differ too much, some individuals may be overselected, causing the algorithm to converge prematurely. Tournament selection is insufficient in maintaining population diversity; if the tournament size is inappropriately set, it may lead to an overconcentration of superior individuals in the population, also causing the algorithm to fall into local optima. Considering this situation, this invention improves the boundary constraint by adopting a flexible boundary handling method: Where e iThe value of an evolved individual in a certain dimension is considered. When an individual goes out of bounds, it is not randomly regenerated, but rather generated within a flexible region near the boundary. The improved strategy selection employs a hybrid selection strategy, combining elite retention and tournament selection. The m individuals with the best fitness in the population are directly retained for the next generation; these individuals are called elite individuals. Then, for the remaining NM individuals, a tournament selection method is used. This improvement uses a flexible boundary handling approach. When an individual goes out of bounds, a new individual is generated in a flexible region near the boundary, preserving effective information near the boundary, reducing redundant searches, maintaining the population's evolutionary trend, and improving search efficiency. The strategy selection uses a hybrid strategy combining elite retention and tournament selection. Elite retention ensures the inheritance of excellent genes, while tournament selection maintains population diversity. The combination of these two approaches prevents premature convergence of the algorithm, making it more capable of optimizing complex problems.
[0027] Furthermore, considering that most existing loss functions focus solely on prediction accuracy, measuring model performance by calculating the difference between predicted and true values, this approach overlooks some details and fails to reflect precise changes. This results in the model not accurately reflecting the actual processing situation, leading to significant deviations between predicted and actual results. The improved loss function in this algorithm is as follows: Where N 样 Let be the total number of samples, and b be the catalyst activity. In the loss function, (1-b) 2 As the denominator, b decreases with the use of catalyst, (1-b) 2 Increase, then This will decrease. When the catalyst is first used and its activity is high, and b is relatively large, the penalty of the loss function on the prediction error is relatively small; however, as the catalyst activity declines, b decreases. As the value decreases, the proportion of the same prediction error in the loss function increases, which means that the requirement for prediction accuracy is getting higher and higher. The model will work harder to reduce the prediction error, which meets the need for improved prediction accuracy as the catalyst degrades with use.
[0028] In the Alpha Evolutionary Algorithm, improvements to the hierarchical sampling during evolutionary matrix construction significantly increase population diversity, providing a rich pool of initial solutions and preventing the algorithm from getting trapped in local optima. Improvements to adaptive basis vector calculation, with dynamic weights that change with iteration counts, actively explore new solution spaces in the early stages and focus on refined searching in the later stages, working in conjunction with evolutionary matrix construction to improve search efficiency. Improvements to the random step size, dynamically adjusted based on the number of function evaluations and population diversity, work in conjunction with the previous two methods, allowing for rapid region localization with large steps in the early stages and precise searching with small steps when close to the optimal solution. The flexible handling of boundary constraints preserves effective information near the boundary, allowing the algorithm to better utilize this information for continuous optimization; the hybrid selection strategy ensures the inheritance of superior genes and population diversity, preventing premature convergence. These improvements, in conjunction with the Alpha Evolutionary Algorithm, optimize the BP neural network structure, determining the number of neurons in the input, hidden, and output layers, the number of hidden layers, connection weights, and biases. The loss function incorporates a catalyst activity decay factor; as catalyst activity decreases, the required prediction accuracy gradually increases. This improvement echoes the entire algorithm optimization process. The enhanced search capability of the algorithm creates conditions for loss function optimization, while the strict requirements of the loss function prompt the algorithm to continuously seek better solutions, thereby comprehensively improving the accuracy and reliability of the prediction model.
[0029] Finally, the constructed model is trained and validated, the best model is output, and the new data is imported into the best model to predict the wastewater treatment effect.
[0030] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments for application in other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A machine learning-based intelligent prediction method for the treatment effect of high-salinity wastewater in coal mines and coal chemical plants, characterized in that, The method comprises the following steps: S1, first collect multi-source data in the process of treating high-salt wastewater in coal mines and coal chemical industry, and pretreat the data; S2, adopt an improved alpha evolutionary algorithm to optimize a BP neural network to construct a machine learning model; the improvement of the alpha evolutionary algorithm comprises improvement of an alpha operator, improvement of a boundary limit, and improvement of a strategy selection; The improvement of the alpha operator is specifically: Step one, the construction of improved evolutionary matrix, first calculate the fitness value of each individual, then sort the individuals from good to bad according to the fitness value, divide the population into k levels, the number of individuals in the jth level is , the total number of individuals is N, the sampling probability of individuals in different levels is , extract individuals from the jth level with a probability of to form the evolutionary matrix E; Step two, the calculation of the adaptive basis vector P is improved, and the improvement is: wherein is a dynamic weight, respectively the evolutionary path of the last iteration, A and B represent matrices; Step three, improve the random step, first improve the damping factor : where are the current function evaluation number and the maximum function evaluation number, respectively, is the population diversity index, calculated by the average Euclidean distance between individuals, is the historical average of the population diversity, then the random step where are the upper and lower bounds of the search space, respectively, and S is the scaling factor, is a real matrix used to generate perturbations; S3, add a catalyst activity attenuation factor to an improved loss function; S4, train and verify the constructed model, and output an optimal model; S5, finally, import new data into the optimal model to predict the wastewater treatment effect.
2. The intelligent prediction method for the treatment effect of coal mine and coal chemical high-salinity wastewater based on machine learning according to claim 1, characterized in that, The multi-source data are input data and predicted output data, wherein the input data are total dissolved solids of influent, chemical oxygen demand of influent, biological oxygen demand of influent, heavy metal ion concentration of influent, suspended solid content of influent, pH value of influent, and treatment process parameters including reaction temperature, reaction time, medicament content, and wastewater residence time, and the predicted output data include total dissolved solids of effluent, chemical oxygen demand of effluent, biological oxygen demand of effluent, heavy metal ion concentration of effluent, suspended solid content of effluent, and pH value of effluent.
3. The intelligent prediction method for the treatment effect of coal mine and coal chemical high-salinity wastewater based on machine learning according to claim 1, characterized in that, The number of individuals in the improvement of the alpha operator in step S2 is satisfies and the sampling probability satisfies .
4. The intelligent prediction method for the treatment effect of coal mine and coal chemical high-salinity wastewater based on machine learning according to claim 1, characterized in that, The dynamic weight in the calculation of the improved adaptive basis vector P in the step S2 is gradually reduced with the increase of the number of iterations, and the specific form is: where T is the maximum number of iterations.
5. The intelligent prediction method for the treatment effect of coal mine and coal chemical high-salinity wastewater based on machine learning according to claim 1, characterized in that, The improvement of the boundary restriction in the step S2 is to adopt an elastic boundary processing mode: wherein is the value of the individual after evolution in a certain dimension, and when the individual is out of the boundary, it is not completely randomly regenerated, but regenerated in an elastic region near the boundary.
6. The intelligent prediction method for the treatment effect of coal mine and coal chemical high-salinity wastewater based on machine learning according to claim 1, characterized in that, The improvement of the strategy selection in the step S2 is a hybrid selection strategy, which adopts a hybrid selection strategy combining elite reservation and tournament selection, directly reserves M individuals with the best fitness in a population to the next generation, the individuals are called elite individuals, then, for the remaining N-M individuals, a tournament selection method is adopted.
7. The intelligent prediction method for the treatment effect of coal mine and coal chemical high-salinity wastewater based on machine learning according to claim 1, characterized in that, The improvement of the loss function in step S3 is: wherein N is the total number of samples and b is the activity of the catalyst.
Citation Information
Patent Citations
Online medicine-adding control method and system for wastewater treatment
CN108408855A
Shared cloud integrated ecology air quality prediction and big data system
CN120067583A