Water body organic dye adsorption and recovery performance prediction method based on machine learning
Through machine learning-based methods, including Monte Carlo simulation, uncertainty analysis, network flow theory and deep learning models, the problems of insufficient data representation and lack of combination optimization strategies in the prediction of water organic dye adsorption and recovery performance are solved, and higher prediction accuracy and adsorption recovery efficiency are achieved.
Patent Information
- Application Number
- CN202510371597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has insufficient data representation in the prediction of adsorption and recovery performance of water organic dyes, difficulty in considering changes in environmental factors, and lack of multi-adsorbent combination optimization strategies, resulting in large deviations in the prediction results and high arbitrary material selection, which in turn reduces the economic and practicality of the adsorption and recovery process.
Using a machine learning-based approach, data diversity and representativeness are enhanced through Monte Carlo simulation and uncertainty analysis, network flow theory is applied to optimize adsorbent combinations, and adsorption recovery performance is predicted using deep learning models.
It improves the accuracy and stability of the prediction of adsorption recovery performance, enhances the ability of multiple adsorbents to work together, reduces the risk of resource waste, and improves the economic and practicality of the adsorption recovery process.
Smart Images

Figure CN120220876A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of adsorption recovery performance prediction, and particularly to a method for predicting the adsorption recovery performance of waterborne organic dyes based on machine learning. Background Art
[0002] The technical field of adsorption recovery performance prediction mainly involves methods for predicting and evaluating the performance of adsorbent materials in adsorbing and recovering target pollutants or substances. Through the establishment of mathematical models, computer simulations, experimental analyses, and data-driven prediction techniques, this field can effectively evaluate the performance of different adsorbent materials, and reveal the adsorption capacity, kinetic characteristics, selectivity, and stability of the materials towards pollutants.
[0003] In the prior art, the recovery efficiency of pollutants is usually inferred by relying on single experimental analysis or simple mathematical models. The data obtained is single and it is difficult to effectively consider the uncertain factors caused by changes in water body conditions. The lack of representativeness of the data leads to large deviations in the prediction results. Moreover, the prior art lacks an effective optimization strategy for multi-adsorbent combinations, and the prediction effect of the cooperation of different adsorbents in treating complex waterborne organic dyes is not ideal, resulting in difficulty in accurately determining the optimal adsorption combination during actual operation, high randomness in material selection, waste of resources, and increased treatment costs, thereby reducing the economy and practicality of the organic dye adsorption recovery process. Summary of the Invention
[0004] The object of the present invention is to solve the drawbacks existing in the prior art, and to propose a method for predicting the adsorption recovery performance of waterborne organic dyes based on machine learning.
[0005] In order to achieve the above object, the present invention adopts the following technical solution. A method for predicting the adsorption recovery performance of waterborne organic dyes based on machine learning includes the following steps: Collect the dye concentration, temperature, and pH value of organic dyes in the water body, and use Monte Carlo simulation to perform multiple random samplings to simulate the situation of the adsorption process, generating Monte Carlo simulation results; based on the Monte Carlo simulation results, perform uncertainty analysis to obtain the recovery efficiency distribution under various conditions, generating uncertainty analysis results; Based on the uncertainty analysis results, apply the network flow theory to optimize the process of multi-adsorbents, establish a flow network model of the interaction between adsorbents and dyes, where the nodes represent adsorbents and the capacity of the edges represents the adsorption capacity, solve the maximum flow problem, find the adsorbent combination, generating the maximum flow optimization results; based on the maximum flow optimization results, perform maximization configuration to obtain the optimized adsorbent configuration; Import the optimized adsorbent configuration into the deep learning model, design and train a neural network to predict the recovery performance, and obtain a deep learning recognition model with reference to the effects of adsorbent properties and dye concentration; evaluate the deep learning recognition model, check the accuracy of the model prediction, and generate a model prediction evaluation result; Judge whether the deep learning recognition model is qualified according to the model prediction evaluation result, and apply the deep learning recognition model to predict the adsorption and recovery performance of organic dyes in water.
[0006] Preferably, the steps for obtaining the Monte Carlo simulation results are as follows: Obtain the dye concentration value, water temperature value and current pH value to obtain triple information; According to the triple information, randomly perturb each set of inputs to generate simulation samples, and calculate the simulated adsorption value. The calculation formula is: ; Wherein, is the simulated adsorption value, represents the dye concentration value collected in the th group, represents the temperature value collected in the th group, represents the pH value collected in the th group, represents the diffusion rate of dye molecules generated by the th perturbation, represents the local water viscosity value after the th perturbation, represents the effective contact area between adsorbent particles in the th simulation, represents the specific polar activity value on the surface of the adsorbent particles, represents the shortest average path length between particles in the th group of samples during simulation; According to the simulation output performance of the simulated adsorption value under different perturbation conditions, collect all the adsorption performance change trends and analyze the stability to generate the Monte Carlo simulation results.
[0007] Preferably, the steps for obtaining the uncertainty analysis result are as follows: Collect the Monte Carlo simulation results, organize the Monte Carlo simulation results to form a data set; Based on the data set, calculate the uncertainty index under each condition. The calculation formula is: ; Wherein, is the uncertainty index, is the simulated adsorption value, is the temperature value, is the pH value, is the concentration of suspended solids in the water body, is the surface area of the adsorbent, is the chemical reaction rate in the sample, is the dissolved oxygen content, is the porosity of the adsorbent, is the surface activity of the adsorbent; According to the uncertainty index, analyze the distribution of the recovery efficiency under different conditions to obtain the uncertainty analysis result.
[0008] Preferably, the steps for obtaining the maximum flow optimization result are as follows: Screen all the connecting edges with non-zero adsorption capacity between the adsorbent nodes and the dye nodes from the uncertainty analysis result, re-number all the adsorbent nodes according to the dye molecular type, and establish a network graph structure composed of the adsorbent and dye nodes; According to the network graph structure, calculate the maximum adsorption flow value from the starting dye node to the ending adsorbent set, and the calculation formula is: ; Among them, is the maximum adsorption flow, is the adsorption flux per unit area between the th dye node and the th adsorbent node, is the adsorption layer thickness of the th adsorbent node, is the th number of pore diameters of the adsorbent, is the th binding frequency between the th dye node and the th adsorbent node, is the molecular polarity response value of the th dye node, is the set of all node pairs with associated edges;
[0009] Preferably, the steps for obtaining the optimized adsorbent configuration are as follows: According to the adsorbent combination path output in the maximum flow optimization result, extract the numbers of all adsorbent nodes and the path structure of the corresponding dye nodes, reconstruct the mapping relationship between the adsorption path and the adsorbent index, and generate a set of candidate adsorbents; According to the set of candidate adsorbents, calculate the flow proportion and the degree of overlap of node distribution borne by each adsorbent in the combined path, arrange and sort the adsorption paths and avoid node conflicts to generate a preliminary configuration structure; According to the preliminary configuration structure, match the deployment order and space occupancy of each adsorbent under the corresponding dye path to generate an optimized adsorbent configuration.
[0010] Preferably, the steps for obtaining the deep learning recognition model are as follows: Call the optimized adsorbent configuration, import the adsorbent node number and space deployment information as structured input parameters into the initial input layer of the deep neural network, and based on the pore size distribution, surface area and surface chemical composition parameters of the adsorbent, construct an input vector and perform dimension standardization to obtain a neural network input vector matrix; According to the neural network input vector matrix, set the dye concentration parameter as a dynamic input variable, construct a training sample sequence, perform parameter initialization, forward propagation and error backpropagation, and synchronously record the error signals under different sample rounds to establish a neural network recovery performance prediction structure; According to the neural network recovery performance prediction structure, perform multiple rounds of iterative convergence detection, combine the adsorbent parameters and dye concentration sample performances under different rounds, screen stable output nodes and freeze the corresponding connection weights to generate a deep learning recognition model.
[0011] Preferably, the steps for obtaining the model prediction evaluation result are as follows: Obtain the output results of the deep learning recognition model, compare and match each group of output results with the corresponding dye recovery results, extract three basic variables of the timestamp, predicted value and actual value of each group of predicted values to generate a model prediction and actual result set; Based on the model prediction and actual result set, calculate the matching accuracy index value; Based on the matching accuracy index value, traverse the prediction and pairing records of all samples, screen the proportion of sample points below the error threshold and record the evaluation fluctuation range to generate a model prediction evaluation result.
[0012] Preferably, judge whether the deep learning recognition model is qualified according to the model prediction evaluation result. The steps for predicting the adsorption and recovery performance of organic dyes in water using the deep learning recognition model are as follows: According to the model prediction evaluation result, screen out the unqualified neural network training rounds, lock the model version with the lowest output error and a stable fluctuation range to generate a deep learning recognition model that can be used for prediction; According to the deep learning recognition model that can be used for prediction, import new adsorbent configuration and dye concentration input parameters, execute the structure mapping and prediction calculation process, and generate the prediction results of the adsorption and recovery performance of organic dyes in water bodies.
[0013] The present invention provides a device for predicting the adsorption and recovery performance of organic dyes in water bodies, including: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, so that the device for predicting the adsorption and recovery performance of organic dyes in water bodies executes the method for predicting the adsorption and recovery performance of organic dyes in water bodies.
[0014] The present invention provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is run, the method for predicting the adsorption and recovery performance of organic dyes in water bodies is realized.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: The present invention realizes multiple random samplings by introducing Monte Carlo simulation, enhances the diversity and representativeness of the data in the process of adsorbing and recovering organic dyes in water bodies, improves the reliability and prediction accuracy of data analysis; on the basis of the Monte Carlo sampling data, further performs uncertainty analysis, clarifies the distribution law of the dye recovery efficiency under different conditions, reduces the uncertainty brought by the change of environmental factors in the recovery prediction, and improves the stability of the decision-making in the adsorption and recovery process; based on the application of network flow theory, constructs a flow network model of the interaction between the adsorbent and the dye, and clarifies the optimal configuration of different adsorbent combinations through the definition of node and edge capacities, enhancing the ability of multiple adsorbents to work together; in addition, uses a deep learning neural network combined with the properties of the adsorbent and the dye concentration to learn the non-linear relationship in the adsorption process, further improving the prediction accuracy, thereby enhancing the overall effect of adsorption and recovery and reducing the risk of resource waste. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] Please refer to Figure 1 , the present invention provides a technical solution, a method for predicting the adsorption and recovery performance of organic dyes in water bodies based on machine learning, including the following steps: Collect the dye concentration, temperature, and pH value of the organic dye in the water body, use Monte Carlo simulation to perform multiple random samplings, simulate the adsorption process, and generate Monte Carlo simulation results; based on the Monte Carlo simulation results, conduct uncertainty analysis to obtain the recovery efficiency distribution under various conditions and generate uncertainty analysis results; Based on the uncertainty analysis results, apply network flow theory to optimize the process of multiple adsorbents, establish a flow network model of the interaction between the adsorbent and the dye, where the nodes represent the adsorbents and the capacity of the edges represents the adsorption capacity, solve the maximum flow problem, find the combination of adsorbents, and generate the maximum flow optimization results; based on the maximum flow optimization results, perform maximization configuration to obtain the optimized adsorbent configuration; Import the optimized adsorbent configuration into the deep learning model, design and train a neural network to predict the recovery performance, refer to the influence of the adsorbent properties and dye concentration, and obtain the deep learning recognition model; evaluate the deep learning recognition model to check the accuracy of the model prediction and generate the model prediction evaluation results; Judge whether the deep learning recognition model is qualified according to the model prediction evaluation results, and apply the deep learning recognition model to predict the adsorption and recovery performance of the organic dye in the water body.
[0019] The steps to obtain the Monte Carlo simulation results are as follows: Obtain the dye concentration value, water body temperature value, and current pH value to get the triple information; According to the triple information, perform random perturbation on each group of inputs to generate simulation samples, and calculate the simulated adsorption value. The calculation formula is: ; where, is the simulated adsorption value, represents the dye concentration value collected in the th group, represents the temperature value collected in the th group, represents the pH value collected in the th group, represents the diffusion rate of dye molecules generated by the th perturbation, represents the local water body viscosity value after the th perturbation, represents the effective contact area between adsorbent particles in the th simulation, represents the specific polar activity value on the surface of the adsorbent particles, represents the shortest average path length between particles in the th group of samples during the simulation; Generate Monte Carlo simulation results by aggregating all the trends of adsorption performance changes and analyzing the stability based on the simulation output performance of the simulated adsorption values under different perturbation conditions.
[0020] Specifically, based on the collected dye concentration values, water temperature values, and current pH values, continuously collect water samples in the specified area and record the detected values each time. The dye concentration values are compared with the range of 0 mg / L to 400 mg / L, the water temperature values are compared with the range of 0 °C to 50 °C, and the pH values are compared with the range of 0 to 14. Whenever any detected value is within the corresponding range, include this detected value in the valid data list; otherwise, regard it as abnormal data and exclude it. Then, view all the valid data item by item and file them in chronological order. At the same time, use the quantitative photometry method for the dye concentration values to obtain the corresponding numerical indicators. After measuring the absorbance with a visible light spectrophotometer, calculate the specific concentration by combining the concentration curve calibrated by this instrument. Then, regularly scan the water temperature values with a standardized thermometer, compare each measured temperature value with the range of 0 °C to 50 °C. If the value is within this range, mark it as a qualified measurement point. At the same time, compare the collected pH values with the range of 0 to 14 through a pH meter, and store all those within this range as valid data. Subsequently, at the same time stamp, combine the concentration, temperature, and pH value at this moment in the form of "(concentration, temperature, pH)" to form complete triple information, and finally form a series of triple data arranged in time sequence for the subsequent invocation and operation of various algorithms to obtain the triple information.
[0021] Formula: , The advantage of the formula is that by introducing multiple parameters characterizing the microscopic characteristics of the water body and the adsorbent, key factors such as dye concentration, water temperature, pH value, local viscosity, and adsorbent surface activity are incorporated into a unified operation structure, which helps to more precisely depict the complex interaction relationships between molecules during the adsorption process, avoid error accumulation caused by ignoring some subtle effects, and make the subsequent adsorption performance analysis more accurate.
[0022] The steps for obtaining the parameters are as follows: First, measure the concentration of the water sample at the specified sampling point and measure the absorbance with a visible light spectrophotometer , and then, combined with the calibration curve of the photometer, through , convert the absorbance into the dye concentration value , where is the instrument calibration coefficient. The process of determining is based on the concentration sequence of the pre-prepared standard dye solution and records its absorbance. After obtaining the optimal fitting slope through the linear regression method, use it as . In an actual detection, if the absorbance , the instrument calibration coefficient , then there is .
[0023] The steps for obtaining the parameter are as follows: Use a digital temperature measuring device to scan the temperature of the current water sample in the water environment at a frequency of once per minute. After recording the stable reading, compare it with the pre-set temperature comparison range (0°C to 50°C). If the measured value is within this range, it is regarded as a valid temperature value. The specific value of is directly read from the result displayed on the measuring device. In one measurement, when the stable reading is 25°C,
[0024] The steps for obtaining the parameter are as follows: Measure the pH value of the water quality using a pH meter with a precision of 0.01. With a standard range of 0 to 14, the instantaneous pH value is obtained immediately after each reading. , if this value falls within the range of 0 to 14, it is included in the valid sample. In one test, the pH meter reading is 7.00, then .
[0025] The steps for obtaining the parameter are as follows: Determine the diffusion rate of dye molecules through molecular dynamics means. Combine the molecular movement paths collected by nanoscale tracking technology with time, and quantify the average distance traveled by each molecule per unit time. Introduce the following calculation: , where represents the displacement of the th molecular trajectory collected at the microscopic level, is the corresponding time interval, represents the total number of molecular trajectories collected. In one measurement, if the average displacement of molecular movement is 1.70 nanometers, the time interval is 10 nanoseconds, and a total of 200 trajectories are collected, then , since each trajectory performs consistently in the detection, the actual value is approximately .
[0026] The steps for obtaining the parameter are as follows: Measure the local viscosity of the water body under microscopic perturbation conditions using a viscometer. Place the water sample in a viscosity measurement container and maintain a fixed shear rate. Record the resistance value by rotating or vibrating methods, and then convert the resistance value into dynamic viscosity with reference to the viscosity conversion table. , and the specific conversion can use , where is the observed shear resistance, is the calibration coefficient. In one test, if , , then .
[0027] The steps for obtaining the parameter are as follows: After the adsorbent particles are evenly dispersed in the water environment, the effective contact area is counted by particle imaging analysis. A batch of adsorbent particles is selected and the field of view is segmented under an electron microscope. The overall contactable area on the particle surface is estimated by pixel counting, and then the effective contact area of a single particle is obtained by using a normalization formula. : , where is the pixel count value on the surface of the th particle, is the total number of particles counted in this time, is the conversion coefficient, which converts the number of pixels into the actual area. In one count, if pixels, , , then .
[0028] The steps for obtaining the parameter are as follows: The specific polar activity value of the adsorbent particles is evaluated by chemical analysis of the particle surface. , and the X-ray photoelectron spectroscopy method can be used to record the surface functional group composition and quantify the and distribution ratios. Then, with the following calculation formula: , where and are the reference coefficients, and represent the relative contents of O and N elements on the surface, represents the surface charge potential index, and are set with reference to the known characteristics of the adsorbent and determined by actual material test comparison. If in one determination , , , , , then .
[0029] The steps for obtaining the parameter are as follows: Image analysis and connectivity detection are performed on the adsorbent particle distribution. The shortest average path length between particles is counted by topological methods. The center points of the particles are regarded as nodes, and the adjacent distances between particles are regarded as edge weights. The shortest distances between all nodes are calculated respectively by the Floyd-Warshall algorithm and the average value is recorded as . If in one particle distribution map, the sum of all shortest paths obtained by the above algorithm is 108000 µm and the total number of nodes is 800, then the average path length: .
[0030] Calculation process: make , , , , , , , , Step 1: Calculate molecular diffusion and viscosity: ; ; This gives the denominator: ; Molecular part: ; So the first item: ; The second step is to calculate the adsorbent surface related terms: ; The third step is to sum and take the absolute value: ; The results show that under the specific parameter values selected, the calculated simulated adsorption value is about 44.48, which has a reference value for the subsequent evaluation of the adsorption performance of the dye under specific conditions. If the subsequent steps need to be compared with other combination conditions, the same calculation process can be used for numerical comparison.
[0031] Based on the simulated adsorption values obtained previously, the numerical sequence of the corresponding adsorption effect is output under different disturbance conditions and arranged in chronological order. All sequence data are compared one by one to see the decrease in dye concentration and the response changes of the corresponding parameters of the adsorbent. If there is a difference between the adsorption rate in any record and the average value of the surrounding records, a secondary test is performed to confirm the accuracy of the value. Then, statistical methods are used to calculate indicators such as the mean, variance and skewness of the data distribution, and the established fluctuation range threshold is used to measure the stability of the distribution. For example, when the variance of the adsorption rate sequence exceeds 3.5, it is considered to have a high fluctuation, and when it does not exceed 3.5, it is considered to have a stable fluctuation. Records with a variance greater than 3.5 are marked separately and listed together, and then the records in the remaining range are merged. Finally, the simulated output performance in the steady-state interval is summarized to form a list of adsorption performance change trends. These trend lists are then proofread one by one and attention is paid to whether local anomalies occur. If there are no anomalies, the overall trend data and its statistical analysis results are aggregated to obtain the Monte Carlo simulation results.
[0032] The steps to obtain the uncertainty analysis results are: Collect the Monte Carlo simulation results, organize the Monte Carlo simulation results, and form a data set; Based on the data set, calculate the uncertainty index under each condition. The calculation formula is: ; where, is the uncertainty index, is the simulated adsorption value, is the temperature value, is the pH value, is the concentration of suspended solids in the water body, is the surface area of the adsorbent, is the chemical reaction rate in the sample, is the dissolved oxygen content, is the porosity of the adsorbent, is the surface activity of the adsorbent; According to the uncertainty index, analyze the distribution of the recovery efficiency under different conditions to obtain the uncertainty analysis results.
[0033] Specifically, based on the obtained Monte Carlo simulation results, compare the time tags of all numerical values with the corresponding sampling conditions, select all data records within a fixed interval and confirm the integrity of their parameters. For example, compare the dye concentration in the range of 0 mg / L to 500 mg / L, the temperature in the range of 0 °C to 60 °C, the pH value in the range of 0 to 14, and the concentration of suspended solids in the range of 0 mg / L to 300 mg / L. When any detected item deviates from the above range, it is excluded. Then, concatenate all the confirmed valid data records in chronological order. Next, extract the dye molecule concentration, local water temperature, real-time pH value, and the simulated adsorption value associated with the original label of the adsorbent for each record, and check their corresponding time tags one by one to ensure no overlap or omission. Number and reorder the non-overlapping records uniformly, so that subsequent query and traceability links can perform continuous indexing. At the same time, match and compare the Monte Carlo simulation outputs of different times to check for entries with the same time tag but conflicting data content. If conflicting entries appear, trace back step by step according to the actual monitoring period and reread the relevant measurement data. After completing the investigation of all conflicting records, summarize these valid and consistent data into a multi-dimensional original mapping list. Finally, use this list to label the corresponding relationships between the dye concentration and the adsorbent-related information, and the temperature and the pH value to form a data set.
[0034] Formula: , The advantage of the formula is that by considering multiple key factors such as simulated adsorption value, temperature, pH, concentration of suspended solids in water, surface area of adsorbent, chemical reaction rate, dissolved oxygen content, and porosity and surface activity of adsorbent simultaneously, and integrating them into a hierarchical and mutually restrictive structure, it can more comprehensively depict the possible fluctuation scale of adsorption efficiency under different external and internal conditions. Thus, when evaluating the distribution range of recovery efficiency subsequently, it is easier to identify the situations with high-frequency fluctuations and conduct targeted treatment.
[0035] The steps for obtaining the parameter are as follows: use the simulated adsorption value obtained previously.
[0036] The steps for obtaining the parameter are as follows: continuously monitor the water environment with a thermometer, and after comparing several readings of the measured instantaneous temperature value, select a relatively stable number as .
[0037] The steps for obtaining the parameter are as follows: read the pH value corresponding to a certain period on a pH meter, and select the actual reading as within the range from 0 to 14. For example, when the pH meter reads 7.0 at the corresponding moment, then .
[0038] The steps for obtaining the parameter are as follows: measure the concentration of suspended solids in water and generate a numerical value denoted as , and use a turbidimeter or particle counting device to measure the same sample multiple times. Each time, a specific turbidity or particle content is obtained, and then use the following formula to convert the turbidity value to mg / L: , where is the value detected by turbidity, and are the reference coefficients calibrated in the laboratory and are obtained by fitting the previous test data. Taking one experiment as an example, if and , when the turbidity detection value , then , and then round it or retain decimals for recording to obtain .
[0039] The steps for obtaining the parameter are as follows: determine the surface area of the adsorbent and quantify it. Usually, during the material preparation stage, a high-precision specific surface area analysis device can be used to measure the specific surface area of the adsorbent by the gas adsorption method (BET method). For example, in one measurement, the specific surface area of a certain activated carbon-based adsorbent is recorded as , and the dosage is , then , after rounding the result, it can be made such that .
[0040] The steps for obtaining the parameter are as follows: First, record the time required for the chemical change of the dye solution at different times and the effective molecular weight participating in the reaction, and then according to the general formula for rate calculation: , where represents the change in the concentration of chemical products or intermediates per unit time, represents the time difference, which is calculated by detecting the actual increase or decrease of the sample chemical products at fixed intervals, and finally the chemical reaction rate is obtained. For example, in the time interval from 0 second to 60 seconds, it is observed that the concentration of a certain intermediate product changes from to , then .
[0041] The steps for obtaining the parameter are as follows: Use a dissolved oxygen measurement probe to scan the dissolved oxygen content in the water body second by second or minute by minute, and record the mg / L values collected in chronological order as . When the real-time dissolved oxygen displayed by the probe is , it can be recorded as .
[0042] The steps for obtaining the parameter are as follows: Measure the porosity of the adsorbent to characterize the internal void distribution. Generally, the pore volume is obtained by mercury intrusion method or nitrogen adsorption-desorption method, and then the ratio of the pore volume to the total volume of the adsorbent is used as the porosity, which is recorded as . For example, when mercury intrusion measurement is performed on a batch of activated carbon samples, the pore volume and the total volume of the material are obtained. Then , and it is recorded as .
[0043] The steps for obtaining the parameter are as follows: Analyze the surface activity of the adsorbent, which can usually be quantified by specific polar activity or surface charge distribution. Here, the method of statistically analyzing the surface chemical element distribution and the number of functional groups of the adsorbent is adopted, and it is obtained by the following formula , where and are reference coefficients, which are obtained by fitting the previous measured samples. is the actually detected surface activity energy value, is the standard activity energy value of the reference sample. Example: If , , , , then .
[0044] Calculation process: Step 1: Prepare measured and simulated parameters: ; Step 2: Calculate the numerator term and the denominator term: ; ; Therefore, the numerator part: ; Step 3: Denominator calculation: ; ; ; ; Therefore, the denominator part is: ; Step 4: Divide the numerator by the denominator and take the absolute value: .
[0045] This result indicates that the uncertainty index has a value of approximately 30.45 under this parameter combination. When this value is large, it means that there is a greater fluctuation range in the adsorption performance among different conditions. When further comparing the values obtained in other time periods and parameter combinations in the future, it is possible to intuitively determine which working conditions belong to the high-fluctuation range and need to be intensively monitored, and which working conditions can be regarded as stable intervals, so as to conduct a distribution analysis of the recovery efficiency in the subsequent links.
[0046] Based on the uncertainty index results obtained above, by comparing the fluctuations corresponding to different temperatures, pH values, suspended solid concentrations, as well as the surface area and porosity of the adsorbent for each time period or batch by batch, the correlation between each uncertainty index and the adsorption rate is identified one by one. The comparison range of the uncertainty index and the adsorption rate is selected as a biaxial interval from 0 to 200. Then, starting from the first record, the difference between the current uncertainty index value and the previous record is compared. If the difference exceeds 12, this record is marked as a significant fluctuation point; if the difference is within 12, it is marked as a gentle fluctuation point. Then, follow-up investigations are carried out for all significant fluctuation points, including screening for numerical mutations that may be caused by mismatched sampling time sequences and instantaneous imbalances caused by changes in the adsorbent dosage. After confirming that there are no missing records or abnormal values, a round of mean calculation and variance calculation are performed on all records to determine the overall fluctuation distribution. The relationship between the variance and 12 is evaluated. If the variance exceeds 12, a prompt is output indicating that the current overall fluctuation degree is high; if it does not exceed 12, it is considered that the distribution is relatively concentrated. On this basis, further median and range analyses are carried out. If the range is significantly higher than 10, it indicates that there are still discrete points caused by local conditions, and they are listed separately for subsequent inspection. Finally, all data are divided into multiple groups according to differences in temperature, pH, etc., and change charts of the corresponding recovery efficiency are generated through numerical comparison as an illustration of the recovery efficiency distribution, obtaining the uncertainty analysis results.
[0047] The steps for obtaining the maximum flow optimization results are as follows: Screen all the connecting edges with non-zero adsorption capacity between the adsorbent nodes and the dye nodes from the uncertainty analysis results, re-number all the adsorbent nodes according to the dye molecular type, and establish a network graph structure composed of adsorbent and dye nodes; According to the network graph structure, calculate the maximum adsorption flow value from the starting dye node to the ending adsorbent set. The calculation formula is: ; Among them, is the maximum adsorption flow, is the adsorption flux per unit area between the th dye node and the th adsorbent node, is the adsorption layer thickness of the th adsorbent node, is the pore number of the th adsorbent, is the binding frequency between the th dye node and the th adsorbent node, is the molecular polarity response value of the th dye node, A set of all node pairs with associated edges; According to the maximum adsorption flow rate, select the adsorption agent combination path with continuous path coverage and the strongest adsorption flow in the graph to generate the maximum flow optimization result.
[0048] Specifically, according to the uncertainty analysis results obtained previously, compare all adsorbent nodes and dye nodes, check the adsorption capacity parameter data between each pair of nodes, filter out the node connection edges with non-zero adsorption capacity, then create corresponding identifiers for each connection edge and map them to the previously integrated parameter list, record the original numbers of the adsorbent nodes and the original numbers of the dye nodes in sequence, and note the current adsorption capacity value in the pairing information. After that, confirm the molecular category of each dye node item by item with reference to the detailed list in the dye molecule type file, and check whether there is a matching item with this dye molecule type in the record of the adsorbent node. If the matching result and the non-zero adsorption capacity hold simultaneously, retain this connection edge and perform further statistics and path operations in the subsequent steps. In order to re-number all adsorbent nodes, it is necessary to scan the adsorbent nodes item by item according to the dye molecule category in this file, map the adsorbent nodes and the dye molecule categories to a set of mapping tables, then combine the old numbers of each adsorbent node and the corresponding dye types together to generate a new number allocation list, and then re-record all adsorbent nodes in a comparison table in the new number order, and sort them in ascending or descending order of the numbers in this table, retain the specific adsorption capacity data of each adsorbent node and the dye node and store them in the same data list. Finally, use this list to construct a network graph structure, regard the dye node as the starting end and the adsorbent node as the target end. In this network graph structure, the connection edges between all nodes carry corresponding adsorption capacity values, and these values will be used as the basic data for path selection and flow calculation in the subsequent steps. After integration, a network graph structure with node and connection information is obtained.
[0049] Formula: , the advantage of the formula is that by simultaneously considering the adsorption flux per unit area, the adsorption layer thickness, the number of pore diameters, and the difference in the polar binding frequency between molecules within the same calculation framework, the true upper limit of the adsorption flow rate of each edge can be effectively characterized, making the subsequent maximum flow search more targeted.
[0050] The steps for obtaining the parameters are as follows: First, use a flow meter to monitor the adsorption amount per unit area between the th dye node and the th adsorbent node in stages, record the adsorption rate of the dye and the rate of the adsorbent surface being covered by dye molecules during each time period, and then fit the monitoring data to a unit area adsorption flux curve by linear regression, obtain the slope of this curve and denote it as ; for example, in a single centralized monitoring, the measured adsorption rate per unit area fluctuates between and . After multiple measurements, a fit is obtained as , indicating that under this condition, every of the adsorbent surface can fix and adsorb dye molecules per minute.
[0051] The steps for obtaining the parameter are as follows: slice or section the adsorbent for testing, obtain the adsorption layer thickness through actual measurement, and then calculate the average value by integrating the measurement results at multiple points and record it as . During the measurement, several equally spaced positions are selected on the surface of the adsorbent and within its penetration layer, and values are obtained by microscopic observation or slice thickness recording. At the same time, multiple batches of adsorbents are repeatedly measured through the same process. If in a certain adsorbent sample, the average of the measured values at all selected positions can reach , then it is recorded as .
[0052] The steps for obtaining the parameter are as follows: count the number of pores in the adsorbent. Here, the pore size imaging detection and segmentation calculation method is used, and the total number of pores is obtained by counting and recording each pore region one by one through software. For example, in a certain test, a total of 350 identifiable pores are found in the adsorbent sample, obtaining .
[0053] The steps for obtaining the parameter are as follows: adopt the method of comparing molecular dynamics simulation with experiments to measure the molecular binding frequency between the rd dye node and the th adsorbent node. First, conduct a small-scale adsorption test on this control group under laboratory conditions, and at the same time simulate the possible collision and binding processes between dye molecules and the adsorbent surface in the computer. The number of successfully bound molecules is counted every fixed time, and then this value is divided by the simulation time to obtain the frequency of molecular binding occurrence. For example, in a 600-second simulation test, it is observed that a total of 450 dye molecules form stable bonds with the adsorbent molecules, then it can be recorded as .
[0054] The steps for obtaining the parameter are as follows: mainly determine its specific value by detecting the polar response degree of the dye molecules themselves. During the experiment, an external electric field or other excitation conditions are applied to the dye solution, and the displacement or orientation change degree of the dye molecules is measured to quantify their reaction intensity to the polar environment. After numericalizing this intensity value, it is recorded as , for example, placing the dye solution in Under the electric field, monitor the proportion of dye molecules whose orientation changes per unit time, through the following formula , where is the reference coefficient calibrated according to the electric field strength, is the proportion of the number of dye molecules whose orientation changes per unit time, is the total number of dye molecules. In one measurement, set , , , then we can get .
[0055] Calculation process: First step, list the required parameter values: ; Second step, calculate the numerator term: ; Third step, calculate : ; Fourth step, calculate the absolute value term : ; Fifth step, sum and then take the square root: ; Get: ; Because this formula will traverse all the connected edges in the network diagram structure and take the minimum value after obtaining a series of values, so only taking the edges of this example as an example, the result is approximately 0.2828.
[0056] This result shows that under the given conditions of adsorption flux per unit area, adsorption layer thickness, number of pore diameters, molecular binding frequency, and dye polarity response value, the upper limit of the maximum achievable adsorption flow is approximately 0.2828. When selecting the adsorption agent combination path subsequently, the values of all candidate edges will be compared and the minimum flow will be taken as the measurement index. If this value is higher than 0.3, it means that this connected edge has greater flow potential in the adsorption channel. If it is lower than 0.1, it indicates that the adsorption capacity may be relatively limited and other edges need to be combined to increase the total flow.
[0057] Based on this maximum adsorption flow value, start selecting each feasible path in the network diagram and compare the flow upper limits of all connecting edges along each path one by one. Record the paths with coverage continuity between all nodes and extract the minimum flow values in each path. Then, calculate the set of these minimum flow values for all feasible paths, and select the path with a larger value from this set as a candidate path with a higher adsorption flow. If the flow value of any connecting edge in a path is significantly lower than 0.1, it may not be able to support a higher overall flow, so it will be regarded as a flow limiting factor in the subsequent path comparison. Finally, number these candidate paths and write them into the comparison table to form a matching relationship with the previously marked dye nodes and adsorbent nodes, and obtain the optimal path list of the adsorbent combination through one batch calculation and comparison, thereby generating the maximum flow optimization result.
[0058] The steps to obtain the optimized adsorbent configuration are as follows: According to the adsorbent combination paths output in the maximum flow optimization result, extract the numbers of all adsorbent nodes and the path structures of the corresponding dye nodes, reconstruct the mapping relationship between the adsorption paths and the adsorbent indexes, and generate a set of candidate adsorbents; Based on the set of candidate adsorbents, calculate the flow ratios and the degree of overlap of node distributions borne by each adsorbent in the combination path, arrange and sort the adsorption paths and avoid node conflicts, and generate a preliminary configuration structure; According to the preliminary configuration structure, match the deployment order and space occupation of each adsorbent under the corresponding dye path to generate the optimized adsorbent configuration.
[0059] Specifically, according to the adsorbent combination path output in the maximum flow optimization result, first read all the path index and node number information attached to the result, disassemble the dye node number and adsorbent node number associated with each path, and compare these numbers in the previously recorded flow data set to confirm that each number does exist in the completion identification list range of 0 to 9999. If it is outside the range, it is marked as invalid data and verified again through independent detection rules. Then, the nodes of the adsorption path are sorted according to the relationship between the number and the flow mark, and all the dye nodes and adsorbent nodes with good sequence connection are taken as candidate combination items. Multiple consecutive node information in the same path are compiled into a valid one. The node chain is identified to avoid confusion between path numbers and node numbers. The internal pressure monitoring records in the range of 0MPa to 2MPa are compared with the temperature detection results. Any pressure fluctuation point exceeding the preset threshold is recorded and bound to the path number to form a prompt. These binding prompts are then summarized. If there are more than 5 fluctuation points, additional observations are made and attempts are made to skip the priority selection of the path in subsequent steps. Finally, a mapping list of adsorption paths and adsorbent indexes is generated based on the number matching results. The position of each adsorbent node in its path and the dye node information connected to it are recorded, and all adsorbent nodes are integrated into an overall set to obtain a set of candidate adsorbents.
[0060] According to the set of candidate adsorbents, the actual flow value and node coordinates are read from the tag data of each adsorbent node and the corresponding path. The working current of the adsorbent is compared item by item in the range of 0A to 5A. If it exceeds 4A, a prompt for subsequent detection is recorded. If it is lower than 4A, the node is directly included in the normal range. Then, the flow ratio of each node in all recorded combined paths is counted. The calculation method is to divide the adsorption flow of the node in the path by the sum of the adsorption flow of all nodes in the path. If the flow ratio of an adsorbent node is higher than 45%, it is recorded as a dominant node. If it is lower than 5%, it is recorded as an auxiliary node. Then, the distribution of different nodes at the spatial coordinate level is evaluated. The calculation method of the distribution overlap can refer to the entropy threshold setting method for interval division. First, a coordinate array m is obtained for the node coordinate set, and then the average value m0 of the coordinate array is calculated. Then a coordinate threshold T=m0 is set to check the proportion of the number of nodes corresponding to the coordinate values greater than T and the coordinate values less than T. If the corresponding number of nodes accounts for more than 50%, it is considered that the distribution overlaps significantly, otherwise it is considered that the distribution is discrete. The results obtained after processing these distribution overlaps are used to arrange and sort the adsorption paths, and the nodes with high overlap and large flow ratio differences in the same area are separated to avoid path conflicts. Finally, a preliminary configuration structure is output and the node conflict avoidance situation is recorded.
[0061] According to the preliminary configuration structure, assign specific positions to each adsorbent in the list of dye paths with the deployment areas marked, and compare the position coordinates with parameters such as the outer diameter size and shape characteristics of the adsorbent. The comparison method is to match level by level according to the outer diameter range from 0 mm to 1000 mm. If the outer diameter of the adsorbent exceeds 500 mm, it is paired with the column marked "Large Space Requirement". If the outer diameter is less than 500 mm, it is matched with the "Small and Medium-sized Space Requirement" to form a deployment order index. During the continuation of the process, read the space occupancy data of the adsorbent in this order and proofread the interval distance between dye paths of the same type to avoid the phenomenon of node conflict caused by being too close. If the interval distance between multiple adsorbents monitored on the same path is less than 100 mm, output a prompt and make a readjustment. If it is still less than 100 mm after adjustment, temporarily record the adsorbent at the end of the candidate queue. Finally, after completing the correspondence between all nodes and space occupancy, generate a fixed position number for each adsorbent and record the deployment order to obtain the optimized adsorbent configuration.
[0062] The steps to obtain the deep learning recognition model are as follows: Call the optimized adsorbent configuration, import the adsorbent node number and space deployment information as structured input parameters into the initial input layer of the deep neural network. Based on the pore size distribution, surface area and surface chemical composition parameters of the adsorbent, construct the input vector and perform dimension standardization to obtain the neural network input vector matrix; According to the neural network input vector matrix, set the dye concentration parameter as a dynamic input variable, construct a training sample sequence, perform parameter initialization, forward propagation and error backpropagation, and synchronously record the error signals in different sample rounds to establish a neural network recovery performance prediction structure; According to the neural network recovery performance prediction structure, perform multiple rounds of iterative convergence detection, combine the adsorbent parameters and dye concentration sample performances in different rounds, screen the stable output nodes and freeze the corresponding connection weights to generate the deep learning recognition model.
[0063] Specifically, call the optimized adsorbent configuration, and import the adsorbent node numbers and spatial deployment information as structured input parameters into the initial input layer of the deep neural network. First, disassemble all the adsorbent numbers and their corresponding deployment coordinates or size data in the configuration file to ensure that these numbers are all between 0 and 99999 in the record to avoid number conflicts. Then, retrieve the pore size distribution range, the calculated surface area value in square micrometers, and the surface chemical composition parameters for each number. By comparing with the previously recorded adsorbent surface test results, confirm whether it falls within the reasonable range of 0 µm² to 100000 µm². Mark the individuals with a surface area higher than 100000 µm² as large-area adsorbents and those lower than 100 µm² as small-area adsorbents. Subsequently, arrange these values into an input vector in a predetermined order, and perform dimensional normalization for all adsorbents. The specific method is to first find the maximum and minimum values of each dimension, and then through linear mapping of each data, make it fall into the standard range of 0 to 1, so as to maintain the balance between input features. For individual nodes where outliers may occur, confirm whether there are measurement errors by reading the previously obtained pore size detection records. If it is confirmed that there are measurement errors, use a reference value for substitution during dimensional normalization and record the reason for substitution in detail. Finally, jointly form the neural network input vector matrix with each adsorbent node number and its mapped feature vector. The rows of this matrix correspond to the adsorbent numbers, and the columns correspond to information such as pore size distribution, surface area, and surface chemical composition. Complete the backend check to ensure that the matrix shape is consistent with the size defined by the initial input layer of the network. If there is a dimension mismatch, adjust the column order during the mapping process and generate the matrix again to obtain the neural network input vector matrix.
[0064] According to the neural network input vector matrix, set the dye concentration parameter as a dynamic input variable. First, segment and label the dye concentration range from 0 mg / L to 500 mg / L, and select several representative concentration points within each segment. Combine these representative concentrations with the adsorbent vector matrix to form a training sample sequence. Each training sample contains an adsorbent vector and a corresponding dye concentration value. Then, perform parameter initialization. During the initialization process, set the connection weights of each layer of the neural network to small random values, and the bias term can be set to a small value between 0 and 0.1 to avoid large fluctuations in the gradient in the initial state. Subsequently, perform the forward propagation process. Starting from the input layer, transfer the adsorbent vector and dye concentration data of each sample to the hidden layer together, and output a temporary prediction value through layer-by-layer linear transformation plus activation function. Compare this prediction value with the target value already registered in the record to obtain an error signal and perform gradient calculation. Correct the calculated gradient on the connection weights and biases of all layers according to the set learning rate. If the error signal generated in this round of training exceeds 10, an additional attention mark is made. If it is lower than 10, it is regarded as the normal range. Parallelly record the error signals and training progress in different sample rounds. Continuously repeat this process. At the end of each iteration, update the neural network parameters and recalculate the new error signal. Through continuous cycling, realize error backpropagation correction and establish a neural network recovery performance prediction structure.
[0065] According to the neural network recovery performance prediction structure, perform multiple rounds of iterative convergence detection. The specific method is to input the same training samples into the network multiple times again and observe whether the error signal gradually decreases after each round of training. If it is found that the error signal fluctuates between 5 and 8 and no longer decreases significantly, it is judged that the network may be close to the convergence state. It is necessary to further confirm whether the performance of each adsorbent parameter and dye concentration sample in different batches tends to be stable. If there are still large error fluctuations, continue to supplement training data and increase the number of training rounds. At the same time, observe the activation value distribution of each node in the network after each round of training. Select the nodes with stable outputs for screening, retain the connection weights of these nodes and freeze them in subsequent training so that they no longer change with the gradient update, and gradually eliminate some nodes or connections with large fluctuations, focusing the learning on the stable area, thereby reducing unnecessary parameter jitter. When the error signal for several consecutive rounds remains below 5, it is regarded as the convergence completed. Finally, save the frozen weights as the finalized parameters, record the number of nodes and connection conditions of all layers for subsequent verification or inference stage calls, and generate a deep learning recognition model.
[0066] The steps to obtain the model prediction evaluation results are as follows: Obtain the output results of the deep learning recognition model, compare and match each group of output results with the corresponding dye recovery results, extract the three basic variables of the timestamp, predicted value, and actual value of each group of predicted values, and generate a model prediction and actual result set; Based on the model prediction and actual result set, calculate the matching accuracy index value. The calculation formula is: ; Among them, is the matching accuracy index value, is the model prediction output value of a single sample, is the true recovery value at the corresponding timestamp, is the length of the current sample data dimension; Based on the matching accuracy index value, traverse the prediction and pairing records of all samples, screen the proportion of sample points below the error threshold and record the evaluation fluctuation range, and generate a model prediction evaluation result.
[0067] Specifically, to obtain the output results of the deep learning recognition model, compare and match each group of output results with the corresponding dye recovery results, extract the three basic variables of the timestamp, predicted value, and actual value of each group of predicted values, and generate a model prediction and actual result set. First, read the predicted values output by the deep learning recognition model at each moment from the record and mark their timestamps. Then, find the true recovery values with the same timestamp item by item from the previously recorded dye recovery data list. Combine the predicted values and the true recovery values to form a complete record and assign a unified number to it. After summarizing all the numbered records, view them by time period. For example, divide all timestamps by minutes and confirm whether there are multiple unreasonable predicted values repeating within the same time period. If it is found that there are multiple predicted values at the same timestamp and the difference exceeds 5, then conduct an investigation to check whether the real-time acquisition values of the adsorbent dosage and dye concentration fluctuate during this time period. Ensure the integrity of the finally paired data through such item-by-item comparison and screening. Subsequently, count the difference between the "predicted value - actual value" in each record. If the difference is within the range of 0 to 1, it is regarded as a record with a small deviation. If it exceeds 1, it is marked in a difference check list as a possible key attention object for follow-up. Then, sort these numbered records by multiple fields and output them in the order of the timestamp. If it is found that there is a large jump in the timestamp within the range of crossing hours, add a prompt beside it to explain that it may be due to the interruption of the detection device resulting in a discontinuous timeline. Incorporate these prompts and records with large deviations into the inspection scope list. Repeatedly check the timestamp and predicted value to confirm whether the corresponding dye recovery result is incorrect. Incorporate the confirmed correct data into the final available set to obtain the model prediction and actual result set.
[0068] Formula: , The advantage of the formula is that by combining various operation forms such as exponential function, cube root, logarithmic operation, and trigonometric function, it can depict the deviation between the predicted value and the actual value from multiple perspectives, not only measuring the degree of prediction deviation but also paying attention to the overall change range at extreme values and small values.
[0069] The steps for obtaining the parameter are as follows: collect a single predicted value output by the deep learning recognition model at a specific timestamp, record this value in the recycling prediction list, and then check it against other monitoring data corresponding to the same time node. If it is confirmed to be between 0 and 100, it is recorded as , if the value exceeds 100, further check whether it is an outlier according to the upper limit information of the dye dosage obtained previously. For example, in a detection, the model output is 45.3, then .
[0070] The steps for obtaining the parameter are as follows: previously, a time series record of the dye recycling results has been made, and each record contains the measured recycling value. By comparing the timestamps, a one-to-one mapping between the measured recycling value and the model output is obtained to get , if the recycling value is monitored to be 52.1 at t = 300 seconds in a on-site detection, it is recorded as .
[0071] The steps for obtaining the parameter are as follows: determine the value by checking the data dimension length of the current sample. When each record consists of multiple items such as "timestamp, predicted value, actual value, adsorbent number, environmental temperature, pH", its data dimension may reach 6 or more. At this time, is 6 or greater; if it is confirmed that each record contains 7 fields in an actual application scenario, then .
[0072] Calculation process: The first step: Substitute the above parameters into the formula: Let , , The second step: Calculate the exponential terms in blocks: ; ; ; ; The third step: Calculate the trigonometric function terms: ; ; ; Step 4: Denominator calculation: ; ; Step 5: Summarize and take absolute value: ; ; This result shows that under this example condition, the matching accuracy index value of a single sample It is about 20.85, which can be calculated with other samples. For comparison, if some records If it is significantly greater than 20, it means that the prediction error is higher or the value fluctuates more violently. If it is less than 5, it means that the prediction result is relatively closer to the actual value. By comparing with multiple samples, the overall matching accuracy distribution can be formed, providing a calculation basis for the next step of counting the proportion of samples below the error threshold.
[0073] Based on the matching accuracy index value, traverse the prediction and pairing records of all samples, screen the proportion of sample points below the error threshold and record the evaluation fluctuation range. First, read the previously calculated one by one value and compare it with the threshold value 10. The value is compared with the threshold. If it is less than 10, it is marked as "small deviation", and if it is higher than 10, it is marked as "large deviation". Then, combined with the timestamp attribute, the number of records with small deviations and the number of records with large deviations in each minute are counted at intervals of minutes. The minute segments with large deviations are marked as fluctuation points, and these fluctuation points are listed on the timeline according to the fluctuation point concentration calculation formula. If there are large deviation records for three consecutive minutes in a certain time period, a prompt is output for subsequent viewing. If most time periods are concentrated in the range of small deviations, it means that the overall fluctuation range is relatively stable. After all data are counted, an overview of the error distribution is obtained and the record list is retained. The list can be used to locate the prediction deviation under the specific timestamp, and finally generate the model prediction evaluation result.
[0074] The qualification of the deep learning recognition model is judged according to the model prediction evaluation results. The steps of using the deep learning recognition model to predict the adsorption and recovery performance of organic dyes in water are as follows: According to the model prediction evaluation results, the neural network training rounds that failed the evaluation were screened out, and the model version with the lowest output error and stable fluctuation range was locked to generate a deep learning recognition model that can be used for prediction; According to the deep learning recognition model available for prediction, new adsorbent configuration and dye concentration input parameters are imported, and the structure mapping and prediction calculation process are executed to generate the prediction results of the adsorption and recovery performance of organic dyes in water bodies.
[0075] Specifically, according to the model prediction evaluation results, the neural network training rounds with unqualified evaluations are screened out, and the model version with the lowest output error and a stable fluctuation range is locked. First, the error signal records of all neural network training history rounds are read and compared with the preset error upper limit, which is obtained from multiple rounds of experiments and the statistics of dye recovery deviation. The rounds with an error signal exceeding 30 are recorded as high-error rounds, and the specific numbers of these rounds are located in chronological order. The corresponding neural network parameters of this batch of records are counted and moved to a separate list, so that these rounds with larger errors can be ignored during subsequent retrieval. At the same time, the rounds with an error within 30 are grouped into the available range. Then, by accessing the comparison table of adsorbent dosage and corresponding dye concentration, the deviation curves of these rounds within the available range are analyzed in segments at different dye concentration intervals. If the fluctuation is greater than 8 in any concentration range, the round is marked as having a large fluctuation for the next step of locking the rounds with stable fluctuations. After traversing all available rounds, the rounds with all fluctuations below 8 are concentrated and screened out. Then, the group with the smallest error value among these rounds is selected for data extraction, and its number is uniformly recorded in the alternative list and its position in the global comparison table is checked. If the same number shows values close to or below 5 in multiple concentration segments, it is regarded as the final available model version. After that, the network structure and frozen weight parameters are exported from this model version file and combined with the previously registered iteration number information to form a single stable neural network structure, obtaining a deep learning recognition model available for prediction.
[0076] According to the deep learning recognition model available for prediction, new adsorbent configuration and dye concentration input parameters are imported, and the structure mapping and prediction calculation process is executed. First, a suitable list of adsorbent numbers is searched in the record, and their pore diameters, surface areas, and surface chemical activities are read item by item. If the surface area of an adsorbent exceeds 400 µm², it is recorded as a large area type; if the pore diameter count is less than 50, it is regarded as a small pore quantity type. These values are packaged and combined with the new dye concentration range into an input batch, and filled into the input layer of the network in matrix form according to the input vector order defined in the previous deep learning recognition model. Immediately afterwards, each data row is calculated forward in sequence with a time resolution of 0.01 seconds. When calculating, the frozen network weights and bias values are read to perform sequential weighting and superposition on the input. At the same time, the operation results of the activation functions of all layers are recorded in the corresponding log for troubleshooting sudden data anomalies. After each batch of data is calculated, the output is immediately summarized, and the output values are recorded in pairs with the adsorbent label and time stamp. If it is detected that any of the predicted result values exceeds 150, it is added to the list shown for subsequent monitoring. Then, all the predicted values within the normal range are arranged and output in chronological order to obtain the prediction result of the adsorption and recovery performance of water body organic dyes.
[0077] The above is only a preferred embodiment of the present invention, and it does not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for predicting the adsorption recovery performance of organic dyes in water based on machine learning, characterized in that: The following steps are involved: Collecting dye concentration, temperature and pH value of organic dyes in water, performing multiple random samplings using Monte Carlo simulation to simulate the adsorption process and generate Monte Carlo simulation results; performing uncertainty analysis based on the Monte Carlo simulation results to obtain the recovery efficiency distribution under various conditions and generate uncertainty analysis results; Based on the uncertainty analysis results, the network flow theory is applied to optimize the process of multiple adsorbents, and a flow network model of the interaction between adsorbents and dyes is established. The nodes represent the adsorbents, and the capacity of the edges represents the adsorption capacity. The maximum flow problem is solved, the adsorbent combination is found, and the maximum flow optimization result is generated; based on the maximum flow optimization result, the maximum configuration is performed to obtain the optimized adsorbent configuration; Importing the optimized adsorbent configuration into the deep learning model, designing and training a neural network to predict recovery performance, and obtaining a deep learning recognition model with reference to the effects of adsorbent properties and dye concentration; evaluating the deep learning recognition model, checking the accuracy of model prediction, and generating a model prediction evaluation result; Whether the deep learning recognition model is qualified is judged according to the model prediction evaluation results, and the deep learning recognition model is applied to predict the adsorption and recovery performance of organic dyes in water.
2. The method for predicting the adsorption recovery performance of organic dyes in water based on machine learning according to claim 1 is characterized in that: The steps for obtaining the Monte Carlo simulation results are: Get the dye concentration value, water temperature value and current pH value to obtain triplet information; According to the triple information, each set of input is randomly perturbed to generate simulated samples, and the simulated adsorption value is calculated. The calculation formula is: ; in, is the simulated adsorption value, Representative The dye concentration values collected by the group, Representative The temperature value collected by the group, Representative The pH value collected by the group, Representative The diffusion rate of dye molecules generated by the secondary disturbance is Representative The local water viscosity value after the disturbance is: Representative The effective contact area between adsorbent particles in each simulation is Represents the specific polar activity value of the adsorbent particle surface, Represents the simulation time the shortest average path length between particles in the group sample; According to the simulated output performance of the simulated adsorption value under different disturbance conditions, all adsorption performance change trends are collected and the stability is analyzed to generate Monte Carlo simulation results.
3. The method for predicting the adsorption recovery performance of organic dyes in water based on machine learning according to claim 1, characterized in that: The steps for obtaining the uncertainty analysis results are: Collecting the Monte Carlo simulation results, and organizing the Monte Carlo simulation results to form a data set; Based on the data set, the uncertainty index under each condition is calculated using the following formula: ; in, is the uncertainty index, is the simulated adsorption value, is the temperature value, is the pH value, is the concentration of suspended matter in water, is the adsorbent surface area, is the chemical reaction rate in the sample, is the amount of dissolved oxygen, is the porosity of the adsorbent, is the surface activity of the adsorbent; According to the uncertainty index, the distribution of recovery efficiency under different conditions is analyzed to obtain uncertainty analysis results.
4. The method for predicting the adsorption recovery performance of organic dyes in water based on machine learning according to claim 1, characterized in that: The steps for obtaining the maximum flow optimization result are: Selecting all the connection edges between the adsorbent nodes and the dye nodes that have non-zero adsorption capacity from the uncertainty analysis results, re-labeling all the adsorbent nodes according to the dye molecule type, and establishing a network graph structure consisting of the adsorbent and dye nodes; According to the network graph structure, the maximum adsorption flow value from the starting dye node to the end adsorbent set is calculated using the following formula: ; in, is the maximum adsorption flow rate, For the The dye node and The adsorption flux per unit area between adsorbent nodes is For the The adsorption layer thickness of the adsorbent node, For the The number of pore sizes of the adsorbent, For the The dye node and The binding frequency between adsorbent nodes, For the The molecular polarity response value of each dye node, is the set of all node pairs with associated edges; According to the maximum adsorption flow rate, an adsorbent combination path with continuous path coverage and the strongest adsorption flow is selected in the graph to generate a maximum flow optimization result.
5. The method for predicting the adsorption recovery performance of organic dyes in water based on machine learning according to claim 1, characterized in that: The steps for obtaining the optimized adsorbent configuration are: According to the adsorbent combination path output in the maximum flow optimization result, the numbers of all adsorbent nodes and the path structure of the corresponding dye nodes are extracted, the mapping relationship between the adsorption path and the adsorbent index is reconstructed, and a set of candidate adsorbents is generated; According to the set of candidate adsorbents, the flow ratio and node distribution overlap of each adsorbent in the combined path are calculated, the adsorption paths are arranged and sorted, and node conflicts are avoided to generate a preliminary configuration structure; According to the preliminary configuration structure, the deployment order and space occupancy of each adsorbent in the corresponding dye path are matched to generate an optimized adsorbent configuration.
6. The method for predicting the adsorption recovery performance of organic dyes in water based on machine learning according to claim 1, characterized in that: The steps for obtaining the deep learning recognition model are: Calling the optimized adsorbent configuration, importing the adsorbent node number and spatial deployment information as structured input parameters into the initial input layer of the deep neural network, constructing and dimensionally normalizing the input vector based on the pore size distribution, surface area and surface chemical composition parameters of the adsorbent, and obtaining a neural network input vector matrix; According to the neural network input vector matrix, the dye concentration parameter is set as a dynamic input variable, a training sample sequence is constructed, parameter initialization, forward propagation and error feedback are performed, and error signals under different sample rounds are synchronously recorded to establish a neural network recovery performance prediction structure; According to the neural network recovery performance prediction structure, multiple rounds of iterative convergence detection are performed, and the adsorbent parameters and dye concentration sample performance in different rounds are combined to screen stable output nodes and freeze the corresponding connection weights to generate a deep learning recognition model.
7. The method for predicting the adsorption recovery performance of organic dyes in water based on machine learning according to claim 1, characterized in that: The steps for obtaining the model prediction evaluation results are as follows: Obtain the output results of the deep learning recognition model, pair and compare each set of output results with the corresponding dye recovery results, extract the three basic variables of the timestamp, predicted value and actual value of each set of predicted values, and generate a set of model predictions and actual results; Calculate the matching accuracy index value based on the model prediction and the actual result set; Based on the matching accuracy index value, the prediction and pairing records of all samples are traversed, the proportion of sample points below the error threshold is screened and the evaluation fluctuation range is recorded to generate the model prediction evaluation result.
8. The method for predicting the adsorption recovery performance of organic dyes in water based on machine learning according to claim 1, characterized in that: According to the model prediction evaluation result, it is judged whether the deep learning recognition model is qualified, and the steps of applying the deep learning recognition model to predict the water body organic dye adsorption recovery performance are as follows: According to the model prediction evaluation results, the neural network training rounds that failed the evaluation were screened out, the model version with the lowest output error and stable fluctuation range was locked, and a deep learning recognition model that can be used for prediction was generated; According to the deep learning recognition model that can be used for prediction, new adsorbent configuration and dye concentration input parameters are imported, and the structure mapping and prediction calculation process is executed to generate the prediction results of the water body organic dye adsorption recovery performance.
9. A device for predicting the adsorption and recovery performance of organic dyes in water, characterized in that: include: A processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, so that the device for predicting the adsorption and recovery performance of organic dyes in water performs the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program or instruction, and when the computer program or instruction is executed, the method according to any one of claims 1 to 8 is implemented.