Method and system for constructing resin adsorption performance prediction model
By constructing a resin adsorption performance prediction model, utilizing quantum chemical calculations and molecular descriptors, and combining the Langmuir adsorption isotherm model and machine learning, the problem of insufficient interpretability in resin adsorption performance prediction is solved. This enables rapid screening and rational design of resin materials in complex pollution scenarios, improving prediction accuracy and engineering application value.
Patent Information
- Application Number
- CN202511663255.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies lack a continuous and measurable quantitative characterization of the pollutant-material compatibility relationship for predicting resin adsorption performance. The models are not interpretable enough, making it difficult to guide rational material design and unable to quickly respond to material selection in complex pollution scenarios.
A predictive model for resin adsorption performance is constructed. Continuous numerical characteristics of pollutants are built through quantum chemical calculations and molecular descriptors. Resin microstructure characteristics are extracted, and polarity matching parameters and size matching parameters are introduced. The Langmuir adsorption isotherm model is used to screen high-quality target variables and train a machine learning model. Interpretive analysis methods are combined to provide a range of material design parameters.
This enables rapid screening and rational design of resin materials in complex pollution scenarios, improves the accuracy and interpretability of adsorption performance prediction, provides quantitative guidance, and enhances the value of engineering applications.
Smart Images

Figure CN121506330A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of adsorption materials, and more specifically, to a method and system for constructing a predictive model of resin adsorption performance. Background Technology
[0002] Porous polymer resins, possessing both designable pore structures and tunable surface functional groups, have been widely used for the adsorption and selective removal of various organic pollutants in complex aquatic bodies. Their adsorption behavior is jointly governed by molecular-material interactions (dispersion / π–π, electrostatic and ion exchange, hydrogen bonding) and mass transfer-confinement effects (micropore ratio, average pore size and distribution, pollutant diffusion characteristics, hydrophilic / hydrophobic partitioning), and is highly sensitive to environmental conditions (pH, temperature, etc.). These factors are coupled across multiple scales, exhibiting high-dimensional nonlinear relationships and significant individual differences. Therefore, in material design and process optimization, constructing a modeling and decision-making framework that is interpretable (consistent with adsorption priors), transferable (extrapolating new pollutants / materials), and implementable (outputting quantitative correspondences of material-pollutant characteristics and parameter control windows) under multi-mechanism, multi-scale, and multi-variable conditions remains a core challenge in this field.
[0003] The concept of a materials genome library emphasizes using standardized, generalizable, high-quality data as the core to construct a cross-scale description and evaluation system for "material structure-application performance" for data-driven material design and optimization. However, in the specific scenario of resin adsorption, existing characterization and modeling practices still fall short of the above concept. For example, CN 118430675 A (publication date 2024-08-02, application number 202410525794.7, applicant Shanghai Jiao Tong University) uses "micropore / mesopore / macropore" to classify pore structure, "strong acid / strong base / nonionic polar / nonionic nonpolar" to characterize surface chemistry, and binary Morgan fingerprints to describe pollutants. This paradigm is simple to implement and semantically intuitive, but it has three fundamental limitations when used for mechanism modeling and design guidance: 1. High cardinality classification and discretization lead to learning and generalization losses: One-Hot encoding results in dimensionality expansion and data sparsity, making the model prone to excessive splitting in the category dimension and memorizing identity codes; moreover, discretizing continuous quantities such as micropore proportion, average pore size, and surface polarity degenerates the decision boundary into piecewise constants, weakening the model's extrapolation ability. 2. Molecular fingerprints are difficult to be isomorphic to the real mechanism: fingerprints only characterize the adjacency relationship of substructures and lack continuous physicochemical quantities directly related to adsorption (such as polarizability, molecular dynamic diameter Dp, pH-dependent logD, polarity / solvation parameters), limiting the quantitative analysis of van der Waals interactions, hydrophilic-hydrophobic partitioning, charge matching, and size confinement. 3. Lack of cross-domain matching relationships: Adsorption essentially depends on the quantitative adaptation of "pollutant-material", such as polarity matching, size matching, charge complementarity, etc. Existing description systems are unable to compress such couplings into continuous, measurable indicators that can be directly used for design, which restricts the stable identification and design translation of thresholds and windows. Summary of the Invention
[0004] To address the lack of continuous and measurable quantitative characterization of pollutant-material compatibility in existing technologies for predicting resin adsorption performance, this application provides a method and system for constructing a resin adsorption performance prediction model, offering quantitative guidance for rapid screening and rational design of resin materials in complex pollution scenarios.
[0005] One aspect of this application provides a method for constructing a resin adsorption performance prediction model, comprising: S1, acquiring pollutant molecular structure data, resin characterization data, and adsorption equilibrium experimental data of different resin-pollutant combinations; S2, extracting features from the pollutant molecular structure data and resin characterization data respectively to obtain a pollutant feature set and a resin feature set; S3, calculating the cross-domain fitting parameters of the pollutant feature set and the resin feature set to obtain a fitting parameter set; S4, using the Langmuir adsorption isotherm model to perform nonlinear fitting on the adsorption equilibrium experimental data of different resin-pollutant combinations to obtain the maximum adsorption capacity of each resin-pollutant combination. Based on maximum adsorption capacity goodness of fit Perform screening to determine the goodness of fit. Maximum adsorption capacity greater than a preset threshold S5: Construct a training set based on the pollutant feature set, resin feature set, adaptation parameter set, and target variable; S6: Use the training set to train the machine learning model to obtain the adsorption performance prediction model.
[0006] Furthermore, it also includes: S7, using interpretability analysis methods to perform mechanistic analysis on the adsorption performance prediction model, obtaining the effect of each characteristic parameter in the pollutant characteristic set, resin characteristic set, and adaptation parameter set on the maximum adsorption capacity. The influence of the resin and the optimal range of each structural parameter in the resin feature set are used as the material design parameter range.
[0007] Furthermore, feature extraction is performed on the pollutant molecular structure data, including: performing molecular dynamics simulations to obtain the lowest energy conformation as the stable conformation of the pollutant molecular structure data; calculating the quantum chemical parameters of the stable conformation using density functional theory to obtain the electronic structure parameters of the pollutant molecular structure data; extracting geometric features of the stable conformation using molecular descriptors to obtain the spatial structure parameters of the pollutant molecular structure data; and calculating the hydrophobicity of the stable conformation under a preset pH condition using a partition coefficient prediction algorithm to obtain the octanol-water partition coefficient at the target pH. ; Including electronic structure parameters, space structure parameters and , as a pollutant characteristic set.
[0008] Furthermore, electronic structure parameters include polarizability and polar surface area; spatial structure parameters include molecular dynamics diameter. Solvent-accessible surface area, number of rotatable bonds, number of hydrogen bond donors, and number of hydrogen bond acceptors.
[0009] Furthermore, feature extraction is performed on the resin characterization data, including: using integrity verification rules to detect missing values in the resin characterization data, and selecting resin data with numerical records for all characterization parameters as a subset of complete characterization data; using condition consistency rules to screen the experimental conditions of the complete characterization data subset, and selecting resin data with the same test temperature, test pH, and pretreatment method or within the allowable deviation range as a subset of comparable characterization data; using box plot method or 3σ criterion to detect outliers for each characterization parameter in the subset of comparable characterization data, and removing outliers to obtain characterization data after removing outliers; and performing parameter analysis on the characterization data after removing outliers to obtain the resin feature set.
[0010] Furthermore, the resin characteristic set includes: ion exchange capacity, micropore ratio, micropore specific surface area, total specific surface area, pore volume, and average pore size. Resin surface polarity and functional group charge parameters.
[0011] Further, in S3, cross-domain adaptation parameters of the pollutant feature set and resin feature set are calculated to obtain an adaptation parameter set, including: extracting from the pollutant feature set... Extract from resin feature set ,calculate and The difference is used to obtain the polarity matching parameter. ,in, Extracting molecular dynamic diameters from pollutant feature sets Average pore size extracted from resin feature set ,calculate and The ratio of the two values yields the size matching parameters. ,in, ;Will and As an adaptation parameter set.
[0012] Furthermore, in S4, the Langmuir adsorption isotherm model was used to perform nonlinear fitting on the adsorption equilibrium experimental data of different resin-pollutant combinations, using the following formula: ;in, To balance the adsorption capacity, For equilibrium concentration, KL is the Langmuir adsorption constant. This represents the maximum adsorption capacity.
[0013] Furthermore, machine learning models include gradient boosting decision trees or random forest algorithms.
[0014] Another aspect of this application provides a system for constructing a predictive model of resin adsorption performance, comprising: a data acquisition module for acquiring pollutant molecular structure data, resin characterization data, and adsorption equilibrium experimental data of different resin-pollutant combinations; a feature extraction module for extracting features from the pollutant molecular structure data and resin characterization data respectively to obtain pollutant feature sets and resin feature sets; an adaptation parameter calculation module for calculating cross-domain adaptation parameters of the pollutant feature sets and resin feature sets to obtain an adaptation parameter set; and a target variable generation module for using the Langmuir adsorption isotherm model to perform nonlinear fitting on the adsorption equilibrium experimental data of different resin-pollutant combinations to obtain the maximum adsorption capacity of each resin-pollutant combination. Based on the maximum adsorption capacity goodness of fit Perform screening to determine the goodness of fit. Maximum adsorption capacity greater than a preset threshold The target variable is used as the training set; the training set construction module constructs the training set based on the pollutant feature set, resin feature set, adaptation parameter set and target variable; the model training module uses the training set to train the machine learning model and obtain the adsorption performance prediction model.
[0015] Compared to existing technologies, the advantages of this application are: To address the shortcomings of existing technologies in predicting resin adsorption performance, such as the lack of continuous and measurable quantitative characterization of pollutant-material compatibility, insufficient model interpretability leading to difficulties in guiding rational material design, and inability to quickly respond to material selection needs in complex pollution scenarios, this application provides a method for constructing a resin adsorption performance prediction model. This method enables the construction of continuous numerical characteristics of pollutants through quantum chemical calculations and molecular descriptors, the extraction of resin microstructure characteristics through multi-layer data quality control, and the introduction of polarity matching parameters. Matching parameters with size This study quantitatively analyzes the compatibility between pollutant molecules and resin structures, uses the Langmuir adsorption isotherm model to screen high-quality target variables and train a machine learning model, and transforms the characteristic influence laws identified by the model into material design parameter ranges with clear physical meaning through interpretability analysis. This provides quantitative guidance for the rapid screening and rational design of resin materials in complex pollution scenarios, and improves the accuracy, interpretability and engineering application value of adsorption performance prediction. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall technical process for the implementation of this application; Figure 2 The core functional modules of the end-to-end computing platform constructed for this application; Figure 3 The statistical distribution of key properties of adsorbent materials and target pollutants included in the database created for this application; Figure 4 This shows the distribution of LogD characteristics of pollutants in the training and test sets of this application; Figure 5 The results are performance evaluations of different models in this application; Figure 6 This is a comprehensive evaluation graph of the performance of the two models in this application; Figure 7 The results of SHAP analysis of the XGBoost-BASE model constructed based on the traditional discretization representation method in this application are shown. Figure 8 The results of SHAP value calculation and feature importance analysis based on the XGBoost-NEW model in this application are shown. Figure 9 This application presents a quantitative structure-activity relationship between adsorption performance and material structural parameters. Figure 10 This application clarifies the optimal operating range for multi-parameter coordinated regulation through interaction dependency analysis; Figure 11 This application presents a high-throughput virtual screening process based on optimized windows with predetermined parameters. Detailed Implementation
[0017] The present application will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0018] Example 1 like Figure 1 As shown, this application constructs a gene library of resin adsorption materials that integrates prior adsorption knowledge and adaptive parameters. It systematically builds a complete computational path encompassing screening, kinetic simulation, and quantum computing, and integrates comprehensive material characterization data, establishing a cross-domain correlation parameter system between molecules and materials. Based on this, multiple regression models are used to deeply analyze the adsorption mechanism, and methods such as SHAP and PDP are combined to quantify characteristic contributions, identify key action regions and patterns, and finally complete application verification under complex pollution scenarios, forming a complete technical closed loop from mechanism analysis to engineering application.
[0019] Step S1: Construct a novel feature description system based on prior knowledge of adsorption.
[0020] The core of this step is to overcome the limitations of traditional discrete classification descriptors and construct a continuous, numerical feature description system that covers the main adsorption mechanisms. Its specific implementation sub-steps include: S1.1: Calculation of the entire chain of pollutant molecules and acquisition of numerical descriptors.
[0021] For each target pollutant molecule, the following standardized computational procedure is performed to obtain high-precision continuous quantum chemical and physicochemical descriptors. Figure 2 The core functional modules and computing processes of the end-to-end computing platform constructed in this application are further clearly demonstrated, such as... Figure 2As shown, (A) molecular models were established for styrene-divinylbenzene copolymers with different monomer ratios, and their polarity logD parameters were calculated (A-1, A-2); (B) modeling and polarity logD calculation were performed for acrylate copolymers with different monomer ratios (B-1, B-2); (C) the dynamic diameter of the molecule was accurately calculated based on the electron density isosurface method (C-1, C-2); (D) the polar surface area and charge distribution characteristics were obtained through surface electrostatic potential analysis (D-1, D-2). This platform realizes integrated analysis of the entire process from molecular structure modeling to physicochemical parameter calculation, providing key technical support for the creation of a description system for the adsorption mechanism of resin materials and pollutants.
[0022] Initial structure acquisition and optimization: Three-dimensional molecular structures were obtained from authoritative databases such as PubChem and ChemSpider. Molecular dynamics conformational search and preliminary geometric optimization were performed to determine the lowest-energy stable conformation as the initial structure. Based on density functional theory, precise geometric optimization and frequency calculations were conducted on the stable conformation at an appropriate basis set (e.g., 6-311G**) level to ensure the structure was located at a potential energy minimum and to obtain accurate electronic wavefunctions. Post-processing, including key descriptor extraction and electronic density isosurface analysis, resulted in the batch extraction of numerical descriptors directly related to the adsorption mechanism. These descriptors were shown to have a significant impact on adsorption capacity. These include: Van der Waals descriptor: polarizability, used to quantify the ability of a molecule's electron cloud to deform under an electric field, and is directly related to the strength of the dispersion interaction (van der Waals force); Electrostatic and hydrogen bonding descriptors include contaminant charge, total surface area, polar surface area, and the sum of the number of hydrogen bond donors and acceptors. Hydrophilicity / hydrophobicity and solubility descriptors: Calculate or query the octanol-water partition coefficient and water solubility of pollutants at the pH of the adsorption environment to assess their partitioning trend between the aqueous phase and the resin phase; Spatial structure and dynamics descriptors: Molecular dynamics diameters are obtained through calculations to evaluate size exclusion effects; the number of rotatable bonds and rings is statistically analyzed to indirectly reflect the rigidity and conformational flexibility of molecules.
[0023] S1.2: Standardized collection and organization of microstructure descriptors for resin adsorbent materials.
[0024] To construct a high-quality gene bank, resin characterization and adsorption isotherm data from scientific literature were first systematically collected. Strict inclusion criteria of "complete characterization, comparable conditions, and meeting quality standards" were established for efficient data cleaning. After systematically reviewing approximately 3600 isotherm data from literature, only about 700 high-standard data were included in the gene bank. For key data in image form (such as pore size distribution curves), precise image-to-data conversion was performed using specialized digitization software to ensure data quality and computability. Finally, continuous microstructure descriptors whose importance had been validated in models were extracted from all standardized data. These descriptors mainly include: BET specific surface area, micropore specific surface area, micropore area ratio, total pore volume, and BET average pore size characterizing pore structure; ion exchange capacity, functional group charge properties, and material surface logD characterizing surface chemistry; and resin particle size characterizing physical morphology. For the Langmuir adsorption isotherms presented in the literature as images, the Qmax value of each set of resin-pollutant pairing data was extracted. To ensure the reliability of the performance label data, a strict goodness-of-fit threshold (determination coefficient R²>0.95) was set. Finally, the verified Qmax value was used as the key quantitative indicator characterizing the adsorption performance of the resin-pollutant pairing and was designated as the performance label for training the machine learning model.
[0025] Step S2: Define and calculate key cross-domain quantitative adaptation parameters.
[0026] This step is the core innovation in connecting contaminants and materials to achieve quantitative matching. Based on the basic descriptor obtained in S1, the following key fitting parameter is calculated: Size matching parameter (Dpratio): Dpratio = Average pore size of the resin / Molecular dynamic diameter of the contaminant. This parameter directly quantifies the degree of size matching between molecules and pores.
[0027] Polarity matching parameter (LogDDifference): LogDDifference = resin surface logD - contaminant logD. This parameter directly indicates the degree of hydrophilicity / hydrophobicity matching between the contaminant and the resin surface.
[0028] Step S3: Construct a machine learning model for predicting adsorption performance.
[0029] Dataset Construction and Preprocessing: All numerical features obtained from S1 and S2 (including pollutant quantum chemical descriptors, resin microstructure descriptors, and cross-domain adaptation parameters) and corresponding adsorption performance data (the maximum adsorption capacity Qmax of the Langmuir adsorption isotherm) were integrated into a structured dataset. Subsequently, the dataset underwent rigorous preprocessing to ensure data quality and model training effectiveness. This primarily included: outlier detection and handling using statistical and model-based methods; and Z-score standardization of all continuous numerical features to eliminate dimensional differences and ensure the stability of model training.
[0030] exist Figure 3 The document illustrates the main properties of the materials and contaminants in this application's database. It uses a combination of box plots and frequency heatmaps to simultaneously display the distribution range and frequency, showcasing the richness and data boundaries of the resin gene library. Clearly defining the data boundaries will also guide subsequent reverse engineering using the resin gene library. Figure 3 As shown, A-1 to A-6 reveal the core parameters of the adsorbent material, covering structural characteristics (specific surface area, average pore size, micropore specific surface area and its proportion) and surface chemical characteristics (ion exchange capacity, surface polarity logD). B-1 to B-6 present the characteristic indicators of pollutants, including molecular size (kinetic diameter, volume, solvent-accessible surface area), physicochemical properties (logD, ion exchange capacity), and their size matching relationship with the material (kinetic diameter to material average pore size ratio). This figure systematically characterizes the breadth of coverage and data quality of the database, laying a solid foundation for subsequent structure-activity relationship studies.
[0031] Stratified Sampling Dataset Partitioning Strategy: To address potential model evaluation biases caused by uneven sample sizes of different pollutants within the dataset, this application employs a stratified sampling strategy to partition the training and test sets. The core of this strategy lies in selecting a key feature highly correlated with adsorption performance as the stratification basis (such as the hydrophilicity / hydrophobicity parameter LogD of the pollutant at the system's pH), and stratifying according to its numerical distribution. This ensures that the partitioned training and test sets have a statistical distribution highly similar to the original dataset in this feature space. Figure 4 As shown, after this method of partitioning, the mean and standard deviation of the pollutant LogDpH characteristics in the training and test sets are extremely close, and the overall distribution pattern also shows a high degree of consistency. This not only proves the representativeness and rationality of the dataset partitioning, but also provides a reliable guarantee for the model to obtain unbiased generalization performance evaluation. It indicates that the dataset partitioning effectively maintains the statistical representativeness of the samples and avoids the problem of decreased model generalization ability caused by data distribution shift.
[0032] Model training and validation: such as Figure 5As shown, ten machine learning models were used to train the adsorption structure-activity relationship in this dataset (including XGBoost, CatBoost, KNNR, LightGBM, RF, MLP, SVM, Ridge, ElasticNet, and LASSO). Triple-fold cross-validation was used to prevent overfitting, grid search was used to determine the optimal hyperparameter configuration, and indicators such as the coefficient of determination and root mean square error were used to comprehensively guide the hyperparameter optimization process and evaluate the model's predictive performance. Figure 5 A indicates that gradient boosting decision trees (such as XGBoost and CatBoost) have strong learning and prediction capabilities for current data systems. Among them, XGBoost has the best prediction performance (R² = 0.856, MAE = 0.299). This is because XGBoost's second-order gradient optimization can accurately determine the splitting direction, effectively capture complex nonlinear relationships and feature interactions, and its built-in L1 / L2 regularization effectively mitigates overfitting.
[0033] Figure 5 Section B compares the modeling performance of the two descriptor construction strategies on the current dataset. XGBoost-BASE and CatBoost-BASE adopt traditional descriptor construction strategies, using the categorical features of resin materials and the molecular structure fingerprint features of pollutants; while XGBoost-NEW and CatBoost-NEW are new descriptor construction strategies proposed in this application. These strategies introduce prior knowledge of adsorption, construct comprehensive numerical features, and build adaptive parameters.
[0034] S4: Verification of the effectiveness of the descriptor system. To empirically demonstrate the significant advantage of the comprehensive numerical feature system constructed in this application in improving the learning ability of the adsorption model mechanism, a traditional descriptor system was designed as a verification experiment. The specific steps are as follows: A baseline model based on traditional discretization and classification characterization methods is constructed as a comparison baseline: On the resin side, qualitative classification labels of "micropore, mesopore, macropore" are used to replace the continuous micropore ratio and average pore size descriptors in this database, and type labels such as "strong base, strong acid, nonionic nonpolar, nonionic polar" are used to replace the continuous surface polarity and ion exchange capacity descriptors, and textual classification variables such as "resin monomer name" and "functional group name" are added; On the contaminant side, binary Morgan fingerprints generated by the RDKit package in Python are used to replace the quantum chemical descriptors obtained by full-chain calculation; Subsequently, all the above classification variables are processed using one-hot encoding; Finally, the same machine learning algorithm as in this application is used to perform large-scale hyperparameter search and tuning on the same training set to construct the baseline models XGBoost-BASE and CatBoost-BASE. The evaluation is based on two dimensions: quantitative comparison of predictive performance and qualitative comparison of mechanistic explanation ability. Performance differences are quantified by comparing key metrics on the test set (see [link to relevant documentation]). Figure 6 ), Figure 6 A presents detailed evaluation results of the prediction performance of the optimal model XGBoost-NEW, constructed based on the new representation system. Figure 6 B presents detailed evaluation results of the predictive performance of the XGBoost-BASE model constructed based on traditional discretization characterization methods. Furthermore, the consistency between the model's decision logic and the adsorption physicochemical mechanism is assessed by analyzing feature importance ranking and partial dependencies (see [link to relevant documentation]). Figure 7 This fully verifies the superiority of the descriptor system in this application. Figure 7 The results of SHAP analysis of the XGBoost-BASE model constructed based on the traditional discretization representation method are presented.
[0035] S5: Construction of a Comprehensive Mechanism Explanation System. To achieve the leap from model prediction to material design, this application constructs a comprehensive mechanism explanation system. This system not only reveals global laws but also guides material optimization in specific scenarios. Specifically, it includes the following steps: S5.1 Global Feature Importance Analysis: Using the SHAP value of the trained optimal machine learning model, the contribution of all input features (including basic descriptors and adaptation parameters) is ranked (see [link to relevant documentation]). Figure 8 This involves identifying key decision-making factors in the model and verifying the consistency between the model's learning decision-making basis and the adsorption of prior knowledge.
[0036] S5.2: Nonlinear Relationship Analysis and Threshold / Platform Identification of Key Features. Based on the key features identified in S5.1, further in-depth analysis is conducted using Partial Dependency Graphs (PDPs) (see [link]). Figure 9 This allows for the precise quantification of the nonlinear relationship between eigenvalues and adsorption performance, and the automatic identification of performance thresholds and performance plateaus.
[0037] S5.3: Design rule translation and material design window generation. The nonlinear relationships and thresholds / plateaus parsed in S5.2 are standardized and translated into quantitative rules and a visual design window for resin material design. Design rules are derived by analyzing cross-domain adaptation parameters such as size matching and polarity matching.
[0038] Generate a two-dimensional design window plot with key adaptation parameters as coordinate axes and adsorption performance dependence values as contour lines (see...). Figure 10 The figure clearly marks the "optimal range" corresponding to high adsorption performance and the "avoidance range" where performance deteriorates sharply, providing intuitive decision support for material design.
[0039] Example 2 The following Example 2 will further verify the advanced nature and application effectiveness of the resin material gene library of this application.
[0040] first, Figure 7 By comparing the determination coefficients and residual distributions of the XGBoost-BASE and XGBoost-NEW models, it is shown that the baseline model (XGBoost-BASE) built based on traditional discretization and classification representation methods has significant limitations in learning ability and poor generalization performance (test set R2=0.622, MAE=0.487). Further mechanistic interpretation of the model is achieved through SHAP value analysis and feature importance ranking. Figure 7 The study found that the model relied too heavily on the apparent classification features of pollutant molecular fingerprints, and failed to deeply understand the intrinsic physicochemical parameters that affect the adsorption process. It had significant shortcomings in the identification of adsorption mechanisms and the understanding of the contributions of key parameters, and could not effectively capture the core mechanisms that affect the adsorption process.
[0041] In contrast, based on the same raw data and modeling methods, the fully numerical feature representation system driven by prior adsorption knowledge proposed in this study significantly improves the predictive performance of the XGBoost-NEW model for adsorption behavior (test set R2=0.856, MAE=0.299). Figure 9As shown, the model further demonstrates a multi-layered and in-depth understanding of the adsorption mechanism, specifically in the following ways: First, the model identifies the polarizability of pollutants as the most critical predictor (SHAP = 0.3864). Polarizability is directly related to the strength of the intermolecular London dispersion force, and its dominance strongly demonstrates that in the organic-resin adsorption system involved in this study, non-specific interactions driven by van der Waals forces constitute the dominant mechanism of the adsorption process. As the second dominant factor, the resin ion exchange capacity directly points to the electrostatic interaction mechanism, indicating that the model has successfully learned that for charged pollutants, ion exchange between them and oppositely charged functional groups on the resin surface is the key pathway to achieve efficient adsorption. Furthermore, the model also accurately resolves the synergistic effects of porous structure parameters. A series of micropore-related structural descriptive parameters (micropore surface area ratio, micropore area, pore volume, and average pore diameter) collectively exhibited high importance (ranking 3rd, 6th, 7th, and 14th, respectively). Among them, the importance of the micropore surface area ratio (0.1304) was even higher than that of the total specific surface area (0.1120). This subtle insight indicates that although the total specific surface area is often used to measure the total number of available adsorption sites, micropores, due to their stronger adsorption potential field and size confinement effect, are more critical for the adsorption of size-matched pollutant molecules, thus revealing that micropore filling is an important adsorption mechanism in this system. Subsequently, the model effectively captured and learned the comprehensive regulation of interfacial chemical interactions. Important features such as pollutant hydrophilicity / hydrophobicity, polar surface area, solvent-accessible surface area, and resin-pollutant polarity matching parameters (LogD, ranked 6th, 8th, 10th, and 13th respectively) were jointly identified, indicating that the model can comprehensively consider the synergistic regulation of adsorption behavior by multiple interfacial forces, including hydrophobic and polar interactions. Finally, cross-domain adaptation features also played a significant role in the model's correct decision-making, with the pore size-molecular diameter adaptation ratio (Dpratio, SHAP = 0.0578, ranked 11th) and the surface polarity matching parameter (LogDDifference, SHAP = 0.0567, ranked 13th) both contributing significantly. These findings suggest that by precisely controlling the structural and polarity matching relationship between materials and pollutants, adsorption selectivity can be effectively enhanced, providing a feasible path for the rational design of high-affinity resin materials.
[0042] The above results collectively demonstrate that the adsorption prior knowledge-driven characterization system constructed in this study enables the model to decode the hierarchical adsorption mechanism: it not only reveals the dominant mechanism, consistent with physicochemical laws, sequentially from van der Waals interactions to electrostatic / ion exchange interactions to micropore filling interactions to interfacial chemical interactions, but also verifies the feasibility of using cross-domain adaptive parameters to achieve quantitative control and rational design of the adsorption process. Thus, the value of the fully numerical characterization system proposed in this study extends beyond improved predictive performance; it provides the model with a cognitive framework consistent with the common sense of adsorption science, successfully reconstructing the adsorption mechanism from generalized non-specific interactions to specific interactions.
[0043] Furthermore, by employing single-feature partial dependency graphs, we gained a deeper understanding of the marginal effects of these important features on model predictions, and systematically revealed the influence of key parameters on adsorption performance, thus providing quantitative guidance for the rational design of resin adsorption materials. Figure 9 A shows the partial dependence analysis results of the ratio of resin pore size to the molecular dynamic diameter of pollutants (Dpratio). When Dpratio increases from 3.5 to 8.5, the adsorption performance shows a stepwise increase, but the adsorption effect rapidly decreases when Dpratio > 8.5. This indicates that there is a clear optimization window between pore size and molecular size. Too small a Dpratio restricts the effective diffusion of pollutant molecules within the resin channels, while too large a Dpratio significantly weakens the superposition effect of the adsorption potential field provided by the pore walls. This result establishes a precise critical point for resin pore size design from a data-driven perspective for the first time. For the polarity matching parameter (LogDDifference), its partial dependence exhibits a unique bipolar characteristic, with peak adsorption performance appearing near -3. Figure 9 (B) It should be noted that this conclusion is based on the current framework for estimating resin surface polarity. The calculations are primarily based on the monomer composition and functional group type of the resin copolymer, achieved through quantum chemical methods, and have not yet systematically incorporated errors that may arise from complex factors such as pore confinement effects, microstructure, and the actual coverage of functional groups. Nevertheless, under the existing calculation system, this bipolar phenomenon still has clear guiding significance: by selectively adjusting the monomer ratio and functional group structure of the copolymer, regulating the LogDDifference between the resin and pollutants to a generalized optimization range around -3 remains a feasible strategy to significantly improve adsorption performance under current understanding. It also provides preliminary parameter basis for moving from empirical screening to rational design of material surface polarity.
[0044] The analysis results of the resin's conventional parameters (BET specific surface area, pore volume) show that ( Figure 9 C Figure 9 D), When the BET specific surface area increased from 500 m² / g to 850 m² / g, the adsorption performance showed a rapid upward trend, confirming that increasing the number of available adsorption sites is a direct and effective way to improve adsorption capacity. The curve then entered a plateau region, indicating a significant diminishing returns effect. Therefore, the optimization threshold for this parameter was controlled within the range of 500–850 m² / g. When the pore volume increased from 0.2 cm³ / g to 0.6 cm³ / g, the adsorption performance significantly improved, indicating that sufficient pore space provided the necessary transport channels for pollutant molecule diffusion. However, when the pore volume exceeded 0.6 cm³ / g, the adsorption performance decreased instead of increasing. This is mainly because an excessively large average pore size leads to a significant reduction in specific surface area and weakens the micropore filling mechanism dependent on a strong adsorption potential field. Therefore, the broad optimization range for pore volume was initially determined to be between 0.2–0.6 cm³ / g. Analysis results of the resin's microporous structure (micropore specific surface area, micropore ratio) showed (…). Figure 9 E, Figure 9 (F) Micropore area is a key determinant of adsorption capacity, and adsorption performance increases with increasing micropore area, especially in the range of 50 m² / g-100 m² / g where the increase is most rapid. This finding is importantly complementary to the optimal window for micropore surface area (around 0.165). Micropore area determines the total basis of adsorption potential, while micropore surface area reflects the efficiency dimension of pore structure. Both need to be closely considered in the design.
[0045] This section, through systematic single-feature partial dependency analysis, establishes for the first time a multi-parameter synergistic quantitative design window for resin adsorption materials from a data-driven perspective: Dpratio should be controlled between 3.5 and 8.5 to balance mass transfer and adsorption potential; LogDDifference may have an extreme value matching optimization range around -3; BET specific surface area (500–850 m² / g) and pore volume (0.2–0.6 cm³ / g) control the number of adsorption sites and mass transfer channels, respectively; micropore area and micropore surface area ratio jointly define the "total amount-efficiency" synergistic criterion of micropore structure. These findings provide quantitative basis and theoretical support for the transition of adsorption materials from empirical screening to rational design.
[0046] Based on univariate analysis, this study further employs dual-feature interactive partial dependency analysis to systematically examine the combined effects of synergistic interactions among key parameters on adsorption performance, in order to deeply analyze the coupling mechanism of multiple parameters in the adsorption process. Figure 10A demonstrates the synergistic regulatory effect of BET specific surface area and pore volume on adsorption performance. Analysis shows that when the BET specific surface area is below 500 m² / g, the adsorption effect is generally insufficient, reflecting a fundamental limitation on the number of available adsorption sites. However, when the BET specific surface area is increased to the optimized range of 500-850 m² / g, pore volume exhibits a crucial regulatory role: maintaining it within the range of 0.4-0.6 cm³ / g rapidly achieves the maximum adsorption capacity under this specific surface area condition. Based on this, a BET specific surface area of 500-850 m² / g and a pore volume of 0.4-0.6 cm³ / g are proposed as the basic framework for high-performance resin materials. Figure 10 B reveals the synergistic regulation mechanism of micropore ratio and pore size matching parameters on adsorption performance. Analysis shows that the optimal size matching window dynamically evolves with changes in micropore ratio: when the micropore ratio is in the moderate range of 0.15-0.225, the resin maintains excellent adsorption performance within a wide Dpratio range (3-8.5), demonstrating good broad-spectrum adaptability. This is mainly due to the ideal micropore-mesopore composite structure of the material in this range, which ensures sufficient high-energy adsorption sites while maintaining efficient mass transfer channels. With further increases in micropore ratio, the system exhibits a complex nonlinear response: the optimal Dpratio range first shrinks to a narrow range of 7-8, and then expands to a wider window of 4-8.
[0047] Figure 10 C revealed the interactive influence of the pollutant's logD value and the polarity matching parameter LogDDifference on the adsorption effect. The study showed that the system exhibited optimal adsorption performance when the logD difference remained in the negative range of -4 to -2, a phenomenon particularly pronounced for hydrophobic pollutants with logD > 0.5 and hydrophilic pollutants with logD < -1.0. This indicates that, under the current resin polarity estimation system, designing the resin with moderate hydrophilic properties (its logD value is approximately 2-4 units lower than the target pollutant) can effectively construct adsorbent materials with broad adaptability, especially suitable for practical water treatment scenarios with a wide range of pollutant polarities. Figure 10 D further analyzed the synergistic effect of micropore structure parameters and determined the optimal design window to be: micropore area > 50 m² / g and micropore ratio maintained in the range of 0.15-0.225. This result indicates that while ensuring a sufficient total number of micropore adsorption sites, controlling the micropore ratio within this specific range can achieve the best balance between adsorption capacity and mass transfer efficiency.
[0048] Based on univariate analysis, a dual-feature interactive partial dependency analysis was used to further investigate the synergistic effects among parameters, ultimately confirming the parameter window for material optimization. First, a basic material framework was established: the optimal combination of BET specific surface area of 500-850 m² / g and pore volume of 0.4-0.6 cm³ / g was determined. Further, the dynamic coupling relationship between micropore proportion and Dpratio was revealed: when the micropore proportion was 0.15-0.225, Dpratio performed excellently across a wide range of 3-8.5. Next, regarding the control of interfacial properties, the logD difference between -4 and -2 was identified as a key optimization window, particularly suitable for contaminants with logD > 0.5 and logD < -1.0, indicating that the moderately hydrophilic resin design has broad applicability. Finally, the optimal design window for the microporous structure was confirmed: micropore area > 50 m² / g and micropore proportion of 0.15-0.225.
[0049] Example 3 The following example 3 will further verify the practical application effect of the method of this patent. For the complex system of multiple pollutants coexisting in the wastewater of antibiotic production in pharmaceutical plants, relying on the constructed high-quality gene library and the adaptive parameter-driven framework, the parallel screening and multi-objective optimization of a large number of virtual resin candidate materials can be completed within a few hours, and the list of preferred material structures that can simultaneously achieve synergistic and efficient removal of multiple pollutants in the scene can be quickly output.
[0050] Pharmaceutical antibiotic production wastewater is a typical complex system with multiple pollutants coexisting. It mainly contains antibiotics such as penicillin and cephalosporins, organic solvents such as ethanol and dimethyl sulfoxide, and unreacted raw materials such as phenylacetic acid and aminophenylacetic acid. These pollutants differ significantly in molecular structure, polarity, and size, forming a complex characteristic of broad polarity distribution and significant size differences, making it difficult for traditional adsorption materials to achieve synergistic and efficient removal.
[0051] Based on the adsorption material gene library and machine learning model constructed in this patent, the following steps are implemented: Step 1: Pollutant Characterization. Using the full-chain computing platform constructed in this application, we obtain continuous quantum chemical descriptors for each pollutant: polarizability, molecular dynamic diameter, hydrophilic / hydrophobic partition coefficient logD at the target pH, polar surface area, solvent-accessible surface area, number of rotational bonds, number of hydrogen bond donors, and number of hydrogen bond acceptors.
[0052] Step Two: Based on the model interpretation and analysis results from the previous section, the initial search window for the resin material is set according to the properties of the pollutants: BET specific surface area: 500-850 m² / g, pore volume: 0.4-0.6 cm³ / g, micropore area: greater than 50 m² / g, micropore surface area ratio: 0.15-0.225, Dpratio range: 3-8.5, LogDDifference range: -4 to -2. Furthermore, optimization weights are set according to the pollution removal requirements of each pollutant, and a resin performance synergistic removal scoring parameter is constructed based on the weights: Q = aQmax1 + bQmax2 + cQmax3 + dQmax4.
[0053] Step 3: Using a high-performance computing cluster, 16,000 virtual resin candidate materials generated based on the resin parameter retrieval window were evaluated in parallel. A grid search combined with Bayesian optimization was used to output the optimal resin structure combination. The materials were synthesized according to the structural requirements and the optimal materials were experimentally verified. The results showed that the optimal materials achieved a 92% removal rate of penicillin, 88% of cephalosporins, 83% of phenylacetic acid, and 85% of dimethyl sulfoxide.
[0054] Example 4 (1) Construct a continuous and measurable pollutant and resin characteristic system. Pollutant characteristics: Use quantum computing descriptors and molecular dynamics calculations to generate the physicochemical characteristics of pollutants, such as polarizability, molecular dynamics diameter, logD, polar surface area, solvent-accessible surface area, number of rotational bonds, number of hydrogen bond donors / acceptors, etc., to ensure continuous and physically consistent parameters that are highly correlated with adsorption behavior.
[0055] Resin characteristics: By characterizing the resin's micropore ratio, micropore specific surface area, pore volume, average pore size, ion exchange capacity, and numerical surface polarity, the resin's microstructure and surface chemical characteristics are quantified into continuous parameters.
[0056] Cross-domain adaptation parameters: construct polarity matching parameter logDDifference, size matching parameter Dpratio, charge complementarity, etc., to explicitly resolve the matching relationship between contaminants and resin, and provide low-dimensional, designable quantitative variables for model input.
[0057] (2) Establish an interpretable model under the framework of adsorption mechanism. Model selection: adopt a supervised learning model with interpretable output (such as gradient boosting tree) and construct a model of the relationship between resin and pollutants through training.
[0058] Mechanism explanation: By analyzing the contribution of SHAP values, PDP curves, and platform description features, the threshold / platform / monotony interval is automatically identified, and combined with cross-domain parameter analysis, it is transformed into actual design parameters.
[0059] Design translation: Transform the mechanistic findings of the model (such as thresholds and plateaus) into a material-condition design window and develop standardized, reproducible experimental guidelines to support rapid validation.
[0060] (3) Rapid response design of materials in complex pollution scenarios, demand identification and target setting: Based on the complex pollution scenario, extract the pollutant structure list and clarify the design target, and formalize it into a multi-objective optimization problem.
[0061] Large-scale virtual combination screening: Large-scale virtual combination screening is conducted through a materials gene library, and multi-dimensional optimization and adjustment are performed within the design window to output candidate resin materials that can achieve simultaneous and efficient removal of multiple pollutants.
[0062] Rapid Response and Validation: By combining Bayesian optimization methods to optimize the resin structure, material solutions that meet the requirements of complex pollution scenarios are quickly generated, and the material gene library is continuously optimized by experimentally verifying and writing back to the gene library.
[0063] (1) This application establishes a mature and referable method for constructing quantum chemical descriptors for organic adsorption, and successfully overcomes the limitations of traditional discretization and categorization methods by constructing a fully numerical feature characterization system driven by prior adsorption knowledge. The prediction accuracy of the machine learning model is significantly improved, with the coefficient of determination on the test set increasing from 0.622 to 0.856 and the mean absolute error decreasing from 0.487 to 0.299. More importantly, the model demonstrates a deep insight into adsorption mechanisms, accurately learning and deconstructing the adsorption mechanism hierarchy of "van der Waals forces - electrostatic interactions - micropore filling - interfacial chemical interactions", providing a new paradigm for understanding complex adsorption processes from a data perspective.
[0064] (2) This application establishes a rational design framework for resin materials and determines the optimization window for key parameters. For the first time, it establishes a multi-level optimization design framework for resin adsorption materials from a data-driven perspective. Through systematic univariate and bivariate partial dependency analysis, a complete parameter optimization system is established, ranging from basic structural parameters (BET specific surface area 500-850 m² / g, pore volume 0.4-0.6 cm³ / g) to microstructural characteristics (micropore area >50 m² / g, micropore ratio 0.15-0.225, size matching window Dpratio 3-8.5), and then to interface properties (LogDDifference of -4 to -2). This realizes the leap from empirical screening to quantitative rational design of adsorption materials.
[0065] (3) This application achieves a breakthrough in efficiency in solving complex engineering problems. In the typical complex system of pharmaceutical wastewater treatment with multiple pollutants, based on the established design framework, the preferred resin was screened from 16,000 virtual materials in just a few hours. Experimental verification shows that the material achieves synergistic removal rates of 92%, 88%, 83%, and 85% for penicillin, cephalosporins, phenylacetic acid, and dimethyl sulfoxide, respectively, successfully solving the problem of synergistic removal in systems with wide polarity distribution, large size and shape differences, and multiple pollutant coexistence. The foregoing illustrative description of the present application and its embodiments is not restrictive and can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. The accompanying drawings are only one embodiment of the present application, and the actual structure is not limited thereto. Therefore, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the present application, such designs should fall within the scope of protection of this application. Furthermore, the word "comprising" does not exclude other elements or steps, and the word "a" preceding an element does not exclude the inclusion of "a plurality" of that element. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
Claims
1. A method for constructing a predictive model for resin adsorption performance, characterized in that, include: S1, to obtain pollutant molecular structure data, resin characterization data, and adsorption equilibrium experimental data of different resin-pollutant combinations; S2, feature extraction is performed on the pollutant molecular structure data and resin characterization data respectively to obtain pollutant feature sets and resin feature sets; S3, calculate the cross-domain adaptation parameters of the pollutant feature set and the resin feature set to obtain the adaptation parameter set; S4. The Langmuir adsorption isotherm model was used to perform nonlinear fitting on the adsorption equilibrium experimental data of different resin-pollutant combinations to obtain the maximum adsorption capacity of each resin-pollutant combination. ; Based on maximum adsorption capacity goodness of fit Perform screening to determine the goodness of fit. Maximum adsorption capacity greater than a preset threshold As the target variable; S5. Construct a training set based on the pollutant feature set, resin feature set, adaptation parameter set, and target variable; S6. Use the training set to train a machine learning model to obtain an adsorption performance prediction model.
2. The method for constructing a resin adsorption performance prediction model according to claim 1, characterized in that: Also includes: S7. The adsorption performance prediction model was analyzed mechanistically using interpretability analysis methods, yielding the relationship between each characteristic parameter in the pollutant characteristic set, resin characteristic set, and adaptation parameter set and the maximum adsorption capacity. The influence of the resin and the optimal range of each structural parameter in the resin feature set are used as the material design parameter range.
3. The method for constructing a resin adsorption performance prediction model according to claim 2, characterized in that: Feature extraction of pollutant molecular structure data includes: Molecular dynamics simulations were performed on the molecular structure data of pollutants to obtain the lowest energy conformation, which was then used as the stable conformation of the pollutant molecular structure data. The quantum chemical parameters of the stable conformation were calculated using density functional theory, and the electronic structure parameters of the pollutant molecular structure data were obtained. Geometric features of stable conformations are extracted using molecular descriptors to obtain spatial structure parameters of pollutant molecular structure data; The hydrophobicity of the stable conformation under a preset pH condition was calculated using a partition coefficient prediction algorithm, and the octanol-water partition coefficient at the target pH was obtained. ; Electronic structure parameters, spatial structure parameters and , as a pollutant characteristic set.
4. The method for constructing a resin adsorption performance prediction model according to claim 3, characterized in that: Electronic structure parameters include: polarizability and polar surface area; Spatial structure parameters include: molecular dynamics diameter Solvent-accessible surface area, number of rotatable bonds, number of hydrogen bond donors, and number of hydrogen bond acceptors.
5. The method for constructing a resin adsorption performance prediction model according to claim 3, characterized in that: Feature extraction was performed on the resin characterization data, including: The resin characterization data were tested for missing values using integrity verification rules. Resin data with numerical records for all characterization parameters were selected as a subset of complete characterization data. The experimental conditions of the complete characterization data subset were screened using the condition consistency rule. Resin data with the same test temperature, test pH and pretreatment method or within the allowable deviation range were selected as comparable characterization data subsets. Outlier detection was performed on each representation parameter in the comparable representation data subset using the box plot method or the 3σ criterion, and outliers were removed to obtain the representation data after outlier removal. The resin feature set is obtained by performing parameter analysis on the characterization data after removing outliers.
6. The method for constructing a resin adsorption performance prediction model according to claim 5, characterized in that: The resin characteristic set includes: ion exchange capacity, micropore ratio, micropore specific surface area, total specific surface area, pore volume, and average pore size. Resin surface polarity and functional group charge parameters.
7. The method for constructing a resin adsorption performance prediction model according to claim 4 or 6, characterized in that: S3, calculate the cross-domain adaptation parameters of the pollutant feature set and the resin feature set to obtain the adaptation parameter set, including: Extracting from pollutant characteristics Extract from resin feature set ,calculate and The difference is used to obtain the polarity matching parameter. ,in, ; Extracting molecular dynamics diameter from pollutant feature set Average pore size extracted from resin feature set ,calculate and The ratio of the two values yields the size matching parameters. ,in, ; Will and As an adaptation parameter set.
8. The method for constructing a resin adsorption performance prediction model according to claim 7, characterized in that: S4. The Langmuir adsorption isotherm model was used to perform nonlinear fitting on the adsorption equilibrium experimental data of different resin-pollutant combinations, using the following formula: ;in, To balance the adsorption capacity, For equilibrium concentration, KL is the Langmuir adsorption constant. This represents the maximum adsorption capacity.
9. The method for constructing a resin adsorption performance prediction model according to claim 8, characterized in that: Machine learning models include: gradient boosting decision tree or random forest algorithms.
10. A system for constructing a predictive model of resin adsorption performance, characterized in that, include: The data acquisition module acquires pollutant molecular structure data, resin characterization data, and adsorption equilibrium experimental data of different resin-pollutant combinations. The feature extraction module extracts features from pollutant molecular structure data and resin characterization data respectively, to obtain pollutant feature sets and resin feature sets; The adaptation parameter calculation module calculates the cross-domain adaptation parameters of the pollutant feature set and the resin feature set to obtain the adaptation parameter set; The target variable generation module uses the Langmuir adsorption isotherm model to perform nonlinear fitting on the adsorption equilibrium experimental data of different resin-pollutant combinations, thereby obtaining the maximum adsorption capacity of each resin-pollutant combination. Based on the maximum adsorption capacity goodness of fit Perform screening to determine the goodness of fit. Maximum adsorption capacity greater than a preset threshold As the target variable; The training set construction module constructs a training set based on the pollutant feature set, resin feature set, adaptation parameter set, and target variable. The model training module uses the training set to train a machine learning model and obtain an adsorption performance prediction model.
Citation Information
Patent Citations
A resin synthesis method, system, medium and product based on machine learning
CN118430675B