Adsorption material screening method and system for transfer learning auxiliary material genome design
By constructing a feature space mapping relationship between the source domain and the target domain, and employing transfer learning methods and materials genome design, the problems of insufficient data and cross-system transfer in the screening of adsorption materials were solved, achieving efficient material screening and synthesis optimization.
Patent Information
- Application Number
- CN202511432108.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-03-03
AI Technical Summary
Existing methods for screening adsorbent materials have low prediction accuracy when there is insufficient data in the target domain and are difficult to transfer knowledge across systems. This results in the need to remodel each time a new adsorbent system is encountered, leading to a waste of computational resources and low efficiency.
By constructing a feature space mapping relationship between the source domain and the target domain, a model is built using transfer learning methods. Combined with materials genome design, the functional group combination is optimized to achieve efficient knowledge transfer and full-process optimization.
It improves prediction accuracy under small sample conditions, shortens material design cycle, reduces computing resource requirements, and realizes cross-system knowledge transfer and full-process automation.
Smart Images

Figure CN121601101A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of machine learning and materials genome engineering, and in particular to a method and system for screening adsorbent materials for materials genome design assisted by transfer learning. Background Technology
[0002] Pollutant adsorption and separation technology plays a vital role in environmental protection and resource recovery, but traditional methods face numerous challenges. While existing technologies such as solvent extraction and chemical precipitation each have their advantages, they still exhibit significant shortcomings in selectivity, efficiency, and environmental friendliness. Adsorption, due to its high efficiency, ease of operation, and low cost, is considered an excellent method for treating polluted wastewater.
[0003] In practice, adsorbent materials can be synthesized using various methods. Adsorbent materials synthesized using different methods exhibit significant differences in adsorption performance. Furthermore, preparing different types of adsorbents and testing their adsorption performance for different types of pollutants is particularly time-consuming and labor-intensive. Utilizing machine learning (ML) to build targeted or generalized models to predict the adsorption performance of adsorbent materials is one of the future development trends in the field of adsorbent material design. In recent years, machine learning technology has shown great potential in the field of material screening. For example, the machine learning-based adsorbent material screening method (CN113963754A) developed by Nanjing University achieves efficient screening of adsorbent materials for specific pollutants by establishing datasets of material descriptors and pollutant descriptors. However, this type of method has obvious limitations: each new system requires a complete training dataset, consuming large computational resources, and it cannot achieve cross-system knowledge transfer. While the semi-coke-based porous carbon heavy metal adsorption efficiency prediction method (CN114530217A) proposed by Xi'an University of Architecture and Technology solves the data missing problem, its model generalization ability is still limited, making it difficult to adapt to the screening needs of different metal ion systems.
[0004] In specific application scenarios, such as the hydrophobic MOF screening method for DMMP adsorption (CN115458073A) and the MOF membrane screening method for H2 separation (CN115620835A) developed by Guangzhou University, both employ single-domain machine learning modeling. While these methods perform well within their respective fields, they lack knowledge transfer mechanisms, requiring remodeling for each new adsorption system and resulting in a significant waste of computational resources. Furthermore, although existing technologies such as CN115544862A and CN116130011A introduce feature importance analysis, they still fail to fundamentally solve the challenge of cross-system prediction.
[0005] To address these technical bottlenecks, this invention proposes an innovative method for screening adsorbent materials in materials genome design assisted by transfer learning. This method overcomes the limitations of traditional machine learning methods by establishing a feature space mapping relationship between the source domain (e.g., rhenium acid data) and the target domain (e.g., chromic acid data), achieving efficient knowledge transfer. In terms of technical implementation, this invention employs an evaluation method that compares the effects of direct modeling (DL) and transfer learning (TL), and introduces group sampling to verify transfer stability, ensuring the reliability of the model. More importantly, this invention organically combines material screening with materials genome design, achieving end-to-end optimization from material screening to synthesis by analyzing common framework structures and designing optimized functional groups. Summary of the Invention
[0006] This invention addresses the problems of low prediction accuracy and difficulties in cross-system knowledge transfer when existing adsorbent screening methods suffer from insufficient target domain data. It proposes a transfer learning-assisted material genome design method and system for adsorbent screening. By constructing a feature space mapping relationship between the source domain and the target domain, knowledge transfer is achieved between systems such as anion-anion, cation-cation, and organic-organic molecules, thus solving the problem of low material development efficiency under small sample conditions.
[0007] To achieve the aforementioned objectives, the technical solution provided by this invention is as follows:
[0008] A method for screening adsorbent materials for genome design of transfer learning aids, comprising the following steps:
[0009] S1. Obtain adsorption material descriptors and performance data of source and target domains to construct adsorption material dataset. For the initial material dataset, calculate the distance of the dual-domain feature distribution through the maximum mean difference and KL divergence, and establish a feature mapping relationship matrix, including constructing a transferable feature subspace based on the feature space similarity of source and target domains.
[0010] S2. Construct a transfer learning model based on the obtained transferable feature subspace, and directly establish a machine learning model based on the adsorption material dataset of the target domain. Compare the prediction effects of transfer learning and direct modeling to evaluate the model prediction performance, and optimize and adjust the transfer learning model to improve the prediction performance of the transfer learning model under small sample conditions.
[0011] S3. Based on the prediction results of the transfer learning model, screen common material framework structure features in the source and target domains. Screening methods include calculating the structural matching degree through cosine similarity algorithm and topological analysis to identify common framework structures of adsorbed materials.
[0012] The common framework structure includes the pore size distribution and metal coordination environment of the target adsorbent material, as well as the surface charge characteristics;
[0013] S4. Based on materials genomics, design and optimize the combination of functional groups, including calculating and selecting active groups according to electronic structure, generating a target adsorption material design scheme, and synthesizing the material according to the design scheme and testing its adsorption performance.
[0014] Further, step S1 includes preprocessing the adsorption material datasets of the source and target domains, wherein the preprocessing includes dimensionality reduction of the adsorption material datasets using PCA and / or t-SNE methods.
[0015] Furthermore, the adsorption material dataset includes structural descriptors such as material specific surface area, pore size distribution, functional group type, and metal coordination environment, as well as performance parameters such as adsorption capacity and adsorption selectivity.
[0016] Further, step S2 involves using random sampling and grouping strategies to evaluate the stability of the transfer learning model, serving as a reference for the optimal adjustment of the transfer learning model;
[0017] The random sampling employs the Bootstrap method for multiple random samplings, and the grouping strategy includes grouping migrations by material type or adsorption conditions.
[0018] Furthermore, step S2's evaluation of the model's predictive performance includes the model's fit coefficient R. 2 Root mean square error (RMSE) and mean absolute error (MAE).
[0019] Furthermore, the topology analysis described in step S3 includes constructing a structure-performance correlation map, and the screening method also includes scoring the structural similarity and performance importance of the target adsorbent material.
[0020] Step S4 includes the optimization design of functional group combinations through molecular docking simulation, density functional theory calculation, and molecular dynamics simulation, and also includes the construction of a functional group combination rule library.
[0021] Furthermore, the grouping strategy refers to grouping data according to the migration of anion data, cation data, and / or organic matter analysis data from the materials genome.
[0022] The present invention also provides a screening system for adsorbent materials for genome design of transfer learning aids, the system being configured to perform the above-described method;
[0023] The system includes:
[0024] The data acquisition module is used to collect adsorption material datasets from the source and target domains.
[0025] The transfer learning module performs maximum mean difference and KL divergence analysis on the adsorption material datasets of the source and target domains to construct a transferable feature subspace and is configured to perform computational processing of the transfer learning model.
[0026] The machine learning module is configured to execute the machine learning model;
[0027] The model validation and evaluation module acquires and compares the prediction results of the models from the transfer learning module and the machine learning module, including calculating the fitting coefficient R of the output transfer learning model. 2 Root mean square error (RMSE) and mean absolute error (MAE);
[0028] The materials genome design module combines functional groups based on a common framework structure, including a function group combination rule verification method.
[0029] The materials synthesis and testing module is used to design test experiments and record test data, including simulation tests or the development of physical test experiment plans.
[0030] Compared to existing technologies, this invention innovatively proposes a transfer learning-assisted solution to address the "data silo" problem in adsorption material development. This overcomes the limitations of traditional methods that require building datasets from scratch when developing new material systems, achieving efficient knowledge transfer. Through the transfer learning mechanism, it significantly improves prediction accuracy under small sample conditions, solving the problem of insufficient target domain data. This invention organically combines material screening with genome design, achieving full-process optimization from theoretical prediction to experimental synthesis. The methods described in this invention have broad applicability and can be applied to transfer learning in various systems such as cation-cation, anion-anion, and organic-organic molecule systems. Attached Figure Description
[0031] Figure 1 This is a schematic diagram illustrating the implementation process of the method described in this invention. Detailed Implementation
[0032] The present invention will be further described below with reference to specific embodiments.
[0033] like Figure 1As shown, this invention provides a transfer learning-assisted method and related apparatus for screening adsorbent materials. The implementation of this invention first requires the collection of source and target domain data. For example, the source domain corresponds to rhenium acid adsorbent materials, and the target domain is chromic acid adsorbent materials. The datasets for both should include the material's structural descriptors (specific surface area, pore size distribution, functional group type, etc.) and performance parameters (adsorption capacity, selectivity, etc.). Given the limited data for the target domain and the small sample size issue, we need to consider performing feature space similarity analysis to establish a feature mapping relationship matrix. Specifically, this includes calculating the feature distribution distance between the source and target domains using the maximum mean difference (MMD) and KL divergence, and selecting low-divergence features with Δ < 0.1 to construct a transfer subspace. A sliding window method is used to dynamically adjust the feature weights. A kernel function method is used to calculate the distance between the two domain feature distributions. The features are mapped to the regenerating kernel Hilbert space (RKHS) using a Gaussian kernel function, and their mean embedding difference is calculated. When the MMD value < 0.3, it is determined to be a transferable feature subspace. This threshold is determined through cross-validation. The formula for calculating the MMD value is:
[0034] MMD 2 (X,Y)=E x,x’~X [k(x,x')]+E y,y’~Y [k(y,y')]-2E x~X,y~Y [k(x,y)]
[0035] In the formula, E x,x’~X [k(x,x')]: This is the expected value of the kernel function for sample pairs (x,x') sampled from distribution X, measuring the similarity between samples within distribution X. E y,y’~Y [k(y,y')] is the expected value of the kernel function of a sample pair (y,y') sampled from distribution Y, which measures the similarity between samples within distribution Y. x~X,y~Y [k(x,y)] is the expected value of the kernel function applied to sample pairs (x,y) sampled from distributions X and Y, respectively, measuring the similarity between distributions X and Y. The kernel function k(x,y) is typically a Gaussian kernel, in the form: Here, σ is the width parameter of the Gaussian kernel, which controls the smoothness of the kernel function.
[0036] Next, we construct prediction models using direct modeling (DL) and transfer learning (TL) methods, respectively. The direct modeling method employs random forest or XGBoost algorithms, while the transfer learning method uses TrAdaBoost or deep transfer learning networks. The structure and expression of the deep transfer learning network are as follows:
[0037] L total =L task +λL MK-MMD
[0038] The above formula indicates that the transfer learning model consists of 3 residual blocks and 1 domain adaptation layer. The model uses the MK-MMD loss function to achieve feature alignment.
[0039] For screening common frameworks, the structural matching degree is calculated using an improved cosine similarity algorithm, and the calculation formula is as follows:
[0040]
[0041] In the formula, α is the weight of the specific surface area ratio, β is the weight of the pore size difference term, γ is the weight of the metal coordination matching degree, and V SA and V SB : Represents the specific surface area of structure A and structure B, respectively. DA and DB represent the pore sizes of structure A and structure B, respectively. σ represents a parameter controlling the smoothness of the pore size difference term. FM represents the metal coordination matching degree, indicating the metal coordination matching status of structure A and structure B. A structure is considered matched when Sim ≥ 0.8. Regarding feature importance evaluation, a comprehensive feature importance evaluation function, Importance = Σ(w_iF_i), was constructed, where w_i is the feature weight obtained through SHAP analysis, F_i is the normalized feature value, and Importance ≥ 0.7 is set as the screening threshold.
[0042] Based on a common framework structure, adsorption functional groups are designed and optimized using materials genomics methods (including molecular docking simulations, density functional theory calculations, and molecular dynamics simulations), leading to the design of adsorbents. Based on the screened common framework structures (ΔG ≤ -20 kJ / mol), combinations of functional groups with strong interactions are selected. Subsequently, electronic structure characteristics are analyzed using density functional theory (DFT) calculations, focusing on the energy level matching degree (≥0.75) between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO) to ensure electron transfer efficiency. Finally, molecular dynamics simulations are used to verify the structural stability of the functional groups in solution (RMSD < 0.2 nm). This method innovatively constructs a combinatorial rule library containing active groups and uses the aforementioned transfer learning model to predict the adsorption performance (including adsorption capacity, adsorption selectivity, and cycling stability) of molecules formed by random combinations on the common framework structure. Based on this, the optimal adsorbent based on gene combinations is selected.
[0043] In addition to the above description, further optimizations are also included. This invention targets a limited dataset of the target domain system (such as the chromic acid adsorption system), and data quality and consistency must be ensured during data acquisition. In the feature space similarity analysis stage, principal component analysis (PCA) and t-SNE dimensionality reduction algorithms are used to map and visualize the feature spaces of the source and target domains. The key to this step is determining the transferability threshold between the two domains; when the feature similarity reaches 0.8 or higher, it indicates that effective knowledge transfer can be performed. Special attention needs to be paid to the uniformity of feature scales and the handling of outliers during the analysis process.
[0044] During the model building phase, a comparative experimental design was employed, simultaneously developing two prediction models: direct modeling (DL) and transfer learning (TL). Direct modeling utilized conventional machine learning algorithms (such as XGBoost and Random Forest), while the transfer learning model was based on a deep transfer learning network architecture. Both models used the same training-test set split ratio (typically 7:3 or 8:2) and employed k-fold cross-validation (k=5) to ensure the reliability of the results.
[0045] In the model validation phase, three key metrics are evaluated: the coefficient of determination (R²). 2 The parameters to be evaluated are root mean square error (RMSE) and mean absolute error (MAE). The validation process requires multiple repeated experiments to confirm the stability of the results, and sensitivity analysis to determine the degree of influence of each input feature on the model output. In particular, it is necessary to compare and analyze the performance improvement effect of the transfer learning model under small sample conditions.
[0046]
[0047] Based on the output of the transfer learning model, common structural features of the source and target domain systems are screened out. This process requires combining materials science principles to identify structural parameters (such as specific pore size ranges and surface functional group combinations) that play a decisive role in adsorption performance. After the common framework is determined, molecular simulation techniques (molecular docking, DFT calculations, etc.) are used to optimize the design of functional groups.
[0048] The following factors need to be considered in material optimization design:
[0049] (1) Molecular characteristics of the target pollutant;
[0050] (2) Adsorption environment conditions (pH, temperature, etc.);
[0051] (3) Synthesizability of the material. Through a multi-objective optimization algorithm, the performance indicators such as adsorption capacity, selectivity and stability are balanced, and the optimal combination of functional groups and material structure parameters are finally determined.
[0052] It should be noted that in the above implementation, the transfer learning model can effectively utilize knowledge from source domain data (such as rhenium acid adsorption data) to improve the predictive performance of the target domain (such as the chromic acid adsorption system), especially when the target domain data is limited (e.g., only 30-50 sets of data). Through feature space similarity analysis, the transferability between different systems can be determined; the transfer learning effect is optimal when the structural similarity is ≥0.8. The model evaluation employs rigorous random sampling and cross-validation methods to ensure transfer stability and reliability. The final material design is based on a common framework structure and optimized functional groups, achieving full-process optimization from theoretical prediction to experimental verification.
[0053] Example 1
[0054] This embodiment of a method for migrating rhenium acid data to a chromic acid data system includes the following steps:
[0055] S10. Source and Target Domain Data Collection: 100 sets of structural descriptors and performance data for rhenium acid adsorbents were collected as the source domain dataset, and 30 sets of structural descriptors and performance data for chromic acid adsorbents were collected as the target domain dataset. The data includes structural descriptors such as specific surface area, pore size distribution, and functional group type, as well as performance parameters such as adsorption capacity and selectivity. Preprocessing includes dimensionality reduction of the adsorbent material datasets using PCA and / or t-SNE methods. 34 features remain.
[0056] S20. Similarity Analysis of Source and Target Domain Feature Spaces: KL divergence was calculated for the feature distributions of the source and target domains, and low-divergence features with Δ < 0.1 were selected to construct a transfer subspace. Features were mapped to the Reproducing Kernel Hilbert Space (RKHS) using a Gaussian kernel function, and their mean embedding differences were calculated. The calculated MMD value was 0.12 < 0.3, indicating a transferable feature subspace. The results show that the two systems have a structural similarity of 0.71 across the main feature dimensions.
[0057] S30. Comparison of DL and TL performance: Predictive models were constructed using direct modeling (DL) and transfer learning (TL) methods, respectively. The DL model employed the XGBoost algorithm, while the TL model used the TrAdaBoost algorithm. The comparison results show that the TL model achieved a higher prediction accuracy (R²). 2 =0.92) is significantly higher than that of the DL model (R = 0.92) 2 =0.78).
[0058] S40. Random Sampling and Grouping to Demonstrate Transfer Stability: A 5-fold cross-validation was performed using a random sampling ratio of 70% training set and 30% test set. Results showed that the model exhibited good transfer stability under different data distributions, with prediction error fluctuations of less than 5%.
[0059] S50. Screening common framework structures: Based on the transfer learning results, common framework structure features of rhenium acid and chromic acid adsorbent materials are screened, including specific pore size distribution patterns and metal coordination environments.
[0060] S60. Functional group design optimization: Based on the common framework structure, the adsorption functional groups were designed and optimized using molecular docking simulation, and the combination of -NH2 and -COOH bifunctional groups was finally determined as the optimal choice.
[0061] S70. Material Synthesis and Performance Testing: Based on the optimization results, the target adsorbent material was synthesized, and the test results showed that the adsorption capacity for chromic acid reached 85 mg / g.
[0062] S80. Device Design and Implementation: Develop a dedicated screening device to achieve fully automated control of the entire process from data acquisition to material design.
[0063] Example 2
[0064] The basic content of this embodiment is the same as that of embodiment 1, except that the method for migrating copper data to the lead data system in this embodiment is as follows.
[0065] S10. Source and Target Domain Data Collection: Collect 150 sets of structural descriptors and performance data of copper adsorbent materials as the source domain dataset, and collect 40 sets of structural descriptors and performance data of lead adsorbent materials as the target domain dataset.
[0066] S20. Similarity Analysis of Source and Target Domain Feature Spaces: KL divergence was calculated for the feature distributions of the source and target domains, and low-divergence features with Δ < 0.1 were selected to construct a transfer subspace. Features were mapped to the Reproducing Kernel Hilbert Space (RKHS) using a Gaussian kernel function, and their mean embedding differences were calculated. The calculated MMD value was 0.12 < 0.3, indicating a transferable feature subspace. The results show that the two systems have a structural similarity of 0.82 across the main feature dimensions.
[0067] S30. Comparison of DL and TL performance: Prediction models were constructed using direct modeling (DL) and transfer learning (TL) methods, respectively. The DL model employed the random forest algorithm, while the TL model used a deep transfer learning network. The comparison results show that the TL model's prediction accuracy (R²) is significantly higher. 2 =0.89) is significantly higher than that of the DL model (R 2 =0.71).
[0068] S40. Random Sampling and Grouping to Demonstrate Transfer Stability: A 5-fold cross-validation was performed using a random sampling ratio of 75% training set and 25% test set. Results showed that the model exhibited good transfer stability under different data distributions, with prediction error fluctuations of less than 6%.
[0069] S50. Screening common framework structures: Based on the results of transfer learning, common framework structure features of copper and lead adsorbent materials are screened out, including specific pore size distribution and surface charge properties.
[0070] S60. Functional group design optimization: Based on the common framework structure, density functional theory was used to calculate and optimize the adsorption functional groups, and the combination of -SH and -COOH bifunctional groups was finally determined as the optimal choice.
[0071] S70. Material Synthesis and Performance Testing: Based on the optimization results, the target adsorbent material was synthesized, and the test results showed that the adsorption capacity for lead reached 120 mg / g.
[0072] Example 3
[0073] The basic content of this embodiment is the same as that of embodiment 1, except that the method for migrating alicyclic data to aromatic data system in this embodiment is as follows.
[0074] S10. Source and Target Domain Data Collection: Collect 200 sets of structural descriptors and performance data of alicyclic compound adsorbents as source domain datasets, and collect 50 sets of structural descriptors and performance data of aromatic compound adsorbents as target domain datasets.
[0075] S20. Similarity analysis of the feature spaces of source and target domains: Principal component analysis (PCA) and t-SNE dimensionality reduction visualization methods were used to analyze the similarity of the feature spaces of alicyclic and aromatic adsorbent materials. The results showed that the two systems had a structural similarity of 0.88 in the main feature dimensions.
[0076] S30. Comparison of DL and TL performance: Prediction models were constructed using direct modeling (DL) and transfer learning (TL) methods, respectively. The DL model employed the XGBoost algorithm, while the TL model used a deep transfer learning network. The comparison results show that the TL model achieved a higher prediction accuracy (R²). 2 =0.94) is significantly higher than that of the DL model (R = 0.94) 2 =0.80).
[0077] S40. Random Sampling and Grouping to Demonstrate Transfer Stability: A 5-fold cross-validation was performed using a random sampling ratio of 80% training set and 20% test set. Results showed that the model exhibited good transfer stability under different data distributions, with prediction error fluctuations of less than 4%.
[0078] S50. Screening common framework structures: Based on the results of transfer learning, common framework structure features of alicyclic and aromatic adsorbent materials are screened, including specific π-π stacking interactions and hydrophobic properties.
[0079] S60. Functional group design optimization: Based on the common framework structure, the adsorption functional groups were designed and optimized using molecular dynamics simulations, and the combination of -SH and -CH3 bifunctional groups was finally determined as the optimal choice.
[0080] S70. Material Synthesis and Performance Testing: Based on the optimization results, the target adsorbent material was synthesized, and the test results showed that the adsorption capacity for aromatic compounds reached 95 mg / g.
[0081] Comparative Example 1: This comparative example uses a direct modeling method for the chromic acid data system. The basic content of this example is the same as Example 1, except that this example uses a direct training method to predict the chromic acid data system. The same 30 sets of chromic acid adsorbent material datasets as in Example 1 were collected, including specific surface areas (500-1200 m²). 2 / g), pore size distribution Surface functional groups (-NH2, -COOH, etc.) and adsorption capacity (20-85 mg / g) were used for structural features and adsorption capacity data. The original 34 features were used directly for modeling without cross-system feature transfer processing. The XGBoost algorithm was used to directly build the prediction model, with key parameters set as follows: learning rate: 0.01, maximum tree depth: 6, subsample ratio: 0.8, feature sampling ratio: 0.7. Model training employed 5-fold cross-validation, with 70% of the data used for training and 30% for testing. Performance test results showed that the prediction accuracy R... 2 =0.78±0.05 Root Mean Square Error (RMSE) = 8.7 mg / g, Mean Absolute Error (MAE) = 6.2 mg / g. Compared with the transfer learning method (R) in Example 1... 2 =0.92), the prediction accuracy decreased by 18%. The prediction bias was particularly significant in the high adsorption capacity region (>70 mg / g), with a maximum bias of 15 mg / g.
[0082] Comparative Example 2: This comparative example uses a direct modeling method for the chromic acid data system. The basic content of this example is the same as that of Example 1, except that: this comparative example does not screen for common framework structures, but directly performs random combination enumeration of 12 common functional groups, generating approximately 1.3 × 10 5 The number of potential material structures leads to an exponential increase in computational complexity. In practice, molecular docking simulations were used to calculate the binding energy (ΔG) of each structure with chromic acid, DFT calculations were used to evaluate the electronic structure matching degree, and the molecular dynamics stability of the top 1% of candidate structures was verified. The entire screening process consumed 2.4 × 10⁻⁶ hours of computation. 6 CPU cores are 8.6 x 10 4GPU computing hours were required, storage demand reached 28TB, and the screening cycle was extended to 98 days, two orders of magnitude longer than the 1-day cycle in Example 1. Ultimately, only 37 effective structures were obtained, with a success rate as low as 0.028%, and the adsorption capacity of the optimal material was 79 mg / g, a 7% decrease compared to 85 mg / g in Example 1.
[0083] The technical advantages of this invention are reflected in several aspects: in transfer learning applications between rhenium acid-chromic acid, copper-lead, and alicyclic-aromatic systems, the screening efficiency is 3-5 times higher than traditional methods, and the material design cycle is shortened by more than 60%. Compared with existing technologies, the transfer learning framework of this invention improves the prediction accuracy of the model in new systems by an average of 25%, while significantly reducing the computational resource requirements. Furthermore, the developed dedicated device automates the entire process from screening to synthesis, filling a gap in existing technologies in this field. These innovations enable this invention to demonstrate significant technical advantages and broad application prospects in the separation of complex metal ion systems.
[0084] The present invention and its embodiments have been described above illustratively. This description is not restrictive, and the data used is only one embodiment of the present invention. The actual combination of data is not limited to this. Therefore, if those skilled in the art are inspired by this description and, without departing from the spirit of the present invention, devise similar embodiments and examples of the technical solution without creative design, all such embodiments and examples should fall within the protection scope of the present invention.
Claims
1. A method for screening adsorbent materials for genome design of transfer learning aids, characterized in that, The implementation steps include: S1. Obtain adsorption material descriptors and performance data of source and target domains to construct adsorption material dataset. For the initial material dataset, calculate the distance of the dual-domain feature distribution through the maximum mean difference and KL divergence, and establish a feature mapping relationship matrix, including constructing a transferable feature subspace based on the feature space similarity of source and target domains. S2. Construct a transfer learning model based on the obtained transferable feature subspace, and directly establish a machine learning model based on the adsorption material dataset of the target domain. Compare the prediction effects of transfer learning and direct modeling to evaluate the model prediction performance, and optimize and adjust the transfer learning model to improve the prediction performance of the transfer learning model under small sample conditions. S3. Based on the prediction results of the transfer learning model, screen common material framework structure features in the source and target domains. Screening methods include calculating the structural matching degree through cosine similarity algorithm and topological analysis to identify common framework structures of adsorbed materials. The common framework structure includes the pore size distribution and metal coordination environment of the target adsorbent material, as well as the surface charge characteristics; S4. Based on materials genomics, design and optimize the combination of functional groups, including calculating and selecting active groups according to electronic structure, generating a target adsorption material design scheme, and synthesizing the material according to the design scheme and testing its adsorption performance.
2. The method for screening adsorbent materials for genome design of transfer learning aids according to claim 1, characterized in that, Step S1 includes preprocessing the adsorption material datasets of the source and target domains. The preprocessing includes dimensionality reduction of the adsorption material datasets using PCA and / or t-SNE methods.
3. The method for screening adsorbent materials for genome design of transfer learning aids according to claim 1, characterized in that, The aforementioned adsorption material dataset includes structural descriptors such as material specific surface area, pore size distribution, functional group type, and metal coordination environment, as well as performance parameters such as adsorption capacity and adsorption selectivity.
4. The method for screening adsorbent materials for genome design of transfer learning aids according to claim 1, characterized in that, Step S2 involves using random sampling and grouping strategies to evaluate the stability of the transfer learning model, serving as a reference for optimal adjustment of the transfer learning model. The random sampling employs the Bootstrap method for multiple random samplings, and the grouping strategy includes grouping migrations by material type or adsorption conditions.
5. The method for screening adsorbent materials for genome design of transfer learning aids according to claim 1 or 4, characterized in that, Step S2, the evaluation of the model's predictive performance, includes the model's fit coefficient R. 2 Root mean square error (RMSE) and mean absolute error (MAE).
6. The method for screening adsorbent materials for genome design of transfer learning aids according to claim 1, characterized in that, The topology analysis described in step S3 includes constructing a structure-performance correlation map, and the screening method also includes scoring the structural similarity and performance importance of the target adsorbent material.
7. The method for screening adsorbent materials for genome design of transfer learning aids according to claim 1, characterized in that, Step S4 includes the optimization design of functional group combinations through molecular docking simulation, density functional theory calculation, and molecular dynamics simulation, and also includes the construction of a functional group combination rule library.
8. The method for screening adsorbent materials for genome design of transfer learning aids according to claim 4, characterized in that, The grouping strategy refers to grouping data according to the migration of anion data, cation data, and / or organic matter analysis data from the materials genome.
9. A screening system for adsorbent materials for genome design of transfer learning aids, characterized in that, The system is used to perform the method as described in any one of claims 1-8; The system includes: The data acquisition module is used to collect adsorption material datasets from the source and target domains. The transfer learning module performs maximum mean difference and KL divergence analysis on the adsorption material datasets of the source and target domains to construct a transferable feature subspace and is configured to perform computational processing of the transfer learning model. The machine learning module is configured to execute the machine learning model; The model validation and evaluation module acquires and compares the prediction results of the models from the transfer learning module and the machine learning module, including calculating the fitting coefficient R of the output transfer learning model. 2 Root mean square error (RMSE) and mean absolute error (MAE); The materials genome design module combines functional groups based on a common framework structure, including a function group combination rule verification method. The materials synthesis and testing module is used to design test experiments and record test data, including simulation tests or the development of physical test experiment plans.
Citation Information
Patent Citations
Adsorption material screening method and device based on machine learning
CN113963754A
Method for predicting heavy metal adsorption efficiency of semi-coke-based porous carbon and related device
CN114530217A
Machine learning method for hydrophobic MOF adsorbing DMMP
CN115458073A
Method and device for screening adsorption performance of metal organic framework and storage medium
CN115544862A
Method for high-throughput screening of metal organic framework membrane based on machine learning assistance
CN115620835A