Screening method for ionic liquids
A neural network and Monte Carlo method, combined with machine learning, efficiently generates and validates ionic liquid candidates for cellulose solubility, addressing the inefficiencies of traditional trial-and-error methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-04-03
AI Technical Summary
Developing optimal ionic liquids (ILs) for cellulose solubility relies on a time-consuming and costly trial-and-error approach, making it impractical to experimentally measure all possible types, which hinders the efficient exploration of IL chemical spaces.
A method combining a neural network with a Monte Carlo method, including a backpropagation step, is used to generate and verify chemically valid ionic liquid candidates, followed by machine learning models to predict cellulose solubility and melting points, thereby screening ILs efficiently.
Enables rapid and efficient screening of ionic liquids soluble in cellulose, exploring undeveloped IL chemical spaces, reducing the time and cost associated with traditional methods.
Smart Images

Figure 2026058214000020 
Figure 2026058214000021 
Figure 2026058214000022
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for screening ionic liquids. [Background technology]
[0002] Cellulose and its derivatives are used in a wide variety of everyday applications, including packaging coatings and thin films, thin-film filtration membranes, cellulose / biopolymer green bio-composite materials, and cellulose-based or composite gels and fibers. However, to maximize the potential of cellulose, it is sometimes necessary to dissolve it as a pretreatment, but cellulose is poorly soluble in conventional organic solvents and water (Non-Patent Documents 1 and 2).
[0003] Ionic liquids (ILs), composed of organic cations and inorganic or organic anions, are considered alternative solvents for cellulose processing. However, developing optimal ILs with high cellulose solubility often relies on a time-consuming and costly trial-and-error approach. There are 10 possible types of ILs. 18 It is impractical to experimentally measure all of the target ILs across different types (Non-Patent Document 3). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Cellulose 26, 35-79 (2019) [Non-Patent Document 2] Chem. Rev. 109, 6712-6728 (2009) [Non-Patent Document 3] Clean Prod. Process. 1, 223-236 (1999) [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] Therefore, in order to facilitate the experimental development process, it is extremely desirable to develop a rapid and efficient computational method for exploring undeveloped IL chemical spaces. That is, the object of the present invention is to provide a rapid and efficient method for screening ionic liquids soluble in cellulose for exploring undeveloped IL chemical spaces. [Means for solving the problem]
[0006] The inventors have found that by combining a neural network with a Monte Carlo method that includes a backpropagation step to verify the chemical validity of the generated structure dataset, a vast number of novel IL candidates can be generated while eliminating chemically invalid structures and neutral molecules. Furthermore, the inventors have found that by screening these novel IL candidates using a relatively rapid machine learning (ML) model to narrow down the candidates to those with the desired cellulose solubility, ionic liquids soluble in cellulose can be screened rapidly and efficiently.
[0007] In other words, the gist of this disclosure concerns the following: (1) A generation step of training a neural network means with existing organic cation and anion structure datasets, and using the neural network means to generate an organic cation structure dataset A and an organic anion structure dataset B that are different from the existing structure datasets, A prediction step is performed to predict the solubility of cellulose for an ionic liquid structure dataset C, which consists of organic cations in structure dataset A and organic anions in structure dataset B, generated in the production step, using a first machine learning model trained with known structural information of an ionic liquid and data on the solubility of cellulose in the ionic liquid. Includes, In the generation process, the generation is guided using a Monte Carlo method that includes a backpropagation step to verify that the generated structure datasets A and B are chemically effective. A method for screening ionic liquids that are soluble in cellulose. (2) The method according to (1), wherein the neural network means is a recurrent neural network (RNN). (3) The method according to (1) above, wherein the Monte Carlo method is Monte Carlo tree search (MCTS). (4) The method according to (1), wherein the prediction step includes predicting the melting point of an ionic liquid structure dataset C consisting of organic cations in structure dataset A and organic anions in structure dataset B generated in the generation step, using a second machine learning model trained with known structural information of the ionic liquid and data on the melting point of the ionic liquid. (5) The method according to (1), further comprising a verification step of verifying the solubility of cellulose in the ionic liquid used for prediction in the prediction step using a Conductor-like Model for Real Solvents (COSMO-RS). [Effects of the Invention]
[0008] The present invention provides a rapid and efficient method for screening ionic liquids that are soluble in cellulose to explore undeveloped IL chemical spaces. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 shows a schematic diagram of a novel cation generation algorithm. (a) Dataset for training the RNN. (b) Function of the trained RNN. (c) Typical loop of the MCTS process. [Figure 2] Figure 2 shows a schematic diagram of an automated workflow for calculating lnγ of cellulose in IL using COSMO-RS. [Figure 3]Figure 3 shows the following: (a) UMAP plot of the newly generated cations (shown as gray dots). Randomly selected cations are highlighted with a "·" along with the symbols listed in (c) for structural representation. (b) UMAP plot of the newly generated anions (shown as gray dots). Randomly selected anions are highlighted with a "·" along with the symbols listed in (c) for structural representation. (c) Structures of randomly selected generated cations and anions. [Figure 4] Figure 4 shows a comparison of the performance of various ML models in predicting the cellulose solubility of IL based on a dataset of 674 data points: (a) RF regression; (b) XGBoost regression; (c) Ridge regression; (d) ANN. [Figure 5] Figure 5 shows a comparison of the performance of various ML models in predicting the cellulose solubility of IL based on a dataset of 379 data points: (a) RF regression; (b) XGBoost regression; (c) Ridge regression; (d) ANN. [Figure 6] Figure 6 shows the following: (a) Accuracy of the ANN model trained to predict cellulose solubility in IL. (b) Distribution of SHAP values for the 10 features that contributed most to the ANN model for predicting cellulose solubility in IL. (c) Accuracy of the RF regression model trained to predict the melting point (Tm) of IL. (d) Distribution of SHAP values for the 10 features that contributed most to the RF regression model for predicting the melting point of IL. [Figure 7] Figure 7 shows a comparison of the performance of various ML models in predicting the melting point of IL. (a) RF regression. (b) XGBoost regression. (c) Ridge regression. (d) ANN. [Figure 8]Figure 8 shows a schematic diagram of the high-throughput virtual screening process and the results represented in a 2D space using UMAP. (a) UMAP plot of the generated novel cations and existing cations tested in the cellulose dissolution experiment. (b) UMAP plot of the generated novel anions and existing anions tested in the cellulose dissolution experiment. (c) UMAP plot of the cations located at AD of the predictive ML model and the cations within the training set. (d) UMAP plot of the anions located at AD of the predictive ML model and the anions within the training set. (e) UMAP plot of the promising cations. (f) UMAP plot of the promising anions. [Figure 9] Figure 9 shows the linear regression between the experimental cellulose solubility and the COSMO-RS calculation lnγ at 80 °C.
Mode for Carrying Out the Invention
[0010] The present invention includes a generation step of training neural network means with a structural dataset of existing organic cations and anions, and generating, by the neural network means, a structural dataset A of organic cations and a structural dataset B of organic anions different from the existing structural dataset; a prediction step of predicting the solubility of cellulose for a structural dataset C of ionic liquids composed of the organic cations in the structural dataset A and the organic anions in the structural dataset B generated in the generation step, using a first machine learning model trained using the structural information of known ionic liquids and data on the solubility of cellulose in the ionic liquids; and guides the generation using a Monte Carlo method including a backpropagation step of verifying that the generated structural datasets A and B are chemically valid in the generation step. The present invention relates to a method for screening ionic liquids having solubility in cellulose.
[0011] Existing organic cation and anion structure datasets can be obtained, for example, from chemical structure databases that record and retrieve structural information of various compounds, as structural information of some or all of the compounds recorded therein. Examples of such chemical structure databases include PubChem, ILThermo, ChemSpider, ChEMBL, MACCS, MA, Available Chemicals Exchange (ACX), SciFinder, ISIS, and ChEB. The structural information of a compound is not particularly limited as long as it includes, for example, the types of atoms contained in the compound, the types of bonds, the presence or absence of ring structures, and other structural characteristics.
[0012] Neural network methods are a type of machine learning that uses neural networks to generate models of complex patterns in data. Examples include recurrent neural networks (RNNs), variational autoencoders (VAEs), artificial neural networks (ANNs), convolutional neural networks (CNNs), generative adversarial networks (GANs), transformers, You Only Look Once (YOLO), single-shot multibox detectors (SSDs), fast regional convolutional neural networks (R-CNNs), fast R-CNNs, and R-CNNs.
[0013] Training a neural network with existing organic cation and anion structure datasets means, for example, having the neural network learn the fundamental patterns and relationships within the structures of the existing organic cations and anions. Here, "fundamental patterns" refer to structural patterns within a molecule, such as the ability of different atoms to form valence bonds, and "relationships" refer to bonding properties between atoms, such as the typical bond selectivity found in organic molecules. This allows the neural network to predict the possibility of the next possible part of structural information (e.g., a set of symbols in a SMILES string) given a portion of structural information (e.g., a set of symbols in a SMILES string), enabling navigation of a vast chemical space. Here, "chemical space" refers to a conceptual representation of a molecule's structure in a higher-dimensional space, where each dimension corresponds to a specific feature of the molecule.
[0014] To generate organic cation structure dataset A and organic anion structure dataset B, which are different from existing structure datasets, using a neural network means, for example, generating compound structure information by repeatedly assigning atoms, bonds, rings, and other structural features to a portion of the provided structure information using a trained neural network, thereby generating a structure dataset containing the structure information of multiple compounds, and then generating organic cation structure dataset A and organic anion structure dataset B, which are different from existing structure datasets, by excluding existing structure datasets and those that are not organic cations or organic anions. However, since the neural network may generate undesirable structural information (e.g., strings) such as chemically invalid structures or neutral molecules, it is necessary to guide the generation using a Monte Carlo method that includes a backpropagation step to verify that the generated structure datasets A and B, as described below, are chemically valid, in order to exclude such structural information. Here, "chemically ineffective structure" refers to a structure that cannot theoretically exist as a compound, "chemically effective structure" refers to a structure that can theoretically exist as a compound, and "chemical effectiveness" or "being chemically effective" refers to a structure that can theoretically exist as a compound.
[0015] In this invention, the Monte Carlo method includes a backpropagation step to verify that the generated structure datasets A and B are chemically valid. The backpropagation step is, for example, a step to verify that the number of chemical bonds attached to each atom is correct using RDKit. The Monte Carlo method, including the backpropagation step, guides the generation of the generation process. Other Monte Carlo methods include, for example, Monte Carlo Tree Search (MCTS), Multiplayer Monte Carlo Tree Search (MP-MCTS), Nested Monte Carlo Search (NMCS), Parallel Monte Carlo Tree Search (Parallel-MCTS), Rapid Action Value Estimation (RAVE), Markov Chain Monte Carlo (MCMC) method, Sequential Monte Carlo method, and Maximum Likelihood Estimation method.
[0016] The generation guidelines include, for example, excluding unintended generated structural information, and include steps to verify the molecular charge, size, or chemical effectiveness, as well as steps to ensure novelty by excluding structural datasets identical to existing ones. For example, the generation guidelines for molecular charge might be that cations are monovalent cations, anions are monovalent anions, size might be less than 80 atoms, chemical effectiveness might be that the structure is chemically effective, and ensuring novelty might be that the structure is not included in existing organic cation and anion structural datasets used in the generation process.
[0017] In the present invention, a known ionic liquid refers to one that contains organic cations and anions, has a water content of 1% by weight or less, and has a reference to its water content.
[0018] Cellulose is not particularly limited as long as it is an organic polymer composed of long chains of β-D-glucose units, and examples include Avicel, microcrystalline cellulose, cellulose, cellulose Iα, cellulose Iβ, and cellulose II.
[0019] In this invention, the solubility of cellulose in an ionic liquid is expressed as the weight (wt)% of cellulose relative to the ionic liquid.
[0020] The solubility data for cellulose in ionic liquids are derived from publicly available information, and each data point includes, for example, structural information indicating the ionic liquid structure (e.g., SMILES string), dissolution temperature, dissolution time, and the type of cellulose used.
[0021] In the present invention, the machine learning model used in the prediction process may include, for example, a first machine learning model, but a second machine learning model may also be used. These are intended, for example, to predict the solubility of cellulose in an ionic liquid or to predict the melting point of an ionic liquid. Examples of machine learning models include Random Forest (RF) regression, Extreme Gradient Boosting (XGBoost) regression, Ridge regression, Artificial Neural Network (ANN), Linear regression, Logistic regression, Support Vector Machine (SVM), Decision Tree, Orthogonal Matching Tracking, Lasso regression, ElasticNet, Regression Tree, LightBGM, TensorFlow, PyTrouch, Recurrent Neural Network, Time Delay Neural Network, Gradient Boosting, Set Classification, Generalized Linear Model (GLM), GLMNET, Gradient Boosting Tree, Bayesian Model, k-Nearest Neighbors, Hierarchical Clustering, Non-Hierarchical Clustering, and Topic Model.
[0022] A first machine learning model, trained using known structural information of ionic liquids and data on the solubility of cellulose in those ionic liquids, is, for example, a machine learning model in which a "machine (computational device)" automatically "learns" from known structural information of ionic liquids and data on the solubility of cellulose in those ionic liquids, enabling it to discover the rules and / or patterns underlying the data. This training enables the prediction of cellulose solubility in ionic liquids. Examples of the first machine learning model include random forest (RF) regression, extreme gradient boosting (XGBoost) regression, ridge regression, and artificial neural networks (ANNs).
[0023] Structural dataset C of the ionic liquid, which consists of organic cations in structural dataset A and organic anions in structural dataset B, is a dataset of structural information for the ionic liquid, which consists of organic cations in structural dataset A and organic anions in structural dataset B, both generated in the production process. This ionic liquid can theoretically exist as a compound.
[0024] The melting point of an ionic liquid is expressed as a value in degrees Celsius (°C).
[0025] The melting point data for ionic liquids is derived from publicly available information. For example, if a dataset contains multiple melting point data points from three different public sources, a dataset containing only one of those points can be used.
[0026] A second machine learning model, trained using known structural information and melting point data of ionic liquids, is, for example, a machine learning model in which a "machine (computational device)" automatically "learns" from known structural information and melting point data of ionic liquids, enabling it to discover the rules and / or patterns behind the data. This training enables the prediction of the melting point of ionic liquids. Examples of second machine learning models include random forest (RF) regression, extreme gradient boosting (XGBoost) regression, ridge regression, and artificial neural networks (ANNs).
[0027] Predicting the solubility or melting point of cellulose involves having a first or second machine learning model related to the prediction process predict the solubility of cellulose in an ionic liquid or the melting point of an ionic liquid based on discovered rules and / or patterns. For the prediction of cellulose solubility, the solubility of cellulose in an ionic liquid is not particularly limited, but for example, values of 5% or more by weight, 8% or more by weight, 10% or more by weight, or 15% or more by weight are adopted. Furthermore, the conditions for predicting cellulose solubility are not particularly limited, but for example, the heating temperature is 50°C to 80°C, 80°C to 100°C, or 100°C to 120°C, the heating time is not particularly limited but for example, 3 to 24 hours, 6 to 15 hours, or 9 to 12 hours, and the type of cellulose is not particularly limited but for example, Avicel, microcrystalline cellulose, cellulose, pulp, or cotton linter. Regarding the prediction of the melting point of an ionic liquid, the melting point of the ionic liquid is not particularly limited, but for example, those with a melting point of less than 80°C, less than 70°C, less than 60°C, or less than 50°C are used.
[0028] Predicting the solubility or melting point of cellulose in an ionic liquid structure dataset C, which consists of organic cations in structure dataset A and organic anions in structure dataset B, generated in the production process, using the first or second machine learning model related to the prediction process, means, for example, having the first or second machine learning model related to the prediction process predict the solubility of cellulose in the ionic liquid that generated the structure dataset of its constituent elements in the production process, or the melting point of the ionic liquid, based on discovered rules and / or patterns. Details of the prediction are as described above.
[0029] The present invention may further include a verification step in which the solubility of cellulose in the ionic liquid used for prediction in the prediction step is verified using a Conductor-like Model for Real Solvents (COSMO-RS).
[0030] The Conductor-Like Model for Real Solvents (COSMO-RS) is a quantum chemical model known to be universally applicable to predicting the properties of ionic liquid systems. Previous studies have demonstrated that it can effectively predict the solubility of cellulose in ionic liquids by calculating the logarithmic activity coefficient (lnγ) (Green Chem. 18, 6246-6254 (2016), Fluid Phase Equilibria 475, 25-36 (2018)).
[0031] In the prediction process, verifying the solubility of cellulose using COSMO-RS for the ionic liquid used for prediction means, for example, determining the solubility of cellulose using the linear correlation between the solubility obtained by the first machine learning model and lnγ for the ionic liquid, and confirming whether it differs significantly from the solubility obtained by the first machine learning model.
[0032] In the present invention, the ionic liquid used for prediction in the prediction step can also be actually synthesized to obtain measured values of cellulose solubility, and it can be confirmed whether these values differ significantly from the solubility obtained by the first machine learning model.
[0033] An ionic liquid that is soluble in cellulose refers to a liquid whose solubility of cellulose is, for example, 5% or more by weight, 8% or more by weight, 10% or more by weight, or 15% or more by weight.
[0034] Screening for ionic liquids that are soluble in cellulose means selecting ionic liquids that have high solubility in cellulose.
[0035] One embodiment of the screening method according to the present invention, which includes a generation step and a prediction step, is not particularly limited, but can be performed by a system consisting of one or more personal computers or cloud computing systems, and the system includes neural network means and means for performing the Monte Carlo method, a machine learning model, and a structured dataset, which are stored as software in the storage device of one of the computers or computing systems, and it is preferable that the generation step and the prediction step are performed by a computing device on one of the computers or computing systems. [Examples]
[0036] (1) Method (A) Method for a novel organic ion generator The process for generating novel organic cations is shown in Figure 1, and a similar approach can be applied to the generation of organic anions. This process can be broadly divided into two stages: (I) Training an RNN on a large dataset of existing ion structures; and (II) Using the MCTS algorithm to guide the RNN toward the generation of novel ion structures.
[0037] In the initial stage, to compile the existing cation / anion dataset, we first obtained, as one approach, a list of 117,700,516 SMILES (Simplified Molecular Input Line Entry System) strings corresponding to all CIDs from the PubChem FTP site (ftp: / / ftp.ncbi.nlm.nih.gov / pubchem / Compound / Extras / CID-SMILES.gz). SMILES compactly encodes molecular graphs as human-readable strings using symbols representing atoms, bonds, rings, and other structural features (J.Chem.Inf.Comput.Sci.28, 31-36(1988)). For example, in SMILES notation, lowercase "c" and uppercase "C" represent aromatic carbon atoms and aliphatic carbon atoms, respectively, and "O" represents oxygen. Symbols such as "-", "=", and "#" correspond to single bonds, double bonds, and triple bonds, respectively. Furthermore, charged atoms can also be represented; for example, negatively charged oxygen is represented as "[O-]", nitrogen with one positive charge as "[NH2+]", and nitrogen with two positive charges as "[N+2]", etc. A filtering step was then applied to select organic cations / anions from this dataset. As a result of this filtering process, the final dataset contained 954,085 organic cation SMILES strings and 486,159 organic anion SMILES strings. All SMILES strings were preprocessed by adding a start symbol "^" and an end symbol "$". They were then vectorized using one-hot encoding for RNN training.
[0038] The RNN was constructed using Keras 2.13.1 (https: / / github.com / keras-team / keras), which has an architecture similar to that of Sci.Technol.Adv.Mater.18, 972-976 (2017). The RNN consists of two stacked gated recurrent units (GRUs) (Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation.arXiv September 2, 2014.https: / / doi.org / 10.48550 / arXiv.1406.1078) each with 256 hidden units, followed by a final output layer with a softmax activation function (Int.J.Eng.Appl.Sci.Technol.04, 310-316 (2020)). To prevent overfitting, a dropout rate of 0.2 was applied to the GRU layer. The network was trained using the Adam optimizer with a learning rate of 0.001 (A Method for Stochastic Optimization.arXiv January 29, 2017.https: / / doi.org / 10.48550 / arXiv.1412.6980) and a categorical cross-entropy loss function (Proceedings of the 40th International Conference on Machine Learning;PMLR, 2023;pp23803-23828). During training, the RNN learns the basic patterns and relationships within the SMILES data of cations and anions. This enables it to predict the likelihood of the next possible symbol given a set of symbols. The complete SMILES string can be generated, as shown in Figure 1b, by repeatedly inputting the symbol with the highest probability predicted in the NumPy 1.22.3 polynomial sampler into an RNN until the termination symbol "$" is predicted or the string length reaches the predefined limit of 80 symbols.
[0039] However, RNNs can generate undesirable strings such as chemically invalid structures or neutral molecules. Therefore, in the second stage, the MCTS algorithm was used to guide the generation. As shown in Figure 1c, MCTS creates a search tree where each node corresponds to one symbol. Starting from the root node "^", the search tree gradually grows by repeating the following four steps. (I) Selection: At each level of the search tree, the node with the highest Upper Limit of Confidence (UCB) score (IEEE Trans.Comput.Intell.AI Games 4, 1-43 (2012), Mach.Learn. 47, 235-256 (2002)) is selected. The UCB score evaluates the merit of a node by considering both its potential for exploration and exploitation, and was calculated based on the previous loop iterations. (II) Expansion: The selected edge nodes are expanded by adding new child nodes. These child nodes are generated using the functionality of the RNN trained in the first stage (predicting the next possible symbol). (III) Simulation: For each child node, perform a rollout by repeatedly running RNN prediction until 10 complete SMILES strings are generated. (IV) Backpropagation: Evaluate the SMILES strings generated in the simulation step. A child node receives a reward of 1 if it generates at least one chemically valid novel cation / anion SMILES; conversely, it receives a reward of 0 if it does not generate any desirable SMILES. Chemical validity is checked using RDKit 2023.9.4, and novelty is verified by identifying whether the SMILES string is included in the PubChem dataset. Subsequently, following the assignment of rewards, all UCB scores of nodes along the path from the current node back to the root node are calculated using the UCB equation (Mach.Learn.47, 235-256(2002)):
number
number
[0040] The MCTS process is repeated in the simulation step until the desired number of cations / anions are generated.
[0041] (B) Methods for predictive ML models To develop a predictive ML model for predicting the solubility and melting point of cellulose, we collected a comprehensive dataset. Regarding the solubility of cellulose, 41 published studies (J.Am.Chem.Soc.124, 4974-4975 (2002), RSC Adv.2, 2476 (2012), Green Chem.13, 2507 (2011), Cellulose 15, 59-66 (2008), J.Mol.Liq.197, 211-214 (2014), Green Chem 12, 268-275 (2010), Green Chem.11, 417 (2009), Chem Commun 47, 511-513 (2011), Green Chem 16, 1326-1335 (2014), Chem.Soc.Rev.41, 1519 (2012), Sci.China) Chem.59, 1421-1429(2016), Green Chem.10, 696-705(2008), Biomacromolecules 7, 3295-3297(2006), Carbohydr.Polym.117, 666-672(2015), Green Chem.10, 44-46(2008), Macromol.Biosci.5, 520-525(2005), Green Chem.9, 1229-1237(2007), Macromol.Biosci.7, 440-445(2007), J.Chem.Technol.Biotechnol.84, 1818-1827(2009), New J.Chem.35, 1596-1606(2011), Macromolecules 51, 4158-4166(2018), Green Chem.14, 304-307(2012), Biotechnol.Bioeng.112, 65-73(2015), RSC Adv.2, 8429-8438(2012), ChemSusChem 5, 388-391(2012), Eur.Polym.J.92, 204-212(2017), J.Mol.Liq.234, 111-116(2017), Carbohydr.Polym.229, 115594(2020), Green Chem.14, 2922-2932(2012), Carbohydr.Polym.130, 18-25(2015), J.Phys.Chem.B 123, 3994-4003(2019), Ind.Eng.Chem.Res.49, 11809-11813(2010), Green Chem.13, 2719-2722(2011), Molecules 25, 3539(2020), Aust.J.Chem.72, 55-60(2018), New J.Chem.43, 2299-2306(2019), New J.Chem.43, 4554-4561(2019), J.Mol.Liq.292, 111353(2019), Macromolecules From 53, 3284-3295 (2020), Ind.Eng.Chem.Res. 58, 16009-16017 (2019), and Phys.Chem.Chem.Phys. 17, 32276-32282 (2015), 674 experimental data points regarding the solubility of cellulose in 332 ILs were compiled. Each data point includes the SMILES string indicating the IL structure, dissolution temperature, dissolution time, and the type of cellulose used (e.g., Avicel, microcrystalline cellulose, and cellulose). In particular, only solubility experiments performed on pure IL (without cosolvent) using normal heating conditions (not conduction heating or microwave heating) were included. For melting point prediction, a unified dataset of 2276 data points after eliminating duplicate melting point data was collected from datasets based on three publications (Ind.Eng.Chem.Res.61, 4683-4706 (2022), Molecules 26, 2454 (2021), Melting Points of Ionic Liquids: Review and Evaluation. Green Energy Environ. (2024) doi:10.1016 / j.gee.2024.01.009).
[0042] To prepare the data for ML modeling, RDKit 2023.9.4 (https: / / www.rdkit.org / ) was used to convert the SMILES strings in IL to numerical descriptors. This invention utilizes a combination of MACCS (Molecular Access System) keys (J.Chem.Inf.Comput.Sci.42, 1273-1280 (2002)) and 210 2D molecular descriptors calculated by RDKit. For cellulose, other features were also incorporated, resulting in a final vector form with 759 features: {MACCS for cations, 2D descriptors for cations, MACCS for anions, 2D descriptors for anions, dissolution temperature, dissolution time, and one-hot encoded cellulose type}. For the melting point model, only MACCS keys and 2D molecular descriptors were involved, resulting in a final vector form with 754 features: {MACCS for cations, 2D descriptors for cations, MACCS for anions, and 2D descriptors for anions}.
[0043] After data preprocessing was complete, model training and evaluation were carried out. For each model, the dataset was randomly divided into two groups: a training set (approximately 80% of the data) used for model development, and a test set (approximately 20% of the data) for evaluating the model's performance. In this invention, tree-based regression models (Random Forest (RF), Extreme Gradient Boosting (XGBoost)) and ridge regression models were implemented using Scikit-learn 1.3.0 (API Design for Machine Learning Software: Experiences from the Scikit-Learn Project.arXiv September 1, 2013: https: / / doi.org / 10.48550 / arXiv.1309.0238) with default hyperparameters. Using Keras 2.13.1, we constructed an artificial neural network (ANN) model with four fully connected feedforward layers activated by the Rectified Linear Unit (ReLU) function (Int.J.Eng.Appl.Sci.Technol.04, 310-316(2020)). The number of units in the input layer and each of the two hidden layers varied based on the dimensionality of the input vector. The output layer contained one unit. To mitigate overfitting, a dropout rate of 0.5 was applied to each hidden layer. The network was trained using an Adam optimizer with a learning rate of 0.0001 and a mean squared error (MSE) loss function.
[0044] Coefficient of determination (R 2 The mean squared error (MSE) and the IL descriptor were used as statistical measures to evaluate the model's performance on the test set. These metrics provide insight into how accurately the model captures the relationship between the IL descriptor and the target property (cellulose solubility or melting point). The corresponding equations are as follows:
number
number
number
number
[0045] While evaluating the performance of a model is important, understanding the model's applicable region (AD) is equally crucial. The AD represents a theoretical region within the chemical space defined by the training set. Data points within the AD are considered to be interpolated with low uncertainty and are therefore highly reliable. Conversely, data points outside the AD represent extrapolations of the model and are highly uncertain (Chemom.Intell.Lab.Syst.145, 22-29(2015), J.Chem.Inf.Model.59, 181-189(2019)). In this invention, we use an approach similar to that of Chemom.Intell.Lab.Syst.145, 22-29(2015) to determine whether a target data point lies within the AD for a given trained ML model. First, all descriptors of the target data point were standardized using the following formula:
number
number
number
number
number
[0046] In addition to model validation and AD determination, the reliability of a predictive model can be increased by rationally explaining the operating principles of the predictive model. While ML models often yield excellent results, interpretability remains a challenge. To address this challenge and gain insight into the relationship between descriptors and predicted properties, particularly the influence (positive or negative) of specific descriptors, this invention utilizes Shapley Additive exPlanations (SHAP) 0.44.1 (A Unified Approach to Interpreting Model Predictions. arXiv November 24, 2017. http: / / arxiv.org / abs / 1705.07874) for model interpretation. SHAP provides a framework for explaining predictions in complex models, enabling an understanding of how each descriptor contributes to the final prediction.
[0047] (C) COSMO-RS Model Method To validate the results obtained from high-throughput virtual screening, we used COSMO-RS (WIREs Comput.Mol.Sci.8, e1338(2018), Annu.Rev.Chem.Biomol.Eng.1, 101-122(2010)), a quantum chemical model known to be universally applicable to predicting the properties of IL systems (Green Chem.18, 6246-6254(2016), Fluid Phase Equilibria 475, 25-36(2018), J.Chem.Eng.Data 48, 475-479(2003), Green Energy Environ.3, 247-265(2018), Phys.Chem.Chem.Phys.19, 11835-11850(2017), Green Chem.12, 2172-2181(2010)). Previous studies have demonstrated the effectiveness of COSMO-RS in predicting the solubility of cellulose in IL by calculating the logarithmic activity coefficient (lnγ) (Green Chem. 18, 6246-6254 (2016), Fluid Phase Equilibria 475, 25-36 (2018)). In this invention, we implemented an automated workflow for calculating the lnγ of cellulose in IL using a custom Python script. As shown in Figure 2, the workflow proceeded as follows: (I) Use RDKit 2023.9.4 to generate the 3D structures of the IL cation and anion separately. (II) Optimize these 3D structures using the Universal Force Field (UFF) of RDKit (J.Am.Chem.Soc.114, 10035-10046 (1992)). (III) Quantum chemical optimization of the UFF-optimized structures of cations and anions in IL, as well as the crystal structure of cellotetraose obtained from Cellulose Builder (https: / / code.google.com / archive / p / cellular-builder / ), using the Gaussian16 package at the BP86 / TZVP level. (IV) Check the vibration frequencies and, if imaginary frequencies exist, re-optimize the structure until imaginary frequencies no longer exist. (V) Using BVP86 / TZVP level theory, the COSMO files for these optimized structures are computed individually. (VI)Calculate the lnγ of cellotetraose using the COSMOtherm package and parameterized BP_TZVP_23. During the calculation, the mole fraction of cellotetraose was set to 0.5, the mole fraction of the cation to 0.25, and the mole fraction of the anion to 0.25.
[0048] (2) Results Example 1 - Generation of Novel IL To explore the vast and undeveloped IL chemical space, we developed a novel organic ion generator by combining RNN and MCTS methods. Using this generator, we generated approximately 900 billion novel IL candidates. These ILs consist of 961,859 novel cations and 957,070 novel anions.
[0049] To investigate the distribution of these cations and anions, their chemical structures are visualized in 2D space using Uniform Manifold Approximation and Projection (UMAP) (UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, arXiv September 17, 2020. https: / / doi.org / 10.48550 / arXiv.1802.03426). Before applying UMAP, the chemical structure of each ion is converted into a 1024-bit vector with an Extended Connected Fingerprint (ECFP4) (J. Chem. Inf. Model. 50, 742-754 (2010)) having a radius of two atoms. ECFP4 essentially encodes the structural information of each ion as a digital fingerprint. UMAP then facilitates the visualization of these high-dimensional fingerprints in a lower-dimensional (usually 2D) space for efficient exploration.
[0050] The results of the UMAP visualization shown in Fig. 3 show the remarkable diversity of the cations (Fig. 3a) and anions (Fig. 3b) generated within the chemical space. The gray dots representing each of the generated ions cover a vast range, emphasizing that extensive exploration has been achieved by the generator of new ions. To further demonstrate this structural diversity, 10 cations and 10 anions were randomly selected and shown in Fig. 3c. The structures of these selected ions are constructed from diverse chemical fragments and ionic cores, not only showing the extensive diversity of the generated IL library but also highlighting the novel properties of the generated ILs by showing completely new structures.
[0051] Example 2 - Development of predictive ML models To facilitate high-throughput virtual screening, two predictive ML models were constructed. The first model aimed to predict the cellulose solubility in ILs and enable the identification of ILs with high cellulose solubility. The second model was developed to predict the melting point of ILs and facilitate the filtering of solid ILs at a specified temperature setting.
[0052] Before developing the cellulose solubility prediction model, an extensive literature survey was conducted to collect experimental data on cellulose solubility in ILs, and a comprehensive dataset containing 674 data points for 332 ILs obtained from the 41 published studies mentioned above was created. Subsequently, this dataset was used to proceed with the development of an ML model for cellulose solubility prediction. Four different ML models, namely RF regression, XGBoost regression, Ridge regression, and ANN, were used for training. Fig. 4 shows the comparison of the predicted values and actual values of the four models for both the training set and the test set. However, none of the models used exceeded a test R 2 of more than 0.7.
[0053] As reported in previous experimental studies (Phys.Chem.Chem.Phys.17, 5767-5775 (2015), Cellulose 16, 207-215 (2009)), the effect of water content on cellulose solubility measurements in ILs was considered, and data points with water content exceeding 1% by weight, or data points where there was no mention of water content in the respective literature, were excluded from the dataset. As a result of this control process, a refined dataset containing 379 data points across 187 ILs was obtained. Figure 5 shows the performance of the four models on this refined dataset. In particular, the performance of all models was significantly improved compared to that observed across the entire dataset (Figures 5a-5d). The ANN model showed the best R, as shown in Figure 6a. 2 It achieved the lowest MSE of 0.879 and the lowest MSE of 8.007, demonstrating the best performance. The RF and XGBoost regression models also showed high accuracy. To further verify the robust properties of each model, 10 cross-validations were performed on each model. As detailed in Table 1, the ANN model again showed superior performance compared to the other models, and the highest R 2 It showed a score of 0.831, and the lowest MSE score was 8.466.
[0054] [Table 1]
[0055] The SHAP method was used to interpret the insights behind the ML model. The SHAP approach addresses the unexplained black-box challenge of ML algorithms by calculating the contribution of features to the model output. Traditional feature importance only indicates the importance of a feature, but does not reveal how it affects the model's predictions. The main advantage of SHAP values is their ability to show the impact of each feature on each sample, revealing whether the impact is positive or negative. Figure 6b shows the distribution of SHAP values for the ANN model. The plot displays the 10 features that have the greatest impact on the model output, ordered from maximum impact to minimum impact. Each example is represented by one dot in the row for each feature. The color of the dots represents the relative feature value from minimum (light black) to maximum (black). The position of the dots on the horizontal axis corresponds to the SHAP value of that feature, reflecting its impact on the model output. The dots "stack" along the row for each feature to show density.
[0056] The heating temperature feature at the top of the plot has the greatest impact on the model output. Almost all of the black dots in this feature's row are located to the right along the horizontal axis, while the lighter black dots are located to the left. This indicates that higher values for this feature have a greater positive impact on the model output. These results suggest that heating temperature has a positive impact on the solubility of cellulose, which is consistent with experimental observations. Furthermore, the Ipc of the cation has a negative impact on the solubility of cellulose. The descriptor Ipc represents the information of the coefficients of the characteristic polynomial in the adjacency matrix of the molecule's hydrogen suppression graph (J. Chem. Phys. 67, 4517-4533 (1977)), reflecting the overall complexity and branching of the molecule. In fact, the structures of the most effective cations are relatively simple, such as 1-ethyl-3-methylimidazolium ([EMIM]) and 1-allyl-3-methylimidazolium ([AMIM]). The descriptor BertzCT is the sum of two terms: one representing the complexity of the bond and another representing the complexity of the heteroatom distribution (J.Am.Chem.Soc.103, 3599-3601 (1981)). This plot shows that BertzCT makes a positive contribution to cellulose solubility for cations and a negative contribution for anions. This observation is somewhat rationalized. For example, the transition from ammonium cations to imidazolium cations usually involves an increase in both the complexity of the bond and the distribution, often resulting in improved cellulose solubility. Conversely, for anions, bis(trifluoromethylsulfonyl)imide ([Tf2N]) shows increased complexity of the bond and heteroatom distribution compared to carboxylate and halogen anions, but with lower cellulose solubility. These SHAP analyses support the general agreement between the ANN model and experimental observations, highlighting the reliability of the developed model. A detailed explanation of the other features is shown in Table 2.
[0057] [Table 2-1] [Table 2-2]
[0058] To develop a melting point prediction model, a dataset was compiled by combining 2276 data points related to melting point from datasets based on the three aforementioned literatures. Then, four different ML models—RF regression, XGBoost regression, ridge regression, and ANN—were used for training. As shown in Figure 7, the RF regression model showed the best performance and the highest R 2 It achieved a minimum MSE of 123.353 with an accuracy of 0.863. This level of accuracy was consistent across 10 cross-validations. As detailed in Table 3, the RF regression model again demonstrated superior performance compared to other models, achieving the highest R 2 It showed a score of 0.858, with the lowest MSE of 523.089.
[0059] [Table 3]
[0060] Furthermore, the SHAP method was used to interpret this RF regression model. In Figure 6d, the feature with the largest contribution to the melting point is the maximum absolute partial charge of the anion. Specifically, this feature shows a positive contribution to the melting point, indicating that the higher the value of the maximum absolute partial charge of the anion, the higher the melting point of IL. Conversely, the minimum partial charge of the anion shows a negative contribution to the melting point of IL. This observation complements the finding related to the maximum absolute partial charge of the anion and suggests a potential pattern: the larger the negative partial charge of the anion, the higher the melting point of IL. This is reasonable, as a high negative partial charge of the anion usually indicates a stronger interaction between the anion and cation, resulting in a higher melting point. The second most influential feature is the cation's BCUT2D_LOGPLOW, which shows a negative contribution to the melting point of IL. BCUT2D_LOGPLOW estimates the logarithm (LogP) of the partition coefficient, also known as the octanol-water partition coefficient, and focuses on regions with low LogP values based on the BCUT2D metric (J. Chem. Inf. Comput. Sci. 29, 225-227 (1989), J. Chem. Inf. Comput. Sci. 41, 402-407 (2001)). A lower LogP value indicates that the surface region of the cation is more hydrophilic, which usually interacts more favorably with the charged anion, resulting in a higher melting point. These SHAP analyses highlight the rationality of the developed model. Additional explanations of other features are provided in Table 2.
[0061] Example 3 - High-throughput virtual screening Following the development of the predictive ML model, a high-throughput virtual screening process was performed on approximately 900 billion generated IL candidates using the predictive ML model, as shown in Figure 8. As shown in Figures 8a and 8b, the UMAP plots show that cations and anions in existing ILs are clustered in small regions within the space of cations and anions in generated ILs, highlighting the vast undeveloped chemical landscape created by the ion generator. First, AD screening was performed for each generated IL. AD represents a theoretical region within the chemical space defined by the training set. Data points within AD are considered more reliable. As a result, 667,637 ILs (including 5,367 cations and 517 anions) present in both the model AD predicting cellulose solubility and melting point were excluded. The UMAP plots also revealed that these ILs occupy a space that closely overlaps with the ILs in the training set used to develop the predictive ML model (see Figures 8c and 8d). This arrangement is logical considering the AD calculation method, which depends on descriptor similarity (see the section "Predictive ML Model Method").
[0062] Next, the cellulose solubility and melting point of each IL located within the AD were predicted using the predictive models developed in the "Generation of Novel ILs" section (ANN for cellulose solubility prediction and RF regression for melting point prediction). Specifically, for each cellulose solubility prediction, the heating temperature was set to 80°C, the heating time to 24 hours, and the cellulose type to microcrystalline cellulose. Finally, 796 promising ILs (containing 108 cations and 106 anions) were identified with predicted cellulose solubility exceeding 15 wt% and a melting point below 80°C. The cations and anions of these promising ILs were found to be densely clustered in one space (Figures 8e and 8f). This indicates a high degree of structural similarity between these ions. Specifically, most of the promising ILs consist of imidazolium cations paired with carboxylate or oleate anions. This is mainly due to the limitations of the training set. During AD screening, most ILs that were not similar to the training set were eliminated. Due to the limited diversity of the training set, the diversity of the screening results is also limited.
[0063] Promising ILs identified by a high-throughput virtual screening process were subsequently subjected to cross-validation using the COSMO-RS model. COSMO-RS calculations for the activity coefficient (γ) have been demonstrated to be effective in predicting the cellulose solubility of ILs. Previous studies have observed a linear correlation between experimentally determined cellulose solubility and lnγ calculated using COSMO-RS (Green Chem. 18, 6246-6254 (2016), Fluid Phase Equilibria 475, 25-36 (2018)). In this invention, to investigate this linear correlation, the lnγ values of 10 ILs with experimentally measured cellulose solubility were calculated using COSMO-RS. Despite being derived from various published studies (J.Phys.Chem.B 123, 3994-4003 (2019), Molecules 25, 3539 (2020), New J.Chem.43, 4554-4561 (2019)), it should be noted that the cellulose solubility measurements of these ILs were performed using microcrystalline cellulose, specifically under conditions similar to the ML model of the present invention, with a heating temperature of 80°C and a heating time of 24 hours. See Table 4 for details of these 10 ILs. The calculated lnγ values were compared with the corresponding experimental cellulose solubility values, as shown in Figure 9, R 2 A linear regression of 0.77 was obtained. Therefore, the cellulose solubility of IL can be calculated using linear regression from the lnγ value of COSMO-RS.
[0064] [Table 4-1] [Table 4-2]
[0065] Example 4 - Verification using COSMO-RS To cross-validate the results of the high-throughput virtual screening, cellulose solubility was recalculated using COSMO-RS. The results showed that of the 796 ILs with ML-predicted solubility exceeding 15% by weight, 544 (approximately 68%) also had a COSMO-RS prediction of exceeding 15% by weight. Table 5 shows detailed information on the 10 ILs with the highest ML-predicted cellulose solubility, along with their solubility obtained using COSMO-RS. This cross-validation demonstrates the effectiveness of the high-throughput virtual screening method in identifying ILs with high cellulose solubility.
[0066] [Table 5-1] [Table 5-2]
[0067] (3) Discussion In Example 3, the cellulose solubility determined by COSMO-RS generally agreed with the cellulose solubility determined experimentally, and in Example 4, the cellulose solubility determined by the ML model generally agreed with the cellulose solubility determined by COSMO-RS. Therefore, it can be inferred that the cellulose solubility determined by the ML model generally agrees with the cellulose solubility determined experimentally. This is because, as shown in Figure 5d, the cellulose solubility determined by the machine learning model for predicting cellulose solubility and the solubility determined experimentally are in R 2 This is supported by the fact that they agree with an accuracy of 0.879. [Industrial applicability]
[0068] This invention enables rapid and efficient screening of ionic liquids soluble in cellulose in undeveloped IL chemical spaces.
Claims
1. A generation step involves training a neural network using existing organic cation and anion structure datasets, and using the neural network to generate organic cation structure dataset A and organic anion structure dataset B, which are different from the existing structure datasets. A prediction step is performed to predict the solubility of cellulose in an ionic liquid structure dataset C, which consists of organic cations in structure dataset A and organic anions in structure dataset B, generated in the production step, using a first machine learning model trained with known structural information of an ionic liquid and data on the solubility of cellulose in the ionic liquid. Includes, In the generation process, the generation is guided using a Monte Carlo method that includes a backpropagation step to verify that the generated structure datasets A and B are chemically effective. A method for screening ionic liquids that are soluble in cellulose.
2. The method according to claim 1, wherein the neural network means is a recurrent neural network (RNN).
3. The method according to claim 1, wherein the Monte Carlo method is Monte Carlo tree search (MCTS).
4. The method according to claim 1, wherein the prediction step includes predicting the melting point of an ionic liquid structure dataset C, which consists of organic cations in structure dataset A and organic anions in structure dataset B, generated in the generation step, using a second machine learning model trained with known structural information of the ionic liquid and data on the melting point of the ionic liquid.
5. The method according to claim 1, further comprising a verification step of verifying the solubility of cellulose in the ionic liquid used for prediction in the prediction step using a Conductor-like Model for Real Solvents (COSMO-RS).