Solvent search system for high polymer, model creation system, solubility prediction system, solvent search method for high polymer, and program
The solvent search system uses molecular dynamics and machine learning to efficiently predict solvent suitability for polymers, addressing inefficiencies in existing solvent search methods and improving prediction accuracy.
Patent Information
- Application Number
- JP2024224669
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-21
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
AI Technical Summary
Existing methods are inefficient in searching for solvents that dissolve polymers, particularly crystalline polymers like cellulose, due to the time-consuming nature of molecular dynamics simulations.
A solvent search system utilizing molecular dynamics simulation, machine learning, and prediction models to efficiently determine solvent solubility by calculating index values and predicting solvent suitability for polymers, including ionic liquids.
Enhances the efficiency of solvent search for polymers by improving prediction accuracy and reducing computational costs, enabling rapid identification of effective solvents.
Smart Images

Figure 2025100487000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a solvent search system for polymers, a model creation system, a solubility prediction system, a method for searching for a solvent for polymers, and a program.
Background Art
[0002] Conventionally, ionic liquids are known as solvents for dissolving cellulose. For example, Non-Patent Document 1 can be cited as a study on the dissolution mechanism. In this study, for example, as a result of analyzing the dissolution of the model cellulose crystal structure in an imidazolium-based ionic liquid by a molecular dynamics approach, it is said that the dissolution involves the cleavage of hydrogen bonds between cellulose molecular chains due to the penetration of the ionic liquid. In addition, in ionic liquids having high dissolving power, both anions and cations contribute to the cleavage of intermolecular hydrogen bonds, but it is said that this cleavage does not occur sufficiently in ionic liquids with low solubility.
Prior Art Documents
Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventionally, it has been difficult to efficiently search for a solvent suitable for a polymer. That is, it takes time to determine the solubility of a polymer in various solvents by molecular dynamics simulation. The present disclosure aims to provide a technique for improving the search efficiency of a solvent that dissolves a polymer.
Means for Solving the Problems
[0005] The present disclosure can be realized by the following aspects. (Aspect 1) A learning data generation process for calculating an index value representing the solubility of a predetermined polymer in a solvent to be learned by molecular dynamics simulation, and A model creation process for creating a prediction model that learns the relationship between a feature amount based on the chemical structure of the solvent to be learned and an index value representing the solubility of the predetermined polymer by machine learning, and outputs a predicted value representing the solubility of the predetermined polymer in a solvent to be predicted; and A search process for determining the solvent to be predicted based on a predetermined algorithm, and repeatedly outputting the predicted value using the feature amount based on the chemical structure of the solvent to be predicted and the prediction model; and A solvent search system for a polymer, including one or more computers that execute the above. (Aspect 2) For the solvent to be predicted for which it is determined that the predicted value satisfies the predetermined criterion, perform the molecular dynamics simulation in the learning data generation process, and repeat the learning data generation process, the model creation process, and the search process The solvent search system according to Aspect 1. (Aspect 3) The molecular dynamics simulation is performed using a force field based on the three-dimensional molecular structure of the solvent to be learned and the predetermined polymer, and Information representing the three-dimensional structure of the solvent to be predicted for which it is determined that the predicted value satisfies the predetermined criterion in the search process is input to the learning data generation process The solvent search system according to Aspect 2. (Aspect 4) The solvent to be predicted contains an ionic liquid The solvent search system according to any one of Aspects 1 to 3 (Aspect 5) The predetermined polymer is a crystalline polymer The solvent search system according to any one of Aspects 1 to 4 (Aspect 6) The predetermined polymer is cellulose and / or its derivative The solvent search system according to any one of Aspects 1 to 5 (Aspect 7) The index value includes at least any one of the number or binding energy of hydrogen bonds formed within or between the predetermined polymers, the self-diffusion coefficient of the solvent to be learned, the mean square displacement of the predetermined polymer in the solvent to be learned, the peak value of the first radial distribution function between the solvent to be learned and the predetermined polymer, the free energy change amount with respect to the peak value of the first radial distribution function, the peak value of the second radial distribution function between two components when the solvent contains a plurality of components, and the free energy change amount with respect to the peak value of the second radial distribution function The solvent search system according to any one of Aspects 1 to 6 (Aspect 8) The predetermined solvent is an ionic liquid, and the index value includes the peak value of the radial distribution function between the cation and the anion, or the free energy change amount with respect to the peak value The solvent search system according to Aspect 7 (Aspect 9) The feature quantity based on the chemical structure of the solvent to be learned and the feature quantity based on the chemical structure of the solvent to be predicted are represented by a molecular descriptor-based model or a graph-based model The solvent search system according to any one of Aspects 1 to 8 (Aspect 10) Calculating an index value representing the solubility of a predetermined polymer in a solvent to be learned by molecular dynamics simulation Learning, by machine learning, the relationship between the feature quantity based on the chemical structure of the solvent to be learned and the index value representing the solubility of the predetermined polymer, and creating a prediction model that outputs a predicted value representing the solubility of the predetermined polymer with respect to the solvent to be predicted. A model creation system including one or more computers that execute the above. (Aspect 11) Obtaining information representing the chemical structure of the solvent to be predicted. A prediction model created by learning, by machine learning, the relationship between the feature quantity based on the chemical structure of the solvent to be learned and the index value representing the solubility of the predetermined polymer with respect to the solvent to be learned, which is calculated by molecular dynamics simulation, and using the feature quantity based on the chemical structure of the solvent to be predicted, outputting a predicted value representing the solubility of the predetermined polymer with respect to the solvent to be predicted. A solubility prediction system including one or more computers that execute the above.
[0006] Note that the content of the means for solving the problem can be provided as a device such as a computer or a system including a plurality of devices, a method executed by one or more computers, or a program to be executed by one or more computers. Note that a recording medium storing the program may be provided.
Effect of the Invention
[0007] According to the disclosed technology, it is possible to provide a technology for improving the search efficiency of a solvent that dissolves a polymer.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0009] Hereinafter, embodiments will be described with reference to the drawings.
[0010] FIG. 1 is a diagram for explaining the outline of an embodiment. In the present embodiment, a prediction model for predicting the solubility of a polymer in a solvent is created by machine learning, and a solvent having good solubility of the polymer is searched using the prediction model. Specifically, learning data generation by the molecular dynamics method (hereinafter referred to as the "MD (Molecular Dynamics) method") (FIG. 1: (1)), prediction model creation using the learning data (FIG. 1: (2)), and solvent search using the prediction model (FIG. 1: (3)) are performed. Further, learning data is created using the found solvent, and the prediction model is reconstructed using the learning data. That is, the processes from (1) to (3) above are repeated (FIG. 1: (4)).
[0011] The polymer is not particularly limited, and may be, for example, a crystalline polymer such as cellulose or its derivative. Since crystalline polymers are poorly soluble in water and general organic solvents and have poor processability, it is particularly useful to develop a solvent that exhibits excellent solubility in crystalline polymers.
[0012] Also, the solvent is not particularly limited, and may be, for example, an ionic liquid, a deep eutectic solvent (DES), or the like. An ionic liquid is a salt of a cation and an anion It is a room-temperature molten salt that is in a liquid state at approximately 100°C or lower. Ionic liquids are structurally diverse and require time to be extensively explored through experimental operations such as MD methods. Therefore, it can be said that it is particularly useful if efficient exploration of ionic liquids can be achieved. Note that the solvent may be a neutral molecule such as an organic solvent, or a mixed solvent of these and an ionic liquid.
[0013] In learning data generation, for example, an index value representing solubility (also referred to as a "parameter related to solubility") is obtained by molecular dynamics simulation (hereinafter referred to as MD simulation). MD simulation can be performed by existing methods. That is, the force fields of the solvent and the polymer are defined, and the simulation is performed by placing the polymer in the solvent. The index value representing solubility may be, for example, at least any one of a parameter representing the interaction between polymers, the self-diffusion coefficient of the solvent, the mean square displacement of the polymer in the solvent, the peak value of the radial distribution function between the solvent and the polymer, the free energy change amount with respect to the peak value, the peak value of the radial distribution function between two components in the solvent, and the free energy change amount with respect to the peak value.
[0014] In prediction model creation, the relationship between the feature amount of the solvent and the index value representing solubility is machine-learned. Machine learning can be performed by existing methods such as deep learning, linear regression, and decision tree regression. The feature amount of the solvent as the explanatory variable can be represented by, for example, a molecular descriptor-based model, a graph-based model, etc. The molecular descriptor-based model may be, for example, a molecular fingerprint that patterns the molecular structure of the solvent. Also, the graph-based model is, for example, a quantification of the feature amount of the molecular structure using a graph convolutional neural network. Also, the index value representing solubility as the objective variable is a value calculated by the above-described learning data generation.
[0015] In solvent exploration, using the created prediction model, a predicted value representing the solubility of a predetermined polymer in the solvent to be predicted is obtained. Further, the solvent to be predicted is generated based on a predetermined algorithm. The predetermined algorithm generates the solvent to be predicted, for example, by recombining side chains to be bonded to the skeleton of a compound that is an initial structure. A plurality of solvents to be predicted are generated, and predicted values for each of them are calculated. By repeating such processing, a solvent with a good predicted value can be searched for. For example, a solvent whose predicted value exceeds a predetermined threshold may be extracted. Further, a preferable solvent may be searched for based on an evaluation value including a plurality of predicted values and scores other than the predicted values.
[0016] By repeating the above-described processing, the prediction accuracy of the prediction model can be improved. That is, for a solvent determined to have a good predicted value, a force field is created based on its structure, and an index value representing solubility is calculated by MD simulation. The calculated index value is further used as learning data. Note that the solvent extracted in solvent exploration can be evaluated based on the calculated index value.
[0017] FIG. 2 is a diagram showing an example of a system according to an embodiment. The system 100 includes a computer 1. The computer 1 further includes a processor 11, a storage device 12, and a user interface (UI) 13. The processor 11 is an arithmetic processing device such as a CPU (Central Processing Unit), and executes programs to perform each process according to the embodiment. Perform processing. The storage device 12 is at least one of a main storage device such as a RAM (Random Access Memory) or a ROM (Read Only Memory), and an auxiliary storage device (secondary storage device) such as an HDD (Hard-disk Drive), an SSD (Solid State Drive), or a flash memory. The main storage device temporarily stores the program read by the processor 11 or secures the working area of the processor 11. The auxiliary storage device stores the program executed by the processor 11 and other data. It is assumed that identification information for specifying a solvent and a predetermined polymer used for generating learning data is stored in the storage device 12 in advance. The UI 13 is an input / output device such as a touch panel, a keyboard, or a pointing device, for example. The UI 13 receives an input based on a user's operation and outputs information to the user.
[0018] <Learning data generation> FIG. 3 is a processing flowchart showing an example of the learning data generation process. The learning data generation process is started at an arbitrary timing based on a user's operation. Also, it is assumed that information defining a predetermined solvent and a predetermined polymer used for simulation is stored in the storage device 12 in advance. The solvent and the polymer can be stored in a database, text data, etc. as a character string expressed by, for example, the SMILES (Simplified Molecular Input Line Entry System) notation. For example, 1-ethyl-3-methylimidazolium is represented by the following character string in the SMILES notation. CCn1cc[n+](C)c1
[0019] When the learning data generation process is started, the processor 11 of the computer 1 creates force fields of the solvent and the polymer (FIG. 3: S1). In this step, the three-dimensional structure of the molecule is converted from the character string representing the solvent, which is described in the SMILES notation, by an existing method. The three-dimensional structure can be represented by data in, for example, the XYZ coordinate (orthogonal coordinate) format, the Z-matrix format, etc. It cuts. The 1-ethyl-3-methylimidazolium described above, when represented in XYZ coordinates, is as shown in, for example, Figure 4. Figure 4 is an example of data in PDB (Protein Data Bank) format containing the molecular structure represented in XYZ coordinates. The PDB format is a file composed of fixed-length records with 80 characters per line. Also, records starting with the identifier "ATOM" or "HETATM" contain the X coordinate, Y coordinate, and Z coordinate (in angstroms) of the atom. Figure 4 excerpts and illustrates the HETATM records describing the atomic coordinates among the PDB format data. From the 31st to 54th characters enclosed by the dashed line, the X coordinate, Y coordinate, and Z coordinate of the atom are included. It is a file composed of.
[0020] Also, for each of the solvent and the polymer, a force field can be created based on the three-dimensional structure. The potential function representing the force field E total is defined, for example, by the following equation (1). E total = E bounded + E nonbounded ···(1) Note that E bounded is the bonding term representing the chemical bonding force of the covalent bond, and is calculated by integrating the term based on the bond length within the molecule, the term based on the bond angle (deviation angle), and the term based on the bond rotation angle (dihedral angle). E nonbounded is the non-bonding term representing the electrostatic force and the intermolecular force, and is calculated by integrating the term based on the van der Waals force within and between molecules and the term based on the electrostatic interaction (electrostatic force). Note that the term based on the bond length within the molecule, the term based on the bond angle, the term based on the bond rotation angle, the term based on the van der Waals force, and the term based on the electrostatic interaction represent the intramolecular interaction. The term based on the van der Waals force and the term based on the electrostatic interaction represent the intermolecular interaction. Also, the creation of the force field can be performed using the functions available in existing software. For example, using GAFF (general AMBER force field) within Amber tools, which is MD simulation software, the chemical bonding force of the covalent bond and, among the non-bonding terms, the van der The Ruval's force can be calculated. Also, the electrostatic force can be calculated by existing quantum chemistry calculation software.
[0021] After S1, the processor 11 performs MD simulation on the solvent (Figure 3: S2). The MD simulation can also be executed using existing methods. In the MD simulation, the equations of motion for individual atoms are created using the above-mentioned force field, and by integrating the equations of motion, the time evolution of the molecular structure (i.e., the positions and velocities of each atom with respect to time change) and the like are simulated. In this step, the simulation is performed on a unit cell composed only of the solvent at normal pressure and at a temperature at which the solvent can diffuse sufficiently.
[0022] After S2, the processor 11 performs MD simulation on the solvent and the polymer (Figure 3: S3). In this step, for example, the polymer is placed in the solvent and the simulation is performed in the same manner as in S2.
[0023] After S3, the processor 11 obtains parameters related to solubility from the results of the MD simulation (Figure 3: S4). According to the process of S2, the particle arrangement in the equilibrium state is created, and in S4, the self-diffusion coefficient of the solvent is obtained by analyzing the MD simulation results. That is, by tracking the trajectory of the molecule, the self-diffusion coefficient can be determined from its mean square displacement. The mean square displacement can be calculated as the average value of the square of the distance between predetermined atoms. In particular, if the simulation is performed at a system setting temperature within a predetermined range or at a predetermined temperature, and the dissolution phenomenon is evaluated by the fluidity of the solvent (i.e., the self-diffusion coefficient), learning data for creating a prediction model suitable for the search for the above-mentioned ionic liquid can be generated.
[0024] Also, in S4, by analyzing the data obtained by the process of S3, as parameters related to solubility, for example, parameters representing intermolecular interactions are obtained. Examples of parameters representing intermolecular interactions between polymers include electrostatic forces and intermolecular forces. The electrostatic force and intermolecular force may be calculated based on the number or energy of hydrogen bonds. A hydrogen bond is considered to occur between an acceptor heavy atom A, a donor hydrogen atom H, and a donor heavy atom D. From the simulation results, the number of hydrogen bonds and the like can be determined. That is, when the distance from A to D is less than the distance threshold (for example, 3 angstroms) and the bond angle of A-H-D is within a predetermined range (for example, greater than 135 degrees and less than 180 degrees), it may be determined that a hydrogen bond is formed. ulation results, the number of hydrogen bonds and the like can be determined. That is, when the distance from A to D is less than the distance threshold (for example, 3 angstroms) and the bond angle of A-H-D is within a predetermined range (for example, greater than 135 degrees and less than 180 degrees), it may be determined that a hydrogen bond is formed.
[0025] Also, in S4, as a parameter related to solubility, the mean square displacement of the polymer in the solvent may be determined.
[0026] Also, in S4, as a parameter related to solubility, at least one of the peak value of the radial distribution function (the existence density at the distance r where the existence probability is the highest) between the polymer and the solvent and the free energy change amount with respect to the peak value may be determined. Generally, the radial distribution function g(r) is defined as a function of the distance r representing the existence probability of other particles existing around a certain particle (atom or molecule). Note that the distance may be based on a group of atoms rather than a single particle. This distance can be the distance between the center of gravity (calculated based on the mass and coordinates of each atom) or the center (charge center calculated based on the atomic charge of each atom) of the atomic group with strong interaction in the repeating unit of the polymer and the center of gravity or the center of the atomic group with strong interaction in the solvent molecule. Note that the atomic group with strong interaction can be appropriately selected based on interactions such as relatively high charge intensity and relatively strong hydrophobic interaction. Also, the LJ (Lennard-Jones) potential is relatively When it is strong, it is preferable to adopt the center of gravity of the related atom (in other words, the geometric center). Also, when the Coulomb force is relatively strong, it is preferable to adopt the charge center of the related atom.
[0027] Specifically, when taking a certain particle i as the center and setting the number of other particles existing in the spherical shell with an inner radius r and an outer radius r + dr (that is, a thickness dr) as n(r), the density of the particles in the spherical shell can be obtained by the following formula (2). n(r) / 4πr 2 dr ···(2) At this time, the radial distribution function g i (r) of a certain particle i can be obtained by the following formula (3). g i (r)=n(r) / 4πr 2 drρ ···(3) That is, the radial distribution function g i (r) of particle i is the density of the particles in the spherical shell divided by the average density ρ of the whole system. And the radial distribution function g(r) of the system can be obtained by the following formula (4). g(r)=<g i (r)>=<n(r)> / 4πr 2 drρ ···(4) That is, the radial distribution function g(r) of the system is the particle average of g i (r) and n(r).
[0028] This radial distribution function can be converted into an index of the free energy change amount F(r) using the following formula (5). F(r)=-kTlog(g(r)) ···(5) Note that k is the Boltzmann constant and T is the temperature. When the distance r is infinite, the radial distribution function g(r) approaches 1, so the free energy change amount F(r) approaches 0. That is, this free energy is the change amount based on infinity. If the minimum value of F(r) (in other words, the free energy change amount corresponding to the peak value of the radial distribution function) is used as a parameter related to solubility, it is strongly correlated with solubility.
[0029] For example, when the polymer is cellulose and the solvent is an ionic liquid having chloride ions as anions, when calculating the radial distribution function g(r), the distance r between the hydrogen atoms of the 2, 3, and 6 hydroxyl groups of cellulose and the chloride ions may be calculated based on the trajectory in the MD simulation. Further, if the minimum value of the free energy change amount F(r) obtained from the radial distribution function g(r) is defined as a parameter related to solubility, it strongly correlates with hydrogen bonding and solubility. 。
[0030] In addition, when the solvent is composed of two or more components, the interaction between solvents may be an important index. Therefore, in S4, as a parameter related to solubility, the peak value of the radial distribution function between solvents or the free energy change amount with respect to this may be obtained. In this case, the peak value of the radial distribution function g(r) and the minimum value of the free energy change amount F(r) expressed using the distance r between the centers of gravity or centers of atomic groups that seem to have strong interactions among the solvent molecules are also important indices in designing solvent molecules. When the solvent is an ionic liquid, the peak value of the radial distribution function g(r) and the minimum value of the free energy change amount F(r) between the charge center of the cation and the charge center of the anion may be used as parameters related to solubility.
[0031] In particular, when the solvent is an ionic liquid composed of an imidazole-based cation and chloride ions, the center of gravity between the two nitrogens in the imidazolium ring is used as the charge center of the cation, and for the distance r to the chloride ion, the radial distribution function g(r) may be calculated based on the trajectory of the MD simulation. Further, if the minimum value of F(r) calculated from the radial distribution function g(r) is defined as the free energy change amount particularly between the cation and the anion of the solvent molecule, although it does not strongly correlate with hydrogen bonding, the solubility increases within a certain range. For example, when the polymer is cellulose and the temperature is 400K, if the free energy change amount from the radial distribution function for the cation-anion pair is in the range of -0.5 kcal / mol to -1.0 kcal / mol, the solubility increases. Note that the temperature changes according to the definition formula, and the threshold values corresponding to the upper and lower limits of the above-mentioned range also change according to the temperature.
[0032] Further, when the polymer contains cellulose and the solvent is an ionic liquid, the radial distribution function g(r) of the distance r between the anion and cellulose may be calculated based on the trajectory of MD simulation. If the minimum value of F(r) calculated from g(r) is defined as the free energy change amount between the solvent molecules, especially between the anion and cellulose, it is correlated with hydrogen bonding to some extent, and the solubility increases only within a certain range.
[0033] Generally, parameters correlated with hydrogen bonding (or solubility) are useful as solubility indices. When the polymer is cellulose, there is a strong correlation between the intermolecular interaction energy of the polymer and hydrogen bonding (or solubility), and the intermolecular interaction energy of the polymer is useful as a solubility index. Also, there is a strong correlation between the diffusion coefficient and hydrogen bonding (or solubility), and the diffusion coefficient is useful as a solubility index. There is a strong correlation between the interaction energy between cellulose and the solvent and hydrogen bonding (or solubility), and the interaction energy between cellulose and the solvent is useful as a solubility index.
[0034] Note that the parameters related to solubility may be determined according to the type of the target polymer. When the user determines that the polymer has hydrogen bonding and is important in intermolecular interaction, hydrogen bonding can be used as one index. In general-purpose polymers, the mean square displacement can be used as an index. Also, the peak value of the radial distribution function between the polymer and the solvent, or the free energy change amount with respect to the peak value (in other words, the minimum value of the free energy change amount) may be used as an index. When the solvent is composed of two or more components, the peak value of the radial distribution function between the solvents (between two components of the solvent), or the free energy change amount with respect to the peak value may be used as an index. Note that in S4, a plurality of parameters may be obtained.
[0035] As described above, according to the learning data generation process, an index value representing the solubility of a predetermined polymer in a solvent can be calculated by simulation. Further, the learning data generation process is repeated to calculate index values representing the solubility of a single predetermined polymer in a plurality of solvents.
[0036] <Prediction model creation> FIG. 5 is a flowchart showing an example of the prediction model creation process. The prediction model creation process is started at an arbitrary timing based on a user operation, for example, after preparing learning data by the learning data generation process. Also, it is assumed that an index value representing the solubility of a predetermined polymer in a solvent is stored in the storage device 12 in advance.
[0037] When the prediction model creation process is started, the processor 11 of the computer 1 performs preprocessing (FIG. 5: S11). In this step, the data to be used as the explanatory variable and the objective variable (teacher data) is converted into a format suitable for machine learning. In the present embodiment, machine learning is performed with the feature amount of the solvent as the explanatory variable and the index representing the solubility of a predetermined polymer in the solvent as the objective variable. Note that in the present embodiment, since the process is performed with a single fixed polymer as the solute, the learning data does not include the feature amount of the predetermined polymer. However, a model that can predict the solubility for any polymer may be created by including the feature amount of the polymer as the solute in the explanatory variable.
[0038] The characteristic quantity of the solvent is, for example, a characteristic quantity based on the chemical structure of the solvent. Specifically, the characteristic quantity may be described using a model based on molecular descriptors such as a molecular fingerprint. A molecular fingerprint represents the presence or absence of specific molecular bonds, functional groups, ring structures, etc. as a binary vector or the number of occurrences as a count vector. For example, Morgan fingerprints obtained by existing software such as RDKit may be used, or descriptors based on physical characteristics such as Mordred may be used. Further, the characteristic quantity representing the chemical structure of the solvent may be a graph-based model such as data obtained by quantifying the characteristic quantity of the molecular structure using a molecular graph convolutional neural network. The explanatory variables are appropriately standardized or normalized.
[0039] In addition, the index value representing solubility, which is the target variable, includes at least one of the number or binding energy of hydrogen bonds formed within a predetermined polymer or between predetermined polymers, the self-diffusion coefficient of the solvent to be learned, the mean square displacement of a predetermined polymer in the solvent to be learned, the peak value of the radial distribution function between the solvent and the polymer, the free energy change amount with respect to the peak value, the peak value of the radial distribution function between two components of the solvent, and the free energy change amount with respect to the peak value. The target variable is also appropriately standardized or normalized.
[0040] After S11, the processor 11 performs machine learning to create a prediction model (Fig. 5: S12). The prediction model can adopt a deep learning model, a linear regression model, a decision tree regression model, etc., and is created by learning the relationship between the characteristic quantity based on the chemical structure of the solvent to be learned and the index value representing the solubility of a predetermined polymer in the solvent. When performing machine learning on a plurality of index values, for example, a prediction model may be created for each of the index values. Also, during machine learning, hyperparameters are appropriately determined for each model. In this step, for example, the weights between nodes in the deep learning model are adjusted.
[0041] According to the prediction model creation process, a prediction model for calculating and outputting a predicted value of the solubility of a predetermined polymer in a solvent can be created by inputting the feature amount of the solvent to be predicted.
[0042] <Solvent search> FIG. 6 is a flowchart showing an example of the solvent search process. The solvent search process is started at an arbitrary timing based on a user operation, for example, after a prediction model is prepared by the prediction model creation process. Further, it is assumed that the prediction model created in the prediction model creation process is stored in the storage device 12 in advance.
[0043] When the solvent search process is started, the processor 11 of the computer 1 sets an initial chemical structure to be used for the search (FIG. 6: S21). In this step, the basic structure of the solvent is input via the UI13 based on a user operation, for example, and held in the storage device 12. When the solvent is a salt, for example, the following chemical structure is input as the basic structure. In the present disclosure, among the structures of the cations in the basic structure when the solvent is a salt, the structure of the part where the conditions are not changed is also referred to as the cation basic skeleton.
Chemical formula
[0044] After S21, the processor 11 sets an evaluation function to be used for the search (FIG. 6: S22). The evaluation value V is defined by, for example, the following formula (6). V = Σ(α i M i ) + βSA ···(6) M is the score of the predicted value by one or more prediction models. That is, the index value representing one or more solubilities predicted by the prediction model is substituted. SA is the SA (Synthetic Accessibility) score. The SA score is a score representing the synthetic possibility with the complexity of the compound as an index and can be calculated by existing methods. α and β are weighting coefficients set according to the degree of importance attached to each score. Thus, in addition to the predicted value representing the solubility of a predetermined polymer, the score representing the synthetic possibility of the solvent may be used to comprehensively evaluate the quality of the solvent. In S22, for example, based on the operation of the user, an evaluation function including weighting coefficients is set. Note that the evaluation value V may be calculated without using the SA score. Also, when there is only one score used in the evaluation function, setting of the weighting coefficient is unnecessary. It is a core and can be calculated by existing methods. α and β are weighting coefficients set according to the degree of importance attached to each score. Thus, in addition to the predicted value representing the solubility of a predetermined polymer, the score representing the synthetic possibility of the solvent may be used to comprehensively evaluate the quality of the solvent. In S22, for example, based on the operation of the user, an evaluation function including weighting coefficients is set. Note that the evaluation value V may be calculated without using the SA score. Also, when there is only one score used in the evaluation function, setting of the weighting coefficient is unnecessary.
[0045] After S22, the processor 11 calculates the evaluation value V for the solvent to be predicted (Fig. 6: S23). In this step, first, R and the anion species are combined with respect to the cation basic skeleton set in S21, and the solvent to be predicted is determined. For example, as the solvent to be predicted, R included in the above-exemplified basic structure may be an organic group composed of a carbon atom and at least a part of a halogen atom such as a hydrogen atom, an oxygen atom, a nitrogen atom, a phosphorus atom, a sulfur atom, and a fluorine atom, and some hydrogen atoms bonded to the carbon atom of the organic group may be substituted by other substituents. R may be linear or may include a branched structure. Also, R may be combined with each other to create a solvent containing a plurality of ring structures. There is no particular limitation on the substituent that R may have, and examples thereof include a group composed of at least a part of a hydrogen atom, a carbon atom, an oxygen atom, a nitrogen atom, a phosphorus atom, a sulfur atom, and a halogen atom. When R is a monovalent group, it may be a functional group such as a hydroxy group, a carboxyl group, an amino group, a phosphoric acid group, a thiol group, or a sulfone group instead of an organic group. Each R and anion species may create various combinations by brute force, or for example, use a genetic algorithm to create a combination expected to have a high evaluation value.
[0046] FIG. 7 is a diagram for explaining a general ionic liquid. An ionic liquid is a room-temperature molten salt composed of an anion and a cation. As shown in FIG. 7, the basic structures with the cation changed include imidazolium salts, pyrrolidinium salts, ammonium salts, pyridinium salts, piperidinium salts, phosphonium salts, etc. Also, the anion species include Cl - , Br - , I - , etc. For the solvent to be predicted created in S23, a basic skeleton based on a known cation is set in S21 so that candidates for ionic liquids can be created by combining a structure (hereinafter also referred to as a "bonding structure") that can bind to the cation basic skeleton. Also, variations of the bonding structure and anion species may be stored in advance and used in S23. Note that the bonding structure in the basic structures exemplified above is R.
[0047] Also, in S23, a predicted value is calculated for the solvent to be predicted using a prediction model. That is, for the solvent to be predicted, the same pretreatment as S11 in FIG. 5 is performed and input into the prediction model created in S12. Also, as a calculation result by the prediction model, a predicted value of an index value representing solubility is calculated and output. Note that when prediction models are created for a plurality of index values respectively, a plurality of predicted values are output.
[0048] Then, in S23, an evaluation value of the solvent to be predicted is calculated using the evaluation function described above. Note that solvents whose evaluation value V exceeds a preset threshold value may be specifically extracted and output.
[0049] After S23, the processor 11 determines whether to end the process (Fig. 6: S24). In this step, it is determined whether to create solvents for further different prediction targets by changing the bonding structure and anion species for the above-described cation basic skeleton. For example, when all combinations of the pre-determined variations of the bonding structure and anion species have been created, it may be determined to end the process. Also, when creating a solvent by a genetic algorithm, it may be determined to end the process when a pre-determined end condition is satisfied. If it is determined not to end (S24: NO), the process returns to S23, and the processor 11 repeats the process for solvents for different prediction targets. On the other hand, if it is determined to end (S24: YES), the solvent search process ends.
[0050] According to the solvent search process, it is possible to search for candidates for solvents that are likely to dissolve a predetermined polymer at a lower computational cost than performing MD simulation.
[0051] <Repeating process> The solvents with high evaluation values found in the solvent search process may be used to verify the performance of the solvents by MD simulation and as inputs for the learning data generation process. That is, for solvents with high evaluation values by prediction, MD simulation may be performed based on their three-dimensional structures, and index values representing solubility may be calculated. The performance of the solvent can be verified by the calculated index values, and the calculated index values are used to reconstruct the prediction model as new learning data. Also, especially when searching for a solvent that is an ionic liquid, the self-diffusion coefficient at a predetermined temperature may be obtained by MD simulation, and it may be confirmed that it is not crystallized.
[0052] As described above, by repeating the learning data generation process, the prediction model creation process, and the solvent search process, the accuracy of the prediction model is gradually improved, and it becomes possible to efficiently search for solvents with high performance for dissolving a predetermined polymer.
[0053] <Modification example> Each configuration in each embodiment, combinations thereof, etc. are merely examples, and within the scope not departing from the gist of the present disclosure, additions, omissions, substitutions, and other changes to the configuration can be made as appropriate. The present disclosure is not limited by the embodiments, but is limited only by the scope of the claims. Also, each aspect disclosed in this specification can be combined with any other features disclosed in this specification.
[0054] The system described above may execute at least a part of learning data generation, prediction model creation, and solvent search. For example, according to the prediction model creation process, by inputting the feature amount of the solvent to be predicted, a prediction model for calculating and outputting the predicted value of the solubility of a predetermined polymer in the solvent can be created. Also, according to the solvent search process, candidates for solvents that are likely to dissolve a predetermined polymer can be searched with a lower computational cost than performing MD simulations.
[0055] At least a part of the functions of the computer 1 may be realized by being distributed among a plurality of devices (in other words, a plurality of processors), or a plurality of devices (in other words, a plurality of processors) may provide the same function in parallel. Also, one computer 1 may have a configuration including a plurality of processors. That is, a system including one or more computers (or one or more processors) for executing the above-described processes may be provided. Also, the computer 1 may be one or more servers connected to a terminal via a communication network. The server receives requests from a terminal operated by a user and performs the processes according to the above-described embodiments. The user's terminal may be a computer such as a PC (Personal Computer), tablet or the like. The terminal also includes a processor, a storage device, and an input / output interface for inputting and outputting information to and from the user. Also, the server and the terminal include a communication interface that is a wired or wireless communication module, and are communicably connected to each other via a network. The communication network is, for example, IP (Internet Protocol) It includes a network, and devices connected to the network can communicate based on a predetermined communication protocol. Part of the network may be a telephone network (fixed telephone network or mobile communication network), an ad hoc network, an intranet, a VPN (Virtual Private Network), a LAN (Local Area Network), a wireless LAN, a W AN (Wide Area Network), or the Internet.
[0056] The present disclosure also includes a method for executing the above-described processing, a computer program, and a computer-readable recording medium on which the program is recorded. By causing a computer to execute the program recorded on the recording medium, the above-described processing becomes possible.
[0057] Here, a computer-readable recording medium refers to a recording medium that accumulates information such as data and programs by an electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer. Examples of removable recording media from a computer include flexible disks, magneto-optical disks, optical disks, magnetic tapes, memory cards, etc. Examples of recording media fixed to a computer include HDDs, SSDs, ROMs, etc.
Explanation of Reference Numerals
[0058] 100: System, 1: Computer, 11: Processor, 12: Storage Device, 13: User Interface
Claims
1. A learning data generation process for calculating an index value representing the solubility of a predetermined polymer in a solvent to be learned by molecular dynamics simulation, and A model creation process for learning, by machine learning, the relationship between a feature amount based on the chemical structure of the solvent to be learned and an index value representing the solubility of the predetermined polymer, and creating a prediction model that outputs a predicted value representing the solubility of the predetermined polymer in a solvent to be predicted; and A search process for determining the solvent to be predicted based on a predetermined algorithm, and repeatedly outputting the predicted value using the feature amount based on the chemical structure of the solvent to be predicted and the prediction model; and A solvent search system for a polymer, including one or more computers that execute the above.
2. For the solvent to be predicted for which it is determined that the predicted value satisfies a predetermined criterion, perform the molecular dynamics simulation in the learning data generation process, and repeat the learning data generation process, the model creation process, and the search process The solvent search system according to claim 1.
3. The molecular dynamics simulation is performed using a force field based on the three-dimensional structures of the solvent to be learned and the predetermined polymer, and Information representing the three-dimensional structure of the solvent to be predicted for which it is determined in the search process that the predicted value satisfies the predetermined criterion is input to the learning data generation process The solvent search system according to claim 2.
4. The solvent to be predicted includes an ionic liquid The solvent search system according to any one of claims 1 to 3.
5. The predetermined polymer is a crystalline polymer The solvent search system according to any one of claims 1 to 3.
6. The predetermined polymer is cellulose and / or its derivative The solvent search system according to any one of claims 1 to 3.
7. The index value includes at least any one of the number or binding energy of hydrogen bonds formed within or between the predetermined polymers, the self-diffusion coefficient of the solvent to be learned, the mean square displacement of the predetermined polymer in the solvent to be learned, the peak value of the first radial distribution function between the solvent to be learned and the predetermined polymer, the free energy change amount with respect to the peak value of the first radial distribution function, the peak value of the second radial distribution function between two components when the solvent to be learned contains a plurality of components, and the free energy change amount with respect to the peak value of the second radial distribution function The solvent search system according to any one of claims 1 to 3.
8. The solvent to be learned is an ionic liquid, and the index value includes the peak value of the radial distribution function between the cation and the anion, or the free energy change amount with respect to the peak value. The solvent search system according to claim 7.
9. The feature quantity based on the chemical structure of the solvent to be learned and the feature quantity based on the chemical structure of the solvent to be predicted are represented by a molecular descriptor-based model or a graph-based model. The solvent search system according to any one of claims 1 to 3.
10. By molecular dynamics simulation, calculating an index value representing the solubility of a predetermined polymer in a solvent to be learned; learning the relationship between the feature quantity based on the chemical structure of the solvent to be learned and the index value representing the solubility of the predetermined polymer by machine learning, and creating a prediction model that outputs a predicted value representing the solubility of the predetermined polymer in the solvent to be predicted; A model creation system including one or more computers that execute the above.
11. Obtaining information representing the chemical structure of the solvent to be predicted; Using the relationship between the feature quantity based on the chemical structure of the solvent to be learned and the index value representing the solubility of a predetermined polymer in the solvent to be learned calculated by molecular dynamics simulation, which is learned by machine learning to create a prediction model, and the feature quantity based on the chemical structure of the solvent to be predicted, outputting a predicted value representing the solubility of the predetermined polymer in the solvent to be predicted. A solubility prediction system including one or more computers that execute the above.
12. By molecular dynamics simulation, calculating an index value representing the solubility of a predetermined polymer in a solvent to be learned; learning the relationship between the feature quantity based on the chemical structure of the solvent to be learned and the index value representing the solubility of the predetermined polymer by machine learning, and creating a prediction model that outputs a predicted value representing the solubility of the predetermined polymer in the solvent to be predicted; repeating determining the solvent to be predicted based on a predetermined algorithm, and outputting a predicted value using the feature quantity based on the chemical structure of the solvent to be predicted and the prediction model. A solvent search method executed by one or more computers.
13. A program for causing one or more computers to execute the method according to claim 12.
Citation Information
Cited By
Method for screening and extracting DES of tea saponin based on molecular simulation and machine learning
CN121768522A
Method for screening deep eutectic solvent for improving cellulose crystallinity based on electrostatic potential descriptor and application thereof
CN122117156A
Method for screening deep eutectic solvents for increasing cellulose crystallinity based on electrostatic potential descriptors and applications thereof
CN122117156B