Catalyst design system, catalyst design program, and catalyst design method

WO2025169864A1PCT designated stage Publication Date: 2025-08-14RIKEN CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/003283
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2025-01-31
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

The prior art is difficult to design catalysts with improved catalytic activity and selectivity through data-driven methods, and lacks effective software and database support.

Method used

Using machine learning technology, catalysts are designed through data-driven methods, databases are used to store high-quality data, and catalyst optimization is carried out in combination with machine learning, including information acquisition, extraction, prediction and identification of key factors in catalytic reactions.

Benefits of technology

The efficient design of the catalyst is achieved, the catalytic activity and selectivity are improved, and the prediction ability of catalytic reactions is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025003283_14082025_PF_FP_ABST
    Figure JP2025003283_14082025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention facilitates catalyst designing. A catalyst design system according to one embodiment of the present invention is a system for designing a catalyst that promotes a specific chemical reaction, and comprises: an acquisition unit that acquires information for identifying a chemical reaction that is specified by a user and that is to be promoted by the catalyst; an extraction unit that extracts, from a database, information pertaining to a molecular structure included in the chemical reaction identified on the basis of the acquired information, and a state quantity in the process of the chemical reaction; a learning unit that performs machine learning using the information pertaining to the molecular structure and the state quantity that have been extracted from the database, to generate a prediction model; a prediction unit that uses the prediction model to predict a state quantity from the information pertaining to the molecular structure specified by the user; and an identification unit that identifies a dominant factor of the predicted state quantity and displays the dominant factor.
Need to check novelty before this filing date? Find Prior Art

Description

Catalyst design system, catalyst design program, catalyst design method

[0001] The present invention relates to a catalyst design system, a catalyst design program, and a catalyst design method.

[0002] Traditionally, catalyst design involves preparing a catalyst, causing a chemical reaction, and evaluating the activity and selectivity of the catalyst. Recently, research into catalyst design has been progressing using machine learning techniques and data collected through experiments.

[0003] 利用对映体决定步骤中的中间体进行分子场分析可提取信息用于不对称催化中数据驱动的分子设计。《日本化学会志》2019年,92卷,1701页。铱 / 硼杂化催化用于立体发散不对称合成的数据驱动催化剂优化。《细胞报告:物理科学》2021年,2卷,100679页。在不对称N-杂环卡宾-铜催化中利用计算筛选数据进行分子场分析以实现数据驱动的计算机辅助催化剂优化。《日本化学会志》2022年,95卷,271页。T. 根施、M. S. 西格曼、A. 阿斯普鲁-古济克等人,“用于催化的有机磷配体的综合发现平台”,《美国化学会志》2022年,144卷,1205页。2022年1月12日。C. W. 科里等人,“开放反应数据库”,《美国化学会志》2021年,143卷,18820页。“融合计算化学与信息化学的合成路线开发”,《计算机辅助化学杂志》,2004年,5卷,26 - 34页

[0004] Regarding molecular catalysts, there is a need for a data-driven catalyst design method that can design catalysts with improved catalytic activity, software for using this methodology, and a catalyst design system for utilizing data that combines this methodology with a database for storing high-quality data created by this method for machine learning.

[0005] Therefore, an object of the present invention is to facilitate the design of catalysts.

[0006] A catalyst design system according to one embodiment of the present invention is a design system for a catalyst that promotes a specific chemical reaction, and includes: an acquisition unit that acquires information for identifying a chemical reaction that is promoted by the catalyst, as specified by a user; an extraction unit that extracts, from a database, information about the molecular structure included in the chemical reaction identified based on the acquired information, and state quantities in the process of the chemical reaction; a learning unit that performs machine learning using the information about the molecular structure and the state quantities extracted from the database to generate a predictive model; a prediction unit that uses the predictive model to predict the state quantity from the information about the molecular structure specified by the user; and an identification unit that identifies a governing factor of the predicted state quantity and displays the governing factor.

[0007] According to the present invention, the design of a catalyst can be facilitated.

[0008] 1 is a diagram showing the overall configuration according to one embodiment of the present invention. FIG. 2 is a sequence diagram showing the overall processing according to one embodiment of the present invention. FIG. 3 is a sequence diagram showing an example of database data expansion processing according to one embodiment of the present invention. FIG. 4 is a diagram showing the functional configuration of a catalyst design system according to one embodiment of the present invention. FIG. 5 is an example of a database <before registration> stored in a database storage unit according to one embodiment of the present invention. FIG. 6 is an example of a database <after registration> stored in a database storage unit according to one embodiment of the present invention. FIG. 7 is an example of a screen (generation of a prediction model) displayed on a user terminal according to one embodiment of the present invention. FIG. 8 is an example of a screen (catalyst design) displayed on a user terminal according to one embodiment of the present invention. FIG. 9 is an example of a screen (catalyst design) displayed on a user terminal according to one embodiment of the present invention. FIG. 10 is a diagram for explaining database combination according to one embodiment of the present invention. FIG. 11 is a diagram showing the hardware configuration of a catalyst design device (server) and a user terminal according to one embodiment of the present invention.

[0009] Hereinafter, each embodiment will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configuration are designated by the same reference numerals, and redundant description will be omitted.

[0010] <Explanation of Terms> A "catalyst" may be any catalyst. For example, a "catalyst" is a homogeneous catalyst. For example, a "catalyst" is a catalyst defined by a single molecular structure. For example, a "catalyst" is a catalyst that controls reaction selectivity, such as stereoselectivity, regioselectivity, or chemoselectivity; a catalyst that controls oxidation or reduction; a catalyst that controls bond formation or cleavage; or a catalyst that has multiple of these functions. For example, a "catalyst" is a metal complex catalyst, an organocatalyst, a supramolecular catalyst, an electrochemical catalyst, a radical catalyst, a peptide catalyst, a redox catalyst, a photoredox catalyst, a polymerization catalyst, an enzyme, an asymmetric catalyst, a cross-coupling catalyst, a metathesis catalyst, an acid-base catalyst, etc. The "information for identifying a chemical reaction promoted by a catalyst" may be any information that can identify a chemical reaction promoted by a catalyst (hereinafter simply referred to as a "catalyzed reaction"). For example, the "information for identifying a chemical reaction promoted by a catalyst" includes at least one of reaction conditions, a reaction name, and a catalyst name. The reaction conditions may include a reaction formula, a molecular structure involved in the reaction, a partial structure of the molecule involved in the reaction, a reaction time, and a reaction temperature. The "information about the molecular structure" may be any information representing a molecular structure (e.g., a transition state structure or an intermediate structure) including a catalyst. For example, the "information about the molecular structure" is information representing the three-dimensional structure of a molecule. For example, the "information about the molecular structure" may be "three-dimensional coordinates of the atoms constituting the molecule (e.g., an xyz file)" and "first-principles calculation conditions for calculating the three-dimensional coordinates." The "information about the molecular structure" may also include descriptors converted from the molecular structure (descriptors are numerical values ​​representing molecular information such as the structure and properties of a molecule). The "information about the molecular structure" is not limited to information about a three-dimensional molecular structure, but may also be information about a two-dimensional molecular structure (structural formula). The "state quantities in the process of a chemical reaction promoted by a catalyst" include at least one of a state quantity of a reaction intermediate and a state quantity of a transition state. For example, "state quantities in the process of a chemical reaction promoted by a catalyst" are activation energy, energy differences between transition states (e.g., corresponding to enantioselectivity, regioselectivity), etc. - "Governing factors of state quantities in the process of a chemical reaction promoted by a catalyst" are factors that increase or decrease the state quantities.For example, a "governing factor of a state quantity in a chemical reaction process promoted by a catalyst" is a spatial position that indicates a molecular structure (transition state, reaction intermediate) including the catalyst that governs the increase or decrease of the state quantity. For example, a "governing factor of a state quantity in a chemical reaction process promoted by a catalyst" is any quantity (e.g., energy level of a molecular orbital, atomic charge, dipole moment) that can be calculated from the transition state or reaction intermediate by first-principles calculation.

[0011] <Overall Configuration> Fig. 1 is a diagram showing the overall configuration according to one embodiment of the present invention. The catalyst design system 1 includes a catalyst design device 10 and user terminals 20A, 20B, and 20C (hereinafter collectively referred to as user terminals 20. Note that the number of user terminals is not limited to three). The catalyst design system 1 and the user terminals 20 can send and receive data via any network.

[0012] The catalyst design system 1 may be realized by only a catalyst design device (e.g., a single personal computer) 10 (i.e., without using the user terminal 20). In this case, the catalyst design device 10 also includes a prediction model setting unit 201 and a catalyst design unit 202, which are described in Fig. 4. For example, a user 21 can search a database stored in the catalyst design device 10 and use the searched data in a catalyst design application within the catalyst design device 10 to design a catalyst.

[0013] <<Catalyst Design Device>> The catalyst design device 10 executes processing for designing a catalyst in response to instructions from the user terminal 20. For example, the catalyst design device 10 is one or more computers (for example, servers).

[0014] <<User Terminal>> The user terminal 20 is a terminal operated by users 21A, 21B, and 21C (hereinafter collectively referred to as users 21. Note that the number of users is not limited to three) who design catalysts. For example, the user terminal 20 is a personal computer, a tablet, a smartphone, or the like.

[0015] In addition, some or all of the processing described in this specification as processing performed by the catalyst design device 10 may be performed by the user terminal 20, and some or all of the processing described in this specification as processing performed by the user terminal 20 may be performed by the catalyst design device 10.

[0016] <Method> The overall process will be described below with reference to FIG. 2, and the data expansion process for the database will be described with reference to FIG.

[0017] FIG. 2 is a sequence diagram showing the overall processing according to one embodiment of the present invention.

[0018] First, we will explain the setting of the predictive model used in catalyst design (steps 101 to 107), then we will explain the design of the catalyst (steps 108 to 113), and then we will explain the data expansion of the database that serves as training data for the predictive model (steps 114 to 117).

[0019] [Setting the Prediction Model] In step 101 (S101), the user 21 inputs "information for identifying a chemical reaction promoted by a catalyst" into the user terminal 20 (e.g., on the screen of S1001 in FIG. 7). For example, the "information for identifying a chemical reaction promoted by a catalyst" includes at least one of the reaction conditions (which may include a reaction formula, a molecular structure involved in the reaction, a partial structure of a molecule involved in the reaction, a reaction time, and a reaction temperature), a reaction name, and a catalyst name.

[0020] In step 102 (S102), the user terminal 20 transmits to the catalyst design device 10 the "information for identifying the chemical reaction promoted by the catalyst" input in S101.

[0021] In step 103 (S103), the catalyst design apparatus 10 extracts from the database 100 "information on molecular structure" and "state quantities in the process of the chemical reaction promoted by the catalyst" for each catalytic reaction, which correspond to the "information for identifying the chemical reaction promoted by the catalyst" received in S102. Note that the database 100 stores, in association with each other, the "catalyst name (including a complex (complex, etc.) with other molecules such as a substrate)" for each catalytic reaction, the "state quantities in the process of the chemical reaction promoted by the catalyst (e.g., enantioselectivity and activation energy)," and "information on molecular structure (e.g., the three-dimensional coordinates (xyz file) of the atoms constituting the molecule and the conditions for first-principles calculation)."

[0022] Specifically, the catalyst design system 10 identifies a catalytic reaction from "information for identifying a chemical reaction promoted by a catalyst." Next, the catalyst design system 10 extracts, from the database 100, "information related to the molecular structure" and "state quantities in the process of the chemical reaction promoted by the catalyst" for each catalytic reaction.

[0023] In step 104 (S104), the catalyst design device 10 transmits to the user terminal 20 the "information on molecular structure" and the "state quantity in the process of the chemical reaction promoted by the catalyst" extracted in S103.

[0024] In step 105 (S105), the user 21 selects, on the user terminal 20 (for example, on the screens of S1002 and 1003 in Figure 7), data to be used as training data for the predictive model used in catalyst design (i.e., which catalytic reaction's "information on molecular structure" and "state quantities in the process of a chemical reaction promoted by a catalyst" will be used to generate the predictive model).

[0025] In step 106 (S106), the user terminal 20 transmits to the catalyst design device 10 the "information on molecular structure" and the "state quantity in the process of the chemical reaction promoted by the catalyst" selected in S105.

[0026] In step 107 (S107), the catalyst design device 10 performs machine learning using the "information related to molecular structure" and the "state quantity in the process of a chemical reaction promoted by a catalyst" received in S106 as training data to generate a prediction model. The prediction model is a trained model (specifically, a regression model (supervised learning), a classification model (supervised learning), dimension reduction, or a clustering model (unsupervised learning)) that has been trained by machine learning so that when "information related to molecular structure" is input, "state quantity in the process of a chemical reaction promoted by a catalyst" is output.

[0027] [Catalyst Design] In step 108 (S108), the user 21 inputs "information regarding molecular structure" into the user terminal 20 (e.g., on the screen of S2001 in Figure 8 or S3001 in Figure 9) (i.e., designs the molecular structure of the catalyst).

[0028] In step 109 (S109), the user terminal 20 transmits the "information relating to molecular structure" input in S108 to the catalyst design device 10 (that is, transmits the molecular structure of the catalyst designed by the user 21).

[0029] In step 110 (S110), the catalyst design device 10 uses the prediction model generated in S107 to predict the "state quantity in the process of a chemical reaction promoted by a catalyst" from the "information regarding molecular structure" received in S109.

[0030] In step 111 (S111), the catalyst design system 10 identifies the governing factors of the "state quantities in the process of a chemical reaction promoted by a catalyst" predicted in S110. The catalyst design system 10 also evaluates the similarity between the identified "governance factors" or the similarity between structures (descriptors).

[0031] Specifically, the catalyst design device 10 identifies the governing factor based on the importance of a descriptor x in a regression and classification model that receives "information about molecular structure" and outputs "a state quantity in the process of a chemical reaction promoted by a catalyst." For example, when "information about molecular structure (e.g., a descriptor of molecular structure)" is defined as x and "a state quantity in the process of a chemical reaction promoted by a catalyst" and a value calculated from that state quantity is defined as y, the governing factor is identified based on the value of βi (regression coefficient / model parameter) in y = f(x) = β1x1 + β2x2 + ... + βnxn. For example, in nonlinear regression and classification models, the governing factor is identified based on a descriptor importance evaluation method, such as importance in a decision tree or SHAP (Shapley Additive exPlanations) value.

[0032] For example, the similarity between governing factors or between structures (descriptors) is calculated and identified by combining visualization through clustering and dimensionality reduction and calculation of distances between descriptors.

[0033] In step 112 (S112), the catalyst design system 10 transmits the "state quantities in the process of the chemical reaction promoted by the catalyst" predicted in S110 and the "controlling factors" identified in S111 to the user terminal 20. In addition, the catalyst design system 10 transmits the similarity between the "controlling factors" identified in S111 or the similarity between structures (between descriptors) to the user terminal 20.

[0034] In step 113 (S113), the user terminal 20 displays the "state quantities in the process of the chemical reaction promoted by the catalyst" and "controlling factors" received in S112 (for example, the screen of S2002 in FIG. 8 and S3002 in FIG. 9).

[0035] [Database Data Expansion] In step 114 (S114), the user 21 inputs into the user terminal 20 (for example, on the screen of S2003 in FIG. 8 or S3003 in FIG. 9) "information on molecular structure" in which part of the "information on molecular structure" in S108 has been modified (i.e., the molecular structure of the catalyst in S108 is modified and redesigned).

[0036] In step 115 (S115), the user terminal 20 calculates, by first-principles calculation, the "state quantity in the process of the chemical reaction promoted by the catalyst" from the "information on the molecular structure" input in S114.

[0037] Although the present specification mainly describes the case where first-principles calculations are used (first-principles calculations incorporating empirical methods such as experimental parameters, or calculations using a quantum computer, may be used), calculations in the present invention may also be performed using a trained model obtained by machine learning instead of first-principles calculations. A structure close to a transition state may be obtained by optimization by fixing the bond length, etc., and calculations may also be performed using the structure close to the transition state.

[0038] In step 116 (S116), the user terminal 20 transmits to the catalyst design device 10 the "information regarding the molecular structure" input in S114, the "information regarding the (optimized) molecular structure" calculated in S115, and the "state quantities in the process of the chemical reaction promoted by the catalyst" (i.e., the user 21 transmits the molecular structure of the catalyst redesigned by the user 21 (including complexes with other molecules such as substrates) and the state quantities in the process of the chemical reaction promoted by the catalyst).

[0039] In step 117 (S117), the catalyst design device 10 adds the "information relating to molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" received in S116 to the database 100. Alternatively, a person other than the user adds the "information relating to molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" to the database 100 from publicly available information such as a paper that describes the use of the catalyst design device 10.

[0040] FIG. 3 is a sequence diagram showing an example of a database data expansion process according to an embodiment of the present invention.

[0041] In step 201 (S201), the user 21 inputs "information for identifying a chemical reaction promoted by a catalyst" into the user terminal 20 (for example, on the screen of S1001 in FIG. 7). For example, the "information for identifying a chemical reaction promoted by a catalyst" includes at least one of the reaction conditions (which may include a reaction formula, a molecular structure involved in the reaction, a partial structure of a molecule involved in the reaction, a reaction time, and a reaction temperature), a reaction name, and a catalyst name.

[0042] In step 202 (S202), the user terminal 20 transmits the "information for identifying the chemical reaction promoted by the catalyst" input in S201 to the catalyst design device 10.

[0043] In step 203 (S203), the catalyst design apparatus 10 extracts from the database 100 "information on molecular structure" and "state quantities in the process of the chemical reaction promoted by the catalyst" for each catalytic reaction, which correspond to the "information for identifying the chemical reaction promoted by the catalyst" received in S202. Note that the database 100 stores, in association with each other, the "catalyst name (including a complex (complex, etc.) with other molecules such as a substrate)" for each catalytic reaction, the "state quantities in the process of the chemical reaction promoted by the catalyst (e.g., enantioselectivity and activation energy)," and "information on molecular structure (e.g., the three-dimensional coordinates (xyz file) of the atoms constituting the molecule and the conditions for first-principles calculation)."

[0044] Specifically, the catalyst design system 10 identifies a catalytic reaction from "information for identifying a chemical reaction promoted by a catalyst." Next, the catalyst design system 10 extracts, from the database 100, "information related to the molecular structure" and "state quantities in the process of the chemical reaction promoted by the catalyst" for each catalytic reaction.

[0045] In step 204 (S204), the catalyst design device 10 transmits to the user terminal 20 the "information on molecular structure" and the "state quantity in the process of the chemical reaction promoted by the catalyst" extracted in S203.

[0046] In step 205 (S205), the user 21 inputs into the user terminal 20 "information regarding molecular structure" in which part of the "information regarding molecular structure" received in S204 has been corrected (i.e., corrects the "information regarding molecular structure" in the database 100 received in S204).

[0047] In step 206 (S206), the user terminal 20 calculates, by first-principles calculation, the "state quantity in the process of the chemical reaction promoted by the catalyst" from the "information on the molecular structure" input in S205.

[0048] Although the present specification mainly describes the case where first-principles calculations are used (first-principles calculations incorporating empirical methods such as experimental parameters, or calculations using a quantum computer, may be used), calculations in the present invention may also be performed using a trained model obtained by machine learning instead of first-principles calculations. A structure close to a transition state may be obtained by optimization by fixing the bond length, etc., and calculations may also be performed using the structure close to the transition state.

[0049] In step 207 (S207), the user terminal 20 transmits to the catalyst design device 10 the "information regarding the molecular structure" input in S205, the "information regarding the (optimized) molecular structure" calculated in S206, and the "state quantities in the process of the chemical reaction promoted by the catalyst" (i.e., the user 21 transmits the molecular structure of the catalyst (including complexes with other molecules such as substrates) whose data in the database 100 has been modified, and the state quantities in the process of the chemical reaction promoted by the catalyst).

[0050] In step 208 (S208), the catalyst design device 10 adds the "information on molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" received in S207 to the database 100. Alternatively, a person other than the user adds the "information on molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" to the database 100 from publicly available information such as a paper that describes the use of the catalyst design device 10.

[0051] <Functional Configuration> FIG. 4 is a diagram showing the functional configuration of the catalyst design system 1 according to one embodiment of the present invention.

[0052] <<Catalyst Design Apparatus>> The catalyst design apparatus 10 can include an acquisition unit 101, an extraction unit 102, a database storage unit 103, a learning unit 104, a prediction unit 105, an identification unit 106, and a registration unit 107. Furthermore, the catalyst design apparatus 10 can function as the acquisition unit 101, the extraction unit 102, the learning unit 104, the prediction unit 105, the identification unit 106, and the registration unit 107 by executing a program.

[0053] The acquisition unit 101 acquires, from the user terminal 20, instructions for generating a prediction model using data in the database 100 stored in the database storage unit 103. Specifically, the acquisition unit 101 acquires "information for identifying a chemical reaction promoted by a catalyst" designated by the user 21. The acquisition unit 101 also acquires information on data selected by the user 21 to be used as training data for the prediction model (i.e., information on which catalytic reaction's "information on molecular structure" and "state quantities in the process of a chemical reaction promoted by a catalyst" are to be used to generate the prediction model).

[0054] The acquisition unit 101 acquires, from the user terminal 20, an instruction for expanding the data in the database 100 stored in the database storage unit 103. Specifically, the acquisition unit 101 acquires "information on molecular structure" in the database 100 in which part of the "information on molecular structure" has been modified by the user 21, and "state quantities in the process of a chemical reaction promoted by a catalyst" calculated by the user terminal 20 using first-principles calculations.

[0055] The extraction unit 102 extracts information about the molecular structure included in a chemical reaction identified based on the "information for identifying a chemical reaction promoted by a catalyst" acquired by the acquisition unit 101, and state quantities in the process of the chemical reaction, from the database 100 stored in the database storage unit 103. Specifically, the extraction unit 102 identifies a catalytic reaction from the "information for identifying a chemical reaction promoted by a catalyst." Next, the extraction unit 102 extracts the "information about the molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" for each catalytic reaction from the database 100, and transmits them to the user terminal 20.

[0056] The database storage unit 103 stores a database 100 (which will be described in detail later with reference to FIGS. 5 and 6).

[0057] The learning unit 104 generates a prediction model by machine learning using the "information about molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" extracted from the database 100 as training data. Note that the "information about molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" selected by the user 21 from the "information about molecular structure" and the "state quantities in the process of a chemical reaction promoted by a catalyst" extracted from the database 100 may be used. The prediction model is a trained model (specifically, a regression model (supervised learning), a classification model (supervised learning), a dimensionality reduction model, or a clustering model (unsupervised learning)) that has been trained by machine learning so that when the "information about molecular structure" is input, the "state quantities in the process of a chemical reaction promoted by a catalyst" are output. The output of the prediction model may include the uncertainty of the predicted value (e.g., the variance used in Bayesian optimization).

[0058] The prediction unit 105 uses the prediction model generated by the learning unit 104 to predict the "state quantity in the process of a chemical reaction promoted by a catalyst" from the "information on the molecular structure" specified by the user 21. For example, the prediction unit 105 converts the molecular structure of the catalyst into a descriptor, inputs the descriptor into the prediction model, and outputs the "state quantity in the process of a chemical reaction promoted by a catalyst."

[0059] An example of a method in which the prediction unit 105 makes a prediction using the prediction model generated by the learning unit 104 will be described below.

[0060] [Regression] For example, the learning unit 104 generates a regression model (e.g., a model for predicting the activation energy value, a model for predicting the energy difference (reaction selectivity)). The prediction unit 105 can use the regression model to predict a value (e.g., the activation energy value) representing a "state quantity in the process of a chemical reaction promoted by a catalyst" from "information related to molecular structure" specified by the user 21.

[0061] [Classification] For example, the learning unit 104 generates a classification model (e.g., a model that predicts whether an activity is high or low, or a model that classifies whether the energy difference (reaction selectivity) is large or small). The prediction unit 105 can use the classification model to predict a class (e.g., two classes, high activity and low activity; it may be classified into three or more classes) that represents a "state quantity in the process of a chemical reaction promoted by a catalyst" from "information on molecular structure" specified by the user 21.

[0062] [Clustering] For example, the learning unit 104 generates a clustering model that clusters "information about molecular structure." The prediction unit 105 uses the clustering model to cluster the "information about molecular structure" specified by the user 21 and calculates the similarity between the "information about molecular structure" and a predetermined catalyst (e.g., a high-performance catalyst). Depending on the situation, the calculation of the similarity may involve visualization by dimensional reduction or calculation of the distance between descriptors, either alone or in combination. When the similarity between the "information about molecular structure" and a predetermined catalyst (e.g., a high-performance catalyst) is high (e.g., above a threshold), the prediction unit 105 calculates the "state quantity in the process of a chemical reaction promoted by a catalyst" of the "information about molecular structure" by first-principles calculation or the like.

[0063] The "information on molecular structure" used in the above clustering, visualization by dimensionality reduction, and calculation of distances between descriptors may be obtained by optimizing a structure close to the transition state by fixing bond lengths, etc., and using the structure close to the transition state.

[0064] The identification unit 106 identifies a dominant factor of the "state quantity in the process of a chemical reaction accelerated by a catalyst" predicted by the prediction unit 105 and displays the dominant factor on the user terminal 20. Specifically, the identification unit 106 identifies the dominant factor based on the importance of a descriptor x of a regression and classification model that outputs a "state quantity in the process of a chemical reaction accelerated by a catalyst" when "information about a molecular structure" is input. For example, when "information about a molecular structure (e.g., a descriptor of the molecular structure)" is x and "a state quantity in the process of a chemical reaction accelerated by a catalyst" and a value calculated from the state quantity is y, the dominant factor is identified based on the value of βi (regression coefficient / model parameter) in y = f(x) = β1x1 + β2x2 + ... + βnxn. For example, in a nonlinear regression and classification model, the dominant factor is identified based on a descriptor importance evaluation method, such as importance in a decision tree or a SHAP (Shapeley Additive exPlanations) value.

[0065] The registration unit 107 registers (adds) new data to the database 100 stored in the database storage unit 103. Specifically, the registration unit 107 stores information about a molecular structure designed by the user 21 based on the governing factors and the state quantities calculated by the user terminal 20 through first-principles calculation in the database 100. The registration unit 107 also stores information about a molecular structure in which part of the information about the molecular structure stored in the database 100 has been corrected, and the state quantities calculated by the user terminal 20 through first-principles calculation in the database 100.

[0066] <<User Terminal>> The user terminal 20 can include a prediction model setting unit 201 and a catalyst design unit 202. Furthermore, the user terminal 20 can function as the prediction model setting unit 201 and the catalyst design unit 202 by executing a program.

[0067] The prediction model setting unit 201 issues instructions for generating a prediction model used in catalyst design. Specifically, the prediction model setting unit 201 transmits "information for identifying a chemical reaction promoted by a catalyst" input by the user 21 to the user terminal 20 to the catalyst design device 10. The prediction model setting unit 201 also receives "information related to molecular structure" and "state quantities in the process of a chemical reaction promoted by a catalyst" corresponding to the "information for identifying a chemical reaction promoted by a catalyst" from the catalyst design device 10, and transmits the "information related to molecular structure" and "state quantities in the process of a chemical reaction promoted by a catalyst" selected by the user 21 to the catalyst design device 10.

[0068] The prediction model setting unit 201 issues an instruction to expand the data in the database 100. Specifically, the prediction model setting unit 201 transmits to the catalyst design device 10 "information on molecular structure" in which a part of the "information on molecular structure" in the database 100 has been modified by the user 21. Furthermore, the prediction model setting unit 201 calculates "state quantities in the process of a chemical reaction promoted by a catalyst" from the "information on molecular structure" in the database 100 in which a part of the "information on molecular structure" has been modified by the user 21, using first-principles calculations (e.g., DFT (Density Functional Theory) calculations), and transmits the calculated state quantities to the catalyst design device 10.

[0069] The catalyst design unit 202 instructs the design of a catalyst. It transmits "information on molecular structure" specified by the user 21 to the catalyst design device 10 and receives "state quantities in the process of a chemical reaction promoted by a catalyst" and "governing factors" predicted from the information on the molecular structure. The catalyst design unit 202 also transmits information on the molecular structure designed by the user 21 based on the governing factors to the catalyst design device 10. The catalyst design unit 202 also calculates the "state quantities in the process of a chemical reaction promoted by a catalyst" from the information on the molecular structure designed by the user 21 based on the governing factors using first-principles calculations (e.g., DFT (Density Functional Theory) calculations), and transmits the calculated values ​​to the catalyst design device 10.

[0070] <Database> The database 100 will be described below with reference to FIGS.

[0071] FIG. 5 shows an example of a database <before registration> stored in the database storage unit 103 according to one embodiment of the present invention.

[0072] The database storage unit 103 stores a database 100. The database 100 stores, for each catalytic reaction, the "catalyst name (including complexes with other molecules such as substrates)," the "state quantities (e.g., enantioselectivity and activation energy) in the process of the chemical reaction promoted by the catalyst," and "information on the molecular structure (e.g., three-dimensional coordinates (xyz file) of atoms constituting the molecule and conditions for first-principles calculations)," all linked together.

[0073] FIG. 6 is an example of a database <after registration> stored in the database storage unit 103 according to one embodiment of the present invention.

[0074] For example, information about a molecular structure designed by a user 21 based on governing factors, and a molecular structure (optimized molecular structure) and state quantities calculated by first-principles calculation from the information about the designed molecular structure are registered in the database 100.

[0075] In this way, in one embodiment of the present invention, various users 21 design the molecular structure of a catalyst based on the governing factors, and data on the molecular structure and state quantities of the catalyst designed based on the governing factors (i.e., the catalyst reflecting the governing factors) is added to the database. This makes it possible to improve the accuracy of prediction models generated using the data in the database.

[0076] For example, information on a molecular structure in which part of the information on the molecular structure stored in database 100 has been modified (for example, by introducing a substituent), and a molecular structure (optimized molecular structure) and state quantities calculated by first-principles calculation from the information on the modified molecular structure are registered in database 100.

[0077] In this way, in one embodiment of the present invention, various users 21 modify part of the molecular structure of the catalyst in the database 100, and the data on the molecular structure and state quantities of the modified catalyst are added to the database, thereby improving the accuracy of the prediction model generated using the data in the database.

[0078] <Screen> Screens displayed on the user terminal 20 will be described below with reference to FIGS.

[0079] FIG. 7 is an example of a screen (generation of a prediction model) displayed on the user terminal 20 according to an embodiment of the present invention.

[0080] In step 1001 (S1001), a screen is displayed for the user 21 to input "information for identifying a chemical reaction promoted by a catalyst." The user 21 inputs "information for identifying a chemical reaction promoted by a catalyst (e.g., reaction conditions, reaction name, catalyst name, etc.)" on the screen in S1001.

[0081] In step 1002 (S1002), the "state quantities in the process of the catalyst-accelerated chemical reaction (e.g., enantioselectivity (energy difference between transition states) and activation energy)" and "information on molecular structure" of each catalytic reaction corresponding to the "information for identifying the catalyst-accelerated chemical reaction" input in S1001 are extracted from database 100 and displayed.

[0082] In step 1003 (S1003), the user 21 selects data to be used as training data for the prediction model.

[0083] Thus, in one embodiment of the present invention, user 21 can select training data for the predictive model to be used in catalyst design, thereby generating a predictive model suitable for the catalyst that the user wishes to design.

[0084] FIG. 8 is an example of a screen (catalyst design) displayed on the user terminal 20 according to one embodiment of the present invention.

[0085] In step 2001 (S2001), a molecular structure (for example, a transition state structure, an intermediate structure) including a catalyst that has been designed by the user 21 (for example, by modifying some molecular structure) is displayed.

[0086] In step 2002 (S2002), the governing factors of the state quantities in the process of the chemical reaction promoted by the catalyst designed in S2001 are displayed. Specifically, using a prediction model, the state quantities are predicted from information on the molecular structure including the catalyst designed in S2001, and the governing factors of the state quantities are identified.

[0087] In step 2003 (S2003), the user 21 adds a partial structure of the molecule to the spatial position of the controlling factor displayed in S2002 (in the example of FIG. 7, the controlling factor is assumed to be a factor that increases enantioselectivity or a factor that decreases activation energy).

[0088] Thereafter, in S2003, information on the molecular structure to which the partial structure of the molecule has been added and the state quantities calculated from the information on the molecular structure by first-principles calculation are added to the database 100.

[0089] In this way, in one embodiment of the present invention, the governing factors of the state quantities of a catalyst designed by the user 21 are visualized, so that the user 21 can redesign the catalyst by adding molecular structures based on the visualized governing factors.

[0090] FIG. 9 is an example of a screen (catalyst design) displayed on the user terminal 20 according to one embodiment of the present invention.

[0091] In step 3001 (S3001), a molecular structure (for example, a transition state structure, an intermediate structure) including a catalyst that has been designed by the user 21 (for example, by modifying some molecular structure) is displayed.

[0092] In step 3002 (S3002), the governing factors of the state quantities in the process of the chemical reaction promoted by the catalyst designed in S3001 are displayed. Specifically, using a prediction model, the state quantities are predicted from information on the molecular structure including the catalyst designed in S3001, and the governing factors of the state quantities are identified.

[0093] In step 3003 (S3003), the user 21 deletes the partial structure of the molecule that is located in the spatial position of the controlling factor displayed in S3002 (in the example of FIG. 8, the controlling factor is assumed to be a factor that reduces enantioselectivity or a factor that increases activation energy).

[0094] Thereafter, in S3003, information on the molecular structure from which the partial structure of the molecule has been deleted and the state quantities calculated from the information on the molecular structure by first-principles calculation are added to the database 100.

[0095] In this way, in one embodiment of the present invention, the governing factors of the state quantities of a catalyst designed by the user 21 are visualized, so that the user 21 can delete molecular structures and redesign the catalyst based on the visualized governing factors.

[0096] <Database Combination> In one embodiment of the present invention, the catalyst design system 10 can combine and manage multiple catalyst-related databases. For example, the catalyst design system 10 may combine multiple databases in table format, or may combine multiple databases in graph format such as RDF (Resource Description Framework).

[0097] For example, the catalyst design system 10 manages the database in an RDF graph format as shown in Fig. 10. Specifically, the catalyst design system 10 manages the relationships between three elements (subject, predicate, and post-object).

[0098] In the database for catalyst A in Figure 10, the subject is transition state a, the predicate is activation energy, and the object is the numerical value of the activation energy. The subject is transition state a, the predicate is catalyst type, and the object is catalyst A. The subject is transition state a, the predicate is coordinate information, and the object is TSa.xyz (the three-dimensional coordinate file of transition state structure a). The subject is transition state a, the predicate is reaction type, and the object is reaction formula 1.

[0099] In the database for catalyst B in Figure 10, the subject is transition state b, the predicate is activation energy, and the object is high activity or low activity. The subject is transition state b, the predicate is catalyst type, and the object is catalyst B. The subject is transition state b, the predicate is coordinate information, and the object is TSb.xyz (the three-dimensional coordinate file for transition state structure b). The subject is transition state b, the predicate is reaction type, and the object is reaction formula 1.

[0100] When reaction formula 1 in the database of catalyst A in Fig. 10 is the same as reaction formula 1 in the database of catalyst B in Fig. 10, the database of catalyst A and the database of catalyst B are linked using reaction formula 1 as a linking key. Note that the linking key is not limited to a reaction formula.

[0101] The multiple databases related to catalysts may include databases related to catalysts other than those designed by the catalyst design system 1 (e.g., a database of basic molecular properties, a database of commercially available compounds, a database of material properties, a database of descriptors, and a database of synthetic routes).

[0102] <Database Search> In one embodiment of the present invention, the database can be searched for catalytic reactions and their transition states, intermediates, and state quantities using the following method. This allows for easy searching of similar reactions (e.g., reactions using the same substrate but a different catalyst). Reaction name, including the name of the entire reaction and the names of the elementary reactions (e.g., the name of the entire reaction: Michael addition, 1,2 addition, etc.; the name of the elementary reaction: oxidative addition, reductive elimination, etc.). Catalyst: Presence or absence of metal, type of metal, type of ligand, catalyst name, ligand name, names of substituents contained in the structure, molecular formula / structural formula, various molecular representations such as SMARTS / SMILES / InChl. Substrate / Reactant: Type of substrate / reactant, substrate name, reactant name, names of substituents contained in the structure, molecular formula / structural formula, various molecular representations such as SMARTS / SMILES / InChl. Other elements included in the reaction, such as additives and solvents, can also be searched by their type, name, molecular formula / structural formula, and SMARTS / SMILES / InChl, as needed. - In addition, when searching using reaction names, etc., it may be possible to select and search for multiple candidates (for example, searching for Ru and Ir as metals at the same time). The same applies to the presence or absence of a metal, the type of metal, the type of ligand, and the type of substrate / reactant.

[0103] <Utilization of Database> In one embodiment of the present invention, a prediction model may be generated by machine learning using, as training data, a character string (e.g., SMILES) representing the structure of molecules used in a reaction and a molecular structure (e.g., a transition state of an elementary reaction). The prediction model is a trained model that has been machine-learned so that, when a "character string representing a molecular structure" is input, a three-dimensional "molecular structure" (a transition state or a reaction intermediate) is output. The prediction model is used to predict a "molecular structure" from a "character string representing a molecular structure" (i.e., a "character string representing a molecular structure" is input to the prediction model, and a "molecular structure" is output). Furthermore, a prediction model may be generated by machine learning using, as training data, a molecular structure and a state quantity in a chemical reaction process promoted by a catalyst. The prediction model is a trained model that has been machine-learned so that, when a "molecular structure" is input, a "state quantity in a chemical reaction process promoted by a catalyst" is output. The prediction model is used to predict the "state quantity in the process of a chemical reaction promoted by a catalyst" from the "molecular structure" (i.e., the "molecular structure" is input into the prediction model, and the "state quantity in the process of a chemical reaction promoted by a catalyst" is output). In this way, the "molecular structure" predicted from a character string representing the molecular structure and the "state quantity in the process of a chemical reaction promoted by a catalyst" predicted from a character string representing the molecular structure may be registered in a database.

[0104] <Simultaneous control of multiple catalyst activities> In one embodiment of the present invention, multiple "state quantities in the process of a chemical reaction promoted by a catalyst" are predicted from "information on molecular structure," and the governing factors for each of the multiple predicted state quantities are identified and displayed. For example, activation energy and selectivity can be predicted from "information on molecular structure," or multiple selectivities can be predicted. For example, catalyst design can be performed by simultaneously displaying the governing factors for activation energy and selectivity from "information on molecular structure."

[0105] <Peptide Catalyst Design, Organic Catalyst Design> The present invention can be applied to the design of catalysts (e.g., peptide catalysts, organic catalysts) that contain multiple conformers. Peptide catalysts have flexible structures and can have a variety of conformations. In other words, there may be multiple structures that are energetically close to the most stable structure of the transition state that significantly contributes to the catalytic reaction. Catalyst design requires extracting information about multiple conformations (i.e., structures energetically close to the most stable structure). However, in experiments, information about all conformations is mixed, making it difficult to interpret the catalytic activity obtained from experiments. In the present invention, each conformation can be separated and analyzed using calculated values. Specifically, the prediction unit 105 predicts state quantities from information about each of multiple molecular structures, and the identification unit 106 identifies the governing factors for each predicted state quantity and displays the governing factors.

[0106] <Effects> As described above, one embodiment of the present invention provides a catalyst design system that combines data-driven catalyst design software with a database for storing the data created thereby, thereby improving the accuracy of catalyst design.

[0107] <Hardware Configuration> FIG. 11 is a diagram showing the hardware configuration of a catalyst design device (server) 10 and a user terminal 20 according to one embodiment of the present invention.

[0108] The catalyst design device 10 and the user terminal 20 may include a control unit 1001, a main memory unit 1002, an auxiliary memory unit 1003, an input unit 1004, an output unit 1005, and an interface unit 1006. Each of these will be described below.

[0109] The control unit 1001 is a processor (for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc.) that executes various programs installed in the auxiliary storage unit 1003 .

[0110] The main memory unit 1002 includes a non-volatile memory (Read Only Memory (ROM)) and a volatile memory (Random Access Memory (RAM)). The ROM stores various programs, data, etc. required for the control unit 1001 to execute various programs installed in the auxiliary memory unit 1003. The RAM provides a working area into which the various programs installed in the auxiliary memory unit 1003 are expanded when executed by the control unit 1001.

[0111] The auxiliary storage unit 1003 is an auxiliary storage device that stores various programs and information used when the various programs are executed.

[0112] The input unit 1004 is an input device through which the operator of the catalyst design system 10 and the user terminal 20 inputs various instructions to the catalyst design system 10 and the user terminal 20 .

[0113] The output unit 1005 is an output device that outputs the internal states of the catalyst design device 10 and the user terminal 20, etc.

[0114] The interface unit 1006 is a communication device for connecting to a network and communicating with other devices.

[0115] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to the specific embodiments described above, and various modifications and changes are possible within the scope of the gist of the present invention.

[0116] This international application claims priority to Japanese Patent Application No. 2024-016371, filed on February 6, 2024, the entire contents of which are incorporated herein by reference.

[0117] REFERENCE SIGNS LIST 1 Catalyst design system 10 Catalyst design device (server) 20 User terminal 21 User 101 Acquisition unit 102 Extraction unit 103 Database storage unit 104 Learning unit 105 Prediction unit 106 Identification unit 107 Registration unit 100 Database 201 Prediction model setting unit 202 Catalyst design unit 1001 Control unit 1002 Main storage unit 1003 Auxiliary storage unit 1004 Input unit 1005 Output unit 1006 Interface unit

Claims

1. A catalyst design system for promoting a specific chemical reaction, comprising: an acquisition unit that acquires information for identifying a chemical reaction designated by a user to be promoted by the catalyst; an extraction unit that extracts, from a database, information on the molecular structure included in the chemical reaction identified based on the acquired information and state quantities in the process of the chemical reaction; a learning unit that performs machine learning using the information on the molecular structure and the state quantities extracted from the database to generate a predictive model; a prediction unit that uses the predictive model to predict the state quantity from the information on the molecular structure designated by the user; and an identification unit that identifies a controlling factor of the predicted state quantity and displays the controlling factor.

2. The catalyst design system of claim 1, further comprising: a registration unit that stores in the database information about the molecular structure designed by the user based on the governing factors, and information about the state quantities and molecular structure calculated by a user terminal from the information about the molecular structure designed by the user based on the governing factors.

3. A catalyst design system as described in claim 1, further comprising: a registration unit that stores in the database information on molecular structures in which a portion of the information on the molecular structures stored in the database has been modified, and information on state quantities and molecular structures calculated by a user terminal from information on molecular structures in which a portion of the information on the molecular structures stored in the database has been modified.

4. The catalyst design system according to claim 1, wherein the database is a database in which multiple databases related to catalysts are combined.

5. The catalyst design system according to claim 4, wherein the plurality of databases relating to catalysts includes a database relating to catalysts other than those designed by the catalyst design system.

6. A catalyst design system according to claim 1 or 2, wherein the catalyst includes a large number of conformational isomers, the prediction unit predicts a state quantity from information relating to each of the plurality of molecular structures, and the identification unit identifies a governing factor for each predicted state quantity and displays the governing factor.

7. A catalyst design system as described in claim 1 or 2, wherein the information regarding the molecular structure is a three-dimensional molecular structure, and the database registers the three-dimensional molecular structure predicted using a machine-learned prediction model so that when a string representing a molecular structure is input, the three-dimensional molecular structure is output.

8. A catalyst design system as described in claim 1 or 2, wherein the prediction unit predicts multiple state quantities simultaneously, and the identification unit identifies a controlling factor for each of the multiple predicted state quantities and displays the controlling factor.

9. The catalyst design system according to claim 1 or 2, wherein the controlling factor is a factor that increases or decreases the state quantity.

10. A catalyst design system according to claim 1 or 2, wherein the information relating to the molecular structure is information representing the three-dimensional structure of the molecule.

11. The catalyst design system according to claim 1 or 2, wherein the state quantities include at least one of a state quantity of a reaction intermediate and a state quantity of a transition state.

12. The catalyst design system according to claim 1 or 2, wherein the catalyst is defined by a single molecular structure.

13. The catalyst design system of claim 1 or 2, wherein the information for identifying a chemical reaction promoted by the catalyst includes at least one of reaction conditions, a reaction name, and a catalyst name, and the reaction conditions include a reaction formula, a molecular structure involved in the reaction, a partial structure of a molecule involved in the reaction, a reaction time, and a reaction temperature.

14. A program for causing a system for designing catalysts that promote specific chemical reactions to execute the following processes: a process of acquiring information for identifying a chemical reaction that is promoted by a catalyst, as specified by a user; a process of extracting from a database information about the molecular structure included in the chemical reaction identified based on the acquired information, and state quantities in the process of the chemical reaction; a process of generating a predictive model through machine learning using the information about the molecular structure and the state quantities extracted from the database; a process of predicting the state quantities from the information about the molecular structure specified by the user using the predictive model; and a process of identifying the governing factors of the predicted state quantities and displaying the governing factors.

15. A method executed by a system for designing catalysts that promote specific chemical reactions, comprising: acquiring information for identifying a chemical reaction that is promoted by a catalyst, specified by a user; extracting from a database information about the molecular structure included in the chemical reaction identified based on the acquired information, and state quantities in the process of the chemical reaction; generating a predictive model through machine learning using the information about the molecular structure and the state quantities extracted from the database; predicting the state quantities from the information about the molecular structure specified by the user using the predictive model; and identifying a controlling factor for the predicted state quantity and displaying the controlling factor.

Citation Information

Patent Citations

  • Information processing apparatus and control method

    JP2024016371A