Catalyst design system, catalyst design program, catalyst design method

The catalyst design system uses machine learning to identify and predict dominant factors in catalyst design, enhancing the accuracy of predictive models and facilitating the redesign of catalysts through data-driven methodologies.

JP7843575B2Active Publication Date: 2026-04-10THE INSTITUTE OF PHYSICAL & CHEMICAL RESEARCH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

There is a need for a catalyst design system that utilizes data-driven methodologies and software for designing molecular catalysts, combining high-quality data storage and machine learning techniques to facilitate the design process.

Method used

A catalyst design system that includes an acquisition unit for identifying chemical reactions, an extraction unit for molecular structure and state quantities, a learning unit for generating predictive models, and an identification unit for dominant factors, using machine learning to predict and display key factors in catalyst design.

Benefits of technology

Facilitates the design of catalysts by improving the accuracy of predictive models through data augmentation and visualization of dominant factors, enabling users to redesign catalysts based on these factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843575000001
    Figure 0007843575000001
  • Figure 0007843575000002
    Figure 0007843575000002
  • Figure 0007843575000003
    Figure 0007843575000003
Patent Text Reader

Abstract

The present invention facilitates catalyst designing. A catalyst design system according to one embodiment of the present invention is a system for designing a catalyst that promotes a specific chemical reaction, and comprises: an acquisition unit that acquires information for identifying a chemical reaction that is specified by a user and that is to be promoted by the catalyst; an extraction unit that extracts, from a database, information pertaining to a molecular structure included in the chemical reaction identified on the basis of the acquired information, and a state quantity in the process of the chemical reaction; a learning unit that performs machine learning using the information pertaining to the molecular structure and the state quantity that have been extracted from the database, to generate a prediction model; a prediction unit that uses the prediction model to predict a state quantity from the information pertaining to the molecular structure specified by the user; and an identification unit that identifies a dominant factor of the predicted state quantity and displays the dominant factor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a catalyst design system, a catalyst design program, and a catalyst design method.

Background Art

[0002] Conventionally, in the design of catalysts, actual catalysts are prepared to cause chemical reactions, and the activity and selectivity of the catalysts are evaluated. Recently, research on the design of catalysts using machine learning techniques and data collected through experiments has also been advanced.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

[0004] Regarding molecular catalysts, there is a need for a catalyst design system that utilizes data, combining a data-driven catalyst design method and its methodology, along with software for using that methodology, and a database for storing the high-quality data generated for machine learning.

[0005] Therefore, the present invention aims to facilitate the design of catalysts. [Means for solving the problem]

[0006] A catalyst design system according to one embodiment of the present invention is a catalyst design system for promoting a specific chemical reaction, comprising: an acquisition unit that acquires information for identifying a chemical reaction to be promoted by the catalyst, as specified by a user; an extraction unit that extracts information regarding the molecular structure included in the chemical reaction identified based on the acquired information, and state quantities in the process of the chemical reaction, from a database; a learning unit that generates a predictive model by machine learning using the information regarding the molecular structure and the state quantities extracted from the database; a prediction unit that predicts state quantities from the information regarding the molecular structure specified by the user using the predictive model; and an identification unit that identifies the dominant factors of the predicted state quantities and displays the dominant factors. [Effects of the Invention]

[0007] According to the present invention, catalyst design can be facilitated. [Brief explanation of the drawing]

[0008] [Figure 1] This diagram shows the overall configuration of one embodiment of the present invention. [Figure 2] This is a sequence diagram showing the overall process related to one embodiment of the present invention. [Figure 3] This is a sequence diagram showing an example of a data augmentation process for a database related to one embodiment of the present invention. [Figure 4] This figure shows the functional configuration of a catalyst design system according to one embodiment of the present invention. [Figure 5] This is an example of a database (before registration) stored in a database storage unit according to one embodiment of the present invention. [Figure 6] This is an example of a database (after registration) stored in a database storage unit according to one embodiment of the present invention. [Figure 7] This is an example of a screen displayed on a user terminal according to one embodiment of the present invention (generation of a predictive model). [Figure 8]This is an example of a screen (catalyst design) displayed on a user terminal related to one embodiment of the present invention. [Figure 9] This is an example of a screen (catalyst design) displayed on a user terminal related to one embodiment of the present invention. [Figure 10] This diagram illustrates the database join related to one embodiment of the present invention. [Figure 11] This figure shows the hardware configuration of a catalyst design apparatus (server) and a user terminal according to one embodiment of the present invention. [Modes for carrying out the invention]

[0009] Each embodiment will be described below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.

[0010] <Explanation of Terms> • The "catalyst" can be any catalyst. For example, the "catalyst" is a homogeneous catalyst. For example, the "catalyst" is a catalyst defined by a single molecular structure. For example, the "catalyst" is a catalyst that controls the selectivity of a reaction, such as stereoselectivity, regioselectivity, or chemoselectivity; a catalyst that controls oxidation or reduction; a catalyst that controls bond formation or cleavage; or a catalyst that has multiple of these functions. For example, the "catalyst" is a metal complex catalyst, an organic molecular catalyst, a supramolecular catalyst, an electrochemical catalyst, a radical catalyst, a peptide catalyst, an oxidation-reduction catalyst, a photoredox catalyst, a polymerization catalyst, an enzyme, an asymmetric catalyst, a cross-coupling catalyst, a metathesis catalyst, an acid-base catalyst, etc. • "Information for identifying a catalyst-accelerated chemical reaction" may be any information that can identify a catalyst-accelerated chemical reaction (hereinafter also simply referred to as a "catalyst reaction"). For example, "Information for identifying a catalyst-accelerated chemical reaction" includes at least one of the following: reaction conditions, reaction name, and catalyst name. The reaction conditions may include the reaction equation, molecular structures involved in the reaction, substructures of molecules involved in the reaction, reaction time, and reaction temperature. · "Information on molecular structure" may be any information representing the molecular structure including a catalyst (e.g., transition state structure, intermediate structure). For example, "information on molecular structure" is information representing the three-dimensional structure of a molecule. For example, "information on molecular structure" is the three-dimensional coordinates of atoms constituting a molecule (e.g., xyz file) and the conditions of first-principles calculations for calculating the three-dimensional coordinates. Further, "information on molecular structure" includes descriptors converted from the molecular structure (a descriptor is a numerical value representing molecular information such as the structure and properties of a molecule). Note that "information on molecular structure" is not limited to information on three-dimensional molecular structures and may be information on two-dimensional molecular structures (structural formulas). · "State quantity in the process of a chemical reaction promoted by a catalyst" includes at least one of the state quantity of a reaction intermediate and the state quantity of a transition state. For example, "state quantity in the process of a chemical reaction promoted by a catalyst" is activation energy, energy difference between transition states (e.g., corresponding to enantioselectivity, regioselectivity), etc. · "Dominant factor of the state quantity in the process of a chemical reaction promoted by a catalyst" is a factor that increases or decreases the state quantity. For example, "dominant factor of the state quantity in the process of a chemical reaction promoted by a catalyst" is the spatial position indicating the molecular structure including a catalyst (transition state, reaction intermediate) that dominates the increase or decrease of the state quantity. For example, "dominant factor of the state quantity in the process of a chemical reaction promoted by a catalyst" is any quantity calculated by first-principles calculations from a transition state or a reaction intermediate (e.g., energy level of a molecular orbital, charge of an atom, dipole moment).

[0011] <Overall configuration> FIG. 1 is a diagram showing the overall configuration according to an embodiment of the present invention. The catalyst design system 1 includes a catalyst design device 10 and user terminals 20A, 20B, 20C (hereinafter, collectively referred to as user terminal 20. Note that the number of user terminals is not limited to three). The catalyst design system 1 and the user terminal 20 can transmit and receive data via an arbitrary network.

[0012] Note that the catalyst design system 1 may be implemented only by the catalyst design device (for example, one personal computer) 10 (that is, the user terminal 20 is not used). In this case, the catalyst design device 10 also includes a prediction model setting unit 201 and a catalyst design unit 202 described in FIG. 4. For example, the user 21 can search the database stored in the catalyst design device 10 and design a catalyst using the retrieved data with the catalyst design application in the catalyst design device 10.

[0013] <<Catalyst Design Device>> The catalyst design device 10 executes a process for designing a catalyst in response to an instruction from the user terminal 20. For example, the catalyst design device 10 is one or more computers (for example, a server).

[0014] <<User Terminal>> The user terminal 20 is a terminal operated by users 21A, 21B, 21C (hereinafter collectively referred to as the user 21. Note that the number of users is not limited to three) who design a catalyst. For example, the user terminal 20 is a personal computer, a tablet, a smartphone, or the like.

[0015] Note that part or all of the processes described as the processes executed by the catalyst design device 10 in this specification may be executed by the user terminal 20, or part or all of the processes described as the processes executed by the user terminal 20 in this specification may be executed by the catalyst design device 10.

[0016] <Method> Hereinafter, the overall process will be described while referring to FIG. 2, and the data expansion process of the database will be described while referring to FIG. 3.

[0017] FIG. 2 is a sequence diagram showing the overall process according to an embodiment of the present invention.

[0018] First, we will explain the setup of the predictive model used in catalyst design (steps 101 to 107), then the catalyst design (steps 108 to 113), and finally, the data augmentation of the database that serves as training data for the predictive model (steps 114 to 117).

[0019] [Setting up the predictive model] In step 101 (S101), user 21 inputs "information for identifying a catalyst-accelerated chemical reaction" into user terminal 20 (for example, on the screen of S1001 in Figure 7). For example, "information for identifying a catalyst-accelerated chemical reaction" includes at least one of the following: reaction conditions (which may include the reaction equation, molecular structures involved in the reaction, substructures of molecules involved in the reaction, reaction time, and reaction temperature), reaction name, and catalyst name.

[0020] In step 102 (S102), the user terminal 20 transmits the "information for identifying the chemical reaction promoted by the catalyst" entered in S101 to the catalyst design device 10.

[0021] In step 103 (S103), the catalyst design apparatus 10 extracts from the database 100 the "information on the molecular structure" and "state variables in the process of the catalyst-promoted chemical reaction" for each catalytic reaction, corresponding to the "information for identifying the catalyst-promoted chemical reaction" received in S102. The database 100 stores the following information linked together for each catalytic reaction: "catalyst name (including complexes with other molecules such as substrates)", "state variables in the process of the catalyst-promoted chemical reaction (e.g., enantioselectivity and activation energy)", and "information on the molecular structure (e.g., 3D coordinates (xyz file) of the atoms constituting the molecule and conditions for first-principles calculations)".

[0022] Specifically, the catalyst design device 10 identifies a catalytic reaction from "information for identifying chemical reactions promoted by the catalyst." Next, the catalyst design device 10 extracts "information on the molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" for each catalytic reaction from the database 100.

[0023] In step 104 (S104), the catalyst design device 10 transmits the "information on molecular structure" and the "state quantities in the process of the chemical reaction promoted by the catalyst" extracted in S103 to the user terminal 20.

[0024] In step 105 (S105), user 21 selects the data to be used as training data for the predictive model used in catalyst design on user terminal 20 (for example, on the screens of S1002 and S1003 in Figure 7) (i.e., which catalytic reaction's "information on the molecular structure" and "state variables in the process of the catalyst-promoted chemical reaction" will be used to generate the predictive model).

[0025] In step 106 (S106), the user terminal 20 transmits the "information regarding molecular structure" and the "state quantities in the process of the chemical reaction promoted by the catalyst" selected in S105 to the catalyst design device 10.

[0026] In step 107 (S107), the catalyst design apparatus 10 uses the "information on molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" received in S106 as training data to perform machine learning and generate a predictive model. The predictive model is a trained model (specifically, a regression model (supervised learning), a classification model (supervised learning), a dimensionality reduction model, or a clustering model (unsupervised learning)) that has been trained by machine learning so that when "information on molecular structure" is input, "state variables in the process of the chemical reaction promoted by the catalyst" are output.

[0027] [Catalyst Design] In step 108 (S108), user 21 inputs "information regarding the molecular structure" into the user terminal 20 (for example, on the screen of S2001 in Figure 8 and S3001 in Figure 9) (i.e., designs the molecular structure of the catalyst).

[0028] In step 109 (S109), the user terminal 20 transmits the "information regarding molecular structure" entered in S108 to the catalyst design device 10 (i.e., transmits the molecular structure of the catalyst designed by user 21).

[0029] In step 110 (S110), the catalyst design apparatus 10 uses the prediction model generated in S107 to predict the state variables in the chemical reaction process promoted by the catalyst from the information on molecular structure received in S109.

[0030] In step 111 (S111), the catalyst design apparatus 10 identifies the dominant factors of the "state variables in the process of the chemical reaction promoted by the catalyst" predicted in S110. The catalyst design apparatus 10 also evaluates the similarity between the identified "dominant factors" or the similarity between structures (descriptors).

[0031] Specifically, the catalyst design device 10 identifies dominant factors based on the importance of descriptors x of regression and classification models, which output "state variables in the process of a chemical reaction promoted by the catalyst" when "information on molecular structure" is input. For example, if "information on molecular structure (e.g., molecular structure descriptors)" is x, and "state variables in the process of a chemical reaction promoted by the catalyst" and the values ​​calculated from those state variables are y, then dominant factors are identified based on the value of βi (regression coefficients, model parameters) in y = f(x) = β1x1 + β2x2 + ... + βnxn. For example, in nonlinear regression and classification models, dominant factors are identified based on descriptor importance evaluation methods such as importance in decision trees or SHAP (SHapley Additive exPlanations) values.

[0032] For example, similarities between governing factors or between structures (descriptors) can be calculated and identified by combining visualization through clustering and dimensionality reduction, and calculation of distances between descriptors.

[0033] In step 112 (S112), the catalyst design device 10 transmits to the user terminal 20 the "state variables in the chemical reaction process promoted by the catalyst" predicted in S110 and the "dominant factors" identified in S111. The catalyst design device 10 also transmits to the user terminal 20 the similarity between the "dominant factors" identified in S111 or the similarity between structures (descriptors).

[0034] In step 113 (S113), the user terminal 20 displays the "state variables in the process of the catalyst-accelerated chemical reaction" and the "governing factors" received in S112 (for example, the screens for S2002 in Figure 8 and S3002 in Figure 9).

[0035] [Database data expansion] In step 114 (S114), user 21 inputs "information on molecular structure" with some of the "information on molecular structure" from S108 modified into the user terminal 20 (for example, on the screen of S2003 in Figure 8, or S3003 in Figure 9) (i.e., modifying and redesigning the molecular structure of the catalyst from S108).

[0036] In step 115 (S115), the user terminal 20 calculates the "state variables in the process of a catalyst-accelerated chemical reaction" from the "information on molecular structure" input in S114 using first-principles calculations.

[0037] In this specification, we will mainly describe the case using first-principles calculations (which may include first-principles calculations incorporating empirical methods such as experimental parameters, or calculations may be performed using a quantum computer). However, in this invention, calculations may be performed using a machine learning-trained model instead of first-principles calculations. Alternatively, a structure close to the transition state may be obtained by optimization by fixing bond lengths, etc., and the calculation may be performed using this structure close to the transition state.

[0038] In step 116 (S116), the user terminal 20 transmits the "information on molecular structure" entered in S114, the "information on (optimized) molecular structure" calculated in S115, and the "state variables in the process of the chemical reaction promoted by the catalyst" to the catalyst design device 10 (that is, it transmits the molecular structure of the catalyst redesigned by user 21 (including complexes with other molecules such as substrates) and the state variables in the process of the chemical reaction promoted by the catalyst).

[0039] In step 117 (S117), the catalyst design device 10 adds the "information on molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" received in S116 to the database 100. Alternatively, someone other than the user adds the "information on molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" to the database 100 from publicly available information such as papers that mention the use of the catalyst design device 10.

[0040] Figure 3 is a sequence diagram showing an example of a database data augmentation process according to one embodiment of the present invention.

[0041] In step 201 (S201), user 21 inputs "information for identifying a catalyst-accelerated chemical reaction" into user terminal 20 (for example, on the screen of S1001 in Figure 7). For example, "information for identifying a catalyst-accelerated chemical reaction" includes at least one of the following: reaction conditions (which may include the reaction equation, molecular structures involved in the reaction, substructures of molecules involved in the reaction, reaction time, and reaction temperature), reaction name, and catalyst name.

[0042] In step 202 (S202), the user terminal 20 transmits the "information for identifying the chemical reaction promoted by the catalyst" entered in S201 to the catalyst design device 10.

[0043] In step 203 (S203), the catalyst design apparatus 10 extracts from the database 100 the "information on the molecular structure" and "state variables in the process of the catalyst-promoted chemical reaction" for each catalytic reaction, corresponding to the "information for identifying the catalyst-promoted chemical reaction" received in S202. The database 100 stores the following information linked together for each catalytic reaction: "catalyst name (including complexes with other molecules such as substrates)", "state variables in the process of the catalyst-promoted chemical reaction (e.g., enantioselectivity and activation energy)", and "information on the molecular structure (e.g., 3D coordinates (xyz file) of the atoms constituting the molecule and conditions for first-principles calculations)".

[0044] Specifically, the catalyst design device 10 identifies a catalytic reaction from "information for identifying chemical reactions promoted by the catalyst." Next, the catalyst design device 10 extracts "information on the molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" for each catalytic reaction from the database 100.

[0045] In step 204 (S204), the catalyst design device 10 transmits the "information on molecular structure" and the "state quantities in the process of the chemical reaction promoted by the catalyst" extracted in S203 to the user terminal 20.

[0046] In step 205 (S205), user 21 inputs the "Molecular Structure Information" received in S204 into user terminal 20, with some of the "Molecular Structure Information" being corrected (i.e., corrects the "Molecular Structure Information" in database 100 received in S204).

[0047] In step 206 (S206), the user terminal 20 calculates the "state variables in the process of a catalyst-accelerated chemical reaction" from the "information on molecular structure" entered in S205 using first-principles calculations.

[0048] In this specification, we will mainly describe the case using first-principles calculations (which may include first-principles calculations incorporating empirical methods such as experimental parameters, or calculations may be performed using a quantum computer). However, in this invention, calculations may be performed using a machine learning-trained model instead of first-principles calculations. Alternatively, a structure close to the transition state may be obtained by optimization by fixing bond lengths, etc., and the calculation may be performed using this structure close to the transition state.

[0049] In step 207 (S207), the user terminal 20 transmits the "information on molecular structure" entered in S205, the "information on (optimized) molecular structure" calculated in S206, and the "state quantities in the process of the chemical reaction promoted by the catalyst" to the catalyst design device 10 (that is, the user 21 transmits the molecular structure of the catalyst (including complexes with other molecules such as substrates) and the state quantities in the process of the chemical reaction promoted by the catalyst, for which the data in the database 100 has been modified by the user 21).

[0050] In step 208 (S208), the catalyst design device 10 adds the "information on molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" received in S207 to the database 100. Alternatively, someone other than the user adds the "information on molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" to the database 100 from publicly available information such as papers that mention the use of the catalyst design device 10.

[0051] <Functional Configuration> Figure 4 is a diagram showing the functional configuration of a catalyst design system 1 according to one embodiment of the present invention.

[0052] <<Catalyst design equipment>> The catalyst design apparatus 10 may include an acquisition unit 101, an extraction unit 102, a database storage unit 103, a learning unit 104, a prediction unit 105, a specification unit 106, and a registration unit 107. Furthermore, the catalyst design apparatus 10 can function as the acquisition unit 101, the extraction unit 102, the learning unit 104, the prediction unit 105, the specification unit 106, and the registration unit 107 by executing a program.

[0053] The acquisition unit 101 receives instructions from the user terminal 20 for generating a predictive model using data stored in the database 100 in the database storage unit 103. Specifically, the acquisition unit 101 acquires "information for identifying catalyst-accelerated chemical reactions" specified by the user 21. The acquisition unit 101 also acquires information selected by the user 21 regarding the data to be used as training data for the predictive model (i.e., which catalytic reaction's "molecular structure information" and "state variables in the process of the catalyst-accelerated chemical reaction" will be used to generate the predictive model).

[0054] The acquisition unit 101 receives instructions from the user terminal 20 for expanding the data in the database 100 stored in the database storage unit 103. Specifically, the acquisition unit 101 acquires "information on molecular structure" in the database 100 that has been partially modified by the user 21, and "state quantities in the process of catalyst-accelerated chemical reactions" calculated by first-principles calculations by the user terminal 20.

[0055] The extraction unit 102 extracts information regarding the molecular structure of a chemical reaction identified based on the "information for identifying a catalyst-accelerated chemical reaction" acquired by the acquisition unit 101, as well as state quantities in the process of that chemical reaction, from the database 100 stored in the database storage unit 103. Specifically, the extraction unit 102 identifies a catalytic reaction from the "information for identifying a catalyst-accelerated chemical reaction." Next, the extraction unit 102 extracts "information regarding the molecular structure" and "state quantities in the process of the catalyst-accelerated chemical reaction" for each catalytic reaction from the database 100 and transmits them to the user terminal 20.

[0056] The database storage unit 103 stores the database 100 (which will be described in detail later with reference to Figures 5 and 6).

[0057] The learning unit 104 generates a predictive model by machine learning using "information on molecular structure" and "state variables in the process of catalyst-accelerated chemical reactions" extracted from the database 100 as training data. Alternatively, the "information on molecular structure" and "state variables in the process of catalyst-accelerated chemical reactions" extracted from the database 100, selected by the user 21, may be used. The predictive model is a pre-trained model (specifically, a regression model (supervised learning), a classification model (supervised learning), a dimensionality reduction model, or a clustering model (unsupervised learning)) that, when "information on molecular structure" is input, outputs "state variables in the process of catalyst-accelerated chemical reactions". The output of the predictive model may include uncertainty in the predicted values ​​(e.g., variance used in Bayesian optimization).

[0058] The prediction unit 105 uses the prediction model generated by the learning unit 104 to predict the "state variables in the process of a catalyst-accelerated chemical reaction" from the "information on molecular structure" specified by the user 21. For example, the prediction unit 105 converts the molecular structure of the catalyst into descriptors, inputs these descriptors into the prediction model, and outputs the "state variables in the process of a catalyst-accelerated chemical reaction."

[0059] The following describes an example of how the prediction unit 105 makes predictions using the prediction model generated by the learning unit 104.

[0060] [Regression] For example, the learning unit 104 generates regression models (e.g., a model that predicts the numerical value of the activation energy, a model that predicts the energy difference (reaction selectivity)). The prediction unit 105 uses the regression models to predict numerical values ​​(e.g., the numerical value of the activation energy) that represent "state variables in the process of a catalyst-promoted chemical reaction" from "information on molecular structure" specified by the user 21.

[0061] [Classification] For example, the learning unit 104 generates a classification model (for example, a model that predicts whether something is highly active or low active, or a model that classifies whether the energy difference (reaction selectivity) is large or small). The prediction unit 105 uses the classification model to predict classes representing "state variables in the process of a catalyst-promoted chemical reaction" from "information on molecular structure" specified by the user 21 (for example, two classes: highly active and low active; however, it may be classified into three or more classes).

[0062] [Clustering] For example, the learning unit 104 generates a clustering model that clusters "information about molecular structure". The prediction unit 105 uses the clustering model to cluster the "information about molecular structure" specified by the user 21 and calculates the similarity between the "information about molecular structure" and a predetermined catalyst (e.g., a high-performance catalyst). Depending on the case, visualization by dimensionality reduction and calculation of distance between descriptors may be used individually or in combination to calculate the similarity. If the similarity between the "information about molecular structure" and the predetermined catalyst (e.g., a high-performance catalyst) is high (e.g., above a threshold), the prediction unit 105 calculates the "state variables in the chemical reaction process promoted by the catalyst" of the "information about molecular structure" using first-principles calculations or the like.

[0063] The "information about molecular structure" used for the clustering, visualization by dimensionality reduction, and calculation of distances between descriptors described above may be obtained by optimizing the structure to be close to the transition state by fixing bond lengths, etc., and then using that structure close to the transition state.

[0064] The identification unit 106 identifies the dominant factors of the "state variables in the process of a catalyst-accelerated chemical reaction" predicted by the prediction unit 105 and displays these dominant factors on the user terminal 20. Specifically, the identification unit 106 identifies the dominant factors based on the importance of descriptors x of regression and classification models that output "state variables in the process of a catalyst-accelerated chemical reaction" when "information on molecular structure" is input. For example, if "information on molecular structure (e.g., a descriptor of molecular structure)" is x, and "state variables in the process of a catalyst-accelerated chemical reaction" and the value calculated from those state variables are y, then the dominant factors are identified based on the value of βi (regression coefficients, model parameters) in y = f(x) = β1x1 + β2x2 + ... + βnxn. For example, in nonlinear regression and classification models, the dominant factors are identified based on descriptor importance evaluation methods such as importance in decision trees or SHAP (SHapley Additive exPlanations) values.

[0065] The registration unit 107 registers (adds) new data to the database 100 stored in the database storage unit 103. Specifically, the registration unit 107 stores in the database 100 information regarding molecular structures designed by user 21 based on governing factors, and state variables calculated by first-principles calculations by user terminal 20. In addition, the registration unit 107 stores in the database 100 information regarding molecular structures with some of the information stored in the database 100 modified, and state variables calculated by first-principles calculations by user terminal 20.

[0066] <<User Terminal>> The user terminal 20 may include a prediction model setting unit 201 and a catalyst design unit 202. Furthermore, the user terminal 20 can function as both the prediction model setting unit 201 and the catalyst design unit 202 by executing a program.

[0067] The prediction model setting unit 201 instructs the generation of a prediction model used in catalyst design. Specifically, the prediction model setting unit 201 transmits to the catalyst design device 10 the "information for identifying the chemical reaction promoted by the catalyst" that user 21 has entered into user terminal 20. The prediction model setting unit 201 also receives "information on molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" corresponding to the "information for identifying the chemical reaction promoted by the catalyst" from the catalyst design device 10, and transmits to the catalyst design device 10 the "information on molecular structure" and "state variables in the process of the chemical reaction promoted by the catalyst" selected by user 21.

[0068] The prediction model setting unit 201 instructs the database 100 to expand its data. Specifically, the prediction model setting unit 201 transmits the "information on molecular structure" in the database 100, which has been partially modified by the user 21, to the catalyst design device 10. The prediction model setting unit 201 also calculates the "state variables in the process of the chemical reaction promoted by the catalyst" from the "information on molecular structure" in the database 100, which has been partially modified by the user 21, using first-principles calculations (for example, DFT (Density Functional Theory) calculations), and transmits this calculation to the catalyst design device 10.

[0069] The catalyst design unit 202 provides instructions for catalyst design. It transmits "information on molecular structure" specified by user 21 to the catalyst design device 10 and receives "state variables in the chemical reaction process promoted by the catalyst" and "governing factors" predicted from the information on molecular structure. The catalyst design unit 202 also transmits information on molecular structure designed by user 21 based on the governing factors to the catalyst design device 10. Furthermore, the catalyst design unit 202 calculates "state variables in the chemical reaction process promoted by the catalyst" from the information on molecular structure designed by user 21 based on the governing factors using first-principles calculations (e.g., DFT (Density Functional Theory) calculations) and transmits this calculation to the catalyst design device 10.

[0070] <database> The following describes database 100 with reference to Figures 5 and 6.

[0071] Figure 5 shows an example of a database (before registration) stored in the database storage unit 103 according to one embodiment of the present invention.

[0072] The database storage unit 103 stores the database 100. The database 100 stores the following information for each catalytic reaction, linked together: the "catalyst name (including complexes with other molecules such as substrates)", "state variables in the chemical reaction process promoted by the catalyst (e.g., enantioselectivity and activation energy)", and "information about the molecular structure (e.g., the three-dimensional coordinates (xyz file) of the atoms constituting the molecule and the conditions for first-principles calculations)".

[0073] Figure 6 shows an example of a database (after registration) stored in the database storage unit 103 according to one embodiment of the present invention.

[0074] For example, information about the molecular structure designed by user 21 based on governing factors, and the molecular structure (optimized molecular structure) and state variables calculated by first-principles calculations from the information about the designed molecular structure are registered in database 100.

[0075] Thus, in one embodiment of the present invention, various users 21 design the molecular structure of a catalyst based on dominant factors, and the molecular structure and state quantity data of the catalyst designed based on those dominant factors (i.e., catalysts that reflect the dominant factors) are added to the database. Therefore, the accuracy of the predictive model generated using the data in the database can be improved.

[0076] For example, information on a molecular structure in which some of the information on the molecular structure stored in database 100 has been modified (e.g., by introducing substituents), and the molecular structure (optimized molecular structure) and state variables calculated by first-principles calculations from the information on the modified molecular structure are registered in database 100.

[0077] Thus, in one embodiment of the present invention, various users 21 modify parts of the molecular structure of catalysts in the database 100, and the data of the modified molecular structure and state quantities of the catalysts are added to the database. Therefore, the accuracy of the predictive model generated using the data in the database can be improved.

[0078] <Screen> The following describes the screens displayed on the user terminal 20, referring to Figures 7 to 9.

[0079] Figure 7 shows an example of a screen displayed on a user terminal 20 according to one embodiment of the present invention (generation of a predictive model).

[0080] In step 1001 (S1001), a screen is displayed for user 21 to input "information to identify the chemical reaction accelerated by the catalyst." On the S1001 screen, user 21 inputs "information to identify the chemical reaction accelerated by the catalyst (e.g., reaction conditions, reaction name, catalyst name, etc.)."

[0081] In step 1002 (S1002), the "state variables in the process of the catalyst-promoted chemical reaction (e.g., enantioselectivity (energy difference between transition states) and activation energy)" and "information regarding the molecular structure" for each catalytic reaction, corresponding to the "information for identifying the catalyst-promoted chemical reaction" entered in S1001, are extracted from database 100 and displayed.

[0082] In step 1003 (S1003), user 21 selects data to be used as training data for the prediction model.

[0083] Thus, in one embodiment of the present invention, user 21 can select training data for the predictive model used in catalyst design, and thus generate a predictive model suitable for the catalyst he or she wants to design.

[0084] Figure 8 shows an example of a screen (catalyst design) displayed on a user terminal 20 according to one embodiment of the present invention.

[0085] In step 2001 (S2001), the molecular structure containing the catalyst (e.g., transition state structure, intermediate structure) designed by user 21 (e.g., by modifying some molecular structure) is displayed.

[0086] In step 2002 (S2002), the dominant factors governing the state variables in the chemical reaction process facilitated by the catalyst designed in S2001 are displayed. Specifically, a predictive model is used to predict the state variables from information about the molecular structure including the catalyst designed in S2001, and the dominant factors governing those state variables are identified.

[0087] In step 2003 (S2003), user 21 adds a molecular substructure to the spatial position of the dominant factor displayed in S2002 (in the example in Figure 7, the dominant factor is assumed to be a factor that increases enantioselectivity or a factor that decreases activation energy).

[0088] Subsequently, information on the molecular structure, including the substructure of the molecule added in S2003, and the state variables calculated from that molecular structure information using first-principles calculations, are added to database 100.

[0089] Thus, in one embodiment of the present invention, the dominant factors of the state quantities of the catalyst designed by user 21 are visualized, so user 21 can redesign the catalyst by adding molecular structures based on these visualized dominant factors.

[0090] Figure 9 shows an example of a screen (catalyst design) displayed on a user terminal 20 according to one embodiment of the present invention.

[0091] In step 3001 (S3001), the molecular structure containing the catalyst (e.g., transition state structure, intermediate structure) designed by user 21 (e.g., by modifying some molecular structure) is displayed.

[0092] In step 3002 (S3002), the dominant factors of the state variables in the chemical reaction process facilitated by the catalyst designed in S3001 are displayed. Specifically, using a predictive model, the state variables are predicted from information on the molecular structure including the catalyst designed in S3001, and the dominant factors of those state variables are identified.

[0093] In step 3003 (S3003), user 21 deletes the molecular substructure located in the spatial position of the dominant factor displayed in S3002 (in the example in Figure 8, the dominant factor is assumed to be a factor that reduces enantioselectivity or a factor that increases activation energy).

[0094] Subsequently, information on the molecular structure from which the molecular substructure was removed in S3003, and state variables calculated from that molecular structure information by first-principles calculations, are added to database 100.

[0095] Thus, in one embodiment of the present invention, the dominant factors of the state quantities of the catalyst designed by user 21 are visualized, so user 21 can redesign the catalyst by deleting the molecular structure based on the visualized dominant factors.

[0096] <Database joining> In one embodiment of the present invention, the catalyst design apparatus 10 can manage multiple databases relating to catalysts by combining them. For example, the catalyst design apparatus 10 may combine multiple tabular databases, or it may combine multiple graph-format databases such as RDF (Resource Description Framework).

[0097] For example, the catalyst design device 10 manages the database in RDF graph format, as shown in Figure 10. Specifically, the catalyst design device 10 manages the relationships between three elements (subject, predicate, and objective).

[0098] In the database for catalyst A in Figure 10, the subject is the transition state a, the predicate is the activation energy, and the object is the numerical value of the activation energy. Also, the subject is the transition state a, the predicate is the type of catalyst, and the object is catalyst A. Also, the subject is the transition state a, the predicate is the coordinate information, and the object is TSa.xyz (3D coordinate file of transition state structure a). Also, the subject is the transition state a, the predicate is the type of reaction, and the object is reaction equation 1.

[0099] In the database for catalyst B in Figure 10, the subject is transition state b, the predicate is activation energy, and the object is high activity or low activity. Also, the subject is transition state b, the predicate is catalyst type, and the object is catalyst B. Also, the subject is transition state b, the predicate is coordinate information, and the object is TSb.xyz (3D coordinate file of transition state structure b). Also, the subject is transition state b, the predicate is reaction type, and the object is reaction equation 1.

[0100] When reaction equation 1 in the database for catalyst A in Figure 10 is the same as reaction equation 1 in the database for catalyst B in Figure 10, the databases for catalyst A and catalyst B are joined using reaction equation 1 as the joining key. Note that the joining key is not limited to a reaction equation.

[0101] Furthermore, the multiple databases related to catalysts may include databases related to catalysts other than those designed by catalyst design system 1 (for example, a database of basic molecular properties, a database of commercially available compounds, a database of material properties, a database of descriptors, and a database of synthesis routes).

[0102] <Database Search> In one embodiment of the present invention, catalytic reactions, as well as their transition states, intermediates, and state quantities, can be searched in a database using the method described below. Therefore, similar reactions (for example, reactions with the same substrate but different catalysts) can be easily searched. • Reaction name. Includes the name of the entire reaction and the names of the elementary reactions (e.g., overall reaction name: Michael addition, 1,2 addition, etc.; elementary reaction names: oxidative addition, reductive elimination, etc.). • [Catalyst] Presence or absence of metal, type of metal, type of ligand, catalyst name, ligand name, substituent names included in the structure, molecular formula / structural formula, various molecular representation methods such as SMARTS / SMILES / InChl. • [Substrate / Reagent] Type of substrate / reagent, substrate name, reagent name, names of substituents included in the structure, molecular formula / structural formula, various molecular representation methods such as SMARTS / SMILES / InChl. Furthermore, if necessary, all elements included in the reaction, such as additives and solvents, may be made searchable by type, name, molecular formula / structural formula, and SMARTS / SMILES / InChl. Furthermore, when searching using reaction names, etc., it may be possible to select and search for multiple candidates (for example, searching for Ru and Ir simultaneously as metals). The same applies to the presence or absence of metals, the type of metal, the type of ligand, and the type of substrate / reactant.

[0103] <Utilizing Databases> In one embodiment of the present invention, a predictive model may be generated by machine learning using a string representing the structure of the molecules used in the reaction (e.g., SMILES) and the molecular structure (e.g., the transition state of an elementary reaction) as training data. This predictive model is a pre-trained model that, when a string representing the molecular structure is input, outputs a three-dimensional molecular structure (transition state or reaction intermediate). This predictive model is used to predict the molecular structure from the string representing the molecular structure (that is, the string representing the molecular structure is input into the predictive model and the molecular structure is output). Alternatively, a predictive model may be generated by machine learning using molecular structure and state variables in the process of a catalyst-accelerated chemical reaction as training data. This predictive model is a pre-trained model that, when "molecular structure" is input, outputs "state variables in the process of a catalyst-accelerated chemical reaction". This predictive model is used to predict "state variables in the process of a catalyst-accelerated chemical reaction" from "molecular structure" (that is, "molecular structure" is input into the predictive model and "state variables in the process of a catalyst-accelerated chemical reaction" is output). It may be possible to register in a database the "molecular structure" predicted from the string representing the molecular structure, and the "state variables in the process of a catalyst-accelerated chemical reaction" predicted from the string representing the molecular structure.

[0104] <Simultaneous control of multiple catalytic activities> In one embodiment of the present invention, it is possible to predict multiple "state variables in the process of a chemical reaction promoted by a catalyst" from "information on molecular structure," identify the dominant factors for each of the predicted multiple state variables, and display the dominant factors. For example, it is possible to predict the activation energy and selectivity, or multiple selectivity, from "information on molecular structure." For example, catalyst design can be performed by simultaneously displaying the dominant factors for activation energy and selectivity from "information on molecular structure."

[0105] <Peptide catalyst design, organic catalyst design> The present invention can be applied to the design of catalysts containing multiple conformational isomers (e.g., peptide catalysts, organocatalysts). Peptide catalysts have a flexible structure and exist in various conformations. That is, there may be multiple structures that are energetically close to the most stable transition state structure that contributes significantly to the catalytic reaction. Catalyst design requires extracting information on multiple conformations (i.e., structures energetically close to the most stable structure), but in experiments, information from all conformations is mixed together, making it difficult to interpret the catalytic activity obtained experimentally. In the present invention, each conformation can be separated and analyzed using calculated values. Specifically, the prediction unit 105 predicts state quantities from information on each of the multiple molecular structures, and the identification unit 106 identifies the dominant factors for each predicted state quantity and displays the dominant factors.

[0106] <Effects> Thus, in one embodiment of the present invention, by providing a catalyst design system that combines data-driven catalyst design software with a database for storing the data generated therefrom, the accuracy of catalyst design can be improved.

[0107] <Hardware Configuration> Figure 11 shows the hardware configuration of a catalyst design apparatus (server) 10 and a user terminal 20 according to one embodiment of the present invention.

[0108] The catalyst design apparatus 10 and user terminal 20 may include a control unit 1001, a main memory unit 1002, an auxiliary memory unit 1003, an input unit 1004, an output unit 1005, and an interface unit 1006. Each of these will be described below.

[0109] The control unit 1001 is a processor (for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc.) that executes various programs installed in the auxiliary storage unit 1003.

[0110] The main memory unit 1002 includes non-volatile memory (ROM (Read Only Memory)) and volatile memory (RAM (Random Access Memory)). The ROM stores various programs, data, etc., necessary for the control unit 1001 to execute various programs installed in the auxiliary memory unit 1003. The RAM provides a work area that is expanded when the various programs installed in the auxiliary memory unit 1003 are executed by the control unit 1001.

[0111] The auxiliary storage unit 1003 is an auxiliary storage device that stores various programs and information used when various programs are executed.

[0112] The input unit 1004 is an input device that allows operators of the catalyst design device 10 and the user terminal 20 to input various instructions to the catalyst design device 10 and the user terminal 20.

[0113] The output unit 1005 is an output device that outputs the internal status of the catalyst design apparatus 10 and the user terminal 20.

[0114] The interface unit 1006 is a communication device for connecting to a network and communicating with other devices.

[0115] Although embodiments of the present invention have been described in detail above, the present invention is not limited to the specific embodiments described above, and various modifications and changes are possible within the scope of the gist of the present invention.

[0116] This international application claims priority based on Japanese Patent Application No. 2024-016371, filed on 6 February 2024, and the entire contents of No. 2024-016371 are incorporated herein by reference. [Explanation of Symbols]

[0117] 1. Catalyst design system 10. Catalyst design equipment (server) 20 User Terminals 21 users 101 Acquisition Department 102 Extraction part 103 Database Storage Unit 104 Learning Department 105 Prediction Section 106 Specific part 107 Registration Department 100 databases 201 Prediction Model Setting Section 202 Catalyst Design Department 1001 Control Unit 1002 Main memory 1003 Auxiliary storage unit 1004 Input section 1005 Output section 1006 Interface section

Claims

1. A catalyst design system for promoting specific chemical reactions, An acquisition unit that acquires information to identify a chemical reaction accelerated by a catalyst, as specified by the user, An extraction unit extracts from a database information regarding the molecular structure included in the chemical reaction identified based on the acquired information, and state quantities during the chemical reaction process. A learning unit that generates a predictive model by machine learning using the information on molecular structure and the state variables extracted from the database, A prediction unit that uses the prediction model to predict state variables from information about the molecular structure specified by the user, A unit that identifies the dominant factors of the predicted state quantities and displays the dominant factors. A catalyst design system equipped with [specific features / features].

2. A registration unit that stores in a database information relating to a molecular structure designed by the user based on the dominant factors, and information relating to state variables and molecular structure calculated by the user terminal from the information relating to a molecular structure designed by the user based on the dominant factors. The catalyst design system according to claim 1, further comprising the above.

3. A registration unit that stores in the database information the molecular structure information obtained by modifying a portion of the molecular structure information stored in the database, and state variables and molecular structure information calculated by the user terminal from the molecular structure information obtained by modifying a portion of the molecular structure information stored in the database, The catalyst design system according to claim 1, further comprising the above.

4. The catalyst design system according to claim 1, wherein the database is a database formed by combining multiple databases relating to catalysts.

5. The catalyst design system according to claim 4, wherein the plurality of databases relating to the catalysts include a database relating to catalysts other than those designed by the catalyst design system.

6. The catalyst comprises numerous conformational isomers, The prediction unit predicts state variables from information about each of the plurality of molecular structures, The catalyst design system according to claim 1 or 2, wherein the identifying unit identifies the dominant factors for each predicted state quantity and displays the dominant factors.

7. The information regarding the molecular structure is a three-dimensional molecular structure. The catalyst design system according to claim 1 or 2, wherein the database registers the three-dimensional molecular structure predicted using a machine learning prediction model that outputs the three-dimensional molecular structure when a string representing the molecular structure is input.

8. The prediction unit predicts multiple state variables simultaneously, The catalyst design system according to claim 1 or 2, wherein the identifying unit identifies the dominant factor for each of the predicted plurality of state quantities and displays the dominant factor.

9. The catalyst design system according to claim 1 or 2, wherein the dominant factor is a factor that increases or decreases the state quantity.

10. The catalyst design system according to claim 1 or 2, wherein the information relating to the molecular structure is information representing the three-dimensional structure of the molecule.

11. The catalyst design system according to claim 1 or 2, wherein the state weight includes at least one of the state weight of a reaction intermediate and the state weight of a transition state.

12. The catalyst design system according to claim 1 or 2, wherein the catalyst is a catalyst defined by a single molecular structure.

13. Information for identifying the chemical reaction promoted by the catalyst includes at least one of the reaction conditions, the reaction name, and the catalyst name. The catalyst design system according to claim 1 or 2, wherein the reaction conditions include a reaction formula, a molecular structure included in the reaction, a partial structure of the molecules included in the reaction, a reaction time, and a reaction temperature.

14. In a system for designing catalysts that promote specific chemical reactions, A process to obtain information to identify a chemical reaction accelerated by a catalyst, as specified by the user, The process involves extracting information from a database regarding the molecular structure of the chemical reaction identified based on the acquired information, and the state quantities during the chemical reaction process. A process to generate a predictive model by machine learning using the information on molecular structure and the state variables extracted from the aforementioned database, Using the aforementioned prediction model, a process is performed to predict state variables from information about the molecular structure specified by the user, A process to identify the dominant factors of the predicted state variables and to display the dominant factors. A program to execute.

15. A method performed by a catalyst design system for accelerating a specific chemical reaction, To obtain information to identify the catalyst-accelerated chemical reaction specified by the user, Based on the information obtained, information regarding the molecular structure of the chemical reaction identified, and the state quantities in the process of the chemical reaction are extracted from the database. Using the information on molecular structure and the state variables extracted from the aforementioned database, machine learning is performed to generate a predictive model. Using the aforementioned prediction model, the state variables are predicted from the information on the molecular structure specified by the user, Identify the dominant factors of the predicted state variables and display the dominant factors. A method that includes this.