Machine learning program, machine learning method, and information processing apparatus

A machine learning program automates catalyst search by analyzing atomic arrangements and chemical reactions to efficiently identify promising catalyst compositions, addressing the challenges of manual search complexity and inaccuracy.

JP7708310B2Active Publication Date: 2025-07-15FUJITSU LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024515220
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-12
Publication Date
2025-07-15
Estimated Expiration
2042-04-12

AI Technical Summary

Technical Problem

The search for promising catalyst compositions is laborious and inaccurate due to the vast number of search axes and potential candidates, making it difficult to automate catalyst search effectively.

Method used

A machine learning program that extracts feature quantities related to the surface structure of a substance based on atomic arrangement, using atomic arrangement information and chemical reaction data to train a model for predicting chemical reactions, thereby automating the search for promising catalyst compositions.

Benefits of technology

Enables rapid and accurate identification of promising catalyst compositions by automatically extracting relevant features and causal relationships, reducing the search time and labor required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708310000001
    Figure 0007708310000001
  • Figure 0007708310000002
    Figure 0007708310000002
  • Figure 0007708310000003
    Figure 0007708310000003
Patent Text Reader

Abstract

This information processing device extracts features related to the surface structure of a substance on the basis of the atomic arrangement of the substance. The information processing device uses training data including, as explanatory variables, atomic arrangement information related to the atomic arrangement of the substance and the extracted features, and including, as an objective variable, information related to the chemical reaction that occurs in the substance, and executes training of a machine leaning model that predicts information related to the chemical reaction that occurs in the substance corresponding to the input explanatory variables.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine learning program, a machine learning method, and an information processing apparatus.

Background Art

[0002] The physical properties of chemical catalysts depend on the composition of the material and the compatibility of the chemical properties with the reactants, and it is difficult to analytically estimate them. In addition, there are many possibilities for selection, such as the combination of catalyst materials, the mixing amount, and the method of creating the surface structure. In order to conduct catalyst exploration to search for promising catalyst compositions, it takes an enormous amount of time to actually conduct chemical experiments or simulate chemical reactions.

[0003] In recent years, due to the development of AI (Artificial Intelligence), a method for estimating the properties of catalysts more simply than simulating chemical reactions has been adopted, and the acceleration of catalyst exploration has been attempted.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] By the way, the possibilities for candidates in catalyst search are diverse, the search range is enormous, and it is difficult to automate catalyst search. For example, there are many types of search axes including elements that can form a catalyst such as catalyst material, mixing ratio, surface shape, etc., and the search range increases exponentially. Therefore, when the number of values that can be taken on each axis is X and the number or type of search axes is n, the search range is X n and becomes an enormous range.

[0007] In the AI for accelerating the above catalyst search, in order to estimate catalyst characteristics, “descriptors” that represent the chemical characteristics of the catalyst, such as the characteristics of the catalyst and reactants, are important as training data for learning appropriate characteristics. Also, in order to narrow down the search range, it is required to learn appropriate feature quantities.

[0008] However, when narrowing down the search axes with AI, candidates for the search axes need to be prepared. However, it is laborious to manually extract the search axes, and due to effects such as extraction omission, human error, and preconceived notions, the accuracy is not high. Thus, it is difficult to accelerate catalyst search for exploring promising catalyst compositions.

[0009] On one aspect, an object is to provide a machine learning program, a machine learning method, and an information processing apparatus that can rapidly search for promising catalyst compositions.

Means for Solving the Problems

[0010] In the first aspect, the machine learning program causes a computer to extract feature quantities related to the surface structure of a substance based on the atomic arrangement of the substance, and includes the atomic arrangement information regarding the atomic arrangement of the substance and the extracted feature quantities as explanatory variables, and executes training of a machine learning model that predicts information regarding a chemical reaction occurring in the substance corresponding to the input explanatory variables using training data including information regarding the chemical reaction occurring in the substance as an objective variable.

Advantages of the Invention

[0011] According to one embodiment, a promising catalyst composition can be searched for at high speed.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Best Mode for Carrying Out the Invention

[0013] Hereinafter, examples of the machine learning program, machine learning method, and information processing apparatus according to the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited by these examples. Also, the respective examples can be appropriately combined within a non - contradictory range.

Examples

[0014] (Description of the Information Processing Apparatus) FIG. 1 is a diagram for explaining an information processing apparatus 10 according to Example 1. The information processing apparatus 10 shown in FIG. 1 is an example of a computer apparatus that efficiently executes catalyst search for exploring a promising catalyst composition from a wide variety of combinations of catalyst materials with many options such as mixing amount and how to form the surface structure.

[0015] This information processing apparatus 10 extracts a feature amount related to the surface structure of a substance based on the atomic arrangement of the substance. The information processing apparatus 10 includes atomic arrangement information regarding the atomic arrangement of the substance and the extracted feature amount as explanatory variables, and uses training data including information regarding a chemical reaction occurring in the substance as an objective variable to train a machine learning model for predicting information regarding a chemical reaction occurring in the substance corresponding to the input explanatory variables. Here, the information processing apparatus 10 extracts the causal relationship between a descriptor used as an explanatory variable that has a large influence on the objective variable and the objective variable. As a result, the information processing apparatus 10 can efficiently execute catalyst search and output the search result.

[0016] For example, as shown in FIG. 1, the information processing apparatus 10 acquires simulation data including the coordinates of catalyst atoms obtained by a generally used chemical simulation, physical property descriptions of catalyst atoms, and physical property descriptions of reactants. The information processing apparatus 10 extracts structure characteristic data representing characteristics related to the surface structure of the catalyst from the simulation data.

[0017] Then, the information processing apparatus 10 generates a machine learning model using the simulation data and the structural characteristic data as feature quantities and information related to chemical reactions in the analysis of catalysts, such as "the magnitude of the reaction energy", as the target variable. For example, the information processing apparatus 10 extracts a combination of descriptors representing the chemical characteristics of the catalyst from the simulation data and the structural characteristic data. In addition, the information processing apparatus 10 generates causal relationship information including the causal relationship between each combination of descriptors used for training the machine learning model and the target variable.

[0018] After that, the information processing apparatus 10 performs a chemical simulation on the catalyst to be predicted and generates structural characteristic data, and extracts descriptors for the catalyst to be predicted. The information processing apparatus 10 extracts the characteristics that have a great influence on the performance of the catalyst to be predicted based on the presence or absence of the combination of descriptors included in the causal relationship information. Note that the results of the chemical simulation may use different systems or existing data.

[0019] In this way, the information processing apparatus 10 automatically extracts features related to the three-dimensional structure of the catalyst from the simulation data evaluated in various catalyst studies as the exploration area of the catalyst researcher, comprehensively verifies the combinations of feature quantities, and extracts the causal relationships for each condition group. Therefore, the information processing apparatus 10 can automatically extract the exploration axis and can rapidly explore promising catalyst compositions. In addition, the explored promising catalyst compositions can be applied to the analysis of catalyst characteristics for discovering new catalysts and reaction mechanisms and to the inference system for inferring the catalyst composition.

[0020] (Functional Configuration) FIG. 2 is a functional block diagram showing the functional configuration of the information processing apparatus 10 according to the first embodiment. As shown in FIG. 2, the information processing apparatus 10 includes a communication unit 11, a storage unit 12, and a control unit 20.

[0021] The communication unit 11 is a processing unit that controls communication with other devices and is realized, for example, by a communication interface. The communication unit 11 receives various instructions from an external device such as an administrator terminal and transmits the prediction results generated by the control unit 20 to an external device such as an administrator terminal.

[0022] The storage unit 12 is a storage device that stores various data and programs executed by the control unit 20 and is realized, for example, by a memory or a hard disk. This storage unit 12 has a simulation data DB 13, a structural characteristic data DB 14, and a prediction target data DB 15.

[0023] The simulation data DB 13 is a database that stores atomic arrangement information of substances that are the results of chemical simulations. The data stored here may be data acquired from an external simulation terminal or data generated by the control unit 20.

[0024] Figure 3 is a diagram for explaining the simulation data DB 13. As shown in Figure 3, the simulation data DB 13 stores, for each "nama" indicating the catalyst to be simulated, the coordinates of the catalyst atoms, the physical property descriptions of the catalyst atoms, the physical property descriptions of the reactants, etc. In the example of Figure 3, descriptors such as "x1, y1, z1,..., target" are stored for "Data1" indicating the catalyst. Here, "x1, y1, z1" indicate the coordinates of atom 1, and "target" indicates a target variable specified in advance such as the magnitude of the reaction energy. Note that it includes not only coordinate data but also characteristic data of the atoms at each lattice point, such as atomic number, electronegativity, ionic radius, etc.

[0025] The structural characteristic data DB 14 is a database that stores structural characteristic data related to the structural characteristics of the catalyst generated by the control unit 20. Figure 4 is a diagram for explaining the structural characteristic data DB 14. As shown in Figure 4, the structural characteristic data DB 14 stores, for each "name" indicating the catalyst, feature amount data indicating the feature amounts of the catalyst.

[0026] In the example of FIG. 4, descriptors such as "kink_1, step_1, vac_1, island_1 ···, target" are stored for "Data1" indicating the catalyst. Here, "kink_1, step_1, vac_1, island_1" are data representing, for each lattice position of the catalyst structure, whether the atoms or spaces present are kink, step, vacancy, or island as a feature quantity (1: Yes (corresponding), 0: No (not corresponding)). "target" indicates a target variable specified in advance such as the magnitude of the reaction energy.

[0027] The prediction target data DB15 is data to be predicted using a machine learning model, and is a database that stores prediction target data regarding the catalyst to be predicted. For example, the data stored in the prediction target data DB15 may be data before input to chemical simulation, or may be data including chemical simulation results and structure characteristic data.

[0028] The control unit 20 is a processing unit that controls the entire information processing apparatus 10, and is realized by, for example, a processor or the like. The control unit 20 includes a simulation execution unit 30, a machine learning unit 40, a causal relationship generation unit 50, and a prediction processing unit 60. Note that the simulation execution unit 30, the machine learning unit 40, the causal relationship generation unit 50, and the prediction processing unit 60 are realized by an electronic circuit such as a processor or a process executed by a processor.

[0029] The simulation execution unit 30 is a processing unit that executes chemical simulation. For example, the simulation execution unit 30 executes an atom-level simulation, generates the coordinates of catalyst atoms, the physical property description of catalyst atoms, and the physical property description of reactants, and stores them in the simulation data DB13. Note that as data for simulation targets, data regarding substances, genetic information of patients, material data targeted for chemical reactions, and the like can be adopted.

[0030] The machine learning unit 40 includes a data extraction unit 41, a combination extraction unit 42, and a model generation unit 43, and is a processing unit that extracts feature quantities related to the surface structure of a substance based on the atomic arrangement of the substance and trains a machine learning model using training data including the atomic arrangement information and feature quantities of the substance.

[0031] The data extraction unit 41 is a processing unit that extracts feature quantity data related to the structural characteristics of a catalyst from the atomic arrangement information of the substance stored in the simulation data DB 13. Specifically, the data extraction unit 41 extracts surface structure information as a catalyst. For example, the data extraction unit 41 extracts surface structure information of a catalyst such as "Island" indicating atoms constituting an island of a certain scale, "Vacancy" indicating lattice points forming holes of a certain scale, "Step" indicating atoms that generate steps on the surface, and "Kink" indicating atoms hitting the corners of steps, and stores it in the structural characteristic data DB 14.

[0032] More specifically, the data extraction unit 41 defines the conditions of surface features and extracts the surface structure information of the catalyst by determining the state of each crystal lattice point. FIG. 5 is a diagram for explaining an extraction example of structural characteristic data. In FIG. 5, an extraction example in the case of taking a lattice structure such as a simple lattice structure or a cubic lattice structure will be described as an example.

[0033] As shown in FIG. 5, after defining that the side where the catalyst crystal exists is the bottom and the surface is the upper side (see (a) of FIG. 5), the data extraction unit 41 determines various conditions such as "Kink" and "Step" and extracts the structural characteristics. For example, when there is no catalyst atom above the target lattice point, and for the four adjacent lattices in the horizontal direction of the target lattice, catalyst atoms exist in two adjacent lattices and no catalyst atoms exist in the other two lattices, the data extraction unit 41 determines it as "Kink". That is, the data extraction unit 41 determines an atom with atoms behind and to the left and no atoms in front and to the right as "Kink".

[0034] In addition, when there is no catalytic atom above the target lattice point, and for the four adjacent lattices in the horizontal direction of the target lattice, there are catalytic atoms in three adjacent lattices and no catalytic atom in the other one lattice, the data extraction unit 41 determines it as "step". That is, the data extraction unit 41 determines an atom with atoms in front, behind, and to the left and no atom to the right as "step".

[0035] The combination extraction unit 42 is a processing unit that extracts combinations of descriptors using descriptor information such as the atomic arrangement information of substances stored in the simulation data DB 13 and the structural characteristic data stored in the structural characteristic data DB 14. Specifically, the combination extraction unit 42 extracts all combination patterns of descriptors (data items) from each descriptor of the atomic arrangement information and each descriptor of the structural characteristic data as hypotheses (knowledge chunks). Note that the descriptors may include information such as the atomic number of atoms arranged at each lattice position, atomic characteristics, the direction and magnitude of the force applied to the atoms, and the interaction between atoms.

[0036] FIG. 6 is a diagram for explaining combinations of descriptors. As shown in FIG. 6, the combination extraction unit 42 uses descriptors included in atomic arrangement data obtained from simulations, structural characteristic data, other atomic characteristic data input from the outside, etc., to extract combinations such as "(coordinates of atom 1) and (kink_2 = 1)" and "(coordinates of atom n) and (island_3 = 0) and (vacancy_5 = 1)". These combinations of descriptors are used not only as explanatory variables of the machine learning model but also for causal relationship analysis. The combination extraction unit 42 extracts the explanatory variables of the machine learning model. The combination extraction unit 42 stores the extracted information in the storage unit 12 and outputs it to the causal relationship generation unit 50. Note that, for example, the method described in Non-Patent Document 1 can be adopted as the combination extraction method.

[0037] The model generation unit 43 is a processing unit that generates a machine learning model using the combination of descriptors generated by the combination extraction unit 42 and the target variable specified in advance or the target variable determined based on the simulation result. This model generation unit 43 comprehensively checks the combinations of descriptors, which are many factors of the analysis target data, and automatically selects combinations that are highly relevant to the target variable indicating, for example, "the amount of energy for the reaction" used in catalyst analysis, and generates a machine learning model (prediction model). Note that the prediction result can be explained based on this combination of factors.

[0038] Figure 7 is a diagram for explaining the generation of a machine learning model. As shown in Figure 7, for each combination of descriptors generated by the combination extraction unit 42, the model generation unit 43 sets information related to chemical reactions in catalyst analysis, such as the magnitude of energy and reaction rate, as the target variable (result), and generates a machine learning model. For example, the model generation unit 43 comprehensively investigates each combination of descriptors using the technology of Non-Patent Document 1 and extracts combinations that have a large impact on the target variable (Target). Then, the model generation unit 43 executes machine learning of the machine learning model using the extracted combination as the explanatory variable and the above target variable. By doing so, a machine learning model with higher accuracy and higher interpretability can be created compared to directly using each descriptor.

[0039] The causal relationship generation unit 50 is a processing unit that generates the causal relationship between explanatory variables used for training the machine learning model or the causal relationship between the explanatory variable (combination of descriptors) and the target variable. Specifically, the causal relationship generation unit 50 comprehensively checks the combination of features of the machine learning model using a technique that analyzes which factor is the cause and which is the result by analyzing the mutual influence when two factors change with each other, and extracts the causal relationship for each condition group.

[0040] More specifically, the causal relationship generation unit 50 generates the causal relationship between the descriptors and the target variable by extracting the degree of influence given by each combination of descriptors for each condition set as the target variable of the machine learning model. For example, the causal relationship generation unit 50 uses, as the grouping rule of the original data, the "combination that has a large influence on the target variable" among the combinations extracted by the combination extraction unit 42. Then, the causal relationship generation unit 50 individually extracts causal relationships that, when viewed as a whole, are mixed and cancel each other out so that they cannot be seen, by analyzing the causal relationships within a specific group.

[0041] Referring to FIG. 7 for explanation, the causal relationship generation unit 50 generates, as causal relationship 1, "the magnitude of the influence of kink_1 = 1 and vac_2 = 1 on reducing the reaction energy of the catalyst is 10.22", etc. The causal relationship generation unit 50 generates, as causal relationship 2, "the magnitude of the influence of element_1 = 44 (atomic number 4: Ru), step_1 = 1, and island_1 = 1 on reducing the reaction energy of the catalyst is 8.74", etc.

[0042] Here, the causal relationships generated by the causal relationship generation unit 50 are represented by the causal relationships between the above-described respective feature amounts. FIG. 8 is a diagram for explaining the causal relationship. As shown in FIG. 8, the causal relationship generation unit 50 analyzes the causal relationship based on the correlation narrowed down by machine learning using a method such as "DirectLiNGAM", for example, and generates a result represented by the causal relationship between each feature amount. That is, the causal relationship generation unit 50 analyzes how much the value of the "result" is affected when the "cause" increases by 1. In the example of FIG. 8, the causal relationship generation unit 50 generates the causal relationship that "when the atom at lattice point No. 53 becomes a kink (changes from kink_53 = 0 to kink_53 = 1), the reaction energy is reduced by 0.19". Also, in FIG. 8, all the causal relationships not connected to energy are causal relationships between explanatory variables.

[0043] In addition, the causal relationship generation unit 50 can output a schematic diagram of the physical property structure of the catalyst, and in the schematic diagram, atoms or lattice points having a causal relationship equal to or greater than a predetermined value can be highlighted according to the content of the causal relationship. For example, the causal relationship generation unit 50 maps the causal relationship of the catalyst structure generated in the analysis of chemical catalyst characteristics to three-dimensional catalyst data, and performs highlighting such as color and shape according to the content of the causal relationship on atoms and lattice points having a particularly strong causal relationship that is equal to or greater than the threshold value.

[0044] FIG. 9 is a diagram for explaining an output example of a causal relationship. As shown in FIG. 9, when "the fact that the 32nd atom is vacant affects the reaction energy of the catalyst", the causal relationship generation unit 50 highlights the position of the 32nd atom of the catalyst structure. As another example, when "the lattice point 1 has a kink structure and the lattice point 2 has no atom and has a vacancy structure, which causes the reduction of the reaction energy of the catalyst", the lattice point 1 representing the kink causality is highlighted in red, the lattice point 2 representing the vacancy causality is highlighted in black, and the other lattice points are displayed in white.

[0045] As described above, the information processing apparatus 10 according to the first embodiment can estimate the influence of descriptors including structural characteristic data and combinations thereof on the target variable using a machine learning model and can illustrate it. Further, the information processing apparatus 10 can also estimate the target variable in the same manner as normal machine learning using this machine learning model.

[0046] The prediction processing unit 60 includes a data generation unit 61 and a prediction unit 62, and is a processing unit that executes prediction processing on prediction target data using a machine learning model. Specifically, the prediction processing unit 60 identifies a corresponding causal relationship among the causal relationships generated by the causal relationship generation unit 50 based on the prediction target data, and thereby identifies, for the prediction target data, for example, a descriptor having a large reaction energy of the catalyst.

[0047] The data generation unit 61 is a processing unit that generates structure characteristic data from prediction target data by the same method as the data extraction unit 41. Further, the data generation unit 61 can also generate atomic arrangement data from the prediction target data. The data generation unit 61 stores each generated data in the storage unit 12 and outputs it to the prediction unit 62.

[0048] The prediction unit 62 is a processing unit that executes prediction processing on prediction target data using a machine learning model. FIG. 10 is a diagram for explaining a prediction example using causal relationships. As shown in FIG. 10, for example, the prediction unit 62 generates a combination of a plurality of descriptors from the atomic arrangement data and the structure characteristic data generated from the prediction target data, for example, using the technique of Non-Patent Document 1. Then, the prediction unit 62 inputs the generated combination of a plurality of descriptors into the machine learning model generated by the machine learning unit 40 and obtains a prediction result. As a result, the prediction unit 62 can obtain the prediction result and calculate the predicted value of the target variable.

[0049] Further, the prediction unit 62 refers to the generated causal relationships generated during the training of the machine learning model and identifies the causal relationship corresponding to the combination of descriptors generated from the prediction target data. Then, the prediction unit 62 outputs the identified causal relationship as a prediction result to a display or the like and transmits it to the administrator terminal. For example, when the combination of "kink_1 = 1" and "vac_2 = 1" is included in the combination of descriptors generated from the prediction target data, the prediction unit 62 determines that it corresponds to causal relationship 1 and predicts that "it includes descriptors that have a large impact on reducing the reaction energy of the catalyst".

[0050] (Flow of processing) FIG. 11 is a flowchart showing the flow of processing. Here, an example in which prediction processing is executed after machine learning processing will be described, but these can be realized by separate flows.

[0051] As shown in FIG. 11, when the machine learning unit 40 of the information processing apparatus 10 starts processing (S101: Yes), it executes a chemical simulation to generate simulation data (S102).

[0052] Subsequently, the machine learning unit 40 extracts structure characteristic data from the simulation data (S103). Then, the machine learning unit 40 generates a combination of descriptors from the simulation data and the structure characteristic data (S104).

[0053] After that, the machine learning unit 40 performs machine learning using the descriptors and their combinations as explanatory variables and the specified target variable to generate a combination of important descriptors and a machine learning model using the same (S105), and generates causal relationship information using this combination of important descriptors (S106).

[0054] After that, when the prediction processing unit 60 acquires prediction target data (S107: Yes), it extracts structure characteristic data from the prediction target data (S108). Then, when the prediction processing unit 60 predicts the target variable from the prediction target data using the machine learning model, it specifies the corresponding causal relationship using the causal relationship information (S109).

[0055] (Effect) As described above, the information processing apparatus 10 can automatically and rapidly extract feature amounts from atomic-level simulation data by extracting only the feature amounts for each atom, as high-order feature amounts that are generally used, rather than, for example, features of a structure composed of a plurality of atoms, temperature, pressure, etc. applied to a lump of atoms.

[0056] The information processing apparatus 10 can apply causal discovery after grouping, as a group having a large influence on the reaction energy that is the output of atomic-level simulation, the feature amounts for each atom, on the condition of the group regarding the feature amounts for each atom, by machine learning. As a result, due to the characteristics of machine learning (for example, wide learning) in which the information processing apparatus 10 can comprehensively examine all combinations of conditions, the information processing apparatus 10 can also find groups of feature amounts for each atom that a person has overlooked, compared to the case where a person gives high-order feature amounts and performs causal discovery.

[0057] By being used for confirming the catalytic reaction process and judging the priority of the range to be explored for catalyst candidates, the information processing apparatus 10 can realize more updated and efficient catalyst exploration at low cost and in a short time. In particular, the information processing apparatus 10 can reduce feature quantity design and exploration axis selection, which are processes that require advanced judgment by experts at the initial stage of the exploration plan, and can explore promising catalyst compositions at high speed. In addition, the information processing apparatus 10 can more accurately discover the causality during the reaction.

[0058] By estimating the characteristics that have a great influence on the catalyst performance, the information processing apparatus 10 narrows down the exploration axis (variables or targets to be tried by changing parameters) in catalyst exploration to those estimated to have a "great influence", thereby reducing the catalyst exploration range and greatly reducing the time and labor required to obtain results. That is, the information processing apparatus 10 can reduce n of X in the exploration range. n

[0059] The information processing apparatus 10 can quickly narrow down the exploration range that is huge and unrealistic by hand and cannot be narrowed down by general AI. The information processing apparatus 10 can also provide such narrowed-down results (causal relationships) to other AIs (machine learning models).

Embodiment

[0060] Now, although the embodiments of the present invention have been described so far, the present invention may be implemented in various different forms other than the above-described embodiments.

[0061] (Numerical values, etc.) The items of the simulation data, descriptors, combinations of descriptors, items of the structural characteristic data, causal relationships, etc. used in the above embodiments are merely examples and can be arbitrarily changed. Also, the processing flows described in each flowchart can be appropriately changed within a non-contradictory range.

[0062] (Input of descriptors) ​For example, in the above embodiment, an example where the information processing apparatus 10 extracts descriptors by chemical simulation or surface structure extraction has been described, but it is not limited thereto. For example, the information processing apparatus 10 can also receive descriptors that cannot be automatically extracted from an experimenter or an evaluator and use them as explanatory variables.

[0063] Further, when the information processing apparatus 10 evaluates specific descriptors such as descriptors specified by the user or descriptors to be evaluated, it can also execute generation of causal relationships using combinations including the specific descriptors. In this case, the information processing apparatus 10 can generate the desired results of the user at high speed compared to the case of generating causal relationships for all descriptors.

[0064] (System) Regarding the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified.

[0065] Also, each component of each illustrated device is a functional concept, and it is not necessarily physically configured as shown in the figure. That is, the specific form of distribution and integration of each device is not limited to that shown in the figure. In other words, all or part of it can be functionally or physically distributed and integrated in arbitrary units according to various loads and usage situations. For example, the simulation execution unit 30, the machine learning unit 40, the causal relationship generation unit 50, and the prediction processing unit 60 can also be realized by separate computers (enclosures).

[0066] Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic.

[0067] (Hardware) FIG. 12 is a diagram for explaining a hardware configuration example. As shown in FIG. 12, the information processing apparatus 10 includes a communication apparatus 10a, a HDD (Hard Disk Drive) 10b, a memory 10c, and a processor 10d. Also, each unit shown in FIG. 12 is connected to each other by a bus or the like.

[0068] The communication apparatus 10a is a network interface card or the like and communicates with other devices. The HDD 10b stores programs and databases for operating the functions shown in FIG. 2.

[0069] The processor 10d reads a program for executing the same processing as each processing unit shown in FIG. 2 from the HDD 10b or the like and expands it in the memory 10c, thereby operating a process for executing each function described in FIG. 2 and the like. For example, this process executes the same functions as each processing unit included in the information processing apparatus 10. Specifically, the processor 10d reads a program having the same functions as the simulation execution unit 30, the machine learning unit 40, the causal relationship generation unit 50, the prediction processing unit 60, etc. from the HDD 10b or the like. Then, the processor 10d executes a process for executing the same processing as the simulation execution unit 30, the machine learning unit 40, the causal relationship generation unit 50, the prediction processing unit 60, etc.

[0070] In this way, the information processing apparatus 10 operates as an information processing apparatus that executes an information processing method by reading and executing a program. Also, the information processing apparatus 10 can read the program from a recording medium by a medium reader and execute the read program to realize the same functions as those in the above-described embodiments. Note that the program in this other embodiment is not limited to being executed by the information processing apparatus 10. For example, the above-described embodiments may be similarly applied when another computer or server executes the program, or when these cooperate to execute the program.

[0071] This program may be distributed via a network such as the Internet. Also, this program may be recorded on a computer-readable recording medium such as a hard disk, flexible disk (FD), CD-ROM, MO (Magneto-Optical disk), DVD (Digital Versatile Disc), etc., and executed by being read from the recording medium by a computer.

Explanation of Signs

[0072] 10 Information processing apparatus 11 Communication unit 12 Storage unit 13 Simulation data DB 14 Structural characteristic data DB 15 Data to be predicted DB 20 Control unit 30 Simulation execution unit 40 Machine learning unit 41 Data extraction unit 42 Combination extraction unit 50 Causal relationship generation unit 60 Prediction processing unit 61 Data generation unit 62 Prediction unit

Claims

1. Cause a computer to Extract feature quantities related to the surface structure of a substance based on the atomic arrangement of the substance, Using training data that includes the atomic arrangement information regarding the atomic arrangement of the substance and the extracted feature quantities as explanatory variables, and includes information regarding the chemical reaction that occurs in the substance as a target variable, execute training of a machine learning model that predicts information regarding the chemical reaction that occurs in the substance corresponding to the input explanatory variables, When the training of the machine learning model is completed, extract the causal relationship between each of the feature quantities set as the explanatory variables and the degree of influence on the target variable, A machine learning program characterized by causing the above processing to be executed.

2. The extraction process Extract feature quantities related to the surface structure of the substance according to the atomic arrangement information obtained by chemical simulation regarding the catalyst and the condition definition of the surface features of the substance, The process of executing the training Generate a combination of descriptors indicating feature quantities representing the chemical properties of the catalyst from the atomic arrangement information and the feature quantities, Generate the machine learning model using the training data with the combination of descriptors as the explanatory variable and the information regarding the chemical reaction as the target variable, The machine learning program according to claim 1, characterized in that.

3. The extraction process Extract features regarding the three-dimensional structure of the catalyst as the feature quantities according to the atomic arrangement information obtained by chemical simulation regarding the catalyst and the condition definition of the surface features of the substance, The machine learning program according to claim 2, characterized in that.

4. The extraction process For each of the descriptors, extract the causal relationship between the descriptor and the degree of influence on the target variable. The machine learning program according to claim 2, characterized in that.

5. Output a schematic diagram of the physical property structure of the catalyst, In the schematic diagram, perform highlighting according to the content of the causal relationship on atoms or lattice points having a causal relationship equal to or greater than a predetermined value, The machine learning program according to claim 2, characterized by causing the computer to execute the process.

6. Extract the feature quantities from the prediction target data regarding the catalyst to be predicted according to the atomic arrangement information and the condition definition of the surface features of the substance, Inputting the atomic sequence information and the feature amount generated from the data to be predicted into the machine learning model to predict the chemical reaction in the analysis of the catalyst for the catalyst to be predicted. The machine learning program according to claim 2, wherein the computer is caused to execute the processing.

7. A computer Extracts feature amounts related to the surface structure of a substance based on the atomic sequence of the substance. Using training data that includes the atomic sequence information regarding the atomic sequence of the substance and the extracted feature amounts as explanatory variables, and includes information regarding the chemical reaction occurring in the substance as a target variable, trains a machine learning model to predict information regarding the chemical reaction occurring in the substance corresponding to the input explanatory variables. When the training of the machine learning model is completed, extracts the causal relationship between each of the feature amounts set for the explanatory variables and the degree of influence on the target variable. A machine learning method characterized by executing the processing.

8. Extracts feature amounts related to the surface structure of a substance based on the atomic sequence of the substance. Using training data that includes the atomic sequence information regarding the atomic sequence of the substance and the extracted feature amounts as explanatory variables, and includes information regarding the chemical reaction occurring in the substance as a target variable, trains a machine learning model to predict information regarding the chemical reaction occurring in the substance corresponding to the input explanatory variables. When the training of the machine learning model is completed, extracts the causal relationship between each of the feature amounts set for the explanatory variables and the degree of influence on the target variable. An information processing apparatus characterized by having a control unit.

Citation Information

Patent Citations

  • Method for evaluating liquefaction reaction activity for coal liquefaction catalyst

    JP1993288665A

  • System for predicting and analyzing economic time sequential data

    JP1993342191A

  • Catalyst searching method

    JP2001264309A

  • Method for searching optical resolution factor, method for judging possibility of optical resolution and method of resolution

    JP2005194254A

  • Information processing method and learning model

    JP2020077206A