Characteristic prediction system, characteristic prediction method, and characteristic prediction program
The characteristic prediction system uses partial structure and composition ratio data to enhance material property prediction by inputting into a machine learning model, addressing the challenge of incomplete chemical structure information.
Patent Information
- Application Number
- JP2021073162
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-04-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-04-23
AI Technical Summary
Existing methods struggle to efficiently predict the properties of materials when only a part of the chemical structure is clear, as they fail to effectively learn and utilize incomplete partial structures.
A characteristic prediction system that generates partial structure input data based on known partial structures and composition ratios, which are then input into a machine learning model to predict material characteristics.
Enables efficient prediction of material properties by reflecting composition ratios in partial structure input data, improving prediction accuracy for multi-component substances with unclear chemical structures.
Smart Images

Figure 0007707626000001 
Figure 0007707626000002 
Figure 0007707626000003
Abstract
Description
Technical Field
[0001] One aspect of the present disclosure relates to a property prediction system, a property prediction method, and a property prediction program.
Background Art
[0002] Conventionally, the structure of a molecule has been obtained in a predetermined format, converted into vector information, and input into a machine learning algorithm to predict properties. For example, a method for predicting the binding property between the three-dimensional structure of a biopolymer and the three-dimensional structure of a compound using machine learning is known (see Patent Document 1 below). In this method, a predicted three-dimensional structure of a complex of a biopolymer and a compound is generated based on the three-dimensional structure of the biopolymer and the three-dimensional structure of the compound, the predicted three-dimensional structure is converted into a predicted three-dimensional structure vector, and the binding property between the three-dimensional structure of the biopolymer and the three-dimensional structure of the compound is predicted by discriminating the predicted three-dimensional structure vector using a machine learning algorithm.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, a technique for predicting the properties of a material using a neural network with data representing a structure such as a molecular graph of a material with a clear structure as input has been known. However, in a case where only a part of the chemical structure that is an input for machine learning is clear, it has not been realized to learn a plurality of incomplete partial structures and predict the properties of the material. Therefore, a mechanism for efficiently predicting the properties of a material with only a part of the chemical structure being clear is desired.
Means for Solving the Problems
[0005] A characteristic prediction system according to an aspect of the present disclosure is a characteristic prediction system that predicts the characteristics of a material based on a plurality of raw materials including known partial structures, and includes at least one processor. The at least one processor receives at least an input of partial structure data that identifies the known partial structures of each of the plurality of raw materials and composition ratio data that represents the composition ratio of each of the plurality of raw materials, generates partial structure input data that represents the known partial structures based on the partial structure data for each of the plurality of raw materials, reflects the composition ratio data regarding the plurality of raw materials in the partial structure input data of the plurality of raw materials, and inputs the input data based on the partial structure input data for each of the plurality of raw materials in which the composition ratio data is reflected into a machine learning model.
[0006] Alternatively, a characteristic prediction method according to another aspect of the present disclosure is a characteristic prediction method executed by a computer including at least one processor, and predicts the characteristics of a material based on a plurality of raw materials including known partial structures. The method includes at least receiving an input of partial structure data that identifies the known partial structures of each of the plurality of raw materials and composition ratio data that represents the composition ratio of each of the plurality of raw materials, generating partial structure input data that represents the known partial structures based on the partial structure data for each of the plurality of raw materials, reflecting the composition ratio data regarding the plurality of raw materials in the partial structure input data of the plurality of raw materials, and inputting the input data based on the partial structure input data for each of the plurality of raw materials in which the composition ratio data is reflected into a machine learning model.
[0007] Alternatively, the characteristic prediction program according to another aspect of the present disclosure is a characteristic prediction program for predicting the characteristics of a material based on a plurality of raw materials including known partial structures, and causes a computer to at least receive an input of partial structure data for specifying each known partial structure of the plurality of raw materials and composition ratio data representing the composition ratio of each of the plurality of raw materials; generate partial structure input data representing the known partial structures based on the partial structure data for each of the plurality of raw materials; reflect the composition ratio data regarding the plurality of raw materials in the partial structure input data of the plurality of raw materials; and input the input data based on the partial structure input data for each of the plurality of raw materials in which the composition ratio data is reflected into a machine learning model.
[0008] According to the above aspect, partial structure input data representing known partial structures is generated based on the partial structure data for each of the plurality of raw materials, the composition ratio of the plurality of raw materials is reflected in the partial structure input data of the plurality of raw materials, and the input data based on the partial structure input data for each of the plurality of raw materials in which the ratio is reflected is input into the machine learning model. As a result, for a material manufactured based on a plurality of raw materials in which only a part of the chemical structure is clear, by processing the input data by machine learning, the characteristics of the material can be efficiently predicted.
Advantages of the Invention
[0009] According to an aspect of the present disclosure, the characteristics of a material manufactured based on a raw material in which only a part of the chemical structure is clear can be efficiently predicted.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the description, the same reference numerals are used for the same elements or elements having the same function, and redundant descriptions are omitted.
[0012] [Outline of the System] The characteristic prediction system 1 according to the embodiment is a computer system that executes prediction processing of the characteristics of a multi-component substance, which is a material manufactured by blending a plurality of raw materials containing known partial structures in various ratios, using a machine learning model. The raw material refers to a chemical substance in which at least a part of the structure in the molecular structure is clear and a chemical substance with an unknown structure used to generate a multi-component substance. For example, it is a monomer, a polymer, or a single molecule such as a low molecular weight additive, a solute molecule, or a gas molecule. One raw material may contain a plurality of types of molecules. A multi-component substance is a chemical substance generated by blending a plurality of raw materials in a predetermined ratio. For example, when the raw material is a monomer or a polymer, it is a polymer alloy; when the raw material is a solute molecule or a solvent, it is a mixed solution; and when the raw material is a gas molecule, it is a mixed gas.
[0013] The object of the prediction process by the property prediction system 1 is the properties of multi-component substances. The properties of multi-component substances refer to, for example, when the multi-component substance is a resin, thermal properties such as glass transition temperature and melting point, mechanical properties, adhesiveness, etc. Also, when the multi-component substance is another type of substance, the properties of the multi-component substance are the efficacy or toxicity of a drug, the risks such as the ignition point of a combustible substance, the properties in appearance, the suitability for a specific use, etc. Machine learning is used for the prediction process of the property prediction system 1. Machine learning is a method of autonomously finding laws or rules based on the given information. The specific method of machine learning is not limited. For example, machine learning may be machine learning using a machine learning model which is a computational model. More specifically, the computational model is a neural network. A neural network refers to a model of information processing that imitates the structure of the human brain's nervous system. As more specific examples of other computational models, in addition to neural networks, SVR (Support Vector Regression), random forest, etc. may also be used.
[0014] [System Configuration] The property prediction system 1 is composed of one or more computers. When using multiple computers, these computers are connected via a communication network such as the Internet or an intranet, and logically, one property prediction system 1 is constructed.
[0015] FIG. 1 is a diagram showing an example of the general hardware configuration of the computer 100 that constitutes the property prediction system 1. For example, the computer 100 includes a processor (e.g., CPU) 101 that executes an operating system, application programs, etc., a main memory unit 102 composed of a ROM and a RAM, an auxiliary storage unit 103 composed of a hard disk, a flash memory, etc., a communication control unit 104 composed of a network card or a wireless communication module, an input device 105 such as a keyboard, a mouse, a touch panel, etc., and an output device 106 such as a monitor, a touch panel display, etc.
[0016] Each functional element of the characteristic prediction system 1 is realized by causing the processor 101 to read a predetermined program on the processor 101 or the main memory unit 102 and execute the program. The processor 101 operates the communication control unit 104, the input device 105, or the output device 106 according to the program, and reads and writes data in the main memory unit 102 or the auxiliary storage unit 103. Data or databases required for processing are stored in the main memory unit 102 or the auxiliary storage unit 103.
[0017] FIG. 2 is a diagram showing an example of the functional configuration of the characteristic prediction system 1. The characteristic prediction system 1 includes an input data generation system 10, a training unit 20, and a predictor 30. These input data generation system 10, training unit 20, and predictor 30 may be constructed on the same computer 100, or some of them may be constructed on another computer 100. First, the functional configuration of the input data generation system 10 will be described. The input data generation system 10 includes an acquisition unit 11, a generation unit 12, a vector conversion unit 13, and a synthesis unit 14 as functional elements.
[0018] The acquisition unit 11 is a functional element that accepts the input of partial structure data regarding known partial structures in each molecule constituting a plurality of raw materials that are the basis of a multi-component substance to be predicted, and mixing ratio data representing the mixing ratio of each raw material when it is assumed that these plurality of raw materials are mixed to produce the multi-component substance, and the number data representing the number of known partial structures in the molecules of the plurality of raw materials. The acquisition unit 11 may acquire these data according to a selection input by the user of the input data generation system 10 from the database in the input data generation system 10, or may acquire them according to a selection by the user from an external computer or the like, or the user may input them directly.
[0019] Specifically, the acquisition unit 11 acquires at least first partial structure data that identifies a partial structure included in the molecule of the first raw material and second partial structure data that identifies a partial structure included in the molecule of the second raw material. These partial structure data are molecular structure information representing partial structures. For example, these partial structure data may be data that identify a molecular structure using numbers, letters, text, vectors, etc., or data that visualize it using two-dimensional coordinates, three-dimensional coordinates, etc., or data that are any combination of two or more of these data. The individual numerical values that make up the partial structure data may be represented in decimal notation, or may be represented by other notations such as binary or hexadecimal notation. More specifically, these partial structure data may be a structural formula, a molecular graph, data in SMILES (Simplified Molecular Input Line Entry System) notation, data in MOL file format, etc.
[0020] FIG. 3 shows an example of a partial structure identified by the first partial structure data, part (a) shows an example of that partial structure, and part (b) shows another example of that partial structure. Thus, the partial structure data of each raw material may include data that identify a plurality of partial structures. The first partial structure data is data that can identify one or more partial structures in the molecule of the first raw material. Similarly, the second partial structure data is data that can identify one or more partial structures in the molecule of the second raw material.
[0021] In addition, the acquisition unit 11 may acquire, as blending ratio data representing the ratio r of a plurality of raw materials, data indicating the ratio itself of each raw material, data indicating the ratio between a plurality of raw materials, or data indicating the blending amount (weight, volume, etc.) of each of the plurality of raw materials in absolute or relative values. For example, the ratio r1 = "0.5" of the first monomer, which is the first raw material, and the ratio r2 = "0.5" of the second monomer, which is the second raw material, are acquired.
[0022] Furthermore, the acquisition unit 11 acquires, as count data regarding the number of known partial structures in the molecules of a plurality of raw materials, for example, data representing the number "1" of the partial structure shown in part (a) of FIG. 3 in the molecule of the first raw material and the number "1" of the partial structure shown in part (b) of FIG. 3 in the first raw material. Similarly, the acquisition unit 11 acquires data representing the number of partial structures in the molecules of each of the plurality of raw materials. Note that the count data acquired here may be data representing a count normalized so that the total number of partial structures in the molecule of the same raw material is "1".
[0023] Based on the partial structure data, blending ratio data, and count data acquired by the acquisition unit 11, the generation unit 12 generates partial structure input data D0 by combining, for each partial structure of the plurality of raw materials, the partial structure data regarding the partial structure contained in the raw material, the data of the ratio r regarding the raw material, and the data of the number regarding the partial structure. Then, the generation unit 12 repeats the generation of the partial structure input data D0 for all the partial structure data of each of the plurality of raw materials.
[0024] The vector conversion unit 13 converts each of all the partial structure input data D0 generated by the generation unit 12 into one vector data. For example, the vector conversion unit 13 refers to each partial structure data included in the partial structure input data D0 and molecularly describes them to obtain the vector V MConvert it. By molecular description, the characteristics of the molecule indicated by the substructure data can be represented as a numerical sequence based on its chemical structure. As this molecular description method, any method can be adopted as long as it is a method for vectorizing the molecular structure. For example, ECFP (Extended Connectivity FingerPrints), MACCS FingerPrints, PubChem FingerPrints, Substructure FingerPrints, Estate FingerPrints, BCI FingerPrints, Molprint2D FingerPrints, Pass base FingerPrints, etc. can be adopted. Furthermore, the vector conversion unit 13 combines the data of the ratio r regarding the raw material containing the corresponding substructure and the data of the number of the corresponding substructures in the raw material for each of the respective substructure data to generate substructure input data D for the vector V M corresponding thereto.
[0025] The synthesis unit 14 combines the vectors V for all substructures for each of a plurality of raw materials, which are converted into vectors by the vector conversion unit 13 M to generate synthesis input data F as one vector data. For example, when there are two substructure input data D 1,1 , D 1,2 corresponding to the first raw material and two substructure input data D 2,1 , D 2,2 corresponding to the second raw material, then the synthesis unit 14 generates synthesis input data F by combining the four vectors V 1,1 , D 1,2 , D 2,1 , D 2,2 corresponding thereto. M
[0026] At this time, the synthesis unit 14 combines the four vectors V 1,1 , D 1,2 , D 2,1 , D 2,2 corresponding to the four substructure input data D MBy reflecting the blending ratio data and the number data corresponding to each partial structure, four vectors V M are weighted-averaged to generate the composite input data F. More specifically, the combining unit 14 multiplies each element of the four vectors V 1,1 , D 1,2 , D 2,1 , D 2,2 corresponding to D M by the ratio r and the number n corresponding to each partial structure, and then adds (or averages) each element of the four vectors V M to generate the composite input data F. For example, for the vector V M corresponding to the partial structure of the first raw material, the value obtained by multiplying the ratio r1 of the first raw material by the number n of that partial structure is multiplied, and for the vector V M corresponding to the partial structure of the second raw material, the value obtained by multiplying the ratio r2 of the second raw material by the number n of that partial structure is multiplied. However, the reflection of the ratio and number data may be performed by adding the value obtained by multiplying the ratio r by the number n to each element of the vector V M , or by concatenating the value obtained by multiplying the ratio r by the number n to each element. More generally, the combining unit 14 may generate a vector using a function that takes as input vectors, blending ratios, and numbers related to all partial structures and outputs a single vector according to a certain rule, or may generate a single vector in a single process without separating the steps of reflecting the blending ratio and adding the vectors.
[0027] Furthermore, the synthesizing unit 14 inputs the generated synthesized input data F to an external machine learning model by outputting it externally. That is, the output synthesized input data F is read by a training unit 20 in a computer connected to the outside of the input data generation system 10. Then, in the training unit 20, the synthesized input data F is input to the machine learning model together with an arbitrary teacher label as an explanatory variable, whereby a learned model is generated. Further, a machine learning model in the predictor 30 is set based on the learned model generated by the training unit 20. However, the training unit 20 and the predictor 30 may be the same functional unit. Then, when the synthesized input data F generated by the input data generation system 10 is input to the machine learning model in the predictor 30, the predictor 30 generates and outputs a prediction result of the characteristics of the multi-component substance. Note that these training unit 20 and predictor 30 may be configured in the same computer as the computer 100 constituting the input data generation system 10, or may be configured in a computer separate from the computer 100.
[0028] In one example, the training unit 20 generates a learned model using a neural network. The learned model is generated by a computer processing teacher data including a large number of combinations of input data and output data. The computer calculates output data by inputting the input data to the machine learning model, and obtains an error between the calculated output data and the output data indicated by the teacher data (that is, the difference between the estimation result and the correct answer). Then, the computer updates a given parameter of the neural network, which is the machine learning model, based on the error. The computer generates a learned model by repeating such learning. The process of generating the learned model can be called a learning phase, and the process of the predictor 30 using the learned model can be called an operation phase.
[0029] [Operation of the System] While explaining the operation of the characteristic prediction system 1 with reference to FIG. 4, the characteristic prediction method according to the present embodiment will also be described. FIG. 4 is a flowchart showing an example of the operation of the characteristic prediction system 1.
[0030] First, when the input data generation process is started triggered by the instruction input of the user of the input data generation system 10, the acquisition unit 11 acquires partial structure data, blending ratio data, and quantity data for each of a plurality of raw materials (step S1). Next, based on each partial structure data, the generation unit 12 generates partial structure input data D0 for each partial structure of the plurality of raw materials (step S2). Then, by the vector conversion unit 13, each of all the partial structure input data D0 is converted into a vector V in a vector format, and data on the ratio r regarding the raw material including the corresponding partial structure and data on the number of the corresponding partial structures in the raw material are combined with the vector V M to generate partial structure input data D (step S3). M
[0031] Next, by the synthesis unit 14, the vectors V corresponding to all the partial structure input data D for each of the plurality of raw materials are combined to generate synthesis input data F (step S4). At this time, by the synthesis unit 14, while reflecting the blending ratio data and the quantity data in each vector V, the weighted average of the vectors V is calculated to generate the synthesis input data F. Then, by the synthesis unit 14, the synthesis input data F is output as input data for machine learning to the training unit 20 (step S5). At this time, the reflection of the ratio and the quantity on the vector V M M is such that for each vector V M M M The value obtained by multiplying the ratio and the number is obtained by multiplication, addition, or concatenation. More generally, the synthesizing unit 14 may generate a vector using a function that outputs a single vector according to a certain rule, taking as input the vectors, mixing ratios, and numbers for all substructures, or may generate a single vector in one process without separating the steps of reflecting the mixing ratio and adding the vectors.
[0032] Next, in the training unit 20, a learning phase is executed, and a learned model is generated by learning using input data and teacher data (step S6). Then, the generated learned model is set in the predictor 30, and an operation phase is executed using the input data newly acquired from the input data generation system 10 by the predictor 30, and prediction results of the characteristics of the multi-component substance are generated and output (step S7).
[0033] [Program] A characteristic prediction program for causing a computer or a computer system to function as the characteristic prediction system 1 includes program codes for causing the computer system to function as the acquisition unit 11, the generation unit 12, the vector conversion unit 13, the synthesizing unit 14, the training unit 20, and the predictor 30. This characteristic prediction program may be provided after being fixedly recorded on a tangible recording medium such as a CD-ROM, a DVD-ROM, or a semiconductor memory. Alternatively, the characteristic prediction program may be provided via a communication network as a data signal superimposed on a carrier wave. The provided characteristic prediction program is stored in, for example, the auxiliary storage unit 103. By the processor 101 reading out and executing the characteristic prediction program from the auxiliary storage unit 103, each of the above functional elements is realized.
[0034] [Effect] As described above, according to the above embodiment, based on the partial structure data for each of a plurality of raw materials, partial structure input data D representing known partial structures is generated. The ratio regarding the plurality of raw materials is reflected in the partial structure input data D for the plurality of raw materials, and the partial structure input data D for each of the plurality of raw materials is combined into synthesis input data F and input into a machine learning model. As a result, for a multi-component substance manufactured based on a plurality of raw materials where only a part of the chemical structure is clear, by learning partial structure information through machine learning, the properties of the substance can be efficiently predicted.
[0035] Also, in the above embodiment, partial structure data for specifying known partial structures in the molecules constituting each of the plurality of raw materials and the number of known partial structures in the molecules are received, and for the partial structure input data D for each of the plurality of raw materials, by reflecting a value obtained by multiplying the blending ratio data for each of the plurality of raw materials by the number of known partial structures, input data to be input into a machine learning model is generated. In that case, when the partial structure and its number in the molecule of the raw material are clear, the ratio of the raw material and the number of partial structures can be reflected in the partial structure input data D. As a result, the properties of the multi-component substance based on the raw material can be accurately predicted.
[0036] Also, in the above embodiment, the partial structure input data D0 is generated as molecular structure data of known partial structures. As a result, the partial structure input data D0 can be efficiently generated.
[0037] Furthermore, in the above embodiment, for the plurality of vectors included in the partial structure input data D for each of the plurality of raw materials, a value based on the blending ratio data for each of the plurality of raw materials is multiplied, added, or concatenated, and the multiplied, added, or concatenated plurality of vectors are combined into one vector to generate input data to be input into a machine learning model. Thereby, the ratio of the raw material can be effectively and simply reflected in the partial structure input data D for each raw material. As a result, the prediction accuracy of the properties of the multi-component substance is improved.
[0038] [Modification Example] As described above, the present invention has been described in detail based on its embodiments. However, the present invention is not limited to the above embodiments. The present invention can be variously modified without departing from its gist.
[0039] In the above embodiment, an example in which the input data generation system 10 combines the partial structure input data D of two raw materials to generate the combined input data F has been shown, but it may function to combine the partial structure input data D of three or more raw materials together with their ratios.
[0040] Also, the certain conversion rule provided in the vector conversion unit 13 of the input data generation system 10 may be other rules.
[0041] FIG. 5 shows the configuration of the characteristic prediction system 1A according to a modified example. In this modified example, the functions of the vector conversion unit 13, the synthesis unit 14, and the predictor 30 are provided in the training unit 20A. And in the training unit 20A, the functions of vectorization, data aggregation, and reflection of blending ratio data and number data in these functional units are realized using a machine learning model of a neural network integrated with the predictor 30. At that time, the training unit 20A inputs the partial structure input data D0 generated by the generation unit 12 of the input data generation system 10A into the machine learning model.
[0042] Here, the training unit 20A inputs the partial structure input data D0 into a machine learning model using a neural network. The partial structure data included in the input partial structure input data D0 is data such as a structural formula, a molecular graph, data in SMILES notation, and data of three-dimensional coordinates. Hereinafter, an example of inputting a molecular graph as the partial structure input data D0 into the machine learning model will be described.
[0043] That is, in the above modification example, the generation unit 12 refers to the molecular graph data for each partial structure, and generates a set of node vectors FV that corresponds one-to-one with the node set V in the molecular graph G = (V, E), and a set of edge vectors FE that corresponds one-to-one with the edge set E in the molecular graph. The node vector is a vector for identifying the atoms of the node. For example, it is a vector element in which numerical values (such as atomic numbers, electronegativities, etc.) representing the characteristics of the atoms constituting the nodes of each element of the set are arranged in order. The edge vector is a vector for identifying the nature of the bonds between nodes. For example, it is a vector element in which numerical values (such as bond order, bond distance, etc.) representing the characteristics of the edges of each element of the set are arranged in order. Further, the generation unit 12 combines the set of node vectors FV and the set of edge vectors FE with the original molecular graph data, the data of the ratio r regarding the raw material containing the corresponding partial structure, and the data of the number of the corresponding partial structures in the raw material to generate the partial structure input data D0. Then, the generation unit 12 repeats the generation of the partial structure input data D0 for all the partial structure data for each of the plurality of raw materials. Further, the generation unit 12 may generate the partial structure input data D0 by combining only the data of the ratio r and the data of the number for the partial structure data that is a molecular graph.
[0044] FIG. 6 shows a specific example of the partial structure input data corresponding to the first raw material generated by the generation unit 12. In the (a) part of FIG. 6, the partial structure input data generated for the partial structure shown in the (a) part of FIG. 3 is shown, and in the (b) part of FIG. 6, the partial structure input data generated for the partial structure shown in the (b) part of FIG. 3 is shown. Thus, when the partial structure shown in the (a) part of FIG. 3 is targeted, the generation unit 12 has the set of node vectors FV 1,1 ={FC α , FC β , FC γ}, and the set of edge vectors FE 1,1 ={FC α C β , FC β C γ}, and generates these data, and for these data, the molecular graph G which is the original partial structure data 1,1 =(V 1,1,E 1,1 ) and the data of the ratio r1 for the first raw material and the number n of the corresponding molecular structures in the molecule of the first raw material 1,1 to generate the partial structure input data D0 combined with the data 1,1 Each element within the parentheses denoted by the initial letter F in the description of the above vector sets FV and FE means the vector element after the vector conversion as described above (the same applies hereinafter). For example, for the node vector set FV 1,1 The element FC α is the vector element after being converted into a numerical value representing the characteristics of the molecule C α that constitutes the node. Also, when the partial structure shown in part (b) of FIG. 3 is targeted, the generation unit 12 sets the node vector set FV 1,2 = {FC δ , FC ε , FC ζ , FC η} and the edge vector set FE 1,2 = {FC δ C ε , FC ε C ζ , FC ζ C η} are generated, and for these data, the molecular graph G 1,2 = (V 1,2 , E 1,2 ) which is the original partial structure data, the data of the ratio r1 for the first raw material, and the data of the number n 1,2 of the corresponding molecular structures in the molecule of the first raw material are combined to generate the partial structure input data D0 1,2 . Further, the generation unit 12 generates partial structure input data D0 2,1 , D0 2,2 ,... for the partial structure corresponding to the second raw material
[0045] The training unit 20A receives all the partial structure input data D0 for each of the plurality of raw materials, and generates composite input data F, which is a single vector, based on those partial structure input data D0 by means of a vector conversion unit 13 and a synthesis unit 14 realized by a machine learning model using a neural network. Then, the composite input data F is input to a predictor 30 within the same machine learning model. Note that, in this modification example, the function of the predictor 30 may be separated from the machine learning model of the training unit 20A and realized by another machine learning model.
[0046] Furthermore, the input data generation system 10 of the above-described embodiment may have a function of merging and expressing the partial structure data, the blending ratio data, and the number data for each partial structure in one molecular graph. In this case, the partial structure input data D0 is generated in the same manner as the generation unit 12 of the property prediction system 1A according to the above-described modification example. Then, the blending ratio data and the number data are reflected in the molecular graph included in the partial structure input data D0 for each partial structure. Specifically, the blending ratio data and the number data are reflected in the feature vectors included in the node vector set FV and the edge vector set FE included in the partial structure input data D0. Further, the molecular graphs included in the partial structure input data D0 for each partial structure in which the blending ratio is reflected are merged into one molecular graph data, thereby generating the composite input data F. In such a case, the input data generation system 10 realizes property prediction by inputting the generated molecular graph for each multi-component material into a machine learning model that can input a molecular graph and generate prediction data.
[0047] Moreover, when the vector conversion unit 13 of the input data generation system 10 of the above-described embodiment converts the partial structure input data D for each of the plurality of raw materials into a one-dimensional vector, it may be configured to reflect a value representing the difference in raw materials in the vector. For example, a vector representing the difference in raw materials by a one-hot vector may be concatenated to the vector V M or a vector representing the difference in raw materials using a distributed representation may be concatenated to the vector V MIt may be connected to. As a result, even when the partial structures are the same, the difference between raw materials can be reflected in the partial structure input data D for each raw material. Consequently, the prediction accuracy of the properties of the multi-component substance is further improved. Also, a vector arranging the blending amounts of a group of raw materials with all unknown structures may be connected to the synthetic input data F. Thereby, even when including raw materials with unknown partial structures, the properties can be predicted.
[0048] On the other hand, when using a neural network that takes a graph as shown in the above modification example as input, the input data generation system 10 of the above embodiment may concatenate the vector representing the difference in raw materials to the node vector of each raw material generated by the generation unit 12 when adding the vector. Note that the vector conversion unit 13 of the property prediction system 1A according to the above modification example may be realized by a machine learning model of a neural network. In that case, it may operate to reflect the vector representing the difference in raw materials in the vector of the intermediate layer of the neural network.
[0049] The processing procedure of the input data generation method executed by at least one processor is not limited to the example in the above embodiment. For example, some of the above-described steps (processes) may be omitted, or the steps may be executed in a different order. Also, any two or more of the above-described steps may be combined, or a part of the steps may be modified or deleted. Alternatively, other steps may be executed in addition to the above steps. For example, the processing of steps S6 and S7 may be omitted.
[0050] In the present disclosure, the expression "at least one processor executes the first process, executes the second process,... executes the nth process." or a corresponding expression indicates a concept including the case where the execution subject (i.e., the processor) of the n processes from the first process to the nth process changes midway. That is, this expression indicates a concept including both the case where all of the n processes are executed by the same processor and the case where the processor changes arbitrarily in the n processes.
Description of Signs
[0051] 1... Characteristic prediction system, 10... Input data generation system, 100... Computer, 101... Processor, 11... Acquisition unit, 12... Generation unit, 13... Vector conversion unit, 14... Synthesis unit, 20... Training unit, 30... Predictor.
Claims
1. A property prediction system for predicting the properties of a material based on a plurality of raw materials including known partial structures, comprising at least one processor, wherein the at least one processor receives at least an input of partial structure data that identifies the known partial structures in the molecules constituting each of the plurality of raw materials, blending ratio data that represents the blending ratio of each of the plurality of raw materials, and the number of the known partial structures in the molecule, generates partial structure input data representing the known partial structures based on the partial structure data for each of the plurality of raw materials, reflects a value obtained by multiplying the number of the known partial structures by the blending ratio data for the plurality of raw materials on the partial structure input data for the plurality of raw materials, and then generates combined input data by combining the partial structure input data for the plurality of raw materials, inputs the combined input data into a machine learning model, Property prediction system.
2. The partial structure input data is molecular structure information representing the structure of the known partial structure, The property prediction system according to claim 1.
3. The at least one processor multiplies, adds, or concatenates a value based on the blending ratio data for each of the plurality of raw materials to a plurality of vectors based on the partial structure input data for each of the plurality of raw materials, combines the multiplied, added, or concatenated plurality of vectors into one vector, and inputs the one vector into the machine learning model, The property prediction system according to claim 1 or 2.
4. A property prediction system for predicting the properties of a material based on a plurality of raw materials including known partial structures, comprising at least one processor, wherein the at least one processor receives at least an input of partial structure data that identifies the known partial structures of each of the plurality of raw materials and blending ratio data that represents the blending ratio of each of the plurality of raw materials, generates partial structure input data representing the known partial structures based on the partial structure data for each of the plurality of raw materials, reflects the blending ratio data for the plurality of raw materials in the partial structure input data for the plurality of raw materials, further reflects a value representing the difference between the raw materials in the data which is the partial structure input data for each of the plurality of raw materials in which the blending ratio data is reflected, combines them into one data, and inputs the one data into a machine learning model, Property prediction system.
5. A property prediction method executed by a computer having at least one processor, for predicting properties of a material based on a plurality of raw materials including known partial structures, receiving at least an input of partial structure data identifying the known partial structures in the molecules constituting each of the plurality of raw materials, composition ratio data representing the composition ratio of each of the plurality of raw materials, and the number of the known partial structures in the molecule; generating partial structure input data representing the known partial structures based on the partial structure data for each of the plurality of raw materials; generating combined input data by combining the partial structure input data of the plurality of raw materials after reflecting a value obtained by multiplying the composition ratio data regarding the plurality of raw materials by the number of the known partial structures; inputting the combined input data into a machine learning model; comprising a property prediction method.
6. A property prediction method executed by a computer having at least one processor, for predicting properties of a material based on a plurality of raw materials including known partial structures, receiving at least an input of partial structure data identifying the known partial structures of each of the plurality of raw materials and composition ratio data representing the composition ratio of each of the plurality of raw materials; generating partial structure input data representing the known partial structures based on the partial structure data for each of the plurality of raw materials; reflecting the composition ratio data regarding the plurality of raw materials in the partial structure input data of the plurality of raw materials; further reflecting a value representing the difference between the raw materials in the data which is the partial structure input data for each of the plurality of raw materials in which the composition ratio data is reflected, combining them into one data, and inputting the one data into a machine learning model; comprising a property prediction method.
7. A property prediction program for predicting properties of a material based on a plurality of raw materials including known partial structures, causing a computer to receive at least an input of partial structure data identifying the known partial structures in the molecules constituting each of the plurality of raw materials, composition ratio data representing the composition ratio of each of the plurality of raw materials, and the number of the known partial structures in the molecule; generating substructure input data representing the known substructures based on the substructure data for each of the plurality of raw materials; for the substructure input data of the plurality of raw materials, after reflecting a value obtained by multiplying the number of the known substructures by the blending ratio data regarding the plurality of raw materials, generating combined input data by combining the substructure input data of the plurality of raw materials; inputting the combined input data into a machine learning model; executing; a property prediction program. **Claim 8** A property prediction program for predicting properties of a material based on a plurality of raw materials including known substructures, causing a computer to at least receive inputs of substructure data identifying known substructures of each of the plurality of raw materials and blending ratio data representing the blending ratio of each of the plurality of raw materials; generating substructure input data representing the known substructures based on the substructure data for each of the plurality of raw materials; reflecting the blending ratio data regarding the plurality of raw materials in the substructure input data of the plurality of raw materials; further reflecting a value representing the difference between the raw materials in the data which is the substructure input data for each of the plurality of raw materials in which the blending ratio data is reflected, combining them into one data, and inputting the one data into a machine learning model; executing; a property prediction program.
Citation Information
Patent Citations
Connectivity prediction method, apparatus, program, recording medium, and production method of machine learning algorithm
JP2019028879A