Characteristic prediction system, characteristic prediction method, and characteristic prediction program

The property prediction system enhances material property prediction by generating input data from known substructures and blending ratios, addressing inefficiencies in existing methods and improving accuracy.

JP7732224B2Active Publication Date: 2025-09-02RESONAC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021073164
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-23
Publication Date
2025-09-02
Estimated Expiration
2041-04-23

AI Technical Summary

Technical Problem

Existing methods struggle to efficiently predict material properties using machine learning when input data representing the structure of materials is limited or small.

Method used

A property prediction system that generates input data by identifying and reflecting blending ratios of known substructures within the structure of raw materials, using a machine learning model to predict material properties.

Benefits of technology

Enables efficient prediction of material properties by focusing on predetermined substructures, improving prediction accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007732224000001
    Figure 0007732224000001
  • Figure 0007732224000002
    Figure 0007732224000002
  • Figure 0007732224000003
    Figure 0007732224000003
Patent Text Reader

Abstract

To cause the characteristic of a material to be efficiently predicted on the basis of the structure of raw materials.SOLUTION: An input data generation system 10 generates input data for machine learning to predict the characteristic of a material based on raw materials with each known structure. The input data generation system includes at least one processor. The at least one processor is configured to: acquire partial structure data that represents partial structures from a database; accept at least inputs of partial structure data that identifies the structures of raw materials and mixture ratio data that represents mixture ratios of the raw materials; generate partial structure input data D that represents the partial structure existing in the structure of raw materials, on the basis of the partial structure data and raw material structure data; cause the mixture ratio data to be reflected in the partial structure input data D of raw materials so as to generate input data; and cause the input data to be inputted into a machine learning model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One aspect of the present disclosure relates to a characteristic prediction system, a characteristic prediction method, and a characteristic prediction program. [Background technology]

[0002] Conventionally, molecular structures are acquired in a predetermined format, converted into vector information, and input into a machine learning algorithm to predict properties. For example, a method for predicting the binding affinity between the 3D structure of a biopolymer and the 3D structure of a compound using machine learning is known (see Patent Document 1 below). In this method, a predicted 3D structure of a complex between the biopolymer and the compound is generated based on the 3D structure of the biopolymer and the 3D structure of the compound, the predicted 3D structure is converted into a predicted 3D structure vector, and the predicted 3D structure vector is discriminated using a machine learning algorithm to predict the binding affinity between the 3D structure of the biopolymer and the 3D structure of the compound. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-28879 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, a technology has been developed that predicts material properties using neural networks that input data representing the structure of materials with known structures, such as molecular graphs. However, when the input data is small, it is difficult to predict material properties using machine learning based on that data. Therefore, a system that can efficiently predict material properties based on the structure of raw materials is desired. [Means for solving the problem]

[0005] A property prediction system according to one embodiment of the present disclosure is a property prediction system that predicts the properties of a material based on a raw material of a known structure, and includes at least one processor. The at least one processor acquires partial structure data representing a partial structure from a database, accepts input of at least raw material structure data that identifies the structure of the raw material and blending ratio data that represents the blending ratio of the raw materials, generates partial structure input data representing a partial structure present in the structure of the raw material based on the partial structure data and the raw material structure data, generates input data by reflecting the blending ratio data in the partial structure input data of the raw material, and inputs the input data into a machine learning model.

[0006] Alternatively, another form of the property prediction method of the present disclosure is a property prediction method executed by a computer having at least one processor, which predicts the properties of a material based on a raw material of known structure, and includes the steps of: acquiring substructure data representing a substructure from a database; receiving input of at least raw material structure data that identifies the structure of the raw material and blending ratio data that represents the blending ratio of the raw materials; generating substructure input data representing a substructure present in the structure of the raw material based on the substructure data and the raw material structure data; generating input data by reflecting the blending ratio data in the partial structure input data of the raw material; and inputting the input data into a machine learning model.

[0007] Alternatively, another form of property prediction program of the present disclosure is a property prediction program that predicts the properties of a material based on a raw material of a known structure, and causes a computer to execute the following steps: acquiring partial structure data representing the partial structure from a database; accepting input of at least raw material structure data that identifies the structure of the raw material and mixture ratio data that represents the mixture ratio of the raw materials; generating partial structure input data representing the partial structure present in the structure of the raw material based on the partial structure data and the raw material structure data; generating input data by reflecting the mixture ratio data in the partial structure input data of the raw material; and inputting the input data into a machine learning model.

[0008] According to the above-described embodiment, substructure input data representing known substructures in the structure of the raw material is generated based on the structural data of the raw material and the structural data of the substructures acquired from the database, and input data is generated by reflecting the blending ratio of the raw material in the substructure input data. The generated input data is then input to a machine learning model. As a result, input data narrowed down to data on predetermined substructures from the entire structure of the raw material is generated, and the input data is processed by machine learning, allowing the properties of the material to be efficiently predicted. [Effects of the Invention]

[0009] According to aspects of the present disclosure, material properties can be efficiently predicted based on the structure of the raw material. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 2 is a diagram illustrating an example of a hardware configuration of a computer that configures the characteristic prediction system according to the embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of a functional configuration of a characteristic prediction system according to an embodiment. [Figure 3] 3 is a diagram showing an example of a molecular structure specified by raw material structure data acquired by the acquisition unit 11 of FIG. 2. FIG. [Figure 4] 4 is a flowchart illustrating an example of an operation of the characteristic prediction system according to the embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of a functional configuration of a characteristic prediction system according to a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the description, the same elements or elements having the same functions will be denoted by the same reference numerals, and redundant description will be omitted.

[0012] [System Overview] The property prediction system 1 according to the embodiment is a computer system that uses a machine learning model to perform a process for predicting the properties of a multi-component substance, which is a material produced by blending multiple raw materials with known structures in various ratios. The raw materials refer to chemical substances with known molecular structures used to produce the multi-component substance, such as monomers, polymers, or single molecules such as low-molecular-weight additives, solute molecules, and gas molecules. A single raw material may contain multiple types of molecules. A multi-component substance is a chemical substance produced by blending multiple raw materials in a predetermined ratio, such as a polymer alloy when the raw materials are monomers or polymers, a mixed solution when the raw materials are solute molecules or solvents, or a mixed gas when the raw materials are gas molecules. However, the input data to be generated does not necessarily have to be a multi-component substance; it may also be a single-component substance produced from a single raw material.

[0013] The property prediction system 1 performs prediction processing on the properties of a multi-component substance. For example, if the multi-component substance is a resin, the properties of the multi-component substance include thermal properties such as glass transition temperature and melting point, mechanical properties, and adhesive properties. For other types of substances, the properties of the multi-component substance include the efficacy or toxicity of a drug, the hazardous properties of a flammable substance such as the ignition point, appearance characteristics, and suitability for a specific application. Machine learning is used for the prediction processing of the property prediction system 1. Machine learning is a technique for autonomously finding laws or rules based on given information. The specific machine learning technique is not limited. For example, the machine learning may be machine learning using a machine learning model, which is a computational model. More specifically, the computational model is a neural network. A neural network is an information processing model that mimics the mechanisms of the human nervous system. More specific examples of computational models other than neural networks include support vector regression (SVR) and random forests.

[0014] [System Configuration] The characteristic prediction system 1 is composed of one or more computers. When multiple computers are used, these computers are connected via a communication network such as the Internet or an intranet, thereby logically constructing a single characteristic prediction system 1.

[0015] 1 is a diagram showing an example of a general hardware configuration of a computer 100 constituting the characteristic prediction system 1. For example, the computer 100 includes a processor (e.g., a CPU) 101 that executes an operating system, application programs, etc., a main memory unit 102 consisting of ROM and RAM, an auxiliary memory unit 103 consisting of a hard disk, flash memory, etc., a communication control unit 104 consisting of a network card or a wireless communication module, an input device 105 such as a keyboard, mouse, or touch panel, and an output device 106 such as a monitor or touch panel display.

[0016] Each functional element of the characteristic prediction system 1 is realized by loading a predetermined program onto the processor 101 or the main memory unit 102 and having the processor 101 execute the program. In accordance with the program, the processor 101 operates the communication control unit 104, the input device 105, or the output device 106, and reads and writes data from and to the main memory unit 102 or the auxiliary memory unit 103. Data or a database required for processing is stored in the main memory unit 102 or the auxiliary memory unit 103.

[0017] 2 is a diagram showing an example of the functional configuration of the property prediction system 1. The property prediction system 1 includes an input data generation system 10, a training unit 20, and a predictor 30. The input data generation system 10, the training unit 20, and the predictor 30 may be constructed on the same computer 100, or some of them may be constructed on different computers 100. First, the functional configuration of the input data generation system 10 will be described. The input data generation system 10 includes an acquisition unit 11, a generation unit 12, a vector conversion unit 13, a synthesis unit 14, and a substructure database 16 as functional elements.

[0018] The substructure database 16 is a data storage means for storing substructure data representing combinations of multiple substructures in known molecules previously registered in the input data generation system 10. The substructure data stored in the substructure database 16 is preferably data (molecular structure information) representing known substructures that affect the properties of a material. Examples of such substructures include structures represented by molecular formulas such as "OH," "CH," "CCC," and "NH." The substructure data may be set in advance from a database within the input data generation system 10, representing a combination of multiple substructures selected in response to a selection input by a user of the input data generation system 10, or may be set in response to a user selection from an external computer or the like. The data representing the substructures in the substructure data (molecular structure information) may be a molecular graph, data representing the molecular structure in text such as a molecular formula, or data representing the molecular structure in an image such as a structural formula. More specifically, the substructure data may be structural formulas, molecular graphs, data in SMILES (Simplified Molecular Input Line Entry System) notation, data in MOL file format, etc.

[0019] However, the substructure data does not necessarily have to be set by user selection. For example, a set of substructures that are predicted to improve the accuracy of prediction for a property to be predicted by the property prediction system 1 (e.g., the glass transition temperature of a polymer) may be determined in advance by an automatic verification process, and the substructure data may be set based on the determined set of substructures. Alternatively, a set of substructures that improves prediction accuracy may be automatically selected in a machine learning process, and the substructure data may be set based on the selected set of substructures. Alternatively, a set of substructures that generally improves prediction accuracy may be set in advance, regardless of the prediction of a specific property for a specific material, and the substructure data may be set based on that set of substructures.

[0020] The acquisition unit 11 is a functional element that receives input of raw material structure data relating to the molecular structures of each of the raw materials that form the basis of the multi-component substance to be predicted, and blending ratio data that represents the blending ratio of each of the multiple raw materials when it is assumed that the multi-component substance is manufactured by blending these multiple raw materials. The acquisition unit 11 may acquire these data from a database within the input data generation system 10 in response to a selection input by a user of the input data generation system 10, or may acquire them from an external computer or the like in response to a selection by the user.

[0021] Specifically, the acquisition unit 11 acquires at least first raw material structure data specifying the molecular structure of a first raw material and second raw material structure data specifying the molecular structure of a second raw material. These raw material structure data are molecular structure information representing molecular structures. For example, these raw material structure data may be data specifying the molecular structure using numbers, letters, text, vectors, etc., or data visualizing the molecular structure using two-dimensional coordinates, three-dimensional coordinates, etc., or data combining any two or more of these data. Individual numerical values ​​constituting the raw material structure data may be expressed in decimal or other notation systems such as binary or hexadecimal. More specifically, these raw material structure data may be structural formulas, molecular graphs, SMILES notation data, MOL file format data, etc. Here, the raw material data does not necessarily have to represent the entire molecule of the raw material, but may represent a partial structure within the molecule of the raw material.

[0022] 3 shows an example of a molecular structure specified by raw material structure data, where part (a) shows an example of a molecular structure specified by first raw material structure data, and part (b) shows an example of a molecular structure specified by second raw material structure data. The first raw material structure data is data that can specify the molecular structure of the first raw material. Similarly, the second partial structure data is data that can specify the molecular structure of the second raw material.

[0023] Furthermore, the acquisition unit 11 may acquire, as the blending ratio data representing the ratio r of the plurality of raw materials, data indicating the ratio of each raw material itself, data indicating the blending ratio between the plurality of raw materials, or data indicating the blending amounts (weight, volume, etc.) of the plurality of raw materials in absolute or relative values. For example, the acquisition unit 11 acquires a ratio r1=“0.5” of the first monomer, which is the first raw material, and a ratio r2=“0.5” of the second monomer, which is the second raw material. Furthermore, the acquisition unit 11 acquires, from the partial structure database 16, partial structure data representing combinations of multiple partial structures in registered known molecules.

[0024] Based on the multiple raw material structure data and substructure data acquired by the acquisition unit 11, the generation unit 12 searches for substructures represented in the substructure data in the molecular structure of the raw material and identifies the number of substructures present in the molecular structure. For example, if the substructures represented in the substructure data include "OH," "CH," "CCC," and "NH" substructures, the generation unit 12 identifies the presence of the substructure "CH" and the substructure "CCC" in the molecular structure represented by the first raw material structure data shown in part (a) of Figure 3, along with their respective numbers of "2" and "1." Similarly, the generation unit 12 identifies the presence of the substructures "OH," "CH," and "CCC" in the molecular structure represented by the second raw material structure data shown in part (b) of Figure 3, along with their respective numbers of "1," "2," and "2." The number of partial structures obtained here may be data representing a number normalized so that the total number of partial structures in the molecules of the same raw material is "1."

[0025] Furthermore, the generation unit 12 generates substructure input data D0 for each substructure contained in the molecules of the multiple raw materials based on the searched substructures. More specifically, based on the molecular structure information corresponding to the substructures searched for by the acquisition unit 11, the generation unit 12 generates substructure input data D0 for each of the multiple raw materials by combining molecular structure information on the substructures contained in the raw material, data on the ratio r for that raw material, and data on the number of that substructure. The generation unit 12 then repeatedly generates substructure input data D0 for all substructures searched for for each of the multiple raw materials.

[0026] The vector conversion unit 13 converts each of all the substructure input data D0 generated by the generation unit 12 into one vector data. For example, the vector conversion unit 13 refers to molecular structure information on each substructure included in the substructure input data D0 and converts it into a molecular description, thereby generating a vector V M By molecular description, the molecular characteristics indicated by the molecular structure information can be expressed as a numerical sequence based on the chemical structure. Any method that vectorizes the molecular structure can be used as the molecular description method, and examples of such methods include ECFP (Extended Connectivity FingerPrints), MACCS FingerPrints, PubChem FingerPrints, Substructure FingerPrints, Estate FingerPrints, BCI FingerPrints, MolPrint2D FingerPrints, and Pass-based FingerPrints. Furthermore, the vector conversion unit 13 converts the vector V generated for each substructure into M For the substructure, the substructure input data D is generated by combining data on the ratio r of raw materials containing the substructure and data on the number of the substructure in the raw material.

[0027] The synthesis unit 14 converts the vector V for each of all partial structures for each of the plurality of raw materials into vectors by the vector conversion unit 13. Minto one vector data to generate synthesized input data F. For example, the synthesis unit 14 synthesizes two partial structure input data D corresponding to the partial structures "CH3" and "CCC" of the first raw material. 1,1 ,D 1,2 and three partial structure input data D corresponding to the partial structures “OH”, “CH3”, and “CCC” of the second raw material. 2,1 ,D 2,2 ,D 2,3 If there are five substructure input data D 1,1 ,D 1,2 ,D 2,1 ,D 2,2 ,D 2,3 Five vectors V corresponding to M The synthesized input data F is generated by combining the above.

[0028] At this time, the synthesis unit 14 synthesizes the five partial structure input data D 1,1 ,D 1,2 ,D 2,1 ,D 2,2 ,D 2,3 Five vectors V corresponding to M By reflecting the compounding ratio data and number data corresponding to each partial structure in M The synthesis unit 14 generates the synthesis input data F by taking a weighted average of the substructure input data D 1,1 ,D 1,2 ,D 2,1 ,D 2,2 ,D 2,3 Five vectors V corresponding to M After multiplying each element of by the ratio r and the number n corresponding to each substructure, five vectors V M By adding (or averaging) the elements of the vector V corresponding to the partial structure of the first raw material, the synthesis unit 14 generates the synthesized input data F. M For the first raw material, the ratio r1 of the first raw material is multiplied by the number n of its partial structures, and a vector V corresponding to the partial structure of the second raw material is obtained. MFor example, the vector V corresponding to the partial structure "CH3" of the first raw material shown in part (a) of FIG. 3 is multiplied by the value obtained by multiplying the ratio r2 of the second raw material by the number n of the partial structures. M For the second raw material, the ratio r1=0.5 is multiplied by the number of partial structures “2” (0.5×2=1.0), and the vector V corresponding to the partial structure “OH” of the second raw material shown in part (b) of FIG. M For the vector V, the ratio r2 = 0.5 is multiplied by the number of the partial structure "1" (0.5 × 1 = 0.5). However, the ratio and number data are reflected in the vector V M This may be done by adding a value obtained by multiplying the ratio r by the number n to each element, or by concatenating values ​​obtained by multiplying the ratio r by the number n to vector elements. More generally, the synthesis unit 14 may generate a vector using a function that receives vectors, compounding ratios, and numbers related to all partial structures as input and outputs a single vector according to a certain rule, or may generate a single vector in a single process without dividing it into steps of reflecting the compounding ratios and adding the vectors.

[0029] Furthermore, the synthesis unit 14 outputs the generated synthetic input data F to an external device, thereby inputting the data to an external machine learning model. That is, the output synthetic input data F is read by a training unit 20 in a computer connected externally to the input data generation system 10. Then, in the training unit 20, the synthetic input data F is input as an explanatory variable to the machine learning model together with an arbitrary teacher label, thereby generating a trained model. Furthermore, a machine learning model in the predictor 30 is set based on the trained model generated by the training unit 20. However, the training unit 20 and the predictor 30 may be the same functional unit. Then, the synthetic input data F generated by the input data generation system 10 is input to the machine learning model in the predictor 30, thereby generating and outputting a prediction result of the properties of the multi-component substance. Note that the training unit 20 and the predictor 30 may be configured in the same computer as the computer 100 constituting the input data generation system 10, or may be configured in a computer separate from the computer 100.

[0030] In one example, the training unit 20 generates a trained model using a neural network. The trained model is generated by a computer processing training data including numerous combinations of input data and output data. The computer calculates output data by inputting the input data into a machine learning model, and determines the error between the calculated output data and the output data indicated by the training data (i.e., the difference between the estimated result and the correct answer). The computer then updates given parameters of the neural network, which is the machine learning model, based on the error. The computer generates the trained model by repeating this type of learning. The process of generating the trained model can be referred to as the training phase, and the process of the predictor 30 using the trained model can be referred to as the operation phase.

[0031] [System Operation] The operation of the characteristic prediction system 1 and the characteristic prediction method according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a flowchart showing an example of the operation of the characteristic prediction system 1.

[0032] First, when the input data generation process is started in response to an instruction input by a user of the input data generation system 10, the acquisition unit 11 acquires raw material structure data and blending ratio data for each of a plurality of raw materials, and acquires substructure data representing a plurality of known substructures from the substructure database 16 (step S1). Next, the generation unit 12 searches for substructures in the molecular structure of the raw material represented by each raw material structure data, thereby identifying the substructures and their number in the molecular structure (step S2). Furthermore, the generation unit 12 generates substructure input data D0 for each substructure in the molecular structure of the plurality of raw materials (step S3). Thereafter, the vector conversion unit 13 converts all of the substructure input data D0 into a vector V in a vector format. M converted to vector V M Then, the data on the ratio r regarding the raw material containing the substructure in question and the data on the number of the substructure in question in the raw material are combined to generate substructure input data D (step S4).

[0033] Next, the synthesis unit 14 synthesizes a vector V corresponding to all the partial structure input data D for each of the plurality of raw materials. M are combined to generate the combined input data F (step S5). M While reflecting the mixture ratio data and number data, vector V M The weighted average of the vectors V is calculated to generate the synthetic input data F. Then, the synthesis unit 14 outputs the synthetic input data F to the training unit 20 as input data for machine learning (step S6). M The ratio and number of vectors V Mis multiplied by the ratio and the number, and then multiplied, added, or concatenated with the result. More generally, the synthesis unit 14 may generate a vector using a function that receives vectors, compounding ratios, and numbers related to all partial structures as input and outputs a single vector according to a certain rule, or may generate a single vector in a single process without dividing it into steps of reflecting the compounding ratios and adding the vectors.

[0034] Next, the training unit 20 executes a learning phase, and generates a trained model by learning using the input data and teacher data (step S7). The generated trained model is then set in the predictor 30, and the predictor 30 executes an operation phase using new input data acquired from the input data generation system 10, and generates and outputs prediction results for the properties of the multi-component substance (step S8).

[0035] [program] The property prediction program for causing a computer or computer system to function as the property prediction system 1 includes program code for causing the computer system to function as an acquisition unit 11, a generation unit 12, a vector conversion unit 13, a synthesis unit 14, a substructure database 16, a training unit 20, and a predictor 30. This property prediction program may be provided by being permanently recorded on a tangible recording medium such as a CD-ROM, a DVD-ROM, or a semiconductor memory. Alternatively, the property prediction program may be provided via a communications network as a data signal superimposed on a carrier wave. The provided property prediction program is stored in, for example, an auxiliary storage unit 103. The processor 101 reads out the property prediction program from the auxiliary storage unit 103 and executes it, thereby realizing each of the above functional elements.

[0036] [effect] As described above, according to the above embodiment, substructure input data D representing known substructures in the structure of the raw material is generated based on the structural data of the raw material and the structural data of the substructures acquired from the database, the ratio of the raw material is reflected in the substructure input data D, and the substructure input data D is compiled to generate input data. The generated input data is then input to a machine learning model. As a result, input data is generated that is narrowed down to data on predetermined substructures from the entire structure of the raw material, and the input data is processed by machine learning, allowing the properties of the material to be efficiently predicted.

[0037] In the above embodiment, raw material structure data for a plurality of raw materials and blending ratio data representing the blending ratios of each of the plurality of raw materials are received as input, and partial structure input data is generated for each of the plurality of raw materials. The blending ratio data for the plurality of raw materials is reflected in the partial structure input data for the plurality of raw materials to generate input data. In this case, partial structure input data D0 representing a known partial structure is generated based on the raw material structure data for each of the plurality of raw materials, and composite input data F is generated by reflecting the ratios of each of the plurality of raw materials in the partial structure input data D0 for each of the plurality of raw materials. The generated composite input data F is then input to a machine learning model. As a result, by processing the input data using machine learning for a multi-component substance manufactured from a plurality of raw materials, the properties of the substance can be efficiently predicted.

[0038] In the above embodiment, the number of partial structures present in the structure of the raw material is identified, and the value obtained by multiplying the blending ratio data by the number of partial structures is reflected in the partial structure input data of the raw material, thereby generating input data. In this case, the number of partial structures in the molecule of the raw material is identified, and the ratio of the raw material and the number of partial structures can be reflected in the partial structure input data D0. As a result, the properties of a multi-component substance based on the raw material can be predicted with high accuracy.

[0039] Furthermore, in the above embodiment, the input data to be input to the machine learning model is generated by multiplying, adding, or concatenating the vectors included in the partial structure input data D by values ​​based on the blending ratio data, and combining the multiplied, added, or concatenated vectors into a single vector. This makes it possible to effectively and easily reflect the ratios of raw materials in the partial structure input data D related to the partial structure of the raw materials. As a result, the prediction accuracy of the properties of multi-component substances is improved.

[0040] [Variations] The present invention has been described in detail above based on the embodiments. However, the present invention is not limited to the above embodiments. Various modifications of the present invention are possible without departing from the spirit and scope of the present invention.

[0041] In the above embodiment, an example was shown in which the input data generation system 10 generates composite input data F by combining the partial structure input data D0 of two raw materials, but the system may also function to combine the partial structure input data D0 of three or more raw materials together with their ratios, or the partial structure input data D of only one raw material may be generated as input data and used to predict the properties of a single-component substance.

[0042] Furthermore, the certain conversion rule provided in the vector conversion unit 13 of the input data generation system 10 may be another rule.

[0043] 5 shows the configuration of a property prediction system 1A according to a modified example. In this modified example, the functions of the vector conversion unit 13, the synthesis unit 14, and the predictor 30 are provided in a training unit 20A. In the training unit 20A, the functions of vectorization, data compilation, and reflection of the mixture ratio data and number data of these functional units are realized using a machine learning model of a neural network integrated with the predictor 30. In this case, the training unit 20A inputs the substructure input data D0 for each substructure generated by the generation unit 12 of the input data generation system 10A to the machine learning model.

[0044] Here, the training unit 20A inputs substructure input data D0 to a machine learning model using a neural network. The substructure information included in the input substructure input data D0 is a structural formula, a molecular graph, SMILES notation data, three-dimensional coordinate data, etc. Below, an example will be described in which a molecular graph is input to the machine learning model as the substructure input data D0.

[0045] That is, in the above-described modified example, the generation unit 12 references molecular graph data, which is molecular structure information for each of the searched substructures, to generate a set FV of node vectors that correspond one-to-one to the node set V in the molecular graph G=(V,E), and a set FE of edge vectors that correspond one-to-one to the edge set E in the molecular graph. The node vector is a vector for identifying the atoms of the node, and is, for example, a vector element in which numerical values ​​(atomic number, electronegativity, etc.) representing the characteristics of the atoms constituting each element node of the set are arranged in order. The edge vector is a vector for identifying the nature of the bond between the nodes, and is, for example, a vector element in which numerical values ​​(bond order, bond distance, etc.) representing the characteristics of each element edge of the set are arranged in order. Furthermore, the generation unit 12 generates substructure input data D0 by combining the node vector set FV and the edge vector set FE with the molecular graph data of the original substructure, data on the ratio r of raw materials containing the corresponding substructure, and data on the number of the corresponding substructures in the raw materials. The generation unit 12 then repeatedly generates the substructure input data D0 for all substructures searched for for each of the multiple raw materials. Alternatively, the generation unit 12 may generate the substructure input data D0 by combining only the data on the ratio r and the data on the number of molecules for the molecular structure information, which is a molecular graph.

[0046] The training unit 20A receives as input all of the substructure input data D0 for each of a plurality of raw materials, and generates composite input data F, which is a single vector, based on the substructure input data D0 using a vector conversion unit 13 and a synthesis unit 14, which are realized by a machine learning model using a neural network. The composite input data F is then input to a predictor 30 within the same machine learning model. Note that in this modification, the function of the predictor 30 may be separated from the machine learning model of the training unit 20A and realized by a separate machine learning model.

[0047] The input data generation system 10 of the above embodiment may also have a function of merging and expressing molecular structure information, blending ratio data, and number data for each substructure into a single molecular graph. In this case, the substructure input data D0 is generated in the same manner as the generation unit 12 of the property prediction system 1A according to the above-described modified example. Then, the blending ratio data and number data are reflected in the molecular graph included in the substructure input data D0 for each substructure. Specifically, the blending ratio data and number data are reflected in the feature vectors included in the node vector set FV and edge vector set FE included in the substructure input data D0. Furthermore, the molecular graphs included in the substructure input data D0 for each substructure, in which the blending ratios have been reflected, are merged into a single molecular graph data to generate composite input data F. In such a case, the input data generation system 10 realizes property prediction by inputting the generated molecular graph for each multi-component material into a machine learning model that can generate predicted data by inputting a molecular graph.

[0048] Furthermore, when the vector conversion unit 13 of the input data generation system 10 of the above embodiment converts the partial structure input data D for each of a plurality of raw materials into a one-dimensional vector, the vector may be configured to reflect a value representing the difference between the raw materials. For example, the vector representing the difference between the raw materials by a one-hot vector may be converted into a vector V M Also, we can use distributed representation to create a vector that represents the difference in ingredients as a vector V MThis allows the differences between raw materials to be reflected in the partial structure input data D for each raw material, even if the raw materials have the same partial structure. As a result, the accuracy of predicting the properties of multi-component substances is further improved. Also, a vector listing the blending amounts of raw materials with unknown molecular structures may be linked to the composite input data F. This makes it possible to predict properties even when raw materials with unknown molecular structures are included.

[0049] On the other hand, when using a neural network that uses a graph as input as shown in the above modification, the input data generation system 10 of the above embodiment may add a vector representing the difference in raw materials by concatenating the vector with the node vector of each raw material generated by the generation unit 12. Note that the vector conversion unit 13 of the property prediction system 1A according to the above modification may be realized by a machine learning model of a neural network. In this case, the vector representing the difference in raw materials may be reflected in the vector in the intermediate layer of the neural network.

[0050] Furthermore, the synthesis unit 14 may select and synthesize data of some substructures from all the substructures for each of the plurality of raw materials searched for by the generation unit 12 according to a predetermined rule, to generate the synthetic input data F. For example, the predetermined rule may be to select some data from data of substructures that are similar to each other.

[0051] Furthermore, in the above embodiment, multiple types of substructure data may be preset in the substructure database. For example, multiple types of substructure data may be set by sequentially deleting one structure from multiple substructures. In this case, the input data generation system 10 has the function of acquiring multiple types of substructure data from the substructure database 16, generating multiple types of input data using the multiple types of substructure data, and inputting the multiple types of input data into multiple machine learning models to construct an ensemble learning module. This allows input data based on the search results for various substructures to be input to multiple learning modules, enabling ensemble learning to be realized in which regression or classification is performed on the prediction results of the multiple learning modules by averaging or majority voting, thereby further improving the prediction accuracy of the properties of multi-component substances.

[0052] In the above embodiment, a feature vector based on data representing the overall molecular structure of the raw material may also be used as the synthetic input data F. For example, such a feature vector may be linked to the synthetic input data F and used for property prediction, or input data based on such a feature vector may be used for ensemble learning.

[0053] The processing procedure of the input data generation method executed by at least one processor is not limited to the example in the above embodiment. For example, some of the steps (processing) described above may be omitted, or the steps may be executed in a different order. Furthermore, any two or more of the steps described above may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be executed in addition to the above steps. For example, the processing of steps S8 and S9 may be omitted.

[0054] In this disclosure, the expression "at least one processor executes a first process, executes a second process, ... executes an nth process" or a corresponding expression indicates a concept including a case where the entity executing the n processes from the first process to the nth process (i.e., the processor) changes midway through. In other words, this expression indicates a concept including both a case where all n processes are executed by the same processor and a case where the processor changes among the n processes according to an arbitrary policy. [Explanation of symbols]

[0055] 1...property prediction system, 10...input data generation system, 100...computer, 101...processor, 11...acquisition unit, 12...generation unit, 13...vector conversion unit, 14...synthesis unit, 16...substructure database, 20...training unit, 30...predictor.

Claims

1. A property prediction system for predicting the properties of a material based on a raw material of known structure, comprising: at least one processor; the at least one processor: Obtaining substructure data representing a substructure from the database; Accepting at least input of raw material structure data that specifies the structure of the raw materials and blending ratio data that indicates the blending ratio of the raw materials; Identifying the number of the partial structures present in the structure of the raw material; generating substructure input data representing the substructure present in the structure of the raw material based on the substructure data and the raw material structure data; generating input data by reflecting a value obtained by multiplying the blending ratio data by the number of the partial structures in the partial structure input data of the raw materials; feeding the input data into a machine learning model; Trait prediction system.

2. The at least one processor Accepting input of the raw material structure data regarding the plurality of raw materials and blending ratio data representing the blending ratio of each of the plurality of raw materials; generating the substructure input data for each of the plurality of raw materials, and generating the input data by reflecting blending ratio data for the plurality of raw materials in the substructure input data for the plurality of raw materials; The property prediction system according to claim 1 .

3. The partial structure input data is molecular structure information representing the structure of the partial structure. The property prediction system according to claim 1 or 2.

4. The at least one processor multiplying, adding, or concatenating a vector based on the partial structure input data by a value based on the blending ratio data, and combining the multiplied, added, or concatenated vectors into one vector, thereby generating the input data; The property prediction system according to any one of claims 1 to 3.

5. A property prediction system for predicting the properties of a material based on a raw material of known structure, comprising: at least one processor; the at least one processor: Obtaining substructure data representing a substructure from the database; Accepting input of raw material structure data relating to the plurality of raw materials and blending ratio data representing the blending ratio of each of the plurality of raw materials; generating substructure input data representing the substructure present in the structure of the raw material for each of a plurality of raw materials based on the substructure data and the raw material structure data; Reflecting blending ratio data regarding the plurality of raw materials in the partial structure input data of the plurality of raw materials; generating input data by further reflecting values ​​representing differences between the raw materials in the plurality of data that are partial structure input data for each of the plurality of raw materials and integrating them into one data; feeding the input data into a machine learning model; Trait prediction system.

6. The at least one processor acquiring a plurality of types of substructure data from the database; generating a plurality of types of the input data using a plurality of types of the partial structure data; constructing an ensemble learning device by inputting the plurality of types of input data into a plurality of machine learning models; The property prediction system according to any one of claims 1 to 5.

7. 1. A property prediction method executed by a computer having at least one processor, for predicting a property of a material based on a raw material of known structure, comprising: obtaining substructure data representing a substructure from a database; a step of receiving at least input of raw material structure data that specifies the structure of the raw materials and blending ratio data that indicates the blending ratio of the raw materials; Identifying the number of the partial structures present in the structure of the raw material; generating substructure input data representing the substructure present in the structure of the raw material based on the substructure data and the raw material structure data; generating input data by reflecting a value obtained by multiplying the blending ratio data by the number of the partial structures in the partial structure input data of the raw material; inputting the input data into a machine learning model; Equipped with Property prediction methods.

8. 1. A property prediction method executed by a computer having at least one processor, for predicting a property of a material based on a raw material of known structure, comprising: obtaining substructure data representing a substructure from a database; receiving input of raw material structure data relating to the plurality of raw materials and blending ratio data representing the blending ratio of each of the plurality of raw materials; generating substructure input data representing the substructure present in the structure of the raw material for each of a plurality of raw materials based on the substructure data and the raw material structure data; a step of reflecting blending ratio data regarding the plurality of raw materials in the partial structure input data of the plurality of raw materials; generating input data by consolidating the plurality of data, which are the partial structure input data for each of the plurality of raw materials, into one data by further reflecting values ​​that represent the differences between the raw materials; inputting the input data into a machine learning model; Equipped with Property prediction methods.

9. A property prediction program for predicting the properties of a material based on a raw material of known structure, On the computer, obtaining substructure data representing a substructure from a database; a step of receiving at least input of raw material structure data that specifies the structure of the raw materials and blending ratio data that indicates the blending ratio of the raw materials; Identifying the number of the partial structures present in the structure of the raw material; generating substructure input data representing the substructure present in the structure of the raw material based on the substructure data and the raw material structure data; generating input data by reflecting a value obtained by multiplying the blending ratio data by the number of the partial structures in the partial structure input data of the raw material; inputting the input data into a machine learning model; Execute Property prediction program.

10. A property prediction program for predicting the properties of a material based on a raw material of known structure, On the computer, obtaining substructure data representing a substructure from a database; receiving input of raw material structure data relating to the plurality of raw materials and blending ratio data representing the blending ratio of each of the plurality of raw materials; generating substructure input data representing the substructure present in the structure of the raw material for each of a plurality of raw materials based on the substructure data and the raw material structure data; a step of reflecting blending ratio data regarding the plurality of raw materials in the partial structure input data of the plurality of raw materials; generating input data by consolidating the plurality of data, which are the partial structure input data for each of the plurality of raw materials, into one data by further reflecting values ​​that represent the differences between the raw materials; inputting the input data into a machine learning model; Execute Property prediction program.

Citation Information

Patent Citations

  • Connectivity prediction method, apparatus, program, recording medium, and production method of machine learning algorithm

    JP2019028879A

  • Synthetic pathway constructing equipment, synthetic pathway constructing method, synthetic pathway constructing program, and processes for manufacturing 3-hydroxypropionic acid, crotonyl alcohol and butadiene

    WO2012081723A1