Input data generation system, input data generation method, and input data generation program

By combining and transforming molecular diagram data of multi-component substances to reflect the mixing ratio, and generating input data for machine learning, the problem of difficulty in predicting the properties of multi-component substances in existing technologies is solved, and high-precision prediction results are achieved.

CN114651309BActive Publication Date: 2025-12-19RESONAC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080077810.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-12
Filing Date
2020-11-10
Publication Date
2025-12-19
Estimated Expiration
2040-11-10

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently predict the properties of multi-component substances, especially since their three-dimensional structures are difficult to know in advance, making it impossible to effectively utilize existing methods for prediction.

Method used

By combining molecular graph data of multiple components, synthetic molecular graph data is generated and transformed into feature vectors that reflect the mixing ratio data. This data is then used as input data for machine learning and processed using neural networks.

Benefits of technology

It enables high-precision prediction of the properties of multi-component substances, especially the properties of polymer alloys formed by mixed monomers, thus improving the accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114651309B_ABST
    Figure CN114651309B_ABST
Patent Text Reader

Abstract

An input data generation system according to an embodiment includes at least one processor configured to accept input of at least first molecule graph data that specifies a molecular graph of a first molecule, second molecule graph data that specifies a molecular graph of a second molecule, and mixture ratio data that indicates a mixture ratio of the first molecule and the second molecule, combine the first molecule graph data and the second molecule graph data, generate synthetic molecule graph data, convert the synthetic molecule graph data into a feature vector, and generate input data for machine learning by reflecting the mixture ratio data in the feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One embodiment of the present application relates to an input data generation system, an input data generation method, and an input data generation program. BACKGROUND

[0002] Conventionally, the structure of a molecule is acquired in a prescribed format, converted into vector information, and input into a machine learning algorithm to predict a property. For example, a method of predicting the bonding of a stereostructure of a biological macromolecule and a stereostructure of a compound using machine learning is known (see Patent Literature 1 below). In this method, a predicted stereostructure of a complex of a biological macromolecule and a compound is generated from the stereostructure of the biological macromolecule and the stereostructure of the compound, the predicted stereostructure is converted into a predicted stereostructure vector, and a machine learning algorithm is used to determine the predicted stereostructure vector, thereby predicting the bonding of the stereostructure of the biological macromolecule and the stereostructure of the compound.

[0003] Prior Art Documents

[0004] Patent Literature

[0005] Patent Literature 1: Japanese Patent Application Publication No. 2019-28879 SUMMARY

[0006] Technical Problem to be Solved by the Invention

[0007] In recent years, a technique of predicting a property of a substance by a neural network with a molecular graph as input is known. However, in this technique, efficient prediction of a property of a multi-component substance in which a plurality of components are mixed at various mixing ratios has not been achieved. Also, regarding a multi-component substance, there is a tendency that it is difficult to know a stereostructure in advance, and therefore it is not possible to predict a property of a multi-component substance using the method of Patent Literature 1 above. Thus, a structure for efficiently predicting a property of a multi-component substance in which a plurality of components are mixed is desired.

[0008] Means for Solving the Technical Problem

[0009] The input data generation system of one embodiment of the present application includes at least one processor that accepts input of at least first molecule graph data that specifies a molecular graph of a first molecule, second molecule graph data that specifies a molecular graph of a second molecule, and mixing ratio data that indicates a mixing ratio of the first molecule and the second molecule, combines at least the first molecule graph data and the second molecule graph data, generates synthetic molecule graph data, converts the synthetic molecule graph data into a feature vector, and generates input data for machine learning by reflecting the mixing ratio data in the feature vector.

[0010] Alternatively, another aspect of the input data generation method of the present invention is executed by a computer having at least one processor, the input data generation method comprising: receiving at least input the first molecular map data for determining a molecular map corresponding to a first molecule, the second molecular map data for determining a molecular map corresponding to a second molecule, and mixing ratio data representing the mixing ratio of the first molecule and the second molecule; combining at least the first molecular map data and the second molecular map data to generate synthetic molecular map data; transforming the synthetic molecular map data into feature vectors; and generating input data for machine learning by reflecting the mixing ratio data in the feature vectors.

[0011] Alternatively, the input data generation program of another aspect of the present invention causes a computer to perform the following steps: at least accepting input of first molecular map data that determines a molecular map corresponding to the first molecule, second molecular map data that determines a molecular map corresponding to the second molecule, and mixing ratio data representing the mixing ratio of the first molecule and the second molecule; at least combining the first molecular map data and the second molecular map data to generate synthetic molecular map data; transforming the synthetic molecular map data into feature vectors; and generating input data for machine learning by reflecting the mixing ratio data in the feature vectors.

[0012] According to the above method, data determining the molecular structure of the first molecule and data determining the molecular structure of the second molecule are combined to generate synthetic molecular map data. This synthetic molecular map data is then transformed into feature vectors, which reflect data representing the mixing ratio of the first and second molecules, generating input data for machine learning. This structure enables the efficient generation of input data related to multi-component substances used as input to a neural network with the molecular map as input. As a result, even for multi-component substances containing multiple components, their properties can be predicted with high accuracy by utilizing neural networks to process the input data.

[0013] Invention Effects

[0014] According to the method of the present invention, the properties of multi-component substances containing multiple components can be predicted with high accuracy. Attached Figure Description

[0015] Figure 1 This is a diagram illustrating an example of the hardware structure of a computer constituting the input data generation system according to the implementation method.

[0016] Figure 2 This is a diagram illustrating an example of the functional structure of the input data generation system involved in the implementation method.

[0017] Figure 3 It means according to Figure 2 An example of a molecular map determined by the molecular map data acquired by the acquisition unit 11.

[0018] Figure 4 is a flowchart showing an example of the action of the input data generation system according to the embodiment. Figure 2 Figure 3 FIG. 1 is a diagram showing an example of a molecular graph of a multi-component substance generated by combining the first molecular graph and the second molecular graph shown in FIGS. 2A and 2B.

[0019] Figure 5 is a flowchart showing an example of the action of the input data generation system according to the embodiment.

[0020] Figure 6 is a diagram showing an example of molecular data processed in the action of the input data generation system according to the embodiment. DETAILED DESCRIPTION

[0021] Hereinafter, the embodiment of the present application will be described in detail with reference to the accompanying drawings. In the description, the same symbols are used for the same elements or elements having the same function, and repetitive description is omitted.

[0022] [Outline of the system]

[0023] The input data generation system 10 according to the embodiment is a computer system that performs generation processing of input data representing a multi-component substance generated by mixing a plurality of components at various mixing ratios. The component refers to a chemical substance having a specific molecular structure used for generating a multi-component substance, such as a monomer, a polymer, or a single molecule of a low-molecular additive, a solute molecule, a gas molecule, and the like. A plurality of molecules can be contained in one component. The multi-component substance refers to a chemical substance generated by mixing a plurality of components at a prescribed mixing ratio, such as a polymer alloy in the case of a component being a monomer, a polymer mixture in the case of a component being a polymer, a mixed solution in the case of a component being a solute molecule or a solvent, and a mixed gas in the case of a component being a gas molecule.

[0024] ​The input data generated by the input data generation system 10 is used as input data for machine learning for predicting a property of a multi-component substance. The property of the multi-component substance, for example, in the case where the multi-component substance is a resin, is a thermal property such as a glass transition temperature and a melting point, a mechanical property, or an adhesive property. Also, in the case where the multi-component substance is another kind of substance, the property of the multi-component substance is a medicinal effect or toxicity of a medicine, a flammability or a risk such as a flash point of a combustible, an appearance property, or suitability for a specific use. Machine learning of inputting the input data refers to a method of autonomously finding a law or a rule by repeatedly learning from given information. The specific method of machine learning is not limited. For example, the machine learning can be machine learning using a machine learning model that is a computational model including a neural network. The neural network refers to an information processing model that simulates the structure of a human brain nervous system. As a more specific example, the machine learning uses at least one of a neural network with a graph as input and a convolutional neural network with a graph as input.

[0025] [Structure of System]

[0026] The input data generation system 10 is constituted by one or more computers. In the case where a plurality of computers are used, the computers are connected via a communication network such as the Internet, an intranet, or the like, and thus one input data generation system 10 is logically constructed.

[0027] Figure 1 is a diagram showing an example of a general hardware structure of the computer 100 that constitutes the input data generation system 10. For example, the computer 100 is provided with a processor (for example, a CPU) 101 that executes an operating system, an application program, or the like, a main storage portion 102 constituted by a ROM and a RAM, an auxiliary storage portion 103 constituted by a hard disk, a flash memory, or the like, a communication control portion 104 constituted by a network card or a wireless communication module, an input device 105 such as a keyboard, a mouse, a touch panel, or the like, and an output device 106 such as a monitor, a touch panel display, or the like.

[0028] Each functional element of the input data generation system 10 is realized by reading a program set in advance into the processor 101 or the main storage portion 102 and causing the processor 101 to execute the program. The processor 101 causes the communication control portion 104, the input device 105, or the output device 106 to act in accordance with the program, and performs reading and writing of data in the main storage portion 102 or the auxiliary storage portion 103. Data or a database required for processing is stored in the main storage portion 102 or the auxiliary storage portion 103.

[0029] Figure 2 is a diagram showing an example of a functional structure of the input data generation system 10. The input data generation system 10 is provided with an acquisition portion 11, a synthesis portion 12, an addition portion 13, a vector conversion portion 14, and a mixing rate reflection portion 15 as functional elements.

[0030] The acquisition unit 11 is a functional unit that accepts input of molecular graph data of a plurality of components and mixing rate data indicating mixing rates of the plurality of components when the plurality of components are mixed to generate a mixture. The acquisition unit 11 can acquire the data from a database in the input data generation system 10 according to selection input of a user of the input data generation system 10, or can acquire the data from an external computer or the like according to selection of the user.

[0031] Specifically, the acquisition unit 11 acquires at least first molecular graph data that determines a molecular graph corresponding to a first molecule included in a first component and second molecular graph data that determines a molecular graph corresponding to a second molecule included in a second component. The molecular graph data is data that determines the structure of an undirected graph in which nodes and edges represent a molecular structure. The molecular graph data can be, for example, data that determines the structure of an undirected graph by numbers, English characters, texts, vectors, or the like, data that visualizes the structure by two-dimensional images, three-dimensional images, or the like, or data that is a combination of any two or more of these data. Each value that constitutes the molecular graph data can be expressed in decimal notation, or can be expressed in binary notation, hexadecimal notation, or the like. In more detail, the acquisition unit 11 acquires at least first molecular graph data that determines a molecular graph of a first monomer that is the first component and second molecular graph data that determines a molecular graph of a second monomer that is the second component.

[0032] In Figure 3 In the (a) part, an example of the structure of the first molecular graph is shown, and in the (b) part, an example of the structure of the second molecular graph is shown. Figure 3 The first molecular graph shown in the (a) part has a structure in which a node N1 of an atom "A" is connected to a node N2 of an atom "B" by an edge E12, and the node N2 is connected to a node N3 of an atom "C" by an edge E23. In the first molecular graph data, node information that determines each node N1 to N3 and edge information that determines each edge E12 and E23 are included. Furthermore, in the first molecular graph, the node N1 and the node N3 are nodes that have a property of being able to be further randomly connected to other nodes. For example, in the case where the first molecular graph is a monomer having a straight chain structure, the nodes N1 and N3 at the ends have a property of being able to be randomly connected. The "able to be randomly connected" referred to here means that connection to other nodes is randomly generated, in other words, a case where connection is generated and a case where connection is not generated. In the case where the first molecular graph has such nodes, connection node information that determines nodes that are able to be further connected (for example, the nodes N1 and N3) is also included in the first molecular graph data. In the connection node information, limiting information that limits the connection destination of the node or the kind (atoms, etc.) of the connection destination node can also be included.

[0033] Similarly,Figure 3 The second molecular graph shown in (b) of FIG. 14 has a structure in which node N4 of atom "D" is combined with node N5 of atom "E" by edge E45, and node N5 is combined with node N6 of atom "F" by edge E56. In the second molecular graph data, node information that specifies each of nodes N4 to N6 and edge information that specifies each of edges E45 and E56 are included. Also, in the second molecular graph, as in the first molecular graph, node N4 and node N6 are nodes that have a property of being able to be further combined with other nodes. In the case where the second molecular graph has such nodes, in the second molecular graph data, combination node information that specifies nodes that are able to be further combined is also included. In the combination node information, information that specifies a combination destination of the node or a kind of node of the combination destination can also be included.

[0034] Also, as the mixing ratio data that represents the mixing ratios r of the plurality of components, the acquisition unit 11 can acquire data that represents the mixing ratios themselves of the respective components, can acquire data that represents mixing ratios between the plurality of components, and can acquire data that represents the amounts (weights, volumes, etc.) of the respective components in absolute values or relative values. For example, the mixing ratio r1 = "0.5" of the first monomer as the first component and the mixing ratio r2 = "0.5" of the second monomer as the second component are acquired.

[0035] The synthesis unit 12 combines the molecular graphs of the plurality of components to generate synthesis molecular graph data that corresponds to the molecular graph of the multi-component substance. Here, the synthesis unit 12 generates synthesis molecular graph data that specifies the molecular graph of the multi-component substance in which the first molecular graph and the second molecular graph are combined, with reference to at least the first molecular graph data and the second molecular graph data. Figure 4 representing a combination Figure 3 An example of the molecular graph of the multi-component substance that is generated from the first molecular graph and the second molecular graph shown in FIG. 15. Thus, the synthesis unit 12 generates synthesis molecular graph data by combining the node information related to nodes N1, N2, and N3 that are specified from the first molecular graph data and the edge information related to edges E12 and E23, and the node information related to nodes N4, N5, and N6 that are specified from the second molecular graph data and the edge information related to edges E45 and E56 as they are. Also, the synthesis unit 12 generates set data V that specifies a set of nodes in the generated synthesis molecular graph data and set data E that specifies a set of edges in the synthesis molecular graph data. For example, the synthesis unit 12 generates set data V = {A, B, C, D, E, F} and set data E = {AB, BC, DE, EF} using identifiers that identify the molecules of the respective nodes in the example of FIG. 15, and generates graph data G = (V, E) that combines these set data V and E as data that represents the synthesis molecular graph data. Figure 4

[0036] ​The addition unit 13 re-generates the synthetic molecular graph data by adding additional edge information that combines two nodes in the molecular graph of the multi-component substance determined from the synthetic molecular graph data to the synthetic molecular graph data. In detail, the addition unit 13 refers to at least the combination node information included in the first molecular graph data and the combination node information included in the second molecular graph data, and extracts a combination of two nodes from the nodes in the first molecular graph that can be further combined and the nodes in the second molecular graph that can be further combined. Also, the addition unit 13 adds additional edge information that combines the extracted combination of nodes to the synthetic molecular graph data. For example, in the example shown in Figure 4 Since the nodes N1, N3, N4, and N6 are designated as nodes that can be further combined, the addition unit 13 adds additional edge information about the edge E13 that combines the node N1 with the node N3, the edge E16 that combines the node N1 with the node N6, the edge E34 that combines the node N3 with the node N4, and the edge E46 that combines the node N4 with the node N6. At this time, the addition unit 13 can refer to the limitation information included in the combination node information to limit the combination that can be combined, or can extract the combination of atoms that can form a chemical bond between the nodes. Figure 4 The molecular graph shown in FIG. 10 is an example of the combination extracted by the addition unit 13 with reference to the limitation information, and is an example in which the combination destination of the node N1 is limited to the nodes N3 and N6, and the combination destination of the node N3 is limited to the nodes N1 and N4, according to the limitation information. Also, the addition unit 13 generates the set data E' by adding the edges shown in the additional edge information to the set data E in the synthetic molecular graph data, and generates the graph data G' = (V, E') that combines the set data V and E' as data that represents the synthetic molecular graph data. For example, according to the example shown in Figure 4 The addition unit 13 generates the set data E' = {AB, AC, AF, BC, CD, DE, DF, EF}.

[0037] The vector transformation unit 14 transforms the graph data G' representing the synthetic molecular graph data generated by the addition unit 13 into a feature vector F. Specifically, the vector transformation unit 14 transforms numerical values representing features of atoms of nodes constituting the set data V included in the graph data G' into vector elements arranged in order when transforming the set data V related to the nodes. The numerical values representing the features of the atoms refer to atomic numbers, electronegativities, and the like. Also, the vector transformation unit 14 transforms numerical values representing features of edges of elements of the set data E' included in the graph data G' into vector elements arranged in order when transforming the set data E' related to the edges. The numerical values representing the features of the edges refer to bond orders, bond distances, and the like. The vector transformation unit 14 generates the feature vector F including the vector elements obtained by transforming the set data V and the vector elements obtained by transforming the set data E' as separate vectors.

[0038] The mixing rate reflection unit 15 reflects the mixing rate data in the feature vector F generated by the vector transformation unit 14, and generates machine learning input data based on the feature vector F in which the mixing rate is reflected. That is, the mixing rate reflection unit 15 reflects the mixing rate r corresponding to the component in the element of the feature vector F corresponding to the node of the molecular graph of the component. For example, the mixing rate reflection unit 15 reflects the mixing rate r1 of the first component constituted by the first molecule in the vector element corresponding to the atom of the node of the first molecular graph, and reflects the mixing rate r2 of the second component constituted by the second molecule in the vector element corresponding to the atom of the node of the second molecular graph. Also, the mixing rate reflection unit 15 reflects the mixing rate corresponding to the component in the element of the feature vector F corresponding to the edge of the molecular graph of the component. For example, the mixing rate reflection unit 15 reflects the mixing rate r1 of the first component constituted by the first molecule in the vector element corresponding to the edge of the first molecular graph, and reflects the mixing rate r2 of the second component constituted by the second molecule in the vector element corresponding to the edge of the second molecular graph. The reflection of the mixing rate is performed by multiplying or adding each element of the vector element by the mixing rate r, or concatenating the elements of the vector element and the mixing rate r.

[0039] Further, the mixing rate reflecting section 15 reflects the mixing rate data in the vector element of the edge corresponding to the additional edge information added by the adding section 13 among the vector elements of the feature vector F. That is, the mixing rate reflecting section 15 reflects the mixing rate r of one or two components corresponding to the molecular graph to which the two nodes combined by the edge belong in the vector element of the edge. That is, the mixing rate reflecting section 15 reflects the product value rixrj of the mixing rates ri, rj of two components in the vector element of the edge in the case where the mixing rate of the component to which one of the nodes belongs is ri and the mixing rate of the component to which the other node belongs is rj. For example, in the case where the corresponding edge is an edge combining the nodes of one molecular graph, the square value of the mixing rate r of the component corresponding to the one molecular graph is reflected in the vector element of the edge, and in the case where the corresponding edge is an edge combining the nodes of two molecular graphs, the product value of the mixing rates r of the two components corresponding to the two molecular graphs is reflected in the vector element of the edge. In other words, in the case where the corresponding edge is an edge combining two nodes within the first molecular graph, only the mixing rate rl of the component composed of the first molecule is reflected in the vector element of the edge, and in the case where the corresponding edge is an edge combining the node of the first molecular graph and the node of the second molecular graph, both the mixing rate rl of the first component composed of the first molecule and the mixing rate r2 of the second component composed of the second molecule are reflected in the vector element of the edge. The reflection of the product value of the mixing rates is performed by multiplying the product value of the mixing rates by each element of the vector element, adding them, or linking the elements of the product value of the mixing rates to the vector element. The reflection of the mixing rates rl, r2 of two components is performed by reflecting the value rixr2 obtained by multiplying the mixing rates of the two components.

[0040] Further, the mixing rate reflecting section 15 outputs the generated input data to the outside. The output input data is read in by the training section 20 in a computer connected to the outside of the input data generation system 10. Further, in the training section 20, the input data is input to the machine learning model as an explanatory variable together with an arbitrary teacher label, thereby generating a learned model. Further, based on the learned model generated by the training section 20, the machine learning model in the predictor 30 is set. However, the training section 20 and the predictor 30 can be the same functional section. Further, by inputting the input data generated by the input data generation system 10 to the machine learning model in the predictor 30, a prediction result of the characteristics of the multi-component substance is generated and output by the predictor 30. In addition, these training section 20 and predictor 30 can be provided in the same computer as the computer 100 constituting the input data generation system 10, or can be provided in a computer separate from the computer 100.

[0041] In one example, the machine learning model generated by the training unit 20 is a completed learning model expected to have the highest calculation accuracy, and thus can be referred to as an "optimal machine learning model". However, it should be noted that the completed learning model is not limited to "optimal in reality". The completed learning model is generated by a computer processing teacher data including a plurality of combinations of input data and output data. The computer calculates the output data by inputting the input data to the machine learning model, and calculates the error of the calculated output data from the output data indicated by the teacher data (i.e., the difference between the calculation result and the correct answer). Also, the computer updates the given parameters of the neural network as the machine learning model according to the error. The computer generates the completed learning model by repeating such learning. The process of generating the completed learning model can be referred to as a learning phase, and the process of the predictor 30 using the completed learning model can be referred to as an application phase.

[0042] [Action of the system]

[0043] Reference Figure 5 and Figure 6 The action of the input data generation system 10 will be described, and the input data generation method according to the present embodiment will be described. Figure 5 is a flowchart showing an example of the action of the input data generation system 10. Figure 6 is a diagram showing an example of molecular data processed in the action of the input data generation system 10.

[0044] First, when the input data generation processing is started as an opportunity of an instruction input by a user of the input data generation system 10, the acquisition unit 11 acquires molecular graph data for each of a plurality of components and mixing ratio data related to each of the plurality of components (step S1). At this time, the acquisition unit 11 acquires at least first molecular graph data indicating a molecular graph of a first molecule included in a first component, second molecular graph data indicating a molecular graph of a second molecule included in a second component, and mixing ratio data related to these first and second components. Figure 6 The (a) part of shows an example of a molecular graph indicated by the first molecular graph data acquired by the acquisition unit 11, Figure 6 The (b) part of shows an example of a molecular graph indicated by the second molecular graph data acquired by the acquisition unit 11. In this example, polypropylene is exemplified as the first molecule, and polybutene is exemplified as the second molecule. For example, as the mixing ratio data, the mixing ratio r1 = "0.5" of polypropylene as the first component and the mixing ratio r2 = "0.5" of polybutene as the second component are acquired.

[0045] Then, the synthetic molecular graph data related to the mixture is generated by the synthesizing section 12 by combining the molecular graph data of the plurality of components, and the set data V which determines the set of nodes in the synthetic molecular graph data is generated by combining the information which identifies the nodes of each molecular graph (step S2). Further, the set data E which determines the set of edges in the synthetic molecular graph data is generated by combining the information which identifies the edges of each molecular graph, and the graph data G = (V, E) which represents the synthetic molecular graph data in which the set data V, E are combined is generated by the synthesizing section 12 (step S3). For example, in the examples of the (a) part and the (b) part in Figure 6 , the set data V1 = {C α , C β , C γ} of nodes shown by the first molecular graph data and the set data V2 = {C δ , C ε , C ζ , C η} of nodes shown by the second molecular graph data are combined, and the set data V = {C α , C β , C γ , C δ , C ε , C ζ , C η} of nodes related to the synthetic molecular graph data is generated. Further, the set data E1 = {C α C β , C β C γ} of edges shown by the first molecular graph data and the set data E2 = {C δ C ε , C ε C ζ , C ζ C η} of edges shown by the second molecular graph data are combined, and the set data E = {C α C β , C β C γ , C δ C ε , C ε C ζ , C ζ C η} of edges related to the synthetic molecular graph data is generated.

[0046] Next, two edges (reaction points) that can be further combined on the molecular graphs of the plurality of components are extracted by the addition unit 13, and an additional edge information that combines the two reaction points is added to the synthetic molecular graph data (step S4). At this time, the set data E' that determines the set of edges in the synthetic molecular graph data is regenerated by adding the edge indicated by the additional edge information to the set data E by the addition unit 13, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. For example, in the examples of the (a) part and the (b) part in FIG. 8, the edge {C Figure 6} indicated by the additional edge information is added, the set data E' = {C α C δ} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. β C δ} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. α C ε} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. β C ε} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. α C β} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. β C γ} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. δ C ε} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. ε C ζ} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. ζ C η} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. α C δ} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. β C δ} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. α C ε} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated. β C ε} is regenerated, and the graph data G' = (V, E') that represents the synthetic molecular graph data in which the set data V and the set data E' are combined is regenerated.

[0047] Further, the graph data G' that represents the synthetic molecular graph data is transformed into the feature vector F by the vector transformation unit 14 with a constant transformation rule (step S5). As the transformation rule, with respect to the elements of the set data V, the features (for example, electronegativity, atomic number) of the atoms that represent the elements are arranged in the vector elements, and with respect to the elements of the set data E', the features (for example, combination frequency, combination distance) of the edges that represent the elements are arranged in the vector elements. The feature vector F is generated by sequentially one-dimensionally linking the vectors transformed from the elements of the graph data G'. α} of the set data V is transformed into the vector [12, 2.55] in which the atomic number and the electronegativity are arranged, and the elements {C α C β} of the set data E' are transformed into the vector [1, 1.53] in which the combination frequency and the combination distance (angstrom) are arranged.

[0048] Then, the mixing rate reflection unit 15 reflects the mixing rate data in the feature vector F, generating the feature vector f. Furthermore, the mixing rate reflection unit 15 combines the feature vector f and the synthesized molecular graph data to generate input data, and outputs this input data to the training unit 20 (step S6). When reflecting the mixing rate, for the elements in the feature vector F corresponding to the nodes and edges of the molecular graph of a certain component, the mixing rate r of that component is reflected; for the elements in the feature vector F corresponding to the edges corresponding to the added edge information, the mixing rate r of the components to which the two nodes connected by the edge belong is reflected. For example, in... Figure 6 In the examples of parts (a) and (b), the mixing ratio r1 = r2 = "0.5" is reflected in the elements other than those corresponding to the edges corresponding to the additional edge information. In the case where the two nodes connected by that edge belong to the same molecular graph, the mixing ratio r1 is reflected in the elements corresponding to the edges corresponding to the additional edge information. 2 (or r2) 2 If the two nodes connected by this edge belong to separate molecular graphs, the mixing ratio r1×r2 = "0.25" is reflected. In this case, the mixing ratio is reflected by multiplying, adding, or connecting the mixing ratio with the vector element. For example, when reflecting the mixing ratio r = "0.5" by multiplying it with the vector element [12, 2.55], it is set to [12×0.5, 2.55×0.5] = [6, 1.275]. And, for example, when reflecting the mixing ratio r = "0.5" by connecting it with the vector element [12, 2.55], it is set to [12, 2.55, 0.5].

[0049] Next, in the training unit 20, a learning phase is performed, and a learned model is generated by repeated training using input data and teacher data (step S7). Then, the generated learned model is set in the predictor 30, and the predictor 30 uses newly acquired input data from the input data generation system 10 to perform an application phase, generating and outputting prediction results of the properties of the multi-component substance (step S8).

[0050] [program]

[0051] The input data generation program for causing a computer or a computer system to function as the input data generation system 10 includes program codes for causing the computer system to function as the acquisition section 11, the synthesis section 12, the addition section 13, the vector conversion section 14, and the mixing ratio reflection section 15. The input data generation program can be provided on the basis of being fixedly recorded in a tangible recording medium such as a CD-ROM, a DVD-ROM, a semiconductor memory, or the like. Alternatively, the input data generation program can also be provided as a data signal superimposed on a carrier wave via a communication network. The provided input data generation program is stored in the auxiliary storage section 103, for example. The processor 101 reads out and executes the input data generation program from the auxiliary storage section 103, thereby realizing the above-described functional elements.

[0052] [Effects]

[0053] As explained above, according to the above-described embodiment, the data that determines the molecular structure of the first molecule and the data that determines the molecular structure of the second molecule are combined to generate the synthetic molecular graph data, the synthetic molecular graph data is converted into a feature vector, data indicating the mixing ratio of the first molecule and the second molecule is reflected in the feature vector, and input data for machine learning is generated. With such a structure, input data related to a multi-component substance for input into a neural network that takes a molecular graph as input can be efficiently generated. As a result, even a multi-component substance containing a plurality of components, the characteristics of the multi-component substance can be predicted with high accuracy by processing the input data using a neural network. In particular, the characteristics of a polymer alloy generated by mixing monomers can be predicted with high accuracy.

[0054] Also, in the above-described embodiment, by reflecting the mixing ratio of a molecule that constitutes a component in node information that is information of an atom of the molecule, input data indicating a multi-component substance can be appropriately generated. As a result, the characteristics of the multi-component substance can be predicted with higher accuracy. In particular, by multiplying, adding, or concatenating the mixing ratio of a component and a vector corresponding to node information of molecular graph data, the mixing ratio can be simply and appropriately reflected in input data indicating a multi-component substance.

[0055] Also, in the above-described embodiment, by reflecting the mixing ratio of a molecule that constitutes a component in edge information that is information of a bond between atoms of the molecule, input data indicating a multi-component substance can be appropriately generated. As a result, the characteristics of the multi-component substance can be predicted with higher accuracy. In particular, by multiplying, adding, or concatenating the mixing ratio of a component and a vector corresponding to edge information of molecular graph data, the mixing ratio can be simply and appropriately reflected in input data indicating a multi-component substance.

[0056] Further, in the above-described embodiment, the bonding information between atoms that can be bonded in a multi-component substance can be generated as additional edge information, and by reflecting the mixing ratio of the molecule in the additional edge information, input data representing the multi-component substance can be appropriately generated. As a result, the characteristics of the multi-component substance can be predicted with higher accuracy. In particular, in a case where a polymer alloy having a random arrangement order of monomers such as a copolymer is the target, it is difficult to construct a molecular graph of the input object in a conventional neural network with a graph as input. In the present embodiment, the chemical bonds between monomers are taken into the molecular graph, and the multi-component substance such as a "polymer alloy" is represented as a graph, and the graph of the multi-component substance can be effectively input to the neural network.

[0057] Further, in the above-described embodiment, the neural network with a graph as input is employed as a model of machine learning. Thereby, the molecular graph data can be input as input, and the characteristics of the multi-component substance can be predicted with high accuracy.

[0058] [Modified Example]

[0059] The present application has been described in detail based on the embodiments thereof. However, the present application is not limited to the above-described embodiments. The present application can be variously modified without departing from the gist thereof.

[0060] In the above-described embodiment, an example in which the input data generation system 10 combines the molecular graphs of two components to generate the molecular graph data and the feature vector related thereto is shown, but a function of combining the molecular graphs of three or more components together with their mixing ratios can also be exerted.

[0061] Further, the constant transformation rule provided in the vector transformation unit 14 of the input data generation system 10 can also be other rules. For example, the feature vector itself can be acquired using machine learning according to the similarity of atoms or bonds. For example, the feature vector can be acquired as a distributed representation using the same method as Word2Vec, which is a neural network used when words are vectorized in natural language processing. Further, the generation of the feature vector can be performed together with the learning stage of the training unit 20.

[0062] The processing order of the input data generation method performed by the at least one processor is not limited to the example in the above-described embodiment. For example, a part of the above-described steps (processes) can be omitted, and each step can be performed in other order. Further, any two or more steps among the above-described steps can be combined, and a part of the steps can be modified or deleted. Alternatively, other steps can be performed in addition to the above-described steps. For example, the processes of steps S7 and S8 can be omitted.

[0063] In the present application, the expression "at least one processor executes a first process, executes a second process,... executes an n-th process," or an expression corresponding thereto indicates the concept including a case where the execution subject (i.e., the processor) of the n processes from the first process to the n-th process is changed in the middle. That is, the expression indicates the concept including both a case where all of the n processes are executed by the same processor and a case where the processor is changed with an arbitrary policy among the n processes.

[0064] Industrial applicability

[0065] One embodiment of the present application can efficiently predict the characteristics of a multi-component substance in which a plurality of components are mixed, as a use, an input data generation system, an input data generation method, and an input data generation program.

[0066] Symbol explanation

[0067] 10 - input data generation system, 100 - computer, 101 - processor, 11 - acquisition unit, 12 - synthesis unit, 13 - addition unit, 14 - vector conversion unit, 15 - mixture rate reflection unit, 20 - training unit, 30 - predictor.

Claims

1. An input data generation system, comprising at least one processor, The at least one processor The system accepts at least the following inputs: first molecular map data corresponding to a first molecule contained in a first component of a multi-component substance; second molecular map data corresponding to a second molecule contained in a second component of the multi-component substance; and mixing rate data representing the mixing ratio of the first and second components. The first molecular map data includes node information identifying nodes of the molecular map corresponding to the first molecule and edge information identifying edges of the molecular map corresponding to the first molecule. The second molecular map data includes node information identifying nodes of the molecular map corresponding to the second molecule and edge information identifying edges of the molecular map corresponding to the second molecule. By combining the node information contained in the first molecular graph data and the node information contained in the second molecular graph data, a node set data is generated. The edge information contained in the first molecular graph data and the edge information contained in the second molecular graph data are combined to generate edge set data. By combining the node set data and the edge set data, synthetic molecular graph data is generated. The synthetic molecular map data is transformed into feature vectors. Input data for machine learning is generated by multiplying, adding, or concatenating the mixing rates of the first and second components shown in the mixing rate data with the features of the feature vector.

2. The input data generation system according to claim 1, wherein, The at least one processor generates the input data by reflecting the mixing rate of the first component in the vector corresponding to the node information of the first molecular map data in the feature vector, and reflecting the mixing rate of the second component in the vector corresponding to the node information of the second molecular map data in the feature vector.

3. The input data generation system according to claim 2, wherein, The at least one processor multiplies, adds, or connects the mixing rates of the first and second components and the vectors corresponding to the node information of the first and second molecular graph data.

4. The input data generation system according to any one of claims 1 to 3, wherein, The at least one processor generates the input data by reflecting the mixing rate of the first component in the vector corresponding to the edge information of the first molecular map data in the feature vector, and reflecting the mixing rate of the second component in the vector corresponding to the edge information of the second molecular map data in the feature vector.

5. The input data generation system according to claim 4, wherein, The at least one processor multiplies, adds, or concatenates the mixing rates of the first and second components and the vectors corresponding to the edge information of the first and second molecular graph data.

6. The input data generation system according to claim 1, wherein, The at least one processor The binding node information of the nodes of the molecular graph that can be bound is further accepted as the first molecular graph data and the second molecular graph data. Additional edge information is generated related to the edge that combines two nodes, namely the node shown in the binding node information contained in the first molecular graph data and the node shown in the binding node information contained in the second molecular graph data, and the additional edge information is added to generate the synthetic molecular graph data. The input data is generated by reflecting the mixing ratio of the first component and the second component in the vector corresponding to the additional edge information in the feature vector.

7. The input data generation system according to claim 1, wherein, The machine learning described is a neural network that takes a graph as input.

8. The input data generation system according to claim 1, wherein, The first molecule and the second molecule are monomers. The mixing ratio data represents the mixing ratio of the first molecule and the second molecule in a polymer alloy generated based on the first molecule and the second molecule.

9. An input data generation method, executed by a computer having at least one processor, the input data generation method comprising: The process includes at least the following steps: accepting first molecular map data that determines a molecular map corresponding to a first molecule contained in a first component of a multi-component substance; second molecular map data that determines a molecular map corresponding to a second molecule contained in a second component of a multi-component substance; and mixing rate data representing the mixing ratio of the first component and the second component. The first molecular map data includes node information that determines the nodes of the molecular map corresponding to the first molecule and edge information that determines the edges of the molecular map corresponding to the first molecule. The second molecular map data includes node information that determines the nodes of the molecular map corresponding to the second molecule and edge information that determines the edges of the molecular map corresponding to the second molecule. The steps are as follows: combining the node information contained in the first molecular graph data and the node information contained in the second molecular graph data to generate node set data, and combining the edge information contained in the first molecular graph data and the edge information contained in the second molecular graph data to generate edge set data. The step of combining the node set data and the edge set data to generate synthetic molecular graph data; The step of transforming the synthetic molecular map data into feature vectors; and The step of generating input data for machine learning by multiplying, adding, or concatenating the mixing rates of the first and second components shown in the mixing rate data with the features of the feature vector.

10. An input data generation program that causes a computer to perform the following steps: The process includes at least the following steps: accepting first molecular map data that determines a molecular map corresponding to a first molecule contained in a first component of a multi-component substance; second molecular map data that determines a molecular map corresponding to a second molecule contained in a second component of the multi-component substance; and mixing rate data representing the mixing ratio of the first component and the second component. The first molecular map data includes node information determining the nodes of the molecular map corresponding to the first molecule and edge information determining the edges of the molecular map corresponding to the first molecule. The second molecular map data includes node information determining the nodes of the molecular map corresponding to the second molecule and edge information determining the edges of the molecular map corresponding to the second molecule. The steps are as follows: combining the node information contained in the first molecular graph data and the node information contained in the second molecular graph data to generate node set data, and combining the edge information contained in the first molecular graph data and the edge information contained in the second molecular graph data to generate edge set data. The step of combining the node set data and the edge set data to generate synthetic molecular graph data; The step of transforming the synthetic molecular map data into feature vectors; and The step of generating input data for machine learning by multiplying, adding, or concatenating the mixing rates of the first and second components shown in the mixing rate data with the features of the feature vector.

Citation Information

Patent Citations

  • Connectivity prediction method, apparatus, program, recording medium, and production method of machine learning algorithm

    JP2019028879A

  • Method For Decomposing Polymer Material, Method For Producing Recycled Resin, And Method For Recovering Inorganic Filler

    CN105778150A

  • Biological sample quick and intelligent recognition method based on molecular map

    CN109870533A