Structured processing device and structured processing method for assisting in generation of structured data representing process

By using a structuring processing device with causal models and graph neural networks, the system addresses the inefficiencies in generating structured data, reducing computational load and improving prediction accuracy through coarse-grained process patterns.

WO2025164006A1PCT designated stage Publication Date: 2025-08-07HITACHI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/038698
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-30
Filing Date
2024-10-30
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing methods for generating structured data representing processes require significant computational resources and often fail to accurately reflect factors external to substances, leading to increased calculation loads and decreased prediction accuracy.

Method used

The system generates structured data using process features based on coarse-grained process patterns, employing a structuring processing device that includes a causal model and graph neural networks to predict process outcomes, reducing computational load and improving prediction accuracy.

Benefits of technology

This approach effectively reduces the computational burden and enhances the accuracy of process prediction by leveraging coarse-grained process patterns and causal models, facilitating efficient generation of structured data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024038698_07082025_PF_FP_ABST
    Figure JP2024038698_07082025_PF_FP_ABST
Patent Text Reader

Abstract

This structured processing device outputs structured data of a process in which a prediction target is predicted on the basis of a process feature amount which is a feature amount based on a coarse granularity such as a process pattern of a treatment applied to a process rather than a fine granularity such as an element in a process having a plurality of elements including a plurality of procedures.
Need to check novelty before this filing date? Find Prior Art

Description

Structured processing device and structured processing method for supporting generation of structured data representing a process

[0001] The present invention generally relates to techniques for assisting in the generation of structured data.

[0002] In recent years, in various fields, there has been an emerging need to support the generation of structured data representing processes consisting of multiple steps using AI (Artificial Intelligence) or other techniques. For example, in the industrial field, AI that recommends operation processes for equipment and processes for equipment failures has been put to practical use, in the medical field, AI that supports diagnosis, treatment, and medication has been put to practical use, and in the materials field, AI that recommends synthesis processes for new materials has been put to practical use.

[0003] Generating structured data for a new process from previously accumulated data generally requires a huge amount of time, specialized knowledge, and trial and error. In response to this, techniques described in Non-Patent Documents 1 and 2 are known.

[0004] Non-Patent Document 1 describes a synthetic route search method in MI (Materials Informatics) that predicts precursors (substances at a stage before a certain chemical substance is produced) based on route template information and structural analysis information using a GNN (Graph Neural Network). Non-Patent Document 2 describes a link estimation method using counterfactual thinking, specifically, a method for estimating whether link connections will change when a change is made to the attributes of a node.

[0005] COLEY, Connor W. GREEN, William H.; ; JENSEN, Klavs F. Machine learning in computer-aided synthesis planning. Accounts of chemical research, 2018, 51.5: 1281-1289. ZHAO, Tong, et al. Learning from counterfactual links for link prediction. In: International Conference on Machine Learning. PMLR, 2022. p. 26911-26926.

[0006] In order to generate structured data representing a process, a prediction support process is performed, which is a process related to the prediction process of the process.

[0007] The technology described in Non-Patent Document 1 focuses on substances, while the technology described in Non-Patent Document 2 focuses on nodes. Therefore, with either of these technologies, there is a risk that the number of combinations in the prediction support process will become enormous, increasing the calculation load of the prediction support process, or that factors external to the substances (nodes) (e.g., reaction conditions and synthesis objectives) will not be reflected, resulting in a decrease in process prediction accuracy.

[0008] The present invention has been made in view of the above-mentioned problems, and aims to reduce the calculation load of the prediction support processing for the process represented by the generated structured data and improve the prediction accuracy of the process.

[0009] A representative example of the present invention is as follows: That is, the structuring processing device outputs structured data of a process in which a prediction target is predicted based on process features, which are features based on coarse granularity such as a process pattern of actions applied to a process, rather than on fine granularity such as elements in a process having multiple elements including multiple procedures.

[0010] According to the present invention, it is possible to reduce the computational load of the process related to the prediction of the process represented by the generated structured data and improve the prediction accuracy of the process. Problems, configurations, and effects other than those described above will become clear from the description of the following embodiments.

[0011] 1 is a diagram illustrating an example of a system according to a first embodiment. FIG. 1 is a diagram illustrating an example of a hardware configuration of a computer. FIG. 2 is a diagram illustrating an example of structured data. FIG. 3 is a diagram illustrating an example of a structured data output screen. FIG. 4 is a diagram illustrating an example of structured data. FIG. 5 is a diagram illustrating an example of a learning process. FIG. 6 is a diagram illustrating an example of an inference process. FIG. 7 is a diagram illustrating an example of a reaction label. FIG. 8 is a diagram illustrating an example of a frequent process pattern [structure] according to a first embodiment. FIG. 9 is a diagram illustrating an example of a frequent process pattern [attribute value] according to a first embodiment. FIG. 10 is a diagram illustrating an example of a process feature [structure] according to a first embodiment. FIG. 11 is a diagram illustrating an example of a process feature [attribute value] according to a first embodiment. FIG. 12 is a diagram illustrating an example of a raw material-product table. FIG. 13 is a diagram illustrating an example of a subgraph list. FIG. 14 is a diagram illustrating an example of the flow of a causal model learning process. FIG. 15 is a diagram illustrating an example of the flow of a frequent process pattern extraction process. FIG. 16 is a diagram illustrating an example of the flow of an attribute value extraction process. FIG. 17 is a diagram illustrating an example of the flow of a process feature extraction process. FIG. 18 is a diagram illustrating an example of the flow of an experiment process information registration process. FIG. 19 is a diagram illustrating an example of the flow of a frequent process pattern matching process according to a first embodiment. FIG. 19 is a diagram illustrating an example of the flow of a graph construction process. FIG. 19 is a diagram illustrating an example of the flow of a process feature inference process. FIG. 19 is a diagram illustrating an example of the flow of an edge prediction process. FIG. 20 is a diagram illustrating an example of a part of a learning process. 1 is a diagram showing an example of experimental process data; FIG. 2 is a diagram showing an example of a frequent process pattern [structure]; FIG. 3 is a diagram showing an example of a frequent process pattern [attribute value]; FIG. 4 is a diagram showing an example of a process feature [attribute value]; FIG. 5 is a diagram showing an example of a part of an inference process; FIG. 6 is a diagram showing an example of a raw material-product table; FIG. 7 is a diagram showing an example of a first subgraph and a second subgraph; FIG. 8 is a diagram showing an example of a process feature [attribute value]; FIG. 9 is a diagram showing an example of measurement process data according to Example 2; FIG. 10 is a diagram showing an example of a frequent process pattern [structure] according to Example 2; FIG. 11 is a diagram showing an example of a frequent process pattern [attribute value] according to Example 2; FIG. 12 is a diagram showing an example of a process feature [structure]; FIG. 13 is a diagram showing an example of a process feature [attribute value]; FIG. 14 is a diagram showing an example of a flow of a frequent process pattern matching process according to Example 2; FIG. 15 is a diagram showing an example of matching (matching of a process pattern and process data) according to Example 1;FIG. 10 is a diagram schematically illustrating an example of matching (matching between a process pattern and process data) according to a second embodiment.

[0012] Hereinafter, several embodiments of the present invention will be described with reference to the drawings. The following description and drawings are examples for explaining the present invention, and some omissions and simplifications have been made as appropriate for clarity of explanation. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.

[0013] In the following description, identical or similar components may be assigned the same reference numerals, and redundant description may be omitted. Furthermore, in the following description, the letter "S" before a reference numeral indicates a processing step. Furthermore, in the following description, various pieces of information may be described using expressions such as "tables," but the various pieces of information may also be expressed using data structures other than these.

[0014] Furthermore, in the following explanation, an example of structuring information about the synthesis process of a material described in an experimental report can be used, but the process to be structured can be applied to various fields and use cases described in the background art.

[0015] 1 is a diagram illustrating an example of a system according to Example 1. FIG. 2 is a diagram illustrating an example of a hardware configuration of a computer 200.

[0016] 1 is composed of a structuring processing device 100 and a user terminal 101. The structuring processing device 100 and the user terminal 101 are connected via a communication network 102 in a state where two-way communication is possible. The communication network 102 is, for example, a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, a public communication network, a dedicated line, etc. The number of user terminals 101 may be two or more.

[0017] The structuring processing device 100 and the user terminal 101 are configured, for example, by a computer 200 as shown in Fig. 2. The computer 200 includes a main memory device 202, an auxiliary memory device 203, an input device 204, an output device 205, a communication device 206, and an arithmetic unit 201 connected to these devices.

[0018] The arithmetic device 201 executes a program stored in the main memory device 202. The arithmetic device 201 is, for example, a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an artificial intelligence (AI) chip, or the like. The arithmetic device 201 operates as a functional unit (module) that realizes a specific function by executing processing according to the program. In the following description, when processing is described using a functional unit as the subject, it indicates that the arithmetic device 201 is executing a program that realizes the functional unit.

[0019] The main memory device 202 stores programs and data executed by the arithmetic device 201. The main memory device 202 is, for example, a non-volatile memory such as a read-only memory (ROM), a random access memory (RAM), or a non-volatile RAM (NVRAM). The main memory device 202 is also used as a work area.

[0020] The auxiliary storage device 203 permanently stores data. The auxiliary storage device 203 is, for example, a solid state drive (SSD) or a hard disk drive. The computer 200 does not necessarily have the auxiliary storage device 203. In this case, the programs and data may be acquired from an optical storage device such as a compact disc (CD) or a digital versatile disc (DVD), an IC card, an SD card, or the like, or may be acquired from a storage area on an externally connected storage system or cloud system. The programs and data stored in the auxiliary storage device 203 are read by the arithmetic device 201 and loaded into the main storage device 202.

[0021] The input device 204 is an interface that accepts input from the outside, and is, for example, a keyboard, a mouse, a touch panel, a card reader, a pen-input tablet, a voice input device, or the like.

[0022] The output device 205 is an interface that outputs various information such as the process progress and results, etc. The output device 205 is, for example, a display device such as a liquid crystal monitor or LCD (Liquid Crystal Display), an audio output device, a printer, etc.

[0023] The computer 200 may not have the input device 204 and the output device 205. In this case, the computer 200 inputs and outputs information via the communication device 206.

[0024] The communication device 206 communicates with other devices, and is, for example, a network interface card (NIC), a wireless communication module, or a USB module.

[0025] The structuring processing device 100 generates structured data. Here, the process is, for example, an experimental process, and is assumed to be composed of multiple procedures. The structured data is data that represents the process as a structure of multiple procedures, and can be, for example, data in Json format, data in XML format, data in RDF format, or data in GraphML format. The present invention is not limited to the data format of the structured data. In this embodiment, the structured data may be data in GraphML format.

[0026] The structuring processing device 100 has an information management unit 110 and a structuring processing unit 120, and also holds a structuring assistance database 130 and a structured data database 140. Structured data is stored in the structured data database 140. Data other than structured data, which is referenced or generated in the structuring processing performed by the structuring processing unit 120, may be stored in the structuring assistance database 130.

[0027] The information management unit 110 manages the data stored in the structuring assistance database 130 and the structured data database 140. The structuring processing unit 120 executes structuring processing and stores the generated structured data in the structured data database 140. The information management unit 110 and the structuring processing unit 120 may be realized as a function of middleware that manages an operating system, a file system, a relational database, and NoSQL such as KVS (Key-Value Store). The structuring processing may include prediction model processing involving a prediction model for process prediction, and processing for generating structured data representing the predicted process.

[0028] The user terminal 101 has a registration unit 170 that displays a screen for accepting registration (an example of a UI (User Interface)), and a display unit 180 that displays a screen for accepting submission and modification of structured data.

[0029] The functions of the structuring processing device 100 may be realized using a computer system configured from multiple computers 200. Furthermore, all or part of the functions of the structuring processing device 100 may be realized using virtualization technology. For example, all or part of the functions of the structuring processing device 100 may be realized using cloud services such as SaaS (Software as a Service), PaaS (Platform as a Service), and IaaS (Infrastructure as a Service). Furthermore, the functions of the user terminal 101 may be aggregated in the structuring processing device 100, and the user terminal 101 may be a user interface device for input and output.

[0030] As shown in Fig. 3, the structured data is data in GraphML format that includes data representing nodes in a graph and edges connecting the nodes. The display unit 180 of the user terminal 101 displays a screen that displays the graph represented by the structured data, such as the screen shown in Fig. 4. Fig. 5 schematically shows an example of a graph represented by the structured data.

[0031] MI technology makes it possible to provide recipes as data including pairs of products (e.g., materials) and raw materials required to produce the products. In this embodiment, the structuring processing unit 120 performs retrosynthesis analysis to predict a process for producing the products represented by the recipes from the raw materials represented by the recipes, and generates structured data representing the process. The retrosynthesis analysis process is included in the structuring process. The structuring process performed in this embodiment can be broadly divided into learning process and inference process. The learning process and inference process will be described below. In the following description, structured data is graph data representing a graph composed of nodes and edges. The graph may be a directed graph or an undirected graph.

[0032] FIG. 6 is a diagram illustrating an example of the learning process.

[0033] The structuring processing unit 120 generates a frequent process pattern [structure] 603 and a frequent process pattern [attribute value] 604 by executing a frequent process pattern extraction process based on the experimental process data 601 and the reaction label 602. The experimental process data 601 is structured data representing a completed experimental process conducted in the past. The reaction label 602 is data representing the correspondence between the experimental process and the pattern applied to the experimental process. The frequent process pattern [structure] 603 is structured data representing a frequent process pattern. A frequent process pattern is a process portion common to multiple experimental processes. The frequent process pattern [attribute value] 604 is attribute value data of the frequent process pattern. The experimental process data 601 may be stored in the structured data database 140. The reaction label 602 may be stored in the structuring support database 130. The frequent process pattern [structure] 603 may be stored in the structured data database 140. The frequent process pattern [attribute value] 604 may be stored in the structuring support database 130. The information management unit 110 may display a screen showing the frequent process pattern [structure] 603 via the display unit 180 of the user terminal 101, and may accept editing of the graph represented by the frequent process pattern [structure] 603 from the user via the screen.

[0034] Next, the structuring processing unit 120 executes a process feature extraction process using the frequent process pattern [structure] 603, the frequent process pattern [attribute value] 604, the experimental process data 601, and the graph feature ZG (graph feature obtained from the experimental process data 601), thereby generating a process feature [structure] 605 and a process feature [attribute value] 606. The process feature [structure] 605 is structured data representing a causal graph of the process feature. The process feature [attribute value] 606 is attribute value data of the process feature. Note that the information management unit 110 may display a screen displaying the process feature [structure] 605 via the display unit 180 of the user terminal 101, and may accept editing of the causal graph of the process feature from the user via the screen. The process feature [structure] 605 may be stored in the structured data database 140. The process feature [attribute value] 606 may be stored in the structuring support database 130.

[0035] Next, the structuring processing unit 120 generates a causal model 607 by performing a causal model generation process based on the generated process feature quantity [structure] 605 and process feature quantity [attribute value] 606. The process feature quantity [attribute value] 606 may be stored in the structuring support database 130.

[0036] Finally, the structuring processing unit 120 trains an encoder 611 that extracts graph features ZG from the experimental process data 601 and a decoder 612 that performs edge prediction between subgraphs using the graph features ZG and the causal model 607. The model as the encoder 611 and the model as the decoder 612 may be stored in the structuring support database 130. The encoder 611 receives structured data and outputs graph features ZG of the input structured data (graph data). The encoder 611 receives the causal model and graph features and outputs the edge prediction table 608. The training of the encoder 611 and the decoder 612 may be performed, for example, as follows: The structuring processing unit 120 trains the encoder 611 and the decoder 612 using data including the experimental process data 601 as training data and data including the edge prediction table 608 as correct data. Thereafter, the structuring processing unit 120 replaces the decoder 612 trained in this process with the causal model 607. Specifically, the model resulting from the learning of the decoder 612 is a model [E=f(ZG)] that predicts an edge from experimental process data, while the causal model is a model [E=f(ZG, process feature)] that predicts an edge using experimental process data and process feature values. By replacing the decoder 612 with the causal model 607, it becomes possible to solve the same problem with more explanatory variables.

[0037] FIG. 7 is a diagram illustrating an example of the inference process.

[0038] The structuring processing unit 120 generates subgraphs inductively by executing a graph construction process based on the raw material-product table 701 (a table representing combinations of raw materials and products that identify an experimental process). As a result, the subgraphs are expanded by connecting them with edges, and as a result, experimental process data 702 is generated as a graph or its subgraph. The raw material-product table 701 may be data provided from the user terminal 101, or may be stored in the structuring support database 130. The experimental process data 702 as a graph (i.e., completed) represents the experimental process for obtaining the product represented by the raw material-product table 701 from the raw materials represented by the same table 701.

[0039] In this graph construction process, the structuring processing unit 120 executes a process feature inference process based on the graph feature Zg (graph feature obtained from the experimental process data 702) and the frequent process pattern [structure] 603 and frequent process pattern [attribute value] 604 generated in the learning process, thereby generating a process feature [attribute value] 706, which is attribute value data of the process feature of the subgraph. The process feature [structure] 605 is the process feature [configuration] 605 generated in the learning process. The structure of the process feature [attribute value] 706 may be the same as the structure of the process feature [attribute value] 606 in the learning process. The process feature [attribute value] 706 may be stored in the structuring support database 130. Next, the structuring processing unit 120 generates a graph feature Zg of the subgraph using the encoder 611 trained in the learning process, and inputs the graph feature Zg and values ​​obtained from the process feature [structure] 605 and the process feature [attribute value] 706 to the decoder 612 trained in the learning process. This generates an edge prediction table 708 that indicates, for each pair of a node in the first subgraph and a node in the second subgraph, a prediction result of whether the nodes constituting the pair are connected by an edge. During the graph construction process, the edge prediction table 708 is applied to the first subgraph and the second subgraph, thereby expanding the subgraph. For example, if the edge prediction table 708 indicates the presence of an edge for a pair of a node in the first subgraph and a node in the second subgraph (i.e., the presence of an edge is predicted), the nodes are connected by an edge, thereby connecting the first subgraph and the second subgraph by an edge, and as a result, the subgraph is expanded. In the graph construction process, the structuring processing unit 120 sets the expanded subgraph as the first subgraph and another subgraph as the second subgraph, performs process feature inference processing and causal model generation processing, and inputs the graph feature Zg generated by the encoder 611 for the first subgraph and the causal model 707 generated by the causal model generation processing to the decoder 612, thereby obtaining a prediction result as to whether or not a node in the first subgraph and a node in the second subgraph are connected by an edge.

[0040] FIG. 8 is a diagram showing an example of the reaction label 602.

[0041] The reaction label 602 has a graph ID column and a pattern name column. The graph ID column registers a graph ID as the ID of the experimental process. The pattern name column registers the name of the reaction to be applied to the experimental process as a pattern name (label). For example, the first record indicates that a radical reaction is applied to the experimental process assigned the graph ID "1."

[0042] Furthermore, in the frequent process pattern extraction process, the structuring processing unit 120 may prepare a flag (or a reaction label with an added flag field) for each action (reaction). The flag field stores a value indicating whether or not the action (reaction) specified by the pattern name exists in the graph. For example, this indicates that a radical reaction is applied to graphs with graph IDs 1 and 3, and that an effect reaction is applied to graphs with graph IDs 2 and 3.

[0043] FIG. 9 is a diagram showing an example of the frequent process pattern [structure] 603.

[0044] The frequent process pattern [structure] is structured data that represents a part of the structure that is common among the experimental processes. For example, in the example shown in the figure, "Cl 2 There is a common pattern that "a radical reaction occurs when certain substances are mixed and then irradiated with light," and the structured data representing this common pattern (graph portion) is the frequent process pattern [structure] 603.

[0045] FIG. 10 is a diagram showing an example of the frequent process pattern [attribute value] 604. As shown in FIG.

[0046] The frequent process pattern [attribute value] 604 has a graph ID column, a T column, an O1 column, an On column, a t1 item column, a t1 value column, a tn item column, a tn value column, a Δ value column, and a Chain column.

[0047] The graph ID field contains a graph ID as the ID of the experimental process. The T field contains the name of the treatment applied to the experimental process as a reaction.

[0048] The O1 column is where the name of the substance to which the reaction is applied is registered, and the On column is where the name of the substance to be generated as a result of applying the reaction is registered.

[0049] The t1 item column is registered with the name of a physical property item of the source substance t1 to be reacted. The t1 value column is registered with the t1 value, which is the physical property value of t1. The tn item column is registered with the name of a physical property item of the target substance tn generated as a result of the reaction. The tn value column is registered with the tn value, which is the physical property value of tn. The "physical property item" may be a structure or attribute item. The "physical property value" may be a value (structure value or attribute value) belonging to a structure item or attribute item.

[0050] In the Δ value column, a Δ value is registered as the difference between the physical property values ​​of the t1 value and the tn value.

[0051] The Chain column has a # column, an S column, and an E column. The # column registers an edge ID as the position of the edge in the frequent process pattern. The S column registers the name of an element (e.g., a substance or an operation) corresponding to a node at one end of the edge (a starting node if the edge is directed). The E column registers the name of an element (e.g., a substance or an operation) corresponding to a node at the other end of the edge (an ending node if the edge is directed).

[0052] For example, the first record shows the following: A radical reaction is applied in the first experimental process, resulting in CH 4 (O1) is CH 3 -CH 3 (On), and the information of Smiles is converted from "C" (t1 value) to "CC" (tn value), so the Δ value (for example, the string edit distance of Smiles) is "1". 4 The edge ID of the edge connecting the node "mixed" to the node "mixed" is "1".

[0053] FIG. 11 is a diagram showing an example of the process feature amount [structure] 605.

[0054] The process feature [structure] 605 is structured data representing a causal graph as the relationship between reactions (explanatory variables) extracted as frequent process patterns and the probability of edge existence between the first subgraph and the second subgraph (objective variable). The process feature [structure] 605 illustrated in the figure is as follows: The explanatory variable T is a radical reaction applied between node vi in ​​the first subgraph G1 (e.g., a subgraph of an experimental process) and node vj in the second subgraph G2. The objective variable A is the probability of edge existence between node vi and node vj to which the radical reaction was applied. Process features are used to predict the objective variable A. Process features include flow features, path features, and chain features. Flow features are features related to the changes that occur when treatment T is applied to material O1 to obtain material On (e.g., a product) in the flow from material O1 (e.g., a raw material) to material On. In this embodiment, the flow feature is based on the treatment (reaction) T and the Δ value representing the change from the O1 value to the On value. The path feature is a feature related to a path in the flow where a change from t1 to tn occurs. In this embodiment, the path feature is based on Pa and Pd. Pa indicates whether subgraph G1 exists on the path between treatment T's source material t1 (element with a t1 value), i.e., whether the target material exists as an ancestor. Pd indicates whether subgraph G2 exists on the path between treatment T's target material tn (element with a tn value), i.e., whether the target material exists as a descendant. The chain feature is based on the positions of subgraphs G1 and G2 in the frequent process pattern. In this embodiment, the chain feature is Chain, which indicates the positions of subgraphs G1 and G2 in the frequent process pattern of treatment T.

[0055] FIG. 12 is a diagram showing an example of the process feature amount [attribute value] 606. As shown in FIG.

[0056] The process feature [attribute value] 606 has a graph ID field, a sub-graph information field, a flow feature field, a path feature field, a chain feature field, and an A field.

[0057] In the graph ID column, a graph ID is registered as an ID of the experiment process.

[0058] The subgraph information column registers information about each listed edge of the experimental process. The subgraph information column has a vi column, a vj column, a G1 column, and a G2 column. The vi column registers the name of the element corresponding to the node vi at one end of the edge (node ​​vi in ​​subgraph G1). The vj column registers the name of the element corresponding to the node vj at the other end of the edge (node ​​vj in subgraph G2). The G1 column registers the feature vector of the subgraph G1 that includes node vi. The G2 column registers the feature vector of the subgraph G2 that includes node vj.

[0059] The flow feature field stores information about the feature of the flow of the change from O1 to On due to the application of a treatment (reaction tail) as a frequent process pattern to the experimental process. The flow feature field includes a T field, an O1 field, an On field, and a Δ value field. The T field stores a T value indicating whether the treatment (reaction) was applied to the experimental process. In the example shown in FIG. 12, the "relevant treatment (reaction)" is a radical reaction. A process feature [attribute value] 606 may be prepared for each treatment (reaction), and the treatment (reaction) corresponding to the T field may be the treatment (reaction) corresponding to the process feature [attribute value] 606. The O1 field stores the name of the substance to which the reaction is applied. The On field stores the name of the substance produced as a result of the application of the reaction. The Δ value field stores a Δ value representing the difference between the t1 value and the tn value.

[0060] The path feature column stores information about path features that are included in the frequent process pattern for each edge listed in the experimental process. The path feature column has a t1 column, a tn column, a Pa column, and a Pd column. The t1 column stores the name of the source substance t1 to be reacted. The tn column stores the name of the target substance tn to be generated as a result of the reaction. The Pa column stores a Pa value indicating whether or not a source substance exists in an ancestor of the subgraph G1. The Pd column stores a Pd value indicating whether or not a target substance exists in a grandchild of the subgraph G2.

[0061] The chain feature column stores information about the chain feature that matches which position in the frequent process pattern. The chain feature column has a Chain column, that is, a # column, an S column, and an E column. The # column stores the edge ID of the edge. The S column stores the name of the element corresponding to the node at one end of the edge. The E column stores the name of the element corresponding to the node at the other end of the edge.

[0062] In the column A, an A value is registered, which indicates whether an edge exists between node vi and node vj in a graph representing a process resulting from application of treatment T of a frequent process pattern to experimental process data. The A value is a binary value indicating whether the probability of edge existence as objective variable A is equal to or greater than a threshold.

[0063] FIG. 13 shows an example of the raw material-product table 701 .

[0064] The raw material-product table 701 has information on pairs of one or more raw materials and one or more products. The raw material-product table 701 has an M column, an M item column, an M value column, a P column, a P item column, and a P value column.

[0065] The M column is where the name of the raw material M (substance) is registered. The M item column is where the name of the physical property item of the raw material is registered. The M value column is where the M value, which is the physical property value of the raw material, is registered.

[0066] The P column is where the name of the product P (substance) is registered. The P item column is where the name of the physical property item (physical property item) of the product is registered. The P value column is where the P value, which is the physical property value of the product, is registered.

[0067] According to the example shown in FIG. 13, CH 4 (Smiles "C", hardness "5°") and Cl (Smiles "Cl") to obtain the product CH 3 -CH 3 The experimental process for producing (Smiles "CCC", hardness "10°") is identified in the inference.

[0068] FIG. 14 shows an example of a subgraph list 1400 .

[0069] The subgraph list 1400 may be provided from the user terminal 101, or may be registered in the structuring support database 130. The subgraph list 1400 has a subgraph ID column and a graph ID column. The subgraph ID column registers the subgraph ID of the subgraph constructed in the inference process. The graph ID column registers the graph ID of the experimental process identified from the raw material-product table 701.

[0070] According to the example shown in FIG. 14, the graph representing the experimental process identified from the raw material-product table 701 is made up of three subgraphs.

[0071] An example of the processing performed in this embodiment will be described below.

[0072] 15 is a diagram showing an example of the causal model learning process. The causal model learning process is a part of the learning process shown in FIG. 6 , from the causal model generation process to the output of a causal model that becomes the final decoder 612.

[0073] The structuring processing unit 120 acquires the reaction labels 602, for example, from the user terminal 101 or the structuring support database 130 (S1501). The structuring processing unit 120 also acquires the experimental process data 601 having each graph ID represented by the reaction labels 602 from the structured data database 140 (S1502).

[0074] If the reaction label 602 acquired in S1501 has one or more unprocessed pattern names (S1503: YES), the structuring processing unit 120 selects one unprocessed pattern name from the one or more unprocessed pattern names. An "unprocessed pattern name" is a pattern name that has not yet been selected in this causal model learning process from the acquired reaction label 602. The structuring processing unit 120 performs a frequent process pattern extraction process (S1504), a process feature extraction process (S1505), and a causal model generation process (S1506) for the selected unprocessed pattern name. Thereafter, the process returns to S1503. A flag column may be added to the reaction label 602, and "1" may be registered in the flag column, indicating that the treatment (reaction) specified by the selected unprocessed pattern name is present in the graph (see FIG. 8 ). This makes it possible to construct a causal graph in which the subgraph information representing the graph structure information for determining whether an edge exists, the flow features including the actions (reactions) corresponding to the pattern names, the path features, and the chain features are explanatory variables, and the probability of edge existence is the objective variable in subsequent processing.

[0075] If there are no unprocessed pattern names in the response label 602 acquired in S1501 (S1503: NO), the structuring processing unit 120 outputs the causal model generated for each selected pattern name (S1507). The causal model for each pattern name is used as the final decoder 612.

[0076] The causal inference described in WO2023 / 141019 can be applied to the causal model generation process (S1506). That is, a conditional probability model A = f(X, Z) is generated. A is the objective variable and is the probability that an edge exists between node vi in ​​subgraph G1 and node vj in subgraph G2. Based on this probability (specifically, whether this probability is equal to or greater than a threshold), the A value (a value representing whether an edge exists) is determined. X is the explanatory variable, which is the node label feature of nodes vi and vj, and the graph feature of G1 and G2. Z is the process feature as a confounding factor, specifically, a flow feature (e.g., a Δ value), a path feature (e.g., a Pa value and a Pd value), and a linkage feature (e.g., an edge ID). The model of the encoder 611 that calculates the graph feature and the model of the decoder 612 may be a GNN (Graph Neural Network), and the edge prediction performed by the decoder 612 using a causal model may be Link Prediction based on causal inference. The decoder 612 itself may be a conditional probability model. That is, when a graph to be predicted as an edge is given as explanatory variables (vi, vj, G1, G2), the structuring processing unit 120 calculates confounding factors (flow feature, path feature, and chain feature) using a frequent process pattern [structure] 603 and a frequent process pattern [attribute value] 604, and inputs the confounding factors into a conditional probability model (A=f(X, Z)) to calculate the existence probability of an edge. The experimental process data input to the encoder 611 includes both subgraphs G1 and G2, not only in the learning process but also in the inference process, and since the experimental process data (graph) including both is input, the probability of whether G1 and G2 are connected by an edge is calculated. In a frequent process pattern, of the edges corresponding to the edge ID, Chain-S may be the node information of the start node (upstream side of the process), and Chain-E may be the node information of the end node (downstream side of the process).

[0077] FIG. 16A is a diagram showing an example of the flow of the frequent process pattern extraction process (S1504 in FIG. 15).

[0078] The structuring processing unit 120 performs a structure extraction process (S1601). Specifically, the structuring processing unit 120 identifies, as a frequent process pattern, a common graph portion (common structure) corresponding to the pattern name selected in S1503 from among all the experimental processes (graphs) represented by all the experimental process data acquired in S1502, and generates a frequent process pattern [structure] 603 representing the identified frequent process pattern. The identification of a frequent process pattern can be performed using GNNExplainer. GNNExplainer is described, for example, in "YING, Zhitao, et al. GnnExplainer: Generating explanations for graph neural networks. Advances in neural information processing systems, 2019, 32." The experimental process data acquired in S1502 may be subjected to generalization processing to remove variations in the spelling of node names, and frequent process patterns may be identified from the generalized experimental process data. For the generalization processing, for example, the technology described in the applicant's prior application (Patent Application No. 2022-200233), which was not published at the time of filing this application, may be used.

[0079] The structuring processing unit 120 performs attribute value extraction processing (S1602). The attribute value extraction processing results in obtaining a frequent process pattern [structure] 604 for the pattern name selected in S1503 (the frequent process pattern identified in S1601).

[0080] The structuring processing unit 120 registers the extracted graph pattern (frequent process pattern [structure] 604) in a frequent process pattern list (S1603), and outputs the frequent process pattern list (S1604). The "frequent process pattern" is a table that contains pairs of frequent process pattern [structure] 603 and frequent process pattern [attribute value] 604, and is used in the inference processing (FIG. 7).

[0081] FIG. 16B is a diagram illustrating an example of the flow of the attribute value extraction process.

[0082] The structuring processing unit 120 determines whether there is any unselected experimental process data among the experimental process data acquired in S1502 (S1611). "Unselected experimental process data" refers to experimental process data that has not yet been selected in this attribute value extraction process from the experimental process data acquired in S1502.

[0083] If there is unselected experimental process data (S1611: YES), the structuring processing unit 120 selects the unselected experimental process data and identifies substances before and after the frequent process pattern identified in S1601 from the experimental process represented by the experimental process data (S1612). In S1612, for the frequent process pattern [attribute value] 604, the graph ID of the selected experimental process data is registered in the graph ID column, the pattern name selected in S1503 is registered in the T column, the names of the previous and next substances (source substance and target substance) identified in S1612 are registered in the O1 column and the On column, the physical property items and physical property values ​​of the previous and next substances identified in S1612 are registered in the t1 item column, the t1 value column, the tn item column and the tn value column, and for each edge in the graph portion including the frequent process pattern identified in S1601 and the previous and next substances identified in S1612, the edge ID, the name of the start node, and the name of the end node are registered in the # column, the S column and the E column. After S1612, the process returns to S1611.

[0084] If there is no unselected experimental process data (S1611: NO), the process proceeds to S1613.

[0085] The structuring processing unit 120 searches for a common structure from the graph portion including the frequent process pattern identified in S1601 and the preceding and following substances identified in S1612 (S1613), and determines whether a common structure exists (S1614). If the t1 item and the tn item are structural items (e.g., "Smiles") and the t1 value and the tn value are structural values, a common structure exists. If a common structure exists (S1614: YES), the structuring processing unit 120 calculates the structural difference between the t1 value and the tn value, and registers the calculated difference as a Δ value in the Δ value column (S1615).

[0086] If there is no common structure (S1614: NO), the structuring processing unit 120 searches for a common attribute from the graph portion including the frequent process pattern and the preceding and following substances identified in S1612 (S1616) and determines whether there is a common attribute (S1617). If the t1 item and the tn item are attribute items (e.g., "hardness") and the t1 value and the tn value are attribute values, there is a common attribute. If there is a common attribute (S1617: YES), the structuring processing unit 120 calculates the attribute difference between the t1 value and the tn value and registers the calculated difference as a Δ value in the Δ value column (S1618).

[0087] If there is no common attribute (S1617: NO), or after S1618, the structuring processing unit 120 outputs the frequent process pattern [attribute value] for the action (reaction) corresponding to the pattern name selected in S1503 (S1619).

[0088] FIG. 17 is a diagram illustrating an example of the flow of the process feature extraction process.

[0089] The structuring processing unit 120 registers the A (edge ​​prediction) node and the T (radical reaction) node in the process feature [structure] 605 (S1701). Note that in S1701, the A node and the T node in the process feature [structure] 605 may be connected to nodes such as O1 and On (nodes exemplified in FIG. 11 ), i.e., nodes corresponding to elements registered in the subgraph information column, flow feature column, path feature column, and chain feature column in the process feature [attribute value] 606. Also, a blank process feature [attribute value] 606 may be prepared.

[0090] Next, the structuring processing unit 120 determines whether there is any unselected experimental process data among the experimental process data acquired in S1502 (S1702). "Unselected experimental process data" refers to experimental process data that has not yet been selected in this process feature extraction processing from the experimental process data acquired in S1502.

[0091] If there is unselected experimental process data (S1702: YES), the structuring processing unit 120 selects the unselected experimental process data and performs experimental process information registration processing, which is registration processing for the selected experimental process data (S1703). For the selected experimental process data, the process feature [attribute value] 606 may register, in the vi column, vj column, O1 column, On column, t1 column, and tn column, the node name of vi, the node name of vj, the name of O1, the name of On, the name of source material t1, and the name of target material tn for each edge in the graph portion including the frequent process pattern identified in S1601 and the preceding and following substances identified in S1612.

[0092] Next, the structuring processing unit 120 determines whether the data obtained as a result of the experimental process information registration process for the selected experimental process data matches the frequent process pattern identified in S1601 (S1704). If the determination result in S1704 is true (S1704: YES), the structuring processing unit 120 registers "1" in the T column of the process feature [attribute value] 606 (S1705) and performs frequent process pattern matching (S1706). Then, the process returns to S1702.

[0093] If the determination result in S1704 is false (S1704: NO), the structuring processing unit 120 sets “0” (or an invalid value) in the T column, Pa column, Pd column, Δ value, and Chain column of the process feature amount [attribute value] 606. Then, the process returns to S1702.

[0094] If there is no unselected experimental process data (S1702: NO), the structuring processing unit 120 outputs the process feature amount [structure] 605 and the process feature amount [attribute value] 606 (S1708).

[0095] FIG. 18 is a diagram showing an example of the flow of the experiment process information registration process.

[0096] The structuring processing unit 120 determines whether there is an unselected node pair (vi, vj) in the process feature [attribute value] 606 (S1801). An "unselected node pair" is a node pair (vi, vj) in the process feature [attribute value] 606 that has not been selected in this experimental process information registration process.

[0097] If there is an unselected node pair (vi, vj) (S1801: YES), the structuring processing unit 120 selects the unselected node pair (vi, vj) and determines whether nodes vi and vj in the selected node pair are connected by an edge (S1802). If the determination result of S1802 is true (S1802: YES), the structuring processing unit 120 registers the A value of the node pair (vi, vj) as "1" in the A column (S1803). If the determination result of S1802 is false (S1802: NO), the structuring processing unit 120 registers the A value of the node pair (vi, vj) as "0" in the A column (S1804).

[0098] After S1803 or S1804, for the node pair (vi, vj), the structuring processing unit 120 registers the feature vector of the subgraph G1 having the node vi in ​​the G1 column (S1805), and registers the feature vector of the subgraph G2 having the node vj in the G2 column (S1806).

[0099] FIG. 19 is a diagram showing an example of the flow of the frequent process pattern matching process.

[0100] The structuring processing unit 120 searches for a node pair (vi, vj) in a frequent process pattern (S1901), and determines whether an edge in the found node pair (vi, vj) is included in the frequent process pattern (S1902).

[0101] If the determination result of S1902 is true (S1902: YES), the structuring processing unit 120 determines whether the ancestor node of node vi in ​​the found node pair (vi, vj) is the node O1 (S1903). If the determination result of S1903 is true (S1903: YES), the structuring processing unit 120 registers the Pa value of the node pair (vi, vj) as "1" in the Pa column (S1904). If the determination result of S1903 is false (S1903: NO), the structuring processing unit 120 registers the Pa value of the node pair (vi, vj) as "0" in the Pa column (S1905).

[0102] After S1904 or S1905, the structuring processing unit 120 determines whether a descendant node of node vj in the found node pair (vi, vj) is an On node (S1906). If the determination result of S1906 is true (S1906: YES), the structuring processing unit 120 registers the Pd value of the node pair (vi, vj) as "1" in the Pd column (S1907). If the determination result of S1906 is false (S1906: NO), the structuring processing unit 120 registers the Pd value of the node pair (vi, vj) as "0" in the Pd column (S1908).

[0103] After S1907 or S1908, the structuring processing unit 120 registers the difference between the physical property value of t1 and the physical property value of tn as a Δ value in the Δ value column (S1909).The structuring processing unit 120 registers the edge ID, the node name of the start node, and the node name of the end node in the # column, S column, and E column as edge information in the frequent process pattern (S1910).

[0104] FIG. 20 is a diagram illustrating an example of the flow of the graph construction process.

[0105] The structuring processing unit 120 acquires the causal model generated in the causal model generation process of the learning process (S2001). A causal model may exist for each treatment (response).

[0106] The structuring processing unit 120 acquires the raw material-product table 701 from, for example, the user terminal 101 or the structuring support database 130 (S2002).

[0107] The structuring processing unit 120 generates experimental process data (subgraph) 702 consisting of one node corresponding to each raw material indicated by the M column in the raw material-product table 701 acquired in S2002, and registers the subgraph IDs of the experimental process data in the subgraph list 1400 (S2003). The subgraph list 1400 may be stored in the structuring support database 130.

[0108] The structuring processing unit 120 determines whether there are two or more unselected subgraphs in the subgraph list 1400 (S2004). An "unselected subgraph" is a subgraph that corresponds to a subgraph ID that has not been selected in this graph construction process, among the subgraph IDs in the subgraph list 1400.

[0109] If there are two or more unselected subgraphs (S2004: YES), the structuring processing unit 120 selects one of the unselected subgraphs and performs process feature inference processing (S2005) and edge prediction processing (S2006). The structuring processing unit 120 updates the subgraph based on the edge prediction result and registers the subgraph ID of the updated subgraph in the subgraph list 1400. This increases the number of unselected subgraphs. Then, the process returns to S2004.

[0110] If there are not two or more unselected subgraphs (S2004: NO), the subgraph list 1400 contains the ID of one unselected subgraph. The one unselected subgraph is a subgraph (a graph constructed as a result of two or more subgraphs connected by edges) expanded as a result of inductive inference in the graph construction process. The structuring processing unit 120 outputs structured data representing the subgraph (graph) as experimental process data (S2007).

[0111] FIG. 21 is a diagram illustrating an example of the flow of the process feature amount inference process.

[0112] The structuring processing unit 120 compares the M column (raw material name) and P column (product name) of the raw material-product table 701 acquired in S2002 with the source substance name (O1 value) and target substance name (On value) identified from the frequent process pattern [attribute value] 604 output in the learning process, identifies a possible action T and its application condition Δ value, and registers the value of action T and the Δ value in the T column and Δ value column of the process feature [attribute value] 706 (S2101). A "possibly applicable action T" is an action (reaction) corresponding to a pair of source substance name and target substance name that matches the pair of raw material name and product name, and "1" is registered as the value of such action T in the T column. An "application condition Δ value" is a Δ value corresponding to a pair of physical property values ​​of the source material and target substance, and is registered in the frequent process pattern [attribute value] 604. In the process feature [attribute value] 706, for a pair of the subgraph G1 selected in S2004 and each subgraph G2 other than the subgraph G1, the node name of the node vi (e.g., a node corresponding to a raw material), the node name of the node vj (e.g., a node corresponding to a product), the feature vector of the subgraph G1, the feature vector of the subgraph G2, O1 (raw material name), On (product name), t1 (raw material name), and tn (product name) may be registered in the vi column, vj column, G1 column, G2 column, O1 column, On column, t1 column, and tn column of the process feature [attribute value] 706.

[0113] Next, the structuring processing unit 120 determines whether a node corresponding to the raw material t1 represented by the raw material-product table 701 is included in the experimental process data (subgraph) 702 (S2102). If the determination result of S2102 is true (S2102: YES), the structuring processing unit 120 registers the Pa value corresponding to the pair of the raw material t1 and product tn as "1" in the Pa column (S2103). If the determination result of S2102 is false (S2102: NO), the structuring processing unit 120 registers the Pa value corresponding to the pair of the raw material t1 and product tn as "0" in the Pa column (S2104).

[0114] After S2103 or S2104, the structuring processing unit 120 determines whether a node corresponding to the product tn represented by the material-product table 701 is included in the experimental process data (subgraph) 702 (S2106). If the determination result of S2106 is true (S2106: YES), the structuring processing unit 120 registers the Pd value corresponding to the pair of the material t1 and the product tn as "1" in the Pd column (S2107). If the determination result of S2106 is false (S2106: NO), the structuring processing unit 120 registers the Pd value corresponding to the pair of the material t1 and the product tn as "0" in the Pd column (S2108).

[0115] After S2107 or S2108, if the frequent process pattern [structure] 603 obtained in the learning process is included in the experimental process data (subgraph) 702, the structuring processing unit 120 registers position information for each edge (edge ​​ID, start node name, and end node name) in the Chain column of the process feature [attribute value] 706 based on the frequent process pattern [attribute value] 604 corresponding to the frequent process pattern [structure] 603 in the frequent process pattern (S2109).

[0116] This process feature inference process generates a process feature [attribute value] 706. The process feature [structure] 605 is data generated in the learning process, as described above.

[0117] FIG. 22 is a diagram illustrating an example of the flow of edge prediction processing.

[0118] The structuring processing unit 120 inputs the experimental process data (subgraph) 702 to the encoder 611 that has been trained in the training process, and thereby acquires graph features from the encoder 611 (S2201).

[0119] The structuring processing unit 120 inputs the graph feature generated in S2201, the process feature [structure] 605 generated in the learning process, and the values ​​(explanatory variables (vi, vj, G1, and G2) and confounding factors) obtained from the process feature [attribute value] 706 generated in the process feature inference process to a decoder 612 (causal model [E=f(ZG, process feature)]) trained in the learning process, thereby acquiring an edge prediction table 706 from the decoder 612 (S2202). The edge prediction table 706 represents two values, indicating whether or not a subgraph G1 (a subgraph including a node corresponding to a raw material) as the experimental process data (subgraph) 702 is connected by an edge to another subgraph G2 (a subgraph including a node corresponding to a product).

[0120] The above processing can be broadly divided into learning processing and inference processing and explained again as follows as an example.

[0121] Fig. 23 is a diagram schematically illustrating an example of a portion of the learning process. Fig. 24A shows an example of experimental process data selected in this example. Fig. 24B shows a frequent process pattern [structure] representing a frequent process pattern (radical reaction) extracted in this example from a portion of the experimental process represented by the experimental process data shown in Fig. 24A. Fig. 24C shows a frequent process pattern [attribute value] generated in this example. Fig. 24D shows a process feature [attribute value] generated in this example.

[0122] The frequent process pattern [structure] corresponds to a record having a graph ID of "1" in the frequent process pattern [attribute value]. 4 ) and On and tn(CH 3 -CH 3 )

[0123] Furthermore, each edge of the experimental process data corresponds to each record of the process feature [attribute value]. The first to fourth records of the process feature [attribute value] correspond to four edges in the graph portion (graph portion shown in FIG. 24B) including the frequent process pattern (radical reaction) and the substances before and after it. Chain 1 to Chain 4 shown in FIG. 24B represent the chain features of the four edges.

[0124] FIG. 25 is a diagram showing a schematic diagram of an example of a part of the inference process. FIG. 26A is a diagram showing an example of a raw material-product table given in this example. FIG. 26B is a diagram showing an example of subgraph G1 and subgraph G2 determined from the raw material-product table. FIG. 26C shows the process feature quantities [attribute values] generated in this example. The frequent process pattern [structure] and frequent process pattern [attribute values] used in the inference process of this example are data generated in the learning process. Candidates for subgraph G2 are (1) operation node candidates (assumed to be a finite number enumerated in advance), (2) substance nodes that appear in the frequent process patterns (for example, Cl in the example of FIG. 9), 2 ), or (3) the substance of the product (P name) in the raw material-product table 701.

[0125] In this example, the raw material (CH 4 ) is a subgraph G1, and one node for each element (e.g., substance or operation) in the frequent process pattern [structure] obtained in the learning process is a subgraph G2. In the inductive reasoning, subgraph G1 (in the example shown in FIG. 26B, initially one CH 4 The next node (subgraph G2) of the first node (subgraph G1) is predicted, and any subgraph G2 is connected to the subgraph G1 with an edge to expand the subgraph G1. This process is repeated until experimental process data representing a predicted experimental process for producing a product from raw materials is obtained.

[0126] Each record of the process feature [attribute value] illustrated in FIG. 26C is 4These correspond to four nodes (four subgraphs G2) that may be connected to the node G1.

[0127] Furthermore, since it is found that the M name and P name in the raw material-product table match the substances before and after the frequently occurring process pattern of the radical reaction, the structuring processing unit 120 registers "1" as the T value, Pa value, and Pd value in the T column, Pa column, and Pd column of the process feature [attribute value] illustrated in FIG. 26C.

[0128] As for the chain feature, the structuring processing unit 120 determines whether the source material (material corresponding to the starting node) of the frequent process pattern is CH 4 In other words, the most likely node among the four records is predicted using Chain 1. In this case, the starting node of G1 (in this case, "CH 4 By searching the Chain-S column of the frequent process pattern [attribute value] 604 for the edge ID=1, it is possible to identify before edge prediction that the next candidate is "mixed." 4 When there are a plurality of "," by setting the edge IDs in order from the one closest to the start node, it is possible to uniquely set candidates in order from the smallest edge ID.

[0129] In Example 2, measurement process data is generated instead of the experimental process data shown in Example 1. In the description of this Example, differences from Example 1 will be mainly described, and descriptions of commonalities with Example 1 will be omitted or simplified.

[0130] Fig. 27 is a diagram illustrating an example of measurement process data according to Example 2. Fig. 28 is a diagram illustrating an example of a frequent process pattern [structure] according to Example 2.

[0131] Example 2 is applied to a task of measuring the performance of a sample (e.g., a product). For example, a measurement department uses a measurement device to measure a sample test piece to see if the target performance is achieved, and creates a report of the measurement results. Generally, the contents of the report vary depending on the person in charge, and there are cases where specifications (details) such as the type of labware used (e.g., the material of the container in which the sample test piece is placed) are not included. When this report is reused as know-how in another measurement task, being able to predict the type of omitted labware improves reusability and is expected to improve work efficiency. Example 2 makes it possible to predict the type of labware in the measurement process.

[0132] The frequent process pattern [structure] is structured data that represents a frequent process pattern as a partial structure common to measurement processes. The frequent process pattern illustrated in FIG. 28 is a know-how pattern that states, "In DSC measurement of PET (polyethylene resin), the type of container to put the test piece in is determined by whether it evaporates, the temperature, and the pressure conditions." DSC stands for Differential Scanning Calorimetry.

[0133] FIG. 29 is a diagram illustrating an example of a frequent process pattern [attribute value] according to the second embodiment.

[0134] The frequent process pattern [attribute value] according to Example 2 has, instead of a Δ value column, an evaporation column, a temperature column, and a pressure column in which the presence or absence of evaporation, the temperature value, and the pressure value are registered as selection conditions for tn. The graph ID is the ID of the measurement process. Furthermore, in Example 1, O1 and t1 are the same substance, and On and tn are the same substance, but in Example 2, O1 and t1 are different, and On and tn are different. Specifically, in Example 2, O1 is the substance, On is the measurement method, t1 is the container, and tn is the type of container.

[0135] For example, according to the first record, in the measurement process corresponding to graph ID "1", by applying calorimetric measurement device know-how, it can be seen that an aluminum cell was selected as the sample container type when performing DSC measurement on PET because the target PET does not evaporate and the temperature is 600°C or less.

[0136] The frequent process pattern [structure] and the frequent process pattern [attribute value] illustrated in FIGS. 28 and 29 may be input by the user via the user terminal 101.

[0137] FIG. 30 is a diagram illustrating an example of a process feature amount [structure] according to the second embodiment.

[0138] According to the process feature [structure] according to the second embodiment, the explanatory variables are explanatory variables extracted from the measurement process (for example, a procedure such as heat measurement). The objective variable is the probability of an edge existing, as in the first embodiment.

[0139] The example shown in Figure 30 is as follows: A (edge ​​existence probability) is predicted when T (calorimetry know-how) is applied between node vi in ​​subgraph G1 of the measurement process and node vj in subgraph G2. It indicates whether evaporation, temperature, and pressure are met for sample O1 and object On. The Pa value indicates whether subgraph G1 exists on the path between prediction target t1 and prediction result tn, i.e., whether prediction target t1 exists as an ancestor. The Pd value indicates whether subgraph G2 exists on the path between prediction target t1 and prediction result tn, i.e., whether prediction result tn exists as a descendant. Chain indicates the positions of subgraphs G1 and G2 in the frequent process pattern of procedure T.

[0140] FIG. 31 is a diagram illustrating an example of a process feature amount [attribute value] according to the second embodiment.

[0141] The process feature amount [attribute value] according to the second embodiment has an evaporation column, a temperature column, and a pressure column instead of the Δ value, similar to the frequent process pattern [attribute value].

[0142] FIG. 32 is a diagram illustrating an example of the flow of the frequent process pattern matching process according to the second embodiment.

[0143] The structuring processing unit 120 searches for a node pair (vi, vj) in a frequent process pattern (S2901), and determines whether an edge in the found node pair (vi, vj) is included in the frequent process pattern (S2902).

[0144] If the determination result of S2902 is true (S2902: YES), the structuring processing unit 120 determines whether the ancestor node of node vi in ​​the found node pair (vi, vj) is the node O1 (S2903). If the determination result of S2903 is true (S2903: YES), the structuring processing unit 120 registers the Pa value of the node pair (vi, vj) as "1" in the Pa column (S2904). If the determination result of S2903 is false (S2903: NO), the structuring processing unit 120 registers the Pa value of the node pair (vi, vj) as "0" in the Pa column (S2905).

[0145] After S2904 or S2905, the structuring processing unit 120 determines whether a descendant node of node vj in the found node pair (vi, vj) is an On node (S2906). If the determination result of S2906 is true (S2906: YES), the structuring processing unit 120 registers the Pd value of the node pair (vi, vj) as "1" in the Pd column (S2907). If the determination result of S2906 is false (S2906: NO), the structuring processing unit 120 registers the Pd value of the node pair (vi, vj) as "0" in the Pd column (S2908).

[0146] After S2907 or S2908, the structuring processing unit 120 registers the presence or absence of evaporation, the temperature value, and the pressure value in the evaporation, temperature, and pressure columns of the process feature [attribute value] (S2909). The presence or absence of evaporation, the temperature value, and the pressure value may be values ​​calculated by comparing each value of the measured process data with the conditions of the frequent process pattern [attribute value]. Since the values ​​of the "evaporation," "temperature," and "pressure" columns of the process pattern are conditions, in S2909, the structuring processing unit 120 determines whether each value of the measured process data matches the conditions. For example, according to the measured process data (during learning) illustrated in FIG. 27, since the "state is gas," evaporation occurs, and therefore "1" may be registered for evaporation. Since the "temperature is 700°C," "2" may be registered for the temperature, as it is 600°C or higher (above the upper limit). Since the "pressure is 0.5 MPa," a "1" may be registered for the pressure, which is 0.3 to 5 MPa (above the lower limit and below the upper limit). In other words, the value for evaporation may be either "0" (no evaporation) or "1" (evaporation), and the values ​​for temperature and pressure may be either "0" (below the lower limit), "1" (between the upper and lower limits), or "2" (above the upper limit).

[0147] The structuring processing unit 120 registers the edge ID, the node name of the start node, and the node name of the end node in the #, S, and E columns as edge information in the frequent process pattern (S2910).

[0148] According to the above-described first and second embodiments, for example, the following is true.

[0149] 33A is a diagram schematically illustrating an example of matching (matching between frequent process patterns and experimental process data) according to Example 1. FIG. 33B is a diagram schematically illustrating an example of matching (matching between frequent process patterns and measurement process data) according to Example 2.

[0150] In the example of the experimental process according to Example 1, in order to predict a pathway from raw materials to a product, O1 / On, which is the reaction objective that causes a chemical reaction, and t1 / tn, which are the two ends of the pathway predicted using the frequent process pattern, are the same. On the other hand, in the example of the measurement process according to Example 2, information other than the predicted pathway is included in the frequent process pattern, and O1 / On and t1 / tn are different.

[0151] In addition, in Example 1, the flow feature may include a Δ value as the difference between the physical property value of t1 and the physical property value of tn, and in Example 2, may include threshold values ​​for the presence or absence of evaporation, the temperature value, and the pressure value.

[0152] As described above, according to the first and second embodiments, a causal model can be constructed using less learning data and pattern descriptions by using frequent process patterns and process features. Furthermore, modeling of process determinants improves the prediction accuracy of unknown processes.

[0153] The present invention is not limited to the above-described embodiments, but includes various modifications. For example, the above-described embodiments are provided to explain the present invention in detail, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, some of the configurations of each embodiment can be added to, deleted from, or replaced with other configurations.

[0154] Furthermore, some or all of the above-described configurations, functions, processing units, processing means, etc. may be implemented in hardware, for example, by designing them as integrated circuits. The present invention can also be realized by software program code that implements the functions of the embodiments. In this case, a storage medium on which the program code is recorded is provided to a computer, and a processor included in the computer reads the program code stored in the storage medium. In this case, the program code itself read from the storage medium implements the functions of the above-described embodiments, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media for providing such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, solid-state drives (SSDs), optical disks, magneto-optical disks, CD-Rs, magnetic tape, non-volatile memory cards, and ROMs.

[0155] Furthermore, the program code that realizes the functions described in this embodiment can be implemented in a wide range of program or script languages, such as assembler, C / C++, perl, Shell, PHP, Python, Java (registered trademark), etc.

[0156] Furthermore, the program code of the software that realizes the functions of the embodiments may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the processor of the computer may read and execute the program code stored in the storage means or the storage medium.

[0157] In the above-described embodiment, the control lines and information lines are those that are considered necessary for the explanation, and not all control lines and information lines are necessarily shown in the product. All components may be interconnected.

[0158] The above description can be summarized, for example, as follows. The following summary may include supplementary explanations and explanations of variations of the above.

[0159] The structuring processing device 100 outputs structured data of a process for which a prediction target is predicted based on process features, which are features based on coarse granularity, such as process patterns of actions applied to a process, rather than on fine granularity, such as elements in a process having multiple elements including multiple procedures. This reduces the number of combinations and takes into consideration factors external to the elements in the prediction, thereby reducing the computational load of processing related to the prediction of the process represented by the generated structured data and improving the prediction accuracy of the process.

[0160] The structuring processing device (e.g., the structuring processing device 100) has a storage device (e.g., a main storage device 202 and an auxiliary storage device 203) and a calculation device (e.g., the calculation device 201). The storage device stores structured action process pattern data (e.g., a frequent process pattern [structure] 603) representing a graph of an action process pattern (e.g., a frequent process pattern) for each of one or more actions, and action process pattern attribute value data (e.g., a frequent process pattern [attribute value] 604) that includes attribute values ​​related to elements represented by the action process pattern structured data. For each of one or more actions, the action process pattern is a process portion having a structure common to two or more processes. For each of one or more actions, the action process pattern attribute value data includes attribute values ​​related to a destination element, which is an element to which the action is applied, a result element, which is an element resulting from the application of the action, a source element, which is an element to which the action is applied, and a target element, which is an element when the action is applied to the source element. The calculation device performs the following steps (A) to (F). (A) for a first subgraph including a source element of a process including an objective element, a result element, and a source element, in which a path from the source element to a target element is to be predicted, a second subgraph that can be an element of the path is selected from one or more second subgraphs. (B) based on the action process pattern structured data and action process pattern attribute value data related to an objective element (e.g., O1) and a result element (e.g., On) that match the objective element and result element of the process, process feature attribute value data (e.g., process feature [structure] 605) related to the process feature of a target process that is a process including the first subgraph (e.g., G1) and the selected second subgraph (e.g., G2) is generated.(C) The graph feature (e.g., graph feature Zg) of the structured data as graph data including the data of the first subgraph and the value obtained from the process feature structured data and the process feature attribute value data (e.g., process feature in the causal model [E = f(ZG, process feature)]) are input into a causal model that determines whether an edge exists between a node in the first subgraph (e.g., vi) and a node in the second subgraph (e.g., vj), thereby obtaining a predicted result (e.g., edge prediction table 708) indicating whether an edge exists between a node in the first subgraph and a node in the second subgraph. (D) It is determined whether prediction of the prediction target has been completed in the updated process to which the result is applied (e.g., S2004 in FIG. 20 ). (E) If the determination result of (D) is false, (A) is performed for the updated process. (F) If the determination result of (D) is true, structured data representing the updated process is output.

[0161] This reduces the number of combinations and allows factors external to the elements to be taken into account in the prediction, thereby reducing the computational load of the process related to the prediction of the process represented by the generated structured data and improving the accuracy of the process prediction.

[0162] The process features may include at least one of flow features, path features, and chain features. The flow features may be features related to changes that occur when a target element is processed to obtain a result element. The path features may be features related to a path including a source element and a target element. The chain features may be features related to a position on the path where a pair of a node in a first subgraph and a node in a second subgraph exists. Such process features contribute to reducing the computational load of processes related to process prediction and improving the accuracy of process prediction.

[0163] For example, the process having a prediction target may be an experimental process. The objective element and the source element may be the same element, i.e., a first substance. The result element and the target element may be the same element, i.e., a second substance. The process feature may include a flow feature. The flow feature may be a value (e.g., a Δ value) representing the difference between a physical property value of a first substance and a physical property value of a second substance. Furthermore, the process feature may include a path feature instead of or in addition to the flow feature. The path feature may include a value representing whether an ancestor of a node in the first subgraph is a source element and a value representing whether a descendant of a node in the second subgraph is a target element. Furthermore, the process feature may include a chain feature instead of or in addition to at least one of the flow feature and the path feature. The chain feature may be a feature representing the position of a node in the first subgraph and a node in the second subgraph.

[0164] Also, for example, the process having the prediction target may be a measurement process. The objective element and the source element may be different elements. The result element and the target element may be different elements. The source element may be an instrument. The target element may be details of the instrument (e.g., specifications such as the type of instrument). The process feature may include a flow feature. The flow feature may be a value for each of one or more conditions related to the relationship between the objective element and the result element.

[0165] The causal model may be a model generated in a learning process and set as a trained decoder (e.g., decoder 612). Generation of a causal model may not be necessary in the inference process, and a causal model set as a decoder may be used. The graph feature may be output data from a trained encoder (e.g., encoder 611) to which structured data as graph data including data of the first subgraph is input. Such a learning process contributes to reducing the computational load of processes related to process prediction and improving the accuracy of process prediction.

[0166] For example, the trained encoder and the trained decoder may be trained using training data and correct answer data in a training process. The training data may be data including process data as graph data. The causal model may be a model based on process feature structured data and process feature attribute value data generated for the process based on action process pattern structured data and action process pattern attribute value data of an action process pattern of an action applied to the process represented by the process data. The correct answer data may be data including node pairs and edge presence / absence on a path from a source element to a target element in the process.

[0167] In the learning process, the arithmetic device may generate, for each of two or more processes, treatment process pattern structured data and treatment process pattern attribute value data of a treatment process pattern that is a process portion common to the two or more processes, based on process data as structured data for each of the two or more processes. For each of the two or more processes, the arithmetic device may generate process feature structured data and process feature attribute value data for the process, based on process data including an objective element, a result element, a source element (e.g., t1), and a target element (e.g., tn), graph features of the process data, and the treatment process pattern structured data and treatment process pattern attribute value data generated for the process. Based on the process feature structured data and process feature attribute value data, the arithmetic device may generate a causal model indicating whether edges exist between element nodes and set it as a decoder. For each of the two or more processes, the arithmetic device may train an encoder and a decoder using training data and correct answer data.

[0168] In (A), the first subgraph may be a subgraph that includes, as a single node, a raw material identified from raw material-product data, which is data representing pairs of raw materials and products. The second subgraph may be a subgraph that includes, as a single node, a substance or operation that may be required to obtain the product identified from the raw material-product data. The raw material may be the first substance. The product may be the second substance. Such retrosynthetic analysis makes it possible to predict an experimental process for producing a product from a raw material.

[0169] The structuring processing device 100 may further include an interface device (e.g., communication device 206) communicatively connected to the destination device. The computing device may provide the destination device with data including structured data representing the updated process as prediction result data, which is data related to the predicted process, via the interface device. For example, the destination device may be a user terminal 101 or an experimental device. For example, if the destination device is an experimental device, it is expected that a practical application will be realized in which the experimental device conducts an experiment based on the prediction result data.

[0170] In the above description, the term "interface apparatus" may refer to one or more interface devices. The one or more interface devices may be at least one of the following: - One or more I / O (Input / Output) interface devices. The I / O (Input / Output) interface device is an interface device for at least one of an I / O device and a remote display computer. The I / O interface device for the display computer may be a communication interface device. The at least one I / O device may be a user interface device, for example, either an input device such as a keyboard and a pointing device, or an output device such as a display device. - One or more communication interface devices. The one or more communication interface devices may be one or more homogeneous communication interface devices (e.g., one or more NICs (Network Interface Cards)) or two or more heterogeneous communication interface devices (e.g., a NIC and an HBA (Host Bus Adapter)).

[0171] Furthermore, the term "memory" refers to one or more memory devices, which are an example of one or more storage devices, and may typically be a primary storage device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.

[0172] Furthermore, the term "persistent storage device" may refer to one or more persistent storage devices, which are an example of one or more storage devices. The persistent storage device may typically be a non-volatile storage device (e.g., an auxiliary storage device), and specifically may be, for example, a hard disk drive (HDD), a solid state drive (SSD), a non-volatile memory express (NVME) drive, or a storage class memory (SCM).

[0173] Furthermore, the "storage device" may be at least one of memory and persistent storage device.

[0174] Furthermore, the "arithmetic unit" may be one or more processor devices. The at least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). The at least one processor device may be a single-core or multi-core. The at least one processor device may also be a processor core. At least one processor device may be a processor device in a broad sense, such as a circuit that is a collection of gate arrays written in a hardware description language that performs some or all of the processing (for example, an FPGA (Field-Programmable Gate Array), a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit)).

[0175] The elements represented by the treatment process pattern structured data may include an element as a purpose and an element as a condition. Specifically, in the example of the radical reaction in Example 1, the raw material (CH 4 ) and the product (CH 3 -CH 3 ) may correspond to the "purpose" (i.e., the purpose is to generate a product from the raw materials), and operations such as mixing and irradiation may correspond to the "conditions" (i.e., the conditions for causing a radical reaction are to perform operations in the order of mixing → irradiation). In the example of measurement know-how in Example 2, PET, DSC measurement, container, and type may correspond to the "purpose", and evaporation, temperature, and pressure connected to the condition node may correspond to the "conditions".

[0176] 100: Structured processing device

Claims

1. A system comprising: a storage device storing structured data of treatment process patterns representing a graph of a treatment process pattern for each of one or more treatments, and attribute value data of treatment process patterns including attribute values of elements represented by the structured data of treatment process patterns; and a computing device connected to the storage device, wherein, for each of the one or more treatments, the treatment process pattern is a process portion having a structure common to two or more processes, and the attribute value data of the treatment process patterns for each of the one or more treatments includes attribute values of an objective element, which is an element to which the treatment is applied, a result element, which is an element resulting from the application of the treatment, a source element, which is an element to which the treatment is applied, and a target element, which is an element when the treatment is applied to the source element, and wherein the computing device (A) selects, for a first subgraph including a source element of a process including an objective element, a result element, and a source element, a second subgraph that can be an element of the path from the source element to the target element, from one or more second subgraphs, (B) based on the disposition process pattern structured data and the disposition process pattern attribute value data related to the objective element and the result element that match the objective element and the result element of the process, generate process feature attribute value data, which is data including attribute values related to elements represented by the process feature structured data related to the process feature of the target process, which is a process including the first subgraph and the selected second subgraph; (C) inputting the graph feature of the structured data as graph data including the data of the first subgraph and the values obtained from the process feature structured data and the process feature attribute value data into a causal model for whether an edge exists between a node in the first subgraph and a node in the second subgraph, thereby obtaining a predicted result of whether an edge exists between a node in the first subgraph and a node in the second subgraph; (D) determining whether the prediction of the prediction target has been completed in the updated process, which is the process to which the result has been applied;(E) If the determination result of (D) is false, perform (A) on the updated process; and (F) If the determination result of (D) is true, output structured data representing the updated process.

2. The structured processing device according to claim 1, wherein the process features include at least one of flow features, path features, and chain features, the flow features being features relating to changes that occur when a target element is processed to obtain a result element, the path features being features relating to a path including a source element and a target element, and the chain features being features relating to a position on the path at which a pair of a node in the first subgraph and a node in the second subgraph exists.

3. The structured processing device according to claim 1, wherein the process having the prediction target is an experimental process, the objective element and the source element are the same element and a first substance, and the result element and the target element are the same element and a second substance.

4. The structuring processing device according to claim 3, wherein the process feature includes the flow feature, and the flow feature is a value representing the difference between a physical property value of the first substance and a physical property value of the second substance.

5. The structured processing device according to claim 1, wherein the process features include the path features, and the path features include a value representing whether an ancestor of a node in the first subgraph is the source element, and a value representing whether a descendant of a node in the second subgraph is the target element.

6. The structuring processing device according to claim 1, wherein the process features include the chain features, and the chain features are features of the positions of nodes in the first subgraph and nodes in the second subgraph.

7. The structured processing device according to claim 1, wherein the process having the prediction target is a measurement process, the objective element and the source element are different elements, the result element and the target element are different elements, the source element is an instrument, and the target element is details of the instrument.

8. The structuring processing device according to claim 7, wherein the process features include the flow features, and the flow features are values for one or more conditions relating to the relationship between the target element and the result element.

9. The structuring processing device according to claim 1, wherein the causal model is a model generated in a learning process and set as a trained decoder, and the graph features are output data from a trained encoder to which structured data as graph data including data of the first subgraph has been input.

10. The structured processing device according to claim 9, wherein the trained encoder and the trained decoder are trained using training data and correct answer data in the training process, the training data is data including process data as graph data, the causal model is a model based on process feature structured data and process feature attribute value data generated for the process based on treatment process pattern structured data and treatment process pattern attribute value data of a treatment process pattern of a treatment applied to the process represented by the process data, and the correct answer data is data including node pairs and edge presence on a path from a source element to a target element in the process.

11. The structuring processing device according to claim 10, wherein in the learning process, the arithmetic device generates, for each of the two or more processes, process pattern structured data and process pattern attribute value data of a process pattern that is a process portion common to the two or more processes, based on process data as structured data for each of the two or more processes; generates process feature structured data and process feature attribute value data for each of the two or more processes, based on process data including an objective element, a result element, a source element, and a target element, graph features of the process data, and the process pattern structured data and process pattern attribute value data generated for the process; generates a causal model indicating whether edges exist between element nodes, based on the process feature structured data and process feature attribute value data; sets the causal model as the decoder; and trains the encoder and the decoder for each of the two or more processes using the training data and the correct answer data.

12. In (A), the first subgraph is a subgraph that includes, as one node, a raw material identified from raw material-product data, which is data representing pairs of raw materials and products; the second subgraph is a subgraph that includes, as one node, a substance or operation that may be required to obtain a product identified from the raw material-product data; the raw material is the first substance; and the product is the second substance. A structured processing device as described in claim 3.

13. The structured processing device according to claim 1, further comprising an interface device communicatively connected to a destination device, wherein the computing device provides, via the interface device, data to the destination device that includes structured data representing the updated process as prediction result data, which is data relating to the predicted process.

14. The structured processing device according to claim 1, wherein the plurality of elements represented by the treatment process pattern structured data include an element as a purpose and an element as a condition.

15. (A) for a first subgraph including a source element of a process that includes an objective element, a result element, and a source element and whose path from the source element to a target element is to be predicted, a second subgraph that can be an element of the path is selected from one or more second subgraphs; (B) based on the action process pattern structured data and action process pattern attribute value data related to the objective element and result element that match the objective element and result element of the process, generate process feature attribute value data that is data including attribute values related to elements represented by process feature structured data related to process features of a target process that is a process that includes the first subgraph and the selected second subgraph; (C) inputting the graph feature of the structured data as graph data including the data of the first subgraph and the values obtained from the process feature structured data and the process feature attribute value data into a causal model that determines whether an edge exists between a node in the first subgraph and a node in the second subgraph, thereby obtaining a predicted result of whether an edge exists between a node in the first subgraph and a node in the second subgraph; (D) determining whether or not prediction of the prediction target has been completed in the updated process, which is the process to which the result has been applied; (E) if the determination result of (D) is false, performing (A) on the updated process; (F) if the determination result of (D) is true, outputting structured data representing the updated process; and the method comprises: for each of one or more actions, action process pattern structured data representing a graph of an action process pattern; and action process pattern attribute value data which is data including attribute values relating to elements represented by the action process pattern structured data; for each of the one or more actions, the action process pattern is a process portion having a structure common to two or more processes;A structured processing method, wherein for each of the one or more actions, the action process pattern attribute value data includes attribute values for a destination element, which is the element to which the action is applied, a result element, which is the element resulting from applying the action, a source element, which is the element to which the action is applied, and a target element, which is the element when the action is applied to the source element.

Citation Information

Patent Citations

  • Device and method for generating structured data relating to document representing plural procedures

    JP2024085624A

  • Estimating the effect of an action using a machine learning model

    WO2023141019A1

  • Scheduling program, scheduling device, and scheduling method

    JP2021149557A

  • Assistance device, generation device, analysis device, assistance method, generation method, analysis method,and program

    WO2021039175A1

  • Material development assistance device, material development assistance method, and material development assistance program

    WO2021124392A1