Factor estimation device, factor estimation method, process execution method, factor estimation system, and terminal device

WO2026204549A1PCT designated stage Publication Date: 2026-10-01JFE STEEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/010318
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-17
Publication Date
2026-10-01

Smart Images

  • Figure JP2026010318_01102026_PF_FP_ABST
    Figure JP2026010318_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This factor estimation device comprises: a knowledge graph acquisition means for acquiring a knowledge graph; a process data acquisition means for acquiring process data items collected from a process; a data connection means for associating, with respective nodes, process data items related to events of the nodes for the knowledge graph; a causal analysis means for calculating an index indicating strength of the causal relationship between the nodes with which the process data items are associated; an edge addition means for adding, to the knowledge graph and on the basis of the index, a second edge connecting nodes having a strong causal relationship, and setting the value of the index to the weight of the second edge; and a causal path presentation means for presenting a path in which the nodes are connected by the second edge as a causal path indicating the causal relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Factor estimation device, factor estimation method, process implementation method, factor estimation system, and terminal device

[0001] The present invention relates to a factor estimation device, a factor estimation method, a process implementation method, a factor estimation system, and a terminal device.

[0002] For products manufactured through complex processes, such as steel products, numerous factors influence quality and defects. Therefore, identifying the factors affecting quality and defects is extremely difficult. Traditional methods relied on data correlation analysis and the knowledge of veteran engineers to estimate these factors, but there is a need for methods that effectively utilize large amounts of data. In response, various methods for factor estimation are being explored using data-driven approaches and knowledge-sharing techniques.

[0003] For example, Patent Document 1 proposes a factor analysis device that allows for easy understanding of the relationships between measurement items in a product manufactured through numerous measurement steps, even without prior knowledge of the product. This device takes data of characteristic values ​​measured in the manufacturing process and their measurement order as input, and selects explanatory variables using multiple regression analysis. This creates a cause-and-effect diagram and automatically visualizes the hierarchical relationships between numerous measurement items, making it possible to intuitively understand how each measurement item influences others.

[0004] Furthermore, Patent Documents 2 and 3 disclose a technology that describes process events using nodes, constructs a knowledge graph (knowledge model) by describing the causal relationships between nodes using a directed graph, and associates anomaly indicators with each node to realize the estimation of the causes when anomalies occur.

[0005] Furthermore, in Patent Documents 2 and 3, when identifying a causal path (causal route), either a causal path consisting of nodes where all process events indicate abnormality is identified, or a causal path is identified based on the sum of the abnormality levels of process events among multiple causal paths from cause to effect.

[0006] Japanese Patent Publication No. 2019-194849, Japanese Patent Publication No. 7578200, Japanese Patent Publication No. 7578199

[0007] In Patent Document 1, when estimating factors from among numerous items in a product manufacturing process, variables with high correlation were selected from the explanatory variables in a multiple regression analysis. Therefore, when using a regression equation in this way to find highly correlated variables, there remained the possibility that the correlation between variables was a spurious correlation.

[0008] In other words, even when there is no causal relationship between variables, the influence of other variables or coincidences could make it appear as if there is a correlation. In such cases, conventional techniques had the problem that the visualized results may not reflect the correct causal relationship.

[0009] Furthermore, when estimating factors, not only data correlations but also domain knowledge held at each site is important. In Patent Documents 2 and 3, such knowledge was estimated by utilizing the causal relationships between process events in a knowledge graph that represents the domain knowledge held in the minds of veteran engineers, and linking it with process event data. However, due to incompleteness in the knowledge graph, such as the relationships between nodes not being recorded or the nodes themselves being missing, there was a possibility of reaching incorrect conclusions.

[0010] Furthermore, when multiple factors interact with each other, even if each factor is within the normal range individually, anomalies may be caused by the interaction effect of these factors combined under specific conditions. However, in Patent Documents 2 and 3, causal paths were identified based on anomaly indicators (anomaly / normal classification determined based on the degree of deviation from the normal state) and anomaly scores (anomaly scores indicating the degree of anomaly) of data linked to nodes in a knowledge graph. As a result, there was a possibility of reaching incorrect conclusions.

[0011] The present invention has been made in view of the above, and aims to provide a factor estimation device, factor estimation method, process implementation method, factor estimation system, and terminal device that can resolve the problem of spurious correlation in multivariate analysis including multiple regression analysis, supplement the incompleteness of knowledge graphs, and improve the accuracy of factor estimation.

[0012] To solve the above-mentioned problems and achieve the objective, the factor estimation device according to the present invention comprises: knowledge graph acquisition means for acquiring a knowledge graph that expresses the causal relationship between phenomena and events in the process, comprising a plurality of nodes representing events occurring in the process and first edges connecting the nodes and representing causal relationships between the nodes; process data acquisition means for acquiring process data collected from the process; data linking means for linking the process data related to the events of the nodes to each node, targeting the knowledge graph; causal analysis means for calculating an index representing the strength of the causal relationship between the nodes to which the process data is linked; edge adding means for adding second edges connecting nodes with strong causal relationships to the knowledge graph based on the index and setting the value of the index as the weight of the second edge; and causal path presentation means for presenting paths in which the nodes are connected by the second edge as causal paths indicating the causal relationship.

[0013] Furthermore, the factor estimation device according to the present invention further comprises a factor estimation means for estimating the factors of a phenomenon in the process by ranking a plurality of causal paths based on the weights of the second edge.

[0014] Furthermore, the factor estimation device according to the present invention further comprises, in the above invention, an optimal condition identification means for identifying the optimal conditions for the second data corresponding to the first data in which the phenomenon does not occur, based on first data which is process data linked to the node corresponding to the phenomenon in the process and second data which is process data linked to the node corresponding to the event that is the cause of the phenomenon, and an output means for outputting the identified optimal conditions to the control device of the process.

[0015] Furthermore, in the factor estimation device according to the present invention, the data linking means uses natural language processing technology or a large-scale language model to link the process data related to the events of the nodes to each node based on the name of the process data.

[0016] Furthermore, in the factor estimation device according to the present invention, the data linking means compresses the dimensionality of the process data to be subjected to causal analysis and links it to each node.

[0017] Furthermore, in the factor estimation device according to the present invention, the data combining means uses principal component analysis or non-negative matrix factorization to reduce the dimensionality of the process data to be subjected to causal analysis.

[0018] Furthermore, in the factor estimation device according to the present invention, the process data acquisition means collects the process data acquired in a plurality of processes corresponding to the length positions of the metal material, and constructs an analysis data matrix Y in which the rows are at a predetermined pitch of length positions and the columns are items of actual data acquired in each process, and the data joining means targets the knowledge graph and links the items of the analysis data matrix Y related to the events of the nodes for each node, and extracts the column groups related to each node from the analysis data matrix Y to form a node-specific analysis data matrix Y node The node-specific analysis data matrix Y is composed of multiple items. node Non-negative matrix factorization is applied to reduce the dimensionality, extracting multiple basis vectors that represent the feature pattern in the length direction, and the causal analysis means performs causal analysis between the basis vectors of each extracted node.

[0019] Furthermore, in the factor estimation device according to the present invention, the data combining means is the node-specific analysis data matrix Y node Applying non-negative matrix factorization to the process data, we decompose it into non-negative matrices W and H, with each column of W being a basis vector representing a feature pattern in the length direction, and each row of H being the coefficient (greater than or equal to 0) of the basis component constituting the basis vector for each item of the process data, and the number of basis vectors k in the non-negative matrix factorization is the reconstruction matrix Y^ node and the original matrix Y node The mean squared error E(k) is evaluated as a function of the basis number k, and the inflection point where the error reduction saturates is determined as the optimal basis number k.

[0020] Furthermore, in the factor estimation device according to the present invention, the causal analysis means calculates the index using one of the following: a Bayesian network, a DAG-GNN, path analysis, or Fisher's exact test.

[0021] Furthermore, in the factor estimation device according to the present invention, the causal path presentation means presents the causal path superimposed on the knowledge graph.

[0022] To solve the above-mentioned problems and achieve the objective, the factor estimation method according to the present invention targets a knowledge graph that expresses the causal relationship between phenomena and events in a process, which is composed of a plurality of nodes representing events occurring in a process and first edges connecting the nodes and representing the causal relationship between the nodes, and includes a data joining step of linking process data related to the events of the nodes to each node; a causal analysis step of calculating an index representing the strength of the causal relationship between the nodes to which the process data is linked; an edge adding step of adding second edges to the knowledge graph that connect nodes with strong causal relationships based on the index and setting the value of the index as the weight of the second edge; and a causal path presentation step of presenting the paths in which the nodes are connected by the second edge as causal paths that indicate the causal relationship.

[0023] Furthermore, the factor estimation method according to the present invention further includes a factor estimation step, after the causal path presentation step, which estimates the factors of the phenomenon in the process by ranking the plurality of causal paths based on the weights of the second edge.

[0024] To solve the above-mentioned problems and achieve the objective, the process implementation method according to the present invention includes, for the causal path presented by the above-mentioned factor estimation method, an optimal condition identification step that identifies the optimal conditions for the second data corresponding to the first data in which the phenomenon does not occur, based on first data which is process data linked to a node corresponding to a phenomenon in the process and second data which is process data linked to the node corresponding to an event that is the cause of the phenomenon, and a process implementation step that implements the process based on the identified optimal conditions.

[0025] To solve the above-mentioned problems and achieve the objective, the factor estimation system according to the present invention includes a factor estimation device and a terminal device, and comprises: knowledge graph acquisition means for acquiring a knowledge graph that expresses the causal relationship between phenomena and events in the process, which is composed of a plurality of nodes representing events occurring in a process and first edges that connect the nodes and represent the causal relationship between the nodes; process data acquisition means for acquiring process data collected from the process; data linking means for linking the process data related to the events of the nodes to each node, targeting the knowledge graph; causal analysis means for calculating an index that represents the strength of the causal relationship between the nodes to which the process data is linked; edge adding means for adding second edges that connect nodes with strong causal relationships to the knowledge graph based on the index, and setting the value of the index as the weight of the second edge; and the node The system includes: causal path presentation means for presenting paths connected by the second edge as causal paths indicating the causal relationship; factor estimation means for estimating the factors of a phenomenon in the process by ranking a plurality of causal paths based on the weight of the second edge; optimal condition identification means for identifying the optimal conditions for the second data corresponding to the first data in which the phenomenon does not occur, based on first data which is process data linked to the node corresponding to the phenomenon in the process and second data which is process data linked to the node corresponding to the event which is the factor of the phenomenon; output means for outputting the identified optimal conditions to the control device of the process; display information output means for outputting information including at least the causal paths and the ranking of the causal paths; display information acquisition means for acquiring the information; and display means for displaying the information.

[0026] To solve the above-mentioned problems and achieve the objective, the terminal device according to the present invention comprises a display information acquisition means for acquiring information from a factor estimation device that includes at least a causal path and a ranking of the causal path, and a display means for displaying the information, wherein the causal path is a path in which nodes linked to process data related to the events of a process are connected by a second edge in a knowledge graph that expresses the causal relationship between phenomena in the process and the events of the events of the node, and the nodes linked to process data related to the events of the node are connected by a second edge, the second edge is an edge in the knowledge graph that connects nodes with strong causal relationships based on an index that represents the strength of the causal relationship between the nodes linked to the process data, and the ranking of the causal path is a ranking of a plurality of causal paths based on the weight of the second edge.

[0027] According to the present invention, by displaying a combination of the results of causal analysis between a knowledge graph and process data, it is possible to resolve the spurious correlation problem in multivariate analysis, including multiple regression analysis, and to compensate for the incompleteness of the knowledge graph, thereby improving the accuracy of factor estimation.

[0028] Figure 1 is a block diagram showing a schematic configuration of a factor estimation system including a factor estimation device according to an embodiment. Figure 2 is a flowchart showing the procedure of a factor estimation method executed by a factor estimation device according to an embodiment of the present invention. Figure 3 is a diagram showing an example of a knowledge graph used in the factor estimation method according to an embodiment of the present invention. Figure 4 is a diagram illustrating the data joining step in the factor estimation method according to an embodiment of the present invention. Figure 5 is a diagram illustrating the edge addition step and the causal path presentation step in the factor estimation method according to an embodiment of the present invention. Figure 6 is a diagram showing an example of the configuration of an analysis data matrix Y (m × n) in which quality data evaluated at a predetermined pitch (e.g., 1 m) in the length direction and operating conditions (actual data or set values) obtained at each process of the same product are joined in correspondence with the length position for a long material (e.g., steel plate coil, etc.). Figure 7 is a node-specific analysis data matrix Y node (i) (m x n)i ) to which non-negative matrix factorization (NMF) is applied, and the basis matrix W (i) (m×k i ) and the coefficient matrix H (i) (k i ×n i ) is a diagram illustrating the concept of decomposition. FIG. 8 is a diagram illustrating the concept of a procedure for evaluating the relationship between basis vectors (each column of W (i) ) extracted by non-negative matrix factorization (NMF) using basis intensity values for each length position. FIG. 9 is a schematic diagram illustrating the determination of adopting a relationship between bases (or between a base and a quality index) when a predetermined threshold is satisfied based on the evaluation results shown in FIG. 8. FIG. 10 is a diagram illustrating an example of causal path ranking in a causal path presentation step in a factor estimation method according to an embodiment of the present invention. FIG. 11 is a diagram illustrating an example of variation in C-camber (quality indicator) in the length direction (or time-converted) and a stable interval with small mean and variance in the factor estimation method according to an embodiment of the present invention. FIG. 12 is a diagram illustrating an example of variation in a basis intensity sequence (a column of W) of basis vectors (e.g., bases 1 to 8) in the factor estimation method according to an embodiment of the present invention. FIG. 13 is a histogram arranged in descending order of the contribution degree of the H matrix to extracted bases in the factor estimation method according to an embodiment of the present invention.

[0029] A factor estimation apparatus, a factor estimation method, a process execution method, a factor estimation system, and a terminal device according to an embodiment of the present invention will be described with reference to the drawings. The present invention is not limited to the following embodiments. In addition, components in the following embodiments include those that can be easily replaced by a person skilled in the art and those that are substantially identical.

[0030] The present invention mainly belongs to the technical field related to product manufacturing processes such as steel plates, and particularly relates to technology for estimating factors related to the quality and defects of final products in manufacturing processes. The technology according to the present invention can be widely used in industrial fields such as the steel material manufacturing industry and steel plate processing industry. In addition, the technology according to the present invention can also be used when performing analysis combining data and knowledge in a wide range of other fields, such as the medical field.

[0031] This invention provides a technology for accurately estimating defect causes using a causal analysis algorithm, by combining process data acquired from on-site equipment with a knowledge graph based on the knowledge of veteran engineers, in specific technical fields such as the manufacturing process of steel products.

[0032] The causal estimation device performs a series of algorithmic processes, including acquiring and analyzing process data, presenting causal paths, and identifying optimal conditions, in conjunction with hardware such as display devices and control devices on terminal equipment. Specifically, it presents the causal paths and optimal operating conditions estimated by the causal analysis algorithm to on-site personnel in an easy-to-understand manner via the display devices on terminal equipment, enabling on-site personnel to make quick decisions and respond to anomalies based on the displayed information.

[0033] Furthermore, the operating conditions extracted by the optimal condition identification means are automatically reflected in the equipment via the control device, contributing to more efficient on-site operation management and improved quality stability. In this way, the present invention specifically solves on-site technical challenges (such as identifying the causes of quality defects, accelerating responses to abnormalities, and optimizing operating conditions).

[0034] (Factor Estimation System) The configuration of the factor estimation system according to this embodiment will be described with reference to Figure 1. The factor estimation system is for inferring the factors of phenomena in a process. Examples of "processes" in this embodiment include the manufacturing process of steel products, the power generation process of power generation equipment, and the transport process of transport equipment.

[0035] As shown in Figure 1, the factor estimation system according to this embodiment comprises a factor estimation device 1, a terminal device 2, a storage device 3, and a control device 4. The factor estimation device 1, terminal device 2, storage device 3, and control device 4 are configured to communicate with each other via a network N. This network N is composed of, for example, a public network such as an internet line network or a mobile phone line network, or a WAN (Wide Area Network). Furthermore, the terminal device 2 and control device 4 of the factor estimation system are located, for example, within a steel mill.

[0036] Furthermore, the factor estimation system according to this embodiment can present causal paths and estimated factor information to on-site operators in real time via the display means (e.g., liquid crystal display, touch panel, etc.) of the terminal device 2. This enables operators to take quick countermeasures based on the displayed information, contributing to the efficiency of quality control and abnormal response in the manufacturing process. In addition, the control device 4 can automatically adjust the set values ​​of the manufacturing equipment (e.g., temperature, pressure, material supply amount, etc.) based on the operating conditions extracted by the optimal condition identification means 17. This reduces the risk of defect occurrence and contributes to the realization of stable product quality.

[0037] (Factor Estimation Device) The factor estimation device 1 is implemented by an information processing device such as a general-purpose computer like a workstation or personal computer, or a server located on the cloud. The information processing device that constitutes the factor estimation device 1 is equipped with a processing unit such as a CPU (Central Processing Unit).

[0038] The factor estimation device 1 functions as a knowledge graph acquisition means 10, a process data acquisition means 11, a data merging means 12, a causal analysis means 13, an edge addition means 14, and a causal path presentation means 15 when the above-mentioned arithmetic processing unit executes a computer program. Furthermore, the factor estimation device 1 functions as a factor estimation means 16, an optimal condition identification means 17, an output means 18, and a display information output means 19 when the above-mentioned arithmetic processing unit executes a computer program.

[0039] The knowledge graph acquisition means 10 acquires a knowledge graph that represents the causal relationships between phenomena and events in a process. This knowledge graph consists of, for example, multiple nodes representing events that occur in the process, and first edges that connect the nodes and represent the causal relationships between the nodes.

[0040] Here, "phenomena" in a process refers to physical or chemical changes or behaviors observed during the process, such as temperature changes, pressure fluctuations, and material reactions in the manufacturing process. These are consequences of the process and directly affect product quality, defects, and production line problems. "Events" in a process refer to occurrences or operations that occur under specific conditions within the process, such as equipment operation, setting changes, and anomaly detection in the manufacturing process. These events occur as the process progresses and influence the "phenomena" of the process.

[0041] The process data acquisition means 11 acquires process data collected from the process. This process data includes, for example, multiple steps in the process, numerous measured values, setpoints, product information, equipment information, etc. The process data may also include state variables calculated by numerical analysis using a physical model that represents the phenomena of the process, based on the acquired measured values ​​and setpoints.

[0042] The data merging means 12 applies the knowledge graph acquired by the knowledge graph acquisition means 10 and links the event data of each node to the associated process data. Here, the data merging means 12 may, for example, use natural language processing technology or a large-scale language model to link the event data of each node to the associated process data based on the name of the process data.

[0043] Furthermore, if there is a large amount of process data related to a node, the data merging means 12 may compress the dimensionality of the process data to be analyzed for causal analysis and link it to each node. Also, when compressing the dimensionality of the process data as described above, the data merging means 12 may use, for example, principal component analysis or non-negative matrix factorization to compress the process data to be analyzed for causal analysis.

[0044] The causal analysis means 13 calculates an index in the knowledge graph that represents the strength of the causal relationship between nodes to which process data is linked. Here, the causal analysis means 13 may calculate the above index using, for example, a Bayesian network, DAG-GNN, path analysis, or Fisher's exact test. When using these methods, the above index may be, for example, the conditional probability between nodes in a Bayesian network, the edge weights between nodes in a DAG-GNN, the path coefficients between nodes in path analysis, or the p-value of Fisher's exact test.

[0045] The edge addition means 14 adds a second edge to the knowledge graph that connects nodes with strong causal relationships, based on an index calculated by the causal analysis means 13. The edge addition means 14 then sets the weight of the second edge to the value of the above index.

[0046] The causal path presentation means 15 presents paths in which nodes are connected by second edges, as causal paths indicating causal relationships, using the edge addition means 14. Here, the causal path presentation means 15 may also present the above causal paths overlaid on a knowledge graph.

[0047] The factor estimation means 16 estimates the factors of phenomena in the process by ranking multiple causal paths based on the weights of the second edge.

[0048] The optimal condition identification means 17 identifies the optimal conditions for the second data corresponding to the first data where the phenomenon does not occur, based on the first data and second data used in the causal analysis. Here, the first data refers to process data linked to the node corresponding to the phenomenon in the process. The second data refers to process data linked to the node corresponding to the event that is the cause of the phenomenon.

[0049] The output means 18 outputs the optimal conditions identified by the optimal condition identification means 17 to the process control device 4. The display information output means 19 outputs information including at least the causal path and the ranking of the causal path to the terminal device 2. Details of the processing performed by each component of the factor estimation device 1 will be described later (see Figures 2 to 6).

[0050] Here, the process data acquisition means 11, data merging means 12, and causal analysis means 13 of the factor estimation device 1 may perform the following processing specific to long materials such as product coils. In this case, the process data acquisition means 11 collects process data acquired in multiple processes corresponding to the length positions of the metal material and constructs an analysis data matrix Y (size: m × n) in which the rows are predetermined pitches of length positions and the columns are items of actual data acquired in each process.

[0051] Next, the data merging means 12 targets the knowledge graph and links the items of the analysis data matrix Y related to the events of the nodes to each node, extracting the column groups related to each node from the analysis data matrix Y to create node-specific analysis data matrix Y node (i) (Size: m x n) i ) constitutes the data merging means 12 also comprises a node-specific analysis data matrix Y composed of multiple items. node (i) Non-negative matrix factorization is applied to reduce the dimensionality, and multiple basis vectors representing the feature pattern in the length direction are extracted.

[0052] The data joining means 12 specifically refers to the node-specific analysis data matrix Y node (i) Non-negative matrix factorization is applied to decompose the data into non-negative matrices W and H. This allows each column of W to be a basis vector representing a feature pattern in the length direction, and each row of H to be the coefficient (greater than or equal to 0) of the basis component constituting the basis vector for each item of the process data. Furthermore, the number of basis vectors k in the non-negative matrix factorization is set to the reconstruction matrix Y^ node (i) and the original matrix Y node (i) Mean squared error E(k) i The result is evaluated as a function of the basis number k, and the inflection point (so-called elbow) where the error reduction saturates is determined as the optimal basis number k.

[0053] Next, the causal analysis means 13 performs causal analysis on the basis vectors of each extracted node. Details of these processes will be explained later in "Factor Estimation Method: Embodiment 2".

[0054] (Terminal device) Terminal device 2 acquires and displays various information from factor estimation device 1 via network N. This terminal device 2 is implemented by an information processing device such as a general-purpose computer like a personal computer or a tablet computer. The information processing device that constitutes terminal device 2 is equipped with a processing unit such as a CPU.

[0055] Terminal device 2 functions as a display information acquisition means 21 when the above-mentioned processing unit executes a computer program. Terminal device 2 also includes a display means 22.

[0056] The display information acquisition means 21 acquires information from the factor estimation device 1 that includes at least causal paths and the ranking of said causal paths. Here, a "causal path" is a path in the knowledge graph in which nodes linked to event nodes and associated process data are connected by second edges. A "second edge" is an edge in the knowledge graph that connects nodes with strong causal relationships based on an index representing the strength of the causal relationship between nodes linked to process data. The "ranking of causal paths" is the ranking of multiple causal paths based on the weights of the second edges.

[0057] The display means 22 is implemented by a display such as a liquid crystal display (LCD), an organic light-emitting diode (OLED), or a touch panel display. The display means 22 displays the information acquired by the display information acquisition means 21. By using the display means 22 of the factor estimation device 1 and the terminal device 2, information including causal paths and the ranking of said causal paths, as well as the estimated optimal conditions, can be presented to the on-site operator in an easy-to-understand manner, enabling rapid decision-making and response to anomalies on-site.

[0058] (Storage Device) The storage device 3 is implemented by recording media such as EPROM (Erasable Programmable ROM), Hard Disk Drive (HDD), and removable media. Examples of removable media include USB (Universal Serial Bus) memory, CD (Compact Disc), DVD (Digital Versatile Disc), and BD (Blu-ray® Disc).

[0059] The storage device 3 can store the operating system (OS), various programs, various tables, various databases, etc. The storage device 3 also stores knowledge graphs. In addition to knowledge graphs, the storage device 3 may also store calculation results and process data from the factor estimation device 1, as needed.

[0060] (Control device) The control device 4 is for controlling the various equipment that executes the process. The control device 4 controls the various equipment and executes the process based on the optimal conditions identified by the optimal condition identification means 17 of the factor estimation device 1. In addition, the control device 4 can automatically adjust the set values ​​of the manufacturing equipment based on the operating conditions extracted by the optimal condition identification means 17. In this way, by automatically adjusting the operating conditions of the equipment via the control device 4, a practical effect is achieved that allows on-site operators to respond immediately.

[0061] Furthermore, the factor estimation device 1 and terminal device 2 of the factor estimation system do not necessarily have to include all of the components shown in Figure 1. Also, the factor estimation device 1 and terminal device 2 may each have components other than those shown in Figure 1. Moreover, the components included in the factor estimation device 1 and terminal device 2 are not limited to the examples in Figure 1. For example, some of the components included in the factor estimation device 1 in Figure 1 may be included in the terminal device 2.Therefore, the factor estimation system comprising the factor estimation device 1 and terminal device 2 may, as a whole, have a configuration comprising a knowledge graph acquisition means 10, a process data acquisition means 11, a data joining means 12, a causal analysis means 13, an edge addition means 14, a causal path presentation means 15, a factor estimation means 16, an optimal condition identification means 17, an output means 18, a display information output means 19, a display information acquisition means 21, and a display means 22.

[0062] (Factor Estimation Method: Embodiment 1) The factor estimation method according to Embodiment 1 will be described with reference to Figures 2 to 5. As shown in Figure 2, the factor estimation method according to the embodiment performs the following steps: knowledge graph acquisition step, process data acquisition step, data merging step, causal analysis step, edge addition step, causal path presentation step, factor estimation step, optimal condition identification step, and output step.

[0063] <Knowledge Graph Acquisition Step> In the knowledge graph acquisition step, the knowledge graph acquisition means 10 acquires the knowledge graph (ontology) as a knowledge database from the storage device 3 in which the knowledge graph is stored (step S1 in Figure 2).

[0064] A knowledge graph represents the causal relationships of phenomena in a process. It consists of multiple nodes representing events that occur in the process, and edges (first edges) that connect the nodes and represent the causal relationships between them. A knowledge graph is a systematically organized representation of factors related to product quality and defects, as well as their causal relationships, based on information extracted from the knowledge held in the minds of veteran engineers, as well as information from numerous internal company documents and literature, in a visually easy-to-understand manner.

[0065] Figure 3 shows an example of a knowledge graph used in this embodiment. The knowledge graph consists of three types of nodes: result nodes, intermediate nodes, and cause nodes, and first edges connecting each node.

[0066] An outcome node is a node that represents a phenomenon that occurs as a result of a process. In Figure 3, the phenomenon "defect X occurs" is described as an outcome node. Only one outcome node is provided in a single knowledge graph. Therefore, multiple knowledge graphs are created in advance for each phenomenon that occurs as a result of a process and stored in the memory device 3. For example, if the phenomenon that occurs as a result of the process is a defect in the steel plate, multiple knowledge graphs are created in advance for each type of defect (defect species) and stored in the memory device 3.

[0067] A cause node is a node that represents the event that causes the phenomenon shown in the result node. In a single knowledge graph, there can be one or more cause nodes. In Figure 3, the event of the state of materials or equipment is described as a cause node. A cause node is connected to one of the intermediate nodes by a first edge (see solid arrow). Note that Figure 3 shows an example where one cause node is connected to one intermediate node, but one cause node may be connected to two or more intermediate nodes.

[0068] An intermediate node is a node that represents an event that is occurring (or is likely to occur) between a phenomenon that occurs as a result of a process (result node) and the event that causes it (cause node). In a single knowledge graph, there may be one or more intermediate nodes. In Figure 3, intermediate nodes describe the reactions of material components, internal states, etc. An intermediate node is connected to at least one of the result node, another intermediate node, or cause node by a first edge (see solid arrow). Also, as shown in Figure 3A1, for example, there may be an intermediate node that is not connected to a cause node but is connected only to another intermediate node. Also, as shown in Figure 3A2, for example, there may be an intermediate node that is connected to a cause node but not to another intermediate node.

[0069] Furthermore, intermediate nodes are classified into nodes that describe events from which measured values ​​or setpoints can be obtained, and nodes that describe events from which measured values ​​or setpoints cannot be obtained. As will be described later, among the intermediate nodes, nodes that describe events from which measured values ​​or setpoints can be obtained are linked to process data. On the other hand, among the intermediate nodes, nodes that describe events from which measured values ​​or setpoints cannot be obtained are not linked to process data.

[0070] In the intermediate node shown in Figure 3, the following are examples of events for which measured values ​​and set values ​​can be obtained: (1) Defect factor Y becomes apparent in the hot rolling process. (2) Defect factor Y becomes apparent in the cold rolling process. (3) Defect factor Y exists inside the steel billet. (4) Component A on the surface reacts with component B in the molten steel to produce C. (5) Internal state G (6) Internal state J (7) Internal state K

[0071] Furthermore, in the intermediate nodes shown in Figure 3, examples of events that make it impossible to obtain measured values ​​or set values ​​include the following. These nodes are created, for example, when creating a knowledge graph, by estimating events that would have occurred between the upper and lower nodes based on the knowledge of veteran engineers, etc. (1) C enters the XX section.

[0072] In the knowledge graph shown in Figure 3, the nodes are arranged from bottom to top in the order of cause node, intermediate node, and effect node. Therefore, the time series in which events and phenomena occur also generally progresses from bottom to top. However, conversely to Figure 3, the knowledge graph may also be arranged from top to bottom in the order of cause node, intermediate node, and effect node.

[0073] <Process Data Acquisition Step> In the process data acquisition step, the process data acquisition means 11 collects a large amount of process data obtained from the manufacturing process (see step S2 in Figure 2). Since this process data consists of multiple processes, numerous measured values, set values, product information, equipment information, etc., it becomes a very high-dimensional data set for analysis.

[0074] <Data merging step> In the data merging step, the data merging means 12 performs an analysis procedure using the knowledge graph and the analysis dataset described above, linking each item (node) of the knowledge graph with the related process data (see step S3 in Figure 2).

[0075] In the data merging step, for example, process data is linked to nodes that make up the knowledge graph, specifically those nodes from which measured values ​​or setpoints of process data related to the events represented by the nodes can be obtained.

[0076] For example, the events in the cause nodes shown in Figure 3 all relate to the state of materials and equipment, and measured values ​​and setpoints can be obtained. Therefore, process data indicating the state of each material and equipment is linked to the cause nodes in Figure 3. Similarly, for the phenomena in the result nodes shown in Figure 3, measured values ​​can be obtained during inspections after product manufacturing. Therefore, process data is linked to the result nodes in Figure 3.

[0077] On the other hand, the events in the intermediate nodes shown in Figure 3 include events for which measured values ​​and setpoints can be obtained, and events for which they cannot. Therefore, in the intermediate nodes of Figure 3, the relevant process data is linked only to nodes that show events for which measured values ​​and setpoints can be obtained.

[0078] The linking of process data in the data merging step may be done manually, for example. On the other hand, if such manual linking of process data is very cumbersome, natural language processing techniques may be used to link the text described as information related to the node's events with the names of the operational items (also simply called "names") of each process data.

[0079] Examples of natural language processing techniques include TF-IDF (Term Frequency-Inverse Document Frequency) and word embedding thesauruses. For example, when applying TF-IDF, information about the events of each node (text) and the names of the operation items (text) may be vectorized, and the similarity using TF-IDF may be calculated to link them together.

[0080] Furthermore, when applying word embedding, word embedding technologies such as Word2Vec and GloVe can be used. Using these word embedding technologies, each word can be converted into a vector, and by averaging the vectors of the entire text, the vectors of information about the node's events (text) and the names of the operation items (text) can be compared, and the similarity can be calculated to link them.

[0081] Furthermore, when using a thesaurus, thesaurus technology, such as WordNe, which is a dictionary of synonyms and related words, can be used. Using such thesaurus technology, the more common synonyms and related words there are between the information about the node's events (text) and the name of the operation item (text), the higher the semantic relationship is judged to be, and the similarity can be calculated and linked based on that semantic relationship.

[0082] Alternatively, in the data merging step, a large-scale language model (generative AI) may be used to link process data related to node events. Specifically, information such as manufacturing knowledge in the product manufacturing process and knowledge about process equipment may be provided to the large-scale language model, and the linking may be performed through fine-tuning or RAG (Retrieval-Augmented Generation).

[0083] In the case of fine-tuning, the accuracy of a pre-trained large-scale language model is improved by retraining (fine-tuning) it using domain-specific information such as manufacturing knowledge and knowledge of process equipment in the product manufacturing process. Then, using the fine-tuned model, information about node events (text) and related process data are linked based on their names (text).

[0084] In the case of RAG, for example, as shown in Figure 4, information related to the text (event) described in the node is retrieved from a knowledge database of manufacturing processes and equipment. The retrieved information is then input into a large-scale language model, which generates a response based on that information. Using the generated response, the information (text) related to the node's event and the associated process data are linked based on their names (text).

[0085] In the data merging step, especially in fields like steel where there are many processes and data items, linking numerous data points to a single node can result in a high-dimensionality dataset. In such cases, causal analysis between process data becomes difficult in the subsequent causal analysis step.

[0086] Therefore, in the data merging step, if there are many process data related to (linkable to) a single node, dimensionality reduction techniques such as principal component analysis (PCA) or non-negative matrix factorization (NMF) are applied as needed to reduce the dimensionality of the process data to be analyzed for each node. The outlines of each dimensionality reduction technique are as follows.

[0087] Principal component analysis (PCA) is a technique for reducing dimensionality by using a linear transformation that maximizes the variance of data. In PCA, the main components are extracted by singular value decomposition of the data's covariance matrix. This makes it possible to capture the overall characteristics of the data.

[0088] Non-negative matrix factorization (NMF) is a technique that decomposes a non-negative data matrix into the product of two non-negative low-rank matrices. Because non-negative matrix factorization preserves the non-negativity of the data, the physical interpretation of the resulting components becomes easier. Furthermore, because non-negative matrix factorization allows for the decomposition of data into partial contributions, it becomes easier to understand the data structure and enables the decomposition and evaluation of data in terms of the influence of specific processes or equipment.

[0089] In the data merging step, if there is a large amount of process data associated with a single node, the method described above can be used to reduce the dimensionality of the process data used for causal analysis while preserving the important characteristics of the process data for each node.

[0090] <Causal Analysis Step> In the causal analysis step, the causal analysis means 13 performs causal analysis using process data (or feature quantities after dimensionality reduction) associated with each node of the knowledge graph, and calculates an index that represents the strength of the causal relationship between the nodes associated with the process data (see step S4 in Figure 2).

[0091] In the causal analysis step, when calculating the above indicators, causal analysis is performed not only between nodes connected by the first edge, but also between all nodes that make up the knowledge graph. For example, the "Internal State J" node shown in A1 of Figure 3 is connected only to the node above it, "Defect Factor Y becomes apparent in the cold rolling process," by the first edge (solid arrow). On the other hand, in the causal analysis step, causal analysis is also performed between the "Internal State J" node and the adjacent "Internal State K" node (see A2 in Figure 3), which is not connected by the first edge. Then, in the causal analysis step, the results of the causal analysis between each node are plotted on the knowledge graph.

[0092] Examples of causal analysis methods used in the causal analysis step include Bayesian networks, DAG-GNN (Directed Acyclic Graph - Graph Neural Network), path analysis, and Fisher's exactness test. An overview of each causal analysis method is as follows.

[0093] Bayesian networks are a method for estimating causal relationships between variables based on probabilistic models. In Bayesian networks, conditional probabilities indicating the strength and direction of causal relationships are assigned to edges.

[0094] DAG-GNN is a method that applies a neural network to data with a graph structure to estimate causal relationships. DAG-GNN can model nonlinear relationships and can be applied to high-dimensional data.

[0095] Path analysis is a statistical method that uses analytical techniques to numerically evaluate causal relationships, such as direct and indirect effects. In path analysis, the strength of a causal relationship can be represented by path coefficients.

[0096] Fisher's exact test is a statistical method for accurately testing the independence between two categorical variables. Fisher's exact test shows the probability, expressed as a p-value, that the observed data agree with the expected value of independence. A lower p-value indicates a stronger denial of independence and a stronger causal relationship.

[0097] The indicators representing the strength of causal relationships between nodes, calculated in the causal analysis step, refer to the conditional probabilities between nodes in the Bayesian network, the edge weights between nodes in the DAG-GNN, the path coefficients between nodes in path analysis, and the p-value of Fisher's exactness test.

[0098] <Edge Addition Step> In the edge addition step, the edge addition means 14 adds an edge (second edge) connecting nodes with strong causal relationships in the knowledge graph, based on an index representing the strength of the causal relationship between nodes obtained in the causal analysis step. Then, the weight (strength) of the second edge is set to the value of the index obtained in the causal analysis step (see step S5 in Figure 2).

[0099] Figure 5 shows an example of a knowledge graph in which a second edge has been added in the edge addition step. In Figure 5, solid arrows connecting nodes represent the first edge, dashed arrows connecting nodes represent the second edge, and the numbers next to the second edge represent the index value (e.g., p-value).

[0100] Furthermore, while both the first and second edges indicate causal relationships between process data at each node, their meanings differ. Specifically, the first edge represents "knowledge-based causal relationships" that are pre-constructed based on knowledge held in the mind of, for example, a veteran engineer. On the other hand, the second edge represents causal relationships revealed through causal analysis of process data, in other words, "data-based causal relationships" that were not apparent when the knowledge graph was constructed.

[0101] In the causal analysis step described above, causal analysis is performed not only on the nodes connected by the first edge, but also on all nodes, and in the edge addition step, a second edge is added based on the results of that causal analysis. Therefore, as shown in Figure 5, if a causal relationship is found between nodes that were initially thought to have no direct causal relationship, those nodes are connected by a second edge. Also, in the edge addition step, there are cases where a cause node and an effect node that were initially thought to have no direct causal relationship, such as the node for "equipment status I" and the node for "defect X occurs," are connected by a second edge.

[0102] In the edge addition step, a second edge may be added between all nodes where a causal relationship was found in the causal analysis step. Alternatively, in the edge addition step, a second edge may be added if the index indicating the strength of the causal relationship between nodes, as determined in the causal analysis step, meets a predetermined condition, such as a predetermined threshold for determining the strength of the causal relationship. By adding only second edges that meet the predetermined threshold in this way, it becomes easier to identify the causal path of events occurring in the process based on the causal analysis.

[0103] <Causal Path Presentation Step> In the causal path presentation step, the causal path presentation means 15 presents paths in which nodes are connected by second edges as causal paths that show the causal relationships of process phenomena (see step S6 in Figure 2). For example, Figure 5 shows an example of presenting three causal paths (1), (2), and (3) in the causal path presentation step.

[0104] In the causal path presentation step, the causal paths may be presented by overlaying them on the knowledge graph, as shown in Figure 5. In the causal path presentation step, the results of the causal analysis between nodes (between process data) in the causal analysis step (information on the strength and direction of causal relationships, i.e., causal paths) are integrated by plotting them on the knowledge graph.

[0105] In the causal path presentation step, the connections between nodes (between process data) are visually presented by drawing second edges obtained from the causal analysis results between each node in the knowledge graph. The strength of the causal relationship between nodes may be indicated by the thickness or color of the second edge, or a numerical value (i.e., an indicator value) representing the strength of the causal relationship may be displayed. This allows users to intuitively grasp the important factors.

[0106] Furthermore, in the causal path presentation step, when presenting the results of the causal analysis, only the top-ranked causal routes, based on the strength of the causal relationships of each causal path, may be presented, or only causal routes where the sum of the strengths of the causal relationships is above a threshold may be presented. This allows users to instantly grasp only the important factors.

[0107] <Factor Estimation Step> In the factor estimation step, the factor estimation means 16 ranks the causal paths based on the weights of the second edges included in the causal paths (strength of causal relationship, index value) and estimates the factors of the phenomena in the process (see step S7 in Figure 2).

[0108] For example, if a Bayesian network is applied as a causal analysis method, and the index representing the strength of the causal relationship between events (nodes) is the conditional probability between nodes on each causal path, then the product of the conditional probabilities for each causal path can be calculated, and the causal paths can be ranked in order from the one with the highest product of conditional probabilities.

[0109] Furthermore, when applying DAG-GNN as a causal analysis method and using the weight of the edges between nodes (strength of connection) as an indicator of the strength of the causal relationship between events (nodes), the sum of the edge weights for each causal path can be calculated, and the causal paths can be ranked in order from those with the highest sum of edge weights.

[0110] Furthermore, when applying path analysis as a causal analysis method, and using path coefficients between nodes as an indicator representing the strength of the causal relationship between events (nodes), the product of the path coefficients for each causal path can be calculated, and the causal paths can be ranked in order from those with the highest product of path coefficients.

[0111] Furthermore, if Fisher's exact test is applied as a method for causal analysis, and the p-value in Fisher's exact test is used as an indicator of the strength of the causal relationship between events (nodes), the average p-value of each causal path can be calculated, and the causal paths can be ranked in descending order of their average p-value.

[0112] <Optimal Condition Identification Step> In the optimal condition identification step, the optimal condition identification means 17 identifies the optimal conditions for the second data corresponding to the first data in which the phenomenon does not occur, based on the first data and second data used in the causal analysis (see step S8 in Figure 2). Here, the first data refers to process data linked to the node (result node) corresponding to the phenomenon in the process. The second data refers to process data linked to the node (intermediate node, cause node) corresponding to the event that is the cause of the phenomenon.

[0113] For example, consider the case where the phenomenon in the process is "the occurrence of defect X," as shown in Figure 3. In this case, the optimal condition identification step identifies the process data that should be associated with the result node, intermediate node, and cause node when defect X does not occur, based on the process data associated with the result node, intermediate node, and cause node of the knowledge graph used in the causal analysis.

[0114] In the optimal condition identification step, the optimal conditions (e.g., operating conditions) for the second data corresponding to the first data where the phenomenon does not occur can be identified by following a procedure such as the following. First, the optimal condition identification means 17 extracts operating conditions (process data) such as material state and equipment state linked to the cause node of each ranked causal path. Next, based on the process data linked to the result node "occurrence of defect X", the optimal condition identification means 17 identifies the optimal operating conditions for manufacturing defect-free products, for example, using data science or machine learning techniques.

[0115] In the optimal conditions identification step, a classification model (e.g., logistic regression, decision tree, random forest, support vector machine, etc.) is constructed to predict the presence or absence of defects. Then, using the constructed model, methods such as grid search or Bayesian optimization are used to identify the optimal operating conditions (process data) for manufacturing a product in which defect X does not occur.

[0116] <Output Step> In the output step, the output means 18 outputs the optimal conditions identified by the optimal condition identification means 17 to the process control device 4 (see step S9 in Figure 2). As a result, the process is carried out under the control of the control device 4 based on the optimal conditions, and the product is manufactured. This makes it possible to manufacture products while improving process phenomena (troubles, quality defects, etc.).

[0117] (Factor Estimation Method: Embodiment 2) The factor estimation method according to Embodiment 2 will be described with reference to Figures 2 and 6 to 9. As shown in Figure 2, the factor estimation method according to the embodiment performs the following steps: knowledge graph acquisition step, process data acquisition step, data merging step, causal analysis step, edge addition step, causal path presentation step, factor estimation step, optimal condition identification step, and output step.

[0118] In this embodiment, while adhering to the framework described above, the focus is on long metal materials such as product coils, and the factors causing variations in material and quality observed along the length are identified based on their causal relationship with operating conditions across multiple processes. In this embodiment, the aim is to reduce quality variations along the entire length of the material and optimize operating conditions.

[0119] In the following, we will explain the process data acquisition step (S2), data merging step (S3), and causal analysis step (S4) of steps S1 to S9 shown in Figure 2, focusing on processes specific to long materials. In particular, in the data merging step (S3), non-negative matrix decomposition (NMF) is applied to determine which of the multiple operating conditions originates from localized material and quality variations.

[0120] NMF is a dimensionality reduction method that is (1) suitable for extracting localization patterns in the length direction, and (2) suitable for separating the distribution (W) of the values ​​of the basis components (magnitude of feature quantities) for each length position from the contribution of those components to the measurement items acquired in each process [magnitude (H) of the coefficients (greater than or equal to 0)]. Here, in non-negative matrix decomposition (NMF), data spanning multiple processes is approximated by non-negative basis matrices and coefficient matrices, and the contribution of each process item is expressed as a linear sum of coefficients greater than or equal to 0. This extracts the spatial patterns of quality and material as a semantically interpretable basis without negative cancellation. Furthermore, in NMF, for example, the matrix Q on the length position side and the matrix H on the process item side are extracted.

[0121] <Process Data Acquisition Step (S2)> In this step, quality data (e.g., strength, hardness, presence or absence of defects, etc.) of long materials (e.g., steel plate coils, etc.) is evaluated in a "predetermined range" in the length direction (e.g., 1m pitch). Then, a data set is acquired as process data, which associates actual data or set values ​​of the operating conditions (composition, temperature, plate thickness, plate feeding speed, tension, etc.) for each process of the same product with the corresponding length position.

[0122] Specifically, as described in reference 1 (Japanese Patent Publication No. 7251687), for example, numerous items collected in sensor time series (in units of time) are converted into units of length using the relationship "plate speed × time = distance". Then, after shaping it into a length series with a fixed period by interpolation, a predetermined range is determined considering the actual swapping of the leading and trailing ends, the swapping of the front and back surfaces, and the cutting position.

[0123] Next, the actual data (operating conditions or quality data) for each process, converted to length units, is scaled to match the material length at the final process exit, and quality and the operating conditions of all processes are associated and combined for each length position (lengthwise position). During the combining process, the cutting off of the leading / trailing ends that occurs at each process is taken into consideration, and the actual slits and coil positions of the metal product are identified by tracing back. Then, for each predetermined range of the final product, the actual data of multiple manufacturing conditions (and quality) of the metal material in all processes is combined, with the actual data aligned to the length units of the metal material, to create an analytical data matrix Y (size: m x n) as a dataset for process data analysis.

[0124] This analytical data matrix Y, as shown in Figure 6 for example, is data in which actual data or set values ​​of operating conditions or quality data are arranged for each position from the leading end to the trailing end in the length direction of the metal material in each process. Figure 6 is a diagram showing an example of the configuration of the analytical data matrix Y (m x n) in which quality data evaluated at a predetermined pitch (e.g., 1 m) in the length direction of a long material (e.g., a steel plate coil) and operating conditions (actual data or set values) obtained in each process of the same product are combined in correspondence with the length position. In Figure 6, the rows represent the length direction position of the metal material, and the columns represent the process data items obtained in each process. In this way, in this step, a dataset is formed in which quality and the operating conditions of multiple processes are associated for each length direction position.

[0125] The manufacturing condition data shown in Figure 6 represents the position l in the longitudinal direction at each process. 1 ,l 2 ,...l m and multiple operating conditions or quality data x measured by sensors at the location in question. 1 1 , x 1 2 ,...x 1 m ,...x n 1 , x n 2 ,...x n m The system has items consisting of the following: Here, m represents the number of positions in the longitudinal direction of the metal material, n represents the number of process data, and the analysis data matrix Y is an m x n matrix.

[0126] <Data merging step (S3)> In this step, the analysis data matrix Y created in the process data acquisition step (S2) is used to link the process data associated with the events of each node constituting the knowledge graph to each node. Specifically, the process data is linked to nodes from which measured values ​​or set values ​​can be acquired, and not to nodes from which data cannot be acquired.

[0127] In this step, process data indicating the state of materials and equipment is assigned to the cause node, and quality data obtained from product inspection is assigned to the result node. For intermediate nodes, only nodes describing events for which measured values ​​or setpoints can be obtained are linked to the relevant process data. This linking can be done manually, or, if the amount of data is large, it can be automated using natural language processing or a large-scale language model based on the similarity between the event information described in the node and the name (text) of each item.

[0128] In this step, assuming an analysis data matrix Y (rows: predetermined pitch of length and position, columns: items of process data acquired in each process [actual data or set values]), only the columns corresponding to the item names of the process data to be associated with each node are extracted, and the node-specific analysis data matrix Y node (i) (Size: m x n) i ) constitutes.

[0129] In other words, in this step, for each node i that describes an event in which a measured value or set value can be obtained, a column corresponding to the item name (including the name of the operation item) associated with that node is selected from the analysis data matrix Y, and extracted in the column direction to Y node (i) The rows (predetermined pitch of length and position) are maintained as they are, and quality labels (presence or absence of defects, material variations, etc.) are placed alongside them as reference columns L (size: m x 1) for matching, as needed. If there are values ​​that do not conform to the non-negative constraint, they are converted to 0 or greater by normalization, and then Y node (i) This is used as input for non-negative matrix factorization (NMF).

[0130] Figure 7 shows the data matrix Y for node-specific analysis. node (i) (m x n) i Applying non-negative matrix factorization (NMF) to the basis matrix W (i) (m x k i ) and coefficient matrix H (i) (k i ×n iThis figure shows the concept of decomposing into ). Also, equation (1) below corresponds to the equation in Figure 7. In NMF, the node-specific analysis data matrix Y of each node. node (i) (Size: m x n) i ) and the number of feature patterns (basis vectors) after compression is k i Let W be a non-negative basis matrix. (i) (Size: m x k) i ) and the coefficient matrix H (i) (Size: k) i ×n i ) is broken down into these two parts.

[0131]

[0132] Here, W (i) H is a matrix where each column consists of non-negative basis vectors. (i) Each data column is W (i) This is a matrix containing the coefficients (weights) for representing the basis vectors of a given expression using non-negative combinations (with respect to column j, Y node (i) The column j is W (i) and H (i) (Approximately represented by the column j).

[0133] Also, matrix W (i) This represents the intensity change of the basis vector (feature pattern) for each length position and serves as an analytical feature used to match the positional correspondence with quality labels such as defect occurrence locations and material variations. Also, matrix H (i) This numerically shows how much each basis vector (feature pattern) relates to (contributes to) each item (temperature, tension, velocity, component, etc.) of the actual data for each process. This allows us to understand "what phenomenon or equipment characteristics does this basis represent?"

[0134] Even for large-scale, high-dimensional data collected from multi-process processes, the use of NMF (Non-negative Matrix Factorization) allows the data to be grouped and compressed in meaningful on-site units such as "for each phenomenon" or "for each piece of equipment". Since NMF limits all values to be greater than or equal to 0, the decomposed elements (bases) are configured by "addition", and the characteristics of each process and each piece of equipment are extracted as they are as meaningful data clusters.

[0135] In the non-negative matrix factorization (NMF) of the present invention, for each node, the node-specific analysis data matrix Y node (i) (size: m×n i ), a non-negative basis matrix W (i) (size: m×k i ) and a non-negative coefficient matrix H (i) (size: k i ×n i ) are obtained. Then, as shown in the following formula (2), the product W (i) of W and H (i) ·H (i) ·H (i) is defined as the reconstruction matrix Y^ node (i) , and the mean squared error E(k node (i) ) between the reconstruction matrix Y^ node (i) and the original matrix Y i ) is evaluated as a function of the number of bases k i . Specifically, "k i = 1, 2, ..., k max (constant)" is scanned to confirm the transition of E(k i ), and the inflection point (so-called elbow) at which the error reduction saturates is determined as the optimal number of bases k i .

[0136]

[0137] The number of bases k iBy determining the error reduction saturation point, it is possible to suppress overfitting while maintaining sufficient reduction of reconstruction errors and ensuring the generalization performance of the model. As a result, W is aggregated into a small number of basis values ​​that are easily correlated with phenomena and equipment, becoming analytical features that can be treated as meaningful units (process / equipment units / operation phases / phenomenon groups / functional blocks) in the field. In addition, H has a sparse and stable contribution with unnecessary weights suppressed. The extracted features are stabilized as features linked to the occurrence of defects and material variations, and the same interpretation becomes possible even for data with different conditions. As a result, it is also effective for monitoring and diagnosis during long-term operation and application to other lines.

[0138] <Causal Analysis Step (S4)> In this step, a non-negative basis matrix W is associated with each node of the knowledge graph (ontology). (i) (Size: m x k) i The basis vectors (W) that constitute the ) (i) Between the columns, the base intensity value (W) along the length position. (i) Using the row elements, the causal direction and causal intensity are calculated through causal analysis. This step is performed for each length of the metal material (from tip to end).

[0139] Here, causal analysis is performed based on the length-position pitch of the metal material (e.g., 1 m), primarily for each position (same row of basis vectors). However, considering the variation in position correspondence due to winding, unwinding, and cutting of thin sheets, as well as the bidirectional propagation of internal structure, "reverse" causality (backward position → forward position) within a nearby range along the length direction is also allowed as part of the analysis, limited to the vicinity of the phenomenon occurrence point. In this case, the process sequence can be used as prior knowledge of position correspondence, but the constraint prohibiting "future → past" on the time axis does not apply.

[0140] The selection of causal relationships is limited to those that meet a predetermined threshold for both causal strength and confidence score (including evaluation of stability and likelihood through resampling, etc.), and only these causal relationships are registered and visualized on the ontology. Furthermore, quality indicators such as the presence or absence of defects and material variations are also included in the causal analysis, and by referring to the W (baseline strength value) at the relevant location, it is possible to identify "which baseline (feature pattern) chained together and contributed to the final defect or material variation." This allows for the evaluation of causality between phenomena that propagate lengthwise across processes, and extends causal analysis to lengthwise positional data in a way that is not limited to conventional time-series causality (future → past prohibition).

[0141] Furthermore, if the causal direction is clear in advance, for example, if fluctuations in the operating conditions of an upstream process affect the quality of a downstream process, the causal direction can be assumed to be given by prior knowledge such as the process sequence. In this case, instead of calculating the causal direction and causal intensity through causal analysis, it may be evaluated as follows: that is, the basis intensity sequence of the basis vectors (W (i) The "strength of association" between series may be evaluated based on the similarity or correlation between the lengthwise variation of the column (of the quality indicators) and the lengthwise variation of quality indicators (presence or absence of defects, material variation, shape value, etc.) or the base strength column of other processes. For example, the correlation coefficient (Pearson correlation coefficient, etc.) between the series of quality indicators and each base strength column can be calculated, and bases whose absolute value of the correlation coefficient is greater than or equal to a predetermined threshold can be extracted as bases that have a strong association with the quality indicator, and these correlation coefficients can be used as a surrogate indicator of causal strength.

[0142] Furthermore, if variations in positional correspondence in the longitudinal direction are expected, the correlation coefficient may be calculated by comparing the baseline strength value and the quality index value in a nearby range of, for example, ±5 m for a length-position pitch (e.g., 1 m), and the strength of the relationship in that nearby range may be evaluated.

[0143] (Causal evaluation of positional synchronization and neighborhood tolerance using base intensity sequences and identification of causal factors) Figure 8 shows the basis vectors (W) extracted by NMF. (i) This diagram illustrates the concept of a procedure for evaluating the relationship between each column using the base intensity value for each length and position.

[0144] The relationship in this step is evaluated by the basis intensity sequence (W) of the basis vectors linked to the nodes (e.g., basis A1 of node A and basis B1 of node B), as shown in Figure 9. (i) The columns are targeted, and the base strength w values ​​at the same length position (same row) are compared in principle. Furthermore, if the causal direction is clear from prior knowledge such as the process sequence, the strength of the relationship may be evaluated based on the similarity or correlation between the base strength column and the lengthwise variation of quality indicators (presence or absence of defects, material variation, shape value, etc.). In addition, if variability is expected in the positional correspondence in the lengthwise direction, the base strength values ​​w are compared within a nearby range of, for example, ±5 m for a lengthwise position pitch (e.g., 1 m), and the causal direction and causal strength within that range are calculated.

[0145] Furthermore, the evaluation of the relationships shown in Figure 8 is primarily performed between basis intensity sequences (i.e., between basis vectors), and quality indicators may be used additionally as reference sequences to evaluate the relationship with the said basis intensity sequences, as needed.

[0146] Figure 9 is a schematic diagram of the decision-making process that adopts relationships between bases (or between bases and quality indicators) when a predetermined threshold is met, based on the evaluation results shown in Figure 8. As shown in Figure 9, in a positional interval of length where an indicator representing the strength of the relationship (causal strength and confidence score, correlation coefficient, etc.) exceeds a predetermined threshold, it is determined that there is a causal relationship between the corresponding bases, and that causal link or correlation is registered on the knowledge graph.

[0147] When identifying causal factors, the coefficient matrix H (i) The contribution of process data items included in (size: k x n) is referenced, and items whose contribution to the base exceeds a threshold are extracted as potential causes. As shown in Figure 13 below, the contributions may be visualized in descending order using a histogram or similar method and presented to the operator. This allows for an intuitive understanding of the cause-effect relationship between cause nodes and effect nodes via causal routes between base elements (groups of items) that are difficult to see through correlations between single variables.

[0148] (Example 1) An example of the factor estimation method according to the embodiment will be described with reference to Figures 3 to 5 and Figure 10. In this embodiment, a method for estimating the factors of surface defects in the manufacturing process of steel products will be described. The process covered in this embodiment is a series of processes from the steelmaking plant (steelmaking process) in the upstream process of a steelworks that manufactures steel, to the final completion of steel products in the hot-dip galvanizing plant. In this embodiment, a causal analysis is performed on a certain defect X that occurs in the downstream process due to the operating conditions of the upstream process.

[0149] The dataset used for causal analysis in this embodiment consists of data on whether or not defect X occurred in each slab cast during the steelmaking process, and operational data of the steelmaking plant collected during the casting of each slab. The data on whether or not defect X occurred is created when the steel products are processed in the final stage, the hot-dip galvanizing plant, and then labeled by quality control personnel to indicate whether or not a defect occurred.

[0150] In addition, along with these datasets, a knowledge graph was prepared as a knowledge database for factor estimation. This graph systematically organizes the knowledge leading to defect occurrence, compiled by experienced engineers. Figure 3 shows the knowledge graph that organizes the defect occurrence mechanism of defect X. Each node in this knowledge graph describes an event, such as "XX is non-uniform" or "XX equipment is in the △△ state."

[0151] Furthermore, the terminal nodes located at the bottom of the knowledge graph are causal nodes corresponding to individual operating conditions that can cause defects, and the causal chain extends upwards. The nodes located in the middle of the knowledge graph are intermediate nodes corresponding to various events. Intermediate nodes become nodes that more directly cause defects as individual factors overlap, and are ultimately connected to the result node at the top of the knowledge graph, which indicates the occurrence of the defect corresponding to the resulting phenomenon.

[0152] This section describes a specific procedure for estimating defect causes using these datasets and knowledge graphs. First, each node on the knowledge graph is linked to the measured process data (measurement data). Figure 4 shows an example of linking process data to nodes. In Figure 4, for example, for the node "Component A on the surface reacts with component B in the molten steel to produce C," process data considered to be related to this phenomenon is extracted from the dataset.

[0153] If the amount of data is very large and manual data extraction is difficult, this linking process may be automated using natural language processing. One method of data extraction using natural language processing is to extract data based on the similarity between the text described in the nodes on the knowledge graph and each operational item. In this case, it is desirable that each operational item be named in a way that clearly indicates what the data measures.

[0154] Alternatively, as a method for extracting data using natural language processing, a large-scale language model (LLM) may be used, which is provided with information equivalent to domain knowledge such as internal company documents, and then extracted using methods such as RAG or fine tuning.

[0155] In this example, a RAG environment was constructed in which internal company documents related to the analysis of defect X were provided as external data to a large-scale language model. Then, as items related to the node "Component A on the surface reacts with component B in the molten steel to produce C," the "amount of component A in equipment U" and the "amount of component B in equipment V" were automatically extracted and linked to the node. In addition, the same procedure was used to link related operational data items for other nodes.

[0156] Furthermore, dimensionality reduction may be performed as a data processing step before causal analysis, if necessary. This dimensionality reduction may not be necessary if the amount of process data is not enormous and spurious correlations between data have been sufficiently eliminated. Also, if the amount of data is enormous, the readability of the causal analysis results when plotted will be poor, so dimensionality reduction may be performed to make the causal analysis results easier to understand.

[0157] Dimensionality reduction techniques such as principal component analysis and non-negative matrix factorization can be used. In this example, since the process data was pre-selected and the number of data points was limited, dimensionality reduction was not performed.

[0158] Next, causal analysis is performed between process data, and the causal relationships are plotted on a knowledge graph. Figure 5 shows an example of a causal relationship plot. In Figure 5, the circles represent process data extracted as being related to each node. If there are multiple data points extracted for each node, the number of circles is plotted. Also, if the number of data points is large and dimensionality reduction is performed, the number of circles is plotted according to the number of compressed features.

[0159] The dashed arrows in Figure 5 represent second edges indicating causal relationships between each process data point or feature. Besides simply examining the correlation between data points, other causal analysis methods such as Bayesian networks, DAG-GNN, path analysis, and Fisher's exact test may also be used.

[0160] Furthermore, the causal analysis of the data may reflect the constraints of knowledge-based causal relationships (i.e., first edges) in the knowledge graph. However, by deliberately not reflecting knowledge-based causal relationships, it is possible that causal relationships that were not organized on the knowledge graph may be highlighted from the data, so either approach can be chosen depending on the purpose. In this embodiment, no constraints based on knowledge-based causal relationships were applied during the causal analysis of the data.

[0161] In Figure 5, data points extracted as related items, including the amounts of component A and component B, are plotted on a knowledge graph, and the causal relationships between each process data are illustrated with dashed arrows (second edges). In this example, the causal relationships between each process data were confirmed using Fisher's exact test. The numbers next to the dashed arrows indicate the p-values ​​in Fisher's exact test. In this example, it was assumed that a p-value of 0.05 or less was statistically significant.

[0162] Figure 5 simultaneously visualizes the causal relationships based on knowledge (first edge), shown by solid arrows, and the causal relationships based on data (second edge), shown by dashed arrows. This eliminates spurious correlations in multivariate analysis between process data, making it possible to perform defect factor analysis that includes causal relationships based on knowledge graphs that cannot be determined from data alone.

[0163] Furthermore, the fact that a second edge, represented by a dashed arrow indicating a causal relationship based on data, connects nodes that are not directly linked in the knowledge graph, such as a cause node "Equipment Status I" and an effect node "Defect X occurs," suggests the following possibilities: (1) The knowledge graph is incomplete due to missing intermediate nodes. (2) There is no data to connect to intermediate nodes due to a lack of sensors, etc.

[0164] Therefore, by overlaying the first edge and the second edge, which shows a significant causal relationship based on the results of causal analysis, on the knowledge graph, the incompleteness of the knowledge graph can be compensated for. In addition, items without process data (measurement data), i.e., nodes to which process data is not linked, can be included in the consideration of causal relationships, and unknown relationships between them can be revealed from the data, allowing for the reconstruction of the knowledge model.

[0165] Furthermore, as a way to more accurately display the connections between data, for example in Figure 5, the thickness and color of the dashed arrows may be changed according to the p-value in Fisher's exact test. That is, the smaller the p-value in Fisher's exact test, the thicker the line may be, and the larger the p-value, the thinner the line may be.

[0166] Alternatively, a threshold may be set in advance, and only second edges with p-values ​​below that threshold may be displayed. For example, in Figure 5, by setting the threshold to 0.04 or less, the second edge with a p-value of 0.05 (see left side of the page) may be hidden. Furthermore, even when a method other than Fisher's exact test is used as the causal analysis method between process data, the thickness and color of the second edge may be changed, or a display threshold may be set, for example, according to the conditional probability in the Bayesian network.

[0167] In this way, by plotting the causal relationships obtained between process data on a knowledge graph, it becomes possible to estimate the causes of defects more reliably and quickly from both the perspective of knowledge-based causal relationships (first edge) and data-based causal relationships (second edge).

[0168] In this example, causal paths from cause nodes to effect nodes were estimated based on an index representing the strength of causal relationships between nodes (here, the p-value in Fisher's exact test). Specifically, in this example, multiple causal paths (1), (2), and (3) tracing back from cause nodes to effect nodes were estimated. Then, each causal path was ranked based on the average value of the index (p-value) representing the strength of causal relationships between nodes on each causal path.

[0169] For example, Figure 5 shows three causal routes: (1) starting from the cause node "Equipment State I", (2) starting from the cause node "Material State E", and (3) starting from "Equipment State D". In this example, these were ranked based on the average p-value.

[0170] Figure 10 shows the average p-values ​​for each causal path and the results of ranking them in descending order of average p-values. In Fisher's exact test, a lower p-value indicates a statistically significant relationship between the result node and the cause node. Therefore, in this embodiment, causal paths were ranked in descending order of average p-values, starting with the lowest average p-values. By ranking causal routes based on the strength of the causal relationship between the result node and the cause node in this way, it becomes possible to display the results of causal analysis in a visually easy-to-understand manner.

[0171] (Example 2) In this example, high-tensile steel sheets (high-strength steel sheets (steel sheets with a predetermined thickness range, for example, about 2 mm thick)) manufactured in a continuous annealing line with a water cooling (rapid cooling) process were targeted, and the factors causing fluctuations in the shape value (e.g., C-curvature), which is a quality indicator, were analyzed based on the processing specific to long materials shown in Example 2.

[0172] First, in the process data acquisition step (S2), quality data (such as warpage) was evaluated in a predetermined range along the length of the target steel plate (e.g., at 1m intervals). Furthermore, actual data or set values ​​of the operating conditions (composition, temperature, plate thickness, feed speed, tension, etc.) for each process of the same product were associated with these length positions to create an analysis data matrix Y (m x n). Next, in the data merging step (S3), columns of process data items associated with each node were extracted from Y to create node-specific analysis data matrices Y node (i) (m x n) i ) was formed.

[0173] Next, using non-negative matrix factorization (NMF), each Y node (i) The basis matrix W (i) (m x k i ) and coefficient matrix H (i) (k i ×n i ) is decomposed into basis vectors (W) representing the feature pattern in the length direction. (i) Each column of was extracted (e.g., k=8). Here, W (i) The basis intensity column represents the variation in the length direction of the basis vector, H (i) This represents the contribution (weight) of the original data (process data items) that make up the basis vector.

[0174] Figure 11 shows an example of the longitudinal variation of the quality index (C-warpage) in the target steel plate, illustrating the interval where both the mean and variance of C-warpage are small (the interval where the shape is stable). Figure 12 shows the base intensity sequence (W) of the extracted basis vectors 1 to 8. (i) This shows examples of variation in the length direction of each column.

[0175] In this embodiment, it was assumed that the causal direction was clear from prior knowledge such as the process sequence (for example, fluctuations in the operating conditions of the upstream process affect the quality of the downstream process). Therefore, instead of calculating the causal direction and causal intensity through causal analysis, the similarity or correlation between the variation in the length direction of the baseline intensity column and the variation in the length direction of the quality index (C-warpage) was used as an alternative index to represent the "strength of the association."

[0176] Specifically, the correlation coefficient (Pearson correlation coefficient, etc.) between the series of C-curvature values ​​and each base intensity series was calculated, and base vectors whose absolute value of the correlation coefficient was above a predetermined threshold were extracted as bases strongly associated with C-curvature. This allows for the quantitative selection of base vectors (feature patterns) that strongly correspond to the appearance of stable intervals of C-curvature, without relying on visualization. If variability is expected in the positional correspondence in the length direction, the correlation coefficient may be calculated by comparing the base intensity values ​​and C-curvature values ​​in a nearby range of, for example, ±5 m for a length-position pitch (e.g., 1 m), and the strength of the association in that nearby range may be evaluated.

[0177] Furthermore, the coefficient matrix H corresponding to the basis vectors extracted above. (i) By referring to the above, process data items with high contribution were extracted as potential causes. Figure 13 shows an example of visualizing the contribution of basis vector 6, shown in Figure 12, which has a high correlation coefficient with the stable interval with small mean and variance among the length-direction fluctuations of the quality index (C warp) shown in Figure 11, in descending order using a histogram, etc. According to the present invention, as shown in Figure 13, it is possible to efficiently grasp the operating conditions (e.g., cooling equipment conditions, plate temperature, etc.) corresponding to the interval in which the quality index (C warp) is stable.

[0178] The factor estimation device, factor estimation method, process implementation method, factor estimation system, and terminal device according to the embodiments described above provide a method for displaying the results of causal analysis of process data in combination with a knowledge graph represented using a knowledge database that stores the knowledge of veterans. This enables simultaneous visualization of the causal relationship between knowledge and process data, making it possible to estimate factors with higher accuracy.

[0179] Furthermore, in the factor estimation apparatus, factor estimation method, process implementation method, factor estimation system, and terminal device according to the embodiment, dimensionality reduction techniques such as principal component analysis (PCA) and non-negative matrix factorization (NMF) are applied to high-dimensional data to reduce the data's dimensionality. In particular, in the manufacturing process of steel products, dimensionality reduction is performed while considering the relationships between each process, thereby simplifying the data while preserving the characteristics of each process.

[0180] Furthermore, in the factor estimation device, factor estimation method, process implementation method, factor estimation system, and terminal device according to the embodiment, Bayesian networks, DAG-GNN, path analysis, etc., are used as causal analysis methods to numerically evaluate the direction and strength of causal relationships. The results based on the causal analysis method are then displayed on a knowledge graph as the strength of the causal relationship between events (nodes), and presented in a way that is easy for the user to understand. Based on the strength of the causal relationship between events (nodes) on the knowledge graph, the causal path at the time of an anomaly is identified.

[0181] Accordingly, according to the factor estimation device, factor estimation method, process implementation method, factor estimation system, and terminal device of the embodiment, by displaying the causal analysis results of knowledge graphs and process data in combination, it is possible to resolve the spurious correlation problem of multivariate analysis, including multiple regression analysis, and to compensate for the incompleteness of knowledge graphs, thereby improving the accuracy of factor estimation.

[0182] Furthermore, the factor estimation apparatus, factor estimation method, process implementation method, factor estimation system, and terminal device according to the embodiment can contribute to stabilizing the quality of products manufactured through complex processes, such as steel products, and reducing defects.

[0183] Furthermore, according to the factor estimation apparatus, factor estimation method, process implementation method, factor estimation system, and terminal device of the embodiment, by performing appropriate dimensionality reduction on high-dimensional data, analysis efficiency is improved, and effective analysis becomes possible even in manufacturing processes with a large number of steps and measurement items.

[0184] Furthermore, according to the factor estimation device, factor estimation method, process implementation method, factor estimation system, and terminal device of the embodiment, the causal analysis results are displayed in an easy-to-understand visual manner, making it easier to interpret the analysis results and enabling rapid decision-making on-site.

[0185] The factor estimation apparatus, factor estimation method, process implementation method, factor estimation system, and terminal device according to the present invention have been specifically described above with reference to embodiments and examples for carrying out the invention. However, the spirit of the present invention is not limited to these descriptions and must be interpreted broadly based on the claims. Furthermore, it goes without saying that various changes and modifications based on these descriptions are also included in the spirit of the present invention.

[0186] 1. Factor estimation device 10. Knowledge graph acquisition means 11. Process data acquisition means 12. Data merging means 13. Causal analysis means 14. Edge addition means 15. Causal path presentation means 16. Factor estimation means 17. Optimal condition identification means 18. Output means 19. Display information output means 2. Terminal device 21. Display information acquisition means 22. Display means 3. Storage device 4. Control device N. Network

Claims

1. A factor estimation device comprising: a knowledge graph acquisition means for acquiring a knowledge graph that expresses the causal relationship between phenomena and events in a process, comprising a plurality of nodes representing events occurring in a process and first edges connecting the nodes and representing causal relationships between the nodes; a process data acquisition means for acquiring process data collected from the process; a data linking means for linking the process data related to the events of the nodes to each node, targeting the knowledge graph; a causal analysis means for calculating an index representing the strength of the causal relationship between the nodes to which the process data is linked; an edge adding means for adding second edges connecting nodes with strong causal relationships to the knowledge graph based on the index and setting the value of the index as the weight of the second edge; and a causal path presentation means for presenting paths in which the nodes are connected by the second edge as causal paths indicating the causal relationship.

2. The factor estimation apparatus according to claim 1, further comprising a factor estimation means for estimating the factors of a phenomenon in the process by ranking a plurality of causal paths based on the weights of the second edge.

3. The factor estimation device according to claim 2, further comprising: an optimal condition identification means for identifying the optimal conditions for the second data corresponding to the first data in which the phenomenon does not occur, based on first data which is process data associated with the node corresponding to the phenomenon in the process, and second data which is process data associated with the node corresponding to the event that is the cause of the phenomenon; and an output means for outputting the identified optimal conditions to the control device of the process.

4. The factor estimation device according to claim 1, wherein the data linking means uses natural language processing technology or a large-scale language model to link the process data related to the events of the nodes to each node based on the name of the process data.

5. The factor estimation device according to claim 1, wherein the data linking means compresses the dimensionality of the process data to be subjected to causal analysis and links it to each node.

6. The factor estimation device according to claim 5, wherein the data merging means reduces the dimensionality of the process data to be subjected to causal analysis using principal component analysis or non-negative matrix factorization.

7. The process data acquisition means collects the process data acquired in multiple processes associated with the length positions of the metal material, and constructs an analysis data matrix Y in which the rows are at a predetermined pitch of length positions and the columns are items of actual data acquired in each process. The data joining means targets the knowledge graph and links the items of the analysis data matrix Y related to the events of the nodes to each node, and extracts the column groups related to each node from the analysis data matrix Y to form a node-specific analysis data matrix Y. node The node-specific analysis data matrix Y, which is composed of multiple items, comprises the above-mentioned items. node The factor estimation device according to claim 5, wherein non-negative matrix factorization is applied to compress the dimension to extract a plurality of basis vectors representing a feature pattern in the length direction, and the causal analysis means performs causal analysis between the basis vectors of each extracted node.

8. The data joining means includes the node-specific analysis data matrix Y node Applying non-negative matrix factorization to the process data, we decompose it into non-negative matrices W and H, with each column of W being a basis vector representing a feature pattern in the length direction, and each row of H being the coefficient (greater than or equal to 0) of the basis component constituting the basis vector for each item of the process data, and the number of basis vectors k in the non-negative matrix factorization is the reconstruction matrix Y^ node and the original matrix Y node The factor estimation device according to claim 7, wherein the mean squared error E(k) with respect to the number of bases k is evaluated as a function of the number of bases k, and the inflection point at which the error reduction saturates is determined as the optimal number of bases k.

9. The causal analysis means calculates the index using any of Bayesian networks, DAG-GNN, path analysis, or Fisher's exact test, according to claim 1.

10. The causal path presentation means presents the causal path overlaid on the knowledge graph, as described in claim 1.

11. A method for estimating factors, comprising: a data joining step of linking process data related to the events of a node to each node, with respect to a knowledge graph that expresses the causal relationship between phenomena and events in a process, which is composed of a plurality of nodes representing events occurring in a process and first edges connecting the nodes and representing causal relationships between the nodes; a causal analysis step of calculating an index representing the strength of the causal relationship between the nodes to which the process data is linked; an edge adding step of adding second edges to the knowledge graph that connect nodes with strong causal relationships based on the index, and setting the value of the index as the weight of the second edge; and a causal path presentation step of presenting the paths in which the nodes are connected by the second edges as causal paths that indicate the causal relationship.

12. The factor estimation method according to claim 11, further comprising a factor estimation step of estimating the factors of a phenomenon in the process by ranking a plurality of the causal paths based on the weights of the second edge, after the causal path presentation step.

13. A process implementation method comprising: an optimal condition identification step for identifying the optimal conditions for the second data corresponding to the first data in which the phenomenon does not occur, based on first data which is process data linked to a node corresponding to a phenomenon in the process and second data which is process data linked to a node corresponding to an event that is a factor of the phenomenon, with respect to a causal path presented by the factor estimation method according to claim 11 or claim 12; and a process implementation step for implementing the process based on the identified optimal conditions.

14. A knowledge graph acquisition means that acquires a knowledge graph expressing the causal relationship between phenomena and events in the process, comprising a factor estimation device and a terminal device, and consisting of a plurality of nodes representing events occurring in the process and first edges connecting the nodes and representing causal relationships between the nodes; a process data acquisition means that acquires process data collected from the process; a data linking means that links the process data related to the events of the nodes to each node, targeting the knowledge graph; a causal analysis means that calculates an index representing the strength of the causal relationship between the nodes to which the process data is linked; an edge adding means that adds second edges connecting nodes with strong causal relationships to the knowledge graph based on the index and sets the value of the index as the weight of the second edge; a causal path presentation means that presents paths in which the nodes are connected by the second edge as causal paths indicating the causal relationship; and a factor estimation means that estimates the factors of phenomena in the process by ranking a plurality of causal paths based on the weight of the second edge. A factor estimation system comprising: an optimal condition identification means for identifying the optimal conditions for the second data corresponding to the first data in which the phenomenon does not occur, based on first data which is process data associated with the node corresponding to the phenomenon in the process and second data which is process data associated with the node corresponding to the event that is the cause of the phenomenon; an output means for outputting the identified optimal conditions to the control device of the process; a display information output means for outputting information including at least the causal path and the ranking of the causal path; a display information acquisition means for acquiring the said information; and a display means for displaying the said information.

15. A terminal device comprising: a display information acquisition means for acquiring information from a factor estimation device, including at least a causal path and a ranking of the causal path; and a display means for displaying the information, wherein the causal path is a path in a knowledge graph representing the causal relationship between phenomena and events in a process, where the nodes associated with the events of the nodes are connected by a second edge, and the second edge is an edge in the knowledge graph that connects nodes with strong causal relationships based on an index representing the strength of the causal relationship between the nodes associated with the process data; and the ranking of the causal path is a ranking of a plurality of causal paths based on the weight of the second edge.