Computer program, information processing device, and information processing method

The computer program and information processing method address the challenge of deriving causal structures in substrate processing systems by using observation data and reference causal structures, resulting in enhanced optimization capabilities for substrate processing.

WO2025110000A1PCT designated stage expired Publication Date: 2025-05-30TOKYO ELECTRON LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/039400
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing technologies lack the capability to derive the causal structure of observation variables in an observation system based on a reference causal structure, which is essential for optimizing substrate processing in substrate processing apparatuses.

Method used

A computer program and information processing method that acquire observation data from a substrate processing apparatus, set priori constraints based on a reference causal structure, and execute learning to derive the causal structure of the observation variable group in the observation system.

Benefits of technology

Enables the derivation of the causal structure of observation variables in the observation system, allowing for improved understanding and optimization of substrate processing by leveraging the causal relationships between observed variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039400_30052025_PF_FP_ABST
    Figure JP2024039400_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a computer program, an information processing device, and an information processing method.  In the present invention, a computer is caused to execute processing for deriving a causal structure of an observation variable group in an observation system by: acquiring observation data corresponding to a plurality of types of observation variables from the observation system; setting a prior constraint on the basis of a causal structure of a variable group in a reference system; and using the acquired observation data to perform learning with the prior constraint imposed thereon.
Need to check novelty before this filing date? Find Prior Art

Description

Computer program, information processing device, and information processing method

[0001] The present invention relates to a computer program, an information processing device, and an information processing method.

[0002] Substrate processing equipment processes substrates based on a process recipe. The recipe consists of multiple steps, and optimal processing results can be obtained by controlling various parameters, such as pressure and temperature, for each step. Since the set values ​​of the various parameters may differ for each step, measurement data from multiple sensors installed in the substrate processing equipment is managed for each substrate.

[0003] Patent Document 1 discloses a technique for visualizing data regarding a process having multidimensional independent variables and dependent variables.

[0004] Special Publication No. 2022-502806

[0005] The present disclosure aims to provide a computer program, an information processing device, and an information processing method that can derive the causal structure of a group of observed variables in an observation system based on a reference causal structure.

[0006] The computer program disclosed herein is a computer program for causing a computer to execute a process of deriving the causal structure of a group of observed variables in an observation system by acquiring observation data corresponding to multiple types of observed variables from an observation system, setting a priori constraints based on the causal structure of a group of variables in a reference system, and performing learning with the a priori constraints imposed using the acquired observation data.

[0007] According to the present disclosure, the causal structure of a group of observed variables in an observation system can be derived based on a reference causal structure.

[0008] FIG. 1 is an explanatory diagram illustrating a configuration of an information processing system according to an embodiment. FIG. 2 is a block diagram illustrating the internal configuration of an information processing device. FIG. 3 is an explanatory diagram illustrating an internal representation of a causal structure. FIG. 4 is an explanatory diagram illustrating an example of setting a confidence interval. FIG. 5 is a schematic diagram illustrating an example of a causal structure including nonlinearity. FIG. 6 is an explanatory diagram illustrating an example of a causal structure obtained by additional learning based on a generic structure. FIG. 7 is a flowchart illustrating a procedure for deriving a causal structure using a generic structure. FIG. 8 is a schematic diagram illustrating an example of a display of a causal structure obtained by additional learning. FIG. 9 is a schematic diagram illustrating an example of a causal structure when a disturbance factor occurs. FIG. 10 is a flowchart illustrating a procedure for estimating an unknown disturbance.

[0009] An embodiment will be described below with reference to the drawings. (Embodiment 1) Fig. 1 is an explanatory diagram illustrating the configuration of an information processing system according to an embodiment. The information processing system according to the embodiment includes an information processing apparatus 100 and a substrate processing apparatus 200 that are communicatively connected.

[0010] The substrate processing apparatus 200 is, for example, a semiconductor manufacturing apparatus including at least one of an exposure apparatus, an etching apparatus, a film forming apparatus, an ion implantation apparatus, an ashing apparatus, a sputtering apparatus, etc. Alternatively, the substrate processing apparatus 200 may be a display manufacturing apparatus that manufactures flat display panels (FDPs) such as liquid crystal display panels and organic electroluminescence (EL) panels.

[0011] At the start of a process, various setting values ​​such as the substrate temperature, the pressure and gas flow rate in the chamber, and the voltage applied from the high-frequency power supply are set in the substrate processing apparatus 200. The substrate processing apparatus 200 is also provided with a plurality of sensors that measure the substrate temperature, the pressure and gas flow rate in the chamber, the voltage applied to the upper electrode and the lower electrode, etc. during the process. The substrate processing apparatus 200 outputs the setting values ​​set at the start of the process and the measurement values ​​measured during the process to the information processing apparatus 100 as observation data.

[0012] The information processing apparatus 100 acquires observation data corresponding to a plurality of types of observation variables from the substrate processing apparatus 200 of the observation system. The information processing apparatus 100 sets a priori constraints based on the causal structure of the reference system, and performs learning with the prior constraints imposed using the observation data obtained from the substrate processing apparatus 200 of the observation system, thereby deriving the causal structure of the group of observation variables in the observation system. In this embodiment, the observation system and the reference system are substrate processing apparatus 200 of the same model. For example, a causal structure created in advance based on user knowledge can be used as the causal structure of the reference system. Alternatively, the causal structure of the reference system may be a causal structure created based on observation data obtained from a specific substrate processing apparatus 200.

[0013] 2 is a block diagram showing the internal configuration of the information processing device 100. The information processing device 100 is, for example, a dedicated or general-purpose computer including a control unit 101, a storage unit 102, a communication unit 103, an operation unit 104, and a display unit 105.

[0014] The control unit 101 includes a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The ROM included in the control unit 101 stores control programs and the like that control the operation of each hardware unit included in the information processing device 100. The CPU in the control unit 101 reads and executes the control programs stored in the ROM and computer programs (described below) stored in the storage unit 102, and controls the operation of each hardware unit, thereby causing the entire device to function as the information processing device of the present disclosure. The RAM included in the control unit 101 temporarily stores data used during execution of calculations.

[0015] In the embodiment, the control unit 101 is configured to include a CPU, a ROM, and a RAM, but the configuration of the control unit 101 is not limited to the above. The control unit 101 may be, for example, one or more control circuits or arithmetic circuits including a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), a quantum processor, volatile or non-volatile memory, etc. The control unit 101 may also have functions such as a clock that outputs date and time information, a timer that measures the elapsed time from when a measurement start instruction is given until when a measurement end instruction is given, and a counter that counts numbers.

[0016] The storage unit 102 includes a storage device such as a hard disk drive (HDD), a solid state drive (SSD), an electronically erasable programmable read-only memory (EEPROM), etc. The storage unit 102 stores various computer programs executed by the control unit 101 and various data used by the control unit 101.

[0017] The computer program (program product) stored in the storage unit 102 includes a causal structure learning program PG1 for causing a computer to execute a process of deriving a causal structure of observation variables from observation data of the substrate processing apparatus 200. The causal structure learning program PG1 may be a single computer program or may be composed of multiple computer programs. Furthermore, the causal structure learning program PG1 may be executed by multiple computers working together. Furthermore, the causal structure learning program PG1 may partially use an existing library.

[0018] A computer program including the causal structure learning program PG1 is provided by a non-transitory recording medium RM on which the computer program is readably recorded. The recording medium RM is a portable memory such as a CD-ROM, a USB memory, a Secure Digital (SD) card, a micro SD card, or a CompactFlash (registered trademark). The control unit 101 reads various computer programs from the recording medium RM using a reading device (not shown) and stores the read various computer programs in the storage unit 102. The computer programs stored in the storage unit 102 may also be provided via communication. In this case, the control unit 101 acquires the computer programs via communication via the communication unit 103 and stores the acquired computer programs in the storage unit 102.

[0019] The communication unit 103 includes a communication interface for transmitting and receiving various data to and from an external device. A communication interface conforming to a communication standard such as a local area network (LAN) can be used as the communication interface of the communication unit 103. The external device may be the substrate processing apparatus 200 described above or a user terminal (not shown). When data to be transmitted is input from the control unit 101, the communication unit 103 transmits the data to the external device as the destination, and when data transmitted from the external device is received, the communication unit 103 outputs the received data to the control unit 101.

[0020] The operation unit 104 includes operation devices such as a touch panel, a keyboard, and switches, and receives various operations and settings from a user, etc. The control unit 101 performs appropriate control based on various pieces of operation information provided by the operation unit 104, and stores setting information in the storage unit 102 as necessary.

[0021] The display unit 105 includes a display device such as a liquid crystal monitor or an organic electroluminescence (EL) display, and displays information to be notified to the user or the like in response to an instruction from the control unit 101 .

[0022] The information processing apparatus 100 in this embodiment may be a single computer, or may be a computer system configured with multiple computers and peripheral devices. The information processing apparatus 100 may be a virtual machine whose entity is virtualized, or may be a cloud. Furthermore, although the information processing apparatus 100 and the substrate processing apparatus 200 are described as separate entities in this embodiment, the information processing apparatus 100 may be provided inside the substrate processing apparatus 200.

[0023] The causal structure generated by the information processing device 100 will be described below. Fig. 3 is an explanatory diagram illustrating the internal representation of the causal structure. The causal structure is drawn, for example, by a directed acyclic graph using nodes representing each observed variable and edges representing causal relationships between the nodes. The upper part of Fig. 3 shows the causal structure drawn by the directed acyclic graph, and the lower part of Fig. 3 shows its internal representation.

[0024] Causal relationships between observed variables can be modeled using structural equation models. One type of structural equation model, the Linear Non-Gaussian Acyclic Model (LiNGAM), uses a linear acyclic model and assumes that the probability distribution of exogenous variables (error variables) is non-Gaussian.

[0025] The DirectLiNGAM algorithm has been proposed as one of the algorithms for exploring structural equation models (causal discovery algorithms) (see, for example, S. Shimizu et al. Journal of Machine Learning Research, 12(Apr): 1225-1248(2011)). Based on the above assumptions, this algorithm optimizes the linear regression coefficients (edge ​​weights) by repeating regression analysis and evaluating the independence of regression residuals. A directed acyclic graph can be drawn by drawing edges between nodes based on the optimized linear regression coefficients.

[0026] In causal search, edges may be drawn from multiple other nodes to one node unless the edges are circular. To avoid such overfitting, a constraint may be added that prohibits edges from being drawn from multiple collinear observation variables to the same observation variable at the same time.

[0027] The observation variables may include settings such as an apparatus recipe that specifies the process execution details, measurements from various sensors taken during the process, observations or measurements related to the state of the substrate processing apparatus 200 or the wear rate of parts, and observations or measurements indicating the results of the substrate processing. The information processing apparatus 200 can acquire, as observation variables, settings specified by its own control device, observations or measurements obtained from various sensors or measuring devices installed in the apparatus, and observations or measurements obtained from measuring devices installed outside the apparatus. For example, in an actual process in the substrate processing apparatus 200, various settings and observations, such as the coating of the light-receiving window, the OES (Optical Emission Spectrometer), wear of the lower electrode, wear of the upper electrode, the voltage setting, the voltage measurement value of the VI sensor, the current measurement value of the VI sensor, and the etching amount, can be treated as observation variables. In the example of FIG. 3 , for simplicity, six observation variables are selected, and only the causal relationships between the six selected observation variables are shown.

[0028] The directed acyclic graph (causal structure) shown in the upper part of Figure 3 is composed of nodes ND1 to ND6 corresponding to six observation variables (observation variables A to F) and multiple edges EG14, EG24, EG36, EG46, and EG56 representing the causal relationships between the observation variables (between nodes). In the example of Figure 3, nodes ND1 to ND6 are represented by regular octagonal icons, but the shape of the icons is not limited to a regular octagon and may be circular or another shape. In the example of Figure 3, the character strings shown inside the icons represent the variable names of each observation variable.

[0029] In the example of Figure 3, an edge EG14 is drawn connecting node ND1 and node ND4. This edge EG14 indicates that there is a causal relationship between the observation variable corresponding to node ND1 (observation variable A) and the observation variable corresponding to node ND4 (observation variable D). Because this edge EG14 is drawn from node ND1 to node ND4, it indicates that observation variable A has an effect on observation variable D. The same applies to the causal relationships between other observation variables.

[0030] A directed acyclic graph can be represented by a matrix. Figure 3 shows the directed acyclic graph representing the causal structure described above, along with the matrix elements of the corresponding matrix. In this diagram, the value of a matrix element indicates whether an edge exists or not, with a value of 0 indicating that an edge does not exist and a value other than 0 indicating that an edge exists. For example, the matrix element in the first row and fourth column has a value of 1.0, which indicates that an edge exists from the node of observation variable A represented by the index in the first row to the node of observation variable D represented by the index in the fourth column. Furthermore, the matrix elements in the first row and columns other than the fourth column have a value of 0, which indicates that no edges exist from the node of observation variable A to observation variables B, C, E, or F. The same applies to rows 2 to 6.

[0031] The values ​​of the matrix elements represent the edge weights (strength of the connections between nodes), and are calculated as linear regression coefficients in the DirectLiNGAM algorithm, for example. Edge weights can also be represented in directed acyclic graphs. For example, a numerical value representing the edge weight can be displayed on the graph, and the edge thickness (line width) or color can be changed depending on the edge weight.

[0032] As shown in the example of FIG. 3 , the edge weight can take a negative value. A negative edge weight indicates a negative correlation between the nodes. In other words, it represents a relationship in which, when the observation variable on the starting point side changes in a positive direction, the observation variable on the end point side changes in a negative direction. In the example of FIG. 3 , the matrix element in the third row and sixth column contains a negative value of -1.2. Therefore, when the observation variable C (the observation variable on the starting point side) changes in a positive direction, the observation variable F (the observation variable on the end point side) changes in a negative direction. This matrix element in the third row and sixth column represents an edge from the node of observation variable C to the node of observation variable F, not an edge from the node of observation variable F to the node of observation variable C. It should be noted that in this embodiment, the direction of the edge does not reverse depending on whether the edge weight is a positive value or a negative value.

[0033] 3, a method of describing a directed acyclic graph using a matrix has been described, but a directed acyclic graph may also be described using an edge list that lists nodes representing the start and end points of edges and edge weights. In either case, the graph is described in list format or array format within the information processing device 100.

[0034] In addition, a confidence interval is set for the edge weight. The confidence interval is set based on the probability distribution of the exogenous variable (error variable) calculated during causal search. FIG. 4 is an explanatory diagram illustrating an example of setting a confidence interval. Assuming that the edge weight has a probability distribution (e.g., a Gaussian distribution) defined by a center value and width, the confidence interval can be set as a range of ±α (α is a value from 0 to ∞) for the original weight. In the example of FIG. 4, the weight of the edge between the nodes corresponding to observation variables A and D is 1.0, and a confidence interval of ±0.1 is set for that weight.

[0035] The same is true when the edge weights are assumed to have a probability distribution such as a sigmoid distribution, and the confidence interval is set based on the probability distribution of the weights. In the example of Figure 4, the weight of the edge between the nodes corresponding to the observation variables C and F is -1.2, but if the weight value is set to at least less than 0 (i.e., if learning can result in a negative value such as -0.5, but not a positive value such as +1.0), the confidence interval is set to less than 0 (denoted as <0 in the figure).

[0036] In the above causal search, it is assumed that the relationship between observed variables is linear, but it is also possible to search for a causal structure that includes nonlinearity. Although the functional form is unknown, an algorithm that allows for searches that include nonlinearity is known, so it is possible to search for a causal structure that includes nonlinearity using this algorithm. When drawing a causal structure that includes nonlinearity, it is also possible to add a node to display the functional form between nonlinear nodes.

[0037] FIG. 5 is a schematic diagram showing an example of a causal structure including nonlinearity. The example in FIG. 5 shows that there is a nonlinear relationship between observed variables A and B and observed variable D, where observed variable D = observed variable A × observed variable B. In this case, node ND7 is added to display the function form A × B, and edges EG17, EG27, and EG74 connect nodes ND1, ND2, and ND4 to node ND7, respectively. To distinguish the nodes displaying the function form from normal observed variable nodes and edges between observed variables, the display mode may be changed by changing the color of the nodes or by displaying the edges connecting the nodes with dashed lines.

[0038] In this embodiment, when the causal structure of the reference system is obtained, a priori constraints are set based on the causal structure of the reference system, and learning with the a priori constraints imposed is performed using observation data obtained from the observation system, thereby deriving the causal structure for the observation system.

[0039] FIG. 6 is an explanatory diagram illustrating an example of a causal structure obtained by additional learning based on a generic structure. The upper part of FIG. 6 shows the generic structure. The generic structure refers to a reference-based causal structure created based on the user's knowledge or a reference-based causal structure derived from observation data of a specific substrate processing apparatus 200. If there is a mechanism that is understood as the user's knowledge without the need to collect observation data, the causal structure created based on the user's knowledge can be used as the generic structure. The user can create a generic structure by using the operation unit 104 of the information processing apparatus 100 to draw nodes corresponding to observation variables and edges connecting the observation variables, and assigning numerical values ​​to the edge weights (the strength of connections between nodes).

[0040] Furthermore, when there are a plurality of substrate processing apparatuses 200 having similar mechanisms, it is considered that the rough causal structure is common among the apparatuses, and therefore the causal structure of a specific substrate processing apparatus 200 created as a representative can be used as a general structure. In this case, the information processing apparatus 100 can create a general structure by collecting observation data from the specific substrate processing apparatus 200 and using the above-mentioned causal search algorithm.

[0041] Furthermore, the causal structure of the reference system created based on the user's knowledge may be corrected based on observation data obtained from the reference system's substrate processing apparatus 200, and the corrected causal structure may be used as a general-purpose structure. That is, the causal structure created by the user may be corrected by additional learning based on observation data actually obtained from the reference system using edge weights set by the user as initial values, and the corrected causal structure may be used as a general-purpose structure.

[0042] The above-mentioned general structure is created in advance, and the parameters of the general structure, including the edge weights and the confidence intervals for the weights (tolerance for variations due to machine differences), are stored in the storage unit 102 of the information processing device 100.

[0043] The lower part of Figure 6 shows an individual structure. The individual structure refers to an individual causal structure obtained by additional learning using the generic structure. By acquiring observation data from each substrate processing apparatus 200 and performing additional learning using the generic structure, it is possible to derive a causal structure (individual structure) for each apparatus. Compared to creating a causal structure from a state without prior knowledge, the amount of observation data required can be reduced, thereby improving the efficiency of learning.

[0044] In this embodiment, when additional learning is performed using a generic structure, a loss according to the confidence interval is imposed to adjust the weight in the individual structure so that it does not deviate significantly from the reference value (original weight). In the example of the generic structure in Figure 6, the weight of the edge between observation variables A and D is 1.0, and the confidence interval of that weight is set to ±0.1. Therefore, a loss of, for example, abs(1.0-w) / 0.1 is imposed to adjust the edge weight w in the individual structure so that a large loss is applied to weights greater than 1.1 or less than 0.9. In the example between observation variables A and D, the edge weight was 1.0 in the generic structure, but was adjusted to 1.03 in the individual structure.

[0045] 6, the edge weight between observation variables C and F is -1.2, and the confidence interval for that weight is set to less than the reference value (<0). Therefore, the edge weight in the individual structure is adjusted so that a large loss is applied to weights that are -1.2 or greater. In the example between observation variables C and F, the edge weight was -1.2 in the general structure, but was adjusted to -3.0 in the individual structure.

[0046] The procedure for deriving a causal structure using a general-purpose structure will be described below. Fig. 7 is a flowchart showing the procedure for deriving a causal structure using a general-purpose structure. The control unit 101 of the information processing device 100 reads out and executes the causal structure learning program PG1 from the storage unit 102, thereby performing the following processing.

[0047] The control unit 101 acquires observation data corresponding to a plurality of observation variables from the substrate processing apparatus 200 of the observation system (step S101). The observation data acquired by the control unit 101 includes data measured by the substrate processing apparatus 200 and data set in the substrate processing apparatus 200, such as the coating of the light-receiving window, the OES, wear of the lower electrode, wear of the upper electrode, the voltage setting value, the voltage measurement value of the VI sensor, the current measurement value of the VI sensor, and the etching amount. The control unit 101 acquires this observation data by communicating with the substrate processing apparatus 200 via the communication unit 103.

[0048] The control unit 101 reads out parameters of the generic structure from the storage unit 102 (step S102). In this embodiment, a causal structure created in advance for a substrate processing apparatus 200 of the same model as the substrate processing apparatus 200 in the observation system is used as the generic structure. In step S102, the control unit 101 reads out parameters for the generic structure, including edge weights and confidence intervals for the weights, from the storage unit 102.

[0049] The control unit 101 sets a priori constraints based on the parameters of the generic structure read from the storage unit 102 (step S103). The parameters of the generic structure include edge weights and confidence intervals for the weights, so the control unit 101 can set a priori constraints based on these parameters. For example, if the edge weight in the generic structure is w0, the confidence interval is ±α, and the weight of the individual structure to be calculated is w, a loss described by abs(w0-w) / α can be set as the priori constraint. Note that the loss is not limited to the above formula, and any significant difference from the prior distribution of the weights can be set as the loss.

[0050] The control unit 101 uses the observation data acquired in step S101 to perform learning with the prior constraints set in step S103 imposed, thereby deriving a causal structure of a group of observed variables in the observation system (step S104). The control unit 101 uses a causal search algorithm to impose prior constraints and search for causal relationships between the observed variables, thereby deriving a causal structure of all observed variables. Furthermore, the prior constraint may allow the creation of new edges, provided that the causal structure maintains directed acyclicity. By performing learning that allows the addition of edges, new knowledge can be obtained when new conditions are added, and structural changes, including unexpected ones, can be learned and visualized.

[0051] For example, when using the DirectLiNGAM algorithm, the control unit 101 optimizes the coefficients of linear regression by repeatedly performing regression analysis and evaluating the independence of regression residuals, assuming that there is linearity between observed variables, that the causal structure is acyclic, that the probability distribution of exogenous variables is non-Gaussian, and that different exogenous variables are independent of each other.The control unit 101 generates a directed acyclic graph by drawing edges between nodes based on the optimized coefficients.This obtains a causal structure between observed variables in which edges from multiple collinear nodes may coexist.

[0052] After deriving the causal structure in step S104, edge pruning may be performed. That is, the control unit 101 detects edges drawn from multiple other collinear nodes to one node from the causal structure derived in step S104, and if such edges exist, the control unit 101 may designate edges other than the most accurate as prohibited edges to add a constraint that prohibits edges from being drawn from multiple collinear observation variables to the same observation variable at the same time, and then continue learning of the causal search. By performing edge pruning, it is possible to prevent misidentification of edges and to learn a causal structure with higher reliability.

[0053] The control unit 101 outputs the derived causal structure (step S105). The control unit 101 may display the derived causal structure on the display unit 105, or may notify a user terminal (not shown) of the derived causal structure via the communication unit 103.

[0054] FIG. 8 is a schematic diagram showing an example of a display of a causal structure obtained by additional learning. When additional learning is performed by allowing the generation of new edges, new edges may be added from the general structure. The example of FIG. 8 shows a state in which edge EG13 is added as a result of additional learning using the causal structure shown in FIG. 3 as the general structure. Also, in the example of FIG. 8, edge EG56, whose edge weight has changed by more than the error, is shown by a thick line. If there are no changes in the equipment or process, it is expected that there will be no addition or disappearance of edges, and no fluctuations beyond the weight error. However, if these fluctuations are observed after additional learning, it can be recognized that there is an abnormality in the equipment or a fluctuation in the mechanism itself, and this can be used to investigate the cause, etc.

[0055] As described above, in the first embodiment, the causal structure of the observation system can be derived based on the observation data obtained from the observation system by utilizing the generic structure obtained in the reference system substrate processing apparatus 200. In the first embodiment, additional learning is performed based on the generic structure, so that the causal structure can be derived based on a small amount of observation data obtained from the observation system, thereby improving the efficiency of learning.

[0056] In the second embodiment, a configuration will be described in which a learned causal structure (general-purpose structure) is incorporated and an unknown disturbance factor is estimated based on learning using new observation data. Note that the overall configuration of the system and the internal configuration of the information processing device 100 are the same as those in the first embodiment, and therefore description thereof will be omitted.

[0057] Figure 9 is a schematic diagram showing an example of a causal structure when a disturbance factor occurs. If a causal structure consisting of nodes ND1 to ND6 as shown in Figure 9 is obtained through causal search, then since only node ND1 is connected upstream of node ND4, the observation variable D of node ND4 should be determined only by the observation variable A of node ND1. Similarly, since only node ND2 is connected upstream of node ND3, the observation variable C of node ND3 should be determined only by the observation variable B of node ND2.

[0058] However, if the fluctuations in the observed variables of nodes ND4 and ND3, which should be determined only by the observed variables A and B of nodes ND1 and ND2, cannot clearly be explained by them alone, it is presumed that a fluctuation factor exists. If it is presumed that a fluctuation factor exists, the unexplained element is extracted and defined as a disturbance node to complement the causal structure. The example in Figure 9 shows a state in which node NDC is complemented as a disturbance node for nodes ND3 and ND4.

[0059] Note that collinear causes may be considered to be common and merged to determine that the differences between observed variables A and B and observed variables D and C are caused by the same cause. In this way, by visualizing the nodes representing the causes, it is possible to confirm common influences and provide suggestions for investigating the causes.

[0060] FIG. 10 is a flowchart showing a procedure for estimating an unknown disturbance. The control unit 101 of the information processing device 100 derives a causal structure using the same procedure as in embodiment 1 and determines whether there is a variation factor that cannot be explained by the derived causal structure (step S201). For example, when a causal structure such as that shown in FIG. 9 is obtained, it is possible to obtain the knowledge that when observed variables A and B are varied, observed variables D and C should vary. However, in an actual experiment, if observed variables D and C do not vary even when observed variables A and B are varied, it can be determined that some variation factor exists. If it is determined that no variation factor exists (S201: NO), the control unit 101 terminates the processing according to this flowchart.

[0061] When the control unit 101 determines that there is a fluctuation factor that cannot be explained from the derived causal structure (S201: YES), the control unit 101 defines the fluctuation factor as a disturbance node (step S202).

[0062] The control unit 101 determines whether multiple disturbance nodes are defined (step S203). If it is determined that one disturbance node is defined (S203: NO), the control unit 101 complements and learns the single disturbance node, and generates a causal structure including the single disturbance node (step S204).

[0063] If it is determined that there are multiple disturbance nodes defined (S203: YES), the control unit 101 determines whether there is collinearity between the disturbance nodes (step S205). If it is determined that there is collinearity (S205: YES), the control unit 101 combines the collinear disturbance nodes into one and defines it as a common cause (step S206). The control unit 101 complements and learns a single or multiple disturbance nodes, and generates a causal structure including a single or multiple disturbance nodes (step S207).

[0064] If it is determined in step S205 that there is no collinearity between the disturbance nodes (S205: NO), the control unit 101 complements and learns the multiple disturbance nodes, and generates a causal structure including the multiple disturbance nodes (step S208).

[0065] As described above, in the second embodiment, by visualizing disturbance nodes, it is possible to confirm common influences and provide the user with suggestions for investigating the causes.

[0066] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any and all combinations, regardless of the reference format. Furthermore, the claims do not use a format in which a claim references two or more other claims (multiple claim format), but this is not limited to this. They may be written using a multiple claim format or a format in which multiple claims (multi-multi claim) reference at least one other multiple claim.

[0067] The embodiments disclosed herein are to be considered in all respects as illustrative and not restrictive. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims.

[0068] For example, in the embodiment, the substrate processing apparatus 200 has been described as an example of the reference system and the observation system. The observation system to be monitored is not limited to the substrate processing apparatus 200, but may be a manufacturing apparatus in which any manufacturing process is carried out for electrical equipment, chemical industrial products, pharmaceuticals, food, chemical industrial products, etc. Furthermore, the reference system and the observation system are not limited to an apparatus or system in which any manufacturing process is carried out, but may be any system that appropriately combines human living environments, economic activities, meteorological environments, etc.

[0069] REFERENCE SIGNS LIST 100 Information processing device 101 Control unit 102 Storage unit 103 Communication unit 104 Operation unit 105 Display unit PG1 Causal structure learning program RM Recording medium 200 Substrate processing device

Claims

1. A computer program for causing a computer to execute a process of acquiring observational data corresponding to multiple types of observational variables from an observation system, setting a priori constraints based on the causal structure of a group of variables in a reference system, and executing learning with the a priori constraints imposed using the acquired observational data, thereby deriving the causal structure of a group of observed variables in the observation system.

2. The computer program of claim 1, for causing the computer to execute the process of: acquiring parameters of the causal structure; and setting the prior constraints based on the acquired parameters; wherein the causal structure of the reference system is a causal structure created in advance based on a user's knowledge.

3. The computer program of claim 1, which causes the computer to execute the following processes: the causal structure of the reference system is a causal structure derived in advance based on observation data obtained by observing the reference system; acquiring parameters of the causal structure; and setting the prior constraints based on the acquired parameters.

4. The computer program of claim 1, which causes the computer to execute the following processes: the causal structure of the reference system is a causal structure created based on a user's knowledge; acquiring information on the causal structure and observational data obtained by observing the reference system; modifying the causal structure based on the acquired observational data; and setting the prior constraints based on parameters of the modified causal structure.

5. The computer program of claim 1, for causing the computer to execute the process of: generating a directed acyclic graph representing the causal structure of the group of observed variables using nodes representing each observed variable and edges representing causal relationships between the nodes; and outputting the generated directed acyclic graph.

6. The computer program according to claim 5, which causes the computer to execute a process of adding a node for displaying a function form between nonlinear nodes when the causal structure of the observation system includes nonlinearity.

7. The computer program according to claim 5, for causing the computer to execute a process of setting a loss according to a confidence interval for the weight between the nodes as the prior constraint.

8. The computer program according to claim 7, wherein the confidence interval is set as a numerical range reflecting a probability distribution of the weights.

9. The computer program of claim 5, wherein the prior constraint allows the creation of new edges on the condition that the causal structure maintains directed acyclicity.

10. The computer program according to claim 5, causing the computer to execute a process of comparing edge weights in the causal structure of the reference system with edge weights in the causal structure of the observation system derived by the learning, and judging that an abnormality in the observation system has been detected if the edge weights have changed by the learning by more than an error.

11. The computer program according to claim 1, for causing the computer to execute a process of estimating disturbance factors not included in the causal structure of the reference system based on observation data of the observation system.

12. The computer program of claim 1, wherein the reference system is a first substrate processing apparatus, and the observation system is a second substrate processing apparatus of the same model as the first substrate processing apparatus.

13. An information processing device comprising at least one processor, which acquires observation data corresponding to a plurality of types of observation variables from an observation system, sets a priori constraints based on a causal structure of a group of variables in a reference system, and performs learning with the a priori constraints imposed using the acquired observation data, thereby deriving the causal structure of a group of observed variables in the observation system.

14. An information processing method in which a computer executes the following process: acquiring observational data corresponding to multiple types of observed variables from an observation system; setting a priori constraints based on the causal structure of a group of variables in a reference system; and performing learning using the acquired observational data with the a priori constraints imposed, thereby deriving the causal structure of a group of observed variables in the observation system.

Citation Information

Patent Citations

  • Probabilistic inference device

    JP2010257269A

  • Device, method, and program for displaying graphs

    JP2020149299A

  • Analysis device, analysis method, and analysis program

    JP2020149301A

  • Non-linear causal modeling based on encoded knowledge

    WO2022104616A1

  • Computer program, information processing device, and information processing method

    WO2024117013A1