Computer program, information processing method, and information processing device
The computer program and information processing apparatus utilize dynamic mode decomposition with control input to accurately model target devices by optimizing model parameters and canceling out noise, addressing the limitations of conventional methods.
Patent Information
- Application Number
- PCT/JP2024/031456
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-09-02
- Publication Date
- 2025-06-05
AI Technical Summary
Existing technologies face challenges in accurately modeling target devices due to noise interference and limited prediction accuracy when using conventional dynamic mode decomposition methods.
A computer program and information processing apparatus that acquires time-series control input data and observation data from a target device, and uses dynamic mode decomposition with control input to calculate parameters for a model that predicts the time-evolution of observation data, thereby improving prediction accuracy and noise cancellation.
The proposed solution enables accurate modeling of target devices by optimizing model parameters to apply across multiple conditions, effectively canceling out noise and improving prediction accuracy.
Smart Images

Figure JP2024031456_05062025_PF_FP_ABST
Abstract
Description
Computer program, information processing method, and information processing device
[0001] The present disclosure relates to a computer program, an information processing method, and an information processing device.
[0002] Patent Document 1 proposes a state determination device that acquires data related to industrial machinery, creates a plurality of partial time-series data by sliding time-series data of physical quantities in the data related to the industrial machinery along the time axis based on the acquired data related to the industrial machinery, extracts a plurality of learning data including the plurality of partial time-series data, and performs machine learning using the extracted learning data to generate a learning model.
[0003] Japanese Patent Application Laid-Open No. 2020-128013
[0004] The present disclosure provides a computer program, an information processing method, and an information processing device that are expected to accurately model a target device or the like.
[0005] A computer program according to one embodiment causes a computer to execute a process of acquiring multiple sets of time-series control input data for a target device obtained by operating the target device under multiple conditions and time-series observation data observing the state of the target device, and calculating model parameters that predict the time evolution data of the observation data in accordance with the control input data and the observation data using dynamic mode decomposition based on the acquired multiple sets of time-series control input data and the observation data.
[0006] According to the present disclosure, it is expected that a target device or the like can be modeled with high accuracy.
[0007] FIG. 1 is a schematic diagram for explaining an overview of an information processing system according to an embodiment. FIG. 2 is a schematic diagram for explaining a difference between calculation of parameters of dynamic mode decomposition by the information processing system according to the embodiment and calculation of parameters of conventional dynamic mode decomposition. FIG. 3 is a block diagram showing a configuration of an information processing device in the embodiment. FIG. 4 is a flowchart showing an example of a procedure of parameter calculation processing performed by the information processing device according to the embodiment. FIG. 5 is a schematic diagram showing an example of a time-series data selection screen. FIG. 6 is a schematic diagram showing an example of a model evaluation screen.
[0008] Specific examples of information processing systems according to embodiments of the present disclosure will be described below with reference to the drawings. Note that the present disclosure is not limited to these examples, but is defined by the claims, and is intended to include all modifications within the meaning and scope of the claims.
[0009] [First Embodiment] <System Overview> Fig. 1 is a schematic diagram for explaining an overview of an information processing system according to this embodiment. The information processing system according to this embodiment is configured to include an information processing device 1 and a target device 3. The target device 3 is various devices (substrate processing devices) used in semiconductor manufacturing, such as a process chamber that performs processing (substrate processing) such as etching or film formation on semiconductor wafers (substrates), or a robot arm that transports wafers. However, the target device 3 may also be various devices other than devices used in semiconductor manufacturing.
[0010] The information processing device 1 is a device that controls and monitors the operation of the target device 3. In this embodiment, the target device 3 includes, for example, an electrostatic chuck for electrostatically fixing a wafer to be processed, a temperature control device such as a chiller (cooling device) and / or a heater (heating device) for controlling the temperature of the wafer, and a sensor for measuring the temperature of the wafer. The information processing device 1 according to this embodiment is connected to the target device 3 via, for example, a communication line or a signal line, and provides control input data (input: U) to the temperature control device of the target device 3 and acquires temperature observation data (output: X) from a sensor of the target device 3. For example, the information processing device 1 determines the control input data according to the observation data acquired from the target device 3 and provides the control input data to the target device 3. As a result, the temperature of the wafer fixed to the electrostatic chuck of the target device 3 is controlled to maintain a desired temperature.
[0011] In this embodiment, the information processing device 1 samples and acquires control input data and observation data at a predetermined frequency, such as once per second or once per minute, while processing is being performed on a wafer in the target device 3. The information processing device 1 stores the acquired control input data and observation data in a database. In this embodiment, the information processing device 1 inputs multiple control values to the target device 3 as control data. The target device 3 is also provided with multiple temperature sensors for observing the temperature distribution of the wafer, and the information processing device 1 acquires temperatures measured by the multiple temperature sensors at multiple locations on the wafer as observation data. Therefore, in this embodiment, the control input data and observation data stored in the database by the information processing device 1 are each multidimensional (two or more dimensions) vector information. The number of dimensions of the control input data and the observation data do not need to be equal. Note that either or both of the control input data and the observation data may be one-dimensional scalar values.
[0012] Although the present embodiment describes the acquisition of control input data and observation data using wafer temperature control as an example, the application of the information processing system according to the present embodiment is not limited to wafer temperature control. The information processing system can be applied to various processes performed by target apparatus 3 such as substrate processing apparatuses, and can acquire control input data and observation data for various processes, store them in a database, and control or predict these processes.
[0013] In this embodiment, after processing of a wafer is completed, the information processing device 1 performs a process of modeling the target device 3 based on the time-series data of control input data and observation data stored in a database, i.e., a process of generating a model that predicts output data based on the input data to the target device 3. In this embodiment, the information processing device 1 uses an improved version of the "dynamic mode decomposition" technique to model the target device 3. Dynamic mode decomposition is a technique that decomposes time-series data into multiple fluctuation factors (modes), and by using the dynamic mode decomposition technique, a model that predicts time-series data can be generated. The generated model accepts, for example, observation data at a certain point in time as input and outputs a predicted value of time-evolving data, such as a first-order differential value of the observation data at a subsequent point in time. The information processing device 1 performs dynamic mode decomposition using the time-series data stored in the database, thereby determining the parameters of the model that outputs this predicted value. Note that dynamic mode decomposition is a known analysis method (see, for example, "Jonathan H. Tu, Clarence W. Rowley, Dirk M. Luchtenburg, Steven L. Brunton, J. Nathan Kutz. On dynamic mode decomposition: Theory and applications. Journal of Computational Dynamics, 2014, 1 (2): 391-421."), so detailed explanation will be omitted. Improvements to dynamic mode decomposition in this embodiment will be described later.
[0014] Furthermore, in this embodiment, the information processing device 1 uses a "dynamic mode decomposition with control input" technique, which enables dynamic mode decomposition to handle additional time-series data of control inputs. By using the dynamic mode decomposition with control input, the information processing device 1 can generate a model that receives control input data and observation data at a certain point in time as input and outputs a predicted value of time-evolved data of the observation data at a subsequent point in time. Note that the dynamic mode decomposition with control input is a known analysis technique (see, for example, "Joshua L. Proctor, Steven L. Brunton, J. Nathan Kutz. "Dynamic mode decomposition with control," SIAM Journal on Applied Dynamical Systems 15.1 (2016)"). Therefore, a detailed description thereof will be omitted. In the following description, even when the term "dynamic mode decomposition" is simply used, this term also includes "dynamic mode decomposition with control input."
[0015] Furthermore, the dynamic mode decomposition method in this embodiment may use "Dynamic mode decomposition with memory" (Ryoji Anzaki, Kei Sano, Takuro Tsutsui, Masato Kazui, and Takahito Matsuzawa, Physical Review E 108, 034216). By using this method, past information can be incorporated into the model, and a better model may be obtained than with ordinary dynamic mode decomposition. Similarly, known methods such as Hankelized dynamic mode decomposition (Hankelized DMD) may be combined.
[0016] The information processing device 1 performs dynamic mode decomposition processing based on control input data input to the target device 3 in a time series and observation data output in a time series by a sensor of the target device 3. This allows the information processing device 1 to generate a model that accepts control input data and observation data at a certain point in time as input and outputs a predicted value of observation data at a subsequent point in time or data obtained by performing some operation on the observation data (such data is referred to as time-evolution data of the observation data). The same model can also be used to make future predictions when the control input data to be applied thereafter is predetermined. In general dynamic mode decomposition, a model that accepts control input data and observation data as input and outputs a predicted value of a first-order difference (differential) value of the observation data as time-evolution data is often used. However, a model that outputs a predicted value as a result of more general time-related transformation, such as the aforementioned dynamic mode decomposition with history, may also be used.
[0017] The time evolution data of the observation data may be measured by separate observations. For example, if the velocity and acceleration of an object can be observed separately, the velocity may be used as the observation data, and the observed acceleration may be used as the time evolution data of the observation data.
[0018] In the information processing system according to this embodiment, the model is assumed to have a linear relationship between the control input data (input: U) and observation data (output: X) and the time-evolution data (Y) of the observation data. The model handled by the information processing system according to this embodiment is expressed, for example, by the following equation (1). In equation (1), U, X, and Y are matrices created by aggregating data within a target time period, with values arranged so that the index of the feature dimension is in the vertical direction and the time of observation or input increases horizontally. In equation (1), A and B are parameter matrices, and in this embodiment, the information processing device 1 calculates these values to obtain a model of the target device 3.
[0019] Y = AX + BU…(1)
[0020] The model generated by the information processing device 1 does not necessarily have to be configured to use control input data and observation data related to the target device 3 as input and output. The model may also use data that has undergone some kind of arithmetic processing (preprocessing) on time series data, such as the average value or differential value of the time series data, as input and output. For example, the model may be configured to accept as input the average value of time series data at a certain point in time and for a predetermined time range prior to that point in time. Furthermore, for example, the model may be configured to accept as input the first differential value of time series data at a certain point in time.
[0021] The information processing device 1 performs dynamic mode decomposition based on time-series control input data U and observation data X that are acquired while the target device 3 is processing a single wafer and stored in a database. This allows the information processing device 1 to calculate the values of internal parameters A and B of a model that accepts the control input data U and observation data X at a certain point in time as input and outputs a predicted value of time-evolved data Y of the observation data at a subsequent point in time. The information processing device 1 can predict, for example, future operations or processing results of the target device 3 using the model based on the calculated parameters A and B, and control the operation of the target device 3 by feeding back the prediction result.
[0022] 2 is a schematic diagram illustrating the difference between the calculation of parameters for dynamic mode decomposition by the information processing system according to the present embodiment and the calculation of parameters for dynamic mode decomposition in the related art. For example, a user of the information processing system may wish to operate the target device 3 under various conditions to collect time-series control input data and observation data, and to generate a model that predicts the operation of the target device 3 under various conditions using the collected time-series data.
[0023] 2 shows a conventional method for calculating the parameters of such a model using dynamic mode decomposition. In this example, the target device 3 is operated under n conditions, from condition 1 to condition n (n is an integer equal to or greater than 2), and the information processing device 1 collects n types of time-series control input data U1, ..., Un and corresponding n types of time-series observation data X1, ..., Xn. Based on the obtained observation data X1, ..., Xn, the information processing device 1 can calculate time-evolution data Y1, ..., Yn of the observation data.
[0024] In the conventional method, dynamic mode decomposition is performed for each condition using time-series control input data U, observation data X, and time-evolution data Y, thereby calculating model parameters A and B corresponding to each condition. That is, in the conventional method, n types of parameters A1, ..., An and parameters B1, ..., Bn are calculated under n types of conditions, and the n types of parameters A1, ..., An and parameters B1, ..., Bn are integrated into one type of parameter A and B by, for example, calculating their average value. However, in such a conventional method, the effects of noise contained in the control input data, observation data, and time-evolution data of the observation data do not cancel each other out well, and the prediction accuracy of the model using the obtained parameters A and B may be low.
[0025] In contrast, the lower part of Figure 2 shows a parameter calculation method used by the information processing system according to this embodiment. The information processing system according to this embodiment performs dynamic mode decomposition processing using n time-series control input data U1, ..., An, observation data A1, ..., An, and time-evolution data Y1, ..., Yn acquired for n types of conditions at once, and directly calculates model parameters A and B corresponding to the n types of conditions. As a result, the values of parameters A and B are optimized to apply to all conditions, and the prediction accuracy of the model using the obtained parameters A and B can be expected to improve. Furthermore, because the contribution of noise between multiple experimental conditions cancels out, the prediction accuracy of the model can also be expected to improve by using data obtained by performing similar or the same experimental conditions multiple times.
[0026] <Device Configuration> Fig. 3 is a block diagram showing the configuration of the information processing device 1 in this embodiment. The information processing device 1 according to this embodiment is configured to include a processing unit 11, a storage unit 12, a communication unit 13, a display unit 14, and an operation unit 15. Note that in this embodiment, the processing will be described as being performed by a single information processing device 1, but the processing of the information processing device 1 may be distributed among a plurality of devices.
[0027] The processing unit 11 is configured using an arithmetic processing device such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit) or a quantum processor, a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The processing unit 11 reads and executes a program 12a stored in the storage unit 12 to perform various processes such as a process of acquiring time-series data related to the processing of the target device 3 and a process of calculating model parameters based on the acquired time-series data.
[0028] The storage unit 12 is configured using a large-capacity storage device such as a hard disk. The storage unit 12 stores various programs executed by the processing unit 11 and various data required for the processing of the processing unit 11. In this embodiment, the storage unit 12 stores a program 12a executed by the processing unit 11. The storage unit 12 also includes a time-series data storage unit 12b that stores and accumulates time-series control input data and observation data collected in conjunction with the operation of the target device 3.
[0029] In this embodiment, the program (computer program, program product) 12a is provided in a form recorded on a recording medium 99 such as a memory card or an optical disc, and the information processing device 1 reads the program 12a from the recording medium 99 and stores it in the storage unit 12. However, the program 12a may also be written to the storage unit 12, for example, during the manufacturing stage of the information processing device 1. Alternatively, the program 12a may be distributed by a remote server device or the like and acquired by the information processing device 1 via communication. For example, the program 12a may be read from the recording medium 99 by a writing device and written to the storage unit 12 of the information processing device 1. The program 12a may be provided in a form distributed via a network or in a form recorded on the recording medium 99.
[0030] The time-series data storage unit 12b of the storage unit 12 stores and accumulates the control input data to the target device 3 and the observation data from the target device 3 that are sampled and acquired by the information processing device 1. The information processing device 1 associates the acquired time-series control input data and observation data with various information such as the date and time when the time-series data was acquired, identification information of the wafer to be processed, and the operating conditions of the target device 3, and stores them in the time-series data storage unit 12b.
[0031] The communication unit 13 is connected to the target device 3 via a cable such as a communication line or a signal line, and transmits and receives data to and from the target device 3 via this cable. In this embodiment, the communication unit 13 transmits control input data provided by the processing unit 11 to the target device 3. The communication unit 13 also receives observation data transmitted from the target device 3 and provides the received observation data to the processing unit 11.
[0032] The display unit 14 is configured using a liquid crystal display or the like, and displays various images, characters, etc. based on the processing of the processing unit 11. In this embodiment, the display unit 14 displays various types of information, such as time-series control input data and observation data related to the target device 3, information such as model parameters calculated based on this time-series data, and prediction results of the operation of the target device 3 using the generated model.
[0033] The operation unit 15 accepts user operations and notifies the processing unit 11 of the accepted operations. For example, the operation unit 15 accepts user operations using input devices such as mechanical buttons or a touch panel provided on the surface of the display unit 14. Furthermore, for example, the operation unit 15 may be input devices such as a mouse and a keyboard, and these input devices may be configured to be detachable from the information processing device 1.
[0034] The storage unit 12 may be an external storage device connected to the information processing device 1. The information processing device 1 may be a multi-computer including multiple computers, or may be a virtual machine virtually constructed by software. The information processing device 1 is not limited to the above configuration, and may not include, for example, the display unit 14 and the operation unit 15.
[0035] In addition, in the information processing device 1 according to this embodiment, the processing unit 11 reads and executes the program 12a stored in the memory unit 12, whereby the control processing unit 11a, data acquisition unit 11b, data selection unit 11c, parameter calculation unit 11d, display processing unit 11e, etc. are realized in the processing unit 11 as software functional units.
[0036] The control processing unit 11a performs processing to control the operation of the target device 3, for example, by generating control input data and transmitting it to the target device 3 according to predetermined conditions for the semiconductor manufacturing process. For example, the control processing unit 11a provides data such as the amount of temperature increase or decrease or the amount of device operation as control input data to one or more temperature control devices such as heaters and chillers provided in the target device 3. This allows the information processing device 1 to control the temperature of wafers processed by the target device 3.
[0037] The data acquiring unit 11b performs processing to acquire control input data input by the control processing unit 11a to the target device 3 and observation data measured by one or more temperature sensors provided in the target device 3. The data acquiring unit 11b acquires these data by sampling at a predetermined period, for example, once per second or once per minute. The data acquiring unit 11b stores the acquired control input data and observation data in the time-series data storage unit 12b, together with information on the operating conditions of the target device 3, the date and time the data was acquired, and identification information of the wafer being processed. The data acquiring unit 11b acquires data at a predetermined period and stores it in the time-series data storage unit 12b, whereby time-series data of the control input data and observation data is accumulated in the time-series data storage unit 12b.
[0038] The data selector 11c selects time-series data to be used in determining model parameters by dynamic mode decomposition from among multiple sets of time-series control input data and observation data stored in the time-series data storage unit 12b. In this embodiment, the data selector 11c displays a list of information about the multiple time-series data stored in the time-series data storage unit 12b, such as the operating conditions when the time-series data was acquired, the acquisition date and time of the time-series data, and identification information of the wafer to be processed, on the display unit 14. The data selector 11c accepts a selection of one or more sets of time-series data from the multiple sets of time-series data displayed in the list from a user based on an operation on the operation unit 15. However, the data selector 11c may select time-series data according to a predetermined rule, such as data acquired within one week or data under specific operating conditions, regardless of the user's selection, or may select all time-series data stored in the time-series data storage unit 12b as the processing target.
[0039] The parameter calculation unit 11d performs a process of calculating model parameters by performing dynamic mode decomposition on the multiple sets of control input data and observation data selected by the data selection unit 11c. In this embodiment, the model is a model that receives time-series observation data X and control input data U as input and outputs a predicted value of time-evolution data Y of the observation data X, and is a model in which there is linear dependency between the input and output as shown in the above equation (1). The parameter calculation unit 11d according to this embodiment can calculate model parameters corresponding to multiple operating conditions by calculating model parameters using multiple sets of control input data U and observation data X that have different operating conditions.
[0040] The parameter calculation unit 11d first calculates time-series time evolution data Y of the observation data X based on each set of time-series observation data X. For example, when the first-order difference value of the observation data X is used as the time evolution data Y, the parameter calculation unit 11d can calculate the time evolution data Y of the first-order difference value by calculating the difference value between the observation value at a certain time point and the observation value at the next time point for the time-series observation data X. However, the time evolution data Y is not limited to the first-order difference value of the observation data X, and various values such as a first-order differential value, a second-order differential value, a second-order differential value, or a non-integer differential value may be used. Furthermore, for example, the time evolution data Y may be a value obtained by advancing the time of the corresponding column of the observation data X by one or more steps. Furthermore, by using the aforementioned history-based dynamic mode decomposition, more general values can be used. The parameter calculation unit 11d calculates the corresponding time evolution data Y for all sets of observation data X used in processing.
[0041] If there is an adjustable parameter φ in the processing of the parameter calculation unit 11d, the loss function corresponding to the optimal A and B matrices may be calculated with φ fixed, and then φ may be adjusted so that the loss function is optimized. Furthermore, the values to be adopted as the time-evolution data of the observation data may be determined based on knowledge of the target device.
[0042] The parameter calculation unit 11d performs dynamic mode decomposition processing using a plurality of sets of time-series data of the obtained control input data U, observation data X, and time-evolution data Y, to calculate model parameters A and B. In this embodiment, the parameter calculation unit 11d calculates the parameters A and B by performing processing to optimize a quadratic function defined as a loss function related to the model of equation (1). As will be described in detail later, the parameter calculation unit 11d calculates a matrix (Jacobian matrix, Jacobian) representing the first-order derivative of the loss function and a matrix (Hessian matrix, Hessian) representing the second-order derivative of the loss function, and is thereby able to calculate the parameters A and B based on these matrices.
[0043] In this embodiment, the parameter calculation unit 11d also calculates the condition number for the calculated parameters A and B. The condition number of a matrix is an index that indicates how much the error contained in the data will expand when a computer performs an operation to solve a linear equation using this matrix. The smaller the value of the condition number, the smaller the error, and the larger the value, the larger the error. In this embodiment, the condition number of the Hessian used in the process of calculating the parameters A and B is calculated, and this is presented to the user as the condition number for the parameters A and B.
[0044] The display processing unit 11e performs processing to display various information such as various information related to the control processing by the control processing unit 11a, the time-series data acquired by the data acquisition unit 11b, and the model parameters calculated by the parameter calculation unit 11d on the display unit 14. Furthermore, when the data selection unit 11c accepts selection of time-series data from the user, the display processing unit 11e displays a list of information related to multiple sets of time-series data stored in the time-series data storage unit 12b on the display unit 14.
[0045] <Parameter Calculation Process> The information processing system according to this embodiment operates the target device 3 under various conditions to acquire time-series control input data and observation data, and performs a process of calculating parameters A and B of a model expressed by equation (1) shown again below. In equation (1), the time-series observation data X is, for example, time-series measurement value data measured by a plurality of sensors mounted on the target device 3, and forms a matrix in which measurement values from one sensor are arranged in time series in the row direction, and the number of columns is equal to the number of sensors. The time-series control input data U and time-evolution data Y also form matrices with similar configurations. The parameters A and B are matrices whose vertical and horizontal sizes are determined according to the configurations of matrices U, X, and Y.
[0046] Y = AX + BU…(1)
[0047] The model given by the above formula (1) is a model in which the dependency from the input of control input data U and observation data X to the output of time-evolution data Y is linear. In order to calculate the parameters A and B of this model, the dynamic mode decomposition process according to this embodiment uses the loss function L shown in the following (2).
[0048] In general dynamic mode decomposition processing, the loss function LDMD shown in (2') below is used. However, since this value is non-negative, the result is the same even if we consider optimizing the loss function shown in (2) below, which is defined as the square of this value. Therefore, in this embodiment, the loss function defined in (2) is used.
[0049] Furthermore, the right-hand side of equation (2) represents the square of the Frobenius norm of (Y-AX-BU), and the loss function L is a quadratic function of the parameters A and B. A typical dynamic mode decomposition process optimizes this loss function L, that is, calculates the parameters A and B that minimize the value of the loss function L.
[0050]
[0051] In the information processing system according to this embodiment, parameters A and B are calculated using time-series data acquired for the target device 3 under multiple operating conditions, so the loss function L in the above equation (2) is expanded to the following equation (3). In equation (3), μ is a subscript that distinguishes between the operating conditions under which the time-series data was acquired and is an integer value ranging from 1 to the number of operating conditions. The loss function L in equation (3) corresponds to the sum of the losses calculated under each operating condition and for all operating conditions.
[0052]
[0053] The parameters A and B that optimize the loss function L in equation (3) are calculated using the following equation (4). The left side of equation (4) is a matrix created by arranging the matrix of parameter A and the matrix of parameter B side by side. J included in the right side of equation (4) is a matrix (Jacobian matrix) representing the first-order derivative of the loss function L, and H is a matrix (Hessian matrix) representing the second-order derivative of the loss function L. While the Jacobian is generally expressed as a vector, in this embodiment, it is written in the above form taking advantage of the characteristics of the loss function. Furthermore, the "+" added to the upper right of H on the right side is a symbol indicating a generalized inverse matrix of the matrix, and the right side of equation (4) indicates that the generalized inverse matrix of H is used in the calculation.
[0054]
[0055] The method for calculating the matrix J in equation (4) is shown in equation (5) below, and the method for calculating the matrix H is shown in equation (6) below.
[0056]
[0057] The information processing device 1 can calculate the matrix [AB] on the left side of equation (4) by using the above equations (4) to (6) based on the control input data U and observation data X acquired for a plurality of operating conditions and the time-evolution data Y generated based on this observation data X. The matrix [AB] is a matrix created by arranging matrix A and matrix B horizontally. The matrix parameters A and B can be obtained by sequentially extracting and reconstructing each component of the matrix [AB].
[0058] Furthermore, in this embodiment, the information processing device 1 calculates the condition number κ(H) of the matrix H calculated by the above equation (6). Since the matrix H is a positive semidefinite symmetric matrix, the condition number κ(H) can be calculated as the ratio of the maximum eigenvalue to the minimum eigenvalue of the matrix H. That is, the condition number κ(H) = (maximum eigenvalue) / (minimum eigenvalue).
[0059] The information processing device 1 stores information such as the calculated parameters A and B and the condition number κ(H) in the storage unit 12, and displays it on the display unit 14 for presentation to the user. The information processing device 1 may also predict time-evolved data Y based on the control input data U and the observation data X using the model of equation (1) constructed based on the calculated parameters A and B, and may display a graph showing the predicted values and actual measured values of the time-series time-evolved data Y, as well as the error between the predicted values and the actual measured values. These information displays enable the user to consider the accuracy of modeling of the target device 3 or whether the data required for modeling is excessive or insufficient, and are expected to collect additional data in order to improve the accuracy of the model, for example.
[0060] 4 is a flowchart showing an example of the procedure of a parameter calculation process performed by the information processing device 1 according to this embodiment. In the information processing system according to this embodiment, the data acquisition unit 11b of the processing unit 11 of the information processing device 1 acquires control input data U and observation data X when the target device 3 is operated, and stores the acquired time-series data in the time-series data storage unit 12b. In addition, the target device 3 is operated under different conditions, and the information processing device 1 acquires time-series data under multiple operating conditions and stores the data in the time-series data storage unit 12b.
[0061] The data selection unit 11c of the processing unit 11 of the information processing device 1 according to this embodiment displays information about the acquired time-series data stored in the time-series data storage unit 12b on the display unit 14 (step S1). The data selection unit 11c accepts from the user a selection of one or more time-series data to be used for calculating parameters A and B from the displayed list of the time-series data (step S2).
[0062] The parameter calculation unit 11d of the processing unit 11 reads one or more pieces of time-series data (control input data U and observation data X) selected in step S2 from the time-series data storage unit 12b (step S3). The parameter calculation unit 11d calculates time-evolution data Y of the observation data X based on the observation data X acquired in step S3 (step S4). It is assumed that the rules or arithmetic expressions for calculating the time-evolution data Y from the observation data X are determined in advance by the user and provided in advance to the information processing device 1 as setting information.
[0063] Next, the parameter calculation unit 11d calculates a matrix representing the first-order derivative of the loss function L of the model, i.e., the Jacobian J, using the above-mentioned equation (5) based on the control input data U and observation data X read in step S3 and the time-evolution data Y calculated in step S4 (step S5). The parameter calculation unit 11d also calculates a matrix representing the second-order derivative of the loss function L of the model, i.e., the Hessian H, using the above-mentioned equation (6) based on the control input data U and observation data X read in step S3 and the time-evolution data Y calculated in step S4 (step S6). The parameter calculation unit 11d calculates model parameters A and B using the above-mentioned equation (4) based on the Jacobian J calculated in step S5 and the Hessian H calculated in step S6 (step S7). The parameter calculation unit 11d also calculates the condition number κ(H) based on the Hessian H calculated in step S6 (step S8).
[0064] The parameter calculation unit 11d stores the model parameters A and B calculated in step S7 and the condition number κ(H) calculated in step S8 in the storage unit 12 (step S9). The display processing unit 11e of the processing unit 11 displays information about the parameters A and B and the condition number κ(H) calculated by the parameter calculation unit 11d on the display unit 14 (step S10), and the processing ends.
[0065] <User Interface> The information processing device 1 according to this embodiment operates the target device 3 under various conditions, samples and acquires the control input data U and observation data X at that time, and stores them as time-series data in the time-series data storage unit 12b. The information processing device 1 stores and accumulates sets of time-series control input data U and observation data X obtained under various operating conditions in the time-series data storage unit 12b. In the information processing system according to this embodiment, the user can select time-series data to be used for calculating model parameters A and B from the time-series data under various conditions accumulated by the information processing device 1.
[0066] FIG. 5 is a schematic diagram illustrating an example of a time-series data selection screen. The information processing device 1 according to the present embodiment displays the illustrated time-series data selection screen on the display unit 14 and accepts a user's selection of time-series data to be used for calculating parameters A and B. On the time-series data selection screen, the information processing device 1 displays a list of information related to the time-series data stored in the time-series data storage unit 12b, such as "operating conditions," "date," and "ID." The "operating conditions" indicate the conditions under which the target device 3 was operated, and may include, for example, names or other information designated by the user for the operating conditions (simply indicated as A1, A2, etc. in the figure). The "date" may be, for example, information such as the date on which the time-series data was acquired or the date on which the target device 3 was operated, and may also include time information. The "ID" is, for example, identification information uniquely assigned to multiple time-series data. In this example, two-digit numerical information is used, but it may also include characters.
[0067] The information processing device 1 also displays check boxes for "learning" and "evaluation" in association with information about each piece of time-series data displayed in the list. The "learning" check box is a check box for selecting time-series data to be used for model learning, i.e., for calculating parameters A and B of the model. The "evaluation" check box is a check box for selecting time-series data to be used when evaluating the model in which the calculated parameters A and B are set. The information processing device 1 uses the time-series data for which these check boxes are checked for calculating parameters A and B and for evaluating the model.
[0068] In addition, in the upper right portion of the time series data selection screen, there are provided a check box for "Evaluate with training data" and a check box for "Select all." The check box for "Evaluate with training data" is a check box for calculating parameters A and B and subsequently evaluating the model using the same time series data. When the check box for "Evaluate with training data" is checked, if the user checks either the "Training" or "Evaluation" check box corresponding to each time series data, the other check box is automatically checked. The check box for "Select all" is a check box for using all the listed time series data for training and evaluation. If the user checks the "Select all" check box, all the "Training" and "Evaluation" check boxes are automatically checked.
[0069] The information processing device 1 reads out one or more pieces of time-series data (control input data U and observation data X) selected for "learning" on the time-series data selection screen from the time-series data storage unit 12b, calculates time-evolution data Y based on the read observation data X, and calculates model parameters A and B using the control input data U, observation data X, and time-evolution data Y. The information processing device 1 also constructs a model based on the calculated parameters A and B, reads out one or more pieces of time-series data selected for "evaluation" from the time-series data storage unit 12b, and evaluates the model. In the evaluation, the information processing device 1 inputs the read control input data U and observation data X into the model, and displays information obtained by comparing the time-evolution data Y (actual measured value of the time-evolution data) calculated by directly calculating a differential value or the like based on the read observation data X with the time-evolution data Y (predicted value of the time-evolution data) output by the model in response to the input.
[0070] FIG. 6 is a schematic diagram showing an example of a model evaluation screen. The information processing device 1 according to the present embodiment displays the illustrated model evaluation screen on the display unit 14 to provide a user with information such as calculation results of model parameters A and B and model evaluation results. The information processing device 1 displays the calculation results of parameters A and B, for example, in the upper part of the model evaluation screen. In the illustrated example, the model parameters A and B are each assumed to be a 48×48 matrix, and the information processing device 1 displays, as calculation results for parameters A and B, graphs in which rectangular regions arranged in a 48×48 matrix are color-coded according to the components of the corresponding parameters A and B, known as heat maps. In the illustrated model evaluation screen, the "coefficient matrix A estimation result" is the heat map of parameter A, and the "coefficient matrix B estimation result" is the heat map of parameter B. In addition to these heat maps, the information processing device 1 also displays information on the condition number calculated for the Hessian matrix H obtained in the process of calculating parameters A and B.
[0071] The information processing device 1 also displays the model evaluation results, for example, at the bottom of the model evaluation screen. In the illustrated example, four graphs are displayed side by side as "prediction evaluation results." In each graph, the actual measured values of the time-evolution data Y are indicated by solid lines, and the predicted values are indicated by dotted lines. In this embodiment, the observation data X output by the target device 3 is data in a matrix format in which vectors are arranged, with multiple values included in the data at a single point in time. The time-evolution data Y of the observation data X is data in a matrix format in which vectors of a size determined by the dimension of the observation data X and the size of the parameter A are arranged. In this example, the four graphs displayed as "prediction evaluation results" are graphs of the actual measured values and predicted values for four components selected from the vector components of the time-evolution data Y. The user can select which components of the time-evolution data Y to display by entering a number associated with each component in a numeric input box labeled "feature to display." Each graph is also labeled with a number associated with the evaluation data to identify which evaluation data it corresponds to. The information processing device 1 also calculates the errors between the actual measured values and the predicted values for the components of the four displayed graphs, calculates the average of the four errors, and displays the numerical value as the error of the predicted evaluation result. In this example, the information processing device 1 calculates the RSSE (Root Sum of Squared Errors) error, but any type of error may be calculated.
[0072] Note that the information displayed as the evaluation result of the model by the information processing device 1 is not limited to the above. On the model evaluation screen, the information processing device 1 can display various information, such as the display of the eigenvalues of the Hessian or their reciprocals, the display of a heat map of the Hessian, or the display of a heat map or graph of the uncertainty of each component of the parameters A and B.
[0073] The uncertainty of each component of parameters A and B can be calculated using the inverse matrix H of the Hessian matrix H obtained in the process of calculating parameters A and B. Note that in the following, if the Hessian is not a regular matrix, the information processing device 1 may use the generalized inverse matrix H of the Hessian instead of the inverse matrix, or may display a message that the estimation error is extremely large and cannot be calculated numerically. The information processing device 1 calculates matrices W and D by diagonalizing matrix H based on the following equation (7), for example. Note that in equation (7), the "*" attached to the upper right of the matrix represents the adjoint matrix.
[0074]
[0075] If the component of interest in the matrix [AB] is the (i, j) component, the uncertainty of that component does not depend on the vertical index (row number) i, but can be described using only the horizontal index (column number) j, and can be calculated, for example, by the following equation (8): where wj is the j-th column vector of the matrix W.
[0076]
[0077] The estimation uncertainty R of the entire matrix [AB] can be calculated using, for example, the diagonal elements λ1, ..., λn of the diagonal matrix D and the vertical size n of the matrix [AB] as shown in the following equation (9).
[0078]
[0079] There are various other methods for calculating the uncertainty of a component. For example, the following formula (10) or (11) can be used for calculation.
[0080]
[0081] Alternatively, functions including the harmonic mean of λ1,...,λn, polynomials, or constants defined by the user may be used. Also, when calculating the uncertainty of the matrix [AB], different calculation methods may be used for each component, such as using different calculation methods for parameters A and B.
[0082] Furthermore, the information processing device 1 may display information such as suggestions for improving the estimation accuracy of the model based on the uncertainty of each calculated component of parameters A and B. For example, when it is determined that a certain component of parameters A and B is uncertain (the calculated uncertainty exceeds a predetermined threshold), the information processing device 1 identifies the component of the control input data X related to this component. If this component is, for example, the 20th component, the information processing device 1 may display a suggestion message such as "The estimation accuracy is low, so please apply an input to the 20th component and conduct the experiment again."
[0083] Based on this information displayed by the information processing device 1, a user can repeatedly conduct experiments to collect data to generate a more accurate model. A model generated based on the time-series data collected by the information processing device 1 is provided, for example, in the target device 3 or the information processing device 1 that controls it, and is used for predicting the operation of infrastructure processing and detecting anomalies. For example, the target device 3 or the information processing device 1 can acquire control input data U input to the target device 3 and observation data X observed by a sensor or the like of the target device 3, and input the acquired control input data U and observation data X into a model to obtain a predicted value of time-evolved data Y. Based on the obtained predicted value of the time-evolved data Y, the target device 3 or the information processing device 1 can predict future operation of the target device 3 and reflect this in control, or can perform processing such as stopping the operation of the target device 3 if an abnormal predicted value is obtained.
[0084] <Summary> In the information processing system according to the present embodiment having the above configuration, the information processing device 1 acquires multiple sets of time-series control input data U and observation data X obtained by operating a target device such as a substrate processing device under multiple conditions. The information processing device 1 calculates parameters A and B of a model that predicts time-evolution data Y of the observation data X in accordance with the control input data U and the observation data X by dynamic mode decomposition based on the acquired multiple sets of control input data U and observation data X. As a result, the information processing system according to the present embodiment can generate a model that predicts the operation of the target device 3 in response to multiple different operating conditions, even when, for example, there are constraints on the input of the device or the influence of noise is significant, and is expected to accurately model the target device 3.
[0085] Furthermore, the information processing system according to this embodiment employs a model in which the dependency from the input of control input data U and observation data X to the output of time-evolved data Y is linear as a model for predicting the operation of the target device 3. This is expected to make it easier for the information processing system according to this embodiment to calculate the parameters A and B of the model based on time-series data under multiple conditions.
[0086] Furthermore, in the information processing system according to this embodiment, the information processing device 1 calculates a matrix J representing a first-order derivative and a matrix H representing a second-order derivative for a quadratic function defined as the loss function L of the model, and calculates parameters A and B based on the calculated matrices J and H. As a result, the information processing system according to this embodiment is expected to calculate parameters A and B easily and quickly.
[0087] Furthermore, in the information processing system according to this embodiment, the information processing device 1 calculates, as the condition number for the calculated parameters A and B, the condition number κ(H) of the Hessian H obtained in the process of calculating the parameters A and B. As a result, the information processing system according to this embodiment can be expected to provide information that enables a user to determine the accuracy of the calculated parameters A and B, etc.
[0088] Furthermore, in the information processing system according to this embodiment, the information processing device 1 displays a list of information about the time series data stored in the time series data storage unit 12b on the display unit 14, accepts selection of one or more time series data from the user, and calculates model parameters A and B based on the selected time series data. The information processing device 1 may also accept selection of time series data to be used for calculating the model parameters A and B, and selection of time series data to be used for evaluating the model using the calculated parameters A and B. As a result, the information processing system according to this embodiment is expected to assist a user who performs modeling based on multiple time series data acquired by operating the target device 3 under various conditions in selecting time series data to be used depending on the necessity of conditions that the model should take into account, etc.
[0089] Furthermore, in the information processing system according to this embodiment, the information processing device 1 displays the calculated matrix values of parameters A and B of the model in heat map format on the display unit 14. As a result, the information processing system according to this embodiment is expected to support the modeling work by the user, for example, when the matrix of parameters A and B has a large number of elements, since the user can grasp, based on the heat map, what values each component of the calculated parameters A and B has and their distribution.
[0090] Furthermore, in the information processing system according to this embodiment, the information processing device 1 displays a comparison between the actual measured value of time evolution data Y based on the observation data X acquired from the target device 3 and the predicted value of the time evolution data Y based on the calculated model of parameters A and B. This allows the user to easily grasp the degree of prediction accuracy of the model based on the calculated parameters A and B, and is therefore expected to assist the user in modeling work.
[0091] [Embodiment 2] In an information processing system according to embodiment 2, it is possible to vary model parameters A and B according to an external parameter θ, such as the ambient temperature when the target device 3 is operating. To this end, the information processing device 1 according to embodiment 2 sets the model parameters A and B as functions A(θ) and B(θ) of the parameter θ, and calculates coefficients of functions that determine these functions A(θ) and B(θ) based on the control input data U and observation data X when the target device 3 is operated, and the external parameter θ. That is, the model according to embodiment 2 is expressed by the following equation (12), and the loss function L is expressed by equation (13).
[0092] Y = A(θ)X + B(θ)U…(12)
[0093]
[0094] Furthermore, in the information processing system according to the second embodiment, the parameters A(θ) and B(θ) can be configured by combining a plurality of basis functions f and g and their corresponding coefficients A and B, as shown in the following equations (14) and (15). A user using the information processing system according to the second embodiment determines in advance the number of basis functions f and g and the definition of the functions of each basis function f and g, and inputs this information to the information processing device 1, for example, as setting information. The basis function f of the parameter A(θ) and the basis function g of the parameter B(θ) may differ in the number and type of functions, etc. The information processing device 1 stores this information input by the user in the storage unit 12.
[0095]
[0096] The parameter θ used in the information processing system according to the second embodiment is, for example, the air temperature when the target device 3 is performing processing, and the information processing device 1 acquires the temperature measured by a thermometer installed in or around the target device 3, and stores the temperature acquired from the thermometer as a measurement value of the parameter θ together with time-series control input data U and observation data X acquired in accordance with the operation of the target device 3. Note that instead of the information processing device 1 acquiring the measurement value of the parameter θ from a device such as a thermometer, for example, the user may input the temperature measured by the thermometer into the operation unit 15, and the information processing device 1 may acquire the value of the parameter θ by accepting the user's input.
[0097] Furthermore, in the second embodiment, the parameter θ is not time-series data, but is a parameter to which one value is assigned for a set of time-series control input data U and observation data X. In other words, the value assigned as the parameter θ is a constant value (or a value that does not fluctuate greatly) in a series of operations of the target device 3. However, multiple values may be assigned as the parameter θ, such as air temperature and water temperature, but each value does not change over time.
[0098] The information processing device 1 acquires time-series control input data U and observation data X and the value of a non-time-series parameter θ under one operating condition, and acquires this information under multiple operating conditions and stores and accumulates it in the time-series data storage unit 12b. The information processing device 1 identifies model parameters A and B as functions of θ (i.e., calculates the coefficients of parameters A(θ) and B(θ)) using multiple sets of control input data U, observation data X, and parameter θ stored in the time-series data storage unit 12b.
[0099] The information processing device 1 according to the second embodiment can calculate parameters A(θ) and B(θ) based on the Jacobian J and Hessian H of the loss function L in the same manner as the information processing device 1 according to the first embodiment. The Jacobian J and Hessian H of the loss function L expressed by the above-mentioned equation (13) are calculated by the following equations (16) to (21).
[0100]
[0101]
[0102] The information processing device 1 can calculate the coefficients of the parameters A(θ) and B(θ) using the following equation (22) based on the calculated Jacobian J and Hessian H. Note that the left side of equation (22) is a vector obtained by sequentially arranging the k coefficients A and m coefficients B shown on the right sides of equations (14) and (15).
[0103]
[0104] The information processing device 1 compares the components of the vector on the left side of equation (22) with the components of the vector on the right side of equation (22) calculated based on the Jacobian J and the Hessian H, and associates the left and right components in order, thereby obtaining the coefficients of the parameters A(θ) and B(θ). Similarly to the information processing device 1 according to embodiment 1, the information processing device 1 according to embodiment 2 calculates the condition number of the Hessian H expressed by equation (17). The information processing device 1 displays the calculated coefficients of the functions of the parameters A(θ) and B(θ) and the calculated number of conditions on the display unit 14, thereby providing the user with information about the model.
[0105] In the second embodiment, the parameter θ is exemplified by air temperature, water temperature, or the like, but the parameter θ is not limited to these. The parameter θ may be, for example, any physical quantity that can be measured inside or outside the target device 3, or may be, for example, any setting value that the user can set for the target device 3, or may be any other value.
[0106] The other configurations of the information processing system according to the second embodiment are the same as those of the information processing system according to the first embodiment, so the same parts are given the same reference numerals and detailed description thereof will be omitted.
[0107] An information processing system according to a third embodiment handles a heat conduction model of a wafer clamping mechanism (substrate clamping mechanism) assuming a wafer clamping mechanism (electrostatic chunk: ESC) of a substrate processing apparatus as the target device 3. In this case, the heat conduction model can be expressed by, for example, the following equation (23).
[0108]
[0109] In equation (23), the variable x is a vector having n temperature measurement values [x1, ..., xn] from n temperature sensors provided in the wafer fixing mechanism of the substrate processing apparatus. The variable u is a vector having k output values [u1, ..., uk] from k heaters provided in the substrate processing apparatus. Equation (23) corresponds to a heat conduction model expressed by using Y in equation (1) as the differential value of x.
[0110] In order to improve the accuracy of this heat conduction model, for example, p temperature sensors for measuring radiant heat from the inner wall of the chamber of the substrate processing apparatus can be added, and a term for radiant heat can be added to equation (23) to obtain the following equation (24). Note that equation (23) is based on the knowledge that radiant heat is proportional to the fourth power of temperature.
[0111]
[0112] In equation (24), the exponent of vector x included in the second term on the right-hand side indicates the exponent of each component included in vector x (the exponent of the Hadamard product). Vector x is a vector containing m temperature measurements [x, ..., x] from m temperature sensors provided in the wafer clamping mechanism and p temperature measurements [x, ..., x] from p temperature sensors measuring radiant heat from the chamber inner wall. Subscripts such as "0" (a x b) in matrices A and A indicate a zero matrix (a matrix with all components zero) of size a x b.
[0113] Furthermore, instead of equation (24), a heat conduction model according to the following equation (25) can be adopted: Equation (24) divides the vector x in equation (23) into a vector x having temperature measurements related to the wafer fixing mechanism and a vector x having temperature measurements related to the chamber inner wall.
[0114]
[0115] Furthermore, instead of equations (24) and (25), a heat conduction model according to the following equation (26) can be adopted: Equation (26) introduces, instead of vector x in equations (23) and (24), vector y having the temperature measurement values [x, ..., x] related to the wafer fixing mechanism and the fourth power of the temperature measurement values related to the chamber inner wall [x, ..., x].
[0116]
[0117] Furthermore, in order to improve the prediction accuracy of the heat conduction model using equation (26), it is possible to include more information in the temperature-related vector y. The following equation (27) is for the case where vector y has a constant term and a quadratic term in addition to the fourth power (quartic term) of the temperature measurement value related to the chamber inner wall.
[0118]
[0119] By using the above equations (23) to (26), the information processing device 1 is expected to create a model that takes nonlinear effects into account. For example, in a wafer temperature adjustment function model, it is expected to incorporate the temperature dependency of thermal conductivity. Also, for example, in a wafer temperature adjustment function model, it is expected to incorporate the contribution of thermal radiation. Also, for example, in a mechanical model including moving parts, it is expected to incorporate higher-order friction (such as air resistance proportional to the square of the velocity). Also, for example, in a fluid model, it is expected to efficiently incorporate the contribution of nonlinear terms in the Navier-Stokes equations. Note that the above models are merely examples and are not limiting, and the present technology can be applied to various models.
[0120] Furthermore, as shown in equations (23) to (26), Y, X, or U in equation (1) may be preprocessed observed quantities. The preprocessing employed may be time-directional preprocessing such as moving average, differentiation, integration, time difference, weighted sum with respect to time, or non-integer differentiation, or may be application of a product, sum, exponentiation, power root, trigonometric function, exponential function, constant multiplication, constant sum, or the like to values at each time (point in time). Furthermore, multiple preprocessing processes may be combined, or a vector combining the results of multiple preprocessing processes may be used.
[0121] For example, when a quantity observed at a certain time is a vector [x, y], it is possible to adopt [x, y, x², y², xy], [x + y, x-y], [x, y, 1 / x, 1 / y, 1, 0], or [sin(x)-1, sin(y), 1+exp(x) sin(y)], etc. Also, when a quantity observed at a certain time t is [x(t), y(t)], it is possible to adopt [x(t), y(t), x(t)-x(t-1), y(t)-y(t-1), y(t)², x(t) y(t)], or [x(t), y(t), x(1)+x(2)+...+x(t-1), y(1)+y(2)+...+y(t-1)], etc.
[0122] Furthermore, the information processing device 1 may treat observation results that have undergone different preprocessing as separate terms with different coefficients, rather than combining them, as shown in the above equations (24) and (25). For example, for the vector x of the observation result, the first term X, the second term X, ... for each component may be modeled as having coefficients A, A, .... In this case, if the model includes up to squared terms, the Hessian is calculated using the following equation (28). Note that, even when terms beyond the squared term are included or when other preprocessing is applied, the desired Hessian can be obtained by inserting rows and columns having preprocessed values into the Hessian.
[0123]
[0124] The type of preprocessing to be performed on the obtained observation results is determined in advance by the user. Alternatively, the information processing device 1 may compare and evaluate the results of applying multiple types of preprocessing, and adopt the preprocessing with the best evaluation result. Indices such as reconstruction error, calculation time, number of parameters, or memory usage during calculation may be used to evaluate the preprocessing. For example, an evaluation function such as "reconstruction error - a x number of model parameters - b x calculation time (a and b are constants determined by the user)" may be used. Furthermore, for example, the Akaike Information Criterion or the Bayes Information Criterion may be used as the evaluation function.
[0125] Other configurations of the information processing system according to the third embodiment are the same as those of the information processing systems according to the first and second embodiments, so the same reference numerals are used for the same parts and detailed description thereof will be omitted.
[0126] [Fourth Embodiment] In the above-described embodiment, the information processing device 1 determines the coefficient matrices A and B by optimizing the loss function L for the model expressed by equation (1) using the Hessian and Jacobian. However, instead of using the Hessian and Jacobian, the information processing device 1 may determine the coefficient matrices A and B by optimizing the arguments θA and θB of the functions A = fA (θA) and B = fB (θB) that generate the coefficient matrices A and B. Note that the Hessian and Jacobian in this case can be obtained by differentiating the loss function L.
[0127] Note that the functions fA and fB are functions predetermined by a user, etc. For example, fA can be a function that receives as input a matrix of the same size as the coefficient matrix A and raises each element of the matrix to the xth power (x is a real number). Furthermore, for example, fA may be implemented by combining common matrix operations such as the product of a matrix and a scalar, the product or sum of matrices, an inverse matrix, a pseudo-inverse matrix, an operation of substituting a predetermined value for a specific element of a matrix, an operation of replacing specific elements with their weighted average value, a magnitude comparison operation, an operation of taking the maximum or minimum value of specific elements, the Kronecker product of a matrix, or a matrix exponential function.
[0128] For example, θA may be a matrix of the same size as the coefficient matrix A, such as A = θA, or it may be specified by scalars a and b, such as A = a × θA2 + b × θA, or after setting A = θA, the (i, i) component of A may be changed to a value that makes the sum of column i zero (i = 1, 2, ... is the column number). The same applies to the coefficient matrix B.
[0129] Furthermore, in the case of some techniques such as Hankel dynamic mode decomposition, there may be parameters other than the coefficient matrices A and B, and parameter optimization may be performed for these parameters as well. For example, in the case of Hankel dynamic mode decomposition, there is an integer parameter with a value of 1 or more that specifies how much past data to refer to during prediction, and this parameter may be optimized within a range of positive integers.
[0130] Furthermore, parameter optimization may be performed by combining existing optimization methods such as Bayesian optimization. These methods are also referred to as dynamic mode decomposition here.
[0131] Furthermore, if coefficient matrices A and B contain components that are to be fixed to predetermined fixed values, the Hessian components corresponding to those components are set to "0." In this case, the Jacobian becomes the Jacobian of the loss function in coefficient matrices A and B, in which the components to be fixed are fixed to fixed values and the other components are set to "0." Furthermore, the Jacobian is replaced with the gradient when the desired components are fixed to fixed values and the other components are set to "0."
[0132] For example, a constraint for fixing a desired component of a coefficient matrix to a fixed value is expressed by the following equation (29). In equation (29), Aij and Bij represent the components of coefficient matrices A and B. In equation (29), A and B with a "-" above them represent matrices having values to be fixed as constraints. In equation (29), Pij indicates whether the corresponding components of A and B are fixed, and for example, Pij = 0 indicates a component with a constraint, and Pij = 1 indicates a component without a constraint.
[0133]
[0134] In order to determine the coefficient matrices A and B that satisfy this constraint, the above equations (4) to (6) are replaced with the following equations (30) to (32). The information processing device 1 according to the fourth embodiment can determine the coefficient matrices A and B based on these equations (30) to (32).
[0135]
[0136] By allowing the components of coefficient matrices A and B to be fixed, it is possible to generate a new model by incorporating user knowledge, theoretical formulas, previous modeling results, etc. For example, the coefficients of degrees of freedom that are known to have no correlation can be set to a fixed value of "0," such as when the heat transfer coefficient between two thermally insulated degrees of freedom in a temperature model should be 0. Furthermore, for example, if the values of specific components of coefficient matrices A and B are known in advance through other experiments, setting those values as fixed values for the corresponding components of coefficient matrices A and B eliminates the need for calculations to calculate those components.
[0137] Furthermore, when the coefficient matrix A is a Toeplitz matrix, it can be generated from a d-dimensional vector a=[a, a, ..., ad]. In this case, the Hessian and Jacobian for the coefficient matrix A can be replaced with the Hessian and Jacobian for the d-dimensional vector a.
[0138] Furthermore, the information processing device 1 can determine the coefficient matrix A by optimizing the difference θ = A - A ref from a known coefficient matrix A ref modeled under slightly different conditions. For example, for multiple target devices 3 with machine differences, the coefficient matrix A can be determined by optimizing θ of the function A = f(θ) + A ref using observation data based on the reference coefficient matrix A ref of the model obtained from a certain reference device. This is expected to simplify modeling, as only the dynamics that originate from machine differences are to be modeled. Note that when there are multiple reference devices, the reference coefficient matrix may be an average or weighted average of these. The reference coefficient matrix may be calculated theoretically, calculated by simulation, or determined by a combination of these methods.
[0139] The information processing device 1 can also determine coefficient matrices for differences in experimental conditions, such as temperature during operation, in a manner similar to that for machine differences. The information processing device 1 stores reference coefficient matrices for each of a plurality of predetermined reference experimental conditions α, β, γ, etc., and, for example, for an experimental condition near the reference experimental condition α, the information processing device 1 can determine the coefficient matrix for this experimental condition using the reference coefficient matrix corresponding to α. For example, in a temperature model for a vacuum chamber, the information processing device 1 stores a reference coefficient matrix for high temperature and normal pressure (100°C, 1 atmosphere), a reference coefficient matrix for high temperature and low pressure (50°C, 0.1 atmosphere), and a reference coefficient matrix for low temperature and low pressure (0°C, 0.1 atmosphere) as conditions for combinations of operating temperature and pressure ranges. If a temperature model at 97°C and 500 hPa is required, a linearly interpolated matrix for pressure between the reference coefficient matrix for high temperature and normal pressure and the reference coefficient matrix for high temperature and low pressure can be used as the reference coefficient matrix, and the information processing device 1 can generate a temperature model by optimizing the calculation of the difference from this reference coefficient matrix. The information processing device 1 may use the nearest reference coefficient matrix instead of interpolation, or may use a reference coefficient matrix stored based on reference coefficient matrices under three or more experimental conditions. Furthermore, for the interpolation, other methods such as spline interpolation may be used instead of linear interpolation.
[0140] Furthermore, the information processing device 1 can align the values of components that should have the same value in a coefficient matrix by using a function that describes components that should have the same value in a coefficient matrix with one parameter. The information processing device 1 can also set a constraint condition, such as that the sum or product of a certain component and another component in a coefficient matrix should be equal to another component.
[0141] More generally, when a constraint described by an arbitrary equation g(A) = 0 relating to a coefficient matrix or its components is known in advance, the information processing device 1 generates a coefficient matrix from a low-dimensional parameter θ using a function A = fA(θ) that satisfies this constraint. The function A = fA(θ) is designed in advance by the user, for example, as a function fA such that g(fA(θ)) = 0 for all θ. For example, in a bilaterally symmetrical temperature control mechanism, if corresponding left and right parts should have the same heat capacity, fA can be a function that copies the heat capacity θ of each part on the right side to create the overall heat capacity and returns it in the form of a matrix. Furthermore, for example, if it is known that coefficient matrix A is an antisymmetric matrix, fA can be a function that constructs coefficient matrix A by inverting the sign of the upper half of the upper triangular matrix excluding the diagonal elements of A and copying it to the lower half. This allows the information processing device 1 to reduce the number of coefficients to be optimized when optimizing a coefficient matrix, which is expected to improve the efficiency and accuracy of optimization calculations.
[0142] The other configurations of the information processing system according to the fourth embodiment are the same as those of the information processing systems according to the first to third embodiments, so the same reference numerals are used for the same parts and detailed description thereof will be omitted.
[0143] Fifth Embodiment An information processing device 1 according to a fifth embodiment accepts from a user settings of ranks rA and rB for coefficient matrices A and B determined by optimization. The information processing device 1 determines coefficient matrices A and B so as to satisfy the set ranks. For example, the information processing device 1 can determine a coefficient matrix when a rank is specified by using a formula for the inverse matrix of a 2×2 block matrix and generating a generalized inverse matrix using singular value decomposition of the specified rank of each block.
[0144] Furthermore, when there are multiple types of observables X or inputs U and they have different coefficient matrices, the information processing device 1 can accept inputs of ranks equal to the number N of coefficient matrices and determine the coefficient matrix in the same manner using the inverse matrix formula for a general N×N block matrix. However, the inverse matrix formula for an N×N block matrix can be obtained by recursively applying the formula for the 2×2 case.
[0145] As a result, in the information processing system according to the fifth embodiment, for example, when generating a temperature model with 50 degrees of freedom, the user can set the rank to 10, which allows the information processing device 1 to generate a low-dimensional model using only the 10 modes with the highest contributions. By generating a low-dimensional model from which modes with low contributions have been deleted, improvements in the interpretability of the model and the processing speed can be expected. Furthermore, for example, by lowering the rank for data with high noise, it can be expected that robust modeling can be achieved.
[0146] The embodiments disclosed herein are to be considered as illustrative in all respects and not restrictive. The scope of the present disclosure is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims.
[0147] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any and all combinations, regardless of the reference format. Furthermore, the claims use a format in which a claim references two or more other claims (multiple claim format), but this is not limited to this. A multiple claim (multi-multi claim) that references at least one other multiple claim may also be used.
[0148] REFERENCE SIGNS LIST 1 Information processing device (computer) 3 Target device 11 Processing unit 11a Control processing unit 11b Data acquisition unit 11c Data selection unit 11d Parameter calculation unit 11e Display processing unit 12 Storage unit 12a Program (computer program) 12b Time-series data storage unit 13 Communication unit 14 Display unit 15 Operation unit
Claims
1. A computer program causing a computer to execute a process of acquiring multiple sets of time-series control input data for a target device obtained by operating the target device under multiple conditions and time-series observation data observing the state of the target device, and calculating model parameters that predict the time evolution data of the observation data in accordance with the control input data and the observation data by dynamic mode decomposition based on the acquired multiple sets of time-series control input data and the observation data.
2. The computer program according to claim 1, wherein the model is a model in which the dependence from the input of the control input data and the observation data to the output of the time evolution data is linear.
3. The computer program according to claim 2, further comprising: calculating a matrix representing a first-order differential and a matrix representing a second-order differential for a quadratic function defined as a loss function of the model; and calculating the parameters based on the calculated matrices.
4. The computer program of claim 1, further comprising the step of: calculating a condition number for the calculated parameters.
5. The computer program of claim 1, which displays a list of information relating to the acquired sets of control input data and observation data, accepts selection of one or more sets from the displayed sets, and calculates parameters of the model based on the selected sets of control input data and observation data.
6. The computer program according to claim 5, further comprising: a selection of a set of the control input data and the observation data to be used for calculating the parameters of the model from among the displayed sets; and a selection of a set of the control input data and the observation data to be used for evaluating the model using the calculated parameters.
7. The computer program according to claim 1, wherein the values of the matrix calculated as the parameters of the model are displayed in a heat map.
8. The computer program according to claim 1, which displays in comparison an actual measured value of the time evolution data based on the acquired observation data and a predicted value of the time evolution data based on a calculated parameter model.
9. The computer program according to claim 1, further comprising: acquiring multiple sets of non-time series data obtained when the target device is operated together with the corresponding time series control input data and observation data; and calculating coefficients of a parameter function for calculating the parameters that change in response to the non-time series data by dynamic mode decomposition based on the acquired multiple sets of the control input data, the observation data and the non-time series data.
10. The computer program according to claim 1, wherein the target device is a substrate processing apparatus, and a set of control input data for the substrate processing apparatus and observation data for observing a state of the substrate processing apparatus is acquired.
11. The computer program according to claim 10, wherein the target device is a substrate fixing mechanism of the substrate processing apparatus, and the model is a heat conduction model of the substrate fixing mechanism including a term relating to radiant heat from an inner wall of a chamber of the substrate processing apparatus.
12. The computer program according to claim 1, wherein at least one of the control input data and the observation data is multi-dimensional data having two or more dimensions.
13. The computer program according to claim 1, which calculates parameters of the model based on data obtained by applying a predetermined arithmetic processing to the acquired observation data.
14. The computer program of claim 1, wherein the model is a model including power terms for components of the observed data.
15. The computer program according to claim 1, wherein some of the parameters are set to predetermined values.
16. The computer program according to claim 1, further comprising: calculating parameters of a function that generates parameters of the model to calculate the parameters of the model.
17. The computer program according to claim 1, further comprising: calculating a difference between parameters of a known model and parameters of a target model, thereby calculating parameters of the target model.
18. The computer program according to claim 17, further comprising: calculating a difference between a parameter of a model for a reference device and a parameter of a model for a target device having a machine difference with respect to the reference device, thereby calculating a parameter of the model for the target device.
19. The computer program according to claim 1, further comprising: receiving a designation of a rank for the parameters; and calculating model parameters according to the received rank.
20. The computer program according to claim 19, which calculates model parameters according to the accepted rank by generating a generalized inverse matrix using singular value decomposition of the specified rank of each block based on the formula for the inverse matrix of a 2x2 block matrix.
21. An information processing method in which an information processing device acquires multiple sets of time-series control input data for a target device obtained by operating the target device under multiple conditions and time-series observation data observing the state of the target device, and calculates model parameters that predict the time evolution data of the observation data in accordance with the control input data and the observation data by dynamic mode decomposition based on the acquired multiple sets of time-series control input data and the observation data.
22. An information processing device comprising a processing unit, which acquires a plurality of sets of time-series control input data for a target device obtained by operating the target device under a plurality of conditions and time-series observation data observing a state of the target device, and calculates parameters of a model that predicts the time evolution data of the observation data in accordance with the control input data and the observation data by dynamic mode decomposition based on the acquired plurality of sets of time-series control input data and the observation data.
Citation Information
Patent Citations
Learning method, sequence analysis method, learning device, sequence analysis device, and program
JP2022070386A
Information processing device, analysis method, control system, and analysis program
JP2023098423A
Performance predictors for semiconductor-manufacturing processes
US20230049157A1