Computer program, information processing method, and information processing device

JPWO2024158019A5Pending Publication Date: 2025-10-09
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024573217
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-07-22
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Conventional machine learning models used in virtual measurement technologies for substrate processing do not account for spatial correlation, leading to inaccurate and difficult-to-interpret predictions, as they treat adjacent locations independently, resulting in distorted prediction results and unclear parameter effectiveness.

Method used

The introduction of a prediction model that incorporates spatial correlation through dimensional mapping, using a feature extraction model to convert feature quantities into target dimensions, allowing for improved accuracy and interpretability by considering the physical dimensions of substrate processing data.

Benefits of technology

This approach significantly enhances prediction accuracy and interpretability by reflecting the actual spatial distribution of substrate processing data, reducing mean square error and enabling better understanding of parameter contributions across locations.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided are a computer program, an information processing method, and an information processing device. The present invention causes a computer to execute a process for: acquiring data relating to substrate processing; extracting feature amounts of the acquired data by using a first trained model that has been trained to output the feature amounts of the data in response to input of the data; converting the extracted feature amounts into feature amounts having a set target dimension; and in response to input of the feature amounts having the target dimension, inputting, to a second trained model that has been trained to output a prediction value relating to the substrate processing, the feature amounts after dimension conversion, and thereby obtaining the prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Computer program, information processing method, and information processing device

[0001] The present invention relates to a computer program, an information processing method, and an information processing device.

[0002] Virtual metrology has been increasingly used in the field of substrate processing. Virtual metrology analyzes measurement data obtained during processing of an object such as a substrate and calculates predicted values ​​for the resulting object.

[0003] Special table 2019-537240 publication

[0004] The present disclosure provides a computer program, an information processing method, and an information processing device that are capable of performing analysis that takes spatial correlation into account using a learning model.

[0005] A computer program according to one embodiment of the present invention is a computer program for causing a computer to execute a process of acquiring data related to substrate processing, extracting features of the acquired data using a first learning model trained to output features of the data in response to input of the data, converting the extracted features into features of a set target dimension, and inputting the dimension-converted features into a second learning model trained to output a predicted value related to substrate processing in response to input of features having the target dimension to obtain a predicted value.

[0006] According to the present disclosure, analysis that takes spatial correlation into account can be performed using a learning model.

[0007] FIG. 1 is an explanatory diagram illustrating a configuration of an information processing system according to an embodiment. FIG. 2 is an explanatory diagram illustrating a prediction method according to embodiment 1. FIG. 3 is a block diagram illustrating the internal configuration of an information processing device. FIG. 4 is a flowchart illustrating a procedure for generating a prediction model. FIG. 5 is a flowchart illustrating a prediction procedure using a prediction model. FIG. 6 is an explanatory diagram for illustrating performance evaluation of a prediction model. FIG. 7 is a graph illustrating a spatial distribution of importance for each piece of observation data. FIG. 8 is a flowchart illustrating a processing procedure executed by an information processing device according to embodiment 2. FIG. 9 is an explanatory diagram illustrating a prediction method according to embodiment 3. FIG. 10 is a flowchart illustrating a processing procedure executed by an information processing device according to embodiment 4. FIG. 11 is a flowchart illustrating a processing procedure executed by an information processing device according to embodiment 5.

[0008] An embodiment will be described below with reference to the drawings. In the description, the same elements or elements having the same functions are designated by the same reference numerals, and redundant description will be omitted.

[0009] 1 is an explanatory diagram illustrating the configuration of an information processing system according to an embodiment. The information processing system according to the embodiment includes an information processing apparatus 100 and a substrate processing apparatus 200 that are communicatively connected.

[0010] The substrate processing apparatus 200 is, for example, a semiconductor manufacturing apparatus including at least one of an exposure apparatus, an etching apparatus, a film forming apparatus, an ion implantation apparatus, an ashing apparatus, a sputtering apparatus, etc. Alternatively, the substrate processing apparatus 200 may be a display manufacturing apparatus that manufactures flat display panels (FDPs) such as liquid crystal display panels and organic electroluminescence (EL) panels.

[0011] When a process is initiated in the substrate processing apparatus 200, various set values ​​are set, such as the substrate temperature, the pressure and gas flow rate in the chamber, and the voltage applied from the high-frequency power supply. The set values ​​are provided, for example, by a process recipe. The substrate processing apparatus 200 is also provided with various sensors and devices for measuring the substrate temperature, the pressure and gas flow rate in the chamber, the voltage applied to the upper electrode and the lower electrode, the plasma emission intensity, and the like, and various measurement values ​​are obtained during the process. In addition to the above-mentioned measurement values, the substrate processing apparatus 200 also collects appropriate time-series data, such as images (RGB data) of the substrate (wafer) before and after the process and process logs, as needed. The substrate processing apparatus 200 outputs the measurement values, images, time-series data, and the like obtained during the process to the information processing apparatus 100 as observation data.

[0012] The information processing apparatus 100 acquires observation data as data related to the substrate processing from the substrate processing apparatus 200. The information processing apparatus 100 calculates predicted values ​​related to the substrate processing based on the acquired observation data.

[0013] Virtual measurement using observation data has been performed for some time. For example, conventionally, some input signal such as a sensor measurement value, image data, or time-series data is input to a machine learning model that corresponds to the input signal, and the machine learning model is operated to obtain the required predicted value.

[0014] However, conventional machine learning models lack accuracy and interpretability because they are not designed to take spatial correlation into account. For example, if spatial correlation is not taken into account, independent predictions are made for each location, which can result in large differences in predicted values ​​even for neighboring locations, potentially resulting in spatially distorted prediction results. Furthermore, if spatial correlation is not taken into account, it is difficult to determine which parameters are most effective in which locations.

[0015] Therefore, in this embodiment, a model that introduces dimensional mapping is proposed as a prediction model MD2 that takes spatial correlation into account. Dimension mapping refers to converting the dimensions of features (variables that serve as clues for prediction) extracted from observed data to match the physical dimensions (target dimensions) for which a predicted value is to be calculated. For example, a machine learning learning model (hereinafter referred to as feature extraction model MD1) is used to extract the features. In this embodiment, by introducing dimensional mapping into a unimodal network structure, spatial correlation is explicitly taken into account, thereby improving accuracy and interpretability.

[0016] 2 is an explanatory diagram illustrating the prediction method according to embodiment 1. The information processing apparatus 100 acquires data related to substrate processing from the substrate processing apparatus 200. The data acquired by the information processing apparatus 100 is arbitrary, and may be observation data including measurement data output from sensors or the like of the substrate processing apparatus 200, image data obtained by capturing an image of the substrate to be processed, and time-series data such as a process log.

[0017] The information processing apparatus 100 receives observation data as input and uses a feature extraction model MD1 (first learning model) that has been trained to output feature quantities of the observation data to extract feature quantities of the observation data acquired from the substrate processing apparatus 200. The feature quantities to be extracted are preferably variables that provide clues for prediction.

[0018] A machine learning learning model including deep learning can be used as the feature extraction model MD1. For example, a learning model based on a convolutional neural network (CNN), a transformer, a recurrent neural network (RNN), a long short-term memory (LSTM), a multi-layer perceptron (MLP), or the like can be used. Alternatively, a learning model other than deep learning, such as an autoregressive model, a moving average model, or an autoregressive moving average model, may be used. The learning model used in the feature extraction model MD1 is set appropriately depending on the input observation data and the features to be extracted.

[0019] The feature extraction model MD1 includes, for example, an input layer, one or more intermediate layers, and an output layer, and is trained to output features from the output layer in response to observation data input to the input layer. Alternatively, a value output from one of the intermediate layers may be used as a feature. The feature extraction model MD1 may be configured to include only an input layer and an output layer, without including an intermediate layer. In this embodiment, the feature output from the feature extraction model MD1 is described as being one-dimensional, but the feature may be two or more-dimensional.

[0020] Next, the information processing device 100 converts (dimension mapping) the dimension of the extracted feature quantity to match the target dimension (physical dimension to be calculated as a predicted value). When it is desired to calculate the etching rate, etching shape (opening width or opening depth), film thickness, etc. at each location on the substrate surface as a predicted value, the dimension of the extracted feature quantity can be converted to two dimensions. The example in FIG. 2 shows dimension mapping from one-dimensional feature quantity to two-dimensional feature quantity. Any dimension can be used before and after conversion, and is set appropriately depending on the observation data used and the predicted value to be calculated. The target dimension may be expanded or reduced, or may be equal to the dimension of the feature quantity before conversion. N feature quantities (N=N x ×N y ), each element is N x ×N y By rearranging (mapping) the one-dimensional feature quantity into a two-dimensional feature quantity, the one-dimensional feature quantity can be converted into a two-dimensional feature quantity.

[0021] The information processing device 100 receives the dimension-mapped feature as an input and uses a prediction model MD2 (second learning model) that has been trained to output a predicted value related to the substrate processing to obtain a predicted value related to the substrate processing.

[0022] A machine learning learning model including deep learning can be used as the prediction model MD2. For example, a learning model based on CNN, Transformer, RNN, LSTM, MLP, etc. can be used. Alternatively, a learning model other than deep learning, such as an autoregressive model, a moving average model, or an autoregressive moving average model, may be used. The learning model used for the prediction model MD2 is set appropriately depending on the target dimension of the input feature and the predicted value to be calculated.

[0023] In this embodiment, for convenience of explanation, the dimension mapping is described as an independent process, but it may be a process executed within the prediction model MD2. For this reason, the prediction model MD2 is also referred to as a dimension mapping model.

[0024] Although the feature extraction model MD1 and the prediction model MD2 are described as independent learning models in this embodiment for convenience, they may be constructed as a single learning model. In this case, feature extraction, dimension mapping, and calculation of prediction values ​​are performed within the single learning model.

[0025] 3 is a block diagram showing the internal configuration of the information processing device 100. The information processing device 100 is, for example, a dedicated or general-purpose computer including a control unit 101, a storage unit 102, a communication unit 103, an operation unit 104, and a display unit 105.

[0026] The control unit 101 includes a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The ROM included in the control unit 101 stores control programs and the like that control the operation of each hardware unit included in the information processing device 100. The CPU in the control unit 101 reads and executes the control programs stored in the ROM and computer programs (described below) stored in the storage unit 102, and controls the operation of each hardware unit, thereby causing the entire device to function as the information processing device of the present disclosure. The RAM included in the control unit 101 temporarily stores data used during execution of calculations.

[0027] In the embodiment, the control unit 101 is configured to include a CPU, a ROM, and a RAM, but the configuration of the control unit 101 is not limited to the above. The control unit 101 may be, for example, one or more control circuits or arithmetic circuits including a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), a quantum processor, volatile or non-volatile memory, etc. The control unit 101 may also have functions such as a clock that outputs date and time information, a timer that measures the elapsed time from when a measurement start instruction is given until when a measurement end instruction is given, and a counter that counts numbers.

[0028] The storage unit 102 includes a storage device such as a hard disk drive (HDD), a solid state drive (SSD), an electronically erasable programmable read-only memory (EEPROM), etc. The storage unit 102 stores various computer programs executed by the control unit 101 and various data used by the control unit 101.

[0029] The computer program (program product) stored in the storage unit 102 includes a prediction processing program PG1 for causing a computer to execute a process for obtaining a predicted value related to substrate processing from observation data of the substrate processing apparatus 200. The prediction processing program PG1 may be a single computer program or a program group consisting of multiple computer programs. The prediction processing program PG1 may be executed by multiple computers in cooperation with each other. Furthermore, the prediction processing program PG1 may partially use an existing library.

[0030] A computer program including the prediction processing program PG1 is provided by a non-transitory recording medium RM on which the computer program is readably recorded. The recording medium RM is a portable memory such as a CD-ROM, USB memory, a Secure Digital (SD) card, a microSD card, or a CompactFlash (registered trademark). The control unit 101 reads various computer programs from the recording medium RM using a reading device (not shown) and stores the read computer programs in the storage unit 102. The computer programs stored in the storage unit 102 may also be provided via communication. In this case, the control unit 101 acquires the computer programs via communication via the communication unit 103 and stores the acquired computer programs in the storage unit 102.

[0031] The storage unit 102 also stores a feature extraction model MD1 used in a process of extracting features from observed data and a prediction model MD2 used in a process of calculating predicted values ​​related to substrate processing from features converted into target dimensions. Alternatively, the feature extraction model MD1 and the prediction model MD2 may be stored in an external device. In this case, the control unit 101 of the information processing device 100 may access the external device via a communication network, transmit the observed data acquired from the substrate processing device 200 to the external device, and acquire the predicted values ​​obtained as a result of calculations performed by the external device via the communication network.

[0032] The communication unit 103 includes a communication interface for transmitting and receiving various data to and from an external device. A communication interface conforming to a communication standard such as a local area network (LAN) can be used as the communication interface of the communication unit 103. The external device may be the substrate processing apparatus 200 described above or a user terminal (not shown). When data to be transmitted is input from the control unit 101, the communication unit 103 transmits the data to the external device as the destination, and when data transmitted from the external device is received, the communication unit 103 outputs the received data to the control unit 101.

[0033] The operation unit 104 includes operation devices such as a touch panel, a keyboard, and switches, and receives various operations and settings from a user, etc. The control unit 101 performs appropriate control based on various pieces of operation information provided by the operation unit 104, and stores setting information in the storage unit 102 as necessary.

[0034] The display unit 105 includes a display device such as a liquid crystal monitor or an organic EL (Electro-Luminescence) monitor, and displays information to be notified to the user or the like in response to an instruction from the control unit 101 .

[0035] The information processing apparatus 100 in this embodiment may be a single computer, or may be a computer system configured with multiple computers and peripheral devices. The information processing apparatus 100 may be a virtual machine whose entity is virtualized, or may be a cloud. Furthermore, although the information processing apparatus 100 and the substrate processing apparatus 200 are described as separate entities in this embodiment, the information processing apparatus 100 may be provided inside the substrate processing apparatus 200.

[0036] The following describes the operation of the information processing apparatus 100. The information processing apparatus 100 according to this embodiment generates a prediction model MD2 in a learning phase before the substrate processing apparatus 200 starts to be put into actual use.

[0037] FIG. 4 is a flowchart showing the procedure for generating the prediction model MD2. Prior to generating the prediction model MD2, training data necessary for learning is collected. For example, when the etching shape at each location on the substrate surface is predicted based on the plasma emission intensity, measurement data of the plasma emission intensity measured by an OES (Optical Emission Spectrometer) and measurement data of the etching shape at each location measured using an optical observation device, an ultrasonic microscope, or the like are collected as training data. The training data is not limited to the measurement data of the plasma emission intensity and the etching shape, but also includes observation data of values ​​used for prediction and actual measured values ​​of the values ​​to be predicted. The collected training data is stored in the storage unit 102 of the information processing device 100. It is assumed that the feature extraction model MD1 has been generated in advance using a known algorithm.

[0038] The control unit 101 reads out training data stored in the storage unit 102 (step S101), and selects a set of training data from the read training data (step S102). The control unit 101 inputs observation data (values ​​used for prediction) included in the selected training data into the feature extraction model MD1, and executes calculations using the feature extraction model MD1 to extract features from the observation data (step S103).

[0039] The control unit 101 converts the dimension of the feature extracted from the observation data into a target dimension (step S104). That is, the control unit 101 performs dimension mapping on the dimension of the extracted feature to match the physical dimension for which a predicted value is to be calculated.

[0040] The control unit 101 inputs the feature quantities converted into the target dimensions into the prediction model MD2 and performs calculations using the prediction model MD2 to obtain predicted values ​​for each location (step S105). It is assumed that initial values ​​are set for the model parameters of the prediction model MD2 before learning begins. Furthermore, although the present flowchart describes the dimension mapping process and the calculation process using the prediction model MD2 as independent processes, the dimension mapping may also be performed within the processing of the prediction model MD2.

[0041] The control unit 101 evaluates the predicted value calculated in step S105 (step S106) and determines whether learning is complete (step S107). A known loss function is used to evaluate the predicted value. If the value of the loss function becomes less than a threshold in the process of optimizing (minimizing) the loss function, the control unit 101 can determine that learning of the prediction model MD2 is complete.

[0042] If it is determined that learning is not complete (S107: NO), the control unit 101 updates the model parameters (weighting coefficients and biases between nodes) in the prediction model MD2 (step S108) and returns the process to step S102.

[0043] If it is determined that learning is complete (S107: YES), a trained model is obtained, and the control unit 101 stores the model in the memory unit 102 as a trained prediction model MD2 (step S109).

[0044] The information processing apparatus 100 performs prediction using the prediction model MD2 in the operation phase after the prediction model MD2 is generated. Fig. 5 is a flowchart showing a prediction procedure using the prediction model MD2. The control unit 101 of the information processing apparatus 100 acquires observation data to be used for prediction from the substrate processing apparatus 200, for example, via the communication unit 103 (step S121).

[0045] The control unit 101 inputs the acquired observation data into the feature extraction model MD1 and executes calculations using the feature extraction model MD1 to extract features from the observation data (step S122).

[0046] The control unit 101 converts the dimension of the feature extracted from the observation data into a target dimension (step S123). That is, the control unit 101 performs dimension mapping on the dimension of the extracted feature to match the physical dimension for which a predicted value is to be calculated.

[0047] The control unit 101 inputs the feature amounts converted into the target dimensions into the prediction model MD2, and performs calculations using the prediction model MD2 to obtain a predicted value for each location (step S124).

[0048] The control unit 101 outputs the prediction result based on the prediction model MD2 (step S125). The control unit 101 may display the prediction result on the display unit 105, or may notify the prediction result to a user terminal or the like via the communication unit 103.

[0049] FIG. 6 is an explanatory diagram for explaining the performance evaluation of the prediction model MD2. Each graph in FIG. 6 shows the in-plane distribution of the etching shape (opening width) when virtually or actually measured. The horizontal axis of each graph corresponds to a first direction in the substrate plane, and the horizontal axis corresponds to a second direction of the substrate perpendicular to the first direction. The shading shown in each graph corresponds to the width of the opening width, with lighter areas indicating wider opening widths and darker areas indicating narrower opening widths. FIG. 6A shows the prediction results (virtual measurement) using a conventional method, FIG. 6B shows the prediction results (virtual measurement) using the method of the present disclosure, and FIG. 6C shows the actual measured values.

[0050] In the actual measurement, a large number of openings were formed on the substrate surface by etching, and the opening width of each opening was measured using measurement devices such as an optical observation device and an ultrasonic microscope. In the virtual measurement, the substrate surface with the same openings formed was imaged with a camera, and the obtained image was used as observation data to predict the opening width. The image was a three-color image of RGB captured by a wafer optical inspection system.

[0051] The design value of the opening width was set to be constant regardless of the location where the opening was formed, but when the opening width of the openings actually formed in the substrate was measured, it was confirmed that the opening width had an in-plane distribution such that the opening width was widest near the center of the substrate surface and narrowed toward the periphery, as shown in Figure 6C.

[0052] On the other hand, when the aperture width was predicted using a conventional method (linear regression in this example), as shown in FIG. 6A , although there was a tendency for the aperture width to be widest near the center of the substrate surface and gradually narrower toward the periphery, the region where the aperture width was the same expanded horizontally on the graph, resulting in a distorted prediction result.

[0053] In contrast, when the opening width was predicted using the method of the present disclosure (prediction model MD2), the predicted results were not distorted in a specific direction, and a uniform distribution in the circumferential direction similar to the actual measurement was obtained, as shown in Figure 6B. The mean square error between the predicted value and the actual measurement value using the conventional method was about 0.8, while the mean square error between the predicted value and the actual measurement value using the method of the present disclosure was about 0.6, indicating a significant improvement in prediction accuracy.

[0054] FIG. 6 shows the prediction results using captured images as observation data, but when the plasma emission intensity and process log were used as observation data to predict the aperture width, it was found that the method disclosed herein improved prediction accuracy compared to conventional methods.

[0055] As described above, in the first embodiment, a method for introducing spatial correlation into a machine learning learning model using dimensional mapping and performing virtual measurement using the learning model (prediction model MD2) has been disclosed. Using spatial correlation makes it easier to interpret the model and allows the actual spatial distribution to be reflected in the prediction. Furthermore, it has been found that the prediction accuracy is significantly improved compared to conventional methods that do not take spatial correlation into account.

[0056] Second Embodiment In a second embodiment, a configuration will be described in which the importance (also referred to as the contribution) of a feature amount is calculated for each location, and a spatial distribution of the calculated importance is output.

[0057] The information processing device 100 according to the second embodiment calculates the importance (contribution) of a feature for each location using a prediction model MD2. The importance calculation uses known methods such as Lime (Local Interpretable Model-Agnostic Explanations), SHAP (Shapely Additive exPlanations), and CAM (Class Activation Mapping). Lime and SHAP are methods that identify how much an output changes when an input is reduced, and determine that the greater the change in output, the higher the importance. CAM is a method that calculates importance using error backpropagation during learning.

[0058] FIG. 7 is a graph showing the spatial distribution of importance for each observation data. FIG. 7A shows the spatial distribution of importance when using plasma emission intensity (OES), FIG. 7B shows a captured image (wafer optical inspection system), and FIG. 7C shows process logs (P-logs) as the observation data. The horizontal axis of each graph corresponds to a first direction in the substrate plane, and the horizontal axis corresponds to a second direction of the substrate perpendicular to the first direction. The shading shown in each graph corresponds to the level of importance. Areas with high shading on the graph indicate areas with high importance, and areas with low shading indicate areas with low importance.

[0059] When the aperture width was predicted using the plasma emission intensity as the observation data, the spatial distribution showed that the importance of the feature based on the plasma emission intensity decreased toward the center of the substrate and increased toward the edge of the substrate (Figure 7A). This graph shows that when the plasma emission intensity was used as the observation data, the aperture width at the edge of the substrate could be predicted well. Similar results were obtained when the process log was used as the observation data (Figure 7C).

[0060] On the other hand, when the aperture width was predicted using the image captured by the wafer optical inspection system as the observation data, the importance of the feature based on the captured image was low in some regions of the substrate periphery (the regions corresponding to the upper right and lower left corners of the graph) and high in other regions, resulting in a spatial distribution (Figure 7B). This graph shows that when the captured image is used, the aperture width can be predicted well in regions other than the substrate periphery.

[0061] As described above, since the spatial distribution of importance differs depending on the type of observation data (feature), when generating the prediction model MD2, training may be performed using a loss function with weights adjusted for each location. For example, when plasma emission intensity or a process log is used as observation data, a prediction model MD2 specialized for the peripheral area may be generated by training using a loss function with a higher weight for the peripheral area. Furthermore, when images captured by a wafer optical inspection system are used as observation data, a prediction model MD2 specialized for the central area may be generated by training using a loss function with a higher weight for the central area.

[0062] Furthermore, in this embodiment, since the contribution of the feature value can be confirmed for each location, it is possible to grasp, for example, which part of the substrate the sensor output value present in the process log contributes to, and by adjusting the process so that the sensor output value changes, it is possible to improve the process. Furthermore, in actual substrate processing, if there are circumstances such as poor yield due to poor process conditions in the peripheral area, a prediction model MD2 specialized for the peripheral area may be created by the above-mentioned method, and the process may be improved by taking into account the prediction results of the prediction model MD2.

[0063] 8 is a flowchart showing the procedure of processing executed by the information processing apparatus 100 according to embodiment 2. The control unit 101 of the information processing apparatus 100 acquires observation data to be used for prediction from the substrate processing apparatus 200, for example, via the communication unit 103 (step S201).

[0064] The control unit 101 calculates a predicted value for each location based on the acquired observation data (step S202). The method of calculating the predicted value is the same as in embodiment 1. That is, the control unit 101 inputs the acquired observation data to a feature extraction model MD1 to extract features, and performs dimension mapping of the extracted features to a target dimension (a physical dimension for which a predicted value is to be calculated). Next, the control unit 101 inputs the dimension-mapped features to a prediction model MD2 and performs calculations to calculate a predicted value for each location.

[0065] The control unit 101 calculates, for each location, the contribution of the observation data to the calculated predicted value (step S203). The contribution is a SHAP value that can be calculated using, for example, the prediction model MD2. The SHAP value is a value corresponding to the difference between a predicted value calculated by inputting multiple pieces of observation data into the prediction model MD2 and a predicted value calculated by the prediction model MD2 when one piece of observation data among the multiple pieces of observation data is not present. The contribution is not limited to the SHAP value, and can be calculated using existing methods such as Lime or CAM.

[0066] The control unit 101 outputs the spatial distribution of the contribution degree (step S204). Based on the contribution degree for each location calculated in step S203, the control unit 101 creates a graph (color contour map) such as those shown in Figures 7A to 7C, and displays it on the display unit 105. The control unit 101 may also transmit the created graph to the user terminal.

[0067] The control unit 101 executes control according to the contribution of each location (step S205). The control unit 101 adjusts the parameters for the control target according to the contribution of each location, and controls the process according to the adjusted parameters. For example, if it is found that the plasma emission intensity of a specific frequency contributes more to the peripheral portion, the gas flow rate can be adjusted to increase the emission intensity, thereby enabling process control to improve in-plane uniformity. The amount of parameter adjustment according to the contribution is determined, for example, on a rule-based basis.

[0068] In the flowchart of FIG. 8, the spatial distribution of the contribution degree is output in step S204, and then control according to the contribution degree is executed in step S205. However, these steps may be reversed, or only one of the steps may be executed.

[0069] As described above, in the second embodiment, the importance (contribution) of a feature is calculated for each location, and the spatial distribution of the calculated importance is output. This makes it possible to understand which parameter is likely to have an effect on which location, which can lead to process improvement and control.

[0070] Third Embodiment In a third embodiment, a configuration for calculating a predicted value from a plurality of types of observation data will be described.

[0071] Typically, there are several measurement points on a single wafer. Rather than calculating these measurement points independently, a highly accurate and easily interpretable model can be realized by extracting features or calculating predicted values ​​based on the physical dimensions of the measurement points.

[0072] FIG. 9 is an explanatory diagram illustrating a prediction method in the third embodiment. In the third embodiment, multimodal virtual measurement that takes spatial correlation into consideration will be described. The information processing device 100 acquires multiple types of observation data. In FIG. 9, inputs 1 to 3 are observation data that are input to feature extraction models MD11, MD12, and MD13, respectively. For example, input 1 is plasma emission intensity by OES, input 2 is an image captured by a wafer optical inspection system, and input 3 is a process log. The number of types of observation data used for prediction is not limited to three, and may be two, four, or more.

[0073] The feature extraction model MD11 is a model corresponding to the feature extraction model MD1 described in the first embodiment, and is trained so that when observation data of input 1 is input, the feature extraction model MD11 outputs the feature of the observation data. The same is true for the feature extraction models MD12 and MD13, which are trained so that when observation data of input 1 and input 2 are input, the feature extraction models MD12 and MD13 output the respective feature. The memory unit 102 of the information processing device 100 stores the trained feature extraction models MD11, MD12, and MD13.

[0074] The information processing device 100 extracts features of inputs 1 to 3 using feature extraction models MD11 to MD13, respectively, and converts the dimension of each extracted feature into a feature of a target dimension. The dimension conversion of the feature is performed using the dimension mapping described in the first embodiment. The feature extracted from the feature extraction model MD11 is converted into a feature of, for example, N x ×N y When converting the feature values ​​into two-dimensional feature values, the feature values ​​extracted from the feature extraction models MD12 and MD13 are also converted into N x ×N y The resulting image is converted into a two-dimensional feature.

[0075] The information processing device 100 connects the feature quantities after the dimension transformation in a connection layer CL. x ×N y When two-dimensional features of are obtained, a channel is added and N x ×N y×C, where C is the number of inputs (number of types of observed data), and in the case of FIG. 9, C=3.

[0076] The information processing device 100 inputs the feature quantities linked by the linking layer CL into a prediction model MD20 to determine a predicted value. The prediction model MD20 is a model corresponding to the prediction model MD2 described in the first embodiment, and is trained to output a predicted value related to substrate processing in response to the input of the feature quantities. The types of models that can be used for the prediction model MD20 and the model training method are the same as those in the first embodiment. The memory unit 102 of the information processing device 100 stores the trained prediction model MD20. The information processing device 100 uses the prediction model MD20 stored in the memory unit 102 to calculate a predicted value at each location on the substrate.

[0077] As described above, in the third embodiment, a method for performing multimodal virtual measurement using a learning model (prediction model MD20) incorporating spatial correlation has been disclosed. By applying the method disclosed in the second embodiment to the prediction model MD20, it is possible to calculate the contribution of features for each modality and location. This makes it possible to understand the locations within the dimensions that each modality excels at, improving interpretability.

[0078] Furthermore, it is possible to explicitly use the location within the dimension in which each modal excels. For example, prediction accuracy can be improved by predicting the substrate periphery using the plasma emission intensity and process log from OES, and predicting the area excluding the substrate periphery using images captured by a wafer optical inspection system. Furthermore, it is possible to analyze which modal has an effect on which location, leading to improvements in the model and process.

[0079] Fourth Embodiment In a fourth embodiment, a configuration for outputting a warning in accordance with a predicted value will be described.

[0080] 10 is a flowchart showing the procedure of processing executed by the information processing apparatus 100 according to embodiment 4. The control unit 101 of the information processing apparatus 100 acquires observation data to be used for prediction from the substrate processing apparatus 200, for example, via the communication unit 103 (step S401).

[0081] The control unit 101 calculates a predicted value for each location based on the acquired observation data (step S402). The method of calculating the predicted value is the same as in embodiment 1. That is, the control unit 101 inputs the acquired observation data into a feature extraction model MD1 to extract features, and performs dimension mapping of the extracted features to the target dimensions. Next, the control unit 101 inputs the dimension-mapped features into a prediction model MD2 and performs calculations to calculate a predicted value for each location. When multiple types of observation data are obtained as the observation data to be used for prediction, the control unit 101 may calculate a predicted value using the prediction model MD20 using the method disclosed in embodiment 3.

[0082] The control unit 101 determines whether or not an alarm needs to be output based on the calculated predicted value (step S403). For example, the control unit 101 compares the calculated predicted value with a preset threshold, and determines that an alarm needs to be output if the predicted value exceeds the threshold (or is less than the threshold). Alternatively, the control unit 101 may determine whether or not the predicted value falls within a preset normal range, and determine that an alarm needs to be output if the predicted value falls outside the normal range. Note that the threshold and normal range may be set for each location to be predicted.

[0083] If it is determined that an alarm output is not necessary (S403: NO), the control unit 101 ends the processing of this flowchart without outputting an alarm.

[0084] If it is determined that an alarm needs to be output (S403: YES), the control unit 101 outputs an alarm (step S404). For example, the control unit 101 outputs the alarm by displaying information that the substrate processing is not normal on the display unit 105. Alternatively, the control unit 101 may notify the information that the substrate processing is not normal via the communication unit 103 to a user terminal or the like.

[0085] In this embodiment, predictions are made using prediction models that take spatial correlation into account (prediction models MD2 and MD20), which allows for more accurate predictions. In this embodiment, such highly accurate predictions are compared with thresholds and normal ranges, allowing for more accurate determination of whether or not to issue an alarm.

[0086] Fifth Embodiment In a fifth embodiment, a configuration for executing control in substrate processing based on predicted values ​​will be described.

[0087] 11 is a flowchart showing the procedure of processing executed by the information processing apparatus 100 according to embodiment 5. The control unit 101 of the information processing apparatus 100 acquires observation data to be used for prediction from the substrate processing apparatus 200, for example, via the communication unit 103 (step S501).

[0088] The control unit 101 calculates a predicted value for each location based on the acquired observation data (step S502). The method of calculating the predicted value is the same as in embodiment 1. That is, the control unit 101 inputs the acquired observation data into a feature extraction model MD1 to extract features, and performs dimension mapping of the extracted features to the target dimensions. Next, the control unit 101 inputs the dimension-mapped features into a prediction model MD2 and performs calculations to calculate a predicted value for each location. When multiple types of observation data are obtained as the observation data to be used for prediction, the control unit 101 may calculate a predicted value using the prediction model MD20 using the method disclosed in embodiment 3.

[0089] The control unit 101 executes control related to the substrate processing in the substrate processing apparatus 200 based on the calculated predicted value (step S503). For example, the control unit 101 compares the calculated predicted value with a preset reference value, and determines a control value for the substrate processing apparatus 200 based on the deviation between the predicted value and the reference value (e.g., a control value that brings the predicted value closer to the reference value). The reference value may be set for each location to be predicted. The control unit 101 performs control related to the substrate processing by outputting a control command including the determined control value to the substrate processing apparatus 200.

[0090] In this embodiment, prediction is performed using prediction models (prediction models MD2 and MD20) that take spatial correlation into account, thereby obtaining more accurate predicted values. In this embodiment, substrate processing is controlled based on such highly accurate predicted values, which can lead to process improvement.

[0091] The embodiments disclosed herein are to be considered in all respects as illustrative and not restrictive. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims.

[0092] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any and all combinations, regardless of the reference format. Furthermore, the claims use a format in which a claim references two or more other claims (multiple claim format), but this is not limited to this. A multiple claim (multi-multi claim) that references at least one other multiple claim may also be used.

[0093] REFERENCE SIGNS LIST 100 Information processing device 101 Control unit 102 Storage unit 103 Communication unit 104 Operation unit 105 Display unit 200 Substrate processing device PG1 Prediction processing program MD1 Feature extraction model MD2 Prediction model RM Recording medium

Claims

1. Obtaining data on substrate processing, extracting features of the acquired data using a first learning model trained to output features of the data in response to input of the data; converting the extracted feature quantities into feature quantities of a target dimension set in accordance with the physical feature to be predicted regarding the substrate processing; A predicted value is obtained by inputting the feature quantity after dimension conversion into a second learning model that has been trained to output a predicted value regarding the physical feature in response to input of the feature quantity having the target dimension. A computer program that causes a computer to execute a process.

2. The target dimension is set by expanding or contracting the dimension of the extracted feature quantity according to the dimension of the physical feature.

2. The computer program product according to claim 1, for causing the computer to execute a process.

3. Outputs data showing the spatial distribution of the feature values ​​after dimension transformation 2. The computer program product according to claim 1, for causing the computer to execute a process.

4. The second learning model is trained using a loss function in which a weight is set for the spatial distribution of the feature amount.

2. The computer program of claim 1.

5. acquiring a plurality of types of data relating to the substrate processing; extracting features from each of the acquired multiple types of data using the first learning model; converting each of the feature quantities extracted from each of the plurality of types of data into a feature quantity of the target dimension; Each of the feature quantities after the dimension transformation is input to the second learning model to obtain a predicted value.

2. The computer program product according to claim 1, for causing the computer to execute a process.

6. calculating a contribution of the feature amount for each location on the substrate to the predicted value; Output the calculation results 2. The computer program product according to claim 1, for causing the computer to execute a process.

7. Calculating the contribution of said data to each location on the substrate; Control of the substrate processing is performed according to the calculation result.

2. The computer program product according to claim 1, for causing the computer to execute a process.

8. A warning is output according to the predicted value obtained using the second learning model.

2. The computer program product according to claim 1, for causing the computer to execute a process.

9. Control of the substrate processing is performed based on the predicted value obtained using the second learning model.

2. The computer program product according to claim 1, for causing the computer to execute a process.

10. The second learning model is trained using a loss function that sets a weight for the spatial distribution of the feature amount.

4. The computer program of claim 3.

11. Obtaining data on substrate processing, extracting features of the acquired data using a first learning model trained to output features of the data in response to input of the data; converting the extracted feature quantities into feature quantities of a target dimension set in accordance with the physical feature to be predicted regarding the substrate processing; Set weights in the loss function for the spatial distribution of the feature values ​​after dimension transformation, A second learning model is generated that outputs a predicted value regarding the physical feature in response to an input of the feature quantity, using a loss function to which a weight is set. A computer program that causes a computer to execute a process.

12. Obtaining data on substrate processing, extracting features of the acquired data using a first learning model trained to output features of the data in response to input of the data; converting the extracted feature quantities into feature quantities of a target dimension set in accordance with the physical feature to be predicted regarding the substrate processing; A predicted value is obtained by inputting the feature quantity after dimension conversion into a second learning model that has been trained to output a predicted value regarding the physical feature in response to input of the feature quantity having the target dimension. An information processing method in which processing is carried out by a computer.

13. an acquisition unit that acquires data related to substrate processing; an extraction unit that extracts feature quantities from the acquired data using a first learning model that has been trained to output feature quantities of the data in response to input of the data; a conversion unit that converts the extracted feature quantity into a feature quantity of a target dimension that is set in accordance with a physical feature to be predicted regarding the substrate processing; a predicted value calculation unit that inputs the feature quantity after dimension conversion into a second learning model that has been trained to output a predicted value regarding the physical feature in response to input of the feature quantity having the target dimension, and calculates a predicted value; An information processing device comprising: