Inference device, inference method, and inference program
Through multiple machine learning network departments processing time series data and adjusting output data in combination with correction parameters, the problem of model multiplexing between different manufacturing processes is solved, high-precision inference is realized and cost and time expenditure is reduced.
Patent Information
- Application Number
- CN202080082402.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-29
- Filing Date
- 2020-11-16
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-11-16
AI Technical Summary
Existing inference technologies are difficult to efficiently reuse models between different manufacturing processes, resulting in higher cost and time spent on model optimization.
Multiple machine learning network units are used to process the time series data group, and the synthesis results are generated through the link unit. The adjustment unit uses correction parameters to adjust the output data to generate high-precision inference results, and reduce individual errors through micro-adjustment functions when applying other processes.
It realizes high-precision inferences regardless of the application object, reduces the cost and time spent on model optimization, and reduces the error caused by individual differences between processes.
Smart Images

Figure CN114746820B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an inference device, an inference method, and an inference program. Background Art
[0002] Conventionally, in the field of various manufacturing processes, an inference technique is known for inferring the state of an object after processing, an event in a process during processing, etc. based on measurement data (a data set of multiple types of time-series data, hereinafter referred to as a time-series data group) measured during the processing of the object.
[0003] As an example, in a semiconductor manufacturing process, a virtual measurement technique for inferring the state of a processed wafer and an anomaly detection technique for inferring the presence or absence of an anomaly in a process during processing are known.
[0004] On the other hand, models (for example, a virtual measurement model, an anomaly detection model) used in these inference techniques need to generate a model for each process and optimize it in order to achieve higher-precision inference, thus taking cost and time.
[0005] In contrast, if a model that has achieved high-precision inference for a specific process can also be applied to other processes of the same type, the cost and time spent on optimizing the model can be reduced.
[0006] <Prior Art Documents>
[0007] <Patent Documents>
[0008] Patent Document 1: Japanese Patent Application Laid-Open No. 2006-163517 Summary of the Invention
[0009] <Problems to be Solved by the Invention>
[0010] The present invention provides an inference device, an inference method, and an inference program that can perform high-precision inference regardless of the application object.
[0011] <Means for Solving the Problems>
[0012] An inference device according to one aspect of the present invention has, for example, the following configuration. That is, it has:
[0013] An acquisition unit that acquires a time-series data group measured during the processing of an object in a prescribed processing unit of a manufacturing process;
[0014] A plurality of network units that have completed machine learning and a connection unit that has completed machine learning, which are generated by a learning unit including the plurality of network units for processing the previously acquired time series data group and the connection unit for synthesizing each output data output by processing using the plurality of network units, and causing the plurality of network units and the connection unit to perform machine learning in such a way that the synthesis result output by the connection unit is close to the inspection data of the result when processing the object; and
[0015] An adjustment unit that adjusts each output data output by processing the acquired time series data group using the plurality of network units that have completed machine learning without synthesizing through the connection unit that has completed machine learning, and synthesizes the adjusted output data to output an inference result,
[0016] The adjustment unit adjusts each output data using a correction parameter corresponding to an error included in the inference result.
[0017] <Effects of the Invention>
[0018] According to the present invention, it is possible to provide an inference device, an inference method, and an inference program that can perform high-precision inference regardless of the application object. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG. is an example showing the overall configuration of a system to which the hypothetical measurement device is applied.
[0020] Figure 2 FIG. is a first diagram showing an example of a prescribed processing unit of a semiconductor manufacturing process.
[0021] Figure 3 FIG. is a second diagram showing an example of a prescribed processing unit of a semiconductor manufacturing process.
[0022] Figure 4 FIG. is a diagram showing an example of the acquired time series data group.
[0023] Figure 5 FIG. is a diagram showing an example of the hardware configuration of the hypothetical measurement device.
[0024] Figure 6 FIG. is a diagram showing an example of the functional configuration of the learning unit of the hypothetical measurement device.
[0025] Figure 7 FIG. is a first diagram showing a specific example of the processing of a branch.
[0026] Figure 8 FIG. is a second diagram showing a specific example of the processing of a branch.
[0027] Figure 9 This is the third figure showing a specific example of the processing of the branch section.
[0028] Figure 10 This is a figure showing a specific example of the processing of the standardization section included in each network section.
[0029] Figure 11 This is the fourth figure showing a specific example of the processing of the branch section.
[0030] Figure 12 This is a figure showing an example of the functional configuration of the inference section of the virtual measurement device.
[0031] Figure 13 This is a flowchart showing the flow of the virtual measurement process of the virtual measurement device.
[0032] Figure 14 This is the first figure showing an example of the functional configuration of the inference section with fine adjustment function of the virtual measurement device.
[0033] Figure 15 This is a flowchart showing the flow of the fine adjustment process of the virtual measurement device.
[0034] Figure 16 This is the second figure showing an example of the functional configuration of the inference section with fine adjustment function of the virtual measurement device. Detailed implementation mode
[0035] Hereinafter, each embodiment will be described with reference to the accompanying drawings. It should be noted that in each of the following embodiments, a case will be described in which a virtual measurement model for inferring the state of a processed wafer or an anomaly detection model for inferring the presence or absence of an anomaly in a process is generated using time series data groups measured along with the processing of a wafer for a specific semiconductor manufacturing process. At this time, in each of the following embodiments, by using a plurality of network sections to process the time series data groups, multi-faceted analysis is performed to generate a model for achieving high-precision inference.
[0036] In addition, in each of the following embodiments, by adding a fine adjustment function to the generated model, when applying the model to other semiconductor manufacturing processes of the same type, the fine adjustment function is used to reduce the error (the error included in the inference result) caused by individual differences between processes.
[0037] Thus, according to each of the following embodiments, it is possible to provide an inference device, an inference method, and an inference program that can perform high-precision inference regardless of the application object. As a result, compared with the case of newly generating a model for other semiconductor manufacturing processes and performing optimization, cost and time can be reduced.
[0038] It should be noted that in the following embodiments, in the first embodiment, a case where a hypothetical measurement model is generated as a model based on a time series data group and a correction matrix is used as a fine-tuning function will be described. In addition, in the second embodiment, a case where a neural network is used instead of the correction matrix as a fine-tuning function will be described. Further, in the third embodiment, a case where an anomaly detection model is generated instead of the hypothetical measurement model as a model based on a time series data group will be described.
[0039] It should be noted that in each embodiment and the accompanying drawings, for components that substantially have the same functional configuration, the same reference numerals are given and repeated descriptions are omitted.
[0040] [First Embodiment]
[0041] [Application Example of Deduction Device]
[0042] First, an application example of a hypothetical measurement device (deduction device) with a fine-tuning function added to the hypothetical measurement model will be described. Figure 1 FIG. is an example showing the overall configuration of a system to which the hypothetical measurement device is applied.
[0043] As Figure 1 shown, the system 100A includes a semiconductor manufacturing process A, time series data acquisition devices 140A_1 to 140A_n, an inspection data acquisition device 150A, and a hypothetical measurement device 160A. In the system 100A, a specific process, i.e., the semiconductor manufacturing process A, is taken as an object, and a hypothetical measurement model for achieving high-precision deduction is generated.
[0044] The system 100B includes a semiconductor manufacturing process B, time series data acquisition devices 140B_1 to 140B_n, an inspection data acquisition device 150B, and a hypothetical measurement device 160B. In the system 100B, the semiconductor manufacturing process B is another process of the same type as the semiconductor manufacturing process A. In the present embodiment, it is the application object to which a hypothetical measurement device (deduction device) with a fine-tuning function added to the hypothetical measurement model generated in the system 100A is applied.
[0045] In the system 100A, the semiconductor manufacturing process A processes an object (pre-processed wafer 110A) in a specified processing unit 120A to generate a product (post-processed wafer 130A). It should be noted that the processing unit 120A mentioned here is an abstract concept and will be described in detail later. In addition, the pre-processed wafer 110A refers to the wafer (substrate) before being processed in the processing unit 120A, and the post-processed wafer 130A refers to the wafer (substrate) after being processed in the processing unit 120A.
[0046] In addition, in system 100A, the time series data acquisition devices 140A_1 to 140A_n measure time series data respectively in association with the processing of the pre-process wafer 110A. The time series data acquisition devices 140A_1 to 140A_n measure different types of measurement items from each other. It should be noted that the number of measurement items measured by the time series data acquisition devices 140A_1 to 140A_n respectively can be one or more. In addition, among the time series data measured in association with the processing of the pre-process wafer 110A, in addition to the time series data measured in the processing of the pre-process wafer 110A, it also includes the time series data measured during the pre-processing and post-processing performed before and after the processing of the pre-process wafer 110A. These processes can include pre-processing and post-processing performed in a state where there is no wafer (substrate).
[0047] The time series data group measured by the time series data acquisition devices 140A_1 to 140A_n is stored as learning data (input data) in the learning data storage unit 163A of the virtual measurement device 160A.
[0048] In addition, in system 100A, the inspection data acquisition device 150A inspects a specified inspection item (for example, ER (Etch Rate)) of the post-process wafer 130A after being processed in the processing unit 120A, and acquires inspection data. The inspection data acquired by the inspection data acquisition device 150A is stored as learning data (correct answer data) in the learning data storage unit 163A of the virtual measurement device 160A.
[0049] In addition, in system 100A, a virtual measurement program including a learning program and an inference program is installed in the virtual measurement device 160A. By executing the virtual measurement program, the virtual measurement device 160A functions as a learning unit 161A and an inference unit 162A.
[0050] The learning unit 161A performs machine learning using the time series data group measured by the time series data acquisition devices 140A_1 to 140A_n and the inspection data acquired by the inspection data acquisition device 150A.
[0051] Specifically, the learning unit 161A uses multiple network units to process the time series data group, and makes the multiple network units perform machine learning in such a way that the combined result of the output data output by the multiple network units is close to the inspection data.
[0052] The inference unit 162A acquires a time series data group measured in association with the processing of a new object (wafer before processing), and inputs the same to a plurality of network units that have undergone machine learning. Thereby, the inference unit 162A infers the inspection data of the wafer after processing based on the time series data acquired in association with the processing of the new wafer before processing, and outputs an inference result (hypothetical measurement data).
[0053] In this way, by using a plurality of network units to process the time series data group measured in association with the processing of the object, the hypothetical measurement device 160A can perform multi-faceted analysis of the time series data group. As a result, compared with the case of using a single network unit to process the time series data group, a hypothetical measurement model (inference unit 162A) for realizing highly accurate inference can be generated.
[0054] On the other hand, in the system 100B, the semiconductor manufacturing process B is the same type of process as the semiconductor manufacturing process A of the system 100A. Further, in the system 100B, the time series data acquisition devices 140B_1 to 140B_n and the inspection data acquisition device 150B respectively correspond to the time series data acquisition devices 140A_1 to 140A_n and the inspection data acquisition device 150A of the system 100A.
[0055] Moreover, in the system 100B, the hypothetical measurement device 160B (inference device) corresponds to the hypothetical measurement device 160A of the system 100A. However, in the case of the hypothetical measurement device 160B of the system 100B, it does not have the learning unit 161A. Further, instead of the inference unit 162A, it has an inference unit 162B with a fine-tuning function (installed with a hypothetical measurement program that does not include a learning program and includes the same inference program as the inference program installed in the hypothetical measurement device 160A).
[0056] In the case of the hypothetical measurement device 160B of the system 100B, instead of newly generating a hypothetical measurement model and performing optimization by using the time series data group for machine learning, the hypothetical measurement model (inference unit 162A) generated in the hypothetical measurement device 160A of the system 100A is applied.
[0057] Here, as described above, although the semiconductor manufacturing process A and the semiconductor manufacturing process B are the same type of process, there are individual differences between them. Therefore, even if the hypothetical measurement model (inference unit 162A) generated in the hypothetical measurement device 160A is applied as it is, errors will be included in the inference result (hypothetical measurement data).
[0058] Then, in the case of the hypothetical measurement device 160B (inference device), an inference unit with a fine-tuning function added to the hypothetical measurement model (inference unit 162A) generated in the hypothetical measurement device 160A is generated. InFigure 1 Among them, the inference unit 162B with a micro-adjustment function in the hypothetical measurement device 160B is an example of an inference unit obtained by adding a micro-adjustment function to the hypothetical measurement model (inference unit 162A) generated in the hypothetical measurement device 160A.
[0059] The inference unit 162B with a micro-adjustment function applies the hypothetical measurement model (inference unit 162A) generated in the hypothetical measurement device 160A (see the dashed line 170). On the other hand, it is equipped with a micro-adjustment function for reducing the error caused by individual differences (the error included in the inference result).
[0060] Specifically, the inference unit 162B with a micro-adjustment function updates the correction parameters (parameters included in the correction matrix used when adjusting each output data. Details will be described later) in such a way that the error between the inference result (hypothetical measurement data) output by processing the time-series data group using a plurality of network units included in the generated hypothetical measurement model, adjusting each output data output by the plurality of network units, and then synthesizing them, and the inspection data obtained by the inspection data acquisition device 150B is reduced.
[0061] As a result, in the hypothetical measurement device 160B, it is possible to implement a model that applies the hypothetical measurement model (inference unit 162A) generated in the hypothetical measurement device 160A for achieving high-precision inference and can also achieve high-precision inference in the application target, i.e., the semiconductor manufacturing process B.
[0062]
[0063] Next, the specified processing units 120A and 120B of the semiconductor manufacturing processes A and B will be described. Figure 2 It is the first figure showing an example of a specified processing unit of the semiconductor manufacturing process. As Figure 2 shown, a semiconductor manufacturing apparatus 200, which is an example of a substrate processing apparatus, has a plurality of chambers (an example of a plurality of processing spaces. In the Figure 2 example, they are "chamber A" to "chamber C"), and wafers are processed in each chamber.
[0064] Among them, Figure 2 2a in shows a case where a plurality of chambers are defined as the processing units 120A and 120B. In this case, the wafers 110A and 110B before processing refer to the wafers before being processed in chamber A, and the wafers 130A and 130B after processing refer to the wafers after being processed in chamber C.
[0065] In addition, in Figure 2In the processing units 120A and 120B of 2a, the time-series data groups measured along with the processing of the wafers 110A and 110B before processing include the time-series data groups measured along with the processing in chamber A (the first processing space), the time-series data groups measured along with the processing in chamber B (the second processing space), and the time-series data groups measured along with the processing in chamber C (the third processing space).
[0066] On the other hand, Figure 2 2b shows a case where one chamber (chamber B in the example of Figure 2 2b) is defined as the processing units 120A and 120B. In this case, the wafers 110A and 110B before processing refer to the wafers before being processed in chamber B (the wafers after being processed in chamber A). In addition, the wafers 130A and 130B after processing refer to the wafers after being processed in chamber B (the wafers before being processed in chamber C).
[0067] In addition, Figure 2 in the processing units 120A and 120B of 2b, the time-series data groups measured along with the processing of the wafers 110A and 110B before processing include the time-series data groups measured along with the processing of the wafers 110A and 110B before processing in chamber B.
[0068] Figure 3 is the second figure showing an example of a specified processing unit in a semiconductor manufacturing process. Similar to Figure 2 the same, the semiconductor manufacturing apparatus 200 has a plurality of chambers, and in each chamber, wafers are processed.
[0069] Among them, Figure 3 3a shows a case where the processing content in chamber B, excluding the pre-processing and post-processing (referred to as "wafer processing"), is defined as the processing units 120A and 120B. In this case, the wafers 110A and 110B before processing refer to the wafers before wafer processing (the wafers after pre-processing), and the wafers 130A and 130B after processing refer to the wafers after wafer processing (the wafers before post-processing).
[0070] In addition, Figure 3 in the processing units 120A and 120B of 3a, the time-series data groups measured along with the processing of the wafers 110A and 110B before processing include the time-series data groups measured along with the wafer processing of the wafers 110A and 110B before processing in chamber B.
[0071] It should be noted that in Figure 3In the example of 3a, the case is shown where the wafer processing is set as processing units 120A and 120B when pre-processing, wafer processing (main processing), and post-processing are performed in the same chamber (chamber B). However, in the case where each process is performed in different chambers (for example, pre-processing is performed in chamber A, wafer processing is performed in chamber B, and post-processing is performed in chamber C), the respective processes in each chamber can be set as processing units 120A and 120B.
[0072] On the other hand, Figure 3 3b of shows the case where one scenario included in the wafer processing in the processing content in chamber B (in the example of Figure 3 3b is "Scenario III") is defined as processing units 120A and 120B. In this case, the pre-process wafers 110A and 110B refer to the wafers before the processing of Scenario III (the wafers after the processing of Scenario II). In addition, the post-process wafers 130A and 130B refer to the wafers after the processing of Scenario III (the wafers before the processing of Scenario IV (not shown)).
[0073] In addition, in Figure 3 the processing units 120A and 120B of 3a, the time series data group measured along with the processing of the pre-process wafers 110A and 110B includes the time series data group measured along with the wafer processing of Scenario III in chamber B.
[0074] <Specific example of time series data group>
[0075] Next, a specific example of the time series data group obtained in the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n will be described. Figure 4 is a diagram showing an example of the obtained time series data group. It should be noted that, in Figure 4 the example, for simplicity of explanation, it is set that the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n respectively measure one-dimensional data. However, one time series data acquisition device can measure two-dimensional data (a data set of multiple types of one-dimensional data).
[0076] Among them, Figure 4 4a of shows that the processing units 120A and 120B are composed of Figure 2 2b of, Figure 3 3a of, Figure 3A time series data group in the case where any of 3b is defined. In this case, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n respectively acquire time series data measured in association with the processing in chamber B. In addition, the time series data measured by the time series data acquisition devices 140A_1 to 140A_n at the same time band are acquired as a time series data group. Similarly, the time series data measured by the time series data acquisition devices 140B_1 to 140B_n at the same time band are acquired as a time series data group.
[0077] On the other hand, Figure 4 In the case of the time series data group in 4b where the processing units 120A and 120B are defined by Figure 2 2a. In this case, the time series data acquisition devices 140A_1 to 140A_3 and 140B_1 to 140B_3 acquire, for example, a time series data group 1 measured in association with the processing of a wafer in chamber A. In addition, the time series data acquisition devices 140A_n - 2 and 140B_n - 2 acquire, for example, a time series data group 2 measured in association with the processing of the wafer in chamber B. In addition, the time series data acquisition devices 140A_n - 1 to 140A_n and 140B_n - 1 to 140B_n acquire, for example, a time series data group 3 measured in association with the processing of the wafer in chamber C.
[0078] It should be noted that, in Figure 4 4a, a case is shown where the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n acquire time series data in the same time range measured in association with the processing of a pre - processed wafer in chamber B. However, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n may acquire time series data in different time ranges measured in association with the processing of a pre - processed wafer in chamber B.
[0079] Specifically, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n may acquire a plurality of time series data measured during the execution of pre - processing as a time series data group 1. In addition, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_ may acquire a plurality of time series data measured during the execution of wafer processing as a time series data group 2. Moreover, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n may acquire a plurality of time series data measured during the execution of post - processing as a time series data group 3.
[0080] Similarly, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n can acquire a plurality of time series data measured during the execution of Scenario I as the time series data group 1. In addition, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n can acquire a plurality of time series data measured during the execution of Scenario II as the time series data group 2. Moreover, the time series data acquisition devices 140A_1 to 140A_n and 140B_1 to 140B_n can acquire a plurality of time series data measured during the execution of Scenario III as the time series data group 3.
[0081] <Hardware Configuration of Hypothetical Measurement Device>
[0082] Next, the hardware configurations of the hypothetical measurement devices 160A and 160B will be described. Figure 5 This is a diagram showing an example of the hardware configuration of the hypothetical measurement device. As Figure 5 shown, the hypothetical measurement devices 160A and 160B include a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503. In addition, the hypothetical measurement device 160 includes a GPU (Graphics Processing Unit) 504. It should be noted that processors (processing circuits, Processing Circuit, Processing Circuitry) such as the CPU 501 and GPU 504 and memories such as the ROM 502 and RAM 503 form a so-called computer.
[0083] Moreover, the hypothetical measurement device 160 includes an auxiliary storage device 505, a display device 506, an operation device 507, an I / F (Interface) device 508, and a drive device 509. It should be noted that the various hardware components of the hypothetical measurement device 160 are connected to each other via a bus 510.
[0084] The CPU 501 is an arithmetic device for executing various programs (e.g., hypothetical measurement programs, etc.) installed in the auxiliary storage device 505.
[0085] The ROM 502 is a non-volatile memory that functions as the main storage device. The ROM 502 is used to store various programs, data, etc. required for the CPU 501 to execute various programs installed in the auxiliary storage device 505. Specifically, the ROM 502 is used to store boot programs such as BIOS (Basic Input / Output System) and EFI (Extensible Firmware Interface).
[0086] The RAM 503 is a volatile memory such as DRAM (Dynamic Random Access Memory) or SRAM (Static Random Access Memory), which functions as the main storage device. The RAM 503 provides a working area for the execution of various programs installed in the auxiliary storage device 505 by the CPU 501.
[0087] The GPU 504 is an arithmetic device for image processing. When the CPU 501 executes the virtual measurement program, high-speed arithmetic operations based on parallel processing are performed on various image data (time series data groups in this embodiment). It should be noted that the GPU 504 is equipped with an internal memory (GPU memory), which temporarily stores information required for parallel processing of various image data.
[0088] The auxiliary storage device 505 is used to store various programs, various data used when the various programs are executed by the CPU 501, and the like.
[0089] The display device 506 is a display device for displaying the internal states of the virtual measurement devices 160A and 160B. The operation device 507 is an input device used by the administrator of the virtual measurement devices 160A and 160B to input various instructions to the virtual measurement devices 160A and 160B. The I / F device 508 is a connection device for connecting to a network (not shown) to perform communication.
[0090] The drive device 509 is a device for accommodating the storage medium 520. The storage medium 520 mentioned here includes media that record information using light, electricity, or magnetism, such as CD-ROMs, floppy disks, and magneto-optical disks. In addition, the storage medium 520 may include semiconductor memories that record information electrically, such as ROMs and flash memories.
[0091] It should be noted that various programs installed in the auxiliary storage device 505 are placed in the drive device 509, for example, through the distributed storage medium 520, and various programs recorded in the storage medium 520 are read by the drive device 509 for installation. Alternatively, various programs installed in the auxiliary storage device 505 can be downloaded through the network for installation.
[0092] <Functional Composition of the Learning Unit>
[0093] Next, the functional composition of the learning unit 161A of the hypothetical measurement device 160A in the system 100A will be described. Figure 6 It is a diagram showing an example of the functional composition of the learning unit of the hypothetical measurement device. The learning unit 161A includes a branch unit 610, first to Mth network units 620_1 to 620_M, a connection unit 630, and a comparison unit 640.
[0094] The branch unit 610 reads the time series data group from the learning data storage unit 163A. In addition, the branch unit 610 processes the time series data group in such a way that it is processed by the multiple network units of the first to Mth network units 620_1 to 620_M.
[0095] The first to Mth network units 620_1 to 620_M are configured based on a convolutional neural network (CNN: Convolutional Neural Network) and have multiple layers.
[0096] Specifically, the first network unit 620_1 has first to Nth layers 620_11 to 620_1N. Similarly, the second network unit 620_2 has first to Nth layers 620_21 to 620_2N. Hereinafter, having the same configuration, the Mth network unit 620_M has first to Nth layers 620_M1 to 620_MN.
[0097] In each of the first to Nth layers 620_11 to 620_1N of the first network unit 620_1, various processes such as normalization processing, convolution processing, activation processing, and pooling processing are performed. In addition, the same various processes are also performed in each of the second to Mth network units 620_2 to 620_M.
[0098] The connection unit 630 synthesizes the output data from the Nth layer 620_1N of the first network unit 620_1 to the output data from the Nth layer 620_MN of the Mth network unit 620_M, and outputs the synthesis result to the comparison unit 640.
[0099] The comparison unit 640 compares the synthesis result output by the connection unit 630 with the inspection data (correct answer data) read from the learning data storage unit 163A, and calculates the error. In the learning unit 161A, the first to Mth network units 620_1 to 620_M and the connection unit 630 perform machine learning in such a way that the error calculated by the comparison unit 640 satisfies a specified condition.
[0100] Thus, the model parameters of the first layer to the Nth layer of each of the first network unit 620_1 to the Mth network unit 620_M and the model parameters of the connection unit 630 are optimized.
[0101] <Details of the processing of each unit of the learning unit>
[0102] Next, specific examples will be given to illustrate the details of the processing of each unit (here, particularly the branch unit 610) of the learning unit 161A of the virtual measurement device 160A in the system 100A.
[0103] (1) Details of the processing of the branch unit 1
[0104] Figure 7 is the first diagram showing a specific example of the processing of the branch unit. In Figure 7 the case of, the branch unit 610 processes the time series data group measured by the time series data acquisition devices 140A_1 to 140A_n according to the first criterion, thereby generating a time series data group 1 (first time series data group), and inputs it to the first network unit 620_1.
[0105] In addition, the branch unit 610 processes the time series data group measured by the time series data acquisition devices 140A_1 to 140A_n according to the second criterion, thereby generating a time series data group 2 (second time series data group), and inputs it to the second network unit 620_2.
[0106] In this way, it is configured to process the time series data group according to different criteria and perform machine learning on the basis of being divided into different network units for processing, so that the time series data group can be analyzed from multiple aspects. As a result, compared with the case of inputting the time series data group into a single network unit for machine learning, a virtual measurement model (inference unit 162A) for achieving high-precision inference can be generated.
[0107] It should be noted that in Figure 7 the example of, a case where two types of time series data groups are generated by processing the time series data group according to two types of criteria is shown, and it is also possible to generate three or more types of time series data groups by processing the time series data group according to three or more types of criteria.
[0108] (2) Details of the processing of the branch unit 2
[0109] Next, the details of another process of the branch unit 610 will be described. Figure 8 is the second diagram showing a specific example of the processing of the branch unit. In Figure 8In the case of, the branch 610 groups the time series data groups measured by the time series data acquisition devices 140A_1 to 140A_n according to the type. Thus, the branch 610 generates a time series data group 1 (first time series data group) and a time series data group 2 (second time series data group). In addition, the branch 610 inputs the generated time series data group 1 into the third network unit 620_3, and inputs the generated time series data group 2 into the fourth network unit 620_4.
[0110] In this way, on the basis of a configuration in which the time series data group is divided into multiple groups according to the data type and different network units are used for processing, machine learning is performed, so that the time series data group can be analyzed in multiple aspects. As a result, compared with the case where the time series data group is input into one network unit for machine learning, a hypothetical measurement model (inference unit 162A) for realizing high-precision inference can be generated.
[0111] It should be noted that in Figure 8 In the example of, although the time series data group is grouped according to different data types based on the different time series data acquisition devices 140A_1 to 140A_n, the time series data group can be grouped according to the time range of data acquisition. For example, in the case where the time series data group is a time series data group measured along with the processing of multiple scenarios, the time series data group can be grouped according to the time range of each scenario.
[0112] (3) Details of the processing of the branch 3
[0113] Next, the details of another processing of the branch 610 will be described. Figure 9 It is the third figure showing a specific example of the processing of the branch. In Figure 9 In the case of, the branch 610 inputs the time series data group obtained by the time series data acquisition devices 140A_1 to 140A_n into both the fifth network unit 620_5 and the sixth network unit 620_6. And in the fifth network unit 620_5 and the sixth network unit 620_6, different processing (normalization processing) is performed on the same time series data group.
[0114] Figure 10 It is a figure showing a specific example of the processing of the normalization unit included in each network unit. As Figure 10 shown, each layer of the fifth network unit 620_5 includes a normalization unit, a convolution unit, an activation function unit, and a pooling unit.
[0115] Figure 10The example of [[ID=]] shows that in the first layer 620_51 among the layers included in the fifth network unit 620_5, there are included a normalization unit 1001, a convolution unit 1002, an activation function unit 1003, and a pooling unit 1004.
[0116] Among them, in the normalization unit 1001, for the time series data group input from the branch unit 610, a first normalization process is performed to generate a normalized time series data group 1 (first time series data group).
[0117] Similarly, Figure 10 The example of [[ID=]] shows that in the first layer 620_61 among the layers included in the sixth network unit 620_6, there are included a normalization unit 1011, a convolution unit 1012, an activation function unit 1013, and a pooling unit 1014.
[0118] Among them, in the normalization unit 1011, for the time series data group input from the branch unit 610, a second normalization process is performed to generate a normalized time series data group 2 (second time series data group).
[0119] In this way, on the basis of a configuration in which multiple network units each including a normalization unit that performs normalization processing by different methods are set to process the time series data group, machine learning is performed, so that the time series data group can be analyzed in multiple aspects. As a result, compared with the case where machine learning is performed by inputting the time series data group into one network unit that performs one normalization process, a hypothetical measurement model (inference unit 162A) for realizing high-precision inference can be generated.
[0120] (4) Details of the processing of the branch unit 4
[0121] Next, the details of another process of the branch unit 610 will be described. Figure 11 is the fourth diagram showing a specific example of the processing of the branch unit. In Figure 11 's case, the branch unit 610 inputs the time series data group 1 (first time series data group) measured along with the processing in chamber A among the time series data groups measured by the time series data acquisition devices 140A_1 to 140A_n into the seventh network unit 620_7.
[0122] In addition, the branch unit 610 inputs the time series data group 2 (second time series data group) measured along with the processing in chamber B among the time series data groups measured by the time series data acquisition devices 140A_1 to 140A_n into the eighth network unit 620_8.
[0123] Thus, based on a configuration in which different network units are set to process each time-series data group measured in association with processing in different chambers (first processing space, second processing space), machine learning is performed, enabling multi-faceted analysis of the time-series data groups. As a result, compared to the case where each time-series data group is input to a single network unit for machine learning, a hypothetical measurement model (inference unit 162A) for achieving highly accurate inference can be generated.
[0124] <Functional configuration of the inference unit of the hypothetical measurement device>
[0125] Next, the functional configuration of the inference unit 162A of the hypothetical measurement device 160A in the system 100A will be described. Figure 12 It is a diagram showing an example of the functional configuration of the inference unit of the hypothetical measurement device. As Figure 12 shown, the inference unit 162A of the hypothetical measurement device 160A includes a branch unit 1210, first to M-th network units 1220_1 to 1220_M, and a connection unit 1230.
[0126] The branch unit 1210 acquires the time-series data groups newly measured by the time-series data acquisition devices 140A_1 to 140A_N. In addition, the branch unit 1210 controls the acquired time-series data groups to be processed using the first to M-th network units 1220_1 to 1220_M.
[0127] The first to M-th network units 1220_1 to 1220_M are formed in such a way that the model parameters of each layer of the first to M-th network units 620_1 to 620_M are optimized through machine learning by the learning unit 161A.
[0128] The connection unit 1230 is formed by the connection unit 630 whose model parameters are optimized through machine learning by the learning unit 161A. The connection unit 1230 synthesizes the output data from the N-th layer 1220_1N of the first network unit 1220_1 to the output data from the N-th layer 1220_MN of the M-th network unit 1220_M, and outputs the hypothetical measurement data.
[0129] <Flow of the hypothetical measurement process>
[0130] Next, the overall flow of the hypothetical measurement process of the hypothetical measurement device 160A in the system 100A will be described. Figure 13 It is a flowchart showing the flow of the hypothetical measurement process of the hypothetical measurement device.
[0131] In step S1301, the learning unit 161A acquires the time-series data groups and the inspection data as learning data.
[0132] In step S1302, the learning unit 161A uses the time series data group in the acquired learning data as input data and the inspection data as correct answer data for machine learning.
[0133] In step S1303, the learning unit 161A determines whether to continue machine learning. When further acquiring learning data and continuing machine learning (when the result in step S1303 is YES), it returns to step S1301. On the other hand, when ending machine learning (when the result in step S1303 is NO), it proceeds to step S1304.
[0134] In step S1304, the inference unit 162A reflects the model parameters optimized through machine learning to generate the first network unit 1220_1 to the Mth network unit 1220_M.
[0135] In step S1305, the inference unit 162A inputs the time series data group measured along with the processing of the new pre-processed wafer 110A to infer the hypothetical measurement data.
[0136] In step S1306, the inference unit 162A outputs the inferred hypothetical measurement data.
[0137] <Functional configuration of the inference unit with micro-adjustment function of the hypothetical measurement device>
[0138] Next, the functional configuration of the inference unit 162B with micro-adjustment function of the hypothetical measurement device 160B in the system 100B will be described. Figure 14 It is a diagram showing an example of the functional configuration of the inference unit with micro-adjustment function of the hypothetical measurement device.
[0139] As Figure 14 shown, the inference unit 162B with micro-adjustment function of the hypothetical measurement device 160B has a branch unit 1210 that functions as an acquisition unit. In addition, the inference unit 162B with micro-adjustment function of the hypothetical measurement device 160B has a first network unit 1220_1 to the Mth network unit 1220_M, a connection unit 1410, an individual adjustment unit 1420, a micro-adjustment unit 1430, and a comparison unit 1440 that function as an inference unit.
[0140] Among them, the branch unit 1210 is the same as the branch unit 1210 of the inference unit 162A, and since it has already been described using Figure 12 it, the description here is omitted. In addition, the first network unit 1220_1 to the Mth network unit 1220_M are also the same as the first network unit 1220_1 to the Mth network unit 1220_M of the inference unit 162A.
[0141] Specifically, the first network unit 1220_1 to the M-th network unit 1220_M are formed as follows: The model parameters of each layer of the first network unit 620_1 to the M-th network unit 620_M are optimized by machine learning performed by the learning unit 161A.
[0142] The connection unit 1410 is formed by the connection unit 630 whose model parameters are optimized by machine learning performed by the learning unit 161A. However, in the case of the connection unit 1410, the output data from the N-th layer 1220_1N of the first network unit 1220_1 to the output data from the N-th layer 1220_MN of the M-th network unit 1220_M are not combined and output.
[0143] The individual adjustment unit 1420 multiplies each output data output from the connection unit 1410 by a coefficient (referred to as "individual sensitivity") based on the individual difference between the processing unit 120A of the semiconductor manufacturing process A and the processing unit 120B of the semiconductor manufacturing process B.
[0144] The fine adjustment unit 1430 multiplies each output data multiplied by the individual sensitivity by the individual adjustment unit 1420 with a correction matrix to calculate a scalar, i.e., the imaginary measurement data.
[0145] The comparison unit 1440 acquires the imaginary measurement data output from the fine adjustment unit 1430 and also acquires the inspection data for the processed wafer 130B. In addition, the comparison unit 1440 calculates the difference between the acquired imaginary measurement data and the inspection data and notifies it to the fine adjustment unit 1430.
[0146] In this way, in the fine adjustment function inference unit 162B, in the semiconductor manufacturing process B, based on the inspection data of the processed wafer 130B for a specified period, the fine adjustment unit 1430 updates the correction parameters (P 1 ~P M ). And in the fine adjustment unit 430 of the fine adjustment function inference unit 162B, the update of the correction parameters (P 1 ~P M ) is continued until the difference between the imaginary measurement data and the inspection data is below a specified threshold.
[0147] Thereby, in the fine adjustment unit 1430, the error (the error included in the inference result) due to the individual difference between the processing unit 120A of the semiconductor manufacturing process A and the processing unit 120B of the semiconductor manufacturing process B can be reduced.
[0148] It should be noted that in the case of the fine adjustment function inference unit 162B, compared with the case of optimizing by re-learning the imaginary measurement model by using the time series data group measured in the semiconductor manufacturing process B as additional data, cost and time can be reduced.
[0149] <Flow of fine adjustment process>
[0150] Next, the flow of the fine adjustment process based on the hypothetical measurement device 160B in the system 100B will be described. Figure 15 It is a flowchart showing the flow of the fine adjustment process based on the hypothetical measurement device.
[0151] In step S1501, the branch unit 1210 of the fine adjustment function inference unit 162B acquires a time series data group measured in association with the processing of the new pre-process wafer 110B in the processing unit 120B of the semiconductor manufacturing process B. In addition, the first to M network units 1220_1 to 1220_M of the fine adjustment function inference unit 162B process the acquired time series data group. As a result, each output data is output from the final layer of the first to M network units 1220_1 to 1220_M.
[0152] In step S1502, the individual adjustment unit 1420 of the fine adjustment function inference unit 162B multiplies each output data output from the final layer of the first to M network units 1220_1 to 1220_M by the individual sensitivity, thereby adjusting each output data.
[0153] In step S1503, the fine adjustment unit 1430 of the fine adjustment function inference unit 162B multiplies each output data multiplied by the individual sensitivity by the correction matrix, thereby calculating the hypothetical measurement data.
[0154] In step S1504, the fine adjustment function inference unit 162B acquires inspection data for the post-process wafer 130B and notifies it to the comparison unit 1440. In addition, the comparison unit 1440 compares the hypothetical measurement data output from the fine adjustment unit 1430 with the notified inspection data and calculates the difference (the error included in the inference result).
[0155] In step S1505, the comparison unit 1440 of the fine adjustment function inference unit 162B determines whether the difference is equal to or less than a specified threshold based on the comparison result, thereby determining whether it is necessary to update the correction parameter.
[0156] In step S1505, when it is determined that the difference exceeds the specified threshold and it is necessary to update the correction parameter (when YES in step S1505), the process proceeds to step S1506.
[0157] In step S1506, the fine adjustment unit 1430 of the fine adjustment function inference unit 162B adjusts the correction parameter (P 1 ~P M) Update. After that, proceed to step S1507.
[0158] On the other hand, in step S1505, when the difference is equal to or less than a specified threshold value and it is determined that the correction parameter does not need to be updated (when it is NO in step S1505), directly proceed to step S1507.
[0159] In step S1507, the micro-adjustment function inference unit 162B determines whether to end the micro-adjustment process. In step S1507, when it is determined that the micro-adjustment process does not end (when it is NO in step S1507), return to step S1501.
[0160] On the other hand, in step S1507, when it is determined that the micro-adjustment process ends (when it is YES in step S1507), end the micro-adjustment process.
[0161] <Summary>
[0162] As described above, it is clearly understood that the virtual measurement device 160A acquires a time series data group measured along with the processing of an object in a specified processing unit of a manufacturing process, and it causes each network unit to perform machine learning so that the combined result of each output data output from each network unit by processing the acquired time series data group is close to the inspection data of the product obtained by processing the object.
[0163] In this way, by processing the time series data group using multiple network units, multi-faceted analysis can be performed. As a result, in the virtual measurement device 160A, a virtual measurement model for realizing highly accurate inference can be generated.
[0164] In addition, the virtual measurement device 160B (inference device) uses multiple network units included in the generated virtual measurement model to process a time series data group measured along with the processing of an object in a specified processing unit of another manufacturing process, thereby outputting each output data, and it performs micro-adjustment on each output data output using a correction parameter and then synthesizes them, thereby inferring virtual measurement data. In addition, it updates the correction parameter according to the error included in the inferred virtual measurement data.
[0165] In this way, in a specified processing unit of a manufacturing process, when applying the virtual measurement model generated using the time series data group to another manufacturing process, in the virtual measurement device 160B, a function for micro-adjusting each output data output from multiple network units is added.
[0166] Accordingly, when applying the hypothetical measurement model to another manufacturing process, it is possible to reduce the error (the error included in the inference result) caused by the individual differences between processes. That is to say, according to the first embodiment, it is possible to provide an inference device, an inference method, and an inference program that can perform high-precision inference regardless of the application object.
[0167] [Second Embodiment]
[0168] In the above first embodiment, it was described that the output data output from the final layer of each network unit is finely adjusted using individual sensitivities and a correction matrix. However, the method of finely adjusting the output data based on the outputs of the inference units with a fine adjustment function is not limited to this. For example, a network unit for fine adjustment can be used to finely adjust each output data.
[0169] Figure 16 It is a second diagram showing an example of the functional configuration of the inference unit with a fine adjustment function of the hypothetical measurement device. The difference from Figure 14 is that Figure 16 in the case of the inference unit with a fine adjustment function 1600B shown, it has a fine adjustment network unit 1610.
[0170] The fine adjustment network unit 1610 is configured based on a convolutional neural network, and by inputting the output data output from the connection unit 1410, it outputs hypothetical measurement data.
[0171] In addition, the fine adjustment network unit 1610 updates the model parameters, that is, the correction parameters, of the fine adjustment network unit 1610 based on the difference notified from the comparison unit 1440 according to the output hypothetical measurement data.
[0172] In this way, in the inference unit with a fine adjustment function 1600B, in semiconductor manufacturing process B, based on the inspection data of the processed wafer 130B for a specified period, the fine adjustment network unit 1610 updates the correction parameters. It should be noted that at this time, the model parameters set for the first network unit 1220_1 to the Mth network unit 1220_M are maintained in a fixed state. And in the fine adjustment network unit 1610 of the inference unit with a fine adjustment function 1600B, the correction parameters are continuously updated until the difference between the hypothetical measurement data and the inspection data becomes below a specified threshold.
[0173] Accordingly, in the fine adjustment network unit 1610, it is possible to reduce the error (the error included in the inference result) caused by the individual differences between the processing unit 120A of semiconductor manufacturing process A and the processing unit 120B of semiconductor manufacturing process B.
[0174] It should be noted that in the case of attaching the micro-adjustment function inference unit 1600B, the possibility of overfitting can be reduced as compared with the case of newly generating a hypothetical measurement model and performing optimization using the time series data group measured in the semiconductor manufacturing process B.
[0175] [Third Embodiment]
[0176] In the above first and second embodiments, although the case of applying the hypothetical measurement model generated by the hypothetical measurement device 160A to another semiconductor manufacturing process B has been described, the model applied to another semiconductor manufacturing process B is not limited to the hypothetical measurement model.
[0177] In the third embodiment, the case of replacing the hypothetical measurement devices 160A and 160B described in the first and second embodiments with the anomaly detection devices 160A and 160B, and applying the anomaly detection model generated by the anomaly detection device 160A to another semiconductor manufacturing process B will be described.
[0178] In the case of the anomaly detection device 160A, the learning unit 161A uses the time series data group as input data and the event (information indicating the presence or absence of an anomaly) as correct answer data, and makes the anomaly detection model (inference unit 162A) perform machine learning. It is assumed that the anomaly detection model (inference unit 162A) has the same configuration as the hypothetical measurement model (inference unit 162A), and only the learning data used for machine learning is different.
[0179] It should be noted that in the case of the anomaly detection device 160A, among the time series data acquisition devices 140A_1 to 140A_n for outputting the time series data group used for machine learning, for example, it includes: an emission spectroscopic analysis device for outputting the time series data group, that is, OES (Optical Emission Spectrometry) data; a process data acquisition device for outputting the sequence data group, that is, process data such as temperature data and pressure data; and a high-frequency power supply device for plasma for outputting the time series data, that is, RF data.
[0180] In addition, in the case of the anomaly detection device 160B (inference device), the micro-adjustment function inference unit 1600B inputs the time series data group and infers the information indicating the presence or absence of an anomaly.
[0181] Note that, in the case of the anomaly detection device 160B, the time series data acquisition devices 140A_1 to 140A_n for outputting the time series data group for inference include, for example: an emission spectroscopic analysis device for outputting a time series data group, i.e., OES (Optical Emission Spectrometry) data; a process data acquisition device for outputting a time series data group, i.e., process data such as temperature data and pressure data; and a high-frequency power supply device for plasma for outputting time series data, i.e., RF data.
[0182] <Summary>
[0183] As described above, it is clearly understood that the anomaly detection device 160A acquires a time series data group (OES data, process data, RF data) measured in association with the processing of an object in a specified processing unit of a manufacturing process, and causes each network unit to perform machine learning in such a manner that the synthesis result of each output data output by each network unit through processing the acquired time series data group using a plurality of network units approaches an event (information indicating the presence or absence of an anomaly) generated in association with the processing of the object.
[0184] In this way, by processing the time series data group using a plurality of network units, multi-faceted analysis can be performed. As a result, in the anomaly detection device 160A, an anomaly detection model for achieving highly accurate inference can be generated.
[0185] In addition, the anomaly detection device 160B (inference device) uses a plurality of network units included in the generated anomaly detection model to process a time series data group (OES data, process data, RF data) measured in association with the processing of an object in a specified processing unit of another manufacturing process, and outputs each output data. And it infers information indicating the presence or absence of an anomaly by synthesizing the outputted each output data after finely adjusting it using a correction parameter. In addition, it updates the correction parameter based on the error included in the inferred information indicating the presence or absence of an anomaly.
[0186] In this way, in a specified processing unit of a manufacturing process, when applying the anomaly detection model generated using the time series data group to another manufacturing process, in the anomaly detection device 160B, a function of finely adjusting each output data output by a plurality of network units is added.
[0187] Thereby, when applying the anomaly detection model to another manufacturing process, the error (error included in the inference result) caused by the individual differences between processes can be reduced. That is, according to the third embodiment, regardless of the application object, an inference device, an inference method, and an inference program capable of performing highly accurate inference can be provided.
[0188] [Other Embodiments]
[0189] In the above-described first and second embodiments, as a method for fine-tuning each output data, the case of using individual sensitivity and a correction matrix, or a network unit for fine-tuning has been described. However, the method for fine-tuning each output data is not limited thereto. For example, a generalized linear mixed model, Gaussian process regression analysis, a Kalman filter, etc. can be used.
[0190] In addition, in the above-described third embodiment, it has been described that the anomaly detection device acquires OES data, process data, and RF data output from the emission spectroscopic analysis device, the process data acquisition device, and the high-frequency power supply device for plasma along with the processing of the object. However, the combination of data acquired by the anomaly detection device is not limited thereto, and it can acquire any one of the data, or a combination of any two of the data.
[0191] In addition, in the above-described embodiments, it has been described that the inference units 162B and 1600B with fine-tuning functions have the first to Mth network units 1220_1 to 1220_M. However, the inference units 162B and 1600B with fine-tuning functions do not need to have all of the first to Mth network units 1220_1 to 1220_M, and are set to have at least any two or more network units.
[0192] In addition, in the above-described embodiments, it has been described that the machine learning algorithm of each network unit of the learning unit 161A is configured based on a convolutional neural network. However, it is not limited thereto, and the machine learning algorithm of each network unit of the learning unit 161A is not limited to a convolutional neural network, and can be configured based on other machine learning algorithms.
[0193] In addition, in the above-described embodiments, it has been described that the virtual measurement device or the anomaly detection device 160A functions as the learning unit 161A and the inference unit 162A. However, the device that functions as the learning unit 161A and the device that functions as the inference unit 162A are not integrated with each other, and can also be separately configured. That is, the virtual measurement device or the anomaly detection device 160A can function as the learning unit 161A without the inference unit 162A, or can function as the inference unit 162A without the learning unit 161A.
[0194] In addition, in the above-described embodiments, it has been described that the virtual measurement device (or anomaly detection device) with a fine-tuning function added to the virtual measurement model (or anomaly detection model) generated in the system 100A is applied to the system 100B. However, the application target of the virtual measurement device (or anomaly detection device) with a fine-tuning function added is not limited to another system, and can also be the self-system.
[0195] For example, in the case of changing a part of the process plan, etc., when the degree of change is small, it can be applied by adding a fine-tuning function to the hypothetical measurement model (or anomaly detection model) generated by the system itself.
[0196] Alternatively, in the device within the system itself, it can be applied when a maintenance operation such as component replacement is performed, when the environment within the device changes due to component wear and tear of the device within the system itself, etc., when the accuracy of the hypothetical measurement model (or anomaly detection model) generated by the system itself decreases.
[0197] It should be noted that the present invention is not limited to the configurations shown here, such as combinations of the configurations and other elements cited in the above embodiments. Regarding these points, changes can be made within the scope not exceeding the gist of the present invention, and can be appropriately determined according to its application mode.
[0198] This application claims priority based on Japanese Patent Application No. 2019-217439 filed on November 29, 2019, and the entire contents of this Japanese patent application are incorporated herein by reference.
[0199] Explanation of Reference Numerals
[0200] 100A, 100B: Systems
[0201] 110A, 110B: Wafers before processing
[0202] 120A, 120B: Processing units
[0203] 130A, 130B: Wafers after processing
[0204] 140A_1 to 140A_n: Time series data acquisition devices
[0205] 140B_1 to 140B_n: Time series data acquisition devices
[0206] 150A, 150B: Inspection data acquisition devices
[0207] 160A, 160B: Hypothetical measurement devices
[0208] 161A: Learning unit
[0209] 162A: Inference unit
[0210] 162B: Inference unit with fine-tuning function
[0211] 200: Semiconductor manufacturing device
[0212] 610: Branch
[0213] 620_1: First Network Department
[0214] 620_11~620_1N: First Layer to Nth Layer
[0215] 620_2: Second Network Department
[0216] 620_21~620_2N: First Layer to Nth Layer
[0217] 620_M: Mth Network Department
[0218] 620_M1~620_MN: First Layer to Nth Layer
[0219] 630: Connection Department
[0220] 640: Comparison Department
[0221] 1001, 1011: Standardization Department
[0222] 1004, 1014: Pooling Department
[0223] 1210: Branch Department
[0224] 1220_1: First Network Department
[0225] 1220_11~1220_1N: First Layer to Nth Layer
[0226] 1220_2: Second Network Department
[0227] 1220_21~1220_2N: First Layer to Nth Layer
[0228] 1220_M: Mth Network Department
[0229] 1220_M1~1220_MN: First Layer to Nth Layer
[0230] 1240: Connection Department
[0231] 1410: Connection Department
[0232] 1420: Individual Adjustment Department
[0233] 1430: Fine Tuning Department
[0234] 1440: Comparison Department
[0235] 1600B: Inference Department with Fine Tuning Function
[0236] 1610: Fine Tuning Network Department
Claims
1. An inference device having: An acquisition unit that acquires a time-series data group measured in association with the processing of an object in a prescribed processing unit of a manufacturing process; A plurality of machine-learned network units and a machine-learned connection unit that are generated by a learning unit including the plurality of network units for processing the time-series data group grouped in advance according to the processing to be performed by the corresponding network unit, and the connection unit for synthesizing each output data output by processing using the plurality of network units, and causing the plurality of network units and the connection unit to perform machine learning so that the synthesis result output by the connection unit approaches the inspection data of the product when processing the object; and An adjustment unit that adjusts each output data output without synthesizing through the machine-learned connection unit after processing the time-series data group grouped according to the processing to be performed by the corresponding network unit acquired by using the plurality of machine-learned network units, and synthesizes the adjusted output data to output an inference result, The adjustment unit adjusts each output data using a correction parameter corresponding to an error included in the inference result.
2. The inference device according to claim 1, wherein the adjustment unit updates the correction parameter in a state where the model parameters of the plurality of machine-learned network units are fixed so as to reduce the error included in the inference result.
3. The inference device according to claim 2, wherein the acquisition unit processes the acquired time-series data group based on a first reference and a second reference respectively to generate a first time-series data group and a second time-series data group, and processes the first time-series data group and the second time-series data group using the plurality of machine-learned network units.
4. The inference device according to claim 2, wherein the acquisition unit groups the acquired time-series data group according to data type or time range, and processes each group after grouping using the plurality of machine-learned network units.
5. The inference device according to claim 2, wherein the acquired time-series data group is processed using the plurality of machine-learned network units each including a normalization unit that performs normalization by different methods.
6. The inference device according to claim 2, wherein the acquisition unit divides the acquired time-series data group into a first time-series data group measured in association with the processing of the object in a first processing space of the prescribed processing unit, and a second time-series data group measured in association with the processing of the object in a second processing space, and processes the first time-series data group and the second time-series data group using the plurality of machine-learned network units.
7. The inference device according to claim 1, wherein the time-series data group is data measured in association with the processing in a substrate processing device.
8. A method of inference, comprising the following steps: An acquisition step of acquiring a time series data group measured in association with the processing of an object in a prescribed processing unit of a manufacturing process; A processing step of processing the acquired time series data group using a plurality of network units that have completed machine learning and a connection unit that has completed machine learning. The plurality of network units that have completed machine learning and the connection unit that has completed machine learning are generated by a learning unit including the plurality of network units for processing the time series data group grouped in advance according to the processing to be performed by the corresponding network unit, and the connection unit for synthesizing each output data output by using the plurality of network units, so that the synthesis result output by the connection unit is close to the inspection data of the product when processing the object, and the plurality of network units and the connection unit are made to perform machine learning; and An adjustment step of adjusting each output data output without synthesizing through the connection unit that has completed machine learning after processing the time series data group grouped in advance according to the processing to be performed by the corresponding network unit using the plurality of network units that have completed machine learning, and synthesizing the adjusted output data to output an inference result, The adjustment step adjusts each output data by using a correction parameter corresponding to the error included in the inference result.
9. A computer-readable storage medium storing an inference program for causing a processor to execute the following steps: An acquisition step of acquiring a time series data group measured in association with the processing of an object in a prescribed processing unit of a manufacturing process; A processing step of processing the acquired time series data group using a plurality of network units that have completed machine learning and a connection unit that has completed machine learning. The plurality of network units that have completed machine learning and the connection unit that has completed machine learning are generated by a learning unit including the plurality of network units for processing the time series data group grouped in advance according to the processing to be performed by the corresponding network unit, and the connection unit for synthesizing each output data output by using the plurality of network units, so that the synthesis result output by the connection unit is close to the inspection data of the product when processing the object, and the plurality of network units and the connection unit are made to perform machine learning; and An adjustment step of adjusting each output data output without synthesizing through the connection unit that has completed machine learning after processing the time series data group grouped in advance according to the processing to be performed by the corresponding network unit using the plurality of network units that have completed machine learning, and synthesizing the adjusted output data to output an inference result, The adjustment step adjusts each output data by using a correction parameter corresponding to the error included in the inference result.
Citation Information
Patent Citations
Abnormality detector
JP2006163517A
Water treatment device
JP2019217439A
Information processor, information processing method and program
CN101622688A
KR20190046018A