Program product, data generation method, learning model generation method, information processing device, and computer-readable recording medium

By acquiring motion images of the substrate processing and generating enhanced data, and using machine learning to generate a learning model, the auxiliary problem of substrate processing motion analysis is solved, improving processing efficiency and controllability.

CN121032886APending Publication Date: 2025-11-28TOKYO ELECTRON LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510631242.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-27
Filing Date
2025-05-16
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies lack effective auxiliary means for analyzing image-based substrate processing actions.

Method used

By acquiring motion images during substrate processing, enhanced data is generated and machine learning is used to create a learning model to assist in the analysis of the substrate processing device's actions.

Benefits of technology

It enables effective analysis and status determination of substrate processing actions, improving processing efficiency and controllability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032886A_ABST
    Figure CN121032886A_ABST
Patent Text Reader

Abstract

The present disclosure provides a program product, a data generation method, a learning model generation method, an information processing device, and a computer-readable recording medium which are expected to assist in operation analysis or the like of image-based substrate processing. A computer program according to the present embodiment causes a computer to execute the following processes: acquiring a moving image obtained by photographing a substrate process; and generating enhancement data on the basis of a frame image at a first point in time included in the acquired moving image and a frame image at a second point in time following the first point in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a program product, a data generation method, a learning model generation method, an information processing device, and a computer-readable recording medium. Background Technology

[0002] Patent Document 1 discloses a substrate processing method comprising the following steps: a holding step, wherein the substrate is moved into and held inside a chamber; a supply step, wherein fluid is supplied to the substrate inside the chamber; an image capture step, wherein a camera sequentially captures images of the interior of the chamber to acquire image data; a condition setting step, wherein a monitored object is determined from a plurality of monitored object candidates within the chamber, and the image conditions are changed based on the monitored object; and a monitoring step, wherein monitoring processing is performed on the monitored object based on the image data having image conditions corresponding to the monitored object.

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2021-190511 Summary of the Invention

[0006] The problem the invention aims to solve

[0007] This disclosure provides a computer program, a data generation method, a learning model generation method, and an information processing apparatus that are expected to assist in motion analysis and other tasks related to image-based substrate processing.

[0008] Solution for solving the problem

[0009] One embodiment of the computer program causes a computer to perform the following processes: acquiring a moving image obtained by photographing a substrate; and generating augmented data based on a frame image of a first time point contained in the acquired moving image and a frame image of a second time point later than the first time point.

[0010] The effects of the invention

[0011] According to this disclosure, it is expected to assist in motion analysis and other aspects of image-based substrate processing. Attached Figure Description

[0012] Figure 1 This is a schematic diagram illustrating a structural example of the substrate processing apparatus according to this embodiment.

[0013] Figure 2 This is a schematic diagram used to illustrate the general outline of the information processing system involved in this embodiment.

[0014] Figure 3This is a block diagram illustrating a structural example of the information processing apparatus according to this embodiment.

[0015] Figure 4 This is a schematic diagram illustrating the data enhancement performed by the information processing apparatus according to this embodiment.

[0016] Figure 5 This is a schematic diagram showing an example of a tag information input screen displayed by the information processing device according to this embodiment.

[0017] Figure 6 This is a flowchart illustrating an example of the learning data generation process performed by the information processing apparatus according to this embodiment.

[0018] Figure 7 This is a schematic diagram illustrating a structural example of a learning model generated by an information processing device.

[0019] Figure 8 This is a flowchart illustrating an example of the learning model generation process performed by the information processing apparatus in this embodiment.

[0020] Figure 9 This is a schematic diagram illustrating a structural example of a learning model generated by the information processing apparatus according to Embodiment 2. Detailed Implementation

[0021] Hereinafter, specific examples of the information processing system according to embodiments of the present disclosure will be described with reference to the accompanying drawings. Furthermore, the present disclosure is not limited to these examples, as illustrated by the claims, which are intended to include all modifications within the meaning and scope equivalent to the claims.

[0022] [Implementation Method 1]

[0023] <System Structure>

[0024] Figure 1 This is a schematic diagram illustrating a structural example of the substrate processing apparatus 1 according to this embodiment. The substrate processing apparatus 1 according to this embodiment is an apparatus for wet etching. Wet etching is a substrate processing method that processes a substrate (e.g., a wafer with an oxide film or nitride film formed on it) into a desired shape by rotating the substrate to be processed while supplying a solution for dissolving the film onto the film. The substrate processing apparatus 1 according to this embodiment is configured to include a chamber 11, a substrate holding mechanism 12, an ejection section 13, and a recovery cup 14, etc.

[0025] Chamber 11 is a sealed reaction vessel that houses a substrate holding mechanism 12, an ejector 13, and a recovery cup 14. An FFU (Fan Filter Unit) 15 is installed at the top of chamber 11. The FFU 15 is used to create a downflow within chamber 11.

[0026] The substrate holding mechanism 12 includes a holding portion 12a, a support portion 12b, and a driving portion 12c. The holding portion 12a is, for example, disk-shaped, and horizontally holds the substrate (wafer) to be processed on the disk. The support portion 12b is connected to the central portion of the lower surface of the holding portion 12a, and is positioned vertically (in the vertical direction). Figure 1 A cylindrical member extending vertically (in the middle) horizontally supports the holding portion 12a. Furthermore, the lower end of the support portion 12b is connected to the drive portion 12c and is supported by the drive portion 12c in a rotatable manner. The drive portion 12c has a prime mover such as a motor, which rotates the support portion 12b about an axis. Thus, the substrate holding mechanism 12 can rotate the holding portion 12a supported on the support portion 12b by rotating the support portion 12b via the drive portion 12c, thereby rotating the substrate held by the holding portion 12a.

[0027] The ejection section 13 ejects a liquid such as a chemical solution or a cleaning solution onto the substrate held in the substrate holding mechanism 12. For example, dilute hydrofluoric acid can be used as the chemical solution, and pure water can be used as the cleaning solution, but the liquid ejected by the ejection section 13 is not limited to these. The ejection section 13 is connected, for example, via a tubular liquid supply path to a liquid supply source 16 located outside the chamber 11, and ejects the liquid supplied from the supply source 16 onto the substrate. In addition, the ejection section 13 is connected to a drive mechanism (not shown) and can move horizontally between the center and the periphery of the substrate. By combining the rotation of the substrate by the substrate holding mechanism 12 with the horizontal movement of the ejection section 13 by the drive mechanism, the substrate processing apparatus 1 can eject liquid from the ejection section 13 to an appropriate position on the substrate to be processed.

[0028] The recovery cup 14 is configured to surround the holding portion 12a of the substrate holding mechanism 12, capturing liquid that splashes from the substrate due to the rotation of the holding portion 12a. A drain port 14a is provided at the bottom of the recovery cup 14, through which the liquid captured by the recovery cup 14 is discharged to the outside of the chamber 11. Additionally, an exhaust port 14b is provided at the bottom of the recovery cup 14, through which gas supplied from the FFU 15 is discharged to the outside of the chamber 11.

[0029] also, Figure 1The substrate processing apparatus 1 shown has a structure with one liquid ejection section 13. The substrate processing apparatus 1 can selectively eject either a chemical solution for dissolving the substrate or a cleaning solution for cleaning the substrate by switching between these two in the supply source 16. However, the substrate processing apparatus 1 may also have a structure with multiple ejection sections 13. For example, the substrate processing apparatus 1 may have only one ejection section 13 for ejecting chemical solution and one ejection section 13 for ejecting cleaning solution.

[0030] Figure 2 This is a schematic diagram illustrating the general outline of the information processing system according to this embodiment. The information processing system according to this embodiment is configured to include the substrate processing apparatus 1, the information processing apparatus 3, and the camera 5 described above. The camera 5 has an imaging element such as a CCD (Charge Coupled Device) or CMOS (Complementary Metal Oxide Semiconductor), and is capable of capturing so-called moving images by taking dozens of shots per second. The camera 5 is, for example, installed inside the chamber 11 of the substrate processing apparatus 1, and captures images of the ejection section 13 during substrate processing. The camera 5 sends the data of the moving images obtained through capture to the information processing apparatus 3. The moving image data is, for example, data obtained by concatenating multiple still images (frame images) in a time sequence. The camera 5 may be a device included in the substrate processing apparatus 1, or it may be a device separate from the substrate processing apparatus 1.

[0031] The information processing device 3 is an apparatus for generating learning data for machine learning based on motion image data obtained from the camera 5, and for generating a learning model using machine learning with the learning data. In this embodiment, the information processing device 3 is configured as an apparatus independent of the substrate processing apparatus 1, but it is not limited thereto and may also be an integrated apparatus with the substrate processing apparatus 1. The information processing device 3 is connected to the camera 5, for example, via a communication cable, and is capable of transmitting and receiving data with the camera 5. The information processing device 3 receives motion image data obtained from the camera 5 by capturing images of the ejector 13, and stores and accumulates the received motion image data in a storage unit.

[0032] The information processing apparatus 3 of this embodiment generates learning data based on frame images contained in motion images acquired and stored from the camera 5. This learning data is used to generate a learning model for determining the state related to the ejection of liquid by the ejector 13. At this time, the information processing apparatus 3 of this embodiment augments the frame images used for machine learning by performing so-called data augmentation, which generates frame images not included in the motion images captured by the camera 5, based on the frame images. For example, the information processing apparatus 3 receives input from the user for each frame image, representing the ejection state of the ejector 13, and uses the resulting data sets, which establish a correspondence between each frame image and the tag information, as learning data for so-called supervised machine learning.

[0033] Furthermore, when the information processing device 3 generates learning data for so-called unsupervised machine learning, it is possible to omit the input of label information and the mapping of label information for each frame image. In this case, the information processing device 3 uses a dataset containing frame images contained in motion images captured by the camera 5 and frame images generated by enhancing those frame images as learning data.

[0034] Furthermore, the information processing device 3 performs machine learning processing using the generated learning data to generate a learning model. For example, when the learning data is data obtained by establishing a correspondence between frame images and label information, the information processing device 3 can generate a learning model by performing supervised machine learning. This learning model takes the frame image as input and classifies the ejection state of the ejection section 13 of the substrate processing device 1 captured in the frame image. The learning model generated by the information processing device 3 is, for example, mounted on a control device that controls the operation of the substrate processing device 1. The control device acquires motion images obtained by the camera 5 capturing images of the ejection section 13 of the substrate processing device 1. The control device inputs the frame images contained in the acquired motion images into the learning model and obtains the classification result of the ejection state output by the learning model. If the obtained classification result indicates, for example, an abnormality, the control device can perform control processing such as stopping the substrate processing of the substrate processing device 1.

[0035] Figure 3This is a block diagram illustrating a structural example of the information processing apparatus 3 according to this embodiment. The information processing apparatus 3 according to this embodiment can be implemented, for example, by installing a specified application program on a general-purpose information processing device such as a personal computer or server computer. The information processing apparatus 3 according to this embodiment is configured to include a processing unit 31, a storage unit 32, a communication unit 33, a display unit 34, and an operation unit 35. Furthermore, in this embodiment, processing is described as being performed by a single information processing apparatus 3, but the processing of the information processing apparatus 3 can also be performed by multiple devices distributed among them.

[0036] The processing unit 31 is constructed using computing devices such as a CPU (Central Processing Unit), MPU (Micro-Processing Unit), GPU (Graphics Processing Unit), or quantum processor, as well as ROM (Read Only Memory) and RAM (Random Access Memory). The processing unit 31 reads and executes the program 32a stored in the storage unit 32 to perform various processes, including generating learning data for machine learning based on motion images acquired from the camera 5, and generating a learning model using the generated learning data.

[0037] The storage unit 32 is constructed using a high-capacity storage device such as a hard disk or an SSD (Solid State Drive). The storage unit 32 stores various programs executed by the processing unit 31 and various data required for processing by the processing unit 31. In this embodiment, the storage unit 32 stores program 32a executed by the processing unit 31. Furthermore, the storage unit 32 is provided with a learning data storage unit 32b for storing the generated learning data and a model information storage unit 32c for storing information related to the generated learning model.

[0038] In this embodiment, the program (computer program, program product) 32a is provided in the form of a recording medium 99 such as a memory card or optical disc, and the information processing device 3 reads the program 32a from the recording medium 99 and stores it in the storage unit 32. However, the program 32a may also be written into the storage unit 32 during the manufacturing stage of the information processing device 3, for example. Alternatively, the program 32a may be obtained by the information processing device 3 via communication from a remote server device or the like. For example, the program 32a recorded on the recording medium 99 may be read by a writing device and written into the storage unit 32 of the information processing device 3. The program 32a may be provided in the form of distribution via a network or in the form of being recorded on the recording medium 99.

[0039] The learning data storage unit 32b stores learning data generated by the information processing device 3 based on motion images captured by the camera 5. The learning data may be, for example, data obtained by establishing a correspondence between frame images and tag information representing the ejection state of the ejection unit 13. The model information storage unit 32c stores information related to the learning model obtained through machine learning. This information may include, for example, information representing the structure of the learning model and the values ​​of internal parameters determined through machine learning.

[0040] Furthermore, in this embodiment, it is assumed that both the generation of learning data and the generation of the learning model are performed by the information processing device 3, but this is not a limitation. The generation of learning data and the generation of the learning model can also be performed by different devices. The device that generates the learning data sends the learning data to the device that generates the learning model, and the device that generates the learning model receives the learning data and performs machine learning processing.

[0041] The communication unit 33 transmits and receives data with the camera 5, for example, via a wired or wireless network N. In this embodiment, the communication unit 33 receives motion image data transmitted from the camera 5 and provides it to the processing unit 31. Furthermore, the communication unit 33 can also send commands to the camera 5, such as commands to control the operation of the camera 5, based on information provided by the processing unit 31.

[0042] The display unit 34 is constructed using a liquid crystal display or the like, and displays various images and characters based on the processing of the processing unit 31. The display unit 34 displays, for example, images (moving images or frame images) captured by the camera 5, screens for receiving input from the user regarding tag information related to the ejection state of the ejection unit 13, or the progress status of machine learning in generating the learning model, and other various information.

[0043] The operation unit 35 accepts user operations and notifies the processing unit 31 of the accepted operations. For example, the operation unit 35 accepts user operations via input devices such as mechanical buttons or a touch panel provided on the surface of the display unit 34. Alternatively, the operation unit 35 may also be an input device such as a mouse and keyboard, which can be detachable from the information processing device 3.

[0044] Furthermore, the storage unit 32 may also be an external storage device connected to the information processing device 3. Additionally, the information processing device 3 may be configured as a multi-computer system including multiple computers, or it may be a virtual machine constructed virtually through software. Furthermore, the information processing device 3 is not limited to the above-described structure; for example, it may include a reading unit for reading information stored on a removable storage medium, and it may not include, for example, a display unit 34 and an operation unit 35.

[0045] Furthermore, in the information processing apparatus 3 according to this embodiment, by reading and executing the program 32a stored in the storage unit 32, the image acquisition unit 31a, data enhancement unit 31b, learning data generation unit 31c, learning processing unit 31d, and display processing unit 31e, which are software functional units, are implemented in the processing unit 31. In addition, in this figure, the functional units of the processing unit 31 that perform processing related to the generation of learning data and the generation of the learning model are shown, and the illustrations of functional units related to other processing are omitted.

[0046] The image acquisition unit 31a performs the following processing: by communicating with the camera 5 via the communication unit 33, it acquires image data of the ejection section 13 of the substrate processing apparatus 1 captured by the camera 5. In this embodiment, the camera 5 is a camera that captures moving images by taking approximately several dozen shots per second. The image data acquired by the image acquisition unit 31a can be in the format of moving images or in the format of still images (frame images) included in the moving images. The image acquisition unit 31a repeatedly acquires images from the camera 5 and stores the images acquired from the camera 5 in the learning data storage unit 32b. Furthermore, by repeatedly acquiring images from the camera 5 during substrate processing, the image acquisition unit 31a can obtain frame images in a time sequence obtained by capturing images of the ejection section 13.

[0047] The data enhancement unit 31b performs the following processing: based on the motion image (multiple frame images in a time series) obtained by the camera 5, it increases the number of frame images by performing data enhancement. In this embodiment, the data enhancement unit 31b performs data enhancement by extracting a frame image at a certain time point and a frame image at the next time point from the multiple frame images in the time series contained in the motion image, and generating a frame image corresponding to the time point between these two time points. The data enhancement unit 31b stores the generated frame image in the learning data storage unit 32b.

[0048] The learning data generation unit 31c performs the following processing: It generates learning data for machine learning based on frame images contained in the motion images acquired by the image acquisition unit 31a and frame images generated by the data enhancement unit 31b through data enhancement. For example, the learning data generation unit 31c processes input from the user, including label information contained in the learning data for supervised machine learning. For example, the learning data generation unit 31c displays frame images on the display unit 34 and accepts input of label information representing the ejection state of the ejector 13 captured in the displayed frame images based on the user's operation on the operation unit 35. The learning data generation unit 31c establishes a correspondence between the displayed frame images and the label information input by the user and stores it in the learning data storage unit 32b. The learning data generation unit 31c accepts input of label information for multiple frame images of the motion images acquired by the image acquisition unit 31a and frame images generated by the data enhancement unit 31b through data enhancement, and uses the dataset composed of multiple sets of frame images and label information as learning data. Furthermore, when generating learning data for unsupervised learning that does not contain label information, the learning data generation unit 31c does not accept the input of label information. Instead, it simply combines the frame images contained in the motion images acquired by the image acquisition unit 31a and the frame images generated by the data augmentation unit 31b through data augmentation into an integrated dataset, which is then used as the learning data.

[0049] The learning processing unit 31d performs the following processing: by performing machine learning processing using the learning data stored in the learning data storage unit 32b, a learning model is generated for prediction based on frame images. For example, the learning processing unit 31d performs supervised machine learning using learning data obtained by establishing a correspondence between frame images and label information representing ejection states, to generate a learning model that takes frame images as input and outputs information representing the ejection state of the ejector 13 captured in the frame image. The learning model can employ structures such as CNN (Convolutional Neural Network) or DNN (Deep Neural Network), but is not limited to these; it can be any structure. The learning processing unit 31d performs machine learning processing using existing methods such as stochastic gradient descent or backpropagation. The structure of the learning model and the machine learning method for generating the learning model are existing technologies, therefore detailed descriptions are omitted in this embodiment.

[0050] The display processing unit 31e performs processing to display various characters and images on the display unit 34. In this embodiment, the display processing unit 31e displays, for example, a motion image acquired by the image acquisition unit 31a or a frame image contained in that motion image on the display unit 34. Furthermore, to accept input regarding the ejection state of the ejector 13 captured in the displayed frame image, the display processing unit 31e displays, for example, multiple selection items indicating the ejection state on the display unit 34. When a user selects one of the multiple selection items displayed on the display unit 34 as the appropriate selection item for the ejection state of the ejector 13 captured in the displayed frame image, the information processing device 3 accepts this operation via the operation unit 35, thereby enabling input of tag information regarding the ejection state of the frame image. Additionally, the display processing unit 31e displays, for example, information such as the number of repetitions of learning or the evaluation value of the learning model as the progress status of machine learning processing on the display unit 34.

[0051] <Learning to Generate and Process Data>

[0052] The information processing apparatus 3 according to this embodiment performs the following processing: based on the frame images contained in the motion image of the ejection section 13 of the substrate processing apparatus 1 obtained by the camera 5, it generates learning data for machine learning, which is used to generate a learning model. At this time, the information processing apparatus 3 can amplify the learning data by performing data augmentation processing based on multiple frame images contained in the motion image.

[0053] Figure 4This is a schematic diagram illustrating the data enhancement performed by the information processing apparatus 3 according to this embodiment. In this diagram, the frame images contained in the moving images captured by the camera 5 are shown in chronological order, such as frame 1, frame 2, and frame 3, with names such as "frame + integer value". In addition, in this diagram, the new frame images generated by the information processing apparatus 3 through data enhancement of these frame images are shown with names such as enhanced frame 1.9 and enhanced frame 2.5, with names such as "enhanced frame + decimal value".

[0054] The information processing device 3 extracts two frame images from the moving image: a frame image at a certain time point and the next frame image in the time sequence. The information processing device 3 calculates the average value of corresponding pixels in the two extracted frame images and generates a frame image with the calculated average value as its pixel value as the enhanced frame image. In the example shown in this figure, enhanced frame 2.5 is generated based on the average value of frame 2 and frame 3.

[0055] Furthermore, the information processing device 3 can generate an enhanced frame image by calculating a weighted average instead of a simple average. In the example shown in this figure, enhanced frame 1.9 is generated based on the average calculated by adding a weight of 1 to 9 to frames 1 and 2. The information processing device 3 can generate an enhanced frame image by, for example, randomly weighting two frame images extracted from a moving image and calculating a weighted average. Additionally, when generating an enhanced frame image from two frame images, the information processing device 3 can use various methods, such as linear interpolation or cubic interpolation, instead of calculating an average or a weighted average.

[0056] The information processing device 3 integrates the frame images contained in the original motion image and the enhanced frame images generated by enhancing those frame images, and stores them as a dataset in the learning data storage unit 32b. This dataset can be used as learning data for unsupervised machine learning, which generates learning models such as autoencoders. Furthermore, the information processing device 3 of this embodiment generates a dataset by attaching label information to each frame image (the original frame image and the enhanced frame image) to form the correct predicted value, which is used as learning data for generating a learning model through supervised machine learning. This learning model is a model that performs image-based predictions such as classification or regression. The information processing device 3 obtains the label information attached to each frame image through user input.

[0057] Figure 5This is a schematic diagram illustrating an example of a tag information input screen displayed by the information processing apparatus 3 according to this embodiment. In this example, it is assumed that the user performs an annotation task, which is the following task: the user selects a tag for the state of droplets falling from the ejection section 13 onto the substrate to be processed, or a tag for the state of droplets not falling, based on a frame image obtained by taking a picture of the ejection section 13 of the substrate processing apparatus 1. However, this is just one example, and arbitrary tag information may also be attached to the frame image.

[0058] In this embodiment, the information processing device 3 appropriately extracts one frame image from multiple frame images (original frame images and enhanced frame images) and displays it on the left side of the label information input screen. On the right side of the screen, a message string "Please select spray state," a button labeled "with droplets," and a button labeled "without droplets" are displayed vertically. The user can input label information for the frame image by clicking or touching either the "with droplets" button or the "without droplets" button using an operation unit 35, such as a mouse or touch panel. The information processing device 3 accepts the user's selection of label information based on which button is pressed and stores the selected label information corresponding to the frame image displayed on the label information input screen in the learning data storage unit 32b. The information processing device 3 accepts label information input from the user sequentially for multiple frame images and stores the accepted label information in the learning data storage unit 32b, thereby enabling the dataset composed of frame images and label information to be used as learning data.

[0059] Figure 6This is a flowchart illustrating an example of the learning data generation process performed by the information processing apparatus 3 according to this embodiment. In this embodiment, the image acquisition unit 31a of the processing unit 31 of the information processing apparatus 3 communicates with the camera 5 via the communication unit 33 to acquire data of moving images captured by the camera 5 during substrate processing by the substrate processing apparatus 1 (step S1). The image acquisition unit 31a stores the data of the moving images acquired in step S1 (including the frame images) in the learning data storage unit 32b of the storage unit 32 (step S2). The image acquisition unit 31a determines whether the substrate processing performed by the substrate processing apparatus 1 has ended (step S3). If the substrate processing has not ended (S3: "No"), the image acquisition unit 31a returns the processing to step S1 and repeats the acquisition and storage of moving images until the substrate processing ends. Furthermore, whether the substrate processing has ended can be determined, for example, by enabling the information processing apparatus 3 to communicate with the substrate processing apparatus 1, or, for example, based on the moving images captured by the camera 5.

[0060] When the substrate processing is completed (S3: "Yes"), the data enhancement unit 31b of the processing unit 31 acquires a frame image at a certain time point (e.g., the first frame image) from the multiple frame images included in the motion image data stored in step S2 (step S4). Additionally, the data enhancement unit 31b acquires the frame image at the next time point in the time series of the frame image acquired in step S4 (step S5). For the frame image at the certain time point and the frame image at the next time point, the data enhancement unit 31b generates a frame image between the certain time point and the next time point, for example, by calculating the average value of the corresponding pixels (step S6). The data enhancement unit 31b stores the enhanced frame image generated in step S6 in the learning data storage unit 32b (step S7). The data enhancement unit 31b determines whether data enhancement has ended for all the frame images acquired and stored in steps S1 and S2 (step S8). If data enhancement has not ended for all frame images (S8: "No"), the data enhancement unit 31b returns the processing to step S5, further acquires the frame image at the next time point, and repeats the same processing.

[0061] If data augmentation has been completed for all frame images (S8: "Yes"), the learning data generation unit 31c of the processing unit 31, for example, will... Figure 5The label information input screen shown is displayed on the display unit 34, and the label information input is accepted for all frame images contained in the motion image and all frame images generated by data augmentation (step S9). The learning data generation unit 31c establishes a correspondence between the label information accepted in step S9 and each frame image, and stores it as learning data in the usage data storage unit 32b (step S10), thus ending the learning data generation process.

[0062] Furthermore, in this embodiment, the information processing device 3 uses both the frame image contained in the motion picture and the enhanced frame image generated by enhancing the frame image as learning data, but it is not limited to this. The information processing device 3 may also use only the enhanced frame image generated by data enhancement processing of the original frame image as learning data.

[0063] <Generation and Processing of Learning Models>

[0064] The information processing apparatus 3 in this embodiment performs the following processing: using learning data generated based on motion images captured by the camera 5, it generates a learning model for making predictions related to substrate processing. Figure 7 This is a schematic diagram illustrating a structural example of the learning model generated by the information processing device 3. In this example, the learning model 101 generated by the information processing device 3 is a learning model that takes an image obtained by photographing the ejection section 13 of the substrate processing device 1 as input and classifies the ejection state of the ejection section 13 as either "with droplets" or "without droplets". This learning model 101 may employ a structure such as CNN. Based on the input image obtained by photographing the ejection section 13, the learning model 101 outputs two values: a value indicating the probability that the ejection state of the ejection section 13 is "with droplets" and a value indicating the probability that the ejection state is "without droplets". The ejection state with the larger value becomes the classification result.

[0065] The information processing device 3 can use the learning data generated in the learning data generation process described above, such as the learning data obtained by mapping frame images obtained by photographing the ejection section 13 of the substrate processing device 1 with label information of "with droplets" or "without droplets", to generate the illustrated learning model 101 by performing supervised machine learning. In supervised machine learning, the information processing device 3 generates the learning model 101 by inputting frame images of the learning data into the learning model 101 and updating the internal parameters of the learning model 101 in a way that makes the information output by the learning model 101 based on the frame images approximate the label information of the learning data.

[0066] Figure 8This is a flowchart illustrating an example of the learning model generation process performed by the information processing device 3 in this embodiment. Furthermore, this figure shows the generation process through supervised machine learning. Figure 7 The process is as shown in the case of the learning model 101. In this embodiment, the learning processing unit 31d of the processing unit 31 of the information processing device 3 sets initial values ​​of the internal parameters of the learning model 101 constructed with a predetermined CNN, for example (step S21). The initial values ​​of the internal parameters may be predetermined values, or they may be determined randomly, or they may be values ​​of the internal parameters of the learning model 101 obtained by some degree of previous learning.

[0067] The learning processing unit 31d acquires one piece of learning data stored in the learning data storage unit 32b (step S22). The learning processing unit 31d inputs the frame image contained in the learning data acquired in step S22 into the learning model 101 (step S23). The learning processing unit 31d acquires the classification result of "with droplets" or "without droplets" output by the learning model 101 based on the input in step S23 (step S24). The learning processing unit 31d calculates the error of the classification result acquired in step S24 based on the label information contained in the learning data acquired in step S22 (step S25). Based on the error calculated in step S25, the learning processing unit 31d updates the internal parameters of the learning model 101, for example, using the error backpropagation method (step S26).

[0068] The learning processing unit 31d determines whether the conditions for ending the machine learning process have been met, such as the number of repetitions exceeding a threshold or the prediction accuracy of the learning model 101 reaching a target value (step S27). If the end conditions are not met (S27: "No"), the learning processing unit 31d returns the process to step S22 and repeats the above process. If the end conditions are met (S27: "Yes"), the learning processing unit 31d stores the internal parameters of the learning model 101 in the model information storage unit 32c (step S28) and ends the learning model generation process.

[0069] Furthermore, in this embodiment, the information output by the learning model 101 is set to two ejection states: "with droplets" or "without droplets," but it is not limited to these. The learning model 101 may also be a structure that outputs information representing three or more ejection states.

[0070] For example, the learning model 101 can be configured to output information representing four ejection states: "with liquid column," "liquid column broken and falling," "with droplets," and "no liquid." The "with liquid column" ejection state is a state where liquid is continuously ejected from the ejector 13, and the columnar liquid connects the lower end of the ejector 13 to the upper surface of the substrate. The "liquid column broken and falling" ejection state is the state when the ejection of liquid from the ejector 13 has just stopped, and there is space between the lower end of the ejector 13 and the upper end of the liquid column, with the liquid column standing upright on the upper surface of the substrate. The "with droplets" ejection state is a state where there is no liquid column between the lower end of the ejector 13 and the upper surface of the substrate, but one or more spherical liquid droplets are present. The "no liquid" ejection state is a state where there is neither a liquid column nor droplets between the lower end of the ejector 13 and the upper surface of the substrate.

[0071] Alternatively, the learning model 101 may not be a classification model that categorizes the ejection state, but rather a regression model that predicts certain values ​​related to ejection. For example, the learning model 101 could be configured to predict and output values ​​such as the size or quantity of droplets captured in a frame image.

[0072] Summary

[0073] In the information processing system according to this embodiment with the above structure, the information processing device 3 acquires a motion image obtained by the camera 5 capturing the substrate processing performed by the substrate processing device 1. The information processing device 3 generates enhanced data (enhanced frame image) based on the frame image at a first time point and the frame image at a second time point later than the first time point contained in the acquired motion image. Thus, in the information processing system according to this embodiment, by utilizing the frame image obtained from the motion image captured by the camera 5 and the enhanced frame image generated by data enhancement, frame images used for machine learning, etc., can be augmented, and it is expected to assist in motion analysis and other image-based substrate processing.

[0074] Furthermore, in the information processing system according to this embodiment, the information processing device 3 acquires a motion image obtained by photographing the ejection section 13 of the substrate processing device 1, which ejects liquid toward the substrate being processed, to generate augmented data. Therefore, it is expected that the information processing system according to this embodiment will generate augmented data for machine learning, which will be used to generate a learning model 101 for determining the ejection state of the ejection section 13 of the substrate processing device 1.

[0075] Furthermore, in the information processing system of this embodiment, the information processing device 3 generates enhanced data based on the average or weighted average of the frame images at the first time point and the frame images at the second time point. Therefore, it is expected that the information processing system of this embodiment can generate enhanced data through simple calculations.

[0076] Furthermore, in the information processing system according to this embodiment, the information processing device 3 generates a learning model by performing machine learning using learning data, which outputs information related to the state of the substrate processing based on frame images contained in motion images obtained by capturing images of the substrate processing device 1. This learning data includes the generated augmented data. Therefore, it is expected that the information processing system according to this embodiment can use the generated learning model to determine the state of the substrate processing and implement control of the substrate processing corresponding to the determined state.

[0077] [Implementation Method 2]

[0078] In Embodiment 1 described above, it is assumed that the learning data generated by the information processing device 3 is used to generate a learning model 101 that accepts images as input. In contrast, the learning data generated by the information processing device 3 in Embodiment 2 is learning data used to generate a learning model that accepts data based on features generated from images as input.

[0079] Figure 9 This is a schematic diagram illustrating a structural example of the learning model generated by the information processing apparatus 3 according to Embodiment 2. The learning model 102 according to Embodiment 2 is a learning model that takes the feature values ​​of an image obtained by photographing the ejection section 13 of the substrate processing apparatus 1 as input, and classifies the ejection state of the ejection section 13 as either "with droplets" or "without droplets". The information processing apparatus 3 according to Embodiment 2 includes a feature extraction model 103 that extracts feature values ​​from frame images contained in moving images obtained by the camera 5. The feature extraction model 103 is a learning model generated in advance by the information processing apparatus 3 or other apparatus through machine learning, and its internal parameters and other information are stored in the model information storage unit 32c. Furthermore, converting images into feature values ​​is a prior art technique; therefore, details regarding the structure and learning method of the learning model are omitted.

[0080] The information processing apparatus 3 of Embodiment 2 generates data that establishes a correspondence between the feature quantities extracted from the frame images contained in the motion image captured by the feature quantity extraction model 103 from the camera 5 and the tag information indicating whether the ejection state of the ejection part 13 captured in the original frame image is "with droplets" or "without droplets". This data is used for processing... Figure 9The learning data used for machine learning in the learning model 102 shown is as follows. The information processing device 3 acquires motion images captured by the camera 5, and acquires frame images contained in the acquired motion images. It then inputs the acquired frame images into the feature extraction model 103 to obtain the feature values ​​output by the feature extraction model 103. Thus, the information processing device 3 can obtain multiple corresponding feature values ​​for multiple frame images contained in the motion image.

[0081] Furthermore, the information processing apparatus 3 according to Embodiment 2 amplifies the feature quantities used as learning data by performing data augmentation on multiple feature quantities extracted from frame images contained in the motion image. For the data augmentation of the feature quantities, similar to the data augmentation of frame images performed in Embodiment 1, methods such as averaging, weighted averaging, linear interpolation, or cubic interpolation can be used. For example, when the feature extraction model 103 converts the image into a vector of N (N is a natural number) dimensional feature quantities, the information processing apparatus 3 can obtain the vector of N dimensional feature quantities obtained by data augmentation by calculating the average of N corresponding values ​​of the feature quantity at a certain time point and the N values ​​of the feature quantity at the next time point.

[0082] The information processing device 3 according to embodiment 2 is capable of processing data with... Figure 5 The same screen as the label information input screen is displayed on the display unit 34. It accepts label information input for frame images from the user and establishes a correspondence between the feature values ​​extracted from the displayed frame images and the label information obtained from the input, using this as learning data. Furthermore, in order to accept input for label information regarding feature values ​​generated through data augmentation, the information processing device 3 can also calculate the average of two frame images corresponding to the two feature values ​​that form the basis for generating the feature values, generate an image for display, and display the generated image on the label information input screen.

[0083] Furthermore, when generating learning data for unsupervised machine learning, the information processing device 3 can use the data obtained by integrating the feature quantities extracted from the frame images contained in the motion image and the feature quantities generated by data augmentation based on the feature quantities as an integrated dataset as learning data.

[0084] In the information processing system according to Embodiment 2 of the above structure, the information processing device 3 converts the frame images contained in the motion image into feature quantity data, and generates enhanced data based on the feature quantity data of the frame images at the first time point and the feature quantity data of the frame images at the second time point. Therefore, it is expected that the information processing system according to this embodiment can use the generated enhanced data to generate a learning model for predicting the state of substrate processing based on feature quantities, and assist in the analysis of image-based substrate processing actions.

[0085] Furthermore, the other structures of the information processing system involved in Embodiment 2 are the same as those of the information processing system involved in Embodiment 1. Therefore, the same reference numerals are used to mark the same parts, and detailed descriptions are omitted.

[0086] [Implementation Method 3]

[0087] In the information processing systems described in Embodiments 1 and 2 above, during wet etching of the substrate processing apparatus 1 as substrate processing, the camera 5 captures images of the ejection section 13 to obtain motion images, and data enhancement is performed on the frame images contained in the acquired motion images to generate learning data. However, the objects of data enhancement and the generation of learning data are not limited to motion images obtained by capturing images of the ejection section 13 during wet etching.

[0088] In the substrate processing performed by the substrate processing apparatus 1, the resist formed on the substrate to be processed is used as a mask for etching or ion implantation, and then the resist that is no longer needed is removed from the substrate. The resist removal is performed using a removal solution such as SPM (Sulfuric acid hydrogen peroxide mixture), a mixture of sulfuric acid and hydrogen peroxide. To improve the resist removal capability, the removal solution is heated to a high temperature and ejected from the ejection section 13 onto the substrate to be processed. In the information processing system according to Embodiment 3, when the substrate processing apparatus 1 removes the resist formed on the substrate, the camera 5 captures images of the ejection section 13 from which the removal solution is ejected to obtain moving images, and data augmentation is performed on the frames contained in the acquired moving images to generate learning data.

[0089] The information processing apparatus 3 according to Embodiment 3 performs data augmentation on the frame images contained in the motion image obtained by capturing the ejection section 13 during the resist removal process, thereby amplifying the generation of learning data. Furthermore, the information processing apparatus 3 according to Embodiment 3 can also, similarly to the information processing apparatus 3 according to Embodiment 2, convert the frame images into feature quantities and perform data augmentation on the feature quantities. Data augmentation of the frame images or feature quantities can be performed, for example, by methods such as averaging, weighted averaging, linear interpolation, or cubic interpolation.

[0090] During the resist removal process, a high-temperature removal liquid is sprayed out, thus generating vapor within chamber 11. The visibility of the moving image captured by camera 5 may be reduced due to the vapor's effect. The enhanced frame image generated by information processing device 3 based on the frame images contained in the moving image is generated through data enhancement based on average values, etc., thus reducing the influence of vapor in the frame image. To achieve this effect, the information processing device 3 according to embodiment 3 may also exclude the frame images contained in the moving image captured by camera 5 from the learning data, and only include the enhanced frame image generated by data enhancement based on the frame images in the learning data.

[0091] Furthermore, the other structures of the information processing system involved in Embodiment 3 are the same as those of the information processing systems involved in Embodiments 1 and 2. Therefore, the same reference numerals are used to mark the same parts, and detailed descriptions are omitted.

[0092] Furthermore, in Embodiments 1 and 2, wet etching was described as an example of substrate processing, and in Embodiment 3, the removal of resist was described as an example. However, the generation of enhanced data by the information processing system involved in this embodiment is not limited to these substrate processing methods and can be applied to various substrate processing methods. In addition, the imaging part of the substrate processing apparatus 1 captured by the camera 5 is not limited to the ejection part 13, but can be various other parts.

[0093] It should be considered that all points in the embodiments disclosed herein are illustrative rather than restrictive. The scope of this disclosure is shown not by the foregoing meaning but by the claims, and is intended to include all modifications within the meaning and scope equivalent to the claims.

[0094] The items described in each embodiment can be combined with each other. Furthermore, the independent and dependent claims described in the claims can be combined with each other in all possible combinations, regardless of their referencing form. Moreover, the claims can be described in a form that refers to two or more other claims (multiple claim form), but are not limited to this. It is also possible to use a form that describes multiple claims that refer to at least one other multiple claim (multiple-referencing-multiple-claims).

[0095] Explanation of reference numerals in the attached figures

[0096] 1: Substrate processing device; 3: Information processing device (computer); 5: Camera; 11: Chamber; 12: Substrate holding mechanism; 12a: Holding part; 12b: Support part; 12c: Drive part; 13: Ejection part; 14: Recovery cup; 14a: Drain port; 14b: Exhaust port; 15: FFU; 16: Supply source; 31: Processing unit; 31a: Image acquisition unit; 31b: Data enhancement unit; 31c: Learning data generation unit; 31d: Learning processing unit; 31e: Display processing unit; 32: Storage unit; 32a: Program (computer program); 32b: Learning data storage unit; 32c: Model information storage unit; 33: Communication unit; 34: Display unit; 35: Operation unit; N: Network.

Claims

1. A program product including a computer program that causes a computer to execute the following processing: acquiring a moving image obtained by photographing a substrate process; and generating enhancement data based on a frame image of a first time point and a frame image of a second time point later than the first time point included in the acquired moving image.

2. The program product according to claim 1, wherein the moving image obtained by photographing an ejection portion of a substrate processing apparatus that ejects a liquid toward a substrate as a processing target is acquired.

3. The program product according to claim 1, wherein the enhancement data is generated based on an average of the frame image of the first time point and the frame image of the second time point.

4. The program product according to claim 1, wherein the frame image is converted into feature quantity data, the enhancement data is generated based on feature quantity data obtained by converting the frame image of the first time point and feature quantity data obtained by converting the frame image of the second time point.

5. The program product according to claim 4, wherein the enhancement data is generated based on an average of the feature quantity data of the first time point and the feature quantity data of the second time point.

6. The program product according to claim 3 or 5, wherein the enhancement data is generated based on a weighted average.

7. The program product according to claim 1, wherein a learning model that outputs information about a state of a substrate process from a frame image included in a moving image obtained by photographing the substrate process is generated by machine learning using the generated enhancement data.

8. A data generation method in which a moving image obtained by photographing a substrate process is acquired by an information processing apparatus, enhancement data is generated by the information processing apparatus based on a frame image of a first time point and a frame image of a second time point later than the first time point included in the acquired moving image.

9. A learning model generation method in which a moving image obtained by photographing a substrate process is acquired by an information processing apparatus, enhancement data is generated by the information processing apparatus based on a frame image of a first time point and a frame image of a second time point later than the first time point included in the acquired moving image, a learning model that outputs information about a state of a substrate process from a frame image included in a moving image obtained by photographing the substrate process is generated by the information processing apparatus by machine learning using the generated enhancement data.

10. An information processing apparatus in which a processing section is provided, a moving image obtained by photographing a substrate process is acquired by the processing section, enhancement data is generated by the processing section based on a frame image of a first time point and a frame image of a second time point later than the first time point included in the acquired moving image.

11. A computer-readable recording medium that records a computer program that causes a computer to execute the following processing: acquiring a moving image obtained by photographing a substrate process; and The enhancement data is generated based on a frame image at a first time point and a frame image at a second time point later than the first time point included in the acquired moving image.

Citation Information

Patent Citations

  • Substrate processing method and substrate processing apparatus

    JP2021190511A