Film thickness estimation from machine learning-based substrate image processing
By training the neural network to use color images to measure substrate thickness, the problem of time-consuming and difficult to achieve high resolution in the prior art is solved, and fast and accurate thickness measurement and mass production control are achieved, and the efficiency and accuracy of the measurement system are improved.
Patent Information
- Application Number
- CN202180014115.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-29
- Filing Date
- 2021-06-25
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-06-25
AI Technical Summary
The prior art has the problem that high resolution measurement is not easy to achieve in substrate thickness measurement, especially in the process of chemical mechanical grinding, which is difficult to quickly and accurately measure the film thickness on the substrate.
Using machine learning methods, we train neural networks to use color images of the substrate to measure thickness, combine an inline optical metering system and a line scanning camera to quickly obtain the thickness information of multiple dies on the substrate, and use deep learning algorithms to perform efficient thickness inference.
It realizes rapid and accurate measurement of substrate thickness without affecting output, with an error of less than 5%, and supports real-time control in mass production, improving the resolution and accuracy of measurement.
Smart Images

Figure CN115104001B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to optical metrology, for example, to detecting the thickness of a layer on a substrate using machine learning methods. Background Art
[0002] Integrated circuits are typically formed on a substrate by sequentially depositing conductive, semiconducting, or insulating layers on a silicon wafer. During integrated circuit fabrication, it may be necessary to planarize the substrate surface in order to remove filler layers or to improve photolithographic planarity.
[0003] Chemical mechanical polishing (CMP) is an acceptable planarization method. This planarization method generally requires mounting the substrate on a carrier or polishing head. The exposed surface of the substrate is typically placed against a rotating polishing pad. The carrier head provides a controlled load to the substrate, pushing the substrate against the polishing pad. A polishing slurry is typically supplied to the surface of the polishing pad.
[0004] The thickness of the substrate layer before and after polishing can be measured, for example, using a variety of optical metrology systems (eg, spectroscopic or ellipsometry) in-line or at a stand-alone metrology station.
[0005] As a parallel problem, the development of hardware resources such as graphics processing units (GPUs) and tensor processing units (TPUs) has led to tremendous progress in deep learning algorithms and their applications. One of the growing areas of deep learning is computer vision and image recognition. Such computer vision algorithms are primarily designed for image classification or segmentation. Summary of the Invention
[0006] In one aspect, a method for training a neural network for a substrate thickness measurement system includes acquiring ground-truth thickness measurements of a top layer of a calibration substrate at a plurality of locations, each location being a defined location of a die fabricated on the substrate. Acquiring a plurality of color images of the calibration substrate, each color image corresponding to a region of a die fabricated on the substrate. Training a neural network to convert the color images of the die region from an in-line substrate imager into thickness measurements of the top layer in the die region. Training is performed using training data comprising the plurality of color images and ground-truth thickness measurements, with each corresponding color image paired with a ground-truth thickness measurement of the die region associated with the corresponding color image.
[0007] In another aspect, a method for controlling lapping includes acquiring a first color image of a first substrate at an inline monitoring station of a lapping system; segmenting the first color image into a plurality of second color images using a die mask such that each second color image corresponds to a region of a die fabricated on the first substrate; generating thickness measurements for one or more locations; and determining lapping parameters for the first substrate or a subsequent second substrate based on the thickness measurements. Each respective one of the one or more locations corresponds to a respective region of a die fabricated on the first substrate. To generate the thickness measurement for a region, the second color image corresponding to the region is processed by a neural network trained using training data, the training data comprising a plurality of third color images of dies of a calibration substrate and ground-truth thickness measurements of the calibration substrate, each respective third color image being paired with a ground-truth thickness measurement of the region of the die associated with the respective color image.
[0008] Implementations may include one or more of the following potential advantages: The thickness of multiple dies on a substrate can be quickly measured. For example, an inline metrology system can determine the thickness of a substrate based on a color image of the substrate without impacting throughput. The estimated thickness can be directly used in multivariate run-to-run control schemes.
[0009] The described method can be used to train a model to produce thickness measurements that are within 5% of the actual film thickness. While thickness measurements can be obtained from color images with three color channels, a hyperspectral camera can be added to the substrate imager system to provide higher-dimensional feature inputs to the model. This can facilitate training more complex models to understand more of the physical properties of the film stack.
[0010] Deep learning in metrology systems can achieve high inference speeds while still enabling high-resolution thickness profile measurements on substrates. This makes metrology systems a fast, low-cost pre- and post-metrology measurement tool for memory applications with higher thickness accuracy.
[0011] The accompanying drawings and the description below set forth the details of one or more embodiments. Other aspects, features, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A diagram showing an example of an in-line optical measurement system.
[0013] Figure 2A Examples showing exemplary images of substrates used for model training.
[0014] Figure 2B A schematic illustration of a computer data storage system.
[0015] Figure 3 A neural network used as part of a controller for a grinding device is shown.
[0016] Figure 4 A flow chart illustrating a method for detecting the thickness of a layer on a substrate using a deep learning method.
[0017] Like reference numbers in the various drawings represent like elements. DETAILED DESCRIPTION
[0018] Due to the variations in polishing rates that occur during the CMP process, thin film thickness measurements using dry metrology systems are often performed during CMP processing. These dry metrology measurement techniques often employ spectroscopic or ellipsometry methods, where variables in an optical model of the film stack are fitted to the collected measurements. These measurement techniques typically require precise alignment of the sensor at the measurement point on the substrate to ensure that the model is appropriate for the collected measurements. Consequently, measuring a large number of points on the substrate can be time-consuming, and collecting high-resolution thickness profiles is not feasible.
[0019] However, using machine learning, it is possible to measure the thickness of films on substrates in a reduced time. By training a deep neural network using color images of dies from the substrate and associated thickness measurements from other reliable metrology systems, the film thickness of the die can be measured by applying the input image to the neural network. This system can be used as a high-throughput and cost-effective solution for low-cost memory applications, for example. In addition to thickness inference, this technology can also be used to classify residue levels on substrates using image segmentation.
[0020] refer to Figure 1 The polishing apparatus 100 includes one or more carrier heads 126, each configured to carry a substrate 10; one or more polishing stations 106; and a transfer station for loading and unloading substrates from the carrier heads. Each polishing station 106 includes a polishing pad 130 supported on a platform 120. The polishing pad 130 may be a two-layer polishing pad having an outer polishing layer and a softer backing layer.
[0021] The carrier head 126 can be suspended from a support 128 and can be moved between the polishing stations. In some embodiments, the support 128 is an overhead track, and each carrier head 126 is coupled to a carriage 108, which is mounted to the track so that each carriage 108 can be selectively moved between the polishing station 124 and the transfer station. Alternatively, in some embodiments, the support 128 is a rotatable turntable, and rotation of the turntable simultaneously moves the carrier head 126 along a circular path.
[0022] Each polishing station 106 of the polishing apparatus 100 may include a port, for example, at the end of the arm 134, to dispense a polishing liquid 136, such as a polishing slurry, onto the polishing pad 130. Each polishing station 106 of the polishing apparatus 100 may also include a pad conditioning device to polish the polishing pad 130 to maintain the polishing pad 130 in a consistent polishing state.
[0023] Each carrier head 126 can be operated to hold the substrate 10 against the polishing pad 130. Each carrier head 126 can independently control polishing parameters, such as the pressure associated with each respective substrate. In particular, each carrier head 126 can include a retaining ring 142 that holds the substrate 10 beneath the flexible membrane 144. Each carrier head 126 can also include a plurality of independently controllable pressurization chambers (e.g., three chambers 146a-146c) defined by the membrane that can apply independently controllable pressure to associated areas on the flexible membrane 144 and, thereby, to the substrate 10. Although for ease of illustration, the carrier head 126 can be independently controlled. Figure 1 Only three chambers are shown in FIG, but there may be one or two chambers, or four or more chambers, for example five chambers.
[0024] Each carrier head 126 is suspended from a support 128 and is connected by a drive shaft 154 to a carrier head rotation motor 156 so that the carrier head can rotate about an axis 127. Optionally, each carrier head 126 can be caused to oscillate laterally, for example, by driving the carriage 108 on a track or by the rotational oscillation of the turntable itself. In operation, the platform rotates about its central axis, and each carrier head rotates about its central axis 127 and moves laterally over the top surface of the polishing pad.
[0025] A controller 190, such as a programmable computer, is connected to each motor to independently control the rotation rate of the platform 120 and the carrier head 126. The controller 190 may include a central processing unit (CPU) 192, a memory 194, and support circuits 196, such as input / output circuits, a power supply, a clock circuit, a cache, and the like. The memory is connected to the CPU 192. The memory is a non-transitory computer-readable medium and may be one or more readily available memories, such as random access memory (RAM), read-only memory (ROM), a floppy disk, a hard disk, or another form of digital storage. In addition, although shown as a single computer, the controller 190 may be a distributed system, for example, including multiple independently operating processors and memories.
[0026] The polishing apparatus 100 also includes an in-line (also referred to as sequential) optical metrology system 160. The color imaging system of the in-line optical metrology system 160 is located within the polishing apparatus 100 but does not perform measurements during polishing operations. Instead, it collects measurements between polishing operations (e.g., when moving a substrate from one polishing station to another, or before or after polishing, such as when moving a substrate from a transfer station to a polishing station or vice versa). Alternatively, the in-line optical metrology system 160 can be located in a fab interface unit or in a module accessible from the fab interface unit to measure substrates after they are removed from a cassette but before they are moved to a polishing unit, or after they are cleaned but before they are returned to the cassette.
[0027] The in-line optical metrology system 160 includes a sensor assembly 161 that provides color imaging of the substrate 10. The sensor assembly 161 may include a light source 162, a light detector 164, and circuitry 166 for sending and receiving signals between a controller 190 and the light source 162 and light detector 164.
[0028] Light source 162 can be operated to emit white light. In one embodiment, the emitted white light comprises light having a wavelength of 200 to 800 nanometers. Suitable light sources are arrays of white light emitting diodes (LEDs), xenon lamps, or xenon-mercury lamps. Light source 162 is oriented to direct light 168 at a non-zero angle of incidence α onto the exposed surface of substrate 10. The angle of incidence α can be, for example, about 30° to 75°, such as 50°.
[0029] The light source can illuminate a substantially linear, elongated area that spans the width of substrate 10. For example, light source 162 can include optics, such as a beam expander, that expand light from the light source to the elongated area. Alternatively or additionally, light source 162 can include a linear array of light sources. Light source 162 itself and the illuminated area on the substrate can be elongated and have a longitudinal axis parallel to the substrate surface.
[0030] Diffuser 170 may be placed in the path of light 168 , or light source 162 may include a diffuser to diffuse the light before it reaches substrate 10 .
[0031] Detector 164 is a color camera that is sensitive to light from light source 162. The camera includes an array of detector elements. For example, the camera may include a CCD array. In some embodiments, the array is a single column of detector elements. For example, the camera may be a line scan camera. A row of detector elements may extend parallel to the longitudinal axis of the elongated area illuminated by light source 162. If light source 162 includes a row of light emitting assemblies, the row of detector elements may extend along a first axis parallel to the longitudinal axis of light source 162. A row of detector elements may include 1024 or more elements.
[0032] The camera 164 is configured with appropriate focusing optics 172 to project a field of view of the substrate onto the array of detector elements. The field of view can be long enough to observe the entire width of the substrate 10, for example, 150 to 300 mm long. The camera 164, including the associated optics 172, can be configured so that individual pixels correspond to regions having a length equal to or less than about 0.5 mm. For example, assuming the field of view is approximately 200 mm long and the detector 164 includes 1024 elements, the image produced by the line scan camera can have pixels having a length of approximately 0.5 mm. To determine the length resolution of the image, the length of the field of view (FOV) can be divided by the number of pixels onto which the FOV is imaged to obtain the length resolution.
[0033] The camera 164 can also be configured so that the pixel width is comparable to the pixel length. For example, an advantage of a line scan camera is its extremely high frame rate. The frame rate can be at least 5 kHz. The frame rate can be set to a frequency such that as the imaging area scans the substrate 10, the pixel width is comparable to the pixel length, for example, equal to or less than about 0.3 mm.
[0034] The light source 162 and the light detector 164 can be supported on a platform 180. If the light detector 164 is a line scan camera, the light source 162 and the camera 164 can be moved relative to the substrate 10 so that the imaging area can scan the length of the substrate. Specifically, the relative motion can be in a direction parallel to the surface of the substrate 10 and perpendicular to a row of detector elements of the line scan camera 164.
[0035] In some embodiments, the platform 182 is stationary while the substrate support moves. For example, the carrier head 126 can be moved, such as by movement of the carriage 108 or rotational oscillation of the turntable, or a robotic arm holding the substrate in a factory interface unit can move the substrate 10 past the line scan camera 182. In some embodiments, the platform 180 can be movable while the carrier head or robotic arm remains stationary for image acquisition. For example, the platform 180 can be moved along a track 184 by a linear actuator 182. In either case, this allows the light source 162 and camera 164 to remain in a fixed position relative to each other as the scanned area moves across the substrate 10.
[0036] A potential advantage of moving the line scan camera and light source together across the substrate is that, for example, the relative angle between the light source and the camera remains constant for different locations on the wafer, compared to conventional 2D cameras. Consequently, artifacts caused by changes in viewing angle can be reduced or eliminated. Furthermore, line scan cameras can eliminate perspective distortion, whereas conventional 2D cameras exhibit inherent perspective distortion that requires image transformation to correct.
[0037] The sensor assembly 161 may include a mechanism to adjust the vertical distance between the substrate 10 and the light source 162 and the detector 164. For example, the sensor assembly 161 may include an actuator to adjust the vertical position of the platform 180.
[0038] Optionally, a polarizing filter 174 may be positioned in the light path, for example, between the substrate 10 and the detector 164. The polarizing filter 174 may be a circular polarizer (CPL). A typical CPL is a combination of a linear polarizer and a quarter-wave plate. Properly orienting the polarization axis of the polarizing filter 174 can reduce haze in the image and sharpen or enhance desired visual features.
[0039] Assuming the outermost layer on the substrate is a semi-transparent layer (e.g., a dielectric layer), the color of the light detected at detector 164 depends on, for example, the composition of the substrate surface, the smoothness of the substrate surface, and / or the amount of interference between light reflected from different interfaces of one or more layers (dielectric layers) on the substrate. As described above, light source 162 and light detector 164 can be connected to a computing device, such as controller 190, which can be operated to control the operation of light source 162 and light detector 164 and receive signals therefrom. The computing device that performs the various functions to convert the color image into a thickness measurement can be considered part of metrology system 160.
[0040] refer to Figure 2A , shows an example of an image 202 of a substrate 10 collected using an inline optical metrology system 160. The inline optical metrology system 160 produces a high-resolution color image 202, such as an image of at least 720 x 1080 pixels, for example, an image of at least 2048 x 2048 pixels, having at least three color channels (e.g., RGB channels). The color at any particular pixel depends on the thickness of one or more layers (including the top layer) in the region of the substrate corresponding to that pixel.
[0041] The image 202 is segmented into one or more regions 208, each corresponding to a die 206 fabricated on the substrate. The portion of the image providing the region 208 may be a predetermined region in the image or may be automatically determined by an algorithm based on the image.
[0042] As an example of a predetermined region in an image, the controller may store a die mask that identifies the location and region in the image for each region 208. For example, for a rectangular region, the region may be defined by the top right and bottom left coordinates in the image. Thus, the mask may be a data file that includes a pair of top right and bottom left coordinates for each rectangular region. In other cases, where the region is non-rectangular, more complex functions may be used.
[0043] In some embodiments, the orientation and position of the substrate can be determined, and the die mask can be aligned relative to the image. The substrate orientation can be determined using a notch detector or by image processing of the color image 202, for example, to determine the angle of a scribe line in the image. The substrate position can also be determined by image processing of the color image 202, for example, by detecting a circular substrate edge and then determining the center of the circle.
[0044] As an example of automatically determining regions 208, an image processing algorithm may analyze image 202 and detect lines. Image 202 may then be segmented into regions between the identified lines.
[0045] By segmenting the initial color image, multiple color images 204 of various regions 208 can be collected from the substrate 10. As described above, each color image 204 corresponds to a die 206 fabricated on the substrate. The collected color images can be output as PNG images, but many other formats may be used, such as JPEG, etc.
[0046] Color image 204 can be provided to an image processing algorithm to generate thickness measurements of the dies shown in color image 204. The image is used as input data for an image processing algorithm that has been trained (e.g., by a supervised deep learning method) to estimate layer thickness based on the color image. The supervised deep learning-based algorithm establishes a model between the color image and the thickness measurements. The image processing algorithm can include a neural network as a deep learning-based algorithm.
[0047] The intensity values for each color channel of each pixel in color image 204 are fed into an image processing algorithm, such as an input neuron of a neural network. Based on this input data, a layer thickness measurement for the color image can be calculated. Thus, inputting color image 204 into the image processing algorithm yields an output with an estimated thickness. This system can be used as a high-throughput and cost-effective solution, for example, for low-cost memory applications. In addition to thickness inference, this technique can also be used to classify residue levels on substrates using image segmentation.
[0048] To train an image processing algorithm (e.g., a neural network) using a supervised deep learning approach, calibration images of the dies of one or more calibration substrates may be acquired as discussed above. That is, each calibration substrate may be scanned with a line scan camera of the inline optical metrology system 160 to generate an initial calibration image, which may be segmented into multiple color images of various regions on the calibration substrate.
[0049] Before or after collecting the initial color calibration image, surface-truth thickness measurements are collected at multiple locations on the calibration substrate using a high-accuracy metrology system (e.g., an in-line or stand-alone metrology system). The high-accuracy metrology system can be a dry optical metrology system. The surface-truth measurements can come from offline reflectometry, ellipsometry, scatterometry, or more advanced TEM measurements, although other techniques may also be suitable. Such systems are available from Nova Measuring Instruments or Nanometrics. Each location corresponds to one of the manufactured dies, i.e., one in each region.
[0050] For example, reference Figure 2B For each volume area on each calibration substrate, a color calibration image 212 is collected using the inline sensor of the optical metrology system 160. Each color calibration image is associated with a ground-truth thickness measurement 214 for the corresponding die on the calibration substrate. The images 212 and the associated ground-truth thickness measurements 214 can be stored in a database 220. For example, the data can be stored as records 210, where each record includes a calibration image 212 and a ground-truth thickness measurement 214.
[0051] The combined dataset 218 is then used to train a deep learning-based algorithm, such as a neural network. When training the model, thickness measurements corresponding to the center of the die, measured using a dry metrology tool, are used as labels for the input images. For example, the model can be trained on approximately 50,000 images collected from five dies on substrates with a wide range of back thicknesses.
[0052] Figure 3 A neural network 320 is shown as being used as part of the controller 190 of the lapping apparatus 100. The neural network 320 may be a deep neural network that performs a regression analysis on the RGB intensity values of an input image of a calibration substrate and ground truth thickness measurements to generate a model for predicting the thickness of a layer in a region of the substrate based on a color image of the region.
[0053] Neural network 320 includes a plurality of input nodes 322. Neural network 320 may include an input node for each color channel associated with each pixel of the input color image, a plurality of hidden nodes 324 (hereinafter also referred to as "intermediate nodes"), and an output node 326, which generates layer thickness measurements. In a neural network with a single layer of hidden nodes, each hidden node 324 may be coupled to each input node 322, and an output node 326 may be coupled to each hidden node 320. However, as a practical matter, a neural network used for image processing may have many layers of hidden nodes 324.
[0054] Typically, hidden node 324 outputs a value that is a nonlinear function of a weighted sum of the values from input node 322 or previous layer hidden nodes to which hidden node 324 is connected.
[0055] For example, the output of a hidden node 324 in the first layer (labeled as node k) can be represented as:
[0056] tanh(0.5*a k1 (I1)+a k2 (I2)+...+a kM (I M )+b k )
[0057] Among them, tanh is the hyperbolic tangent, a kx is the weight of the connection between the kth intermediate node and the xth input node (of M input nodes), and I M is the value at the Mth input node. However, other nonlinear functions can be used instead of tanh, such as the rectified linear unit (ReLU) function and its variants.
[0058] The neural network 320 thus includes an input node 322 for each color channel associated with each pixel of the input color image. For example, where there are J pixels and K color channels, then L=J*K is the number of intensity values in the input color image, and the neural network 320 will include at least input nodes N1, N2, ..., N L .
[0059] Thus, the output H of hidden node 324 (labeled as node k) can be calculated with the number of input nodes corresponding to the number of intensity values in the color image. k Expressed as:
[0060] H k =tanh(0.5*a k1 (I1)+a k2 (I2)+...+a kL (I L )+b k )
[0061] Assume that the row matrix (i1,i2,...,i L ) represents the measured color image S, and the output of the intermediate node 324 (labeled as node k) can be expressed as:
[0062] H k =tanh(0.5*a k1 (V1·S)+ak2(V2·S)+...+a kL (V L ·S)+b k )
[0063] Where V is the weighted value (v1, v2, v L ), V x is the weight of the xth intensity value among the L intensity values of the color image.
[0064] The output node 326 can produce a characteristic value CV, such as thickness, which is a weighted sum of the outputs of the hidden nodes. For example, this can be expressed as
[0065] CV=C1*H1+C2*H2+...+C l *H l
[0066] Among them C k is the weight of the output of the kth hidden node.
[0067] However, the neural network 320 may optionally include one or more other input nodes (e.g., node 322a) to receive additional data. This additional data may come from: previous measurements of the substrate using an in-situ monitoring system, such as pixel intensity values collected early in the processing of the substrate; measurements of a previous substrate, such as pixel intensity values collected during the processing of another substrate; another sensor in the polishing system, such as a measurement of the temperature of the pad or substrate using a temperature sensor; a polishing recipe stored by a controller used to control the polishing system, such as polishing parameters such as carrier head pressure and platen rotation rate used to polish the substrate; variables tracked by the controller, such as a number of substrates after changing the pad; or sensors not part of the polishing system, such as measurements of the thickness of the underlying film using a metrology station. This allows the neural network 320 to take other process or environmental variables into account when calculating the layer thickness measurement.
[0068] The thickness measurement generated at output node 326 is provided to process control module 330. The process control module can adjust process parameters such as carrier head pressure, platen rotation rate, etc. based on the thickness measurement of one or more regions. The adjustments can be made to the polishing process to be performed on the substrate or the next substrate.
[0069] The neural network 320 needs to be configured before being used, for example, for substrate measurement results.
[0070] As part of the configuration procedure, controller 190 may receive a plurality of calibration images. Each calibration image has a plurality of intensity values, e.g., an intensity value for each pixel of each color channel of the calibration image. The controller also receives a characteristic value, e.g., thickness, for each calibration image. For example, a color calibration image may be measured at a particular die fabricated on one or more calibration or test substrates. Additionally, a surface-based measurement of thickness may be performed at a particular die location using a dry measurement device (e.g., a contact surface profilometer or ellipsometer). The surface-based thickness measurement may thus be correlated with a color image of the same die location on the substrate. The plurality of color calibration images may be generated from, e.g., five to ten calibration substrates by segmenting the images of the calibration substrates as discussed above. With respect to the configuration procedure for neural network 320, neural network 320 is trained using the color image and characteristic value for each die fabricated on the calibration substrates.
[0071] V corresponds to one of the color images and is therefore associated with a characteristic value. When the neural network 320 operates in a training mode (e.g., a backpropagation mode), the values (v1, v2, ..., v L ) are provided to the corresponding input nodes N1, N2...N l , and simultaneously provides the characteristic value CV to the output node 326. This process can be repeated for each column. This process sets a in the above formula 1 or 2 k1 Equal values.
[0072] The system is now ready for operation. A color image is measured from the substrate using the inline monitoring system 160. The available row matrix S = (i1, i2, ..., i L ) represents the measured color image, where i j The intensity value representing the j-th intensity value among L intensity values, when the image includes a total of n pixels and each pixel includes three color channels, L=3n.
[0073] When the neural network 320 is used in inference mode, these values (S1, S2, ..., S L ) is provided as input to the corresponding input nodes N1, N2, ... N L Thus, the neural network 320 generates a characteristic value, such as a layer thickness, at an output node 326 .
[0074] The architecture of neural network 320 can vary in depth and width. For example, although neural network 320 is shown with a single row of intermediate nodes 324, it can include multiple rows. The number of intermediate nodes 324 can be equal to or greater than the number of input nodes 322.
[0075] As described above, the controller 190 can associate each color image with different dies on the substrate (see Figure 2A and Figure 2B). The output of each neural network 320 can be classified as belonging to one of the dies based on the position of the sensor on the substrate when the image was collected. This allows the controller 190 to generate a separate sequence of measurements for each die.
[0076] In some embodiments, controller 190 may be configured to have a neural network model structure composed of multiple different types of building blocks. For example, the neural network may be a residual neural network, whose architecture includes residual block features. The residual neural network may utilize skip connections or shortcuts to skip layers. The residual neural network may be implemented, for example, as a ResNet model. In the context of a residual neural network, a non-residual neural network may be described as a normal network.
[0077] In some embodiments, a neural network can be trained to account for the thickness of underlying layers of the stack during calculations, which can improve errors due to variations in thickness measurements. The effects of variations in the underlying thickness of the film stack can be mitigated to improve model performance by providing the intensity values of a color image of the underlying layer's thickness as an additional input to the model.
[0078] The reliability of the calculated thickness measurement can be assessed by comparing it to the measured value and then determining the difference between the calculated value and the original measurement. This deep learning model can then be used to predict thickness in inference mode immediately after scanning a new test substrate. This new approach improves overall system throughput and enables thickness measurements to be performed on all substrates in a batch.
[0079] refer to Figure 4 , an image processing algorithm generated through machine learning techniques for substrate thickness measurement systems. This image processing algorithm accepts RGB images collected by an integrated line scan camera inspection system and provides significantly faster film thickness estimation. Inference time for approximately 2,000 measurement points is on the order of seconds, compared to two hours using dry metrology.
[0080] The method includes a controller combining individual image lines from a light detector 164 into a two-dimensional color image (500). The controller may apply an offset and / or gain adjustment to the intensity values of the image in each color channel (510). Each color channel may have a different offset and / or gain. The image may optionally be normalized (515). For example, a difference between the measured image and a standard predefined image may be calculated. For example, the controller may store a background image for each of the red, green, and blue channels, which may be subtracted from the measured image for each color channel. Alternatively, the standard predefined image may segment the measured image. The image may be filtered to remove low-frequency spatial variations (530). In some embodiments, a filter is generated using the luminance channel and then applied to the red, green, and blue images.
[0081] The image is converted (e.g., resized and / or rotated and / or translated) to a standard image coordinate system (540). For example, the image may be translated so that the center of the die is at the center point of the image, and / or the image may be resized so that the edge of the substrate is at the edge of the image, and / or the image may be rotated so that the x-axis of the image is at a 0° angle to a radial segment connecting the center of the substrate and the substrate orientation feature.
[0082] One or more regions on the substrate are selected and an image is generated for each selected region (550). This step can be performed using the techniques described above, for example, the regions can be predetermined regions or automatically determined by an algorithm to provide portions of region 208.
[0083] The intensity values provided by each color channel of each pixel of the image are treated as input to an image processing algorithm for supervised deep learning training. The image processing algorithm outputs a layer thickness measurement for a specific area (560).
[0084] Each deep model architecture was trained and validated on a small-die test patterned substrate, with the goal of reducing errors in the measurement results. Models that account for underlying characteristics have lower errors. Additionally, preliminary tool-to-tool matching validation was performed by training the model on data collected on one tool and using it to infer data from other tools. The results were compared to training and inferring on data from the same tool.
[0085] In general, data can be used to control one or more operating parameters of a CMP apparatus. These operating parameters include, for example, the platform rotation speed, substrate rotation speed, substrate polishing path, substrate speed on the platform, pressure applied to the substrate, slurry composition, slurry flow rate, and substrate surface temperature. These operating parameters can be controlled in real time and can be automatically adjusted without further human intervention.
[0086] As used in this specification, the term substrate may include, for example, production substrates (e.g., including multiple memory or processor dies), test substrates, bare substrates, and gated substrates. A substrate may be at various stages of integrated circuit fabrication; for example, a substrate may be a bare wafer, or it may include one or more deposited and / or patterned layers. The term substrate may also include circular disks and rectangular sheets.
[0087] However, the color image processing techniques described above can be particularly useful in the context of 3D vertical NAND (VNAND) flash memory. Specifically, the layer stacks used in VNAND manufacturing are so complex that current metrology methods (such as Nova spectroscopy) cannot reliably detect areas of inappropriate thickness. In contrast, color image processing techniques can provide higher reliability in this application.
[0088] The embodiments of the present invention and the functional operations described in this specification may be implemented in a digital electronic circuit system, or in computer software, firmware, or hardware, including the structural components disclosed in this specification and their structural equivalents, or a combination thereof. The embodiments of the present invention may be implemented with one or more computer program products, i.e., one or more computer programs, tangibly present in a non-transitory machine-readable storage medium, executed by a data processing device (e.g., a programmable processor, a computer, or multiple processors or computers) or to control the operation of the data processing device.
[0089] The use of relative position terms refers to the relative positions of the system components to each other (not necessarily with respect to gravity); it is understood that the polishing surface and substrate may be maintained in a vertical orientation or some other orientation.
[0090] Several embodiments have been described. However, it will be appreciated that various modifications may be made. For example
[0091] Instead of a line scan camera, a camera that images the entire substrate can be used. In this case, no movement of the camera relative to the substrate is required.
[0092] The camera may cover less than the entire width of the substrate. In this case, the camera will need to undergo two perpendicular motions (e.g. supported on an XY stage) to scan the entire substrate.
[0093] The light source can illuminate the entire substrate. In this case, the light source does not need to be moved relative to the substrate.
[0094] The light detector may be a spectrometer rather than a color camera; the spectral data may then be reduced to RGB color space.
[0095] The sensing assembly need not be an inline system located between the polishing stations or between the polishing station and the transfer station. For example, the sensing assembly can be located within the transfer station, in the cartridge interface unit, or as a standalone system.
[0096] A uniformity analysis step is optionally used. For example, the image generated by applying a threshold transformation can be provided to a feed-forward process to adjust a subsequent processing step of the substrate, or to a feedback process to adjust a subsequent processing step of the substrate.
[0097] Accordingly, other implementations are within the scope of the following claims.
Claims
1. A non-transitory computer-readable medium encoded with a computer program product comprising instructions for causing one or more processors to perform the following steps: obtaining ground-truth thickness measurements of a top layer of a calibration substrate at a plurality of locations, each location being a defined location of a die fabricated on the calibration substrate; acquiring a plurality of color images of the calibration substrate, each color image corresponding to a region of a die fabricated on the calibration substrate; and training a neural network to convert a color image of a die region from an in-line substrate imager into a thickness measurement of the top layer in the die region, wherein the instructions for training the neural network include instructions for using training data, the training data including the plurality of color images and ground-truth thickness measurements, each respective color image being paired with the ground-truth thickness measurement of the die region associated with the respective color image.
2. The computer-readable medium of claim 1, wherein the instructions to acquire the plurality of color images comprise receiving a scan of the calibration substrate from the inline substrate imager.
3. The computer-readable medium of claim 1, wherein the instructions for acquiring the plurality of color images comprise instructions for receiving a color image of the calibration substrate and segmenting the color image into the plurality of color images based on a die mask.
4. The computer-readable medium of claim 1 , comprising instructions to obtain ground-truth thickness measurements of a top layer of a plurality of calibration substrates and to obtain a plurality of color images of each of the calibration substrates.
5. The computer-readable medium of claim 1, wherein the instructions to obtain the surface-true thickness measurements include instructions to receive a measurement of thickness from an optical surface profiler at each of the plurality of locations.
6. The computer-readable medium of claim 1 , comprising instructions to: obtaining ground-truth thickness measurements of a lower layer of the calibration substrate at the plurality of locations; The neural network is trained to convert a color image of the die area from the in-line substrate imager into a thickness measurement of the underlying layer.
7. The computer readable medium of claim 1, wherein the defined location is a center of the die.
8. The computer readable medium of claim 1 , comprising instructions to obtain measurements from all dies in a substrate.
9. The computer-readable medium of claim 1 , comprising instructions for acquiring measurements from all substrates in a batch before and after chemical mechanical planarization.
10. A method for training a neural network for use in a substrate thickness measurement system, comprising the following steps: obtaining ground-truth thickness measurements of a top layer of a calibration substrate at a plurality of locations, each location being a defined location of a die fabricated on the calibration substrate; acquiring a plurality of color images of the calibration substrate, each color image corresponding to a region of a die fabricated on the calibration substrate; as well as A neural network is trained to convert a color image of a die region from an in-line substrate imager into a thickness measurement of the top layer in the die region, the training being performed using training data comprising the plurality of color images and ground-truth thickness measurements, each respective color image being paired with the ground-truth thickness measurement of the die region associated with the respective color image.
11. A computer program product comprising a non-transitory computer-readable medium encoded with instructions to cause one or more processors to perform the following steps: receiving a first color image of a first substrate from an inline monitoring system of a lapping system; segmenting the first color image into a plurality of second color images using a die mask such that each second color image corresponds to a respective region of a die fabricated on the first substrate; generating thickness measurements for one or more locations, each respective location of the one or more locations corresponding to a respective region of a die fabricated on the first substrate, wherein the instructions for generating the thickness measurements for a region comprise instructions for processing a second color image corresponding to the region using a neural network trained using training data, the training data comprising a plurality of third color images of dies of a calibration substrate and ground-truth thickness measurements of the calibration substrate, each respective third color image being paired with a ground-truth thickness measurement of the region of the die associated with the respective third color image; and A value of a polishing parameter for the first substrate or a subsequent second substrate is determined based on the thickness measurement result.
12. The computer program product of claim 11, comprising instructions for receiving the first color image of the first substrate after the first substrate is ground at a grinding station.
13. The computer program product of claim 12, comprising instructions for determining the polishing parameters of the polishing station for the subsequent second substrate based on the thickness measurement results.
14. The computer program product of claim 11, comprising instructions for receiving the first color image of the first substrate before grinding the first substrate at a grinding station.
15. The computer program product of claim 14, comprising instructions for determining the polishing parameters of the polishing station for the first substrate based on the thickness measurement.
16. The computer program product of claim 11, wherein the milling parameters comprise a pressure of a chamber in a carrier head.
17. A grinding device comprising: a polishing station comprising a platform to support a polishing pad and a carrier head to hold a first substrate against the polishing pad; an inline measurement station having a color camera to generate a color image of the first substrate; and A control system configured to receiving a first color image of the first substrate from an inline monitoring station of the lapping system, segmenting the first color image into a plurality of second color images using a die mask such that each second color image corresponds to a respective region of a die fabricated on the first substrate; generating thickness measurements for one or more locations, each respective location of the one or more locations corresponding to a respective region of a die fabricated on the first substrate, wherein generating the thickness measurements for a region comprises processing a second color image corresponding to the region by a neural network trained using training data, the training data comprising a plurality of third color images of dies of a calibration substrate and ground-truth thickness measurements of the calibration substrate, each respective third color image being paired with a ground-truth thickness measurement of the region of the die associated with the respective third color image, and determining a value of a grinding parameter for the first substrate or a subsequent second substrate based on the thickness measurement result; The polishing station is caused to polish the first substrate or the subsequent second substrate using the determined value of the polishing parameter.
18. The apparatus of claim 17, wherein the milling parameters comprise a pressure of a chamber in the carrier head.
19. The apparatus of claim 17, wherein the control system is configured to receive the first color image of the first substrate after the first substrate is polished at a polishing station.
20. A method for controlling grinding, comprising the steps of: acquiring a first color image of the first substrate at an inline monitoring system of the grinding system; segmenting the first color image into a plurality of second color images using a die mask such that each second color image corresponds to a respective region of a die fabricated on the first substrate; generating thickness measurements for one or more locations, each respective location of the one or more locations corresponding to a respective region of a die fabricated on the first substrate, wherein generating the thickness measurements for a region comprises processing a second color image corresponding to the region using a neural network trained using training data, the training data comprising a plurality of third color images of dies of a calibration substrate and ground-truth thickness measurements of the calibration substrate, each respective third color image being paired with a ground-truth thickness measurement of the region of the die associated with the respective third color image; and A value of a polishing parameter for the first substrate or a subsequent second substrate is determined based on the thickness measurement result.
Citation Information
Patent Citations
Thickness measurement of substrate using color metrology
CN109716494A
Systems and methods for combining optical metrology with mass metrology
CN111066131A