Determining substrate layer thickness by polishing pad wear compensation.

A neural network-based system addresses the challenge of inconsistent polishing endpoints in CMP by compensating for polishing pad wear, enhancing substrate layer thickness determination and uniformity in CMP processes.

JP7723775B2Active Publication Date: 2025-08-14APPLIED MATERIALS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024028235
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-24
Filing Date
2024-02-28
Publication Date
2025-08-14
Estimated Expiration
2041-06-10

AI Technical Summary

Technical Problem

Chemical mechanical polishing (CMP) processes face challenges in accurately determining substrate layer thickness due to variations in polishing pad thickness caused by wear, leading to inconsistent polishing endpoints and non-uniformity across wafers.

Method used

A neural network-based system that compensates for polishing pad wear by using thickness measurements as input to generate corrected thickness values, combining in-situ monitoring with a trained neural network to adjust polishing parameters and endpoints.

Benefits of technology

Improves within-wafer and wafer-to-wafer non-uniformity by accurately determining substrate layer thickness, ensuring consistent polishing endpoints and uniformity across multiple wafers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007723775000001
    Figure 0007723775000001
  • Figure 0007723775000002
    Figure 0007723775000002
  • Figure 0007723775000003
    Figure 0007723775000003
Patent Text Reader

Abstract

To provide a method for training a neural network.SOLUTION: The method for training the neural network comprises: acquiring two ground truth thickness profiles for a test substrate; acquiring two thickness profiles for the test substrate measured by an in-situ monitor system while the test substrate is on polishing pads of different thicknesses; generating a thickness profile estimated for a different thickness value between two thickness values through interpolation between the two profiles; and training the neural network by using the estimated thickness profile.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to in situ monitoring and profile reconstruction during polishing of a substrate. [Background technology]

[0002] Integrated circuits are typically formed on a substrate (eg, a semiconductor wafer) by successively depositing conductive, semiconductive, or insulating layers on the silicon wafer and then processing the layers.

[0003] One manufacturing process involves depositing a filler layer over a non-planar surface and planarizing the filler layer until the non-planar surface is exposed. For example, a conductive filler layer can be deposited over a patterned insulating layer to fill trenches or holes in the insulating layer. The filler layer is then polished until the raised pattern of the insulating layer is exposed. After planarization, portions of the conductive layer remaining between the raised pattern of the insulating layer form vias, plugs, and lines that provide conductive paths between thin-film circuits on the substrate. Additionally, planarization may be used to planarize the substrate surface for lithography.

[0004] Chemical mechanical polishing (CMP) is an accepted method of planarization. This planarization method typically requires a substrate to be mounted on a carrier head. The exposed surface of the substrate is placed against a rotating polishing pad. The carrier head exerts a controllable load on the substrate, pressing it against the polishing pad. A polishing fluid, such as a slurry containing abrasive particles, is supplied to the surface of the polishing pad. To maintain a uniform polishing condition between wafers, the polishing pad is subjected to a conditioning process (e.g., polished by a polishing conditioner disk). Over the course of polishing multiple substrates, the thickness of the polishing pad can change due to the conditioner disk wearing down the polishing pad.

[0005] During semiconductor processing, it may be important to determine one or more characteristics of a substrate or a layer on a substrate. For example, during a CMP process, it is important to know the thickness of a conductive layer so that the process can be terminated at the correct time. Several methods may be used to determine substrate properties. For example, optical sensors may be used to perform in-situ monitoring of a substrate during chemical mechanical polishing. Alternatively (or additionally), eddy current sensing systems may be used to induce eddy currents in conductive regions on the substrate and determine parameters such as the local thickness of the conductive regions. Summary of the Invention

[0006] In one aspect, a method for training a neural network includes, for each test substrate of a plurality of test substrates having different thickness profiles, acquiring a ground truth thickness profile, acquiring a first thickness value; for each test substrate of the plurality of test substrates, acquiring a first measured thickness profile corresponding to the test substrate being measured by an in-situ monitoring system while on a polishing pad at a first thickness corresponding to the first thickness value; acquiring a second thickness value; for each test substrate of the plurality of test substrates, acquiring a second measured thickness profile corresponding to the test substrate being measured by an in-situ monitoring system while on a polishing pad at a second thickness corresponding to the second thickness value; for each test substrate of the plurality of test substrates, generating an estimated third thickness profile for a third thickness value between the first thickness value and the second thickness value by interpolating between the first and second thickness profiles for the test substrate; and, for each test substrate, training the neural network by applying the third thickness and the estimated third thickness profile to a plurality of input nodes and applying the ground truth thickness profile to a plurality of output nodes while the neural network is in a training mode.

[0007] In another aspect, a polishing system includes a platen for supporting a polishing pad, a carrier head for holding a substrate and contacting the substrate with the polishing pad, an in-situ monitor system for generating a signal dependent on the thickness of a conductive layer on the substrate while the conductive layer is being polished by the polishing pad, and a controller. The controller is configured to: receive a pre-polishing thickness dimension of the conductive layer; obtain an initial signal value from the in-situ monitor system at the start of polishing of the conductive layer; determine an expected signal value of the conductive layer based on the pre-polishing thickness; calculate a gain based on the initial signal value and the expected signal value; determine a polishing pad thickness value from the gain using a gain function; receive signals from the in-situ monitor system during polishing of the conductive layer and generate a plurality of measurement signals for a plurality of different locations on the layer; determine a plurality of thickness values for the plurality of different locations on the layer from the plurality of measurement signals; generate a corrected thickness value for the location by processing at least some of the plurality of thickness values through a neural network to provide a plurality of corrected thickness values for each of at least some of the plurality of different locations, wherein at least some of the plurality of thickness values and the polishing pad thickness value are input to the neural network, and the corrected thickness value is output by the neural network; generate a corrected thickness value for the location; and perform at least one of detecting a polishing endpoint or modifying polishing parameters based on the plurality of corrected thickness values.

[0008] In another aspect, a method for controlling polishing includes receiving a pre-polishing thickness dimension of a conductive layer on a substrate; contacting the conductive layer on the substrate with a polishing pad in a polishing system to initiate polishing; obtaining an initial signal value from an in-situ monitor system at the start of polishing of the conductive layer; determining an expected signal value of the conductive layer based on the pre-polishing thickness; calculating a gain based on the initial signal value and the expected signal value; determining a polishing pad thickness value from the gain using a gain function; receiving signals from the in-situ monitor system during polishing of the conductive layer to generate a plurality of measurement signals for a plurality of different locations on the layer; determining a plurality of thickness values for the plurality of different locations on the layer from the plurality of measurement signals; generating a corrected thickness value for each of at least some of the plurality of different locations by processing at least some of the plurality of thickness values through a neural network to provide a plurality of corrected thickness values, wherein at least some of the plurality of thickness values and the polishing pad thickness value are input to the neural network, and the corrected thickness value is output by the neural network; generating the corrected thickness value for the location; and performing at least one of detecting a polishing endpoint or modifying polishing parameters based on the plurality of corrected thickness values.

[0009] In another aspect, a computer program product tangibly embodied in a computer-readable medium includes instructions for causing one or more processors to receive a value representing a polishing pad thickness; receive signals from an in-situ monitor system during polishing of a conductive layer to generate a plurality of measurement signals for a plurality of different locations on the layer; determine a plurality of thickness values for the plurality of different locations on the layer from the plurality of measurement signals; generate a corrected thickness value for each of at least some of the plurality of different locations to provide a plurality of corrected thickness values by processing at least some of the plurality of thickness values through a neural network trained using a plurality of tuples of the thickness values of the test pad, the estimated thickness profile of the test layer, and the ground truth thickness profile of the test layer, wherein at least some of the plurality of thickness values and the polishing pad thickness value are input to the neural network, and the corrected thickness value is output by the neural network; and perform at least one of detecting a polishing endpoint or modifying polishing parameters based on the plurality of corrected thickness values.

[0010] In another aspect, a computer program product tangibly embodied in a computer-readable medium includes instructions for causing one or more processors to receive a value representing a polishing pad thickness, receive a plurality of measurement signals for a plurality of different locations on a layer being polished from an in-situ monitor system, determine a plurality of thickness values for the plurality of different locations on the layer from the plurality of measurement signals, generate a location-related corrected thickness value for each of at least some of the plurality of different locations by processing at least some of the plurality of thickness values through a neural network to provide a plurality of corrected thickness values, and perform at least one of detecting a polishing endpoint or modifying polishing parameters based on the plurality of corrected thickness values. The neural network includes a plurality of input nodes, a plurality of output nodes, and a plurality of intermediate nodes. At least some of the plurality of thickness values are applied to at least some of the input nodes, and the value representing the polishing pad thickness is applied directly to one intermediate node from the plurality of intermediate nodes, and at least some of the plurality of output nodes output a plurality of corrected thickness values.

[0011] Certain embodiments may include one or more of the following advantages: An in-situ monitoring system, such as an eddy current monitoring system, can generate a signal as a sensor scans across a substrate. The system can compensate for distortions in the edge portion of the signal due to, for example, variations in pad thickness from wafer to wafer. The signal can be used for endpoint control and / or closed-loop control of polishing parameters, such as carrier head pressure, thereby providing improved within-wafer non-uniformity (WIWNU) and water-to-wafer non-uniformity (WTWNU).

[0012] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0013] [Figure 1A] 1 is a schematic side view in partial cross section of a chemical mechanical polishing station including an eddy current monitoring system. [Figure 1B] FIG. 1 is a schematic plan view of a chemical mechanical polishing station. [Figure 2] 1 is a schematic top view of a substrate being scanned by a sensor head of a polishing apparatus. [Figure 3A] 1 is a schematic graph of a static formula for determining substrate thickness based on a measured signal. [Figure 3B] 10 is a schematic graph of a function for determining polishing pad thickness based on measured impedance signal gain. [Figure 4] 4 is a schematic graph of measurement signals obtained while monitoring locations on a substrate. [Figure 5] 1 is an exemplary neural network. [Figure 6] FIG. 1 is a flow diagram of an example process for polishing a substrate. [Figure 7] FIG. 10 is a flow diagram of an example process for training a neural network to generate a modified signal for a group of measured signals. DETAILED DESCRIPTION OF THE INVENTION

[0014] Like reference symbols in the various drawings indicate like elements.

[0015] The polishing apparatus can use an in-situ monitoring system, such as an eddy current monitoring system, to detect the thickness of the outer layer being polished on the substrate. The thickness measurements can be used to trigger a polishing endpoint and / or to adjust processing parameters of the polishing process in real time. For example, the substrate carrier head can adjust the pressure on the backside of the substrate to increase or decrease the polishing rate in various zones of the outer layer. The polishing rate can be adjusted so that the zones have substantially the same thickness after polishing and / or so that polishing of the zones is completed at approximately the same time. Such profile control is sometimes referred to as real-time profile control (RTPC).

[0016] In-situ monitoring systems can suffer from signal distortions in measurements close to the substrate edge. For example, eddy current monitoring systems can generate magnetic fields. Near the substrate edge, the magnetic fields only partially overlap with the substrate's conductive layer, artificially reducing the signal. A technique to compensate for these distortions is to feed the thickness measurements into a trained neural network.

[0017] Additionally, the signal from the eddy current monitoring system can artificially increase as the polishing pad thins due to wear from the conditioning disk. As the pad thickness changes, the signal from one or more sensors that read the substrate characteristics through the polishing pad can also change. In particular, as the pad becomes thinner, the distance between the substrate and the eddy current sensor decreases. This can cause an increase in signal strength, artificially increasing the apparent layer thickness and potentially resulting in inconsistent polishing endpoints or wafer-to-wafer non-uniformity. Even with a neural network, the system may not adequately compensate for distortion at the substrate edge if the signal also depends on the pad thickness. However, by training the neural network with layer thickness measurements corresponding to different pad thicknesses, the measured thickness values of the polishing pad can be used as input to generate the modified signal in the neural network.

[0018] 1A and 1B show an example of a polishing apparatus 100. The polishing apparatus 100 includes a rotatable, disk-shaped platen 120 on which a polishing pad 110 is positioned. The platen is operable to rotate about an axis 125. For example, a motor 121 can rotate a drive shaft 124 to rotate the platen 120. The polishing pad 110 can be a two-layer polishing pad having an outer polishing layer 112 and a softer backing layer 114.

[0019] The polishing apparatus 100 may include a port 130 for dispensing a polishing liquid 132, such as a slurry, onto the polishing pad 110. The polishing apparatus may also include a polishing pad conditioner that abrades the polishing pad 110 to maintain the polishing pad 110 in a consistent abrasive state.

[0020] The polishing apparatus 100 includes at least one carrier head 140. The carrier head 140 is operable to hold the substrate 10 against the polishing pad 110. The carrier heads 140 can have independent control of polishing parameters, such as the pressure associated with each substrate.

[0021] In particular, carrier head 140 may include a retaining ring 142 for retaining substrate 10 beneath a flexible membrane 144. Carrier head 140 also includes a plurality of independently controllable, pressurizable chambers defined by the membrane, e.g., three chambers 146a-146c, that can apply independently controllable pressures to associated zones on flexible membrane 144 and, therefore, substrate 10. For ease of illustration, only three chambers are shown in FIG. 1, but there may be one or two chambers, or four or more chambers, e.g., five chambers.

[0022] Carrier head 140 is suspended from a support structure 150, such as a carousel or track, and is connected by a drive shaft 152 to a carrier head rotation motor 154 so that the carrier head can rotate about axis 155. Optionally, carrier head 140 can oscillate laterally, for example, on a slider on the carousel 150 or track, or by rotational oscillation of the carousel itself. In operation, the platen rotates about its central axis 125, and the carrier head rotates about its central axis 155 and translates laterally across the top surface of the polishing pad.

[0023] Although only one carrier head 140 is shown, more carrier heads can be provided to hold additional substrates so that the surface area of the polishing pad 110 can be used efficiently.

[0024] The polishing apparatus 100 also includes an in-situ monitor system 160. The in-situ monitor system 160 generates a series of time-varying values that depend on the thickness of the layer on the substrate. The in-situ monitor system 160 includes a sensor head from which measurements are generated, and due to relative motion between the substrate and the sensor head, the measurements will be taken at different locations on the substrate.

[0025] The in-situ monitoring system 160 may be an eddy current monitoring system. The eddy current monitoring system 160 includes a drive system for inducing eddy currents in a conductive layer on the substrate and a sensing system for detecting the eddy currents induced in the conductive layer by the drive system. The monitoring system 160 includes a core 162 positioned within the recess 128 to rotate with the platen, at least one coil 164 wound around a portion of the core 162, and a drive and sense circuit 166 connected to the coil 164 by wiring 168. The combination of the core 162 and the coil 164 may provide a sensor head. In some embodiments, the core 162 protrudes above the top surface of the platen 120, for example, into the recess 118 in the bottom of the polishing pad 110.

[0026] Drive and sense circuitry 166 is configured to apply an oscillating electrical signal to coil 164 and measure the resulting eddy currents. Various configurations of the drive and sense circuitry and the coil(s) are possible, as described, for example, in U.S. Patent Nos. 6,924,641, 7,112,960, and 8,284,560, and U.S. Patent Publication Nos. 2011-0189925 and 2012-0276661. Drive and sense circuitry 166 could be located in the same recess 128 or a different portion of platen 120, or could be located outside of platen 120 and coupled to components within the platen through a rotary electrical union 129.

[0027] In operation, the drive and sense circuitry 166 drives the coil 164 to generate an oscillating magnetic field. At least a portion of the magnetic field extends through the polishing pad 110 and into the substrate 10. If a conductive layer is present on the substrate 10, the oscillating magnetic field generates eddy currents in the conductive layer. The eddy currents cause the conductive layer to act as an impedance source that is coupled to the drive and sense circuitry 166. Changes in the thickness of the conductive layer cause changes in the raw signal from the sensor head, which can be detected by the drive and sense circuitry 166.

[0028] Additionally, as described above, due to the conditioning process, the thickness of the polishing pad 110 can be reduced between wafers. Because the core 162 and coil 164 can be located within the recess 128 of the polishing pad 110 and the magnetic field can extend through the outer polishing layer 112 into the substrate 10, the distance between the core 162 and the substrate 10 decreases as the thickness of the polishing pad 110 decreases. As a result, the impedance read by the drive and sense circuitry 166 can also change as the thickness of the polishing pad 110 changes.

[0029] Generally, the drive and sense circuit 166 maintains a normalized signal from the core 162 by including a gain parameter on the raw signal from the sensor head. The gain parameter can be used to scale the signal for output to the controller 190. The gain parameter can be at a maximum value when the outer polishing layer 112 of the polishing pad 110 is at its largest, for example, when the polishing pad 110 is new. As the thickness of the outer polishing layer 112 decreases, the gain parameter can be decreased to compensate for the increased signal strength as the sensor moves closer to the substrate.

[0030] Alternatively or additionally, an optical monitoring system, which may function as a reflectometer or interferometer, may be fixed to the platen 120 within the recess 128. If both systems are used, the optical monitoring system and the eddy current monitoring system may monitor the same portion of the substrate.

[0031] The CMP apparatus 100 may also include a position sensor 180, such as an optical interrupter, to sense when the core 162 is under the substrate 10. For example, the optical interrupter could be attached to a fixed point on the opposite side of the carrier head 140. A flag 182 is attached to the periphery of the platen. The attachment point and length of the flag 182 are selected to interrupt the optical signal of the sensor 180 while the core 162 sweeps under the substrate 10. Alternatively or additionally, the CMP apparatus may include an encoder for determining the angular position of the platen.

[0032] A controller 190, such as a general purpose programmable digital computer, receives the intensity signals from the eddy current monitoring system 160. The controller 190 may include a processor, memory, and I / O devices, as well as an output device 192, such as a monitor, and an input device 194 (e.g., a keyboard).

[0033] Signals may be passed from the eddy current monitor system 160 through the rotary electric union 129 to the controller 190. Alternatively, the circuitry 166 could communicate with the controller 190 by wireless signals.

[0034] As the core 162 sweeps under the substrate with each platen revolution, information regarding the thickness of the conductive layer is accumulated in situ and on a continuous, real-time basis (once per platen revolution). The controller 190 can be programmed to sample measurements from the monitoring system when the substrate is generally over the core 162 (as determined by the position sensor). As polishing progresses, the thickness of the conductive layer changes, and the sampled signal changes over time. The sampled signal that changes over time is sometimes referred to as a trace. Measurements from the monitoring system can be displayed on an output device 192 during polishing to allow the equipment operator to visually monitor the progress of the polishing operation.

[0035] During operation, CMP apparatus 100 can use eddy current monitoring system 160 to determine when the bulk of the fill layer has been removed and / or when the underlying stop layer has been substantially exposed. Possible process control and endpoint criteria for the detector logic include local minima or maxima, changes in slope, amplitude or slope thresholds, or combinations thereof.

[0036] The controller 190 may also be connected to a pressure mechanism that controls the pressure applied by the carrier head 140, a carrier head rotation motor 154 that controls the carrier head rotation speed, a platen rotation motor 121 that controls the platen rotation speed, or the slurry distribution system 130 to control the slurry composition supplied to the polishing pad. In addition, the computer 190 may be programmed to divide measurements from the eddy current monitoring system 160 for each sweep under the substrate into multiple sampling zones, calculate the radial position of each sampling zone, and sort the amplitude measurements into radial ranges, as discussed in U.S. Pat. No. 6,399,501. After sorting the measurements into radial ranges, information about film thickness can be provided in real time to a closed-loop controller to periodically or continuously modify the polishing pressure profile applied by the carrier head to provide improved polishing uniformity.

[0037] The controller 190 can use a correlation curve relating the signal measured by the in-situ monitoring system 160 to the thickness of the layer being polished on the substrate 10 to generate an estimated dimension of the thickness of the layer being polished. An example of a correlation curve 303 is shown in FIG. 3A. In the coordinate system depicted in FIG. 3A, the horizontal axis represents the value of the signal received from the in-situ monitoring system 160, while the vertical axis represents values relative to the thickness of the layer on the substrate 10. For a given signal value, the controller 190 can use the correlation curve 303 to generate a corresponding thickness value. The correlation curve 303 can be considered a "static" formula in that it predicts a thickness value for each signal value, regardless of the time or location at which the sensor head acquired the signal. The correlation curve can be represented by various functions, such as a polynomial function or a look-up table (LUT) combined with linear interpolation.

[0038] The controller 190 can also use a gain function that relates the signal measured by the in-situ monitor system 160 to the thickness of the polishing pad 110 to generate an estimated dimension of the thickness of the substrate 10 .

[0039] The gain function can be generated by measuring the "raw" signal from the core 162 with a standard conductive body for different polishing pad thicknesses. Examples of bodies can include a substrate thicker than the impedance penetration depth, or a substrate of known uniform thickness measured by a four-point probe. This ensures a standard expected conductivity measurement from the body. This allows signals from the core 162 measured at various pad thicknesses to be scaled to a constant signal value related to the established conductivity of a standard body. The scaling value required to normalize the core 162 signal at a given pad thickness is the gain parameter.

[0040] For example, the core 162 signal measured for a large pad thickness and a standard substrate body may be lower than the core 162 signal measured for a small pad thickness and the same substrate. A larger pad thickness will reduce the measured core 162 signal in a standard body due to the lower impedance penetration in the standard body. To correct for this reduction, a gain parameter can be used to scale the measured core 162 signal to a normalized value as described above. Alternatively, a smaller pad thickness could result in a lower gain parameter required to scale the core 162 signal to the established normalized value. Each pad thickness and correlation gain value can constitute a correlation point, and multiple correlation points can establish a gain function.

[0041] Alternatively, the conductive layer thickness of the substrate 10 can be accurately measured by a separate metrology station, such as a four-point probe, before it is mounted in the carrier head 140 and moved over the polishing pad 110. Once polishing begins, a sensor can be swept under the substrate 10 to obtain a raw signal from the core 162. This signal is compared to an expected signal based on the measured conductive layer thickness, and a ratio between the given signal and the expected signal is established. This ratio can be used as a gain parameter by the drive and sense circuitry 166.

[0042] 3B shows an exemplary gain function 304. In the coordinate system depicted in FIG. 3B, the vertical axis may represent values of the gain parameter received from the in-situ monitor system 160, and the horizontal axis may represent values for the thickness of the outer polishing layer 112. The pad thicknesses and correlated gain values described above are gain function points 305 on the chart, and the gain function 304 is determined from the correlated points 305.

[0043] The gain function 304 can be determined based on a linear regression of ground truth measurements of the polishing pad 110 thickness correlated with corresponding core 162 signal measurements of a standard body. Figure 3B shows an example gain function 304 generated from four calibration measurements 305 correlating the polishing pad 110 thickness to the gain measured by the drive and sense circuitry 166. Generally, the gain function 304 can be constructed from at least two thickness measurement points 305.

[0044] Generally, to generate the calibration measurement 305, a ground truth measurement of the polishing pad 110's thickness can be determined using a precision instrument external to the system, such as a profilometer. This polishing pad 110 of known thickness is then placed in the polishing apparatus 100 in combination with a calibration substrate. The calibration substrate is a body with a conductive layer of consistent thickness, allowing the same calibration substrate to be used to calibrate multiple polishing tools. The calibration substrate is loaded into the polishing system 100 and moved to a position above the polishing pad 110 on the sensor head. To prevent polishing, a liquid without abrasive particles can be supplied to the polishing pad surface while the in-situ monitoring system measures the calibration substrate. The drive and sense circuitry 166 can then determine the signal strength and correlate it with the ground truth measurement of the polishing pad thickness to generate the calibration measurement 305. More calibration measurements 305 can be generated by repeating this correlation for multiple polishing pads 110 with additional thicknesses.

[0045] Generally, a gain function 304 relating the gain parameter and the thickness of the polishing pad 110 can be determined by regression using calibration measurements 305. In some embodiments, the correlation curve 303 can be a linear regression. In some embodiments, the correlation curve 303 can be a weighted or unweighted linear regression. For example, the weighted or unweighted linear regression can be a Deming, Theil-Sen, or Passing-Bablock linear regression. However, in some embodiments, the correlation curve can be a non-linear function. The gain function 304 can then be used to interpolate an estimated thickness value between two calibration measurements 305 of known gain parameters and pad thicknesses.

[0046] Once the gain function 304 is established, the controller 190 can use the gain function 304 to generate a gain parameter value for a given pad thickness for scaling the core 162 signal. The gain function 304 can be considered a “static” formula in that it predicts a thickness value for each signal value regardless of the time or location at which the sensor head acquires the signal. The gain function can be represented by a variety of functions, such as a LUT combined with linear interpolation.

[0047] 1B and 2, changes in the position of the sensor head relative to the substrate 10 can result in changes in the signal from the in-situ monitoring system 160. That is, as the sensor head scans across the substrate 10, the in-situ monitoring system 160 will measure multiple regions 94, e.g., measurement spots, at different locations on the substrate 10. The regions 94 may partially overlap (see FIG. 2).

[0048] 4 shows a graph 420 illustrating a signal profile 401 from the in-situ monitoring system 160 during a single pass of the sensor head under the substrate 10. The signal profile 401 is composed of a series of individual measurements from the sensor head as it sweeps under the substrate. The graph 420 can be a function of measurement time or the position of the measurement on the substrate, e.g., radial position. In either case, different portions of the signal profile 401 correspond to measurement spots 94 at different locations on the substrate 10 scanned by the sensor head. Thus, the graph 420 shows corresponding measurement signal values from the signal profile 401 for a given location of the substrate scanned by the sensor head.

[0049] 2 and 4, the signal profile 401 includes a first portion 422 corresponding to a location in the edge region 203 of the substrate 10 as the sensor head crosses the leading edge of the substrate 10, a second portion 424 corresponding to a location in the central region 201 of the substrate 10, and a third portion 426 corresponding to a location in the edge region 203 as the sensor head crosses the trailing edge of the substrate 10. The signal may also include a portion 428 corresponding to off-substrate measurements, i.e., the signal generated when the sensor head scans a region beyond the edge 204 of the substrate 10 in FIG.

[0050] The edge region 203 may correspond to a portion of the substrate where the measurement spot 94 of the sensor head overlaps the substrate edge 204. The central region 201 may include an annular anchor region 202 adjacent to the edge region 203 and an interior region 205 surrounded by the anchor region 202. The sensor head may scan these regions on its path 210 and generate a sequence of measurements corresponding to the sequence of locations along the path 210.

[0051] In the first portion 422, the signal strength increases from an initial strength (typically a signal obtained when the substrate and carrier head are not present) to a higher strength by transitioning the monitoring location from initially only slightly overlapping the substrate at the substrate's edge 204 (producing an initial low value) to a monitoring location that almost completely overlaps the substrate (producing a high value). Similarly, in the third portion 426, the signal strength decreases as the monitoring location transitions to the substrate's edge 204.

[0052] Although second portion 424 is shown as flat, this is for simplicity's sake; the actual signal in second portion 424 would likely contain variations due to both noise and layer thickness variations. Second portion 424 corresponds to a monitor location scanning central region 201. Second portion 424 includes subportions 421 and 423 resulting from the monitor location scanning anchor region 202 of central region 201, and subportion 427 resulting from the monitor location scanning interior region 205 of central region 201.

[0053] As discussed above, the variations in signal strength in regions 422, 426 are caused, in part, by the sensor's measurement area overlapping the substrate edge, rather than by inherent variations in the thickness or conductivity of the layer being monitored. As a result, this distortion in signal profile 401 can cause errors in the calculation of characteristic values for the substrate, e.g., layer thickness, near the substrate edge. To address this issue, controller 190 can include a neural network, e.g., neural network 500 of FIG. 5, to generate modified thickness values corresponding to one or more locations on substrate 10 based on the calculated thickness values corresponding to these locations.

[0054] 5, a neural network 500, when properly trained, is configured to generate modified thickness values that reduce and / or eliminate distortion in calculated thickness values near the substrate edge. The neural network 500 receives a group of inputs 504 and processes the inputs 504 through one or more neural network layers to generate a group of outputs 550. The layers of the neural network 500 include an input layer 510, an output layer 530, and one or more hidden layers 520.

[0055] Each layer of neural network 500 includes one or more neural network nodes. Each neural network node in a neural network layer receives one or more node input values (from inputs 504 to neural network 500 or from the outputs of one or more nodes in a previous neural network layer), processes the node input values according to one or more parameter values to generate activation values, and optionally applies a nonlinear transformation function (e.g., a sigmoid function or a hyperbolic tangent function) to the activation values to generate the neural network node's output.

[0056] Each node in the input layer 510 receives one of the inputs 504 to the neural network 500 as a node input value.

[0057] Inputs 504 to the neural network include initial thickness values from the in situ monitor system 160 for a number of different locations on the substrate 10, such as a first thickness value 501, a second thickness value 502, etc., through an nth thickness value 503. The initial thickness values may be individual values calculated from a sequence of signal values in the signal 401 using a correlation curve.

[0058] The input nodes 504 of the neural network 500 may also include one or more state input nodes 546 that receive one or more process state signals 516. In particular, a thickness dimension of the polishing pad 110 may be received as a process state signal 516 at the state input node 546.

[0059] The pad thickness can be a direct measurement of the thickness of the polishing pad 110, such as by a contact sensor at the polishing station. Alternatively, the thickness can be generated from the gain function 304 described above. In particular, the thickness of the conductive layer can be measured before polishing, such as by an in-line or stand-alone metrology system. This thickness can be converted to an expected signal value using the calibration curve 303. The expected signal value can then be compared to the actual signal value at the start of polishing the substrate, and this ratio provides a gain that can be used to determine the pad thickness according to the gain function 304.

[0060] Assuming the neural network has been trained with pad thickness as input, a "pad wear" gain function 304 can be applied to scale the pad thickness. Once applied, the pad thickness and thickness profile (including the appropriate gain after pad wear adjustment) are both used as inputs to the network to output a reconstructed edge profile at that particular pad thickness in real time.

[0061] Generally, the plurality of different locations includes locations within the edge region 203 and anchor region 202 of the substrate 10. In some embodiments, the plurality of different locations is only within the edge region 203 and anchor region 202. In other embodiments, the plurality of different locations spans all regions of the substrate.

[0062] Nodes in the hidden layer 520 and output layer 530 are shown as receiving inputs from all nodes in the preceding layer. This is the case for a fully connected, feedforward neural network. However, neural network 500 may also be a non-fully connected, feedforward neural network or a non-feedforward neural network. Furthermore, neural network 500 may include at least one of one or more fully connected, feedforward layers, one or more non-fully connected, feedforward layers, and one or more non-feedforward layers.

[0063] The neural network generates a group of modified thickness values 550 at nodes or "output nodes" 550 in the output layer 530. In some embodiments, there is an output node 550 for each input thickness value from the in-situ monitor system that is fed to the neural network 500. In this case, the number of output nodes 550 may correspond to the number of signal input nodes 504 in the input layer 510.

[0064] For example, the number of signal input nodes 544 can be equal to the number of measurements in the edge regions 203 and anchor regions 202, and there can be an equal number of output nodes 550. Thus, each output node 550 can generate a modified thickness value corresponding to a respective initial thickness value provided as an input to the signal input node 544, e.g., a first modified thickness value 551 for a first initial thickness value 501, a second modified thickness value 552 for a second initial thickness value 502, and an nth modified thickness value 553 for an nth initial thickness value 503.

[0065] In some implementations, the number of output nodes 550 is less than the number of input nodes 504. In some implementations, the number of output nodes 550 is less than the number of signal input nodes 544. For example, the number of signal input nodes 544 can be equal to the number of measurements in the edge regions 203 and anchor regions 202, while the number of output nodes 550 can be equal to the number of measurements in the edge regions 203. Again, each output node 550 of the output layer 530 generates a modified thickness value corresponding to a respective initial thickness value as the signal input nodes 504, e.g., the first modified thickness value 551 for the first initial thickness value 501, but only for the signal input nodes 554 that receive thickness values from the edge regions 203.

[0066] In some embodiments, one or more nodes in one or more of the hidden layers 520, for example, one or more nodes 572 in the first hidden layer, can directly receive one or more state input nodes 516, such as the thickness of the polishing pad 110. The polishing apparatus 100 can use the neural network 500 to generate modified thickness values. The modified thickness values can then be used as the determined thickness for each location within the first group of locations on the substrate, for example, locations within the edge region (and possibly the anchor region). For example, referring again to FIG. 4, the modified thickness values for the edge region can provide the modified portion 430 of the signal profile 401.

[0067] In some implementations, for a modified thickness value corresponding to a given measurement location, the neural network 500 may be configured such that only input thickness values from measurement locations within a predetermined distance of the given location are used in determining the modified thickness value. For example, for thickness values S1, S2, . . . S M ,…S N is received and corresponds to measurements at N consecutive locations on the path 210, then the Mth location (R M The modified thickness value S' (denoted by M is the modified thickness value S' MTo calculate the thickness value S M-L(min1) ,…S M ,…S M+L(maxN) Only the value of L can be used for a given modified thickness value S' using measurements spaced at most about 2-4 mm apart. M You can choose to generate measurements S M Measurements within about 1-2 mm, e.g., 1.5 mm, of the location can be used. For example, L can be a number in the range of 0-4, e.g., 1 or 2. For example, if measurements within 3 mm are used and the spacing between measurements is 1 mm, L can be 1. If the spacing is 0.5 mm, L can be 2. If the spacing is 0.25, L can be 4. However, this can depend on the polishing equipment configuration and processing conditions. The modified thickness value S' M The values of other parameters, such as pad wear, could still be used in calculating .

[0068] For example, there may be several hidden nodes 570, or "hidden nodes," 570 in one or more hidden layers 520, the number of which is equal to the number of signal input nodes 544, with each hidden node 570 corresponding to a respective signal input node 544. Each hidden node 570 may be disconnected from (or have a parameter value of 0 for) the input node 544 corresponding to measurements at locations greater than a predetermined distance from the location of the corresponding input node's measurement. For example, the Mth hidden node may be disconnected from (or have a parameter value of 0 for) the first through (ML-1)th input nodes 544 and the (M+L+1)th through Nth input nodes 544. Similarly, each output node 560 may be disconnected from (or have a parameter value of 0 for) the hidden node 570 corresponding to modified signals for locations greater than a predetermined distance from the location of the output node's measurement. For example, the Mth output node may be disconnected from the first through (ML-1)th hidden nodes 570 and the (M+L+1)th through Nth hidden nodes (or may have a parameter value of 0 for the hidden nodes).

[0069] In some embodiments, the polishing apparatus 100 can use a static formula to determine the thickness of a first group of locations on the substrate, for example, locations within the edge region. These substrates can be used to generate training data used to train the neural network.

[0070] 6 is a flow diagram of an example process 600 for polishing substrate 10. Process 600 can be performed by polishing apparatus 100.

[0071] The polishing apparatus 100 polishes (602) a layer on the substrate 10 and monitors (604) the layer during polishing to generate measured signal values for different locations on the layer. The locations on the layer may include one or more locations within the edge region 203 of the substrate (corresponding to regions 422 / 426 of the signal 401) and one or more locations within anchor regions 202 on the substrate (corresponding to regions 421 / 423 of the signal). The anchor regions 202 are spaced apart within the central region 201 of the substrate, away from the substrate edge 204, and are therefore not affected by distortions caused by the substrate edge 204. However, the anchor regions 202 may be adjacent to the edge region 203. The anchor regions 202 may also surround the interior region 205 of the central region 201. The number of anchor locations may depend on the measurement spot size and measurement frequency by the in-situ monitor system 160. In some embodiments, the number of anchor locations may not exceed a maximum value (e.g., a maximum of 4).

[0072] The polishing apparatus 100 generates an initial thickness value for each of the different locations from the measured signal values using a static formula (606). To a first approximation, the measured signal values are simply input into the static formula, which outputs a thickness value. However, other processing, such as normalizing the signal based on anchor area or compensating for the conductance of a particular material, can also be performed on the signal as part of generating the initial thickness values.

[0073] The polishing apparatus 100 generates adjusted thickness values using a neural network (608). The inputs to the neural network 500 are the initial thickness values generated by the in-situ monitor system 160 for different locations and the thickness of the polishing pad as a status signal. The output of the neural network 500 is modified thickness values that correspond to the input calculated thickness values, respectively.

[0074] The polishing apparatus 100 detects the polishing endpoint and / or modifies the polishing parameters based on the changed thickness value (610).

[0075] FIG. 7 is a flow diagram of an exemplary process 700 for training the neural network 500 to generate corrected thickness values. Multiple substrates having layers with different thickness profiles are scanned by an in-situ monitoring system while mounted on polishing pads of different thicknesses. The in-situ monitoring system generates estimated thickness dimensions based on a calibration curve (702). For each substrate, the system also obtains ground truth thickness dimensions for each location within a group of locations (704). The system can generate ground truth thickness measurements using electrical impedance measurements, such as a four-point probe. The system also obtains ground truth thickness measurements for the polishing pad using, for example, a profilometer.

[0076] The collected training data is applied to the neural network while the neural network is in training mode (706). In particular, for each substrate profile, estimated measurements of substrate layer thickness and polishing pad thickness are applied to input nodes, and ground truth measurements of substrate layer thickness are applied to output nodes. Training can include calculating a measure of the error between the estimated thickness dimension and the ground truth thickness dimension, and updating one or more parameters of the neural network 500 based on the measure of the error. To do this, the system can use a training algorithm that uses gradient descent with backpropagation.

[0077] Although the above discussion has focused on thickness measurements, the technique is also applicable to other property values, such as electrical conductivity.

[0078] The monitor system can be used in a variety of polishing systems. Either the polishing pad or the carrier head, or both, can move to provide relative motion between the polishing surface and the substrate. The polishing pad can be a circular (or some other shape) pad fixed to the platen, a tape extending between a supply roller and a take-up roller, or a continuous belt. The polishing pad can be fixed to the platen, incrementally advanced over the platen between polishing operations, or continuously driven over the platen during polishing. The pad can also be fixed to the platen during polishing, and a fluid bearing can exist between the platen and the polishing pad during polishing. The polishing pad can be a standard rough pad (e.g., polyurethane with or without fillers), a soft pad, or a fixed-abrasive pad.

[0079] Although the above discussion focuses on eddy current monitoring systems, the modification techniques can be applied to other types of monitoring systems that scan over the edge of a substrate, such as optical monitoring systems. Additionally, although the above discussion focuses on polishing systems, the modification techniques can be applied to other types of substrate processing systems, such as deposition or etching systems, including in-situ monitoring systems that scan over the edge of a substrate.

[0080] A number of embodiments of the invention have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the invention. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. 1. A polishing system comprising: a platen for supporting the polishing pad; a carrier head for holding a substrate and contacting the substrate with the polishing pad; an in-situ monitor system for generating a signal dependent on the thickness of the conductive layer on the substrate while the conductive layer is being polished by the polishing pad; Controller and wherein the controller obtaining a value representative of a polishing pad thickness; receiving signals from the in-situ monitor system while polishing the conductive layer to generate a plurality of measurement signals for a plurality of different locations on the conductive layer; determining a plurality of thickness values for the plurality of different locations on the conductive layer from the plurality of measurement signals; generating, for each location of at least some of the plurality of different locations, a modified thickness value for the location by processing at least some of the plurality of thickness values through a neural network to provide a plurality of modified thickness values, the neural network including a plurality of input nodes, a plurality of output nodes, and a plurality of intermediate nodes, wherein at least some of the plurality of thickness values are applied to at least some of the input nodes, a value representing the polishing pad thickness is applied directly to one intermediate node of the plurality of intermediate nodes, and at least some of the plurality of output nodes output the plurality of modified thickness values; at least one of detecting a polishing endpoint or modifying polishing parameters based on the plurality of corrected thickness values; The polishing system is configured to perform the steps of:

2. The system of claim 1 , wherein the in-situ monitoring system comprises an eddy current monitoring system.

3. 10. The system of claim 1, wherein the controller is configured such that thickness values for the plurality of different locations serve as input values for outputting a corrected thickness value for a particular location of the plurality of different locations.

4. The system of claim 3 , wherein the controller is configured such that not all of the plurality of thickness values are applied to the input node.

5. The controller receiving a dimension of a pre-polish thickness of the conductive layer; obtaining an initial signal value from the in-situ monitor system at the start of polishing the conductive layer; determining an expected signal value for the conductive layer based on the pre-polishing thickness; calculating a gain based on the initial signal value and the expected signal value; determining a value representing the polishing pad thickness from the gain using a gain function; The system of claim 1 configured to execute:

6. 1. A method for training a neural network, comprising: obtaining a ground truth thickness profile for each test substrate of a plurality of test substrates having different thickness profiles; obtaining a first thickness value; For each test substrate of the plurality of test substrates, obtaining a measured first thickness profile corresponding to the test substrate being measured by an in-situ monitor system while on a polishing pad at a first thickness corresponding to a first thickness value; obtaining a second thickness value; For each test substrate of the plurality of test substrates, obtaining a measured second thickness profile corresponding to the test substrate being measured by an in-situ monitor system while on a polishing pad at a second thickness corresponding to a second thickness value; generating, for each test substrate of the plurality of test substrates, an estimated third thickness profile for a third thickness value between the first thickness value and the second thickness value by interpolating between the first thickness profile and the second thickness profile; for each test substrate, while a neural network having a plurality of input nodes and a plurality of output nodes is in a training mode, training the neural network by applying the estimated third thickness profile to some of the plurality of input nodes, applying the third thickness value to one of a plurality of intermediate nodes in the neural network or one of the plurality of input nodes, and applying the ground truth thickness profile to a plurality of output nodes; A method comprising:

7. 7. The method of claim 6, wherein obtaining the first thickness value and the second thickness value comprises measuring a thickness of the polishing pad at the first thickness and measuring a thickness of the polishing pad at the second thickness.

8. 8. The method of claim 7, wherein measuring the thickness of the polishing pad at the first thickness and measuring the thickness of the polishing pad at the second thickness comprises measuring with a profilometer.

9. 7. The method of claim 6, wherein obtaining the first thickness profile or the second thickness profile further comprises placing the test substrate on a polishing pad of the first thickness or the second thickness and scanning the test substrate with the in-situ monitor system.

10. 7. The method of claim 6, wherein training the neural network further comprises applying the first thickness or the second thickness and the measured first thickness profile or the measured second thickness profile to the plurality of input nodes and applying the ground truth thickness profile to a plurality of output nodes while the neural network is in a training mode.

11. 7. The method of claim 6, wherein training the neural network further comprises applying a fourth thickness and an estimated fourth profile to a plurality of input nodes and applying the ground truth thickness profile to a plurality of output nodes while the neural network is in training mode.

12. The method of claim 6 , wherein the interpolation is a linear interpolation.

13. The method of claim 6 , comprising applying the third thickness value directly to the one intermediate node of the plurality of intermediate nodes.

14. The method of claim 6 , wherein the first measured thickness profile and the second measured thickness profile are thicknesses of a conductive layer on the substrate.

15. The method of claim 6 , wherein the in-situ monitoring system comprises an eddy current monitoring system.

16. The method of claim 15 , wherein obtaining the ground truth thickness profile comprises measuring electrical impedance.

17. 17. The method of claim 16, wherein measuring the electrical impedance comprises measuring using a four-point probe.

Citation Information

Patent Citations

  • Integrated endpoint detection system with optical and eddy current monitoring

    JP2004525521A

  • Polishing equipment using neural networks for monitoring

    JP2020518131A

  • Polishing apparatus using machine learning and compensation for pad thickness

    US20190299356A1