Polishing device, information processing system, polishing method and recording medium

By using machine learning models in the grinding device to monitor friction and temperature characteristic quantities in real time, and estimating the substrate film thickness and qualification rate, the problems of missing defective products and reduced processing volume in the prior art are solved, and efficient grinding process management is achieved.

CN113492356BActive Publication Date: 2025-08-19EBARA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110285164.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-19
Filing Date
2021-03-17
Publication Date
2025-08-19
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

When the existing grinding device grinds the substrate, it is difficult to effectively avoid missing defective products, resulting in a decrease in the processing volume and a decrease in the product pass rate, and the measurement time of the film thickness measurer becomes a bottleneck.

Method used

Using the machine learning model, by monitoring the friction signal and temperature characteristic amount during the grinding process, the substrate film thickness and product qualification rate after grinding are estimated, real-time monitoring and prediction of the grinding state is achieved, and the number of film thickness measurements is reduced.

Benefits of technology

The processing volume is increased, the number of film thickness measurements is reduced, the omission of defective products is avoided, the product pass rate is improved, and the grinding state is timely predicted and prevented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113492356B_ABST
    Figure CN113492356B_ABST
Patent Text Reader

Abstract

The present invention is a polishing device, an information processing system, a polishing method and a recording medium. The polishing device can refer to a storage body storing a machine learning model that has completed learning using learning data. The learning data takes as input the characteristic quantity of a signal regarding the friction force between a polishing component and a substrate during polishing or the characteristic quantity of the temperature of the polishing component or the substrate during polishing, and takes as output data regarding the film thickness of the substrate after polishing or parameters regarding the product qualification rate contained in the substrate after polishing. The polishing device has a processor that generates a characteristic quantity based on the signal regarding the friction force between the polishing component and the substrate during polishing or the temperature of the polishing component or the target substrate during polishing, and inputs the generated characteristic quantity into the machine learning model that has completed learning, thereby outputting the data regarding the film thickness of the substrate after polishing or any one of the parameters regarding the product qualification rate contained in the substrate after polishing as an estimated value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to a polishing device, an information processing system, a polishing method, and a program. Background Art

[0002] Polishing devices for polishing substrates (e.g., wafers) are known. For example, Patent Document 1 discloses a polishing device comprising: a rotatable polishing table equipped with polishing components; and a rotatable polishing head facing the polishing table, wherein a substrate can be mounted on a surface facing the polishing table.

[0003] The polishing condition of the polishing device may sometimes deteriorate. The deterioration of the condition in this case includes the consumption of consumable parts of the polishing device (for example, the polishing pad as an example of the polishing part) and the deterioration of the workbench condition. Therefore, when the polishing condition deteriorates, the profile of the film thickness (also called residual film) after the substrate is polished deteriorates (for example, the variation of the film thickness becomes larger). In this case, in order to investigate whether it is a defective product, the film thickness or film thickness profile after polishing is measured for all the polished substrates by a film thickness meter, which takes time. In particular, when there is only one film thickness meter for multiple polishing devices, when the film thickness meter is used to measure all the polished substrates, the measurement time of the film thickness meter becomes a bottleneck, and there is a problem of reduced processing capacity. Although it is also possible to only sample an arbitrary substrate for film thickness measurement, or to reduce the number of measurement points on the substrate to shorten the measurement time of the film thickness meter (ITM), both may miss defective products and affect the product qualification rate, so it is not suitable. Summary of the Invention

[0004] The present technology has been developed in view of the above-mentioned problems, and it is intended to provide a polishing device, an information processing system, and a program that can avoid missing defective products and thereby increase the processing throughput or improve the yield rate.

[0005] (Methods of solving the problem)

[0006] A polishing device in one embodiment is capable of referring to a storage body storing a machine learning model that has been learned using learning data, wherein the learning data has as input a characteristic value of a signal regarding the frictional force between a polishing component and a substrate during polishing, or a characteristic value of the temperature of the polishing component or the substrate during polishing, and outputs data regarding the film thickness of the substrate after polishing, or a parameter regarding the product yield contained in the substrate after polishing. The polishing device is characterized in that it comprises: a polishing table provided with a polishing component and configured to rotate; a polishing head opposite to the polishing table and configured to rotate, and capable of mounting a substrate on a surface opposite to the polishing table; a control unit that controls the polishing head and the polishing table on which the substrate is mounted while pressing the substrate against the polishing component to polish the substrate; and a processor that generates a characteristic value based on the signal regarding the frictional force between the polishing component and the substrate during polishing, or the temperature of the polishing component or the target substrate during polishing, and inputs the generated characteristic value into the machine learning model that has been learned, thereby outputting either the data regarding the film thickness of the substrate after polishing or the parameter regarding the product yield contained in the substrate after polishing as an estimated value.

[0007] An information processing system in one embodiment is capable of referring to a storage body storing a machine learning model that has completed learning using learning data, wherein the learning data takes as input a characteristic value of a signal regarding the friction force between a grinding part and a substrate during grinding, or a characteristic value of the temperature of the grinding part or the substrate during grinding, and outputs data regarding the film thickness of the substrate after grinding, or a parameter regarding the product qualification rate contained in the substrate after grinding. The information processing system is characterized in that it comprises: a generation unit that generates a characteristic value based on a signal regarding the friction force between the grinding part and the substrate during grinding, or the temperature of the grinding part or the target substrate during grinding; and an estimation unit that inputs the generated characteristic value into the machine learning model that has completed learning, thereby outputting any one of the data regarding the film thickness of the substrate after grinding or the parameter regarding the product qualification rate contained in the substrate after grinding as an estimated value.

[0008] A polishing method in one embodiment comprises polishing a substrate by a polishing device, wherein the polishing device is capable of referring to a storage body storing a machine learning model that has completed learning using learning data, wherein the learning data has as input a characteristic value of a signal regarding the friction force between a polishing component and a substrate during polishing, or a characteristic value of the temperature of the polishing component or the substrate during polishing, and outputs data regarding the film thickness of the substrate after polishing, or a parameter regarding the product qualification rate contained in the substrate after polishing. The polishing method is characterized in that the substrate is polished by pressing the substrate against the polishing component while rotating the polishing head and the polishing table on which the substrate is mounted, measuring a characteristic value by generating a signal regarding the friction force between the polishing component and the substrate during polishing, or the temperature of the polishing component or the target substrate during polishing, inputting the generated characteristic value into the machine learning model that has completed learning, and outputting data regarding the film thickness of the substrate after polishing, or any one of the parameters regarding the product qualification rate contained in the substrate after polishing, as an estimated value.

[0009] A program in one embodiment is used to enable a computer to perform the functions of the following elements, wherein the computer is capable of referring to a storage body storing a machine learning model that has completed learning using learning data, wherein the learning data has as input a characteristic value of a signal regarding the friction force between a grinding part and a substrate during grinding, or a characteristic value of the temperature of the grinding part or the substrate during grinding, and has as output data regarding the film thickness of the substrate after grinding, or a parameter regarding the product qualification rate contained in the substrate after grinding, wherein the element includes: a generating unit that generates a characteristic value based on the signal regarding the friction force between the grinding part and the substrate during grinding, or the temperature of the grinding part or the target substrate during grinding; and an estimating unit that inputs the generated characteristic value into the machine learning model that has completed learning, thereby outputting any one of the data regarding the film thickness of the substrate after grinding or the parameter regarding the product qualification rate contained in the substrate after grinding as an estimated value. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a schematic diagram of the information processing system according to the first embodiment.

[0011] Figure 2 This is a schematic diagram showing the overall structure of the polishing device according to the first embodiment.

[0012] Figure 3 This is a schematic diagram of the AI unit of the first embodiment.

[0013] Figure 4 This is an explanatory diagram of the correspondence between the polishing condition of the wafer and the TCM waveform.

[0014] Figure 5 This is a schematic diagram for explaining the waveform of the cut TCM.

[0015] Figure 6 This is a graph showing the correlation coefficient between the maximum value of the residual film and each parameter.

[0016] Figure 7 This is a schematic diagram showing an example of the outline of LightGBM.

[0017] Figure 8 This is a schematic diagram illustrating an example of a learning process and an estimation process.

[0018] Figure 9 This is a graph comparing the measured value of the maximum film thickness and the AI estimated value in the first embodiment.

[0019] Figure 10 This is a graph comparing the measured value of the average film thickness and the AI estimated value in the first embodiment.

[0020] Figure 11 This is a graph comparing measured values and AI estimated values in the film thickness range in the first embodiment.

[0021] Figure 12 This is a flowchart showing an example of a process for stopping the processing of a subsequent substrate that meets the polishing deterioration condition.

[0022] Figure 13 This is a flowchart showing an example of a process for measuring film thickness using a film thickness measuring device within the apparatus when polishing deterioration conditions are met.

[0023] Figure 14 This is a flowchart showing an example of a process for measuring film thickness using a film thickness measuring device within the apparatus when polishing deterioration conditions are met.

[0024] Figure 15 This is a flowchart showing an example of a process for issuing a warning urging maintenance when polishing deterioration conditions are met.

[0025] Figure 16 This is a schematic diagram showing the overall structure of a polishing system according to a second embodiment.

[0026] Figure 17 This is a schematic diagram showing the overall structure of a polishing system according to a third embodiment. DETAILED DESCRIPTION

[0027] Various embodiments are described below with reference to the accompanying drawings. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters and repeated descriptions of substantially identical components may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art.

[0028] The first mode of the present technology is a grinding device that can refer to a storage body storing a machine learning model that has completed learning using learning data, wherein the learning data takes as input a characteristic quantity of a signal regarding the friction force between a grinding component and a substrate during grinding or a characteristic quantity of the temperature of the grinding component or the substrate during grinding, and takes as output data regarding the film thickness of the substrate after grinding or a parameter regarding the product qualification rate contained in the substrate after grinding, and the grinding device is characterized in that it comprises: a grinding table provided with a grinding component and configured to be rotatable; a grinding head, which is opposite to the grinding table and configured to be rotatable, and and is capable of mounting a substrate on a surface opposite to the polishing table; a control unit, which controls the polishing head and the polishing table on which the substrate is mounted to rotate while pressing the substrate against the polishing component to polish the substrate; and a processor, which generates a feature value based on a signal about the friction force between the polishing component and the substrate during polishing, or the temperature of the polishing component or the target substrate during polishing, and inputs the generated feature value into the machine learning model that has completed learning, thereby outputting data about the film thickness of the substrate after polishing or any one of the parameters about the product qualification rate contained in the substrate after polishing as an estimated value.

[0029] When this configuration is adopted, since the polishing device obtains data on the film thickness of the polished substrate or an estimated value of a parameter related to the product pass rate contained in the polished substrate during polishing, the state of the substrate after polishing can be predicted even without measuring the film thickness. This allows the state of the substrate after polishing to be understood even without measuring the film thickness, and reduces the number of film thickness measurements, thereby avoiding the omission of defective products and increasing the throughput. In this way, by omitting film thickness measurements during normal polishing, the throughput can be increased. Furthermore, by estimating parameters related to the pass rate, defects can be detected or predicted. Furthermore, by updating the polishing parameters based on the parameters related to the pass rate, the pass rate can be improved.

[0030] A polishing apparatus according to a second aspect of the present technology is the polishing apparatus according to the first aspect, wherein the processor stops processing of a subsequent substrate when the output estimated value satisfies a predetermined polishing deterioration condition.

[0031] With this configuration, since subsequent substrate processing is stopped when the polishing state deteriorates, maintenance such as replacement of polishing components can be performed, thereby preventing the polishing state from deteriorating further.

[0032] The third type of polishing device of the present technology is such as the polishing device of the first or second type, wherein the device is provided with a film thickness measuring device, which measures the film thickness of the substrate. When the output estimated value meets the predetermined polishing deterioration condition, the processor controls the film thickness measuring device to measure the film thickness of the target substrate after polishing. When the output estimated value does not meet the predetermined polishing deterioration condition, the processor controls the film thickness measuring device not to measure the film thickness of the target substrate after polishing.

[0033] When this configuration is adopted, since the film thickness of the substrate is measured when the polishing state deteriorates, it is possible to determine whether the polishing is progressing well. If the polishing state has not deteriorated, the processing throughput can be improved by not measuring the film thickness of the substrate.

[0034] A polishing apparatus according to a fourth aspect of the present technology is the polishing apparatus according to any one of the first to third aspects, wherein the processor outputs the maintenance timing using a tendency of the estimated value output for substrates polished at a plurality of different times.

[0035] With this configuration, the timing of deterioration of the polishing state can be predicted, and maintenance such as replacement of the polishing components can be performed at that timing, thereby preventing further deterioration of the polishing state.

[0036] A polishing device according to a fifth aspect of the present technology is the polishing device according to any one of the first to fourth aspects, wherein the processor performs control so as to issue a warning urging maintenance when the output estimated value satisfies a predetermined polishing deterioration condition.

[0037] With this configuration, when the polishing condition deteriorates, maintenance such as replacement of the polishing components can be performed, thereby preventing the polishing condition from deteriorating further.

[0038] The sixth mode of the present technology is a grinding device such as any one of the first to fifth modes, wherein the processor adjusts the grinding conditions of subsequent substrates in a manner such that the processor obtains desired data on the film thickness of the substrate after grinding or desired parameters on the product qualification rate contained in the substrate after grinding based on the output estimated value.

[0039] With this configuration, since the polishing conditions of the subsequent substrate can be changed to a good polishing state, the good polishing state can be maintained for a longer period of time.

[0040] A polishing device according to a seventh aspect of the present technology is the polishing device according to any one of the first to sixth aspects, wherein the processor uses the feature value in operation of the polishing device to re-learn the machine learning model.

[0041] When this configuration is adopted, the estimation accuracy can be improved.

[0042] An information processing system according to an eighth embodiment of the present technology is capable of referring to a storage body storing a machine learning model that has completed learning using learning data, wherein the learning data takes as input a characteristic value of a signal regarding the friction force between a grinding part and a substrate during grinding, or a characteristic value of the temperature of a grinding part or a substrate during grinding, and outputs data regarding the film thickness of the substrate after grinding, or a parameter regarding the product qualification rate contained in the substrate after grinding. The information processing system is characterized in that it comprises: a generating unit that generates a characteristic value based on a signal regarding the friction force between a grinding part and a substrate during grinding, or a temperature of a grinding part or a target substrate during grinding; and an estimating unit that inputs the generated characteristic value into the machine learning model that has completed learning, thereby outputting as an estimated value any one of the data regarding the film thickness of the substrate after grinding or the parameter regarding the product qualification rate contained in the substrate after grinding.

[0043] With this configuration, data regarding the film thickness of the polished substrate, or estimated values of parameters related to the product yield contained in the polished substrate, are obtained during polishing by the polishing apparatus. Therefore, even without measuring the film thickness, the substrate's post-polishing condition can be predicted. This allows for understanding the substrate's post-polishing condition even without measuring the film thickness, and reduces the number of film thickness measurements required, thereby preventing the omission of defective products and improving throughput.

[0044] A ninth embodiment of the present technology is a method for polishing a substrate by a polishing device, wherein the polishing device is capable of referring to a storage body storing a machine learning model that has completed learning using learning data, wherein the learning data takes as input a characteristic value of a signal regarding the friction force between a polishing component and a substrate during polishing or a characteristic value of the temperature of the polishing component or the substrate during polishing, and outputs data regarding the film thickness of the substrate after polishing or a parameter regarding the product qualification rate contained in the substrate after polishing. The polishing method is characterized in that the substrate is polished by pressing the substrate against the polishing component while rotating the polishing head and the polishing table on which the substrate is mounted, measuring a characteristic value by generating a signal regarding the friction force between the polishing component and the substrate during polishing or the temperature of the polishing component or the target substrate during polishing, inputting the generated characteristic value into the machine learning model that has completed learning, and outputting data regarding the film thickness of the substrate after polishing or any one of the parameters regarding the product qualification rate contained in the substrate after polishing as an estimated value.

[0045] With this configuration, data regarding the film thickness of the polished substrate, or estimated values of parameters related to the product yield contained in the polished substrate, are obtained during polishing by the polishing apparatus. Therefore, even without measuring the film thickness, the substrate's post-polishing condition can be predicted. This allows for understanding the substrate's post-polishing condition even without measuring the film thickness, and reduces the number of film thickness measurements required, thereby preventing the omission of defective products and improving throughput.

[0046] The program of the tenth mode of the present technology is used to enable a computer to perform the functions of the following elements, which are capable of referring to a storage body storing a machine learning model that has completed learning using learning data, and the learning data takes as input the characteristic quantity of a signal about the friction force between the grinding part and the substrate during grinding or the characteristic quantity of the temperature of the grinding part or the substrate during grinding, and takes as output data about the film thickness of the substrate after grinding or a parameter about the product qualification rate contained in the substrate after grinding, and the elements include: a generating unit, which generates a characteristic quantity based on the signal about the friction force between the grinding part and the substrate during grinding, or the temperature of the grinding part or the object substrate during grinding; and an estimating unit, which inputs the generated characteristic quantity into the machine learning model that has completed learning, thereby outputting any one of the data about the film thickness of the substrate after grinding or the parameter about the product qualification rate contained in the substrate after grinding as an estimated value.

[0047] With this configuration, data regarding the film thickness of the polished substrate, or estimated values of parameters related to the product yield contained in the polished substrate, are obtained during polishing by the polishing apparatus. Therefore, even without measuring the film thickness, the post-polishing state of the substrate can be predicted. This allows for understanding the post-polishing state of the substrate even without measuring the film thickness, and reduces the number of film thickness measurements required, thereby preventing the omission of defective products and improving processing throughput.

[0048] In addition to the above-mentioned problems, there is also the problem of taking time to judge the grinding condition (such as the condition of the workbench).

[0049] Various embodiments use changes in the monitoring waveform during polishing to estimate the polishing status, the film thickness after polishing (also known as residual film thickness), statistical values of the film thickness (average, maximum, or minimum, etc.), or the film thickness profile (also known as film thickness distribution). This allows for timely estimation and management of polishing quality / poor performance and polishing conditions (e.g., table conditions). Therefore, if a polishing quality is poor, subsequent polishing can be omitted and the table conditions can be adjusted. This reduces the number of samples with poor polishing. Various embodiments are described using a wafer as an example of a substrate.

[0050] <First embodiment>

[0051] First, the first embodiment will be described. Figure 1 This is a schematic diagram of the information processing system of the first embodiment. Figure 1 As shown, the information processing system S1 of the first embodiment includes a loading / unloading unit 2 , two polishing devices 10 as an example, a cleaning unit 5 , and a film thickness measuring device 6 .

[0052] The loading / unloading section 2 is equipped with two or more (four in this embodiment) front loading sections 20 for loading and storing wafer cassettes for multiple wafers. The front loading section 20 can be equipped with an open cassette, SMIF (Standard Manufacturing Interface), or FOUP (Front Opening Unified Pod). Here, SMIF and FOUP are cassettes that store wafers inside, and are sealed containers that can maintain an environment independent of the external space by being covered with partition walls. As an example, a FOUP 21 is mounted in one of the front loading sections 20. The wafers are transported from the loading / unloading section 2 to the polishing device 10 by a transfer robot 22 (see Patent Document 1).

[0053] The film thickness measuring device 6 measures the film thickness or the profile of the film thickness (also referred to as film thickness distribution) of the substrate (wafer in this case). The film thickness measuring device 6 is, for example, an optical film thickness measuring device (also referred to as ITM).

[0054] The polishing device 10 includes an AI unit 4. The AI unit 4 outputs data on the film thickness of the substrate after polishing, a profile statistical value of the film thickness of the substrate after polishing, or a parameter related to the product qualification rate (for example, qualification rate) contained in the substrate after polishing as an estimated value. Then, when the estimated value exceeds a predetermined polishing normal condition or satisfies a predetermined polishing deterioration condition, the AI unit 4 causes the film thickness measuring device 6 to measure the film thickness of the target substrate after polishing. For example, Figure 1 When the estimated value of the wafer W1 exceeds the predetermined polishing normal condition or meets the predetermined polishing deterioration condition, as shown by arrow A1, the wafer W1 is cleaned by the cleaning unit 5 and the film thickness is measured by the film thickness measuring device 6. Figure 1 When the estimated value of chip W2 meets the predetermined normal polishing condition, or does not meet the predetermined deteriorated polishing condition, as shown by arrow A2, after chip W2 is cleaned by cleaning unit 5, the film thickness is not measured by film thickness measuring device 6, but returned to FOUP21.

[0055] In addition, when normal polishing conditions are used, the AI unit 4 may be made to learn normal data, and when deteriorated polishing conditions are used, it may be made to learn bad data. Alternatively, the ratio between normal data and bad data may be determined so as to learn normal data and bad data.

[0056] The output of the AI unit 4 can also be classified into three categories: normal, defective, and near defective. When the output is near defective, the film thickness is measured by the film thickness measuring device 6.

[0057] Figure 9 For example, when the AI output estimated value never exceeds the lower limit threshold value, or suddenly exceeds the upper limit threshold value, the AI unit 4 may determine that the value is close to failure and measure the film thickness.

[0058] In addition to this, or alternatively, when the polishing time taken by the AI part 4 exceeds a normal range, it can be determined as being close to defective and the film thickness can be measured.

[0059] Furthermore, the output of the AI unit 4 may be divided into two types: normal and nearly defective.

[0060] Figure 2 1 is a schematic diagram showing the overall structure of the grinding device of the first embodiment. Figure 2 As shown, the polishing device 10 includes: a polishing table 100; and a polishing head 1 as a substrate holding device that holds a substrate to be polished (here, a wafer) and presses the polishing surface on the polishing table 100. The polishing head 1 is also called a top ring. The polishing table 100 is connected to a table rotation motor 102 arranged below it via a table shaft 100a. The polishing table 100 rotates around the table shaft 100a when the table rotation motor 102 rotates. A polishing pad 101 as a polishing member is attached to the top of the polishing table 100. The surface of the polishing pad 101 constitutes a polishing surface 101a for polishing a semiconductor wafer W. Therefore, the polishing device 10 includes: a polishing table 100 provided with a polishing member (here, an example, the polishing pad 101) and configured to be rotatable; and a polishing head 1 that is configured to be rotatable relative to the polishing table 100, and the polishing head 1 can mount a substrate (here, a wafer) on the surface opposite to the polishing table 100.

[0061] A polishing liquid supply nozzle 60 is provided above the polishing table 100 . Polishing liquid (polishing slurry) Q is supplied from the polishing liquid supply nozzle 60 onto the polishing pad 101 on the polishing table 100 .

[0062] The polishing head 1 is basically composed of a top ring body 2 that presses the semiconductor wafer W against the polishing surface 101a; and a retaining ring 3 that serves as a fixing member, holding the outer edge of the semiconductor wafer W to prevent it from being ejected from the polishing head 1. The polishing head 1 is connected to a top ring shaft 111. This top ring shaft 111 moves vertically relative to the top ring head 110 via a vertical movement mechanism 124. The vertical positioning of the polishing head 1 is achieved by raising and lowering the entire polishing head 1 relative to the top ring head 110 as the top ring shaft 111 moves vertically. A rotary joint 26 is attached to the upper end of the top ring shaft 111.

[0063] The vertical movement mechanism 124, which vertically moves the top ring shaft 111 and the polishing head 1, includes a bridge 128 that rotatably supports the top ring shaft 111 via a bearing 126; a ball screw 132 mounted on the bridge 128; a support base 129 supported by support columns 130; and a servo motor 138 mounted on the support base 129. The support base 129, which supports the servo motor 138, is fixed to the top ring head 110 via support columns 130.

[0064] The ball screw 132 includes a screw shaft 132a connected to a servo motor 138 and a nut 132b threadedly engaged with the screw shaft 132a. When the servo motor 138 is driven, the bridge 128 moves vertically via the ball screw 132. This in turn causes the top ring shaft 111 and the polishing head 1, which move vertically integrally with the bridge 128, to move vertically.

[0065] In addition, if Figure 2 As shown, by rotating the top ring rotation motor 114 , the rotation cylinder 112 and the top ring shaft 111 rotate integrally via the timing pulley 116 , the timing belt 115 , and the timing pulley 113 , and the polishing head 1 rotates.

[0066] The top ring head 110 is supported by a top ring head shaft 117 rotatably supported by a frame (not shown). The polishing apparatus 10 includes a control unit 500 connected to various devices within the apparatus, such as the top ring rotation motor 114, the servo motor 138, and the table rotation motor 102, via control lines to control each device. The control unit 500 rotates the polishing head 1 and the polishing table 100, to which a substrate is attached, while simultaneously pressing the substrate against a polishing member (here, the polishing pad 101) to polish the substrate.

[0067] The inputs to the machine learning model described later include table rotation, head rotation, and rotation of a motor (not shown) for shaking the top ring head 110. However, one or more sensor detection values (for example, motor current values) or a calculated torque value calculated from the sensor detection values may also be used.

[0068] The polishing device 10 includes an AI unit 4 connected to a control unit 500 via wiring. Figure 3 This is a schematic diagram of the AI section of the first embodiment. Figure 3 As shown, the AI unit 4 is, for example, a computer, and includes a storage 41 , a memory 42 , an input unit 43 , an output unit 44 , and a processor 45 .

[0069] The memory 41 stores a machine learning model that uses as input a signal characteristic value of the frictional force between the polishing member (here, the polishing pad 101) and the substrate during polishing, and outputs data regarding the film thickness of the polished substrate, or a profile statistic of the film thickness of the polished substrate, or a parameter related to the product yield contained in the polished substrate as learning data. Furthermore, the memory 41 stores a program that is read and executed by the processor 45.

[0070] Here, the signal regarding the frictional force between the polishing member and the substrate is, for example, a signal representing the current value (Table Current Monitor: also referred to as TCM) used to calculate the torque of the table rotation motor 102 during polishing. Here, the signal regarding the frictional force between the polishing member and the substrate may also be a calculated value of the torque converted from the motor current value. Alternatively, the signal regarding the frictional force between the polishing member and the substrate may be a signal representing the driving current value of the top ring rotation motor 114 that rotates the polishing head 1, or a signal representing the driving current value of a motor (not shown) that rotates the top ring head 110 (i.e., the top ring head shaft 117).

[0071] Furthermore, the polishing apparatus 10 may include a load sensor for measuring the frictional force between the polishing member and the substrate. In this case, the signal regarding the frictional force between the polishing member and the substrate may also be the signal from the load sensor. The polishing apparatus 10 may also include a strain sensor for measuring the strain of the substrate. In this case, the signal regarding the frictional force between the polishing member and the substrate may also be the signal from the strain sensor.

[0072] The memory 42 is a medium for temporarily storing information.

[0073] The input unit 43 receives information from the control unit 500 and outputs the information to the processor 45 .

[0074] The output unit 44 receives information from the processor 45 and outputs the information to the control unit 500 .

[0075] The processor 45 reads and executes a program from the storage 41 , thereby fulfilling the functions of the generation unit 451 , the estimation unit 452 , and the determination unit 453 .

[0076] Generator 451 generates a feature value from a signal regarding the frictional force between the polishing member and the substrate during polishing. Here, "during polishing" refers to, for example, the period during which the substrate is pressed against the polishing member while the polishing head 1 and polishing table 100, with the substrate mounted thereon, rotate. This process will be described in detail below.

[0077] The estimation unit 452 inputs the feature value generated by the generation unit 451 into the machine learning model that has completed the learning, thereby outputting data on the film thickness of the substrate after grinding, or any one of the parameters on the product qualification rate contained in the substrate after grinding as an estimated value. This process is described in detail later. Here, the data on the film thickness of the substrate after grinding is, for example, the film thickness of the substrate after grinding, the statistical value of the film thickness profile of the substrate after grinding (such as the average value, maximum value, minimum value, variation range, standard deviation, etc. of the film thickness distribution), the film thickness profile of the substrate after grinding, etc. Here, the film thickness profile is a group of film thickness data (a combination of XY coordinates and film thickness) measured at multiple points by changing the position in the wafer.

[0078] There are multiple chips in a wafer, and defective determination is performed on each chip to calculate parameters related to the chip yield rate in the wafer. The product yield rate included in the polished substrate mentioned above is, for example, the chip yield rate in the wafer.

[0079] Figure 4 This is an explanatory diagram of the correspondence between the polishing condition of the wafer and the TCM waveform. Figure 4 The graph depicts a waveform C1 representing the time-dependent change in TCM, with the vertical axis representing the torque current value (TCM) of the table rotation motor 102 during polishing and the horizontal axis representing time [ms]. The graph also shows a change in the frictional force with the polishing pad 101, as the ratio of the exposed film species changes. Consequently, the TCM value also changes accordingly.

[0080] like Figure 4 As shown, wafer W includes a polished layer 51 mounted opposite a polishing pad 101 and a lower layer 52 disposed above the polished layer 51. The polished layer 51 is removed by the frictional force during polishing. At point P1 on waveform C1, the polished layer 51 is not significantly removed. After a period of time, a portion of the lower layer 52 is exposed at point P2 on waveform C1. Further, after a period of time, the lower layer 52 is fully exposed at point P3 on waveform C1. After the entire lower layer 52 is exposed, the table rotation motor 102 stops, and polishing is complete.

[0081] Figure 4 The portion indicated by arrow A15 is over-ground, causing the lengths of arrows A12 and A13 to be correspondingly shorter than the lengths of arrows A11 and A14.

[0082] The inventors of this application have found that when a film is polished unevenly, the timing of the underlying film being exposed varies within the wafer surface, so the TCM signal ( Figure 5 Therefore, this embodiment generates a feature value based on the TCM signal during a predetermined period before the polishing end point.

[0083] use Figure 5 The following describes how to cut out a portion of the signal and use it to calculate a feature value from the TCM signal. Figure 5 This is a schematic diagram for explaining the waveform of the cut TCM. Figure 5 The vertical axis is TCM and the horizontal axis is time. Figure 5As shown, the waveform of the entire TCM is represented in the graph G1. In the graph G1, the cutout area R1 is the graph G2. Therefore, the generation unit 451 of the present embodiment extracts data of a predetermined time range from the TCM, and calculates a feature quantity from the extracted data. The feature quantity is, for example, the value itself, the differential value, the statistic of the moving average, etc. (for example, the maximum value, the minimum value, the standard deviation, the dispersion, the average value, the central value, the dispersion, the kurtosis, or the strain degree, etc.) of the calculation period (for example, the entire period of the extracted data, a part of the period T1, a part of the period T2 after the part of the period T1, etc.) in the extracted data. The kurtosis here is a number that represents the sharpness of the frequency distribution, and is calculated using a conventional calculation method. The strain degree represents the degree to which the data is not symmetrically distributed around the average value, and is calculated using a conventional calculation method.

[0084] use Figure 6 An example of a feature amount will be described. Figure 6 This is a graph showing the correlation coefficient between the maximum value of the residual film and each parameter. Figure 6 In FIG. 1 , the vertical axis represents each feature quantity, and the horizontal axis represents the correlation coefficient. The correlation coefficient between each feature quantity on the vertical axis and the maximum value of the residual film is 0.5 or more, indicating that each feature quantity is related to each other.

[0085] Figure 6 In the Figure 5 The minimum value of the differential of the 25 data moving average of TCM in the predetermined period T1 during the period of the graph G2, All_min-d_r25 is Figure 5 The minimum value of the differential of the 25 data moving average of TCM in the entire period of the graph G2, the characteristic value T1_min-d_r10 is Figure 5 The minimum value of the differential of the 10-data moving average of TCM in the predetermined period T1 during the period of the graph G2, All_min-d_r10 is Figure 5 The minimum value of the differential of the 10-data moving average of TCM during the entire period of the graph G2. Figure 5 The total value of TCM in period T2 after period T1 in the graph G2 is T1_std-d_r10. Figure 5 The standard deviation of the differential of the moving average of 10 data of TCM in the predetermined period T1 during the period of the graph G2, All_skew-d-r10 is Figure 5 The curve graph G2 shows the differential strain of the moving average of 10 data of TCM during the entire period. All_skew-d-r25 is Figure 5 The graph G2 shows the differential strain of 25 moving average data of TCM in all periods, where T1_len is the number of data.

[0086] Alternatively, the number of types of parameters to be used may be determined from upper-layer parameters having high correlation coefficients (for example, 10 upper-layer parameters), or the applicable conditions may be determined.

[0087] The applicable conditions may be, for example, conditions using parameters equal to or greater than the average value of the correlation coefficient, or conditions using parameters equal to or greater than a value obtained by adding a standard deviation σ to the average value of the correlation coefficient.

[0088] In addition, All_range-d_r25 is Figure 5 The curve G2 shows the differential range of the moving average of 25 data of TCM during the entire period, T1_var-d_r25 is Figure 5 The graph G2 shows the dispersion of the differential of the moving average of 25 data of TCM in the predetermined period T1. All_range-d_r10 is Figure 5 The range of the differential of the moving average of 10 data of TCM in the whole period of the graph G2. Figure 5 The total TCM of the predetermined period T1 in the period of the graph G2.

[0089] T1_mean-d_r10 is Figure 5 The average of the differentials of the moving average of 10 data of TCM in the predetermined period T1 during the period of the graph G2, T1_max is Figure 5 The maximum value of TCM of the predetermined period T1 during the period of the graph G2. Figure 5 The maximum value of TCM during all periods of the graph G2. All_std-d_r10 is Figure 5 The standard deviation of the differential of the moving average of 10 data of TCM during the entire period of graph G2. T1_mean is Figure 5 The average value of TCM of the predetermined period T1 during the period of the graph G2. All_std-d_r25 is Figure 5 The standard deviation of the differential of the moving average of 25 TCM data for all periods of graph G2. All_len is the number of data.

[0090] T1_range-d_r10 is Figure 5 The range of the differential of the moving average of 10 data in the predetermined period T1 during the period of the graph G2 is T1_mean-d_r25. Figure 5 The average of the differentials of the moving average of 25 data in the predetermined period T1 during the period of the graph G2, T1_range-d_r25 is Figure 5 The range of the differential of the moving average of 25 data in the predetermined period T1 during the period of the graph G2 is All_var-d_r10. Figure 5The graph G2 shows the dispersion of the differential of the moving average of 10 data over the entire period, T2_mean-d_r5. Figure 5 In the graph G2, the average of the differentials of the moving average of the five TCM data in the period T2 after the period T1 is obtained. All_var-d_r25 is Figure 5 The graph G2 shows the dispersion of the differential of the 25 moving averages of TCM data for all periods. All_mean is Figure 5 The average TCM of the entire period of the graph G2, T2_mean is Figure 5 The average TCM of the period T2 after the period T1 in the graph G2 is All_skew. Figure 5 The strain of TCM during the entire period of the graph G2, T2_min is Figure 5 The minimum value of TCM is in the period T2 following the period T1 among the periods of the graph G2.

[0091] The AI (artificial intelligence) model used by the estimation unit 452 of the AI unit 4 may be, for example, LightGBM (Light Gradient Boosting Machine) as described in Non-Patent Document 1. LightGBM is a machine learning model based on a decision tree.

[0092] Figure 7 This is a schematic diagram showing an example of the outline of LightGBM. Figure 7 An example is that there are a first decision tree M1, a second decision tree M2, and a third decision tree M3. The first decision tree M1 is used for model training to evaluate the inference result. The "error" between the inference result of the first decision tree M1 and the actual value is used as training data to perform training on the second decision tree M2. Similarly, the "error" between the inference result of the second decision tree M2 and the actual value is used as training data to perform training on the third decision tree M3. After the training is completed, when the feature value is input into the first decision tree M1, the estimated value is output from the third decision tree M3. In the training process of gradient boosting, the decision tree processing method has the so-called "Leaf-wise tree growth" method, and grows according to the leaf of the decision tree.

[0093] Figure 8 This is a schematic diagram illustrating an example of the learning process and the estimation process. Figure 8As shown, the learning process is that the AI unit 4 uses learning data that takes the feature value as input and outputs the estimated film thickness value to learn the machine learning model. Continuing, in the estimation process, when the feature value is input into the learned machine learning model, the machine learning model outputs the estimated film thickness value, for example. In this way, the AI unit 4 outputs the estimated film thickness value for the input feature value using the learned machine learning model. The AI unit 4 can then compare the estimated film thickness value with a set threshold value, for example, to determine whether it is normal or close to defective. If it is determined to be close to defective, the method of transporting the wafer to the film thickness measuring device is controlled. In this way, when it is determined to be close to defective, the film thickness of the wafer is actually measured.

[0094] This embodiment divides data obtained from past polishing into learning data and test data. Only the learning data is used to train the AI. The AI then makes an estimate based on the entire data set and compares it with the actual measured values. The following describes the comparison results for the maximum film thickness, average film thickness, and film thickness range (also referred to as the film thickness range).

[0095] Figure 9 In the first embodiment, a graph is provided comparing the measured value of the maximum film thickness with the AI estimated value. Figure 9 As shown in Figure G11, the vertical axis represents the AI-estimated maximum film thickness, normalized by the maximum film thickness within the allowable limit, while the horizontal axis represents the measured maximum film thickness, normalized by the maximum film thickness within the allowable limit. The AI-estimated values are distributed near the correct answer line, indicating effective estimation. The vertical axis of Figure G12 represents data statistics, while the horizontal axis represents the estimation error (= measured maximum film thickness - AI-estimated maximum film thickness). "Train" represents training data, and "Test" represents testing data. The estimation errors are all within the specified range.

[0096] For example, when the AI estimated value of the maximum film thickness normalized by the maximum film thickness within the allowable limit exceeds a first threshold value, the determination unit 453 may also determine that the film thickness must be measured. Thus, when the first threshold value is exceeded, since there is a possibility that a wafer that has been cut and retained due to the maximum film thickness exceeding the allowable limit is included, the film thickness may also be controlled to be required to be measured. Thus, by using the AI estimated value as a determination value, the conditions for measuring the film thickness can be set without missing defective products. In this example, it can be seen that by using the AI estimated value, only about 25% of the substrates need to be measured. Specifically, the determination unit 453 may also control one or more robots (e.g., conveyor 7, transport robot 22, transport robot 53, refer to patent document 1) in such a way that the wafer is moved to the film thickness measuring device 6 after grinding.

[0097] Figure 10 This is a graph comparing the measured value of the average film thickness and the AI estimated value in the first embodiment. Figure 10As shown, the vertical axis of graph G21 represents the AI-estimated value of the normalized average film thickness, and the horizontal axis represents the measured value of the normalized average film thickness. The AI-estimated values are distributed near the correct answer line, indicating that the estimation is effective. The vertical axis of graph G22 represents the data statistics, and the horizontal axis represents the estimation error (= measured value of the average film thickness - AI-estimated value of the average film thickness). Train represents the training data, and Test represents the testing data.

[0098] Figure 11 This is a graph comparing the measured values of the film thickness range and the AI estimated values in the first embodiment. Figure 11 As shown, the vertical axis of graph G31 represents the AI estimated values for the normalized film thickness range, and the horizontal axis represents the measured values for the normalized film thickness range. The AI estimated values are distributed near the correct answer line, indicating that the estimation is effective. The vertical axis of graph G32 represents the data statistics, and the horizontal axis represents the estimation error (= measured values for the film thickness range - AI estimated values for the film thickness range). Train represents training data, and Test represents testing data.

[0099] Continue to use Figure 12 An example of a process of stopping subsequent processing of a substrate when the estimated value output from the estimating unit 452 satisfies a predetermined polishing deterioration condition will be described. Figure 12 This is a flowchart showing an example of a process for stopping the processing of subsequent substrates when polishing conditions deteriorate.

[0100] Here, the determination unit 453 determines whether the estimated value output by the estimation unit 452 satisfies a predetermined polishing deterioration condition. For example, when the polishing deterioration condition is "the estimated value exceeds a set range," the determination unit 453 determines whether the estimated value output by the estimation unit 452 exceeds the set range. Figures 12 to 14 An example of a polishing deterioration condition is the condition that "the estimated value of the standard deviation of the film thickness profile is greater than or equal to a set threshold value." The determination unit 453 determines whether the estimated value of the standard deviation of the film thickness profile output by the estimation unit 452 is greater than or equal to the set threshold value.

[0101] (Step S110 ) First, the processor 45 obtains a TCM signal when the wafer is polished.

[0102] (Step S120) Next, the generation unit 451 calculates a feature value from the acquired TCM signal.

[0103] (Step S130) Next, the estimating unit 452 inputs the feature value into the learned machine learning model stored in the memory bank 41, and outputs, for example, an estimated value of the standard deviation of the film thickness profile. An example of a learned machine learning model is a model that has learned learning data that takes the feature value of the TCM signal as input and outputs the standard deviation of the film thickness profile.

[0104] (Step S140) Next, the determination unit 453 determines whether the estimated value of the standard deviation of the film thickness profile is greater than or equal to a set threshold. If the standard deviation of the film thickness profile is not greater than or equal to the set threshold (i.e., the standard deviation of the profile is less than the set threshold), the process returns to step S110 and the subsequent processes are repeated.

[0105] (Step S150) If the estimated value of the standard deviation of the film thickness profile is determined to be equal to or greater than the set threshold value in step S140, the determination unit 453 controls the control unit 500 to stop processing the subsequent wafer.

[0106] Therefore, the processor 45 may also stop processing of subsequent substrates when the estimated value output from the estimating unit 452 satisfies a predetermined polishing deterioration condition. Thus, since processing of subsequent substrates is stopped when the polishing condition deteriorates, maintenance such as replacement of polishing components can be performed, thereby preventing further deterioration of the polishing condition.

[0107] Next, a description will be given of a process in which, when the estimated value satisfies a predetermined polishing deterioration condition, the film thickness is measured by a film thickness measuring device within or outside the apparatus. Figure 13 This is a flowchart showing an example of a process for measuring film thickness using a film thickness measuring device within the apparatus when polishing deterioration conditions are met.

[0108] (Step S210) First, the processor 45 obtains a TCM signal when polishing the wafer.

[0109] (Step S220) Next, the generation unit 451 calculates a feature value from the acquired TCM signal.

[0110] (Step S230) Next, the estimating unit 452 inputs the feature value into the learned machine learning model stored in the memory bank 41, and outputs, for example, an estimated value of the standard deviation of the film thickness profile. An example of a learned machine learning model is a model that has learned learning data that takes the feature value of the TCM signal as input and outputs the standard deviation of the film thickness profile.

[0111] (Step S240 ) Next, the determination unit 453 determines whether the estimated value of the standard deviation of the film thickness profile is equal to or greater than a set threshold value.

[0112] (Step S250) When it is determined in step S240 that the estimated value of the standard deviation of the film thickness profile is greater than the set threshold, the processor 45 controls one or more robots (for example, the conveyor 7, the transport robot 22, the transport robot 53, refer to patent document 1) in order to measure the film thickness of the polished chip with the film thickness measuring device 6.

[0113] (Step S260) In step S240, when the estimated value of the standard deviation of the film thickness profile is not greater than the set threshold value (that is, when the standard deviation of the profile is less than the set threshold value), the processor 45 controls one or more robots (for example, the conveyor 7, the transport robot 22, the transport robot 53, refer to patent document 1) in such a manner that the film thickness measuring device 6 does not perform measurement but returns the chip to the FOUP.

[0114] Therefore, when the estimated value output by the estimating unit 452 satisfies the predetermined polishing deterioration condition, the processor 45 controls the film thickness measuring device 6 to measure the film thickness of the target substrate after polishing. When the estimated value output does not satisfy the predetermined polishing deterioration condition, the processor 45 controls the film thickness measuring device 6 not to measure the film thickness of the target substrate after polishing. Thus, when the polishing condition deteriorates, the film thickness of the substrate is measured, thereby determining whether polishing is being performed effectively. When the polishing condition has not deteriorated, the throughput can be increased by not measuring the film thickness of the substrate.

[0115] Next, a description will be given of a process in which, when the estimated value satisfies a predetermined polishing deterioration condition, the film thickness is measured using a film thickness measuring device within or outside the apparatus. Figure 14 This is a flowchart showing an example of a process for measuring film thickness using a film thickness measuring device within the apparatus when polishing deterioration conditions are met.

[0116] (Step S310) First, the processor 45 obtains a TCM signal when polishing the wafer.

[0117] (Step S320) Next, the generation unit 451 calculates a feature value from the acquired TCM signal.

[0118] (Step S330) Next, the estimating unit 452 inputs the feature value into the learned machine learning model stored in the memory bank 41 and outputs, for example, an estimated value of the standard deviation of the film thickness profile. An example of a learned machine learning model is a model that has learned learning data that takes the feature value of the TCM signal as input and outputs the standard deviation of the film thickness profile.

[0119] (Step S340) Next, the determination unit 453 determines, for example, whether the estimated value of the standard deviation of the film thickness profile is greater than or equal to a set threshold. If the estimated value of the standard deviation of the film thickness profile is not greater than or equal to the set threshold (that is, if the standard deviation of the profile is less than the set threshold), the process returns to step S310 and the subsequent processes are repeated.

[0120] (Step S350) Furthermore, if the estimated value of the standard deviation of the film thickness profile is determined to be greater than a set threshold value in step S340, the processor 45 controls the output of a warning urging maintenance. This warning may be a voice message, or the content of the warning may be displayed on a display device. When a light source (e.g., PATLITE (registered trademark)) with multiple colors (e.g., red, yellow, and green) is provided, a PATLITE of a specific color (e.g., yellow) may be illuminated (or extinguished). Alternatively, vibration may be generated. Alternatively, the user of the polishing device 10 may be notified by sending an email to the user's address in a manner that automatically contacts the user. Any combination of these may also be used.

[0121] Therefore, the processor 45 controls the output of a maintenance warning when the estimated value output by the estimating unit 452 satisfies a predetermined polishing deterioration condition. This allows maintenance such as replacing polishing parts to be performed when the polishing condition deteriorates, thereby preventing the polishing condition from deteriorating further.

[0122] Continue to use Figure 15 This section explains how to issue a warning urging maintenance when grinding deterioration conditions are met. Figure 15 This is a flowchart showing an example of a process for issuing a warning urging maintenance when polishing deterioration conditions are met.

[0123] (Step S410) First, the processor 45 obtains a TCM signal when polishing the wafer.

[0124] (Step S420) Next, the generation unit 451 calculates a feature value from the acquired TCM signal.

[0125] (Step S430) Next, the estimating unit 452 inputs the feature value into the learned machine learning model stored in the memory bank 41 and outputs, for example, an estimated value of the standard deviation of the film thickness profile. An example of a learned machine learning model is a model that has learned learning data that takes the feature value of the TCM signal as input and outputs the standard deviation of the film thickness profile.

[0126] (Step S440 ) Next, the estimating unit 452 stores the estimated value in the memory bank 41 .

[0127] (Step S450) Next, the determination unit 453 determines whether a predetermined number of estimated values have been stored, for example. If the predetermined number of estimated values have not been stored, the process returns to step S410 and the subsequent processes are repeated.

[0128] (Step S460) Furthermore, if it is determined in step S450 that a predetermined number of estimated values have been stored, the processor 45 refers to the estimated values output for multiple substrates polished at different times stored in the memory 41 and outputs a maintenance timing using the trend of the estimated values output for the multiple substrates polished at different times. The maintenance timing may be output as "Maintenance recommended after 0 hours." The processor 45 may also notify the user of the polishing device of the maintenance timing. This automatically notifies the user of the maintenance timing. This notification method may also include displaying the maintenance timing on a web screen or application, or sending an email to the user's address.

[0129] Alternatively, the processor 45 can notify the user of the polishing device when the maintenance time arrives. This automatically notifies the user of the maintenance time. This notification method can also display the maintenance details on a web screen or application, or send an email to the user's address.

[0130] Specifically, for example, the processor 45 can store the estimated standard deviation of the film thickness profile at a set time interval, calculate the change in the estimated value per unit time by dividing the difference between the estimated values by the set time interval, and output the timing when the estimated value exceeds a set threshold as the maintenance timing. This allows the timing of polishing condition deterioration to be predicted, allowing maintenance such as polishing component replacement to be performed at that time, thereby preventing further deterioration of the polishing condition.

[0131] Furthermore, the processor 45 can also adjust the polishing conditions for subsequent substrates by obtaining desired data regarding the film thickness of the polished substrate or desired parameters regarding the product yield of the polished substrate based on the estimated value output by the estimating unit 452. This allows the polishing conditions of subsequent substrates to be modified to improve the polishing state, thereby maintaining a good polishing state for a longer period of time.

[0132] The processor 45 can also use the characteristic values of the polishing device during operation to re-learn the machine learning model, thereby improving the estimation accuracy.

[0133] As described above, the polishing device 10 of the first embodiment can refer to a memory 41 storing a machine learning model learned using learning data. The learning data takes as input a characteristic value of a signal regarding the frictional force between the polishing member and the substrate during polishing, or a characteristic value of the temperature of the polishing member or substrate during polishing, and outputs data regarding the film thickness of the polished substrate, or a parameter regarding the product yield contained in the polished substrate. The polishing device 10 includes: a polishing table 100 rotatably provided with a polishing member; a polishing head 1 rotatably facing the polishing table 100, and capable of mounting a substrate on a surface facing the polishing table 100; and a control unit 500 that controls the polishing head and the polishing table to rotate the substrate mounted thereon while pressing the substrate against the polishing member to polish the substrate.

[0134] The polishing device 10 includes a processor 45, which generates a feature value based on a signal about the friction force between the polishing part and the substrate during polishing, or the temperature of the polishing part or the target substrate during polishing, inputs the generated feature value into a machine learning model that has completed learning, and outputs data about the film thickness of the substrate after polishing or any one of the parameters about the product qualification rate contained in the substrate after polishing as an estimated value.

[0135] With this configuration, since the polishing device obtains data on the film thickness of the polished substrate or an estimated value of a parameter related to the product qualification rate contained in the polished substrate during polishing, the state of the substrate after polishing can be predicted even without measuring the film thickness. Thus, even without measuring the film thickness, the state of the substrate after polishing can be grasped, and the number of film thickness measurements can be reduced, thereby avoiding the omission of defective products and increasing the processing capacity. In this way, by omitting the film thickness measurement during normal polishing, the processing capacity can be increased. In addition, by estimating the parameters related to the qualification rate, defects can be detected or predicted. In addition, by updating the polishing parameters based on the parameters related to the qualification rate, the qualification rate can be improved.

[0136] Alternatively, the AI unit 4 can be installed in a gateway within the factory, with the polishing equipment connected to the gateway via a network line. This gateway is preferably located near the polishing equipment. When high-speed processing is required (e.g., when the sampling rate is less than 100ms), edge computing can be performed by the AI unit 4 in the polishing equipment or by the AI unit 4 installed in the gateway. The AI unit 4 in the polishing equipment can also be installed in the equipment PC or controller.

[0137] <Second embodiment>

[0138] The second embodiment will be described below. The polishing apparatus 10 of the first embodiment includes the AI unit 4. However, the second embodiment differs in that the AI unit 4 is not provided in the polishing apparatus but is provided in a factory management room or a clean room within the factory.

[0139] Figure 16 : is a schematic diagram showing the overall structure of the grinding system of the second embodiment. Figure 16 As shown, a polishing system S2 according to the second embodiment includes polishing apparatuses 10-1 to 3-N and an AI unit 4 located within the same factory as the polishing apparatuses 10-1 to 10-N or within a factory management office. The AI unit 4 and the polishing apparatuses 10-1 to 3-N can communicate via a local network NW1. The AI unit 4 is, for example, installed in a computer (e.g., a server or fog computing).

[0140] When the AI unit 4 is installed in the polishing device or gateway, the learned machine learning model can be executed through edge computing, enabling high-speed processing. For example, it can be processed at high speed on time (real time).

[0141] Furthermore, when the AI unit 4 is installed in a server or fog computing system within the factory, data from multiple polishing machines within the factory can be aggregated to update the machine learning model. Furthermore, data from multiple polishing machines within the factory can be aggregated and analyzed, and the analysis results can be reflected in the polishing parameter settings.

[0142] <Third embodiment>

[0143] The third embodiment will now be described. The polishing apparatus 10 of the first embodiment includes the AI unit 4. However, the third embodiment differs in that the AI unit 4 is not located in the polishing apparatus but in the cloud.

[0144] Figure 17 : is a schematic diagram showing the overall structure of the grinding system of the third embodiment. Figure 17 As shown, a polishing system S3 according to the third embodiment includes polishing devices 10-1 to 10-N installed in multiple factories and an AI unit 4 installed in the cloud. The AI unit 4 can communicate with the polishing devices 10-1 to 10-N via a global network NW2 and a local network NW1. The AI unit 4 is, for example, a computer (e.g., a server).

[0145] Therefore, by placing the AI unit 4 in the cloud, physically separate from the polishing equipment, it can be shared across multiple factories, improving the maintainability of the AI unit 4. Furthermore, by utilizing polishing data from multiple factories, the machine learning model can be retrained with a large amount of data, which can further improve the estimation accuracy.

[0146] Furthermore, data from multiple polishing machines across multiple factories (e.g., a large amount of data) can be aggregated to update the machine learning model. Furthermore, data from multiple polishing machines across multiple factories (e.g., a large amount of data) can be aggregated for analysis, and the analysis results can be reflected in polishing parameter settings.

[0147] In addition, the AI unit 4 may be located not in the cloud but in an analysis center that performs centralized analysis.

[0148] Regarding the installation location of the AI unit 4, it can also be (1) inside the polishing device, and / or (2) a gateway near the polishing device, and / or (3) a computer (PC, server, fog computing, etc.) inside the factory (for example, in the factory management room).

[0149] The AI unit 4 may be installed in (1) a polishing device, and / or (2) a gateway near the polishing device, and / or (4) a computer in the cloud (or analysis center).

[0150] Regarding the installation location of the AI unit 4, it can also be (1) inside the polishing device, and / or (2) a gateway near the polishing device, and / or (3) a computer in the factory (for example, in the factory management room), and / or (4) a computer in the cloud (or analysis center).

[0151] In addition, the various components of the AI unit 4 can also be dispersed and configured in (1) the polishing device, and / or (2) a gateway near the polishing device, and / or (3) a computer (PC, server, fog computing, etc.) in the factory (for example, in the factory management room), and / or (4) a computer in the cloud (or analysis center).

[0152] In addition, the input to the machine learning model of various embodiments is a characteristic quantity of a signal related to the frictional force between the polishing member and the substrate during polishing, but this is not limited to this. The input to the machine learning model can also be a characteristic quantity of the temperature of the polishing member (here, the polishing pad 101) or the substrate during polishing. Therefore, if the frictional force between the polishing member and the substrate increases during polishing, the amount of heat generated by the polishing member or substrate also increases. Since the temperature of the polishing member or substrate increases, there is a positive correlation between the temperature of the polishing member or substrate and the frictional force between the polishing member and the substrate during polishing.

[0153] That is, the storage body 41 may also store a machine learning model that completes learning using learning data, wherein the learning data takes as input the temperature characteristic quantity of the grinding part or substrate during grinding, and takes as output data on the thickness of the substrate film after grinding, or the profile statistics of the thickness of the substrate film after grinding, or parameters related to the product qualification rate contained in the substrate after grinding.

[0154] At this time, the generating unit 451 can also rotate the polishing head 1 and the polishing table 100 on which the target substrate is installed, and press the target substrate against the polishing component, and generate characteristic quantities from the signal of the friction force between the polishing component and the substrate during polishing, or the temperature of the polishing component or substrate during polishing.

[0155] Furthermore, at least a portion of the AI unit 4 described in the above embodiment may be implemented as hardware or software. In the case of software implementation, a program implementing at least a portion of the functions of the AI unit 4 may be stored on a recording medium such as a floppy disk or CD-ROM, and read and executed by a computer. The recording medium is not limited to a removable structure such as a magnetic disk or optical disk; it may also be a fixed recording medium such as a hard disk device or memory.

[0156] Alternatively, a program that implements at least a portion of the functions of the AI unit 4 may be distributed via a communication line such as the Internet (including wireless communication). Furthermore, the program may be distributed in an encrypted, modulated, or compressed state via a wired or wireless line such as the Internet, or stored on a recording medium.

[0157] Furthermore, the AI unit 4 may function by one or more information processing devices. When using multiple information processing devices, one of the information processing devices may be used as a computer, and the computer may execute a predetermined program to realize the function of at least one means of the AI unit 4.

[0158] In addition, in the invention of the method, all processes (steps) can also be realized by computer automatic control. In addition, each process can also be implemented by a computer, and control between processes can be performed manually. In addition, further, at least part of all processes can also be implemented manually.

[0159] The present technology is not limited to the aforementioned embodiments in their entirety. During implementation, the components may be modified and embodied without departing from the gist of the present technology. Furthermore, various inventions may be formed by appropriately combining multiple components disclosed in the aforementioned embodiments. For example, some components may be deleted from all components disclosed in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

[0160]

Explanation of symbols

[0161] 1: Grinding head

[0162] 100:Grinding table

[0163] 100a: axis

[0164] 101: polishing pad

[0165] 101a: Grinding surface

[0166] 102: Rotary motor

[0167] 110: Top ring head

[0168] 111: Top ring shaft

[0169] 112: Rotating drum

[0170] 113: Timing pulley

[0171] 114: Rotation motor for top ring

[0172] 115: Timing belt

[0173] 116: Timing pulley

[0174] 117: Top ring shaft

[0175] 124: Up and down movement mechanism

[0176] 126: Bearing

[0177] 128: Bridge

[0178] 129: Support platform

[0179] 130: Pillar

[0180] 132: Ball Screw

[0181] 132a: spiral shaft

[0182] 132b: Nut

[0183] 138:Servo motor

[0184] 20: Front loading unit

[0185] 21:FOUP

[0186] 22: Transport Robot

[0187] 26: Rotary joint

[0188] 3: retaining ring

[0189] 4:AI Department

[0190] 41: Storage

[0191] 42: Memory

[0192] 43: Input unit

[0193] 44: Output unit

[0194] 45: Processor

[0195] 451: Generation Department

[0196] 452: Presumption

[0197] 453: Judgment Department

[0198] 5: Cleaning Department

[0199] 500: Control Department

[0200] 53: Transport Robot

[0201] 6: Film thickness measuring device

[0202] 7:Transmitter

[0203] S1~S3: Grinding system

Claims

1. A polishing device capable of referencing a memory storing a machine learning model learned using learning data, wherein the learning data has as input a characteristic value of a signal related to the frictional force between a polishing member and a substrate during polishing, or a characteristic value of the temperature of the polishing member or substrate during polishing, and outputting data related to the film thickness of the polished substrate or a parameter related to the product yield contained in the polished substrate, The grinding device is characterized by comprising: A grinding table provided with a grinding member and configured to rotate; a polishing head, the polishing head being rotatable opposite to the polishing table and capable of mounting a substrate on a surface thereof opposite to the polishing table; a control unit configured to control the polishing head and the polishing table to rotate while the polishing head and the polishing table are mounted thereon, so as to press the substrate against the polishing member to polish the substrate; a processor that generates a feature value based on a signal regarding the frictional force between the polishing member and the substrate during polishing, or the temperature of the polishing member or the target substrate during polishing, and inputs the generated feature value into the machine learning model that has completed learning, thereby outputting data regarding the film thickness of the substrate after polishing or a parameter regarding the product yield included in the substrate after polishing as an estimated value; and a film thickness measuring device for measuring the film thickness of the substrate, When the output estimated value satisfies a predetermined polishing deterioration condition, the processor controls the film thickness measuring device to measure the film thickness of the target substrate after polishing. When the output estimated value does not satisfy the predetermined polishing deterioration condition, the processor controls the film thickness measuring device not to measure the film thickness of the target substrate after polishing.

2. The grinding device according to claim 1, wherein The processor stops subsequent processing of the substrate when the output estimated value satisfies a predetermined polishing deterioration condition.

3. The grinding device according to claim 1, wherein The processor outputs a maintenance timing using a tendency of the estimated value output for substrates polished at a plurality of different times.

4. The grinding device according to claim 1, wherein The processor performs control so as to issue a warning urging maintenance when the output estimated value satisfies a predetermined polishing deterioration condition.

5. The grinding device according to claim 1, wherein The processor adjusts polishing conditions for subsequent substrates so as to obtain desired data on the film thickness of the polished substrate or desired parameters on the product yield of the polished substrate based on the output estimated value.

6. The grinding device according to claim 1, wherein The processor re-learns the machine learning model using the feature value during operation of the polishing device.

7. An information processing system capable of referencing a memory storing a machine learning model learned using learning data, the learning data having as input a characteristic value of a signal related to the frictional force between a polishing member and a substrate during polishing, or a characteristic value of the temperature of the polishing member or substrate during polishing, and outputting data related to the film thickness of the polished substrate or a parameter related to product yield contained in the polished substrate, The information processing system is characterized by comprising: a generating unit that generates a feature value based on a signal regarding a frictional force between a polishing member and a substrate during polishing, or a temperature of the polishing member or a target substrate during polishing; an estimating unit that inputs the generated feature value into the learned machine learning model to output, as an estimated value, either data on the film thickness of the polished substrate or a parameter related to a product yield included in the polished substrate; and The control component controls the film thickness measuring device to measure the film thickness of the target substrate after grinding when the output estimated value satisfies the predetermined grinding deterioration condition, and controls the film thickness measuring device not to measure the film thickness of the target substrate after grinding when the output estimated value does not satisfy the predetermined grinding deterioration condition.

8. A polishing method comprising polishing a substrate using a polishing apparatus, the polishing apparatus being capable of referencing a memory storing a machine learning model learned using learning data, the learning data having as input a characteristic value of a signal related to the frictional force between a polishing member and a substrate during polishing, or a characteristic value of the temperature of the polishing member or substrate during polishing, and outputting data related to the film thickness of the polished substrate or a parameter related to a product yield contained in the polished substrate. The grinding method is characterized in that The substrate is polished by pressing the substrate against the polishing member while rotating the polishing head and the polishing table on which the substrate is mounted. A characteristic value is generated by measuring a signal related to the friction force between the polishing member and the substrate during polishing, or the temperature of the polishing member or the target substrate during polishing, The generated feature quantity is input into the machine learning model that has completed learning. Outputting as an estimated value either data on the film thickness of the polished substrate or a parameter on the product yield included in the polished substrate, Determine whether the output estimated value satisfies a predetermined polishing deterioration condition. When the predetermined polishing deterioration condition is satisfied, the film thickness of the target substrate is measured by a film thickness measuring device after polishing. When the output estimated value does not satisfy the predetermined polishing deterioration condition, the film thickness of the target substrate is not measured by the film thickness measuring device after polishing.

9. The grinding method according to claim 8, wherein: It is determined whether the output estimated value satisfies a predetermined polishing deterioration condition, and if the predetermined polishing deterioration condition is satisfied, subsequent processing of the substrate is stopped.

10. The grinding method according to claim 8, wherein The maintenance timing is output using the tendency of the estimated values output for substrates polished at a plurality of different times.

11. The grinding method according to claim 8, wherein It is determined whether the output estimated value satisfies a predetermined polishing deterioration condition, and when the predetermined polishing deterioration condition is satisfied, a warning urging maintenance is issued.

12. The grinding method according to claim 8, wherein The polishing conditions of the subsequent substrate are adjusted so as to obtain desired data on the film thickness of the substrate after polishing or desired parameters on the product yield included in the substrate after polishing based on the output estimated value.

13. The grinding method according to claim 8, wherein The machine learning model is relearned using the feature values during operation of the polishing device.

14. A recording medium, characterized in that A program is stored for causing a computer to function as the following element, the computer being able to refer to a storage body storing a machine learning model that has completed learning using learning data, the learning data having as input a characteristic quantity of a signal regarding the frictional force between a polishing member and a substrate during polishing, or a characteristic quantity of the temperature of the polishing member or substrate during polishing, and outputting data regarding the film thickness of the substrate after polishing, or a parameter regarding a product yield included in the substrate after polishing, The elements include: a generating unit that generates a feature value based on a signal regarding a frictional force between a polishing member and a substrate during polishing, or a temperature of the polishing member or a target substrate during polishing; an estimating unit that inputs the generated feature value into the learned machine learning model to output, as an estimated value, either data on the film thickness of the polished substrate or a parameter related to a product yield included in the polished substrate; and The control component controls the film thickness measuring device to measure the film thickness of the target substrate after grinding when the output estimated value satisfies the predetermined grinding deterioration condition, and controls the film thickness measuring device not to measure the film thickness of the target substrate after grinding when the output estimated value does not satisfy the predetermined grinding deterioration condition.

Citation Information

Patent Citations

  • Tool state estimation apparatus and machine tool

    CN108500736A

  • Machine Learning Systems for Monitoring of Semiconductor Processing

    US20190286111A1

  • Automated system check for metrology unit

    US8825444B1