A method and system for data classification of a machine
By performing multi-level processing and classification of machine data, a curved data set with the same or similar data change trends is generated, the problem of massive data processing pressure and cost is solved, more timely and accurate fault detection is achieved, monitoring costs are reduced, and monitoring is adapted to different production environments.
Patent Information
- Application Number
- CN202411885988.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-12-19
AI Technical Summary
When the prior art faces massive data points, it is difficult to effectively reduce the pressure and processing costs of data processing, resulting in low monitoring delay and data analysis efficiency.
A machine data classification method is proposed, which generates numerical curves by fitting data points, performs multi-level data processing and classification, including calculating change trends, amplifying change trends, first and second classification operations, and selects curve data sets with the same or similar data change trends.
Pre-classification reduces the pressure of subsequent data analysis, improves the timeliness and accuracy of fault detection, reduces monitoring costs, and provides a mechanism for autonomous update of classification levels to adapt to different production lines and production stages.
Smart Images

Figure CN119691623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wafer manufacturing, and particularly relates to a method and system for classifying machine data. Background Art
[0002] In the manufacturing process of semiconductor devices, each process may cause some unexpected structures on the wafer. Among them, those that cause the circuits on the chip to malfunction are called wafer defects. To ensure the wafer yield and production capacity, the maintenance of wafer processing equipment becomes crucial.
[0003] In wafer processing, it is necessary to control the quality of the wafer. When defective products appear, it is also necessary to analyze and process the wafer defect data. During the quality control process, the operating state of the machine is usually monitored. For example, when a certain operating parameter of the machine exceeds the set upper threshold or is lower than the lower threshold, it indicates that the machine may have a running fault. However, on the one hand, the traditional threshold monitoring method has a huge data monitoring pressure, often requiring a traversal comparison of millions of data points; on the other hand, this threshold monitoring method also has great limitations in monitoring capabilities. It can often only be detected after the fault occurs, that is, there is a certain delay in detection.
[0004] To improve the accuracy of data monitoring, a method of processing raw data using frequency domain algorithms or statistical probability distributions has been proposed in the prior art.
[0005] For example, patent application CN111223799A discloses a process control method, device, system and storage medium, which acquires the raw data of a sensor; extracts features from the raw data to obtain feature data; performs a correlation analysis on the feature data and the measurement data of a detection device to obtain target data; inputs the target data into an error detection and classification FDC system so as to output a control strategy through the FDC system; and performs process control of the machine according to the control strategy.
[0006] For another example, patent application CN113255840A discloses a fault detection and classification method, device, system and storage medium. This method attempts to apply frequency domain algorithms and / or statistical probability distribution algorithms to extract features from the segmented data to obtain feature data that is relevant to at least one characteristic of the product; and perform fault detection and classification on the process for the product based on the feature data.
[0007] However, when facing millions or even hundreds of millions of data points, this data processing method will face extremely high data processing pressure.
[0008] Therefore, there is an urgent need for a method that can reduce data processing pressure and processing costs. Summary of the Invention
[0009] The object of the present invention is to provide a data classification method and system for a machine tool, which can partially solve or alleviate the above deficiencies in the prior art, and can pre-classify the massive data of the machine tool, so that the wafer engineer can intervene and process it specifically, in order to discover more timely and accurately the possible fault problems that may occur during the operation of the machine tool.
[0010] In order to solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions:
[0011] In the first aspect of the present invention, there is provided a data classification method for a machine tool, including the steps of:
[0012] S100, obtaining a plurality of data points from at least one machine tool, and the plurality of data points are used to reflect at least one working index of the machine tool;
[0013] S101, respectively fitting a plurality of first numerical curves corresponding to the machine tool according to the plurality of data points;
[0014] S102, performing data processing on the plurality of first numerical curves to correspondingly obtain a plurality of second numerical curves, wherein the data processing method includes: a first type of processing method for calculating the change trend of the first numerical curve, and / or a second type of processing method for amplifying the change trend of the first numerical curve;
[0015] S103, performing a first classification operation on the plurality of second numerical curves to select a type of curve from the plurality of second numerical curves, wherein the similarity between the type of curve and a pre-stored first reference curve is less than or equal to a set first similarity threshold;
[0016] S104, obtaining a plurality of fault types, and respectively determining corresponding plurality of data change types according to the fault types, wherein the fault types include: a first typical fault type, and the data change type is used to describe the first characteristic trend of the data change in the type of curve, and at least one data change type corresponds to one first typical fault type;
[0017] S105, performing a second classification operation on the plurality of type of curves according to the plurality of data change types to divide the type of curves into a plurality of curve data sets, and one curve data set includes: a plurality of the type of curves, and the plurality of the type of curves have the same or similar data change trend.
[0018] In some embodiments, at least one test result corresponds to one of the first typical fault types, and the method further includes:
[0019] S106. Calculate a plurality of correlation number sets between at least one of the curve data sets and at least one of the test results; wherein, one of the curve data sets corresponds to one of the test results to form one of the correlation number sets, and the correlation number set includes: a plurality of correlations obtained by calculating the plurality of first-class curves and the test result respectively through a correlation calculation method;
[0020] S107. Calculate the difference between at least two of the correlations in the correlation data set;
[0021] S108. When the difference is greater than a first set difference, output a first prompt signal to prompt the user to perform a reclassification operation on the curve data set.
[0022] In some embodiments, at least one test result corresponds to one of the first typical fault types, including the steps of:
[0023] Calculate the correlation between the first-class curve and the test result; wherein, when the correlation is lower than a set first correlation value, it is considered that the first-class curve is not relevant to the test result, otherwise, it is considered that the first-class curve is relevant to the test result;
[0024] When the number of first-class curves relevant to the test result is less than a set first relevant quantity, mark the corresponding test result as an unknown result, and correspondingly issue a second prompt signal to prompt the user to perform a correction process on the first numerical curve.
[0025] In some embodiments, the steps of performing a correction process on the first numerical curve include:
[0026] (1) Obtain the first numerical curve processed by the first-class processing method, and process the first numerical curve by the second-class processing method to obtain a new second numerical curve;
[0027] (2) Determine whether the similarity between the new second numerical curve and the second reference curve is less than or equal to a set second similarity threshold;
[0028] If not, proceed to (3);
[0029] (3) Calculate the correlation between the second numerical curve and the test result marked as an unknown result;
[0030] (4) If the correlation is greater than a set second correlation value, classify the second numerical curve into a new curve data set.
[0031] In some embodiments, it further includes the steps of:
[0032] Obtain the fault type of the new curve data set, and label the fault type as the second typical fault type.
[0033] In some embodiments, it further includes the steps of:
[0034] Record the data change type corresponding to the second typical fault type, so as to update the fault type and the data change type in S104.
[0035] In some embodiments, it includes the steps of: selecting the first typical fault type from the historical database of the machine platform.
[0036] In some embodiments, the first type of processing method includes one or more of the following: calculus processing; and / or, the second type of processing method includes one or more of the following: exponential processing.
[0037] In some embodiments, it further includes the steps of:
[0038] Select the corresponding data processing method according to the data label of the first type of curve; wherein, the data label includes: machine platform type, or the typical fault type of the machine platform; and at least one preferred data processing method is marked corresponding to the machine platform type or the typical fault type.
[0039] The present invention also provides a data classification system for a machine platform, including:
[0040] A data acquisition module, configured to acquire a plurality of data points from at least one machine platform, and the plurality of data points are used to reflect at least one working index of the machine platform;
[0041] A data fitting module, configured to respectively fit a plurality of first numerical curves corresponding to the machine platform according to the plurality of data points;
[0042] A data processing module, configured to perform data processing on the plurality of first numerical curves to correspondingly obtain a plurality of second numerical curves, wherein the data processing method includes: a first type of processing method, the first type of processing method is used to calculate the change trend of the first numerical curve, and / or, a second type of processing method, the second type of processing method is used to amplify the change trend of the first numerical curve;
[0043] A first classification module, configured to perform a first classification operation on the plurality of second numerical curves to sort out a first type of curve from the plurality of second numerical curves, wherein the similarity between the first type of curve and a pre-stored first reference curve is less than or equal to a set first similarity threshold;
[0044] A type determination module, configured to obtain multiple fault types and respectively determine corresponding multiple data change types according to the fault types, where the fault types include: a first typical fault type, and the data change types are used to describe a first characteristic trend of data change in the one type of curves, and at least one of the data change types corresponds to one first typical fault type;
[0045] A second classification module, configured to perform a second classification operation on the multiple one type of curves according to the multiple data change types, so as to divide the one type of curves into multiple curve data sets, and one curve data set includes: multiple one type of curves, and the multiple one type of curves have the same or similar data change trends.
[0046] Beneficial technical effects:
[0047] For the massive data points output on a wafer production line, the present invention proposes a multi-level verification and classification mechanism for data based on a one type or two type of processing method.
[0048] First, the present invention proposes a collaborative classification method based on curve types and typical fault types to reasonably set the classification degree. Thus, the reasonable setting of the classification degree can, on the one hand, avoid over-classification, that is, avoid prematurely ignoring the possible associations between different curve data sets, increasing the difficulty for wafer engineers to accurately locate faults in the later stage, or prolonging the judgment time for associated faults; on the other hand, it can also avoid too low classification degree (equivalent to ineffective classification), resulting in wafer engineers still needing a large degree of manual intervention to judge the correlation between data and perform manual classification, which will not only increase the work pressure of wafer engineers, but also put higher requirements on the engineering experience of wafer engineers. For example, for some relatively junior wafer engineers, it is very likely that they will draw wrong conclusions in data processing.
[0049] Furthermore, the present invention also provides a classification mechanism with autonomous update of the classification degree. For example, the present invention can evaluate the correlation of test results within a curve data set to determine whether it is necessary to further disassemble the preliminary classification result, such as dividing a curve data set into multiple sub-curve data sets. Another example is that the present invention can also evaluate the association between all one type of curves (i.e., all curve data sets) and test results to determine whether there is a situation of missed classification. In other words, the present invention provides a collaborative classification scheme based on three dimensions of fault types, curve change types, and test results of faults to comprehensively evaluate the classification degree and classification accuracy.
[0050] Furthermore, the present invention can also analyze the accuracy and applicability of the real-time classification degree, that is, obtain the latest typical fault types, so as to timely adjust the classification degree (i.e., the number of curve data sets classified in S104) in the next classification process, thereby ensuring that the classification mechanism has a high adaptability to the actual operation of the wafer production line. In other words, this classification mechanism with autonomous update of the classification degree also makes the method have stronger versatility when facing different production lines, or different production stages of the same production line and other scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 is a schematic flowchart of the classification method in the first exemplary embodiment of the present invention;
[0053] Figure 2 is a schematic flowchart of the classification method in the second exemplary embodiment of the present invention;
[0054] Figure 3 is a schematic flowchart of the reprocessing step method in the third exemplary embodiment of the present invention;
[0055] Figure 4 is a schematic diagram of the module structure of the classification system in an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0057] In this article, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of the description of the present invention, and they have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably.
[0058] In this text, the orientation or positional relationship indicated by terms such as "upper", "lower", "inner", "outer", "front", "rear", "one end", "the other end", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0059] In this text, unless otherwise clearly specified and defined, terms such as "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0060] As used in this text, "and / or" includes any and all combinations of one or more of the listed related items.
[0061] As used in this text, "a plurality of" means two or more, that is, it includes two, three, four, five, etc.
[0062] As used in this specification, the term "about" typically represents + / - 5% of the value, more typically + / - 4% of the value, more typically + / - 3% of the value, more typically + / - 2% of the value, even more typically + / - 1% of the value, and even more typically + / - 0.5% of the value.
[0063] In this specification, certain embodiments may be disclosed in a format within a certain range. It should be understood that this description of "within a certain range" is only for convenience and brevity and should not be construed as a rigid limitation on the disclosed range. Therefore, the description of the range should be considered to have specifically disclosed all possible sub-ranges and the individual numerical values within this range. For example, the description of the range 1 - 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and the individual numbers within this range, such as 1, 2, 3, 4, 5, and 6. The above rules apply regardless of the breadth of the range.
[0064] In this text, "machine tool (or processing machine tool)" is also referred to as equipment (or processing equipment), which can refer to any one or more processing modules or processing devices on the production and processing line of wafers.
[0065] In this article, the "working indicator" refers to the status parameters output (or collected) by the machine during the process of processing wafers. For example, the status parameters can be the working current and working voltage of the machine itself. Correspondingly, the "first numerical curve" can be the numerical curve of the current of the machine changing with time, or the "first numerical curve" can be the numerical curve of the voltage of the machine changing with time. Another example is that the status parameter can be the pressure applied to the wafer by the machine during wafer processing. Another example is that the status parameter can also refer to the working temperature of the machine, etc. These working indicators are usually the original data of the machine operation.
[0066] In this article, "fault" can also be referred to as machine fault or equipment fault, which refers to equipment or process problems that may affect processes such as wafer processing and detection, thereby affecting the yield or production capacity of wafers.
[0067] For example, the fault can be a mechanical fault. For example, due to the loosening of the robotic arm of the machine, scratches may be generated on the wafer. For example, due to the poor voltage stabilization effect of the machine, the voltage of the wafer may fluctuate too much, which may lead to corresponding defects in the circuit structure formed on the wafer. Another example is that due to the failure of the cleaning machine, residual cleaning agent will appear, which will form white impurity-like faults on the wafer.
[0068] Another example is that the fault can also be a process fault. For example, due to problems in the processing recipe setting of the machine, the wafer production capacity is insufficient.
[0069] Taking the current as an example, the processing process of the working indicator is explained as follows: In a wafer production line, a wafer may go through about 800 process steps, and each process step is executed by a corresponding machine. Correspondingly, after the wafer is processed through the production line, a machine will output at least tens of thousands of current data points, and the entire production line will output at least millions of data points. When these data points are abnormal, it may indicate that the machine has a fault, which will also lead to corresponding defects on the wafer, resulting in the yield being affected. Therefore, the analysis of data points helps wafer engineers understand the operation status of the production line. However, millions or even hundreds of millions of data points will bring great data processing pressure to wafer engineers.
[0070] In response to this, the present invention proposes a data pre - processing method, that is, to perform reasonable classification processing on a large amount of data in advance. Specifically, the present invention classifies a large number of data points generated during the wafer processing process according to the actual operation conditions of the wafer fab through mechanisms such as classification and matching, so that subsequent engineers can, under the guidance of the preliminary classification, rely on professional experience to perform more refined data analysis and processing on each data set (such as specifically analyzing the causes of faults and repair solutions with the help of a wafer data analysis platform or manual experience), so as to relieve the data processing pressure of engineers and reduce the monitoring cost of the wafer production line.
[0071] In other words, the present invention proposes a method for simplifying a large number of data points.
[0072] Embodiment 1
[0073] First, as shown in Figure 1 The present invention provides a data classification method for a machine tool, including the steps:
[0074] S100, obtaining a plurality of data points from at least one machine tool, and the plurality of data points are used to reflect at least one working index of the machine tool.
[0075] In some embodiments, the data points can be collected through an FDC (Failure Date Collection, defect data collection system).
[0076] S101, respectively fitting a plurality of first numerical curves corresponding to the machine tool according to the plurality of data points.
[0077] S102, performing data processing on the plurality of first numerical curves to correspondingly obtain a plurality of second numerical curves, wherein the data processing method includes: a first - type processing method for calculating the change trend of the first numerical curve, and / or a second - type processing method for amplifying the change trend of the first numerical curve.
[0078] S103, performing a first - stage classification operation on the plurality of second numerical curves to sort out a type of curves from the plurality of second numerical curves.
[0079] In some embodiments, the second numerical curves are classified by comparing the similarity between the second numerical curves and a reference value, so as to divide them into a type of curves and a second - type curves.
[0080] For example, in some embodiments, calculate the similarity between the second numerical curve and the reference value, and classify the second numerical curve with a similarity less than or equal to a set similarity threshold as a type of curve. On the contrary, when the similarity is greater than the similarity threshold, it is classified as a second - type curve.
[0081] For example, in some embodiments, the reference value can be set as a first reference curve, which is used to reflect the fluctuation trend of the conventional working index and can be the reference data preset by the wafer engineer in combination with the actual operation of the wafer production line. Correspondingly, the similarity between the first type of curve and the pre-stored first reference curve is less than or equal to a set first similarity threshold; the similarity between the second type of curve and the first reference curve is greater than the first similarity threshold.
[0082] In this embodiment, the second numerical curve can preferably be classified for the first time according to the first reference curve, that is, it is divided into the first type of curve with abnormal fluctuations and the second type of curve with normal fluctuations. For example, the second type of curve refers to the second numerical curve with no fluctuations or very small fluctuation differences compared with the first reference curve.
[0083] Alternatively, in some embodiments, the reference value can be set as one or more detected numerical points at one or more moments, and the curves are classified by comparing the similarity between the multiple detected numerical points and the second numerical curve (such as the difference between the numerical points and the corresponding values on the second numerical curve). Among them, the similarity can be the statistical result of the similarities calculated by multiple detected data points. For example, the statistical result can be the mean, maximum value, median, mode, total number, etc. of multiple data (such as multiple similarities).
[0084] S104. Obtain multiple fault types, and respectively determine corresponding multiple data change types according to the fault types, where the fault types include: the first typical fault type, and the data change type is used to describe the first characteristic trend of the data change in the first type of curve, and at least one data change type corresponds to one first typical fault type;
[0085] S105. Perform a second classification operation on the multiple first type of curves according to the multiple data change types to divide the first type of curves into multiple curve data sets, and one curve data set includes: multiple first type of curves, and the multiple first type of curves have the same or similar data change trends (or data change types). For example, in some embodiments, the second numerical curves that conform to the same data change type are divided into one curve data set.
[0086] Correspondingly, based on the two classifications, on the one hand, most of the invalid data can be filtered out, and at the same time, the first type of curves in the same curve data set have similar characteristics (or it is highly probable that they originate from the same or a type of fault), so it is also convenient to correspondingly assign them to different wafer engineers for subsequent analysis.
[0087] In some embodiments, the first typical failure type may be a failure that occurs frequently and has a certain universality in the wafer production line. For these known failures, the fluctuations of the working indicators caused by them also have certain characteristics. These characteristics are recorded as different data change types to facilitate the secondary classification of a type of curve with abnormal fluctuations.
[0088] Specifically, the present invention classifies the second numerical curve into different curve data sets according to possible failure types through the first classification operation and the second classification operation in sequence. Thus, by jointly applying the curve type and the failure type, it is possible to facilitate a more reasonable classification of a large number of data points.
[0089] In other words, through the collaborative classification of the curve type and the typical failure type, the classification degree can be reasonably set. Thus, the reasonable setting of the classification degree can, on the one hand, avoid over-classification, that is, avoid prematurely ignoring the possible associations between different curve data sets, which increases the difficulty for wafer engineers to accurately locate the failure in the later stage; on the other hand, it can also avoid too low a classification degree (equivalent to ineffective classification), resulting in wafer engineers still needing a large degree of manual intervention to judge the correlation between data and perform manual classification, which also puts higher requirements on the engineering experience of wafer engineers.
[0090] In particular, in the present invention, the typical failure type and the second numerical curve are used in cooperation to complete the classification and analysis of data points, which can ensure that the classification degree and classification results have better generality and reliability in multiple application scenarios (such as different wafer production lines).
[0091] Preferably, in some embodiments, the first typical failure type may include one or more of the following types: cleaning agent residue failure, voltage stabilization failure, probe failure, etching by-product failure, etc. The applicant notes that these types of failures have a certain universality in the wafer processing factory, so preferably they are set as the initial classification conditions for curve classification.
[0092] In some embodiments, the probe failure may generally be caused by unstable voltage or aging of the power supply line. Specifically, it may have a specific fluctuation trend (or called characteristic trend) on the second numerical curve of voltage or current.
[0093] In some embodiments, when there are certain errors in the cleaning process (such as cleaning parameters such as temperature, pressure, flow rate of the cleaning valve, time, etc.), it may lead to poor cleaning effect, and then cause cleaning agent residue on the wafer, resulting in a cleaning agent residue failure. Correspondingly, specific fluctuation trends may also be formed among the numerical curves of its temperature, pressure, flow rate of the cleaning valve, pH value of the residual water, etc.
[0094] For example, when the nozzle of the cleaning valve is blocked or filled, it will form a specific trend on the flow curve (such as the flow rate gradually decreasing to a certain value and remaining unchanged).
[0095] Preferably, these selected typical faults usually cause fluctuations with a certain degree of recognizability on the first numerical curve or the second numerical curve. For example, the fluctuation may be manifested as the value being far above or below the reference value for a long time. Another example is that the fluctuation may be manifested as the value fluctuating periodically within a certain period of time. Another example is that the change trend of the value may change periodically within a certain period of time, such as the slope of the numerical curve formed by it fluctuating up and down within a certain period of time. In this embodiment, the graphic change characteristics reflected by the data fluctuation on the numerical curve are distinguished and classified.
[0096] Preferably, in some embodiments, the first typical fault type includes multiple primary faults. For example, the primary faults can be one or more of the faults such as cleaning agent residue fault, voltage stabilization fault, and probe fault, and this type of fault can be further divided into multiple secondary faults. For example, taking the probe fault as an example, the probe fault can specifically be divided into one or more of the following faults: unstable voltage, aging of the power supply line (equivalent to multiple secondary faults). Another example is that taking the cleaning agent residue fault as an example, it can specifically be divided into one or more of the following faults: temperature control fault, pressure control fault, cleaning valve blockage fault (equivalent to multiple secondary faults), and so on.
[0097] In some embodiments, based on the results of these classifications, that is, multiple curve data sets, the wafer engineer can analyze in combination with engineering experience to predict or evaluate the real fault causes, and then perform timely maintenance on the machine.
[0098] Alternatively, in some other embodiments, the wafer engineer can also choose to classify these multiple curve data sets into new bin codes respectively and input them into the YMS (Yield Management System), so that the YMS can adaptively update the data processing mechanism based on the new classification results.
[0099] The YMS is a professional data analysis tool integrating data management, data analysis, visualization, and standardization functions. In the processes of semiconductor design, wafer manufacturing, packaging and testing, etc. (especially in the mass production stage), it can help customer engineers greatly improve the data analysis efficiency, quickly analyze one or more types of data, find the key points to improve the yield, and further promote the stability and controllability of the entire production process, improve the product quality, and reduce the enterprise cost. In this embodiment, through the reasonable preliminary classification of a large number of data points, it is also beneficial for the YMS to further analyze and judge different data sets using its adaptive learning ability.
[0100] In some embodiments, steps S100 - S103 and S104 can be executed synchronously or sequentially.
[0101] In some embodiments, at least one of the test results corresponds to one of the first typical fault types. Refer to Figure 2 as shown, the method further includes:
[0102] S106, calculating a plurality of correlation datasets between at least one of the curve datasets and at least one of the test results; wherein, one of the curve datasets and one of the test results form one of the correlation datasets, and the correlation datasets include: a plurality of correlations obtained by calculating the plurality of first - type curves and the test results respectively through a correlation calculation method.
[0103] For example, in some embodiments, when initially evaluating that a curve dataset is related to Fault I, the correlation between the curve dataset and the test result corresponding to Fault I is calculated. In a specific embodiment, a curve dataset includes: n first - type curves, and n correlations between the n first - type curves and the test result are calculated correspondingly.
[0104] For example, in some embodiments, the correlation refers to the correlation coefficient r between two curves.
[0105] For example, in some embodiments, the correlation calculation method can adopt a statistical method or a machine learning algorithm, and the present invention does not limit this.
[0106] S107, calculating the difference between at least two of the correlations in the correlation dataset.
[0107] S108, when the difference is greater than a first set difference, a first prompt signal is output to prompt the user to perform a re - classification operation on the curve dataset.
[0108] For example, in some embodiments, when among the values of a plurality of correlations formed by a curve dataset, the difference between two or more correlations deviates greatly, it is considered that the primary fault corresponding to the curve dataset may be generated by two or more secondary faults. Therefore, the classification degree of the curve dataset can be refined again. For example, the current curve dataset is further divided into two or more sub - curve datasets, and among the plurality of correlations corresponding to one of the sub - curve datasets, the difference between any two correlations is less than or equal to a second set difference. In some embodiments, the first set difference and the second set difference may be equal.
[0109] For example, in this embodiment, the test results can include one or more of the following: yield parameters, defect parameters, measurement parameters, device electrical parameters, etc.
[0110] Taking the defect parameter as an example, this embodiment is exemplarily explained as follows: When it is initially analyzed that the fault I reflected in one of the curve data sets may be the main cause of the defect A, then calculate the n correlations between the n first-class curves in the curve data set and the defect parameter of the defect A. If the results of the n correlations are close to being consistent, it can be preliminarily considered that the classification result of the curve data set is relatively reliable, and the classification result can be directly output to the user (such as a wafer engineer) for subsequent analysis. If the results of the n correlations show significant differences, preferably, the curve data set will be further subdivided so that the user can analyze that the fault I may be a first-level fault formed by multiple second-level faults, and the user may be advised to analyze these second-level faults one by one.
[0111] In some embodiments, at least one test result corresponds to one of the first typical fault types. Correspondingly, the method includes the steps of:
[0112] Calculate the correlation between the first-class curve and the test result; wherein, when the correlation is lower than the set first correlation value, it is considered that the first-class curve is not relevant to the test result, otherwise, it is considered that the first-class curve is relevant to the test result;
[0113] When the number of first-class curves related to the test result is less than the set first correlation quantity, mark the corresponding test result as an unknown result, and correspondingly send out a second prompt signal to prompt the user to correct the first numerical curve.
[0114] In this embodiment, when a fault is shown in the test result at the test end, but no meaningful fault data (i.e., the first-class curve) can be matched from the classification result, preferably, an attempt can be made to reprocess the second-class curves screened out in S103 to avoid omission of abnormal data due to improper previous data processing.
[0115] In some embodiments, as shown in Figure 3 The steps of correcting the first numerical curve include:
[0116] (1) Obtain the first numerical curve processed by the first-class processing method, and process the corresponding first numerical curve with the second-class processing method to obtain a new second numerical curve;
[0117] (2) Determine whether the similarity between the new second numerical curve and the pre-stored second reference curve is less than or equal to the set second similarity threshold;
[0118] If not, proceed to (3);
[0119] (3) Calculate the correlation between the second numerical curve and the test result marked as the unknown result;
[0120] (4) If the correlation is greater than the set second correlation value, classify the second numerical curve into a new curve data set.
[0121] That is to say, in this embodiment, when there is a test result (or a fault result) that does not match a relevant type of curve, the first numerical curve filtered out can also be re - processed for trend amplification to avoid ignoring some fluctuations due to their too small degree.
[0122] Similar to the first reference curve, the second reference curve can also be preset by the wafer engineer in combination with experience.
[0123] Specifically, in this embodiment, preferably, the first numerical curve processed by a type of processing method is re - processed to efficiently screen out possible missing data results.
[0124] In some embodiments, it further includes the steps of: obtaining the fault type of the new curve data set and marking the fault type as the second typical fault type.
[0125] In some embodiments, it further includes the steps of: recording the data change type corresponding to the second typical fault type to update the fault type and the data change type in S104. For example, in a new data classification process, the second numerical curve will be classified as a whole with the first typical fault type and the second typical fault type.
[0126] It can be understood that this embodiment also provides a classification mechanism with autonomous update of the classification degree, that is, by analyzing the accuracy and applicability of the real - time classification degree, obtaining the latest typical fault type, and being able to timely adjust the classification degree (i.e., the number of curve data sets classified in S104) in the next classification process to ensure that the classification mechanism has a high adaptability to the actual operation of the wafer production line. In other words, this classification mechanism with autonomous update of the classification degree also makes the method have flexible versatility in the face of different production lines or different production stages of the same production line and other scenarios.
[0127] For example, in some embodiments, before the method is actually applied, historical data for a period of time in the current wafer production line (including data points collected at the operation end and test results collected at the test end) can be selected to apply this historical data to execute S100 - S105 (or the method or steps in any embodiment of the present invention can be executed) to determine the typical fault type suitable for the current wafer production line (i.e., determine the classification degree of the preliminary classification), and then balance the problems of over - classification and invalid classification.
[0128] In some embodiments, the method further includes the step of selecting a first typical fault type from the historical database of the machine tool.
[0129] In some embodiments, the first type of processing method includes one or more of the following: calculus processing.
[0130] For example, in some embodiments, calculus processing refers to taking the derivative of a first numerical curve. In other words, the second numerical curve is the result of taking the derivative of the first numerical curve.
[0131] In some embodiments, the second type of processing method includes one or more of the following: exponential processing. Exponential processing refers to a method of representing data by converting the original data into an exponential form through a mathematical formula.
[0132] For example, in some embodiments, the first numerical curve obtained is: Y1 = A1x + B1; the second numerical curve obtained after exponential processing is: Y2 = A2e kx + B2e l(x-1) + …… + N2e 0 .
[0133] It should be noted that other processing methods can also be used for the second type of processing method in this embodiment. The main point is to amplify the degree of data change. Therefore, as long as the selected method can amplify the degree of data change to facilitate capture, it is acceptable.
[0134] In some embodiments, a corresponding data processing method is selected according to the data label of the first type of curve; wherein, the data label includes: the type of machine tool, or the typical fault type of the machine tool; and at least one preferred data processing method is correspondingly marked for the type of machine tool or the typical fault type.
[0135] For example, for some specific fault types, the fluctuations caused by them on the first numerical curve may be very small. Preferably, the second type of processing method is used to amplify the fluctuation trend to a certain extent to capture abnormal conditions more sensitively. Correspondingly, for these specific fault types, the recommended data processing methods will be correspondingly marked to further improve the reliability and accuracy of the classification degree.
[0136] Embodiment 2
[0137] Correspondingly, the present invention also provides a data classification system for a machine tool, as shown in Figure 4 The system includes:
[0138] A data acquisition module 10, configured to acquire a plurality of data points from at least one machine tool, and the plurality of data points are used to reflect at least one working index of the machine tool;
[0139] A data fitting module 11, configured to respectively fit multiple first numerical curves corresponding to the machine according to multiple said data points;
[0140] A data processing module 12, configured to perform data processing on multiple said first numerical curves to correspondingly obtain multiple second numerical curves, wherein the data processing method includes: a first type of processing method, the first type of processing method is used to calculate the change trend of the first numerical curve, and / or, a second type of processing method, the second type of processing method is used to amplify the change trend of the first numerical curve;
[0141] A first classification module 13, configured to perform a first classification operation on multiple said second numerical curves to select a type of curves from multiple said second numerical curves, wherein the similarity between the type of curves and a pre-stored first reference curve is less than or equal to a set first similarity threshold;
[0142] A type determination module 14, configured to obtain multiple fault types and respectively determine corresponding multiple data change types according to the fault types, wherein the fault types include: a first typical fault type, the data change type is used to describe a first characteristic trend of data change in the type of curves, and at least one said data change type corresponds to one said first typical fault type;
[0143] A second classification module 15, configured to perform a second classification operation on multiple said type of curves according to the multiple data change types to divide the type of curves into multiple curve data sets, and one said curve data set includes: multiple said type of curves, and multiple said type of curves have the same or similar said data change types.
[0144] In some embodiments, at least one test result corresponds to one said first typical fault type. Correspondingly, the system further includes:
[0145] A first calculation module 16, configured to calculate multiple correlation data sets between at least one said curve data set and at least one said test result; wherein one said curve data set forms one said correlation data set corresponding to one said test result, and the correlation data set includes: multiple correlations obtained by respectively calculating multiple said type of curves and the test result through a correlation calculation method;
[0146] A second calculation module 17, configured to calculate the difference between at least two said correlations in the correlation data set;
[0147] A prompt module 18, configured to output a first prompt signal when the difference is greater than a first set difference to prompt the user to perform a re-classification operation on the curve data set.
[0148] In some embodiments, at least one test result corresponds to one of the first typical fault types. Correspondingly, the system further includes:
[0149] A correlation calculation module, configured to calculate the correlation between the one type of curve and the test result; wherein, when the correlation is lower than a set first correlation value, it is considered that the one type of curve is not relevant to the test result, otherwise, it is considered that the one type of curve is relevant to the test result;
[0150] A correction module, configured to, when the number of one type of curves relevant to the test result is less than a set first relevant quantity, mark the corresponding test result as an unknown result, and correspondingly send out a second prompt signal to prompt the user to perform a correction process on the first numerical curve.
[0151] In some embodiments, the correction module further includes:
[0152] A reprocessing unit, configured to obtain the first numerical curve processed by the one type of processing method, and process the first numerical curve by using the second type of processing method to obtain a new second numerical curve;
[0153] A judgment unit, configured to judge whether the similarity between the new second numerical curve and a second reference curve is less than or equal to a set second similarity threshold;
[0154] If not, enter the correlation calculation unit;
[0155] A correlation calculation unit, configured to calculate the correlation between the second numerical curve and the test result marked as an unknown result;
[0156] An update unit, configured to, if the correlation is greater than a set second correlation value, classify the second numerical curve into a new curve data set.
[0157] In some embodiments, it further includes: a first fault update module, configured to obtain the fault type of the new curve data set, and mark the fault type as a second typical fault type.
[0158] In some embodiments, it further includes: a second fault update module, configured to record the data change type corresponding to the second typical fault type, so as to update the fault type and the data change type in the type determination module 14.
[0159] In some embodiments, it further includes: a selection module, configured to select a first typical fault type from the historical database of the machine.
[0160] It can be understood that the system in the embodiments of the present invention can be used to implement the methods or steps in any of the above embodiments, which will not be elaborated here.
[0161] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods in the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a computer terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0163] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A data classification method for a machine, characterized in that: Includes steps: S100, acquiring a plurality of data points from at least one machine, wherein the plurality of data points are used to reflect at least one working indicator of the machine; S101, fitting a plurality of first numerical curves corresponding to the machine according to the plurality of data points; S102, performing data processing on the plurality of first numerical curves to obtain a plurality of second numerical curves correspondingly, wherein the data processing method comprises: a first processing method, the first processing method is used to calculate the change trend of the first numerical curve, and / or a second processing method, the second processing method is used to amplify the change trend of the first numerical curve; S103, performing a first classification operation on the plurality of second value curves to select a type of curves from the plurality of second value curves, wherein the similarity between the type of curves and the pre-stored first reference curve is less than or equal to a set first similarity threshold; S104, acquiring multiple fault types, and determining corresponding multiple data change types according to the fault types, respectively, wherein the data change type is used to describe a first characteristic trend of data changes in the type of curve, and the fault type includes: a first typical fault type, and one of the first typical fault types corresponds to at least one of the data change types; S105, performing a second classification operation on the plurality of the first-class curves according to the plurality of data change types, so as to classify the first-class curves into a plurality of curve data sets, wherein one curve data set includes: a plurality of the first-class curves, and the plurality of the first-class curves have the same or similar data change types.
2. The data classification method according to claim 1, characterized in that: One of the first typical fault types corresponds to at least one test result, and the method further includes: S106, calculating a plurality of correlation number sets between at least one of the curve data sets and at least one of the test results; wherein one of the curve data sets forms one of the correlation number sets corresponding to one of the test results, and the correlation number set includes: a plurality of correlations between a plurality of the one type of curves and the test results respectively calculated by a correlation calculation method; S107, calculating the difference between at least two of the correlations in the correlation data set; S108: When the difference is greater than a first set difference, a first prompt signal is output to prompt the user to perform a reclassification operation on the curve data set.
3. The data classification method according to claim 2, characterized in that: One of the first typical fault types corresponds to at least one test result, comprising the steps of: Calculating the correlation between the one type of curve and the test result; wherein, when the correlation is lower than a set first correlation value, it is considered that the one type of curve is not correlated with the test result, otherwise, it is considered that the one type of curve is correlated with the test result; When the number of the type of curves related to the test result is less than a set first correlation amount, the corresponding test result is marked as an unknown result, and a second prompt signal is issued accordingly to prompt the user to correct the corresponding first numerical curve.
4. The data classification method according to claim 3, characterized in that: The step of correcting the first numerical curve includes: (1) obtaining the first numerical curve processed by the first type of processing method, and processing the corresponding first numerical curve by the second type of processing method to obtain a new second numerical curve; (2) determining whether the similarity between the new second numerical curve and the pre-stored second reference curve is less than or equal to a set second similarity threshold; If not, proceed to (3); (3) calculating the correlation between the second numerical curve and the test result marked as an unknown result; (4) If the correlation is greater than a set second correlation value, the second numerical curve is classified as a new curve data set.
5. The data classification method according to claim 4, characterized in that: Also includes the steps: The fault type of the new curve data set is obtained, and the fault type is marked as a second typical fault type.
6. The data classification method according to claim 5, characterized in that: Also includes the steps: The data change type corresponding to the second typical fault type is recorded to update the fault type and the data change type in S104.
7. The data classification method according to claim 1, characterized in that: The method comprises the steps of: selecting a first typical fault type from a historical database of the machine.
8. The data classification method according to claim 1, characterized in that: The first type of processing method includes one or more of the following: calculus processing; and / or, the second type of processing method includes one or more of the following: exponential processing.
9. The data classification method according to claim 1, characterized in that: Also includes the steps: A corresponding data processing method is selected according to a data label of a type of curve; wherein the data label includes: a machine type, or a typical failure type of a machine; and the machine type or the typical failure type is correspondingly marked with at least one preferred data processing method.
10. A data classification system for a machine, characterized in that: include: A data acquisition module, used to acquire a plurality of data points from at least one machine, wherein the plurality of data points are used to reflect at least one working indicator of the machine; A data fitting module, used for fitting a plurality of first numerical curves corresponding to the machine according to the plurality of data points; a data processing module, configured to perform data processing on the plurality of first numerical curves to obtain a plurality of second numerical curves correspondingly, wherein the data processing method comprises: a first type of processing method, the first type of processing method is used to calculate the change trend of the first numerical curve, and / or a second type of processing method, the second type of processing method is used to amplify the change trend of the first numerical curve; a first classification module, configured to perform a first classification operation on the plurality of the second value curves, so as to select a type of curves from the plurality of the second value curves, wherein the similarity between the type of curves and the pre-stored first reference curve is less than or equal to a set first similarity threshold; A type determination module, used to obtain multiple fault types, and determine corresponding multiple data change types according to the fault types, wherein the data change type is used to describe a first characteristic trend of data changes in the first type of curve, and the fault type includes: a first typical fault type, and one of the first typical fault types corresponds to at least one of the data change types; The second classification module is used to perform a second classification operation on the multiple first-class curves according to the multiple data change types to divide the first-class curves into multiple curve data sets, and one curve data set includes: multiple first-class curves, and the multiple first-class curves have the same or similar data change types.
Citation Information
Patent Citations
Process control method, device and system and storage medium
CN111223799A
Fault detection and classification method, device and system and storage medium
CN113255840A
Automatic sequencing system and method for wafer machine fault processing
CN117575253A
Dynamic analysis of event data
WO2014144893A1