Root cause analysis method, system and program product
By using big data tracing and analysis methods, combined with fault prediction algorithms and the Apache Spark platform, we have achieved rapid and accurate root cause localization in the semiconductor manufacturing process. This solves the problems of data dispersion and long analysis cycles in existing technologies, and improves analysis efficiency and accuracy.
Patent Information
- Application Number
- CN202610062771.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-03-20
AI Technical Summary
In the semiconductor manufacturing process, existing root cause analysis methods rely on human experience, have scattered data sources, are difficult to integrate, have long analysis cycles, slow response speeds, and cannot quickly locate anomalies.
By employing big data source tracing analysis methods, and acquiring semiconductor production line data, we utilize fault prediction algorithms and the Apache Spark platform to perform non-parametric preprocessing and AI collaborative analysis to quickly identify high-probability root cause regions.
It improves the efficiency and accuracy of root cause analysis, enabling automated root cause identification of large-scale manufacturing data within minutes to hours, reducing the workload of engineers and providing fast and reliable analysis results.
Smart Images

Figure CN121707429A_ABST
Abstract
Description
[0001] Divisional application This application is a divisional application of the Chinese Invention Patent Application with the application number 2025115813205 and the application date of October 31, 2025, the invention name of "a semiconductor big data traceability analysis method and system". TECHNICAL FIELD
[0002] The present application relates to the technical field of semiconductor manufacturing, in particular to a root cause analysis method, system and program product. BACKGROUND
[0003] In the field of semiconductor manufacturing, especially in 12-inch wafer manufacturing plants (hereinafter referred to as Fab plants), root cause analysis (RCA) method is an important means of quality management and process control. When there are problems such as yield decline, product defects or customer complaints in the production process, systematic analysis is needed to identify the root cause of the problem and take measures to prevent it from happening again.
[0004] Currently, the RCA method commonly used in Fab plants mainly relies on the experience judgment of engineers and manual data collection. The specific process usually includes: extracting relevant data from multiple independent data systems such as FDC (fault detection and classification), SPC (statistical process control), WAT (electrical property test), Inline (online measurement), etc., and then comparing and analyzing.
[0005] For example, patent application CN202310893076.0 discloses a FDC traceability analysis method based on distributed parallel computing and storage medium, which includes obtaining various information of failed devices; obtaining FDC data and configuration files in the semiconductor wafer production and manufacturing process; preprocessing the FDC data set to obtain data for algorithm calculation structure; performing traceability analysis to find out the impact factors; based on the traceability analysis results, summarize and take out the top N impact factors of each group; compare the test results with the known failure modes to determine the possible failure causes; according to the feedback of multiple possible failure causes, further test and diagnosis operation is carried out to verify and determine the root cause of failure; according to the root cause of failure, take corresponding measures.
[0006] However, semiconductor production data suffers from the following problems: 1. Scattered data sources and difficulty in integration: Existing systems have varying data formats, structures, and storage methods, lacking a unified data platform. Efficient interconnection and fusion between different data sources are difficult, leading to low data analysis efficiency and the potential for missing crucial information. 2. Long analysis cycles and slow response times: Traditional root cause analysis heavily relies on manual intervention. Engineers must manually export, clean, process, and analyze data, a time-consuming process that typically takes weeks or even a month to complete a single RCA analysis, failing to meet the need for rapid anomaly localization. Therefore, a more efficient semiconductor root cause tracing method is urgently needed. Summary of the Invention
[0007] The purpose of this invention is to provide a big data source tracing and analysis method that partially solves or alleviates the aforementioned shortcomings of existing technologies, thereby improving the efficiency and accuracy of data analysis. To address the aforementioned technical problems, this invention specifically adopts the following technical solution: A first aspect of the present invention is to provide a semiconductor big data traceability and analysis method, comprising the following steps: S101, acquire semiconductor production line data, the semiconductor production line data including at least: (1) information on multiple processing equipment experienced by multiple semiconductor products, (2) process parameters used when the semiconductor products experience at least one of the processing equipment, and / or category information of the semiconductor products, (3) measured values, and the processing equipment information including at least one of the following: site information, machine information, chamber information, the process parameters including at least one of the following: formula, sensor type, the category information including at least one of the following: raw material source, semiconductor model; S102, call a data analysis scheme for the semiconductor production line data, the data analysis scheme including: fault prediction algorithm; S103, preprocess the semiconductor production line data according to the data analysis scheme to generate preprocessing results, the preprocessing The results include: the predicted root causes of anomalies, and / or the anomaly weights of the root causes; wherein, S103 includes the following steps: S1031, obtaining extraction rules, the extraction rules including: equipment type and configuration type, wherein the equipment type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; S1032, identifying multiple production sequences from the semiconductor production line data according to the extraction rules, wherein the production sequences are correspondingly recorded with the measured values; S1033, using a fault prediction algorithm to predict the root causes of anomalies based on the differences between at least one production sequence and multiple other production sequences; S104, inputting the preprocessing results and the corresponding semiconductor production line data into a preset big data platform, wherein the big data platform outputs the corresponding analysis results.
[0008] In some embodiments, the fault prediction algorithm includes: a one-way ANOVA algorithm, a history analysis algorithm, and / or a similarity analysis algorithm. In some embodiments, the big data platform is the Apache Spark big data platform. In some embodiments, the production sequence records at least the processing sequence of at least two processing devices that the semiconductor product has undergone, as well as the corresponding measurement values; correspondingly, S1033 further includes: S10331, obtaining at least one production sequence group, and the production sequences in the same group have the same equipment type and configuration type; S10332, using the fault prediction algorithm to calculate the difference between a production sequence and other multiple production sequences in the same group, in order to predict the root cause of the anomaly.
[0009] In some embodiments, when the fault prediction algorithm is the history analysis algorithm, step S10332 includes the steps of: identifying the data type of the measured value, wherein the data type includes: discrete and / or continuous; selecting a corresponding mutual information algorithm according to the data type; and calculating the mutual information of the production sequence according to the mutual information algorithm. The mutual information is used to represent the information contained in the random variable Y about the features. The amount of information; calculate the root cause weight of at least one processing device in the production sequence based on the mutual information index.
[0010] In some embodiments, before S10332, the method further includes the steps of: obtaining the number of sequences in a set of product sequences; determining whether the number of sequences is greater than a preset number; if so, proceeding to step S10332.
[0011] In some embodiments, prior to S1031, the method further includes the steps of: selecting at least one set of pre-extraction rules from a grouping rule base; generating multiple pre-production sequences based on at least a portion of the semiconductor generation data using the pre-extraction rules; dividing the multiple pre-production sequences into multiple sub-production sequences, each sub-production sequence having a preset length; obtaining the analysis results of the multiple sub-production sequences and the actual results input by the user; and scoring the pre-extraction rules based on the difference between the analysis results and the actual results.
[0012] In some embodiments, when the data type is discrete, the mutual information is calculated as follows: ;in, Representation of features The value, Represents the measured value of category Y; Features and categories The joint probability distribution of ; Features The marginal probability distribution; is the marginal probability distribution of the feature refers to the feature in the i-th production sequence, and the feature refers to the equipment type or configuration type; the category refers to the category of the measurement value of a production sequence.
[0013] In some embodiments, when the data type is continuous, the mutual information indicator is calculated as follows: ; wherein, represents the value of the feature , represents the measurement value of the category Y; is the joint probability distribution of the feature and the category ; is the marginal probability distribution of the feature ; is the marginal probability distribution of the category Y; wherein, the feature refers to the feature in the i-th production sequence, and the feature refers to the equipment type or configuration type; the category refers to the category of the measurement value of a production sequence.
[0014] In some embodiments, the root cause weight is calculated as follows: ; represents the feature of the j-th sub-production sequence in the i-th production sequence, is the proportion of the number of wafers that have experienced the i-th production sequence in the total number of wafers, is the number of production sequences in a product sequence group.
[0015] Another aspect of the present application also provides a semiconductor big data traceability analysis system, comprising: a data acquisition module, configured to acquire semiconductor production line data, wherein the semiconductor production line data at least includes: (1) a plurality of semiconductor products that have experienced a plurality of processing equipment information, (2) process parameters used when the semiconductor products have experienced at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; a scheme calling module, configured to call a data analysis scheme for the semiconductor production line data, wherein the data analysis scheme includes a fault prediction algorithm; a preprocessing module configured to preprocess the semiconductor production line data according to a data analysis scheme to generate a preprocessing result, the preprocessing result comprising: a predicted abnormal root cause, and / or an abnormal weight of the abnormal root cause; wherein the preprocessing module comprises: a rule acquisition unit configured to acquire an extraction rule, the extraction rule comprising: a device type and a configuration type, the device type comprising at least one piece of processing device information, and the configuration type comprising at least one piece of process parameter information and / or at least one piece of the category information; a sequence identification unit configured to identify a plurality of production sequences from the semiconductor production line data according to the extraction rule; and a fault prediction unit configured to predict an abnormal root cause according to a difference between at least one production sequence and other production sequences using a fault prediction algorithm; an analysis module configured to input the preprocessing result and corresponding semiconductor production line data into a preset big data platform, the big data platform corresponding to an output analysis result.
[0016] Beneficial technical effects: It should be noted that in the process of semiconductor manufacturing, it may need to go through thousands of processes, and the parameters of each process (such as the formula set by the engineer, or the working state of the device itself, etc.) may also differ. Therefore, a semiconductor will record a large amount of production line data. Moreover, once the yield of a semiconductor decreases, it is very difficult to analyze the root cause from the vast amount of production line data, even experienced engineers may need to spend several days or even weeks to locate the fault root.
[0017] To this end, the present application proposes a preprocessing and AI processing coordinated big data traceability analysis method (i.e. root cause analysis method), that is, first, the root cause is roughly located (such as filtering out the area where the root cause is highly likely to exist) through preprocessing to obtain a preprocessing result (such as a possible abnormal root cause, or an abnormal weight of the abnormal root cause) and corresponding detailed information (i.e. production line data) input into a preset AI platform, and the production line data is analyzed in depth by the AI platform under the guidance of the preprocessing result. This coordinated analysis based on the preprocessing result and the production line data can effectively improve the efficiency of AI analysis.
[0018] Moreover, it should be noted that the traditional root cause analysis often uses an analysis scheme based on complete process parameters, such as combining complete process parameters, such as FDC data sets (which often contain relatively complete device parameters such as temperature, pressure, flow, voltage, and current) to trace the root cause. That is, the traditional root cause analysis is a parameter-based analysis.
[0019] On the contrary, the present application adopts a non-parametric analysis scheme, such as the present application only needs to extract production sequences (which do not need to contain detailed process parameters, such as only need to record process sequence and process category) from the FDC data set, and predict possible root cause types according to the production sequence. And this non-parametric analysis scheme helps to realize faster preliminary analysis, and to a certain extent, improve the analysis efficiency.
[0020] Further, the preliminary results obtained by the comprehensive non-parametric analysis means, combined with the corresponding complete data set, can be based on the AI platform to analyze the possible root cause in detail, thereby efficiently providing the user with relatively reliable and complete analysis results.
[0021] Specifically, for the non-parametric analysis means, the present application actually provides a piecewise data preprocessing method, specifically, the present application adopts grouping and segmentation to obtain sub-production sequences with analysis value, and the grouping and segmentation includes the following steps: 1) grouping the semiconductor production line data according to the equipment type and configuration type, and extracting the production sequence; 2) cutting the production sequence to obtain multiple segments (i.e. sub-production sequences). BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. In all the drawings, similar elements or parts are generally identified by similar reference signs. In the drawings, each element or part is not necessarily drawn according to the actual proportion. Obviously, the drawings described below are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0023] Figure 1 Method flowchart in an exemplary embodiment of the present application; Figure 2 Production line correlation diagram in an exemplary embodiment; Figure 3 Product flow diagram in another exemplary embodiment; Figure 4 System module architecture diagram in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0025] Herein, the suffix such as "module", "part" or "unit" used to represent an element is only for the convenience of description of the present application, and has no specific meaning by itself. Therefore, "module", "part" or "unit" can be mixedly used.
[0026] Herein, the terms "upper", "lower", "inner", "outer", "front", "back", "one end", "the other end" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of description of the present application and simplification of the description, and do not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.
[0027] Herein, unless otherwise explicitly specified and limited, the terms "mount", "provided with", "connected" and the like should be understood broadly, for example, "connected" can be fixedly connected, or detachably connected, or integrally connected; can be mechanically connected, can be directly connected, or indirectly connected through an intermediate medium, can be the communication inside two elements. For a person of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0028] Herein, "and / or" includes any and all combinations of one or more listed related items.
[0029] Herein, "a plurality of" means two or more, that is, it includes two, three, four, five, etc.
[0030] In the present specification, the term "about" typically means + / - 5% of the stated value, more typically + / - 4% of the stated value, more typically + / - 3% of the stated value, more typically + / - 2% of the stated value, even more typically + / - 1% of the stated value, even more typically + / - 0.5% of the stated value.
[0031] In this specification, certain embodiments can be disclosed in a format that is a range. It is to be understood that such a range format is used only for convenience and brevity and should be interpreted in a flexible manner to include not only the numerical values explicitly recited as the limits of a range, but also to include all the individual numerical values corresponding to the range bounds. For example, a range of 1 to 6 should be interpreted to include not only the explicitly recited limits of 1 and 6, but also the individual numbers 2, 3, 4, 5, and 6. The same applies to ranges including these individual numbers. A range of 1 to 5 should be interpreted to include not only the explicitly recited limits of 1 and 5, but also the individual numbers 2, 3, 4, and 5. The same applies to ranges including these individual numbers.
[0032] In this specification, a "wafer" refers to a silicon wafer used for manufacturing a silicon semiconductor integrated circuit. Since the shape of the silicon wafer is generally circular, it is also referred to as a wafer. Various circuit element structures can be processed on the silicon wafer to become an IC product having a specific electrical function.
[0033] In this specification, a "lot" refers to a group of wafers in a basic batch. For example, generally, 1 lot is 12 wafers.
[0034] Embodiment One Referring to Figure 1 As shown in the drawings, the present application provides a semiconductor big data traceability analysis method, comprising the steps of: S101, acquiring semiconductor production line data, the semiconductor production line data at least comprising: (1) a plurality of processing equipment information experienced by a plurality of semiconductor products, (2) process parameters adopted when the semiconductor products experience at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information comprises at least one of the following: site information, machine table information, chamber information, the process parameters comprise at least one of the following: formula, sensor type, and the category information comprises at least one of the following: raw material source, semiconductor model; For example, in some embodiments, the semiconductor production line data can come from an FDC system, or can come from a WAT system. The semiconductor production line data can record the processing information experienced by a semiconductor product in the whole process from entering the production line to leaving the production line, such as the model, category, number, etc. of the processing equipment (such as site, machine table, chamber) experienced, and the process parameters (equivalent to process parameters) of each processing equipment, such as temperature, pressure, voltage, current, etc.
[0035] For example, in some embodiments, the measurement value can be the result of measuring the yield indicators such as structure, size or performance of the semiconductor product after passing through any one processing equipment.
[0036] Or, in other embodiments, the measurement value can also be a process parameter of the processing equipment, such as voltage, current, etc., measured when the semiconductor product is processed by any one processing equipment, which can reflect whether the equipment is in a healthy running state, or reflect the probability of defects or faults on the semiconductor product.
[0037] S102, calling a data analysis scheme for the semiconductor production line data, the data analysis scheme comprising: a fault prediction algorithm; S103, pre-processing the semiconductor production line data according to the data analysis scheme to generate a pre-processing result, the pre-processing result comprising: a predicted abnormal root cause, and / or an abnormal weight of the abnormal root cause; S104, inputting the pre-processing result and the corresponding semiconductor production line data into a preset big data platform, the big data platform corresponding to output an analysis result.
[0038] In some embodiments, the fault prediction algorithm comprises: a single factor variance analysis algorithm, a history analysis algorithm, and / or a similarity analysis algorithm.
[0039] In some embodiments, the big data platform is configured with one or more fault prediction AI models (also referred to as AI platform) trained in advance, which can analyze abnormal points through the semiconductor production line data, thereby predicting the root cause of the fault.
[0040] In some embodiments, the big data platform is an Apache Spark big data platform.
[0041] In this embodiment, the pre-processing result and the corresponding semiconductor production line data are synchronously input into the big data platform, so that the big data platform can perform localized focused analysis on the massive semiconductor production line data based on the pre-processing result, thereby improving the efficiency of fault prediction and reducing the requirement of analysis computing power for massive semiconductor production line data.
[0042] Specifically, the present application proposes a big data traceability analysis method (i.e. root cause analysis method) of pre-processing and AI processing in cooperation, that is, first roughly positioning the root cause (such as screening out the area with high probability of root cause) through pre-processing to obtain the pre-processing result (such as possible abnormal root cause, or abnormal weight of abnormal root cause) and the corresponding detailed information (i.e. production line data) input into the preset AI platform, and then the AI platform analyzes the production line data in depth under the guidance of the pre-processing result. This way of cooperative analysis based on the pre-processing result and the production line data can effectively improve the efficiency of AI analysis.
[0043] And, it is worth noting that the traditional root cause analysis often uses an analysis scheme based on complete process parameters, such as the need to combine complete process parameters, such as FDC data sets (which often contain relatively complete device parameters such as temperature, pressure, flow, voltage, and current). That is, the traditional root cause analysis is a parameter-based analysis.
[0044] In contrast, the present application uses a non-parametric analysis scheme, such as the present application only needs to extract the production sequence (which does not need to include detailed process parameters, such as only needs to record the process sequence and process category) from the FDC data set, and predicts the possible root cause type according to the production sequence. This non-parametric analysis scheme helps to achieve faster preliminary analysis and improves analysis efficiency to some extent.
[0045] Further, the preliminary results obtained by the non-parametric analysis method are combined with the corresponding complete data set, and the possible root cause can be analyzed in detail based on the AI platform, thereby efficiently providing the user with relatively reliable and complete analysis results.
[0046] In the following, the non-parametric analysis method proposed by the present application will be further described in detail: In some embodiments, S103 includes the steps of: S1031, obtaining an extraction rule, the extraction rule comprising: a device type and / or a configuration type, and the device type comprising at least one piece of processing equipment information, and the configuration type comprising at least one piece of process parameter information and / or at least one piece of the category information; Preferably, the extraction rule comprises: a device type and a configuration type; S1032, identifying a plurality of production sequences from the semiconductor production line data according to the extraction rule, and the production sequences corresponding to the measurement values recorded; S1033, using a fault prediction algorithm to predict an abnormal root cause according to the difference between at least one production sequence and other production sequences.
[0047] In this embodiment, the extraction rule can also be regarded as a grouping rule, and data meeting the same extraction rule is recorded in a production sequence, and a plurality of production sequences meeting the same rule are regarded as a production sequence group.
[0048] Referring to Figure 2 When the extraction rule includes: site, machine, and chamber, the history path of wafer1 product that experiences the same site, machine, and chamber can be extracted as a production sequence, and the history path of wafer2 product can be extracted as another production sequence. That is, the site, machine, and chamber information is equivalent to a classification factor.
[0049] It can be understood that the more rules (or classification factors) in the extraction rule, the finer the granularity of the grouping. Specifically, the extraction rule can be adaptively selected according to actual analysis requirements.
[0050] In some embodiments, the production sequence at least records the processing sequence of at least two processing equipment experienced by the semiconductor product, and the corresponding measurement value; correspondingly, S1033 further includes: S10331, obtaining at least one production sequence group, and the production sequences in the same group have the same equipment type and configuration type; S10332, using a fault prediction algorithm to calculate the difference between a production sequence and other production sequences in the same group to predict the abnormal root cause.
[0051] For example, in some embodiments, a group of production sequences includes abnormal sequences and normal sequences, where the abnormal sequence refers to a production sequence corresponding to a semiconductor product that is not qualified in yield or measurement value, and the normal sequence refers to a production sequence corresponding to a semiconductor product that is normal in yield or qualified in measurement value. By calculating the difference between the abnormal sequence and the normal sequence in the production condition, the root cause of the problem that may occur can be preliminarily predicted.
[0052] For example, in some embodiments, a production sequence includes semiconductor products passing through n processing equipment of the same type (that is, at least two processing equipment are configured under at least one equipment type). For example, a process can include multiple identical stations, such as station 1, station 2, … (that is, station 1 and station 2 can be used to perform the same process). For example, a station can include multiple identical parallel chambers, such as chamber 1 and chamber 2, and the chamber numbers of different chambers need to be recorded. The parameterized analysis method provided by the present application can directly analyze the root cause area (such as the machine that may cause failure) from the equipment processing sequence, equipment type and difference, without getting into specific parameters such as voltage and current abnormality identification. This non-parametric root cause area prediction and step-by-step analysis method in cooperation with the AI platform are beneficial to more quickly and accurately analyze the fault.
[0053] In some embodiments, when the fault prediction algorithm is the history analysis algorithm, S10332 includes the steps of: (1) identifying the data type of the measurement value, wherein the data type includes discrete and / or continuous; (2) selecting a corresponding mutual information algorithm according to the data type; (3) calculating the mutual information of the production sequence according to the mutual information algorithm wherein the mutual information is used to represent an amount of information contained in the random variable Y about ; wherein, refers to a feature of the production sequence, such as may represent a device type or a configuration type of the i-th production sequence.
[0054] (4) calculating a root cause weight of at least one processing device in the production sequence according to the mutual information.
[0055] For example, in some embodiments, a root cause weight of a production sequence can be calculated according to the mutual information, or a root cause weight of a specific processing device in a production sequence can also be calculated according to the mutual information.
[0056] In some embodiments, before S10332, further comprising the steps of: obtaining a sequence number within a group of product sequences; determining whether the sequence number is greater than a preset number, and if so, proceeding to step S10332.
[0057] For example, in some embodiments, the preset number is 1.
[0058] In some embodiments, before S1031, further comprising the steps of: selecting at least one set of pre-extraction rules from a rule library; generating a plurality of pre-production sequences according to at least a part of the semiconductor production data using the pre-extraction rules; dividing the plurality of pre-production sequences into a plurality of sub-production sequences, the sub-production sequences having a preset length; obtaining analysis results of the plurality of sub-production sequences and actual results input by a user; scoring the pre-extraction rules according to a difference between the analysis results and the actual results.
[0059] For example, in some embodiments, the preset length can also be selected / updated, wherein the selection / update of the preset length comprises: (1) cutting the pre-production sequences using the preset length and a test length respectively to obtain a first cutting group and a second cutting group correspondingly; wherein the cutting group includes a plurality of sub-production sequences after cutting; for example, in some embodiments, the preset length can be a default length preset by a user, or the preset length can be a cutting length used in the last cutting of the pre-production sequences. The test length can be a new cutting length obtained by increasing or decreasing the preset length, which can be manually input by a user, or can be randomly generated according to the preset length.
[0060] It should be noted that the cutting of the production sequence in this embodiment is equivalent to dividing the production process into multiple smaller stages.
[0061] (2) Obtain the first prediction accuracy of the first cutting group and the second prediction accuracy of the second cutting group; wherein the prediction accuracy refers to the difference between the analysis result based on the fault prediction algorithm and the actual result input by the user; (3) Calculate the prediction difference between the first prediction accuracy and the second prediction accuracy; (4) When the prediction difference is greater than a preset prediction threshold (which can be set or adjusted by the user), then go to step (5); (5) Obtain the first prediction time corresponding to the first cutting group and the second prediction time of the second cutting group; (6) Calculate the time difference between the first prediction time and the second prediction time; (7) When the time difference is less than a preset time threshold, then suggest updating the test length to a new preset length.
[0062] It should be noted that in order to improve the overall fault analysis efficiency, the present application introduces a cutting group analysis mode for the early non-parametric analysis process, that is, by dividing a large production sequence into multiple small sequences (i.e. sub-production sequences), the parallel operation of multiple small sequences is realized to quickly locate the root cause area.
[0063] At the same time, the present embodiment also provides a restrictive cutting size adjustment scheme for non-parametric analysis, which limits the cutting length based on prediction time consumption and prediction accuracy to restrictively update / adjust the cutting length, thereby being able to quickly predict the area suspected to have a fault within a relatively limited time, while ensuring that the prediction result has a certain reliability.
[0064] For example, for the same type of wafer manufacturing task or the same type of wafer production line, a pre-test method can be used to screen the appropriate cutting size. For example, before the step-by-step analysis method is formally applied, the appropriate cutting size can be screened through the initial trial stage.
[0065] In some embodiments, when the data type is discrete, the mutual information indicator is calculated as follows: ; wherein, represents the value of the feature (such as site category, machine category, etc.), represents the measured value of category Y; is the feature and category a joint probability distribution of the features and the categories, which is used to describe or represent the probability of the random variables and the categories the probability law of the simultaneous occurrence. is the marginal probability distribution of the features and the marginal probability distribution is used to describe or represent the prior probability of the specific feature in all possible states after integrating (or ignoring) the influence of all other features and category information. is the marginal probability distribution of the categories Y, and the marginal probability distribution is used to describe or represent the prior distribution or base rate of each category in the population after integrating (or ignoring) the specific values of all other features; wherein the features refer to the features in the i-th production sequence; denotes the measurement value of the category Y.
[0066] In some embodiments, when the data type is continuous, the mutual information is calculated as follows: ; wherein denotes the value of the feature , denotes the measurement value; is the joint probability distribution of the features and the categories ; is the marginal probability distribution of the features ; is the marginal probability distribution of the categories Y; wherein the features refer to the features in the i-th production sequence; denotes the measurement value of the category Y.
[0067] In some embodiments, the root cause weight is calculated as follows: ; represents the features of the j-th sub-production sequence (or the j-th stage) in the i-th production sequence, refers to the mutual information of the j-th sub-production sequence in the i-th production sequence, is a penalty factor, wherein denotes the wafer proportion (i.e., the proportion of the number of wafers that have undergone the i-th production sequence in the total number of wafers), is the number of production sequences in a product sequence group.
[0068] It should be noted that the above probability-based failure prediction algorithm is only an exemplary analysis method provided by the present application, and other non-parametric failure analysis methods can also be used according to specific analysis needs.
[0069] In some embodiments, the probability density is estimated using distance information of K-nearest neighbors, and exemplary calculation steps are as follows: Given a sample point , the distance of the kth nearest neighbor of the point is calculated; The number of other points within this distance range is counted (for estimating local density); Using these local density information, the mutual information contribution of each sample point is estimated; The final mutual information estimate is obtained by averaging all sample points.
[0070] For example, assume that the collected data has m site groups, each group has s stages (e.g., a sub-production sequence can be divided based on a stage), each stage has site chamber combinations, the number of wafers in the current group is wafer_ratio of the total number of wafers, and the mutual information algorithm I is called when =1, i.e., all wafers have passed through a chamber, and at this time, it is impossible to separate, .
[0071] First exemplary embodiment: Collecting production line data based on user input information, and adding a column value: corresponding to the wafer user value information.
[0072] Grouping based on user input grouping fields, each group is analyzed separately, for example, in the case of history chamber analysis, the results of a group are as follows.
[0073] Or History analysis of a group, each wafer is sorted according to time from front to back, divided into different stages, and the grouping results about Figure 2 can be shown in the following table: Furthermore, based on the above grouping results, a difference analysis is performed: If, in any given stage, for continuous values (i.e., values), those with high values appear in one chamber and those with low values in another, and for discrete values, such as 0 values appearing in one chamber and 1 values in another, these values can be completely separated, the greater the likelihood that this situation will be identified as the root cause. This embodiment introduces a mutual information analysis method to measure the difference in the history analysis.
[0074] Second exemplary embodiment: The following section uses another example of production line data to illustrate the implementation of the history analysis algorithm: Resume Information Form The path diagram corresponding to the above resume information table is as follows: Figure 3 As shown. STAGE1 contains machine tool M1, and STAGE2 contains all the chambers of machine tool M2. The path from STAGE1 to STAGE2 is shown as follows. Figure 3 The connection in the diagram, VALUE (i.e., the measured value) is discrete here, 0 is GOOD marked in green, and 1 is BAD marked in red. The flow diagram can intuitively show the path flow of GOOD and BAD wafers.
[0075] Step 1: Before calculating the mutual information between STAGE1 and VALUE, because the STAGE1 column is text type, it needs to be encoded using natural numbers (a simple encoding that maps categorical data to a sequence of natural numbers (0, 1, 2, 3,...)). The encoded STAGE1 column is as follows: Correspondingly, = STAGE1_encoding; Similarly, calculate =STAGE2_encoding.
[0076] The encoded STAGE2 column is as follows: Step 2: Calculate mutual information, where y is the VALUE column information, which is of numeric type.
[0077] = ; .
[0078] Third step: calculate the weight of the site, assuming all wafer quantities are 16 pieces, the wafer ratio of the group is 0.5, and based on the history table, the weight of each stage chamber combination can be calculated .
[0079] According to the weight calculation formula, the weight of the site is as follows: Finally, all the weight parameters of all sites of all groups are summarized, and the root cause weight ranking from large to small is obtained. The output result is as follows: It should be noted that for a 12-inch fab, it involves TB level data every month. With the help of Apache Spark, a big data analysis engine, batch processing, stream processing, graph computing and machine learning can be supported. The core advantage of Spark is in-memory computing, which is faster than Hadoop, while providing high fault tolerance and easy-to-use API (supporting Scala, Java, Python and R).
[0080] The main modules of the Spark big data platform include SparkCore, SparkSQL, SparkStreaming, MLlib (machine learning) and GraphX (graph computing). In this application, three kinds of traceability analysis algorithms are designed and run on spark, which are suitable for application in FDC, WAT, Inline, wafer history information and other data sources, so that large amount of data analysis becomes possible, and the result generation speed is realized in the fastest minute level and the slowest hour level.
[0081] That is to say, one of the purposes of the present application is to provide a root cause analysis method based on 12-inch semiconductor Fab manufacturing data, so as to solve the problems of scattered data sources, low analysis efficiency, dependence on manual judgment and limited analysis granularity in the prior art.
[0082] To achieve the above purpose, the present application constructs three kinds of key algorithm models based on Apache Spark big data platform, including data grouping analysis algorithm, history analysis algorithm and similarity analysis algorithm, which are suitable for different types of input data (such as discrete labels, continuous output values, abnormal characteristic parameters, etc.), so as to realize unified analysis and efficient traceability of various data sources such as FDC, WAT, Inline and wafer history.
[0083] The application can complete the automatic root cause identification of large-scale manufacturing data in the time range of minutes to hours, improve the speed and accuracy of root cause positioning, reduce the work intensity of engineers, and provide strong support for process optimization and quality control.
[0084] Embodiment two Next, taking a data grouping analysis algorithm as an example, an exemplary analysis method provided by the application is introduced and described: The user gives a batch of WAFER a GOOD and BAD label, and locates the root cause from a selected time range, and the particle size of the location can be adjusted, such as (product + site + machine + Sensor_Name) or (product + site + machine + chamber + Sensor_Name). The specific analysis process includes the following steps: Step 1: Select the data source to be analyzed, such as FDC, WAT, Inline, upload wafer information, mark “GOOD” and “BAD”, select the time range for tracing, and if further refinement is needed, filter through product, site, machine, chamber, and Lot information.
[0085] Step 2: Select the grouping analysis particle size to determine the degree of positioning. For example, for FDC data source, the tracing particle size is selected as three kinds: 1. site + machine + chamber 2. site 3. site + machine, and Parameter is the smallest parameter for tracing. For example, FDC Parameter is SensorName#Window#stats (the statistical characteristics of a group of consecutive points in the specified set window, such as ‘max’: maximum value, ‘min’: minimum value, ‘min’: average value). Click to run and submit the analysis task.
[0086] Step 3: The platform submits the task information to the spark big data analysis engine of the platform, calls the data grouping analysis algorithm, and outputs the results of the algorithm to the database after running.
[0087] Step 4: The platform displays the results of the task in the platform result viewing.
[0088] Among them, the core algorithm part steps are as follows: Step 1: Collect data set: select data from the user-specified data source with the user-provided waferlist, good, and bad label information. About wafer column assignment, the ‘label’ column takes value 0 to represent ‘GOOD’ and value 1 to represent ‘BAD’.
[0089] Step 2: Select product, machine, chamber, site, Parameter, wafer, 'label' tag column data, remove rows with missing values.
[0090] Step 3: Grouping based on product + grouping information, score all Parameters for each group, labeled as "weight". For example, Group 1, grouping fields are site, machine, chamber.
[0091] Step 4: If there is a significant difference in the label group, parameter, the parameter is more likely to be the root cause. One case of root cause: because the sensor corresponding to the equipment accessory at that moment appears abnormal, leading to the collection of data too high or too low. There is a large separation phenomenon in the form of expression. In the post-analysis, one-way analysis of variance is used to judge the difference of wafer in parameter between good and bad groups.
[0092] One-way analysis of variance, also known as one-dimensional variance analysis, is used to analyze whether the mean of the dependent variable is significantly different when different levels of a single control factor are taken. Different levels correspond to the two groups of GOOD and BAD, and the control factor is a single parameter.
[0093] Specifically, the core calculation process of step 4 is as follows: First, give the hypothesis definition: H0 (Null Hypothesis): The means of all groups are equal (μ1=μ2=...=μ k ), k is the number of groups.
[0094] I1 (Alternative Hypothesis): At least one group mean is different from the others.
[0095] Calculate the total sum of squares SST (total sum of squares): the sum of squares of the deviation of all observations from the overall mean: : the jth observation of the ith group.
[0096] ..: the total average of all data.
[0097] Between-group sum of squares (SSB): The between-group sum of squares measures the difference between the means of each group and the overall mean: : the sample size of the ith group.
[0098] Mean of the ith group.
[0099] Sum of squares within (SSW): measures the variation of data points within each group from the mean of the group: Degrees of freedom for calculation: Total degrees of freedom: (N is the total number of samples).
[0100] Degrees of freedom between groups: (k is the number of groups).
[0101] Degrees of freedom within groups: .
[0102] Mean square error (MSE) for calculation: Mean square between (MSB): ; Mean square within (MSW): ; Calculate F-statistic for comparing variance between groups and within groups: Determine the significance level, set to 0.05, calculate the p-value, which is calculated according to the F-distribution.
[0103] Cumulative distribution function of F-distribution, which represents the probability of appearing less than or equal to the calculated F-statistic under the given degrees of freedom between groups and degrees of freedom within groups: ; The p-value is the probability of the opposite event.
[0104] ; Determine significance: When , the difference is significant, reject the original hypothesis, and choose the alternative hypothesis.
[0105] Otherwise, the difference is not significant, accept the original hypothesis.
[0106] The smaller the p-value is in the case of less than 0.05, the greater the difference, and the F-statistic value is the greater the better.
[0107] Specifically, in step 3, if the number of rows of "GOOD" or "BAD" in the grouping + Parameter is less than 3, the sample size is too small, the analysis is not reliable, and the root cause weight is set to 0.
[0108] Otherwise, call the variance analysis algorithm, given the total number of groups m, the total number of parameters in the group n, and calculate the F-statistic of all parameters in the group, The F-statistic for the j-th parameter in the i-th group has the following weight formula: This represents the proportion of Badwafer fragments within this group to the total number of Badwafer fragments. It indicates that the more Badwafer fragments there are in a group, the greater the likelihood that it is the root cause.
[0109] By summing the weights of all parameters for all groups and sorting them from largest to smallest, we obtain the root cause weight ranking from largest to smallest. The output is shown below: Example 3 The following example, using a similarity analysis algorithm, further illustrates the analysis method provided by this invention: The analysis method in this embodiment is applicable to situations where a set of outlier parameters is known. It searches other data sources, such as FDC, WAT, and Inline, for a set of parameters related to the wafer and checks for linear, square, or radical relationships. If any one of these relationships exists, it can be used to set the root cause weights. The algorithm designed here uses linear regression and then evaluates the R² of the goodness of fit as the root cause weight.
[0110] The steps for similarity analysis in business analysis are as follows: Step 1: Select the data source for similarity analysis and comparison, such as FDC, WAT, or Inline. Select a batch of wafers with anomaly characteristics, such as FDC, WAT, or Inline. Select the time range for tracing. If further refinement is needed, filtering can be done by product, site, machine, chamber, or lot information.
[0111] Step 2: Select Groups: The granularity of the analysis determines the level of detail in the localization. For example, for the FDC data source, the granularity of tracing can be selected in three ways: 1. Site + Machine + Chamber; 2. Site; 3. Site + Machine. Parameter is the smallest parameter for tracing. For example, FDCParameter is SensorName#Window#stats (the statistical characteristics calculated for a group of consecutive data points within the specified window, such as 'max': maximum value, 'min': minimum value, 'min': average value). Click Run to submit the analysis task.
[0112] Step 3: The platform submits the task information to the platform's Spark big data analysis engine, calls the data similarity analysis algorithm, and outputs the results of the algorithm to the database after the operation is completed.
[0113] Step 4: The platform displays the results of the task in the platform results view.
[0114] The specific similarity analysis algorithm may include the following steps: Step 1: Data Collection: Collect the necessary columns from the data source to be compared for similarity, such as a parameter of an anomaly characterization column that is inlined. The data source being compared is FDC. Example results are shown below: Step 2: Based on the defined data source for comparison, such as FDC, and the grouping fields are site + machine + chamber, compare the Parameter with the anomaly representation value. Group the data from Step 2 according to site + machine + chamber, and check if there is a linear relationship, a square relationship, or a square root relationship among them; satisfying any one of these three conditions is sufficient. Train a linear regression on the Parameter and anomaly representation value columns, and evaluate the goodness of fit between the true and predicted values as the similarity score.
[0115] The principle of step 2 is as follows: Linearity detection: Using linear regression, known... The last column is the intercept column, all of which are 1. The sample size is N, the number of features is d, and the slope vector is... (Including intercept), target vector , For the intercept term, construct the model: ; in, For a linear regression model, X represents the FDC / WAT / Inline parameter to be detected (e.g., the parameter column of FDC in the grouping), and R represents a real number.
[0116] The loss function is: ; Minimize the loss to obtain the optimal solution: ; Here, Y represents the anomaly representation values of a selected set of wafers, see the value column in the table.
[0117] Square relation test: Consider the simplest case, d=1, where there is only one-dimensional feature x=>{ }; Then, by substituting the solutions into linear regression, the optimal solution can be obtained.
[0118] Square root relation check: Equivalent to and There exists a linear case where, for any x, if x is less than 0, we transform x by subtracting the minimum value from each element of x, so that x as a whole is greater than or equal to 0. Then, we substitute this into a linear regression model.
[0119] Methods for determining the weights in root cause analysis: Given a one-dimensional x and a one-dimensional y as input; Detect the linear relationship and construct a linear regression model of y and x; Detecting square relationships: Constructing y and x, The regression model; Detecting square root relationships: Constructing The regression model with x.
[0120] The goodness of fit between the predicted and actual values is calculated sequentially; it represents the model's ability to explain changes in the target variable. ; in, The predicted value of the model. The mean of the target variable. Coefficient of determination. The closer to 1, the better; coefficient of determination A value of 0 is equivalent to mean prediction; the coefficient of determination is... A value less than 0 indicates that the model performs worse than "predicting using the mean".
[0121] For the existence of a determination coefficient less than 0 Change it to 0, then find the largest of the three. , as a similarity score.
[0122] Assume the data has a total of m groups, and each group has The algorithm has one parameter and is based on linear, square, or radical relationships; if any one of these three conditions is met, it is considered an abnormal algorithm. If the sample size is less than or equal to 3, the sample size is too small and the data is unreliable. 0; Otherwise, use the similarity scoring algorithm to calculate: ; For training linear regression models, square regression models, and radical regression models, This is the goodness-of-fit formula.
[0123] Step 3: Summarize the weight parameters of all sites in all groups, sort them from largest to smallest, and you will get the root cause weight ranking from largest to smallest. For example, the output result is as follows: It should be noted that the data generated during the manufacturing process in a fab is massive and diverse; on the other hand, traditional analysis methods and technical architectures are difficult to effectively cope with the challenges of such high-dimensional, heterogeneous, and real-time data.
[0124] This invention addresses the challenges of high-dimensional, heterogeneous, and real-time data by providing an improved solution that overcomes many shortcomings of existing technologies and offers a method for large-scale data tracing and analysis suitable for 12-inch fabs.
[0125] Alternatively, this invention provides a root cause analysis method based on manufacturing data from a 12-inch semiconductor fab, to solve the problems of scattered data sources, low analysis efficiency, reliance on manual judgment, and limited analysis granularity in the source tracing analysis process of existing technologies.
[0126] To achieve the above objectives, this invention constructs three key algorithm models based on the Apache Spark big data platform: a data grouping analysis algorithm, a history analysis algorithm, and a similarity analysis algorithm. These algorithms are applicable to different types of input data (such as discrete labels, continuous output values, and anomaly representation parameters), thereby enabling unified analysis and efficient tracing of multiple data sources, including FDC, WAT, Inline, and wafer history.
[0127] This invention can automate the root cause identification of large-scale manufacturing data within a time range of minutes to hours, improve the speed and accuracy of root cause localization, reduce the workload of engineers, and provide strong support for process optimization and quality control.
[0128] Example 4 See Figure 4 As shown, the present invention also provides a semiconductor big data traceability and analysis system, comprising: The data acquisition module 101 is used to acquire semiconductor production line data, which includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; The scheme invocation module 102 is used to invoke a data analysis scheme for the semiconductor production line data, the data analysis scheme including: a fault prediction algorithm; Preprocessing module 103 is used to preprocess the semiconductor production line data according to the data analysis scheme to generate preprocessing results, the preprocessing results including: predicted root causes of anomalies, and / or anomaly weights of the root causes of anomalies; wherein, the preprocessing module includes: The rule acquisition unit 1031 is used to acquire extraction rules, the extraction rules including: device type and configuration type, wherein the device type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; The sequence recognition unit 1032 is used to identify multiple production sequences from the semiconductor production line data according to the extraction rules; The fault prediction unit 1033 is used to predict the root cause of the anomaly based on the difference between at least one production sequence and multiple other production sequences using a fault prediction algorithm. The analysis module 104 is used to input the preprocessing results and the corresponding semiconductor production line data into a preset big data platform, and the big data platform outputs the analysis results accordingly.
[0129] It should be noted that the system in this embodiment can implement the methods or steps in any of the above embodiments, which will not be repeated here.
[0130] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a computer terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0132] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A root cause analysis method, characterized in that, Including the following steps: S101, acquire semiconductor production line data, wherein the semiconductor production line data includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; S102, invoke a data analysis scheme for the semiconductor production line data, the data analysis scheme including: a fault prediction algorithm; S103, preprocess the semiconductor production line data according to the data analysis scheme to generate preprocessing results, the preprocessing results including: predicted root causes of anomalies, and / or anomaly weights of the root causes of anomalies; wherein, S103 includes the following steps: S1031, Obtain extraction rules, the extraction rules include: equipment type and configuration type, wherein the equipment type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; S1032, Multiple production sequences are identified from the semiconductor production line data according to the extraction rules, and the production sequence is correspondingly recorded with the measured value; S1033, a fault prediction algorithm is used to predict the root cause of the anomaly based on the difference between at least one production sequence and multiple other production sequences. S104, the preprocessing results and the corresponding semiconductor production line data are input into a preset big data platform, and the big data platform outputs the corresponding analysis results.
2. The method according to claim 1, characterized in that, The fault prediction algorithm includes: one-way ANOVA algorithm, history analysis algorithm and / or similarity analysis algorithm.
3. The method according to claim 1, characterized in that, The big data platform mentioned is the Apache Spark big data platform.
4. The method according to claim 1, characterized in that, The production sequence includes abnormal sequences and normal sequences.
5. The method according to claim 4, characterized in that, The abnormal sequence refers to the production sequence in which the yield of the corresponding semiconductor product is unqualified or the measurement value is unqualified, while the normal sequence refers to the production sequence in which the yield of the corresponding semiconductor product is normal or the measurement value is qualified.
6. The method according to claim 1, characterized in that, One of the production sequences corresponds to a semiconductor product processed by n identical types of processing equipment.
7. A root cause analysis system, characterized in that, include: The data acquisition module is used to acquire semiconductor production line data, which includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; The scheme invocation module is used to invoke a data analysis scheme for the semiconductor production line data, and the data analysis scheme includes: a fault prediction algorithm; A preprocessing module is used to preprocess the semiconductor production line data according to a data analysis scheme to generate preprocessing results, the preprocessing results including: predicted root causes of anomalies, and / or anomaly weights of the root causes of anomalies; wherein, the preprocessing module includes: The rule acquisition unit is used to acquire extraction rules, the extraction rules including: device type and configuration type, wherein the device type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; A sequence recognition unit is used to identify multiple production sequences from the semiconductor production line data according to extraction rules; The fault prediction unit is used to predict the root cause of anomalies based on the differences between at least one production sequence and multiple other production sequences using a fault prediction algorithm. The analysis module is used to input the preprocessing results and the corresponding semiconductor production line data into a preset big data platform, and the big data platform outputs the analysis results accordingly.
8. The system according to claim 7, characterized in that, The fault prediction algorithm includes: one-way ANOVA algorithm, history analysis algorithm and / or similarity analysis algorithm.
9. The system according to claim 7, characterized in that, The big data platform mentioned is the Apache Spark big data platform.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements a root cause analysis method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
A FDC (Fault-Derived Conversion) Cause Analysis Method and Storage Medium Based on Distributed Parallel Computing
CN116629707B
Abnormal root cause analysis method and device
CN118056189A
Defective root cause analysis method and system based on comprehensive analysis framework
CN119558541A
Semiconductor anomaly detection method and system based on physical causal relationship modeling
CN120781269A
Failure diagnosing system and method for semiconductor manufacturing device
JP2010015205A