A root cause analysis method, system, and program product

CN121707429BActive Publication Date: 2026-09-25SHENZHEN EXX IND AUTOMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610062771.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-09-25
Estimated Expiration
2045-10-31

AI Technical Summary

Benefits of technology

[0016]有益技术效果:需要说明的是,在半导体加工制造的过程中,其可能需要历经上千次的工序,且每一步工序的参数(如工程师所设定的配方,或者设备本身的工作状态等等)也可能存在着差异。因此,一个半导体将记录有海量的产线数据。并且,一旦半导体的良率出现降低,要从海量的产线数据中分析出根因的难度非常高,即使是经验丰富的工程师也可能需要花费数天甚至数周的时间,才能够定位故障根源。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707429B_ABST
    Figure CN121707429B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of semiconductor data analysis, in particular to a root cause analysis method, system and program product, the method comprising: acquiring semiconductor production line data, the semiconductor production line data at least comprising: (1) a plurality of processing equipment information experienced by a plurality of semiconductor products, (2) process parameters adopted when the semiconductor products experience at least one processing equipment and / or category information of the semiconductor products, and (3) measurement values; calling a data analysis scheme for the semiconductor production line data; the data analysis scheme comprising: a fault prediction algorithm; pre-processing the semiconductor production line data according to the data analysis scheme to generate a pre-processing result; inputting the pre-processing result and corresponding semiconductor production line data into a preset big data platform; and the big data platform corresponding outputs an analysis result. The big data traceability analysis method provided by the application can improve data analysis efficiency, reduce the requirement for computing power in the analysis process, and reduce the implementation cost.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Divisional application This application is a divisional application of Chinese invention patent application No. 2025115813205, filed on October 31, 2025, entitled "A Semiconductor Big Data Traceability Analysis Method and System". Technical Field

[0002] This invention relates to the field of semiconductor manufacturing technology, and more specifically to a root cause analysis method, system, and program product. Background Technology

[0003] In the semiconductor manufacturing industry, especially in 12-inch wafer fabrication plants (hereinafter referred to as fabs), root cause analysis (RCA) is an important tool for quality management and process control. When problems such as declining yield, product defects, or customer complaints occur during the production process, it is necessary to identify the root cause of the problem through systematic analysis and take measures to prevent its recurrence.

[0004] Currently, the RCA method commonly used in fab plants mainly relies on engineers' experience and manual data collection. The specific process usually includes: extracting relevant data from multiple independent data systems such as FDC (Fault Detection and Classification), SPC (Statistical Process Control), WAT (Wireless Testing), and Inline (Online Measurement), and then comparing and analyzing them.

[0005] For example, patent application CN202310893076.0 discloses a FDC (Failure Data Conversion) cause analysis method and storage medium based on distributed parallel computing. This method includes: acquiring various information about failed devices; acquiring FDC data and configuration files during the semiconductor wafer manufacturing process; preprocessing the FDC dataset to obtain data on the algorithm's computational structure; performing cause analysis to identify influencing factors; summarizing the cause analysis results and extracting the top N influencing factors from each group; comparing test results with known failure modes to determine possible failure causes; conducting further testing and diagnostic operations based on multiple possible failure causes to verify and determine the root cause of the failure; and taking corresponding measures based on the root cause of the failure.

[0006] However, semiconductor production data suffers from the following problems: 1. Scattered data sources and difficulty in integration: Existing systems have varying data formats, structures, and storage methods, lacking a unified data platform. Efficient interconnection and fusion between different data sources are difficult, leading to low data analysis efficiency and the potential for missing crucial information. 2. Long analysis cycles and slow response times: Traditional root cause analysis heavily relies on manual intervention. Engineers must manually export, clean, process, and analyze data, a time-consuming process that typically takes weeks or even a month to complete a single RCA analysis, failing to meet the need for rapid anomaly localization. Therefore, a more efficient semiconductor root cause tracing method is urgently needed. Summary of the Invention

[0007] The purpose of this invention is to provide a big data source tracing and analysis method that partially solves or alleviates the aforementioned shortcomings of existing technologies, thereby improving the efficiency and accuracy of data analysis. To address the aforementioned technical problems, this invention specifically adopts the following technical solution: A first aspect of the present invention is to provide a semiconductor big data traceability and analysis method, comprising the following steps: S101, acquire semiconductor production line data, the semiconductor production line data including at least: (1) information on multiple processing equipment experienced by multiple semiconductor products, (2) process parameters used when the semiconductor products experience at least one of the processing equipment, and / or category information of the semiconductor products, (3) measured values, and the processing equipment information including at least one of the following: site information, machine information, chamber information, the process parameters including at least one of the following: formula, sensor type, the category information including at least one of the following: raw material source, semiconductor model; S102, call a data analysis scheme for the semiconductor production line data, the data analysis scheme including: fault prediction algorithm; S103, preprocess the semiconductor production line data according to the data analysis scheme to generate preprocessing results, the preprocessing The results include: the predicted root causes of anomalies, and / or the anomaly weights of the root causes; wherein, S103 includes the following steps: S1031, obtaining extraction rules, the extraction rules including: equipment type and configuration type, wherein the equipment type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; S1032, identifying multiple production sequences from the semiconductor production line data according to the extraction rules, wherein the production sequences are correspondingly recorded with the measured values; S1033, using a fault prediction algorithm to predict the root causes of anomalies based on the differences between at least one production sequence and multiple other production sequences; S104, inputting the preprocessing results and the corresponding semiconductor production line data into a preset big data platform, wherein the big data platform outputs the corresponding analysis results.

[0008] In some embodiments, the fault prediction algorithm includes: a one-way ANOVA algorithm, a history analysis algorithm, and / or a similarity analysis algorithm. In some embodiments, the big data platform is the Apache Spark big data platform. In some embodiments, the production sequence records at least the processing sequence of at least two processing devices that the semiconductor product has undergone, as well as the corresponding measurement values; correspondingly, S1033 further includes: S10331, obtaining at least one production sequence group, and the production sequences in the same group have the same equipment type and configuration type; S10332, using the fault prediction algorithm to calculate the difference between a production sequence and other multiple production sequences in the same group, in order to predict the root cause of the anomaly.

[0009] In some embodiments, when the fault prediction algorithm is the history analysis algorithm, step S10332 includes the steps of: identifying the data type of the measured value, wherein the data type includes: discrete and / or continuous; selecting a corresponding mutual information algorithm according to the data type; and calculating the mutual information of the production sequence according to the mutual information algorithm. The mutual information is used to represent the information contained in the random variable Y about the features. The amount of information; calculate the root cause weight of at least one processing device in the production sequence based on the mutual information index.

[0010] In some embodiments, before S10332, the method further includes the steps of: obtaining the number of sequences in a set of product sequences; determining whether the number of sequences is greater than a preset number; if so, proceeding to step S10332.

[0011] In some embodiments, prior to S1031, the method further includes the steps of: selecting at least one set of pre-extraction rules from a grouping rule base; generating multiple pre-production sequences based on at least a portion of the semiconductor generation data using the pre-extraction rules; dividing the multiple pre-production sequences into multiple sub-production sequences, each sub-production sequence having a preset length; obtaining the analysis results of the multiple sub-production sequences and the actual results input by the user; and scoring the pre-extraction rules based on the difference between the analysis results and the actual results.

[0012] In some embodiments, when the data type is discrete, the mutual information is calculated as follows: ;in, Representation of features The value, Represents the measured value of category Y; Features and categories The joint probability distribution of ; Features The marginal probability distribution; Let Y be the marginal probability distribution of category Y; where, feature This refers to a feature in the i-th production sequence, where the feature refers to the device type or configuration type; category This refers to the category of measurements for a production sequence.

[0013] In some embodiments, when the data type is continuous, the mutual information index is calculated as follows: ;in, Representation of features The value, Represents the measured value of category Y; Features and categories The joint probability distribution of ; Features The marginal probability distribution; Let Y be the marginal probability distribution of category Y; where, feature This refers to a feature in the i-th production sequence, where the feature refers to the device type or configuration type; category This refers to the category of measurements for a production sequence.

[0014] In some embodiments, the root cause weights are calculated as follows: ; This represents the characteristics of the j-th sub-production sequence within the i-th production sequence. This represents the percentage of the total number of wafers that have undergone the i-th production sequence. The number of production sequences in a product sequence group.

[0015] Another aspect of the present invention provides a semiconductor big data traceability and analysis system, comprising: The data acquisition module is used to acquire semiconductor production line data, which includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; The scheme invocation module is used to invoke a data analysis scheme for the semiconductor production line data, and the data analysis scheme includes: a fault prediction algorithm; A preprocessing module is used to preprocess the semiconductor production line data according to a data analysis scheme to generate preprocessing results, the preprocessing results including: predicted root causes of anomalies, and / or anomaly weights of the root causes of anomalies; wherein, the preprocessing module includes: a rule acquisition unit, used to acquire extraction rules, the extraction rules including: equipment type and configuration type, wherein the equipment type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; a sequence identification unit, used to identify multiple production sequences from the semiconductor production line data according to the extraction rules; and a fault prediction unit, used to predict the root causes of anomalies using a fault prediction algorithm based on the differences between at least one production sequence and multiple other production sequences; The analysis module is used to input the preprocessing results and the corresponding semiconductor production line data into a preset big data platform, and the big data platform outputs the analysis results accordingly.

[0016] Beneficial Technical Effects: It should be noted that the semiconductor manufacturing process may involve thousands of steps, and the parameters of each step (such as the formula set by engineers or the operating status of the equipment itself) may vary. Therefore, a single semiconductor will record a massive amount of production line data. Furthermore, once the yield of a semiconductor decreases, it is extremely difficult to analyze the root cause from this vast amount of production line data. Even experienced engineers may need to spend days or even weeks to pinpoint the cause of the failure.

[0017] To address this, this invention proposes a big data source tracing analysis method (i.e., root cause analysis method) that combines preprocessing and AI processing. First, preprocessing is used to roughly locate the root cause (e.g., filtering out areas with a high probability of root causes). The preprocessing results (e.g., possible abnormal root causes, or abnormal weights of abnormal root causes) and corresponding detailed information (i.e., production line data) are then input into a pre-defined AI platform. Guided by the preprocessing results, the AI ​​platform performs in-depth analysis of the production line data. This collaborative analysis based on preprocessing results and production line data effectively improves the efficiency of AI analysis.

[0018] Furthermore, it is worth noting that traditional root cause analysis often employs analytical methods based on complete process parameters. This requires combining complete process parameters, such as FDC datasets (which often contain relatively complete equipment parameters such as temperature, pressure, flow rate, voltage, and current), for root cause tracing. In other words, traditional root cause analysis is parameter-based.

[0019] In stark contrast, this invention employs a non-parametric analysis approach. For example, it only needs to extract production sequences from the FDC dataset (which does not require detailed process parameters, such as only recording the process sequence and process category) and predict possible root cause types based on the production sequences. This non-parametric analysis approach facilitates faster preliminary analysis and improves analysis efficiency to some extent.

[0020] Furthermore, by combining the preliminary results obtained from integrated nonparametric analysis methods with the corresponding complete dataset, a detailed analysis of potential root causes can be conducted based on the AI ​​platform, thereby efficiently providing users with relatively reliable and complete analysis results.

[0021] Specifically, for non-parametric analysis methods, this invention actually provides a fragmented data preprocessing method. Specifically, this invention uses a grouping and segmentation method to obtain sub-production sequences with analytical value. The grouping and segmentation method includes the following steps: 1) grouping semiconductor production line data according to equipment type and configuration type and extracting production sequences; 2) cutting the production sequences to obtain multiple segments (i.e., sub-production sequences). Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. The elements or parts in the drawings are not necessarily drawn to scale. Obviously, the drawings described below are some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0023] Figure 1 This is a schematic diagram of the method flow in an exemplary embodiment of the present invention; Figure 2 This is a production line association diagram in an exemplary embodiment; Figure 3 This is a product flow diagram in yet another exemplary embodiment; Figure 4 This is a system module architecture diagram of an exemplary embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0025] In this document, suffixes such as "module," "part," or "unit" used to denote elements are used only for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, "module," "part," or "unit" may be used interchangeably.

[0026] In this document, the terms "upper," "lower," "inner," "outer," "front," "rear," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the present invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0027] In this document, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," "connected," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0028] In this document, "and / or" includes any and all combinations of one or more of the listed related items.

[0029] In this article, "multiple" means two or more, that is, it includes two, three, four, five, etc.

[0030] As used in this specification, the term "about" typically means + / -5% of the value, more typically + / -4% of the value, more typically + / -3% of the value, more typically + / -2% of the value, even more typically + / -1% of the value, and even more typically + / -0.5% of the value.

[0031] In this specification, certain embodiments may be disclosed in a range-bound format. It should be understood that this "range-bound" description is merely for convenience and brevity and should not be construed as a rigid limitation on the disclosed range. Therefore, the description of a range should be considered as having specifically disclosed all possible subranges and the individual numerical values ​​within those ranges. For example, a description of the range 1-6 should be considered as having specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and the individual numbers within those ranges, such as 1, 2, 3, 4, 5, and 6. This rule applies regardless of the breadth of the range.

[0032] In this article, "wafer" refers to the silicon wafer used in the fabrication of silicon semiconductor integrated circuits. Because its shape is usually circular, it is called a wafer. Various circuit element structures can be fabricated on silicon wafers to become IC products with specific electrical functions.

[0033] In this article, "lot" refers to a basic batch of wafers. For example, 1 lot typically consists of 12 wafers.

[0034] Example 1 See Figure 1 As shown, this invention provides a semiconductor big data traceability and analysis method, including the following steps: S101, acquire semiconductor production line data, wherein the semiconductor production line data includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; For example, in some embodiments, semiconductor production line data may come from an FDC system or a WAT ​​system. This semiconductor production line data can record the processing information a semiconductor product undergoes throughout its entire journey from entering the production line to leaving it, such as the model, type, and serial number of the processing equipment (e.g., stations, machines, chambers), and the process parameters (equivalent to process parameters) of each processing equipment, such as temperature, pressure, voltage, and current.

[0035] For example, in some embodiments, the measured value can be the result of measuring the yield indicators such as the structure, size or performance of the semiconductor product after it has passed through any processing equipment.

[0036] Alternatively, in other embodiments, the measured values ​​may be the process parameters of the processing equipment, such as voltage and current, measured when the semiconductor product is processed by any processing equipment. These parameters can reflect, to some extent, whether the equipment is in a healthy operating state or the probability of defects or malfunctions in the semiconductor product.

[0037] S102, invoke a data analysis scheme for the semiconductor production line data, the data analysis scheme including: a fault prediction algorithm; S103, the semiconductor production line data is preprocessed according to the data analysis scheme to generate preprocessing results, the preprocessing results including: the predicted root causes of anomalies, and / or the anomaly weights of the root causes of anomalies; S104, the preprocessing results and the corresponding semiconductor production line data are input into a preset big data platform, and the big data platform outputs the corresponding analysis results.

[0038] In some embodiments, the fault prediction algorithm includes: a one-way ANOVA algorithm, a history analysis algorithm, and / or a similarity analysis algorithm.

[0039] In some embodiments, the big data platform is configured with one or more pre-trained fault prediction AI models (and is therefore also referred to as an AI platform), which can analyze semiconductor production line data to identify anomalies and thereby predict the root cause of the fault.

[0040] In some embodiments, the big data platform is the Apache Spark big data platform.

[0041] In this embodiment, the preprocessing results and the corresponding semiconductor production line data are synchronously input into the big data platform, enabling the big data platform to perform localized focused analysis on massive amounts of semiconductor production line data based on the preprocessing results. This improves the efficiency of fault prediction and reduces the computational power requirements for analysis of massive amounts of semiconductor production line data.

[0042] Specifically, this invention proposes a big data source tracing analysis method (i.e., root cause analysis method) that combines preprocessing and AI processing. First, preprocessing is used to roughly locate the root cause (e.g., filtering out areas with a high probability of root causes). The preprocessing results (e.g., possible abnormal root causes, or abnormal weights of abnormal root causes) and corresponding detailed information (i.e., production line data) are input into a pre-defined AI platform. Guided by the preprocessing results, the AI ​​platform performs in-depth analysis of the production line data. This collaborative analysis based on preprocessing results and production line data effectively improves the efficiency of AI analysis.

[0043] Furthermore, it is worth noting that traditional root cause analysis often employs analytical methods based on complete process parameters. This requires combining complete process parameters, such as FDC datasets (which often contain relatively complete equipment parameters such as temperature, pressure, flow rate, voltage, and current), for root cause tracing. In other words, traditional root cause analysis is parameter-based.

[0044] In stark contrast, this invention employs a non-parametric analysis approach. For example, it only needs to extract production sequences from the FDC dataset (which does not require detailed process parameters, such as only recording the process sequence and process category) and predict possible root cause types based on the production sequences. This non-parametric analysis approach facilitates faster preliminary analysis and improves analysis efficiency to some extent.

[0045] Furthermore, by combining the preliminary results obtained from integrated nonparametric analysis methods with the corresponding complete dataset, a detailed analysis of potential root causes can be conducted based on the AI ​​platform, thereby efficiently providing users with relatively reliable and complete analysis results.

[0046] The nonparametric analysis method proposed in this invention will be further described in detail below: In some embodiments, S103 includes the step of: S1031, Obtain extraction rules, the extraction rules include: equipment type and / or configuration type, wherein the equipment type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; Preferably, the extraction rules include: device type and configuration type; S1032, Multiple production sequences are identified from the semiconductor production line data according to the extraction rules, and the production sequence is correspondingly recorded with the measured value; S1033, a fault prediction algorithm is used to predict the root cause of the anomaly based on the difference between at least one production sequence and multiple other production sequences.

[0047] In this embodiment, the extraction rule can also be regarded as a grouping rule. Data that meets the same extraction rule is recorded into a production sequence, and multiple production sequences that meet the same rule are regarded as a production sequence group.

[0048] See Figure 2 As shown, when the extraction rules include: site, machine, and chamber, the history path of wafer1 products that have experienced the same site, machine, and chamber can be extracted as a production sequence, and the history path of wafer2 products can be extracted as another production sequence. In other words, information such as site, machine, and chamber also serves as a classification factor.

[0049] Understandably, the more rules (or classification factors) there are in the extraction rules, the finer the granularity of the grouping. Specific extraction rules can be adaptively selected based on actual analytical needs.

[0050] In some embodiments, the production sequence records the processing order of at least two processing devices through which the semiconductor product undergoes processing, and the corresponding measurement values; correspondingly, S1033 further includes: S10331, Obtain at least one production sequence group, and the production sequences in the same group have the same equipment type and configuration type; S10332, a fault prediction algorithm is used to calculate the difference between a production sequence and multiple other production sequences in the same group in order to predict the root cause of the anomaly.

[0051] For example, in some embodiments, a set of production sequences includes abnormal sequences and normal sequences. Abnormal sequences refer to production sequences where the yield or measurement value of the corresponding semiconductor product is unacceptable, while normal sequences refer to production sequences where the yield or measurement value of the corresponding semiconductor product is normal. By calculating the differences in production conditions between abnormal and normal sequences, the root causes of potential problems can be preliminarily predicted.

[0052] For example, in some embodiments, a production sequence includes semiconductor products processed by n identical types of processing equipment (i.e., at least two processing equipment are configured under at least one equipment type). The same process may include multiple identical stations, such as station 1, station 2, etc. (i.e., station 1 and station 2 can be used to perform the same process). As another example, a station may include multiple identical parallel chambers, such as chamber 1 and chamber 2, in which case the chamber numbers of the different chambers need to be recorded. The parametric analysis method provided by this invention can directly perform a holistic analysis of root cause regions (such as machines that may cause failures) from the perspective of equipment processing sequence, equipment type, and differences, without getting bogged down in specific parameters, such as voltage and current anomaly identification. This non-parametric root cause region prediction and step-by-step analysis method in collaboration with an AI platform facilitates faster and more accurate fault analysis.

[0053] In some embodiments, when the fault prediction algorithm is the history analysis algorithm, S10332 includes the following steps: (1) Identify the data type of the measured value, wherein the data type includes: discrete and / or continuous; (2) Select the corresponding mutual information algorithm according to the data type; (3) Calculate the mutual information of the production sequence according to the mutual information algorithm. The mutual information is used to represent the information contained in the random variable Y about... The amount of information; in, This refers to the characteristics of the production sequence, such as... It can represent the device type or configuration type of the i-th production sequence.

[0054] (4) Calculate the root cause weight of at least one processing equipment in the production sequence based on mutual information.

[0055] For example, in some embodiments, the root cause weight of a production sequence can be calculated based on mutual information, or the root cause weight of a specific processing equipment in a production sequence can be calculated based on mutual information.

[0056] In some embodiments, prior to S10332, the following step is further included: Get the number of sequences within a set of product sequences; Determine whether the number of sequences is greater than a preset number. If so, proceed to step S10332.

[0057] For example, in some embodiments, the preset quantity is 1.

[0058] In some embodiments, prior to S1031, the following step is further included: Select at least one set of pre-extracted rules from the grouping rule base; Multiple pre-production sequences are generated based on at least a portion of the semiconductor generation data using the pre-extraction rules. Multiple pre-production sequences are divided into multiple sub-production sequences, and each sub-production sequence has a preset length. Obtain the analysis results of multiple sub-production sequences and the actual results input by the user; The pre-extraction rules are scored based on the difference between the analysis results and the actual results.

[0059] For example, in some embodiments, a preset length can also be selected / updated, wherein the preset length selection / updating step includes: (1) The pre-production sequence is cut using a preset length and a test length respectively to obtain a first cutting group and a second cutting group; wherein the cutting group includes multiple sub-production sequences after cutting; for example, in some embodiments, the preset length can be a default length preset by the user, or the preset length can be the cutting length used in the last cutting of the pre-production sequence. The test length can be a new cutting length obtained by increasing or decreasing the preset length, such as it can be manually input by the user, or it can be randomly generated according to the preset length.

[0060] It should be noted that, in this embodiment, the production sequence is cut, which is actually equivalent to dividing the production process into multiple smaller stages.

[0061] (2) Obtain the first prediction accuracy of the first segmentation group and the second prediction accuracy of the second segmentation group; wherein, the prediction accuracy refers to the degree of difference between the analysis results based on the fault prediction algorithm and the actual results input by the user; (3) Calculate the prediction difference between the first prediction accuracy and the second prediction accuracy; (4) When the prediction difference is greater than the preset prediction threshold (which can be set or adjusted by the user), proceed to step (5). (5) Obtain the first prediction duration corresponding to the first segmentation group and the second prediction duration of the second segmentation group; (6) Calculate the duration difference between the first prediction duration and the second prediction duration; (7) When the duration difference is less than the preset duration threshold, it is recommended to update the test length to the new preset length.

[0062] It should be noted that, in order to improve the overall efficiency of fault analysis, this invention introduces a segmentation and grouping analysis mode for the early non-parametric analysis process. That is, by dividing a large production sequence into multiple small sequences (i.e., sub-production sequences), the root cause region can be quickly located through the parallel operation of multiple small sequences.

[0063] Meanwhile, this embodiment also provides a restrictive cutting scale adjustment scheme for non-parametric analysis. This size adjustment updates / adjusts the cutting length in a restrictive manner based on prediction time and prediction accuracy. This allows for rapid prediction of areas suspected of having faults within a relatively limited time, while ensuring that the prediction results have a certain degree of reliability.

[0064] For example, for the same type of wafer manufacturing task or the same type of wafer production line, a pre-testing method can be used to screen for suitable dicing sizes. Similarly, before the formal application of this step-by-step analysis method, a suitable dicing size can be screened through an initial trial phase.

[0065] In some embodiments, when the data type is discrete, the mutual information index is calculated as follows: ; in, Representation of features Values ​​(such as site category, machine category, etc.). This represents the measurement value for category Y; Features and categories The joint probability distribution, which is used to describe or represent the characteristics of random variables. and categories The probability of simultaneous occurrence. Features The marginal probability distribution is used to describe or represent the prior probability of a particular feature in all possible states after integrating (or ignoring) the effects of all other features and category information. The marginal probability distribution of category Y, which is used to describe or represent the specific values ​​after integrating (or ignoring) all other features; the prior distribution or basic ratio of each category in the population; Among them, features This refers to the features in the i-th production sequence; This represents the measurement value for category Y.

[0066] In some embodiments, when the data type is continuous, the mutual information is calculated as follows: ; in, Representation of features The value, Indicates the measured value; Features and categories The joint probability distribution of ; Features The marginal probability distribution; The marginal probability distribution of category Y; Among them, features This refers to the features in the i-th production sequence; This represents the measurement value for category Y.

[0067] In some embodiments, the root cause weights are calculated as follows: ; This represents the characteristics of the j-th sub-production sequence (or the j-th stage) within the i-th production sequence. This refers to the mutual information of the j-th sub-production sequence within the i-th production sequence. As the penalty factor, where, This represents the wafer proportion (i.e., the percentage of wafers that have gone through the i-th production sequence in the total number of wafers). The number of production sequences in a product sequence group.

[0068] It should be noted that the above probability-based fault prediction algorithm is only an exemplary analysis method provided by the present invention. Depending on the specific analysis needs, the present invention can also use other methods to perform non-parametric fault analysis.

[0069] In some embodiments, the probability density is estimated using distance information from the K nearest neighbors, and an exemplary calculation step is as follows: Given sample points Calculate the distance to the k-th nearest neighbor of the point; Count how many other points are within this distance range (to estimate local density); Using this local density information, the mutual information contribution of each sample point is estimated; The final mutual information estimate is obtained by averaging all sample points.

[0070] For example, suppose the collected data has m site groups, each group has s stages (e.g., a sub-production sequence can be divided based on a stage), and each stage has The combination of chambers at each machine station site, the ratio of the number of wafers in the current group to the total number of wafers at the station (wafer_ratio), is determined by calling the mutual information algorithm I. When the value is 1, meaning all wafers have passed through a single chamber, they cannot be separated. .

[0071] First exemplary embodiment: Data is collected from the production line based on user input, and a column named "value" is added: the value information corresponding to the wafer user.

[0072] Based on the grouping field input by the user, groups are formed, and each group is analyzed separately. For example, in the case of resume chamber analysis, the result of one group is shown below.

[0073] or A grouped resume analysis, sorting each job application chronologically from beginning to end, divided into different stages, regarding... Figure 2 The grouping results can be shown in the table below: Furthermore, based on the above grouping results, a difference analysis is performed: If, in any given stage, for continuous values ​​(i.e., values), those with high values ​​appear in one chamber and those with low values ​​in another, and for discrete values, such as 0 values ​​appearing in one chamber and 1 values ​​in another, these values ​​can be completely separated, the greater the likelihood that this situation will be identified as the root cause. This embodiment introduces a mutual information analysis method to measure the difference in the history analysis.

[0074] Second exemplary embodiment: The following section uses another example of production line data to illustrate the implementation of the history analysis algorithm: Resume Information Form The path diagram corresponding to the above resume information table is as follows: Figure 3 As shown. STAGE1 contains machine tool M1, and STAGE2 contains all the chambers of machine tool M2. The path from STAGE1 to STAGE2 is shown as follows. Figure 3 The connection in the diagram, VALUE (i.e., the measured value) is discrete here, 0 is GOOD marked in green, and 1 is BAD marked in red. The flow diagram can intuitively show the path flow of GOOD and BAD wafers.

[0075] Step 1: Before calculating the mutual information between STAGE1 and VALUE, because the STAGE1 column is text type, it needs to be encoded using natural numbers (a simple encoding that maps categorical data to a sequence of natural numbers (0, 1, 2, 3,...)). The encoded STAGE1 column is as follows: Correspondingly, = STAGE1_encoding; Similarly, calculate =STAGE2_encoding.

[0076] The encoded STAGE2 column is as follows: Step 2: Calculate mutual information, where y is the VALUE column information, which is of numeric type.

[0077] = ; .

[0078] Step 3: Calculate the weight of this site. Assuming the total number of wafers is 16, the wafer ratio of this group is 0.5. Based on the history table, the chamber combination for each stage can be calculated. .

[0079] Based on the weight calculation formula, the weight of this site is as follows: Finally, by summing the weight parameters of all sites in all groups and sorting them from largest to smallest, we obtain the root cause weight ranking from largest to smallest. The output results are shown below: It's worth noting that a 12-inch fab involves terabytes (TB) of data per month. Leveraging the Apache Spark big data analytics engine, it supports various computing models, including batch processing, stream processing, graph computing, and machine learning. Spark's core advantage lies in in-memory computing, offering faster execution speeds than Hadoop, while providing high fault tolerance and an easy-to-use API (supporting Scala, Java, Python, and R).

[0080] The main modules of the Spark big data platform include SparkCore, SparkSQL, Spark Streaming, MLlib (machine learning), and GraphX ​​(graph computation). In this application, three source tracing analysis algorithms were designed and run on Spark, adapted for use with data sources such as FDC, WAT, Inline, and wafer record information, enabling large-scale data analysis and achieving result generation speeds ranging from minutes to hours.

[0081] In other words, one of the objectives of this invention is to provide a root cause analysis method based on manufacturing data from a 12-inch semiconductor fab, in order to solve the problems of scattered data sources, low analysis efficiency, reliance on manual judgment, and limited analysis granularity in the source tracing analysis process in the prior art.

[0082] To achieve the above objectives, this invention constructs three key algorithm models based on the Apache Spark big data platform: a data grouping analysis algorithm, a history analysis algorithm, and a similarity analysis algorithm. These algorithms are applicable to different types of input data (such as discrete labels, continuous output values, and anomaly representation parameters), thereby enabling unified analysis and efficient tracing of multiple data sources, including FDC, WAT, Inline, and wafer history.

[0083] This invention can automate the root cause identification of large-scale manufacturing data within a time range of minutes to hours, improve the speed and accuracy of root cause localization, reduce the workload of engineers, and provide strong support for process optimization and quality control.

[0084] Example 2 The following section uses a data grouping analysis algorithm as an example to illustrate an exemplary analysis method provided by this invention: Users assign GOOD and BAD tags to a batch of WAFERs and locate the root cause within a selected time range. The granularity of root cause location can be adjusted, such as (product + site + machine + sensor_name) or (product + site + machine + chamber + sensor_name). The specific analysis process includes the following steps: Step 1: Select the data source to be analyzed, such as FDC, WAT, Inline, upload wafer information, mark "GOOD" or "BAD", select the time range for traceability, and if further refinement is needed, filter by product, site, machine, chamber, or lot information.

[0085] Step 2: Select the granularity of the group analysis to determine the level of detail in the localization. For example, for the FDC data source, the granularity of tracing can be selected in three ways: 1. Site + Machine + Chamber; 2. Site; 3. Site + Machine. Parameter is the smallest parameter for tracing. For example, FDCParameter is SensorName#Window#stats (the statistical characteristics calculated for a group of consecutive data points within the specified window, such as 'max': maximum value, 'min': minimum value, 'min': average value). Click Run to submit the analysis task.

[0086] Step 3: The platform submits the task information to the platform's Spark big data analysis engine, calls the data grouping analysis algorithm, and outputs the results of the algorithm to the database after the operation is completed.

[0087] Step 4: The platform displays the results of the task in the platform results view.

[0088] The core algorithm steps are as follows: Step 1: Collect the dataset: Select data from the user-specified data source using the user-provided waferlist, good, and bad label information. Assign the 'label' value to the wafer column, with 0 representing 'GOOD' and 1 representing 'BAD'.

[0089] Step 2: Select the data for the product, machine, chamber, site, parameter, wafer, and 'label' column, and remove rows with missing values.

[0090] Step 3: Group based on product and grouping information, and score all parameters of each group, marking them as "weight". For example, Group 1 has the grouping fields: Site, Machine, Chamber.

[0091] Step 4: If there are significant differences in parameters across label groups, the parameter is more likely to be the root cause. One possible root cause is a malfunction in the device component corresponding to the sensor at a certain moment, leading to excessively high or low data collection. This manifests as a significant separation. In post-analysis, one-way ANOVA is used to determine the differences in parameters between the "good" and "bad" wafer groups.

[0092] One-way ANOVA, also known as one-dimensional ANOVA, is used to analyze whether there are significant differences in the mean of the dependent variable when a single control factor is taken at different levels. The different levels correspond to the GOOD and BAD groups, with a single control factor.

[0093] Specifically, the core calculation process of step 4 is as follows: First, let's define the assumptions: H0 (Null Hypothesis): The means of all groups are equal (μ1=μ2=...=μ...). k ), k is the number of groups.

[0094] I1 (Alternative Hypothesis): There is at least one set of means that differ from the others.

[0095] Calculate the total sum of squares (SST): the sum of squared deviations of all observations from the population mean. : The j-th observation in the i-th group.

[0096] ...: The overall average of all data.

[0097] Between-group sum of squares (SSB): The between-group sum of squares measures the difference between the group means and the population mean. : The number of samples in the i-th group.

[0098] : The average value of the i-th group.

[0099] Within-group sum of squares (SSW): Measures the difference between each data point within a group and the group mean. Calculate the degrees of freedom: Total degrees of freedom: (N is the total number of samples.)

[0100] Between-group degrees of freedom: (k is the number of groups).

[0101] Intragroup freedom: .

[0102] Calculate the mean square error: Between-group mean square (MSB): ; Within-group mean square (MSW): ; Calculate the F-statistic to compare variances between and within groups: Determine the significance level, set it to 0.05, and calculate the p-value based on the F-distribution.

[0103] The cumulative distribution function of the F-distribution represents the probability that, given the between-group degrees of freedom and the within-group degrees of freedom, the calculated F-statistic is less than or equal to the given F-statistic. ; The p-value represents the probability of the opposite event occurring.

[0104] ; Determine significance: when The difference is significant, so we reject the null hypothesis and choose the alternative hypothesis.

[0105] Otherwise, if the difference is not significant, the null hypothesis is accepted.

[0106] When p is less than 0.05, the smaller the value, the greater the difference; and the larger the F-statistic, the better.

[0107] Specifically, in step 3, if the number of rows for "GOOD" or "BAD" in the group + Parameter is less than 3, the data sample size is too small and the analysis is unreliable, so the root cause weight is set to 0.

[0108] Otherwise, the analysis of variance algorithm is invoked. Given a total number of groups of m and a total number of parameters in each group of n, the F-statistic for all parameters in that group is calculated. The F-statistic for the j-th parameter in the i-th group has the following weight formula: This represents the proportion of Badwafer fragments within this group to the total number of Badwafer fragments. It indicates that the more Badwafer fragments there are in a group, the greater the likelihood that it is the root cause.

[0109] By summing the weights of all parameters for all groups and sorting them from largest to smallest, we obtain the root cause weight ranking from largest to smallest. The output is shown below: Example 3 The following example, using a similarity analysis algorithm, further illustrates the analysis method provided by this invention: The analysis method in this embodiment is applicable to situations where a set of outlier parameters is known. It searches other data sources, such as FDC, WAT, and Inline, for a set of parameters related to the wafer and checks for linear, square, or radical relationships. If any one of these relationships exists, it can be used to set the root cause weights. The algorithm designed here uses linear regression and then evaluates the R² of the goodness of fit as the root cause weight.

[0110] The steps for similarity analysis in business analysis are as follows: Step 1: Select the data source for similarity analysis and comparison, such as FDC, WAT, or Inline. Select a batch of wafers with anomaly characteristics, such as FDC, WAT, or Inline. Select the time range for tracing. If further refinement is needed, filtering can be done by product, site, machine, chamber, or lot information.

[0111] Step 2: Select Groups: The granularity of the analysis determines the level of detail in the localization. For example, for the FDC data source, the granularity of tracing can be selected in three ways: 1. Site + Machine + Chamber; 2. Site; 3. Site + Machine. Parameter is the smallest parameter for tracing. For example, FDCParameter is SensorName#Window#stats (the statistical characteristics calculated for a group of consecutive data points within the specified window, such as 'max': maximum value, 'min': minimum value, 'min': average value). Click Run to submit the analysis task.

[0112] Step 3: The platform submits the task information to the platform's Spark big data analysis engine, calls the data similarity analysis algorithm, and outputs the results of the algorithm to the database after the operation is completed.

[0113] Step 4: The platform displays the results of the task in the platform results view.

[0114] The specific similarity analysis algorithm may include the following steps: Step 1: Data Collection: Collect the necessary columns from the data source to be compared for similarity, such as a parameter of an anomaly characterization column that is inlined. The data source being compared is FDC. Example results are shown below: Step 2: Based on the defined data source for comparison, such as FDC, and the grouping fields are site + machine + chamber, compare the Parameter with the anomaly representation value. Group the data from Step 2 according to site + machine + chamber, and check if there is a linear relationship, a square relationship, or a square root relationship among them; satisfying any one of these three conditions is sufficient. Train a linear regression on the Parameter and anomaly representation value columns, and evaluate the goodness of fit between the true and predicted values ​​as the similarity score.

[0115] The principle of step 2 is as follows: Linearity detection: Using linear regression, known... The last column is the intercept column, all of which are 1. The sample size is N, the number of features is d, and the slope vector is... (Including intercept), target vector , For the intercept term, construct the model: ; in, For a linear regression model, X represents the FDC / WAT / Inline parameter to be detected (e.g., the parameter column of FDC in the grouping), and R represents a real number.

[0116] The loss function is: ; Minimize the loss to obtain the optimal solution: ; Here, Y represents the anomaly representation values ​​of a selected set of wafers, see the value column in the table.

[0117] Square relation test: Consider the simplest case, d=1, where there is only one-dimensional feature x=>{ }; Then, by substituting the solutions into linear regression, the optimal solution can be obtained.

[0118] Square root relation check: Equivalent to and There exists a linear case where, for any x, if x is less than 0, we transform x by subtracting the minimum value from each element of x, so that x as a whole is greater than or equal to 0. Then, we substitute this into a linear regression model.

[0119] Methods for determining the weights in root cause analysis: Given a one-dimensional x and a one-dimensional y as input; Detect the linear relationship and construct a linear regression model of y and x; Detecting square relationships: Constructing y and x, The regression model; Detecting square root relationships: Constructing The regression model with x.

[0120] The goodness of fit between the predicted and actual values ​​is calculated sequentially; it represents the model's ability to explain changes in the target variable. ; in, The predicted value of the model. The mean of the target variable. Coefficient of determination. The closer to 1, the better; coefficient of determination A value of 0 is equivalent to mean prediction; the coefficient of determination is... A value less than 0 indicates that the model performs worse than "predicting using the mean".

[0121] For the existence of a determination coefficient less than 0 Change it to 0, then find the largest of the three. , as a similarity score.

[0122] Assume the data has a total of m groups, and each group has The algorithm has one parameter and is based on linear, square, or radical relationships; if any one of these three conditions is met, it is considered an abnormal algorithm. If the sample size is less than or equal to 3, the sample size is too small and the data is unreliable. 0; Otherwise, use the similarity scoring algorithm to calculate: ; For training linear regression models, square regression models, and radical regression models, This is the goodness-of-fit formula.

[0123] Step 3: Summarize the weight parameters of all sites in all groups, sort them from largest to smallest, and you will get the root cause weight ranking from largest to smallest. For example, the output result is as follows: It should be noted that the data generated during the manufacturing process in a fab is massive and diverse; on the other hand, traditional analysis methods and technical architectures are difficult to effectively cope with the challenges of such high-dimensional, heterogeneous, and real-time data.

[0124] This invention addresses the challenges of high-dimensional, heterogeneous, and real-time data by providing an improved solution that overcomes many shortcomings of existing technologies and offers a method for large-scale data tracing and analysis suitable for 12-inch fabs.

[0125] Alternatively, this invention provides a root cause analysis method based on manufacturing data from a 12-inch semiconductor fab, to solve the problems of scattered data sources, low analysis efficiency, reliance on manual judgment, and limited analysis granularity in the source tracing analysis process of existing technologies.

[0126] To achieve the above objectives, this invention constructs three key algorithm models based on the Apache Spark big data platform: a data grouping analysis algorithm, a history analysis algorithm, and a similarity analysis algorithm. These algorithms are applicable to different types of input data (such as discrete labels, continuous output values, and anomaly representation parameters), thereby enabling unified analysis and efficient tracing of multiple data sources, including FDC, WAT, Inline, and wafer history.

[0127] This invention can automate the root cause identification of large-scale manufacturing data within a time range of minutes to hours, improve the speed and accuracy of root cause localization, reduce the workload of engineers, and provide strong support for process optimization and quality control.

[0128] Example 4 See Figure 4 As shown, the present invention also provides a semiconductor big data traceability and analysis system, comprising: The data acquisition module 101 is used to acquire semiconductor production line data, which includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; The scheme invocation module 102 is used to invoke a data analysis scheme for the semiconductor production line data, the data analysis scheme including: a fault prediction algorithm; Preprocessing module 103 is used to preprocess the semiconductor production line data according to the data analysis scheme to generate preprocessing results, the preprocessing results including: predicted root causes of anomalies, and / or anomaly weights of the root causes of anomalies; wherein, the preprocessing module includes: The rule acquisition unit 1031 is used to acquire extraction rules, the extraction rules including: device type and configuration type, wherein the device type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; The sequence recognition unit 1032 is used to identify multiple production sequences from the semiconductor production line data according to the extraction rules; The fault prediction unit 1033 is used to predict the root cause of the anomaly based on the difference between at least one production sequence and multiple other production sequences using a fault prediction algorithm. The analysis module 104 is used to input the preprocessing results and the corresponding semiconductor production line data into a preset big data platform, and the big data platform outputs the analysis results accordingly.

[0129] It should be noted that the system in this embodiment can implement the methods or steps in any of the above embodiments, which will not be repeated here.

[0130] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a computer terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0132] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A root cause analysis method, characterized in that, Including the following steps: S101, acquire semiconductor production line data, wherein the semiconductor production line data includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; S102, invoke a data analysis scheme for the semiconductor production line data, the data analysis scheme including: a fault prediction algorithm; S103, the semiconductor production line data is preprocessed according to the data analysis scheme to generate preprocessing results, the preprocessing results including: predicted root causes of anomalies, and / or the anomaly weights of the root causes of anomalies; wherein, the anomaly weights refer to the anomaly weights of the production sequence; S103 includes the following steps: S1031, Obtain extraction rules, the extraction rules include: device type and configuration type, wherein the device type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; S1032, Multiple production sequences are identified from the semiconductor production line data according to the extraction rules, and the production sequence is correspondingly recorded with the measured value; S1033, a fault prediction algorithm is used to predict the root cause of the anomaly based on the difference between at least one production sequence and multiple other production sequences; the production sequence records the processing order of at least two processing devices through which the semiconductor product has undergone, and the corresponding measurement values; correspondingly, S1033 further includes: S10331, Obtain at least one production sequence group, and the production sequences in the same group have the same equipment type and configuration type; S10332, A fault prediction algorithm is used to calculate the difference between a production sequence and multiple other production sequences in the same group to predict the root cause of the anomaly; S104, The preprocessing results and the corresponding semiconductor production line data are input into a preset big data platform, and the big data platform outputs the corresponding analysis results; The big data platform is equipped with one or more pre-trained fault prediction AI models.

2. The method according to claim 1, characterized in that, The fault prediction algorithm includes: one-way ANOVA algorithm, history analysis algorithm and / or similarity analysis algorithm.

3. The method according to claim 1, characterized in that, The production sequence includes abnormal sequences and normal sequences.

4. The method according to claim 3, characterized in that, The abnormal sequence refers to the production sequence in which the yield of the corresponding semiconductor product is unqualified or the measurement value is unqualified, while the normal sequence refers to the production sequence in which the yield of the corresponding semiconductor product is normal or the measurement value is qualified.

5. The method according to claim 1, characterized in that, One of the production sequences corresponds to a semiconductor product processed by n identical types of processing equipment.

6. A root cause analysis system, characterized in that, include: The data acquisition module is used to acquire semiconductor production line data, which includes at least: (1) information on multiple processing equipment that multiple semiconductor products have undergone, (2) process parameters used when the semiconductor products undergo at least one of the processing equipment, and / or category information of the semiconductor products, (3) measurement values, and the processing equipment information includes at least one of the following: site information, machine information, chamber information, the process parameters include at least one of the following: formula, sensor type, and the category information includes at least one of the following: raw material source, semiconductor model; The scheme invocation module is used to invoke a data analysis scheme for the semiconductor production line data, and the data analysis scheme includes: a fault prediction algorithm; A preprocessing module is used to preprocess the semiconductor production line data according to a data analysis scheme to generate preprocessing results. The preprocessing results include: predicted root causes of anomalies, and / or anomaly weights of the root causes; wherein the anomaly weights refer to the anomaly weights of the production sequence; wherein the preprocessing module includes: The rule acquisition unit is used to acquire extraction rules, the extraction rules including: device type and configuration type, wherein the device type includes at least one piece of processing equipment information, and the configuration type includes at least one piece of process parameter information and / or at least one piece of category information; A sequence recognition unit is used to identify multiple production sequences from the semiconductor production line data according to extraction rules; A fault prediction unit is configured to use a fault prediction algorithm to predict the root cause of an anomaly based on the differences between at least one production sequence and multiple other production sequences; the production sequence records the processing order of at least two processing devices through which the semiconductor product undergoes processing, and the corresponding measurement values; correspondingly, the fault prediction unit is also configured to perform: At least one production sequence group is obtained, and the production sequences in the same group have the same device type and configuration type; A fault prediction algorithm is used to calculate the difference between a production sequence and several other production sequences in the same group in order to predict the root cause of the anomaly. An analysis module is used to input the preprocessing results and the corresponding semiconductor production line data into a preset big data platform, and the big data platform outputs the corresponding analysis results. The big data platform is equipped with one or more pre-trained fault prediction AI models.

7. The system according to claim 6, characterized in that, The fault prediction algorithm includes: one-way ANOVA algorithm, history analysis algorithm and / or similarity analysis algorithm.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements a root cause analysis method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • A FDC (Fault-Derived Conversion) Cause Analysis Method and Storage Medium Based on Distributed Parallel Computing

    CN116629707B

  • Defective root cause analysis method and system based on comprehensive analysis framework

    CN119558541A