DDR wafer test fault detection method and system based on association rules
Patent Information
- Application Number
- CN202610379669.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-26
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-03-26
AI Technical Summary
然而,这种仅依据单个故障点或其简单分布密度的故障检测方式较为简单,容易导致对晶圆上根本缺陷的定位不够精准,影响后续复测效率和晶圆良率分析的准确性
[0006]本发明通过获取原始电性测试数据并进行测试项目间的关联关系挖掘,得到具有空间位置关联的故障码序列,使后续处理能够基于测试数据内部的深层关联规律展开;通过对故障码序列进行频繁项集搜索与强关联规则生成,输出与空间分布关联的初始故障关联规则集合,使挖掘出的规则同时反映故障码的统计相关性与晶圆空间分布特征;通过根据初始故障关联规则集合进行动态故障推理,将实时故障码与规则集合进行匹配与置信度累积计算,生成包含故障传播路径及故障根源候选节点的晶圆级故障演化图谱,使故障发展脉络得以动态追踪和量化表征;最终基于图谱中的故障根源候选节点及其置信度累积值进行优先级排序,生成包含待优先复测晶粒坐标及关联测试项目标识的故障根因定位指令并发送至晶圆测试机台以触发复测操作,使复测资源能够聚焦于最可能存在根本缺陷的晶粒及其关联项目,从而提升了DDR晶圆测试故障检测的精准定位能力与检测效率。
Smart Images

Figure CN122286697B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of semiconductor testing and data processing, and more specifically, to a method and system for detecting faults in DDR wafer testing based on association rules. Background Technology
[0002] In the semiconductor manufacturing field, DDR wafer testing fault detection technology is used to identify and analyze abnormal electrical parameters of memory cells during the wafer testing stage, thereby locating defective dies and their associated test items on the wafer. Currently, common DDR wafer testing fault detection methods typically acquire raw electrical test data generated by the wafer testing equipment, extract test items exceeding preset specification limits as isolated fault points, and then determine the fault area on the wafer based on the distribution density of these fault points or the recurrence frequency of individual fault items. However, this fault detection method, which relies solely on a single fault point or its simple distribution density, is relatively simplistic and can easily lead to inaccurate localization of fundamental defects on the wafer, affecting subsequent retesting efficiency and the accuracy of wafer yield analysis. Summary of the Invention
[0003] This invention provides a method and system for detecting faults in DDR wafer testing based on association rules.
[0004] In a first aspect, embodiments of the present invention provide a DDR wafer test fault detection method based on association rules, including: Obtain the original electrical test data set of the DDR wafer test items generated after the electrical parameter measurement operation is performed on each memory cell of the wafer by the probe station during the wafer testing phase; The original electrical test data set is processed to mine the correlation between test items, and a fault code sequence with spatial location correlation within the wafer is obtained, which includes the co-occurrence relationship of fault codes between different test items and the order of occurrence of fault codes. Frequent itemset iterative search and strong association rule generation based on minimum support threshold and minimum confidence threshold are performed on the fault code sequence to obtain an initial fault association rule set associated with the spatial distribution; Dynamic fault reasoning is performed based on the initial fault association rule set. The real-time fault codes in the real-time acquired DDR wafer online test data stream are matched with the initial fault association rule set and confidence accumulation is calculated to generate a wafer-level fault evolution map containing fault propagation paths and candidate nodes of fault root causes. Based on the candidate nodes of the fault root cause in the wafer-level fault evolution map and their associated cumulative confidence values, priority sorting is performed to generate a fault root cause location instruction containing the coordinates of the die to be retested first and the identifier of the associated test item. The fault root cause location instruction is then sent to the wafer test instrument to trigger the retest operation for the target die.
[0005] Secondly, embodiments of the present invention provide a computer device, including: a memory storing a computer program; and a processor for loading the computer program to implement the DDR wafer test fault detection method based on association rules as described above.
[0006] This invention acquires raw electrical test data and mines the correlations between test items to obtain fault code sequences with spatial location associations, enabling subsequent processing to be based on the deep correlation patterns within the test data. By performing frequent itemset search and generating strong association rules on the fault code sequences, an initial set of fault association rules related to spatial distribution is output, allowing the mined rules to simultaneously reflect the statistical correlation of fault codes and wafer spatial distribution characteristics. Through dynamic fault reasoning based on the initial set of fault association rules, real-time fault codes are matched with the rule set and confidence scores are accumulated to generate a wafer-level fault evolution map containing fault propagation paths and candidate nodes for fault root causes, enabling dynamic tracking and quantitative characterization of fault development. Finally, based on the candidate nodes for fault root causes in the map and their accumulated confidence scores, priority is ranked, generating a fault root cause location command containing the coordinates of the die to be retested and the identifiers of associated test items, which is then sent to the wafer testing equipment to trigger the retest operation. This allows retest resources to focus on the die most likely to have fundamental defects and its associated items, thereby improving the accuracy and efficiency of DDR wafer test fault detection. Attached Figure Description
[0007] Figure 1 This is a flowchart of a DDR wafer test fault detection method based on association rules provided in an embodiment of the present invention.
[0008] Figure 2 This is a schematic diagram of the composition of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0010] Please see Figure 1 The flowchart below illustrates a DDR wafer testing fault detection method based on association rules, as provided in an embodiment of the present invention. This method can be executed by a computer device and may include the following steps: Step S100: Obtain the original electrical test data set of the DDR wafer test items generated after the electrical parameter measurement operation is performed on each memory cell of the wafer by the probe station during the wafer testing stage.
[0011] A DDR wafer refers to a circular semiconductor substrate that has completed the front-end semiconductor process, forming a large array of DDR memory cells, but has not yet been diced. The wafer testing stage specifically refers to the process of verifying the electrical performance of each die on the wafer using automated testing equipment after all manufacturing processes are completed and before packaging. A memory cell is the smallest structural unit in a DDR chip used for storing data, and can consist of a single access transistor and a storage capacitor; its electrical parameters directly determine the chip's function and performance.
[0012] The electrical parameter measurement operation involves the measurement unit inside the test equipment applying a set voltage or current excitation signal to the storage unit through a probe according to a preset test program. Then, a high-precision sampling circuit measures the response signal of the storage unit under the excitation, such as the voltage change, current magnitude, or charging and discharging time. These analog response signals are then converted into digital quantities by an analog-to-digital converter to obtain electrical parameter values that reflect the physical characteristics of the device.
[0013] DDR wafer test items can include, for example, memory cell access time, standby current, refresh cycle, input leakage current, and connectivity tests of various pins. Each test item corresponds to one or more electrical parameters that need to be measured. The raw electrical test data set refers to the massive data set generated by the test equipment after completing a comprehensive electrical parameter measurement of all measurable memory cells on the entire wafer, without any post-processing. The data in this set is stored in the form of data files, and each record contains at least the identifier of the test item, the physical coordinates of the memory cell on the wafer, and the numerical value of the electrical parameter measurement result corresponding to that test item.
[0014] In an exemplary scenario, the wafer test equipment first loads a test program containing all test items for the DDR wafer under test. This program defines the type, amplitude, duration, and sampling timing window of the excitation signal applied for each test item. Then, the probe station precisely moves the first memory cell under test on the wafer to below the probe card according to a preset stepping path. The probes on the probe card move vertically, applying a certain overdrive pressure to the pads of the memory cell to form a stable electrical contact. After contact is established, the test equipment sequentially executes the electrical parameter measurement operations for each test item according to the instructions of the test program. For example, when performing the access time test item, the test equipment first applies a high-level pulse to the word line of the memory cell to turn on the access transistor, and simultaneously applies a pre-charge voltage to the bit line. Then, it measures the time required for the voltage on the bit line to drop from the pre-charge voltage to a reference voltage; this time is used as the electrical parameter value of the access time. After completing all test item measurements for each memory cell, the probe station lifts the probe and moves to the position of the next memory cell, repeating the above contact and measurement process until all the specified test memory cells on the wafer are covered. During this process, the real-time operating system of the test equipment adds metadata tags to each measured electrical parameter value, including the corresponding test item name, the row and column coordinates of the storage unit, and the timestamp of test completion. Then, these structured data records are written into the test result file one by one, ultimately forming a raw electrical test data set covering the entire wafer and containing the measurement results of all test items.
[0015] Step S200: Perform correlation mining on the original electrical test data set to obtain a fault code sequence with spatial location correlation within the wafer, which includes the co-occurrence relationship of fault codes among different test items and the timing order of fault codes.
[0016] In one implementation, step S200 may specifically include the following steps S210 to S260: Step S210: Perform wafer map mapping processing on the original electrical test data set. Based on the physical coordinates of each storage cell in the wafer, project the test result data of the test item onto the wafer space grid to generate a wafer test data spatial distribution matrix with row and column coordinates.
[0017] Specifically, the basic wafer layout information can be parsed from the original electrical test data set, including the total number of rows and columns of memory cells on the wafer, as well as the physical spacing between adjacent memory cells in the row and column directions. Then, a two-dimensional matrix data structure with the exact same dimensions as the wafer layout is created, with each element initialized to empty. Next, each test result record in the original electrical test data set is read sequentially, and the row and column coordinates of the memory cell to which the record belongs, as well as the test item name and measurement value corresponding to the record, are extracted from each record. Based on the extracted row and column coordinates, the corresponding element position in the two-dimensional matrix is located, and the name of the test item and its measurement value are stored as key-value pairs in the data structure of this matrix element. When a memory cell corresponds to multiple test items, this matrix element will eventually become a set containing multiple key-value pairs. By traversing and filling all records in the original data set in this way, the originally scattered data, existing in the form of records, is reorganized into a structured wafer test data spatial distribution matrix closely corresponding to the wafer's spatial location. Each cell in the matrix completely contains the measurement results of the corresponding memory cell for all test items.
[0018] Step S220: Analyze the test result data of each storage unit in the wafer test data spatial distribution matrix, extract the fault code identifier corresponding to the test item whose test result data exceeds the preset specification limit, and the row and column coordinate information of the storage unit in the wafer space grid.
[0019] For example, starting from the first cell of the wafer test data spatial distribution matrix (the cell where both row and column coordinates are the initial values), the system reads the test result data for all test items stored within it. For each test item in this cell, its measured value is compared with the corresponding preset specification limit. For example, for an access time test item, its preset specification limit might be a maximum allowable time. If the measured access time exceeds this maximum allowable time, the test item is deemed abnormal. When a test item's measured value exceeds its preset specification limit, the system immediately retrieves the corresponding fault code identifier from a pre-built fault code lookup table based on the test item's internal identifier. Simultaneously, the row and column coordinates of the current cell are recorded. After extracting the fault code identifier and row and column coordinate information, they are stored as a temporary data pair. After determining the test results for all test items in a cell, the system continues to traverse the next cell in the matrix, repeating the above process of numerical comparison, fault code retrieval, and coordinate recording until all cells containing test result data in the wafer test data spatial distribution matrix have been traversed. Ultimately, for all memory cells on the entire wafer that exceeded the test parameters, one or more corresponding fault codes and the row and column coordinates of the memory cell itself were extracted.
[0020] Step S230: Perform noise filtering processing on the extracted fault code identifiers and row and column coordinate information based on probe contact status to remove isolated fault records caused by momentary poor contact, and generate a cleaned list of fault code identifiers and their corresponding row and column coordinates.
[0021] Specifically, the probe contact status monitoring data, synchronously recorded and stored by the probe station during the electrical parameter measurement operation, is first acquired. This data typically includes the contact resistance value between the probe and the pad at each measurement moment or a flag indicating contact stability. For each temporary data pair containing a fault code identifier and its row and column coordinates extracted in step S220, the moment when the fault record was generated is located based on its row and column coordinates, and the corresponding contact status parameters at that moment are extracted from the probe contact status monitoring data. If the contact status parameters at that moment show that the contact resistance exceeds the upper limit of the preset normal contact resistance range, or the contact stability flag indicates a momentary fluctuation, it is determined that the measurement was interfered with by poor probe contact. For fault records determined to be interfered with, their spatial distribution characteristics are further examined. That is, with the row and column coordinates of the fault record as the center, within a preset neighborhood range, such as adjacent memory cells on the top, bottom, left, and right, it is searched to see if there are other fault records in the cleaned list of fault code identifiers and their corresponding row and column coordinates. If there are no other fault records in the neighborhood range, the fault record is confirmed as an isolated fault record caused by momentary poor contact, and it is removed from the list to be processed. If other fault records exist within the neighborhood, even if the current record is slightly affected by poor contact, it may reflect the actual defect because it is in a fault cluster and therefore retained. After performing the above dual judgment based on contact status and spatial isolation on all temporary data pairs, the final generated list is the cleaned fault code identifier and row and column coordinate correspondence list.
[0022] Step S240: Based on the cleaned list of fault code identifiers and their corresponding row and column coordinates, construct an initial fault code transaction set with a single storage unit as the basic transaction unit. Each transaction record in the initial fault code transaction set contains at least one fault code identifier belonging to the same row and column coordinates.
[0023] For example, first, an empty transaction set is created to store the transaction records to be generated. Then, the cleaned list of fault code identifiers and their corresponding row and column coordinates is traversed. Each record in this list is associated with a specific row and column coordinate and a fault code identifier. Since a storage unit may generate multiple fault codes, there may be multiple records sharing the same row and column coordinate in the list. During processing, the row and column coordinates are used as the grouping key to merge all fault code identifiers with the same row and column coordinates together. For each unique row and column coordinate, all corresponding fault code identifiers are deduplicated and combined into a set. This set is a transaction record corresponding to that storage unit. The generated transaction record is added to the transaction set, and the row and column coordinates corresponding to this transaction are recorded so that the spatial location of each fault code set can be traced in subsequent steps. After traversing and merging all records in the list, the resulting transaction set is the initial fault code transaction set. Each transaction record in this set completely describes all concurrent fault types that occurred on a failed storage unit.
[0024] Step S250: Sort the transaction records in the initial fault code transaction set according to the physical scan order of the storage cells in the wafer to generate a transaction sequence with time order attributes and spatial adjacency relationships.
[0025] First, the physical scan path algorithm and starting point information actually used by the probe station during this wafer testing process are obtained. This information defines the recursive rule from the row and column coordinates of the starting memory cell to the row and column coordinates of the next memory cell. Based on this rule, a mapping function from row and column coordinates to their order number in the scan path can be established. Each transaction record in the initial fault code transaction set is traversed, and its corresponding row and column coordinates are extracted. The order number of this coordinate on the entire wafer scan path is calculated using the mapping function. Then, using the calculated order number as the sorting key, a sorting algorithm is used to sort all transaction records in the transaction set in ascending order. After sorting, the originally unordered transaction records become a list that strictly follows the wafer testing sequence, i.e., a transaction sequence. In this sequence, the record with the smaller index represents the memory cell tested first, and the record with the larger index represents the memory cell tested later. Furthermore, due to the spatial continuity of the scan path, the physical positions of consecutive records on the wafer are often adjacent or close, thus making the sequence contain both temporal order and spatial adjacency information.
[0026] Step S260: Perform sequential correlation analysis on the fault code identifiers of adjacent transaction records in the transaction sequence, identify the fault code combination patterns that appear sequentially on the same wafer or adjacent die locations, and generate a fault code sequence with wafer spatial location correlation based on the fault code combination patterns, which includes the co-occurrence relationship of fault codes between different test items and the order of occurrence of fault codes.
[0027] In one implementation, step S260 may specifically include the following steps S261 to S266: Step S261: Using a single transaction record as a sliding window unit, perform a sliding window scan operation with a preset step size on the transaction sequence to extract the fault code identifiers of at least two transaction records contained in the window and the row and column coordinates of the corresponding transaction records.
[0028] For example, first determine the size of the sliding window, which is the number of transaction records and is set to contain at least two transaction records. Simultaneously determine the sliding step size, which can be set to one transaction record to achieve dense window sliding. Starting from the first record in the transaction sequence, place the window on the sequence such that the left boundary of the window aligns with the first record, and the right boundary covers subsequent consecutive records according to the window size. At this point, read all transaction records within the window. For each record within the window, extract its fault code identifier set and the corresponding row and column coordinates, and combine these extracted information into a window dataset. After completing the data extraction for the current window, move the starting position of the window one transaction record backward according to the preset step size. At this point, the window will cover multiple consecutive records starting from the second record. At the new window position, perform the extraction operation of fault code identifiers and row and column coordinates again. Repeat the above moving and extraction process until the starting position of the window moves to the last window size plus one record in the sequence. After the last extraction, the scan ends. Through this sliding window scan, the originally linear transaction sequence is decomposed into a series of overlapping window data blocks containing local spatial and sequential information.
[0029] Step S262: Perform local spatial density clustering on the fault code identifiers and their row and column coordinates extracted by the sliding window, merge multiple fault code identifiers with a spatial distance smaller than the preset clustering radius into a composite fault event identifier, and generate a clustered fault code identifier set.
[0030] For example, first, obtain all fault code identifiers and their corresponding row and column coordinates within the current sliding window. Using these coordinates as input points, apply a density-based spatial clustering algorithm. The core logic of the algorithm is: starting from the first unprocessed point, search for other points within a preset clustering radius, centered on that point. If at least one other point exists, mark the center point and all found points as belonging to the same temporary cluster. Then, for each newly added point within this temporary cluster, repeat the operation of searching for unmarked points within the radius centered on that point, continuously expanding the cluster's range until no new points can be added. After constructing a cluster, generate a unique composite fault event identifier for that cluster. After traversing all points, all fault code identifiers in the original window are divided into several clusters and some isolated points not assigned to any cluster. For each cluster, replace all existing fault code identifiers within the cluster with its corresponding composite fault event identifier. For isolated points not clustered, their original fault code identifiers remain unchanged. After the above processing, the list of fault code identifiers in the original window is transformed into a clustered set of fault code identifiers, which may contain both single fault code identifiers and composite fault event identifiers representing multiple faults.
[0031] Step S263: Calculate the Manhattan distance in the spatial dimension of different fault code identifiers in the clustered fault code identifier set. Determine whether fault code identifiers belonging to the same window come from spatially adjacent storage units based on whether the Manhattan distance is less than or equal to the preset adjacency distance threshold.
[0032] In the specific execution process, the spatial location represented by each element in the clustered fault code identifier set can be determined first. If the element is a composite fault event identifier, its spatial location is taken as the geometric center coordinates of the row and column coordinates of the memory cells corresponding to all the original fault codes contained in the composite event. The calculation method is to add the row coordinates of each point and divide by the number of points to obtain the center row coordinate, and add the column coordinates of each point and divide by the number of points to obtain the center column coordinate. If the element is a single fault code identifier, its spatial location is the row and column coordinates of the memory cell corresponding to it. After obtaining the spatial location of each element, the Manhattan distance between any two different elements in the set is calculated. That is, the absolute value of the difference between the row coordinates and the absolute value of the difference between the column coordinates of the two center coordinates are calculated respectively, and then the two absolute values are added together. For each pair of Manhattan distances calculated, it is compared with a preset adjacency distance threshold. If the distance is less than or equal to the preset adjacency distance threshold, the fault events corresponding to the two fault code identifiers are determined to be spatially adjacent; if the distance is greater than the preset adjacency distance threshold, they are determined to be spatially non-adjacent. This determination will serve as the basis for selecting valid fault code combinations in subsequent steps.
[0033] Step S264: Perform time-order labeling on the fault code identifiers determined to be from spatially adjacent storage units, and add the order number of the transaction record in the transaction sequence to each fault code identifier to generate a set of candidate fault code combinations with spatial adjacency constraints and time order.
[0034] In this embodiment of the invention, the set of fault code identifiers from the same sliding window, obtained after steps S262 and S263, is first traversed. For each fault code identifier in the set (whether single or compound), the original transaction record from which it originated is found. Since a compound fault event identifier corresponds to multiple original fault codes, it is necessary to backtrack to the transaction record containing each original fault code constituting the compound event. From these transaction records, their order numbers in the transaction sequence are extracted. If a fault code identifier corresponds to multiple original records (in the case of compound events), the smallest and largest order numbers among these records are taken as the start and end times of the fault event. Then, this fault code identifier and its time information (e.g., the start order number) are combined to form a new element with a timestamp. After all fault code identifiers determined to be spatially adjacent within the sliding window have completed the above-mentioned time information appending, these elements with timestamps are arranged according to their order of appearance in the sequence to form an ordered list. This list, along with its corresponding sliding window information, is added to the candidate fault code combination set as a candidate fault code combination. Perform this operation for each sliding window, and the final set contains a large number of candidate combinations with spatial adjacency constraints and a clear temporal order among the elements.
[0035] Step S265: Count the frequency of occurrence of each fault code combination in the candidate fault code combination set, and select the fault code combinations whose occurrence frequency exceeds the preset frequency threshold as high-frequency adjacent fault code combination patterns.
[0036] In practical applications, an empty mapping table is first created to record various fault code combination patterns and their frequency of occurrence. Each candidate combination in the candidate fault code combination set is then traversed. For each candidate combination, a key is generated to uniquely identify the combination pattern based on its list of fault code identifiers with timestamps. The key generation method needs to consider both the element type and the element order. For example, each element in the list can be concatenated into a string in sequence, with elements connected by a set delimiter. After generating the key, it is searched in the mapping table. If the key already exists in the mapping table, the corresponding count is incremented by one. If the key does not exist in the mapping table, it is added to the mapping table, and its count is initialized to one. After traversing the entire candidate fault code combination set, the mapping table records the total frequency of each fault code combination pattern. Next, each entry in the mapping table is traversed, and the count of each entry is compared with a preset frequency threshold. If the count is greater than or equal to the preset frequency threshold, the fault code combination pattern corresponding to that entry is determined to be a high-frequency pattern and extracted. All the extracted high-frequency patterns together form a set of high-frequency adjacent fault code combination patterns.
[0037] Step S266: Analyze the occurrence order of fault code identifiers in the high-frequency adjacent fault code combination mode, extract the fixed order relationship between the fault code identifiers that appear first and the fault code identifiers that appear later, and generate a fault code sequence with spatial location association within the wafer based on the fixed order relationship, which includes the co-occurrence relationship of fault codes between different test items and the timing order of fault codes.
[0038] For example, first, traverse each pattern in the set of high-frequency adjacent fault code combination patterns. Each pattern is itself an ordered list. From this list, all sequential relationship pairs between adjacent elements can be extracted. For example, for the list [Identifier X, Identifier Y, Identifier Z], two sequential relationships can be extracted: (Identifier X -> Identifier Y) and (Identifier Y -> Identifier Z). For each extracted sequential relationship, the predecessor identifier and successor identifier are recorded. Then, these sequential relationship pairs are used as a basic unit. All sequential relationship pairs extracted from high-frequency patterns constitute a directional association graph. In this graph, nodes are fault code identifiers (including composite event identifiers), and directed edges represent that after one fault occurs, another fault will immediately occur in a spatially adjacent position. Finally, based on this directed graph and combined with the spatial location information in the original transaction sequence, the final output is generated. This output can be represented as a set where each element is a triple in the form of (previous fault code identifier, subsequent fault code identifier, spatial association description). The spatial association description indicates in which adjacent row and column coordinate regions the two fault codes usually appear, or the relative positional offset pattern between them. This generates a fault code sequence with spatial location associations within the wafer, containing the co-occurrence relationships of fault codes among different test items and the timing order of fault code occurrence.
[0039] Step S300: Perform frequent itemset iterative search and strong association rule generation on the fault code sequence based on minimum support threshold and minimum confidence threshold to obtain an initial fault association rule set associated with the spatial distribution.
[0040] In one implementation, step S300 may specifically include the following steps S310 to S360.
[0041] Step S310: Use the fault code sequence as the input transaction database for association rule mining, perform a single scan of the input transaction database, count the frequency of occurrence of each fault code identifier as a single item, and filter out the fault code identifiers whose frequency of occurrence meets the preset minimum support threshold as a frequent item set.
[0042] For example, first define an empty counting map table, where the key is the fault code identifier and the value is the frequency count of that identifier. Then, begin traversing the input transaction database. For each transaction record in the database, read all the fault code identifiers contained in that record. For each read fault code identifier, search for that identifier in the counting map table. If an entry with that identifier as the key already exists in the map table, increment the count value of that entry by one; if the key does not exist in the map table, add a new entry with that identifier as the key and initialize its count value to one. After a single pass through the entire transaction database, the counting map table records the total frequency of each fault code identifier. Next, obtain a preset minimum support threshold, which can be a specific number. Traverse each key-value pair in the counting map table, comparing the count value of each key-value pair with the minimum support threshold. If the count value is greater than or equal to the minimum support threshold, add the fault code identifier corresponding to that key-value pair to the result list of the frequent itemset. The final list is the frequent itemset, which contains all single fault code types that occur with a sufficiently high frequency.
[0043] Step S320: Normalize the frequency of each fault code identifier in the frequent item set based on the wafer region distribution density to eliminate the influence of the difference in test density between the wafer center region and the edge region on the support calculation, and generate the normalized frequent item set.
[0044] Specifically, a wafer region density distribution baseline map is first constructed. This baseline map can be obtained by analyzing a large amount of historical wafer test data, statistically analyzing the average probability or total frequency of any fault code occurring at each grid location or macroscopic region (such as the central region, intermediate region, and edge region) on the wafer, and using this as the background density of that region. For each fault code identifier in the frequent one-dimensional set, it is necessary to trace back to the original transaction database, find all transaction records containing that fault code identifier, and obtain the row and column coordinates of the storage unit corresponding to each record. Based on these row and column coordinates, the wafer region to which each fault occurrence location belongs is determined. Then, based on the background density of that region, a normalization coefficient is calculated. For example, the value of this coefficient can be inversely proportional to the background density of the region, that is, the higher the background density of the region, the smaller its normalization coefficient. Next, the original occurrence frequency of the fault code identifier is multiplied by its average normalization coefficient across all occurrence locations to obtain the normalized frequency of the fault code identifier. Alternatively, a more refined method can be used, calculating a weight for each occurrence separately and then summing them to obtain the total normalized frequency. The original occurrence frequencies of all fault code identifiers are replaced with the calculated normalized frequencies, while retaining those identifiers whose normalized frequencies meet the adjusted minimum support threshold. This generates a normalized frequent itemset. This normalized frequent itemset will be used for subsequent itemset generation, making the support calculation more fair and accurate.
[0045] Step S330: Perform recursive join and pruning operations based on the normalized frequent itemsets. Generate candidate k-itemsets by joining two frequent itemsets containing k-1 items, and calculate the support count of the candidate k-itemsets in the input transaction database. Select candidate k-itemsets whose support count meets the preset minimum support threshold as frequent k-itemsets, until no new frequent itemsets can be generated.
[0046] In the specific execution process, we first set k to two, using the normalized frequent itemsets generated in the previous step as the initial set of frequent one-itemsets. Then, we enter a loop to generate candidate k-itemsets. Specifically, for each itemset A in the frequent k-1 itemsets, we pair it with another itemset B, checking if the first k-2 fault code identifiers of A and B are completely identical, and if the last identifier of A is lexicographically less than the last identifier of B. If the conditions are met, we merge A and B to generate a new candidate set C containing k fault code identifiers. For each generated candidate set C, we perform a pruning operation: we check all subsets of C containing k-1 items to see if they all exist in the frequent k-1 itemsets. If any k-1 item subset is not in the frequent k-1 itemsets, then according to the property, C cannot be frequent, so C is removed from the candidate set. After generating and pruning all candidate sets, we obtain the final list of candidate k-itemsets. Next, we need to calculate the support count of each itemset in the candidate k-itemsets. Therefore, the input transaction database is scanned again. For each transaction record in the database, it is checked which itemsets from the candidate k-itemsets it contains. A common approach is to generate all subsets of size k for each record and then match them with the candidate k-itemsets, incrementing the count of each successfully matched candidate itemset by one. After traversing the entire database, the actual support count for each candidate k-itemset is obtained. Finally, the support count of each candidate k-itemset is compared with a preset minimum support threshold, and itemsets with counts greater than or equal to the threshold are retained; these are the frequent k-itemsets. k is incremented by one, and the above process of joining, pruning, scanning, counting, and filtering is repeated until no candidate k-itemsets can be generated in a certain loop, or all candidate k-itemsets fail the support filtering, at which point the algorithm terminates. The final set of all frequent itemsets, frequent 2-itemsets, up to frequent N-itemsets is the output of this step.
[0047] Step S340: Extract all non-empty subsets from the generated frequent k-itemsets. For each frequent itemset and its non-empty subset, calculate the confidence of the association rule derived from the non-empty subset. The confidence of the association rule is the ratio of the support count of the frequent itemset to the support count of its non-empty subset.
[0048] In one implementation, step S340 may specifically include the following steps S341 to S346: Step S341: For each frequent k-itemset, list all fault code identifiers contained in the frequent k-itemset as elements of the universal set, and generate all non-empty proper subsets consisting of the elements of the universal set, excluding the empty set and the universal set itself, through a recursive backtracking algorithm.
[0049] For example, first, all fault code identifiers from a frequent k-item set are placed into a list, and the length of the list is determined by k. Then, an empty temporary list is initialized to store the subset currently being built. A recursive backtracking function is called, which takes the index of the currently considered element as an argument. Inside the function, if the current index equals k, it means all elements have been considered. If the temporary list is not empty and its length is less than k, a copy of the temporary list is recorded as a non-empty proper subset, and then the function is returned. If the current index is less than k, two branches are chosen. The first branch selects the element not pointed to by the current index, i.e., the function is recursively called to process the next index. The second branch selects the element pointed to by the current index, i.e., the element is added to the temporary list, and then the function is recursively called to process the next index. After the recursion returns, the element just added to the temporary list needs to be removed to allow for exploration of other branches. In this way, the recursive function can traverse all possible combinations. After the recursive function finishes executing, it obtains a set of all non-empty proper subsets of the frequent k-itemset, where each subset is a list containing one to k-1 fault code identifiers.
[0050] Step S342: Perform a redundancy check on each generated non-empty true subset, and remove low-association subsets whose difference from the overall support count of frequent itemsets exceeds a preset difference threshold, to obtain the filtered effective non-empty true subsets.
[0051] In the specific execution process, the support count of the currently processed frequent k-itemsets is first obtained, denoted as S. full Then, iterate through each non-empty proper subset generated in step S341. For each subset, obtain the support count of the subset itself as an itemset in the input transaction database. This count can be calculated and stored during the frequent itemset mining process in step S330 and can be obtained by looking up previous results. Let the support count of the subset be S. sub Next, calculate S. full With S sub The degree of difference between them can be measured, for example, by calculating the absolute difference S. sub Subtract S full Alternatively, calculate the relative difference, i.e., S. full Divide by S subThe difference value is compared to a preset difference threshold. The preset difference threshold can be set as a range, for example, requiring that the support count of the subset not be too close to or too far from the support count of frequent itemsets. If the calculated difference falls within the range allowed by the preset difference threshold, the subset is considered valid and retained. If the difference exceeds the allowed range, for example, S... sub With S full Almost equal, or S sub Much smaller than S full If a subset is found to be low-association, it is removed from the candidate list. After performing this check on all non-empty proper subsets, the remaining subsets are the filtered valid non-empty proper subsets, which will be used for subsequent rule generation.
[0052] Step S343: For each valid non-empty true subset, use the valid non-empty true subset as the set of antecedent fault codes for the association rule, and use the remaining part of the frequent k-item set after removing the valid non-empty true subset as the set of consequent fault codes for the association rule.
[0053] First, a subset is selected from the filtered list of valid non-empty proper subsets. This subset is directly marked as the antecedent of the rule. Then, the complete list of fault code identifiers for the current frequent k-itemset is obtained. From this complete list, each fault code identifier is checked one by one. If the identifier exists in the antecedent subset, it is skipped; if the identifier does not exist in the antecedent subset, it is added to a new set. After traversal, this new set contains all remaining fault code identifiers in the complete frequent itemset excluding the antecedent portion, and these are marked as the consequent of the rule. For example, for a frequent three-itemset {A,B,C}, if the current valid non-empty proper subset is {A,B}, then the consequent is {C}. If the subset is {A}, then the consequent is {B,C}. This constructs a candidate association rule of the form "antecedent -> consequent" for each valid non-empty proper subset, and this rule is recorded for confidence calculation.
[0054] Step S344: Count the number of transaction records that appear simultaneously in the fault code sequence as a whole for this valid non-empty true subset, and use this count as the predecessor support count.
[0055] For the currently processed valid non-empty true subset, it is first confirmed that it is also an itemset. Since this subset was generated from a frequent k-itemset in step S341, it must also be a frequent itemset, or at least an itemset whose support has been calculated in step S330. Therefore, it is not necessary to scan the entire transaction database again. The support count corresponding to this subset can be directly looked up in a mapping table maintained during the mining of frequent itemsets in step S330. The key of this mapping table is the itemset (which can be represented as a sorted list of fault code identifiers), and the value is the support count of the itemset in the entire transaction database. By converting the current valid non-empty true subset into a key according to the same sorting rule, and then searching in the mapping table, the support count of the subset can be directly obtained. This count is the antecedent support count of the association rule, denoted as S_antecedent. If, for some reason, this count is not pre-stored, a database scan for this specific subset is required to count its occurrences.
[0056] Step S345: Divide the count of transaction records in which frequent k-itemsets appear simultaneously in the fault code sequence by the count of previous support to obtain the preliminary confidence value of the association rule derived from the effective non-empty true subset.
[0057] First, obtain the support count of the currently processed frequent k-itemset, S_consequent_union, from the result of step S330. Then, obtain the antecedent support count of the current valid non-empty proper subset, S_antecedent, from step S344. Perform a division operation on these two values. The result of this operation, the quotient, is the preliminary confidence value of the association rule derived from the valid non-empty proper subset for the entire frequent k-itemset. This preliminary value is associated with and stored along with the antecedent and consequent information of the current rule for further correction and final confidence determination.
[0058] Step S346: Perform confidence decay correction processing on the preliminary confidence value of the association rule corresponding to each valid non-empty true subset. The degree of decay is positively correlated with the average distance between the antecedent fault code set and the consequent fault code set in spatial distribution. Determine the final confidence value of the association rule based on the corrected preliminary confidence value of the association rule.
[0059] In the specific execution process, it is first necessary to determine the average distance in spatial distribution between the antecedent fault code set and the consequent fault code set. To do this, we trace back to the input transaction database and find all transaction records that simultaneously contain both the antecedent and consequent sets (i.e., the entire frequent itemset). For each such record, we record the row and column coordinates of the storage units corresponding to each fault code in the antecedent set, and the row and column coordinates of the storage units corresponding to each fault code in the consequent set. We calculate the spatial center point of the antecedent set, for example, by using the average of the row coordinates of all antecedent coordinates as the center row and the average of the column coordinates as the center column. Similarly, we calculate the spatial center point of the consequent set. We then calculate the Manhattan distance or Euclidean distance between these two center points. For all transaction records containing all items of this rule, we average this distance value to obtain the average distance. Next, we define a decay function that takes the average distance as input and outputs a decay coefficient between zero and one. The characteristic of this function is that when the average distance is small, the decay coefficient is close to one; as the average distance increases, the decay coefficient gradually decreases, approaching zero. Multiply the initial confidence value calculated in step S345 by this attenuation coefficient to obtain the corrected confidence value. This corrected value is the final confidence value of the association rule, which reflects both the strength of the statistical association and the tightness of spatial propagation.
[0060] Step S350: Filter association rules whose confidence level is greater than or equal to the preset minimum confidence level threshold, extract the antecedent fault code set and consequent fault code set of the association rule, and record the spatial location distribution information of the antecedent fault code set and consequent fault code set in the fault code sequence.
[0061] In one implementation, step S350 may specifically include the following steps S351 to S356: Step S351: Compare the calculated confidence value of each association rule with the preset confidence threshold, and retain the association rules whose confidence values are greater than or equal to the preset confidence threshold as candidate strong association rules.
[0062] First, a list is created to store the strong association rules to be output. Then, all candidate association rules generated in step S340 are traversed. Each candidate rule is associated with its antecedent, consequent, and final confidence level after spatial decay correction. For the currently traversed rule, its final confidence level value is extracted and compared with a preset confidence threshold. If the confidence level value is greater than or equal to the preset confidence threshold, the rule is determined to be a candidate strong association rule, and its antecedent fault code set, consequent fault code set, and confidence level value are stored in the candidate strong association rule list. If the confidence level value is less than the preset confidence threshold, the rule is determined to be insufficiently strong and is discarded. After traversing all candidate rules, the resulting candidate strong association rule list is the initially selected set of rules that meet the minimum reliability requirements.
[0063] Step S352: Perform clustering and merging processing on the retained candidate strong association rules based on spatial distribution overlap, merge rules with highly overlapping spatial distribution areas of the antecedent fault code set and the consequent fault code set into rule clusters, and generate a merged candidate strong association rule set.
[0064] Specifically, the spatial distribution information of each candidate strongly associated rule is initially calculated. This can be based on the row and column coordinates of all transaction records corresponding to all antecedents and consequents that appear in the rule. For example, the centroid coordinates and convex hull regions of the antecedent set, and the centroid coordinates and convex hull regions of the consequent set can be calculated. Then, a clustering algorithm, such as density-based clustering or hierarchical clustering, is used, with rules as the basic unit and the spatial distribution similarity between rules as the distance metric. The spatial distribution similarity between two rules can be defined as the sum of the reciprocals of the centroid distances of their antecedents and consequents, or more precisely, considering the degree of overlap between their antecedent and consequent convex hulls. For each cluster formed after clustering, a new rule needs to be generated to represent the cluster. The antecedent of the representative rule can be the union of the antecedent sets of all rules in the cluster, or the combination of the most frequently occurring antecedent fault codes; the consequent is similar. The confidence of the representative rule can be the weighted average of the confidence of all rules in the cluster, with the weights related to the support of each rule. Ultimately, all the representative rules generated by the clusters, along with the isolated rules that were not assigned to any cluster, together constitute the merged set of candidate strongly correlated rules.
[0065] Step S353: For each candidate strongly associated rule in the merged candidate strongly associated rule set, parse at least one fault code identifier contained in its predecessor fault code set, and backtrack the set of row and column coordinates corresponding to the transaction records in the fault code sequence that simultaneously contain all fault code identifiers in the predecessor fault code set.
[0066] For example, a rule is taken from the merged set of candidate strongly associated rules, and its predecessor fault code set is read. Then, an empty list is created to store all matched row and column coordinates. Next, the original fault code sequence, i.e., the transaction sequence generated in step S250, is scanned. Each record in this sequence contains the row and column coordinates corresponding to the transaction and all fault code identifiers in that transaction. For each transaction record in the sequence, it is checked whether the set of fault code identifiers contained in the record completely contains the predecessor fault code set of the current rule. If it completely contains them, i.e., the predecessor set is a subset of the current record, the row and column coordinates corresponding to the record are added to the previously created list. If the record does not contain all the fault codes of the predecessor, it is skipped. After performing this check on all transaction records in the fault code sequence, the accumulated coordinate points in the list are the set of all historical occurrences of the predecessor of this rule. This set of row and column coordinates is associated with the rule and stored for subsequent spatial region boundary calculations.
[0067] Step S354: Perform convex hull calculation on the set of row and column coordinates to generate the smallest convex polygon that surrounds all occurrence positions of the previous fault code set as the spatial distribution region boundary of the previous fault code set.
[0068] For example, first obtain all row and column coordinate points collected in step S353 as the antecedent of the current rule, and treat these points as a set of points on a two-dimensional plane. Then, apply the convex hull algorithm to process this set of points. A classic convex hull algorithm is the Graham scan algorithm. The execution flow of the algorithm is as follows: First, find the point with the smallest y-coordinate in the set of points. If there are multiple points, take the leftmost one and use it as the starting point. Then, calculate the angle of the vector formed by the starting point and all other points relative to the horizontal direction, and sort all points according to this angle. Next, initialize an empty stack and push the starting point and the first sorted point onto the stack. After that, traverse the other sorted points. For each new point, check whether the direction formed by the top two points of the stack and the current new point is a left turn or a right turn. If it forms a right turn, it means that the top point of the stack is not a point on the convex hull, so pop it from the stack, and continue to check the direction of the new top two points of the stack and the current point until the direction becomes a left turn or there is only one point left in the stack. When the direction becomes a left turn, push the current point onto the stack. After traversing all points, the remaining points in the stack, when connected sequentially, form the smallest convex polygon enclosing all input points. The list of vertices of this polygon represents the spatial distribution boundary of the antecedent fault code set. This list of polygon vertices is then appended as an attribute to the current rule.
[0069] Step S355: Parse at least one fault code identifier contained in the consequent fault code set of each candidate strong association rule, and backtrack the set of row and column coordinates corresponding to the transaction records in the fault code sequence that contain any fault code identifier in the consequent fault code set.
[0070] First, the set of consequent fault codes is read from the currently processed candidate strong association rules. Then, a new empty list is created to store the coordinates of all occurrences of the consequent. Next, the original fault code sequence is scanned again. For each transaction record in the sequence, it is checked whether the set of fault code identifiers contained in the record intersects with the set of consequent fault codes; that is, whether at least one fault code identifier appears in both the record set and the consequent set. If an intersection exists, the record is determined to have a consequent fault, and the corresponding row and column coordinates are added to the list prepared for the consequent. If the fault codes in the record do not overlap with the consequent set, they are skipped. After scanning the entire fault code sequence, the accumulated coordinates in the list represent the set of all historical occurrences of the consequent of the rule. This set may be very large, and multiple clusters may form between the points, providing input data for the next step of density clustering analysis.
[0071] Step S356: Perform density clustering analysis on the row and column coordinate sets corresponding to the set of subsequent fault codes to identify at least one high-density clustering region, and describe the spatial distribution area information of the set of subsequent fault codes with the center coordinates and coverage radius of each high-density clustering region.
[0072] For example, first, obtain the set of all row and column coordinate points collected in step S355 for the current rule. Apply a density-based clustering algorithm, such as the DBSCAN algorithm, to process this set of points. The algorithm requires two preset parameters: neighborhood radius and minimum number of points. The neighborhood radius defines the size of the neighborhood to consider when determining whether a point is a core point, and the minimum number of points defines the minimum number of points a core point must contain in its neighborhood. The algorithm execution flow is as follows: Traverse all points, mark each point as a core point, that is, check whether the number of points contained in the circle centered on the core point and with the neighborhood radius as the distance reaches the minimum number of points. Then, starting from an unvisited core point, create a new cluster and add the core point and all points in its neighborhood to the cluster. Next, for each newly added point in the cluster, if it is also a core point, add the unvisited points in its neighborhood to the cluster, and continue to expand the range of the cluster in this way until there are no more new points to add. After that, continue to find the next unvisited core point and start building the next cluster. All points marked as noise (i.e., points not belonging to any cluster) are ignored. After the algorithm completes, several clusters are obtained. For each identified cluster, its geometric center coordinates are calculated, which are the average row and column coordinates of all points within the cluster. Then, the maximum Manhattan or Euclidean distance from all points within the cluster to this center coordinate is calculated as the coverage radius of the cluster. Finally, for the current rule, the spatial distribution information of its consequent fault code set is described as a list, where each element is a tuple containing the center coordinates and coverage radius of the cluster. Appending this information to the rule completes a comprehensive description of the rule's spatial attributes.
[0073] Step S360: Based on the set of preceding fault codes, the set of succeeding fault codes, and the spatial location distribution area information, construct an initial fault association rule set associated with spatial distribution, with association rules as nodes and spatial location distribution area information as node attributes.
[0074] Specifically, first, an empty set object is created to store the final association rules. Then, each rule processed in step S350 is traversed. These rules already possess attributes such as the antecedent fault code set, the consequent fault code set, the confidence score corrected for spatial decay, the spatial distribution boundary convex hull of the antecedent, and the center and radius of multiple high-density clustering regions of the consequent. For each rule, this scattered information is integrated into a unified data structure, such as a record or an object. In this structure, the antecedent set, consequent set, confidence score, antecedent convex hull vertex list, and consequent clustering region list (each region includes the center row coordinates, center column coordinates, and coverage radius) of the rule are clearly identified. After integrating all rules, these rule nodes are added one by one to the previously created set object. The final set obtained is the initial fault association rule set associated with spatial distribution required in step S360, which provides a complete rule template for subsequent online inference.
[0075] Step S400: Perform dynamic fault reasoning based on the initial fault association rule set. Match the real-time fault codes in the real-time acquired DDR wafer online test data stream with the initial fault association rule set and calculate the confidence level to generate a wafer-level fault evolution map containing fault propagation paths and candidate nodes for fault root causes.
[0076] As one implementation method, step S400 may specifically include the following steps S410 to S460: Step S410: Receive the DDR wafer online test data stream generated by the wafer tester for the current wafer during real-time testing, and extract the real-time fault code generated by each tested memory cell on the current wafer and its corresponding real-time row and column coordinates from the online test data stream.
[0077] First, a data receiving service is initiated on the computing system running the fault detection method. This service establishes a connection with the wafer testing equipment via a network protocol and subscribes to test data publishing topics. When the testing equipment begins testing the current wafer, it continuously sends data packets to the network. Each data packet encapsulates the results of testing one or more memory cells. The data receiving service continuously receives these data packets and unpacks and parses each packet. For each memory cell described in the data packet, its test completion flag and the judgment results of all test items are extracted. For each test item judged as failed, a corresponding fault code identifier is generated or obtained by looking up a table according to a preset fault code encoding rule. Simultaneously, the row and column coordinates of the memory cell are parsed from the data packet. All extracted real-time fault code identifiers and the row and column coordinates of the cell are recorded as a whole and pushed to the next stage of the data processing pipeline. This process continues until all memory cells of the current wafer have been tested.
[0078] Step S420: Assemble the extracted real-time fault codes and real-time row and column coordinates into a real-time fault code stream according to the order in which the tests are completed. The real-time fault code stream includes the fault code identifier, its real-time location information within the wafer, and a real-time timestamp.
[0079] Specifically, while extracting the fault code set and row / column coordinates for each memory cell from the online test data stream, the arrival time of the data packet or the test completion timestamp contained within the data packet is recorded. This timestamp is combined with the extracted fault code set and row / column coordinates to form a data tuple. This tuple is added to a queue maintained in chronological order. Since the test equipment tests memory cells sequentially according to the physical scan order, this queue naturally maintains the temporal relationship of fault occurrence. As testing progresses, new tuples are continuously added to the tail of the queue, while the tuple at the head of the queue corresponds to the earliest tested memory cell. This real-time growing, ordered sequence of tuples constitutes the real-time fault code stream. Each element in this stream precisely describes which specific fault type occurred at a specific location on the wafer at a specific time.
[0080] Step S430: Traverse each association rule in the initial fault association rule set, and match the fault code identifiers appearing in the real-time fault code stream with the preceding fault code set of the association rule. If the fault code identifiers appearing in the real-time fault code stream completely contain the preceding fault code set of a certain association rule, then activate that association rule.
[0081] In the actual execution process, a set is first maintained, which is the union of all different fault code identifiers that have appeared in the real-time fault code stream. This set expands continuously as new fault codes are added. Then, for each rule in the initial fault association rule set, its predecessor fault code set is retrieved. It is checked whether this predecessor set is a subset of the union of currently appearing fault code identifiers. If so, the predecessor condition of the rule is considered satisfied. At this point, the event of the rule's activation is recorded, including the activated rule itself, the time of activation (which could be the current moment), and the fault code information on which the activation is based. This rule then transitions from a dormant state to an active state and participates in subsequent dynamic reasoning. If the predecessor set is not a subset of the current union, the rule remains inactive, waiting for subsequent fault codes to appear.
[0082] Step S440: For an activated association rule, determine the activation position of the association rule based on the real-time location information of the fault code identifier in the previous fault code set appearing in the real-time fault code stream, and predict the candidate location area where the subsequent fault code may appear in the area that has not yet appeared in the real-time fault code stream or in the area that may appear in subsequent tests based on the subsequent fault code set.
[0083] In one implementation, step S440 may specifically include the following steps S441 to S446: Step S441: Parse at least one fault code identifier contained in the antecedent fault code set of the activated association rule, extract the real-time row and column coordinates of all storage cells containing these fault code identifiers from the real-time fault code stream, and calculate the geometric center point of these real-time row and column coordinates as the activation center coordinates of the association rule.
[0084] For example, first, obtain the set of preceding fault codes from the activated rule. Then, create an empty list to store all coordinates that meet the conditions. Iterate through each record in the real-time fault code stream. For each record, check if its set of fault code identifiers intersects with the set of preceding fault codes, i.e., whether any of the preceding fault codes have appeared. If so, add the row and column coordinates of that record to the list. Since the rule activation requires all fault codes in the preceding set to have appeared, the list will eventually contain all locations where preceding fault codes have appeared. After collecting all coordinates, sum all row coordinates in the list and divide by the list length to obtain the average row coordinate; sum all column coordinates and divide by the list length to obtain the average column coordinate. The new coordinate point formed by these two averages is the activation center coordinate. Use this coordinate as the spatial anchor point for this activation of the rule on the current wafer.
[0085] Step S442: Extract the spatial distribution area information of the subsequent fault code set from the spatial location distribution area information of the association rule. The spatial distribution area information includes the center coordinates and coverage radius of at least one high-density clustering area of the subsequent fault code set in the historical wafer test data.
[0086] Specifically, the first step is to locate the data structure in memory for the currently activated rule. Within this structure, there is a field specifically for storing consequent spatial distribution information. Reading this field returns a list, where each element corresponds to a historical high-density clustering region. Each element itself is a structure containing three values: the center row coordinates, the center column coordinates, and the coverage radius of the region. This extracted information is temporarily stored as a baseline template for the next step of location prediction.
[0087] Step S443: Perform centroid drift correction processing on the spatial distribution area information of the extracted fault code set, dynamically adjust the center coordinates of the predicted area according to the boundary of the area that has been tested on the current wafer, and generate the corrected predicted center coordinates.
[0088] In the specific execution process, the boundary information of the areas that have been tested on the current wafer is first obtained. This can be roughly determined by recording the row and column coordinates of the last tested memory cell. Then, historical data shows that subsequent failures usually occur in a certain directional offset relative to the previous active position. To calculate this offset, all instances of this rule in the historical data are reviewed. For each instance, the average offset vector from the previous active center to the center of the subsequent cluster area is calculated. This offset vector represents a historical pattern. On the current wafer, the coordinates of the active center have already been calculated. Adding this historical average offset vector to the active center coordinates yields a preliminary, offset-based predicted center. Next, this predicted center is compared with the boundary of the tested area. If the predicted center falls within the tested area, it means the predicted position has been missed and fine-tuning is needed according to the scan direction, for example, moving the predicted center forward to near the boundary of the area to be tested. The final coordinates obtained after this adjustment are the corrected predicted center coordinates.
[0089] Step S444: Using the active center coordinates as the reference point, calculate the offset vector between the reference point and each corrected predicted center coordinate, and superimpose the offset vectors onto the spatial coordinate system of the current wafer to obtain the predicted center coordinates where the subsequent fault codes may occur in the current wafer.
[0090] Assume that after step S443, one or more corrected predicted center coordinates are obtained, denoted as P. corrected Step S441 has yielded the coordinates of the activation center, denoted as C. activate For each P corrected Calculate its relationship with C activate The offset vector between them, i.e., vector V=(P corrected row coordinates minus C activate row coordinates, P corrected column coordinates minus C activate (Column coordinates). This vector V represents the standard offset of the consequent cluster center relative to the predecessor activation center in historical data. Now, on the current wafer, with the current actual activation center C... activate Starting from point P, we directly apply this offset vector V to obtain a new point P. predicted =(C activate The row coordinates of C are added to the row coordinates of V. activate (The column coordinates of P plus the column coordinates of V). predicted This refers to the center coordinates where the predicted subsequent fault code might occur in the current wafer coordinate system. This process is performed for each corrected predicted center coordinate to obtain a set of predicted center coordinates.
[0091] Step S445: Using the predicted center coordinates as the center and the coverage radius of the historical high-density cluster area as the reference radius, and combining the boundary of the untested area of the current wafer, perform outward expansion correction processing to generate candidate location areas where subsequent fault codes may appear in the current wafer.
[0092] For example, first obtain the coverage radius R of the corresponding prediction center from the historical spatial distribution information of the rules. Then, use the prediction center coordinates P calculated in step S444. predicted A circular region is defined on the wafer grid with center R and radius R. Then, the boundaries of the untested regions of the current wafer are obtained, typically determined by the location of the last tested memory cell and the scan path direction. The untested region can be defined as a set containing all cells with row numbers greater than a certain threshold, or row numbers equal to the threshold but column numbers greater than the threshold. A geometric intersection operation is performed between the previously defined circular region and this untested region to obtain one or more intersecting sub-regions. To ensure prediction continuity, if an intersecting sub-region is fragmented, its minimum bounding rectangle or convex hull can be used. Furthermore, if the area of the intersecting region is much smaller than the typical area of historical regions, the region can be appropriately expanded along the scan direction, for example, extending the region a certain distance along the scan direction until a preset minimum area requirement is reached. After these reduction and expansion processes, the resulting one or more geometric regions represent the candidate locations where subsequent fault codes may occur within the current wafer. The boundary descriptions of these regions, such as a list of vertex coordinates constituting the region boundaries, are stored as prediction information.
[0093] Step S446: Assign an initial occurrence probability that is positively correlated with the confidence level of the association rule to each candidate location region, and store the candidate location region and its initial occurrence probability as prediction information for subsequent fault codes.
[0094] Specifically, first, obtain the final confidence value of the currently activated association rule, denoted as Conf. For each candidate location region generated in step S445, calculate an initial probability P. initial The simplest way is to let P... initial It equals Conf. Conf can also be fine-tuned based on the size of the region or its distance from the prediction center; for example, the closer the region is to the prediction center, the higher the probability. Once P is determined... initial Then, a data structure describing the geometry of the region (e.g., a list of polygon vertices) and P will be used. initialThe values are linked together to form a prediction record. All such prediction records are aggregated and stored in a temporary prediction database as the consequential failure prediction information generated by this rule's activation. This database is indexed by the rule identifier or activation event identifier. This prediction information will be continuously monitored during subsequent testing to verify the accuracy of the predictions and update the confidence level.
[0095] Step S450: Based on the active location and the predicted candidate location region, draw directional connection lines from the active location to the candidate location region in the wafer space grid, and assign the confidence of the association rule to each connection line as the initial confidence of the fault propagation path, thereby generating an initial version of the wafer-level fault evolution map containing at least one fault propagation path.
[0096] In practice, an empty data structure is first created to represent the wafer-level fault evolution graph. Essentially, it's a directed graph where nodes are specific locations on the wafer (which could be fault points or the center of a prediction region), and edges are weighted directed connections. Once a rule is activated and the activation center coordinates and several candidate location regions are calculated, operations begin on the graph. For each candidate location region, a new graph node is created, identified by its center coordinates, and stores its detailed boundaries and initial probability information. Then, a directed edge is created, starting at the node represented by the activation center coordinates (created if it doesn't already exist) and ending at the newly created candidate region node. The weight (initial confidence) of this edge is set to the confidence value used to activate the rule. This edge is then added to the graph. After performing this operation on all candidate regions, all fault propagation paths corresponding to that rule are added to the graph. As testing continues, more and more rules are activated, and more nodes and edges are added to the graph, forming an evolving initial version of the graph reflecting the overall wafer fault dynamics.
[0097] Step S460: Receive newly added real-time fault codes in the subsequent test data stream in real time. When a newly added real-time fault code matches a set of predicted consequent fault codes, perform cumulative enhancement processing on the confidence level of the fault propagation path corresponding to the association rule that activates the prediction. Based on the cumulative enhanced confidence level, select path nodes from all fault propagation paths whose cumulative confidence level exceeds the preset path confidence level threshold as candidate fault root source nodes.
[0098] In one implementation, step S460 may specifically include the following steps S461 to S466: Step S461: Continuously monitor the subsequent test data stream sent by the wafer testing machine, and parse the newly added real-time fault codes and their corresponding newly added real-time row and column coordinates from the subsequent test data stream.
[0099] Step S462: Compare the newly added real-time fault code with the set of consequent fault codes of all activated association rules one by one. If the newly added real-time fault code completely matches the set of consequent fault codes of a certain activated association rule, then the prediction of that association rule is verified.
[0100] In the specific execution process, a list is first maintained, containing all the associated rules that have been activated during the current wafer testing process, along with their respective sets of consequent fault codes and their corresponding fault propagation paths. When a new fault information (potentially containing one or more fault codes) is obtained from step S461, this list of activated rules is traversed. For each rule in the list, its consequent fault code set is compared with the newly added fault code set. The comparison condition is that the newly added fault code set must be exactly equal to the rule's consequent fault code set; there cannot be more or less. If the newly added fault code happens to be a superset of the consequent set but contains additional fault codes that do not belong to the consequent set, this is generally not considered a complete match unless there is special processing logic. If the complete match condition is met, the rule is marked as "verified," and the subsequent confidence enhancement process is triggered. After traversing all rules, if no match is found, no processing is performed, and the system continues to wait for the next newly added fault code.
[0101] Step S463: Perform confidence enhancement calculation on the newly added real-time fault codes that have been successfully matched. Assign different enhancement levels according to the order of the occurrence of the newly added real-time fault codes relative to the prediction time window, and generate dynamic enhancement values.
[0102] First, for the rule currently being validated, the activation time (or test number) recorded when it was activated needs to be obtained. Simultaneously, the prediction time window length corresponding to this rule needs to be obtained; this length can be determined based on the average time interval between the occurrence of the preceding and succeeding events in historical data. Then, the occurrence time of the newly added fault code needs to be obtained. The time difference from the activation time to the current validation time is calculated. This time difference is compared with the prediction time window. If the time difference is within the window, a higher enhancement factor is set, such as 1.2; if the time difference exceeds the window but is still within a certain maximum tolerance range, a lower enhancement factor is set, such as 0.8; if the time difference exceeds the maximum tolerance range, the enhancement factor can be set to 1, meaning no enhancement or reduction. Finally, the rule's base confidence (or a portion thereof) is multiplied by this enhancement factor to obtain the dynamic enhancement value contributed by this validation. This enhancement value is a positive number and will be used to update the cumulative confidence value of the path.
[0103] Step S464: For the verified association rule, locate the fault propagation path generated by the association rule in the wafer-level fault evolution map, increase the initial confidence of the connection line from the activation center coordinate to the candidate location region on the fault propagation path according to the dynamic enhancement value, and update the cumulative confidence value of the fault propagation path.
[0104] When a rule is activated and a propagation path is generated, a mapping relationship should be established to quickly locate the graph edges it generates based on the rule identifier and activation time. After step S462 determines that a rule has been validated, this mapping relationship is used to find the one or more fault propagation paths generated by that rule. For each found path (i.e., a directed edge in the graph), its current cumulative confidence value is obtained. Then, the dynamic enhancement value calculated in step S463 is added to the current cumulative value to obtain a new cumulative value. This new cumulative value replaces the original confidence attribute of that edge in the graph. In this way, the confidence of this path is increased because it has been validated. If the same rule has multiple different prediction regions, and only the predictions of some regions are validated, then only the confidence of the edges pointing to the validated regions will be enhanced.
[0105] Step S465: Traverse all fault propagation paths in the wafer-level fault evolution map, compare the cumulative confidence value of each fault propagation path with the preset path confidence threshold, and select fault propagation paths whose cumulative confidence value is greater than or equal to the preset path confidence threshold.
[0106] First, retrieve a list of all directed edges from the data structure storing the wafer-level fault evolution map. Initialize an empty list to store the selected high-confidence paths. Iterate through each directed edge and read its current cumulative confidence value. Compare this value with a preset path confidence threshold. If the cumulative value is greater than or equal to the threshold, add the edge's information (including the starting node, ending node, and cumulative value) to the high-confidence path list. If the cumulative value is less than the threshold, ignore it. After iteration, the high-confidence path list contains all fault propagation paths that have reached the threshold and are worth further analysis.
[0107] Step S466: Extract the starting point of the selected fault propagation path as the candidate fault root source location, and parse the test item identifier contained in the antecedent fault code set associated with these fault propagation paths. Combine the candidate fault root source location and the corresponding test item identifier into a fault root source candidate node.
[0108] For each high-confidence directed edge selected in step S465, the starting node of this edge is first obtained. The starting node stores the row and column coordinates of that location, which are the candidate fault root cause locations. Then, using the metadata carried by this edge, or through the previously established mapping relationship between rules and edges, the association rule that generated this edge is found. From this rule, its set of antecedent fault codes is parsed. Each fault code identifier in this set corresponds to a specific test item. These test item identifiers are organized into a list. Finally, a data structure is created to encapsulate the row and column coordinates of the candidate fault root cause locations and the corresponding list of test item identifiers together, forming a complete fault root cause candidate node. All these nodes generated by the selected paths together constitute the set of the most noteworthy potential fault root cause points on the current wafer.
[0109] Step S500: Based on the candidate nodes of the fault root cause in the wafer-level fault evolution map and their associated cumulative confidence values, perform priority sorting processing, generate a fault root cause location instruction containing the coordinates of the die to be retested first and the identifier of the associated test item, and send the fault root cause location instruction to the wafer test machine to trigger the retest operation for the target die.
[0110] In one implementation, step S500 may specifically include the following steps S510-S560: Step S510: Extract all candidate nodes of the root cause of the fault from the wafer-level fault evolution map. Each candidate node of the root cause of the fault contains the row and column coordinates of the corresponding die in the wafer and the identifier of at least one associated test item, and obtain the cumulative confidence value of the fault propagation path to which each candidate node of the root cause of the fault belongs.
[0111] First, iterate through all nodes marked as "fault root cause candidates" in the wafer-level fault evolution graph. For each such node, read its stored row and column coordinates and the list of associated test item identifiers. Then, find all directed edges originating from this node in the graph. Collect the cumulative confidence values of all these edges. Calculate the overall confidence score for the node according to a preset strategy, for example, taking the maximum value among these cumulative values. Associate this overall confidence score with the node. Finally, summarize all extracted nodes, their coordinates, test item lists, and overall confidence scores into a single list, which serves as input for subsequent priority ranking.
[0112] Step S520: Perform local clustering analysis based on physical distance on the extracted candidate nodes of the root cause of the fault, identify the node clusters that are less than the preset clustering radius in the wafer space grid, and assign a cluster identifier to each node cluster.
[0113] First, obtain a list of row and column coordinates for all candidate root cause nodes. Apply a density-based spatial clustering algorithm, such as DBSCAN, using these coordinates as input. Set the neighborhood radius of the clustering algorithm to a preset cluster radius and set the minimum number of points within a cluster to a small value, such as one or two. After the algorithm runs, it outputs the cluster number to which each point belongs, as well as which points are isolated noise points. For nodes that are grouped into the same cluster, assign them a new, unique cluster identifier. For isolated nodes that are not clustered, they either form their own cluster or remain independent individuals. Finally, each candidate root cause node will be marked with its cluster identifier or as an isolated point.
[0114] Step S530: Sort the cumulative confidence values of all candidate nodes for root cause of failure in descending order to generate a priority list of candidate nodes for root cause of failure.
[0115] Sort the cumulative confidence values of all candidate root cause nodes in descending order. This means arranging the entire candidate node list in descending order based on the comprehensive confidence score of each node calculated in step S510. A priority list of candidate root cause nodes is generated. This list is the final ranking result used to guide retesting decisions. The node at the top is currently considered the most likely root cause and should be addressed first.
[0116] Step S540: Analyze the preset number of candidate fault root cause nodes ranked first in the priority list of candidate fault root cause nodes, and extract the row and column coordinates of these nodes as the coordinates of the grains to be retested first.
[0117] In one implementation, step S540 may specifically include the following steps S541 to S546: Step S541: Read the priority list of candidate nodes for root cause of the fault, and obtain the candidate node with the highest cumulative confidence value at the top of the list as the first candidate node.
[0118] In the specific execution process, the element with index zero, i.e., the first node in the priority list generated in step S530, is directly obtained. All attributes of this node are read, especially its row and column coordinates and overall confidence score. This node is marked as the first candidate node, and a local area to be retested is planned starting from it.
[0119] Step S542: Using the row and column coordinates of the first candidate node as the center, define a local region within the wafer space grid that includes the first candidate node and its surrounding adjacent row and column coordinates.
[0120] Specifically, obtain the row and column coordinates of the first candidate node, denoted as (row, col). Set a radius parameter for a local region, for example, a radius r equal to 2, meaning that in the row direction, from row minus r to row plus r, and in the column direction, from (col-r, col+r), a square grid region with a side length of 2r+1 is formed. All row and column coordinate points within this region, that is, every integer coordinate point in the range from (row-r, col-r) to (row+r, col+r), are defined as this local region. Record this set of coordinate points as the initial set to be filtered.
[0121] Step S543: Filter all row and column coordinates within the defined local area based on wafer map availability markers, excluding grain coordinates that have been marked as permanently failed or have undergone retesting, and generate a set of valid coordinates after filtering.
[0122] Specifically, a dynamic data structure maintaining the current wafer state is first accessed. This structure records various markers for each memory cell on the wafer, such as "permanently failed" and "retested" markers. Each row and column coordinate within the local region defined in step S542 is traversed. For each coordinate, its corresponding marker is searched in the wafer state data. If the coordinate has no marker prohibiting retesting (i.e., it is neither permanently failed nor marked as retested), the coordinate is added to a new list called the valid coordinate set. If the coordinate has any marker prohibiting retesting, it is ignored. After traversing all coordinates within the region, the resulting valid coordinate set represents the candidate points for retesting.
[0123] Step S544: Check whether the valid coordinate set contains other fault root cause candidate nodes that are ranked higher in the priority list. If it does, merge the row and column coordinates of all fault root cause candidate nodes in the valid coordinate set into a retest area coordinate set.
[0124] Specifically, first, obtain the nodes ranked higher in the priority list besides the first candidate node, such as the nodes ranked second to fifth. For each such node, obtain its row and column coordinates and check whether these coordinates are already included in the valid coordinate set generated in step S543. If the coordinates of a node exist in the valid coordinate set, mark the node as "covered". After traversing all the nodes to be checked, merge all the coordinates in the valid coordinate set and the coordinates of all nodes marked as "covered" into a new, unique coordinate list. This new list is the retest area coordinate set, which represents one or more areas around the first candidate node that are worth retesting.
[0125] Step S545: Calculate the minimum bounding rectangle of the coordinate set of the retest area, and use the coordinates of the four vertices of the minimum bounding rectangle as the boundary description of the coordinates of the grains to be retested first.
[0126] Specifically, obtain the row coordinate list and column coordinate list of all points in the coordinate set of the re-measured area. Find the minimum value (min) in the row coordinate list. row and maximum value max row and the minimum value min in the column coordinate list. col and maximum value max col Therefore, the coordinates of the four vertices of the minimum bounding rectangle are: top left corner (min... row ,min col ), top right corner (min row ,max col ), bottom right corner (max) row ,max col ), bottom left corner (max) row ,min col These four vertex coordinates completely define a rectangular region that can cover all the points to be remeasured. This rectangular region is used as the boundary description of the grain coordinates to be preferentially remeasured.
[0127] Step S546: Sort all row and column coordinate points in the retest area coordinate set according to their physical scan paths to generate a continuous retest path coordinate sequence as the final list of grain coordinates to be retested first.
[0128] First, determine the default scanning path mode of the probe station, such as a bow-shaped scan, which involves scanning the first row from left to right, then moving to the next row and scanning from right to left, and so on. Using the minimum bounding rectangle determined in step S545 as the boundary, generate all the grain coordinate points within the rectangle that need to be retested (these points come from the retest area coordinate set). Then, sort these points according to the scanning path mode. For example, starting from the first row of the rectangle, extract the points in that row that exist in the retest area coordinate set from left to right; then move to the second row, and if the scanning direction is reversed, extract the points in that row that exist in the set from right to left. Continue in this manner until all rows within the rectangle have been processed. The resulting ordered list is a continuous and efficient retesting path. Use this path list as the final list of grain coordinates to be prioritized for retesting.
[0129] Step S550: Based on the extracted coordinates of the grains to be retested, trace back the associated test item identifiers and bind each coordinate of the grains to be retested with the corresponding test item identifier to form a set of tuples for the retested tasks.
[0130] First, create an empty list to store the tuples of tasks to be retested. Iterate through each coordinate point in the retest path coordinate list generated in step S546. For the current coordinate point, find all previously stored records that have that coordinate as the root candidate node. Extract the associated test item identifiers from these records, and merge and deduplicate all item identifiers to form a list of retest items for that coordinate point. Then, create a tuple containing the row and column coordinates of that coordinate point, as well as the merged list of retest items. Add this tuple to the list of the task tuple set. After traversing all coordinate points, the resulting list is the complete set of task tuples to be retested.
[0131] Step S560: Encapsulate the set of task tuples to be retested into a data frame format that conforms to the wafer test instrument communication protocol, and generate a fault root cause location instruction containing the coordinates of the die to be retested first and the identifier of the associated test item.
[0132] For example, the communication protocol details for receiving external retest commands are determined according to the technical manual of the wafer testing machine. For instance, the protocol might specify that the command begins with a specific byte, followed by a command type byte (indicating this is a retest command), then the number of dies to be retested, and then, for each die, sequentially sending its row coordinate high octets, row coordinate low octets, column coordinate high octets, column coordinate low octets, and the number of retest items, followed by the code for each item. Based on this format, all information from the set of retest task tuples generated in step S550 is filled into the corresponding fields of the data frame. Finally, the checksum of the entire data frame is calculated and appended to the end of the frame. The assembled byte array is the final fault root cause location command. This command is sent to the wafer testing machine through the established communication interface, and the machine, after parsing it, will trigger the corresponding retest operation.
[0133] In one implementation, after step S500, the following steps S600~S1100 may also be included: Step S600: Receive the retest result data stream returned by the wafer test equipment after performing retest on the target die in response to the fault root cause location command. The retest result data stream contains the retest electrical parameters corresponding to the associated test item identifier on the target die.
[0134] After sending the fault root cause location command, the system enters a waiting state. Once the testing equipment completes the retest of the specified die, it will package the retest results into data packets and send them back according to the agreed protocol. The system receives these data packets and parses them according to the protocol. From the parsed data, the row and column coordinates of each retested die, as well as the electrical parameter values measured for the test items specified in the command, are extracted. This information is then summarized to form a retest result data stream for subsequent analysis.
[0135] Step S700: Perform data integrity verification on the received retest result data stream, verify whether the returned retest electrical parameters cover all associated test item identifiers in the fault root cause location instruction, and generate retest result data that has passed the verification.
[0136] First, create a mapping table with each task tuple (grain coordinates, test item) in the instruction as the key, initializing the expected data receiving status to false. Then, iterate through each record in the received retest result data stream. For each record, search the mapping table for the corresponding entry based on its row and column coordinates and test item identifier. If found, mark the entry's status as true and record the measurement value. After iterating through all retest data, check the status of all entries in the mapping table. If all entries are true, the data is complete, and all collected retest data is organized into a new data structure as the validated retest result data. If any entries are false, the data is incomplete, and an error handling process may need to be triggered.
[0137] Step S800: Analyze the verified retest result data, extract the retest electrical parameters and compare them with the original electrical parameters of the same test items for the corresponding grains in the original electrical test data, and calculate the degree of difference between the retest electrical parameters and the original electrical parameters.
[0138] Specifically, for each data point in the verified retest results data, that is, for a certain test item of a certain grain, its retest value V is obtained. retest Then, from the initially stored set of raw electrical test data, the original measured value V is retrieved based on the same grain coordinates and test item identifier. original Next, calculate the degree of difference, D, between these two values. A common way to calculate it is that D equals the absolute value (V). retest -V original Another way is to calculate the relative difference, where D equals the absolute value (V). retest -V original Divide by V original The absolute value of the difference. The calculated degree of difference D is then correlated with this data point for use in the next step of threshold comparison.
[0139] Step S900: Compare the degree of difference with the preset retest verification threshold and generate a comparison result identifier.
[0140] In the specific execution process, for each degree of difference D calculated in step S800, a pre-set retest verification threshold T is obtained. D is compared with T. If D is less than or equal to T, a result marked "verification passed" or "valid" is generated. If D is greater than T, a result marked "verification failed" or "false alarm" is generated. This comparison result is associated with the corresponding grain coordinates and test items for subsequent processing.
[0141] Step S1000: If the comparison result indicates that the degree of difference exceeds the preset retest verification threshold, the original fault detection result is determined to be valid, and the cumulative confidence value of the candidate node of the fault root cause is permanently marked in the wafer-level fault evolution map.
[0142] When the comparison result in step S900 is marked as a "false alarm," the processing flow for this step begins. First, the wafer-level fault evolution map is located, and the node corresponding to the grain coordinates of the verified false alarm and marked as a candidate fault root cause is found. In the node's data structure, the "Is it solidified?" attribute is set to "Yes" or "True." Simultaneously, the node's current cumulative confidence value is written into a "Solidified Confidence" field, indicating that this value is practically verified and reliable. Furthermore, this node can be highlighted with a prominent color in the map display to distinguish it from other unverified candidate nodes.
[0143] Step S1100: If the comparison result indicates that the degree of difference does not exceed the preset retest verification threshold, the original fault detection result is determined to be a false alarm. Based on the location information of the candidate node of the fault root cause, all fault propagation paths in the wafer-level fault evolution map that pass through the node are traced in reverse. The cumulative confidence values of these fault propagation paths are globally attenuated.
[0144] In one implementation, step S1100 may specifically include the following steps S1110 to S1160: Step S1110: Starting from the row and column coordinates of the candidate fault root cause node that is determined to be a false alarm, perform a reverse breadth-first search operation in the wafer-level fault evolution map to identify all directed connections that end at or pass through the node.
[0145] Reverse breadth-first search (RBFS) is a graph traversal algorithm that reverses the direction of breadth-first search. Starting from a given starting node, it traverses the graph along the reverse direction of directed edges (from the endpoint of an edge to the starting node), finding all other nodes and edges that can reach the starting node. An edge that directly points to the starting node is considered an endpoint. Edges along the path from a more distant node, passing through the starting node, and then pointing to other nodes are considered paths. By performing a reverse search, the confidence information of all nodes in the entire graph that depend on the false positive node can be completely identified.
[0146] Specifically, first find the graph node corresponding to the node identified as a false alarm in the wafer-level fault evolution graph, denoted as N. false Then, initialize a queue and set N false Add the node to the queue. Simultaneously, initialize an empty set to store all identified directed connections. Begin a loop; while the queue is not empty, remove a node N from the head of the queue. current Find all instances of N in the graph. current A directed edge that terminates at N (i.e., points to N) current For each found edge E, add it to the set of directed connections. Also, obtain the starting node N of this edge E. start If N start If it has not been visited, then N start Add to the tail of the queue so that the reverse search can continue from N. start The starting edge. After the loop ends, the set of directed connections contains all paths that can reach N via the reverse path. false The edges that directly or indirectly point to N are... false , or from N false An edge that originates from N but is traced in the reverse direction (actually from N). false The upstream points to N false (The edges). This set is the target set of edges that need to be decayed.
[0147] Step S1120: Dynamically adjust the attenuation level of the searched directed connection lines based on the propagation distance. Assign a greater attenuation magnitude to the connection lines that are farther away from the starting point, and generate a personalized attenuation level for each connection line.
[0148] First, determine the node N that was falsely reported. false The propagation distance to the origin (or edge itself) of each identified directed connection E. This distance can be defined as the distance from N in the reverse breadth-first search tree. false The number of edges traversed to reach the starting point E. For example, a direct path to N. falseEdges pointing to each other have a distance of 1. Edges pointing to these edges have a distance of 2, and so on. Then, a decay function is defined. This function takes the distance d as input and outputs a decay coefficient α, where α is a subtractive function of d; that is, the larger d is, the smaller α is. For example, α can be defined as 1 divided by d, or α can be defined as the reciprocal of 1 raised to the power of d. For each edge, the decay coefficient α is calculated using the decay function based on its distance d. This α value is then multiplied by the current confidence level of the edge to obtain the updated confidence level. The decay degree is 1 minus α.
[0149] Step S1120: Extract the fault propagation path corresponding to each identified directed connection line, and record the complete node sequence and connection line confidence of each fault propagation path.
[0150] For each directed edge E identified in step S1110, path tracing is performed starting from the origin of E and proceeding in the forward direction (i.e., from the origin to the destination). Since these edges are found through reverse search, N can be reached by following the forward direction. false During the tracing process, every node and edge visited, along with the current cumulative confidence value of each edge, is recorded. Ultimately, a path is obtained starting from point E, passing through a series of intermediate nodes, and finally reaching point N. false The node sequence and the corresponding edge confidence list are stored. This path information is used for subsequent confidence updates.
[0151] Step S1140: For each fault propagation path, starting from the starting node and proceeding along the node sequence direction, multiply the current confidence cumulative value of each connection line by its corresponding personalized attenuation degree to generate an updated connection line confidence cumulative value.
[0152] Specifically, for each path recorded in step S1130, its node sequence and edge confidence list are obtained. For the first edge on the path (i.e., the upstream edge), the personalized decay level α1 corresponding to that edge is obtained from the result of step S1120. The current cumulative confidence value C1 of this edge is multiplied by α1 to obtain the new confidence C1. new Then, for the second edge on the path, obtain its decay level α2, and multiply its current value C2 by α2 to get C2. new This process continues until the last edge on the path has been processed. Ultimately, all edges on the path receive updated cumulative confidence values.
[0153] Step S1150: Based on the updated cumulative confidence values of the connecting lines, recalculate the cumulative confidence values of all nodes on these fault propagation paths. The cumulative confidence value of a node is the sum of the cumulative confidence values of all incoming connecting lines pointing to that node.
[0154] For all nodes on all paths involved in step S1130, their confidence scores need to be recalculated. First, determine the set of these nodes. Then, for each node N in the set, find all directed edges (incoming edges) ending at N in the graph. Sum the updated cumulative confidence scores of all these incoming edges. This sum is the new cumulative confidence score of node N. If a node has no incoming edges, its confidence score can be defined as zero or remain unchanged. Update the node's attributes with the calculated new value, replacing the old confidence score value.
[0155] Step S1160: Write the updated cumulative confidence values of the connection lines and nodes back to the wafer-level fault evolution map, replacing the original confidence information, and complete the global attenuation processing of the false alarm associated paths.
[0156] In the specific execution process, the new cumulative confidence value of each directed edge calculated in step S1140 is obtained, the corresponding edge is located in the graph data structure, and its confidence attribute is updated to the new value. Similarly, the new cumulative confidence value of each affected node calculated in step S1150 is obtained, the corresponding node is located in the graph, and its confidence attribute is updated. After all updates are completed, the wafer-level fault evolution graph reflects the latest confidence state after false alarm verification and correction. Subsequently, any subsequent analysis based on this graph, such as re-screening candidate nodes for fault root causes, will use these corrected and more accurate confidence values.
[0157] Please see Figure 2 , Figure 2 This is a schematic diagram of a computer device provided in an embodiment of the present invention. The computer device includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer device, capable of parsing various instructions and processing various data within the computer device. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer device. The memory 103 is a storage device in the computer device used to store programs and data. It is understood that the memory 103 here can include the computer device's built-in memory, or it can include extended memory supported by the computer device. The memory 103 provides storage space, which stores the computer device's operating system; this invention does not limit the storage space.
[0158] In one embodiment, the processor 101 executes the DDR wafer test fault detection method based on association rules provided in the above embodiments of the present invention by running a computer program in the memory 103.
Claims
1. A method for fault detection in DDR wafer testing based on association rules, characterized in that, The method includes: Obtain the original electrical test data set of the DDR wafer test items generated after the electrical parameter measurement operation of each memory cell of the wafer is performed by the probe station during the wafer testing phase; The original electrical test data set is processed to mine the correlation between test items, and a fault code sequence with spatial location correlation within the wafer is obtained, which includes the co-occurrence relationship of fault codes between different test items and the order of occurrence of fault codes. The fault code sequence is subjected to frequent itemset iterative search and strong association rule generation based on minimum support threshold and minimum confidence threshold to obtain an initial fault association rule set associated with spatial distribution; Dynamic fault reasoning is performed based on the initial fault association rule set. The real-time fault codes in the real-time acquired DDR wafer online test data stream are matched with the initial fault association rule set and confidence accumulation is calculated to generate a wafer-level fault evolution map containing fault propagation paths and candidate nodes of fault root causes. Based on the candidate nodes of the fault root cause in the wafer-level fault evolution map and their associated cumulative confidence values, priority sorting is performed to generate a fault root cause location instruction containing the coordinates of the die to be retested and the identifier of the associated test item. The fault root cause location instruction is then sent to the wafer test instrument to trigger the retest operation for the target die.
2. The method according to claim 1, characterized in that, The process of mining the correlation between test items in the original electrical test data set yields a fault code sequence with spatial location correlation within the wafer, containing the co-occurrence relationship of fault codes among different test items and the temporal order of fault codes. The original electrical test data set is subjected to wafer map mapping processing. Based on the physical coordinates of each storage cell in the wafer, the test result data of the test item is projected onto the wafer space grid to generate a wafer test data spatial distribution matrix with row and column coordinates. The test result data of each storage unit in the wafer test data spatial distribution matrix is analyzed, and the fault code identifier and the row and column coordinate information of the storage unit in the wafer space grid are extracted for the test result data that exceeds the preset specification limit. The extracted fault code identifiers and row and column coordinate information are subjected to noise filtering based on probe contact status to remove isolated fault records caused by momentary poor contact, and a cleaned list of fault code identifiers and corresponding row and column coordinates is generated. Based on the cleaned list of fault code identifiers and their corresponding row and column coordinates, an initial fault code transaction set is constructed with a single storage unit as the basic transaction unit. Each transaction record in the initial fault code transaction set contains at least one fault code identifier belonging to the same row and column coordinates. The transaction records in the initial fault code transaction set are sorted according to the physical scanning order of the storage units within the wafer to generate a transaction sequence with time order attributes and spatial adjacency relationships. The fault code identifiers of adjacent transaction records in the transaction sequence are analyzed for sequential correlation to identify fault code combination patterns that appear sequentially on the same wafer or adjacent die locations. Based on the fault code combination patterns, a fault code sequence with spatial location correlation within the wafer is generated, which includes the co-occurrence relationship of fault codes between different test items and the order of occurrence of fault codes.
3. The method according to claim 2, characterized in that, The step involves performing a sequential correlation analysis on the fault code identifiers of adjacent transaction records in the transaction sequence to identify fault code combination patterns that appear sequentially at the same wafer or adjacent die locations. Based on these fault code combination patterns, a fault code sequence with intra-wafer spatial location correlation is generated, containing co-occurrence relationships of fault codes between different test items and the temporal order of fault code appearance. This includes: Using a single transaction record as a sliding window unit, a sliding window scanning operation with a preset step size is performed on the transaction sequence to extract the fault code identifiers of at least two transaction records contained in the window and the row and column coordinates of the corresponding transaction records; Local spatial density clustering is performed on the fault code identifiers and their row and column coordinates extracted by the sliding window. Multiple fault code identifiers with a spatial distance smaller than the preset clustering radius are merged into composite fault event identifiers to generate a clustered fault code identifier set. Calculate the Manhattan distance in the spatial dimension of different fault code identifiers in the clustered fault code identifier set, and determine whether fault code identifiers belonging to the same window come from spatially adjacent storage units based on whether the Manhattan distance is less than or equal to a preset adjacency distance threshold. The fault code identifiers determined to be from spatially adjacent storage units are subjected to time-order labeling. Each fault code identifier is appended with the order number of the transaction record in the transaction sequence, generating a set of candidate fault code combinations with spatial adjacency constraints and time order. The frequency of occurrence of each fault code combination in the candidate fault code combination set is counted, and fault code combinations whose occurrence frequency exceeds a preset frequency threshold are selected as high-frequency adjacent fault code combination patterns. The occurrence order of fault code identifiers in the high-frequency adjacent fault code combination mode is analyzed, the fixed order relationship between the fault code identifiers that appear first and the fault code identifiers that appear later is extracted, and a fault code sequence with spatial location association within the wafer is generated based on the fixed order relationship, which includes the co-occurrence relationship of fault codes between different test items and the temporal occurrence order of fault codes.
4. The method according to claim 1, characterized in that, The process of performing frequent itemset iterative search and strong association rule generation on the fault code sequence based on minimum support threshold and minimum confidence threshold yields an initial set of fault association rules associated with the spatial distribution, including: The fault code sequence is used as the input transaction database for association rule mining. The input transaction database is scanned once to count the frequency of occurrence of each fault code identifier as a single item. Fault code identifiers whose frequency of occurrence meets the preset minimum support threshold are selected as frequent item sets. The frequency of each fault code identifier in the frequent one-item set is normalized based on the wafer region distribution density to eliminate the influence of the difference in test density between the wafer center region and the edge region on the support calculation, and a normalized frequent one-item set is generated. Based on the normalized frequent itemsets, recursive join and pruning operations are performed. By joining two frequent itemsets containing k-1 items, a candidate k-itemset is generated. The support count of the candidate k-itemset in the input transaction database is calculated. Candidate k-itemsets whose support count meets the preset minimum support threshold are selected as frequent k-itemsets until no new frequent itemsets can be generated. Extract all non-empty subsets from the generated frequent k-items set. For each frequent itemset and its non-empty subset, calculate the confidence of the association rule derived from the non-empty subset. The confidence of the association rule is the ratio of the support count of the frequent itemset to the support count of its non-empty subset. Filter association rules whose confidence level is greater than or equal to a preset minimum confidence threshold, extract the antecedent fault code set and consequent fault code set of the association rule, and record the spatial location distribution information of the antecedent fault code set and consequent fault code set in the fault code sequence; Based on the set of preceding fault codes, the set of succeeding fault codes, and the spatial location distribution area information, an initial fault association rule set associated with spatial distribution is constructed, with association rules as nodes and spatial location distribution area as node attributes.
5. The method according to claim 4, characterized in that, The step of extracting all non-empty subsets from the generated frequent k-itemsset, and for each frequent itemset and its non-empty subsets, calculating the confidence of the association rule derived from the non-empty subsets, includes: For each frequent k-itemset, list all fault code identifiers contained in the frequent k-itemset as elements of the universal set, and generate all non-empty proper subsets consisting of the elements of the universal set, excluding the empty set and the universal set itself, through a recursive backtracking algorithm. Redundancy checks are performed on each generated non-empty true subset, and low-association subsets that differ from the overall support count of frequent itemsets by more than a preset difference threshold are removed, thus obtaining the filtered effective non-empty true subsets. For each valid non-empty true subset, the valid non-empty true subset is used as the set of antecedent fault codes for the association rule, and the remaining part of the frequent k-item set after removing the valid non-empty true subset is used as the set of consequent fault codes for the association rule. The number of transaction records that appear simultaneously in the fault code sequence as a whole, representing the previous support count, is calculated. Divide the number of transaction records in which frequent k-itemsets appear simultaneously in the fault code sequence by the count of the predecessor support to obtain the preliminary confidence value of the association rule derived from the effective non-empty true subset. For each valid non-empty true subset, the initial confidence value of the association rule is subjected to confidence decay correction processing. The degree of decay is positively correlated with the average distance between the antecedent fault code set and the consequent fault code set in spatial distribution. The final confidence value of the association rule is determined based on the corrected initial confidence value of the association rule.
6. The method according to claim 4, characterized in that, The process involves filtering association rules whose confidence level is greater than or equal to a preset minimum confidence threshold, extracting the antecedent fault code set and consequent fault code set of the association rule, and recording the spatial distribution information of the antecedent and consequent fault code sets in the fault code sequence, including: The confidence score of each association rule is calculated and compared with a preset confidence threshold. Association rules with confidence scores greater than or equal to the preset confidence threshold are retained as candidate strong association rules. The retained candidate strongly correlated rules are clustered and merged based on spatial distribution overlap. Rules with highly overlapping spatial distribution areas of the antecedent fault code set and the consequent fault code set are merged into rule clusters, generating a merged set of candidate strongly correlated rules. For each candidate strongly associated rule in the merged set of candidate strongly associated rules, parse at least one fault code identifier contained in its predecessor fault code set, and backtrack the set of row and column coordinates corresponding to the transaction records in the fault code sequence that simultaneously contain all fault code identifiers in the predecessor fault code set. Perform convex hull calculation on the set of row and column coordinates to generate the smallest convex polygon that surrounds all occurrence positions of the set of previous fault codes, which serves as the boundary of the spatial distribution region of the set of previous fault codes. Parse at least one fault code identifier contained in the consequent fault code set of each candidate strong association rule, and backtrack the set of row and column coordinates corresponding to the transaction record in the fault code sequence that contains any fault code identifier in the consequent fault code set; Density clustering analysis is performed on the row and column coordinate sets corresponding to the set of subsequent fault codes to identify at least one high-density clustering region, and the spatial distribution information of the set of subsequent fault codes is described by the center coordinates and coverage radius of each high-density clustering region.
7. The method according to claim 1, characterized in that, The step of performing dynamic fault reasoning based on the initial fault association rule set involves matching and accumulating confidence scores between real-time fault codes in the real-time acquired DDR wafer online test data stream and the initial fault association rule set to generate a wafer-level fault evolution map containing fault propagation paths and candidate fault root cause nodes, including: Receive the DDR wafer online test data stream generated by the wafer test equipment for the current wafer during real-time testing, and extract the real-time fault code generated by each tested memory cell on the current wafer and its corresponding real-time row and column coordinates from the online test data stream; The extracted real-time fault codes and real-time row and column coordinates are assembled into a real-time fault code stream according to the order in which the tests are completed. The real-time fault code stream includes the fault code identifier, its real-time location information within the wafer, and a real-time timestamp. Traverse each association rule in the initial fault association rule set, and match the fault code identifiers appearing in the real-time fault code stream with the preceding fault code set of the association rule. If the fault code identifiers appearing in the real-time fault code stream completely contain the preceding fault code set of a certain association rule, then activate that association rule. For an activated association rule, the activation position of the association rule is determined based on the real-time location information of the fault code identifier in the previous fault code set in the real-time fault code stream. Based on the subsequent fault code set, the candidate location area where the subsequent fault code may appear is predicted in the area that has not yet appeared in the real-time fault code stream or in the area that may appear in subsequent tests. Based on the activated location and the predicted candidate location region, directional connection lines are drawn in the wafer space grid from the activated location to the candidate location region, and each connection line is assigned the confidence of the association rule as the initial confidence of the fault propagation path, generating an initial version of the wafer-level fault evolution map containing at least one fault propagation path. When a new real-time fault code is received in the test data stream in real time, and the new real-time fault code matches a set of predicted consequent fault codes, the confidence level on the fault propagation path corresponding to the association rule that activates the prediction is cumulatively enhanced. Based on the cumulatively enhanced confidence level, path nodes whose cumulative confidence level exceeds the preset path confidence level threshold are selected from all fault propagation paths as candidate nodes for the root cause of the fault.
8. The method according to claim 7, characterized in that, For an activated association rule, the activation position of the association rule is determined based on the real-time location information of the fault code identifier in the preceding fault code set appearing in the real-time fault code stream. Furthermore, based on the following fault code set, candidate location regions where the following fault codes may appear are predicted in areas not yet appearing in the real-time fault code stream or in areas that may appear in subsequent tests. These regions include: Parse at least one fault code identifier contained in the antecedent fault code set of the activated association rule, extract the real-time row and column coordinates of all storage cells containing these fault code identifiers from the real-time fault code stream, and calculate the geometric center point of these real-time row and column coordinates as the activation center coordinates of the association rule. Extract the spatial distribution area information of the subsequent fault code set from the spatial location distribution area information of the association rule. The spatial distribution area information includes the center coordinates and coverage radius of at least one high-density clustering area of the subsequent fault code set in the historical wafer test data. The spatial distribution information of the extracted fault code set is subjected to centroid drift correction processing. The center coordinates of the predicted area are dynamically adjusted according to the boundary of the area that has been tested on the current wafer, and the corrected predicted center coordinates are generated. Using the active center coordinates as a reference point, calculate the offset vector between the reference point and each corrected predicted center coordinate, and superimpose the offset vector onto the spatial coordinate system of the current wafer to obtain the predicted center coordinates where the subsequent fault codes may occur in the current wafer. Using the predicted center coordinates as the center and the coverage radius of the historical high-density cluster area as the reference radius, the outer expansion correction process is performed in combination with the boundary of the untested area of the current wafer to generate the candidate location area where the subsequent fault code may appear in the current wafer. Each candidate location region is assigned an initial occurrence probability that is positively correlated with the confidence level of the association rule, and the candidate location region and its initial occurrence probability are stored as prediction information for subsequent fault codes.
9. The method according to claim 7, characterized in that, When a newly added real-time fault code in the test data stream after real-time reception matches a predicted consequent fault code set, the confidence level on the fault propagation path corresponding to the association rule that activated the prediction is cumulatively enhanced. Based on the enhanced confidence level, path nodes whose cumulative confidence level exceeds a preset path confidence threshold are selected from all fault propagation paths as candidate fault root source nodes, including: Continuously monitor the subsequent test data stream sent by the wafer testing machine, and parse the newly added real-time fault codes and their corresponding newly added real-time row and column coordinates from the subsequent test data stream; The newly added real-time fault code is compared one by one with the set of consequent fault codes of all activated association rules. If the newly added real-time fault code completely matches the set of consequent fault codes of a certain activated association rule, then the prediction of that association rule is verified. For newly added real-time fault codes that have been successfully matched, a dynamic confidence enhancement calculation is performed. Different enhancement levels are assigned according to the order of the occurrence of the newly added real-time fault codes relative to the prediction time window, and dynamic enhancement values are generated. For the verified association rule, locate the fault propagation path generated by the association rule in the wafer-level fault evolution map, increase the initial confidence of the connection line from the activation center coordinate to the candidate location region on the fault propagation path according to the dynamic enhancement value, and update the cumulative confidence value of the fault propagation path. Traverse all fault propagation paths in the wafer-level fault evolution graph, compare the cumulative confidence value of each fault propagation path with the preset path confidence threshold, and filter out fault propagation paths whose cumulative confidence value is greater than or equal to the preset path confidence threshold. The starting point of the selected fault propagation path is extracted as the candidate fault root source location, and the test item identifier contained in the antecedent fault code set associated with these fault propagation paths is parsed. The candidate fault root source location and the corresponding test item identifier are combined into a fault root source candidate node.
10. A computer device, characterized in that, include: A memory, wherein a computer program is stored; A processor for loading the computer program to implement the DDR wafer test fault detection method based on association rules as described in any one of claims 1-9.
Citation Information
Patent Citations
Aero-engine fault diagnosis method and system based on extension association rule mining
CN113761032A
Software fault repair method and system fused with intelligent analysis
CN121326627A