Probe test yield fluctuation multi-dimensional data correlation analysis method
By constructing a multi-dimensional data correlation analysis method for probe test yield fluctuations, the problem of not being able to identify the root cause of probe test yield fluctuations in existing technologies has been solved. This enables in-depth tracing and accurate identification of the source of fluctuations, thereby improving the stability and efficiency of the production process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WEINAN MUWANG INTELLIGENT TECH CO LTD
- Filing Date
- 2026-06-11
- Publication Date
- 2026-07-14
AI Technical Summary
Existing technologies cannot effectively identify the true root cause of fluctuations in probe test yield, leading to repeated yield fluctuations during the production process, which affects product qualification rate and cost.
By constructing a multi-dimensional data correlation analysis method for probe test yield fluctuations, including obtaining the evolution of test batches, constructing a set of preceding trigger scenarios, extracting and sorting events, generating event perturbation transmission sequences and preliminary classification results, and finally constructing a root cause localization model to identify abnormal root causes.
It significantly improves the depth and accuracy of tracing the source of fluctuations, can automatically distinguish different types of fluctuation sources, provides verifiable causal logic support, and enhances the credibility and operability of the analysis conclusions.
Smart Images

Figure CN122388980A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data correlation analysis technology, specifically to a multidimensional data correlation analysis method for probe test yield fluctuations. Background Technology
[0002] In semiconductor manufacturing, probe testing is a critical step in ensuring the electrical performance and reliability of chips. This test involves direct contact between precision probes and chip pads to apply signals and acquire responses. However, probe testing yields during production often experience unexpected fluctuations, which directly impact the final product's pass rate and production costs. Therefore, developing an analytical method that can efficiently and accurately pinpoint the root cause of these fluctuations is essential for ensuring stable operation of test lines and improving overall productivity.
[0003] A common traditional analysis method is based on the frequency statistics of anomalies in a single batch. When a yield decline is detected in a test batch, this method collects all alarm logs, parameter exceedance records, and other anomalies recorded during the batch's execution. The system then counts the frequency of each anomaly type in the batch. Based on the frequency values, all anomaly categories are sorted from highest to lowest, generating a list of suspected root causes, and the anomaly category with the highest frequency is listed as the highest priority candidate root cause for investigation.
[0004] However, this frequency-based statistical method cannot effectively distinguish the true technical role of anomalies based on their dynamic behavior throughout the entire test production sequence, such as their location on the timeline, the duration of their impact, and their aggravating effect on fluctuations. This results in maintenance measures failing to fundamentally solve the problem, and yield fluctuations may occur repeatedly. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method for multidimensional data correlation analysis of probe test yield fluctuations, thus solving the problems in the background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for multidimensional data correlation analysis of probe test yield fluctuations, comprising the following steps: Step S1: Obtain the evolutionary context of the test batch, and construct a set of preceding trigger scenarios for yield fluctuations based on the evolutionary context of the test batch; Step S2: Extract and sort the events from the preceding trigger scene set to form a batch event sequence. Perform correlation analysis on the batch event sequence to generate an event disturbance transmission sequence and preliminary classification results.
[0007] Step S3: Based on the event disturbance propagation sequence and preliminary classification results, generate cross-batch anomaly propagation chain or sudden anomaly characteristic analysis results.
[0008] Step S4: Construct a root cause localization model based on verification-driven and role profiling, input the cross-batch anomaly propagation chain or sudden anomaly feature analysis results into the root cause localization model, and output the anomaly root cause results.
[0009] Preferably, obtaining the evolutionary trajectory of the test batch includes: Extract target test batches from production data The identification information and the specific time points when yield fluctuations occurred. By time point Using this as a baseline, we trace back along the timeline to obtain data from point [time point]. At the appointed time All test batches completed within the time window constitute a candidate preceding batch set. Candidate Pre-batch Set Contains multiple batches, denoted as ,in This indicates the total number of candidate batches in the set; Candidate preceding batch sets are organized according to a predetermined set of rules closely integrated with the characteristics of probe testing. The batches are filtered and sorted to reconstruct the target test batch. The evolution of the test batches that had an impact.
[0010] Preferably, the set of preceding triggering scenarios for yield fluctuations, constructed based on the evolution of test batches, includes: Identify the evolution of test batches Each preceding batch For target batch Potential impact roles, through analysis of each preceding batch This is achieved through metadata and process data, with the analysis focusing on indicators directly related to probe testing. The first category consists of continuous batches that constitute the accumulation of contact states, and their roles are denoted as... The second category is the transition batch that introduces program switching disturbances, and its role is denoted as... The third category is the induced batch that amplifies environmental shifts, and its role is denoted as... ; The defining characteristic is a significant drift in environmental parameters such as temperature and humidity relative to process control limits during the batch's execution, which is crucial for understanding the evolution of the test batch. Each batch It is given a specific role label. The value of this tag is a set containing One element in the set will determine the evolution of the test batch. Character tags corresponding to each batch The combination constitutes the set of preceding triggering scenarios for subsequent analysis. .
[0011] Preferably, the process of extracting and sorting events from the preceding triggering scenario set to form a batch event sequence includes: From the preceding trigger scene set Each batch From the corresponding process logs, equipment monitoring data, and environmental sensor data, a series of key events are extracted, and each event is abstracted into a quintuple. ,in, Indicates the timestamp of the event; This indicates the type of event, and its value comes from the set of event types mentioned above. Indicates the intensity or magnitude of an event; This indicates the test batch identifier to which the event belongs. The field is directly inherited from the batch to which the event belongs. In the preceding trigger scene set The role assigned to the middle ; For each batch Events extracted internally, according to their timestamps Arrange in ascending order to form a batch event sequence
[0012] Preferably, the association analysis performed on the batch event sequence to generate an event disturbance propagation sequence and preliminary classification results includes: In-batch micro-analysis, for each batch event sequence Only pairwise correlation analysis of events within the sequence is performed to identify succession and amplification relationships within a batch; in cross-batch macro-analysis, the batch is used as the basic unit, starting from the event sequence of each batch. The three most representative events within a batch, or those at the beginning or end in terms of time, are selected as the characteristic events of that batch; then, a global time list is used. The global temporal location of the feature event is determined, and the delayed release and coupling triggering relationship are determined only for two batches of feature events that meet the conditions. Traverse and analyze the time-ordered list of events Construct a directed graph structure To represent the perturbation propagation relationship between all events, where, It is a collection of nodes, where each node corresponds to an event. ; It is a set of edges, where each directed edge originates from an event. Pointing to event , and with weight labels; From a directed graph Extract all pre-triggered scene sets The last batch Directed path with associated yield fluctuation events as the endpoint After calculating the ranking scores of all directed paths, they are sorted from highest to lowest score. The path with the highest ranking score is selected as the main event perturbation propagation sequence, forming the final event perturbation propagation sequence. The event disturbance propagation sequence The events in the data are categorized by batch to obtain preliminary classification results.
[0013] Preferably, based on the event disturbance propagation sequence and preliminary classification results, the analysis results of cross-batch anomaly propagation chains or sudden anomaly characteristics include: Received event disturbance propagation sequence Based on the preliminary classification results, a differentiated analysis strategy is adopted, if the event disturbance transmission sequence... Classified as transitive, it tracks the continuous cross-batch propagation path of anomalous features along multiple dimensions, generating a cross-batch anomalous propagation chain; if the event perturbation propagation sequence Classified as sudden type, generating sudden anomaly feature analysis results.
[0014] Preferably, constructing a root cause localization model based on validation-driven and role profiling includes: The root cause localization model based on verification-driven and role profiling adopts a modular and process-oriented processing architecture. Its core processing logic consists of three sequentially executed core processing sub-modules. First, the multi-dimensional role feature extraction module calculates quantitative features representing the behavioral patterns of each anomalous factor based on multiple batches of input test data. Second, the verification-driven evidence calculation module searches for subsequent corrective events and calculates quantitative evidence for yield recovery. Finally, the fusion decision and root cause determination module integrates the outputs of the preceding modules and makes a comprehensive decision based on preset rules. These three modules work together sequentially to complete the derivation from data to root cause conclusions.
[0015] Preferably, the multi-dimensional role feature extraction module includes: In the next model, character profiling is achieved by calculating quantified scores across three key dimensions, thus forming a priori character feature vector. The three dimensions are origin, persistence, and amplification. Each anomalous factor receives a quantified character feature vector. ; In Mode 2, the sudden causal association score is directly adopted. The core quantitative characteristics of abnormal factors are used as the basis for calculation; simultaneously, the instantaneous deviation magnitude of sudden abnormal events in four dimensions is calculated to form a sudden feature vector. .
[0016] Preferably, the verification-driven evidence calculation module includes: The verification-driven evidence calculation module begins its work. This module searches the event log for the first corrective event directly related to the current anomaly. If a related event is found, it calculates the key verification metric, known as the inverse convergence coefficient. The calculation formula is as follows: ; in, It is a core sub-item used to quantify yield recovery speed; It is the normalized value of the standard deviation of batch yield after the corrective event, used to measure the stability of the recovery; and These are weighting coefficients, representing the importance of recovery speed and stability in the overall evaluation; Mode 1 involves a corrective event following the end of the search propagation chain; Mode 2 involves searching the target batch. List of sudden abnormal events following yield fluctuation events The first maintenance or correction operation directly related to the event type.
[0017] Preferably, the abnormal root cause results include: The output of the root cause results includes a list of root causes, which are labeled with different source categories depending on the analysis mode: For cross-batch root causes, such as the trend of deteriorating contact resistance of probe cards across batches, drift of test instrument calibration parameters across multiple batches, and cumulative effects caused by defects in test program versions, corresponding inverse convergence coefficients are provided. The values serve as quantitative verification evidence and describe the specific scope of influence of each root cause; For sudden root causes within the batch itself, such as sudden abnormal modifications to test program parameters, sudden transient failures of environmental controls, and sudden equipment alarms leading to test interruptions and resets, a sudden causal correlation score is provided. As quantitative evidence and each root cause in the target batch Description of the time range within.
[0018] This invention provides a method for multidimensional data correlation analysis of probe test yield fluctuations, which involves deep learning technology and has the following beneficial effects: (1) By constructing a cross-batch anomaly propagation chain from four key dimensions—probe card, test equipment, test program, and product category—this method can effectively identify the transmission and evolution path of anomaly characteristics between consecutive test batches. This approach expands the analysis scope from isolated results of a single batch to reflect the complete process of dynamic evolution of potential unstable states in multiple batches and multiple test contexts. This solves the problem of identifying surface anomalies that occur in the current batch but whose true causes are hidden in previous batches, and significantly improves the depth and accuracy of tracing the source of fluctuations.
[0019] (2) Through the classification analysis and adaptive classification mechanism of the event disturbance transmission sequence, the method realizes automatic differentiation and differentiated processing of different fluctuation source types. After generating the event disturbance transmission sequence in step S2, the method automatically determines the source type of the oscillation disturbance by analyzing the position of the first occurrence of the key amplified or coupled event in the sequence: if the key disturbance first appears in the preceding order, it is determined to be a descending order transmission; if the first significant disturbance occurs within the target batch itself, it is determined to be a sudden disturbance. This adaptive mechanism enables the method to flexibly respond to different scenarios and avoids the misjudgment and efficiency loss caused by the forced tracing of history when dealing with sudden problems in traditional methods.
[0020] (3) By designing a verification-driven evidence calculation module and defining a backward convergence coefficient, this method adds a powerful causal verification step to the role profiling. This invention provides objective empirical data support for root cause determination by associating subsequent corrective events and quantifying the speed and stability of yield recovery after corrective measures are implemented. For cross-batch-transmitted root causes, the verification evidence is reflected in the speed and stability of recovery after correction; for sudden root causes within a batch, the verification evidence is reflected in the recovery effect verification of immediate corrective operations after fluctuations. This makes the final root cause localization no longer solely dependent on prior correlation analysis, but possesses verifiable causal logic, thereby significantly improving the credibility and operability of the analysis conclusions and providing a direct and clear decision-making basis for production maintenance. Attached Figure Description
[0021] Figure 1 This is a flowchart of a multidimensional data correlation analysis method for probe test yield fluctuation proposed in this invention.
[0022] Figure 2 This is a hierarchical diagram of the root cause localization model in the multidimensional data correlation analysis method for probe test yield fluctuation proposed in this invention.
[0023] Figure 3 This is a hierarchical diagram of abnormal root cause results obtained in a multidimensional data correlation analysis method for probe test yield fluctuations proposed in this invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Please see Figures 1-3 This invention provides a technical solution: a method for multidimensional data correlation analysis of probe test yield fluctuations. Specifically, it provides the following method for multidimensional data correlation analysis of probe test yield fluctuations; please refer to [link to relevant documentation]. Figure 1 The method includes the following steps: Step S1: Obtain the evolution path of the test batch and construct a set of preceding trigger scenarios for yield fluctuations based on the evolution path of the test batch.
[0026] This step aims to establish a data foundation with temporal evolution and causal correlation attributes for subsequent correlation analysis. The core idea of this method is to understand yield fluctuations in semiconductor probe testing as the endpoint of a dynamic evolutionary process closely related to probe contact state, test equipment interaction, and manufacturing process background, rather than an isolated anomaly. Therefore, this step needs to focus on the target test batch. The system systematically traces and reconstructs the key historical test context prior to its formation, thereby constructing a structured set of pre-triggered scenarios rich in probe test scenario information. .
[0027] First, extract the target test batch from the production data. The identification information and the specific time points when yield fluctuations occurred. Next, based on time points... Using this as a baseline, we trace back along the timeline to obtain data from point [time point]. At the appointed time All test batches completed within the time window constitute a candidate preceding batch set. Candidate Pre-Batch Set Contains multiple batches, denoted as ,in This represents the total number of candidate batches in the set.
[0028] The start of the time window The threshold is set to 8 hours prior to the occurrence of the current abnormal batch. This threshold was derived through a systematic source analysis of historical yield fluctuation events. Actual data shows that the vast majority of root causes leading to drastic yield fluctuations, such as probe card replacements, major program switches, or drastic environmental parameter disturbances, occur within the 8-hour period before the anomaly manifests. Using this time span ensures coverage of critical preceding events that could trigger failures while eliminating historical data with weak correlation due to its age, thus balancing depth of root cause analysis with data processing efficiency.
[0029] Subsequently, the candidate preceding batch set was organized according to a predetermined set of organizational rules closely integrated with the characteristics of probe testing. The batches are filtered and sorted to reconstruct the target test batch. The evolution of the most likely impactful test batches is outlined below. The priority order and purpose of these organizational rules are as follows: Prioritize test equipment to capture the impact of compatibility between specific equipment and probe cards, as well as the continuity of equipment status; prioritize probe cards to directly focus on the evolution of contact state accumulation, wear, or contamination effects of the same set of physical probes; prioritize test procedures to isolate disturbances caused by test procedure logic switching in probe contact condition judgment; prioritize product categories to eliminate background noise from electrical parameters due to product design differences, thus allowing for a purer analysis of probe contact issues; and prioritize production shifts to correlate environmental factors that may affect probe test stability within the same time period, such as temperature and humidity fluctuations. Based on these rules, each candidate batch... Calculate its relationship with the target batch Association weight .
[0030] Relevance weight The calculation method can be expressed as: ; in, The candidate batch is in the set Index in; Represents the dimension index in the organization rules, when Time represents the dimension of the test machine. Time represents the probe card dimension, when Time represents the dimension of the test program. Time represents the product category dimension, when Time represents the production shift dimension; It is an indicator function, when candidate batches With the target batch In the When all dimensions are consistent, its function value is 1; otherwise, it is 0. It is the first Priority weight coefficients were assigned to each dimension. Long-term historical test data covering multiple product categories and testing equipment were collected. Batches were grouped according to consistency across each dimension, and Pearson correlation analysis was used to calculate the correlation coefficient between the consistency status of each dimension and the final yield of the batch. The results showed that batches using the same physical probe card had the most significant correlation with yield fluctuations, with a correlation coefficient of 0.72; the correlation coefficient for batches using the same testing equipment was 0.68; these two constituted the core physical factors affecting yield. The correlation coefficients for test procedures and product categories were 0.45 and 0.22, respectively, and were considered secondary factors. The above correlation coefficients were normalized and fine-tuned based on engineering experience, and the final weight coefficients were assigned values. This quantitatively reflects the relative importance of each dimension's impact on test stability.
[0031] It should be noted that the indicator function The rules for determining the value: When When testing a test machine, if the candidate batch and the target batch are run on the exact same physical machine with the same machine asset number, then the function value is... Otherwise ;when When using a probe card, if the candidate batch and the target batch use the same physical probe card with the same unique identifier, the function value is... Otherwise ;when When testing the program, if the candidate batch and the target batch use the same full version number of the test program, the function value is... Otherwise ;when When considering product categories, if the complete product part number of the chips processed in the candidate batch is the same as that in the target batch, then the function value is... Otherwise ;when When the production shift is specified, if the candidate batch and the target batch are completed by the same production team within the same scheduling cycle, then the function value is... Otherwise .
[0032] Then, based on the calculated correlation weights Sort all candidate batches in descending order and select those with weights higher than a threshold. The batches constitute a list of highly correlated batches after initial screening. Threshold The value is 0.5. The significance of setting this threshold is that a batch must be consistent with the target batch simultaneously in at least the two most core physical dimensions: the testing equipment and the probe card. Alternatively, given consistency in one core dimension, at least two secondary dimensions must also be consistent, such as... Only batches selected from these groups can be considered highly correlated. This ensures that the selected batches have a significant and non-accidental correlation with the target batch within the context of probe testing.
[0033] Next, the list of highly correlated batches... The batches are arranged according to their completion time, forming a time-based evolutionary trajectory of the test batches, denoted as... .in Representing the evolutionary path The total number of highly correlated batches included in it meets the following requirements. Within this context, The earliest highly correlated batch in terms of time, For the closest to the target batch Highly correlated batches.
[0034] Finally, and this is the key step, is to identify the evolutionary path of the test batch. Each preceding batch For target batch The potential impact role. By analyzing each preceding batch. It utilizes metadata and process data, with analysis focusing on metrics directly related to probe testing. The influencing roles are primarily categorized into three types.
[0035] The first type is a continuation batch that constitutes the accumulation of successive states, and its role is denoted as... The criterion is that after the batch is completed, its selected key indicators, such as the average probe connection resistance, are compared with the previous batch. Connecting the moving averages of each batch, a monotonically increasing trend is observed for three or more consecutive batches, with the growth rate of the most recent batch exceeding the baseline. ; The second type is a transition batch that introduces program switching disturbances, whose role is denoted as... The criterion is that when this batch is executed, the MD5 hash value of the test program recorded by the system is different from that of the previous batch that is immediately adjacent to it, and the change involves the parameter field of probe contact timing or signal judgment threshold.
[0036] The third type is the induced batch that amplifies environmental shifts, and its role is denoted as... The criterion was the weighted comprehensive temperature and humidity index recorded by environmental sensors during the execution of this batch. The duration exceeding the process control limit exceeds the total duration of the batch. ,in The calculation formula is: , This is the normalized temperature value. This is the normalized value for humidity.
[0037] For the evolution of the test batch Each batch If the criteria for the corresponding role are met, then a specific role label is assigned. The value of this tag is a set. One of the elements in [the dataset]. The evolutionary trajectory of the test batch. Character tags corresponding to each batch The combination constitutes the set of preceding trigger scenarios for subsequent analysis. Its mathematical expression is In this way, yield fluctuations are transformed from a single point of result into a time-field spectrum closely related to the probe testing scenario, marked with preliminary causal annotations, thus establishing a clear background for subsequent in-depth analysis of the dynamic evolution of probe-related anomalies.
[0038] At this point, the generated set of preceding triggering scenarios has been completed. The target batch was described Key historical context preceding the fluctuations. Target batch. Its own abnormal characteristics will be discussed in subsequent step S3, along with... By comparing the characteristics exhibited by each batch in terms of continuity and trend, the source of fluctuations can be effectively distinguished. Is it a sudden, isolated problem of its own, or the final outbreak of a gradual instability that had already emerged in previous batches and spread along the dimensions of probe cards, testing equipment, etc.
[0039] Step S2: Extract and sort the events from the preceding trigger scene set to form a batch event sequence. Perform correlation analysis on the batch event sequence to generate an event disturbance transmission sequence and preliminary classification results.
[0040] Continuing from the pre-triggered scenario set constructed in step S1 This step aims to extract key events from the scenario and construct a sequence describing the dynamic relationships between these events, i.e., an event perturbation propagation sequence. The core idea of this step is to understand yield fluctuations in semiconductor probe testing as the result of multiple small perturbation events related to probes, equipment, procedures, and the environment, which are gradually superimposed, propagated, and ultimately amplified at different time levels through specific relationships. Therefore, this step needs to transform the discrete event points identified in step S1 into a dynamic process chain depicting how perturbations gradually evolve into significant fluctuations around the core aspects of probe testing.
[0041] First, from the preceding trigger scene set Each batch A series of key events were extracted from the corresponding process logs, equipment monitoring data, and environmental sensor data.
[0042] The extraction process is executed based on a predefined event rule base, which clarifies the specific triggering conditions and filtering criteria for various events. For example: only alarm events of the test equipment with an alarm level of severe or major are extracted; environmental parameter events must meet the following conditions to be recorded: the measured value continuously deviates from the process center value by more than 20% of the specified tolerance for more than 2 minutes; test program switching events only record situations where the version number has changed and the change involves key parameter files such as probe contact force, test timing, or electrical judgment thresholds; the batch average probe contact count exceeding the baseline value in the baseline event is defined as the statistical process control upper limit of the average contact count of the previous 25 normal batches. This ensures that all extracted events are operations or state changes that have a potentially substantial impact on the stability of probe testing.
[0043] These events are behavioral or state changes that may disturb probe contact stability, test signaling integrity, or test condition consistency. They mainly include, but are not limited to, the following categories: probe card replacement events, probe cleaning events, test equipment alarm clearing events, test program switching events, test equipment unplanned shutdown and reset events, batch average probe contact counts exceeding the baseline, and test environment temperature and humidity parameters exceeding control limits. Each event is abstracted into a quintuple. .
[0044] in, Indicates the timestamp of the event; This indicates the type of event, and its value comes from the set of event types mentioned above. This indicates the intensity or magnitude of an event, such as the duration of a probe cleaning event, the degree of deviation from the ambient temperature event, or the specific number of probe contact retries. Indicates the test batch identifier to which the event belongs; The field is directly inherited from the batch to which the event belongs. In the preceding trigger scene set The role assigned to the middle .
[0045] In subsequent event correlation analysis, based on the event quintuple... Differentiate fields from For events in the (program switching disturbance) batch, the test program switching event is used as the core hub node of the relationship network, and its connection or amplification relationship with other events is searched first; for events from... (Environmental offset amplification) For events in the batch, forcibly extract events with environmental parameters exceeding limits, and prioritize analyzing the amplification relationship between these events and probe contact anomaly events; for events from (Contact State Accumulation) events of a batch are preferentially associated with probe contact events of adjacent batches in cross-batch analysis.
[0046] To build a time-series view of event analysis, first analyze each batch. Events extracted internally, according to their timestamps Arrange in ascending order to form a batch event sequence This operation preserves the batch-level causal context established in step S1, ensuring the integrity of the micro-perturbation logic within each batch. Simultaneously, it sorts the events of all batches by global timestamp, forming a global time list. It serves as a timing reference framework for cross-batch event association.
[0047] The above hierarchical structure is used collaboratively in subsequent event correlation analysis in the following ways: In-batch micro-analysis, for each batch event sequence This involves performing pairwise correlation analysis only within the sequence to identify succession and amplification relationships within batches. The scope of this analysis is strictly limited to the same... Internally, ensure that physically unrelated events from different batches are not mistakenly classified as causal relationships.
[0048] In cross-batch macro analysis, using batches as the basic unit, we first start with the event sequence of each batch. The three most representative events within a batch, or those at the beginning or end in terms of time, are selected as the characteristic events of that batch; then, a global time list is used. The global temporal location of the feature events is determined, and the delayed release and coupling triggering relationship is determined only for two batches of feature events that meet the following conditions: Condition 1 is that the two batches are assigned high correlation weights in step S1. Condition two is that the two batches are in the global time list. The time distances are within the corresponding time thresholds of the relationships. This hierarchical approach preserves the logical integrity of micro-perturbations within batches, significantly reduces the computational load of cross-batch analysis through feature event extraction and dual-condition filtering, and effectively avoids spurious relationships introduced by global blind associations.
[0049] Subsequently, a limited-scope correlation analysis is performed on the events in the above hierarchical structure to identify only physically meaningful event relationships. This step defines four core event relationships: succession relationship, amplification relationship, delayed release relationship, and coupling triggering relationship. For relationships within a batch, for each For internal event pairs, it is determined whether there is a successor relationship or an amplification relationship. For cross-batch relationships, for feature event pairs from different batches, under the dual conditions of high correlation weight and time-series threshold, it is determined whether there is a delayed release relationship or a coupling triggering relationship. This limited analysis strategy ensures the physical validity and computational feasibility of relationship identification.
[0050] To identify these relationships, specific judgment logic and quantification thresholds need to be set. For any two events that occur sequentially in time... and Among them, the events timestamp Smaller than event timestamp Determine whether there is a relationship between them.
[0051] The first type of relationship is a succession relationship, denoted as... When the event It is an event The direct follow-up steps or events in the operational process. The occurrence of this is in response to or handling of an event. When the state is represented, the succession relationship is determined to be established. For example, a cleaning event targeting a specific probe card. Immediately afterwards, a test program switching event occurred. To verify the cleaning effect of the probe card, a connection exists between the two. Its formal condition is that the event... The type belongs to event The standard set of subsequent response operations, and the time interval between two events. Less than a preset process response time threshold The threshold Set to 2 minutes, this threshold It is determined by analyzing the time interval distribution of alarm-reset, program loading-ready, and other event pairs with known process succession relationships in historical event logs, and taking the 95th percentile, aiming to cover the vast majority of real process response situations.
[0052] The second type of relationship is the amplification relationship, denoted as... The criteria for judgment are further divided into the following two situations: For events of the same type, amplification is applied when the event... and When events belong to the same type and their corresponding value fields are all continuous numerical values, the value difference formula is used for determination. Specifically, if the following conditions are met... If the amplification relationship holds, then the formula is considered valid. It is a very small constant used to prevent the denominator from being zero; To amplify the decision coefficient, it is set to... This means that when the intensity of a subsequent event is greater than that of a preceding event... When the above is true, it is considered an amplification relationship.
[0053] For amplifying heterogeneous events, when the event and They belong to different types, but based on domain knowledge, they are... yes When the known physical cause is known, it can also be determined as an amplification relationship. For example, an environmental temperature shift event usually causes an event where the average number of probe contacts per batch exceeds the limit. The determination criterion is: in this batch sequence, the event... The magnitude compared to when it did not occur The average baseline intensity of similar events increased by more than At this point, the weight of the edge is taken as the growth rate plus one. This extended rule expands the scope of the amplification effect identification from similar events to dissimilar events with causal relationships.
[0054] The third type of relationship is the delayed release relationship, denoted as... Its formal determination requires that the following three conditions be met simultaneously: First, the incident and The type is defined in the process knowledge base as having a potential delayed causal relationship. For example, a predefined association rule states that a probe card collision event will cause a change in the contact resistance value of the probe card; secondly, the time interval between the two must meet specific constraints, namely, greater than a minimum delay threshold and less than a maximum delay threshold. The maximum delay threshold is set at 3 hours, a value determined based on statistical analysis of the time required for slow-changing factors such as probe contamination accumulation and equipment parameter drift to have an impact; the minimum delay threshold is set at a unit work cycle, and this lower limit is set to avoid confusion with succession relationships; thirdly, in the event The batch in which it occurred contained items related to the incident. The directly related performance metrics exhibit quantifiable degradation, and the degradation trend is significant. Specifically, during the final period of the maximum delay threshold window, the slope of the linear fitting trend of the relevant metrics is greater than 0.05, and the goodness of fit is greater than 0.6.
[0055] The fourth type of relationship is the coupling triggering relationship, denoted as... When two events occur within a preset coupling time window and When they appear together, analyze their impact on the event. The impact. A two-stage strategy is used for determination: When sufficient statistics are available, if the combination of events occurs at least 10 times in the historical database, then the probability gain ratio formula is used for determination: ; In this formula, Representative event type Marginal probability within a recent statistical window; represent and Co-occurrence conditions within the coupling time window The conditional probability of occurrence; The coupling determination threshold is set to 2.0. This value was determined by optimizing historical data with the goal of maximizing the detection of the true coupling mode and controlling the false alarm rate.
[0056] In cases of insufficient statistics, if historical samples are insufficient to calculate reliable probabilities, the immediate intensity comparison rule is used: when... and If they appear together The intensity compared to historical levels or When any one appears alone If the average intensity of the event is more than 50% higher, the coupling triggering relationship is also determined to be valid.
[0057] Then, based on the above relationship judgment logic, the time-ordered event list is traversed and analyzed. : Construct a directed graph structure This represents the perturbation propagation relationship between all events. It is a collection of nodes, where each node corresponds to an event. ; It is a set of edges, where each directed edge originates from an event. Pointing to event Each directed edge is assigned a weight label, which is a set of one or more of the aforementioned relation types. This is used to quantify the contribution of the relationship to the final yield fluctuation. To ensure the additivity and comparability of the weights of different types of relationships in subsequent path rankings, all weights are mapped to the [0, 10] interval using a unified rule: Relationship: Weight This represents the basic path transmission effect; Delayed release relationship: weight Because of its inherent time lag, it has a stronger indicative significance for tracing the root cause; Amplification relationship: weight When the intensity increase reaches 100% or more, the weight is close to the full score of 10; Coupling triggering relationship: weights The probability gain ratio is linearly mapped with an upper limit of 5, so that the weight of the strong coupling is close to the full score of 10.
[0058] Finally, from the directed graph Extract all pre-triggered scene sets The last batch Directed paths terminated by associated yield fluctuation events. To select the critical path that best reflects the perturbation propagation and roof effect from among many paths, these paths need to be ranked. The Rank score for each path is calculated using the following formula: ; in, This is the geometric mean of the weights of all directed edges along the path. This represents the total number of event nodes included in the path. Using the geometric mean effectively suppresses ranking bias caused by excessively high weights on one side, and better reflects the overall relationship strength of the path; logarithmic terms... This formula ensures that the contribution of path length exhibits diminishing marginal returns, avoiding an over-preference for long paths. It effectively filters out critical paths with generally high relationship strength and a certain evolutionary depth.
[0059] After calculating the ranking scores of all directed paths, they are sorted in descending order of score. The path with the highest ranking score is selected as the primary event perturbation propagation sequence. Simultaneously, if other paths exist whose ranking scores differ from the highest score by less than 10%, they are also included as candidate propagation sequences. The final event perturbation propagation sequence is then determined. It consists of a main sequence and alternative sequences to avoid missing key root cause chains when there are multiple competing transmission hypotheses.
[0060] The final output event perturbation propagation sequence It not only provides a timeline of key events, but also clearly defines the target batch. Position and role in the perturbation propagation path. Based on this sequence, this step performs the following preliminary classification to provide analytical guidance for step S3: Event perturbation propagation sequence The events in the data are categorized and analyzed by their respective batches: If the key amplification event or coupled triggering event that causes yield fluctuations first appears In a previous batch, and in If the fluctuations are merely a delayed release or continuation of these preceding disturbances, it can be preliminarily determined that the yield fluctuations belong to the type of potential instability propagation across batches, and the real cause is hidden in the preceding batches. At this point, Marked as a transitive sequence and associated with the set of the primary preceding batches. ; If the event disturbance propagation sequence In The previous event chain relationship weight was too low, for example, below 1.5, and the first significant amplification relationship or coupling triggering relationship occurred during Internally, the initial assessment is that the issue stems from a sudden problem within the batch itself. At this point, Marked as burst-type sequences and associated Internal list of unexpected and abnormal events .
[0061] This preliminary classification result is consistent with the event disturbance propagation sequence. The data is input into step S3, which will then conduct targeted analysis: for transitive sequences, the cross-batch evolution trend of anomalous features will be traced back along dimensions such as probe cards and testing machines; for burst-type sequences, the focus will be on analyzing the target batch. The combined characteristics of its own sudden abnormal events and their causal relationship with the sharp drop in yield.
[0062] Step S3: Based on the event disturbance propagation sequence and preliminary classification results, generate cross-batch anomaly propagation chain or sudden anomaly characteristic analysis results.
[0063] The event perturbation propagation sequence output from step S2 Based on the preliminary classification results, the core task of this step is to adopt differentiated analysis strategies according to the preliminary classification results. If the event disturbance propagation sequence... If classified as transitive, it indicates that the root cause of the quality fluctuations may be hidden in previous batches. This step will trace the continuous cross-batch propagation path of the anomalous features along multiple dimensions to generate a cross-batch anomaly propagation chain. If the event disturbance propagation sequence... If it is classified as sudden, it means that the fluctuation is more likely to originate from the target batch. For internal sudden disturbances, this step will focus on analyzing the correlation characteristics of internal sudden events to generate sudden anomaly characteristic analysis results. Through this strategy, the focus of analysis can be effectively expanded from a single batch of instantaneous anomalies to a complete framework that combines multi-batch dynamic evolution with single-batch sudden characteristics.
[0064] First, the core analysis objects are determined based on the preliminary classification results.
[0065] For transitive mode, directly obtain the output of step S2 and the target batch. The main preceding batch set with causal transmission relationship Combine this set with the target batch. Merging them together forms the core batch set to be analyzed. ; For the burst-type pattern, the analysis object is the target batch. itself, at this time, the core batch Subsequent analysis will focus on the internal components of this batch; Next, focusing on the core batch set Construct multidimensional batch sequences for analysis.
[0066] To improve efficiency, this step refers to... The physical carriers associated with the events are used for dimensional focusing. If the amplification or delayed release relationship in the sequence mainly involves a specific probe card, then the batch sequence of the probe card dimension is constructed first; if it mainly involves a specific test equipment, then the batch sequence of the test equipment dimension is constructed first. If the event relationship is not significantly biased towards a certain dimension, then the batch sequences of all four dimensions are constructed as follows: The first dimension is the probe card dimension, which is extracted from the target batch. All test batches processed consecutively using the same physical probe card before and after the occurrence point are sorted by test completion time to form a probe card dimension batch sequence. .in, This indicates the total number of batches in the batch sequence of this probe card dimension. This sequence is used to track the evolution of contact performance of the same set of probes.
[0067] The second dimension is the testing equipment dimension, which extracts data from the target batch. All test batches performed by the same testing machine before and after the occurrence of the event, and sorted by test completion time, constitute the test machine-based batch sequence. .in, This represents the total number of batches in the batch sequence for this test equipment dimension. This sequence is used to analyze the impact on the system stability of a specific test equipment.
[0068] The third dimension is the test program dimension, which extracts data from the target batch. All test batches using the same version of the test program before and after the occurrence point are sorted by test completion time to form a test program-dimensional batch sequence. .in, This indicates the total number of batches in the batch sequence of this test program dimension. This sequence is used to observe the consistency impact of the test program logic on probe test behavior.
[0069] The fourth dimension is the product category dimension, extracted from the target batch. All test batches belonging to the same product category before and after the occurrence time point are sorted by test completion time to form a batch sequence based on product category. .in, This indicates the total number of batches in the batch sequence for this product category. This sequence is used to control for background noise caused by product design differences.
[0070] Subsequently, trend analysis was performed on the batch sequence of the aforementioned dimension; For each dimension batch sequence in the transitive pattern (e.g.) The values of selected key performance indicators (such as the average number of probe contact retries per batch) for each batch were calculated to obtain the indicator change sequence. To quantify the persistence of anomalous characteristics, linear fitting was used to assess the significance of the change trend. For the sequence... Its trend slope The calculation formula is: ; in, Indicates the number of consecutive batches used for trend analysis; It is the sequential index of the batch in the analyzed subsequence, from 1 to... It is the first The formula is used to calculate the slope of the linear trend of the abnormal feature value as the batch sequence index changes, with positive values indicating an upward trend and negative values indicating a downward trend.
[0071] It should be noted that, The value is adaptively determined based on the characteristics of the dimension sequence: the minimum is 3, the maximum is the total number of batches in the dimension sequence within the previous time window, and the default value is 5. This default value is based on the analysis of historical anomalous propagation cases—in more than 85% of confirmed cross-batch propagation cases, the process from the initial detectable trend to significant deterioration occurs within 5 batches. If the total number of preceding batches in the dimension sequence is less than 5, then all available preceding batches are used; If the calculated trend slope Greater than a set positive trend threshold If so, it is determined that the anomalous feature exhibits a significant upward and deteriorating trend in the batch sequence of that dimension. Trend threshold The value is set to 0.1. This threshold was determined through statistical analysis of a large amount of batch data from a historical stable production phase, i.e., a period of three consecutive months without significant yield fluctuations. The specific calculation process includes: statistically analyzing the slope distribution of the natural fluctuation trends of various key performance indicators across any five consecutive batches, and selecting the 90th percentile as the final threshold. This percentile was chosen to strike a balance between the sensitivity and specificity requirements for anomaly detection in engineering practice. Validation results on historical stable production data show that this threshold achieves approximately 90% specificity, meaning that only 10% of stable sequences will be falsely identified as trending; simultaneously, it achieves a recall rate of over 95% for known anomaly cases. This setting aims to effectively avoid over-alarms while promptly capturing genuine anomalies.
[0072] For emergency response modes, perform internal emergency analysis. The list of sudden abnormal events output from step S2 In the process, the type, intensity, and occurrence time of each event are extracted, and the relationship between each event and the time of occurrence is calculated. The time interval at which yield begins to decline; dimensional instantaneous performance analysis for... The benchmark performance data for each of the four context dimensions—probe card used, test equipment, test program version, and product category—is extracted. For example, the average contact resistance and average number of retries of the probe card in previous normal batches are used to calculate... The deviation of the corresponding indicator from the benchmark; List of sudden abnormal events For each event, calculate its sudden causal association score with the yield decline. : ; in, The time when the event occurred. The time when the yield rate begins to decline. For the intensity of the event, This serves as the baseline value for this dimension. To prevent extremely small constants with a denominator of zero, the scoring takes into account both temporal tightness and intensity significance.
[0073] Finally, this step outputs the corresponding results based on the analysis mode: Transitive pattern output, set of cross-batch exception propagation chains This set represents propagation chains identified across four dimensions, exhibiting anomalous characteristics and a continuous deterioration trend across batches. Each propagation chain... It is associated with a specific batch list and the main anomalies identified. The data on the changing trend of this feature, including the trend slope. and the number of consecutive batches involved And the dimensional type to which this propagation chain belongs. Among them, This represents the total number of identified propagation chains, and its value is less than or equal to 4.
[0074] Burst-type output, burst anomaly characteristic analysis results The result includes the target batch. Internal list of unexpected and abnormal events Scoring of the sudden causal relationship of each event And instantaneous deviation data in four dimensions.
[0075] The two products described above complement each other, meaning that a corresponding output is only generated when the judgment of a certain pattern is met. This provides a direct and structured input for the root cause localization model in the subsequent step S4. Step S4 will activate the corresponding analysis mode based on the specific type of input, namely, the propagation chain analysis mode or the sudden problem analysis mode.
[0076] Step S4: Construct a root cause localization model based on verification-driven and role profiling, input the cross-batch anomaly propagation chain or sudden anomaly feature analysis results into the root cause localization model, and output the anomaly root cause results.
[0077] Following the analysis results output in step S3, this step aims to construct a dedicated root cause localization model based on validation-driven and persona-based approaches. This model supports two analysis modes and automatically initiates the corresponding processing flow based on the output type of S3: Mode 1 is the propagation chain analysis mode: when step S3 outputs a set of cross-batch anomaly propagation chains. At that time, the model takes these propagation chains as the main input, performs multi-dimensional role profiling and verification-driven evidence calculation on the identified anomalous factors, and outputs cross-batch propagation root cause results.
[0078] Mode 2 is the sudden problem analysis mode: when step S3 outputs the sudden anomaly characteristic analysis results. At that time, the model uses the list of sudden abnormal events in the result. and its sudden causal relationship score As the primary input, combined with the target batch The process data outputs the root cause results of the batch's own sudden occurrences.
[0079] This root cause localization model, based on verification-driven and role profiling, employs a modular and process-oriented architecture. Its core processing logic consists of three sequentially executed sub-modules. First, the multi-dimensional role feature extraction module calculates quantitative features representing the behavioral patterns of each anomalous factor based on multiple batches of input test data. Second, the verification-driven evidence calculation module searches for subsequent corrective events and calculates quantitative evidence for yield recovery. Finally, the fusion decision and root cause determination module integrates the outputs of the preceding modules and makes a comprehensive judgment based on preset rules. These three modules collaborate sequentially to complete the derivation from data to root cause conclusions.
[0080] First, the input to the model includes the structured results from step S3: In mode one, the input is a set of cross-batch exception propagation chains. In Mode 2, the input is the result of the sudden anomaly feature analysis. In addition, when calculating internal metrics, the model also needs to use relevant historical and real-time batch performance data, as well as event logs including probe card operation records, equipment maintenance records, etc., as auxiliary inputs.
[0081] The processing flow begins with the multi-dimensional role feature extraction module. In the first mode, this module calculates the quantitative scores of the portrait in three key dimensions to form a priori role feature vector. The three dimensions are origin, persistence and amplification. The first dimension is starting point, and the starting point score is... The calculation formula is: ; in, Indicates abnormal characteristics In the cross-batch abnormal propagation chain The index of the first detected position in the batch sequence, with the index counting starting from 0 (i.e., the index of the first batch in the sequence is 0); Indicates the propagation chain The total number of batches included. If there are abnormal characteristics... Only in the target batch This was the first time it was detected in the sample, and it had not appeared in any previous batches. Values Corresponding score This represents the minimum starting score. The purpose of this formula is to quantify the relative timing of the appearance of anomalies in the probe test sequence, and the score is calculated accordingly. The higher the value, the earlier it appeared, and the more likely it is to be the source triggering factor in the transmission chain.
[0082] The second dimension is persistence, specifically the persistence score. The calculation formula is: ; in, This indicates an abnormal propagation chain across batches. In the middle, abnormal features The number of batches detected. The purpose of this formula is to calculate the coverage ratio of this anomaly throughout the entire transmission chain, in order to assess the repeatability and persistence of its impact in continuous probe testing. (Score) The higher the value, the stronger the persistence.
[0083] The third dimension is magnification, and the magnification score is... The calculation formula is: ; in, Indicates abnormal features The average absolute value of the yield difference between adjacent batches after the first detection; This represents the average absolute value of the yield difference between adjacent batches before the feature is first detected. This is the average yield fluctuation during a historically stable production phase, used as a baseline for comparison. This baseline is taken from the target batch. The average absolute value of the yield difference between adjacent batches is calculated using all normal batches within the same product category and the same testing machine, that is, batches that did not trigger S1 yield fluctuation detection within 30 days. It is a very small constant, taking the value of This is used to prevent the denominator from being zero. The purpose of this formula is to quantify the exacerbating effect of abnormal factors on the fluctuation of inter-batch probe test yield. If the calculation result... This indicates that the factor has an amplifying effect, and The larger the value, the more significant the amplification effect. At this point, each anomaly has obtained a quantified role feature vector. The initial portrait was completed.
[0084] In Mode 2, the processing of the multi-dimensional role feature extraction module is different: because the analysis object of the sudden problem is the target batch. Since the event itself is an unexpected event and does not have the concept of a starting point or continuity across batches, this module directly adopts the sudden causal relationship score output from step S3. The core quantitative characteristics of anomalous factors.
[0085] Simultaneously, the instantaneous deviation magnitude of the sudden abnormal event in four dimensions is calculated to construct the sudden feature vector. Among them, the deviation magnitude value Dev represents the target batch. The standard deviation of the corresponding indicator relative to the dimensional baseline.
[0086] Subsequently, the verification-driven evidence calculation module begins its work, aiming to provide decisive empirical support for the character profile. This module searches the event log for the first corrective event directly related to the current anomaly. If a relevant event is found, a key verification metric, called the inverse convergence coefficient, is calculated. The calculation formula is as follows: ; in, It is a core sub-item used to quantify yield recovery speed; It is the normalized value of the standard deviation of batch yield after the corrective event, used to measure the stability of the recovery; and These are weighting coefficients, representing the importance of recovery speed and stability in the overall evaluation. Based on the analysis of engineering experience and historical recovery patterns, these two weighting coefficients were set as follows: and This assignment is based on statistical analysis of yield recovery data after correction of confirmed root cause cases: in cases where the true root cause was corrected, the yield recovery speed... The association strength with root cause confirmation (measured by a dotted bicollinear correlation coefficient of 0.78) was significantly higher than the association strength with recovery stability.
[0087] It should be noted that the sub-terms in the formula Calculated using the following formula: ; in, This represents the average yield of the last 5 affected batches prior to the occurrence of the associated corrective event. This represents the average yield of the first 5 batches completed after the corrective event occurred; It is the target yield or the historical baseline yield, taking the average yield of the same product and the same batch on the same machine during the historical stable production stage; It is a very small constant used to prevent the denominator from being zero.
[0088] Mode 1 involves a corrective event following the end of the search propagation chain; Mode 2 involves searching the target batch. List of sudden abnormal events following yield fluctuation events The first maintenance or correction operation directly related to the event type.
[0089] Finally, the fusion decision and root cause determination module makes the final ruling on each anomalous factor. This module integrates the outputs from the first two sub-modules, namely the role feature vector. and inverse convergence coefficient The system applies a set of pre-defined judgment rules. These rules are a multi-layered decision tree based on thresholds. The specific content and thresholds were determined through statistical analysis of a large number of historically confirmed root cause cases, combined with optimization using grid search and cross-validation methods. The core logic of the rules is to comprehensively judge based on the dimensional scores of the role profile and verification evidence. Specifically, the judgment rules are divided into two sets: The propagation chain pattern determination rule for Pattern 1 is: the starting point score is determined when an anomaly occurs. Amplification score And the inverse convergence coefficient At that time, this factor was identified as the verified source trigger gene root. This rule aims to screen for root causes that appear early in the probe test sequence, have a significant deteriorating effect, and show obvious recovery after correction. This corresponds to the anomalous feature appearing in the first 20% of the propagation chain. This indicates that the fluctuation range has increased by more than 20% compared to the baseline. but and When identified as a verified root cause of enhanced coupling, secondary causes that are not initial but significantly amplify the problem and are verified are identified. Furthermore, verification evidence could not be obtained, but the score remained consistent. and Factors that do not meet the above criteria are marked as highly probable factors to be verified. Factors that do not meet the above criteria are considered as concomitant factors of the outcome.
[0090] The rules for determining the emergency problem mode in Mode 2 are as follows: when the emergency factor of a certain emergency abnormal event is correlated with the score. When the event type has been identified as a root cause in the historical root cause database, and the instantaneous deviation of the event from the corresponding baseline value exceeds 50% in all four dimensions, it is determined to be a verified batch-emergent root cause. However, if none of the above conditions are met, it is marked as a highly probable sudden factor requiring verification. When this occurs, it is considered a factor accompanying the outcome.
[0091] The thresholds (0.8, 0.2, 0.7, 0.3, 0.1, etc.) were determined using the following method: On a historical case set consisting of yield fluctuation cases of confirmed root causes within the past three years, the goal was to maximize the F1-score for root cause localization, i.e., the harmonic mean of precision and recall. Five-fold cross-validation was used to optimize the threshold combinations using a grid search. The partitioning of the search space and the setting of initial values referenced the distribution differences of corresponding scores in confirmed root cause cases and associated cases. Finally, the threshold combination that yielded the highest F1-score was selected as the basic parameter for the decision rule. The specific definitions of precision and recall, the specific size of the training set, and the detailed settings of cross-validation are specific implementation parameters of this method. Those skilled in the art can fine-tune the thresholds based on the above method framework and their own data accumulation.
[0092] After processing by the three sub-modules described above, the root cause localization model based on verification-driven and role profiling completes its calculations and produces the final output. The output of the abnormal root cause results includes a list of root causes, which are labeled with different source categories depending on the analysis mode: For root causes that are transferred across batches, such as the trend of deteriorating contact resistance of probe cards across batches, drift of test instrument calibration parameters across multiple batches, and cumulative effects caused by defects in test program versions, corresponding inverse convergence coefficients are provided. The values serve as quantitative verification evidence, along with a description of the specific impact range of each root cause, such as clearly indicating which probe cards on which machines are affected, and how many batches are covered by the impact over a given timeframe.
[0093] For sudden root causes within the batch itself, such as sudden abnormal modification of test program parameters within the batch, sudden temporary failure of environmental control within the batch, or sudden equipment alarms leading to test interruption and reset within the batch, a sudden causal correlation score is attached. As quantitative evidence and each root cause in the target batch The time range is described, such as the time when the event occurred, the time when the yield started to decline, and the duration.
[0094] Through this step, the method completes the entire analysis from identifying the abnormal propagation path to building a deep role profile for the abnormal factors, and then to integrating verification evidence to make accurate judgments. Regardless of whether the root cause of yield fluctuations is cross-batch cumulative transmission or a sudden occurrence within the batch itself, it can output a definite abnormal root cause result that is closely related to the probe test field.
[0095] This invention proposes a multi-dimensional data correlation analysis method for probe test yield fluctuations. It achieves precise root cause localization through four rigorous steps, automatically distinguishing whether the fluctuation originates from cross-batch cumulative transmission or a sudden occurrence within the target batch itself. The specific steps are as follows: First, constructing a front-end trigger scenario chain. Tracing back along the timeline around the target batch, the correlation weight is calculated according to five dimensions, including time cards and test equipment. Highly correlated batches are selected to construct the test batch evolution path, analyzing the potential roles of preceding batches, and transforming yield fluctuations into a time scenario chain with preliminary causal annotations. Second, generating an event disturbance transmission sequence and performing preliminary classification. Key events from each batch log are extracted to construct a directed graph and critical path. If a key event first appears in a preceding batch, it is determined to be cross-batch transmission; if it occurs within the target batch, it is determined to be a sudden occurrence within the batch itself. This classification determines the subsequent analysis strategy. Next, performing differential analysis. In the transmission mode, a batch sequence is constructed along four dimensions, and the continuous deterioration pattern of abnormal characteristics is identified through trend slope analysis, forming a set of cross-batch abnormal propagation chains. In the sudden mode, the focus is on events within the target batch, calculating the sudden causal correlation score and analyzing the instantaneous deviation performance of the four dimensions, forming the sudden abnormal characteristic analysis results. Finally, root cause localization is performed based on validation-driven approaches and role profiling. A modular model is constructed, and role characteristics and validation indicators are calculated through propagation chain analysis or emergency problem analysis. The fusion decision output includes conclusions containing root cause identification, type, validation evidence, and scope of impact.
[0096] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the statement "including a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0097] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their likenesses.
Claims
1. A method for multidimensional data correlation analysis of probe test yield fluctuations, characterized in that, Includes the following steps: Step S1: Obtain the evolutionary context of the test batch, and construct a set of preceding trigger scenarios for yield fluctuations based on the evolutionary context of the test batch; Step S2: Extract and sort the events from the preceding trigger scene set to form a batch event sequence. Perform correlation analysis on the batch event sequence to generate an event disturbance transmission sequence and preliminary classification results. Step S3: Based on the event disturbance propagation sequence and preliminary classification results, generate cross-batch anomaly propagation chain or sudden anomaly characteristic analysis results; Step S4: Construct a root cause localization model based on verification-driven and role profiling, input the cross-batch anomaly propagation chain or sudden anomaly feature analysis results into the root cause localization model, and output the anomaly root cause results.
2. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 1, characterized in that, Obtain the evolutionary trajectory of the test batch, including: Extract target test batches from production data The identification information and the specific time points when yield fluctuations occurred. By time point Using this as a baseline, we trace back along the timeline to obtain data from point [time point]. At the appointed time All test batches completed within the time window constitute a candidate preceding batch set. Candidate Pre-batch Set Contains multiple batches, denoted as ,in This indicates the total number of candidate batches in the set; Candidate preceding batch sets are organized according to a predetermined set of rules closely integrated with the characteristics of probe testing. The batches are filtered and sorted to reconstruct the target test batch. The evolution of the test batches that had an impact.
3. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 2, characterized in that, Based on the evolution of test batches, a set of preceding trigger scenarios for yield fluctuations is constructed, including: Identify the evolution of test batches Each preceding batch For target batch Potential impact roles, through analysis of each preceding batch This is achieved through metadata and process data, with the analysis focusing on indicators directly related to probe testing. The first category consists of continuous batches that constitute the accumulation of contact states, and their roles are denoted as... The second category is the transition batch that introduces program switching disturbances, and its role is denoted as... The third category is the induced batch that amplifies environmental shifts, and its role is denoted as... ; The defining characteristic is a significant drift in temperature and humidity environmental parameters relative to process control limits during the batch's execution, which is relevant to understanding the evolution of the test batch. Each batch It is given a specific role label. The value of this tag is a set containing One element in the set will determine the evolution of the test batch. Character tags corresponding to each batch The combination constitutes the set of preceding triggering scenarios for subsequent analysis. .
4. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 3, characterized in that, The preceding trigger scene set is used to extract and sort events to form a batch event sequence, including: From the preceding trigger scene set Each batch From the corresponding process logs, equipment monitoring data, and environmental sensor data, a series of key events are extracted, and each event is abstracted into a quintuple. ,in, Indicates the timestamp of the event; This indicates the type of event, and its value comes from the set of event types mentioned above. Indicates the intensity or magnitude of an event; This indicates the test batch identifier to which the event belongs. The field is directly inherited from the batch to which the event belongs. In the preceding trigger scene set The role assigned to the middle ; For each batch Events extracted internally, according to their timestamps Arrange in ascending order to form a batch event sequence .
5. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 4, characterized in that, The batch event sequences are subjected to correlation analysis to generate event disturbance propagation sequences and preliminary classification results, including: In-batch micro-analysis, for each batch event sequence Only pairwise correlation analysis of events within the sequence is performed to identify succession and amplification relationships within a batch; in cross-batch macro-analysis, the batch is used as the basic unit, first from the event sequence of each batch. The three most representative events within a batch, or those at the beginning or end in terms of time, are selected as the characteristic events of that batch; then, a global time list is used. The global temporal location of the feature event is determined, and the delayed release and coupling triggering relationship are determined only for two batches of feature events that meet the conditions. Traverse and analyze the time-ordered list of events Construct a directed graph structure To represent the perturbation propagation relationship between all events, where, It is a collection of nodes, where each node corresponds to an event. ; It is a set of edges, where each directed edge originates from an event. Pointing to events And with weight labels; from a directed graph Extract all pre-triggered scene sets The last batch For directed paths ending at associated yield fluctuation events, calculate the ranking score of all directed paths, sort them from highest to lowest score, and select the path with the highest ranking score as the main event perturbation propagation sequence, thus forming the final event perturbation propagation sequence. The sequence of event disturbance propagation The events in the data are categorized by batch to obtain preliminary classification results.
6. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 5, characterized in that, Based on the event disturbance propagation sequence and preliminary classification results, the analysis results of cross-batch anomaly propagation chains or sudden anomaly characteristics are generated, including: Received event disturbance propagation sequence Based on the preliminary classification results, a differentiated analysis strategy is adopted, if the event disturbance transmission sequence... Classified as transitive, it tracks the continuous cross-batch propagation path of anomalous features along multiple dimensions, generating a cross-batch anomalous propagation chain; if the event perturbation propagation sequence It is classified as a sudden type, and the sudden anomaly feature analysis results are generated.
7. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 6, characterized in that, Constructing a root cause localization model based on validation-driven and role profiling includes: The root cause localization model based on verification-driven and role profiling adopts a modular and process-oriented processing architecture. Its core processing logic consists of three sequentially executed core processing sub-modules. First, the multi-dimensional role feature extraction module calculates quantitative features representing the behavioral patterns of each anomalous factor based on multiple batches of input test data. Second, the verification-driven evidence calculation module searches for subsequent corrective events and calculates quantitative evidence for yield recovery. Finally, the fusion decision and root cause determination module integrates the outputs of the preceding modules and makes a comprehensive decision based on preset rules. These three modules work together sequentially to complete the derivation from data to root cause conclusions.
8. The method for multidimensional data correlation analysis of probe test yield fluctuation according to claim 7, characterized in that, The multi-dimensional role feature extraction module includes: In this model, character profiling is achieved by calculating quantified scores across three key dimensions: origin, persistence, and amplification. Each anomalous factor receives a quantified character feature vector. ; In Mode 2, the sudden causal association score is directly adopted. As the core quantitative characteristic of abnormal factors; the instantaneous deviation magnitude of sudden abnormal events in four dimensions is calculated to form a sudden feature vector. .
9. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 8, characterized in that, The verification-driven evidence calculation module includes: The verification-driven evidence calculation module begins its work. This module searches the event log for the first corrective event directly related to the current anomaly. If a related event is found, it calculates the key verification metric, known as the inverse convergence coefficient. The calculation formula is as follows: ; in, It is a core sub-item used to quantify yield recovery speed; It is the normalized value of the standard deviation of batch yield after the corrective event, used to measure the stability of the recovery; and These are weighting coefficients, representing the importance of recovery speed and stability in the overall evaluation; Mode 1 involves a corrective event following the end of the search propagation chain; Mode 2 involves searching the target batch. List of sudden abnormal events following yield fluctuation events The first maintenance or correction operation directly related to the event type.
10. The method for multidimensional data correlation analysis of probe test yield fluctuations according to claim 9, characterized in that, Abnormal root cause results include: The output of the root cause results includes a list of root causes, which are labeled with different source categories depending on the analysis mode: For cross-batch transfer of root causes, including the trend of deteriorating contact resistance of probe cards across batches, drift of test instrument calibration parameters across multiple batches, and cumulative effects caused by defects in test program versions, corresponding inverse convergence coefficients are provided. The values serve as quantitative verification evidence and describe the specific scope of influence of each root cause; For sudden root causes within the batch itself, such as sudden abnormal modification of test program parameters, sudden temporary failure of environmental control within the batch, and sudden equipment alarms leading to test interruption and reset, a sudden causal correlation score is attached. As quantitative evidence and each root cause in the target batch Description of the time range within.