A chip detection failure mode root cause positioning method

CN122570995BActive Publication Date: 2026-09-29SHANGHAI JUYUE INSPECTION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611056770.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-29
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

[0002]随着集成电路芯片制程不断缩小、集成度持续提升,芯片内部电路耦合关系愈发复杂,工作过程中易出现时序违例、电源短路、电迁移、闩锁效应等各类失效问题

Benefits of technology

[0011]基于以上方面,本申请实施例通过时序条件熵构建单向熵涌特征,通过前置、后置双向条件熵量化节点状态不确定性的时序演化趋势,结合熵涌正负属性与幅值特征,区分原生失效萌芽节点与受扰节点,解决传统技术根因与被动失效混淆、早期隐性失效漏检的问题,从而提升芯片早期失效根因识别的精准度;同时,通过引入时间滞后核、构双向时滞传递熵谱,通过最优传播时延提取核心失效因果特征,让节点间失效驱动关系的量化结果更符合芯片真实工况;通过构建传递熵不对称系数与净信息流出量量化体系,结合第二阈值判别多节点耦合失效工况,区分主根因与次根因位置从而实现单根因和多根因精准定位。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122570995B_ABST
    Figure CN122570995B_ABST
Patent Text Reader

Abstract

The application provides a chip failure mode root cause positioning method, continuous state time series data of a chip in a failure process is acquired through a monitoring node in the chip, time series data sets corresponding to each monitoring node are constructed, one-way entropy surge indexes of each monitoring node are acquired according to conditional entropies of each monitoring node at different moments, a first threshold is combined for screening to acquire a candidate node set, a time lag kernel is introduced to candidate nodes in the candidate node set, time lag parameters in a time delay interval are traversed to acquire forward transfer entropy and reverse transfer entropy, and a bidirectional transfer entropy spectrum is constructed, a maximum value of bidirectional transfer entropy is extracted from the bidirectional transfer entropy spectrum and a transfer entropy asymmetry coefficient is acquired, net information outflow of each candidate node is acquired through the transfer entropy asymmetry coefficient, a chip failure root cause position is acquired according to the net information outflow, and the chip failure mode root cause positioning method improves the accuracy of root cause positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chip testing technology, and more specifically, to a root cause localization method for chip detection failure modes. Background Technology

[0002] As integrated circuit chip manufacturing processes continue to shrink and integration levels continue to increase, the internal circuit coupling relationships of chips become increasingly complex, making them prone to various failure problems such as timing violations, power supply short circuits, electromigration, and latch-up effects during operation.

[0003] Traditional chip failure location technologies often rely on static circuit testing, single timing feature analysis, or manual fault diagnosis, which often have some technical shortcomings. These include the inability to capture the timing evolution characteristics of chip failures, difficulty in distinguishing between passively disturbed failures and original root cause failures, lack of adaptability to chip failure propagation delay characteristics, and neglect of the hysteresis characteristics of failure propagation across nodes. They can only achieve single failure point location and cannot identify multi-root cause coupled failure scenarios.

[0004] Therefore, this invention proposes a chip failure root cause localization method based on unidirectional entropy surge coarse screening and time-delay bidirectional transfer entropy spectrum, which effectively solves the technical problems of low localization accuracy and inability to adapt to coupled failures in traditional technologies. Summary of the Invention

[0005] In view of the aforementioned technical problems of low positioning accuracy and inability to adapt to coupling failures, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a root cause localization method for chip detection failure modes, the method comprising:

[0006] The continuous state timing data of the chip during the failure process is obtained based on the monitoring nodes inside the chip, and the timing dataset corresponding to each monitoring node is constructed.

[0007] Based on the conditional entropy of each monitoring node at different times, the unidirectional entropy surge index of each monitoring node is obtained, and a candidate node set is obtained by combining it with the first threshold.

[0008] A time lag kernel is introduced into the candidate nodes in the candidate node set. The time lag parameters within the time delay interval are traversed to obtain the forward and reverse propagation entropy between candidate nodes, and a bidirectional propagation entropy spectrum between each candidate node is constructed.

[0009] Extract the maximum value of the bidirectional transfer entropy from the bidirectional transfer entropy spectrum, and obtain the transfer entropy asymmetry coefficient between candidate nodes based on the maximum value;

[0010] The net information outflow of each candidate node is obtained based on the transfer entropy asymmetry coefficient, and the root cause location of chip failure is obtained based on the net information outflow.

[0011] Based on the above, this application embodiment constructs a unidirectional entropy surge feature through temporal conditional entropy, quantifies the temporal evolution trend of node state uncertainty through pre- and post-bidirectional conditional entropy, and distinguishes between primary failure initiation nodes and disturbed nodes by combining the positive and negative attributes and amplitude characteristics of entropy surge, thus solving the problems of confusion between traditional technical root causes and passive failures, and missed detection of early latent failures, thereby improving the accuracy of early failure root cause identification of chips; at the same time, by introducing a time lag kernel and constructing a bidirectional time-delayed propagation entropy spectrum, core failure causal features are extracted through optimal propagation delay, making the quantification results of failure driving relationships between nodes more consistent with the actual operating conditions of the chip; by constructing a quantification system of propagation entropy asymmetry coefficient and net information outflow, and combining a second threshold to judge multi-node coupled failure conditions, the location of primary root cause and secondary root cause is distinguished, thereby achieving accurate location of single root cause and multiple root causes. Attached Figure Description

[0012] Figure 1 This is a flowchart of the steps of a chip failure mode detection root cause localization method according to the present invention;

[0013] Figure 2 This is a flowchart of the steps for obtaining the bidirectional transfer entropy spectrum in the root cause localization method for chip detection failure modes of the present invention. Detailed Implementation

[0014] The present invention will be further described in detail below through specific embodiments. The following embodiments are merely descriptive and not limiting, and should not be used to limit the scope of protection of the present invention.

[0015] like Figure 1 , Figure 2 As shown, a root cause localization method for chip detection failure modes includes the following steps:

[0016] Step S1: Obtain continuous state timing data of the chip during the failure process based on the monitoring nodes within the chip, and construct the timing dataset corresponding to each monitoring node.

[0017] Step S1 includes:

[0018] Step S1-1: Set up monitoring nodes in the chip's critical circuit paths, power consumption core areas, ports of easily failed devices, and upstream and downstream connection paths.

[0019] Specifically, this step involves strategically selecting high-risk, highly sensitive areas within the chip that are prone to failure propagation and deploying monitoring nodes. The critical circuit paths are the core pathways for chip signal transmission and instruction execution; the core power consumption areas are regions where chip power consumption is concentrated, heat generation is high, and electrical anomalies are likely to occur; and the ports of easily failing devices and upstream / downstream connection pathways are critical locations where failures are frequent and chain propagation is likely. By strategically deploying monitoring nodes, the state changes throughout the entire process of chip failure initiation, propagation, and evolution can be accurately captured, avoiding redundant and ineffective monitoring data.

[0020] In some possible embodiments, taking a commonly used vehicle control chip as an example, 12 monitoring nodes are deployed on the chip's internal power supply path, core logic operation circuit path, input and output ports of easily failed components such as MOSFETs and capacitors, as well as upstream and downstream connection lines. This covers the chip's core power consumption area and main signal transmission path, ensuring that any electrical abnormality occurring at any location on the chip can be captured in real time by the corresponding monitoring node.

[0021] Step S1-2: Based on the real-time acquisition of continuous state timing data during the chip failure evolution process by each monitoring node, the continuous state timing data includes node level status, signal transition timing, and real-time power consumption status.

[0022] Specifically, this step involves continuous time-series sampling of the chip's complete evolution process from normal operation, abnormal emergence, failure propagation to failure solidification, based on the pre-arranged monitoring nodes. The three types of state data collected correspond to the chip's electrical state, signal transmission state, and power consumption state, respectively, comprehensively covering the characterization features of chip failure and fully reflecting abnormal changes in different dimensions of the chip.

[0023] It should be noted that the reasons for selecting node level status, signal transition timing, and real-time power consumption status as continuous state timing data in this invention include the fact that the essential causes of early chip failure initiation and chain propagation are abnormal circuit electrical properties and abnormal signal transmission. The level status can directly reflect the on / off state and electrical deviation of the node circuit. The signal transition timing can accurately characterize failure features such as signal transmission delay, distortion, and disorder. The real-time power consumption status can reflect abnormal phenomena such as local circuit short circuit, leakage, and overload. These three types of data can completely cover the core intrinsic characteristics of chip failure.

[0024] In some possible embodiments, continuing the above case, 12 monitoring nodes synchronously collect the high and low levels of the chip during its operation, the time nodes and frequency of signal transitions, and the real-time power consumption values ​​of each area, continuously collecting the timing data of the entire process from normal operation to local abnormality and finally circuit failure.

[0025] For example, node level status timing data includes: under normal operating conditions, the node continuously outputs a standard logic level of 1.8V high and 0V low, with the timing data stable at [1.8V, 1.8V, 1.8V...]; during the chip anomaly initiation stage, this data exhibits intermittent level drops, and the timing becomes [1.8V, 1.8V, 1.2V, 0.6V...]; after complete failure, the level remains consistently low or floating, stabilizing at a fixed abnormal value; signal transition timing data includes: under normal operating conditions, the signal transition period is stable and the timing is regular, fixed at every 10ms. A transition is completed; during the failure evolution process, anomalies such as transition delay, transition disorder, and missed transitions occur, manifested as transition time intervals that fluctuate in length and shortness, with no transition signal at certain times, forming a disordered transition timing record; real-time power consumption status timing data includes that under normal operating conditions, the power consumption in the node area is stable at around 20mW with minimal timing fluctuations; when there is a local abnormality in the chip, leakage and short circuit problems cause the power consumption to continuously rise, and the timing data gradually rises to [20mW, 23mW, 28mW, 35mW...], and after complete failure, the power consumption tends to a fixed peak value or drops sharply to 0.

[0026] Steps S1-3 involve performing time-series alignment and noise reduction preprocessing on the collected continuous state time-series data according to a uniform sampling frequency.

[0027] Specifically, due to slight differences in the hardware acquisition accuracy of each monitoring node, and the presence of noise such as electromagnetic interference and environmental interference during chip operation, the original acquired data suffers from issues such as timing misalignment, abnormal noise, and invalid data. This step aligns the original time-series data of all monitoring nodes by unifying and fixing the sampling frequency, ensuring that the time dimension of the data from all nodes is consistent. At the same time, noise reduction processing is used to remove abnormal noise and invalid data caused by random interference, avoiding the distortion of entropy value calculation and feature discrimination errors caused by noisy data.

[0028] In some possible embodiments, continuing the above case, the original timing data collected from 12 monitoring nodes are uniformly set to a sampling frequency of 100Hz to complete timing alignment and correct the timing deviation of each node's data; at the same time, invalid noise data such as instantaneous level jumps and power consumption changes caused by electromagnetic interference are removed to obtain smooth, time-uniform, and continuous state timing data without abnormal interference.

[0029] Steps S1-4: Based on the preprocessed continuous state time series data, construct the time series dataset for each monitoring node.

[0030] Specifically, this step builds a standardized data foundation for subsequent algorithm operations. For each monitoring node, the preprocessed continuous state time series data is independently integrated to form a structured time series dataset exclusive to each node. Each dataset is stored and processed independently, which can accurately correspond to the state evolution characteristics of a single monitoring node, avoid data confusion and interference from multiple nodes, and provide data support for subsequent steps.

[0031] In some possible embodiments, continuing the above case, 12 independent time series datasets are constructed for the preprocessed data of the 12 monitoring nodes. Each dataset contains the level, signal transition, and power consumption time series data of the corresponding node at consecutive moments.

[0032] For example, taking the time-series dataset of monitoring node No. 1 in the failure area as an example, the dataset uses the sampling time as the time-series index, and each time point corresponds to a set of three-dimensional state samples. The structure of a single set of time-series data can be concretely represented as time 1: (1.8V, normal jump, 20mW), time 2: (1.8V, normal jump, 20mW), time 3: (1.6V, slight delayed jump, 24mW), time 4: (1.2V, disordered jump, 29mW), time 5: (0.7V, no jump, 36mW), and so on, forming a continuous time-series sequence. Following the same scheme in this step, the remaining 11 monitoring nodes are integrated, and their respective time-series evolution samples are stored. The data of normal nodes maintains a stable numerical sequence in the long term, while the data of abnormal nodes continuously deteriorates with the failure evolution, ensuring that each set of datasets can independently and completely represent the real working state and time-series abnormal trend of the corresponding monitoring node.

[0033] Step S2: Based on the conditional entropy of each monitoring node at different times, obtain the unidirectional entropy surge index of each monitoring node, and filter it in combination with the first threshold to obtain a candidate node set.

[0034] Step S2 includes:

[0035] Step S2-1: Based on each monitoring node, obtain the precondition entropy and postcondition entropy at different times within the preset time window.

[0036] Specifically, the conditional entropy is used to quantify the uncertainty changes in node states and is a characteristic of chip abnormal emergence. In this step, for the independent time series dataset of each monitoring node, continuous time series data within a preset fixed time window are extracted, and the preconditional entropy and postconditional entropy corresponding to each sampling moment are calculated respectively. Among them, the preconditional entropy represents the uncertainty constraint of the node's historical state on the current state, and the postconditional entropy represents the uncertainty constraint of the node's current state on the future state. By using the two-way conditional entropy, the evolution law of the node's time series state can be completely captured, and the characteristics of node abnormal emergence can be accurately identified.

[0037] Furthermore, the steps for obtaining the preconditional entropy and postconditional entropy include: for the preprocessed time series dataset of each monitoring node, extracting time series data of fixed duration frame by frame in a sliding window manner to ensure that each frame window contains multiple sets of node state samples that are continuous and time-aligned; for the time series sequence within the current window, extracting the state variables of historical moments and the state variables of the current moment, fitting the probability distribution from the historical state to the current state based on a probabilistic statistical model, and substituting it into the conditional entropy formula to calculate the preconditional entropy, which is used to characterize the uncertainty of the constraint of the historical time series on the current node state; within the same time window, extracting the state variables of the current moment and the state variables of the future moment, fitting the probability distribution of the evolution from the current state to the future state, and similarly calculating the postconditional entropy, which is used to characterize the uncertainty constraint of the current node state on the evolution of the future state, and traversing all sampling points within the window time by time to complete the bidirectional conditional entropy solution in the full time series dimension.

[0038] Understandably, the above conditional entropy formula can be specifically expressed as:

[0039] ;

[0040] Among them, the Represented as conditional entropy, the Represented as the prior probability of historical states, the This is represented as the conditional probability of state transition.

[0041] Preconditional entropy is expressed as The postconditional entropy is expressed as .

[0042] In some possible embodiments, taking the actual sampling time-series data of the aforementioned monitoring node No. 1 as an example, normalized state samples from five consecutive sampling times are extracted for calculation. The node state sequence at each time point includes t1 (normal steady state), t2 (normal steady state), t3 (slight anomaly), t4 (moderate anomaly), and t5 (severe anomaly). Taking t3 as an example, the prior probability of obtaining the historical state of t2 as normal steady state is 1.0, the conditional probability of maintaining steady state under the premise of obtaining normal steady state is 0.85, and the conditional probability of transitioning to slight anomaly is 0.15. Substituting these values ​​into the equation... The conditional entropy formula yields a preconditional entropy of 0.61. Similarly, the prior probability of obtaining a historical state where t3 is a normal steady state is 0.85, the prior probability of obtaining a historical state where t3 is slightly abnormal is 0.15, the conditional probability of maintaining the steady state under the premise of obtaining a normal steady state at the next moment is 0.85, the conditional probability of transitioning to a slightly abnormal state is 0.15, the conditional probability of maintaining the slightly abnormal state under the premise of obtaining a slightly abnormal state is 0.85, the conditional probability of transitioning to a moderately abnormal state is 0.15, and substituting these into the conditional entropy formula yields a postconditional entropy of 0.64.

[0043] It should be noted that all prior probabilities and conditional probabilities of states involved in this step are obtained from the frequency statistics of real time-series state samples within the current sliding time window. The specific acquisition process includes: statistically analyzing the occurrence frequency of different node states within the window, including normal steady state, slight anomaly, moderate anomaly, and severe anomaly, and obtaining the prior probability of each state after normalization; statistically analyzing the pairwise state transition frequency of historical states and current states within the window, and obtaining the corresponding conditional probabilities of state transition after normalization. All probability values ​​are consistent with the actual time-series evolution of the chip.

[0044] Similarly, the ternary joint probability in step S3-2 of this invention is also obtained through the same method described above.

[0045] Step S2-2: Obtain the unidirectional entropy surge index based on the preconditional entropy and postconditional entropy.

[0046] Specifically, the one-way entropy surge index is used to quantify the temporal evolution trend of node state uncertainty and distinguish between passively disturbed nodes and natively anomalous nodes. In this step, the one-way entropy surge index is calculated by the difference between the post-conditional entropy and the pre-conditional entropy. The sign and magnitude of the entropy surge value directly reflect the anomalous attributes of the node. A positive entropy surge indicates that the node state uncertainty has increased and belongs to the passively affected anomaly. A negative entropy surge indicates that the node state uncertainty has converged and solidified and belongs to the native anomaly of failure initiation.

[0047] In some possible embodiments, continuing the above case, the unidirectional entropy surge of 12 monitoring nodes is obtained by the difference between the post-conditional entropy and the pre-conditional entropy. The entropy surge value of normal nodes is close to 0, the entropy surge of nodes affected by abnormalities is positive, and the entropy surge of nodes where original failures occur is negative.

[0048] For example, for the t3 failure initiation critical point, the preconditional entropy of 0.61 represents a stable steady state within the window, with strong constraints from historical states and low overall uncertainty; the postconditional entropy of 0.64 represents a small number of initial anomalies within the window, after which the probability of future state degradation of the node increases significantly, and the overall time-series uncertainty increases slightly; finally, the unidirectional entropy surge is 0.03, showing a positive weak entropy surge characteristic, which accurately matches the evolution law of chip failure initiation, low probability anomaly initiation, and slow instability.

[0049] Step S2-3: Set the first threshold and filter the candidate nodes by combining the unidirectional entropy surge index.

[0050] Specifically, monitoring nodes whose unidirectional entropy surge index is greater than the first threshold are marked as affected area nodes, and monitoring nodes whose unidirectional entropy surge index is less than the opposite of the first threshold are marked as failed area nodes. Monitoring nodes that are biased towards the failed area and the boundary between the affected area and the failed area, i.e., monitoring nodes whose unidirectional entropy surge index is greater than or equal to the opposite of the first threshold and less than 0, are selected to form a candidate node set.

[0051] Specifically, this step completes the hierarchical screening of all nodes through quantification thresholds, eliminating normal nodes and invalid passive nodes, and obtaining candidate nodes suspected of being the root cause of failure; combined with the corrected one-way entropy surge index pattern of this invention, the one-way entropy surge index of passively disturbed nodes is positive, and the one-way entropy surge index of the original failure initiation root cause node is negative; a reasonable first threshold adapted to this scenario is set to 0.05, which can accurately adapt to the early weak entropy surge characteristics and distinguish various types of nodes: nodes with entropy surge greater than 0.05 are judged as passively affected area nodes, and nodes with entropy surge less than -0.05 are judged as completely failed area nodes; In the transition boundary between the two types of regions, the interval [0,0.05] belongs to the weakly disturbed transition sub-interval, representing that the node is in the early stage of failure budding, only generating weak uncertainty fluctuations, without complete failure or passive strong disturbance; the interval [-0.05,0) indicates that the node is not passively disturbed, but its own state uncertainty gradually converges and defects slowly take shape, with an active failure budding trend, belonging to the typical initial primary root cause node; therefore, the node with entropy surge in the interval [-0.05,0) is selected as the candidate node, the node in this interval is the source of failure budding, the entropy surge is weak and negative, and is the most likely root cause of failure.

[0052] It should be noted that the entropy surge value of a normally functioning chip node approaches 0 infinitely, and the fluctuation amplitude is generally less than 0.05; the entropy surge of a passively strongly disturbed node is stably greater than 0.05, and the entropy surge of a completely failed and solidified node is stably less than -0.05; therefore, this invention selects 0.05 as the first threshold for unidirectional entropy surge screening. The threshold of 0.05 can accurately distinguish between four states: normal steady-state fluctuation, weak nascent primary failure, passive disturbance anomaly, and completely failed and solidified. This can filter out the small numerical fluctuation interference during the normal operation of the chip, avoid normal nodes being mistakenly screened as candidate nodes, and at the same time find the most likely root cause node of failure.

[0053] In some possible embodiments, continuing the above case, and combining the actual one-way entropy surge index of the instance, the first threshold of 0.05 is adjusted and adapted to obtain passively affected nodes with a one-way entropy surge index greater than 0.05 and completely failed nodes with a one-way entropy surge index less than -0.05; further, nodes with a one-way entropy surge index in the range of [-0.05, 0) and with a weak negative entropy surge are selected as candidate node sets to obtain the original root cause nodes in the early stage of chip failure.

[0054] For example, the unidirectional entropy surge index at monitoring node t3 is 0.03, which falls within the (0, 0.05) range and does not belong to the candidate node set. It should be noted that a unidirectional entropy surge index greater than 0.05 corresponds to passively affected nodes. These nodes have no primary failure initiation and are only affected by the propagation disturbances of surrounding failed nodes, resulting in increased state uncertainty and a significantly positive entropy surge characteristic. For example, peripheral signal path nodes of a chip exhibit slight timing disorder after abnormal propagation from the core node, with entropy surge values ​​such as 0.08 and 0.12, significantly deviating from 0. A unidirectional entropy surge index less than -0.05 corresponds to completely failed nodes. These nodes have completed failure initiation and degradation solidification, with stable circuit defects, significantly reduced state uncertainty, and fully formed failure characteristics. For example, severely failed chip nodes exhibit entropy surge values ​​such as -0.07 and -0.10, significantly deviating from 0. A unidirectional entropy surge index in the [-0.05, 0.05] range indicates a transitional period of failure initiation. The entropy surge of 0.03 at time t3 in this example belongs to the weak disturbance transition sub-interval of [0, 0.05], indicating that the node is in the early stage of failure budding, with only weak uncertainty fluctuations, no complete failure, and no passive strong disturbances. This is a typical characteristic of early primary failure budding. The unidirectional entropy surge index is in the range of [-0.05, 0). Taking the failure scenario of the automotive chip as an example, the core power switch node of the chip is in the early stage of failure budding. The timing state is still mainly in a normal steady state, but the internal circuit has generated weak electromigration latent defects. The unidirectional entropy surge calculated at time t4 of this node is -0.028, which falls in the range of [-0.05, 0). A total of 3 candidate nodes were selected. This negative weak entropy surge indicates that the node is not passively disturbed, but its own state uncertainty gradually converges and defects slowly form, which has an active failure budding trend. It belongs to a typical early primary root cause node.

[0055] Step S3: Introduce time lag kernels to the candidate nodes in the candidate node set, traverse the time lag parameters within the time delay interval, obtain the forward and reverse propagation entropy between candidate nodes, and construct the bidirectional propagation entropy spectrum between each candidate node.

[0056] Step S3 includes:

[0057] Step S3-1: Based on any pair of candidate nodes in the candidate node set, preset the maximum delay value, set the delay interval based on the maximum delay value, introduce a time lag kernel, and traverse all time lag parameters within the delay interval.

[0058] Specifically, chip failure propagation involves physical delay and is not instantaneous synchronous transmission. This step targets candidate node pairs after coarse screening, presets a maximum delay value to define the effective propagation delay range of chip failure, and forms a complete delay interval. A time lag kernel is introduced to correct the propagation entropy calculation logic to adapt to the lag characteristics of failure propagation. By traversing all delay parameters within the interval, all possible delay scenarios of failure propagation are covered.

[0059] In some possible embodiments, continuing the above case, for the three selected candidate nodes, the maximum latency value is preset to 10, the latency interval is set to [1,10], and the time lag kernel correction calculation logic is introduced to sequentially traverse the lag time = 1, 2, 3...10, thereby covering all effective latency scenarios of chip failure propagation.

[0060] It should be noted that, considering the circuit transmission characteristics, sampling frequency, and failure propagation physics of the vehicle control chip, this invention sets the maximum delay value to 10 and the effective delay interval to [1, 10]. The sampling frequency of this scheme is 100Hz, with a single sampling interval corresponding to a timing accuracy of 10ms. Considering the inherent hysteresis characteristics of signal transmission and the chain propagation of electrical defects within the chip, the effective delay from the local germination of an anomaly to the completion of cross-node failure propagation is generally concentrated in the 10ms-100ms range, which corresponds precisely to 1-10 sampling steps. When the delay is less than 1... When the time interval is too short, the node state does not evolve effectively, there are no effective failure propagation characteristics, and the calculation results are of no reference value. When the time delay is greater than 10, it exceeds the effective time window for the rapid propagation of conventional chip failures. Most of the time delays are steady-state fluctuations after the failures have solidified in the later stages, resulting in a large amount of invalid and redundant data, which can easily interfere with the accuracy of causal relationship quantification. Therefore, the time delay interval [1, 10] is set, which can not only fully cover all the effective time delay characteristics of the early failure initiation and chain propagation of the chip, but also eliminate invalid time sequence interference and reduce computing power, thereby ensuring the accuracy of causal quantification of entropy.

[0061] Step S3-2: Based on the transfer entropy formula, obtain the forward transfer entropy and the reverse transfer entropy between any pair of candidate nodes.

[0062] The transfer entropy formula is expressed as:

[0063] ;

[0064] Among them, the Represented as propagation entropy, indicating time lag is... At that time, candidate nodes For candidate nodes The state evolution provides an effective causal information gain; Represented as discrete sampling time; the Represented as a time lag parameter, it indicates the chip failure state from the candidate node. propagate to candidate nodes The physical propagation delay; Indicates candidate nodes At any moment Continuous state time-series data; the Represented as candidate nodes At any moment Continuous state time-series data; the Represented as candidate nodes At the time delay Continuous state time-series data; the The ternary joint probability is expressed as the statistical probability of the simultaneous occurrence of three sets of continuous state time series data; This is a second-order conditional probability, expressed as follows: given... In this state, The conditional probability of the state occurring; the Let be a first-order conditional probability, expressed as known. In this state, The conditional probability of a state occurring.

[0065] Specifically, this step quantifies the bidirectional causal propagation strength between any two candidate nodes using the forward propagation entropy formula. The forward propagation entropy characterizes the failure-driving ability of node i to node j, and the backward propagation entropy characterizes the failure-driving ability of node j to node i. The formula is constructed based on the preprocessed node time series data. By calculating the difference between the ternary joint probability and the conditional probability, the interference of the node's own time series fluctuations is removed, the effective information gain between nodes is extracted, and the unidirectional propagation causal relationship of failure is quantified.

[0066] In some possible embodiments, taking the node pair consisting of the selected candidate root cause node i and the passively disturbed node j as an example, the propagation entropy formula is calculated by traversing the time delay interval τ=1-10 one by one, and the actual quantized values ​​under different time delays are extracted. When the time lag parameter τ=3, the ternary joint probability and conditional probability are solved by substituting into the node time sequence state probability distribution, and the forward propagation entropy is obtained as 0.426, which means that the original failed node i has a very strong failure causal driving force and information gain for node j; and the reverse propagation entropy under the corresponding simultaneous delay is obtained as 0.038, which has a very low reverse information gain and almost no reverse driving capability; similarly, when τ=2, the forward propagation entropy is obtained as 0.315 and the reverse propagation entropy under the corresponding simultaneous delay is obtained as 0.032; when τ=4, the forward propagation entropy is obtained as 0.382 and the reverse propagation entropy under the corresponding simultaneous delay is obtained as 0.041, etc.

[0067] It should be noted that the process of obtaining the forward and reverse propagation entropy includes: statistically analyzing the proportion of samples in the full time series that simultaneously satisfy the three states of node i at time t (slight anomaly), node j at time t (normal steady state), and node i deteriorating to moderate anomaly at time t+3), representing the synchronous occurrence probability of this set of causal time series states, i.e., the ternary joint probability of 0.82; and considering the conditional probability of state transition from slight anomaly at time t and normal steady state at time j, the state transition probability of node i deteriorating to moderate anomaly at time t+3, reflecting the failure evolution probability under multi-node state coupling. That is, the second-order conditional probability is 0.94; given that i is slightly abnormal at time t, the independent transition probability of i autonomously deteriorating to moderate abnormality at time t+3, after removing the coupling influence of node j, reflects the failure evolution trend of a single node itself, i.e., the first-order conditional probability is 0.31; the contribution value of this main term is obtained as 1.313 through the transfer entropy formula; and the contribution value of the remaining non-dominant state summation terms is obtained as -0.887, and the final result of the global summation is 0.426; and the reverse transfer entropy under the same time delay is obtained as 0.038 using the same method as above;

[0068] Furthermore, using the same method as in this step, obtain the forward propagation entropy and directional propagation entropy under other time delays.

[0069] Step S3-3: Integrate the forward and reverse propagation entropies corresponding to each delay to obtain the forward and reverse propagation entropy sequences, and construct the bidirectional propagation entropy spectrum between candidate nodes.

[0070] In some possible embodiments, continuing the above case, all forward and reverse transfer entropy values ​​corresponding to τ=1 to τ=10 are integrated to generate forward transfer entropy sequences and reverse transfer entropy sequences for each node pair, and the bidirectional transfer entropy spectrum between the three candidate nodes is constructed based on the two sets of sequences.

[0071] Step S4: Extract the maximum value of the bidirectional transfer entropy from the bidirectional transfer entropy spectrum, and obtain the transfer entropy asymmetry coefficient between candidate nodes based on the maximum value.

[0072] Step S4 includes:

[0073] Step S4-1: Based on the bidirectional transfer entropy spectrum, traverse the bidirectional transfer entropy corresponding to all time lag parameters respectively, and extract the forward maximum transfer entropy and the reverse maximum transfer entropy.

[0074] Specifically, the propagation intensity varies with different time delays in the bidirectional propagation entropy spectrum. The maximum propagation entropy corresponds to the optimal time delay and the strongest driving intensity of failure propagation between nodes. This step traverses all bidirectional propagation entropy values ​​corresponding to time delays in the entropy spectrum, extracts the peak values ​​of forward propagation and reverse propagation respectively, eliminates invalid weak propagation features, retains the most core failure propagation features between nodes, and avoids weak time delay data interfering with the accuracy of causal discrimination.

[0075] In some possible embodiments, continuing the above case, the selected candidate nodes i, j, and k are selected to form three candidate node pairs. The bidirectional propagation entropy spectrum data with full time delay from τ=1 to 10 is traversed, and the forward maximum propagation entropy and reverse maximum propagation entropy of each node pair are extracted to lock the strongest failure propagation feature of each node pair. For example, for candidate nodes i and j, when traversing the full time delay entropy spectrum data, the forward propagation intensity reaches its peak at τ=3, and the forward maximum propagation entropy is 0.426; the corresponding reverse entropy values ​​at each delay are extremely low, and the reverse peak value is obtained at τ=3, with a reverse maximum propagation entropy of 0.038; for candidate nodes i and k, the optimal propagation delay is also τ=3, the forward maximum propagation entropy is 0.392, and the reverse maximum propagation entropy is 0.035; for candidate nodes k and j, the overall propagation intensity is relatively weak, the optimal delay is τ=4, the forward maximum propagation entropy is 0.113, and the reverse maximum propagation entropy is 0.096.

[0076] Step S4-2: Obtain the asymmetric coefficient of the transfer entropy between candidate nodes based on the difference of the bidirectional maximum transfer entropy.

[0077] Specifically, chip failure propagation exhibits unidirectional causal asymmetry, with the driving strength from the root cause node to surrounding nodes being much greater than the reverse propagation strength. This step calculates the propagation entropy asymmetry coefficient by using the difference between the bidirectional maximum propagation entropy. The positive or negative sign of the coefficient directly represents the causal direction of propagation, and the absolute value of the coefficient represents the causal driving strength, thereby achieving quantitative modeling of the causal relationship of failure propagation between nodes.

[0078] In some possible embodiments, continuing the above example, the asymmetry coefficient is obtained by the difference between the forward maximum transfer entropy and the reverse maximum transfer entropy. This quantifies the asymmetry strength of the causal drive of node failures. For example, the transfer entropy asymmetry coefficient of candidate node pair (ij) is 0.426 - 0.038 = 0.388; the transfer entropy asymmetry coefficient of candidate node pair (ik) is 0.392 - 0.035 = 0.357; and the transfer entropy asymmetry coefficient of candidate node pair (kj) is 0.113 - 0.096 = 0.017. By examining the asymmetric coefficients of the transfer entropy of the candidate node pairs, we found that the asymmetric coefficient of candidate node pair (ij) is significantly positive, proving that node i has a very strong unidirectional failure driving ability for node j and is the active failure source. The asymmetric coefficient of the transfer entropy of candidate node pair (ik) also shows a significant positive asymmetric characteristic, further confirming the broad-spectrum failure propagation ability of node i. The asymmetric coefficient of the transfer entropy of candidate node pair (kj) is close to 0, proving that the causal driving differences between the nodes in this group are minimal and there is no obvious unidirectional propagation dominant relationship.

[0079] Step S5: Obtain the net information outflow of each candidate node based on the transfer entropy asymmetry coefficient, and obtain the chip failure root cause location based on the net information outflow.

[0080] Step S5 includes:

[0081] Step S5-1: Traverse the asymmetric coefficients of the propagation entropy between each candidate node and the other candidate nodes in the candidate node set, remove the negative propagation entropy asymmetric coefficients, and retain the positive propagation entropy asymmetric coefficients.

[0082] It should be noted that positive asymmetric coefficients represent the current node's effective failure driving ability on other nodes, while negative asymmetric coefficients represent the current node being driven by other nodes, which are passively disturbed features and have no root cause reference value. Therefore, this step iterates through all asymmetric coefficients between nodes, removes negative interference coefficients, retains only the effective positive driving coefficients, and filters the outward propagation driving features of candidate nodes to avoid passive features interfering with root cause identification.

[0083] In some possible embodiments, continuing the above case, the asymmetric coefficients between each pair of the three candidate nodes are traversed. If all the asymmetric coefficients of the candidate node pairs are found to be valid positive driving coefficients, then there is no need to eliminate them, and all valid asymmetric coefficients are retained.

[0084] Step S5-2: Accumulate all forward propagation entropy asymmetry coefficients to obtain the net information outflow of candidate nodes.

[0085] Specifically, the net information outflow is a comprehensive quantitative indicator of the overall failure propagation driving capability of a node. This step obtains the total external driving strength of a node by accumulating all positive asymmetric coefficients of the node. The larger the net information outflow value, the stronger the node's ability to drive abnormal disturbances to surrounding nodes, and the greater the probability that it is the root cause node of failure initiation, thus realizing the quantitative characterization of the overall failure attributes of a single node.

[0086] In some possible embodiments, continuing the above case, the asymmetric coefficient of the propagation entropy of candidate node pair (ij) is 0.426-0.038=0.388; the asymmetric coefficient of the propagation entropy of candidate node pair (ik) is 0.392-0.035=0.357; and the asymmetric coefficient of the propagation entropy of candidate node pair (kj) is 0.113-0.096=0.017. The positive asymmetric coefficients of the three candidate nodes are accumulated one by one to accurately calculate the net information outflow of each node and quantify the overall failure propagation driving capability of each node; candidate node i drives externally... Both sets of asymmetric coefficients for moving nodes j and k are positive and valid values, therefore the net information outflow of candidate node i is 0.745; candidate node j has no externally valid positive driving coefficient, therefore the net information outflow of candidate node j is 0; candidate node k has only a weak positive asymmetric coefficient driving node j, therefore the net information outflow of candidate node k is 0.017; it is found that the net information outflow of node i is much higher than the other two nodes, possessing a very strong active failure propagation driving capability, node k has only a very weak driving capability, and node j has no autonomous failure driving capability.

[0087] Step S5-3: Compare the net information outflow of all candidate nodes, select the candidate node with the largest net information outflow, and determine the candidate node corresponding to the largest net information outflow as the root cause of chip failure.

[0088] It should be noted that the net information outflow is the total gain of failure causal information output by the node, which quantitatively represents the comprehensive driving ability of the node to actively induce anomalies in surrounding nodes, thereby determining the node corresponding to the maximum net information outflow as the main root cause node.

[0089] Understandably, in the chip failure propagation system, the root cause node is the initial source of abnormal disturbances. No external node can effectively drive its failure, but the latent defects that emerge within it will continuously output abnormal information and drive the surrounding normal nodes to deteriorate step by step. It has the characteristics of unidirectional, active, and broad-spectrum failure propagation and can accumulate to form a huge positive net information outflow. On the other hand, passively disturbed nodes and weakly disturbed transition nodes can only receive external abnormal disturbances and have extremely weak autonomous external driving capabilities, with net information outflow approaching 0. Therefore, among all candidate nodes, the node corresponding to the maximum net information outflow has the strongest external failure propagation driving capability and the highest priority for initiating abnormalities. It is the core source that first initiates failure and drives the global failure spread, and can be accurately determined as the location of the main root cause of chip failure.

[0090] In some possible embodiments, continuing the above case, the measured values ​​of the net information outflow of the three candidate nodes are compared. The net information outflow of candidate node i is 0.745, the net information outflow of candidate node k is 0.017, and the net information outflow of candidate node j is 0. Among them, the net information outflow value of node i is much larger than that of the other two candidate nodes, and it has the dominant ability to drive the failure to propagate outward. Therefore, the chip circuit location corresponding to node i can be accurately determined as the root cause location of this chip failure.

[0091] Step S5-4: Set a second threshold. If the difference between the net information outflow and the maximum net information outflow in the remaining candidate nodes is less than or equal to the second threshold, it is determined to be a multi-root cause coupling failure, and the corresponding candidate node is determined to be the secondary root cause location of the chip failure.

[0092] Understandably, net information outflow directly characterizes the node's ability to autonomously propagate failures. In a single root cause failure scenario, only the primary root cause node has a very large net information outflow, while the net outflow of other passive and weakly disturbed nodes is close to 0, with a large difference from the maximum net outflow. If the difference between the net information outflow of other candidate nodes and the maximum net information outflow is very small and meets the threshold condition, it indicates that the node is not a passive failure caused by the disturbance of the primary root cause node, but rather it also has its own independent latent defects and autonomous failure initiation and propagation capabilities, participating in failure evolution synchronously with the primary root cause node and jointly driving the abnormal diffusion of the entire chip. Therefore, it can be determined as a multi-root cause coupled failure, and the corresponding node is the location of the secondary root cause of the failure.

[0093] It should be noted that, in order to adapt to the optimal critical threshold for early multi-node weak coupling failure of the chip, the second threshold is set to 0.05 in this invention; the fluctuation amplitude of the net outflow error caused by the chip's normal operating conditions and passive disturbances is less than 0.05; if the difference between the node's net outflow and the maximum value is less than or equal to 0.05, it indicates that the node's driving capability is very close to that of the main root cause node, and it does not belong to the weak failure characteristics of passive disturbance, and has independent root cause attributes; if the difference is greater than 0.05, it indicates that the node's external driving capability is much weaker than that of the main root cause node, and it is only a secondary failure node that is passively propagated, and has no independent root cause value.

[0094] In some possible embodiments, continuing the above case, the second threshold is set to 0.05, and the operating mode is determined by combining the actual net information outflow value of the candidate nodes:

[0095] If, as in step S5-3, the maximum net information outflow is 0.745, and the difference between it and other net information outflows is greater than the second threshold of 0.05, it proves that other candidate nodes do not have independent strong failure driving capabilities. Therefore, this failure is determined to be a single primary cause failure, and only node i is the location of the primary cause of failure.

[0096] If the net information outflow of candidate node m is 0.745, the net information outflow of candidate node n is 0.712, and the net information outflow of candidate node v is 0; the maximum net information outflow is 0.745. The difference between the secondary root cause nodes is calculated as 0.745-0.712=0.033, which is less than 0.05, satisfying the second threshold judgment condition. This proves that the external failure driving capabilities of candidate node n and primary root cause node m are similar. Node n is not passively disturbed and has its own independent hidden failure defect, which can autonomously drive the abnormal spread of surrounding nodes and jointly induce the chip's global failure with node m. Therefore, it can be determined as a dual root cause coupling failure, with node m as the primary root cause position and node n as the secondary root cause position.

[0097] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0098] It should be understood that, in the embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0099] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A root cause localization method for chip detection failure modes, characterized in that, It includes the following steps: The continuous state timing data of the chip during the failure process is obtained based on the monitoring nodes inside the chip, and the timing dataset corresponding to each monitoring node is constructed. Based on the conditional entropy of each monitoring node at different times, the unidirectional entropy surge index of each monitoring node is obtained, and a candidate node set is obtained by combining it with the first threshold. Specifically, based on each monitoring node, the preconditional entropy and postconditional entropy at different times within a preset time window are obtained; the one-way entropy surge index is obtained based on the preconditional entropy and postconditional entropy; this step calculates the one-way entropy surge index by the difference between the postconditional entropy and the preconditional entropy. A time lag kernel is introduced into the candidate nodes in the candidate node set. The time lag parameters within the time delay interval are traversed to obtain the forward and reverse propagation entropy between candidate nodes, and a bidirectional propagation entropy spectrum between each candidate node is constructed. Extract the maximum value of the bidirectional transfer entropy from the bidirectional transfer entropy spectrum, and obtain the transfer entropy asymmetry coefficient between candidate nodes based on the maximum value; Specifically, based on the bidirectional transfer entropy spectrum, the bidirectional transfer entropy corresponding to all time lag parameters is traversed to extract the forward maximum transfer entropy and the reverse maximum transfer entropy; the transfer entropy asymmetry coefficient between candidate nodes is obtained based on the difference between the bidirectional maximum transfer entropy. The net information outflow of each candidate node is obtained based on the transfer entropy asymmetry coefficient, and the location of the root cause of chip failure is obtained based on the net information outflow. Among them, the asymmetric coefficient of the propagation entropy between each candidate node and the other candidate nodes in the candidate node set is traversed, the negative propagation entropy asymmetric coefficient is eliminated, and the positive propagation entropy asymmetric coefficient is retained; The net information outflow of candidate nodes is obtained by summing up all forward propagation entropy asymmetry coefficients.

2. The root cause localization method for chip detection failure modes according to claim 1, characterized in that, Based on the monitoring nodes within the chip, continuous state timing data of the chip during the failure process is acquired, and a timing dataset corresponding to each monitoring node is constructed, including: Monitoring nodes are deployed along the critical circuit paths of the chip, the core power consumption areas, the ports of easily failed devices, and the upstream and downstream connection paths. Based on the real-time acquisition of continuous state timing data during the chip failure evolution process by each monitoring node, the continuous state timing data includes node level status, signal transition timing, and real-time power consumption status. The collected continuous state time-series data are preprocessed with time alignment and noise reduction according to a uniform sampling frequency; Based on the preprocessed continuous state time series data, a time series dataset for each monitoring node is constructed.

3. The root cause localization method for chip detection failure modes according to claim 1, characterized in that, Based on the conditional entropy of each monitoring node at different times, the unidirectional entropy surge index of each monitoring node is obtained, and then filtered using a first threshold to obtain a candidate node set, including: Set a first threshold and use it in conjunction with the unidirectional entropy surge index to filter and obtain a set of candidate nodes; Specifically, monitoring nodes whose unidirectional entropy surge index is greater than the first threshold are marked as affected area nodes, and monitoring nodes whose unidirectional entropy surge index is less than the opposite of the first threshold are marked as failed area nodes. Monitoring nodes that are biased towards the failed area and the boundary between the affected area and the failed area, i.e., monitoring nodes whose unidirectional entropy surge index is greater than or equal to the opposite of the first threshold and less than 0, are selected to form a candidate node set.

4. The root cause localization method for chip detection failure modes according to claim 1, characterized in that, A time lag kernel is introduced into the candidate nodes in the candidate node set. The time lag parameters within the time delay interval are traversed to obtain the forward and backward propagation entropies between candidate nodes. A bidirectional propagation entropy spectrum between each candidate node is constructed, including: Based on any pair of candidate nodes in the candidate node set, a maximum delay value is preset, a delay interval is set based on the maximum delay value, a time lag kernel is introduced, and all time lag parameters within the delay interval are traversed. Based on the transfer entropy formula, the forward transfer entropy and the reverse transfer entropy between any pair of candidate nodes are obtained respectively. By integrating the forward and reverse propagation entropies corresponding to each delay, the forward and reverse propagation entropy sequences are obtained, and a bidirectional propagation entropy spectrum between candidate nodes is constructed.

5. The root cause localization method for chip detection failure modes according to claim 4, characterized in that, The formula for transfer entropy includes: The transfer entropy formula is expressed as: ; Among them, the Represented as propagation entropy, indicating time lag is... At that time, candidate nodes For candidate nodes The state evolution provides an effective causal information gain; Represented as discrete sampling time; the Represented as a time lag parameter, it indicates the chip failure state from the candidate node. propagate to candidate nodes The physical propagation delay; Indicates candidate nodes At any moment Continuous state time-series data; the Represented as candidate nodes At any moment Continuous state time-series data; the Represented as candidate nodes At the time delay Continuous state time-series data; the The ternary joint probability is represented as the statistical probability of the simultaneous occurrence of three sets of continuous state time series data; the... This is a second-order conditional probability, expressed as follows: given... In this state, The conditional probability of the state occurring; the Let be a first-order conditional probability, expressed as known. In this state, The conditional probability of a state occurring.

6. The root cause localization method for chip detection failure modes according to claim 1, characterized in that, The net information outflow of each candidate node is obtained based on the transfer entropy asymmetry coefficient, and the root cause location of chip failure is obtained based on the net information outflow, including: Compare the net information outflow of all candidate nodes, filter out the candidate node with the largest net information outflow, and determine the candidate node corresponding to the largest net information outflow as the root cause of chip failure. A second threshold is set. If the difference between the net information outflow and the maximum net information outflow in the remaining candidate nodes is less than or equal to the second threshold, it is determined to be a multi-root cause coupling failure, and the corresponding candidate node is determined to be the secondary root cause location of chip failure.

Citation Information

Patent Citations

  • Chip-level integrated circuit fault diagnosis method based on improved wavelet entropy

    CN119322249A

  • SoC fault diagnosis method, system, device and medium

    CN121212034A