Deep learning based integrated circuit fault diagnosis and early warning system
Patent Information
- Application Number
- CN202511894588.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-12-16
AI Technical Summary
[0002]随着电子技术的飞速发展,数字集成电路已广泛应用于工业控制、通信、汽车电子及消费电子等关键领域,成为现代信息社会的核心基础;随着工艺节点不断微缩与集成度持续提升,芯片在制造和运行过程中面临日益严峻的可靠性挑战;例如,制造环节中的光刻偏差、蚀刻不均等工艺波动,运行阶段的温度波动、电压漂移等环境变化,以及长期使用中的元件老化、电磁干扰等因素,均可能诱发各类电路故障;在5G通信、自动驾驶等对系统可靠性要求极高的应用场景中,传统的事后维修模式已难以满足对芯片故障实时预警与精准诊断的迫切需求
(1)本发明通过将ATE测试异常结果与DfT传感器数据进行特征级融合,将标准化处理转化为异构图节点的多维特征向量,弥补了单一数据的局限性,对ATE的批量测试数据与DfT的实时监测数据形成互补,还通过异构图结构保留了标准单元、互连线与电源网络的物理关联,使模型能够充分挖掘不同硬件组件间的潜在故障关联,从而显著提高了对纳米工艺下软故障、间歇性故障的识别精度。
Smart Images

Figure CN122045988B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated circuit fault diagnosis and early warning technology, specifically to an integrated circuit fault diagnosis and early warning system based on deep learning. Background Technology
[0002] With the rapid development of electronic technology, digital integrated circuits have been widely used in key fields such as industrial control, communication, automotive electronics, and consumer electronics, becoming the core foundation of modern information society. As process nodes continue to shrink and integration levels continue to increase, chips face increasingly severe reliability challenges during manufacturing and operation. For example, process fluctuations such as photolithography deviations and uneven etching in the manufacturing process, environmental changes such as temperature fluctuations and voltage drift during operation, and factors such as component aging and electromagnetic interference during long-term use can all induce various circuit failures. In application scenarios such as 5G communication and autonomous driving, which have extremely high requirements for system reliability, traditional post-repair methods are no longer sufficient to meet the urgent need for real-time early warning and accurate diagnosis of chip failures.
[0003] Current mainstream integrated circuit fault detection methods still have significant limitations. On the one hand, traditional output response comparison-based testing methods can only determine whether a fault exists, but cannot pinpoint the specific fault location. Furthermore, the number of test vectors increases exponentially with circuit size, resulting in low efficiency. On the other hand, existing methods are mostly based on fixed fault models, such as the stuff-at model, which is difficult to effectively cover complex fault types such as soft faults and intermittent faults. Although design-for-test (CFT) techniques such as boundary scan and built-in self-test have improved test coverage to some extent, they often come at the cost of increased chip area and power consumption, and their fault diagnosis capabilities for complex structures such as interconnect networks are limited. In addition, existing solutions mostly rely on a single data source: while automated test equipment can provide electrical and functional test results for batches of chips, it lacks dynamic sensing capabilities during operation; and while DfT sensors can collect physical parameters such as temperature and voltage in real time, they are not effectively correlated with the chip's physical structure, making it difficult to trace fault origins. Therefore, there is an urgent need for an intelligent diagnosis and early warning mechanism that integrates multi-source heterogeneous data, combines chip physical topology, and can model fault propagation characteristics to improve the accuracy, real-time performance, and interpretability of integrated circuit fault detection. Summary of the Invention
[0004] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides an integrated circuit fault diagnosis and early warning system based on deep learning, which solves the problems mentioned in the background technology.
[0005] (II) Technical Solution To achieve the above objectives, the present invention is implemented through the following technical solution: A deep learning-based integrated circuit fault diagnosis and early warning system, comprising: The data acquisition and preprocessing module acquires ATE test results and DfT sensor data, including physical parameters such as temperature, voltage, and signal delay. It also abstracts the standard cells, interconnects, and power networks in the IC physical layout into a heterogeneous graph. The ATE test results and DfT sensor data are fused together and used as the initial features of the heterogeneous graph nodes for quantification, outputting standardized graph structure data. The graph neural network model training module acquires historical IC fault data, simulates the propagation of faults in graph structure data through message passing mechanism, and uses Grad-CAM technology to generate heat maps of IC fault-sensitive areas to construct the HGNN model. The fault diagnosis and early warning module automatically triggers a processing flow when it receives new ATE test anomaly results and DfT sensor data. It outputs fault graph structure data, inputs the fault graph structure data into the HGNN model, and outputs the fault probability value of each node. Based on the node fault probability value and the coordinates in the IC physical layout, it uses a non-maximum suppression algorithm to filter out the area with the highest fault probability. A fault probability threshold is preset. When the HGNN model predicts that the fault probability of a certain area exceeds the threshold, an early warning is automatically triggered.
[0006] Furthermore, the specific steps of obtaining the ATE test results and DfT sensor data are as follows: Integrated circuits are tested using automated test equipment to obtain ATE test results; By using multiple sensors to monitor the temperature of different areas of the chip in real time, the power supply voltage of each module, and to obtain the transmission delay information of signals along different paths within the chip, DfT sensor data is obtained.
[0007] Furthermore, the process of abstracting the standard cells, interconnects, and power networks in the IC physical layout into a heterogeneous diagram is as follows: The standard cells, interconnects, and power networks in the IC physical layout are respectively regarded as nodes in the heterogeneous graph; When there is a connection between a standard cell and an interconnect, an edge is established between the corresponding standard cell node and the interconnect node. The attributes of the edge represent the direction of signal transmission and the reliability parameters of the connection. The edge reflects the flow path of the signal between the standard cell and the interconnect. For interconnected nodes, edges are established to represent connectivity. The attributes of the edges include signal transfer loss information and topological characteristics of the connection. Different topological characteristics have different effects on signal transmission. The signal transmission in complex cabling networks is analyzed to identify signal interference and crosstalk caused by improper interconnection. When a standard cell is powered by a specific power network node, the edges of the power network node are constructed. The attributes of the edges include the magnitude of the supply current and the stability index of the power supply. By analyzing the attributes of the edges, the power supply status of the standard cell is analyzed, and it is determined whether the abnormal operation of the chip is caused by a power supply problem.
[0008] Furthermore, the process of outputting standardized graph structure data is as follows: S101: Clean and filter the ATE test results and DfT sensor data, and establish a correspondence with the nodes in the heterogeneous graph; S102: Obtain the functional characteristics of different types of nodes in the heterogeneous graph, and fuse the ATE test results and DfT sensor data into a multi-dimensional feature vector of the node based on the functional characteristics. S103: The fused multidimensional feature vector is quantized to convert non-numerical information into numerical information and standardize it; S104: Assign the quantized multidimensional feature vectors to the corresponding nodes in the heterogeneous graph, while preserving the graph's topology, node types, edge connections, and attributes, to form standardized graph structure data.
[0009] Furthermore, the process of simulating the propagation of faults in graph-structured data through message passing mechanism is as follows: S201: Collect and organize fault data generated during the historical operation and testing of integrated circuits, and clarify the correspondence between fault types and fault locations; S202: Combine historical fault data with the constructed heterogeneous graph to form a graph structure input for simulating fault propagation; S203: Utilize the message passing mechanism of graph neural networks to simulate the propagation process of faults between nodes and edges in heterogeneous graphs and capture the patterns of faults; S204: After completing the message passing simulation, record the state changes of each node during the fault propagation process, and analyze the propagation path and impact range of the fault.
[0010] Furthermore, the process of constructing the HGNN model is as follows: Train a basic fault diagnosis model to generate a Grad-CAM heatmap; use Grad-CAM to generate a heatmap of IC fault-sensitive areas; extract the sensitivity weights of each node or region in the heatmap and add them as new feature dimensions to the initial feature vectors of the heterogeneous graph nodes; construct an HGNN model based on the optimized heterogeneous graph, enhanced node features, and adjusted topological relationships.
[0011] Furthermore, the process of outputting the fault diagram structure data is as follows: When new ATE test anomaly results and DfT sensor data are received, preprocessing is performed. Based on the preprocessed data, a graph structure data for fault diagnosis is constructed. The new ATE test anomaly results and DfT sensor data are fused into feature vectors of each node in the heterogeneous graph, preserving the original topology, node types, edge connections and attributes of the heterogeneous graph, thus forming fault graph structure data.
[0012] Furthermore, the process of outputting the failure probability value of each node is as follows: The generated fault graph structure data is input into the trained HGNN model. The HGNN model analyzes the fault graph structure data through its internal message passing mechanism and heterogeneous node processing module. By considering the characteristics of different types of nodes, standard cells, interconnects, power networks, and edge connection relationships, it simulates the propagation and impact of faults in the graph. The HGNN model is used to evaluate each node in the graph and outputs the fault probability value of each node.
[0013] Furthermore, the process of selecting the area with the highest probability of failure is as follows: The fault probability values of each node output by HGNN are bound to their actual coordinates in the IC physical layout to form a set of fault candidate points. The candidate point set is sorted from high to low fault probability, the area with the highest current probability is retained, and other low probability areas with high overlap with this area are suppressed. This process is repeated until all non-redundant high probability areas are selected. After NMS filtering, the regions in the result list are the non-overlapping core regions with the highest failure probability, which are the failure probability and physical coordinates of the nodes.
[0014] (III) Beneficial Effects This invention provides a deep learning-based integrated circuit fault diagnosis and early warning system, which has the following beneficial effects: (1) This invention integrates the abnormal results of ATE test with the DfT sensor data at the feature level, transforms the standardized processing into a multi-dimensional feature vector of heterogeneous graph nodes, makes up for the limitations of single data, complements the batch test data of ATE and the real-time monitoring data of DfT, and retains the physical relationship between standard units, interconnects and power networks through the heterogeneous graph structure, so that the model can fully explore the potential fault relationships between different hardware components, thereby significantly improving the identification accuracy of soft faults and intermittent faults under nanotechnology.
[0015] (2) This invention abstracts the IC physical layout into a heterogeneous graph, with nodes corresponding to various hardware components, such as standard cells, interconnects, power networks, etc., and edges representing the connection relationships between components, accurately restoring the circuit topology and laying a solid data foundation for fault propagation modeling. At the same time, based on the message passing mechanism of the HGNN model, it can simulate the propagation path of faults between different components and clearly present the diffusion law of faults. On this basis, the non-maximum suppression algorithm effectively filters the core fault area, greatly compresses the number of candidate fault sites, and greatly narrows the scope of fault investigation. Furthermore, combined with the fault sensitive area heat map generated by Grad-CAM technology, the diagnostic results can be transformed from black box output into intuitively understandable physical location markings, transforming fault investigation from traditional blind detection to precise positioning, and significantly improving the efficiency of fault investigation. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the system of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figure 1 This embodiment provides a deep learning-based integrated circuit fault diagnosis and early warning system, which includes: The data acquisition and preprocessing module acquires ATE test results and DfT sensor data, including physical parameters such as temperature, voltage, and signal delay, and abstracts the standard cells, interconnects, and power networks in the IC physical layout into a heterogeneous graph. The ATE test results and DfT sensor data are fused and quantified as the initial features of the heterogeneous graph nodes, and standardized graph structure data is output. ATE test results: The data and information obtained after testing integrated circuits using automated test equipment; during the integrated circuit manufacturing process, various tests are performed on the chips to ensure that they meet design specifications and quality requirements; ATE equipment can automatically execute a series of test items, such as electrical parameter testing: measuring the basic electrical parameters of integrated circuits such as voltage, current, and resistance to determine whether they are within the specified range. For example, it can detect whether the power supply voltage of the chip is stable within the allowable fluctuation range to ensure that the chip can operate normally.
[0019] Based on the chip's design functions, functional tests are performed. Specific stimulus signals are input, and the output signals of the chip are checked to see if they meet expectations. For example, for a processor chip, the tests are conducted to see if its arithmetic functions are correct and if the instructions are executed accurately.
[0020] Performance testing evaluates a chip's performance under different operating conditions, such as operating frequency, power consumption, and heat dissipation. For mobile phone chips, it is necessary to test their heat generation and power consumption when running applications such as games under high load, as well as whether they can maintain a stable operating frequency.
[0021] The data and results obtained through testing constitute the ATE test results, which serve as the basis for determining whether an integrated circuit is qualified and whether it has any faults.
[0022] DfT sensor data: DfT stands for Design for Testability. DfT sensor data is collected by various sensors that are specifically added during the integrated circuit design stage for easy testing. These sensors include DfT temperature sensors, DfT voltage sensors, and DfT signal delay sensors. By testing the temperature data of the heat generated by the chip during operation, the DfT temperature sensor is used to monitor the temperature of different areas of the chip in real time. Excessive temperature may affect the chip performance or even cause failure. For example, in high-performance processors, the temperature of the core area will rise rapidly as the computing task increases. By collecting temperature data through the DfT temperature sensor, it can be determined whether the chip is overheating and timely heat dissipation measures or adjustment of the chip's operating frequency can be taken.
[0023] DfT voltage sensors can monitor the power supply voltage of various modules in a chip. A stable power supply voltage is the foundation for the normal operation of the chip. When the voltage fluctuates abnormally, such as excessively high voltage, it may damage chip components, while excessively low voltage may prevent the chip from running instructions normally. At this time, the collected voltage data can be used to detect problems in time and adjust or warn the power management module.
[0024] Inside an integrated circuit, it takes time for a signal to travel from one component to another. DfT signal delay sensors can acquire information about the transmission delay of a signal along different paths within the chip. For example, in high-speed data transmission circuits, if the signal delay exceeds the design range, it may lead to problems such as data transmission errors and communication interruptions. Signal delay data can be used to analyze the integrity and stability of signal transmission and to investigate possible causes of failure such as unreasonable wiring or degraded component performance.
[0025] The process of abstracting the IC physical layout into a heterogeneous layout is as follows: By converting standard cells, interconnects, and power networks in the IC physical layout into nodes in the graph, and establishing edges based on the actual connection relationships, a heterogeneous graph containing various types of nodes and edges is constructed.
[0026] Define the nodes of the graph: Standard cells in an IC physical layout are basic circuit modules with specific logic functions, such as AND gates, OR gates, and flip-flops. Each standard cell is considered a node in a heterogeneous graph. Each node is assigned attributes related to the standard cell, such as cell type, number of input / output ports, physical coordinates in the layout, and electrical parameters, such as drive capability and delay time.
[0027] Interconnects are used to connect different standard units to achieve signal transmission; they can be considered as nodes in a heterogeneous diagram; their attributes include line length, line width, number of wiring layers, signal transmission delay, parasitic capacitance, and inductance; the attributes of interconnects play an important role in analyzing the integrity and interference of signals during transmission, and in fault diagnosis, they help determine whether signal transmission abnormalities are caused by interconnect problems.
[0028] The power network provides a stable power supply to the standard cells and is also abstracted as nodes. The attributes of the power network nodes can include power type, such as the voltage value of the DC power supply, power supply range, i.e., which standard cells the power node is responsible for powering, the resistance and capacitance characteristics of the power network, which affect the stability and transient response of the power supply, etc., which are used for in-depth troubleshooting when analyzing chip power consumption anomalies, power fluctuations and other faults.
[0029] Define the edges of the graph: When standard cells and interconnects are connected, meaning signals can be transmitted between standard cells or between standard cells and other circuit components via interconnects, edges are established between the corresponding standard cell nodes and interconnect nodes. The attributes of the edges represent the direction of signal transmission and the reliability parameters of the connection, such as connection resistance, which reflects whether the connection is good; excessive resistance may indicate a potential connection failure. The edges clearly reflect the flow path of signals between standard cells and interconnects, providing information for fault location. For example, when a signal error occurs, the problematic component can be traced along the direction of the edge.
[0030] For interconnected nodes, edges are established to represent connectivity. The attributes of the edges can include signal transfer loss information and topological characteristics of the connection, such as whether it is a T-connection, a cross-connection, etc. Different topologies may have different effects on signal transmission. This is used to analyze the signal transmission in complex cabling networks and to troubleshoot signal interference, crosstalk, and other problems caused by improper interconnection.
[0031] When a standard cell is powered by a specific power network node, the edges of the power network node are constructed. The attributes of the edges include the magnitude of the supply current and the stability indicators of the power supply. By analyzing the attributes of the edges, the power supply status of the standard cell is analyzed, and it is determined whether the abnormal operation of the chip is caused by a power supply problem.
[0032] Standardized graph structure data output: S101: Clean and filter the ATE test results and DfT sensor data, and establish a correspondence with the nodes in the heterogeneous graph to ensure that the data matches the nodes; Outliers caused by equipment errors in ATE testing are removed, such as voltage and current data exceeding the physical reasonable range, as well as noise data collected by DfT sensors, such as temperature with excessive instantaneous fluctuations and signal delay values. Missing data is filled in; for example, if the DfT temperature data of a certain standard cell is missing, it can be filled by interpolation based on the temperature trend of adjacent cells. Based on the coordinate information and node attributes of the IC physical layout, ATE test results are associated with the corresponding nodes; for example, the output voltage deviation of a certain standard cell in ATE testing is accurately matched to the node of that standard cell in the heterogeneous diagram; similarly, the temperature of a certain area and the signal delay of a certain interconnect collected by the DfT sensor are mapped to the standard cell node, interconnect node, or power network node covered by that area in the heterogeneous diagram.
[0033] S102: For different types of nodes in a heterogeneous graph, such as standard cells, interconnects and power networks, the ATE test results and DfT sensor data are fused into a multi-dimensional feature vector for the node, taking into account its functional characteristics.
[0034] Feature fusion of standard unit nodes: ATE test data is obtained from the unit's input / output logic function test results, such as binary flags indicating normal or abnormal function, quiescent current, and dynamic power consumption; DfT sensor data is obtained from temperature sensor data near the unit, power supply voltage fluctuation data, and internal signal transmission delay; the above data are integrated into a vector, for example, functional test results, 0 or 1, quiescent current mA, temperature °C, power supply voltage V, and signal delay ps.
[0035] Feature fusion of interconnect nodes: ATE test data, through the continuity test results of interconnects, such as short circuit or open circuit markings, parasitic capacitance or inductance measurements; DfT sensor data, through the real-time signal attenuation of interconnects, crosstalk noise values, interference from adjacent interconnects; fused into vectors, such as continuity, 0 or 1, parasitic capacitance pF, signal attenuation dB, crosstalk noise mV.
[0036] Feature fusion of power network nodes: ATE test data includes voltage drop test values and short-circuit test results of the power supply network; DfT sensor data includes the real-time current load and voltage ripple fluctuation amplitude of the network; these are fused into vectors, such as quiescent current, short-circuit test 0 or 1, current load, and voltage ripple.
[0037] S103: The fused multidimensional feature vectors are quantized to convert non-numerical information into numerical values and standardize them, eliminating dimensional differences and ensuring that the graph structure data meets the model input requirements.
[0038] Boolean results are quantized, such as normal = 1, abnormal = 0, short circuit = 1, open circuit = 0, and directly represented by binary values. For physical quantity data, such as temperature, voltage, and delay, their original values are retained. For classification information, such as standard unit types of AND gates and OR gates, one-hot encoding is used to convert them into numerical vectors. For example, an AND gate is represented as [1,0,0], and an OR gate is represented as [0,1,0].
[0039] Using Min-Max normalization or Z-score standardization maps features of different magnitudes to a unified interval, such as [0,1] or a mean of 0 and a variance of 1, to avoid the model becoming overly sensitive to certain types of features due to differences in units, such as temperature being in °C and current being in A.
[0040] For example, after performing Min-Max normalization on temperature data [50℃, 80℃, 60℃], it can be converted to [0, 1, 0.33], assuming that the normal operating temperature range is 50~80℃.
[0041] S104: Assign the quantized multidimensional feature vectors to the corresponding nodes in the heterogeneous graph, while preserving the graph's topology, node types, edge connections, and attributes, ultimately forming standardized graph structure data.
[0042] Node feature assignment: The attribute fields of each node are updated to quantized feature vectors. For example, the features of a standard cell node are expanded from type AND gate and coordinates (x,y) to a vector containing fused features of ATE and DfT. Edge attributes are preserved. The connection relationships of edges, such as the connection between standard cells and interconnects, the power supply relationship between power networks and standard cells, and the original attributes, such as signal transmission direction and connection resistance, remain unchanged to ensure the integrity of the graph topology. The output format is graph data, such as PyTorchGeometric's Data object or DGL's DGLGraph storage, which includes node type, node feature matrix, edge index matrix, edge attribute matrix, etc., and can be directly input into the graph neural network model for training.
[0043] The graph neural network model training module acquires historical IC fault data, simulates the propagation of faults in graph structure data through message passing mechanism, and uses Grad-CAM technology to generate heat maps of IC fault-sensitive areas to construct the HGNN model. Simulate the process of fault propagation in graph-structured data: S201: Collect and organize fault data generated during the historical operation and testing of integrated circuits, and clarify the correspondence between fault types and fault locations; Test results of the faulty chip are extracted from the ATE test history, and abnormal data such as temperature, voltage, and signal delay at the time of the fault are filtered from the historical data of the DfT sensor. Combined with the physical failure analysis of the chip, such as microscopic imaging of the failure area and electrical testing, the root cause of the fault is determined, such as the failure of a standard cell or a short circuit of an interconnect. The collected historical data is labeled to clarify the fault type corresponding to each fault sample, such as the fault location of logic function error, power supply noise fault, signal crosstalk fault, etc., and the corresponding node in the heterogeneous graph, such as a certain AND gate cell or a certain interconnect node. S202: Combine historical fault data with the previously constructed heterogeneous graph to form a graph structure input for simulating fault propagation; For each historical fault sample, initial fault features are set on the corresponding node in the heterogeneous graph according to its fault location. For example, if a standard cell has a logic function error, the functional test result attribute in the feature vector of that node is marked as an abnormal value, changing from the normal "1" to "-1" indicating a fault. At the same time, combined with the abnormal temperature, voltage and other information of that node in the DfT sensor data, the feature representation of the initial fault node is constructed. For nodes in the heterogeneous graph that have not experienced a fault, the feature vector after the normal ATE test result and DfT sensor data are fused is retained to simulate the real state of the chip when the fault occurs.
[0044] S203: Utilize the message passing mechanism of graph neural networks to simulate the propagation process of faults between nodes and edges in heterogeneous graphs and capture the patterns of faults; For each node in the graph, its own state is calculated and updated using a message passing function based on the state of its neighboring nodes, including whether there is a fault and the severity of the fault. The message passing function includes message generation, message aggregation, and node update. Message generation: Nodes generate messages to pass on to their neighbors based on their own fault characteristics and connection attributes with their neighbors, such as the signal transmission direction of the edge and the connection resistance. For example, a faulty standard cell node will generate a message containing the fault signal strength and propagation direction based on its connection relationship with the interconnection node.
[0045] Message aggregation: Each node collects messages from all its neighbors and integrates them using aggregation functions such as summation, averaging, and maximum value to obtain the impact of neighbors on the node.
[0046] Node update: Combining its previous state with aggregated neighbor messages, the node's new state is calculated through an update function, such as a neural network layer, to simulate the impact of a fault on the node.
[0047] Multiple rounds of message passing are performed to simulate the gradual spread of a fault in a graph structure. After each round of message passing, the state of the nodes is updated to reflect the progress of the fault propagation. For example, after the first round of passing, the state of the neighboring interconnect nodes of the faulty standard cell changes. After the second round of passing, the state of other standard cell nodes connected to the interconnect is also affected, and so on, until the fault propagation reaches a stable state or a preset number of rounds.
[0048] S204: After completing the message passing simulation, record the state changes of each node during the fault propagation process, and analyze the propagation path and impact range of the fault.
[0049] Based on the order and degree of node state changes, the path of fault propagation from the initial node to other nodes is extracted; for example, a fault starts from standard cell A, passes through interconnection B, and propagates to standard cell C, this path is one of the fault propagation paths; the number and type of nodes affected during the fault propagation process are counted to assess the severity and potential harm of the fault; these simulated fault propagation data, including the initial fault node, propagation path, and finally affected nodes, are used as training data for subsequent training of the graph neural network model.
[0050] The process of building an HGNN model: Step 1: Train the basic fault diagnosis model and generate Grad-CAM heatmap: The basic model for preliminary IC fault diagnosis is trained and serves as the carrier of Grad-CAM technology. It acquires the fault-sensitive areas of interest to the model and uses standardized graph structure data with labeled fault types and locations, including node features, topological relationships, or image data of the IC physical layout, such as converting the layout into pixel images and labeling fault areas. If based on graph structure data, a basic GNN, such as GCN or GAT, can be trained. The input is the node features and topological relationships of the heterogeneous graph, and the output is the fault type, such as standard cell failure, interconnect short circuit, or fault probability. This can accurately identify faults and provide reliable gradient information for Grad-CAM to extract sensitive areas.
[0051] Step 2: Generate a heat map of IC fault-sensitive areas using Grad-CAM: Grad-CAM calculates the gradient of the model output with respect to intermediate layer features, locating the sensitive areas that the model focuses on when making decisions. It intuitively reflects the areas with the highest correlation to the fault. For a pre-trained basic fault diagnosis model, the input is IC data containing the fault, i.e., a graph structure or image. The model outputs the gradient of the fault probability with respect to the intermediate layer feature map or graph node features. The magnitude of the gradient reflects the contribution of the region or node to the fault diagnosis result. The larger the gradient, the more the model focuses on the region and the stronger the correlation with the fault.
[0052] For nodes in the heterogeneous graph, the gradient of the output with respect to each node's features is calculated to obtain the node importance weight. The higher the weight, the more fault-sensitive the node is. For pixels in the layout image, the gradient of the output with respect to the convolutional layer feature map is calculated. The feature map weights are obtained through global average pooling, and then the feature maps are weighted and combined to generate a heatmap. The darker the color in the heatmap, such as red, the more fault-sensitive the pixel area is. The importance weights obtained from the gradient calculation are mapped to the actual structure of the IC physical layout, and the node importance weights are labeled on the corresponding nodes in the heterogeneous graph, such as standard cells and interconnect nodes, forming a node-level heatmap that visually shows which nodes have the greatest impact on fault decisions. Step 3: Optimize the node features and topology of the heterogeneous graph based on the heatmap: Extract the sensitive weights of each node or region in the heatmap and add them as new feature dimensions to the initial feature vector of the heterogeneous graph nodes.
[0053] For example, if the sensitivity weight of a certain standard unit node in the heatmap is 0.8, ranging from 0 to 1, then the attribute of sensitivity weight 0.8 will be added to its original feature vector, which contains ATE and DfT data. For interconnects and power network nodes, add corresponding sensitive weights so that node features include not only test data but also fault-sensitive information; adjust the edge attributes of the heterogeneous graph according to the connection relationship of sensitive areas in the heatmap.
[0054] For example, if the heatmap shows two highly sensitive nodes, such as the interconnection between the fault source node and the directly affected node, which is crucial for fault propagation, the propagation weight attribute of that edge can be increased, such as from 0.2 to 0.9, to strengthen the connection between the two in the graph. For edges in low-sensitivity regions, reduce their weights or keep them unchanged to reduce the interference of irrelevant information on the model.
[0055] Step 4: Build and train the HGNN model: Based on the optimized heterogeneous graph, including enhanced node features and adjusted topological relationships, a heterogeneous graph neural network model, namely the HGNN model, is constructed to achieve more accurate fault diagnosis and location.
[0056] HGNN is used to process various types of nodes and edges in heterogeneous graphs, which is achieved through the following modules: For different types of nodes, design independent feature transformation subnetworks, such as MLP, to map the original features, including sensitive weights, to an embedding space of a unified dimension, while preserving the differences between node types; Different message passing functions are designed according to the type of edge, such as signal transmission edge and power supply edge. For example, for the edge between standard cell and interconnection line, message passing needs to consider signal direction and sensitivity weight; for the edge between power network and standard cell, power stability characteristics and sensitivity weight need to be combined to ensure that different types of associated information are effectively captured. The final node embedding is obtained by aggregating the node's own embedding and the messages passed by its neighbors. Then, the failure probability of each node is output through the classification header, such as whether it is faulty and the type of fault.
[0057] Using labeled fault data as supervision signals, such as whether a node is faulty, the parameters of HGNN are optimized through backpropagation. During training, the sensitive weights introduced by the heatmap guide the model to pay more attention to highly correlated areas, thereby improving the accuracy of fault location. Combined with the visualization results of the heatmap, it is verified that the model focuses on physically fault-sensitive areas. If there is a deviation, such as the model focusing on irrelevant areas, the heatmap mapping or model structure is adjusted retrospectively.
[0058] By visualizing the decision logic of the model through Grad-CAM, sensitive areas of IC faults are revealed and encoded into the node features and topological relationships of the heterogeneous graph. The constructed HGNN model utilizes heterogeneous information and fault-sensitive features to learn fault modes and propagation laws, thereby improving the accuracy and interpretability of integrated circuit fault diagnosis.
[0059] The fault diagnosis and early warning module automatically triggers a processing flow when it receives new ATE test anomaly results and DfT sensor data. It outputs fault graph structure data, inputs the fault graph structure data into the HGNN model, and outputs the fault probability value of each node. Based on the node fault probability value and the coordinates in the IC physical layout, it uses a non-maximum suppression algorithm to filter out the area with the highest fault probability. A fault probability threshold is preset. When the HGNN model predicts that the fault probability of a certain area exceeds the threshold, an early warning is automatically triggered.
[0060] Output fault diagram structure data: When the system receives new ATE test anomaly results and DfT sensor data, it performs preprocessing to remove outliers in the ATE test and DfT sensor data, such as voltage, temperature, and noise data that are outside the reasonable range. Based on the coordinate information and node attributes of the IC physical layout, the ATE test anomaly results are associated with the corresponding heterogeneous graph nodes, and the DfT sensor data, such as temperature, voltage, and signal delay, are mapped to the corresponding standard cells, interconnects, or power network nodes in the heterogeneous graph.
[0061] Based on the preprocessed data, a graph structure data for fault diagnosis is constructed. The new ATE test anomaly results and DfT sensor data are fused into feature vectors of each node in the heterogeneous graph. For example, the feature vector of a standard cell node needs to include information such as its ATE test functional anomaly marker and the abnormal temperature value collected by the DfT sensor. The original topology of the heterogeneous graph, node type, edge connection relationship and attributes are preserved to form fault graph structure data. This data fully reflects the fault-related status of the current IC.
[0062] Output the failure probability value for each node: The generated fault graph structure data is input into a pre-trained HGNN model, i.e., a heterogeneous graph neural network model. The HGNN model analyzes the fault graph structure data through its internal message passing mechanism and heterogeneous node processing module. By considering the characteristics of different types of nodes, standard cells, interconnects, power networks, and edge connections, it simulates the propagation and impact of faults in the graph. The HGNN model evaluates each node in the graph and outputs a fault probability value for each node; for example, the fault probability of a standard cell node is 0.85, and the fault probability of an interconnect node is 0.3, etc., representing the likelihood of that node failing.
[0063] Filter out the areas with the highest probability of failure: The fault probability values of each node output by HGNN are bound to their actual coordinates in the IC physical layout to form a fault candidate point set. The physical coordinates of each node are obtained from the node attributes of the heterogeneous graph, such as the center coordinates (x1, y1) of a standard cell, the start or end coordinates (x2, y2)-(x3, y3) of an interconnect, and the coverage area coordinates of a power network node. A fault candidate region is generated for each node. The region range is defined according to the node type and associated with its fault probability. Typically, it is a rectangular region based on the cell size of the layout design, such as 10μm×10μm, with the center coordinates (x, y) as the reference and the probability being the fault probability value of that node. It is a line segment or strip region with the width being the line width, the range being the coordinates of its wiring path, and the probability being the fault probability of that interconnect node. It is a region covering its power supply range, such as a polygon, with the probability being the fault probability of the power network node. This forms a candidate point set list containing the region coordinate range and the fault probability.
[0064] The candidate point set is sorted from highest to lowest fault probability to ensure that the most likely faulty regions are processed first. For example, candidate regions A have a probability of 0.92, B has a probability of 0.88, C has a probability of 0.754, and so on. Arranged in descending order of probability, subsequent algorithms will first retain high-probability regions and then suppress low-probability regions that overlap with them. A non-maximum suppression algorithm is used to filter core regions, retaining the region with the highest current probability and suppressing other low-probability regions with high overlap with it. This process is repeated until all non-redundant high-probability regions are selected. The region with the highest probability from the sorted candidate point set is selected as the seed region and added to the final result list. This region is then removed from the candidate set, and the overlap ratio (IoU) between the remaining regions in the candidate point set and the seed region is calculated.
[0065] By using the formula IoU = Intersection area of two regions / Union area of two regions, an overlap threshold can be set, such as 0.5, which can be adjusted according to the accuracy of the IC layout: If the IoU of a certain region and the seed region is greater than or equal to the threshold, it means that the two regions are highly overlapping in physical location and belong to the redundant candidates of the same fault region.
[0066] Low-probability regions in the candidate point set whose overlap with the seed region exceeds a threshold are removed. The region with the highest probability is selected again from the remaining candidate point set as the new seed region. Steps 2 to 3 are repeated until the candidate point set is empty. After NMS filtering, the regions in the result list are the non-overlapping core regions with the highest failure probability. Combining the node's failure probability and physical coordinate position, it reflects the location where the IC is most likely to fail.
[0067] For example, after filtering, two core regions may be obtained: Region 1, a standard cell cluster with a probability of 0.92 and a coordinate range of (x1, y1) - (x2, y2), and Region 2, an interconnection intersection area with a probability of 0.85 and a coordinate range of (x3, y3) - (x4, y4). The two regions do not overlap and correspond to different potential fault points. NMS solves the problem of high fault probability and overlapping physical locations of multiple adjacent nodes in the IC layout. For example, a core fault point may cause multiple surrounding nodes to show a high probability, avoiding duplicate labeling, and the output region is closer to the actual physical location of the fault.
[0068] The process of triggering an alert is as follows: Based on the reliability requirements of the integrated circuit type and application scenario, reasonable fault probability thresholds are set. For example, for automotive chips, consumer electronics chips, and industrial control chips, by analyzing the fault probability distribution of real fault areas in historical fault cases, the threshold is set in the range where the probability of most real fault areas exceeds the value, while the probability of normal areas rarely exceeds it. For example, the optimal threshold is determined based on the ROC curve. Combined with the test engineer's understanding of common chip fault modes, the threshold is fine-tuned, such as setting a lower threshold for power network faults because they may cause cascading failures. The determined thresholds are further subdivided according to fault type, such as a short circuit fault threshold of 0.5 and a logic error threshold of 0.4, and stored in the system configuration file as a benchmark for early warning judgment.
[0069] After the system receives new ATE test anomaly data and DfT sensor data, it generates fault graph structure data through processing, inputs it into the HGNN model to obtain the fault probability value of each node, and then uses a non-maximum suppression algorithm to filter out the core fault areas and their corresponding highest probability values. Each core area corresponds to a comprehensive probability, such as the highest probability or average probability of all nodes in that area.
[0070] By analyzing each of the selected core fault areas, the fault probability value P of each core area is compared with the preset corresponding type threshold T: If P > T: the area is determined to be a high-risk fault area, and the early warning triggering condition is met; If P≤T: It is determined to be a low-risk area, and no warning will be triggered for the time being. Only the information of this area will be recorded for subsequent trend analysis.
[0071] When the probability of failure in at least one core area exceeds the threshold, the system automatically activates the early warning mechanism, generates an early warning identifier (e.g., a binary signal indicating an early warning trigger), and records key information such as the trigger time, the corresponding chip batch or number, and the coordinates of the core area. The faulty area is highlighted on the monitoring interface, overlaid on the IC physical layout, and a pop-up window displays the area coordinates, the probability of failure, and the possible fault type, such as the risk of short circuit in the interconnect, with a probability of 0.72 > the threshold of 0.5. The early warning information is sent to test engineers and the production line control system via email, SMS, or industrial bus, including the fault location, urgency level, and is graded according to the magnitude of the probability exceeding the threshold.
[0072] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0073] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0074] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A deep learning-based integrated circuit fault diagnosis and early warning system, characterized in that: The system includes: The data acquisition and preprocessing module acquires ATE test results and DfT sensor data, including physical parameters such as temperature, voltage, and signal delay. It also abstracts the standard cells, interconnects, and power networks in the IC physical layout into a heterogeneous graph. The ATE test results and DfT sensor data are fused together and used as the initial features of the heterogeneous graph nodes for quantification, outputting standardized graph structure data. The process of outputting standardized graph structure data is as follows: S101: Clean and filter the ATE test results and DfT sensor data, and establish a correspondence with the nodes in the heterogeneous graph; S102: Obtain the functional characteristics of different types of nodes in the heterogeneous graph, and fuse the ATE test results and DfT sensor data into a multi-dimensional feature vector of the node based on the functional characteristics. S103: The fused multidimensional feature vector is quantized to convert non-numerical information into numerical information and standardize it; S104: Assign the quantized multidimensional feature vectors to the corresponding nodes in the heterogeneous graph, while preserving the graph's topology, node types, edge connections, and attributes to form standardized graph structure data. The graph neural network model training module acquires historical IC fault data, simulates the propagation of faults in graph structure data through message passing mechanism, and uses Grad-CAM technology to generate heat maps of IC fault-sensitive areas to construct the HGNN model. The process of simulating the propagation of faults in graph-structured data through message passing mechanism is as follows: S201: Collect and organize fault data generated during the historical operation and testing of integrated circuits, and clarify the correspondence between fault types and fault locations; S202: Combine historical fault data with the constructed heterogeneous graph to form a graph structure input for simulating fault propagation; S203: Utilize the message passing mechanism of graph neural networks to simulate the propagation process of faults between nodes and edges in heterogeneous graphs and capture the patterns of faults; S204: After completing the message passing simulation, record the state changes of each node during the fault propagation process, and analyze the propagation path and impact range of the fault; The fault diagnosis and early warning module automatically triggers a processing flow when it receives new ATE test anomaly results and DfT sensor data. It outputs fault graph structure data, inputs the fault graph structure data into the HGNN model, and outputs the fault probability value of each node. Based on the node fault probability value and the coordinates in the IC physical layout, it uses a non-maximum suppression algorithm to filter out the area with the highest fault probability. A fault probability threshold is preset. When the HGNN model predicts that the fault probability of a certain area exceeds the threshold, an early warning is automatically triggered.
2. The integrated circuit fault diagnosis and early warning system based on deep learning according to claim 1, characterized in that: The specific steps for obtaining the ATE test results and DfT sensor data are as follows: Integrated circuits are tested using automated test equipment to obtain ATE test results; By using multiple sensors to monitor the temperature of different areas of the chip in real time, the power supply voltage of each module, and to obtain the transmission delay information of signals along different paths within the chip, DfT sensor data is obtained.
3. The integrated circuit fault diagnosis and early warning system based on deep learning according to claim 2, characterized in that: The process of abstracting the standard cells, interconnects, and power networks in the IC physical layout into a heterogeneous graph is as follows: The standard cells, interconnects, and power networks in the IC physical layout are respectively regarded as nodes in the heterogeneous graph; When there is a connection between a standard cell and an interconnect, an edge is established between the corresponding standard cell node and the interconnect node. The attributes of the edge represent the direction of signal transmission and the reliability parameters of the connection. The edge reflects the flow path of the signal between the standard cell and the interconnect. For interconnected nodes, edges are established to represent connectivity. The attributes of the edges include signal transfer loss information and topological characteristics of the connection. Different topological characteristics have different effects on signal transmission. The signal transmission in complex cabling networks is analyzed to identify signal interference and crosstalk caused by improper interconnection. When a standard cell is powered by a specific power network node, the edges of the power network node are constructed. The attributes of the edges include the magnitude of the supply current and the stability index of the power supply. By analyzing the attributes of the edges, the power supply status of the standard cell is analyzed, and it is determined whether the abnormal operation of the chip is caused by a power supply problem.
4. The integrated circuit fault diagnosis and early warning system based on deep learning according to claim 1, characterized in that: The process of constructing the HGNN model is as follows: Train a basic fault diagnosis model to generate a Grad-CAM heatmap; use Grad-CAM to generate a heatmap of IC fault-sensitive areas; extract the sensitivity weight of each node or region in the heatmap and add it as a new feature dimension to the initial feature vector of the heterogeneous graph nodes; An HGNN model is constructed based on the optimized heterogeneous graph, enhanced node features, and adjusted topological relationships.
5. The integrated circuit fault diagnosis and early warning system based on deep learning according to claim 1, characterized in that: The process of outputting the fault diagram structure data is as follows: When new ATE test anomaly results and DfT sensor data are received, preprocessing is performed. Based on the preprocessed data, a graph structure data for fault diagnosis is constructed. The new ATE test anomaly results and DfT sensor data are fused into feature vectors of each node in the heterogeneous graph, preserving the original topology, node types, edge connections and attributes of the heterogeneous graph, thus forming fault graph structure data.
6. The integrated circuit fault diagnosis and early warning system based on deep learning according to claim 5, characterized in that: The process of outputting the failure probability value of each node is as follows: The generated fault graph structure data is input into the trained HGNN model. The HGNN model analyzes the fault graph structure data through its internal message passing mechanism and heterogeneous node processing module. By considering the characteristics of different types of nodes, standard cells, interconnects, power networks, and edge connection relationships, it simulates the propagation and impact of faults in the graph. The HGNN model is used to evaluate each node in the graph and outputs the fault probability value of each node.
7. The integrated circuit fault diagnosis and early warning system based on deep learning according to claim 1, characterized in that: The process of selecting the area with the highest probability of failure is as follows: The fault probability values of each node output by HGNN are bound to their actual coordinates in the IC physical layout to form a set of fault candidate points. The candidate point set is sorted from high to low fault probability, the area with the highest current probability is retained, and other low probability areas with high overlap with this area are suppressed. This process is repeated until all non-redundant high probability areas are selected. After NMS filtering, the regions in the result list are the non-overlapping core regions with the highest failure probability, which are the failure probability and physical coordinates of the nodes.
Citation Information
Patent Citations
Integrated circuit test method and system
CN118884191A
Display control drive circuit fault diagnosis system based on deep learning
CN119360758A