Digital chip fault detection method and system based on big data

By establishing multi-dimensional feature associations of chip failure failure analysis data, and building a single fault positioning model and multiple fault positioning model, the problem of low chip failure failure positioning accuracy in the existing technology is solved, and more efficient and accurate fault positioning is achieved.

CN120197067APending Publication Date: 2025-06-24SUZHOU XINLIAN ZHILIAN TECHNOLOGY CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510274899.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing chip failure failure analysis methods are not very accurate, especially when facing multiple failure failure points, the positioning accuracy is low.

Method used

By collecting historical hierarchical data of chip failure failure analysis, extracting feature data at each stage, and establishing multi-dimensional feature associations of failure failure analysis data. Based on the failure multiple root causes of single failure and associated interaction, a single fault location model and a multi-fault location model are established respectively.

Benefits of technology

Accurate positioning of single and multiple failure faults is achieved, and the efficiency and accuracy of chip failure fault location is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197067A_ABST
    Figure CN120197067A_ABST
Patent Text Reader

Abstract

The invention discloses a digital chip fault detection method and system based on big data, and the method comprises the steps: collecting the historical hierarchy data of chip failure fault analysis, extracting the feature data of each stage, and building the multi-dimensional feature association of the failure fault analysis data based on the feature data; synchronizing a single fault positioning model of the chip failure fault established based on the single failure root cause and a multi-fault positioning model of the chip failure fault established based on the associated interaction failure multiple root causes; failure fault positioning is carried out through a single fault positioning model and a multi-fault positioning model, a failure fault positioning result is output, and finally, the fault positioning output results of the single fault positioning model and the multi-fault positioning model are verified and reproduced. According to the method, the corresponding fault positioning model is constructed by using the judgment result of the failure root cause, and the problem of low positioning precision caused by various influence factors and interaction of an existing single fault positioning method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital chip fault detection, and particularly to a method and system for digital chip fault detection based on big data. Background Technique

[0002] As the core component of modern electronic devices, the reliability and stability of chips are crucial for the normal operation of the entire system. During the R & D, production, and use of chips, various failure faults will occur. And the analysis based on chip failure faults is a key technology in the electronics industry, which involves the detection, diagnosis, and repair of faulty chips to ensure the normal operation and reliability of electronic devices. It is particularly important to accurately analyze the causes of chip failures, which not only helps improve product quality, provides feedback for chip design and manufacturing, but also is crucial for ensuring the normal operation of devices.

[0003] Existing methods for analyzing chip failure faults include visual inspection, electrical testing, thermal analysis, X-ray detection, and comparative analysis, etc. And the reasons for chip failure faults are diverse, including overheating, unstable power supply voltage, circuit design problems, material aging, environmental factors, and human operation errors, etc. And due to the high cost of equipment for chip failure fault analysis, most failure fault location analyses are carried out through a single failure fault analysis device. And the failure fault location methods of these single failure fault analysis devices are mainly based on the matching and location of failure modes based on fixed failure phenomena. However, the influencing factors of chip failure faults often interact with each other, resulting in low accuracy of existing failure fault analysis methods and low accuracy of failure fault location when there are multiple failure fault points. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for digital chip fault detection based on big data to solve the problems raised in the above background technique.

[0005] To solve the above technical problems, the present invention provides the following technical solution: A method for digital chip fault detection based on big data, including the following steps: Step 1: Collect historical hierarchical data of chip failure fault analysis, extract characteristic data of each stage of the chip, and establish a multi-dimensional feature association of failure fault analysis data based on the characteristic data; Preferably, the historical hierarchical data collected for chip failure fault analysis is a plurality of information data related to chip failure fault analysis collected and recorded starting from the chip manufacturing stage; the levels are divided according to different stages the chip has experienced, and the levels include the chip manufacturing stage, the chip detection stage, and the chip production output stage. The data in different stages have different effects on chip failure analysis and fault location; the characteristic data of each stage of the chip is extracted as: extracting the characteristic data of different stages from the historical hierarchical data collected for chip failure fault analysis, and different data collection methods for failure fault analysis are used in different stages.

[0006] Preferably, establishing multi-dimensional feature associations for failure fault analysis data includes: using an association algorithm to associate the characteristic data related to chip failure fault analysis collected in different stages with the corresponding dimensions. The corresponding dimensions include the process parameter dimension, electrical characteristic dimension, and material property dimension related to chip failure fault analysis. The process parameter dimension uses a quantum annealing algorithm to optimize the feature association to obtain the optimal association mapping of process parameters and locate the root cause of failure. The electrical characteristic dimension cross-level couples the characteristic data of different stages through a Bayesian network to identify hidden and interacting multi-root cause associations of failures. The material property dimension uses a cross-modal graph convolutional network to construct a heterogeneous graph structure from material SEM images, XRD spectra, and physical property parameters. Through heterogeneous data fusion, an association model of material microstructure - macroscopic properties is established, and the edge weights are calculated through a cross-modal attention mechanism to capture the non-linear association between lattice defects and electrical parameters.

[0007] Step 2: Based on the single root cause of failure and the multi-root causes of failure with associated interactions, establish a single fault location model and a multi-fault location model for chip failure faults respectively; Preferably, the specific steps of establishing a single fault location model and a multi-fault location model for chip failure faults respectively include: judging the type of root cause of failure corresponding to the association relationship between different dimensions and the characteristic data. When it is judged as a single root cause of failure association type, a single fault location model for chip failure faults is established according to the single root cause of failure. When it is judged as an associated and interacting multi-root cause of failure association type, a multi-fault location model for chip failure faults is established according to the multi-root causes of failure with associated interactions.

[0008] Preferably, the failure fault corresponding to the single root cause of failure is positioned as the input under the single fault location model. That is, when scanning and detecting chip failure faults starting from the initial stage of the chip manufacturing process, the single fault location model detects and marks the failure fault location points based on the scanning logic where one scanning detection point corresponds to one fault point. When there are multiple fault points at one scanning detection point, if each fault point has an independent scanning detection point, the single fault model is still used. When any one scanning detection point and two equivalent scanning detection points correspond to at least two fault points, a bionic pulse neural network location model is utilized to construct a three-layer pulse network with synaptic plasticity. The input layer receives the characteristic pulse sequence, the hidden layer dynamically adjusts the synaptic weights using the STDP learning rule, and the firing frequency of the output layer represents the fault probability. The synaptic reinforcement function is defined as , and the calculation formula is: Among them, represents the change in synaptic weight between neurons to , is the learning rate that controls the weight update amplitude, , respectively represent the firing times of the presynaptic neuron i and the postsynaptic neuron j, is the time decay constant, which is used to determine the decay speed of the influence of the pulse time difference on the weight. When the spatio-temporal correlation of the multi-fault signal exceeds the defined threshold, synchronous firing is triggered. Based on the dynamic adjustment of the synaptic weights, the detection of multi-root cause coupling is realized, and the multi-fault location model of the chip failure fault established based on the associated and interacting failure multi-root causes is analyzed to determine whether the fault points belong to the same characteristic dimension, and multi-label classification is performed according to the judgment result to obtain the failure indication of each characteristic data.

[0009] Step Three: Use the single fault location model to perform single failure fault location; Preferably, the specific operation process of the single fault location model includes: Step a1: Input the chip failure-related characteristic data with consistency, and determine the associated dimension based on the scanning detection points and the corresponding fault points marked by this characteristic data; Step a2: Use the known historical data corresponding to the failure types and characteristic data, perform training and matching based on the deep learning algorithm, output the mapping relationship between the characteristic data and the failure types, and determine the failure types corresponding to the characteristic data; Step a3: According to the determined failure types, locate the failure fault positions using static causal diagnosis, dynamic causal diagnosis, and causal reverse diagnosis.

[0010] Step 4: Use the multi-fault location model to locate multi-failure faults; Preferably, the specific operation process of the multi-fault location model includes: Step b1: Synchronously map the input chip failure-related feature data with consistency to the associated dimensions, independently train a binary classification model for each failure root cause, infer the interactive causality of multiple root causes, determine the failure type corresponding to the feature data, and extract the common features across failure types; Step b2: Calculate the linear relationship between the feature data and the failure type using the Pearson correlation coefficient, capture the non-linear relationship through mutual information, and generate the association matching rules between the failure type and the fault point based on the extracted common features; Step b3: Use a machine learning algorithm to detect outliers in the input chip failure-related feature data with consistency. When an abnormal feature data is detected, infer according to the mined association matching rules. If feature data 1 is abnormal and meets the conditions of rule 1, it is inferred that the fault point B is likely to appear at feature data 1; if feature data 3 is abnormal and meets the conditions of rule 2, it is inferred that the equivalent feature data 4 may also have the corresponding fault point A; Step b4: Finally, generate a fault point location network diagram corresponding to the anomaly detection, mark the feature data and associated dimensions corresponding to each fault point, and update the marked status of the network diagram through message passing to locate the multi-fault combination of concurrent failures.

[0011] Step 5: Verify and reproduce the fault location output results of the single-fault location model and the multi-fault location model.

[0012] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: By collecting the historical hierarchical data of chip failure fault analysis, extracting the feature data at each stage, and establishing the multi-dimensional feature association of the failure fault analysis data; based on the single failure root cause and the multiple failure root causes of associated interaction, a single-fault location model and a multi-fault location model are respectively established, which can provide accurate fault location methods for different failure situations, not only can accurately identify the location of single-failure faults, but also can locate the locations of multiple failure faults at the same time, solve the problem of low location accuracy caused by diverse and interacting influencing factors in the existing single-fault location method, and improve the efficiency and accuracy of chip failure fault location. Description of the Drawings

[0013] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings: Figure 1 It is a schematic diagram of the method step flow provided by the embodiment of the present invention; Figure 2 It is a schematic diagram of the system module composition provided by the embodiments of the present invention. Specific embodiments

[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0015] Please refer to Figure 1 , the embodiments of the present invention provide the following technical solutions: A digital chip fault detection method based on big data, comprising the following steps: Step 1: Collect historical hierarchical data of chip failure fault analysis, extract feature data at each stage, and establish a multi-dimensional feature association of failure fault analysis data based on the feature data; In this embodiment, the historical hierarchical data of chip failure fault analysis collected refers to: multiple information data related to chip failure fault analysis collected and recorded starting from the chip manufacturing stage, and at the same time, the levels are divided according to different stages experienced by the chip. This level includes the chip manufacturing stage, the chip detection stage, and the chip production output stage. The data in different stages have different effects on chip failure analysis and fault location. Similarly, the chip failure-related data collected in the chip manufacturing stage is used to determine the timeliness of chip failure; and the chip failure-related data collected in the chip detection stage is used to determine whether it is a direct chip failure detection problem or a hidden detection problem; among them, the direct detection problem is the corresponding failure mode matching situation of the failure phenomenon defined in the existing failure fault analysis equipment, and the hidden detection problem is the chip failure situation formed based on the later stage of the detection stage; finally, the chip failure-related data collected in the chip production output stage can synchronously feedback the timeliness of the chip failure data and determine the type of chip failure detection problem.

[0016] Exemplarily, characteristic data of different stages are extracted from the historical hierarchical data collected for chip failure analysis. Different failure analysis data collection methods are used in different stages, that is, different data collection dimensions are emphasized. In the initial stage of chip manufacturing, automated appearance detection is used to collect the appearance, morphology, and surface defects of the chip. Therefore, the characteristic data in the initial stage of chip manufacturing includes appearance data, morphological data, and surface defect data related to chip failure analysis. In the chip detection stage, X-Ray detection is used to penetrate the chip to observe its internal structure, and various performance tests are used to analyze the characteristics of the chip. Therefore, the characteristic data in the chip detection stage includes layer structure data, wire bonding data, and packaging defect data inside the chip related to chip failure analysis. Among them, the packaging defect data includes chip crack packaging defect data, uneven dispensing data, wire breakage data, wire bridging data, internal bubble data, etc. Secondly, the characteristic data related to performance tests includes leakage current value data, threshold voltage data, capacitance resistance data, etc. related to chip failure analysis. Finally, in the chip production output stage, through physical characteristic analysis and structural analysis, etc., the characteristic data in this stage includes line width data, doping concentration data, etc. related to chip failure analysis.

[0017] Furthermore, the correlation algorithm is used to correlate the characteristic data related to chip failure analysis collected in different stages with the corresponding dimensions. The corresponding dimensions include process parameter dimensions, electrical characteristic dimensions, material property dimensions, etc. related to chip failure analysis. The process parameter dimension uses hierarchical clustering algorithms and decision tree algorithms to map cross-stage characteristic data onto a multi-dimensional correlation network, that is, the process parameter dimension is used to horizontally correlate the characteristic data in different stages to locate the root cause of failure. The electrical characteristic dimension couples the characteristic data in different stages across levels through a Bayesian network to identify hidden interactive multi-root cause correlations of failures. Exemplarily, a quantum annealing algorithm is introduced in the process parameter dimension to optimize feature correlation. The process parameters are mapped to qubit energy states, and the local optimum problem that is difficult to break through by traditional hierarchical clustering is solved through the quantum tunneling effect. By quantifying the cross-stage process parameter differences as a Hamiltonian matrix , where is the parameter offset, that is, the offset between the actual value and the standard value of the th process parameter, is the coupling coefficient, , is the coupling term of different process parameters, indicating the mutual influence between parameters. Among them, , the ground state of the Hamiltonian is found through the quantum annealing algorithm. The ground state is the lowest energy state, so as to obtain the optimal correlation mapping of process parameters and break through the local optimum limit of traditional clustering algorithms. Add a cross-modal graph convolutional network (CM-GCN) in the material property dimension, construct a heterogeneous graph structure from the material SEM images, XRD patterns, and physical property parameters, and establish a correlation model between the microstructure and macroscopic properties of the material through heterogeneous data fusion. That is, define the node type mapping function as: Where, represents the multi-modal feature vector of node , is a three-dimensional convolution operation to extract the microstructure features of the material SEM image , is a long short-term memory network for processing the temporal diffraction peak information of the X-ray diffraction pattern , is a multi-layer perceptron for encoding the physical property parameters of the material , is a cross-modal feature splicing operator including channel splicing or weighted fusion, which calculates the edge weights through a cross-modal attention mechanism to capture the non-linear correlation between lattice defects and electrical parameters.

[0018] Step 2: Based on the single failure root cause and the associated interactive failure multiple root causes, establish a single fault location model and a multiple fault location model for chip failure faults respectively; In this embodiment, based on the characteristic data extracted from the failure fault analysis at different stages and the associated dimensions obtained according to the association algorithm, the failure fault location corresponding to the relevant data for failure fault analysis can be performed according to different stages, that is, judge the type of failure root cause corresponding to the association relationship between different dimensions and characteristic data. When it is judged as the single failure root cause association type, establish a single fault location model for chip failure faults according to the single failure root cause. When it is judged as the interactive failure multiple root cause association type, establish a multiple fault location model for chip failure faults according to the associated interactive failure multiple root causes; Exemplarily, the failure fault location corresponding to the single failure root cause is the input under the single fault location model. That is, when scanning and detecting chip failure faults starting from the initial stage of the chip manufacturing stage, the single fault location model detects and marks the failure fault location points based on the scanning logic that one scanning detection point corresponds to one fault point. When there are multiple fault points at one scanning detection point, if each fault point has an independent scanning detection point, the single fault model is still used. When any scanning detection point and two equivalent scanning detection points correspond to at least two fault points, use the bionic pulse neural network location model to construct a three-layer pulse network with synaptic plasticity. The input layer receives the characteristic pulse sequence, the hidden layer dynamically adjusts the synaptic weights using the STDP learning rule, and the output layer pulse firing frequency represents the fault probability. Define the synaptic reinforcement function as , and the calculation formula is: Among them, represents the change in synaptic weight from neuron to ; is the learning rate that controls the amplitude of weight update, , respectively represent the spike times of the presynaptic neuron i and the postsynaptic neuron j, is the time decay constant, which is used to determine the decay speed of the influence of the pulse time difference on the weight. When the spatio-temporal correlation of the multi-fault signal exceeds the defined threshold, synchronous firing is triggered. Based on the dynamic adjustment of the synaptic weight, the detection of multi-root cause coupling is realized, and the multi-fault location model of the chip failure fault established according to the failed multi-root causes of the associated interaction is analyzed to determine whether the fault points belong to the same feature dimension, and multi-label classification is performed according to the judgment result to obtain the failure direction of each feature data.

[0019] Step 3: Perform single failure fault location through the established single fault location model of the chip failure fault; In this embodiment, the feature data in different stages and different dimensions are cleaned to remove noise data, and standardized processing is performed, that is, the feature data with different dimensions are normalized, so that the specific data input into the model has consistency; Exemplarily, the specific operation process of the single fault location model includes: Step a1: Input the chip failure-related feature data with consistency, and determine the associated dimension based on the scan detection points and the corresponding fault points marked by the feature data; Step a2: Use the known failure types and the historical data corresponding to the feature data, perform training and matching based on the deep learning algorithm, output the mapping relationship between the feature data and the failure types, and determine the failure types corresponding to the feature data; Step a3: According to the determined failure types, use static causal diagnosis, dynamic causal diagnosis, and causal reverse diagnosis to locate the failure fault positions; In this embodiment, static causal diagnosis simulates all failure faults to obtain a complete failure fault dictionary, and then calculates the scores of each fault position , and the calculation formula is: Among them, is the fault detection weight, is the no-fault confirmation weight, is the number of simulation times of test failure, TF is the total number of test failures, is the number of simulation times of test pass, is the total number of test passes. When gets closer to 1, the simulation can accurately reproduce the measured fault. When When it approaches 1, the simulation has very few false alarms when there is no fault, that is, finally the higher the score, the greater the possibility of a fault; dynamic causal diagnosis screens out candidate faults through pre-tests, that is, using the formula: , where v is the fault value, that is, SA1 / SA0, and SA1 and SA0 respectively represent the fault value when a certain node is fixed to 1, and the fault value when a certain node is fixed to 0, is the parity expression of the number of logic inversions, with an even number of inversions being 0 and an odd number being 1, and f is the actual output value when the detection fails. It is judged whether the fault is added to the candidate list according to whether the input characteristic data satisfies the above formula, and then simulation is performed based on the candidate faults to generate a partial fault dictionary. Further, the simulation failure value SF and the test failure value TF are compared according to the above calculation formula; Furthermore, in combination with the path tracing algorithm, trace back from the scan detection point where the failure fault is found to the input end of the model. At the same time, introduce the photon path integral optimization algorithm to model the signal path as a photon propagation path. In this embodiment, the specific operation process of identifying key signals and locating failure faults in combination with the path tracing algorithm includes: Exemplarily, when the model input changes, the signal whose output changes is the key signal. The identification of the key signal includes: the input of the model is , and the output is , for each input , the criticality score is , , where is the change amount of the output , is the change amount of the input . When is greater than or equal to the set preset threshold , it is judged that is a key signal. When is less than the set preset threshold, it is judged that is a non-key signal; Furthermore, after identifying the key signal, continue to trace back to the input. When the change of a certain input does not affect the output, randomly select an input as the key signal and continue to trace back. That is, define the current scan detection point as , and its corresponding input path is , where represents the first iteration and the nth iteration. The iterative process of path tracing is: , where represents the th input signal detected in the th iteration, represents meeting the condition All input signals The set of, update the key signal set after each iteration, and continue to trace back to the input end; Specifically, after continuously tracing back to each scan detection point, delete the found failure faults. The real failure faults are the intersection of the faults detected by each vector, that is, the set of failure faults found through different scan detection points is , where each is a set of faults, the real failure faults The intersection calculation is , take the failure faults detected by coincidence as possible failure faults, and then use the correlation score to judge and delete the non-coincident failure faults, that is, when the correlation score is less than the set deletion threshold, delete the non-coincident failure faults , and finally compare the remaining failure faults with SF and TF to obtain the final failure fault position; At the same time, based on the introduced photon path integral optimization algorithm, define the weight of the photon path as , where is the inverse temperature parameter, used to control the randomness of path selection, and , is the Lagrangian, including the signal delay parameter and the noise interference parameter , , where is the noise influence coefficient, generate the three-dimensional key path probability cloud map of the final failure fault position through Monte Carlo optical path sampling, locate the key fault propagation path based on the final failure fault position, and calculate the probability of the path integral , where is the functional integral over all possible paths, is the action, when the probability of the path integral exceeds the defined critical value, it is determined as a high-probability fault propagation path.

[0020] Step 4: Perform multi-failure fault location through the established multi-failure fault location model of chip failure faults; In this embodiment, in different stages of the chip, various defects are introduced, resulting in various failure faults. By establishing a multi-failure fault location model of chip failure faults, based on a multi-dimensional association network, a hypergraph association rule mining engine is constructed, and the multi-fault features are constructed into a hypergraph structure H=(V,E) connecting multiple nodes with hyperedges. Among them, the hyperedge e∈E contains ≥3 associated feature nodes, and use the hypergraph Apriori algorithm: first generate k hyperedge candidate sets, and calculate the support , confidence , where is the support of the hyperedge e, N is the total number of chip test cases, DB is the database, T is a transaction in the database DB, representing the complete test data record of a single chip. In the chip fault detection scenario, each transaction corresponds to a set of all feature data collected during the entire process from chip manufacturing to testing. All node features of the hyperedge e appear simultaneously in the transaction T. is the rule → is the confidence of, indicating when appears the conditional probability. When it is detected that the hyperedge confidence exhibits the characteristics of quantum entanglement, a quantum Bayesian network is started for deviation correction calculation, calculating the non-classical correlation between multi-fault features, and mining and screening out meaningful association rules. These association rules are used to describe the correlation between the feature data corresponding to different fault points, that is, when "fault point A appears at feature data 1, then fault point B is very likely to also appear at feature data 1; and when feature data 3 has fault point A, the corresponding equivalent feature data 4 may also have the corresponding fault point A". Therefore, when an abnormality is detected at a certain feature data point, multiple possible fault points can be further inferred using the association rules.

[0021] Exemplarily, the specific operation process of the multi-fault location model includes: Step b1: Synchronously map the input chip failure-related feature data with consistency to the associated dimensions, and independently train a binary classification model for each failure root cause to infer the multi-root cause interaction causality, determine the failure type corresponding to the feature data, and extract the common features across failure types; Step b2: Calculate the linear relationship between the feature data and the failure type using the Pearson correlation coefficient, and capture the non-linear relationship through mutual information. At the same time, generate the association matching rules between the failure type and the fault point based on the extracted common features; Exemplarily, mutual information is a measure used in information theory to quantify the amount of information shared between two random variables. It measures the reduction in uncertainty of a random variable due to knowing another random variable. Mutual information is applicable not only to linear relationships but also to capture non-linear relationships. Among them, the failure type association matching rules include: Rule 1: When fault point A appears at feature data 1, then fault point B is very likely to also appear at feature data 1; Rule 2: When feature data 3 has fault point A, the corresponding equivalent feature data 4 may also have the corresponding fault point A; Step b3: Use a machine learning algorithm to perform outlier detection on the input chip failure-related feature data with consistency. When an abnormal feature data is detected, make an inference according to the mined association matching rules. If feature data 1 is abnormal and meets the conditions of rule 1, it is inferred that the fault point B is likely to appear at feature data 1; if feature data 3 is abnormal and meets the conditions of rule 2, it is inferred that the equivalent feature data 4 may also have the corresponding fault point A; Step b4: Finally, generate a fault point location network diagram corresponding to the anomaly detection, mark the feature data and associated dimensions corresponding to each fault point, and update the network diagram marking status through message passing to locate the multi-fault combination of concurrent failures.

[0022] Step Five: Verify and reproduce the fault location output results of the above single-fault location model and multi-fault location model respectively.

[0023] In this embodiment, use the test data to run the single-fault location model, record the fault point location information output by the single-fault location model, compare the fault point location output by the single-fault location model with the actual location of the known fault point, calculate the location error including distance error and area error, and evaluate the accuracy, stability and reliability of the single-fault location model based on the location error. When the location error is greater than the set error threshold, continuously iterate and optimize the single-fault location model; for the multi-fault location model, input the validation set containing the feature data of multiple fault points, record the location information of the multiple fault points output and the applied association rules, compare the locations of the multiple fault points output by the multi-fault location model with the actual locations of the known multiple fault points, calculate the location error of each fault point and the accuracy of the applied association rules, and evaluate the stability and reliability of the multi-fault location model based on the location error. When the location error is greater than the set error threshold, continuously iterate and optimize the multi-fault location model.

[0024] Combined with Figure 2 , the embodiment of the present invention provides a digital chip fault detection system based on big data, including the following functional modules: A multi-dimensional feature data acquisition and association module, which is used to collect feature data related to chip failure fault analysis and associate the feature data with the corresponding dimensions by using an association algorithm; A failure root cause type judgment module, which is used to judge the failure root cause type corresponding to the association relationship between different dimensions and feature data; A single-fault location model construction module, which is used to establish a single-fault location model for chip failure faults based on a single failure root cause; A multi-fault location model construction module, which is used to establish a multi-fault location model for chip failure faults according to the associated and interacting multiple failure root causes; A fault location output module, configured to output the final fault point location based on a single fault location model and a multi-fault location model; A fault location output result verification module, configured to verify and optimize the final fault location points output by the single fault location model and the multi-fault location model.

[0025] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0026] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, they can still modify the above-described technical solutions or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A digital chip fault detection method based on big data, characterized in that: The following steps are involved: Step 1: Collect historical hierarchical data of chip failure analysis, extract feature data of each stage of the chip, and establish multi-dimensional feature association of failure analysis data based on the feature data; The establishment of multi-dimensional feature association of failure analysis data includes: using an association rule mining algorithm to associate chip failure analysis-related feature data collected at different stages with corresponding dimensions, wherein the corresponding dimensions include process parameter dimensions, electrical feature dimensions, and material property dimensions related to chip failure analysis; Step 2: Based on the single failure root cause and the multiple failure root causes with correlation and interaction, a single fault location model and a multiple fault location model of chip failure are established respectively; Step 3: locating a single failure fault using the single fault location model; The single fault location model locates the failure fault position according to the determined failure type using static causal diagnosis, dynamic causal diagnosis and causal reverse diagnosis. The causal reverse diagnosis and location uses a path tracing algorithm and a photon path integral optimization algorithm to trace the source from the scanning detection point where the failure fault is found to the model input end, identify key signals and locate the failure fault, and calculate the fault propagation path based on the failure fault. Step 4: locating multiple faults using the multiple fault location model; The multi-fault location model is based on a multi-dimensional association network, constructs a hypergraph association rule mining engine, constructs multi-fault features into a hypergraph structure with hyperedges connecting multiple nodes, and uses the hypergraph Apriori algorithm to mine and filter out meaningful association rules. The association rules are used to describe the correlation between feature data corresponding to different fault points; Step 5: Verify and reproduce the fault location output results of the single fault location model and the multiple fault location model.

2. A digital chip fault detection method based on big data according to claim 1, characterized in that: The process parameter dimension uses quantum annealing algorithm to optimize feature association, maps process parameters to quantum bit energy states, and quantizes cross-stage process parameter differences into Hamiltonian matrices. ,in is the parameter offset, i.e. The deviation between the actual value and the standard value of a process parameter, is the coupling coefficient, , is the coupling term of different process parameters, indicating the mutual influence between parameters, where , through the quantum annealing algorithm, the ground state of the Hamiltonian is found. The ground state is the lowest energy state, the optimal correlation mapping of the process parameters is obtained, and the root cause of the failure is located; The electrical characteristic dimension couples the characteristic data of different stages across levels through a Bayesian network to identify hidden interactive multi-root causes of failure; The material property dimension uses a cross-modal graph convolutional network to construct a heterogeneous graph structure of material SEM images, XRD spectra and physical property parameters. By fusing heterogeneous data, a correlation model of material microstructure-macro performance is established, that is, the node type mapping function is defined as: ,in, Representation Node The multimodal feature vector of Extract material SEM images for 3D convolution operations The microstructural characteristics of Long short-term memory network for processing X-ray diffraction patterns Time series diffraction peak information, It is a multi-layer perceptron used to encode material properties , For cross-modal feature concatenation operators including channel concatenation or weighted fusion, edge weights are calculated through a cross-modal attention mechanism to capture the nonlinear association between lattice defects and electrical parameters.

3. A digital chip fault detection method based on big data according to claim 2, characterized in that: The establishing of a single fault location model and a multiple fault location model for chip failure faults respectively specifically includes: determining the failure root cause type corresponding to the correlation relationship between different dimensions and the feature data; when it is determined to be a single failure root cause correlation type, establishing a single fault location model for chip failure fault according to the single failure root cause; when it is determined to be an interactive multiple failure root cause correlation type, establishing a multiple fault location model for chip failure fault according to the interactive multiple failure root causes.

4. A digital chip fault detection method based on big data according to claim 3, characterized in that: The step 2 also includes: the failure fault location corresponding to the single failure root cause is used as the input under the single fault location model, that is, when scanning and detecting chip failure faults from the early stage of chip manufacturing, the single fault location model performs detection marking of the failure fault location point based on the scanning logic that one scanning detection point corresponds to one fault point, and when there are multiple fault points in one scanning detection point, if each fault point has an independent scanning detection point, the single fault model is still used, and when any scanning detection point and two equivalent scanning detection points correspond to at least two fault points, a bionic pulse neural network positioning model is used to construct a three-layer pulse network with synaptic plasticity, the input layer receives the characteristic pulse sequence, the hidden layer uses the STDP learning rule to dynamically adjust the synaptic weight, the output layer pulse emission frequency represents the fault probability, and the synaptic reinforcement function is defined as , the calculation formula is: in, Representing neurons arrive The change in synaptic weight, To control the learning rate of the weight update amplitude, , represent the pulse firing time of presynaptic neuron i and postsynaptic neuron j respectively, is the time decay constant, which is used to determine the decay speed of the influence of the pulse time difference on the weight. When the spatiotemporal correlation of multiple fault signals exceeds the defined threshold, the coordinated discharge is triggered. The multi-root cause coupling detection is realized based on the dynamic adjustment of the synaptic weights. The multi-fault location model of the chip failure fault established according to the multiple root causes of the associated interaction is analyzed to determine whether the fault point belongs to the same feature dimension. Multi-label classification is performed according to the judgment result to obtain the failure direction of each feature data.

5. A digital chip fault detection method based on big data according to claim 4, characterized in that: The specific operation process of the single fault location model includes: Step a1: input consistent chip failure related feature data, and determine the associated dimensions based on the scan detection points marked by the feature data and the corresponding fault points; Step a2: Use the known failure types and historical data corresponding to the feature data to perform training matching based on the deep learning algorithm, output the mapping relationship between the feature data and the failure type, and determine the failure type corresponding to the feature data; Step a3: According to the determined failure type, the failure fault position is located using static causal diagnosis, dynamic causal diagnosis and causal reverse diagnosis.

6. A digital chip fault detection method based on big data according to claim 5, characterized in that: The specific method of the static causal diagnosis is: by simulating all failure faults, a complete failure fault dictionary is obtained, and the score of each fault location is calculated. , the calculation formula is: ,in, is the fault detection weight, is the no-fault confirmation weight, is the number of test failure simulations, TF is the total number of test failures, is the number of simulations passed by the test, is the total number of test passes, when When it is closer to 1, the simulation can accurately reproduce the measured fault. The closer it is to 1, the fewer false alarms there are when the simulation is fault-free. The higher the score, the greater the possibility of failure; The specific method of the dynamic causal diagnosis is: screen out candidate faults through pre-testing, and use the formula: , where v is the fault value, i.e., SA1 / SA0. SA1 and SA0 represent the fault value of a node being fixed to 1 and the fault value of a node being fixed to 0, respectively. is the parity expression of the number of logic inversions, where an even number of inversions is 0 and an odd number of inversions is 1. f is the actual output value that failed the test. Whether the fault is added to the candidate list is determined based on whether the input feature data satisfies the above formula. Then, simulation is performed based on the candidate faults to generate a partial fault dictionary. The simulation failure value SF and the test failure value TF are further compared based on the above calculation formula. The causal reverse diagnosis positioning is as follows: based on the path tracing algorithm, the source is traced from the scanning detection point where the failure is discovered to the model input end, the key signal is identified and the failure is located, and at the same time, a photon path integral optimization algorithm is introduced to model the signal path as a photon propagation path.

7. A digital chip fault detection method based on big data according to claim 6, characterized in that: The specific operation process of combining the path tracking algorithm to identify key signals and locate failure faults includes: When the model input changes, the signal of output change is the key signal. The identification of the key signal includes: the input of the model is , the output is , for each input , the criticality score is , ,in, is the output The amount of change, Yes Input The change in Greater than or equal to the preset threshold When judging As a key signal, When it is less than the preset threshold, It is a non-critical signal; After identifying the key signal, continue to trace the input. When a change in a certain input does not affect the output, randomly select an input as the key signal to continue tracing the source, that is, define the current scanning detection point as , the corresponding input path is ,in Represents the first iteration and the nth iteration, and the iterative process of path tracing is: ,in, Indicates The first detected Input signal, Indicates that the conditions are met All input signals After each iteration, the key signal set is updated and the traceability continues to the input end; After tracing the source of each scanning and detection point, the failure faults found are deleted. The real failure fault is the intersection of the faults measured by each vector, that is, the failure fault set found by different scanning and detection points is , where each is a set of faults, real failure faults The intersection is calculated as , take the failure detected by coincidence test as the possible failure, and then use the correlation score The non-overlapping failures are deleted and judged, that is, when the correlation score When it is less than the set deletion threshold, the non-overlapping failure faults are deleted. ,Finally, compare SF and TF for the remaining failure faults to obtain the final failure fault location; At the same time, based on the photon path integral optimization algorithm, the weight of the photon path is defined as ,in is the inverse temperature parameter, which is used to control the randomness of path selection, and , is the Lagrangian, including the signal delay parameter and noise interference parameters , ,in The noise influence coefficient is obtained by Monte Carlo optical path sampling to generate the final three-dimensional critical path probability cloud map of the failure fault location. Based on the final failure fault location, the critical fault propagation path is located and the probability of path integral is calculated. ,in is the functional integral over all possible paths, is the action, when the probability of the path integral When the defined critical value is exceeded, it is determined to be a high-probability fault propagation path.

8. The method for digital chip fault detection based on big data according to claim 7, characterized in that: The use of the hypergraph Apriori algorithm to mine and filter out meaningful association rules includes: Based on the multi-dimensional association network, a hypergraph association rule mining engine is constructed. The multi-fault features are constructed into a hypergraph structure H=(V,E) with hyperedges connecting multiple nodes, where the hyperedge e∈E contains ≥3 associated feature nodes. The hypergraph Apriori algorithm is used: first, k hyperedge candidate sets are generated, and the support is calculated. , confidence ,in is the support of hyperedge e, N is the total number of chip test cases, DB is the database, T is a transaction in the database DB, which represents the complete test data record of a single chip. In the chip fault detection scenario, each transaction corresponds to a set of all feature data collected from the entire process of manufacturing to testing of a chip. All node features of hyperedge e appear simultaneously in transaction T, For rules → The confidence level, indicating When it appears When the hyperedge confidence is detected to show quantum entanglement characteristics, the quantum Bayesian network is started to perform correction calculations, calculate the non-classical correlations between multiple fault features, and mine and screen out meaningful association rules. These association rules are used to describe the correlation between the feature data corresponding to different fault points.

9. A digital chip fault detection method based on big data according to claim 8, characterized in that: The specific operation process of the multi-fault location model includes: Step b1: synchronously map the input consistent chip failure-related feature data to the associated dimensions, independently train a binary classification model for each failure root cause, infer the interactive causality of multiple root causes, determine the failure type corresponding to the feature data, and extract common features across failure types; Step b2: The linear relationship between the feature data and the failure type is calculated using the Pearson correlation coefficient, and the nonlinear relationship is captured through mutual information. At the same time, the association matching rules between the failure type and the fault point are generated based on the extracted common features; Step b3: Use a machine learning algorithm to detect abnormal points on the input consistent chip failure-related feature data. When a feature data is detected to be abnormal, inference is made based on the mined association matching rules. If feature data 1 is abnormal and meets the conditions of rule 1, it is inferred that fault point B is likely to appear in feature data 1; if feature data 3 is abnormal and meets the conditions of rule 2, it is inferred that the corresponding fault point A may also appear in equivalent feature data 4; Step b4: Finally, a fault point location network diagram corresponding to the anomaly detection is generated, the characteristic data and associated dimensions corresponding to each fault point are marked, and the marking status of the network diagram is updated through message passing to locate the multi-fault combination of concurrent failures.

10. A digital chip fault detection system based on big data, characterized in that: Includes the following functional modules: The multi-dimensional feature data acquisition and association module is used to collect feature data related to chip failure analysis and associate the feature data with the corresponding dimensions using an association algorithm; A failure root cause type judgment module is used to judge the failure root cause type corresponding to the correlation between different dimensions and feature data; A single fault localization model building module is used to build a single fault localization model of chip failure based on a single failure root cause; A multi-fault localization model building module is used to establish a multi-fault localization model for chip failures based on multiple root causes of failures that are correlated and interacted; A fault location output module is used to output the final fault point location based on a single fault location model and a multi-fault location model; The fault location output result verification module is used to verify and optimize the final fault location points output by the single fault location model and the multiple fault location models.

Citation Information

Cited By

  • Chip function detection method based on Internet protocol (IP), electronic equipment and medium

    CN120870834A

  • Observable container cluster fault operation and maintenance method, device, equipment and medium

    CN121037239A

  • Method for testing memory chip with redundancy function

    CN121506228A

  • A method for implementing redundant functional memory chip testing

    CN121506228B

  • Classification decision and intelligent management and control method for chip failure analysis path

    CN121805822A