Chip performance evaluation and fault diagnosis optimization test method
By generating a test vector library with the best coverage and a multi-dimensional performance feature model, combined with the path characteristic map, the problems of insufficient coverage and difficulty in fault location in chip testing are solved, and efficient fault diagnosis and accurate performance evaluation are achieved.
Patent Information
- Application Number
- CN202510630911.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-02
AI Technical Summary
The existing chip testing methods have insufficient coverage, low abnormal identification granularity, low fault diagnosis efficiency, and difficult to accurately identify and locate potential implicit faults in complex operating scenarios.
Based on the functional modules and circuit structure characteristics of chip design, a test vector library with the best coverage is generated, dynamic performance parameters are collected in real time, multi-dimensional performance feature models are built, and a fault function module is reverse traceable with the path characteristic diagram.
It significantly improves the comprehensiveness and efficiency of the test, can accurately characterize the performance characteristics of the chip in multiple scenarios, quickly locate the fault module, and improves the accuracy and reliability of fault diagnosis.
Smart Images

Figure CN120577671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chip performance evaluation and testing, and in particular to a chip performance evaluation and fault diagnosis optimization testing method. Background Art
[0002] With the increasing complexity of chip design and the rapid development of integrated circuit processes, modern chips are gradually moving towards high performance, high density, and multi-functionality. However, this also brings more severe challenges in performance evaluation and fault diagnosis. The performance stability and fault tolerance of chips under various operating conditions are directly related to their reliability and lifespan in actual applications. Therefore, how to comprehensively and accurately evaluate chip performance and quickly diagnose potential hidden faults has become a core issue that needs to be addressed in the field of chip design and testing.
[0003] Existing chip testing and fault diagnosis methods have the following shortcomings:
[0004] Insufficient test coverage: Traditional testing methods rely on static test vector libraries, which usually cannot fully cover all critical paths in chip operation, especially under boundary conditions, which can easily lead to the omission of potential hidden faults.
[0005] Low anomaly identification granularity: Existing anomaly detection algorithms are mostly based on single parameters or simple rules, which make it difficult to capture the correlation characteristics between multi-dimensional performance data in complex operating scenarios. This results in low anomaly identification accuracy and difficulty in effectively classifying abnormal patterns.
[0006] Low fault diagnosis efficiency: For detected anomalies, traditional methods mainly rely on manual analysis or coarse-grained fault speculation, which has low diagnostic efficiency and makes it difficult to trace the root cause of the anomaly. Especially in complex chips, the high coupling between functional modules further increases the difficulty of fault location. Summary of the Invention
[0007] The present invention provides a chip performance evaluation and fault diagnosis optimization test method.
[0008] The chip performance evaluation and fault diagnosis optimization test method includes the following steps:
[0009] S1, intelligent test vector library generation: Generates a test vector library with optimal coverage based on the functional module distribution and circuit structure characteristics of the chip design;
[0010] S2, dynamic performance acquisition: In the chip operating environment, test conditions are sequentially input using the test vector library, while dynamic performance parameters are collected in real time. The dynamic performance parameters include voltage fluctuation, power consumption curve, heat dissipation, and signal delay, and a dynamic performance data set is constructed based on the dynamic performance parameters;
[0011] S3, Multidimensional Performance Model Construction: Based on the dynamic performance dataset, a multidimensional performance characteristic model is established, including performance fluctuation characteristics, critical state thresholds, and latent fault characteristics. The multidimensional performance characteristic model is optimized using a nonlinear fitting algorithm to characterize the performance characteristics of the chip in multiple scenarios.
[0012] S4, fault module analysis: Based on the multi-dimensional performance characteristic model combined with the circuit structure characteristics, reverse trace the correlation between the chip functional modules and locate the specific faulty functional module.
[0013] Optionally, the S1 specifically includes:
[0014] S11, Chip Functional Module Analysis: Extract the functional module layout of the chip design, identify the logical connections and operation paths of each functional module, and mark the key nodes in the functional module, including input ports, output ports and data transmission paths, to provide a structured reference for test vector generation;
[0015] S12, circuit structure feature extraction: Combine the chip's circuit design files and layout information to obtain the circuit's topological characteristics, signal transmission delay characteristics, power consumption hotspots, and frequency distribution. Build a path characteristic graph R between functional modules and identify the critical paths that affect chip performance.
[0016] S13, multi-scenario operation simulation: Based on the functional modules and circuit structure characteristics, multiple test scenarios of the chip are simulated, including normal load mode, high load mode, low power mode and boundary conditions, to generate an initial test scenario set;
[0017] Generate the initial test scene set S={S1,S2,...,S i}, each scene S i Including operation mode and input parameter configuration.
[0018] S14, test vector library construction: using the key path in the path characteristic graph, generate an initial test vector set for each test scenario. The test vector set is expressed as: V = {v1, v2, ..., v i} represents that, where each test vector v i Including specific signal input combination: input voltage V in 、Input current I in , clock frequency f, data input signal S data , ensuring that the paths of different scenarios are covered, integrating the test vectors of all test scenarios to form a complete test vector library V 库 .
[0019] Optionally, the S2 specifically includes:
[0020] S21, test vector input and test condition loading: select test vector V from the test vector library i , where V i ={V in ,I in ,f,S data}, load test condition C i The test conditions are defined by the operating scenario, including external environmental parameters (temperature, humidity) and load power L load , the test vector V i and test condition C i Synchronous input chip test environment to ensure the chip performs under specified operating conditions;
[0021] S22, deploy a multi-dimensional sensor network to collect dynamic performance parameters in real time, including voltage fluctuation ΔV(t), power consumption curve P(t), heat dissipation T(x,y,t), and signal delay τ(t);
[0022] S23, dynamic performance data set construction: The collected multi-dimensional dynamic performance parameters are integrated into a data set D(t) = {ΔV(t), P(t), T(x, y, t), τ(t)} in time series, where the parameters ΔV, P, T, and τ correspond to voltage fluctuation, power consumption, heat dissipation, and signal delay, respectively;
[0023] S24, data annotation: according to the input test condition C i Add labels to dynamic performance datasets to indicate the corresponding test vectors and running scenarios.
[0024] Optionally, in S22:
[0025] The voltage fluctuation ΔV(t) is calculated by collecting the voltage changes of each functional module of the chip through the voltage sensor: ΔV(t) = V max (t)-V min (t); V max (t) is the peak voltage, V min (t) is the voltage valley value, t represents time;
[0026] The power consumption curve P(t) uses a power monitor to record the overall power consumption P(t) of the chip and the power consumption P(t) of the local module in real time. i (t), t represents time;
[0027] The heat dissipation T(x,y,t) is recorded by a thermocouple array at the time t at each coordinate point (x,y) on the chip layout to form a spatiotemporal heat dissipation map;
[0028] The signal delay τ(t) is measured by the timing analysis tool to measure the delay of data transmission. The delay value τ(t) = t arrival -t sendIndicates the difference between the signal arrival time and the sending time.
[0029] Optionally, the S3 multi-dimensional performance model construction specifically includes:
[0030] S31, normalize the dynamic performance data set D(t) and map each performance parameter to a uniform range [0,1];
[0031] S32, performance fluctuation feature extraction: Calculate the fluctuation characteristics of dynamic performance parameters, including mean μ, variance σ 2 and instantaneous rate of change r(t), the extracted fluctuation characteristics μ, σ 2 , r(t) is integrated into the eigenvector F 波动 ;
[0032] S33, critical state threshold calculation: Based on the historical data of dynamic performance parameters, define the critical state threshold T 临界 , which indicates the safe range of the chip in normal operation, including the voltage fluctuation threshold T ΔV , power consumption curve threshold T P , heat dissipation threshold T T And the signal delay threshold T τ :
[0033] S34, Hidden fault feature extraction: Use anomaly detection algorithm to analyze data that deviates from the critical state threshold and extract hidden fault features F 隐性 , such as overload failure, high temperature signal offset, etc.;
[0034] S35, the performance fluctuation feature F 波动 , critical state threshold T 临界 and hidden fault characteristics F 隐性 , integrated into a multi-dimensional performance characteristic model U, expressed as: U = {F 波动 ,T 临界 ,F 隐性}, Model U describes the comprehensive characteristics of chip performance in multiple scenarios.
[0035] Optionally, S3 also includes nonlinear fitting optimization: taking the dynamic performance data set D(t) as input, nonlinear fitting optimization is performed on the multidimensional performance characteristic model U, with the goal of minimizing the error between the model output and the actual performance parameters.
[0036] Optionally, the anomaly detection algorithm in S34 adopts principal component analysis, specifically including:
[0037] S341, Covariance Matrix Calculation: Based on the normalized dynamic performance data set, a covariance matrix is established to describe the linear relationship between the various performance data. The principal component eigenvalues and principal component eigenvectors are calculated for the covariance matrix, and the principal components with the first few cumulative contributions reaching 95% are selected.
[0038] S342, principal component dimensionality reduction: Use the selected principal components to project the data into a low-dimensional space to obtain a low-dimensional representation;
[0039] S343, Reconstructing Data and Error Calculation: Reconstructing the original data using the low-dimensional representation generated by the principal components, and calculating the reconstruction error between the actual data and the reconstructed data. The larger the error, the higher the probability of anomaly.
[0040] S344, outlier determination: Based on the calculated reconstruction error, define the Euclidean norm of the reconstruction error as an anomaly score, and determine the outlier based on the anomaly score;
[0041] S345, Hidden Fault Feature Extraction: Extract relevant dynamic performance parameters (such as voltage fluctuations at specific time points, power consumption curve deviations, abnormal heat dissipation distribution, etc.) from the data marked as abnormal points to form a hidden fault feature set to describe the chip problem.
[0042] Optionally, determining an anomaly point based on the anomaly score in S344 specifically includes: taking the 99% quantile of the reconstruction error distribution as a determination point, comparing the anomaly score with the 99% quantile of the reconstruction error distribution, and determining that a point exceeding the determination point is an anomaly point;
[0043] Optionally, the fault module analysis in S4 specifically includes:
[0044] Extract abnormal performance characteristics from the multidimensional performance characteristic model U, including performance fluctuation characteristics, deviation from critical state thresholds, and latent fault characteristics. Map the abnormal performance characteristics to specific functional modules based on the characteristic attribution of the multidimensional performance characteristic model U. Each abnormal performance characteristic corresponds to one or more abnormal functional modules.
[0045] Use the path characteristic graph R to trace the critical path including the abnormal functional module, confirm the path of the affected functional module, and mark the modules in the path with increased delay, abnormal power consumption, or disturbed signal transmission.
[0046] Fault module screening: Comprehensive abnormal performance characteristics and path characteristics to screen faulty functional modules:
[0047] Functional modules that multiple abnormal performance characteristics focus on;
[0048] The starting point of anomaly propagation or the node affected by the accumulation on the path.
[0049] Beneficial effects of the present invention:
[0050] The present invention dynamically generates a test vector library with optimal coverage by combining the functional module distribution and circuit structure characteristics of the chip design, covering a variety of operating scenarios (such as high load, low power consumption mode) and boundary conditions. This dynamic generation method solves the problem of insufficient coverage in traditional testing methods, effectively reduces test blind spots, and significantly improves the comprehensiveness and efficiency of testing. Combined with dynamic performance acquisition technology, it records multi-dimensional performance parameters of the chip under operating conditions (such as voltage fluctuations, power consumption curves, heat dissipation and signal delays) in real time, ensuring the timeliness and accuracy of test data, and providing high-quality data support for subsequent performance evaluation and fault analysis.
[0051] The present invention, based on a method of a multidimensional performance characteristic model, realizes a panoramic description of chip performance for the first time by comprehensively modeling data of performance fluctuation characteristics, critical state thresholds and latent fault characteristics. The model is optimized through a nonlinear fitting algorithm and can accurately characterize the performance characteristics of the chip in multiple scenarios. Through the feature attribution method in the performance characteristic model, abnormal features can be quickly mapped to specific functional modules, and the source of the abnormality can be tracked in combination with the path characteristic diagram. This module-level fine positioning breaks through the granularity limitations of traditional anomaly detection and provides innovative technical support for chip design optimization and improved operational stability.
[0052] This paper combines anomaly detection results with a multidimensional performance characteristic model and path characteristic diagrams to propose a logically clear fault root cause analysis method. By using anomaly characteristics to locate specific functional modules and trace the source of anomaly propagation in critical paths, it can quickly classify and identify fault types (such as overload failure and signal offset). Compared to traditional methods that rely on manual inference, this automated root cause analysis method can significantly shorten fault diagnosis time and improve diagnostic accuracy and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0054] Figure 1 A schematic flow chart of an evaluation and diagnostic test method according to an embodiment of the present invention;
[0055] Figure 2 Schematic diagram of the anomaly detection algorithm flow in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It is also noted that, to provide a more detailed description, the following embodiments are best and preferred embodiments, and those skilled in the art may employ alternative methods for implementing certain known technologies. Furthermore, the accompanying drawings are intended only to provide a more detailed description of the embodiments and are not intended to limit the present invention.
[0057] It should be noted that references in the specification to "one embodiment," "an embodiment," "an exemplary embodiment," "some embodiments," etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not every embodiment necessarily includes such specific features, structures, or characteristics. In addition, when specific features, structures, or characteristics are described in conjunction with an embodiment, it is within the knowledge of persons skilled in the relevant art to implement such features, structures, or characteristics in conjunction with other embodiments (whether or not explicitly described).
[0058] In general, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending at least in part on the context, allow for the presence of other factors that are not necessarily explicitly described.
[0059] like Figure 1-Figure 2 As shown, the chip performance evaluation and fault diagnosis optimization test method includes the following steps:
[0060] S1, intelligent test vector library generation: Generates a test vector library with optimal coverage based on the functional module distribution and circuit structure characteristics of the chip design;
[0061] S2, dynamic performance acquisition: In the chip operating environment, test conditions are sequentially input using a test vector library, while dynamic performance parameters are collected in real time. Dynamic performance parameters include voltage fluctuations, power consumption curves, thermal dissipation, and signal delays. A dynamic performance dataset is then constructed based on these dynamic performance parameters.
[0062] S3, Multidimensional Performance Model Construction: Based on the dynamic performance dataset, a multidimensional performance characteristic model is established, including performance fluctuation characteristics, critical state thresholds, and latent fault characteristics. The multidimensional performance characteristic model is optimized using a nonlinear fitting algorithm to characterize the performance characteristics of the chip in multiple scenarios.
[0063] S4, fault module analysis: Based on the multi-dimensional performance characteristic model combined with the circuit structure characteristics, reverse trace the correlation between the chip functional modules and locate the specific faulty functional module.
[0064] S1 specifically includes:
[0065] S11, Chip Functional Module Analysis: Extract the functional module layout of the chip design, identify the logical connections and operation paths of each functional module, and mark the key nodes in the functional module, including input ports, output ports and data transmission paths, to provide a structured reference for test vector generation;
[0066] The functional modules in chip design are represented by a set as M = {M1, M2, ..., M n}, where M i Represents the i-th functional module.
[0067] Mark key nodes:
[0068] Input port set: I={I1,I2,…,I p};
[0069] Output port set: O={O1, O2,…, O q};
[0070] Data transmission path matrix: D ij Indicates that the module M i To module M j The data transmission path, if there is a path, then D ij =1, otherwise D ij =0.
[0071] Form a logical topology structure between functional modules, which serves as the basis for test vector generation.
[0072] S12, circuit structure feature extraction: Combine the chip's circuit design files and layout information to obtain the circuit's topological characteristics, signal transmission delay characteristics, power consumption hotspots, and frequency distribution. Build a path characteristic graph R between functional modules and identify the critical paths that affect chip performance.
[0073] By parsing the design file, the circuit connection characteristics of the chip are extracted and represented by a directed graph G = (V, E), where V represents the node set in the circuit (functional modules, input ports, output ports), and E represents the signal transmission path set.
[0074] The signal transmission delay characteristics are represented by matrix T, where T ij represents the signal propagation delay from node i to node j;
[0075] Power consumption hotspot identification: Combined with the running power consumption data P(x,y), use the two-dimensional coordinates (x,y) to mark the power consumption distribution in the chip layout, and extract the high power consumption area P(x,y)>P threshold ;P threshold is the power consumption threshold;
[0076] Frequency distribution analysis: Count the operating frequencies of each module during operation to form a frequency distribution F(M i ), where F represents module M i Frequency range;
[0077] Construct a path characteristic diagram R = (G, T, P, F) to provide a key reference for subsequent test scenario simulation;
[0078] S13, multi-scenario operation simulation: Based on the functional modules and circuit structure characteristics, multiple test scenarios of the chip are simulated, including normal load mode, high load mode, low power mode and boundary conditions, to generate an initial test scenario set;
[0079] Test scenario definition:
[0080] Normal load scenario: Assume input voltage V in =V nom (nominal voltage), operating frequency f=f nom ;
[0081] High load scenario: Assume the load power L load =αL nom , where α>1;
[0082] Low power mode: reduce the input voltage V in =βV nom , where 0<β<1;
[0083] Boundary Condition Scenario: Setting V in,max =V nom +ΔV,T max =T nom +ΔT, where ΔV and ΔT represent the magnitude of the change in voltage and temperature;
[0084] Path selection rule: Prioritize high-power consumption and long-delay paths in the path characteristic graph R. High-power consumption and long-delay paths are considered critical paths.
[0085] Generate the initial test scene set S={S1,S2,...,S i}, each scene S i Including operation mode and input parameter configuration.
[0086] S14, test vector library construction: using the key path in the path characteristic graph, generate an initial test vector set for each test scenario. The test vector set is expressed as: V = {v1, v2, ..., v i} represents that, where each test vector v i Including specific signal input combination: input voltage V in 、Input current I in, clock frequency f, data input signal S data , ensuring that the paths of different scenarios are covered, integrating the test vectors of all test scenarios to form a complete test vector library V 库 ;
[0087] Vector generation rule: select the corresponding input signal combination according to the key path in the path characteristic graph R;
[0088] Each test scenario S i Mapped into a set of test vectors V(S i ), by adjusting V in ,f,S data Generate test inputs that cover critical paths;
[0089] The vector sets of all test scenarios are integrated into the test vector library: A complete test vector library covering multiple scenarios, multiple paths, and multiple boundary conditions is formed for chip performance evaluation and fault diagnosis.
[0090] S2 specifically includes:
[0091] S21, test vector input and test condition loading: select test vector V from the test vector library i , where V i ={V in ,I in ,f,S data}, load test condition C i , the test conditions are defined by the operating scenario, including external environmental parameters (temperature, humidity) and load power L load , the test vector V i and test condition C i Synchronous input chip test environment to ensure the chip performs under specified operating conditions;
[0092] S22, deploy a multi-dimensional sensor network to collect dynamic performance parameters in real time, including voltage fluctuation ΔV(t), power consumption curve P(t), heat dissipation T(x,y,t), and signal delay τ(t);
[0093] S23, dynamic performance data set construction: The collected multi-dimensional dynamic performance parameters are integrated into a data set D(t) = {ΔV(t), P(t), T(x, y, t), τ(t)} in time series, where the parameters ΔV, P, T, and τ correspond to voltage fluctuation, power consumption, heat dissipation, and signal delay, respectively;
[0094] The data is stored in matrix form: Where n is the number of collection time points;
[0095] S24, data annotation: according to the input test condition Ci Add labels to dynamic performance datasets to indicate the corresponding test vectors and running scenarios.
[0096] Label content: Identifies the test strip Ci and test vector Vi corresponding to each set of dynamic performance parameters (such as voltage fluctuation, power consumption curve, etc.), including:
[0097] Test scenarios;
[0098] Input parameters;
[0099] External environment parameters.
[0100] In S22:
[0101] The voltage fluctuation ΔV(t) is calculated by collecting the voltage changes of each functional module of the chip through the voltage sensor: ΔV(t) = V max (t)-V min (t); V max (t) is the peak voltage, V min (t) is the voltage valley value, t represents time;
[0102] The power consumption curve P(t) uses a power monitor to record the overall power consumption P(t) of the chip and the power consumption P(t) of the local module in real time. i (t), t represents time;
[0103] The heat dissipation T(x,y,t) is recorded by a thermocouple array at the time t at each coordinate point (x,y) on the chip layout to form a spatiotemporal heat dissipation map;
[0104] The signal delay τ(t) is measured by the timing analysis tool to measure the delay of data transmission. The delay value τ(t) = t arrival -t send Indicates the difference between the signal arrival time and the sending time.
[0105] The S3 multi-dimensional performance model construction specifically includes:
[0106] S31, normalize the dynamic performance data set D(t) and map each performance parameter to a uniform range [0,1], using the formula: Where X is any performance parameter, X min and X max are the minimum and maximum values of the parameters respectively;
[0107] S32, performance fluctuation feature extraction: Calculate the fluctuation characteristics of dynamic performance parameters, including mean μ, variance σ 2 and the instantaneous rate of change r(t), where: X(t) is any performance parameter at time t, X(t-1) is any performance parameter at time t-1, Δt is the time change, and the extracted fluctuation features μ, σ2 , r(t) is integrated into the eigenvector F 波动 ;
[0108] S33, critical state threshold calculation: Based on the historical data of dynamic performance parameters, define the critical state threshold T 临界 , which indicates the safe range of the chip in normal operation, including the voltage fluctuation threshold T ΔV , power consumption curve threshold T P , heat dissipation threshold T T And the signal delay threshold T τ :
[0109] Voltage fluctuation threshold T ΔV =[ΔV min ,ΔV max ];
[0110] Power consumption curve threshold T P =[P min ,P max ];
[0111] Thermal dissipation threshold T T =[T min ,T max ];
[0112] Signal delay threshold T τ =[τ min ,τ max ];
[0113] S34, Hidden fault feature extraction: Use anomaly detection algorithm to analyze data that deviates from the critical state threshold and extract hidden fault features F 隐性 , such as overload failure, high temperature signal offset, etc.;
[0114] S35, the performance fluctuation feature F 波动 , critical state threshold T 临界 and hidden fault characteristics F 隐性 , integrated into a multi-dimensional performance characteristic model U, expressed as: U = {F 波动 ,T 临界 ,F 隐性}, Model U describes the comprehensive characteristics of chip performance in multiple scenarios.
[0115] S3 also includes nonlinear fitting optimization: using the dynamic performance data set D(t) as input, it performs nonlinear fitting optimization on the multidimensional performance characteristic model U. The goal is to minimize the error between the model output and the actual performance parameters. The optimization formula is: in, is the optimized multi-dimensional performance characteristic model, D i is the i-th dynamic performance data, is the corresponding target performance value, and θ is the model parameter.
[0116] The anomaly detection algorithm in S34 uses principal component analysis, which includes:
[0117] S341, Covariance Matrix Calculation: Based on the normalized dynamic performance data set, a covariance matrix is established to describe the linear relationship between the various performance data. The principal component eigenvalues and principal component eigenvectors are calculated for the covariance matrix, and the principal components with the first few cumulative contributions reaching 95% are selected.
[0118] The eigenvalues are λ1,λ2,…,λ i and the corresponding eigenvectors e1,e2,…,e m , the eigenvalue represents the contribution of each principal component to the total variance of the data, and the contribution rate is calculated: a single principal component λ i The contribution rate is calculated as:
[0119] in It is the sum of all eigenvalues (i.e. the total variance of the data), and the contribution rate indicates the proportion of the principal component in explaining the variance of the data;
[0120] Calculate the cumulative contribution rate: The cumulative contribution rate is the sum of the contribution rates of the first Q principal components, which can be expressed as follows: Select the first few principal components with a cumulative contribution rate of 95% or higher as the number of retained principal components for dimensionality reduction;
[0121] S342, principal component dimensionality reduction: Use the selected principal components to project the data into a low-dimensional space to obtain a low-dimensional representation;
[0122] S343, Reconstructing Data and Error Calculation: Reconstructing the original data using the low-dimensional representation generated by the principal components, and calculating the reconstruction error between the actual data and the reconstructed data. The larger the error, the higher the probability of anomaly.
[0123] S344, outlier determination: Based on the calculated reconstruction error, define the Euclidean norm of the reconstruction error as an anomaly score, and determine the outlier based on the anomaly score;
[0124] S345, latent fault feature extraction: Extract relevant dynamic performance parameters (such as voltage fluctuations at specific time points, power consumption curve deviations, abnormal heat dissipation distribution, etc.) from data marked as abnormal points to form a latent fault feature set to describe chip problems. The abnormal point here refers to a single sampling point (data sample) of the chip performance data, and identify data samples that deviate from the normal performance distribution.
[0125] Determining anomalies based on anomaly scores in S344 specifically includes: using the 99th percentile of the reconstruction error distribution as a determination point, comparing the anomaly score with the 99th percentile of the reconstruction error distribution, and determining if an outlier exceeds the determination point; determining an outlier based on the 99th percentile of the reconstruction error distribution means: sorting the reconstruction errors of all data by size, finding the maximum error value of the data in the top 99% of the sorted data (i.e., the vast majority of normal data), and setting this value as the anomaly determination threshold. If the reconstruction error (i.e., anomaly score) of a piece of data exceeds this determination threshold, it is considered an outlier because it belongs to the very small number (1%) of data in the distribution that deviates from the normal range. In simple terms, the 99th percentile is the boundary between "normal" and "abnormal", and data exceeding this boundary are determined to be anomalies;
[0126] Data reconstruction error E recon The calculation is expressed as: E recon =D-(D PCA ·[e1,e2,...e m ,] TZ ), where D PCA is a low-dimensional representation, e1,e2,…,e m is the selected principal component eigenvector, e m is the eigenvector corresponding to the □th principal component in principal component analysis, D is the normalized performance dataset, and TZ represents the transpose operation.
[0127] The fault module analysis in S4 specifically includes:
[0128] Extract abnormal performance characteristics from the multidimensional performance characteristic model U, including performance fluctuation characteristics, deviation from critical state thresholds, and latent fault characteristics. Map the abnormal performance characteristics to specific functional modules based on the characteristic attribution of the multidimensional performance characteristic model U. Each abnormal performance characteristic corresponds to one or more abnormal functional modules.
[0129] Use the path characteristic graph R to trace the critical path including the abnormal functional module, confirm the path of the affected functional module, and mark the modules in the path with increased delay, abnormal power consumption, or disturbed signal transmission.
[0130] Fault module screening: Comprehensive abnormal performance characteristics and path characteristics to screen faulty functional modules:
[0131] Functional modules that multiple abnormal performance characteristics focus on;
[0132] The starting point of anomaly propagation or the node affected by the accumulation on the path.
[0133] Compare the suspected faulty modules identified with the time and scenario of the performance anomaly to confirm the final faulty module. This abnormal functional module, unlike the anomaly points obtained through the previous principal component analysis, focuses on the chip's functional modules, such as the power module, processing module, and communication module, to determine the specific functional module causing the performance anomaly and its location within the path.
[0134] Feature attribution refers to clarifying the relationship between each performance parameter (voltage fluctuation, power consumption, heat dissipation, signal delay) and the chip functional module in the multi-dimensional performance feature model U. For example:
[0135] Voltage fluctuations are mainly related to the power module or signal driver module;
[0136] Abnormal power consumption is attributed to high-load computing modules;
[0137] Abnormal signal delay is attributed to the data transmission path or interface module.
[0138] When an abnormality in performance characteristics is detected, these attribution relationships are used to quickly locate possible faulty functional modules.
[0139] Feature attribution relationships can be established in the following ways:
[0140] Performance parameter characteristic analysis of functional modules:
[0141] Analyze the performance characteristics of each functional module in the chip design:
[0142] Power module: closely related to voltage fluctuation and power consumption;
[0143] Data processing module: related to signal delay and computing load;
[0144] Thermal management module: related to abnormal heat dissipation.
[0145] Path characteristic graph correlation: Use the path characteristic graph R to determine the signal and power distribution paths between modules.
[0146] Mapping abnormal features to functional modules is as follows:
[0147] Based on feature attribution mapping:
[0148] When a performance characteristic (such as voltage fluctuation or power consumption) is abnormal, the attribution relationship is directly searched and mapped to the corresponding functional module. For example:
[0149] Abnormal power consumption characteristics may be mapped to the power module or execution module;
[0150] Signal delay anomalies may be mapped to the communication module or interface module.
[0151] The mapping results are further verified using the path characteristic graph R:
[0152] If a module has multiple abnormal characteristics (such as power consumption and signal delay) at the same time, it is likely to be a faulty module.
[0153] If an abnormal feature propagates along a path, it can be traced back to the root module.
[0154] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.
[0155] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A chip performance evaluation and fault diagnosis optimization test method, characterized in that: The following steps are involved: S1, intelligent test vector library generation: Generates a test vector library with optimal coverage based on the functional module distribution and circuit structure characteristics of the chip design; S2, dynamic performance acquisition: In the chip operating environment, test conditions are sequentially input using the test vector library, while dynamic performance parameters are collected in real time. The dynamic performance parameters include voltage fluctuation, power consumption curve, heat dissipation, and signal delay, and a dynamic performance data set is constructed based on the dynamic performance parameters; S3, Multidimensional Performance Model Construction: Based on the dynamic performance dataset, a multidimensional performance characteristic model is established, including performance fluctuation characteristics, critical state thresholds, and latent fault characteristics. The multidimensional performance characteristic model is optimized using a nonlinear fitting algorithm to characterize the performance characteristics of the chip in multiple scenarios. S4, fault module analysis: Based on the multi-dimensional performance characteristic model combined with the circuit structure characteristics, reverse trace the correlation between the chip functional modules and locate the specific faulty functional module.
2. The chip performance evaluation and fault diagnosis optimization test method according to claim 1, characterized in that: Said S1 specifically includes: S11, Chip Functional Module Analysis: Extract the functional module layout of the chip design, identify the logical connections and operation paths of each functional module, and mark the key nodes in the functional module, including input ports, output ports and data transmission paths; S12, circuit structure feature extraction: Combine the chip's circuit design files and layout information to obtain the circuit's topological characteristics, signal transmission delay characteristics, power consumption hotspots, and frequency distribution. Build a path characteristic graph R between functional modules and identify the critical paths that affect chip performance. S13, multi-scenario operation simulation: Based on the functional modules and circuit structure characteristics, simulate multiple test scenarios of the chip, including normal load mode, high load mode, low power mode and boundary conditions, and generate the initial test scenario set: S = {S1, S2, ..., S i }, each scene S i Including operation mode and input parameter configuration; S14, test vector library construction: using the key path in the path characteristic graph, generate an initial test vector set for each test scenario. The test vector set is expressed as: V = {v1, v2, ..., v i } represents that, where each test vector v i Including specific signal input combination: input voltage V in 、Input current I in , clock frequency f, data input signal S data , integrate the test vectors of all test scenarios to form a complete test vector library V 库 .
3. The chip performance evaluation and fault diagnosis optimization test method according to claim 2, characterized in that: The S2 specifically includes: S21, test vector input and test condition loading: select test vector V from the test vector library i , where V i ={V in ,I in ,f,S data }, load test condition C i , the test conditions are defined by the operating scenario, including external environmental parameters and load power L load , the test vector V i and test condition C i Synchronous input chip test environment; S22, deploy a multi-dimensional sensor network to collect dynamic performance parameters in real time, including voltage fluctuation ΔV(t), power consumption curve P(t), heat dissipation T(x,y,t), and signal delay τ(t); S23, dynamic performance data set construction: The collected multi-dimensional dynamic performance parameters are integrated into a data set D(t) = {ΔV(t), P(t), T(x, y, t), τ(t)} in time series, where the parameters ΔV, P, T, and τ correspond to voltage fluctuation, power consumption, heat dissipation, and signal delay, respectively; S24, data annotation: according to the input test condition C i Add labels to dynamic performance datasets to indicate the corresponding test vectors and running scenarios.
4. The chip performance evaluation and fault diagnosis optimization test method according to claim 3, characterized in that: In said S22: The voltage fluctuation ΔV(t) is calculated by collecting the voltage changes of each functional module of the chip through the voltage sensor, where t represents time; The power consumption curve P(t) uses a power monitor to record the overall power consumption P(t) of the chip and the power consumption P(t) of the local module in real time. i (t), t represents time; The heat dissipation T(x,y,t) is recorded by a thermocouple array at the time t at each coordinate point (x,y) on the chip layout to form a spatiotemporal heat dissipation map; Signal delay τ(t) is measured by timing analysis tools to measure the delay in data transmission.
5. The chip performance evaluation and fault diagnosis optimization test method according to claim 2, characterized in that: The S3 multi-dimensional performance model construction specifically includes: S31, normalize the dynamic performance data set D(t) and map each performance parameter to a uniform range [0,1]; S32, performance fluctuation feature extraction: Calculate the fluctuation characteristics of dynamic performance parameters, including mean μ, variance σ 2 and instantaneous rate of change r(t), the extracted fluctuation characteristics μ, σ 2 , r(t) is integrated into the eigenvector F 波动 ; S33, critical state threshold calculation: Based on the historical data of dynamic performance parameters, define the critical state threshold T 临界 , which indicates the safe range of the chip in normal operation, including the voltage fluctuation threshold T ΔV , power consumption curve threshold T P , heat dissipation threshold T T And the signal delay threshold T τ ; S34, Hidden fault feature extraction: Use anomaly detection algorithm to analyze data that deviates from the critical state threshold and extract hidden fault features F 隐性 ; S35, the performance fluctuation feature F 波动 , critical state threshold T 临界 and hidden fault characteristics F 隐性 , integrated into a multi-dimensional performance characteristic model U, expressed as: U = {F 波动 ,T 临界 ,F 隐性 }.
6. The chip performance evaluation and fault diagnosis optimization test method according to claim 5, characterized in that: The S3 also includes nonlinear fitting optimization: taking the dynamic performance data set D(t) as input, nonlinear fitting optimization is performed on the multidimensional performance characteristic model U, with the goal of minimizing the error between the model output and the actual performance parameters.
7. The chip performance evaluation and fault diagnosis optimization test method according to claim 5, characterized in that: The anomaly detection algorithm in S34 adopts principal component analysis, which specifically includes: S341, Covariance Matrix Calculation: Based on the normalized dynamic performance data set, a covariance matrix is established to describe the linear relationship between the various performance data. The principal component eigenvalues and principal component eigenvectors are calculated for the covariance matrix, and the principal components with the first few cumulative contributions reaching 95% are selected. S342, principal component dimensionality reduction: Use the selected principal components to project the data into a low-dimensional space to obtain a low-dimensional representation; S343, Reconstructing Data and Error Calculation: Reconstructing the original data using the low-dimensional representation generated by the principal components, and calculating the reconstruction error between the actual data and the reconstructed data. The larger the error, the higher the probability of anomaly. S344, outlier determination: Based on the calculated reconstruction error, define the Euclidean norm of the reconstruction error as an anomaly score, and determine the outlier based on the anomaly score; S345, Hidden Fault Feature Extraction: Extract relevant dynamic performance parameters from the data marked as abnormal points to form a hidden fault feature set to describe the chip problem.
8. The chip performance evaluation and fault diagnosis optimization test method according to claim 7, characterized in that: Determining anomalies based on anomaly scores in S344 specifically includes: using the 99% percentile of the reconstruction error distribution as a determination point, comparing the anomaly score with the 99% percentile of the reconstruction error distribution, and determining if anomalies exceed the determination point.
9. The chip performance evaluation and fault diagnosis optimization test method according to claim 8, characterized in that: The fault module analysis in S4 specifically includes: Extract abnormal performance characteristics from the multidimensional performance characteristic model U, including performance fluctuation characteristics, deviation from critical state thresholds, and latent fault characteristics. Map the abnormal performance characteristics to specific functional modules based on the characteristic attribution of the multidimensional performance characteristic model U. Each abnormal performance characteristic corresponds to one or more abnormal functional modules. Use the path characteristic graph R to trace the critical path including the abnormal functional module and confirm the path of the affected functional module; Fault module screening: Comprehensive abnormal performance characteristics and path characteristics to screen faulty functional modules: Functional modules that multiple abnormal performance characteristics focus on; The starting point of anomaly propagation or the node affected by the accumulation on the path.
Citation Information
Cited By
Chip efficient detection system and method
CN121254044A
Chip efficient detection system and method
CN121254044B
Automatic alarm method and system for integrated circuit test
CN121348044A