Fault model construction method for adaptive testing of system-level chip outliers

Through multi-dimensional test data acquisition and deep learning algorithms, combined with clustering and machine learning, an adaptive fault model of system-level chips is built, solving the problems of outlier testing and diagnosis in complex chips, and achieving high accuracy and fast response fault diagnosis capabilities.

CN119716501BActive Publication Date: 2025-05-13HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510237403.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The prior art is difficult to comprehensively and accurately test and analyze outliers in complex system-on-chips, and lacks adaptive adjustment capabilities, resulting in low accuracy and reliability of fault diagnosis.

Method used

A multi-dimensional test data acquisition method is adopted, combined with statistical analysis and deep learning algorithms, outlier features are screened and extracted, and fault models are constructed through clustering and machine learning algorithms. The system has adaptive adjustment function and can dynamically adjust the model according to the real-time working conditions of the chip.

Benefits of technology

It realizes accurate capture of system-level chip outliers and effective modeling of fault modes, improves the accuracy and reliability of fault diagnosis, reduces the risk of misjudgment and misjudgment, and saves time and resources for testing and model reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119716501B_ABST
    Figure CN119716501B_ABST
Patent Text Reader

Abstract

The present invention discloses a fault model construction method for adaptive testing of abnormal values ​​of system-level chips; it relates to the technical field of fault construction of chip testing, tests chips on a large scale under different conditions such as temperature, voltage, frequency and load, collects performance parameters such as signal delay and power consumption, uses Z-score or box plot method to screen abnormal values, and processes with sliding windows; extracts features and normalizes through CNN or RNN; classifies fault features by K-Means or hierarchical clustering, with an accuracy of ≥80%; constructs a model by SVM or decision tree, can predict faults, self-update, adaptively adjust according to chip parameter changes, uses independent data to verify optimization, and evaluates by leave-one-out method and gradient descent optimization. The present invention can perform multi-condition testing on system-level chips, accurately screen abnormal values, extract features and classify faults, construct a self-updatable fault model, realize adaptive adjustment, verify optimization to improve accuracy, and improve chip testing and fault diagnosis efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault model construction systems, and in particular to a fault model construction method for system-level chip abnormal value adaptive testing. Background Art

[0002] With the rapid development of semiconductor technology, system-on-chip (SoC) has been widely used in various electronic devices, from consumer electronics such as smartphones and tablets to high-end fields such as industrial control, automotive electronics, and aerospace. System-on-chip integrates multiple functional modules, and its complexity and integration are increasing, which brings great challenges to chip testing and fault diagnosis. Traditional chip testing methods mainly focus on functional testing and performance testing, and verify the chip through pre-set test vectors and test programs to ensure that the chip meets basic functional and performance indicators. However, these methods gradually expose their limitations when dealing with complex system-on-chip. On the one hand, in a complex working environment and under various working conditions, the chip may exhibit various abnormal behaviors. Traditional testing methods are difficult to cover all possible abnormal situations, especially those that only occur under a specific combination of working conditions. On the other hand, the testing and analysis of outliers often rely on manual experience and post-diagnosis, and lack systematic and automated testing and analysis methods.

[0003] In actual applications, system-level chips may generate abnormal values ​​due to changes in factors such as temperature, voltage, frequency, and load. These abnormal values ​​may be caused by a variety of fault sources, such as physical defects, manufacturing process deviations, circuit aging, or electromagnetic interference. Due to the highly complex internal structure of system-level chips, these abnormal values ​​may be difficult to accurately capture and interpret, resulting in difficulties in fault diagnosis. Traditional fault diagnosis methods are usually based on empirical rules and simple threshold judgments, which are difficult to adapt to the rapid development of chip design and manufacturing processes.

[0004] Existing fault models are usually built based on specific chip types and fixed working conditions, and lack adaptability to different chip architectures and working conditions. Moreover, the screening and analysis of outliers are not accurate and detailed enough, and it is difficult to establish an effective connection between outliers and specific failure modes. In addition, as chip manufacturing technology enters the nanoscale or even smaller process, traditional fault analysis technology cannot meet the increasing reliability and stability requirements, which may result in the inability to quickly locate and solve problems when chip products fail in the market, causing huge economic losses and reputation risks to enterprises.

[0005] At the same time, the current chip testing and fault diagnosis tools are mostly static in their handling of abnormal values ​​and cannot be adaptively adjusted according to the real-time working conditions of the chip. When the chip working conditions change, a large amount of testing and model building work needs to be re-performed, which consumes a lot of time and resources. In addition, in the process of fault model construction, there is a lack of effective data processing and feature extraction technology, making it difficult to extract valuable information from a large amount of test data, and it is also difficult to accurately classify and model different types of faults, resulting in low accuracy and reliability of fault prediction and diagnosis.

[0006] In order to solve these problems, there is an urgent need for a fault model construction method and system that can adapt to the complex working environment of the system-level chip, can comprehensively and accurately test and analyze abnormal values, and can adaptively adjust according to the real-time working conditions of the chip, so as to improve the test efficiency of the system-level chip, the accuracy and reliability of fault diagnosis, and ensure the normal operation of the chip in various complex scenarios. Summary of the invention

[0007] The present invention proposes a fault model construction method for system-level chip abnormal value adaptive testing to solve the problems mentioned in the above-mentioned prior art.

[0008] In order to achieve the above object, the present invention adopts the following technical solution: a fault model construction method for system-level chip outlier adaptive testing, comprising:

[0009] Test data collection steps: Test the system-level chip under different working conditions, including different temperatures, voltages, frequencies and loads, and use test instruments to collect the performance parameters of the chip. During the test, the performance parameters are collected through data channel multiplexing technology, and the analog-to-digital converter ADC converts the analog signal into a digital signal, including data under normal and abnormal working conditions of the chip;

[0010] Outlier screening steps: Use statistical analysis methods to analyze the collected test data and screen out outliers that are beyond the normal range; for the Z-score method, set the Z value threshold to ±3; for the box plot method, determine the upper and lower limits as outlier screening thresholds based on the quartiles Q1 and Q3 and 1.5 times the interquartile range IQR. For different performance parameters, determine them separately according to the statistical distribution characteristics, and the proportion of screened outliers in the total data volume is not less than 5%;

[0011] Fault feature extraction steps: The fault features are extracted from the outlier data through the learning algorithm convolutional neural network CNN or recurrent neural network RNN. For the CNN algorithm, convolutional layers and pooling layers are used, the convolution kernel size is set to 3x3 or 5x5, and the activation function uses the ReLU function. Features at different levels are extracted by stacking convolutional layers. For the RNN algorithm, LSTM or GRU units are used, and the number of hidden layer units is determined according to the time series length of the data.

[0012] Fault mode classification step: clustering algorithms, including K-Means clustering or hierarchical clustering, are used to classify the extracted fault features; for K-Means clustering, the optimal number of clusters K is determined by the elbow rule, and for hierarchical clustering, the full connection method is used to calculate the inter-cluster distance to generate a more discriminative clustering result;

[0013] Fault model construction steps: Based on the classified fault modes, a fault model is constructed through machine learning algorithms, including support vector machine (SVM) or decision tree algorithm. For the SVM algorithm, a radial basis kernel function is selected and hyperparameters are optimized through grid search and cross-validation techniques. For the decision tree algorithm, information gain or Gini index is used as the splitting criterion, and the depth of the tree is limited to prevent overfitting. Different fault modes are modeled, and the constructed fault model predicts and identifies chip faults that occur under different working conditions.

[0014] Adaptive adjustment steps: During the operation of the system-level chip, the chip performance parameters are continuously monitored. If the chip working conditions change, the fault model is adaptively adjusted according to the new working conditions. Through the real-time feedback mechanism, the real-time performance parameters of the chip are fed back to the fault model, and the performance parameter change rate is calculated and compared with the preset threshold. If the change rate exceeds the threshold, the model adjustment is triggered;

[0015] Model verification and optimization steps: Use an independent test data set to verify the constructed fault model, and evaluate the model performance by calculating the confusion matrix, accuracy, and recall rate indicators; use the leave-one-out method or K-fold cross-validation method to evaluate the model. For the confusion matrix, calculate the ratio of diagonal elements to off-diagonal elements to accurately evaluate the classification performance of the model; optimize the fault model according to the verification results, and adjust the model parameters through the gradient descent algorithm.

[0016] Further, the method further comprises the following steps:

[0017] Outlier enhancement step: perform data enhancement operations on the filtered outlier data, including adding noise, rotation, and scaling, to expand the amount of outlier data; for the noise addition operation, the noise type includes Gaussian noise, and for the rotation and scaling operations, the rotation angle is within the range of ±15°.

[0018] Cross-platform verification steps: The constructed fault model is applied to different types of system-level chip platforms for verification, including but not limited to chip platforms of ARM architecture, x86 architecture and RISC-V architecture, and the verification data on the platform is collected, and the fault model is adjusted according to the verification results.

[0019] Furthermore, in the outlier screening step, timestamps and location information are added to the screened outlier data to trace the time when the outlier occurred and the specific location inside the chip. Through the address mapping technology inside the chip, the outlier is mapped to the specific circuit module and storage unit of the chip.

[0020] Furthermore, in the fault feature extraction step, the extracted fault features are weighted in combination with the hardware architecture information of the chip, including the circuit topology and logic unit layout; weights are assigned according to the paths and logic units of the hardware architecture, and the weight coefficient ranges from 1.2 to 2.0.

[0021] Furthermore, in the fault mode classification step, an ensemble learning method is used to integrate the results of the classifiers, including the fusion of the results of different classifiers using the voting method or the weighted average method; for the voting method, when a data point of the classifier is judged to be a certain fault mode, it is finally judged to be the fault mode; for the weighted average method, different weights are assigned according to the performance of different classifiers.

[0022] Furthermore, in the fault model construction step, risk levels are set for different fault modes, and divided into high, medium and low risk levels according to the impact of the fault on chip performance. The risk level is determined according to the number of performance parameters affected by the fault, the frequency of the fault, and the factors affecting the function, providing a priority reference for subsequent fault handling.

[0023] A system using the fault model construction method for system-level chip abnormal value adaptive testing comprises:

[0024] Test data acquisition unit: equipped with various test instruments, flexible setting of test conditions, including temperature range, voltage range, frequency range, to meet the test requirements of different system-level chips, collect chip performance parameters, and use solid-state drives (SSDs) to store data;

[0025] Data processing unit: including outlier screening module, fault feature extraction module, fault mode classification module, fault model construction module, with computing capabilities, using processors and GPUs to accelerate computing;

[0026] Adaptive adjustment module: It makes real-time adjustments to the fault model according to changes in chip working conditions. It has a data communication interface and uses a PCIe4.0 interface to interact with the test data acquisition unit and the data processing unit in real time.

[0027] Furthermore, it also includes:

[0028] Model verification unit: Verify and optimize the fault model through independent test data sets, provide performance evaluation reports, including accuracy, recall, and F1 value indicators, and use automated test scripts to periodically verify the model.

[0029] Furthermore, it also includes:

[0030] Visual display module: The construction process of the fault model, the verification results and the test data of the chip are displayed on the display terminal in the form of charts, curves and three-dimensional models. It supports user interaction and users can view different levels of information through mouse and keyboard operations; including fault mode distribution and feature extraction results.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] First, in terms of the accuracy of fault detection and diagnosis, through multi-dimensional test data collection and sophisticated outlier screening, feature extraction, and advanced classification and modeling algorithms, it is possible to accurately capture the outliers of the chip under different working conditions, establish a close connection between the outliers and the fault mode, and make the prediction accuracy of the fault model no less than 75%, which significantly improves the accuracy of fault diagnosis and reduces the risk of misjudgment and missed judgment.

[0033] Secondly, the adaptive adjustment function of this patent is very outstanding. When the working conditions of the chip change, the fault model can be adjusted in real time within 1 second, so that it always adapts to the real-time working environment of the chip, without the need for large-scale re-testing and model reconstruction, saving a lot of time and resources, and ensuring the continuous monitoring and fault diagnosis capabilities of the chip in a dynamic environment.

[0034] Furthermore, through technologies such as data enhancement, cross-platform verification and feature weighting, the versatility and generalization ability of the fault model are improved, and it can be applied to a variety of chip architectures and working conditions. It not only improves the applicability to different chip types, but also avoids the problem of model overfitting, ensuring the reliability and effectiveness of the fault model on different platforms.

[0035] In addition, the system provides detailed visualization and performance evaluation reports, which facilitate users to intuitively understand the process and results of chip testing and fault diagnosis, providing convenience for users. At the same time, the risk level of different fault modes can enable maintenance personnel to handle faults according to their priority, which helps to ensure the key functions of the chip, improve the overall reliability and stability of the system-level chip, and provide a more efficient and scientific solution for fault management during chip development, production and use.

[0036] Finally, through the optimized verification and optimization process, the performance of the model is continuously improved, providing a strong guarantee for the long-term stable operation of the chip, meeting the semiconductor industry's needs for chip reliability and high performance, and promoting the further development of chip technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic block diagram of a fault model building system for adaptive testing of system-level chip outliers proposed by the present invention;

[0038] Figure 2 The present invention is a schematic block diagram of a fault model construction method for adaptive testing of system-level chip outliers proposed by the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0041] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, and it can be the internal connection of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0042] Reference Figure 1-2 :A fault model construction method for system-level chip abnormal value adaptive testing includes the following steps:

[0043] S1. Test data collection steps: First, place the system-level chip to be tested in a specially designed test environment equipped with various high-precision test instruments, such as temperature controllers, voltage sources, signal generators, power meters, etc. These instruments can accurately adjust and stably provide different working conditions, including temperature ranges from -40°C to 125°C (in 5°C intervals), voltages from 0.8V to 1.2V (in 0.1V intervals), frequencies from 100MHz to 1GHz (in 50MHz intervals), and loads from 10% to 100% (in 10% intervals). During the test, multiple test instruments are connected to different test points of the chip using multiplexing technology, and various performance parameters of the chip, such as signal delay, power consumption, current, etc., are collected at the same time. The analog signal is converted into a digital signal through a high-precision analog-to-digital converter (ADC), ensuring that the data accuracy reaches more than 16 bits and the sampling frequency is not less than 1000Hz. In order to ensure that the collected data covers the data under normal and abnormal working conditions of the chip, a large number of tests will be carried out, and the amount of collected data is not less than 100GB. During the test, the collected data is stored in a large-capacity storage device to form a comprehensive test data set.

[0044] S2. Outlier screening step: Read the test data set from the storage device and use statistical analysis methods to screen outliers. For the Z-score method, first calculate the mean and standard deviation of each performance parameter, and then use the formula Calculate the Z value, where x is the data point, is the mean, is the standard deviation, and the data points with Z values ​​exceeding ±3 are marked as outliers; for the box plot method, the quartiles (Q1 and Q3) are calculated first, and then the interquartile range (IQR=Q3-Q1), and the data points less than Q1-1.5*IQR or greater than Q3+1.5*IQR are regarded as outliers. At the same time, the corresponding screening thresholds are determined for different performance parameters according to their statistical distribution characteristics. In the screening process, the sliding window technology is used to process the data in segments, and the window size is set according to the time series characteristics of the data. For example, for periodic data, the window size can be set to an integer multiple of the period to reduce the impact of data fluctuations on outlier screening. The outliers finally screened out should account for no less than 5% of the total data volume, and these outliers are stored in a special outlier data set to provide a basis for the subsequent fault model construction.

[0045] S3. Fault feature extraction step: Input the screened outlier data set into the deep learning algorithm to extract fault features. Taking the convolutional neural network (CNN) as an example, a network structure containing multiple convolutional layers and pooling layers is constructed, where the convolution kernel size is 3x3 or 5x5, and the activation function uses the ReLU function. In the first convolution layer, the input data is convolved with the convolution kernel, and the convolution result is nonlinearly activated by the ReLU function, and then downsampled through the pooling layer. This process is repeated, and multiple convolutional layers are stacked to extract features at different levels. For the recurrent neural network (RNN), LSTM or GRU units are used, and the number of hidden layer units is determined according to the time series length of the data to effectively capture the time series information. The extracted fault features include the time series pattern of the outliers, the spatial distribution characteristics, and the difference characteristics from the normal data. In order to make the features comparable, the feature normalization method is used to map the feature values ​​to the [0,1] interval. Use a large amount of data for training to ensure that the extracted fault features can effectively characterize the potential fault mode of the chip, and the accuracy of feature extraction is not less than 85%.

[0046] S4. Fault mode classification step: For the extracted fault features, a clustering algorithm is used for classification. When using K-Means clustering, the elbow rule is first used, that is, the sum of squared errors (SSE) under different K values ​​is calculated, and the optimal number of clusters K is determined according to the inflection point of the SSE curve. Then the K-Means++ algorithm is used to initialize the cluster center, and the data points are assigned to the nearest cluster center. The cluster center is continuously updated iteratively until the cluster center no longer changes or the predetermined number of iterations is reached. For hierarchical clustering, the full connection method is used to calculate the inter-cluster distance, that is, the maximum distance between all point pairs in the two clusters is calculated, and the clusters with the closest distance are gradually merged, and finally the fault features are divided into multiple fault modes. Different fault modes will have clear feature descriptions, and the accuracy of the classification results will not be less than 80%, providing a basis for subsequent fault diagnosis and repair.

[0047] S5. Fault model construction steps: Based on the classified fault mode, a fault model is constructed using a machine learning algorithm. For the support vector machine (SVM), a radial basis kernel function is selected, and hyperparameters such as the penalty parameter C and the kernel function parameter gamma are optimized through grid search and cross-validation techniques. In the grid search, different value ranges of C and gamma are set, and the data set is divided into multiple subsets, one part as a training set and the other part as a validation set. The hyperparameters are continuously adjusted to optimize the performance on the validation set. For the decision tree algorithm, information gain or Gini index is used as the splitting criterion. Starting from the root node, the attribute that maximizes the information gain or Gini index is selected as the splitting attribute, and the maximum depth of the tree is limited to prevent overfitting. The constructed fault model should be able to accurately predict and identify possible faults of the chip under different working conditions, with a prediction accuracy of not less than 75%. In addition, the fault model has the ability to self-update and optimize. When there is new test data in the future, the model parameters are updated in real time using an online learning algorithm (such as stochastic gradient descent) to improve the adaptability of the model.

[0048] S6. Adaptive adjustment step: During the normal operation of the system-level chip, the chip performance parameters are continuously monitored through special monitoring equipment. When the chip working conditions change, the real-time performance parameters of the chip are fed back to the fault model using a real-time feedback mechanism. The performance parameter change rate is calculated and compared with the preset threshold. When the change rate exceeds the threshold, the model adjustment is triggered. For example, for the performance parameter P, the change rate can be calculated as ,If the rate of change exceeds the threshold, the system will input the ,newly collected data into the fault model, and use online learning methods ,such as the stochastic gradient descent algorithm to update the model ,parameters, and adjust the response time to no more than 1 second, ensuring that the fault model can ,adapt to the real-time working environment of the chip.

[0049] S7. Model validation and optimization steps: Use an independent test data set to validate the constructed fault model. Use the leave-one-out method or K-fold cross-validation method. For example, for K-fold cross-validation, randomly divide the data set into K subsets, select one of the subsets as the validation set each time, and the rest as the training set, and repeat K times. In each validation, calculate the confusion matrix, and the elements of the confusion matrix are It represents the number of samples whose actual category is i but predicted to be category j. The classification performance of the model is evaluated by calculating the ratio of elements on the diagonal to elements on the off-diagonal. At the same time, the accuracy rate, recall rate and other indicators are calculated. According to the verification results, the model parameters are adjusted using the gradient descent algorithm. The accuracy of the optimized model on the verification data set is improved by at least 5%.

[0050] The present invention also discloses a fault model building system for adaptive testing of abnormal values ​​of system-level chips, including:

[0051] 1. Test data acquisition unit: The data acquisition unit is equipped with a high-precision temperature controller that can accurately adjust the temperature range (-50℃ to 150℃), a high-precision voltage source that provides a voltage range (0.8V to 1.5V), and a signal generator that can adjust the frequency range (100MHz to 5GHz), etc., to meet the testing needs of different system-level chips. These instruments are connected to the test pins of the chip and transmit the collected chip performance parameters to the storage device through a high-speed data acquisition card. The storage device uses a high-speed solid-state drive (SSD) with a storage speed of not less than 500MB / s and a storage capacity of not less than 1TB to store a large amount of test data.

[0052] 2. Data processing unit: The data processing unit includes an outlier screening module, a fault feature extraction module, a fault mode classification module and a fault model building module. The unit is equipped with a multi-core processor and NVIDIA's high-end GPU computing card. The number of GPU computing cores is not less than 1,000, and large-scale test data is processed in parallel. The processing speed is more than 10 times faster than that of traditional CPU single-core processing. The outlier screening module performs the above-mentioned outlier screening steps and passes the screened outliers to the fault feature extraction module. The fault feature extraction module uses a deep learning algorithm to extract fault features and passes the extracted features to the fault mode classification module for classification. The fault mode classification module uses clustering algorithms and ensemble learning methods to classify and integrate fault features, and finally passes the results to the fault model building module, and uses machine learning algorithms to build a fault model.

[0053] 3. Adaptive adjustment module: It is connected to the test data acquisition unit and the data processing unit through the PCIe4.0 high-speed data communication interface to receive the chip performance parameters in real time. When the chip working conditions change, the information is passed to the fault model construction module through the real-time feedback mechanism according to the performance parameter change rate, and the fault model is adjusted in real time using the online learning algorithm, with the adjustment delay not exceeding 1 second.

[0054] 4. Model verification unit: Use independent test data sets to verify and optimize the fault model. It uses automated test scripts to verify the model regularly (no more than 24 hours), calculate accuracy, recall, F1 value and other indicators, and generate detailed performance evaluation reports to provide a basis for model optimization.

[0055] 5. Visualization display module: The fault model construction process, verification results and chip test data are displayed on the display terminal in the form of charts, curves, 3D models, etc. WebGL technology is used to achieve 3D visualization. Users can view different levels of detailed information, such as fault mode distribution, feature extraction results, etc., through mouse and keyboard operations. Users can intuitively see the performance of the fault model and various information of chip testing, which is convenient for analysis and decision-making.

[0056] In the present invention, the implementation of cross-platform verification applies the constructed fault model to different types of system-level chip platforms, such as chip platforms of ARM architecture, x86 architecture and RISC-V architecture. The above test and model building process is applied to chips of different platforms, and verification data on different platforms are collected. The fault model is adjusted according to the verification results to make it applicable to a wider range of chip types, ensuring that the fault model has good versatility and portability.

[0057] In the present invention, the implementation of the outlier enhancement step performs data enhancement operations on the outlier data that have been screened out. When adding Gaussian noise, the noise standard deviation is determined based on 1% to 5% of the data range, and the noise is added to the outlier data; for rotation and scaling operations, the rotation angle is within the range of ±15°, and the scaling ratio is between 0.8 and 1.2. The amount of data after enhancement is at least twice the amount of the original outlier data to expand the outlier data set and avoid the problem of model overfitting caused by insufficient outlier data.

[0058] In the present invention, the weighted processing of fault features is implemented in the fault feature extraction step, and the extracted fault features are weighted in combination with the hardware architecture information of the chip, such as circuit topology, logic unit layout, etc. A higher weight is given according to the critical path and important logic unit of the hardware architecture, and the weight coefficient ranges from 1.2 to 2.0, thereby improving the characterization ability of the fault feature to the fault mode.

[0059] In the present invention, the implementation of ensemble learning is in the fault mode classification step, and the results of multiple classifiers are integrated by using ensemble learning methods. When the voting method is used, the classification results of multiple classifiers for each data point are counted. When most classifiers determine a data point as a certain fault mode, it is finally determined as the fault mode; for the weighted average method, different weights are assigned according to the performance of different classifiers, and the results of different classifiers are weighted averaged, so that the classification accuracy after integration is improved by more than 15% compared with a single classifier.

[0060] In the present invention, the implementation of fault risk level division is in the fault model construction step, and risk levels are set for different fault modes. According to the degree of impact of the fault on chip performance, such as the number of performance parameters affected by the fault, the frequency of fault occurrence, and the impact on key functions, the fault is divided into high, medium, and low risk levels. In the fault handling process, high-risk faults are handled first to ensure that the key functions of the chip are not affected. Through the above specific implementation method, adaptive testing of system-level chip abnormal values ​​can be achieved, and a reliable, universal, and optimizable fault model can be constructed. At the same time, a complete system is provided for realizing the full process operation from data collection, processing, model construction, verification to visualization, ensuring the accuracy and efficiency of system-level chip fault detection and diagnosis.

[0061] The above are only preferred specific implementation modes of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solutions and inventive concepts of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A fault model construction method for system-level chip abnormal value adaptive testing, characterized in that: The following steps are involved: Test data collection steps: Test the system-level chip under different working conditions. During the test, collect performance parameters through data channel multiplexing technology; Outlier screening steps: Use statistical analysis methods to analyze the collected test data and screen out outliers that are beyond the normal range; Fault feature extraction step: The outlier data is filtered out and the fault features are extracted through the learning algorithm convolutional neural network or recurrent neural network; Fault mode classification steps: clustering algorithm is used to classify the extracted fault features; for K-Means clustering, the optimal number of clusters K is determined by the elbow rule, and for hierarchical clustering, the full connection method is used to calculate the distance between clusters; Fault model construction steps: Based on the classified fault modes, a fault model is constructed through a machine learning algorithm. For the SVM algorithm, a radial basis kernel function is selected and hyperparameters are optimized through grid search and cross-validation techniques. For the decision tree algorithm, information gain or Gini index is used as the splitting criterion to model different fault modes. Adaptive adjustment steps: During the operation of the system-level chip, the real-time performance parameters of the chip are fed back to the fault model through the real-time feedback mechanism. The performance parameter change rate is calculated and the model adjustment is triggered if the change rate exceeds the threshold. Model validation and optimization step: Use an independent test data set to validate the constructed fault model.

2. A fault model construction method for system-level chip abnormal value adaptive testing according to claim 1, characterized in that: Also includes: Outlier enhancement step: perform data enhancement operations on the filtered outlier data, including adding noise, rotating, scaling, and expanding the amount of outlier data; For the add noise operation, the noise type includes Gaussian noise, and for the rotate and scale operations, the rotation angle is in the range of ±15°.

3. A fault model construction method for system-level chip abnormal value adaptive testing according to claim 1, characterized in that: Also includes: Cross-platform verification steps: The constructed fault model is applied to different types of system-level chip platforms for verification, including but not limited to chip platforms of ARM architecture, x86 architecture and RISC-V architecture, and the verification data on the platform is collected, and the fault model is adjusted according to the verification results.

4. The fault model construction method for system-level chip abnormal value adaptive testing according to claim 1 is characterized in that: In the outlier screening step, timestamps and location information are added to the screened outlier data to trace the time when the outlier occurred and the specific location inside the chip. Through the address mapping technology inside the chip, the outlier is mapped to the specific circuit module and storage unit of the chip.

5. The fault model construction method for system-level chip abnormal value adaptive testing according to claim 1 is characterized in that: In the fault feature extraction step, the extracted fault features are weighted by combining the hardware architecture information of the chip, including the circuit topology and logic unit layout; Weights are assigned based on the paths and logic units of the hardware architecture, with the weight coefficient ranging from 1.2 to 2.

0.

6. A fault model construction method for system-level chip abnormal value adaptive testing according to claim 1, characterized in that: In the fault mode classification step, an ensemble learning method is used to integrate the results of multiple classifiers, specifically including the use of voting method or weighted average method; among them, for the voting method, when a data point of the classifier is judged as a certain fault mode, it is finally judged as a fault mode; for the weighted average method, different weights are assigned according to the performance of different classifiers.

7. The fault model construction method for system-level chip abnormal value adaptive testing according to claim 1 is characterized in that: In the fault model construction step, risk levels are set for different fault modes, and they are divided into high, medium, and low risk levels according to the impact of the fault on chip performance. The risk level is determined based on the number of performance parameters affected by the fault, the frequency of the fault, and the factors affecting the function, providing a priority reference for subsequent fault handling.

8. A system for implementing the fault model construction method for system-level chip abnormal value adaptive testing according to any one of claims 1 to 7, characterized in that: include: Test data acquisition unit: equipped with various test instruments, flexible setting of test conditions, including temperature range, voltage range, frequency range, to meet the test requirements of different system-level chips, collect chip performance parameters, and use solid-state drives (SSDs) to store data; Data processing unit: including outlier screening module, fault feature extraction module, fault mode classification module, fault model construction module, with computing capabilities, using processors and GPUs to accelerate computing; Adaptive adjustment module: It makes real-time adjustments to the fault model according to changes in chip working conditions. It has a data communication interface and uses a PCIe4.0 interface to interact with the test data acquisition unit and the data processing unit in real time.

9. The system according to claim 8, characterized in that Also includes: Model verification unit: Verify and optimize the fault model through an independent test data set, provide a performance evaluation report, including accuracy, recall, and F1 value indicators, and use automated test scripts to periodically verify the model; use the leave-one-out method or K-fold cross-validation method to evaluate the model, optimize the fault model based on the verification results, and adjust the model parameters through the gradient descent algorithm.

10. The system according to claim 8, characterized in that Also includes: Visual display module: The construction process of the fault model, the verification results and the test data of the chip are displayed on the display terminal in the form of charts, curves and three-dimensional models. It supports user interaction and users can view different levels of information through mouse and keyboard operations. Including failure mode distribution and feature extraction results.

Citation Information

Patent Citations

  • Integrated circuit chip testing method and device and storage medium

    CN114089153A

  • Integrated circuit fault recording system

    CN117074916A