Storage unit yield analysis method based on sparse representation, electronic equipment and storage medium
Through the memory cell yield analysis method based on sparse representation, the covariance matrix descriptor and the improved Hill estimator, combined with the ANN accelerated classification model, the problem of low efficiency of high-dimensional parameter space and tail-part distribution analysis in the prior art is solved, and the efficient accuracy of the memory circuit yield analysis is achieved.
Patent Information
- Application Number
- CN202510324531.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, when performing yield analysis of storage circuits, especially when facing high-dimensional parameter space and tail distribution, computing resources and time overhead are too large, making it difficult to meet the needs of rapid evaluation.
The memory cell yield analysis method based on sparse representation is adopted, and local features are extracted through covariance matrix descriptors, and improved Hill estimator and generalized Pareto distribution fitting are used to construct it in combination with ANN accelerated classification model to improve analysis accuracy and robustness.
The efficiency and accuracy of memory circuit yield analysis are significantly improved, especially in scenarios where the tail sample is insufficient and the data distribution is inconsistent, providing efficient and accurate tools, overcoming the disadvantages of traditional methods under the characteristics of high-dimensional data distribution.
Smart Images

Figure CN120197578A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for analyzing the yield of memory cells based on sparse representation, an electronic device, and a storage medium, and belongs to the technical field of integrated circuit design automation. Background Art
[0002] Yield analysis is the cornerstone of ultra-large-scale integrated circuit design and optimization. Advanced yield analysis tools can effectively and quickly guide the design of the chip backend. The reduction of circuit size and the decrease of operating voltage make it more difficult to accurately evaluate the impact of process parameter fluctuations on circuit performance and yield. Especially in SRAM circuits, due to the existence of a large number of repetitive units, the failure rate of each unit must be extremely low to ensure the overall yield, thus increasing the need to evaluate extremely low-probability events. In the context of wide-voltage design, the sensitivity of circuit performance to process variations increases significantly, and this sensitivity is further amplified by the extensive replication and size reduction of memory cells. For example, process variations such as threshold voltage mismatch, oxide thickness fluctuations, and random doping effects can cause the performance distribution to deviate from the Gaussian distribution, showing non-Gaussian characteristics such as tail distributions. These inconsistencies not only increase the complexity of circuit design but also have a significant adverse impact on the reliability and yield of memory circuits. Especially in the operating region close to the threshold, minor electrical parameter deviations may lead to significant performance degradation, thus causing a greater impact on the overall performance and reliability of the circuit.
[0003] Therefore, during the yield analysis of a circuit, these factors must be carefully analyzed to ensure the stability and reliability of the circuit under various operating conditions. Although the traditional Monte Carlo method (MC) is highly reliable in yield analysis, its computational resources and time overhead increase with the increase of parameter dimensions and yield levels. This huge computational cost significantly limits the applicability of the MC method in high-dimensional parameter spaces and is difficult to meet the requirements of rapid evaluation in the circuit design process.
[0004] To further improve the efficiency of memory circuit analysis, researchers have proposed various improvement methods: The importance sampling (IS)-based method migrates the process variation center to near the failure boundary and samples through Monte Carlo method (MC) combined with SPICE (Simulation Program with Integrated Circuit Emphasis) simulation, which improves the sampling efficiency and is applicable to high-dimensional settings with complex failure distributions. However, it requires accurate failure boundary estimation and has numerical instability problems. Assuming that the circuit performance follows a Gaussian distribution, it is difficult to accurately capture the tail behavior of the long-tailed distribution; The classifier-based method proposes to train a classifier to learn the failure boundary, which can focus on exploring the boundary of the fault region and handle complex boundary conditions. When facing the tail distribution, the analysis efficiency is affected by the choice of the tail fraction, and the choice of the tail fraction will have a significant impact on the accuracy and computational cost.
[0005] It can be seen that the existing importance sampling method and classifier-based method have not achieved ideal results when analyzing extremely low-probability events based on the tail distribution. Therefore, it is urgent to develop an efficient method to overcome such problems. Summary of the Invention
[0006] Objective: To overcome the deficiencies in the prior art, the present invention provides a memory cell yield analysis method, an electronic device, and a storage medium based on sparse representation. Aiming at high-dimensional problems, the present invention proposes a sparse representation classification framework to evaluate the probability distribution of the tail in the long-tailed distribution, uses an ANN (Artificial Neural Network) to accelerate the construction of the classification model, solves the circuit yield by fitting the Generalized Pareto Distribution (GPD), and improves the accuracy and robustness of the yield analysis through the Hill estimator and tail resampling means.
[0007] Technical Solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] In a first aspect, a memory cell yield analysis method based on sparse representation specifically includes:
[0009] Simulate the stored circuit netlist, and collect an equal number of samples from the failure region and the non-failure region of the simulation results as the sample data set.
[0010] According to the sample data set, obtain k nearest neighbor samples, and calculate the covariance matrix of the samples based on the sample data set and the corresponding k nearest neighbor samples.
[0011] A method of obtaining the optimal tail fraction using an estimator, taking the optimal tail fraction as a criterion to measure whether the training samples in the sample dataset are tail samples, and classifying the test samples in the sample dataset using the covariance matrix of the tail samples to obtain the scores of the classified samples.
[0012] Simulate the stored circuit netlist corresponding to the optimal tail fraction, re-collect the tail training samples, and calculate the corresponding parameters of the tail distribution according to the tail training samples.
[0013] Substitute the scores of the classified samples and the corresponding parameters of the tail distribution into the storage cell yield calculation formula to obtain the storage cell yield.
[0014] As a preferred solution, samples with comparable acquisition quantities are collected from the failure area and non-failure area of the Latin hypercube sampling simulation results as the sample dataset.
[0015] As a preferred solution, the method for obtaining the k nearest neighbor samples specifically includes:
[0016] According to the hash table Calculate the hash key value of each sample in the sample dataset and map the sample to the corresponding hash bucket according to the hash key value.
[0017] For the query sample perturb the hash key value , where , .
[0018] Calculate the difference between the hash key value of the query sample in the hash bucket corresponding to the perturbed hash key value and the hash key value of the sample and the sample .
[0019] Merge the query samples and whose hash key value differences fall within the interval as the candidate sample set, where represents the sample in the hash table corresponding to the th hash function hash key value.
[0020] Search for the nearest neighbor samples in the candidate sample set to obtain the k nearest neighbor samples.
[0021] As a preferred solution, the covariance matrix expression of the sample is as follows:
[0022]
[0023] Where is the number of nearest neighbor samples, denotes the -th nearest neighbor sample, denotes the mean of neighboring samples, T denotes the transpose of a matrix, denotes the sample covariance matrix.
[0024] As a preferred solution, the method for obtaining the optimal tail fraction using an estimator specifically includes:
[0025] Construct a Hill estimator and calculate the estimated value of the tail index. Among them, the expression of the Hill estimator is as follows:
[0026]
[0027] where, denotes the estimated value of the tail index, k is the highest statistical order of the tail region used for estimation, denotes the -th sample corresponding to the probability density magnitude, denotes the -th sample corresponding to the probability density magnitude.
[0028] Substitute the estimated value of the tail index into the optimal tail fraction optimization model, and solve the optimal tail fraction optimization model to obtain the optimal tail fraction. Among them, the expression of the optimal tail fraction optimization model is as follows:
[0029]
[0030] where, denotes the estimated value of the tail index calculated based on the first samples, represents the size of the neighborhood window covered by candidate samples, denotes the estimated value of the tail index calculated based on the first samples, denotes the variable value when argmin makes the objective function reach the minimum value, and k is the highest statistical order of the tail region used for estimation.
[0031] As a preferred solution, the method of using the optimal tail fraction as a criterion for determining whether a training sample in a sample dataset is a tail sample, and classifying test samples in the sample dataset using the covariance matrix of tail samples to obtain the scores of classified samples. Specifically includes:
[0032] Use the optimal tail fraction as a criterion for determining whether a training sample in a sample dataset is a tail sample, and obtain tail samples from the training samples.
[0033] Obtain the covariance matrix of each sample in the tail samples and establish a set of covariance matrices.
[0034] Calculate the residual of the i-th class target of each test sample based on the sparse vector represented by the covariance matrix of each test sample in the set of covariance matrices.
[0035] Input the residual of the i-th class target of each test sample into the prediction class model, and select the class with the smallest residual as the class of each test sample.
[0036] Classify the test samples according to the class of each test sample to obtain the scores of the classified samples.
[0037] Among them, the expression of the sparse vector represented by the covariance matrix of each test sample in the set of covariance matrices is as follows:
[0038]
[0039] Among them, represents the sparse vector represented by the covariance matrix of each test sample in the set of covariance matrices, represents the coefficient vector, represents the N-dimensional vector space, represents the L1 norm, represents the L2 norm, represents the covariance matrix of each test sample, represents the set of covariance matrices of the tail samples, represents the regularization parameter, represents the kernel mapping function, (∙) represents the linear logarithmic Euclidean kernel, T represents the transpose of the matrix, and min represents the value of the variable when the objective function reaches the minimum.
[0040] The expression of the residual of the i-th class target of each test sample is as follows:
[0041]
[0042] Among them, represents the residual of the i-th class target of each test sample, (∙) represents the class mask operation, which is used to extract the sparse coefficients of a specific class.
[0043] The expression of the prediction class model is as follows:
[0044]
[0045] Among them, represents the prediction class model.
[0046] As a preferred solution, the method for calculating the corresponding parameters of the tail distribution according to the tail training samples specifically includes:
[0047] Obtain a set of circuit metrics exceeding the tail fraction according to the tail training samples .
[0048] According to the set of circuit metrics exceeding the tail fraction , calculate the parameter , the parameter .
[0049] Among them, the expression of the parameter is as follows:
[0050]
[0051] Among them, N represents the number of elements in the set of circuit metrics exceeding the tail fraction, i represents the serial number of the circuit metric exceeding the tail fraction, represents the i-th circuit metric exceeding the tail fraction.
[0052] The expression of the parameter is as follows:
[0053]
[0054] Among them, max represents the value of the variable when the objective function reaches the maximum value.
[0055] As a preferred solution, the expression of the storage cell yield calculation formula is as follows:
[0056]
[0057] Among them, represents the storage cell yield, represents the probability distribution function of the sample, represents the tail fraction of the probability distribution function of the sample.
[0058] In the formula: represents the classification sample fraction.
[0059] In the formula: The expression of
[0060]
[0061] Among them, represents the location parameter, that is, the starting threshold of the tail region, represents the specific value of the random variable, corresponding to the output of the sample, represents the failure threshold of the sample output.
[0062] In a second aspect, a computer-readable storage medium stores a computer program which, when executed by a processor, implements a method for analyzing the yield of a storage cell based on sparse representation as described in any one of the first aspects.
[0063] In a third aspect, a computer device includes:
[0064] a memory for storing instructions;
[0065] a processor for executing the instructions, such that the computer device performs operations of a method for analyzing the yield of a storage cell based on sparse representation as described in any one of the first aspects.
[0066] Advantageous effects: A method for analyzing the yield of a storage cell based on sparse representation, an electronic device, and a storage medium provided by the present invention, when facing the problem of tail low-probability events, first use a covariance matrix descriptor to summarize the characteristics of a local area to solve the covariance shift problem of the traditional sparse representation classification (SRC) method; use an improved Hill estimator to derive the cumulative probability density of the generalized Pareto distribution (GPD) to calculate the circuit yield; and propose a quasi-maximum likelihood (quasi-ML) method to ensure the stability and robustness of the GPD distribution parameter estimation.
[0067] The present invention overcomes the poor performance of traditional sparse representation in the face of yield analysis problems, combines multi-probe hashing, high-dimensional manifold feature extraction, tail sample optimization, and a sparse classification model, provides a new solution for the high-dimensional data distribution characteristics, and particularly shows significant advantages in scenarios with insufficient tail samples and inconsistent data distributions, provides an efficient and accurate tool for reliability analysis and yield prediction. Compared with the prior art, its advantages are as follows:
[0068] (1) It is proposed to construct a class-balanced training set, which overcomes the problem that the tail samples are too few under the ordinary sampling method, resulting in inaccurate classification of the SRC (Sparse Representation-based Classifier) model.
[0069] (2) Use the MP-LSH (Multi-Probe Locality-Sensitive Hashing) method to search for the nearest neighbor samples, construct a data structure, accelerate the sample search process through dimensionality reduction and random space partitioning, solve the efficiency bottleneck of traditional linear search and the original data structure in high-dimensional scenarios, and propose a covariance matrix descriptor. Extract regional features through the logarithmic Euclidean kernel method on the SPD (Symmetric Positive Definite) manifold, and use LASSO (Least Absolute Shrinkage and Selection Operator) regression and kernel tricks to achieve sparse representation of test samples and calculation of classification residuals, which not only preserves the internal structure of the covariance matrix but also improves the classification efficiency and accuracy.
[0070] (3) Propose a method for selecting tail fractions that balances variance and bias through the Hill estimator and stability measure, ensuring sufficient and effective samples are obtained in the tail region, while avoiding estimation biases caused by overly high or low tail fractions. At the same time, propose a quasi-maximum likelihood method based on tail samples for estimating the relevant parameters of the Generalized Pareto Distribution (GPD), improving the estimation accuracy of the tail failure probability. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is a flowchart of the yield analysis of storage cells based on sparse representation. DETAILED DESCRIPTION OF THE INVENTION
[0072] The following describes clearly and completely the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0073] The following further illustrates the present invention with specific embodiments.
[0074] Embodiment 1:
[0075] This embodiment introduces a method for yield analysis of storage cells based on sparse representation, which specifically includes the following steps:
[0076] Step S01: Construct a class-balanced data set. For a circuit netlist, a set of circuit process parameters (inputs) and the corresponding circuit performance (outputs) obtained through simulation are used as a sample. During simulation, Latin Hypercube Sampling (LHS) is used to collect an equal number of samples in the failure region and non-failure region of the simulation results to ensure the uniformity of data distribution. Finally, a class-balanced data set is constructed using the two types of samples.
[0077] Furthermore, due to the limitation of the traditional SRC method by the imbalance of the training set, the present invention uses Latin Hypercube Sampling (LHS) to sample an equal number of samples from the failure region and the non-failure region. This method is suitable for learning high-quality representations of long-tailed distributions.
[0078] Step S02: Based on the obtained dataset with class balance, calculate the covariance matrix. For a single sample in the dataset and the samples in the nearby space of the sample, use the Multi-Probe Locality-Sensitive Hashing (MP-LSH) method to reduce the dimension and randomly partition the space of the sample, reconstruct the data structure of the sample, and efficiently map and store the proximity relationship of the sample into the hash bucket, thereby significantly improving the sample retrieval speed and calculating the hash key value. Conduct approximate nearest neighbor (ANN) sample search through the hash key value, calculate the covariance using the searched nearest neighbor samples, and finally construct the covariance matrix for the entire dataset.
[0079] Furthermore, when extracting the features of the covariance matrix, the traditional linear search method and the original data structure become inefficient when applied to the k-nearest neighbor algorithm in high-dimensional scenarios. The present invention uses the Multi-Probe Locality-Sensitive Hashing (MP-LSH) method as an approximate nearest neighbor (ANN) search, and constructs a data structure based on dimensionality reduction and random space partitioning. The main concept is to select a hash function that can map samples from the original high-dimensional space to different positions in one or more hash buckets.
[0080] Step 2.1: Use k hash functions to form a hash table , where the hash function is selected as , where is a d-dimensional random vector, each component is independently sampled from the standard Gaussian distribution, is sampled from a uniform distribution offset, is the interval length, representing the coverage range of the hash bucket, so that nearby samples are more likely to be mapped to the same hash bucket than samples farther away. After obtaining the hash table , for each sample in the dataset, use to calculate their respective hash key values, map the samples to the corresponding hash buckets according to the hash key values, and finally, during the query, the samples in the corresponding hash bucket can be found using the hash key of the query sample.
[0081] To reduce the deviation of the projected object, an AND-OR strategy is adopted to enhance the accuracy and check multiple hash buckets that may contain the nearest neighbors of the query point. Each hash function first maps Projected onto a one-dimensional line segment, which is divided into several intervals, and the hash key value is the serial number of the interval where it is located. For a sample , because the hash key value difference roughly follows a Gaussian distribution. Since the potential nearest neighbor samples will fall within and in a certain interval, for multi-probe locality-sensitive hashing (MP-LSH), set , and apply perturbations to the hash key values of the query samples, which can cover neighboring buckets to improve the recall rate. Finally, merge the samples in the hash buckets corresponding to all the perturbed hash key values to obtain a candidate sample set. Search for the nearest neighbor samples in the candidate sample set, which can effectively accelerate the search for k nearest neighbor samples. This extension provides a solution that can locate more hash buckets that may contain the nearest neighbors of the query object. Such a search method can significantly improve the search efficiency. To alleviate the prediction differences caused by covariate shift, introduce the covariance matrix features constructed by each sample and its neighboring samples into the classification framework of SRC. Given the training sample and its nearest neighbor samples , the covariance matrix can be obtained by the following formula:
[0082]
[0083] where: is the number of nearest neighbor samples, represents the th nearest neighbor sample, represents the mean value of the neighboring samples within the region , T represents the transpose of the matrix, represents the covariance matrix of the training sample .
[0084] Step S03: Train a classifier to classify the tail samples.
[0085] For the problem of storage cell yield analysis that follows a tail distribution, using the tail fraction , the probability distribution function can be re-expressed as:
[0086]
[0087] The corresponding yield calculation can be expressed by the following formula:
[0088]
[0089] where, It can be approximately replaced by estimating the tail sample scores in a large-scale test sample through a classification method. In this step, the classification method is used to estimate this part. represents the probability distribution function of the sample, represents the tail score of the sample, represents the random variable output by the sample, represents the yield of the storage unit.
[0090] First, among all samples, the Hill estimator is used to select an appropriate tail score , and the samples are classified into tail samples and non-tail samples. Then, the geodesic distance between the covariance matrices of the test samples and the training samples is tested using the Log-Euclidean Metric (LEM). Based on the geodesic distance, a linear Log-Euclidean kernel function is defined to map the covariance matrix to a high-dimensional Reproducing Kernel Hilbert Space (RKHS). Finally, the LASSO regression method is used to construct a classifier. In the RKHS, the features of the test samples are represented by the sparse vectors of the training set, and the residuals of each category are calculated by combining the linear Log-Euclidean kernel function. Then, the classification of the samples is completed by selecting the category corresponding to the minimum residual.
[0091] Furthermore, in the modeling of the Generalized Pareto Distribution, the tail score has a significant impact on the accuracy of the yield analysis. Sufficient samples are required in the tail region to construct an effective SRC framework. To select an appropriate in a compromising way, the present invention uses the Hill estimator to calculate the estimate of the tail score, and obtains an optimal that balances the variance and bias of the Hill estimator. Given a sequence sorted in ascending order, where represents the magnitude of the probability density corresponding to the -th sample. The Hill estimator can be expressed as:
[0092]
[0093] where: represents the estimated value of the tail exponent, k is the highest statistical order of the tail region used for estimation, represents the magnitude of the probability density corresponding to the -th sample, represents the magnitude of the probability density corresponding to the -th sample.
[0094] The tail score can be solved as an optimization problem and expressed as:
[0095]
[0096] Among them, represents the estimated value of the tail index calculated based on the previous samples, represents the candidate sample covering the size of the neighborhood window, represents the estimated value of the tail index calculated based on the previous samples.
[0097] After obtaining the best tail sample score, the covariance matrix is used to classify the test samples . If the covariance matrices of two regions are highly similar, it is assumed that the two regions are similar. Therefore, the distribution difference between the training set and the test set can be characterized by extracting the covariance matrix. Also, because the covariance matrix is a symmetric positive definite matrix (Symmetric Positive Definite SPD), the log-Euclidean kernel can be used to measure the geodesic distance between SPD covariance matrices, and this method is used to map the SPD manifold to a high-dimensional reproducing kernel Hilbert space (RKHS). For two given SPD matrices , the distance measured by the log-Euclidean metric (LEM)
[0098]
[0099] can be expressed as: represents the F-norm of the matrix. Inspired by the inner product and the log-Euclidean distance, the linear log-Euclidean kernel in the SPD manifold is defined as:
[0100]
[0101] Among them, represents the implicitly defined kernel space, and tr() represents the trace of the matrix.
[0102] For a given kernel mapping function , the covariance feature of the test sample can be sparsely represented as a sparse vector , and the sparse vector is represented on the set of covariance matrices obtained from training samples. At the same time, the LASSO regression method is introduced to construct a kernel sparse model on the manifold and simplify it through the kernel trick:
[0103]
[0104] Wherein: is the Log-Euclidean kernel mapping, represents the sparse representation coefficient, represents the covariance matrix of the test sample, represents the set of covariance matrices of the training samples, represents the coefficient vector, represents the regularization parameter. Thus, for a test sample , its covariance feature is represented as a sparse vector . Then the residual and predicted category of the i-th class target of the test sample can be expressed as:
[0105]
[0106]
[0107] Calculate the residual of each category , and select the category with the smallest residual as the prediction result. Finally, by classifying the test samples, the ratio of the tail samples to the non-tail samples in the obtained classified samples can be used to replace in Formula 1.
[0108] Step S04: Fit the tail distribution and finally calculate the circuit yield.
[0109] For the part mentioned above , in this step, fit to solve. After obtaining the tail fraction , simulate the circuit netlist and re-collect the tail samples. For the collected tail samples, use the quasi-maximum likelihood (qusai-MLE) method to fit the shape parameter and scale parameter in the generalized Pareto distribution (GPD), and use the GPD to calculate , and finally calculate the circuit yield.
[0110] Furthermore, after obtaining the tail fraction, for the resampling of the tail samples, the corresponding parameters of the tail distribution can be calculated. The present invention proposes a quasi-maximum likelihood (quasi-ML) method to estimate the corresponding parameters:
[0111]
[0112]
[0113] Wherein represents the circuit metric exceeding the tail fraction. in Formula 1It can be obtained from the following formula:
[0114]
[0115] Wherein, represents the position parameter, i.e., the starting threshold of the tail region, represents the specific value of the random variable, corresponding to the output of the sample, represents the failure threshold of the sample output. Use to replace , and finally substitute the result into Formula 1 to calculate the specific yield level.
[0116] Example 2:
[0117] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for analyzing the yield of a storage unit based on sparse representation as described in any one of Example 1.
[0118] Example 3:
[0119] A computer device, comprising:
[0120] A memory for storing instructions.
[0121] A processor for executing the instructions, so that the computer device performs the operations of a method for analyzing the yield of a storage unit based on sparse representation as described in any one of Example 1.
[0122] Example 4:
[0123] This example introduces the working principle of a method for analyzing the yield of a storage unit based on sparse representation. The present invention uses the covariance matrix of each training sample and its k nearest neighbor samples to implement classifier training. These covariance matrix features can describe the local features of the samples and help capture the geometric relationships between the samples. After obtaining the covariance matrix features, the Log-Euclidean kernel is used to measure the geodesic distance between the covariance matrices to effectively process the geometric relationships of the covariance matrices. In addition, in order to construct a sparse representation model, the LASSO regression method is used to perform sparse representation on the covariance matrix features. Through this method, the most representative features can be extracted from the training samples and redundant information can be reduced. After establishing the sparse model, the method of minimizing the residual is used for sample classification. For each test sample to be classified, calculate its residual with the training samples, and determine the category of the test sample by minimizing the residual, so as to improve the accuracy and efficiency of classification.
[0124] In the Generalized Pareto Distribution (GPD) modeling, the tail fraction The choice directly affects the modeling effect. A larger will result in too few tail samples, increasing the variance of the estimation, while a smaller will introduce too many tail samples, thus causing estimation bias. The present invention uses the Hill estimator to select the optimal tail fraction. The Hill estimator is based on the sorting of tail samples, approximates the tail index by calculating the logarithmic differences of the sequence, and then finds a stable by minimizing the variation of the tail samples to achieve a compromise between variance and bias. After obtaining the GPD modeling, tail resampling will be performed, and the quasi-maximum likelihood (qusai-MLE) method will be used to estimate the parameters and related to the failure probability. These parameters are adjusted by the maximum value of the tail samples, and the obtained parameters are used to further estimate the failure probability. Specifically, the reshuffled tail samples are used to calculate the shape parameter and the scale parameter of the tail distribution, and a more accurate estimate is obtained through the quasi-maximum likelihood method, thereby obtaining the final failure probability.
[0125] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A storage cell yield analysis method based on sparse representation, characterized in that: Specifically include: Simulating the stored circuit netlist, and collecting equal numbers of samples from the failure area and the non-failure area of the sampling simulation result as a sample data set; According to the sample data set, k nearest neighbor samples are obtained, and the covariance matrix of the sample is calculated according to the sample data set and the corresponding k nearest neighbor samples; A method for obtaining the best tail fraction using an estimator, using the best tail fraction as a criterion for measuring whether a training sample in a sample data set is a tail sample, and classifying a test sample in the sample data set using the covariance matrix of the tail sample to obtain the score of the classified sample; Simulating the stored circuit netlist corresponding to the best tail number, recollecting the tail training samples, and calculating the corresponding parameters of the tail distribution according to the tail training samples; Substitute the scores of classified samples and the corresponding parameters of tail distribution into the storage cell yield calculation formula to obtain the storage cell yield.
2. The storage unit yield analysis method based on sparse representation according to claim 1, characterized in that: An equal number of samples are collected from the failure area and the non-failure area of the simulation results using Latin hypercube sampling as the sample data set.
3. The storage unit yield analysis method based on sparse representation according to claim 1, characterized in that: The method for obtaining the k nearest neighbor samples specifically includes: According to the hash table Calculate each sample in the sample data set The hash key value of the sample Map the hash key value to the corresponding hash bucket; For query sample Perturb the hash key value ,in, , ; Calculate the query sample in the hash bucket corresponding to the hash key value after adding the disturbance With sample The hash key value difference; The merge hash key value difference falls within and Query samples within the interval As a candidate sample set, Represents a hash table Middle Samples corresponding to hash functions Hash key value; Search for neighbor samples in the candidate sample set and obtain k nearest neighbor samples.
4. The storage unit yield analysis method based on sparse representation according to claim 1, characterized in that: The covariance matrix expression of the sample is as follows: ; in, is the number of nearest neighbor samples, Indicates nearest neighbor samples, represents the mean of neighboring samples, T represents the transpose of the matrix, Representation sample The covariance matrix of .
5. The storage unit yield analysis method based on sparse representation according to claim 1, characterized in that: The method of using an estimator to obtain an optimal tail fraction specifically includes: Construct the Hill estimator and calculate the estimated value of the tail exponent; the expression of the Hill estimator is as follows: ; in, represents the estimated value of the tail index, k is the highest statistical order of the tail region used for estimation, Indicates The size of the probability density corresponding to the sample, Indicates The size of the probability density corresponding to the sample; Substitute the estimated value of the tail index into the optimal tail fraction optimization model, solve the optimal tail fraction optimization model to obtain the optimal tail fraction; the optimal tail fraction optimization model expression is as follows: ; in, Indicates that based on the previous The tail exponential estimate calculated from samples, Represents the size of the neighborhood window covered by the candidate sample, Indicates that based on the previous The tail exponential estimate calculated from samples, represents the value of the variable when argmin minimizes the objective function, and k is the highest statistical order of the tail region used for estimation.
6. The storage unit yield analysis method based on sparse representation according to claim 1, characterized in that: The method of using the best tail fraction as a standard for measuring whether a training sample in a sample data set is a tail sample, and using the covariance matrix of the tail samples to classify the test samples in the sample data set to obtain the scores of the classified samples specifically includes: The best tail fraction is used as a criterion for measuring whether a training sample in a sample data set is a tail sample, and a tail sample is obtained from the training sample; Obtain the covariance matrix of each sample in the tail sample and establish a covariance matrix set; Calculate the residual of the i-th target of each test sample according to the sparse vector represented by the covariance matrix of each test sample in the covariance matrix set; Input the residual of the i-th target of each test sample into the prediction category model, and select the category with the smallest residual as the category of each test sample; Classify the test samples according to the category of each test sample and obtain the score of the classified samples; Among them, the sparse vector expression represented by the covariance matrix of each test sample in the covariance matrix set is as follows: ; in, Represents the sparse vector represented by the covariance matrix of each test sample in the covariance matrix set, represents the coefficient vector, represents an N-dimensional vector space, represents the L1 norm, represents the L2 norm, represents the covariance matrix of each test sample, represents the set of covariance matrices of the tail samples, represents the regularization parameter, represents the kernel mapping function, ( ∙ ) represents the linear logarithmic Euclidean kernel, T represents the transpose of the matrix, and min represents the variable value that minimizes the objective function; The expression of the residual of the i-th target of each test sample is as follows: ; in, represents the residual of the i-th target of each test sample, (∙) represents the category mask operation; The expression of the prediction category model is as follows: ; in, Represents a predictive class model.
7. The storage unit yield analysis method based on sparse representation according to claim 1, characterized in that: The method for calculating corresponding parameters of the tail distribution according to the tail training samples specifically includes: According to the tail training samples, obtain the circuit metric set exceeding the tail number ; Based on the circuit metric set exceeding the tail number , calculation parameters ,parameter ; Among them, the parameters The expression is as follows: ; Where N represents the number of elements in the circuit metric set that exceeds the tail fraction, i represents the sequence number of the circuit metric that exceeds the tail fraction, represents the circuit metric of the i-th circuit that exceeds the tail fraction; parameter The expression is as follows: ; Among them, max represents the variable value when the objective function reaches the maximum value.
8. The storage unit yield analysis method based on sparse representation according to claim 1, characterized in that: The storage unit yield calculation formula is as follows: ; in, represents the storage cell yield, represents the probability distribution function of the sample, Indicates the tail number The probability distribution function of the sample; Where: Represents the classification sample score; Where: The expression is as follows: ; in, Represents positional parameters, represents the output of the sample, Indicates the failure threshold of the sample output.
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, a storage unit yield analysis method based on sparse representation as described in any one of claims 1 to 8 is implemented.
10. A computer device, characterized in that: include: A memory for storing instructions; A processor is used to execute the instructions so that the computer device performs the operation of the storage cell yield analysis method based on sparse representation as described in any one of claims 1 to 8.