Failure Region Assessment Method Applied to High Sigma Scenarios in Electronic Engineering
The failure area of the integrated circuit is evaluated through the weighted rank correlation algorithm, and the computing resources and accuracy problems of the existing technology in the high Sigma scenario are solved, and efficient and accurate failure area evaluation is achieved, which is suitable for large-scale integrated circuits and high-reliability systems.
Patent Information
- Application Number
- CN202411598645.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In the high Sigma scenario, it is difficult for the existing technology to efficiently and accurately evaluate the failure area of an integrated circuit. The traditional Monte Carlo method has high computing resources and time costs, which accelerates the failure probability estimation deviation of the Monte Carlo method during large-scale variable processing. The high Sigma Monte Carlo method ignores the failure area sorting accuracy, and the maximum and minimum rank method lacks flexibility, making it difficult to distinguish the failure area from the ordinary area.
Weighted rank correlation algorithm is used to establish initial samples, predict models, screen high-risk samples, simulate high-risk samples, calculate the offset and Spearman rank coefficient, give different weights to samples in different regions, perform weighted summing, and improve the ranking accuracy of failed regions.
It significantly improves the sorting accuracy of the failure area, reduces unnecessary simulation times, saves computing resources, and is suitable for failure area evaluation of large-scale integrated circuits and high-reliability systems.
Smart Images

Figure CN119538830B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of methods for evaluating rare events in electronic engineering, and in particular relates to a failure area evaluation method applied to high Sigma scenarios in electronic engineering. Background Art
[0002] Integrated circuit (IC) component production in high Sigma scenarios can only tolerate a few defects in hundreds of millions or billions of instances. This is because these components are usually replicated in large arrays, so producing a properly functioning product requires a large number of replicated components to work correctly. Once fluctuations occur in the manufacturing process, this can lead to performance deviations or even failures. Therefore, analyzing IC yield has become a critical task.
[0003] The traditional Monte Carlo (MC) method is a common technique for evaluating IC yield. It collects a large amount of circuit data by simulating deviations in the circuit process, then simulates all circuits to determine the number of failed circuits, and finally obtains the circuit yield calculation. Its application efficiency is low in high Sigma scenarios because a large number of simulations are required to capture rare failure events, resulting in huge computing resources and time costs. To solve this problem, the industry has proposed a variety of accelerated Monte Carlo methods (such as importance sampling and model substitution methods), as well as high Sigma Monte Carlo (HSMC) methods. Although these methods have made some progress in reducing the number of samples and improving simulation efficiency, existing methods often ignore the importance of alternative models in the sorting of failure areas, resulting in inaccurate capture of failure events.
[0004] These methods are analyzed in detail as follows: 1. Traditional Monte Carlo method: Although the results are reliable and applicable to circuit designs of different dimensions, the number of samples required is extremely large in high Sigma scenarios, resulting in huge computing time and resource consumption; 2. Accelerated Monte Carlo method: Although the importance sampling method can reduce the number of simulation samples, it performs poorly when dealing with large-scale variables and can easily lead to failure probability estimation errors; 3. High Sigma Monte Carlo method (HSMC): HSMC reduces the number of simulations by giving priority to failed samples, but relies on global error minimization, ignores the sorting accuracy of the failure area, and cannot accurately capture rare events; 4. Maximum and minimum rank method: Although it improves the sorting accuracy of the model, it lacks flexibility in the allocation of sample weights, and it is difficult to accurately distinguish samples from failure areas and ordinary areas.
[0005] Weighted Rank is a ranking method that takes into account the weights of different factors or features in a ranking or scoring system. Different from simple ranking, weighted rank allows certain factors to have a greater impact on the final ranking while other factors have a smaller impact. Applying the algorithm of weighted rank to this evaluation method, ranking the different influencing factors in the rare event area by weight, and then evaluating the occurrence probability of various rare events. The applicant believes that this can obviously overcome the limitations of the prior art. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a failure area evaluation method applied to high-Sigma scenarios in electronic engineering. This evaluation method evaluates the ranking deviation of the surrogate model in the rare event area through a weighted rank-related algorithm, improving the accuracy of capturing failure samples.
[0007] To solve the above technical problems, the present invention adopts the following technical solutions: A failure area evaluation method applied to high-Sigma scenarios in electronic engineering, including the following steps:
[0008] The first step: Establish an initial sample Sam 初始 : Randomly select the initial sample Sam 总 from the total sample set Sam 初始 . The quantity Qua 初始 of the initial sample Sam 初始 satisfies the formula (1): Qua 初始 = max(Qua 总 / 1000, 1000) (1), where Qua_total is the total sample quantity, and the initial sample quantity Qua 初始 is one-thousandth of the total sample quantity Qua 总 , and is not less than 1000. This quantity requirement not only ensures that the initial sample Sam 初始 covers the basic characteristics of the total sample set Sam 总 , but also controls the calculation amount.
[0009] The second step: Establish a prediction model: Simulate the selected initial sample Sam 初始 to obtain the simulation result YS 初始 . Use the simulation result YS 初始 to train a machine learning equivalent surrogate model to generate a prediction model.
[0010] Preferably, the simulation tool adopts a Spice circuit simulation tool.
[0011] Preferably, alternative models such as the FFX algorithm, linear regression, Lasso regression, and LightGBM can be used to establish the prediction model.
[0012] Step 3: Screening high-risk samples Sam 风险 :Use the prediction model to predict the remaining samples Sam 剩余 Make a prediction and get the prediction result YP 剩余 , where the remaining sample Sam 剩余 Sam is the total sample set 总 Remove the initial sample Sam 初始 The following sample set, the remaining sample Sam 剩余 The predicted results of YP 剩余 Arrange from large to small, and select high-risk samples Sam from large to small 风险 , high-risk sample Sam 风险 The number of Qua 风险 Satisfying formula (2): Qua 风险 =max(Qua 总 / 2000,500) (2), where Qua 总 is the total number of samples, high-risk sample Sam 风险 The number of Qua 风险 is the total sample size Qua 总 0.05% and no less than 500%, high-risk sample Sam 风险 The predicted value is YP 风险 .
[0013] Using the initial sample Sam 初始 Train the prediction model and then use it to predict the large number of remaining samples Sam 剩余 Make predictions, rank the prediction results, and select a certain number of remaining samples Sam with high rankings 剩余 As a high-risk sample Sam 风险 Further analysis and calculation are performed, and the remaining samples Sam ranked at the bottom 剩余 The failure rate is very low, so it is not studied.
[0014] Step 4: Simulate high-risk sample Sam 风险 :For high-risk sample Sam 风险 Perform simulation, obtain simulation results, and arrange the simulation results from large to small to obtain simulation results YS 风险 .
[0015] Step 5: Calculate the numerical offset Δd and the average numerical offset Δd 平均 : Arrange the simulation results YS from large to small 风险 Calculate the difference between two adjacent prediction values one by one, that is, Δd i =(YS 风险 ) i -(YS 风险 )i+1 (3), where i is each natural number ≤ the quantity Qua of the high-risk samples Sam 风险 of -1; and the average numerical offset of the entire risk sample Sam 风险 of each natural number -1; and the average numerical offset of the entire risk sample Sam 风险 of the average numerical offset
[0016] Step 6: Determine the samples Sam in the near-failure region 接近 and the samples Sam in the sub-near-failure region 次接近 : Set a simulation value threshold Thr 实际 , then the threshold Thr of the samples Sam in the near-failure region 接近 is Thr 接近 = Thr 实际 - 3Δd 平均 (5), and the threshold Thr of the samples Sam in the sub-near-failure region 次接近 is Thr 次接近 = Thr 实际 - 6Δd 平均 (6).
[0017] Samples with a simulation value greater than or equal to Thr 接近 are used as the samples Sam in the near-failure region 接近 , and samples with a simulation value greater than or equal to the threshold Thr of the sub-near-failure region 次接近 and less than the threshold Thr of the near-failure region samples 接近 are used as the samples Sam in the sub-near-failure region 次接近 , then the quantity of the samples Sam in the near-failure region 接近 is Qua 接近 , and the quantity of the samples Sam in the sub-near-failure region 次接近 is Qua 次接近 .
[0018] Step 7: Determine the samples Sam in the normal region 普通 : Samples remaining after removing the samples Sam in the near-failure region 风险 and the samples Sam in the sub-near-failure region 接近 from the high-risk samples Sam 次接近 are the samples Sam in the normal region 普通 , then the quantity Qua 普通 of the samples Sam in the normal region 普通 satisfies formula (9), Qua 普通 = Qua 风险 - Qua 接近 - Qua 次接近 (7).
[0019] Step 8: Calculate the sorting offset Δrank of the prediction result: Compare the high-risk samples Sam风险 The simulation result YS of each sample data in 风险 and the prediction result YP 风险 Sort the differences to obtain the sorting offset Δrank. The sorting offset Δrank satisfies formula (8):
[0020] Δrank i = |rank(YS 风险 ) i - rank(YP 风险 ) i | (8), where i is each natural number ≤ the quantity Qua 风险 of the high - risk samples Sam 风险 rank(YS 风险 ) is the ranking serial number of the simulation result YS 风险 rank(YP 风险 ) is the ranking serial number of the prediction result YP 风险 ;
[0021] Step 9: Calculate the Spearman rank coefficient ρ1 of the samples Sam 接近 in the near - failure region: As in formula (9), that is where the value of Δrank i ' in the formula is as follows:
[0022] where Thr1 and Thr2 are the artificially set primary threshold Thr1 and secondary threshold Thr2 for the samples in the near - failure region.
[0023] When Δd ≤ Thr1, it is considered that the sorting of this sample is very accurate and close to the failure region. Therefore, the sorting relative offset Δrank i ' is set to 0, indicating that these samples play the most important role in the evaluation and do not require sorting offset processing in the evaluation; while when Thr1 < Δd 平均 ≤ Thr2, the sorting error is small, and the sorting relative offset Δrank' is assigned Δrank / 2, which means that these samples perform well within the failure region, but their importance is relatively reduced; while when Δd 平均 > Thr2, Δrank′ is assigned Δrank, which means that the sorting deviation of this sample is large, but it is still within the expansion range of the failure region. Therefore, its sorting offset is directly assigned.
[0024] Step 10: Calculate the Spearman rank coefficient ρ2 of the samples Sam 次接近 in the sub - near - failure region and the Spearman rank coefficient ρ3 of the samples Sam 普通 in the normal region: Then ρ2 is as in formula (10), that is ρ3 is as shown in formula (11), that is All the data in these two formulas are known or can be calculated. Substituting them can obtain ρ2 and ρ3.
[0025] Step 11: Comprehensive weighted summation of ρ 总 : Define the weight of the samples in the near-failure region as α1, the weight of the samples in the second-near-failure region as α2, and the weight of the samples in the normal region as α3. Then ρ 总 The calculation formula of is as shown in formula (12), that is ρ 总 = α1·ρ1 + α2·ρ2 + α3·ρ3 (12), where α3 = 1 - α1 - α2, and α1 > α2 > α3.
[0026] In the final evaluation, we calculate the weighted values of the samples in different regions and comprehensively consider the contributions of the samples in the near-failure region Sam 接近 , the samples in the second-near-failure region Sam 次接近 and the samples in the normal region Sam 普通 . Different weights are assigned to each region. Among them, the samples in the near-failure region Sam 接近 and the samples in the second-near-failure region Sam 次接近 will be assigned greater weights, while the weight of the samples in the normal region Sam 普通 is usually set to a smaller value to reduce its impact on the overall evaluation result.
[0027] Preferably, we set α1 to 0.6 - 0.7 to ensure the leading role of the samples in the failure region Sam 接近 in the overall evaluation; and set α2 to 0.2 - 0.3. Although the samples in the second-near-failure region Sam 次接近 are not as crucial as the samples in the near-failure region Sam 接近 , they are still in a relatively important position; α3 is usually set to a smaller value, such as 0.1 - 0.2, indicating that the samples in the normal region Sam 普通 have little impact on the overall evaluation.
[0028] Preferably, the high-risk samples Sam 风险 are divided into n regions according to the situation. By calculating the Spearman rank correlation coefficient of each region and defining the weight of each region, the final weighted rank correlation evaluation result is calculated. Its calculation is as shown in formula (13), where Pfina1 is the final weighted rank correlation evaluation result, ρi is the Spearman rank correlation coefficient of the samples in the i-th region, and α i is the weight of the i-th region, and it satisfies
[0029] More preferably, according to the actual situation, the high-risk samples Sam 风险It is divided into 4 regions according to the situation, namely the near failure region, the sub - near failure region, the extended region and the normal region, with weights of α1, α2, α3 and α4 respectively, and Spearman rank coefficients of ρ1, ρ2, ρ3 and ρ4 respectively. Then the final weighted rank result ρ final = α1·ρ1 + α2·ρ2 + α3·ρ3 + α4·ρ3 (14), where α1 = 0.5, α2 = 0.2, α3 = 0.2, α4 = 0.1.
[0030] Preferably, the value ratios and the minimum value numbers in formula (1) and formula (2) change according to the actual situation.
[0031] Preferably, the value of Δrank i ' in formula (9) changes according to the actual situation.
[0032] The beneficial effects of the present invention are as follows:
[0033] This evaluation method uses weighted rank correlation to evaluate the sorting function of the surrogate model, especially significantly improving the accuracy in the sorting performance of the high - Sigma failure region; through more accurate model evaluation, it reduces unnecessary simulation times, improves efficiency, and saves time and computing resources.
[0034] By adopting a dynamic weighting formula, regions can be freely divided and weights can be dynamically adjusted. According to the needs of the actual engineering scenario, the evaluation model can be flexibly adjusted, and the evaluation result is more in line with the real situation; by introducing multi - region division and weighted calculation, it helps to improve the accuracy of rare event capture, especially in the evaluation of the high - Sigma failure region, and can better utilize computing resources; the adoption of a general formula can be applied to any number of divided regions and can be widely used in large - scale integrated circuits, high - reliability systems and other scenarios requiring high - Sigma failure regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is the overall process schematic diagram of the present invention;
[0036] Figure 2 is the calculation flow chart of the weighted rank correlation of model evaluation in the present invention;
[0037] Figure 3 is the division definition diagram of each region in the risk sample of the present invention;
[0038] Figure 4 is the schematic diagram of the calculation of the number of each region in the risk sample of the present invention;
[0039] Figure 5 The schematic diagram of Comparative Example 1 in the present invention;
[0040] Figure 6It is a schematic diagram of Comparative Example 2 of the present invention. Detailed implementation manners
[0041] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the structure of the present invention.
[0042] Embodiment 1: As shown in the specification appendix Figure 1-2 If the total sample set Sam 总 contains 1,000,000 sample data, then an initial sample Sam 初始 is randomly selected from it. According to the formula Qua 初始 = max(Qua 总 / 1000, 1000) (1), the number of initial samples Qua 初始 is 1000. The selected initial sample Sam 初始 is simulated using a Spice circuit simulation tool to obtain a simulation result YS 初始 . The FFX algorithm is used to train a machine learning equivalent substitution model for the simulation result YS 初始 to generate a prediction model.
[0043] The sample set remaining after removing the initial sample Sam 总 from the total sample set Sam 初始 constitutes the remaining sample Sam 剩余 . Then the number of the remaining sample Sam 剩余 is 999,000. The prediction model is used to predict these 999,000 remaining samples Sam 剩余 to obtain a prediction result YP 剩余 . The prediction result YP 剩余 is arranged from largest to smallest. We need to screen out high-risk samples Sam 风险 . The number Qua 风险 of the high-risk samples Sam 风险 satisfies the formula (2): Qua 风险 = max(Qua 总 / 2000, 500) (2). Then the top 500 ranked are screened as high-risk samples Sam 风险 . The prediction values of the high-risk samples Sam 风险 are arranged from largest to smallest to obtain a prediction result YP 风险 . These are the data with the highest risk in the total sample set Sam 总 . We will focus on analyzing them.
[0044] To improve the accuracy of the evaluation method, we will continue to partition these 500 high-risk samples Sam 风险 and divide them into samples close to the failure area according to the actual situation接近 , the samples Sam in the sub-closest-to-failure region 次接近 and the samples Sam in the normal region 普通 , and assign different weights to them for weighted calculation. Taking Figure 3 as an example, after simulating the high-risk samples Sam 风险 , we sort their simulation values. The samples with a more forward risk ranking are the samples Sam in the closest-to-failure region 接近 , while the samples with the last risk ranking are the samples Sam in the normal region 普通 , and the samples Sam in the sub-closest-to-failure region are located between the two 次接近 . From Figure 3 we can also see that the sorting of the simulation results and prediction results of the same sample is different.
[0045] Among them, the number of samples Sam in the closest-to-failure region 接近 , the samples Sam in the sub-closest-to-failure region 次接近 and the samples Sam in the normal region 普通 each contain how many, and their calculation methods are as Figure 4 shown. For the sake of brief expression, we only assume 10 samples as the risk samples Sam 风险 .
[0046] We first determine the number of samples in the closest-to-failure region. Here, we will use the numerical offset Δd and the average numerical offset Δd 平均 : Calculate the difference between two adjacent predicted values one by one for the simulation results YS 风险 arranged from largest to smallest, that is, Δd i =(YS 风险 ) i -(YS 风险 ) i+1 (3), where i is each natural number ≤ the quantity Qua 风险 of the high-risk samples Sam 风险 minus 1; and the average numerical offset 风险 of the entire risk sample Sam
[0047] is among Figure 4 , Δd1 = 68 - 58 = 10, Δd2 = 58 - 44 = 14...... Δd9 = 6 - 4 = 2. The specific results are shown in the Δd column, and the average numerical offset Δd 平均 is calculated to be 7.11.
[0048] We set a simulation value threshold Thr 实际 , then the threshold Thr 接近 of the samples Sam in the closest-to-failure region 接近 = Thr 实际-3Δd 平均 (5), the threshold Thr of the samples Sam in the sub - close - to - failure region 次接近 is 次接近 =Thr 实际 -6Δd 平均 (6).
[0049] Here, we set the simulation value threshold Thr 实际 to be 52. Then, for the samples Sam in the close - to - failure region, the threshold Thr 接近 is 接近 =52 - 3*7.11=30.67. For the samples Sam in the sub - close - to - failure region, the threshold Thr 次接近 is 次接近 =52 - 6*7.11=9.34. Then, all samples with simulation values greater than or equal to 30.67 are selected as the samples Sam in the close - to - failure region 接近 , that is Figure 4 the red - colored area part in ; while samples with simulation values greater than or equal to 9.34 and less than 30.67 are selected as the samples Sam in the sub - close - to - failure region 次接近 , that is Figure 4 the yellow - colored area part in ; the remaining samples are the samples Sam in the normal region 普通 , that is Figure 4 the green - colored area part in . That is, among these 10 samples, those ranked 1 - 4 are the samples Sam in the close - to - failure region 接近 , and those ranked 5 - 8 are the samples Sam in the sub - close - to - failure region 次接近 , and the remaining samples ranked 9 and 10 are the samples Sam in the normal region 普通 .
[0050] After dividing the high - risk samples Sam 风险 into three regions, we need to perform a weighted sum on the three regions. Among them, the weights α1, α2, and α3 of the samples Sam in the close - to - failure region 接近 , the samples Sam in the sub - close - to - failure region 次接近 and the samples Sam in the normal region 普通 are defined artificially. However, the Spearman rank coefficients ρ1, ρ2, and ρ3 of each region need to be calculated.
[0051] In the Spearman rank - coefficient formula, the sorting offset Δrank is used: The sorting difference between the simulation result YS 风险 of each sample data and its prediction result YP 风险 is obtained, that is, the sorting offset Δrank. The sorting offset Δrank satisfies formula (8):
[0052] Δrank i =|rank(YS 风险 )i - rank(YP 风险 ) i | (8), where i is each natural number less than or equal to the quantity Qua of high - risk samples Sam 风险 of 风险 rank(YS 风险 ) is the ranking serial number of the simulation result YS 风险 rank(YP 风险 ) is the ranking serial number of the prediction result YP 风险 .
[0053] Calculate the Spearman rank coefficient ρ1 of the samples Sam 接近 close to the failure region: As shown in formula (9), that is where the value of the sorting relative offset Δrank1’ in the formula is as follows: where Thr1 and Thr2 are the artificially set primary threshold Thr1 of the samples close to the failure region and the secondary threshold Thr2 of the samples close to the failure region respectively.
[0054] When Δd ≤ Thr1, it is considered that the sorting of this sample is very accurate and close to the failure region. Therefore, the sorting relative offset Δrank i ’ is set to 0, indicating that these samples play the most important role in the evaluation and do not require sorting offset processing in the evaluation; while when Thr1 < Δd 平均 ≤ Thr2, the sorting error is small, and Δrank’ is assigned Δrank / 2, which means that these samples perform well within the failure region, but their importance is relatively reduced; while when Δd 平均 > Thr2, Δrank’ is assigned Δrank, which means that the sorting deviation of this sample is large, but it is still within the extended range of the failure region. Therefore, directly assign its sorting offset Δrank.
[0055] Still taking the data in Figure 4 as an example, the calculation method of Δrank is: Δrank1 = 3 - 1 = 2, Δrank2 = 6 - 2 = 4...... Δrank 10 = 10 - 7 = 3, specifically see the last column data in Figure 4 , and then substitute Δrank i into formula (11) one by one to obtain ρ1.
[0056] The Spearman rank coefficient ρ2 of the samples Sam 次接近 close to the failure region for the second time and the Spearman rank coefficient ρ3 of the samples Sam 普通 in the normal region: Then ρ2 is as shown in formula (10), that is ρ3 is as shown in formula (11), that is Its calculation method only needs to substitute the sample data Δrank in each region i for calculation.
[0057] In this way, we have obtained the weights and Spearman rank coefficients of the three regions of these 500 high-risk samples Sam 风险 respectively. We only need to perform weighted summation on them to obtain the final evaluation result. The comprehensive weighted summation ρ 总 is calculated as shown in formula (12), that is, ρ 总 =α1·ρ1 + α2·ρ2 + α3·ρ3 (12), where α3 = 1 - α1 - α2, and α1 > α2 > α3.
[0058] Example 2: Obviously, according to different actual situations, the high-risk samples Sam 风险 are divided into n regions as the case may be. By calculating the Spearman rank coefficient of each region and defining the weight of each region, the final weighted rank correlation evaluation result is calculated as shown in formula (13). Among them, ρ final is the final weighted rank correlation evaluation result, ρ i is the Spearman rank correlation coefficient of the samples in the i-th region, and α i is the weight of the i-th region, and it satisfies
[0059] For example, the high-risk samples Sam 风险 are divided into 4 regions as the case may be, namely the near-failure region, the sub-near-failure region, the extended region, and the normal region. Their weights are α1, α2, α3, and α4 respectively, and their Spearman rank coefficients are ρ1, ρ2, ρ3, and ρ4 respectively. Then the final weighted rank result ρ final =α1·ρ1 + α2·ρ2 + α3·ρ3 + α4·ρ3 (14), where α1 = 0.5, α2 = 0.2, α3 = 0.2, and α4 = 0.1.
[0060] Obviously, the value ratios and the minimum number of values in formulas (1) and (2) change according to the actual situation, and the value of Δrank i ' in formula (9) changes according to the actual situation.
[0061] Example 3: In actual applications, the characteristics of different samples are also different. In addition to determining the samples Sam 实际 in the near-failure region by setting the threshold Thr 接近 , we can also determine the samples Sam 接近 in the near-failure region by directly setting the number of samples Qua 接近 in the near-failure region. For example, set the thresholds S1 and S2.
[0062] Δd 平均 When ≤S1, the sample Sam is close to the failure area 接近 The number of Qua 接近 Satisfies formula (15):
[0063] Qua 接近 =max(Qua 总 / 20000,50) (15),
[0064] S1<Δd 平均 When ≤S2, the sample Sam is close to the failure area 接近 The number of Qua 接近 Satisfies formula (16):
[0065] Qua 接近 =max(Qua 总 / 5000,200) (16),
[0066] When Δd>S2, the sample Sam is close to the failure area 接近 The number of Qua 接近 Satisfies formula (17):
[0067] Qua 接近 =max(Qua 总 / 2000,500) (17),
[0068] Select simulation result YS 风险 Ranking from large to small Quad 接近 The samples are taken as the samples close to the failure area Sam 接近 , then the last sample Sam selected into the close failure area 接近 The simulation value is close to the failure area sample threshold Thr 接近 In this way, we determine the sample Sam close to the failure area 接近 The smaller the sorting error is, the smaller the failure area sample Sam is selected. 接近 The fewer the number of failure areas, and correspondingly, the larger the sorting error, the more failure area samples Sam is selected. 接近 The more.
[0069] Obviously, Thr2>Thr1, when Δd 平均 When ≤Thr1, it is considered that the ranking of the samples by simulation and prediction is very accurate and perfectly captures the tail failure area. Then we select a smaller proportion of samples as samples close to the failure area Sam 接近 It can better reflect the total sample set Sam 总 The minimum number is 50; and when Δd 平均>When Thr1, we select samples Sam close to the failure region 接近 The proportion should be larger, and the minimum number is 200; while when Δd 平均 >When Thr2, the minimum number is 500, that is, the processing results have a large gap and the probability of failure is large.
[0070] In this calculation method, Qua 次接近 = Qua 接近 - 3Δd 平均 (18), and successively determine the samples Sam close to the second failure region 次接近 and the samples Sam in the normal region 普通 .
[0071] Comparative example 1: As Figure 5 shown, it shows the necessity of setting the high-risk sample Sam 风险 in partitions. Suppose two prediction models A and B are used to predict the high-risk sample Sam 风险 , and two groups of data Y_PredictA and Y_PredictB are obtained. Among them, for the samples with higher rankings in Y_PredictA, the ranking values have a huge gap with the simulation value rankings, while for the sample data in other dangerous regions, the rankings are smaller than the simulation value rankings; while for the samples with higher rankings in Y_PredictB, the ranking values have a smaller gap with the simulation value rankings, or even the same, and then for the sample data in other dangerous regions, the gap with the simulation value rankings is larger. Through calculation, we get that the Spearman coefficient of Y_PredictA is 0.458, while the Spearman coefficient of Y_PredictB is -0.458. Obviously, 0.458 > -0.458. If only a single formula is used for calculation, the equivalent ranking accuracy of model A is higher than that of model B. However, in fact, we are more concerned about the ranking of samples close to the failure region, and the ranking in this region can best reflect the accuracy of the equivalent substitution model. Therefore, to avoid such situations, we divide the samples with higher rankings into samples Sam close to the failure region 接近 , assign higher weights to them, and then set more regions.
[0072] Comparative example 2: As Figure 6 shown, it shows the necessity of setting Δrank' when calculating the Spearman coefficient ρ1 of samples close to the failure region. Suppose two prediction models A and B are used to obtain the rankings of two groups of predicted values Y_PredictA and Y_PredictB. Among them, for the samples Sam 接近 close to the failure region in Y_PredictA, the predicted value rankings are from 100 to 91, while for the samples Sam 接近The predicted values are sorted from 1 to 10. If the proportional relationship between Δrank’ and Δrank is not set, assuming the value of α is 0.7, the Spearman coefficient rank ρ is calculated in two cases A = ρ B = 1.
[0073] However, the sorting effects of these two predictions are obviously not equivalent. Obviously, the prediction result B is better because when calculating the Spearman coefficient, we will calculate the Spearman coefficients of two regions, namely the region close to the failure area and other dangerous areas respectively. The sorting serial numbers within the region are relative and cannot reflect the sorting advantages and disadvantages of the whole. Therefore, we need to introduce a new index, the relative offset of sorting Δrank’. When the numerical offset Δd is less than a certain value Thr1, generally 20% of the number of samples in the region close to the failure area, we consider the sorting to be good and set the relative offset of sorting Δrank’ to 0 when calculating; when the numerical offset Δd is between Thr1 and Thr2, generally between 20% and 50% of the number of samples in the region close to the failure area, the relative offset of sorting Δrank’ is set to half of the original; in other cases, the sorting offset Δrank itself is taken. Then, through calculation, the prediction result is consistent with the actual analysis result.
[0074] The above is only the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A method for evaluating failure regions in high Sigma scenarios in electronic engineering, characterized in that: Including the following steps: (S1)Establish the initial sample Sam 初始 : Randomly select the initial sample Sam 总 from the total sample set Sam 初始 . The quantity Qua 初始 of the initial sample Sam 初始 satisfies the formula (1): Qua 初始 = max(Qua 总 / 1000, 1000) (1), where Qua 总 is the total sample quantity, and the quantity Qua 初始 of the initial sample is one-thousandth of the total sample quantity Qua 总 , and is not less than 1000; (S2) Establish a prediction model: For the selected initial sample Sam 初始 Perform simulation to obtain the simulation result YS 初始 , and use the simulation result YS 初始 Train a machine learning equivalent substitution model to generate a prediction model; (S3)Screen high-risk sample Sam 风险 : Use the prediction model to predict the remaining samples Sam 剩余 to obtain the prediction result YP 剩余 , where the remaining samples Sam 剩余 is the sample set after removing the initial samples Sam 总 from the total sample set Sam 初始 . Arrange the prediction results YP 剩余 of the remaining samples Sam 剩余 in descending order, and select high-risk samples Sam 风险 from largest to smallest, where the number Qua 风险 of the high-risk samples Sam 风险 satisfies the formula (2): Qua 风险 = max(Qua 总 / 2000, 500) (2), where Qua 总 is the total number of samples, and the number Qua 风险 of the high-risk samples Sam 风险 is five ten-thousandths of the total number of samples Qua 总 , and is not less than 500. The predicted value of the high-risk samples Sam 风险 is YP 风险 ; (S4) Simulate the high-risk sample Sam 风险 : Simulate the high-risk sample Sam 风险 to obtain the simulation results, and arrange the simulation results from largest to smallest to obtain the simulation result YS 风险 ; (S5) Calculate the numerical offset Δd and the average numerical offset Δd 平均 : For the simulation results YS arranged from largest to smallest 风险 Calculate the difference between two adjacent predicted values one by one, that is, Δd i = (YS 风险 ) i - (YS 风险 ) i+1 (3), where i is each natural number ≤ the number Qua of high-risk samples Sam 风险 minus 1; and the average numerical offset Δd of the entire risk sample Sam 风险 = 风险 平均 / (Qua 风险 minus 1) (4); (S6) Determine the sample Sam of the near-failure region 接近 and the sample Sam of the sub-near-failure region 次接近 : Set a simulation value threshold Thr 实际 Then, for the sample Sam of the near-failure region 接近 the threshold Thr 接近 = Thr 实际 - 3Δd 平均 (5), for the sample Sam of the sub-near-failure region 次接近 the threshold Thr 次接近 = Thr 实际 - 6Δd 平均 (6). The simulation value is greater than or equal to Thr 接近 The samples are used as the samples Sam in the near-failure region 接近 , and the simulation value is greater than or equal to the threshold Thr of the samples in the second near-failure region 次接近 and less than the threshold Thr of the samples in the near-failure region 接近 The samples are used as the samples Sam in the second near-failure region 次接近 , then the number of samples Sam in the near-failure region 接近 is Qua 接近 , and the number of samples Sam in the second near-failure region 次接近 is Qua 次接近 ; (S7) Determine the normal region sample Sam 普通 : High-risk sample Sam 风险 Remove the samples close to the failed region Sam 接近 and the samples close to the sub-failed region Sam 次接近 The remaining samples are the normal region samples Sam 普通 , then the number Qua 普通 of the normal region samples Sam 普通 satisfies formula (7), Qua 普通 = Qua 风险 - Qua 接近 - Qua 次接近 (7); (S8) Calculate the sorting offset Δrank: Compare each sample data in the high-risk sample Sam 风险 with the simulation result YS in step (S4) 风险 and its predicted result YP in step (S3) 风险 to obtain the sorting difference, that is, the sorting offset Δrank. The sorting offset Δrank satisfies formula (8): Δrank i = (8), where i is a natural number less than or equal to the number Qua of high - risk samples Sam 风险 of 风险 each, rank(YS 风险 ) is the ranking serial number of the simulation result YS 风险 , rank(YP 风险 ) is the ranking serial number of the prediction result YP 风险 ; (S9) Calculate the Spearman rank coefficient ρ1 of the samples Sam close to the failure region: As shown in Equation (9), 接近 That is , (9), where Δrank in the formula i ’ takes values as follows: , where Thr1 and Thr2 are respectively the first-level threshold Thr1 of the samples in the near-failure region and the second-level threshold Thr2 of the samples in the near-failure region set artificially; (S10) Calculate the Spearman rank coefficient ρ2 of the sub-closest-to-failure region sample Sam 次接近 and the Spearman rank coefficient ρ3 of the normal region sample Sam 普通 : Then ρ2 is as shown in formula (10), that is (10), and ρ3 is as shown in formula (11), that is (11); (S11)Comprehensive weighted summation ρ 总 : Define the weight of samples in the near-failure region as α1, the weight of samples in the sub-near-failure region as α2, and the weight of samples in the normal region as α3. Then the calculation formula of ρ 总 is as shown in formula (12). That is (12), where α3 = 1 - α1 - α2, and α1 > α2 > α3.
2. The failure area evaluation method applied to the high Sigma scenario in electronic engineering according to claim 1, wherein: 0.6≤α1<0.7,0.2≤α2<0.3。 3. A method for evaluating a failure area applied to a high Sigma scenario in electronic engineering according to claim 1, characterized in that: High-risk sample Sam 风险 It is divided into n regions according to the situation. By calculating the Spearman rank coefficient of each region and defining the weight of each region, the final weighted rank correlation evaluation result is calculated. Its calculation is as shown in formula (13). (13), where ρ final is the final weighted rank correlation evaluation result, and ρ i is the Spearman rank correlation coefficient of the samples in the i-th region, and α i is the weight of the i-th region, which is obtained by artificial definition and satisfies , n ≥ 2.
4. A method for evaluating a failure region applied to a high Sigma scenario in electronic engineering according to claim 3, characterized in that: Divide the high-risk sample Sam 风险 into 4 regions according to the situation, namely the near-failure region, the sub-near-failure region, the extended region, and the normal region, with their weights being α1, α2, α3, and α4 respectively, and their Spearman rank coefficients being ρ 1、 ρ 2、 ρ3 and ρ4, then the final weighted rank result (14), where α1 = 0.5, α2 = 0.2, α3 = 0.2, α4 = 0.
1.
5. A method for evaluating a failure area in a high Sigma scenario applied in electronic engineering according to claim 1, characterized in that: The simulation tool uses the Spice circuit simulation tool.
6. The failure area evaluation method applied to high Sigma scenarios in electronic engineering according to claim 1, characterized in that: Use the FFX algorithm, linear regression, Lasso regression, or LightGBM surrogate model to establish a prediction model.
7. The failure area evaluation method applied to high Sigma scenarios in electronic engineering according to claim 1, characterized in that: The value ratios and minimum value numbers in formulas (1) and (2) change according to the actual situation.
8. A method for evaluating a failure area in a high Sigma scenario applied in electronic engineering according to claim 1, characterized in that: Δrank in formula (9) i ’ The value changes according to the actual situation.
Citation Information
Patent Citations
Method and system of circuit yield analysis for evaluating rare failure events
CN108694273A
Extra-long tunnel construction site risk assessment method and system
CN118246744A