A method for screening processor performance
By using the k-means clustering algorithm in processor performance screening, the reference performance value is automatically selected, which solves the problem of frequent manual setting of benchmark performance in the prior art, and improves the automation efficiency and accuracy of the screen.
Patent Information
- Application Number
- CN202110381442.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-04-09
AI Technical Summary
When running diversified tasks in a diverse environment, the existing technology requires frequent manual setting of benchmark performance, resulting in interruption of false screening and sieve processes, and the operation of restoring the sieve of the sieve processor is complicated.
The k-means clustering algorithm is used to cluster the processor according to the performance value, select the performance average value of the category with the largest number of samples as the benchmark performance value, and broadcast the reference value to all processors through set communication to filter out processors whose performance exceeds or is below the threshold.
The automation of benchmark performance selection is realized, which reduces human intervention, improves the automation efficiency and accuracy of screens, and reduces the possibility of misoperation.
Smart Images

Figure CN114253705B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for screening processor performance, belonging to the technical field of chip screening. Background Art
[0002] Before a processor is submitted to users for use, it needs to go through a set of chip screening processes to eliminate chips with instability, abnormal results, and low performance. Performance screening is an important part of it. Chip screening is usually highly automated. To ensure that processors with occasional errors are screened out to the greatest extent, multiple rounds of chip screening processes need to be repeatedly executed. When a processor is eliminated due to low performance, its chip screening process is interrupted and it will no longer participate in subsequent chip screening cycles.
[0003] Performance screening is an important step in the processor screening process, and the selection of performance benchmarks greatly affects the screening results. When individual processors have abnormal configurations or timing anomalies, it may lead to extremely high performance, and then lead to inappropriate selection of performance benchmarks, resulting in a large number of normal processors being eliminated, greatly affecting the chip screening efficiency, and bringing a heavy workload to chip screening personnel.
[0004] There are usually two methods for selecting performance benchmarks in current performance screening. One is to pre-give the performance that should be obtained when running a topic, and the other is to select the maximum or average value of a batch of performance results as the performance benchmark each time. The main disadvantage of the first method is the lack of flexibility. Whenever the processor configuration, task, and running parameters change, the benchmark performance will change, and it is necessary to regenerate the benchmark performance again. When running multiple different tasks with multiple different running parameters under multiple processor configurations, the workload increases sharply, the process is difficult to automate, the screening process is complex and changeable, and it is easy to introduce human misoperations. The main disadvantage of the second method is the lack of robustness. This method assumes that the processor configurations are completely consistent and the processor performance calculations are correct. However, when individual processors have timing anomalies or configuration anomalies, it may cause the performance of individual processors to be extremely high or extremely low, resulting in inappropriate benchmark selection, causing the performance of all normal points in a batch of tests to deviate significantly from the benchmark, and causing a large number of normal processors to be screened out; this process can only be intervened manually, manually eliminating processors with extremely high or low performance, putting the mis-screened processors back into the chip screening process, modifying the chip screening database, etc., interrupting the automated process, affecting the chip screening progress and the automated efficiency of chip screening, increasing a heavy burden on chip screening personnel, and the work is complicated and error-prone. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for screening processor performance to solve the problems of frequent manual setting of benchmark performance in a diverse environment when running diverse topics, and problems such as mis-screening caused by processors with extremely high performance, interruption of the chip screening process, and complex operations for restoring the screening of mis-screened processors.
[0006] To achieve the above object, the technical solution adopted by the present invention is: to provide a method for screening processor performance, including the following steps:
[0007] S1. Divide the processors to be screened into n subsets, each subset contains m processors, and the processors in subset i (0 ≤ i ≤ n - 1) are represented as Pij, where 0 ≤ j ≤ m - 1;
[0008] S2. Record the start time of performance screening, the total time of performance screening up to now, and the current screening round. Determine whether the performance screening is completed according to whether the total time of performance screening or the total number of rounds is greater than or equal to the preset requirement. If completed, end the performance screening; otherwise, execute S3;
[0009] S3. Simultaneously execute the tasks for performance screening on n subsets. Each processor records the time for executing the performance screening task and calculates its own task performance, and denote the performance of processor Pij as Aij;
[0010] S4. Designate a certain processor r in subset i as the root node, and execute the gather operation in collective communication to collect the performance values of all processors in subset i to the root node r as the samples to be classified;
[0011] S5. Select the number of classes k, set the number of classes of k - means to k, set the classification termination condition to k_thres according to the classification accuracy requirement and the convergence speed requirement, and sort Aij;
[0012] If k = 2, select the minimum value Amin and the maximum value Amax as the initial centroids of the two classes;
[0013] If k ≥ 3, divide the sorted Aij into k - 2 segments, and select the minimum value Amin, the maximum value Amax, and the average value Aavg[x] of each of the k - 2 segments as the initial centroids of the k classes;
[0014] Denote the initial centroids as Centroid_update[t], where 0 ≤ x ≤ k - 3, 0 ≤ t ≤ k - 1, use num[t] to represent the current number of samples in the t - th class, and use S[t] to represent the samples included in the t - th class;
[0015] S6. Let Centroid[t]=Centroid_update[t], initialize num[t] to 0, and initialize S[t] to an empty set;
[0016] S7. For each sample Aij, calculate its Euclidean distance to each initial centroid, assign it to the class t where the nearest centroid Centroid[t] is located, and record it in S[t];
[0017] S8. For each category t, recalculate the mean of the samples S[t] belonging to this category and use it as the new centroid of category t, denoted as Centroid_update[t].
[0018] S9. For each category t, calculate the Euclidean distance by which the centroid moves when Centroid[t] is updated to Centroid_update[t]. If the distance by which any category centroid moves exceeds k_thres, then go back to S6.
[0019] S10. Update Centroid[t] to Centroid_update[t]. Select the centroid of the category with the largest num[t] among the k categories as the baseline performance value Pbase. If the number of samples in several categories is the same and the largest, then select the centroid with the larger value as the baseline performance value Pbase.
[0020] S11. The root node r uses the broadcast operation in collective communication to send the baseline performance value to each processor.
[0021] S12. According to the preset performance screening threshold P_thres, calculate the upper limit Pceil and lower limit Pfloor of the normal performance interval. If the performance of processor Pij exceeds Pceil, then output the FAST flag and record the processor number.
[0022] If the performance of processor Pij is lower than Pfloor, then output the SLOW flag and record the processor number, and remove Pij as a processor with abnormal performance.
[0023] Otherwise, there is no output.
[0024] S13. Use the remaining processors after this round of screening as the processors to be screened, accumulate the total screening time or the number of screening rounds, and go back to S1.
[0025] The further improved solutions in the above technical solutions are as follows:
[0026] 1. In the above solution, in S3, calculate its own task performance by the method of "task computation amount divided by task running time" or directly use "the reciprocal of task running time".
[0027] 2. In the above solution, in S5, select the number of clusters k according to prior knowledge, actual requirements, the elbow method, or classification effect.
[0028] 3. In the above solution, in S12, calculate the upper limit Pceil and lower limit Pfloor of the normal performance interval according to the formulas Pceil = Pbase * (1 + P_thres) and Pfloor = Pbase * (1 - P_thres).
[0029] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0030] A method for screening processor performance according to the present invention has strong flexibility and good robustness, significantly reduces human intervention, facilitates the automation of chip screening, reduces the work burden and possible misoperations of chip screening personnel, and helps to improve the efficiency and effect of chip screening. Description of the Drawings
[0031] Appendix Figure 1 It is a schematic diagram of a method for screening processor performance according to the present invention. Detailed Embodiment
[0032] Embodiment: The present invention provides a method for screening processor performance, which specifically includes the following steps:
[0033] S1. Divide the processors to be screened into n subsets, each subset contains m processors, and the processors in subset i (0 ≤ i ≤ n - 1) are represented as Pij, where 0 ≤ j ≤ m - 1;
[0034] S2. Record the start time of performance screening, the total time of performance screening up to now, and the current screening round. Determine whether the performance screening is completed according to whether the total time of performance screening or the total number of rounds is greater than or equal to the preset requirement. If it is completed, end the performance screening; otherwise, execute S3;
[0035] S3. Simultaneously execute tasks for performance screening on n subsets. Each processor records the time for executing the performance screening task, and calculates its own task performance by the method of "task calculation amount divided by task running time" or directly uses "the reciprocal of task running time", and record the performance of processor Pij as Aij;
[0036] S4. Designate a certain processor r in subset i as the root node, and execute the collect operation in collective communication to collect the performance values of all processors in subset i to the root node r as the samples to be classified;
[0037] S5. Select the number of classes k according to prior knowledge, actual requirements, elbow method or classification effect. Set the number of classes of k-means to k, set the classification termination condition to k_thres according to the classification accuracy requirement and convergence speed requirement, and sort Aij;
[0038] If k = 2, select the minimum value Amin and the maximum value Amax as the initial centroids of the two classes;
[0039] If k ≥ 3, divide the sorted Aij into k - 2 segments, and select the minimum value Amin, the maximum value Amax, and the average value Aavg[x] of each segment of the k - 2 segments as the initial centroids of the k classes;
[0040] Denote the initial centroid as Centroid_update[t], where 0 ≤ x ≤ k - 3 and 0 ≤ t ≤ k - 1. Use num[t] to represent the current number of samples in the t-th category, and use S[t] to represent the samples included in the t-th category;
[0041] S6. Let Centroid[t] = Centroid_update[t], initialize num[t] to 0, and initialize S[t] to an empty set;
[0042] S7. For each sample Aij, calculate its Euclidean distance to each initial centroid, assign it to the category t where the centroid Centroid[t] with the closest distance is located, and record it in S[t];
[0043] S8. For each category t, recalculate the average value of the samples S[t] belonging to this category, and use it as the new centroid of category t, denoted as Centroid_update[t];
[0044] S9. For each category t, calculate the Euclidean distance by which the centroid moves when Centroid[t] is updated to Centroid_update[t]. If the moving distance of any category centroid exceeds k_thres, then go back to S6;
[0045] S10. Update Centroid[t] to Centroid_update[t], select the centroid of the category with the largest num[t] among the k categories as the benchmark performance value Pbase. If the number of samples in several categories is the same and the largest, then select the one with the larger centroid as the benchmark performance value Pbase;
[0046] S11. The root node r uses the broadcast operation in set communication to send the benchmark performance value to each processor;
[0047] S12. According to the preset performance screening threshold P_thres, calculate the upper limit Pceil and lower limit Pfloor of the normal performance interval according to the formulas Pceil = Pbase * (1 + P_thres) and Pfloor = Pbase * (1 - P_thres). If the performance of processor Pij exceeds Pceil, then output the FAST flag and record the processor number to remind the sieve personnel to pay attention to inspection;
[0048] If the performance of processor Pij is lower than Pfloor, then output the SLOW flag and record the processor number, and remove Pij as a processor with abnormal performance;
[0049] Otherwise, there is no output;
[0050] S13. Take the remaining processors after this round of screening as the processors to be screened, accumulate the total screening time or the number of screening rounds, and return to S1.
[0051] The further explanations of the above embodiments are as follows:
[0052] After the processors are produced, they need to be screened before being submitted to users for use to ensure that users can run stably on different processors and obtain correct results and consistent performance. Performance screening is an important step among them. In performance screening, the determination of the benchmark performance value is crucial, which directly affects the screening results.
[0053] If the method of determining the benchmark by pre-executing tasks on trusted processors and using the performance value of this execution as the benchmark performance is adopted, it will face the following problems:
[0054] 1) To determine the trusted processors for specific tasks, developers need to conduct repeated experiments, select processors, and count the results, which is very time-consuming;
[0055] 2) Whenever changes that may affect performance such as processor configuration, task type, and running parameters occur, it is necessary to regenerate the benchmark performance. This makes the workload increase sharply when running multiple different tasks with multiple different running parameters under multiple processor configurations, and the process is difficult to automate;
[0056] 3) The screening personnel need to properly switch the benchmark performance according to the current environmental parameters, which is easy to introduce human errors and cause mis-screening, affecting the screening progress and effect.
[0057] Another approach is to use the highest value or average value of the performance among the processors screened in the same batch as the benchmark performance. If individual processors have abnormally high or low performance due to abnormal configuration, abnormal timing, or serious problems with their own performance, there may be cases where the performance of some processors is extremely high or low, resulting in improper selection of the benchmark and causing a large number of normal processors to be screened out, affecting the screening progress. If it is necessary to resume the screening process of the mis-screened processors, manual intervention is required: manually removing the processors with extremely high or low performance, putting the mis-screened processors back into the screening process, modifying the screening database, etc., which interrupts the automated process, and the work is complicated and error-prone.
[0058] To solve the above problems, this patent introduces the k-means clustering algorithm in chip performance screening. This algorithm is based on unsupervised learning and has good flexibility. It obtains all processor performance values through the collect operation in the set operation, clusters the processors into multiple categories according to the performance values through the k-means clustering algorithm, takes the average performance of the category with the largest number of processors as the benchmark for performance screening, and uses the broadcast operation in the set operation to send the benchmark performance value to all processors, and screens out the processors whose performance exceeds the upper threshold and is lower than the lower threshold;
[0059] On the one hand, it realizes the automation of benchmark selection, solves the flexibility problem of performance screening. On the other hand, points with extremely high and extremely low individual performances will be grouped into separate categories respectively, reminding the sieve personnel to pay attention to checking the configuration. The remaining processors will be grouped into the category with the largest number of samples, and their average value will be used as the screening benchmark. Points with performance below the benchmark by more than the threshold range will be excluded, avoiding the influence of abnormal configurations and timing anomalies of individual processors on the selection of performance benchmarks, and solving the robustness problem of performance screening, thereby improving the efficiency and effect of sieving.
[0060] The specific operations are as follows:
[0061] 1. Divide the processors to be screened into n subsets of appropriate sizes, with each subset containing m processors. The processors in subset i (0 ≤ i ≤ n - 1) are denoted as Pij, where 0 ≤ j ≤ m - 1.
[0062] 2. Determine whether the total time or the total number of rounds of performance screening reaches the preset requirements. If it does, end the performance screening; otherwise, execute step 3.
[0063] 3. Simultaneously execute the tasks for performance screening on the n subsets. Each processor records the time for executing the task and calculates its own task performance. The performance of Pij is denoted as Aij.
[0064] 4. Designate a certain processor r in subset i as the root node, and execute the gather operation in collective communication to collect the performances of all processors within the subset to r. These performance values are the samples to be classified.
[0065] 5. Select the number of clusters k according to prior knowledge, actual requirements, the "elbow method", or classification effect, and set the number of clusters of k-means to k; set the classification termination condition to k_thres according to the classification accuracy requirement and convergence speed requirement. Taking k = 3 and k_thres = 0.0 as an example, select the minimum value Amin, average value Aavg, and maximum value Amax from Aij as the initial centroids of the 3 clusters, denoted as Centroid_update[t], where 0 ≤ t ≤ k - 1. Use num[t] to represent the current number of samples in the t-th cluster, and use S[t] to represent the samples included in the t-th cluster.
[0066] 6. Let Centroid[t] = Centroid_update[t], initialize num[t] to 0, and initialize S[t] to an empty set.
[0067] 7. For each sample Aij, calculate its Euclidean distance to each centroid, and assign it to the category t of the centroid Centroid[t] with the closest distance, and record it in S[t].
[0068] 8. For each category t, recalculate the mean of the samples S[t] belonging to that category and use it as the new centroid of category t, denoted as Centroid_update[t].
[0069] 9. For each category t, calculate the Euclidean distance by which the centroid moves from Centroid[t] to Centroid_update[t]. If the distance by which any category centroid moves exceeds k_thres, go back to step 6.
[0070] 10. Update Centroid[t] to Centroid_update[t], and select the centroid of the category with the largest num[t] among the k categories as the benchmark performance value. If the number of samples in several categories is the same and the largest, select the one with the larger centroid as the benchmark performance value Pbase.
[0071] 11. The root node r uses the broadcast operation in collective communication to send the benchmark performance value to each processor.
[0072] 12. According to the preset performance screening threshold P_thres, calculate the upper limit Pceil and lower limit Pfloor of the normal performance range. If the performance of processor Pij exceeds Pceil, output the FAST flag and record the processor number to remind the screening personnel to pay attention to the inspection; if the performance of processor Pij is lower than Pfloor, output the SLOW flag and record the processor number, and remove Pij as a processor with abnormal performance; otherwise, there is no output.
[0073] 13. Take the remaining processors after this round of screening as the processors to be screened, accumulate the total screening time or the number of screening rounds, and go back to step 1.
[0074] When adopting the above-mentioned processor performance screening method, it has strong flexibility and good robustness, significantly reduces human intervention, facilitates screening automation, reduces the work burden and possible misoperations of the screening personnel, and helps to improve the screening efficiency and effect.
[0075] To facilitate a better understanding of the present invention, the terms used in this article will be briefly explained below:
[0076] Processor performance: The ratio of the workload of a processor to complete a specific task to the processing time is called the processor performance.
[0077] Performance consistency: The performance of processors will not be exactly the same. When the performance fluctuations of a batch of processors meet a given threshold range, these processors are said to be performance-consistent.
[0078] Performance screening: A processor can have multiple processor performance values for different application scenarios. However, for a specific task, the processor performance should be consistent. The purpose of performance screening is to filter out processors with low performance and provide users with processors with normal performance.
[0079] Processor configuration: Specifically refers to the configuration that affects the performance of the processor in this patent, such as the processor main frequency, network bandwidth, memory access frequency, etc.
[0080] Sieve: Processor screening.
[0081] Collective communication: Communication operations performed on a set of processes.
[0082] Gather: A type of collective communication that collects the data of each process to the root node.
[0083] Bcast: A type of collective communication that distributes the data of the root node to each process.
[0084] Centroid: The center of mass, which refers to the sample mean of a category in the k-means clustering algorithm.
[0085] Euclidean distance: The straight-line distance between two points in Euclidean space. For point A(x 1 ,x 2 ,…,x n ) and point B(y 1 ,y 2 ,…,y n ), the formula for calculating the Euclidean distance between the two points is .
[0086] The above embodiments are only used to illustrate the technical concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. However, the protection scope of the present invention cannot be limited thereby. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for screening processor performance, characterized in that, it includes the following steps: S1. Divide the processors to be screened into n subsets, each subset contains m processors, and the processors in subset i (0 ≤ i ≤ n - 1) are represented as Pij, where 0 ≤ j ≤ m - 1; S2. Record the start time of performance screening, the total time of performance screening up to now, and the current screening round. Determine whether the performance screening is completed according to whether the total time of performance screening or the total number of rounds is greater than or equal to the preset requirement. If completed, end the performance screening; otherwise, execute S3; S3. Simultaneously execute tasks for performance screening on n subsets. Each processor records the time for executing the performance screening task and calculates its own task performance. Denote the performance of processor Pij as Aij; S4. Designate a certain processor r in subset i as the root node, and execute the gather operation in collective communication to gather the performance values of all processors within subset i to the root node r as the samples to be classified; S5. Select the number of classes k, set the number of classes of k-means to k, set the classification termination condition to k_thres according to the classification accuracy requirement and the convergence speed requirement, and sort Aij; If k = 2, select the minimum value Amin and the maximum value Amax as the initial centroids of the two classes; If k ≥ 3, divide the sorted Aij into k - 2 segments, and select the minimum value Amin, the maximum value Amax, and the average value Aavg[x] of each of the k - 2 segments as the initial centroids of the k classes; Denote the initial centroids as Centroid_update[t], where 0 ≤ x ≤ k - 3, 0 ≤ t ≤ k - 1. Use num[t] to represent the current number of samples in the t-th class, and use S[t] to represent the samples included in the t-th class; S6. Let Centroid[t] = Centroid_update[t], initialize num[t] to 0, and initialize S[t] to an empty set; S7. For each sample Aij, calculate its Euclidean distance to each initial centroid, and assign it to the class t where the nearest centroid Centroid[t] is located, and record it in S[t]; S8. For each class t, recalculate the average value of the samples S[t] belonging to this class and use it as the new centroid of class t, denoted as Centroid_update[t]; S9. For each class t, calculate the Euclidean distance by which the centroid moves from Centroid[t] to Centroid_update[t]. If the moving distance of any class centroid exceeds k_thres, then go back to S6; S10. Update Centroid[t] to Centroid_update[t], and select the centroid of the class with the largest num[t] from the k classes as the reference performance value Pbase. If the number of samples in several classes is the same and the largest, then select the centroid with the larger value as the reference performance value Pbase; S11. The root node r uses the broadcast operation in collective communication to send the reference performance value to each processor; S12. Calculate the upper limit Pceil and the lower limit Pfloor of the normal performance range according to the preset performance screening threshold P_thres. If the performance of the processor Pij exceeds Pceil, output the FAST flag and record the processor number; If the performance of the processor Pij is lower than Pfloor, output the SLOW flag and record the processor number, and remove Pij as a processor with abnormal performance; Otherwise, there is no output; S13. Use the remaining processors after this round of screening as the processors to be screened, accumulate the total screening time or the number of screening rounds, and return to S1.
2. A method for screening the performance of a processor according to claim 1, characterized in that: In S3, calculate its own task performance by the method of "task computation amount divided by task running time" or directly use "the reciprocal of task running time".
3. A method for screening the performance of a processor according to claim 1, characterized in that: In S5, select the number of clusters k according to prior knowledge, actual requirements, the elbow method or classification effect.
4. A method for screening the performance of a processor according to claim 1, characterized in that: In S12, calculate the upper limit Pceil and the lower limit Pfloor of the normal performance range according to the formulas Pceil = Pbase * (1 + P_thres) and Pfloor = Pbase * (1 - P_thres).
Citation Information
Patent Citations
Improved K-means clustering algorithm based on density radius
CN108549913A
Feature selection method based on rough set and swarm intelligence
CN108875895A