Compiler evaluation method and device, electronic equipment and readable medium
Through the improvement of the compiler performance evaluation method, the test subset of acceleration ratio matching is screened out using clustering technology, which solves the limitations and low flexibility of compiler performance evaluation in the existing technology, and achieves a more efficient and accurate evaluation effect.
Patent Information
- Application Number
- CN202510072245.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
The existing compiler performance evaluation methods are highly limited and have poor flexibility, and cannot adapt to compiler performance evaluation requirements under different optimization levels, optimization options and compilation settings.
By determining performance-related features for pre-collected original assemblies, clustering raw and default test programs, filtering out test subsets of acceleration ratio matching for performance evaluation of compilers.
It realizes more flexible and accurate compiler performance evaluation, reduces resource usage and evaluation time, and improves compiler optimization iteration efficiency.
Smart Images

Figure CN120066516A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of compilation technology, and particularly to a compiler evaluation method, device, electronic device, and readable medium. Background Art
[0002] Currently, with the continuous development of compilation technology, compilers have been more and more widely used. A compiler can compile source code into executable code that can run. In order to ensure the normal operation of the compiler, it is often necessary to evaluate the performance of the compiler.
[0003] In the related art, the default program set provided by the evaluation tool is used to evaluate the performance of the compiler. Since only the default program set provided by the evaluation tool can be used for testing, the limitation is high and the flexibility is poor. Summary of the Invention
[0004] Embodiments of the present invention provide a compiler evaluation method, device, electronic device, and readable medium, which can solve the problems of high limitation and poor flexibility.
[0005] To solve the above problems, embodiments of the present invention disclose a compiler evaluation method, the method comprising:
[0006] Determining performance-related features for a pre-collected original program set as first related features;
[0007] Performing a clustering operation on the original test programs in the original program set and the default test programs based on the feature values of the first related features corresponding to the original test programs in the original program set and the feature values of the first related features corresponding to the default test programs in the default program set;
[0008] Screening the original test programs matching each of the default test programs based on the clustering result to form a test subset;
[0009] When the speedup ratio of the test subset matches the speedup ratio of the default program set, evaluating the performance of the compiler based on the test subset.
[0010] On the other hand, embodiments of the present invention disclose a compiler evaluation device, the device comprising:
[0011] A first determination module, configured to determine performance-related features for a pre-collected original program set as first related features;
[0012] A clustering module, configured to perform a clustering operation on the original test programs in the original program set and the default test programs based on the feature values of the first related features corresponding to the original test programs in the original program set and the feature values of the first related features corresponding to the default test programs in the default program set;
[0013] A screening module, configured to screen out the original test programs matched by each of the default test programs based on the clustering result, and form a test subset.
[0014] An evaluation module, configured to perform performance evaluation on the compiler based on the test subset when the speedup ratio of the test subset matches the speedup ratio of the default program set.
[0015] In another aspect, an embodiment of the present invention discloses an electronic device, including: a processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the foregoing method.
[0016] An embodiment of the present invention also discloses a machine-readable medium, on which instructions are stored, and when executed by one or more processors, cause the processors to execute the method as described above.
[0017] The embodiments of the present invention have the following advantages: In the compiler evaluation method provided by the embodiments of the present invention, performance-related features are determined for a pre-collected original program set as the first related features. Based on the feature values of the first related features corresponding to the original test programs in the original program set and the feature values of the first related features corresponding to the default test programs in the default program set, clustering operations are performed on the original test programs and the default test programs in the original program set. The original test programs matched by each default test program are screened out based on the clustering result to form a test subset. When the speedup ratio of the test subset matches the speedup ratio of the default program set, performance evaluation is performed on the compiler based on the test subset. In this way, it is possible to perform performance evaluation on the compiler using a test subset screened from a pre-collected original program set, reducing the limitations of performance evaluation and improving the flexibility of performance evaluation. At the same time, by further screening a test subset whose speedup ratio matches the speedup ratio of the default program set from the original test program set, it can be ensured that the screened test subset can more accurately represent the default program set, and thus ensure the evaluation effect when using the test subset to replace the default program set for performance evaluation. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0019] Figure 1It is a flowchart of the steps of a compiler evaluation method provided by an embodiment of the present invention;
[0020] Figure 2 It is a schematic diagram of a processing flow provided by an embodiment of the present invention;
[0021] Figure 3 It is a block diagram of a compiler evaluation device provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] Figure 1 It is a flowchart of the steps of a compiler evaluation method provided by an embodiment of the present invention. As Figure 1 shown, the compiler evaluation method may include the following steps:
[0025] Step 101: Determine performance-related features for a pre-collected original program set as the first related features.
[0026] Step 102: Perform a clustering operation on the original test programs in the original program set and the default test programs based on the feature values of the first related features corresponding to the original test programs in the original program set and the feature values of the first related features corresponding to the default test programs in the default program set.
[0027] Step 103: Screen the original test programs matching each of the default test programs based on the clustering result to form a test subset.
[0028] Step 104: When the speedup ratio of the test subset matches the speedup ratio of the default program set, perform performance evaluation on the compiler based on the test subset.
[0029] Among them, the original program set may include multiple pre-collected original test programs. The performance-related features are features that affect program performance, and the first related features are performance-related features collected during the running process of the original program set. The performance-related features may specifically be determined based on performance events collected during the running process.
[0030] The default assembly can be a test assembly provided by the evaluation tool. Based on the original test program and the eigenvalue of each corresponding first related feature, clustering operations are performed on the original test program and the default test program in the original assembly, so that the original test program and the default test program in the original assembly can be aggregated into the same class, facilitating the screening of the original test program matched by the default test program. Further, the original test programs matched by all default test programs can be combined into a set to obtain a test subset. The speedup ratio of the test subset matches that of the default assembly, which can mean that the change trends of their speedup ratios match, and the difference between their speedup ratios is within a preset difference range. Correspondingly, if the speedup ratio of the test subset matches that of the default assembly, it can be considered that the change in the speedup ratio performance of the test subset is consistent with that of the default assembly, and the test subset can be used to represent the default assembly for performance evaluation.
[0031] In the embodiments of the present invention, the code size of the test subset is smaller than that of the default assembly, and the number of test programs included in the test subset is not greater than the number of test programs included in the default assembly. Specifically, the original test program can be a standardized test sample collected in advance for performance evaluation of the compiler. The original test program can be collected manually in advance or automatically obtained from the preset code based on a preset script. Exemplarily, compute-intensive benchmark test samples can be obtained from the preset code and added to the original assembly. The original assembly can include original test programs of different types (such as integer, floating-point, scalar, and vector, etc.), algorithms, and data structures to evaluate the performance of the compiler under different circumstances.
[0032] In an actual scenario, the default assemblies in the related art are mainly used to evaluate the performance of a central processing unit (CPU) and a memory system, rather than being a test suite specifically for evaluating compiler performance. Exemplarily, the evaluation tool in the related art can be the SPEC tool. In this application scenario, the default assembly can be the SPEC test suite provided by the SPEC tool, and this default assembly can also be referred to as the SPEC component. Although the accuracy of using the default assembly for performance evaluation is relatively good, the number of source codes of the test programs in the default assembly is large, the scale is large, and the complexity is high, which is not suitable for evaluating compiler performance. Therefore, when using the default assembly for performance evaluation, the resource occupancy is too large, and the compilation and running time is too long, which seriously hinders the optimization and iteration efficiency of the compiler. Moreover, in the actual application scenario, there are various combinations among compiler types, compiler versions, optimization levels, optimization options (i.e., compilation options), and compilation settings. It is often necessary to evaluate the performance of different versions and different types of compilers separately under different optimization levels, different optimization options, and different compilation settings. The method of using the default assembly for performance evaluation in the related art cannot be applied to this application scenario.
[0033] In the embodiments of the present invention, by pre-collecting standardized test cases for evaluating compiler performance as the original test program, the test program can be made more suitable for evaluating compiler performance. The code scale of the standardized test cases is smaller. Therefore, it can be ensured that the code scale of the test subset is smaller than that of the default assembly. Using the test subset screened from the original test program for performance evaluation can avoid the problems of large resource occupancy and too long compilation and running time. Furthermore, it can reduce the time cost and improve the optimization and iteration efficiency of the compiler, enabling the use of the test subset to evaluate the performance of different versions and different types of compilers under different optimization levels, different optimization options, and different compilation settings.
[0034] Further, when evaluating the performance of a compiler, the compiler can be used to compile the test programs in the test subset, and the compilation duration can be determined, and also, the running duration of the executable code obtained by the test compilation can be determined. Among them, the compilation duration refers to the time required for the entire process from when the compiler starts processing the source code to generating the executable file, and the running duration refers to the time required for the compiled program to execute during runtime. The compilation duration and the running duration can be used as performance evaluation parameters (i.e., performance data) for the compiler performance. If the compilation duration and the running duration are longer, it can be determined that the quality and execution efficiency of the compiled code are lower, the compilation and running efficiency are lower, and the performance is worse. On the contrary, if the compilation duration and the running duration are shorter, it can be determined that the quality and execution efficiency of the compiled code are higher, the compilation and running efficiency are higher, and the performance is better. It should be noted that a performance analysis tool, such as the Perf tool, can also be used to evaluate the performance of the executable code obtained by compilation during runtime. The embodiments of the present invention do not limit this.
[0035] Specifically, under different optimization levels, different optimization options, and different compilation settings, different versions and different types of compilers can be used to compile the test programs in the test subset, and then the obtained executable files can be run, so as to obtain the performance evaluation parameters of different versions and different types of compilers for the same test program under different optimization levels, the performance evaluation parameters of different optimization options for the same test program, and the performance evaluation parameters of different compilation settings for the same test program, thereby facilitating the comparison of the compilation and running efficiency of different versions and different types of compilers under different optimization levels, different optimization options, and different compilation settings, that is, facilitating the comparison of the performance of different versions and different types of compilers under different optimization levels, different optimization options, and different compilation settings.
[0036] Further, by evaluating the performance of the compiler, it is convenient to optimize and iterate the compiler. For example, the performance of the compiler can be evaluated first. If the compilation duration and the running duration are greater than a preset threshold, the compiler can be optimized. Then, the performance evaluation is performed again. If the compilation duration and the running duration of the optimized compiler are still greater than the preset threshold, the compiler can be continuously optimized until the compilation duration and the running duration of the optimized compiler are not greater than the preset threshold. Correspondingly, the higher the efficiency of the performance evaluation, the more beneficial it is to the optimization and iteration of the compiler.
[0037] In summary, in the compiler evaluation method provided by the embodiments of the present invention, performance-related features are determined for the pre-collected original program set as the first related features. Based on the feature values of the first related features corresponding to the original test programs in the original program set and the feature values of the first related features corresponding to the default test programs in the default program set, clustering operations are performed on the original test programs and the default test programs in the original program set. The original test programs matching each default test program are screened based on the clustering results to form a test subset. When the speedup ratio of the test subset matches the speedup ratio of the default program set, the performance of the compiler is evaluated based on the test subset. In this way, it is possible to use the test subset screened from the pre-collected original program set to evaluate the performance of the compiler, reducing the limitations of performance evaluation and improving the flexibility of performance evaluation. At the same time, by further screening the test subset with a speedup ratio matching that of the default program set from the original test program set, it can be ensured that the screened test subset can more accurately represent the default program set, thereby ensuring the evaluation effect when using the test subset to replace the default program set for performance evaluation.
[0038] Optionally, in the embodiments of the present invention, the step of determining performance-related features for the pre-collected original program set as the first related features may specifically include:
[0039] Step 1011: Obtain the performance events generated during the running of the original test program as the first performance events.
[0040] Step 1012: Determine the performance-related features possessed by the original program set based on the first performance events as the original related features.
[0041] Step 1013: Determine the importance scores of the original related features, and determine the original related features with importance scores not less than the preset score threshold as the first related features.
[0042] Among them, the number of the first performance events collected can be related to the collection duration. Exemplarily, the longer the collection duration, the more first performance events tend to be collected. In the embodiments of the present invention, the pre-specified collection duration can be read, and then the original test program is run until the running duration reaches the collection duration. Finally, the number of performance events counted for the original test program in the performance event counter can be read. Among them, the performance events can also be called program performance characteristic events, and the performance event counter will count the number of performance events that occur during the running of the original test program. The performance-related characteristics can also be called program characteristics. Alternatively, a performance analysis tool (e.g., Perf) can also be used to collect the number of various performance events generated during the running of the original test program. Accordingly, the performance events with a non-zero number can be regarded as the performance events generated during the running of the original test program, and then the first performance events can be obtained.
[0043] Among them, a first performance event can represent a performance-related characteristic possessed by the original program set. Accordingly, the obtained first performance event can be determined as the original related characteristic. It should be noted that in the actual application scenario, the pre-set conventional performance-related characteristics can also be directly determined as the performance-related characteristics possessed by the original program set, and the embodiments of the present invention do not limit this. Among them, the conventional performance-related characteristics can be the performance events that will be generated during the pre-collected program running.
[0044] Furthermore, based on the importance scores of the original related characteristics, the original related characteristics can be filtered to retain the characteristics with good correlation with performance as the first related characteristics, and filter out the characteristics with low correlation with performance, little or no impact on performance. Specifically, when the importance score is not less than the preset score threshold, it can be determined that the original related characteristic has good correlation with performance. On the contrary, when the importance score is less than the preset score threshold, it can be determined that the original related characteristic has low correlation with performance. In an optional example, the first related characteristics can be as shown in Table 1 below:
[0045]
[0046]
[0047] Table 1
[0048] Among them, " / " in Table 1 represents a ratio operator. Specifically, the importance score of the original relevant feature can be determined according to the number of events of the first performance event corresponding to the original relevant feature. Among them, the importance score can be positively correlated with the number of events of the first performance event corresponding to the original relevant feature. The number of events of the first performance event corresponding to the original relevant feature can be used as the input of the preset importance score function, and then the output of the function can be obtained as the importance score of the original relevant feature. The preset score threshold can be set as needed. Exemplarily, in the case of a longer collection duration, a higher preset score threshold can be set to avoid retaining too many features. In the case of a shorter collection duration, a lower preset score threshold can be set to avoid retaining too few features.
[0049] Among them, the preset importance score function can be pre-fitted through machine learning. Exemplarily, multiple groups of training data can be preset. Among them, a group of training data includes the number of events of a performance-related feature and a standard importance score. The number of events of the performance-related feature is used as the independent variable, and the standard importance score is used as the dependent variable to train the initial score function until the accuracy of the initial score function reaches the preset requirement, then the initial score function can be used as the preset importance score function. Alternatively, correlation analysis can also be performed on the original relevant feature and the cycles per instruction (CPI) to determine the importance score of the original relevant feature. Exemplarily, the correlation score between the original relevant feature and the CPI can be calculated based on the Pearson correlation coefficient formula as the importance score of the original relevant feature.
[0050] After obtaining the importance scores of each original relevant feature, the importance scores of each original relevant feature can be compared with the preset score threshold. If the importance score of the original relevant feature is not less than the preset score threshold, the original relevant feature is retained as the first relevant feature. Otherwise, the original relevant feature is filtered out.
[0051] In the embodiments of the present invention, the performance events generated during the operation of the original test program are obtained as the first performance events. Based on the first performance events, the performance-related features possessed by the original program set are determined as the original relevant features. The importance scores of each original relevant feature are determined, and the original relevant features with importance scores not less than the preset score threshold are determined as the first relevant features. In this way, based on the performance events generated during the actual operation of the programs in the original program set, the original relevant features are determined, and the original relevant features with importance scores not less than the preset score threshold are selected as the first relevant features, which can ensure that the finally determined first relevant features can more accurately reflect the performance-related features actually possessed by the original program set. At the same time, the number of the first relevant features can be prevented from being too large, which is convenient for subsequent processing.
[0052] In the embodiments of the present invention, the original assembly can also be maintained and updated. Exemplarily, some test programs therein can be deleted, new test programs can be added to the original assembly, and so on. Optionally, the embodiments of the present invention may further include the following steps:
[0053] Step S21: Determine performance-related features for the default assembly as the second related features.
[0054] Step S22: Search for second related features that do not belong to the first related features as the features to be supplemented.
[0055] Step S23: Obtain the original test programs with the features to be supplemented and add them to the original assembly.
[0056] Among them, determining performance-related features for the default assembly as the second related features may be to obtain performance events generated during the operation of the default test program as the second performance events. Determine the performance-related features possessed by the default assembly based on the second performance events as the original related features of the default assembly. Determine the importance scores of the original related features of the default assembly, and determine the original related features with importance scores not less than the preset score threshold as the second related features. The implementation manners of the steps for determining the second related features can refer to the implementation manners of the steps for determining the first related features described above, and will not be elaborated here.
[0057] Furthermore, each of the second related features can be compared with the first related features. If the first related features include this second related feature, it can be determined that this second related feature belongs to the first related features. On the contrary, if the first related features do not include this second related feature, it can be determined that this second related feature does not belong to the first related features, and the original assembly lacks test programs with this feature. Accordingly, this second related feature can be used as the feature to be supplemented, and test programs with this feature to be supplemented can be added to the original assembly.
[0058] Specifically, the feature to be supplemented can be used as the input of the preset program generation model, and the output of the preset program generation model can be obtained as the original test program with the feature to be supplemented. Among them, the preset program generation model can be a pre-trained model for generating a program with the input feature. Exemplarily, multiple sets of training data can be obtained. Among them, a set of training data includes a performance-related feature and a test program with the performance-related feature. The initial generation model is trained using these multiple sets of training data until the initial generation model reaches the preset loss rate or the number of training times reaches the preset number threshold, then the initial generation model can be used as the preset program generation model. Alternatively, the feature to be supplemented can also be sent to the user to remind the user to provide a program with the feature to be supplemented. Accordingly, the program with the feature to be supplemented input by the user can be received as the original test program with the feature to be supplemented and added to the original program set.
[0059] In the embodiment of the present invention, by determining the performance-related feature for the default program set as the second related feature. Search for the second related feature that does not belong to the first related feature as the feature to be supplemented. Obtain the original test program with the feature to be supplemented and add it to the original program set. In this way, by expanding the original program set, the original program set has richer program-related features, and then the test subset screened from the original program set can better simulate the test effect of the default program set, ensuring the accuracy of the test.
[0060] Optionally, the step of clustering the original test programs in the original program set and the default test program based on the feature values of the original test programs in the original program set corresponding to the first related feature and the feature values of the default test program in the default program set corresponding to the first related feature may specifically include:
[0061] Step 1021: For any one of the default test programs, form a program group to be clustered by the default test program and all the original test programs in the original program set.
[0062] Step 1022: For any one of the program groups to be clustered, perform a clustering operation on the test programs in the program group to be clustered based on the feature values of the test programs in the program group to be clustered corresponding to the first related feature, and obtain multiple clustering clusters corresponding to the program group to be clustered; where the clustering cluster includes the test programs in the program group to be clustered.
[0063] Specifically, the number of program groups to be clustered is the same as the number of default test programs. That is, for a default test program, all the original test programs in the original program set can be grouped with this default test program to form a program group to be clustered. By performing clustering operations on this program group to be clustered respectively, it can be known which original test programs belong to the same category as this default test program, thus facilitating the search for the original test programs that match this default test program. Among them, the multiple clusters corresponding to the program group to be clustered are the clustering results. Suppose there are 10 default test programs, then 10 program groups to be clustered can be obtained, and each of the 10 program groups to be clustered includes a default test program.
[0064] Clustering operations are performed on each program group to be clustered. Exemplarily, multiple clusters corresponding to each of the 10 program groups to be clustered can be obtained. Specifically, when performing clustering operations on any program group to be clustered, a test program in this program group to be clustered can be represented as a data point, and this data point is the feature value of the first relevant feature corresponding to this test program. Exemplarily, suppose there are 20 first relevant features, then a data point can be represented as (x1, x2,..., x20). The data points of all the test programs in this program group to be clustered can be used as the input of the clustering algorithm, and the clustering algorithm performs clustering analysis based on the input data points to generate the number of clusters of clusters, and the clusters include the test programs in the program group to be clustered. The goal of clustering analysis is to discover the internal structure of the data, group similar objects into one category, and separate dissimilar objects. Specifically, the clustering algorithm can adopt the hierarchical clustering method. Exemplarily, the Euclidean distance between each data point and other data points can be calculated first to obtain a distance matrix. Suppose there are N test programs in the program group to be clustered, then an N×N distance matrix can be obtained. For this distance matrix, the element at the (i, j) position is the Euclidean distance between the test program represented by the i-th row and the test program represented by the j-th row. Suppose the test program represented by the i-th row is P(q1, q2,..., qn), and the test program represented by the j-th row is Q(p1, p2,..., pn), then the Euclidean distance between the two can be calculated by the following formula:
[0065]
[0066] Next, each test program can be regarded as a separate cluster, and the distances between this test program and other test programs can be determined based on the distance matrix (for example, the values of the elements in the row where this test program is located can be read, and the distances between this test program and other individual test programs can be extracted therefrom). Then, the test program with the minimum distance (i.e., the closest to this test program) is merged with this test program into a cluster. Then, according to the data points of the test programs in the same cluster, the cluster center of this cluster is determined. For example, the data points of the test programs in the same cluster can be averaged to obtain the data point representing this cluster. Among them, when averaging, the eigenvalue of the same feature can be averaged. Exemplarily, in the case of including n first relevant features, the average value of the eigenvalues corresponding to each of the n first relevant features can be obtained. Next, according to the data points of the cluster, the Euclidean distance between each cluster is calculated.
[0067] The distance matrix is updated based on the Euclidean distance between each cluster. For the updated distance matrix, the element at the (i, j) position is the Euclidean distance between the cluster represented by the i-th row and the cluster represented by the j-th row. For any cluster, based on the updated distance matrix, the distances between this cluster and other clusters are determined. Then, the cluster with the minimum distance from this cluster among other clusters is selected, and the selected cluster is merged with this cluster into a cluster. Then, return to the step of determining the cluster center of this cluster according to the data points of the test programs in the same cluster and start to execute, so that the test programs with close distances are grouped into one category until the number of clusters reaches the clustering number. Among them, the clustering number can be set in advance as needed, and the clustering number is a positive integer. The specific value of the clustering number is not limited in the embodiments of the present invention. It should be noted that in the embodiments of the present invention, it can also be repeatedly executed until all data points are merged into one cluster. Correspondingly, a clustering dendrogram can be generated accordingly, and the clustering number is determined based on the clustering dendrogram. Exemplarily, the inflection points in the clustering dendrogram can be identified, that is, the points where the inter-cluster distance suddenly changes. The number of inflection points is used as the clustering number. Then, a target horizontal line is set, and the target horizontal line is a horizontal line whose number of intersection points with the clustering dendrogram is the clustering number. For each intersection point, the test programs below this intersection point are grouped into one cluster.
[0068] Optionally, in the embodiments of the present invention, the eigenvalues of the first relevant features corresponding to each test program in the program group to be clustered can also be dimensionally reduced first, and clustering operations are performed based on the dimensionally reduced eigenvalues.
[0069] Specifically, before clustering the test programs in the program group to be clustered based on the eigenvalues of the first relevant features corresponding to each test program in the program group to be clustered, the following steps can also be included:
[0070] Step S31: Based on the eigenvalues of the first performance-related features corresponding to each test program in the program group to be clustered, construct an original feature matrix.
[0071] Step S32: Perform principal component analysis on the original feature matrix to reduce the dimensionality of the eigenvalues of the first performance-related features corresponding to each test program.
[0072] Among them, principal component analysis (PCA) refers to transforming a set of variables into a set of linearly independent variables through orthogonal transformation, and the transformed set of variables is called the principal components. Specifically, one row of the matrix can be set to correspond to a test program in a program group to be clustered, and one column of the matrix can correspond to a first related feature. A first related feature can be regarded as a feature variable. For any first related feature, use the eigenvalue of the test program corresponding to the first performance-related feature as the element at the position of the row corresponding to the test program and the column corresponding to the first performance-related feature, thus obtaining the original feature matrix. Suppose there are 100 test programs in the program group to be clustered and 20 first performance-related features, then an original feature matrix of 100 rows × 20 columns can be obtained. The element value at the (i, j) position in this original feature matrix is the eigenvalue of the first performance-related feature represented by the jth column corresponding to the test program represented by the ith row.
[0073] Next, the PCA algorithm can be used to perform principal component analysis on the original feature matrix. Specifically, the data in the original feature matrix can be standardized first to eliminate the deviation caused by different measurement units and value ranges of different first related feature events, so that different feature variables in the original feature matrix are comparable. Exemplarily, the mean and standard deviation can be calculated by column. For each element in a column, calculate the difference between the element and the mean of the column, and then update the element to the ratio of the difference to the standard deviation of the column, that is, the ratio of the difference to the standard deviation of the column is the element after standardization processing.
[0074] Then, the covariance matrix of the original feature matrix after data standardization can be calculated. Among them, covariance can describe the correlation degree between multiple groups of feature data. If the covariance > 0, it indicates a positive correlation between the two; if the covariance < 0, it indicates a negative correlation between the two; if the covariance = 0, it indicates that the two are not correlated. Next, the eigenvalues and eigenvectors of this covariance matrix can be calculated. Specifically, the covariance matrix can be decomposed by eigenvalues to obtain the eigenvalues and corresponding eigenvectors of the covariance matrix. Among them, the eigenvectors of the covariance matrix represent the principal component directions. An eigenvector of the covariance matrix describes a principal component, and each eigenvector is orthogonal and describes the data changes in different directions. The eigenvalue represents the variance size in each principal component direction. The larger the eigenvalue, the greater the variance of the data in the corresponding eigenvector direction, so this eigenvector has a stronger ability to explain the data.
[0075] Furthermore, the eigenvalues of the covariance matrix can be sorted in descending order, and the eigenvectors corresponding to the top K largest eigenvalues are selected to form a matrix to obtain the target feature matrix. Among them, a row in the target feature matrix corresponds to a test program in the program group to be clustered, a column corresponds to an eigenvector, and an eigenvector is a principal component. In this way, it is equivalent to projecting the original data onto the selected principal components to obtain the data after dimensionality reduction. The eigenvalue of the first performance-related feature corresponding to the test program represented by a row in the target feature matrix after dimensionality reduction can be used. When performing clustering operations, the eigenvalues of the first performance-related features corresponding to each test program after dimensionality reduction can be used. Among them, the larger the eigenvalue corresponding to the principal component, the greater the variance of the original data retained by the principal component, and the stronger the ability to explain the original data. For example, the first principal component with the largest corresponding eigenvalue retains the largest variance of the original data, and the second principal component retains the second largest variance of the original data. In the embodiments of the present invention, dimensionality reduction is performed through principal component analysis, which can reduce the amount of data while retaining as much original information as possible. It should be noted that the variance ratio to be retained can also be set in advance for the PCA algorithm. Exemplarily, the set variance ratio can be 0.9 to control that the data after dimensionality reduction can retain 90% of the variance information of the original data.
[0076] After performing principal component analysis, it is equivalent to performing principal component dimensionality reduction on the original feature matrix, and a 100×K matrix can be obtained as the target feature matrix. Correspondingly, when performing clustering operations, the data points of each test program are K-dimensional. Assuming K is 8, then it is equivalent to reducing the original 20-dimensional matrix to 8 dimensions, and the 8 dimensions can better reflect the change trend of the original 20 dimensions. The data points of each test program can be expressed as (x1, x2,..., x8), corresponding to the values of 8 principal components of the test program in the target feature matrix.
[0077] In the embodiments of the present invention, before clustering the test programs in the program group to be clustered based on the feature values of the first related features corresponding to each test program in the program group to be clustered, an original feature matrix is first constructed based on the feature values of each first performance-related feature corresponding to each test program in the program group to be clustered. Principal component analysis is performed based on the original feature matrix to reduce the dimensionality of the feature values of the first performance-related features corresponding to each test program. In this way, the data processing volume in the subsequent clustering operation can be reduced, which is convenient for processing and thus improves the processing efficiency.
[0078] Optionally, in the embodiments of the present invention, the step of screening the original test programs matched by each of the default test programs based on the clustering result to form a test subset may specifically include:
[0079] Step 1031: For any one of the program groups to be clustered, determine the cluster that includes the default test program among the multiple clusters corresponding to the program group to be clustered as the target cluster.
[0080] Step 1032: Select the original test program with the highest similarity to the default test program from the target cluster as the original test program matched by the default test program.
[0081] Step 1033: Form a program set from all the original test programs matched by the default test programs to obtain the test subset.
[0082] Specifically, the program group to be clustered consists of an original test program and 1 default test program. Through the clustering operation, the program group to be clustered will be divided into one cluster. Correspondingly, for any program group to be clustered, the cluster that includes the default test program in the clustering result (i.e., the multiple clusters corresponding to the program group to be clustered) of the program group to be clustered can be used as the target cluster. Exemplarily, assume that the multiple clusters corresponding to the program group to be clustered are: Cluster 1, Cluster 2, and Cluster 3, where Cluster 1 includes the default test program, then Cluster 1 can be used as the target cluster.
[0083] Furthermore, the original test program with the closest Euclidean distance to the included default test program in the target cluster can be determined as the original test program matched by the default test program. That is, the highest similarity means the smallest Euclidean distance. Finally, the original test programs matched by the default test programs selected from all the target clusters can be formed into a program set to obtain the test subset. Assume there are Z clustering program groups, then finally Z target clusters can be obtained. Correspondingly, the original test programs respectively matched by the Z default test programs can be obtained, and the test subset can include these Z original test programs.
[0084] In an embodiment of the present invention, for any program group to be clustered, the cluster containing the default test program among the multiple clusters corresponding to the program group to be clustered is determined as the target cluster. The original test program with the highest similarity to the default test program is selected from the target cluster as the original test program matched with the default test program. All the original test programs matched with the default test programs are grouped into a program set to obtain a test subset. In this way, the original test programs matched with each default test program can be added to the program set, ensuring that the test subset can include all the original test programs matched with the default test programs, and further ensuring the representativeness of the test subset.
[0085] Optionally, the embodiment of the present invention may further include the following steps:
[0086] Step S41: When the speedup ratio of the test subset does not match the speedup ratio of the default program set, update the clustering parameters of the clustering operation.
[0087] Step S42: Re-enter the step of performing a clustering operation on the original test programs in the original program set and the default test programs based on the eigenvalues of the first related features corresponding to the original test programs in the original program set and the eigenvalues of the first related features corresponding to the default test programs in the default program set.
[0088] Among them, the clustering parameters may be the parameters of the above clustering algorithm. Exemplarily, the clustering parameters may include the number of clusters. Correspondingly, updating the clustering parameters of the clustering operation may include reducing the number of clusters according to a preset value, or increasing the number of clusters according to a preset value. Specifically, in the embodiment of the present invention, when the speedup ratio of the test subset does not match the speedup ratio of the default program set, the program similarity between the test subset and the default program set may be calculated first. If the program similarity is less than a preset similarity threshold, new original test programs may be added to the original program set. Then return to step 101 above to start execution, so as to perform clustering operations again subsequently to find a new test subset. If the program similarity is not less than the preset similarity threshold, the clustering parameters of the clustering operation may be updated, and directly re-enter the step of performing a clustering operation on the original test programs in the original program set and the default test programs based on the eigenvalues of the first related features corresponding to the original test programs in the original program set and the eigenvalues of the first related features corresponding to the default test programs in the default program set to start execution, until a test subset with a speedup ratio matching that of the default program set is screened out.
[0089] In an embodiment of the present invention, when the speedup ratio of the test subset does not match the speedup ratio of the default program set, the clustering parameters of the clustering operation are updated. Then, re-enter the step of performing a clustering operation on the original test programs and the default test programs in the original program set based on the eigenvalues of the first relevant features corresponding to the original test programs in the original program set and the eigenvalues of the first relevant features corresponding to the default test program in the default program set. In this way, by starting to execute the step of performing the clustering operation again, a test subset can be re-determined to screen a test subset whose speedup ratio matches the speedup ratio of the default program set.
[0090] Optionally, the embodiment of the present invention may further include the following steps:
[0091] Step S31: Determine multiple first speedup ratios corresponding to multiple optimization levels of the test subset, and determine multiple second speedup ratios corresponding to multiple optimization levels of the default program set.
[0092] Step S32: If the change trends of the multiple first speedup ratios match the change trends of the multiple second speedup ratios, and the difference between the first speedup ratio and the second speedup ratio corresponding to the same optimization level is within a preset difference range, it is determined that the speedup ratio of the test subset matches the speedup ratio of the default program set.
[0093] Correspondingly, if the change trends of the multiple first speedup ratios do not match the change trends of the multiple second speedup ratios, and the difference between the first speedup ratio and the second speedup ratio at the same optimization level is within the preset difference range, it can be determined that the test subset does not match the default program set.
[0094] For any optimization level, the speedup of the test subset is defined as the execution time of the program for the test subset without optimization divided by the execution time of the program for the test subset after optimization at this optimization level. The speedup of the default assembly is defined as the execution time of the program for the default test program without optimization divided by the execution time of the program for the default test program after optimization at this optimization level. Specifically, for any optimization level, the speedup of each original test program in the test subset can be calculated first. Among them, the speedup of the original test program at this optimization level is: the execution time of the original test program without optimization divided by the execution time of the original test program after optimization at this optimization level, and then the average value of the speedups of all original test programs at this optimization level is calculated. Specifically, the weight can be set for the original test program according to the number of programs included in the cluster where the original test program is located. Among them, the weight is proportional to the number of programs, and the higher the number of programs, the higher the weight. Exemplarily, the ratio of the number of programs included in the cluster where the original test program is located to the number of programs in the program group to be clustered can be calculated to obtain the weight of the original test program. The product of the speedup of the original test program at this optimization level and the weight of the original test program is calculated to obtain the weighted speedup of the original test program at this optimization level. Then, based on the weighted speedups of all original test programs in the test subset at this optimization level, the geometric mean is calculated to obtain the speedup of the test subset corresponding to this optimization level, that is, a first speedup is obtained.
[0095] Furthermore, for any optimization level, the speedup of each default test program in the default assembly can be calculated first. Among them, the speedup of the default test program at this optimization level is: the execution time of the default test program without optimization divided by the execution time of the default test program after optimization at this optimization level. Then, based on the speedups of all default test programs at this optimization level, the geometric mean is calculated to obtain the speedup of the default assembly corresponding to this optimization level, that is, a second speedup is obtained.
[0096] For one of the multiple optimization levels, a first speedup and a second speedup can be obtained. Among them, the multiple optimization levels can be preset, and the optimization level can refer to the optimization level of the compiler, such as O1, O2, O3, Ofast, etc. The test program obtained without enabling the optimization level can be run, and the program execution time can be counted to obtain the execution time of the test program without optimization. The test program obtained after enabling this optimization level can be run, and the program execution time can be counted to obtain the execution time of the test program after optimization at this optimization level. Among them, the test program can be the above-mentioned original test program or default test program.
[0097] Specifically, a measurement parameter of multiple first speedup ratios and multiple second speedup ratios can be calculated. If the measurement parameter is not greater than a preset threshold, it can be determined that the change trends of the multiple first speedup ratios are similar to the change trends of the multiple second speedup ratios, and further, it can be determined that the change trends of the multiple first speedup ratios match the change trends of the multiple second speedup ratios. Conversely, it can be determined that the change trends of the multiple first speedup ratios do not match the change trends of the multiple second speedup ratios. Among them, the measurement parameter is used to measure whether the change trends of the two match. Specifically, the measurement parameter can be the mean square error, Pearson correlation coefficient, Dynamic Time Warping (DTW) distance, etc. Or, a line graph can be drawn based on the multiple first speedup ratios and the multiple second speedup ratios. Then, the line formed by the multiple first speedup ratios and the line formed by the multiple second speedup ratios are analyzed through the line graph. If the included angle between the two lines is less than the preset threshold, it can be determined that the change trends of the multiple first speedup ratios match the change trends of the multiple second speedup ratios.
[0098] Further, the preset difference range can be set in advance according to the actual situation. Exemplarily, the preset difference range can be (-5, 5), and the embodiments of the present invention are not limited thereto. Suppose there are 3 optimization levels: optimization level 1, optimization level 2, and optimization level 3. Then, multiple first speedup ratios corresponding to multiple optimization levels of the test subset can be obtained: a1, a2, and a3, and multiple second speedup ratios corresponding to multiple optimization levels of the default assembly: b1, b2, and b3. Among them, a1 and b1 correspond to the same optimization level: optimization level 1, a2 and b2 correspond to the same optimization level: optimization level 2, and a3 and b3 correspond to the same optimization level: optimization level 3. The mean square error of the multiple first speedup ratios and the multiple second speedup ratios can be calculated: ((a1 - b1) 2 +(a2 - b2) 2 +(a3 - b3) 2 ) / 3. If the mean square error is not greater than the preset mean square error threshold, it can be determined that the change trends of the multiple first speedup ratios match the change trends of the multiple second speedup ratios. At the same time, it can be determined whether a1 - b1 is within the preset difference range, whether a2 - b2 is within the preset difference range, and whether a3 - b3 is within the preset difference range. If all are within the preset difference range, the test subset can be used to evaluate the performance of the compiler. Among them, the compiler can be any compiler that needs to be evaluated for performance.
[0099] In an embodiment of the present invention, first, multiple first speedup ratios corresponding to multiple optimization levels of a test subset are determined, and multiple second speedup ratios corresponding to multiple optimization levels of a default program set are determined. If the change trends of the multiple first speedup ratios match the change trends of the multiple second speedup ratios, and the difference between the first speedup ratio and the second speedup ratio corresponding to the same optimization level is within a preset difference range, it is determined that the speedup ratio of the test subset matches the speedup ratio of the default program set. In this way, it can be ensured that a test subset having the same change trend as the default program set and with a difference in speedup ratio within the preset difference range and a relatively small deviation in the speedup ratio value is used for performance evaluation, thereby ensuring the evaluation accuracy of using the test subset for performance evaluation.
[0100] Optionally, in an embodiment of the present invention, each original test program can also be split into multiple modules according to a preset format. The preset format may include: source code, compilation settings and link commands, input files required for running, and runtime command parameters. Accordingly, through splitting, each original test program is split into: the source code of the original test program, compilation settings and link commands, input files required for running, and runtime command parameters. Exemplarily, splitting can be achieved by manual splitting by a person, or automatically extracting each part of the original test program according to the preset format. In this way, the original test programs in the original program set all adopt a unified format. Due to the unified format, it is convenient for other users to understand and use, and at the same time, it is convenient to uniformly process different original test programs.
[0101] Exemplarily, a developer can pre-establish a script framework for unified processing of compilation and running at the top level. When it is necessary to compile a certain compiler, compiler parameters and compilation options can be passed into the script framework, and the script framework can uniformly compile the test programs in the test subset through the compiler according to the passed compiler parameters and compilation options to obtain binary executable files. Further, the generated binary executable files can be uniformly run through a running script, and then performance evaluation parameters are generated, and the performance evaluation parameters of each test program in the test subset are summarized and output.
[0102] Figure 2 is a schematic diagram of a processing flow provided by an embodiment of the present invention, as Figure 2As shown, the original assembly collected in advance can be obtained first. Then, the first relevant features are determined for the original assembly. Next, based on the default test program and all the original test programs in the original assembly, a program group to be clustered is formed. Principal component analysis is performed based on the feature values of the first performance-related features corresponding to each test program in the program group to be clustered to achieve dimensionality reduction. Then, clustering operations are performed based on the feature values of the first performance-related features corresponding to each test program after dimensionality reduction to screen the test subset. It is determined whether the speedup ratio of the test subset matches the speedup ratio of the default assembly. If so, the test subset is used to evaluate the performance of the compiler. If not, the clustering parameters are modified and the clustering operations are returned to be performed again.
[0103] In the embodiments of the present invention, by performing clustering operations on the original test programs and the default test program in the original assembly, the mapping relationship between the original assembly and the default test program is found, that is, the original test program matched by each default test program is found. Based on the original test programs matched by each default test program, a test subset is formed, and the test subset is used to represent the entire default assembly. The performance of the compiler is evaluated based on the test subset. The original test programs matched by different default test programs may be the same, that is, different default test programs are mapped to the same original test program, and one original test program can represent multiple default test programs. In this way, the number of programs in the test subset can be less than the number of programs in the default assembly. Since the number of programs in the test subset is smaller, the time required for compilation and testing can be further reduced, the time cost of performance evaluation can be reduced, the performance evaluation of the compiler (Compiler Performance Evaluation) can be accelerated, the evaluation efficiency of the compiler can be improved, and thus the optimization efficiency of the compiler can be improved. At the same time, when the speedup ratio of the test subset matches the speedup ratio of the default assembly, the test subset is used to represent the entire default assembly, and the performance of the compiler is evaluated based on the test subset. In this way, the accuracy of the evaluation can be ensured to a certain extent. At the same time, since the number of programs in the test subset is smaller, it is convenient to use the test subset to evaluate the performance of different versions and different types of compilers under different optimization levels, different optimization options, and different compilation settings respectively.
[0104] Refer to Figure 3 , a block diagram of a compiler evaluation device provided by an embodiment of the present invention is shown. As Figure 3 shown, the compiler evaluation device may specifically include:
[0105] A first determination module 201, configured to determine performance-related features for the original assembly collected in advance as the first relevant features;
[0106] The clustering module 202 is configured to perform a clustering operation on the original test programs in the original program set and the default test programs based on the feature values of the first related features corresponding to the original test programs in the original program set and the feature values of the first related features corresponding to the default test programs in the default program set;
[0107] The screening module 203 is configured to screen the original test programs matched by the default test programs based on the clustering result to form a test subset;
[0108] The evaluation module 203 is configured to perform a performance evaluation on the compiler based on the test subset when the speedup ratio of the test subset matches the speedup ratio of the default program set.
[0109] Optionally, the clustering module 202 is specifically configured to:
[0110] For any one of the default test programs, form a program group to be clustered by the default test program and all the original test programs in the original program set;
[0111] For any one of the program groups to be clustered, perform a clustering operation on the test programs in the program group to be clustered based on the feature values of the first related features corresponding to the test programs in the program group to be clustered, and obtain multiple clustering clusters corresponding to the program group to be clustered; wherein, the clustering clusters include the test programs in the program group to be clustered.
[0112] Optionally, the screening module 203 is specifically configured to:
[0113] For any one of the program groups to be clustered, determine the clustering cluster including the default test program among the multiple clustering clusters corresponding to the program group to be clustered as the target cluster;
[0114] Select the original test program with the highest similarity to the default test program from the target cluster as the original test program matched by the default test program;
[0115] Form a program set with all the original test programs matched by the default test programs to obtain the test subset.
[0116] Optionally, the device further includes:
[0117] The construction module is configured to construct an original feature matrix based on the feature values of the first performance-related features corresponding to the test programs in the program group to be clustered before the clustering module performs a clustering operation on the test programs in the program group to be clustered based on the feature values of the first related features corresponding to the test programs in the program group to be clustered;
[0118] An analysis module, for a module, for performing principal component analysis based on the original feature matrix to reduce the dimensionality of the eigenvalues of each of the test programs corresponding to the first performance-related features.
[0119] Optionally, the device further includes:
[0120] A second determination module, for determining a plurality of first speedup ratios of the test subset corresponding to a plurality of optimization levels, and determining a plurality of second speedup ratios of the default program set corresponding to a plurality of optimization levels;
[0121] A third determination module, for determining that the speedup ratio of the test subset matches the speedup ratio of the default program set if the change trends of the plurality of first speedup ratios match the change trends of the plurality of second speedup ratios, and the difference between the first speedup ratio and the second speedup ratio corresponding to the same optimization level is within a preset difference range.
[0122] Optionally, the device further includes:
[0123] A fourth determination module, for determining performance-related features for the default program set as second related features;
[0124] A search module, for searching for second related features that do not belong to the first related features as features to be supplemented;
[0125] An acquisition module, for acquiring original test programs having the features to be supplemented and adding them to the original program set.
[0126] Optionally, the first determination module 201 is specifically configured to:
[0127] Acquire performance events generated during the running of the original test program as first performance events;
[0128] Determine performance-related features possessed by the original program set based on the first performance events as original related features;
[0129] Determine the importance scores of the original related features, and determine the original related features with importance scores not less than a preset score threshold as the first related features.
[0130] Optionally, the device further includes:
[0131] An update module, for updating the clustering parameters of the clustering operation in the case where the speedup ratio of the test subset does not match the speedup ratio of the default program set;
[0132] A processing module, configured to re-enter the eigenvalue of each original test program in the original program set corresponding to the first related feature and the eigenvalue of the default test program in the default program set corresponding to the first related feature, and perform a clustering operation on the original test programs and the default test programs in the original program set.
[0133] In summary, in the compiler evaluation device provided in the embodiment of the present invention, performance-related features are determined for a pre-collected original program set as the first related features. Based on the eigenvalues of each original test program in the original program set corresponding to the first related feature and the eigenvalues of the default test programs in the default program set corresponding to the first related feature, a clustering operation is performed on the original test programs and the default test programs in the original program set. Based on the clustering result, the original test programs matching each default test program are screened to form a test subset. When the speedup ratio of the test subset matches the speedup ratio of the default program set, the performance of the compiler is evaluated based on the test subset. In this way, it is possible to use a test subset screened from a pre-collected original program set to evaluate the performance of the compiler, reducing the limitations of performance evaluation and improving the flexibility of performance evaluation. At the same time, by further screening a test subset from the original test program set whose speedup ratio matches the speedup ratio of the default program set, it can be ensured that the screened test subset can more accurately represent the default program set, and thus ensure the evaluation effect when using the test subset to replace the default program set for performance evaluation.
[0134] Refer to Figure 4 , which is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. As Figure 4 shown, the electronic device includes: a processor, a memory, a communication interface, and a communication bus.
[0135] The processor, the memory, and the communication interface complete mutual communication through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the compiler evaluation method of the foregoing embodiment. The executable instructions can form a program.
[0136] An embodiment of the present invention provides a machine-readable medium, on which instructions are stored, and when executed by one or more processors, enable the processors to execute the compiler evaluation method of the foregoing embodiment.
[0137] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same and similar parts among the embodiments can be referred to each other.
[0138] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and with the authorization given by the owner of the corresponding device.
[0140] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0141] These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing terminal device to work in a predictive manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0143] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0144] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0145] Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element.
[0146] The above has introduced in detail a compiler evaluation method, a compiler evaluation device, an electronic device and one or more readable media provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A compiler evaluation method, characterized in that: The method comprises: determining performance-related features for the pre-collected original program set as first-related features; Based on the feature value of each original test program in the original program set corresponding to the first relevant feature and the feature value of the default test program in the default program set corresponding to the first relevant feature, clustering the original test programs in the original program set and the default test programs; Based on the clustering results, the original test programs matching the default test programs are screened to form a test subset; When the speedup ratio of the test subset matches the speedup ratio of the default program set, a performance evaluation is performed on the compiler based on the test subset.
2. The method according to claim 1, characterized in that The clustering operation of the original test programs in the original program set and the default test programs based on the feature value of the first relevant feature corresponding to each original test program in the original program set and the feature value of the first relevant feature corresponding to the default test program in the default program set comprises: For any of the default test programs, the default test program and all the original test programs in the original program set are combined into a program group to be clustered; For any of the program groups to be clustered, based on the feature values of each test program in the program group to be clustered corresponding to the first related feature, a clustering operation is performed on the test programs in the program group to be clustered to obtain multiple clustering clusters corresponding to the program group to be clustered; wherein the clustering cluster includes the test programs in the program group to be clustered.
3. The method according to claim 2, characterized in that The original test programs matching the default test programs are screened based on the clustering results to form a test subset, including: For any of the program groups to be clustered, determining a cluster including the default test program among the multiple clusters corresponding to the program group to be clustered as a target cluster; Selecting an original test program having the highest similarity to the default test program from the target cluster as the original test program matched by the default test program; All original test programs that match the default test program are grouped into a program set to obtain the test subset.
4. The method according to claim 2, characterized in that: Before clustering the test programs in the group of programs to be clustered based on the feature value of each test program in the group of programs to be clustered corresponding to the first relevant feature, the method further includes: constructing an original feature matrix based on the feature values of each test program in the group of programs to be clustered corresponding to the first performance-related feature; A principal component analysis is performed based on the original feature matrix to reduce the dimensionality of the eigenvalues of the first performance-related features corresponding to each of the test programs.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Determine a plurality of first speedup ratios corresponding to a plurality of optimization levels for the test subset, and determine a plurality of second speedup ratios corresponding to a plurality of optimization levels for the default program set; If the changing trend of the multiple first acceleration ratios matches the changing trend of the multiple second acceleration ratios, and the difference between the first acceleration ratio and the second acceleration ratio corresponding to the same optimization level is within a preset difference range, it is determined that the acceleration ratio of the test subset matches the acceleration ratio of the default program set.
6. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: determining a performance-related characteristic for the default program set as a second related characteristic; Searching for a second related feature that does not belong to the first related feature as a feature to be supplemented; The original test program having the features to be supplemented is obtained and added to the original program set.
7. The method according to any one of claims 1 to 4, characterized in that: The determining of performance-related features for the pre-collected original program set as first related features includes: Acquire a performance event generated during the running of the original test program as a first performance event; Determine, based on the first performance event, performance-related features possessed by the original program set as original related features; The importance score of each of the original relevant features is determined, and the original relevant features whose importance scores are not less than a preset score threshold are determined as the first relevant features.
8. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: In the case where the speedup ratio of the test subset does not match the speedup ratio of the default program set, updating clustering parameters of the clustering operation; Re-enter the step of clustering the original test programs in the original program set and the default test programs based on the feature values of the first relevant features corresponding to each original test program in the original program set and the feature values of the first relevant features corresponding to the default test programs in the default program set.
9. A compiler evaluation device, characterized in that: The device comprises: A first determination module, configured to determine a performance-related feature for a pre-collected original program set as a first related feature; A clustering module, configured to perform clustering operations on the original test programs in the original program set and the default test programs based on the feature values of the first relevant features corresponding to the original test programs in the original program set and the feature values of the first relevant features corresponding to the default test programs in the default program set; A screening module, used for screening original test programs matching each of the default test programs based on the clustering results to form a test subset; An evaluation module is used to perform performance evaluation on the compiler based on the test subset when the speedup ratio of the test subset matches the speedup ratio of the default program set.
10. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store executable instructions, and the executable instructions enable the processor to execute the device according to any one of claims 1 to 8.
11. One or more machine-readable media, characterized in that Instructions are stored thereon, which, when executed by one or more processors, cause the processors to execute the apparatus according to any one of claims 1-8.