Heterogeneous CPU cluster scheduling method based on program similarity analysis
By building a control flow diagram and combining Weisfeiler-Lehman kernel method for program similarity analysis, the resource matching and load balancing problems of task scheduling under the heterogeneous CPU architecture are solved, and efficient resource utilization and performance improvement of heterogeneous CPU clusters are achieved.
Patent Information
- Application Number
- CN202510471429.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional scheduling methods are difficult to effectively deal with the problem of task characteristics and resource matching and load balancing under heterogeneous CPU architecture, resulting in low resource utilization and poor performance.
By constructing a control flow graph, node feature vectors are extracted and normalized, similarity analysis is performed using Weisfeiler-Lehman kernel method, and combined with program global features and resource usage information, the scheduling decisions of tasks in heterogeneous CPU clusters are optimized.
The resource utilization and overall performance of heterogeneous CPU clusters are improved, and the computing efficiency is improved through precise scheduling decisions.
Smart Images

Figure CN120353595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of GPU cluster scheduling, and particularly to a heterogeneous CPU cluster scheduling method based on program similarity analysis. Background Art
[0002] With the continuous development of cloud computing, driven by hardware iteration and diverse service requirements, cloud service providers are prompted to retain CPUs with different architectures in multiple generations in the data center to meet diverse requirements for performance and cost.
[0003] Traditional scheduling methods use optimization algorithms, such as genetic algorithms and ant colony algorithms, etc., to search for the optimal nodes. However, due to the very large performance differences of different architectures when processing different types of tasks, and the performance changes dynamically with load, temperature, etc., it is difficult to effectively handle the matching of task characteristics and architectures, as well as resource competition and load balancing. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a scheduling method based on program similarity analysis, aiming to optimize task allocation in a heterogeneous architecture cloud environment through cross-architecture program feature comparison.
[0005] To achieve the above purpose, the present invention provides the following solution:
[0006] A scheduling method based on program similarity analysis, comprising the following steps:
[0007] Step 1. Construct a control flow graph of the target program. Obtain the assembly code of the target program through a disassembler tool, use the functions in the assembly code as nodes, and add edges to the nodes according to the jump and call relationships between the functions to form a control flow graph;
[0008] Step 2. Extract the feature vectors of the control flow graph nodes, including classifying and counting the assembly instructions to generate node feature vectors and the overall feature vector of the program;
[0009] Step 3. Normalize the node feature values to eliminate the influence of program scale differences on the feature distribution;
[0010] Step 4. Perform hierarchical mapping on the feature values to map the high-dimensional feature vectors to a low-dimensional space and generate a low-dimensional embedding representation of the program;
[0011] Step 5. Calculate the similarity between programs. Perform similarity analysis on the control flow graph through a graph kernel method, and combine the global feature information of the programs to comprehensively evaluate the similarity between programs;
[0012] Step 6. According to the program similarity analysis results and combined with the real-time status information of the cluster nodes, select the reference program with the highest similarity to the target program, and schedule the task to the cluster node that best matches the resource usage preference of the reference program.
[0013] Advantages of the present invention: By combining the control flow graph and the Weisfeiler-Lehman kernel method, the present invention can quickly analyze the resource usage preference of programs, provide more accurate guidance for scheduling decisions, and thus improve resource utilization and overall performance. Brief Description of the Drawings
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments.
[0015] Figure 1 It is a schematic diagram of the technical route provided by the embodiment of the present invention;
[0016] Figure 2 It is a flowchart of the method provided by the embodiment of the present invention. Detailed Embodiments
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0018] Such as Figure 1As shown in the figure, the present invention first constructs a control flow graph for the target program. By disassembling the program to extract basic blocks and functions, analyzing the jump and call relationships in the assembly code, and constructing a control flow graph with functions as nodes and call relationships as edges, the control flow structure of the program can be fully expressed. Then, the node features of the control flow graph are extracted. The assembly instructions are classified into multiple categories (such as GP, GP_EXT, FPU, SIMD, etc.), the number of instructions in each category and other features (such as the total number of instructions, in-degree, out-degree, etc.) are counted, a node feature vector is generated, and the feature values are processed by a normalization method to reduce the influence of program scale differences. Further, combined with an embedding mapping strategy based on eigenvalue grading, the high-dimensional feature vector is mapped into a low-dimensional space to optimize the feature distribution and improve the subsequent calculation efficiency. The Weisfeiler-Lehman Kernel method is used to calculate the similarity of the control flow graph. By iteratively updating the node labels to capture the structural features, the rapid comparison of large-scale graphs is completed with a relatively low time complexity. On this basis, combined with the cosine similarity of the program global feature vector and the Jaccard similarity of shared library functions and system calls, the similarity between programs is comprehensively evaluated. Finally, the program similarity analysis results are applied to the task scheduling of heterogeneous CPU clusters. By accurately predicting the resource requirements of the program and optimizing the task allocation according to the similarity calculation results, the utilization rate and calculation efficiency of the scheduling system for the resources of heterogeneous CPU clusters are improved.
[0019] In some embodiments: The construction of the control flow graph includes:
[0020] Use the objdump tool to disassemble the target program to obtain the assembly code in the program. Each function in the assembly code is used as a node, the instructions in each function code block are counted, the number of instructions in each category, the total number of instructions of the node, and the in-degree and out-degree of the node are extracted, and these statistical metrics are used as the feature vector of the node. And according to the jump and call relationships between each function code block, edges are added to the nodes in the control flow graph to construct the control flow graph (CFG) of the program.
[0021] In some embodiments: Instruction classification and feature extraction include:
[0022] Based on the hardware characteristics, the assembly instructions are divided into multiple categories independent of the CPU architecture, including GP, GP_EXT, GP_IN_OUT, FPU, MMX, STATE, SIMD, SSE, SCALAR, CRYPTO_HASH, AVX, AVX512, etc.
[0023] When traversing the assembly code to construct the feature vector and control flow graph of the node, the total number of instructions in each category and the total number of instructions in the program are simultaneously counted, and these metrics are used as the overall feature vector of the program.
[0024] In some embodiments, the normalization process of control flow graph node features includes:
[0025] For programs of different scales, in order to further eliminate the influence of program scale differences on feature distribution, an intermediate representation is designed to normalize the feature values of nodes into proportional values relative to the program scale, representing the importance of the current node features in the entire program. Record the maximum value and the minimum value tr i representing the value of the i-th dimension of the feature.
[0026] For the possible quantity differences between instruction categories, proportional processing is performed to make the features of each category evenly distributed.
[0027] In some embodiments, the hierarchical mapping of feature values includes:
[0028] For the non-uniform distribution characteristics of feature values (such as exponential distribution), a hierarchical partitioning strategy is designed. Each feature value range is divided into multiple intervals in an exponentially increasing manner, and the length of each interval is defined as Each interval range is The number of intervals is Lv i .
[0029] In some embodiments, the embedding generation of program feature vectors includes:
[0030] The embedding of the program is a general low-dimensional digital vector embeding, representing the characteristics of the entire program. The embedding dimension is calculated by the following formula:
[0031] Based on the results of hierarchical mapping, traverse each node in the control flow graph, and then traverse the feature vector curVec of each node to calculate the difference between each dimension of the vector curVec and the minimum value: Then find dif through binary search i At the position pos of the corresponding feature partition interval of the vector, let idx += pos * Lv i , where idx is the position where the embedding needs to be updated, and finally update the embedding, let embeding[i] += 1.
[0032] In some embodiments, the control flow graph similarity calculation includes:
[0033] For the embeddings embeding1 and embeding2 of two programs, calculate their similarity. For the values val1 and val2 of the two vectors on each dimension feature, use the following formula to calculate the similarity of the feature values: Obtain the similarity score of the control flow graph embedding vector.
[0034] In some embodiments, the normalized control flow graph embedding similarity includes:
[0035] Supplement of global information:
[0036] To make up for the defect that the Weisfeiler-Lehman kernel method may ignore global structural information, the feature vector of the overall program, shared libraries, and system call information are added.
[0037] Use the ldd tool to find the shared libraries called by the program, calculate the similarity of shared libraries using the Jaccard similarity, calculate the similarity of the feature vectors of the overall program using the cosine similarity, and calculate the weighted sum of the results with the control flow graph embedding similarity similarity to form the final program similarity score.
[0038] In some embodiments, the task deployment includes:
[0039] Compare the tasks to be deployed with the programs running in the cluster in parallel, and find the program with the highest similarity score as the reference program.
[0040] Collect the disk utilization, memory utilization, CPU utilization, network usage of each node in the cluster through Prometheus, and the usage of these four types of resources by the reference program.
[0041] Use the Z-score normalization method for processing, and normalize the extracted resource usage. Calculate the most suitable node through the cosine similarity, and deploy the task to the corresponding node.
[0042] Figure 2 For the method flow chart provided by the embodiments of the present invention, as Figure 1 shown, this embodiment provides a heterogeneous CPU cluster scheduling method, including:
[0043] Step 100 disassembles the target program through the objdump tool. The disassembled assembly code consists of multiple functions, and each assembly code function is represented as a node in the control flow graph. By analyzing the jump and call relationships between the functions in the assembly code, edges are added to the nodes of the control flow graph. The connections between the nodes reflect the jump and call relationships between the functions, and thus the control flow graph is constructed.
[0044] Step 200 To better capture the characteristics of the program, the assembly instructions are classified into general instructions, extended general instructions, input / output instructions, floating-point unit instructions, multimedia extension instructions, status control instructions, single instruction multiple data instructions, stream SIMD extension instructions, scalar instructions, cryptographic hash instructions, advanced vector extension instructions, and advanced vector extension 512 instructions according to hardware characteristics.
[0045] This classification method based on hardware implementation and technology extension is more accurate than the traditional classification based on instruction function or operation behavior. For example, for programs that require a large amount of floating-point calculations, such as scientific computing applications, FPU and AVX instructions dominate, while in image processing programs, SIMD and MMX instructions are more common.
[0046] This classification method helps to more clearly understand the utilization of hardware resources by different programs. By counting the assembly code blocks of each node, extracting the number of various instructions, the total number of instructions, as well as the in-degree and out-degree of the node, a feature vector of the node is formed. While traversing, the total number of various instructions and the total number of program instructions are counted, and these metrics are used as the overall feature vector of the program.
[0047] Step 300 Since the number of instructions in different categories varies greatly, and the scale of different programs (i.e., the number of basic blocks) varies significantly, it is unreasonable to directly compare the number of instructions. To address this issue, this embodiment conducts a more in-depth eigenvalue distribution analysis and designs a hierarchical mapping strategy.
[0048] This embodiment finds that the eigenvalues of most nodes are concentrated near the minimum value, showing an obvious exponential distribution. For this reason, an intermediate representation method is designed to normalize the feature vector of each node into a value independent of the scale of the entire program. Specifically, the eigenvalue of a node is obtained by dividing the sum of the same-type features of all nodes in the program by the eigenvalue of the current node, forming a relative ratio. This ratio represents the relative importance of the current node compared to other nodes in the entire program. A hierarchical partitioning strategy for eigenvalues is introduced, dividing each eigenvalue into several intervals to adapt to the non-uniform distribution of program sizes. Specifically:
[0049] For each feature tr i , its value range to is divided into 2 n -1 parts, and the length of each part Each interval range is The number of intervals is Lv i .
[0050] Step 400 traverses the feature vectors curVec of each node, maps them to the corresponding rank intervals, finds the positions of the feature intervals where each node's features are located, and sums all the feature positions pos i Multiply by Lv i The sum of which is used as the position of this node in the embedding, and then the value at the corresponding position is incremented. Calculate the similarity between the two programs by traversing each dimension of embeding1 and embeding2. Similarity And val2 are the values of the dimensional features at the corresponding dimensions.
[0051] In each iteration, for each pair of non-zero features val1 and val2, take the absolute value of the difference, add one, and then take the negative. This ensures that the closer the two feature values are, the higher the similarity. The calculated metric value is accumulated into sum. The main reason for using one as a scaling factor here is to prevent division by zero.
[0052] Integrate the similarity results of all feature dimensions to obtain the similarity score of the embedding vectors. Here, there are three comparison dimensions, and the values of each comparison dimension need to be scaled to between 0 and 1. Divide the accumulated similarity sum and similarity by intersection intersection represents the number of non-zero corresponding positions in the embedding of the representative program. The comprehensive measurement method based on feature intersection and feature difference is adopted here. By calculating the non-zero intersection of two vectors, the algorithm captures the structural similarity of the programs. If the embedding vectors of two programs have more shared features, they are considered to have a similar control flow structure to some extent.
[0053] For shared features, the algorithm further considers the numerical differences between them. The smaller the difference in feature values, the more similar the two programs are in that feature dimension. By using this measurement method, the algorithm balances the structural similarity and value similarity of the vectors. By weighted averaging these two parts, the comprehensive similarity between the two vectors is finally obtained
[0054] In step 500, since the WL Kernel depends on the local expansion of node labels, it may ignore global structure information. In this embodiment, the overall feature dimensions of the program and shared libraries are added for comparison. Use the ldd tool to find the shared libraries and system calls called by the program. Calculate the similarity of shared libraries and system calls using Jaccard similarity, and calculate the similarity of the overall feature vectors of the program using cosine similarity. Finally, take the weighted average of these three values as the similarity of the program.
[0055] Step 600 uses Prometheus with a 5-minute time window. The five-minute time window finds a good balance between smoothing data fluctuations and providing real-time feedback. It can effectively filter short-term anomalies while ensuring data timeliness, making it suitable for most scenarios. At the same time, the scheduler not only focuses on the resource usage of nodes but also integrates a program similarity analyzer to find the program most similar to the current task from historical data.
[0056] To uniformly compare the node resource usage information with the hardware requirements of the program, the resource requirements of the program (obtained from the historical running data of similar programs) and the real-time resource usage data of the nodes are processed using the Z-score normalization method. Finally, the cosine similarity is used to find the most suitable node. Because the smaller the cosine similarity value (close to 0 or negative), the larger the angle between the two vectors, indicating a greater difference in their directions and a lower similarity.
[0057] In this article, specific examples are used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A heterogeneous CPU cluster scheduling method based on program similarity analysis, characterized in that It includes the following steps: Step 1. Construct the control flow graph of the target program. Obtain the assembly code of the target program through a disassembler tool. Take the functions in the assembly code as nodes, and add edges to the nodes according to the jump and call relationships between functions to form a control flow graph; Step 2. Extract the feature vectors of the control flow graph nodes, including classifying and counting the assembly instructions to generate the node feature vectors and the overall feature vectors of the program; Step 3. Normalize the node feature values to eliminate the influence of program scale differences on the feature distribution; Step 4. Perform hierarchical mapping on the feature values to map the high-dimensional feature vectors to a low-dimensional space and generate the low-dimensional embedding representation of the program; Step 5. Calculate the similarity between programs. Perform similarity analysis on the control flow graph through the graph kernel method, and combine the global feature information of the programs to comprehensively evaluate the similarity between programs; Step 6. According to the results of the program similarity analysis, combine the real-time status information of the cluster nodes, select the reference program with the highest similarity to the target program, and schedule the task to the cluster node that best matches the resource usage preferences of the reference program.
2. The heterogeneous CPU cluster scheduling method according to claim 1, wherein In the said Step 2, the classification of the assembly instructions includes but is not limited to the following categories: general-purpose instructions GP, extended general-purpose instructions GP_EXT, input / output instructions GP_IN_OUT, floating-point unit instructions FPU, multimedia extension instructions MMX, single instruction multiple data instructions SIMD, stream SIMD extension instructions SSE, scalar instructions SCALAR, cryptographic hash instructions CRYPTO_HASH, advanced vector extension instructions AVX, and advanced vector extension 512 instructions AVX512.
3. The heterogeneous CPU cluster scheduling method according to claim 1, wherein In the said Step 3, the normalization process of the node feature values is achieved by dividing the feature values of the nodes by the sum of the same type of features of all nodes in the program to obtain the relative importance of each node feature in the entire program.
4. The heterogeneous CPU cluster scheduling method according to any one of claims 1 to 3, characterized in that In the said Step 4, the hierarchical mapping strategy of the feature values includes: dividing the feature value range into multiple intervals according to exponential growth, and mapping each feature value to the corresponding rank interval to adapt to the non-uniform distribution of the program scale.
5. The heterogeneous CPU cluster scheduling method according to claim 4, wherein In the said Step 5, the graph kernel method is the Weisfeiler-Lehman kernel method, which captures the structural features of the program by iteratively updating the node labels.
6. The heterogeneous CPU cluster scheduling method according to claim 4, wherein, In the said Step 5, the global feature information of the program includes the overall feature vectors of the program, shared library information, and system call information. Among them, the similarity of the shared libraries is calculated by the Jaccard similarity, and the similarity of the overall feature vectors of the program is calculated by the cosine similarity.
7. The heterogeneous CPU cluster scheduling method according to claim 6, wherein In the said Step 6, the real-time status information of the cluster nodes includes disk utilization, memory utilization, CPU utilization, and network usage, which is collected through the Prometheus tool, and the Z-score normalization method is used to process the resource usage.
8. The heterogeneous CPU cluster scheduling method according to claim 6, wherein In the said Step 6, task scheduling is achieved by calculating the similarity between the task and the programs running in the cluster, and deploying the task to the node that best matches the resource usage preferences of the reference program to optimize resource utilization and improve program performance.
9. The heterogeneous CPU cluster scheduling method according to claim 1, characterized in that, In the above-mentioned step 1, the construction of the control flow graph further includes extracting the in-degree and out-degree information of the nodes and using them as part of the node feature vector.
10. The heterogeneous CPU cluster scheduling method according to claim 1, wherein, In the above-mentioned step 4, the low-dimensional embedded representation of the program feature vector maps the eigenvalue to a specific position in the embedded space through a hash operation and updates the embedded representation to form the low-dimensional embedded representation of the program.