A Distributed Intelligent Decomposition Method for Function-Level Jobs

The function-level job distribution method optimizes job distribution in distributed computing by predicting node execution times and using graph algorithms to enhance computational efficiency and adaptability.

CN114997417BActive Publication Date: 2025-07-15ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210620107.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-07-15
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

The existing distributed computing technology mainly relies on the accumulation of computing power, and lacks effective solutions to finely match jobs and computing resources, resulting in limited improvement in computing efficiency.

Method used

Through the distributed intelligent decomposition method of function-level jobs, the machine learning prediction model is used to analyze the function call relationship, form a call relationship diagram, and calculate the key path based on the graph algorithm, insert the Ray modifier to complete the decomposition, realizing function-level decomposition and scheduling.

Benefits of technology

The computing efficiency of distributed computing systems is improved, the efficiency of computing power is improved, and the dependence on expert experience is reduced. The performance of the decomposition scheme is improved with the increase in the number of job decomposition records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114997417B_ABST
    Figure CN114997417B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical fields of computer operating systems and machine learning modeling, and discloses a method for distributed intelligent decomposition of function-level jobs. By analyzing the source code of the jobs, a directed acyclic graph based on function call relationships is formed. After combining features such as function input and output, estimating the running time of each function, the optimal decomposition plan for the jobs is obtained, and modifiers are inserted into the code to complete the decomposition. The formation of the decomposition plan does not rely too much on the expert experience in the field of job decomposition. The jobs decomposed by this plan have better computing efficiency in a distributed computing system than those without decomposition, and the performance of the decomposition method will improve as the number of job decompositions and running records increases. The present invention realizes the intelligent decomposition and automatic scheduling of large jobs, speeds up the operation speed of computationally intensive jobs, and improves the utilization efficiency of the computing power of the distributed system. Its implementation method is flexible and reliable, has very little intrusion into the job source code, and achieves high computing efficiency at a low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer operating systems and machine learning modeling, and particularly relates to a method for distributed intelligent decomposition of function-level jobs. Background Art

[0002] With the development of big data and artificial intelligence technologies, the storage and use of massive amounts of data have led to a huge increase in the demand for computer computing power in the industry and academia. Moreover, the growth rate of the demand for computing power in related fields is much higher than the development rate of computing power itself. This makes computing power, to some extent, one of the bottlenecks in the development of artificial intelligence, especially in the field of deep learning. Driven by such demands, various distributed computing technologies have emerged. By connecting more computing devices, they bring more computing power to users: Grid computing and volunteer computing have established a computing power sharing mechanism, enabling organizations and individuals to share their own idle computing power with organizations in need of more computing power, thus realizing the effective utilization of computing power resources; the popularization and rise of commercial cloud computing enable consumers to rent computing resources in an economical and convenient way, alleviating the pressure brought by computing power demands and hardware costs.

[0003] Generally speaking, the above-mentioned distributed technologies mainly rely on the accumulation of computing power in terms of quantity. However, there is no effective solution for how to maximize the role of computing power and improve computing efficiency by precisely matching jobs with computing resources. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for distributed intelligent decomposition of function-level jobs to solve the above-mentioned technical problems.

[0005] To solve the above technical problems, the specific technical solution of a method for distributed intelligent decomposition of function-level jobs of the present invention is as follows:

[0006] A method for distributed intelligent decomposition of function-level jobs includes the following steps:

[0007] Step 1: The user selects an existing prediction model in the system, or obtains a job running time prediction model in an offline training manner in advance based on sandbox experimental data and machine learning regression algorithms for subsequent steps;

[0008] Step 2: The user submits the job source code to the job decomposition system;

[0009] Step 3: The job decomposition system analyzes the function call relationships in the job source code and the data sources of each stage of the job to form a call relationship graph;

[0010] Step 4: Based on the feature information of each node and the model generated or selected by the user in Step 1, predict the expected running time of each node in Step 3 as the weight of each edge in the call relationship graph;

[0011] Step 5: Use graph algorithms to calculate the critical path from start to end in the call graph;

[0012] Step 6: Based on the calculation results of Step 4, group the functions that are on the paths connecting the critical path nodes but not on the critical path into groups. The functions within each group will be scheduled to run on the same computing node, thereby forming a function-level decomposition scheme;

[0013] Step 7: Insert Ray modifiers at the corresponding positions in the code to complete the decomposition.

[0014] Furthermore, when there is an existing function running time prediction model in the system in Step 1, the user can directly select the existing model for subsequent steps, and Step 1 ends accordingly.

[0015] Furthermore, when there is no ready-made model available in Step 1, or the user wishes to train the model by themselves, the model is obtained through the following steps:

[0016] Step 1.1: Model training data preparation: When there is no real historical data of any job running in the system and for the user, the system relies on sandbox experiments to obtain some function input, output, code volume, code structure, and corresponding running time records to solve the cold start problem; when there are existing job running records in the system, the system uses this information to retrain the model and develop a model with higher prediction accuracy based on the existing model.

[0017] Step 1.2: Select a regression algorithm and train: Arbitrarily select one or more from machine learning regression algorithms for training. The specific method is to use the training data in Step 1.1 to fit and validate the selected algorithm under one or more sets of hyperparameters, and obtain the best-performing algorithm and parameter combination as the finally selected prediction model.

[0018] Furthermore, each node in the call graph in Step 3 represents a function or an input data set, and the nodes are connected by directed edges, representing the sequential relationship of node operations; at the same time, the system collects the input, output, code volume, code structure information, and function running time of each node and archives them.

[0019] Furthermore, Step 4 specifically includes the following steps:

[0020] Step 4.1: The system captures the input, output, code volume, code structure, and function running time information of the functions in the current job during the actual operation process, uses the data to predict the running time of each function node in real time after processing according to certain specifications, and archives it for offline retraining of the model;

[0021] Step 4.2: The system processes and transforms the data collected in Step 4.1 in real time until the data format is exactly the same as the input format used for model training in Step 1. Then, the fully processed data is input into the model selected or trained in Step 1 to obtain the predicted running time of the corresponding node output by the model.

[0022] Step 4.3: The user manually or the system automatically retrains the function running time prediction model at regular intervals based on the archived data in Step 4.1 using the original algorithm or a new regression algorithm, continuously improving the accuracy of the function running time prediction model through repeated iterations.

[0023] Further, the said Step 5 includes the following specific steps:

[0024] Step 5.1: If there is only one end node in the call relationship graph to be decomposed, directly calculate the critical path from the start node to the end node.

[0025] Step 5.2: If there are multiple end nodes in the call relationship graph to be decomposed, segment the call relationship graph by the end nodes. Each segment includes an end node and all its preceding nodes. Different segments may have the same repeated nodes. Calculate the critical path of each segment using the method in Step 5.1, and integrate them as the critical path of the entire call relationship graph.

[0026] A function-level job distributed intelligent decomposition method of the present invention has the following advantages: The formation of the decomposition scheme of the present invention does not rely too much on the expert experience in the field of job decomposition. The jobs decomposed by this scheme have better computing efficiency in the distributed computing system than those without decomposition, and the performance of the decomposition method will improve as the number of job decompositions and operation records increases. The present invention uses the open-source software Ray to achieve the intelligent decomposition and automatic scheduling of large jobs, accelerating the operation speed of compute-intensive jobs and improving the utilization efficiency of the computing power of the distributed system. Its implementation method is flexible and reliable, with minimal intrusion into the job source code, achieving high computing efficiency at a low cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flowchart of the function-level job distributed intelligent decomposition method of the present invention;

[0028] Figure 2 It is an example diagram of the decomposition of the function call relationship graph of the present invention;

[0029] Figure 3 It is the principle diagram of the prediction algorithm and model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To better understand the purpose, structure, and function of the present invention, the following further describes in detail a method for distributed intelligent decomposition of function-level jobs of the present invention in conjunction with the accompanying drawings.

[0031] As Figure 1 shown, a method for distributed intelligent decomposition of function-level jobs of the present invention includes the following steps:

[0032] Step 1: The user selects a pre-existing prediction model in the system or obtains a job running time prediction model in an offline training manner in advance based on sandbox experiment data and machine learning regression algorithms for subsequent steps.

[0033] 1.1: When there is a function running time prediction model in the system, the user can directly select the existing model for subsequent steps, and Step 1 ends here; when there is no ready-made model available or the user wishes to train the model by themselves, the model is obtained through training by the following steps.

[0034] 1.2: Preparation of model training data: When there is no real historical data of any job running in the system and the user, the system can rely on sandbox experiments to obtain some function input, output, code volume, code structure, and corresponding running time records to solve the cold start problem; when there are existing job running records in the system, the system can use this information to retrain the model and develop a model with higher prediction accuracy based on the existing model.

[0035] 1.3: Select a regression algorithm and train: The core feature of the regression algorithm is that its prediction target is a continuous variable. Therefore, any one or more of linear regression, LASSO, random forest regression, XGBoost regression, or other machine learning regression algorithms can be selected to train a function running time prediction model. According to different difficulties, workload inputs, output model accuracies, etc., the following three different training methods are used for training: (1) Simple training based on one of the selected machine learning regression algorithms; (2) Training after debugging various hyperparameters for the selected algorithm and selecting the best hyperparameters based on the same data; (3) Selecting the best algorithm + hyperparameter combination for combined training after selecting multiple algorithm and hyperparameter combinations. After training, select the best model for subsequent steps, and Step 1 ends here.

[0036] Step 2: The user submits the job source code to the job decomposition system.

[0037] Step 3: The job decomposition system analyzes the function call relationships in the job source code and the data sources of each stage of the job to form a call relationship diagram, as Figure 2As shown, where each node represents a function or an input data set, and the nodes are connected by directed edges, representing the sequential relationship of node execution; meanwhile, the system collects the input, output, code volume, code structure information, and function running time of each node and archives them;

[0038] Step 4: Based on the feature information of each node and the model generated or selected by the user in Step 1, predict the expected running time of each node in Step 3 as the weight of each edge in the call relationship graph; note that a function is usually defined only once but can be called multiple times, so a function may appear as a node multiple times in the call relationship graph with edges of different weights;

[0039] 4.1: During the actual operation of the system, the input, output, code volume, code structure, and function running time information of the functions in the current job are captured for two purposes: 1. To predict the running time of each function node in real time after processing the data according to certain specifications, and 2. To archive for offline retraining of the model;

[0040] 4.2: As Figure 3 shown, the system processes and transforms the data collected in 4.1 in real time until the data format is exactly the same as the input format used for model training in Step 1. Input the fully processed data into the model selected or trained in Step 1, and the predicted running time value of the corresponding node output by the model can be obtained;

[0041] 4.3: The user manually or the system automatically retrains the function running time prediction model regularly based on the archived data in 4.1 using the original algorithm or a new regression algorithm, continuously improving the accuracy of the function running time prediction model through repeated iterations, thereby improving the overall performance of the job decomposition system.

[0042] Step 5: Use graph algorithms to calculate the critical path from start to end in the call relationship graph;

[0043] 5.1: If there is only one end node in the call relationship graph to be decomposed, directly calculate the critical path from the start node to the end node;

[0044] 5.2: If there are multiple end nodes in the call relationship graph to be decomposed, segment the call relationship graph by end nodes. Each segment includes an end node and all its preceding nodes (the same repeated nodes can exist in different segments). Calculate the critical path of each segment using the method in 5.1 and integrate them as the critical path of the entire call relationship graph.

[0045] Step 6: Based on the calculation results in Step 4, group the functions that are on the paths connecting the nodes of the critical path but not on the critical path. The functions within each group will be scheduled to run on the same computing node, thereby forming a function-level decomposition scheme;

[0046] Step Seven: Insert the Ray modifier at the corresponding position of the code to complete the decomposition.

[0047] It can be understood that the present invention is described by way of some embodiments. Those skilled in the art will know that, without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A distributed intelligent decomposition method for function-level jobs, characterized in that, It includes the following steps: Step 1: The user selects an existing prediction model in the system or obtains a job running time prediction model in an offline training manner in advance based on sandbox experiment data and machine learning regression algorithms for subsequent steps; Step 2: The user submits the job source code to the job decomposition system; Step 3: The job decomposition system analyzes the function call relationships in the job source code and the data sources of each stage of the job to form a call relationship graph; Step 4: Based on the feature information of each node and the model generated or selected by the user in Step 1, predict the expected running time of each node in Step 3 as the weight of each edge in the call relationship graph; Step 5: Use a graph algorithm to calculate the critical path from start to end in the call relationship graph; Step 6: Based on the calculation result of Step 5, group the functions that are on the path connecting the nodes of the critical path but not on the critical path. The functions within each group will be scheduled to run on the same computing node, thereby forming a function-level decomposition scheme; Step 7: Insert Ray modifiers at the corresponding positions in the code to complete the decomposition.

2. The function-level job distributed intelligent decomposition method according to claim 1, wherein When there is an existing job running time prediction model in the system in Step 1, the user can directly select the existing model for subsequent steps, and Step 1 ends here.

3. The function-level job distributed intelligent decomposition method according to claim 1, wherein When there is no ready-made model to select in Step 1, or the user hopes to train the model by themselves, the model is obtained through the following steps: Step 1.1: Model training data preparation: When there is no real historical data of any job running in the system and for the user, the system relies on sandbox experiments to obtain some function input, output, code volume, code structure, and corresponding running time records to solve the cold start problem; When there are existing job running records in the system, the system uses this information to retrain the model and develop a model with higher prediction accuracy based on the existing model; Step 1.2: Select a regression algorithm and train: Arbitrarily select one or more from machine learning regression algorithms for training. The specific method is to use the training data in Step 1.1 to fit and verify the selected algorithm under one or more sets of hyperparameters, and obtain the best-performing algorithm and parameter combination as the finally selected prediction model.

4. The function-level job distributed intelligent decomposition method according to claim 1, wherein Each node in the call relationship graph in Step 3 represents a function or an input data set, and the nodes are connected by directed edges, representing the sequential relationship of node operations; Meanwhile, the system collects the input, output, code volume, code structure information, and function running time of each node and archives them.

5. The function-level job distributed intelligent decomposition method according to claim 1, wherein Step 4 specifically includes the following steps: Step 4.1: The system captures the input, output, code volume, code structure, and function running time information of the functions in the current job during actual operation, which is used to predict the running time of each function node in real time after data is processed according to certain specifications and archived for offline retraining of the model; Step 4.2: The system processes and transforms the data collected in Step 4.1 in real time until the data format is exactly the same as the input format used for model training in Step 1, and inputs the fully processed data into the model selected or trained in Step 1 to obtain the running time prediction value of the corresponding node output by the model; Step 4.3: The user manually or the system automatically and regularly retrains the job running time prediction model based on the archived data in Step 4.1 using the original algorithm or a new regression algorithm, and continuously improves the accuracy of the function running time prediction model in repeated iterations.

6. The function-level job distributed intelligent decomposition method according to claim 1, characterized in that The said Step 5 includes the following specific steps: Step 5.1: If there is only one end node in the call relationship graph to be decomposed, directly calculate the critical path from the start node to the end node; Step 5.2: If there are multiple end nodes in the call relationship graph to be decomposed, segment the call relationship graph by end nodes. Each segment includes an end node and all its preceding nodes. Different segments may have the same repeated nodes. Calculate the critical path of each segment using the method in Step 5.1, and integrate them as the critical path of the entire call relationship graph.

Citation Information

Patent Citations

  • Distributed streaming data processing method and system based on microkernel operating system

    CN110532072A

  • Method and device for predicting running time in batch processing operation and electronic equipment

    CN112906971A