A method and system for model training task allocation in a heterogeneous computing power environment
By building a hierarchical deep reinforcement learning model and resource perception module, optimizing split point selection and resource allocation, the problem of uneven allocation of computing resources in heterogeneous computing power environment is solved, and efficient computing task allocation and model training is achieved.
Patent Information
- Application Number
- CN202411975476.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the heterogeneous computing power environment, improper selection of split points in the existing methods leads to uneven allocation of computing resources, resulting in inefficient training of large models.
A hierarchical deep reinforcement learning model is built, and through multiple sub-strategy networks and resource perception modules, the split point selection and resource allocation between the terminal, edge layer and cloud computing layer are optimized, combined with a heuristic algorithm to initially estimate the split point range, and iteratively find local and overall optimal split points.
Efficient computing task allocation and model training in a heterogeneous computing power environment are realized, training efficiency is improved, resource utilization is optimized, irrelevant location exploration is reduced, and overall optimal task allocation is achieved.
Smart Images

Figure CN119376958B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed training, and specifically relates to a method and system for model training task allocation in a heterogeneous computing power environment. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] In the era of large models, the model parameters and data scale grow exponentially. Although the total amount of computing power resources develops rapidly, the computing power provided by each computing power provider is unevenly distributed. Under such non-uniform computing power distribution conditions, how to better support the training of large models and give full play to the advantages of current computing power resources has very important research significance and value.
[0004] Split computing is a new computing mode that splits a large model and deploys it on a distributed cluster for parallel training. It can improve the training efficiency of large models and make full use of fragmented computing resources, which is the current technical development trend for large model training under distributed computing power conditions. However, in existing methods, the selection of split points varies, resulting in uneven distribution of computing resources and thus low training efficiency. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a method and system for model training task allocation in a heterogeneous computing power environment. The present invention establishes a hierarchical deep reinforcement learning (DRL) model. Taking the terminal-edge layer as an example, different sub-policy networks learn and optimize for different computing power resource requirements respectively, and realize the selection of split points between levels by combining deep reinforcement learning, achieving overall serial and parallel between layers. The global policy is updated after each training cycle to realize real-time task allocation and optimization. The key to the research is how to effectively combine deep reinforcement learning and a resource perception module to optimize the split point selection and task allocation strategy between the terminal and edge layers and achieve overall optimality.
[0006] According to some embodiments, the first solution of the present invention provides a method for model training task allocation in a heterogeneous computing power environment, adopting the following technical solution:
[0007] A method for model training task allocation in a heterogeneous computing power environment, comprising:
[0008] Under the computing architecture composed of the terminal layer - edge layer - cloud computing layer, obtain multiple model training requests of the terminal layer, and group multiple terminals according to the size relationship between the data volume of the model training requests and a set threshold;
[0009] Based on the computing power information at the edge layer and the computing power requirements of the model training request, use the comprehensive evaluation function to find the best edge matching node for each terminal;
[0010] According to the grouping situation of the model training request, determine the preliminary selection range of the splitting point in the model. Within the determined preliminary selection range, based on the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the splitting point, iteratively search for the local optimal splitting point in a loop;
[0011] According to the grouping situation of the model training request, determine the preliminary selection range of the splitting point in the model. Within the determined preliminary selection range, based on the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the splitting point, iteratively search for the edge-cloud optimal splitting point in a loop;
[0012] Overall plan the local optimal splitting point and the edge-cloud optimal splitting point to realize the allocation of the training task.
[0013] Furthermore, the comprehensive evaluation function is specifically:
[0014] ;
[0015] Among them, is the computing resource of the edge node ; is the computing requirement of the current terminal task for the node, is the network bandwidth from the terminal to the edge node, is the network delay, is the adjustment parameter, which is used to balance the weights of computing power, bandwidth, delay and energy consumption, and is for the current user request , There is an edge node within its search range ; Load balancing factor , is the energy consumption generated by selecting the current node, node selection frequency ;
[0016] The larger the value of the comprehensive evaluation function, the higher the matching degree between the terminal and the edge node.
[0017] Furthermore, use a hierarchical deep reinforcement learning model to solve the iterative search process. Among them, the hierarchical deep reinforcement learning model includes a state space, an action space, and an overall reward function;
[0018] The state space includes the computing resource state, network bandwidth and delay of the terminal layer; the computing resource state of the edge layer and the connection state with the cloud server, and the computing resource state of the cloud computing layer;
[0019] The action space includes corresponding splitting schemes for various condition divisions;
[0020] The overall reward function includes the reward function of the edge and cloud computing layers and the reward function of the terminal and edge layers.
[0021] Furthermore, the multiple terminals are grouped according to the magnitude relationship between the data volume of the model training request and the set threshold, specifically:
[0022] Calculate the data volume of the model training request for each terminal in turn;
[0023] If the data volume of the model training request is greater than the set threshold, it is a high computing power demand level task and is classified into the high computing power demand set;
[0024] If the data volume of the model training request is less than the set threshold, it is a low computing power demand level task and is classified into the low computing power demand set;
[0025] Thus, the model training requests of all terminals in the terminal layer are divided into two groups: the high computing power demand set and the low computing power demand set.
[0026] Furthermore, the preliminary selection range of the splitting point is determined in the model according to the grouping situation of the model training request. Within the determined preliminary selection range, with the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the splitting point, iteratively loop to find the local optimal splitting point, specifically:
[0027] If the data volume of the model training request is less than the set threshold, it is a low computing power demand level task, and the range after the splitting point in the middle layer of the model structure is used as the preliminary selection range of the splitting point; otherwise, it is a high computing power demand level task, and the range before the splitting point in the middle layer of the model structure is used as the preliminary selection range of the splitting point;
[0028] Taking one splitting point as an example within the preliminary selection range, calculate the delay and energy consumption between the terminal and the edge matching node at the current splitting point;
[0029] Iteratively loop through all splitting points within the preliminary selection range, with the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the splitting point, to determine the local optimal splitting point.
[0030] Furthermore, the preliminary selection range of the splitting point is determined in the model according to the grouping situation of the model training request. Within the determined preliminary selection range, with the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the splitting point, iteratively loop to find the edge-cloud optimal splitting point, specifically:
[0031] If the data volume of the model training request is less than the set threshold, it is a task with low computing power requirement level, and the range after the splitting point in the middle layer of the model structure is used as the preliminary selection range of the splitting point; otherwise, it is a task with high computing power requirement level, and the range after the splitting point in the middle layer of the model structure is used as the preliminary selection range of the splitting point;
[0032] Within the preliminary selection range, taking a splitting point as an example, calculate the latency and energy consumption between the cloud computing layer and the edge matching node when calculating the current splitting point;
[0033] Iteratively loop through all splitting points within the preliminary selection range, and determine the locally optimal splitting point with the goal of minimizing the latency and energy consumption between the edge matching node and the cloud computing layer at the splitting point.
[0034] Furthermore, the overall coordination of the locally optimal splitting point and the edge-cloud optimal splitting point is used to allocate the training task, specifically as follows:
[0035] Based on the locally optimal splitting point, split the corresponding inter-layer position of the model's structure layer;
[0036] Based on the edge-cloud optimal splitting point, split the corresponding inter-layer position of the model's structure layer;
[0037] Divide the model's structure layer into three structure blocks, and further split the same model training request task into three segments for parallel training to achieve the allocation of the training task.
[0038] According to some embodiments, the second solution of the present invention provides a model training task allocation system in a heterogeneous computing power environment, adopting the following technical solution:
[0039] A model training task allocation system in a heterogeneous computing power environment, comprising:
[0040] A task computing power identification module, configured to obtain multiple model training requests at the terminal layer under the computing architecture composed of the terminal layer - edge layer - cloud computing layer, and group multiple terminals according to the size relationship between the data volume of the model training request and the set threshold;
[0041] A matching node determination module, configured to find the best edge matching node for each terminal by using a comprehensive evaluation function based on the computing power information of the edge layer and the computing power requirements of the model training request;
[0042] A terminal-edge splitting module, configured to determine the preliminary selection range of the splitting point in the model according to the grouping situation of the model training request, and within the determined preliminary selection range, iteratively loop to find the locally optimal splitting point with the goal of minimizing the latency and energy consumption between the terminal and the edge matching node at the splitting point;
[0043] An edge-cloud splitting module, configured to determine a preliminary selection range of splitting points in a model according to the grouping situation of model training requests, and within the determined preliminary selection range, iteratively search for the optimal edge-cloud splitting point with the goal of minimizing the latency and energy consumption between edge matching nodes and the cloud computing layer at the splitting point;
[0044] A training task allocation module, configured to coordinate the local optimal splitting point and the optimal edge-cloud splitting point to achieve the allocation of training tasks.
[0045] According to some embodiments, the third aspect of the present invention provides a computer-readable storage medium.
[0046] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a method for allocating model training tasks in a heterogeneous computing power environment as described in the first aspect above.
[0047] According to some embodiments, the fourth aspect of the present invention provides a computer device.
[0048] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in a method for allocating model training tasks in a heterogeneous computing power environment as described in the first aspect above.
[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0050] The present invention constructs a hierarchical deep reinforcement learning architecture for optimizing the selection of splitting points and resource allocation between the terminal layer, the edge layer, and the cloud computing layer, so as to achieve efficient computing task allocation and model training. Multiple sub-policy networks are designed to handle tasks with different computing power requirements. A resource awareness module is combined to measure the computing resources, network bandwidth, latency, and other states of each layer. A heuristic algorithm is used to initially estimate the range of splitting points, reduce the exploration of irrelevant positions, and optimize the search efficiency. In the first stage, splitting points are roughly selected to screen out better regions; in the second stage, these regions are finely optimized, thereby gradually improving the training accuracy. Through a reward signal integration mechanism, the rewards of high-computing-power and low-computing-power tasks are aggregated, and the global policy is updated after each control cycle. Real-time task allocation and optimization are achieved through a joint policy, enabling the system to achieve global optimality in a multi-level architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention, and the schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0052] Figure 1 It is a schematic diagram of the hierarchical DRL model calculation architecture in an embodiment of the present invention;
[0053] Figure 2 It is a flowchart for selecting the optimal splitting point in an embodiment of the present invention. Detailed implementation manners
[0054] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0055] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0056] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0057] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0058] Term explanation:
[0059] Heterogeneous computing power: There are differences in the computing power that computing power providers can offer, and there are differences in the model training tasks requested by users.
[0060] Embodiment 1
[0061] This embodiment provides a method for allocating model training tasks in a heterogeneous computing power environment. In this embodiment, taking the application of this method to a server as an example, it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is realized through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. In this embodiment, the method includes the following steps:
[0062] Under the computing architecture composed of the terminal layer, edge layer, and cloud computing layer, obtain multiple model training requests from the terminal layer, and group multiple terminals according to the size relationship between the data volume of the model training request and the set threshold;
[0063] Based on the computing power information of the edge layer and the computing power requirements of the model training request, use a comprehensive evaluation function to find the best edge matching node for each terminal;
[0064] According to the grouping situation of the model training request, determine the initial selection range of the split point in the model. Within the determined initial selection range, based on the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the split point, iteratively and cyclically search for the local optimal split point;
[0065] According to the grouping situation of the model training request, determine the initial selection range of the split point in the model. Within the determined initial selection range, based on the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the split point, iteratively and cyclically search for the optimal edge-cloud split point;
[0066] Overall plan the local optimal split point and the optimal edge-cloud split point to achieve the allocation of training tasks.
[0067] The respective computing delays of the terminal layer, edge layer, and cloud computing layer are affected by the parameters and quantities of the deep neural network model and their respective computing capabilities. The resources required for neural network model computing are measured by the number of floating-point operations. In the selection of the split point, mainly adopt the method of deep reinforcement learning combined with a resource perception module, and find the optimal split point through structural optimization to achieve the effective splitting of the model and complete efficient joint training on both ends.
[0068] This embodiment is described from the following three aspects, specifically including:
[0069] 1. Construction of a hierarchical deep reinforcement learning model.
[0070] Taking the terminal layer and the edge layer as an example:
[0071] First, define its state space, including information such as the computing resource status of each layer (such as CPU usage rate, memory utilization rate), network bandwidth, delay, and the current computing task requirements. Represents the computing power of the terminal layer (number of CPU cores, utilization rate), network bandwidth, and delay. Represents the computing resource status of the edge server and the connection status with the cloud server, Represents the computing resource status of the cloud computing layer.
[0072] Secondly, define its action space or search space, which covers the corresponding splitting schemes divided according to various conditions. According to the diversity of its structure and the characteristics of global search, explore as many possible splitting points as possible. Specifically, the current state is three-layer, and the optimal splitting points for the two splitting points of "terminal-edge" and "edge-cloud" need to be found. Here, based on resource awareness, select a suitable search method, and reasonably select splitting points according to the characteristics of unbalanced computing power distribution to achieve real-time global optimization, and tentatively determine the current optimal splitting point.
[0073] In each model training cycle, the model is split according to the splitting point set in this cycle, and the splitting performance is evaluated for this splitting point, including aspects such as system performance, communication overhead, and computing efficiency after splitting. Set a reward function to evaluate the value generated by each current action, and continuously update the parameters of the model. After the control cycle composed of multiple training cycles is completed, update its policy function.
[0074] Such as Figure 1 shown, this embodiment specifically designs a multi-level DRL agent architecture (Deep Reinforcement Learning, DRL). Specifically, the multi-agent or hierarchical DRL method in DRL can be used to optimize the splitting points and resource allocation in the multi-level architecture. When measuring the resources consumed by model calculation, it is in units of the number of floating-point operations (FLOPS).
[0075] Similarly, taking the terminal-edge layer as an example, first set a splitting point at this layer , indicating that this splitting point is the splitting point of the model selected by the th user device between the terminal (UE) and the edge layer in its th control cycle. Among them , here is the set of layers. When the splitting point is selected at the extreme of the terminal layer at the splittable position, that is , in this case, all the computing tasks are handed over to the edge layer after the splitting point. Taking the DNN model training task as an example, the terminal layer has no computing for the DNN model training task. At this time, the computing delay of the terminal layer is 0, and all the computing tasks are handed over to the edge layer. And if , then the splitting point is placed at the end, that is, all the computing tasks are handed over to the terminal layer.
[0076] Before distinguishing the splitting points by latency, the selection of splitting points can be determined according to a certain current request volume at the terminal layer, that is, a heuristic algorithm is used for preliminary estimation before training to quickly find a preliminary range of splitting points. By this heuristic method, the search space is reduced, and the training of the model is concentrated within the possible optimal range of splitting points, thereby reducing the exploration time for irrelevant positions.
[0077] 2. Training process:
[0078] The first stage
[0079] By dividing the training process into multiple stages, each stage focuses on different splitting point precisions. In the first stage, the splitting point positions are roughly selected to screen out better regions.
[0080] If the data volume requested for model training is less than the set threshold, it is a task with low computing power demand level, and the preliminary selection range of the splitting point is after the splitting point in the middle layer of the model structure;
[0081] Conversely, it is a task with high computing power demand level, and the preliminary selection range of the splitting point is before the splitting point in the middle layer of the model structure.
[0082] Here, it means that first, multiple terminals are grouped according to the size relationship between the data volume requested for model training and the set threshold, divided into a high computing power demand set and a low computing power demand set, and then according to this grouping situation, the preliminary selection range of the splitting point is divided before or after the splitting point in the middle layer of the entire structure layer of the model.
[0083] It can be understood that in the model training request sent by the terminal, each layer structure layer of the model serves as a splitting point, and by comparing the latency and energy consumption of different splitting points, it is measured whether the current splitting point is better or worse than other selected splitting points.
[0084] Taking one splitting point as one training cycle, all the splitting points within this preliminary selection range form a control cycle, that is, one control cycle contains multiple training cycles.
[0085] The second stage
[0086] (1)Calculate latency
[0087] If the splitting point is selected between the terminal-edge levels, the transmission latency of the terminal layer can be expressed as:
[0088] (1);
[0089] Among them, is the data volume of the user device in the layer of the model structure, That is the total data to be processed for the entire terminal layer, indicating the number of FLPOS that each core in the terminal layer can process, indicating the number of cores within this control cycle table. That is, the computing latency is obtained by dividing the total number of FLOPS required to complete the task by the number of FLOPS that can be processed simultaneously per second in the current control cycle.
[0090] The computing task part is handed over to the edge layer. The computing latency of the edge layer can be expressed as:
[0091] (2);
[0092] Among them, is the total data to be processed for the entire edge layer, indicating the number of FLPOS that each core in the edge layer can process, indicating the number of cores within this control cycle table. It should be noted that the model structure layer involved in splitting is only used for splitting, not for computing, in order to avoid duplicate computing.
[0093] Similarly, the computing latency of the edge-cloud computing layer also has a certain relationship with the choice of its splitting point. Therefore, the computing latency of the cloud computing layer is specifically:
[0094] (3);
[0095] Among them, is the total data to be processed for the entire cloud server in the cloud computing layer, is the number of FLPOS that each core providing computing in the cloud server of the current cloud computing layer can process, then indicates the number of cores provided according to its policy function within this control cycle.
[0096] (2) Transmission latency
[0097] The transmission latency between layers is the second point to be considered. Similarly, starting from the terminal-edge layer.
[0098] The latency generated by the uplink transmission from the terminal-edge layer is:
[0099] (4);
[0100] Among them, is the user equipment the uplink transmission latency generated at the terminal-edge, indicating that on this terminal side based on The amount of data transmitted at this splitting point represents the uplink transmission rate on the terminal side while [[0000256]] is the inherent delay during the uplink transmission
[0101] Within each training cycle, in addition to the forward transmission of data from the previous layer to the next layer, the next layer also needs to return a certain amount of feedback to the previous layer to optimize the model parameters. Here, the layers refer to the terminal layer, the edge layer, and the cloud computing layer
[0102] (5);
[0103] Among them, is the downlink transmission delay generated by the terminal - edge while [[0000266]] is the amount of data sent from the edge node under this splitting point to the terminal layer and [[0000268]] is the downlink transmission rate of this edge layer while [[0000269]] is the inherent delay in the downlink transmission between this edge layer
[0104] Similarly, the uplink transmission delay between the edge - cloud computing layer is:
[0105] (6);
[0106] Among them, is the uplink transmission delay generated by the edge - terminal while [[0000279]] represents the amount of data transmitted at this splitting point on this edge side and [[0000281]] represents the uplink transmission rate of this edge side while [[0000282]] is the inherent delay during the uplink transmission;
[0107] The downlink transmission delay between the edge - cloud computing layer is:
[0108] (7);
[0109] Among them, is the downlink transmission delay generated by the edge layer - cloud computing layer while [[0000292]] is the amount of data sent from the cloud computing layer under this splitting point to the edge layer and [[0000294]] is the downlink transmission rate of this cloud computing layer while [[0000295]] is the inherent delay in the downlink transmission of this cloud computing layer
[0110] During a complete training (epoch) cycle, the model will perform a forward pass and a backward pass on the entire dataset, and each pass will call the above four formulas once. It should be noted that if the dataset is divided into multiple batches for transmission during this training cycle, then each transmission needs to call the uplink and downlink formulas for model data transmission.
[0111] After the training of a control cycle consisting of multiple training cycles is completed, it is necessary to summarize and evaluate the results of multiple training tasks. By transmitting the finally updated model parameters to the end-user device in the terminal layer, the user device can have the latest trained model to achieve parameter synchronization.
[0112] (8);
[0113] Among them, is the total amount of data to be processed by the entire cloud server in the cloud computing layer, means that the split point between the edge and the cloud is not at the end but in the middle of the layer, is an update delay for one control cycle of the model between the edge-cloud computing layer;
[0114] The above formula means that if the split point between the edge-cloud computing layer is before the last layer of the DNN model, that is, there is a need for transmission in the cloud computing layer, the model strategy that needs to be transmitted and updated included in the cloud computing layer is transmitted to the edge layer, and divided by the current downlink transmission rate And then considering the current inherent transmission delay . On the contrary, there is no need for transmission and the data transmission delay is 0. Different from the previous downlink transmission, which is to find the local optimum of a single training cycle, the purpose here is to overall consider multiple training results to find the global optimum.
[0115] Similarly, when transmitting content from the edge layer to each user device in the terminal layer, specifically:
[0116] (9);
[0117] Among them, means that the split point between the terminal layer and the edge layer is not at the end but in the middle of the layer, is an update delay for one control cycle of the model between the terminal-edge layer;
[0118] Through the update of the model parameters after DNN model training, parameter synchronization is achieved.
[0119] User device The total delay generated by the user device
[0120] (10);
[0121] Among them, indicates that a control period contains training cycles. Whether it is to find the splitting point between the terminal layer and the edge layer or the splitting point between the edge layer and the cloud computing layer, it is to sum up the corresponding energy consumption and then find the minimum sum of the total delay generated within a training cycle as the goal.
[0122] (3) Energy consumption
[0123] If delay is the efficiency that is continuously updated and pursued, energy consumption is a constraint condition that needs to be considered when pursuing efficiency. The energy consumption generated by the system is also a key point of concern. The energy consumption mainly consists of two aspects, namely computing energy consumption and transmission energy consumption. This embodiment only considers the dynamic power consumption of the device during execution. Taking the terminal layer as an example:
[0124] , after arrangement, it can be obtained:
[0125] (11);
[0126] Among them, indicates the energy consumption calculated according to the splitting point. Specifically, is the capacitive load of the processor, is the operating voltage of the processor, is the processing frequency of the processor. is the ratio of the data volume to the uplink transmission, that is, the time consumed by data transmission, is the transmission power between the current terminal layer and the edge layer, represents the energy consumption generated by the terminal layer in one training cycle.
[0127] Similarly, the energy consumption of the edge layer can be expressed as:
[0128] (12);
[0129] Different from the terminal layer which only needs to consider its own computing energy consumption and a section of uplink transmission energy consumption, represents the energy consumption generated by the edge layer in one training cycle. The edge layer includes its own computing energy consumption, downlink transmission energy consumption with the terminal layer, uplink transmission energy consumption with the cloud computing layer, and a for realizing synchronization in the current training cycle, indicates that the energy consumption in the current training cycle is of the entire control cycle, is the transmission power between the current edge layer and the cloud computing layer.
[0130] Similarly, the energy consumption of the cloud computing layer can also be expressed as:
[0131] (13);
[0132] represents the energy consumption generated by the cloud computing layer in one training cycle. Whether it is to find the splitting point between the terminal layer and the edge layer or the splitting point between the edge layer and the cloud computing layer, the goal is to find the minimum sum of the corresponding energy consumptions.
[0133] 3. Node and splitting point selection:
[0134] As Figure 2 shown, through the requests of the terminal layer, the computing power requirement level of the task is identified based on the data volume (FLOPS). That is, before the DRL model starts training, the input tasks are preprocessed and classified into two categories: high computing power requirements and low computing power requirements.
[0135] (14);
[0136] Among them, contains the set of all user requests, contains the set of all user requests, is the threshold set for the data volume.
[0137] Based on formula (14), multiple terminals are grouped according to the size relationship between the data volume of the model training request and the set threshold. Specifically:
[0138] Calculate the data volume of the model training request for each terminal in turn;
[0139] If the data volume of the model training request is greater than the set threshold, it is a task with high computing power requirements and is classified into the high computing power requirement set;
[0140] If the data volume of the model training request is less than the set threshold, it is a task with low computing power requirements and is classified into the low computing power requirement set;
[0141] Thus, the model training requests of all terminals in the terminal layer are divided into two groups: the high computing power requirement set and the low computing power requirement set.
[0142] Different from the latency generated by the calculation, before transmission, an edge node that matches the current computing power requirement of the terminal needs to be found. That is, in view of the uneven distribution of computing power, a resource perception model is established so that the terminal device can evaluate the current computing power, network bandwidth, and latency of each edge node, and a function is defined to help find the edge nodes within the range.
[0143] (15);
[0144] Wherein, is the computing resource of the edge node of is the computing requirement (task volume) of the current terminal task for the node, is the network bandwidth from the terminal to the edge node, is the network latency, is a tuning parameter for balancing the weights of computing power, bandwidth, latency, and energy consumption for the current user request and has within its search range. To avoid always selecting a node with high computing power resources rather than a suitable one, a load balancing factor is added to the evaluation function to prevent individual nodes from being overloaded, represents the node which is the ratio of the total computing volume of all currently assigned tasks to the processable task volume. is the energy consumption generated by selecting the current node. In addition, to avoid an edge node from being selected by multiple users, is introduced, which is obtained by adding the ratio of the current real-time load of the current edge node to the maximum load volume and the node load weighted by , , focuses on the resource occupancy rate, which is used to measure the proportion of the current load of a certain type of resource (such as bandwidth or computing). A higher value indicates that it is close to its maximum load capacity and the penalty is greater. Similarly, represents the proportion of the current request volume. A higher value indicates that the server is close to its maximum concurrent processing capacity and the penalty is greater. By updating in real time, nodes that are less used are preferentially selected, making the system tend to select edge nodes more evenly to participate in the calculation.
[0145] The idea of this embodiment is as follows: The optimal edge node is selected to match the terminal device through the comprehensive evaluation function , that is, after grouping requests with different computing power requirements, training is carried out separately, which can avoid unordered search in a large range. For high-computing-power requests, the model directly learns to split near the front end of the hierarchy; while for low-computing-power requests, the split point can be kept at the back to reduce transmission latency. This makes the position of the split point clearer and more definite, helping to quickly find the optimal position for each type of request.
[0146] For multiple user requests in the terminal layer simultaneously attempt to select the best set of edge nodes , considering the proposed The function calculates the matching degree of the current selection. For more flexible implementation, The weights assigned to different devices reflect their priorities. Satisfy
[0147] (18);
[0148] Among them, is the sum of the functions of all terminal devices.
[0149] When maximizing the comprehensive evaluation function, the following constraints need to be satisfied:
[0150] (16);
[0151] (17);
[0152] (18);
[0153] Among them, is the tolerance range of load balancing, which limits the differences of all nodes within this range to achieve the purpose of load balancing.
[0154] Construct a hierarchical DRL model, including two sub-policy networks, which are respectively for high-computing-power-demand tasks and low-computing-power-demand tasks, and learn with different focuses. Specifically, taking the terminal-edge layer as an example, the high-computing-power-demand sub-network will preferentially learn to place the splitting point at a more forward position and focus on computing at the edge nodes, while the low-computing-power-demand sub-network will tend to place the splitting point at a more backward position and utilize edge and cloud computing resources, and learn according to this focus. In each training cycle, select the corresponding sub-policy for training according to the task classification result. Through this division, the search for the best splitting point position for the training of the two is also focused. The former is in this range, and the latter is in this range. By reducing the search range, higher training accuracy can be achieved with the same number of training times.
[0155] Provide a return value according to the final value, and continuously update parameters to form a closed loop.
[0156] By setting up a reward function
[0157] (19);
[0158] Similarly,
[0159] (20);
[0160] Among them, m and n are the high-computing-power request sets The task requests and the set of low-computing-power requests in The task requests in is the computing delay of the task requests in the set of high-computing-power requests based on this splitting point, is the transmission delay of the task requests in the set of high-computing-power requests based on this splitting point, is the transmission delay of the task requests in the set of low-computing-power requests based on this splitting point, is the computing delay of the task requests in the set of low-computing-power requests based on this splitting point, and are weight factors respectively;
[0161] Use the reward-based integration mechanism to aggregate the reward signals of high-computing-power and low-computing-power tasks, and update the global policy after each control cycle. Use different weight parameters to balance the importance of the two types of tasks.
[0162] (21);
[0163] Among them, is the reward function between the terminal layer and the edge layer, is the weight of high-computing-power tasks, is the weight of the current terminal-edge link in the overall link in the high-computing-power set, is the weight of the current terminal-edge link in the overall link in the low-computing-power set.
[0164] It can be regarded as separately learning and finding the optimal splitting point within its current training cycle at the initial stage of training, and coordinating the two types of computing-power tasks in the later stage of training, and using a joint strategy for real-time task allocation and splitting point selection. and are the weights of the current task, which can handle tasks more flexibly. is the weight of high-computing-power tasks, reflecting the influence of high-computing-power tasks in the reward function, If the value of is greater than, the model will be more inclined to select the splitting point position that optimizes high-computing-power tasks, otherwise it will pay more attention to optimizing the splitting point position of low-computing-power tasks. Each subsequent learning process is only related to the previous one, reflecting the Markov property.
[0165] Combine it with the edge-cloud computing layer to achieve overall optimization. First, design the reward function between the edge and the cloud computing layer
[0166] (22);
[0167] Among them, is the reward function between the edge layer and the cloud computing layer, , and is a weight factor, is the weight of the current best edge node in the set of best edge nodes and the link between the edge node and the cloud computing layer in the overall link.
[0168] Redesign the overall reward function:
[0169] (23);
[0170] Among them, is the weight factor of the reward function between the terminal and the edge layer, is the weight factor between the edge and the cloud computing layer.
[0171] Through this multi-dimensional reward mechanism, combined with the reward signals of high and low computing power tasks, the iterative update and optimization of the system are realized. The reward function provides a feedback signal for policy update. The decision function here uses the PPO algorithm, and by restricting the policy update amplitude, the new policy will not deviate significantly from the old policy.
[0172] The main idea of this method is to construct a hierarchical deep reinforcement learning architecture for selecting the optimization splitting points and resource allocation between the terminal layer, the edge layer and the cloud computing layer, so as to achieve efficient computing task allocation and model training. Multiple sub-policy networks are designed to handle tasks with different computing power requirements. A resource awareness module is combined to measure the status of computing resources, network bandwidth and latency of each layer. The heuristic algorithm is used to initially estimate the range of splitting points, reduce the exploration of irrelevant positions, and optimize the search efficiency. In the first stage, the splitting points are roughly selected to screen out better regions; in the second stage, these regions are finely optimized, so as to gradually improve the training accuracy. Through the reward signal integration mechanism, the rewards of high and low computing power tasks are aggregated, and the global policy is updated after the end of each control cycle. Real-time task allocation and optimization are achieved through the joint policy, so that the system can achieve the overall optimum in the multi-level architecture.
[0173] Based on the local optimal splitting point, split the corresponding inter-layer position of the structural layer of the model;
[0174] Based on the edge-cloud optimal splitting point, split the corresponding inter-layer position of the structural layer of the model;
[0175] Divide the structural layer of the model into three structural blocks, and then split the same model training request task into three segments for parallel training to achieve the allocation of training tasks.
[0176] Embodiment 2
[0177] This embodiment provides a model training task allocation system in a heterogeneous computing power environment, including:
[0178] The task computing power recognition module is configured to obtain multiple model training requests at the terminal layer under the computing architecture composed of the terminal layer - edge layer - cloud computing layer, and group multiple terminals according to the size relationship between the data volume of the model training requests and the set threshold;
[0179] The matching node determination module is configured to find the best edge matching node for each terminal by using a comprehensive evaluation function based on the computing power information of the edge layer and the computing power requirements of the model training requests;
[0180] The terminal-edge splitting module is configured to determine the initial selection range of the splitting point in the model according to the grouping situation of the model training requests. Within the determined initial selection range, based on the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the splitting point, iteratively loop to find the local optimal splitting point;
[0181] The edge-cloud splitting module is configured to determine the initial selection range of the splitting point in the model according to the grouping situation of the model training requests. Within the determined initial selection range, based on the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the splitting point, iteratively loop to find the edge-cloud optimal splitting point;
[0182] The training task allocation module is configured to coordinate the local optimal splitting point and the edge-cloud optimal splitting point to achieve the allocation of training tasks.
[0183] The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the first embodiment above. It should be noted that the above modules can be executed in a computer system such as a set of computer executable instructions as part of the system.
[0184] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0185] The proposed system can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above module division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0186] Embodiment Three
[0187] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a model training task allocation method in a heterogeneous computing power environment as described in the first embodiment above.
[0188] Embodiment Four
[0189] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a model training task allocation method in a heterogeneous computing power environment as described in Embodiment 1 above.
[0190] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.
[0191] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0194] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0195] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A method for allocating model training tasks in a heterogeneous computing power environment, characterized in that Including: Construct a hierarchical deep reinforcement learning model. Under the computing architecture composed of the terminal layer - edge layer - cloud computing layer, obtain multiple model training requests at the terminal layer, and group multiple terminals according to the size relationship between the data volume of the model training request and the set threshold; Based on the computing power information of the edge layer and the computing power requirements of the model training request, use the comprehensive evaluation function to find the best edge matching node for each terminal; Determine the preliminary selection range of the splitting point in the model according to the grouping situation of the model training request. Within the determined preliminary selection range, based on the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the splitting point, iteratively loop to find the local optimal splitting point; Determine the preliminary selection range of the splitting point in the model according to the grouping situation of the model training request. Within the determined preliminary selection range, based on the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the splitting point, iteratively loop to find the best edge-cloud splitting point; Use the hierarchical deep reinforcement learning model to solve the iterative loop search process. Among them, the hierarchical deep reinforcement learning model includes a state space, an action space, and an overall reward function; The state space includes the computing resource status, network bandwidth and delay of the terminal layer; the computing resource status of the edge layer and the connection status with the cloud server, and the computing resource status of the cloud computing layer; The action space includes the corresponding splitting schemes divided by various conditions; The overall reward function includes the reward function between the edge and the cloud computing layer and the reward function between the terminal and the edge layer; The reward function provides a feedback signal for policy update; Overall plan the local optimal splitting point and the best edge-cloud splitting point to achieve the allocation of training tasks; The overall planning of the local optimal splitting point and the best edge-cloud splitting point to achieve the allocation of training tasks is specifically as follows: Based on the local optimal splitting point, split the corresponding inter-layer position of the structure layer of the model; Based on the best edge-cloud splitting point, split the corresponding inter-layer position of the structure layer of the model; Divide the structure layer of the model into three structure blocks, and then split the same model training request task into three segments for parallel training to achieve the allocation of training tasks.
2. The model training task allocation method in a heterogeneous computing power environment according to claim 1, wherein The comprehensive evaluation function is specifically as follows: ; Among them, is the computing resource of the edge node ; is the computing requirement of the current terminal task for the node is the network bandwidth from the terminal to the edge node is the network latency is the adjustment parameter for balancing the weights of computing power, bandwidth, latency, and energy consumption for the current user request , has within its search range; the load balancing factor , is the energy consumption generated by selecting the current node, and the node selection frequency ; The larger the value of the comprehensive evaluation function, the higher the matching degree between the terminal and the edge node.
3. The model training task allocation method in a heterogeneous computing power environment according to claim 1, characterized in that The grouping of multiple terminals according to the size relationship between the data volume of the model training request and the set threshold is specifically as follows: Calculate the data volume of the model training request for each terminal in turn; If the data volume of the model training request is greater than the set threshold, it is a high computing power requirement level task and is classified into the high computing power requirement set; If the data volume of the model training request is less than the set threshold, it is a low computing power requirement level task and is classified into the low computing power requirement set; Thus, the model training requests of all terminals in the terminal layer are divided into two groups: the high computing power requirement set and the low computing power requirement set.
4. The model training task allocation method in a heterogeneous computing power environment according to claim 1, wherein, Determine the initial selection range of split points in the model according to the grouping situation of the model training request. Within the determined initial selection range, iterate and loop to find the local optimal split point with the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the split point. Specifically: If the data volume of the model training request is less than the set threshold, it is a task with low computing power requirement level, and the range after the split point in the middle layer of the model structure is used as the initial selection range of the split point; otherwise, it is a task with high computing power requirement level, and the range before the split point in the middle layer of the model structure is used as the initial selection range of the split point; Within the initial selection range, take a split point as an example to calculate the delay and energy consumption between the terminal and the edge matching node at the current split point; Iterate and loop through all split points within the initial selection range, and determine the local optimal split point with the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the split point.
5. The model training task allocation method in a heterogeneous computing power environment according to claim 1, characterized in that, Determine the initial selection range of split points in the model according to the grouping situation of the model training request. Within the determined initial selection range, iterate and loop to find the optimal edge-cloud split point with the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the split point. Specifically: If the data volume of the model training request is less than the set threshold, it is a task with low computing power requirement level, and the range after the split point in the middle layer of the model structure is used as the initial selection range of the split point; otherwise, it is a task with high computing power requirement level, and the range before the split point in the middle layer of the model structure is used as the initial selection range of the split point; Within the initial selection range, take a split point as an example to calculate the delay and energy consumption between the cloud computing layer and the edge matching node at the current split point; Iterate and loop through all split points within the initial selection range, and determine the local optimal split point with the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the split point.
6. A model training task allocation system in a heterogeneous computing power environment, characterized in that, Including: A task computing power identification module, configured to obtain multiple model training requests at the terminal layer under the computing architecture composed of the terminal layer - edge layer - cloud computing layer, and group multiple terminals according to the size relationship between the data volume of the model training request and the set threshold; A matching node determination module, configured to find the best edge matching node for each terminal by using a comprehensive evaluation function based on the computing power information of the edge layer and the computing power requirements of the model training request; A terminal-edge splitting module, configured to determine the initial selection range of split points in the model according to the grouping situation of the model training request. Within the determined initial selection range, iterate and loop to find the local optimal split point with the goal of minimizing the delay and energy consumption between the terminal and the edge matching node at the split point; An edge-cloud splitting module, configured to determine the initial selection range of split points in the model according to the grouping situation of the model training request. Within the determined initial selection range, iterate and loop to find the optimal edge-cloud split point with the goal of minimizing the delay and energy consumption between the edge matching node and the cloud computing layer at the split point; Solve the process of iterative loop search using a hierarchical deep reinforcement learning model, where the hierarchical deep reinforcement learning model includes a state space, an action space, and an overall reward function; The state space includes the computing resource status, network bandwidth, and latency of the terminal layer; the computing resource status of the edge layer and the connection status with the cloud server, and the computing resource status of the cloud computing layer; The action space includes corresponding splitting schemes for various conditions; The overall reward function includes a reward function for the edge and cloud computing layers and a reward function for the terminal and edge layers; The reward function provides a feedback signal for policy update; A training task allocation module, configured to coordinate the local optimal splitting point and the edge-cloud optimal splitting point to achieve the allocation of training tasks; The coordinating the local optimal splitting point and the edge-cloud optimal splitting point to achieve the allocation of training tasks is specifically: Based on the local optimal splitting point, split the corresponding inter-layer position of the structure layer of the model; Based on the edge-cloud optimal splitting point, split the corresponding inter-layer position of the structure layer of the model; Divide the structure layer of the model into three structure blocks, and then split the same model training request task into three segments for parallel training to achieve the allocation of training tasks.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a method for allocating model training tasks in a heterogeneous computing power environment as described in any one of claims 1-5.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the steps in a method for allocating model training tasks in a heterogeneous computing power environment as described in any one of claims 1-5.