Edge deployment method of large language model
By adopting pipeline parallel deployment of large language models in edge computing clusters, the problem of transmission delay and privacy issues in traditional deployment methods is solved, and efficient and accurate model deployment and inference computing are achieved.
Patent Information
- Application Number
- CN202510179840.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
The existing technology is difficult to effectively deploy large language models in edge computing clusters, resulting in unsolvable transmission delay and privacy issues, and model compression methods will lead to performance degradation.
By obtaining the sub-layer parameters and edge computing cluster device information of the large language model, the pipeline parallel deployment method is adopted to allocate multi-layer sub-layers to each device, and iteratively adjust the deployment plan with the optimization goal of minimizing the parallel processing delay in the cluster.
It realizes efficient deployment of large language models in edge computing clusters, reducing transmission delay and privacy leakage risks, while maintaining the model's inference accuracy and improving the inference computing efficiency of edge computing clusters.
Smart Images

Figure CN120123079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, specifically to the field of distributed edge computing, and more specifically to an edge deployment method for a large language model. Background Art
[0002] In recent years, a series of large language models (LLMs) represented by ChatGPT have developed rapidly in the field of intelligence. With its powerful understanding and generation capabilities, large language models can solve practical problems of algorithm models in a series of application scenarios, including text generation, translation, and question and answer. However, due to the limitations of existing computing resources, most large language models are deployed in the cloud. This widely used deployment strategy cannot fully protect user privacy, and the centralized computing architecture will generate serious bandwidth burden and transmission delay problems during the service process. These problems are particularly prominent in application fields such as autonomous driving and drone clusters.
[0003] Edge computing is a computing method that uses network edge devices close to data sources (edge devices are usually embedded devices with limited size, power consumption and computing power) to process tasks. By deploying large language models on edge devices and cloud devices for edge-cloud collaborative reasoning, the transmission delay in the traditional cloud reasoning mode can be effectively improved. However, during collaborative reasoning, the process of transmitting intermediate feature data to the cloud will still produce a certain end-to-end delay, and it still cannot effectively solve the privacy issues during the transmission process.
[0004] In addition, the model compression method can effectively reduce the weight size of the large language model and reduce the computing resource requirements. Therefore, when deployed to the edge device, a single edge device can meet the resource requirements, so that a single edge device has the ability to infer the large language model. However, this method will lead to a decrease in model performance. Specifically, the quality of the text generated by the large model after compression will be greatly reduced, and compared with the original model, the robustness and adaptability of the compressed model will be damaged to varying degrees.
[0005] An edge computing cluster is a computing cluster composed of multiple local edge devices or embedded devices. The above problems can be effectively improved by reasoning large language models through edge computing clusters. First of all, this local reasoning method can greatly reduce data transmission delays, shield privacy leaks, and cluster neural networks through model parallelism and other methods without compromising the existing reasoning accuracy of the model. However, existing edge cluster deployment methods are often oriented towards traditional DNNs, which mainly include data parallelism, tensor parallelism, and pipeline parallelism. These three methods can effectively solve the edge computing cluster deployment problem of traditional DNN models, but they are not applicable when facing large language models.
[0006] For the data parallel deployment method, please refer to Figure 1 , which is a schematic diagram of the data parallel deployment model. The core idea of this method is to deploy a complete neural network model in each edge device of the cluster. The memory capacity of modern edge devices is generally about 4GB - 16GB, and the weight size of traditional deep neural network models is generally at the 100MB level, which can be fully tolerated in edge devices. However, the weight size of large language models is at the 10GB or even 100GB level, and it is difficult for a single edge device to deploy a complete large language model. That is, data parallelism will cause a huge weight redundancy.
[0007] For the tensor parallel method, please refer to Figure 2 , which is a schematic diagram of the tensor parallel deployment model. The core consideration of this method is to longitudinally split the neural network model through the Map-Reduce method and send it to the edge computing cluster for calculation. When performing cluster inference in this way, Map-Reduce operations need to be performed on each divided layer. The Map operation shown in the figure divides the network layer calculation tasks and distributes them to other cluster devices for collaborative calculation, and then the Reduce operation is performed to recover the results of other devices in the cluster to prepare for the division of the next network layer task. This operation will not cause a large communication pressure when inferring traditional neural networks. However, when processing large language model layers, the communication delay caused by the huge data transmission volume per layer will become the performance bottleneck of the entire system inference. That is, the Map-Reduce method of tensor parallelism will exacerbate the communication burden.
[0008] For the pipeline parallel method, please refer to Figure 3 , which is a schematic diagram of the pipeline parallel deployment model. The characteristic of this method is that after horizontally dividing the original network model into layers and deploying each layer in the cluster, it can overcome the defects of data parallelism and tensor parallelism to a certain extent. However, in the existing pipeline parallel methods, since the weight size of sub-layers of traditional network models is at the MB level, which is much smaller than the device storage space, while the weight size of sub-layers of large language models is at the GB level, the existing solutions are not applicable to deploying large language models.
[0009] In summary, in traditional edge cluster collaborative inference methods, although the pipeline parallel scheme can overcome the respective defects of data parallel and tensor parallel schemes to a certain extent, due to the huge number of parameters of large language models and the limitations of the existing pipeline parallel scheme modeling, traditional multiple edge cluster collaborative inference methods are not applicable to the problem of deploying large language models.
[0010] It should be noted that: This background technology is only used to introduce relevant information of the present invention to facilitate understanding of the technical solution of the present invention, but it does not necessarily mean that the relevant information is prior art. The relevant information is submitted and disclosed together with the solution of the present invention. Without evidence showing that the relevant information was publicly available before the filing date of the present invention, the relevant information should not be regarded as prior art. Summary of the Invention
[0011] Therefore, the object of the present invention is to overcome the above-mentioned defects of the prior art and provide a method for edge deployment of a large language model.
[0012] The object of the present invention is achieved by the following technical solutions:
[0013] According to a first aspect of the present invention, there is provided a method for edge deployment of a large language model, the method comprising: S1, obtaining sub-layer parameters of the large language model, including multiple sub-layers, the storage requirements of each sub-layer, and the output data volume of each sub-layer; S2, obtaining device information of multiple edge devices included in the edge computing cluster, including the storage space of each device and the bandwidth of each device; S3, obtaining a preset constraint condition based on the device information, adopting a pipelined parallel deployment method, and allocating the multiple sub-layers to each device according to the sub-layer parameters and the constraint condition to obtain a deployment plan, which includes one layer or consecutive multiple layers of sub-layers allocated to each device, and taking minimizing the parallel processing latency in the cluster as the optimization goal, iteratively adjusting the deployment plan; S4, deploying the multiple sub-layers of the large language model to each device according to the deployment plan adjusted in S3.
[0014] In some embodiments of the present invention, in the S3, the method of allocating the multiple sub-layers to each device includes: sorting each device in descending order according to the bandwidth size of each device, and sequentially allocating the multiple sub-layers to each device sorted in descending order according to the sub-layer parameters and the constraint condition; wherein, between two adjacent sorted devices, the previous device inputs the output data of the last sub-layer allocated to it into the first sub-layer allocated to the next device.
[0015] In some embodiments of the present invention, in the S3, the method of iteratively adjusting the deployment plan includes: obtaining the operation time of each initial sub-layer on each device, and presetting a constraint condition based on the storage requirements of each sub-layer and the storage space of each device; constructing a sub-function for calculating the processing delay time of each device under the deployment plan based on the operation time of each sub-layer on each device, the output data volume of each sub-layer, and the bandwidth of each device; constructing an objective function for calculating the parallel processing delay in the cluster under the deployment plan based on the constraint condition and the sub-function; and adjusting the deployment plan based on the operation time of each initial sub-layer on each device to minimize the value of the objective function to obtain a preliminarily adjusted deployment plan.
[0016] In some embodiments of the present invention, the manner of adjusting the deployment scheme to minimize the value of the objective function includes: adopting a dynamic programming algorithm to predict the parallel processing delay of each possible deployment scheme, determining the deployment scheme with the minimum parallel processing delay. The dynamic programming algorithm is as follows:
[0017] ,
[0018] wherein, represents the sub-layer number, represents the device number, represents the minimized parallel processing delay when the sub-layers of the 0th to layer of the model are deployed to the 0th to devices, represents determining the deployment scheme that minimizes the parallel processing delay among all possible deployment schemes, represents the number of sub-layers allocated to the th device, represents the minimized parallel processing delay when the sub-layers of the 0th to layer of the model are deployed to the 0th to devices, represents the operation time accumulation auxiliary variable, represents the operation time of the sub-layer of the th layer on the th device, represents the bandwidth of the th device, represents the bandwidth of the th device, represents taking the minimum value.
[0019] In some embodiments of the present invention, the manner of iteratively adjusting the deployment scheme further includes performing one or more iterations using an iterative greedy algorithm, and using the deployment scheme of the last iteration as the finally adjusted deployment scheme. Wherein, the process of each iteration includes: correcting the operation time of each sub-layer on each device in the current iteration according to a preset manner to obtain the operation time of each sub-layer on each device after correction in the current iteration; based on the operation time of each sub-layer on each device in the current iteration, adopting the dynamic programming algorithm to predict the deployment scheme with the minimum parallel processing delay in the current iteration.
[0020] In some embodiments of the present invention, the preset method includes: based on the obtained deployment plan, counting the operation time of each device under this deployment plan, where the initial iteration uses the initially adjusted deployment plan, and each iteration after the first uses the deployment plan obtained from the previous iteration; obtaining the updated operation time of each sub-layer assigned to each device per layer on this device according to the quotient of the counted operation time of each device and the number of sub-layers assigned to each device; using the updated operation time of each sub-layer assigned to each device per layer on this device as the corrected operation time of the corresponding sub-layer on this device.
[0021] In some embodiments of the present invention, the sub-function is as follows:
[0022] ,
[0023] where, represents the processing delay time of the th device under the deployment plan, represents taking the maximum value, represents the communication time of the th device under the deployment plan, represents the operation time of the th device under the deployment plan.
[0024] In some embodiments of the present invention, the objective function is as follows:
[0025] ,
[0026] ,
[0027] where, represents the device number, represents the total number of devices, represents the parallel processing delay in the cluster, represents the preset constraint condition, represents the number of the last sub-layer among all sub-layers assigned to the th device under the deployment plan, represents the number of the last sub-layer among all sub-layers assigned to the th device under the deployment plan, represents the storage requirement of the th sub-layer, represents the storage space of the th device.
[0028] According to a second aspect of the present invention, there is provided an electronic device, including: one or more processors; and a memory, where the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method according to the first aspect of the present invention by executing the executable instructions.
[0029] Compared with the prior art, the advantages of the present invention are as follows:
[0030] The present invention takes into account the particularity of the large language model itself, collects device information of the edge computing cluster, such as device bandwidth and storage resources, so as to preset constraint conditions for model deployment calculation to construct a deployment plan, making the constructed deployment plan more suitable for deploying the large language model. In addition, the present invention adopts a pipeline parallel deployment method, distributes multiple sub-layers of the model to each device according to the sub-layer parameters and constraint conditions of the large language model to obtain a deployment plan, and takes minimizing the parallel processing latency in the cluster as the optimization goal to iteratively adjust the deployment plan, making the obtained deployment plan more suitable for deploying the large language model. At the same time, it minimizes the parallel processing latency and improves the inference calculation efficiency of the edge computing cluster for the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The following further describes the embodiments of the present invention with reference to the drawings, where:
[0032] Figure 1 It is a schematic diagram of a data parallel deployment model according to an embodiment of the present invention;
[0033] Figure 2 It is a schematic diagram of a tensor parallel deployment model according to an embodiment of the present invention;
[0034] Figure 3 It is a schematic diagram of a pipeline parallel deployment model according to an embodiment of the present invention;
[0035] Figure 4 It is a schematic diagram of the edge deployment method process of a large language model according to an embodiment of the present invention;
[0036] Figure 5 It is a schematic diagram of the processing process of a single device in the cluster under the pipeline parallel deployment method according to an embodiment of the present invention;
[0037] Figure 6 It is a schematic diagram of the processing process of all devices in the cluster under the pipeline parallel deployment method according to an embodiment of the present invention;
[0038] Figure 7 It is a schematic diagram of the complete execution process of the edge deployment method of a large language model according to an embodiment of the present invention;
[0039] Figure 8Schematic diagram of experimental comparison results between the method of the present invention and other methods according to an embodiment of the present invention. Detailed implementation manners
[0040] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below through specific embodiments with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] As mentioned in the background art section, in the traditional edge cluster collaborative inference method, although the pipeline parallelism scheme can overcome the respective defects of the data parallelism and tensor parallelism schemes to a certain extent, due to the huge number of parameters of large language models and the limitations of the existing pipeline parallelism scheme modeling, the traditional multiple edge cluster collaborative inference methods are not applicable to the problem of deploying large language models.
[0042] To address the above problems, the inventors propose an edge deployment method for large language models. On the one hand, the method of the present invention takes into account the particularity of the large language model itself. For example, the parameters of the sub-layers of the large language model are at the extremely large GB level, and the device information of the edge computing cluster, such as device bandwidth and storage resources, is collected, so as to preset constraints for model deployment calculation to construct a deployment plan, making the constructed deployment plan more suitable for deploying large language models. On the other hand, the method of the present invention adopts a pipeline parallel deployment method. According to the sub-layer parameters and constraints of the large language model, the multi-layer sub-layers of the model are allocated to each device to obtain a deployment plan, and the optimization goal is to minimize the parallel processing delay in the cluster, and the deployment plan is iteratively adjusted to make the obtained deployment plan more suitable for deploying large language models. Therefore, the method of the present invention not only overcomes the respective defects of the data parallelism and tensor parallelism schemes, but also obtains a deployment plan that minimizes the parallel processing delay of the edge computing cluster and is more suitable for large language models, improving the inference calculation efficiency of the edge computing cluster for the model.
[0043] According to an embodiment of the present invention, refer to Figure 4 , which is a schematic flowchart of an edge deployment method for a large language model. The method includes steps S1, S2, S3, and S4. To better understand the present invention, the following will specifically describe each step in detail with reference to specific embodiments.
[0044] In step S1, obtain the sub-layer parameters of the large language model, including multiple sub-layers, the storage requirements of each sub-layer, and the output data volume of each sub-layer.
[0045] According to an embodiment of the present invention, record the sub-layer parameters of the large language model as the data required for subsequent deployment plans. For example, the number of sub-layers of the large language model is N, and each obtained sub-layer is set as layer in the operation order , = 0, 1, 2, 3, ……, N - 1, and obtain the storage requirements of each sub - layer accordingly, denoted as , and obtain the output data volume of each sub - layer, denoted as , and in bytes.
[0046] In step S2, obtain the device information of multiple edge devices included in the edge computing cluster, which includes the storage space of each device and the bandwidth of each device.
[0047] According to an embodiment of the present invention, collect the device information of multiple edge devices included in the edge computing cluster. Schematically, assume that the number of all edge devices in the edge computing cluster is M (M > 1), and each device is denoted as device = 0, 1, 2, 3, ……, M - 1, and obtain the storage space of each device accordingly. The storage space is denoted as , and the storage space is in bytes. Obtain the bandwidth of each device, and the bandwidth is denoted as , and in Mbps.
[0048] In step S3, obtain the constraint conditions preset based on the device information, adopt a pipelined parallel deployment method, and allocate multiple sub - layers to each device according to the sub - layer parameters and the constraint conditions to obtain a deployment plan, which includes one layer or consecutive multiple layers of sub - layers allocated to each device, and take minimizing the parallel processing latency in the cluster as the optimization goal, and iteratively adjust the deployment plan.
[0049] According to an embodiment of the present invention, although there are many types of large language models, their structures have something in common. They all include multiple decoder layers, usually including dozens to hundreds of decoder layers. The pipelined parallel deployment method: After horizontally stratifying the large language model, each sub - layer after stratification is deployed in each edge device in the edge computing cluster. When the model input data arrives, it sequentially passes through each device according to the pipeline method to execute the calculations of the sub - layers deployed on it, and finally completes the calculation process of the entire model for the input data. When several input data arrive continuously, each input data is processed according to the pipeline method and each device executes in parallel to achieve the processing of multiple input data in a pipelined parallel manner. Among them, the method of horizontally stratifying the large language model can regard each decoder layer as one sub - layer.
[0050] According to an embodiment of the present invention, in step S3, the method of allocating multiple sub-layers to each device includes: sorting each device in descending order according to the bandwidth size of each device, and sequentially allocating multiple sub-layers to each device sorted in descending order according to sub-layer parameters and constraint conditions; wherein, between two adjacent sorted devices, the previous device inputs the output data of the last sub-layer it is allocated to the first sub-layer allocated to the next device. Schematically, in the case of meeting the constraint conditions, the sub-layers from layer 0 to layer 10 of the model are allocated to the device ranked first, and the sub-layers from layer 11 to layer 22 are allocated to the device ranked second. The first device inputs the output data of layer 10 it is allocated to layer 11 allocated to the second device to achieve pipelined operation. The technical solution of this embodiment can at least achieve the following beneficial technical effects: Without sorting in descending order, if there is a device with a bandwidth of 1000M sandwiched between two devices with a bandwidth of 10M, then the gigabit bandwidth device can only communicate at a bandwidth of 10M, resulting in low communication efficiency. The present invention sorts each device in descending order according to the bandwidth size and sequentially allocates multiple sub-layers to each device sorted in descending order, which can maximally prevent the communication of low-bandwidth devices from restricting high-bandwidth devices, thereby maximizing the transmission efficiency and computing efficiency of the edge computing cluster.
[0051] According to an embodiment of the present invention, in step S3, the method of iteratively adjusting the deployment plan includes the following steps a1, a2, a3, and a4:
[0052] Step a1: Obtain the operation time of each initial sub-layer on each device, and preset constraint conditions based on the storage requirements of each sub-layer and the storage space of each device.
[0053] According to an embodiment of the present invention, the operation time of each initial sub-layer on each device is obtained online. By simultaneously running N sub-layers of the large language model on each device, the operation time of different sub-layers on different devices can be obtained online, and the operation time is in milliseconds (ms). Among them, the operation time of the th sub-layer on the th device is denoted as , = 0, 1, 2, 3, ……, N - 1, = 0, 1, 2, 3, ……, M - 1.
[0054] According to an embodiment of the present invention, the method of presetting constraint conditions includes: for any device, ensuring that the total storage requirement of all sub-layers allocated to it does not exceed the storage space of the device, that is, the preset constraint condition is expressed in the following form:
[0055] , (1)
[0056] Among them, represents the device number, represents the total number of devices, represents the storage requirement of the th sub-layer, represents the number of the last sub-layer of all sub-layers allocated to the th device under the deployment scheme, represents the number of the last sub-layer of all sub-layers allocated to the th device under the deployment scheme, represents the th device's storage space.
[0057] Step a2: Based on the operation time of each sub-layer on each device, the output data volume of each sub-layer, and the bandwidth of each device, construct a sub-function for calculating the processing delay time of each device under the deployment scheme.
[0058] According to an embodiment of the present invention, the operation time and communication time within a single device are arranged in a pipeline, and the arrangement result can be seen in Figure 5 , which is a schematic diagram of the processing process of a single device in the cluster under the pipeline parallel deployment method. In the figure, the pipeline situation within a single device includes: in the initial period, first receive data 1 from the previous device; after this moment, the device is in a full-load state of the pipeline, that is, in the subsequent period, receive data 2 from the previous device, and at the same time, run the model sub-layer to calculate and infer data 1, and data 1 is processed in the device; receive data 3 from the previous device, and at the same time, run the model sub-layer to calculate and infer data 2, and data 2 is processed in the device; receive data 4 from the previous device, and at the same time, run the model sub-layer to calculate and infer data 3, and data 3 is processed in the device, and so on to process the received data in a pipeline manner. Therefore, after the pipeline is full, the processing delay time of a single device is the maximum value of the operation time and communication time within the device.
[0059] According to an embodiment of the present invention, based on the arrangement result and analysis of the operation time and communication time within a single device in the above embodiment, the sub-function can be constructed in the following form:
[0060] , (2)
[0061] Among them, represents the processing delay time of the th device under the deployment scheme, represents taking the maximum value, represents the communication time of the th device under the deployment scheme, represents the operation time of the th device under the deployment scheme.
[0062] According to an embodiment of the present invention, the communication time of the nth device is calculated as follows:
[0063] , (3)
[0064] where represents the output data volume of the last sub-layer allocated to the nth device under the deployment scheme, represents the number of the last sub-layer allocated to the nth device under the deployment scheme, represents the bandwidth of the nth device, represents the bandwidth of the nth device. Among them, is the output data volume size of the last layer allocated to the previous device. Since the unit is byte (Byte), multiply the number of bytes by 8 to unify it into the unit bit (bit). And the unit of bandwidth is megabit per second (Mbps), 1Mbps = 1000000bps. It is necessary to multiply the bandwidth by 1000 to unify the communication time unit into milliseconds.
[0065] According to an embodiment of the present invention, the operation time of the nth device is calculated as follows:
[0066] , (4)
[0067] where represents the number of the last sub-layer allocated to the nth device under the deployment scheme, represents the number of the sub-layer, represents the operation time of the mth sub-layer in the nth device.
[0068] Step a3: Based on the constraint conditions and the sub-functions, construct an objective function for calculating the parallel processing delay in the cluster under the deployment scheme.
[0069] According to an embodiment of the present invention, arrange the processing delay times between all devices in the cluster in a pipeline, and the obtained arrangement result can be seen in Figure 6 , which is a schematic diagram of the processing process of all devices in the cluster under the pipeline parallel deployment mode. In the figure, when counting the pipeline situation in the cluster system, it is assumed that the cluster contains 4 devices, and the pipeline arrangement result includes multiple processing stages. Among them, Stage 1: Device 1 processes Data 1; Stage 2: Device 1 processes Data 2, and Device 2 processes Data 1; Stage 3: Device 1 processes Data 3, Device 2 processes Data 2, and Device 3 processes Data 1; when the processing of the third stage ends, after this moment, the cluster system is in a full-load state of the pipeline, that is, all devices in the subsequent stages process data in parallel. For example, in Stage 4: Device 1 processes Data 4, Device 2 processes Data 3, Device 3 processes Data 2, and Device 4 processes Data 1, and Data 1 is processed completely in the cluster system; Stage 5: Device 1 processes Data 5, Device 2 processes Data 4, Device 3 processes Data 3, and Device 4 processes Data 2, and Data 2 is processed completely in the cluster system; Stage 6: Device 1 processes Data 6, Device 2 processes Data 5, Device 3 processes Data 4, and Device 4 processes Data 3, and Data 3 is processed completely in the cluster system; the received data is processed in parallel in this pipeline manner. Therefore, when the pipeline is full, the parallel processing delay of the cluster system is the maximum value of the processing delay time of the devices in the system.
[0070] According to an embodiment of the present invention, based on the above embodiment, after arranging the processing delay time of all devices in the cluster in a pipeline, the obtained arrangement result and analysis can construct the objective function in the following form:
[0071] , (5)
[0072] ,
[0073] Among them, represents the parallel processing delay, represents the device number, represents the total number of devices, represents the parallel processing delay in the cluster, represents the preset constraint condition, represents the storage requirement of the layer sub-layer, represents the storage space of the th device. Schematically, can be the DDR memory size of the deep learning dedicated accelerator configured on the th device. For example, if the device uses a graphics card of the NVIDIA series as the deep learning dedicated accelerator, then represents the video memory size of the
[0074] The technical solutions of steps a1 - a3 in the above embodiments can at least achieve the following beneficial technical effects: By constructing the objective function through the process of steps a1 - a3, under the pipeline parallel deployment scheme, the communication time and operation time inside the device and the parallel processing delay between devices are arranged and analyzed in a pipeline manner, so as to better realize the objective function modeling of the calculation method of the parallel processing delay of the cluster system, facilitating the subsequent calculation of more accurate parallel processing delay and a better deployment scheme.
[0075] Step a4: Based on the operation time of each initial sub - layer on each device, adjust the deployment scheme to minimize the value of the objective function, and obtain the preliminarily adjusted deployment scheme.
[0076] According to an embodiment of the present invention, the method of adjusting the deployment scheme to minimize the value of the objective function includes: adopting a dynamic programming algorithm to predict the parallel processing delay of all possible deployment schemes respectively, and determining the deployment scheme with the minimum parallel processing delay. Among them, the dynamic programming algorithm is as follows:
[0077] , (6)
[0078] Among them, represents the sub - layer number, represents the device number, represents the minimum parallel processing delay when the sub - layers from the 0th to the th layer of the model are deployed to the 0th to the rd device, represents determining the deployment scheme that minimizes the parallel processing delay among all possible deployment schemes, represents the th number of sub - layers allocated to the represents the minimum parallel processing delay when the sub - layers from the 0th to the th layer of the model are deployed to the 0th to the rd device, represents the auxiliary variable for accumulating the operation time, represents the th layer sub - layer's operation time on the rd device, represents taking the minimum value.
[0079] Among them, this dynamic programming algorithm is a state - transfer equation, which aims to find the minimum parallel processing delay when the sub - layer sequence ending with the i - th layer sub - layer is deployed in the device sequence ending with the th device. By comprehensively considering various factors to determine this minimum value, that is, comprehensively considering the parallel processing delay situation of the sub - layer deployment under the previous device sequence from 0 to : , the operation time of the device: Communication time with the device: . The technical solution of this embodiment can at least achieve the following beneficial technical effects: During the calculation of the dynamic programming algorithm (i.e., the state transition equation), the influence of the previous device deployment situation on the current state is reflected. For example, when calculating , it is necessary to refer to the sub-layer deployment delay situation under the previous device sequence , which helps to find the optimal solution from the existing states.
[0080] According to an embodiment of the present invention, the pseudocode of the dynamic programming algorithm is as follows:
[0081] 1. Input:
[0082] 2. Sub-layer parameters of the large language model, including the storage requirements of each sub-layer and the output data volume of each sub-layer ;
[0083] 3. Device information of multiple edge devices included in the edge computing cluster and the operation time of the sub-layers in the devices , the device information includes the storage space of each device sorted in descending order of bandwidth size and the device bandwidth ;
[0084] 4. Output:
[0085] 5. Deployment plan of the large language model ;
[0086] 6. (0 <= i < N, 0 <= j < M)
[0087] 7. / / Initialize the state
[0088] 8. for i = 0 to N - 1 do:
[0089] 9. for j = 0 to M - 1 do:
[0090] 10. dp[i][j] = INF
[0091] 11. / / Initialize the state of the first device
[0092] 12. for i = 0 to N - 1 do:
[0093] 13. dp[i][0] = 0
[0094] 14. mem = 0
[0095] 15. for layer_index = 0 to i do:
[0096] 16. mem += weight[layer_index]
[0097] 17. if mem > memory[0] do:
[0098] 18. dp[i][0] = INF
[0099] 19. break
[0100] 20. dp[i][0] += infer[layer_index][0]
[0101] 21. / / try_arrange[i][j] represents the number of sub - layers assigned to the - th device (denoted as device ) in the deployment plan corresponding to the minimum parallel processing delay of the device sequence ending with device for the sub - layer sequence ending with layer[i];
[0102] 22. try_arrange[i][0] = i + 1
[0103] 23. / / Dynamic programming
[0104] 24. for j = 1 to M - 1 do:
[0105] 25. for i = 0 to N - 1 do:
[0106] 26. for k = 0 to i do:
[0107] 27. mem = 0
[0108] 28. current_device_infer = 0
[0109] 29. for layer_index = k + 1 to i do:
[0110] 30. mem += weight[layer_index]
[0111] 31. current_device_infer += infer[layer_index][j]
[0112] 32. if mem > memory[j] do:
[0113] 33. continue
[0114] 34. current_device_latency=max(current_device_infer,comm[k][j])
[0115] 35. if dp[i][j]>max(dp[k][j-1],current_device_latency) do:
[0116] 36. dp[i][j]=dp[k][j-1],current_device_latency
[0117] 37. try_arrange[i][j]=ik
[0118] 38. / / Get the allocation plan
[0119] 39. unarranged_layer_nums=N
[0120] 40. for device_index=n-1 to 0 do:
[0121] 41. arranged_layers=arrange[unarranged_layer_nums-1][node_index]
[0122] 42. arrange[device_index]=unarranged_layer_nums-1
[0123] 43. unarranged_layer_nums-=arranged_layers
[0124] 44. / / Return result
[0125] 45. Return Arrange
[0126] The pseudo code of the dynamic programming algorithm is explained as follows:
[0127] Lines 1-3: Indicates that the input includes: The model number is The sublayer ( =0~N-1) storage requirement , Output data volume ; Number is ( =0~M-1) storage space of the device ,bandwidth and the model's sublayer numbered i is in the Devices (denoted as device[ The inference operation time on ;
[0128] Lines 4 - 5: It indicates that the output is: , It represents the last sub - layer number of the set of model sub - layers allocated on device ; try_arrange[i][j] represents the deployment plan tried to be set for the th device, that is, it represents the number of sub - layers allocated to the th device in the deployment plan of deploying sub - layers 0 to i on devices 0 to j. After the algorithm is completed, the final allocation plan arrange output by the algorithm can be obtained according to the value of try_arrange;
[0129] Lines 8 - 10: It indicates that the state variables are initialized to infinity for facilitating subsequent state transition calculations; among them, dp[i][j] in the state transition process is the state storage variable to be transferred. Specifically, dp[i][j] represents the minimized parallel processing delay when the set of sub - layers with model sub - layer numbers from 0 to i is deployed on devices numbered from 0 to j;
[0130] Lines 12 - 22: It indicates the initial update of the dp state, mainly setting the state of dp[0~N - 1][0], where:
[0131] Lines 15 - 20: It indicates judging whether sub - layers 0 to i can be deployed on device j. If not, directly jump out of the loop and set dp[i][0] to inf, that is, infinity; otherwise, accumulate the running time of this layer on device 0 into dp[i][0];
[0132] Line 22: It indicates updating try_arrange[i][0] to i + 1. The explanation for this update principle is: Since the current plan is to deploy sub - layers 0 to i on the th device, and try_arrange stores the number of layers, so 1 needs to be added;
[0133] Lines 24 - 38: It indicates performing state transition. Among them, the innermost loop from lines 28 - 38 is executing the state transition formula. That is, for dp[i][j], which is the deployment plan of sub - layers 0 to i on devices 0 to j, it is obtained by state transitions from the deployment of layer 0 on devices 0 to j - 1 and sub - layers 1 to i on the th device, the deployment of sub - layers 0 to 1 on devices 0 to j - 1 and sub - layers 2 to i on the th device... These state transitions result in the deployment plan with the minimum parallel processing delay, and the value of try_arrange will be continuously updated during the transition;
[0134] Lines 39 - 43: Represent the deployment plan arrange of the nth device calculated by try_arrange;
[0135] Lines 44 - 45: Output the deployment plan arrange obtained by the dynamic programming algorithm.
[0136] It should be noted that the initially adjusted deployment plan obtained in the above - mentioned embodiment is obtained under theoretical circumstances. Since there are many types of large - language models, but their biggest common feature is that they include multiple identical decoder layer sub - layers, and these decoder layers can be understood as completely identical operator layers from a computational perspective. Therefore, existing methods usually only count the operation time of a single - layer sub - layer on each device and predict the operation time of multiple - layer sub - layers by multiplication. That is, in the initially adjusted deployment plan obtained by calculation, this method is used to determine the operation time of multiple - layer sub - layers on the device. Schematically, if the operation time of a single layer on the second device is 2s and 10 layers of sub - layers are deployed for this device, then the operation time of 10 layers of sub - layers on the second device is 2×10 = 20s.
[0137] However, it is found in actual computer operation that this method of only counting the operation time of a single - layer sub - layer on each device and multiplying cannot accurately obtain the operation time of multiple - layer sub - layers on the device. Therefore, according to an embodiment of the present invention, the method of iteratively adjusting the deployment plan further includes step a5: performing one or more iterations using the iterative greedy algorithm, taking the deployment plan of the last iteration as the finally adjusted deployment plan, and in each iteration process, the operation time of each layer of sub - layers on the device is statistically analyzed online with the deployment plan of the previous iteration to correct and obtain the operation time of each layer of sub - layers on the device in the actual situation, and then the deployment plan is adjusted. The technical solution of this embodiment can at least achieve the following beneficial technical effects: by continuously correcting the operation time of the sub - layers and further adjusting the deployment plan based on the corrected operation time, the theoretical effect of the model deployment plan is continuously approximated to the actual effect, so as to adjust and obtain a model deployment plan close to the optimal one.
[0138] According to an embodiment of the present invention, in step a5, the process of each iteration includes the following steps a51 and a52:
[0139] Step a51: Correct the operation time of each sub - layer on each device in the current iteration according to a preset method to obtain the operation time of each sub - layer on each device after the current iteration is corrected.
[0140] According to an embodiment of the present invention, the preset method includes: based on the obtained deployment plan, counting the operation time of each device under this deployment plan, where the initial iteration uses the preliminarily adjusted deployment plan, and each iteration after the first uses the deployment plan obtained from the previous iteration; according to the quotient of the counted operation time of each device and the number of sub-layers allocated to each device, obtaining the updated operation time of each sub-layer allocated to each device per layer on this device; taking the updated operation time of each sub-layer allocated to each device per layer on this device as the corrected operation time of the corresponding sub-layer on this device.
[0141] Illustratively, if in the deployment plan obtained from the previous iteration, 12 sub-layers are allocated to the first device, and the operation time of each sub-layer in the first 12 sub-layers on this device before the current correction is 2 s, and if the online statistics show that the overall operation time for the first device to complete the 12 sub-layers allocated to it is 18 s, then the operation time of each sub-layer in the 12 sub-layers on this device is 18÷12 = 1.5 s, and the operation time of each sub-layer on the first device is updated to 1.5 s.
[0142] Step a52: Based on the operation time of each sub-layer on each device in the current iteration, use the dynamic programming algorithm to predict the deployment plan with the minimum parallel processing delay in the current iteration.
[0143] According to an embodiment of the present invention, based on the operation time of each sub-layer on each device after the current correction, use the same dynamic programming algorithm as in step a4 to obtain the deployment plan with the minimum parallel processing delay in the current iteration.
[0144] The technical solutions of the above embodiments can at least achieve the following beneficial technical effects: In the iterative greedy algorithm of the present invention, in each iteration, according to the deployment plan obtained from the previous iteration, the actual operation time of the device is counted, and then the total time is divided by the number of layers to update the operation time statistics of each sub-layer on the device, so as to correct the multi-layer operation time, making the theoretical effect of the model partitioning scheme continuously approach the actual effect, thereby obtaining a nearly optimal model deployment plan and reducing the parallel processing delay of the cluster for inferring large language models.
[0145] According to an embodiment of the present invention, during each iteration, it further includes: obtaining a preset error threshold, and when the difference between the minimum parallel processing delays of two adjacent iterations is less than or equal to the preset threshold or the number of iterations reaches the preset number, using the deployment plan in the current iteration as the finally adjusted deployment plan. The technical solutions of this embodiment can at least achieve the following beneficial technical effects: Since the more accurate the prediction of the overall operation time of multiple sub-layers on the device, the closer the parallel processing delays obtained in two consecutive iterations will be, thus making the finally adjusted deployment plan closer and closer to the actual optimal deployment plan.
[0146] According to an embodiment of the present invention, the complete processing procedure of the iterative greedy algorithm is as follows:
[0147] First, define the error threshold as Θ (0 < Θ < 1) and the upper limit of the number of iterations as H (2 < H). The reference deployment scheme is arrange_last (initially random or a preliminarily adjusted deployment scheme), the parallel processing delay of the reference deployment scheme is τ_last (initially infinite), and the number of iterations is η (initialized to 0).
[0148] Secondly, for the first iteration, use the preliminarily adjusted deployment scheme as the reference deployment scheme, count the computing time of each device under this reference deployment scheme, and update infer[i][j] (0 <= i < N, 0 <= j < M) according to the statistical results. Then, execute according to step a52 to obtain the deployment scheme adjusted after the first iteration and the parallel processing delay of this scheme, denoted as τ_cur. Calculate the error θ between τ_cur and τ_last, and increment the iteration number η by 1.
[0149] Then, each iteration after the first iteration is executed according to the following process:
[0150] If θ > Θ and η <= H, count the computing time of each sub-layer on each device under the deployment scheme obtained in the current iteration in the manner of step a51, and correct the computing time infer[i][j] (0 <= i < N, 0 <= j < M) of each sub-layer on each device in the current iteration according to the statistical results. Then, execute according to step a52 to obtain the deployment scheme adjusted after the current iteration and the parallel processing delay of this scheme, denoted as the current τ_cur. Calculate the error θ between the current τ_cur and the parallel processing delay of the previous deployment scheme, and increment the iteration number η by 1.
[0151] Finally, if η > H or θ <= Θ obtained after each of the above iterations, determine the deployment scheme obtained in the η-th iteration as the finally adjusted deployment scheme and stop the iteration.
[0152] In step S4, according to the deployment scheme adjusted in step S3, deploy multiple sub-layers of the large language model to each device.
[0153] According to an embodiment of the present invention, if the model includes 30 sub-layers numbered from 0 to 29, and the deployment scheme adjusted in step S3 is: allocate the 0th to 22nd layers of the model to the 0th device, the 23rd to 26th layers of the model to the 1st device, and the 27th to 29th layers of the model to the 2nd device, then, the 0th to 22nd sub-layers are deployed to the 0th device, the 23rd to 26th sub-layers are deployed to the 1st device, and the 27th to 29th sub-layers are deployed to the 2nd device.
[0154] Generally speaking, according to an embodiment of the present invention, seeFigure 6 , which is a schematic diagram of the complete execution process of the edge deployment method of the large language model. The complete execution process in the figure includes the following 7 steps from 01 to 07. Now, the 7 steps are described as follows with specific parameter examples:
[0155] Step 01, obtain the sub-layer parameters of the large language model.
[0156] According to an embodiment of the present invention, taking the model ChatGLM6b as the tuning object, this model mainly includes 30 processing layers, namely one starting layer (word_embedding), 28 decoding layers (decoder layer), and 1 prediction layer (lm_head). The storage requirements (weight) of each of the 30 sub-layers are 1069286556, 402763422, 402763422, ……, 402763422, 1069286556. Among them, token processing does not occupy memory, and the output data volume (output) of each sub-layer is 4096, 4096, ……, 4096, 4096, 1.
[0157] Step 02, obtain the device information of multiple edge devices included in the edge computing cluster.
[0158] According to an embodiment of the present invention, the edge computing cluster includes two sets of edge intelligent modules (NVIDIA Jetson NX 16G) and an edge server (NVIDIA L4) equipped with an edge acceleration card. Among them, the storage space (memory) of each device is 16, 16, and 24, and the bandwidth (bandwidth) of each device is 100, 100, and 1000.
[0159] Step 03, sort the devices in descending order according to the bandwidth of each edge device.
[0160] According to an embodiment of the present invention, sort the devices in descending order according to the bandwidth of each device in the cluster, number the devices according to the sorted result, and arrange their storage spaces: The multiple devices after descending order sorting and their numbers are as follows:
[0161] Number 0: NVIDIA L4;
[0162] Number 1: NVIDIA Jetson Orin NX 16G;
[0163] Number 2: NVIDIA Jetson Orin NX 16G;
[0164] The storage spaces of the devices numbered 0 - 2 are 24, 16, and 16 in sequence.
[0165] Step 04, obtain the operation time of the sub-layer on each device online.
[0166] According to an embodiment of the present invention, all devices in the cluster simultaneously perform inference on the model sub-layer to statistically calculate the operation time (infer) of different sub-layers on different devices. The statistical data is recorded in the following form:
[0167] [[0.0297, 1.4969, 0.092], [0.204, 10.66, 0.689], [0.204, 10.66, 0.689]...], where a group of three time data forms the operation time of one sub-layer on the three devices given in the above embodiment.
[0168] Step 05, construct the objective function according to the process of steps a1 - a3 in the above embodiment, and adjust the deployment plan by minimizing the value of the objective function.
[0169] Step 06, use the dynamic programming algorithm to solve the deployment plan corresponding to the minimum value of the objective function, and obtain the initially adjusted deployment plan.
[0170] According to an embodiment of the present invention, taking the pseudo-code shown in the above embodiment as an example, the weight, output, memory, bandwidth, and infer obtained in 01 - 04 are used as the input of the dynamic programming algorithm and the dynamic programming algorithm is run. The initial pipeline parallel deployment plan of ChatGLM6b in the edge computing cluster is obtained as arrange_cur = [22, 25, 29]. This deployment plan means that the 0th device deploys layers 0 - 22, the 1st device deploys layers 23 - 25, and the 2nd device deploys layers 25 - 29. And the parallel processing delay under this plan is: 36.638ms.
[0171] Step 07, adopt the iterative greedy algorithm to obtain the finally adjusted deployment plan of the model. This includes correcting the operation time of multiple sub-layers of the model on the device, and based on the corrected operation time, adjusting the deployment plan in combination with the dynamic programming algorithm in step 06.
[0172] According to an example of the present invention, a schematic process of the iterative greedy algorithm is as follows:
[0173] First, preset the error threshold as Θ = 0.05, the iteration upper limit as η = 10, the reference deployment plan as arrange_last (initially the initially adjusted deployment plan), and set the parallel processing delay of this plan as τ_last = inf.
[0174] Secondly, obtain the deployment plan obtained after each iteration in the manner of the above embodiment, and calculate the difference between the parallel processing latency corresponding to the deployment plan after each iteration and its previous parallel processing latency. Denote this difference as the error θ. Among them, the error θ calculated in the first iteration = inf, and record the iteration number = 1;
[0175] Then, if the error θ > Θ and the iteration number is 1 <= η, that is, it does not exceed the iteration number limit, according to the deployment plan after the first iteration above, re - perform online statistics on the operation time of multiple sub - layers on the device, and correct the sub - layer running time according to the statistical results. Based on the corrected time, use step 06 to obtain the deployment plan arrange_cur[20, 24, 29] again. The parallel processing latency under this plan is: 50.86497ms. Record the iteration number = 2, and calculate the error between the second parallel processing latency and the first parallel processing latency as θ = 0.3883:
[0176] Since θ > Θ and the iteration number is 2 <= η, that is, it does not exceed the iteration number limit, according to the deployment plan after the second iteration above, re - perform online statistics on the operation time of multiple sub - layers on the device, and correct the sub - layer running time according to the statistical results. Based on the corrected time, use step 06 to obtain the deployment plan arrange_cur[20, 24, 29] again. The parallel processing latency under this plan is: 50.91197ms. Record the iteration number = 3. Calculate the error between the third parallel processing latency and the second parallel processing latency as θ = 0.047:
[0177] Since the θ in the third time <= Θ, determine that the finally adjusted deployment plan of the model is arrange_cur = [20, 24, 29].
[0178] To verify the beneficial effects of the present invention, the inventor conducts the following comparative experiments:
[0179] 1) Denote the method of the present invention as the deployment method of multi - node pipeline, and select Method 1 for comparison with the method of the present invention: random allocation and Method 2: multi - node without pipeline. Among them, random allocation means: randomly allocate each sub - layer of the large - language model to each device for inference, and each device does not form a pipeline for parallel execution; multi - node without pipeline means: each device does not form a pipeline for parallel execution. First, both the above formulas (2) and (5) are adopted in the summation (sum) manner, and then a non - pipeline objective function is constructed based on the summation manner, and in the dynamic programming formula (6), also becomes the summation manner , to optimize the deployment plan by minimizing the non - pipeline objective function.
[0180] 2) Use the dataset WikiText2, randomly select a prompt, and under the same model and the same edge computing cluster, use the deployment solutions generated by the method of the present invention, Method 1, and Method 2 respectively for deployment and large language model inference (LLM inference) to obtain the comparison results among the various methods. See Figure 8 , which is a schematic diagram of the experimental comparison results between the method of the present invention and other methods. In the figure, the abscissa represents the number of token generation rounds, that is, during the LLM inference process, the number of times the model gradually generates tokens. In each round, the model predicts the next token based on the previously generated tokens and the input prompt. The ordinate represents the inference time (unit: millisecond), that is, when the model generates each token, it counts the time interval between it and the previously generated token. According to Figure 8 the experimental results, compared with the random allocation of Method 1, the method of the present invention improves the system throughput by 7.11 times, and compared with the multi-node non-pipelined of Method 2, the method of the present invention improves the system throughput by 1.26 times. It shows that the present invention greatly improves the inference computing efficiency of the edge computing cluster for the model.
[0181] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even the order can be changed as long as the required functions can be achieved.
[0182] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.
[0183] The computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the above.
[0184] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technological improvements in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A method for edge deployment of a large language model, characterized in that: Methods include: S1. Obtain sub-layer parameters of the large language model, including multiple sub-layers, storage requirements of each sub-layer, and output data volume of each sub-layer; S2. Obtain device information of multiple edge devices contained in the edge computing cluster, including the storage space and bandwidth of each device; S3, obtaining constraints preset based on device information, adopting a pipeline parallel deployment method, allocating multiple sub-layers to each device according to sub-layer parameters and constraints, and obtaining a deployment plan, which includes one or multiple consecutive sub-layers allocated to each device, and iteratively adjusting the deployment plan with minimizing the parallel processing delay in the cluster as the optimization goal; S4. According to the deployment plan adjusted in S3, the multi-layer sub-layers of the large language model are deployed to each device.
2. The deployment method according to claim 1, characterized in that: In S3, the manner of allocating multiple sub-layers to each device includes: Sort the devices in descending order according to their bandwidths, and assign multiple sub-layers to the devices in descending order in sequence according to sub-layer parameters and constraints; Among them, between two adjacently ordered devices, the former device inputs the output data of the last sub-layer allocated to it into the first sub-layer allocated to the latter device.
3. The deployment method according to claim 1, characterized in that: In S3, the method of iteratively adjusting the deployment plan includes: Get the initial computing time of each sub-layer on each device, based on the storage requirements of each sub-layer and the storage space constraints of each device; Based on the operation time of each sub-layer on each device, the output data volume of each sub-layer and the bandwidth of each device, a sub-function is constructed to calculate the processing delay time of each device under the deployment scheme; Based on the constraint condition and the sub-function, construct an objective function for calculating the parallel processing delay in the cluster under the deployment scheme; Based on the initial computing time of each sub-layer on each device, the deployment plan is adjusted to minimize the value of the objective function, and a preliminary adjusted deployment plan is obtained.
4. The deployment method according to claim 3, characterized in that: The method of adjusting the deployment scheme to minimize the value of the objective function includes: A dynamic programming algorithm is used to predict the parallel processing delays of all possible deployment schemes and determine the deployment scheme with the minimum parallel processing delay. The dynamic programming algorithm is as follows: , in, Indicates the sublayer number, Indicates the device number. Indicates the model from 0 to Layer sublayers are deployed from 0 to Minimized parallel processing latency in 1 device, It means to determine the deployment scheme that minimizes the parallel processing delay among all possible deployment schemes. Indicates The number of sublayers allocated to each device, Indicates the model from 0 to Layer sublayers are deployed from 0 to Minimized parallel processing latency in 1 device, Indicates the auxiliary variable for the accumulated operation time. Indicates The sublayer is in the The computing time of each device, Indicates The bandwidth of each device, Indicates The bandwidth of each device, Indicates taking the minimum value.
5. The deployment method according to claim 4, characterized in that: The method of iteratively adjusting the deployment plan also includes using an iterative greedy algorithm to perform one or more iterations, and taking the last deployment plan as the final adjusted deployment plan, wherein the process of each iteration includes: Correcting the operation time of each sub-layer on each device in a preset manner to obtain the corrected operation time of each sub-layer on each device; Based on the operation time of each sub-layer on each device at that time, the dynamic programming algorithm is used to predict the deployment plan with the minimum parallel processing delay at that time.
6. The deployment method according to claim 5, characterized in that: The preset method includes: Based on the obtained deployment plan, the operation time of each device under the deployment plan is counted. The deployment plan after preliminary adjustment is adopted in the first iteration, and the deployment plan obtained in the previous iteration is adopted in each iteration after the first iteration. According to the quotient of the statistical operation time of each device and the number of sub-layers allocated to each device, the updated operation time of each sub-layer allocated to each device on the device is obtained; The updated operation time of each sub-layer allocated to each device on the device is used as the corrected operation time of the corresponding sub-layer on the device.
7. The deployment method according to claim 3, characterized in that: The sub-function is as follows: , in, Indicates the deployment plan The processing delay of each device, Indicates taking the maximum value, Indicates the deployment plan The communication time of each device, Indicates the deployment plan The operation time of a device.
8. The deployment method according to claim 7, characterized in that: The objective function is as follows: , , in, Indicates the device number. Indicates the total number of devices. represents the parallel processing delay in the cluster, represents the preset constraints, Indicates that the deployment plan is The number of the last sublayer of all sublayers assigned to the device. Indicates that the deployment plan is The number of the last sublayer of all sublayers assigned to the device. Indicates The storage requirements of the sub-layers, Indicates storage space of each device.
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1 to 8.
10. An electronic device, characterized in that: include: one or more processors; as well as A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of any one of claims 1-8 by executing the executable instructions.
Citation Information
Patent Citations
Traveler stops for spinning rings.
IN009313B
Method for producing metals or alloys poor in carbon and silicon in electrical furnaces
IN009616B
Improvements in or relating to chlorinators.
IN009717B
Improvements in means for moistening warps during weaving.
IN010020B