A Large Language Model Optimization Method and System for Edge Devices
By constructing edge device networks, pruning, and allocating resources, large language models are optimized for edge devices, overcoming computational and memory constraints to ensure efficient and consistent performance.
Patent Information
- Application Number
- CN202510416611.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Large language models are not feasible to run directly on edge devices with limited computing power, memory capacity and battery capacity, resulting in reduced generation capacity and accuracy, which is difficult to meet the needs.
By building an edge device network, dividing it into multiple clusters, extracting device parameter information, pruning, obtaining a compressed model, and splitting and loading based on fixed and dynamic resources, optimizing the deployment of large language models.
Efficient data processing is realized on edge devices, ensuring generation capabilities and accuracy, improving task processing speed, and adapting to the hardware characteristics of different devices.
Smart Images

Figure CN119918622B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large language models, and particularly relates to a method and system for optimizing large language models for edge devices. Background Art
[0002] A large language model (LLM) is a complex system based on deep learning. By pre-training on a large scale of text data, it learns rich language knowledge and context understanding capabilities. These models usually adopt the Transformer architecture and can generate high-quality natural language text, which are widely used in natural language processing tasks such as text generation, machine translation, question answering systems, and sentiment analysis. Through multi-layer self-attention mechanisms and feed-forward neural networks, large language models can capture complex dependencies when processing long texts and provide accurate and coherent outputs.
[0003] With the wide application of large language models in various applications, their huge parameter scale and computational complexity pose great challenges to the deployment on edge devices. Due to the limitations of edge devices in computing power, memory capacity, and battery capacity, it is often infeasible to directly run a complete large language model on edge devices. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for optimizing large language models for edge devices, aiming to solve the problem that it is often infeasible to directly run a complete large language model on edge devices due to the limitations of edge devices in computing power, memory capacity, and battery capacity.
[0005] The present invention is implemented as follows. A method for optimizing a large language model for an edge device, the method comprising:
[0006] Construct an edge device network, and the edge device network is overall managed by a cloud server;
[0007] Divide the edge devices to obtain multiple edge device clusters, and extract the device parameter information of each edge device;
[0008] Prune the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and divide the resources of the edge device clusters to obtain fixed resources and dynamic resources;
[0009] Based on the fixed resources, split and load the compressed model, process the data processing tasks, and based on the dynamic resources, perform dynamic local repeated loading on the compressed model to complete the processing of subsequent tasks.
[0010] Preferably, the step of partitioning the edge devices to obtain multiple edge device clusters and extracting the device parameter information of each edge device specifically includes:
[0011] Statistically analyze the hardware data of each edge device, quantify the computing power resources of the edge device to obtain a quantization result;
[0012] Cluster the edge devices according to the minimum computing power resources required to run the large language model to obtain multiple edge device clusters, and the sum of the computing power resources of all edge devices within the edge device cluster is greater than the minimum computing power resources required to run the large language model;
[0013] Extract the device parameter information of all edge devices within each edge device cluster, and the device parameter information includes at least computing power and memory capacity.
[0014] Preferably, the step of pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model and partitioning the resources of the edge device cluster to obtain fixed resources and dynamic resources specifically includes:
[0015] Determine the pruning ratio layer by layer, describe the overall performance under different pruning ratios through a metric function to obtain a pair of ratio scores, and embed the pair of ratio scores into a continuous representation space;
[0016] In the continuous representation space, start from the initialization point and iteratively update along the gradient direction of the evaluator to determine the optimal pruning ratio, and prune the large language model to obtain a compressed model;
[0017] Partition the resources of the edge device cluster based on the computing power resources required by the compressed model to obtain fixed resources and dynamic resources.
[0018] Preferably, the step of splitting and loading the compressed model based on the fixed resources to process the data processing task, and dynamically and locally repeatedly loading the compressed model based on the dynamic resources to complete the processing of subsequent tasks specifically includes:
[0019] Split the compressed model into multiple independent model units, and load the independent model units on the corresponding edge devices, and the number of edge devices is the same as the number of independent model units;
[0020] Obtain the data processing task, import the data processing task into each independent model unit, and monitor the data processing results of the edge device to obtain a task processing monitoring result;
[0021] Evaluate the data processing speed of each part of the compressed model based on the task processing monitoring result, and load the corresponding independent model unit according to the evaluation result to assist in processing subsequent tasks.
[0022] Preferably, the pruning ratio is expressed as , and the metric function is expressed as:
[0023] ;
[0024] wherein, is the zero-shot perplexity, which is used to quantify the generation ability, and respectively represent the edge-device-specific latency and energy consumption limits, and respectively represent the latency and energy consumption corresponding to a given ratio, is the indicator function, which returns 1 if the condition holds, otherwise returns 0, and are both penalty factors, is the th pruning ratio.
[0025] Another object of the present invention is to provide a large language model optimization system for edge devices, the system comprising:
[0026] A device network construction module for constructing an edge device network, which is overall managed by a cloud server;
[0027] A device division module for dividing edge devices to obtain multiple edge device clusters, and extracting device parameter information of each edge device;
[0028] A model loading module for pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and dividing the resources of the edge device cluster to obtain fixed resources and dynamic resources;
[0029] A dynamic compensation module for splitting and loading the compressed model based on the fixed resources, processing data processing tasks, and dynamically locally reloading the compressed model based on the dynamic resources to complete the processing of subsequent tasks.
[0030] Preferably, the device division module includes:
[0031] A computing power quantization unit for counting the hardware data of each edge device, quantifying the computing power resources of the edge device to obtain a quantization result;
[0032] A cluster division unit for clustering edge devices according to the minimum computing power resources required to run the large language model to obtain multiple edge device clusters, and the sum of the computing power resources of all edge devices within the edge device cluster is greater than the minimum computing power resources required to run the large language model;
[0033] The device information extraction unit is used to extract the device parameter information of all edge devices within each edge device cluster, and the device parameter information includes at least computing power and memory capacity.
[0034] Preferably, the model loading module includes:
[0035] The pruning ratio determination unit is used to determine the pruning ratio layer by layer, describe the overall performance under different pruning ratios through a metric function, obtain a ratio score pair, and embed the ratio score pair into a continuous representation space;
[0036] The model pruning unit is used to iteratively update along the gradient direction of the evaluator starting from the initialization point in the continuous representation space, determine the optimal pruning ratio, prune the large language model, and obtain a compressed model;
[0037] The resource partitioning unit is used to partition the edge device cluster based on the computing power resources required by the compressed model to obtain fixed resources and dynamic resources.
[0038] Preferably, the dynamic compensation module includes:
[0039] The model loading unit is used to split the compressed model into multiple independent model units, load the independent model units on the corresponding edge devices, and the number of edge devices is the same as the number of independent model units;
[0040] The task monitoring unit is used to obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge devices, and obtain task processing monitoring results;
[0041] The dynamic assistance unit is used to evaluate the data processing speed of each part of the compressed model based on the task processing monitoring results, load the corresponding independent model units according to the evaluation results, and assist in processing subsequent tasks.
[0042] Preferably, the pruning ratio is expressed as , and the metric function is expressed as:
[0043] ;
[0044] where is the zero-shot perplexity, which is used to quantify the generation ability, and respectively represent the latency and energy consumption limits specific to the edge device, and respectively represent the latency and energy consumption corresponding to the given ratio, is an indicator function. If the condition holds, it returns 1; otherwise, it returns 0. and are both penalty factors. is the th pruning ratio.
[0045] A method for optimizing large language models for edge devices provided by the present invention adapts to compress the large language model, performs compression in different ratios according to different edge devices, thereby obtaining a customized compressed model, splits the compressed model and loads it on each edge device, and repeats the loading of part of the compressed model by the edge device to ensure the balance of data processing, improve the data processing speed, and enable edge devices with lower performance to also run large language models. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flowchart of a method for optimizing large language models for edge devices provided by an embodiment of the present invention;
[0047] Figure 2 is a flowchart of the step of dividing edge devices to obtain multiple edge device clusters and extracting device parameter information of each edge device provided by an embodiment of the present invention;
[0048] Figure 3 is a flowchart of the step of pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and dividing the resources of the edge device cluster to obtain fixed resources and dynamic resources provided by an embodiment of the present invention;
[0049] Figure 4 is a flowchart of the step of splitting and loading the compressed model based on the fixed resources, processing the data processing task, and dynamically locally repeating the loading of the compressed model based on the dynamic resources to complete the processing of subsequent tasks provided by an embodiment of the present invention;
[0050] Figure 5 is an architecture diagram of a system for optimizing large language models for edge devices provided by an embodiment of the present invention;
[0051] Figure 6 is an architecture diagram of a device division module provided by an embodiment of the present invention;
[0052] Figure 7 is an architecture diagram of a model loading module provided by an embodiment of the present invention;
[0053] Figure 8 is an architecture diagram of a dynamic compensation module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0055] As Figure 1 shown, it is a flowchart of a method for optimizing a large language model for edge devices provided by an embodiment of the present invention. The method includes:
[0056] S100, construct an edge device network, and the edge device network is overall managed by a cloud server.
[0057] In this step, an edge device network is constructed, a cloud server is set up, and the cloud server provides management services for all edge devices. When an edge device needs to use a large language model, it is added to the edge device network. That is, all edge devices in the edge device network have the need to use the large language model, and the edge devices can communicate with each other through the cloud server to facilitate collaborative operations.
[0058] S200, divide the edge devices to obtain multiple edge device clusters, and extract the device parameter information of each edge device.
[0059] In this step, the edge devices are divided. For edge devices, their computing power is limited, and the memory of the devices themselves is also limited, and they cannot directly run a large language model. Although a large language model with pruning operations can run, its generation ability and accuracy will be greatly reduced, making it difficult to meet the requirements. In the present invention, when an edge device needs to use a large language model, it is added to the edge device network. At this time, multiple edge devices are divided into an edge device cluster according to the situation of the edge devices themselves. An edge device cluster is used to jointly run a large language model. After the division of the edge device cluster is completed, the device parameter information of all edge devices in the edge device cluster is extracted. This is because the complete large language model has a large volume and is relatively cumbersome, and a customized large language model can be generated for each edge device cluster by pruning, without affecting the accuracy, to improve the task processing speed.
[0060] S300, perform pruning processing on the large language model to be loaded according to the device parameter information of the edge devices to obtain a compressed model, and divide the resources of the edge device cluster to obtain fixed resources and dynamic resources.
[0061] In this step, the large language model to be loaded is pruned according to the device parameter information of the edge device. Traditional model compression methods often aim to reduce the number of parameters and the amount of computation, but may ignore the impact on the generation performance. In the present invention, model compression is regarded as a generation task again. The goal is to maximize the generation ability of the large language model under the premise of meeting the resource constraints of the edge device. Considering the memory constraint of the edge device and the size of the model to be deployed, the overall reduction ratio of the large language model is first determined, and an exploration-exploitation strategy is used to determine the pruning ratio layer by layer. The score situation under each pruning ratio, that is, the ratio-score pair, is determined. An encoder-evaluator-decoder framework is constructed. This architecture includes a single-layer long short-term memory (LSTM) network as the encoder and decoder, and a feedforward neural network as the evaluator. The ratio-score pair is embedded into a continuous representation space to determine the best pruning scheme to obtain a compressed model. The above compressed model is customized according to the edge device cluster and still has a high generation ability under the condition of meeting the computing power constraint. Subsequently, in order to facilitate the split loading of the compressed model, the resources of the edge device cluster are partitioned, and the resources of all edge devices are quantified, including memory and computing power. The resources are partitioned into fixed resources and dynamic resources. The fixed resources come from each edge device and are used to jointly run the compressed model, while the dynamic resources are used to dynamically load the corresponding partial content of the compressed model according to the running situation of the compressed model in each edge device, adjust the overall task processing progress, and ensure the consistency of the operation of the compressed model, avoiding the situation that the overall task processing speed is reduced due to insufficient local resources.
[0062] S400, split and load the compressed model based on the fixed resources, process the data processing task, and perform dynamic local repeated loading of the compressed model based on the dynamic resources to complete the processing of subsequent tasks.
[0063] In this step, the compressed model is split and loaded based on the fixed resources. The compressed model is split into multiple parts according to the number of edge devices, that is, independent model units. Each edge device is used to run an independent model unit. When the computing power of the edge device is weak or the number is large, multiple edge devices can jointly run an independent model unit. At this time, the fixed resources of the edge device are allocated, and all tasks initiated by the edge devices are processed through the jointly formed compressed model. During the task processing, the processing time of each task in each independent model unit is counted, and the independent model unit that limits the overall data processing speed of the compressed model is determined according to the statistical results. Then, the dynamic resources are used to jointly load this independent model unit, and the processing speed of this part of the task is increased through one or more identical independent model units to improve the overall task processing speed.
[0064] As Figure 2 shown, as a preferred embodiment of the present invention, the step of dividing the edge devices to obtain multiple edge device clusters and extracting the device parameter information of each edge device specifically includes:
[0065] S201, count the hardware data of each edge device, quantify the computing power resources of the edge device to obtain a quantization result.
[0066] In this step, when counting the hardware data of each edge device, since the computing capabilities and memory bandwidths of edge devices are different, it is unreasonable to simply rely on the compression ratio of the available memory evaluation model. The present invention introduces two evaluation indicators - the first Token return time (TTFT) and the generation time of each subsequent Token (TPOP) to evaluate the computing performance of the target device.
[0067] S202, cluster the edge devices according to the minimum computing power resources required to run the large language model to obtain multiple edge device clusters, and the sum of the computing power resources of all edge devices within the edge device cluster is greater than the minimum computing power resources required to run the large language model.
[0068] In this step, when clustering the edge devices according to the minimum computing power resources required to run the large language model, the sum of the computing power resources of the selected edge devices should be greater than the minimum computing power resources required by the large language model, so as to ensure that the large language model can run completely. When dividing, the computing power resources included in the edge device cluster can also be multiple times the minimum computing power resources required by the large language model, so that multiple large language models can be run simultaneously by the edge device cluster in the subsequent process.
[0069] S203, extract the device parameter information of all edge devices within each edge device cluster, and the device parameter information at least includes computing power and memory capacity.
[0070] In this step, when extracting the device parameter information of all edge devices within each edge device cluster, in order to facilitate pruning the large language model and constructing a customized compression model for the edge device cluster, it is necessary to collect the device information of the edge cluster for analysis and generating pruning strategies.
[0071] As Figure 3 shown, as a preferred embodiment of the present invention, the step of pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compression model and partitioning the resources of the edge device cluster to obtain fixed resources and dynamic resources specifically includes:
[0072] S301. Determine the pruning ratio layer by layer, describe the overall performance under different pruning ratios through a metric function, obtain ratio-score pairs, and embed the ratio-score pairs into a continuous representation space.
[0073] S302. In the continuous representation space, start from the initialization point and iteratively update along the gradient direction of the evaluator to determine the optimal pruning ratio, prune the large language model, and obtain a compressed model.
[0074] In this step, when determining the pruning ratio layer by layer, considering the memory limitation of the edge device and the size of the model to be deployed, first determine the overall reduction ratio of the large language model. For fine-grained optimization, an exploration-exploitation strategy is adopted to determine the pruning ratio layer by layer. The pruning ratio is expressed as , and a heuristic-based pruning method is used to generate high-quality ratios, and the random pruning ratio is used as part of the exploration process. Compared with the prior art that mainly focuses on reducing the model size and computational overhead, the present invention integrates a wider range of criteria, including generation ability, inference latency, and energy consumption, and defines an overall metric function to describe the overall performance of a given . The goal is to optimize for resource-constrained edge devices:
[0075] ;
[0076] where is the zero-shot perplexity, which is used to quantify the generation ability. A lower value indicates that the model prediction is more accurate. and represent the latency and energy consumption limits specific to the edge device respectively. and represent the latency and energy consumption corresponding to the given ratio respectively. is an indicator function that returns 1 if the condition holds, otherwise returns 0. Configurations exceeding the thresholds and will be penalized by the factors and specified by the developer. Finally, we obtain comprehensive ratio-score pairs, expressed as .
[0077] Construction of the continuous customization space: The prior art usually adopts a discrete pruning space, which limits the diversity of configurations. To more comprehensively explore the customization optimization space, the present invention establishes an encoder-evaluator-decoder framework. This architecture includes a single-layer long short-term memory (LSTM) network as the encoder and decoder, and a feed-forward neural network as the evaluator to embed the ratio-score pairs into a continuous representation space.
[0078] Gradient-based Customized Optimization: In the trained representation space, use gradient-based optimization methods to identify the best pruning configuration. Starting from the initialization point and perform iterative updates along the gradient direction of the evaluator:
[0079] ;
[0080] where represents the representation of the best pruning configuration, is the step size used in the gradient update. (1.5) Best Customized Generation: Based on the best representation , generate the corresponding pruning ratio through the trained decoder. Adopt a beam search strategy to systematically explore potential configurations until the best pruning scheme that meets the performance requirements is generated. Finally, through structured pruning, remove unnecessary model components to achieve efficient compression of the large language model.
[0081] S303. Divide the resources of the edge device cluster based on the computing power resources required by the compressed model to obtain fixed resources and dynamic resources.
[0082] In this step, divide the resources of the edge device cluster based on the computing power resources required by the compressed model. When dividing, divide according to the resources that each edge device can allocate. For example, if the total computing power resources that edge device A can allocate is M, then divide it into two parts according to the ratio, namely fixed resources and dynamic resources. The fixed resources are used to load fixed independent model units, and the dynamic resources of each edge device are used to jointly load one or more independent model units.
[0083] As Figure 4 shown, as a preferred embodiment of the present invention, the steps of splitting and loading the compressed model based on the fixed resources, processing the data processing task, and dynamically and locally repeating the loading of the compressed model based on the dynamic resources to complete the processing of subsequent tasks specifically include:
[0084] S401. Split the compressed model into multiple independent model units, and load the independent model units on the corresponding edge devices. The number of edge devices is the same as the number of independent model units.
[0085] In this step, split the compressed model into multiple independent model units. When the number of independent model units is the same as the number of edge devices, load the independent model units into each edge device. When the number of edge devices is greater than the number of independent model units, multiple edge devices jointly load the same independent model unit.
[0086] S402. Obtain a data processing task, import the data processing task into each independent model unit, monitor the data processing results of the edge device, and obtain a task processing monitoring result.
[0087] In this step, obtain a data processing task, import the data processing task into each independent model unit, and execute the data processing task according to the data processing flow of each independent model unit. For example, if a compression model contains 4 independent model units, namely a, b, c, and d, and the execution processes are a, b, c, and d respectively, then the data processing task is imported into a, and after being processed by a, the output is imported into b until the final processing result is obtained through d. During this process, the time taken for the task to pass through each independent model unit is counted to obtain a task processing monitoring result.
[0088] S403. Evaluate the data processing speed of each part of the compression model based on the task processing monitoring result, and load the corresponding independent model unit according to the evaluation result to assist in processing subsequent tasks.
[0089] In this step, evaluate the data processing speed of each part of the compression model based on the task processing monitoring result, number each task, determine the processing speed of each task passing through each independent model unit. The processing speed S is represented by the ratio of the number of tokens H corresponding to each task to the time T, that is, S = H / T. Construct a module processing speed coordinate (n, S), where n is the task number. For example, if the task is executed 5 times, each independent model unit will correspond to five module processing speed coordinates. Based on this, construct a processing speed prediction function. Specifically, import the five groups of module processing speed coordinates into function fitting software such as matlab to fit and obtain the speed prediction function. According to the speed prediction function, when executing the next task, determine the processing speed of each independent model unit, and select the independent model unit with the slowest processing speed (assumed to be a), and separately load one such independent model unit a through the idle dynamic resources, that is, there are two such independent model units a at this time. When the computing power resources of the dynamic resources are sufficient, multiple independent model units with slow processing speeds can be selected for loading to ensure the fastest global processing speed of each task and avoid affecting the global processing speed due to the slow processing speed of a certain independent model unit.
[0090] Large language models are usually composed of multiple decoder layers with the same structure stacked together, but the actual contributions of these layers are heterogeneous: the first few layers are crucial for extracting input Prompt features, while the last few layers play a more significant role in generating the output. Due to the heterogeneity of the input and output data streams, the generation effects also vary between different layers. Finally, existing compression technologies often do not fully consider hardware characteristics, resulting in the compressed model being unable to fully utilize hardware resources.
[0091] Based on the hardware resources of edge devices and the characteristics of large models, the present invention designs a dedicated compression technology adapted to target devices, which significantly reduces the number of model parameters and computational complexity while ensuring model performance, making it more suitable for edge devices with limited resources. First, a heuristic algorithm is used to comprehensively evaluate the compression scheme to ensure that the service level agreement (SLO) is not violated during model inference and to maximize the inference efficiency of the device. Subsequently, through ablation experiments, the actual contribution of each layer to generating tokens and its impact on the inference speed are analyzed on multiple datasets. Based on these analyses, the compression ratio of each layer is determined to maximize the accuracy of the model within a given compression ratio range and improve the inference efficiency. Finally, through a co-optimization method, the hardware characteristics of edge devices are combined with the model compression strategy to establish a mapping relationship between the hardware characteristics and the model layer characteristics. For different hardware configurations, the compression strategy is dynamically adjusted to ensure that the compressed model achieves optimal performance on edge devices. In this process, factors such as the cache size, memory bandwidth, and computing unit architecture of the device are comprehensively considered to optimize the computational graph and memory access pattern of the model, giving full play to the hardware advantages and ultimately achieving efficient edge inference.
[0092] As Figure 5 shown, a large language model optimization system for edge devices provided by an embodiment of the present invention includes:
[0093] A device network construction module 100 for constructing an edge device network, and the edge device network is overall managed by a cloud server.
[0094] In this system, the device network construction module 100 constructs an edge device network, sets up a cloud server, and provides management services for all edge devices through the cloud server. When an edge device needs to use a large language model, it joins the edge device network. That is, all edge devices in the edge device network have the need to use the large language model, and edge devices can communicate with each other through the cloud server to facilitate collaborative operations.
[0095] A device partitioning module 200 for partitioning edge devices to obtain multiple edge device clusters and extracting device parameter information of each edge device.
[0096] In this system, the device division module 200 divides edge devices. For edge devices, their computing power is limited and the memory of the devices themselves is also limited, so they cannot directly run large language models. Although the large language model with pruning operations can run, its generation ability and accuracy will be greatly reduced, making it difficult to meet the requirements. In the present invention, when an edge device needs to use a large language model, it joins the edge device network. At this time, multiple edge devices are divided into an edge device cluster according to the situation of the edge device itself. An edge device cluster is used to jointly run a large language model. After the division of the edge device cluster is completed, the device parameter information of all edge devices in the edge device cluster is extracted. This is because the complete large language model is large in volume and cumbersome, and a customized large language model can be generated for each edge device cluster by means of pruning processing without affecting the accuracy, so as to improve the task processing speed.
[0097] The model loading module 300 is used to perform pruning processing on the large language model to be loaded according to the device parameter information of the edge device, obtain a compressed model, and divide the resources of the edge device cluster to obtain fixed resources and dynamic resources.
[0098] In this system, the model loading module 300 prunes the large language model to be loaded according to the device parameter information of the edge device. Traditional model compression methods often aim to reduce the number of parameters and the amount of computation, but may ignore the impact on the generation performance. In the present invention, model compression is regarded as a generation task again. The goal is to maximize the generation ability of the large language model under the premise of meeting the resource constraints of the edge device. Considering the memory constraints of the edge device and the size of the model to be deployed, the overall reduction ratio of the large language model is first determined. An exploration-exploitation strategy is adopted to determine the pruning ratio layer by layer, and the score situation under each pruning ratio, that is, the ratio-score pair, is determined. An encoder-evaluator-decoder framework is constructed. This architecture includes a single-layer long short-term memory (LSTM) network as the encoder and decoder, and a feedforward neural network as the evaluator. The ratio-score pair is embedded into a continuous representation space to determine the best pruning scheme to obtain the compressed model. The above compressed model is customized according to the edge device cluster. Under the condition of meeting the computing power constraints, it still has a high generation ability. Subsequently, in order to facilitate the split loading of the compressed model, the edge device cluster is resource-partitioned, and the resources of all edge devices are quantified, including memory and computing power. The resources are partitioned into fixed resources and dynamic resources. The fixed resources come from each edge device and are used to jointly run the compressed model, while the dynamic resources are used to dynamically load the corresponding parts of the compressed model according to the running situation of the compressed model in each edge device, adjust the overall task processing progress, ensure the consistency of the operation of the compressed model, and avoid the situation where the overall task processing speed is reduced due to local resource shortages.
[0099] The dynamic compensation module 400 is used to split and load the compressed model based on the fixed resources, process the data processing tasks, and perform dynamic local repeated loading of the compressed model based on the dynamic resources to complete the processing of subsequent tasks.
[0100] In this system, the dynamic compensation module 400 splits and loads the compressed model based on the fixed resources, splits the compressed model into multiple parts according to the number of edge devices, that is, independent model units. Each edge device is used to run an independent model unit. When the computing power of the edge device is weak or the number is large, multiple edge devices can jointly run an independent model unit. At this time, the fixed resources of the edge device are allocated, and all tasks initiated by the edge devices are processed through the jointly constructed compressed model. During the task processing, the processing time of each task in each independent model unit is statistically analyzed, and the independent model unit that limits the overall data processing speed of the compressed model is determined according to the statistical results. Then, the dynamic resources are used to jointly load this independent model unit, and the processing speed of this part of the task is increased through one or more identical independent model units to improve the overall task processing speed.
[0101] As Figure 6 shown, as a preferred embodiment of the present invention, the device partitioning module 200 includes:
[0102] A computing power quantization unit 201, configured to count the hardware data of each edge device, quantify the computing power resources of the edge device, and obtain a quantization result.
[0103] In this module, the computing power quantization unit 201 counts the hardware data of each edge device. Since the computing capabilities and memory bandwidths of edge devices are different, it is unreasonable to solely rely on the available memory to evaluate the compression rate of the model. The present invention introduces two evaluation metrics - the first Token return time (TTFT) and the generation time of each subsequent Token (TPOP) - to evaluate the computing performance of the target device.
[0104] A cluster partitioning unit 202, configured to partition edge devices into multiple edge device clusters according to the minimum computing power resources required to run a large language model, where the sum of the computing power resources of all edge devices within the edge device cluster is greater than the minimum computing power resources required to run the large language model.
[0105] In this module, the cluster partitioning unit 202 partitions edge devices according to the minimum computing power resources required to run a large language model. The sum of the computing power resources of the selected edge devices should be greater than the minimum computing power resources required by the large language model, so as to ensure the complete operation of the large language model. During partitioning, the computing power resources included in the edge device cluster can also be multiples of the minimum computing power resources required by the large language model, so that multiple large language models can be run simultaneously by the edge device cluster in subsequent processes.
[0106] A device information extraction unit 203, configured to extract the device parameter information of all edge devices within each edge device cluster, where the device parameter information includes at least computing power and memory capacity.
[0107] In this module, the device information extraction unit 203 extracts the device parameter information of all edge devices within each edge device cluster. To facilitate pruning the large language model and constructing a customized compression model for the edge device cluster, it is necessary to collect the device information of the edge cluster for analysis and generating pruning strategies.
[0108] As Figure 7 shown, as a preferred embodiment of the present invention, the model loading module 300 includes:
[0109] The pruning ratio determination unit 301 is used to determine the pruning ratio layer by layer, describe the overall performance under different pruning ratios through a metric function, obtain ratio score pairs, and embed the ratio score pairs into a continuous representation space.
[0110] The model pruning unit 302 is used to iteratively update along the gradient direction of the evaluator starting from the initialization point in the continuous representation space, determine the optimal pruning ratio, prune the large language model, and obtain a compressed model.
[0111] In this module, when determining the pruning ratio layer by layer, considering the memory limitation of the edge device and the size of the model to be deployed, first determine the overall reduction ratio of the large language model. In order to perform fine-grained optimization, an exploration-exploitation strategy is adopted to determine the pruning ratio layer by layer. The pruning ratio is expressed as , and a heuristic-based pruning method is used to generate high-quality ratios, and the random pruning ratio is used as part of the exploration process. Compared with the prior art that mainly focuses on reducing the model size and computational overhead, the present invention integrates a wider range of criteria, including generative ability, inference latency, and energy consumption, and defines an overall metric function to describe the overall performance of a given . The goal is to optimize for resource-constrained edge devices:
[0112] ;
[0113] where, is the zero-shot perplexity, which is used to quantify the generative ability. A lower value indicates that the model prediction is more accurate. and represent the edge device-specific latency and energy consumption limits respectively. and represent the latency and energy consumption corresponding to the given ratio respectively. is an indicator function, which returns 1 if the condition holds, otherwise returns 0. Configurations exceeding the thresholds and will be penalized by the factors and specified by the developer. Finally, we obtain comprehensive ratio score pairs, expressed as .
[0114] Construction of the continuous customization space: The prior art usually adopts a discrete pruning space, which limits the diversity of configurations. To more comprehensively explore the customization optimization space, the present invention establishes an encoder-evaluator-decoder framework. This architecture includes a single-layer long short-term memory (LSTM) network as the encoder and decoder, and a feed-forward neural network as the evaluator, embedding the ratio score pairs into a continuous representation space.
[0115] Gradient-based Customized Optimization: In the trained representation space, use gradient-based optimization methods to identify the best pruning configuration. Starting from the initialization point and iteratively update along the gradient direction of the evaluator:
[0116] ;
[0117] where represents the representation of the best pruning configuration, is the step size used in the gradient update. (1.5) Best Customized Generation: Based on the best representation generate the corresponding pruning ratio through the trained decoder. Adopt a beam search strategy to systematically explore potential configurations until the best pruning scheme that meets the performance requirements is generated. Finally, through structured pruning, remove unnecessary model components to achieve efficient compression of the large language model.
[0118] The resource partitioning unit 303 is used to partition the edge device cluster based on the computing power resources required by the compressed model to obtain fixed resources and dynamic resources.
[0119] In this module, the resource partitioning unit 303 partitions the edge device cluster based on the computing power resources required by the compressed model. When partitioning, it partitions according to the resources that each edge device can allocate. For example, if the total computing power resources that edge device A can allocate is M, then it is divided into two parts according to the ratio, namely fixed resources and dynamic resources. The fixed resources are used to load fixed independent model units, and the dynamic resources of each edge device are used to jointly load one or more independent model units.
[0120] As Figure 8 shown, as a preferred embodiment of the present invention, the dynamic compensation module 400 includes:
[0121] The model loading unit 401 is used to split the compressed model into multiple independent model units and load the independent model units on the corresponding edge devices. The number of edge devices is the same as the number of independent model units.
[0122] In this module, the model loading unit 401 splits the compressed model into multiple independent model units. When the number of independent model units is the same as the number of edge devices, the independent model units are loaded into each edge device. When the number of edge devices is greater than the number of independent model units, the same independent model unit is jointly loaded by multiple edge devices.
[0123] The task monitoring unit 402 is used to obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge device, and obtain task processing monitoring results.
[0124] In this module, the task monitoring unit 402 obtains data processing tasks, imports the data processing tasks into each independent model unit, and executes the data processing tasks according to the data processing flow of each independent model unit. For example, if the compression model contains 4 independent model units, namely a, b, c, and d, and the execution processes are a, b, c, and d respectively, then the data processing task is imported into a, and after being processed by a, the output is imported into b until the final processing result is obtained through d. During this process, the time taken for the task to pass through each independent model unit is counted to obtain the task processing monitoring results.
[0125] The dynamic assistance unit 403 is used to evaluate the data processing speed of each part of the compression model based on the task processing monitoring results, and load the corresponding independent model unit according to the evaluation results to assist in processing subsequent tasks.
[0126] In this module, the dynamic assistance unit 403 evaluates the data processing speed of each part of the compression model based on the task processing monitoring results, numbers each task, determines the processing speed of each task passing through each independent model unit. The processing speed S is represented by the ratio of the number of tokens H corresponding to each task to the time T, that is, S = H / T. A module processing speed coordinate (n, S) is constructed, where n is the task number. For example, if 5 tasks are executed, each independent model unit will correspond to five module processing speed coordinates. Based on this, a processing speed prediction function is constructed. Specifically, the five groups of module processing speed coordinates are imported into function fitting software such as matlab, and the speed prediction function is obtained by fitting. According to the speed prediction function, when executing the next task, the processing speed of each independent model unit is determined, and the independent model unit with the slowest processing speed (assumed to be a) is selected, and an additional independent model unit a of this type is loaded separately through the idle dynamic resources. That is, there are two independent model units a at this time. When the computing power resources of the dynamic resources are sufficient, multiple independent model units with slow processing speeds can be selected for loading to ensure the fastest global processing speed of each task and avoid affecting the global processing speed due to the slow processing speed of a certain independent model unit.
[0127] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for optimizing large language models for edge devices, characterized in that, The method includes: Construct an edge device network, and the edge device network is overall managed through a cloud server; Divide the edge devices to obtain multiple edge device clusters, and extract the device parameter information of each edge device; Prune the large language model to be loaded according to the device parameter information of the edge devices to obtain a compressed model, and divide the resources of the edge device clusters to obtain fixed resources and dynamic resources; Based on the fixed resources, split and load the compressed model, process the data processing tasks, and based on the dynamic resources, perform dynamic local repeated loading on the compressed model to complete the processing of subsequent tasks; In the step of dividing the edge devices to obtain multiple edge device clusters and extracting the device parameter information of each edge device, the calculation performance of the target device is evaluated by using the Token return time and the generation time of each subsequent Token as evaluation indicators, and the edge devices are clustered according to the minimum computing power resources required to run the large language model. The sum of the computing power resources of the selected edge devices is greater than the minimum computing power resources required by the large language model, and the device parameter information includes at least computing power and memory capacity; In the step of partitioning edge devices to obtain multiple edge device clusters and extracting device parameter information of each edge device, different pruning ratios are determined , by defining a metric function to describe the overall performance of a given , the metric function is aimed at optimizing resource-constrained edge devices, and the overall performance is expressed as: ; Among them, is the zero-shot perplexity, and represent the edge device-specific latency and energy consumption limits respectively, and represent the latency and energy consumption corresponding to the given ratio respectively, is the indicator function, which returns 1 if the condition holds, otherwise returns 0; configurations exceeding the thresholds and will penalize the specified factors and to obtain the ratio score pair, denoted as ; Embed the ratio score into the continuous representation space by constructing an encoder-evaluator-decoder framework. In the representation space, use the gradient-based optimization method to identify the best pruning configuration, and perform structured pruning on the large language model based on the obtained best pruning configuration; The step of splitting and loading the compressed model based on the fixed resources, processing the data processing tasks, and performing dynamic local repeated loading on the compressed model based on the dynamic resources to complete the processing of subsequent tasks specifically includes: Split the compressed model into multiple independent model units, and load the independent model units on the corresponding edge devices. The number of edge devices is the same as the number of independent model units; Obtain the data processing tasks, import the data processing tasks into each independent model unit, and monitor the data processing results of the edge devices to obtain the task processing monitoring results; Evaluate the data processing speed of each part of the compressed model based on the task processing monitoring results, and load the corresponding independent model unit according to the evaluation results to assist in processing subsequent tasks.
2. The large language model optimization method for edge devices according to claim 1, wherein The step of dividing the edge devices to obtain multiple edge device clusters and extracting the device parameter information of each edge device specifically includes: Count the hardware data of each edge device, quantify the computing power resources of the edge device to obtain a quantization result; Cluster the edge devices according to the minimum computing power resources required to run the large language model to obtain multiple edge device clusters. The sum of the computing power resources of all edge devices in the edge device cluster is greater than the minimum computing power resources required to run the large language model; Extract the device parameter information of all edge devices in each edge device cluster, and the device parameter information includes at least computing power and memory capacity.
3. The large language model optimization method for edge devices according to claim 1, characterized in that The steps of pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and partitioning the resources of the edge device cluster to obtain fixed resources and dynamic resources specifically include: Determine the pruning ratio layer by layer, describe the overall performance under different pruning ratios through a metric function, obtain a pair of ratio scores, and embed the pair of ratio scores into a continuous representation space; In the continuous representation space, start from the initialization point and perform iterative updates along the gradient direction of the evaluator to determine the optimal pruning ratio, and prune the large language model to obtain a compressed model; Partition the resources of the edge device cluster based on the computing power resources required by the compressed model to obtain fixed resources and dynamic resources.
4. A large language model optimization system for edge devices, characterized in that, The system includes: A device network construction module for constructing an edge device network, and the edge device network is overall managed by a cloud server; A device partitioning module for partitioning edge devices to obtain multiple edge device clusters, and extracting the device parameter information of each edge device; A model loading module for pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and partitioning the resources of the edge device cluster to obtain fixed resources and dynamic resources; A dynamic compensation module for splitting and loading the compressed model based on the fixed resources, processing the data processing task, and performing dynamic local repeated loading on the compressed model based on the dynamic resources to complete the processing of subsequent tasks; The device partitioning module uses the Token return time and the generation time of each subsequent Token as evaluation metrics to evaluate the computing performance of the target device, and partitions the edge devices into clusters according to the minimum computing power resources required to run the large language model. The sum of the computing power resources of the selected edge devices is greater than the minimum computing power resources required by the large language model, and the device parameter information includes at least computing power and memory capacity; The belonging model loading module determines different pruning ratios , and defines a metric function to describe the overall performance of a given . The goal of the metric function is to optimize for resource-constrained edge devices, and the overall performance is expressed as: ; Among them, is the zero-shot perplexity, and represent the edge device-specific latency and energy consumption limits respectively, and represent the latency and energy consumption corresponding to the given ratio respectively, is the indicator function, which returns 1 if the condition holds, otherwise returns 0; configurations exceeding the thresholds and will penalize the specified factors and to obtain the ratio score pair, denoted as ; Embed the pair of ratio scores into the continuous representation space by constructing an encoder-evaluator-decoder framework. In the representation space, use a gradient-based optimization method to identify the best pruning configuration, and perform structured pruning on the large language model based on the obtained best pruning configuration; The dynamic compensation module includes: A model loading unit for splitting the compressed model into multiple independent model units, and loading the independent model units on the corresponding edge devices. The number of edge devices is the same as the number of independent model units; A task monitoring unit for obtaining data processing tasks, importing the data processing tasks into each independent model unit, and monitoring the data processing results of the edge devices to obtain task processing monitoring results; A dynamic assistance unit for evaluating the data processing speed of each part of the compressed model based on the task processing monitoring results, and loading the corresponding independent model unit according to the evaluation results to assist in processing subsequent tasks.
5. The large language model optimization system for edge devices according to claim 4, wherein The device partitioning module includes: A computing power quantization unit for counting the hardware data of each edge device, quantizing the computing power resources of the edge device to obtain a quantization result; A cluster division unit, which is used to divide edge devices into clusters according to the minimum computing power resources required to run a large language model, so as to obtain multiple edge device clusters, and the sum of the computing power resources of all edge devices within the edge device cluster is greater than the minimum computing power resources required to run the large language model; A device information extraction unit, which is used to extract the device parameter information of all edge devices within each edge device cluster, and the device parameter information at least includes computing power and memory capacity.
6. The large language model optimization system for edge devices according to claim 4, wherein The model loading module includes: A pruning ratio determination unit, which is used to determine the pruning ratio layer by layer, describe the overall performance under different pruning ratios through a metric function, obtain a ratio score pair, and embed the ratio score pair into a continuous representation space; A model pruning unit, which is used to iteratively update along the gradient direction of the evaluator starting from the initialization point in the continuous representation space, determine the optimal pruning ratio, prune the large language model, and obtain a compressed model; A resource division unit, which is used to divide the resources of the edge device cluster based on the computing power resources required by the compressed model to obtain fixed resources and dynamic resources.
Citation Information
Patent Citations
Edge computing scheduling system based on knowledge graph
CN119065852A