Large language model optimization method and system for edge device
By building networks, dividing clusters, pruning and loading large language models on edge devices, the problem of resource limitation of edge devices is solved, and the effect of efficiently running large language models on edge devices is achieved.
Patent Information
- Application Number
- CN202510416611.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Due to the limitations of edge devices in terms of computing power, memory capacity and battery capacity, directly running a complete large language model is often not feasible on edge devices.
By building an edge device network, dividing edge device clusters, extracting device parameter information, pruning large language models, splitting and loading models, and dynamic local repeated loading to adapt to resource limitations of edge devices.
It realizes the possibility of running large language models on edge devices, improves task processing speed and generation capabilities, and ensures that edge devices with lower performance can also effectively run large language models.
Smart Images

Figure CN119918622A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large language models, and in particular, relates to a large language model optimization method and system for edge devices. Background Art
[0002] Large Language Model (LLM) is a complex system based on deep learning. It learns rich language knowledge and contextual understanding capabilities by pre-training on large-scale text data. These models usually adopt the Transformer architecture and can generate high-quality natural language text. They are widely used in natural language processing tasks such as text generation, machine translation, question-answering systems, and sentiment analysis. Through multi-layer self-attention mechanisms and feedforward neural networks, large language models can capture complex dependencies when processing long texts and provide accurate and coherent output.
[0003] With the widespread use of large language models in various applications, their huge parameter size and computational complexity have brought great challenges to the deployment of edge devices. Due to the limitations of edge devices in computing power, memory capacity, and battery capacity, it is often not feasible to directly run a complete large language model on edge devices. Summary of the invention
[0004] The purpose of the present invention is to provide a large language model optimization method for edge devices, aiming to solve the problem that it is often not feasible to directly run a complete large language model on edge devices due to the limitations of edge devices in computing power, memory capacity and battery capacity.
[0005] The present invention is implemented as follows: a large language model optimization method for edge devices, the method comprising: Building an edge device network, which is centrally managed by a cloud server; Divide the edge devices to obtain multiple edge device clusters, and extract device parameter information of each edge device; Prune the large language model that needs to be loaded according to the device parameter information of the edge device to obtain a compressed model, and divide the resources of the edge device cluster to obtain fixed resources and dynamic resources; The compression model is split and loaded based on fixed resources to process data processing tasks, and the compression model is dynamically and locally repeatedly loaded based on dynamic resources to complete the processing of subsequent tasks.
[0006] Preferably, the step of dividing the edge devices to obtain multiple edge device clusters and extracting device parameter information of each edge device specifically includes: Collect statistics on the hardware data of each edge device, quantify the computing resources of the edge device, and obtain quantitative results; Clustering edge devices according to minimum computing resources required to run the large language model to obtain multiple edge device clusters, wherein the sum of computing resources of all edge devices in the edge device cluster is greater than the minimum computing resources required to run the large language model; The device parameter information of all edge devices in each edge device cluster is extracted, where the device parameter information includes at least computing power and memory capacity.
[0007] Preferably, the steps of pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, dividing the resources of the edge device cluster to obtain fixed resources and dynamic resources specifically include: Determine the pruning ratio of each layer, describe the overall performance under different pruning ratios through the metric function, obtain the ratio score pair, and embed the ratio score pair into the continuous representation space; In the continuous representation space, iterative updates are performed along the gradient direction of the estimator starting from the initialization point to determine the optimal pruning ratio, and the large language model is pruned to obtain a compressed model. Based on the computing power resources required by the compression model, the edge device cluster is divided into fixed resources and dynamic resources.
[0008] Preferably, the steps of splitting and loading the compression model based on fixed resources, processing the data processing task, and dynamically and locally repeatedly loading the compression model based on dynamic resources to complete the processing of subsequent tasks specifically include: Splitting the compressed model into multiple independent model units, loading the independent model units on corresponding edge devices, where the number of edge devices is the same as the number of independent model units; Obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge device, and obtain the task processing monitoring results; Based on the task processing monitoring results, the data processing speed of each part of the compression model is evaluated, and the corresponding independent model unit is loaded according to the evaluation results to assist in the processing of subsequent tasks.
[0009] Preferably, the pruning ratio Expressed as , the metric function is expressed as: ; in, is the zero-shot perplexity, used to quantify the generative ability, and represent the latency and energy consumption constraints specific to edge devices, respectively. and Represent the delay and energy consumption corresponding to a given ratio, respectively. It is an indicator function. If the condition is met, it returns 1, otherwise it returns 0. and are penalty factors, For the Pruning ratio.
[0010] Another object of the present invention is to provide a large language model optimization system for edge devices, the system comprising: A device network construction module is used to construct an edge device network, and the edge device network is centrally managed by a cloud server; The device partitioning module is used to partition edge devices to obtain multiple edge device clusters and extract device parameter information of each edge device; The model loading module is used to prune the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and divide the resources of the edge device cluster to obtain fixed resources and dynamic resources; The dynamic compensation module is used to split and load the compression model based on fixed resources, process the data processing tasks, and dynamically and locally repeatedly load the compression model based on dynamic resources to complete the processing of subsequent tasks.
[0011] Preferably, the device division module includes: The computing power quantification unit is used to count the hardware data of each edge device, quantify the computing power resources of the edge device, and obtain the quantified results; A cluster division unit, used to cluster edge devices according to the minimum computing power resources required to run the large language model to obtain multiple edge device clusters, wherein the sum of computing power resources of all edge devices in the edge device cluster is greater than the minimum computing power resources required to run the large language model; The device information extraction unit is used to extract device parameter information of all edge devices in each edge device cluster, and the device parameter information at least includes computing power and memory capacity.
[0012] Preferably, the model loading module includes: A pruning ratio determination unit is used to determine the pruning ratio of each layer, describe the overall performance under different pruning ratios through a metric function, obtain a ratio score pair, and embed the ratio score pair into a continuous representation space; The model pruning unit is used to iteratively update the estimator in the continuous representation space from the initialization point along the gradient direction of the estimator, determine the optimal pruning ratio, prune the large language model, and obtain a compressed model; The resource partitioning unit is used to partition the resources of the edge device cluster based on the computing resources required by the compression model to obtain fixed resources and dynamic resources.
[0013] Preferably, the dynamic compensation module includes: A model loading unit is used to split the compressed model into multiple independent model units, and load the independent model units on corresponding edge devices, where the number of edge devices is the same as the number of independent model units; The task monitoring unit is used to obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge device, and obtain the task processing monitoring results; The dynamic auxiliary unit is used to evaluate the data processing speed of each part of the compression model based on the task processing monitoring results, load the corresponding independent model unit according to the evaluation results, and perform auxiliary processing on subsequent tasks.
[0014] Preferably, the pruning ratio Expressed as , the metric function is expressed as: ; in, is the zero-shot perplexity, used to quantify the generative ability, and represent the latency and energy consumption constraints specific to edge devices, respectively. and Represent the delay and energy consumption corresponding to a given ratio, respectively. It is an indicator function. If the condition is met, it returns 1, otherwise it returns 0. and are penalty factors, For the Pruning ratio.
[0015] The present invention provides a large language model optimization method for edge devices, which adaptively compresses the large language model and compresses it in different proportions according to different edge devices to obtain a customized compression model. The compression model is split and loaded on each edge device, and some compression models are repeatedly loaded through the edge device to ensure the balance of data processing and improve the data processing speed, so that edge devices with lower performance can also run the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flowchart of a large language model optimization method for edge devices provided by an embodiment of the present invention; Figure 2A flowchart of the steps of dividing edge devices to obtain multiple edge device clusters and extracting device parameter information of each edge device provided in an embodiment of the present invention; Figure 3 A flowchart of the steps of pruning a large language model to be loaded according to device parameter information of an edge device to obtain a compressed model, dividing resources of an edge device cluster to obtain fixed resources and dynamic resources provided in an embodiment of the present invention; Figure 4 A flowchart of the steps of splitting and loading a compression model based on fixed resources, processing a data processing task, dynamically and locally repeatedly loading the compression model based on dynamic resources, and completing the processing of subsequent tasks provided by an embodiment of the present invention; Figure 5 An architecture diagram of a large language model optimization system for edge devices provided by an embodiment of the present invention; Figure 6 An architecture diagram of a device partition module provided by an embodiment of the present invention; Figure 7 An architectural diagram of a model loading module provided by an embodiment of the present invention; Figure 8 An architectural diagram of a dynamic compensation module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0018] like Figure 1 As shown, it is a flowchart of a large language model optimization method for edge devices provided by an embodiment of the present invention, and the method includes: S100, building an edge device network, wherein the edge device network is centrally managed by a cloud server.
[0019] In this step, an edge device network is built and a cloud server is set up to provide management services for all edge devices through the cloud server. When the edge device needs to use a large language model, if it is in the edge device network, that is, all edge devices in the edge device network have the need to use a large language model, the edge devices can communicate with each other through the cloud server to facilitate collaborative work.
[0020] S200, dividing the edge devices to obtain multiple edge device clusters, and extracting device parameter information of each edge device.
[0021] In this step, the edge devices are divided. For edge devices, their computing power is limited, and the memory of the device itself is also limited. They cannot directly run the large language model. Although the large language model can be run through the pruning operation, the generation capacity and accuracy will be greatly reduced, which is difficult to meet the needs. In the present invention, when the edge device needs to use the large language model, it is added to the edge device network. At this time, multiple edge devices are divided into an edge device cluster according to the situation of the edge device itself. An edge device cluster is used to jointly run a large language model. After the division of the edge device cluster is completed, the device parameter information of all edge devices in the edge device cluster is extracted. This is because the complete large language model is large in size and relatively cumbersome. Through pruning, a customized large language model can be generated for each edge device cluster without affecting the accuracy, so as to improve the speed of task processing.
[0022] S300, pruning the large language model that needs to be loaded according to the device parameter information of the edge device to obtain a compressed model, and dividing the resources of the edge device cluster to obtain fixed resources and dynamic resources.
[0023] In this step, the large language model to be loaded is pruned according to the device parameter information of the edge device. Traditional model compression methods often aim to reduce the number of parameters and the amount of calculation, but may ignore the impact on the generation performance. In the present invention, model compression is re-considered as a generation task. The goal is to retain the generation capability of the large language model to the greatest extent while meeting the resource limitations of the edge device. Considering the memory limitations of the edge device and the size of the model to be deployed, the overall reduction ratio of the large language model is first determined, and an exploration-exploitation strategy is used to determine the pruning ratio layer by layer, and the score under each pruning ratio is determined, that is, the ratio score pair, to construct an encoder-estimator-decoder framework, which includes a single-layer long short-term memory (LSTM) network as an encoder and decoder, and a feedforward neural network as an evaluator. , embed the proportional score pairs into the continuous representation space to determine the best pruning scheme to obtain the compression model. The above compression model is customized according to the edge device cluster. Under the condition of meeting the computing power limit, it can still have a high generation capacity. Then, in order to facilitate the splitting and loading of the compression model, the edge device cluster is divided into resources, and the resources of all edge devices are quantified, including memory and computing power. The resources are divided into fixed resources and dynamic resources. The fixed resources come from each edge device and are used to run the compression model together, while the dynamic resources are used to dynamically load part of the corresponding compression model according to the operation of the compression model in each edge device, adjust the overall task processing progress, ensure the consistency of the compression model operation, and avoid the overall task processing speed being reduced due to local resource shortages.
[0024] S400, splitting and loading the compression model based on fixed resources, processing the data processing task, dynamically and locally repeatedly loading the compression model based on dynamic resources, and completing the processing of subsequent tasks.
[0025] In this step, the compression model is split and loaded based on fixed resources. The compression model is split into multiple parts, namely independent model units, according to the number of edge devices. Each edge device is used to run an independent model unit. When the computing power of the edge devices is weak or the number is large, an independent model unit can be jointly run by multiple edge devices. At this time, the fixed resources of the edge devices are allocated, and all tasks initiated by the edge devices are processed by the jointly constructed compression model. During the task processing process, the processing time of each task in each independent model unit is counted, and the independent model unit that limits the overall data processing speed of the compression model is determined according to the statistical results, so as to jointly load the independent model unit with dynamic resources, and improve the processing speed of this part of the task through one or more identical independent model units, so as to improve the overall character processing speed.
[0026] like Figure 2 As shown, as a preferred embodiment of the present invention, the step of dividing the edge devices to obtain multiple edge device clusters and extracting device parameter information of each edge device specifically includes: S201, counting the hardware data of each edge device, quantifying the computing resources of the edge device, and obtaining a quantified result.
[0027] In this step, the hardware data of each edge device is counted. The computing power and memory bandwidth of edge devices vary. It is unreasonable to evaluate the compression rate of the model based solely on available memory. This paper introduces two evaluation indicators - the first token return time (TTFT) and the subsequent token generation time (TPOP) to evaluate the computing performance of the target device.
[0028] S202, clustering the edge devices according to the minimum computing resources required to run the large language model to obtain multiple edge device clusters, wherein the sum of the computing resources of all edge devices in the edge device cluster is greater than the minimum computing resources required to run the large language model.
[0029] In this step, the edge devices are clustered according to the minimum computing resources required to run the large language model. The sum of the computing resources of the selected edge devices should be greater than the minimum computing resources required by the large language model. This ensures that the large language model can run completely. When dividing, the computing resources contained in the edge device cluster can also be multiple times the minimum computing resources required by the large language model, so that in the subsequent process, multiple large language models can be run simultaneously through the edge device cluster.
[0030] S203, extracting device parameter information of all edge devices in each edge device cluster, where the device parameter information at least includes computing capability and memory capacity.
[0031] In this step, the device parameter information of all edge devices in each edge device cluster is extracted. In order to facilitate the pruning of the large language model and build a customized compression model for the edge device cluster, it is necessary to collect the device information of the edge cluster for analysis and generation of pruning strategies.
[0032] like Figure 3 As shown, as a preferred embodiment of the present invention, the steps of pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, dividing the resources of the edge device cluster to obtain fixed resources and dynamic resources specifically include: S301, determine the pruning ratio of each layer, describe the overall performance under different pruning ratios through a metric function, obtain ratio score pairs, and embed the ratio score pairs into a continuous representation space.
[0033] S302, in the continuous representation space, iteratively update along the gradient direction of the evaluator starting from the initialization point, determine the optimal pruning ratio, prune the large language model, and obtain a compressed model.
[0034] In this step, the pruning ratio of each layer is determined. Considering the memory limit of the edge device and the size of the model to be deployed, the overall reduction ratio of the large language model is first determined. In order to perform fine-grained optimization, the exploration-exploitation strategy is adopted to determine the pruning ratio of each layer. Expressed as , using a heuristic-based pruning method to produce high-quality proportions, and using random pruning proportions as part of the exploration process. Compared with the prior art work that mainly focuses on reducing model size and computational overhead, this invention integrates a wider range of criteria, including generation capacity, inference latency, and energy consumption, and defines an overall metric function To describe a given The overall performance of the ,target is to optimize for resource-constrained edge devices: ; in, is the zero-sample perplexity, which is used to quantify the generative ability. Lower values indicate more accurate model predictions. and represent the latency and energy consumption constraints specific to edge devices, respectively. and Represent the delay and energy consumption corresponding to a given ratio, respectively. It is an indicator function. If the condition is met, it returns 1, otherwise it returns 0. and The configuration will be developer-specified factors and Finally, we obtain the overall proportional score pair, expressed as .
[0035] Construction of continuous customization space: Existing technologies usually adopt discrete pruning spaces, which limits the diversity of configurations. In order to explore the customized optimization space more comprehensively, this paper establishes an encoder-estimator-decoder framework. The architecture includes a single-layer long short-term memory (LSTM) network as encoder and decoder, and a feedforward neural network as an evaluator to embed proportional score pairs into a continuous representation space.
[0036] Gradient-based custom optimization: Use gradient-based optimization methods in the trained representation space to identify the best pruning configuration. First, iterate and update along the gradient direction of the estimator: ; in, represents the representation of the optimal pruning configuration, is the step size used in gradient updates. (1.5) Best Custom Generation: Based on the best representation , the corresponding pruning ratio is generated by the trained decoder A beam search strategy is used to systematically explore potential configurations until an optimal pruning solution that meets performance requirements is generated. Finally, structured pruning is used to remove unnecessary model components, achieving efficient compression of large language models.
[0037] S303, dividing the resources of the edge device cluster based on the computing resources required by the compression model to obtain fixed resources and dynamic resources.
[0038] In this step, the edge device cluster is divided into resources based on the computing resources required for the compression model. When dividing, the resources are divided according to the resources that each edge device can allocate. For example, if the computing resources that edge device A can allocate are M, they are divided into two parts according to the proportion, namely fixed resources and dynamic resources. The fixed resources are used to load fixed independent model units, and the dynamic resources of each edge device are used to jointly load one or more independent model units.
[0039] like Figure 4 As shown, as a preferred embodiment of the present invention, the steps of splitting and loading the compression model based on fixed resources, processing the data processing task, and dynamically and locally repeatedly loading the compression model based on dynamic resources to complete the processing of subsequent tasks specifically include: S401, split the compression model into multiple independent model units, load the independent model units on corresponding edge devices, and the number of edge devices is the same as the number of independent model units.
[0040] In this step, the compression model is split into multiple independent model units. When the number of independent model units is the same as the number of edge devices, the independent model units are loaded in each edge device. When the number of edge devices is greater than the number of independent model units, the same independent model unit is loaded through multiple edge devices.
[0041] S402, obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge device, and obtain task processing monitoring results.
[0042] In this step, the data processing task is obtained, and the data processing task is imported into each independent model unit. The data processing task is executed according to the data processing flow of each independent model unit. For example, the compression model contains 4 independent model units, namely a, b, c and d, and the execution flows are a, b, c and d respectively. Then the data processing task is imported into a, and after processing by a, the output is imported into b, until the processing result is finally obtained through d. In this process, the time for the task to pass through each independent model unit is counted to obtain the task processing monitoring result.
[0043] S403, based on the task processing monitoring results, the data processing speed of each part of the compression model is evaluated, and the corresponding independent model unit is loaded according to the evaluation results to assist in the processing of subsequent tasks.
[0044] In this step, the data processing speed of each part of the compression model is evaluated based on the task processing monitoring results, each task is numbered, and the processing speed of each task through each independent model unit is determined. The processing speed S is represented by the ratio of the number of tokens H corresponding to each task to the time T, that is, S=H / T, and the module processing speed coordinate (n, S) is constructed, where n is the number of the task. If 5 tasks are executed, each independent model unit will correspond to five module processing speed coordinates, and a processing speed prediction function is constructed accordingly. Specifically, the five sets of module processing speed coordinates are imported into the function fitting software, such as MATLAB, and the speed prediction function is obtained by fitting. According to the speed prediction function, the processing speed of each independent model unit when executing the next task is determined, and the independent model unit with the slowest processing speed (assuming it is a) is selected. One of the independent model units a is loaded separately through idle dynamic resources, that is, there are two independent model units a at this time. When the computing power resources of the dynamic resources are sufficient, multiple independent model units with slow processing speed can be selected for loading to ensure that the global processing speed of each task is the fastest, and avoid affecting the global processing speed due to the slow processing speed of a certain independent model unit.
[0045] Large language models are usually stacked with multiple decoder layers of the same structure, but the actual contributions of these layers are heterogeneous: the first few layers are crucial in extracting input prompt features, while the latter layers have a more significant effect on generating outputs. Due to the heterogeneity of input and output data streams, the generation effect also varies between different layers. Finally, existing compression technologies often do not fully consider hardware characteristics, resulting in the inability of compressed models to fully utilize hardware resources.
[0046] Starting from the hardware resources of edge devices and the characteristics of large models, this paper designs a special compression technology adapted to the target device. While ensuring the performance of the model, it significantly reduces the number of model parameters and the amount of calculation, making it more suitable for edge devices with limited resources. First, the compression scheme is comprehensively evaluated through a heuristic algorithm to ensure that the service level agreement (SLO) is not violated during model reasoning and the reasoning efficiency of the device is maximized. Subsequently, through ablation experiments, the actual contribution of each layer to the generation of tokens and its impact on the reasoning speed are analyzed on multiple data sets. Based on these analyses, the compression rate of each layer is determined to maximize the accuracy of the model within the established compression ratio range and improve the reasoning efficiency. Finally, through a collaborative optimization method, the hardware characteristics of the edge device are combined with the model compression strategy to establish a mapping relationship between the hardware characteristics and the model hierarchical characteristics. For different hardware configurations, the compression strategy is dynamically adjusted to ensure that the compressed model achieves the best performance on the edge device. In this process, the calculation graph and memory access mode of the model are optimized by comprehensively considering factors such as the cache size, memory bandwidth and computing unit architecture of the device, giving full play to the hardware advantages, and finally achieving efficient edge reasoning.
[0047] like Figure 5 As shown, a large language model optimization system for edge devices provided by an embodiment of the present invention includes: The device network construction module 100 is used to construct an edge device network, and the edge device network is centrally managed by a cloud server.
[0048] In this system, the device network construction module 100 constructs an edge device network and sets up a cloud server to provide management services for all edge devices through the cloud server. When the edge device needs to use a large language model, if it is in the edge device network, that is, all edge devices in the edge device network have the need to use a large language model, the edge devices can communicate with each other through the cloud server to facilitate collaborative work.
[0049] The device division module 200 is used to divide the edge devices to obtain multiple edge device clusters and extract device parameter information of each edge device.
[0050] In this system, the device division module 200 divides the edge devices. For edge devices, their computing power is limited, and the memory of the device itself is also limited. They cannot directly run the large language model. Although the large language model that performs pruning operations can run, the generation capacity and accuracy will be greatly reduced, which is difficult to meet the needs. In the present invention, when the edge device needs to use a large language model, it is added to the edge device network. At this time, multiple edge devices are divided into an edge device cluster according to the situation of the edge device itself. An edge device cluster is used to jointly run a large language model. After the division of the edge device cluster is completed, the device parameter information of all edge devices in the edge device cluster is extracted. This is because the complete large language model is large in size and relatively cumbersome. Through pruning processing, a customized large language model can be generated for each edge device cluster without affecting the accuracy, so as to improve the speed of task processing.
[0051] The model loading module 300 is used to prune the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and divide the resources of the edge device cluster to obtain fixed resources and dynamic resources.
[0052] In this system, the model loading module 300 prunes the large language model to be loaded according to the device parameter information of the edge device. Traditional model compression methods often aim to reduce the number of parameters and the amount of calculation, but may ignore the impact on the generation performance. In the present invention, model compression is re-considered as a generation task. The goal is to retain the generation capability of the large language model to the greatest extent while meeting the resource limitations of the edge device. Considering the memory limitations of the edge device and the size of the model to be deployed, the overall reduction ratio of the large language model is first determined, and an exploration-utilization strategy is used to determine the pruning ratio layer by layer, and the score under each pruning ratio is determined, that is, the ratio score pair, to construct an encoder-evaluator-decoder framework, which includes a single-layer long short-term memory (LSTM) network as an encoder and decoder, and a feedforward neural network as a decoder. As an evaluator, the proportional score pairs are embedded into the continuous representation space to determine the best pruning scheme to obtain the compression model. The above compression model is customized according to the edge device cluster and can still have a high generation capacity while meeting the computing power limit. Then, in order to facilitate the splitting and loading of the compression model, the edge device cluster is divided into resources, and the resources of all edge devices are quantified, including memory and computing power. The resources are divided into fixed resources and dynamic resources. The fixed resources come from each edge device and are used to run the compression model together, while the dynamic resources are used to dynamically load part of the corresponding compression model according to the operation of the compression model in each edge device, adjust the overall task processing progress, ensure the consistency of the compression model operation, and avoid the overall task processing speed being reduced due to local resource shortages.
[0053] The dynamic compensation module 400 is used to split and load the compression model based on fixed resources, process the data processing task, and dynamically and locally repeatedly load the compression model based on dynamic resources to complete the processing of subsequent tasks.
[0054] In this system, the dynamic compensation module 400 splits and loads the compression model based on fixed resources, and splits the compression model into multiple parts according to the number of edge devices, namely independent model units. Each edge device is used to run an independent model unit. When the computing power of the edge devices is weak or the number is large, an independent model unit can be jointly run by multiple edge devices. At this time, the fixed resources of the edge devices are allocated, and all tasks initiated by the edge devices are processed by the jointly constructed compression model. During the task processing process, the processing time of each task in each independent model unit is counted, and the independent model unit that limits the overall data processing speed of the compression model is determined according to the statistical results, so as to jointly load the independent model unit by using dynamic resources, and improve the processing speed of this part of the task through one or more identical independent model units, so as to improve the overall character processing speed.
[0055] like Figure 6 As shown, as a preferred embodiment of the present invention, the device division module 200 includes: The computing power quantification unit 201 is used to count the hardware data of each edge device, quantify the computing power resources of the edge device, and obtain a quantification result.
[0056] In this module, the computing power quantization unit 201 counts the hardware data of each edge device. The computing power and memory bandwidth of edge devices are different. It is unreasonable to evaluate the compression rate of the model based solely on available memory. The present invention introduces two evaluation indicators - the first token return time (TTFT) and the subsequent token generation time (TPOP) to evaluate the computing performance of the target device.
[0057] The cluster division unit 202 is used to cluster the edge devices according to the minimum computing power resources required to run the large language model to obtain multiple edge device clusters, and the sum of the computing power resources of all edge devices in the edge device cluster is greater than the minimum computing power resources required to run the large language model.
[0058] In this module, the cluster division unit 202 clusters the edge devices according to the minimum computing resources required to run the large language model. The sum of the computing resources of the selected edge devices should be greater than the minimum computing resources required by the large language model. This ensures that the large language model can run completely. When dividing, the computing resources contained in the edge device cluster can also be multiple times the minimum computing resources required by the large language model, so that in the subsequent process, multiple large language models can be run simultaneously through the edge device cluster.
[0059] The device information extraction unit 203 is used to extract device parameter information of all edge devices in each edge device cluster, where the device parameter information at least includes computing capability and memory capacity.
[0060] In this module, the device information extraction unit 203 extracts the device parameter information of all edge devices in each edge device cluster. In order to facilitate the pruning of the large language model and build a customized compression model for the edge device cluster, it is necessary to collect the device information of the edge cluster for analysis and generation of pruning strategies.
[0061] like Figure 7 As shown, as a preferred embodiment of the present invention, the model loading module 300 includes: The pruning ratio determination unit 301 is used to determine the pruning ratio of each layer, describe the overall performance under different pruning ratios through a metric function, obtain a ratio score pair, and embed the ratio score pair into a continuous representation space.
[0062] The model pruning unit 302 is used to iteratively update along the gradient direction of the evaluator starting from the initialization point in the continuous representation space, determine the optimal pruning ratio, prune the large language model, and obtain a compressed model.
[0063] In this module, the pruning ratio of each layer is determined. Considering the memory limitation of the edge device and the size of the model to be deployed, the overall reduction ratio of the large language model is first determined. In order to perform fine-grained optimization, the exploration-exploitation strategy is adopted to determine the pruning ratio of each layer. Expressed as , using a heuristic-based pruning method to produce high-quality proportions, and using random pruning proportions as part of the exploration process. Compared with the prior art work that mainly focuses on reducing model size and computational overhead, this invention integrates a wider range of criteria, including generation capacity, inference latency, and energy consumption, and defines an overall metric function To describe a given The overall performance of the ,target is to optimize for resource-constrained edge devices: ; in, is the zero-sample perplexity, which is used to quantify the generative ability. Lower values indicate more accurate model predictions. and represent the latency and energy consumption constraints specific to edge devices, respectively. and Represent the delay and energy consumption corresponding to a given ratio, respectively. It is an indicator function. If the condition is met, it returns 1, otherwise it returns 0. and The configuration will be developer-specified factors and Finally, we obtain the overall proportional score pair, expressed as .
[0064] Construction of continuous customization space: Existing technologies usually adopt discrete pruning spaces, which limits the diversity of configurations. In order to explore the customized optimization space more comprehensively, this paper establishes an encoder-estimator-decoder framework. The architecture includes a single-layer long short-term memory (LSTM) network as encoder and decoder, and a feedforward neural network as an evaluator to embed proportional score pairs into a continuous representation space.
[0065] Gradient-based custom optimization: Use gradient-based optimization methods in the trained representation space to identify the best pruning configuration. First, iterate and update along the gradient direction of the estimator: ; in, represents the representation of the optimal pruning configuration, is the step size used in gradient updates. (1.5) Best Custom Generation: Based on the best representation , the corresponding pruning ratio is generated by the trained decoder A beam search strategy is used to systematically explore potential configurations until an optimal pruning solution that meets performance requirements is generated. Finally, structured pruning is used to remove unnecessary model components, achieving efficient compression of large language models.
[0066] The resource partitioning unit 303 is used to partition the resources of the edge device cluster based on the computing resources required by the compression model to obtain fixed resources and dynamic resources.
[0067] In this module, the resource division unit 303 divides the resources of the edge device cluster based on the computing resources required for the compression model. When dividing, the resources are divided according to the resources that each edge device can allocate. For example, if the computing resources that edge device A can allocate are M, they are divided into two parts in proportion, namely fixed resources and dynamic resources. The fixed resources are used to load fixed independent model units, and the dynamic resources of each edge device are used to jointly load one or more independent model units.
[0068] like Figure 8 As shown, as a preferred embodiment of the present invention, the dynamic compensation module 400 includes: The model loading unit 401 is used to split the compression model into multiple independent model units, and load the independent model units on corresponding edge devices, where the number of edge devices is the same as the number of independent model units.
[0069] In this module, the model loading unit 401 splits the compression model into multiple independent model units. When the number of independent model units is the same as the number of edge devices, the independent model units are loaded in each edge device. When the number of edge devices is greater than the number of independent model units, the same independent model unit is loaded through multiple edge devices.
[0070] The task monitoring unit 402 is used to obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge device, and obtain task processing monitoring results.
[0071] In this module, the task monitoring unit 402 obtains the data processing task, imports the data processing task into each independent model unit, and executes the data processing task according to the data processing flow of each independent model unit. For example, the compression model contains 4 independent model units, namely a, b, c and d, and the execution flows are a, b, c and d respectively. Then the data processing task is imported into a, and after being processed by a, the output is imported into b, until the processing result is finally obtained through d. In this process, the time for the task to pass through each independent model unit is counted to obtain the task processing monitoring result.
[0072] The dynamic auxiliary unit 403 is used to evaluate the data processing speed of each part of the compression model based on the task processing monitoring results, load the corresponding independent model unit according to the evaluation results, and perform auxiliary processing on subsequent tasks.
[0073] In this module, the dynamic auxiliary unit 403 evaluates the data processing speed of each part of the compression model based on the task processing monitoring results, numbers each task, and determines the processing speed of each task through each independent model unit. The processing speed S is represented by the ratio of the number of tokens H corresponding to each task to the time T, that is, S=H / T, and the module processing speed coordinate (n, S) is constructed, where n is the number of the task. If 5 tasks are executed, each independent model unit will correspond to five module processing speed coordinates, and the processing speed prediction function is constructed accordingly. Specifically, the five groups of module processing speed coordinates are derived It is imported into function fitting software, such as MATLAB, and fitted to obtain a speed prediction function. The processing speed of each independent model unit when executing the next task is determined according to the speed prediction function. The independent model unit with the slowest processing speed (assuming it is a) is selected, and one of the independent model units a is loaded separately through idle dynamic resources. That is, there are two independent model units a at this time. When the computing power of the dynamic resources is sufficient, multiple independent model units with slow processing speeds can be selected for loading to ensure that the global processing speed of each task is the fastest, avoiding the impact of the slow processing speed of a certain independent model unit on the global processing speed.
[0074] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A large language model optimization method for edge devices, characterized in that: The method comprises: Building an edge device network, which is centrally managed by a cloud server; Divide the edge devices to obtain multiple edge device clusters, and extract device parameter information of each edge device; Prune the large language model that needs to be loaded according to the device parameter information of the edge device to obtain a compressed model, and divide the resources of the edge device cluster to obtain fixed resources and dynamic resources; Split and load the compression model based on fixed resources to process data processing tasks, and dynamically and locally repeatedly load the compression model based on dynamic resources to complete the processing of subsequent tasks; The edge devices are divided into multiple edge device clusters. In the step of extracting device parameter information of each edge device, the computing performance of the target device is evaluated by using the token return time and each subsequent token generation time as an evaluation indicator. The edge devices are clustered according to the minimum computing power resources required to run the large language model. The sum of the computing power resources of the selected edge devices is greater than the minimum computing power resources required by the large language model. The device parameter information includes at least computing power and memory capacity. Divide the edge devices to obtain multiple edge device clusters, and determine different pruning ratios in the step of extracting device parameter information of each edge device. , by defining the metric function To describe a given The overall performance of the metric function The goal is to optimize for resource-constrained edge devices, and the overall performance is expressed as: ; in, is the zero-shot perplexity, and represent the latency and energy consumption constraints specific to edge devices, respectively. and Represent the delay and energy consumption corresponding to a given ratio, respectively. It is an indicator function. If the condition is met, it returns 1, otherwise it returns 0. and The configuration will specify the factor and Penalization, get the proportional score pair, expressed as ; By constructing an encoder-estimator-decoder framework, the proportional score pairs are embedded into a continuous representation space. In the representation space, a gradient-based optimization method is used to identify the optimal pruning configuration, and structured pruning is performed on the large language model based on the obtained optimal pruning configuration.
2. The large language model optimization method for edge devices according to claim 1, characterized in that: The step of dividing the edge devices to obtain multiple edge device clusters and extracting device parameter information of each edge device specifically includes: Collect statistics on the hardware data of each edge device, quantify the computing resources of the edge device, and obtain quantitative results; Clustering edge devices according to minimum computing resources required to run the large language model to obtain multiple edge device clusters, wherein the sum of computing resources of all edge devices in the edge device cluster is greater than the minimum computing resources required to run the large language model; The device parameter information of all edge devices in each edge device cluster is extracted, where the device parameter information includes at least computing power and memory capacity.
3. The large language model optimization method for edge devices according to claim 1, characterized in that: The steps of pruning the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, dividing the resources of the edge device cluster to obtain fixed resources and dynamic resources specifically include: Determine the pruning ratio of each layer, describe the overall performance under different pruning ratios through the metric function, obtain the ratio score pair, and embed the ratio score pair into the continuous representation space; In the continuous representation space, iterative updates are performed along the gradient direction of the estimator starting from the initialization point to determine the optimal pruning ratio, and the large language model is pruned to obtain a compressed model. Based on the computing power resources required by the compression model, the edge device cluster is divided into fixed resources and dynamic resources.
4. The large language model optimization method for edge devices according to claim 1, characterized in that: The steps of splitting and loading the compression model based on fixed resources, processing the data processing task, and dynamically and locally repeatedly loading the compression model based on dynamic resources to complete the processing of subsequent tasks specifically include: Splitting the compressed model into multiple independent model units, loading the independent model units on corresponding edge devices, where the number of edge devices is the same as the number of independent model units; Obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge device, and obtain the task processing monitoring results; Based on the task processing monitoring results, the data processing speed of each part of the compression model is evaluated, and the corresponding independent model unit is loaded according to the evaluation results to assist in the processing of subsequent tasks.
5. A large language model optimization system for edge devices, characterized in that: The system comprises: A device network construction module is used to construct an edge device network, and the edge device network is centrally managed by a cloud server; The device partitioning module is used to partition edge devices to obtain multiple edge device clusters and extract device parameter information of each edge device; The model loading module is used to prune the large language model to be loaded according to the device parameter information of the edge device to obtain a compressed model, and divide the resources of the edge device cluster to obtain fixed resources and dynamic resources; The dynamic compensation module is used to split and load the compression model based on fixed resources, process the data processing tasks, and dynamically and locally repeatedly load the compression model based on dynamic resources to complete the processing of subsequent tasks; The device classification module evaluates the computing performance of the target device by using the token return time and the subsequent token generation time as the evaluation index, and clusters the edge devices according to the minimum computing power resources required to run the large language model. The sum of the computing power resources of the selected edge devices is greater than the minimum computing power resources required by the large language model, and the device parameter information includes at least computing power and memory capacity; The model loading module determines different pruning ratios , by defining the metric function To describe a given The overall performance of the metric function The goal is to optimize for resource-constrained edge devices, and the overall performance is expressed as: ; in, is the zero-shot perplexity, and represent the latency and energy consumption constraints specific to edge devices, respectively. and Represent the delay and energy consumption corresponding to a given ratio, respectively. It is an indicator function. If the condition is met, it returns 1, otherwise it returns 0. and The configuration will specify the factor and Penalization, get the proportional score pair, expressed as ; By constructing an encoder-estimator-decoder framework, the proportional score pairs are embedded into a continuous representation space. In the representation space, a gradient-based optimization method is used to identify the optimal pruning configuration, and structured pruning is performed on the large language model based on the obtained optimal pruning configuration.
6. The large language model optimization system for edge devices according to claim 5, characterized in that: The equipment division module comprises: The computing power quantification unit is used to count the hardware data of each edge device, quantify the computing power resources of the edge device, and obtain the quantified results; A cluster division unit, used to cluster edge devices according to the minimum computing power resources required to run the large language model to obtain multiple edge device clusters, wherein the sum of computing power resources of all edge devices in the edge device cluster is greater than the minimum computing power resources required to run the large language model; The device information extraction unit is used to extract device parameter information of all edge devices in each edge device cluster, and the device parameter information at least includes computing power and memory capacity.
7. The large language model optimization system for edge devices according to claim 5, characterized in that: The model loading module includes: A pruning ratio determination unit is used to determine the pruning ratio of each layer, describe the overall performance under different pruning ratios through a metric function, obtain a ratio score pair, and embed the ratio score pair into a continuous representation space; The model pruning unit is used to iteratively update the estimator in the continuous representation space from the initialization point along the gradient direction of the estimator, determine the optimal pruning ratio, prune the large language model, and obtain a compressed model; The resource partitioning unit is used to partition the resources of the edge device cluster based on the computing resources required by the compression model to obtain fixed resources and dynamic resources.
8. The large language model optimization system for edge devices according to claim 5, characterized in that: The dynamic compensation module comprises: A model loading unit is used to split the compressed model into multiple independent model units, and load the independent model units on corresponding edge devices, where the number of edge devices is the same as the number of independent model units; The task monitoring unit is used to obtain data processing tasks, import the data processing tasks into each independent model unit, monitor the data processing results of the edge device, and obtain the task processing monitoring results; The dynamic auxiliary unit is used to evaluate the data processing speed of each part of the compression model based on the task processing monitoring results, load the corresponding independent model unit according to the evaluation results, and perform auxiliary processing on subsequent tasks.
Citation Information
Patent Citations
System and method for distributed learning of wireless edge dynamics
CN114930347A
Model structured pruning method and device based on edge computing architecture
CN118070868A
Edge computing scheduling system based on knowledge graph
CN119065852A
Image reader, image forming apparatus, and file management method
US20070177225A1
Similarity-based quantization selection for federated learning with heterogeneous edge devices
US20240111607A1