End-side operating system-oriented large model adaptive quantitative deployment method and device
By calculating the layer importance of each layer of the large model and adjusting the quantization strategy based on the memory adaptive adjustment of the terminal-side device, the large model is quantified by per-channel, which solves the problem of deploying the large model with limited resources on the terminal-side device and realizes efficient and accurate model deployment.
Patent Information
- Application Number
- CN202411897344.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-27
AI Technical Summary
On resource-constrained end-side devices, how to efficiently deploy large models without losing the inference speed and accuracy of the model, especially due to the huge number of parameters of large models, traditional GPU memory is difficult to meet the needs.
By calculating the layer importance indicators of each layer of the large model, using Jaccard coefficients, and adaptively adjusting the quantization accuracy strategy based on the memory capacity of the end-side device, using per-channel quantization to quantify the large model, ensuring efficient deployment of the model in different hardware platforms and usage scenarios.
While maintaining the accuracy of the large model, the storage needs of the model are significantly reduced, and the efficient deployment of the model on different hardware platforms and usage scenarios is achieved, which improves the inference speed and efficiency of the model on the end-side device.
Smart Images

Figure CN120046753A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and particularly to a large model adaptive quantization deployment method and device for end-side operating systems. Background Art
[0002] With the wide application of large models (such as Gemini, Llama, etc.) in natural language processing, code generation, and other fields, it has brought infinite innovation space to the application ecosystem of operating systems. However, how to efficiently deploy these large models on resource-constrained end-side devices has become an important technical challenge. Large models usually have billions of parameters and require a large amount of computing and memory resources during deployment. Traditional GPU memory (such as the 24GB video memory of NVIDIA GeForce RTX 4090) often fails to meet the requirements. For example, the Llama-2-13B model requires approximately 25GB of memory when loaded in half-precision, far exceeding the memory limit of current consumer-grade GPUs. Therefore, how to compress the scale of these models, reduce memory occupancy, and at the same time ensure the model inference speed and accuracy is the focus of current research.
[0003] Existing compression techniques include quantization, pruning, and distillation, etc. Quantization reduces the storage requirements of the model and improves the inference speed by converting floating-point weights and activation values in the model into low-precision integers. Pruning reduces the model complexity by removing unimportant weights or neurons, but may result in accuracy loss. Distillation trains a small model to imitate the behavior of a large model. Although it can effectively reduce the model size, it requires additional training and resources.
[0004] However, existing quantization methods usually perform uniform quantization on the entire model, without considering the importance differences between different layers, nor dynamically adjusting the quantization strategy according to specific hardware resources and task requirements. In order to more efficiently deploy large models on end-side devices and integrate them with operating systems, the present application proposes a large model adaptive quantization deployment method for end-side operating systems to solve this problem. Summary of the Invention
[0005] The embodiments of the present application provide a large model adaptive quantization deployment method and device for end-side operating systems.
[0006] On the one hand, a large model adaptive quantization deployment method for end-side operating systems is provided, and the method includes:
[0007] Calculating the layer importance index of each layer in the large model according to the Jaccard coefficient, and the layer importance index is used to indicate the quantization accuracy strategy corresponding to each layer;
[0008] Adapting and adjusting the quantization accuracy strategy according to the memory capacity of the end-side device;
[0009] Perform per-channel quantization on the layers to be quantized in the large model according to the adjusted quantization precision strategy with the corresponding precision;
[0010] Deploy the quantized large model to the edge device for task processing.
[0011] Optionally, calculating the layer importance index of each layer in the large model according to the Jaccard coefficient includes:
[0012] Perform forward propagation on the input text through the model() method of the large model, obtain the hidden layer states (outputs.hidden_states) of each layer and store them in the hiddens list;
[0013] Extract the hidden state H of the last time step of the input of the i-th layer from the hiddens list i,in and the hidden state H of the last time step of the output of the i-th layer i,out ;
[0014] Perform matrix multiplication on the hidden state H i,in of each layer and the hidden state H i,out with the word embedding matrix Embedding of the model to obtain the projection of each word in the vocabulary, where the projections correspond to the indices of each word respectively;
[0015] Determine the decoded word sets of the input and output of each layer by selecting the indices of the top k projection values with the largest values, and label them as the input word set C i,in of each layer and the output word set C i,out , where i represents the i-th layer;
[0016] Calculate the Jaccard coefficient according to the input word set C i,in of each layer and the output word set C i,out ;
[0017] Convert the Jaccard coefficient to obtain the layer importance index of each layer.
[0018] Optionally, the input word set C i,in is expressed as
[0019]
[0020] The output word set C i,out is expressed as
[0021]
[0022] where, is the transpose of the word embedding matrix Embedding, and TopK represents the top K words corresponding to the maximum projection value.
[0023] Optionally, the calculation formula of the Jaccard coefficient is:
[0024]
[0025] where J i represents the Jaccard coefficient of the i-th layer of the large language model.
[0026] Optionally, the conversion of the Jaccard coefficient to obtain the layer importance index for each layer includes:
[0027] The Jaccard coefficient is converted by taking the inverse and adding one to obtain the layer importance index for each layer, denoted as I i = 1 - J i where I i is the layer importance index of the i-th layer, and the magnitude of I i is positively correlated with the importance of the i-th layer.
[0028] Optionally, the quantization precision strategy is adaptively adjusted according to the memory capacity of the edge device, including:
[0029] The resource detection module traverses the memory information of all GPUs on the edge device, where the resource detection module is used to evaluate the free memory size of each GPU on the edge device;
[0030] Set an environment variable to select a GPU with the most abundant memory as the target GPU to deploy the quantized large model;
[0031] Adjust the quantization precision strategy according to the memory size of the target GPU.
[0032] Optionally, the adjustment of the quantization precision strategy according to the memory size of the target GPU includes:
[0033] Evaluate the memory size of the current GPU of the edge device and compare it with the target GPU;
[0034] In response to the current GPU of the edge device being equal to the target GPU, there is no need to quantize the large model, and directly execute the content of deploying the quantized large model to the edge device for task processing;
[0035] In response to the current GPU of the edge device not being equal to the target GPU but being able to support the memory required by the INT8 quantization model, perform the quantization precision strategy adjustment of quantizing each layer of the large model to INT8;
[0036] When the current GPU of the edge device does not support the memory required by the INT8 quantization model but can support the memory required by the INT4 quantization model, sort each layer according to the layer importance index to obtain an importance list importance_list of layers;
[0037] Perform a quantization precision strategy adjustment to select the x layers with lower importance in the importance list importance_list for INT4 quantization, and perform a quantization precision strategy adjustment to retain INT8 quantization for the remaining (N - x) layers, where the x layers with lower importance obtain the corresponding layer numbers through slicing operations, and N is the number of layers of the large model.
[0038] Optionally, performing per-channel quantization on the layers to be quantized in the large model according to the adjusted quantization precision strategy includes:
[0039] For the layers to be quantized in the large model, obtain the weight matrix W of each layer FP16 , and calculate the maximum value max{|W FP16 |} of the absolute value of the weight matrix W FP16 of each channel, where FP16 represents 16-bit floating-point numbers, and the weight matrix W FP16 of each channel refers to different channel dimensions of the weight matrix W FP16 ;
[0040] Calculate the quantization scaling factor scale according to the quantization bit width n corresponding to each channel. The quantization scaling factor scale is used to map the maximum value max{|W FP16 |} of the floating-point value range to the integer value range. The quantization bit width n is 8 in INT8 quantization and 4 in INT4 quantization;
[0041] After calculating the quantization scaling factor scale of each channel, divide the weight matrix W FP16 of each channel by the quantization scaling factor scale of each channel and perform rounding to obtain the integer weight W INTn of each channel. The integer weight W INTn of each channel is used to generate and then be deployed to the edge device for task processing.
[0042] Optionally, the calculation formula of the quantization scaling factor scale is:
[0043] where 2 n-1 -1 is the maximum value of integers;
[0044] The integer weight WINTn The calculation formula is as follows:
[0045]
[0046] On the other hand, a large model adaptive quantization deployment device for the end-side operating system is also provided. The device includes:
[0047] A first processing module, configured to calculate the layer importance index of each layer in the large model according to the Jaccard coefficient, and the layer importance index is used to indicate the quantization precision strategy corresponding to each layer;
[0048] A second processing module, configured to adaptively adjust the quantization precision strategy according to the memory capacity of the end-side device;
[0049] A third processing module, configured to perform per-channel quantization on the layers to be quantized in the large model according to the adjusted quantization precision strategy;
[0050] A fourth processing module, configured to deploy the quantized large model to the end-side device for task processing.
[0051] On the other hand, a computer-readable storage medium is provided. The storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the large model adaptive quantization deployment method for the end-side operating system as described in the above aspect.
[0052] On the other hand, a computer program product is also provided. The computer program product stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the large model adaptive quantization deployment method for the end-side operating system as described in the above aspect.
[0053] The technical effects brought by this application are at least as follows.
[0054] The large model adaptive quantization deployment method for the end - side operating system provided by the present body first evaluates the layer importance index of each layer through the Jaccard coefficient, adaptively adjusts the quantization accuracy strategy according to the memory capacity of the end - side device, dynamically formulates a quantization scheme based on the layer importance index and device resource conditions, and finally performs per - channel quantization on the layers to be quantized in the large model according to the adjusted quantization accuracy strategy and deploys the model on the end - side device. Through this intelligent adaptive quantization and deployment method, the model storage requirements can be significantly reduced while maintaining the accuracy of the large model as much as possible, thereby achieving efficient deployment of the model on different hardware platforms and different usage scenarios of the same hardware platform. During the quantization process of the large model, the importance of the layer is considered, and a more flexible and efficient quantization strategy is adopted in combination with the device resource conditions. Compared with the traditional unified quantization method, it can be dynamically optimized according to the importance of different layers and hardware resource conditions, so as to reduce the model storage requirements and accelerate inference without significantly losing accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 FIG. shows a flowchart of a large model adaptive quantization deployment method for an end - side operating system provided by an exemplary embodiment of the present application;
[0056] Figure 2 shows the corresponding Figure 1 processing framework schematic diagram;
[0057] Figure 3 FIG. is a schematic flowchart of the specific process of constructing the input - output word sets of each layer of the large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the accompanying drawings.
[0059] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0060] Embodiment 1
[0061] Please refer to Figure 1 , which shows a flowchart of a large model adaptive quantization deployment method for an end - side operating system provided by an exemplary embodiment of the present application, Figure 2 shows the corresponding Figure 1 processing framework schematic diagram. The method includes:
[0062] Step 1, calculate the layer importance index of each layer in the large model according to the Jaccard coefficient, and the layer importance index is used to indicate the quantization precision strategy corresponding to each layer.
[0063] In a possible implementation, Step 1 includes Steps S11 to S16.
[0064] Step S11, perform forward propagation on the input text (prompt) through the model() method of the large model to obtain outputs, obtain the hidden layer states (outputs.hidden_states) of each layer and store them in the hiddens list.
[0065] hiddens is a list containing multiple tensors, and each tensor represents a hidden layer state. hiddenstates is an important output of the model, which contains the hidden state information of each layer in the model.
[0066] Step S12, extract the hidden state H of the last time step of the input of the i-th layer from the hiddens list i,in and the hidden state H of the last time step of the output of the i-th layer i,out .
[0067] These hidden states reflect the encoding information of the model for the last word in the input sequence at this layer.
[0068] Step S13, perform matrix multiplication on the hidden state H of each layer i,in and the hidden state H i,out with the word embedding matrix Embedding of the model to obtain the projection of each word in the vocabulary, where the projections correspond to the indices of each word respectively.
[0069] Multiply these hidden states with the transpose of the word embedding matrix Embedding responsible for mapping discrete vocabulary to a continuous vector space through the matrix multiplication operator @, where the word embedding matrix Embedding can be obtained by model.get_input_embeddings().weight.detach(). The multiplication operation produces the projection of each word on the vocabulary, and these projections correspond to the indices of each word respectively.
[0070] Step S14, determine the decoded word sets of the input and output of each layer by selecting the indices of the top k projection values with the largest values, and mark them as the input word set C i,in and the output word set C i,out of each layer, where i represents the i-th layer.
[0071] Finally, decode these indexes into actual words to construct the word sets of the input and output of the i-th layer. The construction process of the word sets is as Figure 3 shown.
[0072] The input word set C iin is represented as
[0073]
[0074] The output word set C i,out is represented as
[0075]
[0076] where is the transpose of the word embedding matrix Embedding, and TopK represents the top K words corresponding to the maximum projection value.
[0077] Step S15, calculate the Jaccard coefficient according to the input word set C i,in and the output word set C i,out of each layer;
[0078] The calculation formula of the Jaccard coefficient is:
[0079]
[0080] where J i represents the Jaccard coefficient of the i-th layer of the large language type.
[0081] After obtaining the two word sets C i,in and C i,out it is necessary to calculate the Jaccard coefficient of the two word sets to quantify the semantic change degree of each layer in the large model. The Jaccard coefficient is widely used in multiple fields, including machine learning and text mining, to evaluate the similarity between sample sets or feature sets.
[0082] In addition, it should be noted that only the Jaccard similarity is used as the evaluation index in the present invention. Existing technologies also use other indexes such as cos similarity as evaluation indexes, including calculating the cosine similarity between the input and output of each layer in LLMs to determine the importance, and analyzing the distribution of weights within the layer (the number of weights far greater than the average value) to determine the importance, but the actual measured effect is not as good as the Jaccard similarity.
[0083] Specifically, the present invention selects the Jaccard similarity as the indicator for evaluating the importance of layers in a large language model (LLM), mainly based on its effectiveness in capturing the similarity between input and output word sets and its semantic transformation ability. The Jaccard similarity can directly measure the similarity between two word sets and is suitable for evaluating the degree of "semantic change" of each layer to the input information. Specifically, by measuring the ratio of the intersection to the union of the word sets in the input and output layers, the Jaccard similarity can effectively reflect whether significant semantic transformation occurs when the layer processes the input information, thus providing strong support for ranking the importance of layers.
[0084] In contrast, although the commonly used cosine similarity in the prior art can also measure the similarity in the vector space, it mainly depends on calculating the angle between vectors and cannot directly reflect the differences and semantic changes of the vocabulary sets. Therefore, the effect of cosine similarity in layer importance evaluation is relatively weak, especially when dealing with large models, it may ignore the details of semantic changes.
[0085] In addition, some techniques also use the analysis of the weight distribution within the layer (for example, the number of weights greater than the average) as the basis for importance evaluation. Although this method can provide certain information in some scenarios, it mainly focuses on the magnitude and distribution characteristics of the weight values and may not fully capture the dynamic characteristics of the model in semantic processing and information transformation, resulting in the evaluation results may not be as accurate as the Jaccard similarity.
[0086] Based on the actual test results, the Jaccard similarity adopted by the present invention as the evaluation indicator can better reflect the semantic transformation ability of the layer to the input information compared with other methods (such as cosine similarity and weight distribution analysis). Therefore, it has higher reliability and practicality in the quantization scheme of the present invention.
[0087] Step S16, convert the Jaccard coefficient to obtain the layer importance indicator for each layer.
[0088] The Jaccard coefficient is converted by taking the inverse and adding one to obtain the layer importance indicator for each layer, denoted as I i = 1 - J i , where I i is the layer importance indicator for the i-th layer.
[0089] I i is positively correlated with the importance of the i-th layer. When I i decreases, J iAn increase indicates an enhanced similarity between the two word sets, suggesting that the layer may not have achieved significant semantic transformation when processing the input information. Therefore, it can be inferred that the contribution of this layer in terms of semantic transformation ability is relatively small, and thus its importance is considered low.
[0090] Step 2: Adaptively adjust the quantization precision strategy according to the memory capacity of the edge device.
[0091] In a possible implementation, Step 2 includes Steps S21 to S23.
[0092] Step S21: The resource detection module traverses the memory information of all GPUs on the edge device. Among them, the resource detection module is used to evaluate the free memory size of each GPU on the edge device.
[0093] In this process, it is first necessary to evaluate the hardware resource situation of the edge device, especially the memory situation of the GPU. To achieve this goal, a resource detection module is introduced. This module can traverse and evaluate the memory status of each GPU to select the GPU with the most free memory for model deployment.
[0094] Among them, the resource detection module uses the Wikipedia dataset as input to measure the layer importance of the large model. The importance coefficient I of the i-th layer obtained for each prompt i is accumulated, and finally divided by the total length of the dataset to obtain the final importance coefficient I i and the quantization strategy of the large model will be formulated. During the calculation process, the hyperparameter K of TopK can be set as a trade-off value of 20.
[0095] Specifically, an overly small K value (such as K = 5) may result in the inability to fully capture the semantic changes between the input and output layers because a small K value only considers the most prominent few words in the input and output, which may ignore some important information. An overly large K value (such as K = 100) will increase the computational overhead because more words need to be processed when evaluating the semantic transformation of each layer. A larger K may introduce unnecessary noise, leading to accuracy loss and increased computational complexity. K = 20 is a reasonable trade-off value. It can better capture the main semantic changes of each layer while maintaining computational efficiency without introducing too much noise. Those skilled in the art can make custom selections according to requirements.
[0096] Step S22: Set an environment variable to select a GPU with the most abundant memory as the target GPU to deploy the quantized large model.
[0097] When quantifying large models, it is crucial to select a GPU with sufficient memory to ensure the efficient operation of the quantized model. By evaluating the memory capacity of each GPU, we can choose the GPU with the most available memory as the target device. For example, when there are multiple GPUs on the device, we can specify the GPU number by setting the environment variable "CUDA_VISIBLE_DEVICES", and then deploy the model on the specified GPU subsequently.
[0098] Environment variables are used in computer systems to store configuration information. In the present invention, by setting specific environment variables, the system can be helped to identify and select the GPU with the most abundant memory. This method can not only achieve automated device selection, but also simplify the model deployment process, ensure that the quantized large model is deployed on a GPU with sufficient memory, and thus improve the system operation efficiency. After the resource detection module determines the memory status of each GPU, the system automatically sets the environment variable according to the evaluation result, and selects the GPU with the most abundant memory as the target device for deployment. Through the setting of environment variables, the system can seamlessly achieve the automatic deployment of the quantized model and ensure its operation on the device with optimal resources.
[0099] Step S23, adjust the quantization precision strategy according to the memory size of the target GPU.
[0100] Evaluate the memory size of the current GPU on the edge device and compare it with the target GPU.
[0101] In response to the current GPU on the edge device being equal to the target GPU, there is no need to quantize the large model, and directly execute the content of deploying the quantized large model to the edge device for task processing.
[0102] In response to the current GPU on the edge device not being equal to the target GPU but being able to support the memory required by the INT8 quantization model, adjust the quantization precision strategy of each layer of the large model to INT8.
[0103] In response to the current GPU on the edge device not supporting the memory required by the INT8 quantization model but being able to support the memory required by the INT4 quantization model, sort each layer according to the layer importance index to obtain an importance list importance_list of the layers.
[0104] Perform quantization precision strategy adjustment for the x layers with lower importance in the importance list importance_list for INT4 quantization, and perform quantization precision strategy adjustment for the remaining (N - x) layers to retain INT8 quantization, where the x layers with lower importance obtain the corresponding layer numbers through slicing operations, that is, the corresponding layer numbers to be quantized with INT4 are obtained through the slicing operation. N is the number of layers of the large model, and x can also be understood as having x layers with lower importance, and N - x can also be understood as having (N - x) remaining layers. Among them, the slicing operation is the slicing operation of the importance list importance_list, and the corresponding operator for the slicing operation is importance_list[x:].
[0105] This application suggests that x should be as small as possible without exceeding the target GPU memory limit, so as to ensure that the maximum number of layers are quantized with INT8 to maintain the model accuracy.
[0106] Step 3, perform per-channel quantization on the layers to be quantized in the large model according to the adjusted quantization precision strategy with the corresponding precision.
[0107] In a possible implementation manner, Step 3 includes Steps S31 to S23.
[0108] Step S31, for the layers to be quantized in the large model, obtain the weight matrix W of each layer FP16 , and calculate the maximum value max{|W FP16 |} of the absolute value of the weight matrix W of each channel, where FP16 represents 16-bit floating-point numbers, and the weight matrix W of each channel FP16 refers to different channel dimensions of the weight matrix W FP16 . FP16 For example, in the convolutional layer, the channel usually refers to the feature map channel in each layer, and in the fully connected layer, the channel usually refers to each column of the matrix. For a large model, the weight matrix of each layer has multiple channels.
[0109] This application quantizes each channel independently, so as to adjust the quantization precision more precisely. Per-channel quantization assigns independent quantization parameters to each channel of the weight matrix, which enables the quantization process to more precisely capture the dynamic range of the data in each channel. Therefore, per-channel quantization can reduce the error in the quantization process, thereby improving the accuracy of the model.
[0110] This application quantizes each channel independently, so as to adjust the quantization precision more precisely. Per-channel quantization assigns independent quantization parameters to each channel of the weight matrix, which enables the quantization process to more precisely capture the dynamic range of the data in each channel. Therefore, per-channel quantization can reduce the error in the quantization process, thereby improving the accuracy of the model.
[0111] Step S32, calculate the quantization scaling factor scale according to the quantization bit width n corresponding to each channel. The quantization scaling factor scale is used to scale the maximum value max{|W FP16|} is mapped to the integer value range. The quantization bit width n is 8 under INT8 quantization and 4 under INT4 quantization;
[0112] The calculation formula for the quantization scaling factor scale is:
[0113] where 2 n-1 -1 is the maximum value of an integer. The maximum value of INT8 is 127 (the corresponding value range is -128 to 127), and the maximum value of INT4 is 7.
[0114] The role of the quantization scaling factor scale is to convert the value range of the floating-point weights into the integer value range to meet the requirements of quantization.
[0115] Step S33, after calculating the quantization scaling factor scale for each channel, divide the weight matrix W of each channel FP16 by the quantization scaling factor scale of each channel and perform rounding to obtain the quantized integer weight W of each channel INTn , and the quantized integer weight W of each channel INTn is used to generate and is later deployed to the edge device for task processing.
[0116] The superscript of W represents the corresponding data type. For example, W FP16 represents that the weight matrix is stored and calculated in the 16-bit floating-point (FP16) format. Another example is W INTn , where n is the quantization bit width, indicating that the weight matrix has been quantized to an integer type, specifically INT8 or INT4. This means that the representation of the weights has changed from floating-point numbers to integers to further reduce memory occupancy and improve computational efficiency.
[0117] The calculation formula for the integer weight W INTn is:
[0118]
[0119] After the above steps, the weights of each channel are converted into integer form, thus completing the quantization process.
[0120] Step 4, deploy the quantized large model to the edge device for task processing.
[0121] After deployment, the quantized model will run on the edge device for actual inference tasks. Since the storage and computational requirements of the quantized model are optimized, the edge device can execute tasks with lower computational costs and less memory occupancy while maintaining high performance.
[0122] The deployment in this step is not just about loading the model onto the device. It also involves intelligently adjusting the quantization precision strategy according to the device's resource situation. For example, if the GPU memory of the target device is sufficient, the model may be deployed with INT8 precision, which helps maintain good inference accuracy. If the memory is limited, the system will deploy according to the adjusted quantization precision strategy (such as using INT4 for quantization of some layers), thus further saving memory and ensuring performance.
[0123] In summary, by combining the operations in Steps 1 to 3, the core task of Step 4 is to achieve the efficient deployment of the quantized model. While ensuring accuracy, quantization reduces the memory requirements of the model, making the deployment process more flexible and efficient. After deployment, the quantized model can quickly respond to user requests and provide real-time inference results, meeting the requirements of edge devices for low latency and high-performance computing. Deploy the quantized model to the target GPU and start executing the actual task. At this time, the edge device can process the input data and provide inference results based on the quantized large model. Ensure that the quantized large model can perform efficient task processing on the edge device, maximizing the utilization of model accuracy and computing resources.
[0124] The following Figure 2 is explained. The offline and online regions respectively indicate two different stages in the quantization and deployment process of the large model.
[0125] Due to the differences in layer importance among various LLMs and the different GPU resources of different edge devices or the same edge device under various conditions, the offline part needs to determine the most appropriate quantization strategy based on these situations. In the online stage, according to the selected quantization strategy, the LLM is quantized to the corresponding bit width. Then the quantized LLM is deployed to the target device.
[0126] In the offline region, the large model to be deployed and the edge resource device show the large model to be quantized and the edge device, ready for quantization deployment in subsequent steps. The two boxes of layer importance detection and device resource detection respectively handle two key tasks. Layer importance detection evaluates the layer importance of each layer of the model based on the Jaccard coefficient, and device resource detection evaluates the GPU memory situation of the edge device to ensure the effective deployment of the quantized model. The large model quantization strategy formulation indicates formulating a suitable quantization strategy according to the layer importance and device resource situation, deciding which layers use high precision (such as INT8) and which layers use low precision (such as INT4).
[0127] In the online region, the deployment of the quantized large model shows that the quantized model is deployed to edge devices for task processing. The quantized model will be adaptively deployed according to the strategy. The process of large model quantization transforms the quantization strategy in the offline region into specific quantization operations to ensure the efficient operation of the model on different hardware devices.
[0128] During the offline stage, importance evaluation and quantization strategy formulation are carried out, and during the online stage, the deployment of the quantized model is performed. Through the offline functions of "layer importance detection" and "device resource detection", the system can flexibly adjust the quantization accuracy of the model, and finally deploy the appropriate model to edge devices to optimize memory usage and model performance. Ensure that the model can maintain good accuracy and efficiency under different hardware conditions.
[0129] In summary, the large model adaptive quantization deployment method provided by the present invention for the edge operating system first evaluates the layer importance index of each layer through the Jaccard coefficient, adaptively adjusts the quantization accuracy strategy according to the memory capacity of the edge device, dynamically formulates the quantization scheme according to the layer importance index and device resource conditions, and finally performs per-channel quantization on the layers to be quantized in the large model according to the adjusted quantization accuracy strategy and deploy the model on the edge device. Through this intelligent adaptive quantization and deployment method, it is possible to significantly reduce the model storage requirements while maintaining the accuracy of the large model as much as possible, and then achieve the efficient deployment of the model on different hardware platforms and different usage scenarios of the same hardware platform. During the quantization process of the large model, the importance of the layer is considered, and combined with the device resource conditions, a more flexible and efficient quantization strategy is adopted. Compared with the traditional unified quantization method, it can be dynamically optimized according to the importance of different layers and the hardware resource conditions, so as to reduce the model storage requirements and accelerate the inference without significantly sacrificing accuracy.
[0130] Furthermore, an inherent feature of the large model is used to more finely evaluate the importance of each layer of the model, so as to effectively capture the correlation between the semantic information of each layer, reveal the features and aspects that the model focuses on at different levels, and achieve an accurate evaluation of the layer importance.
[0131] Furthermore, by combining the layer importance index of the large model and the real-time resources of the edge device, an optimal quantization strategy is formulated, and based on this, the large model is quantized for specific layers and specific bit widths, and the deployment of the quantized model can be completed while maintaining the optimal model performance as much as possible.
[0132] Furthermore, compared with the existing quantization methods of the same granularity, the quantization method of the present invention is more effective in maintaining the performance of the large model, can lose less key information during the quantization process, and has better effects on the accuracy of zero-shot downstream tasks and the perplexity index of the model.
[0133] An exemplary embodiment of the present application further provides a large model adaptive quantization deployment device for an end-side operating system, and the device includes:
[0134] A first processing module, configured to calculate the layer importance index of each layer in the large model according to the Jaccard coefficient, and the layer importance index is used to indicate the quantization precision strategy corresponding to each layer;
[0135] A second processing module, configured to adaptively adjust the quantization precision strategy according to the memory capacity of the end-side device;
[0136] A third processing module, configured to perform per-channel quantization on the layers to be quantized in the large model according to the adjusted quantization precision strategy;
[0137] A fourth processing module, configured to deploy the quantized large model to the end-side device for task processing.
[0138] Optionally, the first processing module includes:
[0139] A first processing unit, configured to perform forward propagation on the input text through the model() method of the large model, obtain the hidden layer states (outputs.hidden_states) of each layer and store them in the hiddens list;
[0140] A second processing unit, configured to extract the hidden state H of the last time step of the input of the i-th layer from the hiddens list i,in and the hidden state H of the last time step of the output of the i-th layer i,out ;
[0141] A third processing unit, configured to perform matrix multiplication on the hidden state H of each layer i,in and the hidden state H i,out with the word embedding matrix Embedding of the model to obtain the projection of each word in the vocabulary, where the projections respectively correspond to the indexes of each word;
[0142] A fourth processing unit, configured to determine the decoded word sets of the input and output of each layer by selecting the indexes of the top k projection values, and mark them as the input word set C i,in and the output word set C i,out , where i represents the i-th layer;
[0143] A fifth processing unit, configured to calculate the Jaccard coefficient according to the input word set C i,in and the output word set C i,out of each layer;
[0144] The sixth processing unit is used to convert the Jaccard coefficient to obtain the layer importance index of each layer.
[0145] Optionally, the input word set C i,in is represented as
[0146]
[0147] The output word set C iout is represented as
[0148]
[0149] where is the transpose of the word embedding matrix Embedding, and TopK represents the top K words corresponding to the maximum projection value.
[0150] Optionally, the calculation formula of the Jaccard coefficient is:
[0151]
[0152] where J i represents the Jaccard coefficient of the i-th layer of the large language model.
[0153] Optionally, the sixth processing unit is further used to convert the Jaccard coefficient by taking the inverse and adding one to obtain the layer importance index of each layer, denoted as I i = 1 - J i , where I i is the layer importance index of the i-th layer, and the magnitude of I i is positively correlated with the importance of the i-th layer.
[0154] Optionally, the second processing module includes:
[0155] The seventh processing unit is used to traverse the memory information of all GPUs on the edge device through the resource detection module, where the resource detection module is used to evaluate the free memory size of each GPU on the edge device;
[0156] The eighth processing unit is used to set environment variables to select a GPU with the most abundant memory as the target GPU to deploy the quantized large model;
[0157] The ninth processing unit is used to adjust the quantization precision strategy according to the memory size of the target GPU.
[0158] Optionally, the ninth processing unit is further configured to evaluate the memory size of the current GPU of the edge device and compare it with the target GPU; in response to the current GPU of the edge device being equivalent to the target GPU, there is no need to perform quantization of the large model, and directly execute the content of deploying the quantized large model to the edge device for task processing; in response to the current GPU of the edge device not being equivalent to the target GPU but being able to support the memory required by the INT8 quantization model, perform a quantization precision strategy adjustment for each layer of the large model to the quantization precision of INT8; in response to the current GPU of the edge device not supporting the memory required by the INT8 quantization model but being able to support the memory required by the INT4 quantization model, sort each layer according to the layer importance index to obtain an importance list importance_list of layers; perform a quantization precision strategy adjustment for the x layers with lower importance in the importance list importance_list to perform INT4 quantization, and perform a quantization precision strategy adjustment for the remaining (N - x) layers to retain INT8 quantization, where the x layers with lower importance obtain the corresponding layer numbers through slicing operations, and N is the number of layers of the large model.
[0159] Optionally, the third processing module includes:
[0160] A tenth processing unit, configured to obtain the weight matrix W of each layer for the layers that need to be quantized in the large model FP16 , and calculate the maximum value max{|W FP16 |} of the absolute value of the weight matrix W of each channel, where FP16 represents 16-bit floating-point numbers, and the weight matrix W of each channel FP16 refers to different channel dimensions of the weight matrix W FP16 ; FP16
[0161] An eleventh processing unit, configured to calculate the quantization scaling factor scale according to the quantization bit width n corresponding to each channel, and the quantization scaling factor scale is used to map the maximum value max{|W FP16 |} of the floating-point value range to the integer value range, the quantization bit width n is 8 in INT8 quantization, and the quantization bit width n is 4 in INT4 quantization;
[0162] A twelfth processing unit, configured to, after calculating the quantization scaling factor scale of each channel, divide the weight matrix W of each channel FP16 by the quantization scaling factor scale of each channel and perform rounding processing to obtain the integer weight W of each channel after quantization INTn , and the integer weight W of each channel after quantization INTn is used to be generated and then deployed to the edge device for task processing.
[0163] Optionally, the calculation formula for the quantization scaling factor scale is as follows:
[0164] where 2 n-1 -1 is the maximum value of an integer;
[0165] The calculation formula for the integer weight W INTn is as follows:
[0166]
[0167] On the other hand, a computer-readable storage medium is provided. The storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the large model adaptive quantization deployment method for the end-side operating system as described in the above aspects.
[0168] On the other hand, a computer program product is further provided. The computer program product stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the large model adaptive quantization deployment method for the end-side operating system as described in the above aspects.
[0169] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0170] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, etc.
[0171] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A large model adaptive quantization deployment method for a terminal-side operating system, characterized in that: The method comprises: Calculate the layer importance index of each layer in the large model according to the Jaccard coefficient, and the layer importance index is used to indicate the quantization accuracy strategy corresponding to each layer; Adaptively adjust the quantization accuracy strategy according to the memory capacity of the terminal device; According to the adjusted quantization precision strategy, the layers that need to be quantized in the large model are per-channel quantized according to the corresponding precision; The quantized large model is deployed to the end-side device for task processing.
2. The method according to claim 1, characterized in that The layer importance index of each layer in the large model is calculated according to the Jaccard coefficient, including: The input text is forward propagated through the model() method of the large model, and the hidden layer states of each layer (outputs.hidden_states) are obtained and stored in the hiddens list; Extract the hidden state H of the last time step of the i-th layer input from the hiddens list i,in and the hidden state H of the last time step output by the i-th layer i,out ; The hidden state H of each layer i,in and the hidden state H i,out Perform matrix multiplication with the model's word embedding matrix Embedding to obtain the projection of each word in the vocabulary, where the projection corresponds to the index of each word; By selecting the index with the largest projection value of the first k, the decoded word set of each layer input and output is determined, which is marked as the input word set C of each layer. i,in and the output word set C i,out , where i represents the i-th layer; According to the input word set C of each layer i,in and the output word set C i,out Calculate the Jaccard coefficient; The Jaccard coefficient is converted to obtain the layer importance index of each layer.
3. The method according to claim 2, characterized in that The input word set C iin It is expressed as, The output word set C i,out It is expressed as, in, It is the transpose of the word embedding matrix Embedding, and TopK represents the top K words corresponding to the maximum projection value.
4. The method according to claim 3, characterized in that The calculation formula of the Jaccard coefficient is: Among them J i Represents the Jaccard coefficient of the i-th layer of the large language type.
5. The method according to claim 4, characterized in that The Jaccard coefficient is converted to obtain the layer importance index of each layer, including: The Jaccard coefficient is converted by taking the inverse and adding one to obtain the layer importance index of each layer, which is recorded as I i =1-J i , where I i is the layer importance index of the i-th layer, I i The size of is positively correlated with the importance of the i-th layer.
6. The method according to claim 1, characterized in that The adaptively adjusting the quantization accuracy strategy according to the memory capacity of the terminal device includes: Traversing memory information of all GPUs on the end-side device through a resource detection module, wherein the resource detection module is used to evaluate the free memory size of each GPU on the end-side device; Set environment variables to select a GPU with the most memory as the target GPU to deploy the quantized large model; The quantization precision strategy is adjusted according to the memory size of the target GPU.
7. The method according to claim 6, characterized in that The adjusting the quantization precision strategy according to the memory size of the target GPU includes: Evaluate the memory size of the current GPU of the client device and compare it with the target GPU; In response to the current GPU of the end-side device being equivalent to the target GPU, there is no need to quantize the large model, and directly executing the content of deploying the quantized large model to the end-side device for task processing; In response to the current GPU of the end-side device not being equal to the target GPU but being able to support the memory required by the INT8 quantization model, adjusting the quantization accuracy strategy of each layer of the large model to INT8; In response to the current GPU of the end-side device not supporting the memory required by the INT8 quantization model but being able to support the memory required by the INT4 quantization model, sorting the layers according to the layer importance index to obtain a layer importance list importance_list; Select x layers with lower importance in the importance list importance_list to adjust the quantization precision strategy for INT4 quantization, and adjust the quantization precision strategy for the (Nx) remaining layers to retain INT8 quantization, wherein the x layers with lower importance obtain the corresponding layer number through slicing operation, and N is the number of layers of the large model.
8. The method according to claim 1, characterized in that: The step of performing per-channel quantization on the layers to be quantized in the large model according to the corresponding precision according to the adjusted quantization precision strategy includes: For the layers that need to be quantized in the large model, obtain the weight matrix W of each layer FP16 , and calculate the weight matrix W of each channel FP16 The maximum absolute value of max{|W FP16 |}, where FP16 represents a 16-bit floating point number, and the weight matrix W of each channel FP16 Refers to the weight matrix W FP16 Different channel dimensions; The quantization scaling factor scale is calculated according to the quantization bit width n corresponding to each channel, and the quantization scaling factor scale is used to convert the maximum value max{|W FP16 |} is mapped to an integer value domain, the quantization bit width n is 8 under INT8 quantization, and the quantization bit width n is 4 under INT4 quantization; After calculating the quantization scaling factor scale of each channel, the weight matrix W of each channel is FP16 Divide by the quantization scaling factor scale of each channel and round off to get the quantized integer weight W of each channel INTn , the quantized integer weights W of each channel INTn After being generated, it is deployed to the end-side device for task processing.
9. The method according to claim 8, characterized in that The calculation formula of the quantization scaling factor scale is: Among them, 2 n-1 -1 is the maximum value of an integer; The integer weight W INTn The calculation formula is:
10. A large model adaptive quantization deployment device for a terminal-side operating system, characterized in that: The device comprises: The first processing module is used to calculate the layer importance index of each layer in the large model according to the Jaccard coefficient, and the layer importance index is used to indicate the quantization accuracy strategy corresponding to each layer; A second processing module, configured to adaptively adjust the quantization accuracy strategy according to the memory capacity of the terminal device; The third processing module is used to perform per-channel quantization on the layers that need to be quantized in the large model according to the corresponding precision according to the adjusted quantization precision strategy; The fourth processing module is used to deploy the quantized large model to the terminal side device for task processing.