Layered hybrid structured compression method and device for large language model

By allocating the compression ratio according to the importance of each layer of the large language model and performing specific compression processing for different hierarchies, the problem of poor compression effect of large language models in the prior art is solved, and the compression rate and performance balance is achieved.

CN120087419APending Publication Date: 2025-06-03INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510544956.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to reasonably allocate the compression ratios of each layer of a large language model, resulting in reduced model performance or waste of resources after compression.

Method used

By determining the compression ratio of each layer according to the importance of each layer of the target model, and weighted low-rank decomposition and differentiated singular value allocation are performed for the multi-head attention sublayer, channel pruning and compensation processing are used for the feedforward network sublayer to achieve differentiated compression ratio allocation.

Benefits of technology

The compression rate and performance balance of large language models is achieved, ensuring better compression effect at the same overall compression rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087419A_ABST
    Figure CN120087419A_ABST
Patent Text Reader

Abstract

The invention provides a hierarchical hybrid structured compression method and device for a large language model, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining the compression ratio of each layer of a target model based on the importance degree of each layer of the target model; based on the compression ratio of each layer, compressing the target model to obtain a compressed model; wherein the compression of the target model comprises the following steps: for a current processing layer, calculating a two-norm of an input activation value corresponding to each weight of the current processing layer on a calibration data set to obtain a characteristic norm; under the condition that the current processing layer is the multi-head attention sub-layer, performing weighted low-rank decomposition and differential singular value distribution processing on the current processing layer based on the characteristic norm of the current processing layer and the compression ratio of each layer; and under the condition that the current processing layer is the feedforward network sub-layer, performing channel pruning and compensation processing on the current processing layer based on the characteristic norm of the current processing layer and the compression ratio of each layer. Therefore, the compression ratio and the performance balance of the LLM are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a hierarchical hybrid structured compression method and device for a large language model. Background Art

[0002] A large language model (LLM) refers to a language model with billions to hundreds of billions of parameters. These models are pre-trained using a large amount of text data, enabling them to exhibit powerful performance in natural language understanding and natural language generation fields such as machine translation, reading comprehension, and question answering systems. However, due to the huge parameter scale of the LLM, it requires a large amount of storage and computing resources for deployment and inference, which severely limits the availability and popularity of the LLM in practical applications. Studying the structured compression of large models can effectively reduce the resource requirements of large models during the application stage. Therefore, researching efficient and accurate structured compression of large models has strong practical significance.

[0003] Model compression usually uses modules with fewer parameters to approximately replace the weight matrices in the model. In the actual model compression process, the contribution degrees of different layers to the model performance are different. How to reasonably allocate the compression ratios of each layer so that it can not only achieve the expected overall compression goal but also maximize the preservation of model performance is an important research issue. Over-compressing the key layers will lead to a sharp decline in model performance, while conservative compression of some unimportant layers will cause waste of compression resources. Therefore, a reasonable allocation of the compression ratio based on the importance of the layers can obtain better compression effects under the same overall compression rate. Previous compression methods usually use the same processing method for two sub-layers with different structures (feed-forward network sub-layer FFN, multi-head self-attention sub-layer MHA) in the Transformer architecture model, thus ignoring the differences between different sub-layer structures. The weight matrix in the MHA sub-layer shows stronger low-rank properties, and its internal minimum dependence structure is a single attention head; while the weight matrix in the FFN sub-layer shows weaker low-rank properties, and its minimum dependence structure is a single channel. In addition, since previous weight importance estimation methods rely on gradients, but the calculation and storage of gradients on the LLM require a large amount of resources, it is difficult to directly use the gradient-based weight importance estimation method. Summary of the Invention

[0004] The present invention provides a hierarchical hybrid structured compression method and device for a large language model to solve the defect in the prior art that it is difficult for the LLM to reasonably allocate the compression ratios of each layer so that it can not only achieve the expected overall compression goal but also maximize the preservation of model performance, and to achieve the balance between the compression rate and performance of the LLM.

[0005] The present invention provides a hierarchical hybrid structured compression method for large language models, including the following steps: Determine the compression ratio of each layer of the target model based on the importance degree of each layer of the target model; Compress the target model based on the compression ratio of each layer to obtain a compressed model; Among them, compressing the target model based on the compression ratio of each layer includes: For the currently processed layer, calculate the second norm of the input activation value corresponding to each weight of the currently processed layer on the calibration dataset to obtain the feature norm; When the currently processed layer is a multi-head attention sublayer, based on the feature norm of the currently processed layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value assignment processing on the currently processed layer to complete the compression of the currently processed layer; When the currently processed layer is a feed-forward network sublayer, based on the feature norm of the currently processed layer and the compression ratio of each layer, perform channel pruning and compensation processing on the currently processed layer to complete the compression of the currently processed layer.

[0006] According to the hierarchical hybrid structured compression method for large language models provided by the present invention, performing weighted low-rank decomposition and differential singular value assignment processing on the currently processed layer based on the feature norm of the currently processed layer and the compression ratio of each layer includes: Based on the feature norm, weight the original weight matrix of the currently processed layer to obtain a weighted matrix; Perform singular value decomposition on the weighted matrix, determine the relationship between the approximate matrix corresponding to different numbers of retained singular values and the perplexity, and determine the strength of the low-rankness of each projection matrix of the currently processed layer; Based on the strength of the low-rankness of different projection matrices and the compression ratio of each layer, determine the number of singular values retained by the projection matrix of the currently processed layer.

[0007] According to the hierarchical hybrid structured compression method for large language models provided by the present invention, performing channel pruning and compensation processing on the currently processed layer based on the feature norm of the currently processed layer and the compression ratio of each layer includes: Based on the feature norm of the currently processed layer and the original weight matrix, determine the importance score of each weight of the currently processed layer; Aggregate the importance scores of each weight to obtain the channel importance score; Group the channels based on the dependency structure, and use the sum of the channel importance scores within the group as the group importance score of each group of channels; Based on the group importance score and the compression ratio of each layer, perform pruning processing on each group of channels; Based on the optimal brain surgery OBS optimal pruning framework, adjust the parameters of each group of channels after pruning to compensate for the pruning loss.

[0008] According to a hierarchical hybrid structured compression method for large language models provided by the present invention, the method further includes: Based on lightweight fine-tuning LoRA, perform knowledge restoration on the compressed model to obtain the final compressed model.

[0009] According to a hierarchical hybrid structured compression method for large language models provided by the present invention, the calibration dataset is obtained in the following manner: Extract a certain number of samples from the natural language processing dataset; After being processed by the tokenizer of the target model, obtain the token sequence; Intercept the obtained token sequence to a specified length to obtain the calibration dataset.

[0010] According to a hierarchical hybrid structured compression method for large language models provided by the present invention, determining the compression ratio of each layer of the target model based on the importance of each layer of the target model includes: Based on the change situation of the input activation value and the output activation value of each layer of the target model, determine the importance of each layer of the target model; the change situation includes the amplitude change situation and the direction change situation; Based on the importance of each layer of the target model, determine the compression ratio of each layer of the target model; among them, the higher the importance, the lower the compression ratio assigned to the layer, and the lower the importance, the higher the compression ratio assigned to the layer.

[0011] The present invention also provides a hierarchical hybrid structured compression device for large language models, including the following modules: A determination module for determining the compression ratio of each layer of the target model based on the importance of each layer of the target model; A compression module for compressing the target model based on the compression ratio of each layer to obtain a compressed model; Among them, compressing the target model based on the compression ratio of each layer includes: For the currently processed layer, calculate the two-norm of the input activation value corresponding to each weight of the currently processed layer on the calibration dataset to obtain the feature norm; When the currently processed layer is a multi-head attention sub-layer, based on the feature norm of the currently processed layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value assignment processing on the currently processed layer to complete the compression of the currently processed layer; When the current processing layer is a feed-forward network sub-layer, based on the feature norm of the current processing layer and the compression ratio of each layer, channel pruning and compensation processing are performed on the current processing layer to complete the compression of the current processing layer.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the hierarchical hybrid structured compression method of the large language model as described in any one of the above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the hierarchical hybrid structured compression method of the large language model as described in any one of the above is implemented.

[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the hierarchical hybrid structured compression method of the large language model as described in any one of the above is implemented.

[0015] The hierarchical hybrid structured compression method and device of the large language model provided by the present invention determine the compression ratio of each layer of the target model according to the importance of each layer of the target model, realize the allocation of differential compression ratios, and perform weighted low-rank decomposition and differential singular value allocation processing on the multi-head attention sub-layer for the current processing layer during the compression process, and adopt channel pruning and compensation processing for the feed-forward network sub-layer, so as to complete the compression of the target model, and can achieve the balance between the compression rate and performance of the LLM. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic flowchart of the hierarchical hybrid structured compression method of the large language model provided by the present invention.

[0018] Figure 2 It is an overall schematic diagram of a hierarchical hybrid structured compression method for a large language model provided by the present invention.

[0019] Figure 3 It is a schematic diagram of fine-tuning after compression provided by the present invention.

[0020] Figure 4Schematic diagram of the hierarchical hybrid structured compression device for the large language model provided by the present invention.

[0021] Figure 5 Schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0023] To better understand the hierarchical hybrid structured compression method and device for the large language model provided by the present invention, a brief introduction to the related technologies of the present invention will be made first.

[0024] The matrix low-rank approximation technology is a method of approximating a high-dimensional matrix using the product of low-rank matrices. In recent years, the low-rank decomposition approximation technology has been widely applied in fields such as collaborative filtering, image compression, and text mining, and also has a wide application space in the field of model compression. However, due to the differences in the importance of parameters in the weight matrix, directly performing low-rank approximation on the original weight matrix will cause serious performance losses. Therefore, the weighted low-rank decomposition method is proposed. However, the singular value decomposition (SVD) of the weighted matrix is also a problem that is difficult to directly solve, and obtaining the exact solution requires a huge amount of computing resources. Regarding the problem of weighting the weights, many methods adopt gradient-based estimation methods, but these methods are usually experimented on relatively small-scale models and do not consider that the calculation and storage of gradients on the LLM both consume a large amount of computing resources. Directly adopting the gradient-based weighting method will greatly reduce the efficiency of the compression algorithm. In addition, regarding the singular value decomposition problem of the weighted matrix, an approximate solution is usually obtained according to the approximation method, but this method needs to balance between accuracy and computational efficiency. Therefore, designing a gradient-independent weighting method and a simple and efficient weighted matrix decomposition method for the weight matrix is of great significance for the compression of the weight matrix.

[0025] Model pruning is an effective technique for compressing models. By removing redundant parameters and connections in the model, it reduces the size and computational complexity of the model, thereby improving the inference efficiency. During the pruning process, the granularity of the pruning unit is extremely important. Fine-grained pruning can more precisely control the size and complexity of the model to meet specific resource constraints or performance requirements, while coarse-grained pruning can bring significant inference acceleration. During the pruning process, it is first necessary to estimate the weight importance on a certain calibration dataset and aggregate it into granularity importance through a reasonable aggregation method. Essentially, model pruning will prune the least important part of the weights according to the granularity importance. Therefore, model pruning is of great significance for realizing the structured compression of models.

[0026] Figure 1 is a schematic flowchart of the hierarchical hybrid structured compression method for large language models provided by the present invention, as Figure 1 shown, the method includes the following steps: Step 100: Determine the compression ratio of each layer of the target model based on the importance degree of each layer of the target model.

[0027] Step 101: Compress the target model based on the compression ratio of each layer to obtain a compressed model.

[0028] Among them, compressing the target model based on the compression ratio of each layer includes: For the currently processed layer, calculate the second norm of the input activation value corresponding to each weight of the currently processed layer on the calibration dataset to obtain the feature norm; When the currently processed layer is a multi-head attention sublayer, based on the feature norm of the currently processed layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value assignment processing on the currently processed layer to complete the compression of the currently processed layer; When the currently processed layer is a feed-forward network sublayer, based on the feature norm of the currently processed layer and the compression ratio of each layer, perform channel pruning and compensation processing on the currently processed layer to complete the compression of the currently processed layer.

[0029] Specifically, in the embodiments of the present invention, the target model is an LLM model that needs to be compressed.

[0030] First, the compression ratio of each layer can be determined according to the importance degree of each layer of the target model. It can be understood that a lower compression ratio is assigned to the layer with higher importance to retain more parameters, while a higher compression ratio is assigned to the layer with lower importance to reduce redundant parameters. This step is to ensure that key features can be retained to the greatest extent during the compression process, thereby minimizing the impact on the model performance.

[0031] A hierarchical hybrid structured compression method for large language models provided by the present invention determines the compression ratio of each layer of the target model based on the importance level of each layer of the target model, including: Determine the importance level of each layer of the target model based on the change situation of the input activation value and the output activation value of each layer of the target model; the change situation includes the amplitude change situation and the direction change situation; Determine the compression ratio of each layer of the target model based on the importance level of each layer of the target model; among them, the lower the compression ratio is allocated to the layer with higher importance, and the higher the compression ratio is allocated to the layer with lower importance.

[0032] Specifically, in the embodiment of the present invention, for each layer of the target model, first determine its importance level and allocate different compression ratios according to this importance level.

[0033] The embodiment of the present invention can quantify the importance level of each layer according to the change situation of the input activation value and the output activation value of each layer in the target model, and these changes can be analyzed from two dimensions: amplitude and direction.

[0034] Input or output activation value is a high-dimensional vector and can be expressed as , where represents the number of samples, represents the sequence length, represents the feature dimension. To better understand its change situation, first perform normalization processing on the activation value to convert it into two independent components: amplitude and direction. Among them, the amplitude component is represented by , and the direction component is represented by a unit vector, that is .

[0035] On this basis, the importance level of the model layer is quantified by calculating the change situation of the input activation value and the output activation value. The degree of change of the current layer can be measured from two dimensions: The first metric measures the change in the size of the activation value, and its calculation formula is:

[0036] Among them, and respectively represent the unit vectors of the input and output activation values, reflecting the change in the activation amplitude.

[0037] The second metric reflects the change in the direction of the activation value, and its calculation formula is:

[0038] Here, Represents the cosine similarity between the input and output activation value directions, measuring the change in activation direction. and , which can fully characterize the influence of this layer in the model.

[0039] Next, in order to obtain a unified evaluation standard, the input and output changes of all layers can be normalized through normalization to obtain a comprehensive importance index for each layer. This importance index takes into account the amplitude and direction changes of each layer and quantifies the role of each layer in the model.

[0040] Finally, based on these importance indicators, the embodiment of the present invention can allocate the compression ratio of each layer through a reasonable compression strategy.

[0041] For layers with higher importance, the compression ratio will be lower to preserve the functions and information of the layer as much as possible, thereby avoiding excessive loss of its contribution to model performance during the compression process. For layers with lower importance, a higher compression ratio can be assigned to reduce computing and storage overhead. This compression strategy can ensure that the layers that have a greater impact on model performance are retained to the greatest extent possible, while effectively compressing less important layers, thereby achieving efficient model compression.

[0042] Next, the target model can be compressed according to the compression ratio determined according to the importance. In the specific compression process, for the current processing layer, the binary norm of the input activation value corresponding to each weight of the current processing layer on the calibration data set is first calculated to obtain the feature norm. The calculation of the feature norm can reflect the influence of the weight on the input data, which can be used to guide the compression strategy.

[0043] Among them, the calibration dataset refers to a small-scale dataset used to evaluate and adjust model parameters during the model compression process.

[0044] After obtaining the feature norm, differentiated strategies can be adopted for different types of the current processing layer: For the Multi-Head Attention (MHA) sublayer, weighted low-rank decomposition and differentiated singular value allocation can be performed according to the feature norm of the layer and the corresponding compression ratio. Among them, weighted low-rank decomposition reduces parameter redundancy by matrix decomposition while retaining the main information structure as much as possible. Differentiated singular value allocation adjusts the assigned singular values ​​according to the importance of each part during the decomposition process, so that the features of the key parts are retained, while the unimportant parts are weakened or removed. This method ensures that the effectiveness of the multi-head attention mechanism is maintained while reducing the computational complexity.

[0045] For the Feed-Forward Neural Network (FFN) sublayer, channel pruning and compensation processing can be performed according to the feature norm and compression ratio of this layer. During the channel pruning process, neurons with low contribution degrees can be removed based on the importance of the feature norm, thereby reducing the computational burden. The introduction of compensation processing is to make up for the information loss caused by pruning. Certain fine-tuning strategies or compensation weight adjustments can be adopted to reduce the impact of pruning on the overall performance of the model. Thus, while reducing the number of parameters, the FFN sublayer can still maintain a high expressive ability.

[0046] The hierarchical hybrid structured compression method for large language models provided by the present invention determines the compression ratio of each layer of the target model according to the importance degree of each layer of the target model, realizes the allocation of differential compression ratios. For the current processing layer during the compression process, weighted low-rank decomposition and differential singular value allocation processing are performed on the multi-head attention sublayer, and channel pruning and compensation processing are adopted for the FFN sublayer, thereby completing the compression of the target model, and the compression rate and performance balance of the LLM can be achieved.

[0047] According to a hierarchical hybrid structured compression method for large language models provided by the present invention, based on the feature norm of the current processing layer and the compression ratio of each layer, weighted low-rank decomposition and differential singular value allocation processing are performed on the current processing layer, including: Based on the feature norm, the original weight matrix of the current processing layer is weighted to obtain a weighted matrix; Perform singular value decomposition on the weighted matrix, determine the relationship between the approximate matrix corresponding to different numbers of retained singular values and the perplexity, and determine the strength degree of the low-rankness of each projection matrix of the current processing layer; Based on the strength degree of the low-rankness of different projection matrices and the compression ratio of each layer, determine the number of singular values retained by the projection matrix of the current processing layer.

[0048] Specifically, in the method provided by the embodiments of the present invention, in order to overcome the problem that traditional weight matrix decomposition methods ignore the difference in weight importance, for the weight matrix of the multi-head attention sublayer, a weighted low-rank decomposition method is adopted for optimization.

[0049] In terms of specific operations, first calculate the feature norm of the current processing layer and use it to weight the original weight matrix to obtain a weighted matrix. The purpose of weighting is to ensure that during the subsequent decomposition process, weights with different importance levels can be differentially processed, so that the weights of key parts can be better retained, while redundant parts can be more effectively compressed.

[0050] Next, perform singular value decomposition (SVD) on the weighted matrix, and determine the relationship between the approximate matrices corresponding to retaining different numbers of singular values and the perplexity based on the decomposition results. Perplexity is an important indicator for measuring the quality of a language model, which reflects the model's predictive ability for data.

[0051] In this step, analyze the strength of the low-rank property of the weighted matrix to determine the number of singular values that should be retained by different projection matrices during the compression process.

[0052] Specifically, first obtain the decomposition result of the weighted matrix through singular value decomposition. Then, determine the strength of the low-rank property of different projection matrices by observing how the retention of different numbers of singular values affects the change in perplexity of the model on the natural language processing dataset.

[0053] Use the input activation value on the th feature as the weighted value of the weights of the multi-head attention sublayer. At this time, the weighted low-rank reconstruction error of the matrix can be expressed as:

[0054] where represents the element in the th row and th column of the weight matrix , and represent the matrices obtained by decomposing the weight matrix . According to the calculation method of the feature norm, it is found in this embodiment that the weights in the same column of the weight matrix have the same importance score. Therefore, this embodiment can simplify the problem accordingly. Assuming a diagonal matrix

[0055] , the weighted low-rank reconstruction error of the matrix can be expressed as: At this time, this problem can be solved by performing SVD decomposition on the weighted matrix , where , are orthogonal matrices obtained by SVD decomposition, containing left and right singular vectors respectively; is a diagonal matrix containing singular values, representing the main components of the data.

[0056] The conclusion of the weighted low-rank reconstruction error problem is: . In order to achieve the effect of compressing the weight matrix, this embodiment of the present invention retains the matrix and the first r eigen - components of the elements to obtain . Finally, the embodiment of the present invention obtains through matrix decomposition: .

[0057] This method ensures that the low - rank approximation of the weight matrix can more accurately retain key information while effectively reducing the parameter scale.

[0058] When compressing different modules of the compressed multi - head attention sub - layer, the singular value retention strategy of the weight matrix can be further optimized. Experiments show that the Q matrix (query matrix) and the K matrix (key matrix) can maintain good performance while retaining fewer singular values, while the V matrix (value matrix) and the O matrix (output matrix) need to retain more singular values to ensure that the model performance does not drop significantly. Therefore, under the condition of a given total number of parameters, a differential allocation strategy is adopted, so that the approximation of the V matrix and the O matrix uses more parameters, while the Q matrix and the K matrix use fewer parameters.

[0059] In some embodiments, the V matrix and the O matrix can be regarded as a group, and the Q matrix and the K matrix as another group. The two matrices within the same group use the same number of parameters, and different groups use different numbers of parameters. The ratio of the number of parameters of the two groups is set to , that is, the parameter allocation ratio of the V and O matrices is higher. Then, under the condition of a specified number of parameters , the calculation methods of the number of parameters of the two groups are as follows:

[0060] where is the number of parameters of the Q matrix and the K matrix; is the number of parameters of the V matrix and the O matrix.

[0061] In this way, under the condition of limited number of parameters, the storage and computing efficiency of the model are optimized as much as possible while maintaining the prediction ability of the model. Finally, through this weighted low - rank decomposition and differential singular value allocation method, the refined compression of the multi - head attention sub - layer is realized, and the computing efficiency and storage utilization rate of the model are improved.

[0062] According to a hierarchical hybrid structured compression method of a large - language model provided by the present invention, based on the feature norm of the current processing layer and the compression ratio of each layer, channel pruning and compensation processing are performed on the current processing layer, including: Based on the feature norm of the current processing layer and the original weight matrix, determine the importance scores of each weight of the current processing layer; Aggregate the importance scores of each weight to obtain the channel importance scores; Based on the dependency structure, group the channels, and use the sum of the channel importance scores within the group as the group importance score of each group of channels; Prune each group of channels based on the group importance score and the compression ratio of each layer; Based on the optimal pruning framework of optimal neurosurgery OBS, the parameters of each group of channels after pruning are adjusted to compensate for the pruning loss.

[0063] Specifically, in the method provided in the embodiment of the present invention, in order to efficiently compress the feedforward network sublayer, a strategy combining channel group pruning and compensation processing is adopted.

[0064] First, based on the feature norm of the current processing layer and the original weight matrix, the importance score of each weight of the layer is calculated.

[0065] In some implementations, for a single weight in a feedforward network sublayer , its importance score It can be calculated by multiplying the absolute value of the weight by its corresponding feature norm, that is:

[0066] This calculation method takes into account the numerical size of the weight itself and the impact of its corresponding input features, and can more accurately evaluate the contribution of the weight in the network.

[0067] After calculating the importance scores of each weight of the layer, the importance scores of individual weights can be aggregated into channel importance scores. No. OK The corresponding The importance score of each channel It can be expressed by calculating the bi-norm of the importance scores of all weights in the channel, that is:

[0068] in, Represents the dimension of the input features (i.e. the number of columns in the weight matrix), which determines the width of the channel.

[0069] This approach can measure the overall contribution of the entire channel to the feedforward network calculation, making pruning decisions more robust.

[0070] On this basis, the channels in the network are grouped according to the dependency structure, the interdependent channels are regarded as a whole, and the importance score of each group of channels is calculated. Specifically, for a group of channels, its importance score is equal to the sum of the importance scores of all channels in the group, that is, Group Channel Importance score It can be calculated according to the following formula:

[0071] in, Indicates the upper network The weight row vector of channels; Indicates the gating mechanism The weights associated with each channel; Indicates the lower layer network Column vector of weights for channels.

[0072] The basis for channel grouping is the dependency structure, that is, considering the information transmission relationship between channels to ensure that the key information flow is not destroyed during pruning.

[0073] After determining the group importance score, based on the compression ratio set for each layer, channels with lower group importance scores are selected for pruning. The principle of pruning is to retain channels that have a greater impact on the model as much as possible, while removing channels with lower group importance scores, while ensuring that a portion of the channels with the lowest importance are still retained to prevent excessive pruning from affecting the expressiveness of the model.

[0074] After obtaining the pruned channel set M, the index matrix of the pruned channels can be obtained , the Optimal Brain Surgeon (OBS) pruning framework is used to adjust the parameters of each group of channels after pruning to compensate for the loss caused by pruning. The remaining weights are updated under this framework, and the weight update increment The specific formula is as follows:

[0075] The core idea of ​​the OBS framework is based on the Hessian matrix The weights are adjusted so that the pruned network still maintains the best possible performance. Specifically, the remaining weight matrix after pruning is optimized and adjusted under the OBS framework so that it can compensate for the information loss caused by pruning to the greatest extent, thereby reducing the impact of pruning on model performance. The introduction of the OBS method allows the model to maintain a high level of accuracy even after pruning, while significantly reducing computing and storage requirements.

[0076] Finally, after completing channel pruning and parameter adjustment, the feedforward network sublayer is still a dense model, but the intermediate dimensions are significantly reduced, achieving more efficient computing and storage while minimizing the degradation of model performance. This method combines channel importance evaluation, dependency structure grouping, optimal pruning strategy, and OBS compensation optimization to achieve refined pruning of the feedforward network sublayer, allowing large language models to significantly reduce computational complexity, improve inference speed, and reduce storage overhead without significantly losing performance.

[0077] A hierarchical hybrid structured compression method for large language models provided by the present invention further includes: Based on lightweight fine-tuning LoRA, knowledge recovery is performed on the compressed model to obtain the final compressed model.

[0078] Specifically, the compressed model usually suffers from a certain performance loss because some redundant parameters and structures are removed during the compression process, which may affect the model's ability to represent input data. To better recover the performance of the compressed model, this method introduces lightweight fine-tuning technology (Low-Rank Adaptation, LoRA). LoRA is a method for recovering knowledge by fine-tuning a small number of parameters on the compressed model, which can effectively improve the model's performance on specific tasks without the need for large-scale retraining of the entire network.

[0079] The process of LoRA fine-tuning includes adding some low-rank adaptation modules to the compressed model, which only adjust a small part of the model's parameters, avoiding the high computational cost of full retraining. Through LoRA fine-tuning, fine-grained optimization can be performed on the performance bottleneck of the compressed model, making its performance on specific tasks close to the original model. In some cases, the task-specific ability of the model can even be restored or improved through more refined parameter adjustment.

[0080] In some embodiments, LoRA fine-tuning can be applied to train on the publicly available dataset Alpaca, and the number of training epochs is set to two. The main purpose of the training in these two epochs is to restore the performance of the compressed model on the Alpaca dataset through fine-tuning. This fine-tuning process only needs to adjust a small number of parameters in the LoRA module, so the training efficiency is relatively high.

[0081] Through this fine-tuning method, the compressed model can gradually recover the lost knowledge and optimize its performance on specific tasks, enabling the final compressed model to achieve a better balance between computational efficiency and task performance.

[0082] A hierarchical hybrid structured compression method for large language models provided by the present invention, the calibration dataset is obtained in the following manner: Extract a certain number of samples from the natural language processing dataset; After being processed by the tokenizer of the target model, a token sequence is obtained; The obtained token sequence is truncated to a specified length to obtain the calibration dataset.

[0083] Specifically, the calibration dataset can be obtained by extracting a certain number of samples from the publicly available natural language processing dataset.

[0084] First, randomly select a number of samples from publicly available natural language processing datasets. These samples represent the text data that may be encountered in actual use. In some embodiments, a number of samples can be randomly selected from the C4 dataset. The selected samples are processed by the tokenizer of the target model to convert the original text into a sequence of tokens. The tokenizer converts elements such as words and symbols in the text into tokens that can be processed by a computer, providing the basic data for subsequent model compression.

[0085] After obtaining the token sequences, these sequences can be truncated to ensure that the length of each sample's token sequence meets the standards required during the compression process.

[0086] In some embodiments, each token sequence can be truncated according to a preset length requirement. If the length of the current token sequence already meets the length requirement, the first n tokens of the sequence are directly taken as valid samples. If the length of the token sequence is insufficient to meet the requirement, new samples are selected from the dataset again until a sufficient number of valid samples that meet the length requirement are obtained. This process is repeated until the specified number of valid samples is collected.

[0087] The calibrated dataset is composed of the screened and truncated samples and will be used in the subsequent model compression process. The calibrated dataset not only provides the necessary input data for the compression process but also helps to statistically analyze relevant information of the compressed model to evaluate the impact of the compression strategy on the model performance.

[0088] The following further elaborates on the hierarchical hybrid structured compression method for large language models provided by the present invention through embodiments in specific application scenarios.

[0089] Figure 2 The overall schematic diagram of a hierarchical hybrid structured compression method for large language models provided by the present invention. As Figure 2 shown, the method includes the following steps: Step 1: Acquisition of calibration data: Randomly select samples from the publicly available dataset C4. Process the samples through the tokenizer to obtain token sequences. Determine whether the sequence is longer than the required sequence length. If it meets the requirement, the first n tokens can be intercepted on the token sequence as valid samples; if it does not meet the requirement, select samples again. Repeat the above operations until the specified number of valid samples is obtained. In the subsequent compression process, these samples will be used as the calibrated dataset and sent into the model for statistical analysis of relevant information.

[0090] Step 2: For the input or output activation values obtained during the forward calculation process , they can be arranged as and decomposed into amplitude and direction into two independent components:

[0091] To quantify the importance of the model layer, this embodiment needs to examine the changes in the input and output. Based on this, this embodiment can calculate the degree of change of the current layer from two dimensions:

[0092] where, and respectively represent the unit vectors of the input and output activation values, represents the cosine similarity between the directions of the input and output activation values, measures the change in the magnitude of the activation value, while reflects the change in direction. These two indicators jointly characterize the influence degree of this layer on the network. Next, in order to obtain a unified evaluation criterion, this embodiment calculates the importance of all layers through a series of normalizations based on the change amounts of the input and output of all layers :

[0093] Finally, based on the importance index calculated above, this embodiment can reasonably allocate the compression ratio of each layer . Specifically, the higher the importance of a layer, the lower the compression rate will be allocated to more completely retain its capabilities:

[0094] where, represents the base compression rate, represents the adjustment coefficient, represents the average value of the importance scores of all layers , ensuring that the compression ratio is within a reasonable range.

[0095] Step 3: To address the problem that it is difficult to calculate and store the gradients of large models, a method for estimating the importance of weights based on activation values is adopted. To reduce the memory requirements, this embodiment adopts a layer-by-layer compression method, independently compressing one Transformer layer each time. Let the weight matrix in the current Transformer layer be , and the input activation value statistically obtained during the forward calculation process is , where represents the number of samples, represents the sequence length, Denote the feature dimension. For any single weight in the current layer , the input activation value on the -th feature corresponding to it on the calibration dataset is The L2 norm of , which is called the feature norm.

[0096] Step 4: To address the problem that directly decomposing the weight matrix ignores the differences in weight importance, perform weighted low-rank approximation on the multi-head attention sub-layer. Use the feature norm as the weighting value for the weights of the multi-head attention sub-layer. At this time, the weighted low-rank reconstruction error of the matrix can be expressed as:

[0097] where and represent the matrices obtained through matrix decomposition. According to the calculation method of the feature norm, this embodiment finds that the weights in the same column of the weight matrix W have the same importance score. Therefore, this embodiment can simplify the problem accordingly. Assume the diagonal matrix , the weighted low-rank reconstruction error of the matrix can be expressed as:

[0098] At this time, this problem can be solved by performing SVD decomposition on the weighted matrix , , where is the orthogonal matrix obtained by SVD decomposition, containing the left and right singular vectors respectively; is a diagonal matrix containing singular values, representing the main components of the data.

[0099] The conclusion of the weighted low-rank reconstruction error problem is: . To achieve the effect of compressing the weight matrix, this embodiment retains the first r eigen-components of the elements of the matrices and , obtaining . Finally, this embodiment obtains through matrix decomposition: .

[0100] Step 5: In Table 1, in this embodiment, different amounts of parameters are used to approximate the weight matrices of different modules in the multi-head attention sub-layer. After approximation, the perplexity of the model on WikiText2 is used to reflect the change in model performance. From the results in Table 1, it can be found that the Q matrix and the K matrix can maintain good performance while retaining fewer singular values (using fewer parameters), while the V matrix and the O matrix need to retain more singular values (using more parameters) to maintain performance. Therefore, under the condition of a given total amount of parameters, more parameters should be used for the approximation of the V matrix and the O matrix. In this embodiment, the V matrix and the O matrix are regarded as a group, and the Q matrix and the K matrix are regarded as another group. The two matrices within the same group use the same amount of parameters, and different groups use different amounts of parameters. The ratio of the amounts of parameters of the two groups is set to . Then, under the condition of specifying the amount of parameters , the calculation methods of the amounts of parameters of the two groups are as follows:

[0101] Table 1: Influence of different amounts of parameter compression of different modules on the perplexity of the model on WikiText2

[0102] Step 6: For the feed-forward network sub-layer, channel group pruning is performed. First, estimate the importance score of a single weight in the feed-forward network sub-layer , and use the product of the absolute value of the weight and its corresponding eigen-norm as the importance score of the weight:

[0103] Then aggregate the single weight importance scores into channel importance scores. In this embodiment, the two-norm of the in-channel weight importance scores is used to represent the th row corresponding th channel importance score :

[0104] Among them, represents the dimension of the input feature (i.e., the number of columns of the weight matrix), which determines the width of the channel.

[0105] Group the channels in the network according to the dependency structure, regard the mutually dependent channels as a group of channels, and use the sum of the in-group channel importance scores as the group importance score. The specific method is as follows. The importance score of the th group of channels can be calculated according to the following formula:

[0106] Among them, represents the weight row vector of the -th channel in the upper network; represents the weight related to the -th channel in the gating mechanism; represents the weight column vector of the -th channel in the lower network.

[0107] According to the group importance score and the specified compression ratio, select to cut off the weights of the part with lower group importance, but retain the weights of the part with the lowest group importance. The specific formula is as follows:

[0108] Among them, is the retention ratio parameter, retaining the groups with the top importance scores; the groups with the last 1% of the importance scores in the ranking are also retained.

[0109] Step 7: After obtaining the pruned channel set M, the index matrix of the pruned channels can be obtained. In this embodiment, the OBS optimal pruning framework is introduced, and the update of the remaining weights in is completed under this framework. The specific formula for the weight update increment is as follows:

[0110] Among them represents the Hessian matrix. After completing the channel pruning and the update of the remaining parameters for the feed-forward network sublayer, the obtained model is still a dense model, but the intermediate dimension will be significantly reduced.

[0111] Step 8: When each layer of the Transformer in the model has gone through Steps 3 - 7, the compression of the entire model is completed. The compressed model usually has a certain performance loss, and further fine-tuning can effectively help the compressed model to recover knowledge. Figure 3 is a schematic diagram of the post-compression fine-tuning provided by the present invention. In this embodiment, LORA fine-tuning is used to train for two epochs on the public dataset Alpaca to recover the knowledge of the compressed model, as Figure 3 shown. The forward calculation process of the model obtained after training is as follows:

[0112] Among them, represents the forward calculation of the feed-forward neural network (FFN) of the input tensor ,​ Represents the input tensor Forward computation of multi-head attention (MHA); is the original weight matrix of the feed-forward layer in the pre-trained model, is the original weight matrix of the attention layer in the pre-trained model; is an adaptation matrix for inserting trainable parameters through LoRA to fine-tune the model behavior, in and the forms are respectively and .

[0113] The performance of this embodiment on the LLaMA-7B model is shown in Table 2. The evaluation datasets include the perplexity on the WikiText2 and PTB datasets and the average accuracy on seven inference tasks: BooLQ, PIQA, HellaSwag, WinoGrande, ARC-E, ARC-C, and OBQA. The selected comparison methods include multiple large model compression methods: Table 2: Results of this embodiment and other comparison methods on the test set

[0114] It can be seen from the test results in the table that in the zero-shot perplexity and zero-shot question-answering tasks, the present invention can greatly improve the performance of the compressed model, especially with more obvious advantages at high compression ratios, and has significant advantages for compressing LLMs.

[0115] Next, the hierarchical hybrid structured compression device for large language models provided by the present invention will be described. The hierarchical hybrid structured compression device for large language models described below can be correspondingly referred to the hierarchical hybrid structured compression method for large language models described above.

[0116] Figure 4 is a schematic structural diagram of the hierarchical hybrid structured compression device for large language models provided by the present invention. As Figure 4 shown, the device includes the following modules: Determination module 400, configured to determine the compression ratio of each layer of the target model based on the importance of each layer of the target model; Compression module 410, configured to compress the target model based on the compression ratio of each layer to obtain a compressed model; Among them, compressing the target model based on the compression ratio of each layer includes: For the current processing layer, calculate the two-norm of the input activation value corresponding to each weight of the current processing layer on the calibration dataset to obtain the feature norm; When the current processing layer is a multi-head attention sub-layer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value allocation processing on the current processing layer to complete the compression of the current processing layer; When the current processing layer is a feed-forward network sub-layer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform channel pruning and compensation processing on the current processing layer to complete the compression of the current processing layer.

[0117] According to a hierarchical hybrid structured compression device for a large language model provided by the present invention, based on the feature norm of the current processing layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value allocation processing on the current processing layer, including: Based on the feature norm, weight the original weight matrix of the current processing layer to obtain a weighted matrix; Perform singular value decomposition on the weighted matrix, determine the relationship between the approximate matrix corresponding to different numbers of retained singular values and the perplexity, and determine the strength of the low-rankness of each projection matrix of the current processing layer; Based on the strength of the low-rankness of different projection matrices and the compression ratio of each layer, determine the number of singular values retained by the projection matrix of the current processing layer.

[0118] According to a hierarchical hybrid structured compression device for a large language model provided by the present invention, based on the feature norm of the current processing layer and the compression ratio of each layer, perform channel pruning and compensation processing on the current processing layer, including: Based on the feature norm of the current processing layer and the original weight matrix, determine the importance score of each weight of the current processing layer; Aggregate the importance scores of each weight to obtain the channel importance score; Group the channels based on the dependency structure, and use the sum of the channel importance scores within the group as the group importance score of each group of channels; Based on the group importance score and the compression ratio of each layer, perform pruning processing on each group of channels; Based on the optimal brain surgeon (OBS) optimal pruning framework, adjust the parameters of each group of pruned channels to compensate for the pruning loss.

[0119] According to a hierarchical hybrid structured compression device for a large language model provided by the present invention, the device further includes: A recovery module for performing knowledge recovery on the compressed model based on lightweight fine-tuning (LoRA) to obtain the final compressed model.

[0120] According to a hierarchical hybrid structured compression device for a large language model provided by the present invention, the calibration dataset is obtained in the following manner: Extract a certain number of samples from the natural language processing dataset; After being processed by the tokenizer of the target model, a token sequence is obtained; The obtained token sequence is truncated to a specified length to obtain a calibration dataset.

[0121] According to a hierarchical hybrid structured compression device for a large language model provided by the present invention, based on the importance level of each layer of the target model, the compression ratio of each layer of the target model is determined, including: Based on the change situation of the input activation value and the output activation value of each layer of the target model, the importance level of each layer of the target model is determined; the change situation includes the amplitude change situation and the direction change situation; Based on the importance level of each layer of the target model, the compression ratio of each layer of the target model is determined; among them, the higher the importance, the lower the compression ratio assigned to the layer, and the lower the importance, the higher the compression ratio assigned to the layer.

[0122] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 5 shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the hierarchical hybrid structured compression method for the large language model provided by each of the above methods, and the method includes: Based on the importance level of each layer of the target model, the compression ratio of each layer of the target model is determined; Based on the compression ratio of each layer, the target model is compressed to obtain a compressed model; Among them, based on the compression ratio of each layer, compressing the target model includes: For the current processing layer, calculate the second norm of the input activation value corresponding to each weight of the current processing layer on the calibration dataset to obtain a feature norm; When the current processing layer is a multi-head attention sublayer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value allocation processing on the current processing layer to complete the compression of the current processing layer; When the current processing layer is a feed-forward network sublayer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform channel pruning and compensation processing on the current processing layer to complete the compression of the current processing layer.

[0123] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0124] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the hierarchical hybrid structured compression method of the large language model provided by the above-mentioned various methods. The method includes: Determine the compression ratio of each layer of the target model based on the importance of each layer of the target model; Compress the target model based on the compression ratio of each layer to obtain a compressed model; Among them, compressing the target model based on the compression ratio of each layer includes: For the current processing layer, calculate the second norm of the input activation value corresponding to each weight of the current processing layer on the calibration data set to obtain a feature norm; When the current processing layer is a multi-head attention sub-layer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value assignment processing on the current processing layer to complete the compression of the current processing layer; When the current processing layer is a feed-forward network sub-layer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform channel pruning and compensation processing on the current processing layer to complete the compression of the current processing layer.

[0125] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the hierarchical hybrid structured compression method of the large language model provided by the above-mentioned various methods. The method includes: Determine the compression ratio of each layer of the target model based on the importance of each layer of the target model; Compress the target model based on the compression ratio of each layer to obtain a compressed model; Among them, based on the compression ratio of each layer, compressing the target model includes: For the current processing layer, calculate the second norm of the input activation value corresponding to each weight of the current processing layer on the calibration dataset to obtain the feature norm; When the current processing layer is a multi-head attention sublayer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform weighted low-rank decomposition and differential singular value allocation processing on the current processing layer to complete the compression of the current processing layer; When the current processing layer is a feed-forward network sublayer, based on the feature norm of the current processing layer and the compression ratio of each layer, perform channel pruning and compensation processing on the current processing layer to complete the compression of the current processing layer.

[0126] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0128] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A hierarchical hybrid structured compression method for a large language model, characterized in that: include: Determining a compression ratio of each layer of the target model based on the importance of each layer of the target model; Based on the compression ratio of each layer, compressing the target model to obtain a compressed model; The target model is compressed based on the compression ratio of each layer, including: For the current processing layer, calculating the binary norm of the input activation value corresponding to each weight of the current processing layer on the calibration data set to obtain a feature norm; In the case where the current processing layer is a multi-head attention sublayer, based on the feature norm of the current processing layer and the compression ratio of each layer, weighted low-rank decomposition and differentiated singular value allocation processing are performed on the current processing layer to complete the compression of the current processing layer; In the case where the current processing layer is a feedforward network sublayer, channel pruning and compensation processing are performed on the current processing layer based on the feature norm of the current processing layer and the compression ratio of each layer to complete the compression of the current processing layer.

2. The hierarchical hybrid structured compression method for a large language model according to claim 1, characterized in that: Based on the feature norm of the current processing layer and the compression ratio of each layer, weighted low-rank decomposition and differentiated singular value allocation processing are performed on the current processing layer, including: Based on the feature norm, weighting the original weight matrix of the current processing layer to obtain a weighted matrix; Performing singular value decomposition on the weighted matrix, determining the relationship between the approximate matrix corresponding to different numbers of retained singular values ​​and the perplexity, and determining the strength of the low rank of each projection matrix of the current processing layer; Based on the low-rank properties of different projection matrices and the compression ratio of each layer, the number of singular values ​​retained by the projection matrix of the current processing layer is determined.

3. The hierarchical hybrid structured compression method for a large language model according to claim 1, characterized in that: Based on the feature norm of the current processing layer and the compression ratio of each layer, channel pruning and compensation processing are performed on the current processing layer, including: Determining the importance score of each weight of the current processing layer based on the feature norm of the current processing layer and the original weight matrix; Aggregating the importance scores of the weights to obtain a channel importance score; The channels are grouped based on the dependency structure, and the sum of the channel importance scores within the group is used as the group importance score of each group of channels; Based on the group importance score and the compression ratio of each layer, pruning each group of channels; Based on the optimal pruning framework of optimal neurosurgery OBS, the parameters of each group of channels after pruning are adjusted to compensate for the pruning loss.

4. The hierarchical hybrid structured compression method for a large language model according to any one of claims 1 to 3, characterized in that: The method further comprises: Based on lightweight fine-tuning LoRA, knowledge recovery is performed on the compressed model to obtain the final compressed model.

5. The hierarchical hybrid structured compression method for a large language model according to claim 1, characterized in that: The calibration data set is obtained according to the following method: Extract a certain number of samples from the natural language processing dataset; After being processed by the word segmenter of the target model, a word segmentation sequence is obtained; The obtained word segmentation sequence is truncated to a specified length to obtain a calibration data set.

6. The hierarchical hybrid structured compression method for a large language model according to claim 1, characterized in that: The step of determining the compression ratio of each layer of the target model based on the importance of each layer of the target model includes: Determine the importance of each layer of the target model based on changes in input activation values ​​and output activation values ​​of each layer of the target model; the changes include amplitude changes and direction changes; Based on the importance of each layer of the target model, the compression ratio of each layer of the target model is determined; wherein, the compression ratio allocated to the layer with higher importance is lower, and the compression ratio allocated to the layer with lower importance is higher.

7. A hierarchical hybrid structured compression device for a large language model, characterized in that: include: A determination module, used for determining the compression ratio of each layer of the target model based on the importance of each layer of the target model; A compression module, used for compressing the target model based on the compression ratio of each layer to obtain a compressed model; The target model is compressed based on the compression ratio of each layer, including: For the current processing layer, calculating the binary norm of the input activation value corresponding to each weight of the current processing layer on the calibration data set to obtain a feature norm; In the case where the current processing layer is a multi-head attention sublayer, based on the feature norm of the current processing layer and the compression ratio of each layer, weighted low-rank decomposition and differentiated singular value allocation processing are performed on the current processing layer to complete the compression of the current processing layer; In the case where the current processing layer is a feedforward network sublayer, channel pruning and compensation processing are performed on the current processing layer based on the feature norm of the current processing layer and the compression ratio of each layer to complete the compression of the current processing layer.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the hierarchical hybrid structured compression method for a large language model according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the hierarchical hybrid structured compression method for a large language model according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the hierarchical hybrid structured compression method for a large language model according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Model code generation method and device

    CN120301432A

  • Voice model compression method, device and equipment and readable storage medium

    CN120356473A

  • Large language model reasoning method and device, electronic equipment, storage medium and program product

    CN120952152A