Model compression method and device and electronic equipment

By evaluating the importance of the weight matrix of the neural network model and splitting layer by layer, combining pruning and quantization processing, the problem of poor model compression effect is solved, and better compression effect and smaller output loss is achieved. It is suitable for devices with weak storage and computing power.

CN120373370APending Publication Date: 2025-07-25FANXING INTELLIGENT COMPUTING TECHNOLOGY (BEIJING) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510395992.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the model compression effect is poor, a single compression method is difficult to meet user requirements, and it is easy to cause damage to the inter-layer relationship of the original model.

Method used

By obtaining the weight matrix of the neural network model to be compressed, the importance evaluation is performed based on the original input data, the weight matrix is split into important parameter matrix and non-important parameter matrix, and the compression process is carried out layer by layer, including pruning and quantization.

Benefits of technology

The weight parameters of each target network layer are refined, which avoids inter-layer relationship damage, increases model compression rate and reduces output losses, making the model deployment more efficient on devices with weak storage and computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373370A_ABST
    Figure CN120373370A_ABST
Patent Text Reader

Abstract

The invention discloses a model compression method and device and electronic equipment, belongs to the field of artificial intelligence, and is used for solving the problem of poor model compression effect in related technologies. Comprising the following steps: acquiring a first weight matrix of each target network layer in a to-be-compressed neural network model; for each target network layer, according to the original input data of the target network layer and the first weight matrix, carrying out importance evaluation on each weight parameter in the first weight matrix to obtain an evaluation result; splitting the first weight matrix into an important parameter matrix and a non-important parameter matrix according to the evaluation result and a preset splitting strategy; and from the first target network layer in the target network layers, performing compression processing on the to-be-compressed neural network model layer by layer according to the important parameter matrix and the non-important parameter matrix until all the target network layers are compressed, thereby obtaining the compressed neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence, and specifically relates to a model compression method, device and electronic equipment. Background Art

[0002] Generative AI technology is in the stage of large-scale implementation, but the huge number of parameters and computational complexity of large models make it impossible to efficiently deploy and reason the models on terminal devices. Therefore, model compression has become a necessary task for the implementation of generative AI technology on the terminal side. At present, the compression methods for trained models are mainly divided into quantization, pruning, knowledge distillation, and weight sharing. However, a single compression method has limited compression of the model and is difficult to meet user requirements.

[0003] In this regard, a model compression method proposed in the related art is to perform end-to-end compression and fine-tune the entire model. Although this model compression method increases the degree of model compression, it is necessary to update all weight parameters of the entire model when iteratively updating the parameters. It lacks consideration of the weight parameters of each layer in the original model and is likely to destroy the inter-layer relationship of the original model, thereby affecting the output effect of the model on other data. In addition, the related art also proposed an automated convolutional neural network quantization pruning method based on reinforcement learning. This method uses reinforcement learning to search for the best convolutional neural network pruning and quantization strategy, but it is not universal for other network models. Therefore, it is particularly necessary to provide a general model compression scheme with better compression effect. Summary of the invention

[0004] The embodiments of the present application provide a model compression method, device and electronic device, which can solve the problem of poor model compression effect in related technologies.

[0005] In a first aspect, an embodiment of the present application provides a model compression method, including: obtaining a first weight matrix of each target network layer in a neural network model to be compressed; for each of the target network layers, performing an importance evaluation on each weight parameter in the first weight matrix according to the original input data of the target network layer and the first weight matrix to obtain an evaluation result; according to the evaluation result and a preset splitting strategy, splitting the first weight matrix into an important parameter matrix and an unimportant parameter matrix; starting from the first target network layer in each of the target network layers, compressing the neural network model to be compressed layer by layer according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed to obtain a compressed neural network model.

[0006] Second aspect, an embodiment of the present application provides a model compression device, including: an acquisition module, configured to acquire a first weight matrix of each target network layer in a neural network model to be compressed; an evaluation module, configured to, for each of the target network layers, evaluate the importance of each weight parameter in the first weight matrix according to the original input data of the target network layer and the first weight matrix, to obtain an evaluation result; a splitting module, configured to split the first weight matrix into an important parameter matrix and an unimportant parameter matrix according to the evaluation result and a preset splitting strategy; a compression processing module, configured to start from the first target network layer among the target network layers, and perform layer-by-layer compression processing on the neural network model to be compressed according to the important parameter matrix and the unimportant parameter matrix until all the target network layers are compressed, to obtain a compressed neural network model.

[0007] Third aspect, an embodiment of the present application provides an electronic device, which includes a processor; and a memory arranged to store computer-executable instructions, the computer-executable instructions being configured to be executed by the processor, and the computer-executable instructions being executed by the processor to implement the steps of the model compression method as described in the first aspect.

[0008] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the steps of the model compression method as described in the first aspect.

[0009] Fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and the computer program, when executed by a processor, implements the steps of the model compression method as described in the first aspect.

[0010] Sixth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run executable instructions to implement the steps of the model compression method as described in the first aspect.

[0011] In the embodiments of the present application, by obtaining the first weight matrix of each target network layer in the neural network model to be compressed, for each target network layer, according to the original input data of the target network layer and the first weight matrix, the importance of each weight parameter in the first weight matrix is evaluated to obtain an evaluation result, and according to the evaluation result and the preset splitting strategy, the first weight matrix is split into an important parameter matrix and an unimportant parameter matrix. Furthermore, starting from the first target network layer among the target network layers, the neural network model to be compressed is compressed layer by layer according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed, and a compressed neural network model is obtained. It can be seen that this technical solution can finely split the weight matrix of each target network layer in the model based on the importance of each weight parameter, realizing the refined management of the weight parameters of each target network layer. Moreover, by layer-by-layer and specifically compressing the split weight parameter matrices, compared with the model compression solutions that adjust multiple layers or all weight parameters simultaneously, the refined processing of each weight parameter matrix is realized, avoiding the situation of damaging the inter-layer relationship, and while increasing the model compression rate, ensuring that the model output has a small loss, making the model compression effect better, which is beneficial for the model to be deployed and used on devices with weak storage and computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic flowchart of a model compression method provided by an embodiment of the present application; Figure 2 is a schematic diagram of an importance matrix provided by an embodiment of the present application; Figure 3 is a schematic diagram of a weight matrix provided by an embodiment of the present application; Figure 4 is a schematic flowchart of a model compression method provided by another embodiment of the present application; Figure 5 is a schematic diagram of the network structure of a neural network model provided by an embodiment of the present application; Figure 6 is a schematic diagram of the structure of a model compression device provided by an embodiment of the present application; Figure 7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0014] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0015] Generative artificial intelligence technology is in the stage of large-scale implementation and application. However, the huge number of parameters and computational complexity of large models make it impossible to efficiently deploy and infer on terminal devices. Therefore, model compression has become an essential task for the implementation of generative artificial intelligence technology on the edge side. Currently, the compression methods for trained models mainly include quantization, pruning, knowledge distillation, weight sharing, etc. However, a single compression method has limited compression for the model and is difficult to meet user requirements.

[0016] In response to this, a model compression method proposed by related technologies is as follows: The first neural network model processes the first data to be processed to obtain a first output; the third neural network model processes the first data to be processed to obtain a second output; based on the first output and the second output, a first target loss is determined, and the second neural network model is updated based on the first target loss to obtain an updated second neural network model; the updated second neural network model is compressed to obtain a target neural network model. This method performs end-to-end compression and fine-tuning on the entire model. Each iteration requires updating all the model parameters of the second neural network, lacking consideration of the weight parameters of each layer in the original model, and it is also easy to damage the inter-layer relationship of the original model, thereby affecting the output effect of the model on other data. Related technologies also propose an automated convolutional neural network quantization and pruning method based on reinforcement learning: First, obtain a dataset of images, pre-train the images using the initialized model to obtain the average rank of the feature maps output by each filter, and combine the average rank with the global importance ranking of the filters to obtain filter importance information; through reinforcement learning, implement automated neural network model quantization and pruning operations, obtain a neural network model compression strategy with the highest model accuracy, and obtain the final neural network model after pruning. This method uses reinforcement learning to search for the best convolutional neural network pruning and quantization strategy, with high implementation complexity, and can only be used for compressing convolutional neural networks and cannot be used under other network structures. Therefore, it is particularly necessary to provide a general and better model compression scheme.

[0017] In response to the above problems, the embodiments of the present application provide a model compression method, apparatus, and electronic device, which can finely split the weight matrix of each target network layer in the model based on the importance of each weight parameter, realizing refined management of the weight parameters of each target network layer. Moreover, by performing compression processing on each of the split weight parameter matrices layer by layer, compared with the model compression scheme that adjusts multiple layers or all weight parameters simultaneously, it realizes refined processing of each weight parameter matrix, avoids damaging the inter-layer relationship, can ensure a smaller loss in the model output while increasing the model compression rate, making the model compression effect better, and is conducive to the deployment and use of the model on devices with weak storage and computing power. In addition, this technical solution can implement compression processing on neural network models, rather than only being able to compress convolutional neural networks, so it has universality.

[0018] The following combines the accompanying drawings to elaborate in detail on the model compression method, apparatus, and electronic device provided by the embodiments of the present application through specific embodiments and their application scenarios.

[0019] Figure 1A model compression method provided by an embodiment of the present application is shown. This method can be executed by an electronic device, which can include: a server and / or a terminal device, where the terminal device can be, for example, an in-vehicle terminal or a mobile phone terminal, etc. In other words, this method can be executed by software or hardware installed in the electronic device. The method includes the following steps: Step 102: Obtain the first weight matrix of each target network layer in the neural network model to be compressed.

[0020] Among them, based on different prediction services to be implemented, the types of neural network models are also different. Exemplarily, the neural network model can be a model for implementing computer vision tasks / image recognition tasks, such as a convolutional neural network, etc.; the neural network model can be a model for implementing natural language processing tasks, such as a recurrent neural network, a long short-term memory network, a feedforward fully-connected deep neural network, etc. In addition, the neural network model can also be a combination of multiple neural networks, for example, simultaneously including a recurrent neural network and a convolutional neural network.

[0021] A neural network model usually includes multiple network layers, and each network layer includes multiple weight parameters, and these weight parameters constitute the weight matrix corresponding to this network layer. The target network layer can be a network layer that can be compressed among these network layers, such as a linear layer, a convolutional layer, etc. In a neural network model, the number of target network layers can be multiple.

[0022] Step 104: For each target network layer, according to the original input data of the target network layer and the first weight matrix, perform importance evaluation on each weight parameter in the first weight matrix to obtain an evaluation result.

[0023] Among them, the original input data of each target network layer in the neural network model can be obtained by acquiring a validation data set and inputting the data in the validation data set into the neural network model to be compressed and outputting the result.

[0024] Step 106: According to the evaluation result and a preset splitting strategy, split the first weight matrix into an important parameter matrix and an unimportant parameter matrix.

[0025] Among them, the unimportant parameter matrix can include a pruning parameter matrix and a fine-tuning parameter matrix.

[0026] Step 108: Starting from the first target network layer among each target network layer, perform compression processing on the neural network model to be compressed layer by layer according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed to obtain a compressed neural network model.

[0027] Among them, the compression processing can include pruning processing and quantization processing.

[0028] In an embodiment of the present application, by obtaining the first weight matrix of each target network layer in the neural network model to be compressed, for each target network layer, according to the original input data and the first weight matrix of the target network layer, importance evaluation is performed on each weight parameter in the first weight matrix to obtain an evaluation result, and according to the evaluation result and a preset splitting strategy, the first weight matrix is split into an important parameter matrix and an unimportant parameter matrix. Furthermore, starting from the first target network layer among the target network layers, the neural network model to be compressed is processed layer by layer according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed to obtain a compressed neural network model. It can be seen that this technical solution can perform a detailed split on the weight matrix of each target network layer in the model based on the importance of each weight parameter, realizing refined management of the weight parameters of each target network layer. Moreover, by performing targeted compression processing on each weight parameter matrix obtained by splitting layer by layer, compared with the model compression scheme that adjusts multiple layers or all weight parameters simultaneously, refined processing of each weight parameter matrix is achieved, avoiding the situation of damaging the inter-layer relationship, and while increasing the model compression rate, ensuring that the model output has a small loss, making the model compression effect better, which is beneficial for the model to be deployed and used on devices with weak storage and computing power.

[0029] In one implementation, obtaining the first weight matrix of each target network layer in the neural network model to be compressed (i.e., step 102) can be performed as the following steps A1 - A3: Step A1, obtain a validation data set.

[0030] Among them, according to the type distribution of the model training data, at least 1 piece of data of each type can be obtained to form a validation data set.

[0031] Step A2, input the validation data set into the neural network model to be compressed, and output the original input data and the original output data of each network layer.

[0032] In this step, before inputting the validation data set into the neural network model to be compressed, data preprocessing needs to be performed on the data in the validation data set. The steps of data preprocessing include but are not limited to: processing the data according to the input method of the model and splicing it into 1 piece of data as the input of the model.

[0033] Among them, processing the data set according to the input method of the model can be adding the template used during training to all the data and performing tokenization processing.

[0034] Step A3, for each target network layer, determine the first weight matrix of the target network layer according to the original input data and the original output data of the target network layer.

[0035] It can be understood that the first weight matrix is the original weight matrix of the target network layer, and the model compression process is to perform operations such as pruning and quantization on the first weight matrix.

[0036] Optionally, n the original input data and the original output data of the target network layer and the first weight matrix n have a functional relationship as shown in formula (1). The original output data of the target network layer n and the original input data of the target network layer of the

[0037] . (1) . (2) Among them, is the transformation n from the output of the n -th layer to the input of the

[0038] -th layer.

[0039] In one implementation, according to the original input data and the first weight matrix of the target network layer, an importance evaluation is performed on each weight parameter in the first weight matrix to obtain an evaluation result (i.e., step 104), which can be performed as follows: steps B1 - B2: Step B1, calculate the mean value of the original input data of the target network layer on each column to obtain the mean value matrix corresponding to the original input data.

[0040] Among them, the form of the original input data is generally a matrix, and the constituent elements of the matrix can be positive numbers, negative numbers, zeros, etc. Therefore, in order to perform a more accurate importance evaluation on the first weight matrix, the absolute value matrix corresponding to the original input data can be determined first, and then the mean value of each column of the absolute value matrix is calculated to obtain the mean value matrix corresponding to the original input data.

[0041] Step B2, determine the importance matrix corresponding to the first weight matrix according to the mean value matrix and the first weight matrix.

[0042] Among them, each element in the importance matrix is used to characterize the importance of the weight parameter at the corresponding position in the first weight matrix.

[0043] Optionally, the relationship among the mean matrix, the first weight matrix, and the importance matrix is shown in formula (3).

[0044] .(3) Among them, represents taking the absolute value of the first weight matrix. is the mean value of the absolute value of the original input data of the target network layer on the m dimension, that is, the above-mentioned mean matrix. represents and are multiplied element by element on the m -1st dimensional channel. is the importance matrix, which is a positive matrix. Thus, in each channel of the m -1st dimension, the importance of the corresponding weight parameter in this channel can be judged according to the magnitude of the value.

[0045] Exemplarily, as Figure 2 shown, in the original input data of the target network layer is a 3-row × 4-column matrix , and the first weight matrix is a 4-row × 4-column matrix , the can be determined to be . If is set to 1, then is , that is, the mean value of the absolute value of on each column. Thus, by multiplying and element by element on each row channel, can be obtained as .

[0046] In this embodiment, by calculating the mean value of the original input data of the target network layer on each column, the mean matrix corresponding to the original input data is obtained. Thus, according to the mean matrix and the first weight matrix, the importance matrix corresponding to the first weight matrix is determined, clarifying a method for determining the importance matrix, providing an evaluation criterion for realizing the importance evaluation of the weight parameters of the target network layer, and being beneficial to the accurate execution of the importance evaluation process.

[0047] In step 106, the non-important parameter matrix may include a pruning parameter matrix and a fine-tuning parameter matrix. The preset splitting strategy may include an important parameter threshold, a fine-tuning parameter threshold, a pruning parameter threshold, etc. Thus, when splitting the first weight matrix, the important parameter matrix and the non-important parameter matrix can be split according to the important parameter threshold, and then the non-important parameter matrix can be split into a pruning parameter matrix and a fine-tuning parameter matrix according to the fine-tuning parameter threshold and / or the pruning parameter threshold.

[0048] Among them, each parameter threshold may be a numerical threshold set for the importance value corresponding to each weight parameter in the evaluation result (i.e., each element in the above-mentioned importance matrix). Or it is a selection ratio threshold set for the sorted result after sorting the importance values of each row of channels in the evaluation result by size.

[0049] In this embodiment, the n first weight matrix of the layer target network layer , important parameter matrix , fine-tuning parameter matrix and pruning parameter matrix

[0050] . (4) Among them, according to each parameter threshold, the important parameter matrix can be obtained by setting the fine-tuning weight parameter and the pruning weight parameter to zero, and then the pruning parameter matrix can be obtained by setting the fine-tuning weight parameter to zero, and the fine-tuning parameter matrix can be obtained by setting the pruning weight parameter to zero.

[0051] Continuing with Figure 2 the importance matrix shown, when the selection ratio threshold of the important parameter threshold is 50%, that is, selecting the weight parameters corresponding to the first 50% of the importance values in each row of channels as important weight parameters, and the selection ratio threshold of the fine-tuning parameter threshold is 25%, that is, selecting among the weight parameters other than the important weight parameters in each row of channels, the weight parameters corresponding to the first 25% of the importance values as fine-tuning weight parameters, and the preset splitting strategy is: according to the important parameter threshold, splitting the first weight matrix into an important parameter matrix and a non-important parameter matrix, and according to the fine-tuning parameter threshold, splitting the non-important parameter matrix into a pruning parameter matrix and a fine-tuning parameter matrix, in this case, as Figure 3 shown, the first weight matrix can be split into the sum of 3 matrices, that is: .

[0052] In this embodiment, by splitting the first weight matrix into an important parameter matrix, a pruning parameter matrix, and a fine-tuning parameter matrix according to the evaluation result and the preset splitting strategy, a detailed splitting of the weight matrix of the target network layer is achieved, providing a data basis for more carefully processing the weight matrix to compress the target network layer. This is beneficial to increasing the compression rate while minimizing the model compression loss and avoiding affecting the output of each network layer of the model, thereby affecting the model prediction result.

[0053] In one implementation, the non-important parameter matrix may include a pruning parameter matrix and a fine-tuning parameter matrix. Starting from the first target network layer in each target network layer, according to the important parameter matrix and the non-important parameter matrix, the neural network model to be compressed is processed layer by layer (i.e., step 108), which can be performed as steps C1 - C4 as follows: Step C1, for each target network layer, prune the target network layer according to the pruning parameter matrix to obtain the pruned target network layer.

[0054] Among them, the pruned target network layer can be obtained by setting the pruning parameter matrix to zero. In this way, the weight matrix of the pruned target network layer is the sum of the important parameter matrix and the fine-tuning parameter matrix.

[0055] Step C2, determine the first pruning error corresponding to the target network layer according to the first output data of the pruned target network layer and the original output data of the target network layer.

[0056] It can be understood that if this target network layer is the first target network layer, then the first pruning error is the pruning error introduced after this target network layer is pruned; if this target network layer is not the first target network layer, then the first pruning error includes not only the pruning error introduced by this target network layer, but also the pruning errors and quantization errors introduced by multiple previous target network layers.

[0057] Step C3, determine the optimal fine-tuning parameter matrix corresponding to the target network layer according to the first pruning error.

[0058] In this step, considering avoiding affecting the inter-layer relationship of the model and reducing the model compression loss, it is necessary to minimize the output error of each network layer of the model. Since the pruning process of the target network layer introduces pruning errors, it is necessary to adjust the fine-tuning parameter matrix in the non-important parameter matrix to compensate for the pruning errors and obtain the optimal fine-tuning parameter matrix that can minimize the output error of the target network layer.

[0059] Step C4, perform quantization processing on the important parameter matrix and the optimal fine-tuning parameter matrix of the target network layer to obtain the second weight matrix corresponding to the target network layer.

[0060] In this embodiment, the relationship among the second weight matrix n corresponding to the layer target network layer, the important parameter matrix and the optimal fine-tuning parameter matrix is shown in formula (5).

[0061] (5) Among them, the quantization processing method can be any quantization method that can be quantized layer by layer, such as AWQ (Activation-aware Weight Quantization). It can be understood that after quantization processing, quantization errors will be introduced. In the embodiments of the present application, the calculation and minimization process of quantization errors are not taken as the focus. The quantization errors will be implicitly reflected in the first pruning errors of the second target network layer to the last target network layer of the model. The process of adjusting the fine-tuning parameter matrix is the process of compensating for the pruning errors of the current target network layer, the pruning errors of multiple previous target network layers, and the quantization errors.

[0062] In this embodiment, a neural network model compression method that combines quantization and pruning is implemented. Pruning is performed by setting the pruning parameter matrix to zero layer by layer, the fine-tuning parameter matrix is adjusted to compensate for the compression error, and quantization operations are performed on the important parameter matrix and the optimized fine-tuning parameter matrix to further increase the compression rate. The refined management of weight parameters is realized, the output error of each layer can be ensured to be minimized, thereby ensuring the minimization of the compression loss of the entire model, and ensuring that the inter-layer relationship is not affected as much as possible.

[0063] In one implementation, when the target network layer is not the first target network layer, according to the first output data of the target network layer after pruning processing and the original output data of the target network layer, the first pruning error corresponding to the target network layer (i.e., step C2) can be determined, and the following steps D1 - D3 can be executed: Step D1, determine the output data of the previous target network layer according to the input data of the previous target network layer and its corresponding second weight matrix.

[0064] Optionally, referring to the above formula (1), the input data of the previous target network layer can be substituted into , and the second weight matrix corresponding to the previous target network layer can be substituted into to calculate the output data of the previous target network layer.

[0065] Step D2, determine the first output data of the target network layer after pruning processing according to the output data of the previous target network layer, and the important parameter matrix and the fine-tuning parameter matrix of the target network layer.

[0066] Optionally, referring to the above formula (2), the output data of the previous target network layer can be substituted into , and the input data of the target network layer can be calculated. Then, referring to the above formula (1), the input data of the target network layer is substituted into , and the important parameter matrix and fine-tuning parameter matrix of the target network layer are substituted into , and the first output data of the target network layer after pruning processing is calculated.

[0067] Step D3: Calculate the loss value between the first output data and the original output data to obtain the first pruning error corresponding to the target network layer.

[0068] Optionally, the loss value can be the L2 norm of the difference between the first output data and the original output data. In practical applications, the loss value can also be the L2 loss, difference, etc. It can be understood that the first pruning error here will include the pruning error introduced by this target network layer, as well as the pruning errors and quantization errors introduced by multiple previous target network layers.

[0069] If the L2 norm of the difference is used to determine the loss value between the first output data and the original output data, then the n first pruning error corresponding to the th layer target network layer can be calculated by formula (6).

[0070] (6) where is the important parameter matrix of the n th layer target network layer, and is the fine-tuning parameter matrix of the n th layer target network layer. is the output data of the previous target network layer. It can be understood that the previous target network layer has been compressed. Calculating can obtain the input data of the n th layer target network layer. is the original output data of the n th layer target network layer. The first pruning error is the L2 norm of the difference between the output data obtained by processing the input data after pruning the first weight matrix of the n th layer target network layer and the original output data of this layer target network layer, expressed as a function of . Since the current target network layer is not the first target network layer, 0, then includes three parts: the pruning error of this layer, n- the pruning error and quantization error brought by the

[0071] In one implementation, when the target network layer is the first target network layer, the first pruning error corresponding to the target network layer is determined according to the first output data of the target network layer after pruning and the original output data of the target network layer (i.e., step C2), and the following steps E1 - E2 can be executed: Step E1, determine the first output data of the target network layer after pruning according to the original input data of the target network layer, and the important parameter matrix and fine-tuning parameter matrix of the target network layer.

[0072] Among them, referring to the above formula (1), the original input data of the target network layer can be substituted into and the important parameter matrix and fine-tuning parameter matrix of the target network layer can be substituted into to calculate the first output data of the target network layer after pruning.

[0073] Step E2, calculate the loss value between the first output data and the original output data to obtain the first pruning error corresponding to the target network layer.

[0074] Among them, similar to the above step D3, the loss value can be the L2 norm of the difference between the first output data and the original output data. The first pruning error corresponding to the first target network layer can be calculated by formula (7).

[0075] . (7) Among them, by calculating the first output data of the target network layer after pruning is obtained, and the L2 norm of the difference between the first output data and the original output data of the n th layer target network layer is the first pruning error . Since the current target network layer is the first target network layer, 0, then it is only the pruning error of this layer.

[0076] In one implementation, according to the first pruning error, the optimal fine-tuning parameter matrix corresponding to the target network layer is determined (i.e., step C3), and the following steps F1 - F4 can be executed: Step F1, when the first pruning error is greater than the preset error threshold, iteratively adjust each fine-tuning parameter in the fine-tuning parameter matrix of the target network layer according to the preset parameter adjustment strategy to obtain the adjusted fine-tuning parameter matrix.

[0077] Among them, since the fine-tuning parameters are non-important weight parameters, in order to avoid excessive changes in the parameters, a change threshold can be set for each fine-tuning parameter so that the parameter change does not exceed the threshold range. The preset parameter adjustment strategy can be a strategy regarding the adjustment amplitude, adjustment direction (i.e., increase or decrease) of each fine-tuning parameter, which is not specifically limited in the embodiments of the present application.

[0078] In this embodiment, if the first pruning error is less than or equal to the preset error threshold, it indicates that after the pruning process, the error between the output of this layer of the model and the original output is not large, meeting the requirements of model compression. Then, the fine-tuning parameter matrix can be not adjusted.

[0079] Step F2: Determine the second output data of the target network layer after pruning according to the input data of the target network layer, as well as the important parameter matrix and the adjusted fine-tuning parameter matrix of the target network layer.

[0080] Among them, if the target network layer is the first target network layer, the input data is the original input data. If the target network layer is not the first target network layer, the input data is determined based on the output data of the previous target network layer.

[0081] In this embodiment, referring to the above formula (1), the input data of the target network layer can be substituted into and the important parameter matrix and the adjusted fine-tuning parameter matrix of the target network layer can be substituted into to calculate the second output data of the target network layer after pruning.

[0082] Step F3: Determine the second pruning error corresponding to the target network layer according to the second output data and the original output data of the target network layer.

[0083] Among them, similar to the above steps D3 and E2, the second pruning error corresponding to the target network layer can be determined by calculating the L2 norm of the difference between the second output data and the original output data of the target network layer. Specifically, when the target network layer is not the first target network layer, referring to the above formula (6), the second pruning error corresponding to the target network layer can be calculated; when the target network layer is the first target network layer, referring to the above formula (7), the second pruning error corresponding to the target network layer can be calculated.

[0084] Step F4: When the second pruning error meets the iteration termination condition, determine the current adjusted fine-tuning parameter matrix as the optimal fine-tuning parameter matrix corresponding to the target network layer.

[0085] Optionally, the iteration termination condition can be: reaching the maximum number of iterations or the second pruning error being less than or equal to the preset error threshold.

[0086] In practical applications, the gradient descent method can be used to minimize the pruning error, and the optimization objective is as shown in formula (8).

[0087] (8) Wherein, represents the optimal fine-tuning parameter matrix after error minimization.

[0088] In this embodiment, by adjusting the fine-tuning parameters of each layer, the pruning error of this layer, the pruning error of the previous layer, and the quantization error can be compensated, and the minimization of the output error before and after the compression of the weight matrix of each layer can be ensured, avoiding the situation of damaging the inter-layer relationship when adjusting multiple layers or all weight matrices simultaneously.

[0089] Figure 4 is a schematic flowchart of a model compression method provided by another embodiment of the present application. As Figure 4 shown, the model compression method may include the following steps: Step 401, obtain the first weight matrix of each target network layer in the neural network model to be compressed.

[0090] Step 402, for each target network layer, according to the original input data of the target network layer and the first weight matrix, evaluate the importance of each weight parameter in the first weight matrix to obtain an evaluation result.

[0091] Step 403, according to the evaluation result and the preset splitting strategy, split the first weight matrix into an important parameter matrix, a pruning parameter matrix, and a fine-tuning parameter matrix.

[0092] Step 404, starting from the first target network layer among the target network layers, perform pruning processing on the current target network layer according to the pruning parameter matrix to obtain the current target network layer after pruning processing.

[0093] Step 405, according to the first output data of the current target network layer after pruning processing and the original output data of the current target network layer, determine the first pruning error corresponding to the current target network layer.

[0094] Step 406, according to the first pruning error, determine the optimal fine-tuning parameter matrix corresponding to the current target network layer.

[0095] Step 407, perform quantization processing on the important parameter matrix and the optimal fine-tuning parameter matrix of the current target network layer to obtain the second weight matrix corresponding to the current target network layer.

[0096] Step 408, when the compression of all target network layers is completed, obtain the compressed neural network model.

[0097] The specific processes of the above steps 401 to 408 have been described in detail in the above embodiments, and will not be repeated here.

[0098] In the embodiment of the present application, by obtaining the first weight matrix of each target network layer in the neural network model to be compressed, for each target network layer, according to the original input data of the target network layer and the first weight matrix, the importance of each weight parameter in the first weight matrix is evaluated to obtain an evaluation result, and according to the evaluation result and the preset splitting strategy, the first weight matrix is split into an important parameter matrix, a pruning parameter matrix, and a fine-tuning parameter matrix. Furthermore, starting from the first target network layer among the target network layers, according to the important parameter matrix, the pruning parameter matrix, and the fine-tuning parameter matrix, the neural network model to be compressed is pruned and quantized layer by layer until all target network layers are compressed to obtain the compressed neural network model. It can be seen that this technical solution can finely split the weight matrix of each target network layer in the model based on the importance of each weight parameter, realizing the refined management of the weight parameters of each target network layer. Moreover, by performing pruning and quantization processing on the split weight parameter matrices layer by layer, compared with the model compression scheme that adjusts multiple layers or all weight parameters simultaneously, it realizes the refined processing of each weight parameter matrix and avoids damaging the inter-layer relationship. By integrating two compression methods of quantization and pruning, while increasing the model compression rate, it can ensure that the model output has a small loss, making the model compression effect better and facilitating the deployment and use of the model on devices with weak storage and computing power.

[0099] Exemplarily, the model compression method provided in the embodiment of the present application can be used to compress the llama2 large language model based on the Transformer-Decoder architecture. The network structure of llama2 is as Figure 5 shown, including InputEmbedding (input embedding layer), Transformer Block Layer (Transformer layer), RMSNorm (normalization layer), Linear (linear layer), and Softmax (activation function). Among them, there are multiple Transformer layers, and each Transformer layer includes a normalization layer RMSNorm, a linear layer Linear for processing the Q, K, V matrices, RoPE (relative position encoding layer), activation function Softmax, SiLU (activation function), and multiple linear layers Linear.

[0100] For Figure 5 the model shown, the model compression method may include the following steps: Step 1, prepare a validation dataset.

[0101] For example, 256 pieces of data in various fields can be selected to ensure that there is at least 1 piece of data for each type, and templates used during training are added to all the data. After that, tokenization is performed, and then all the data is concatenated into 1 piece of data to form a validation dataset, which serves as the input to the model.

[0102] Step 2: Input the validation dataset into the model, and the original input data and original output data of each network layer are obtained as outputs. For each target network layer, based on the original input data and original output data of the target network layer, the first weight matrix of the target network layer is determined.

[0103] Such as Figure 5 in the model shown, the target network layer can be the Linear layer. In this step, the original input data and the original output data of all Linear layers can be recorded, and with reference to the above formulas (1) and (2), the first weight matrix of each Linear layer is calculated.

[0104] Step 3: Calculate the importance of each weight parameter in the first weight matrix, and based on the importance, decompose the weight matrix of each layer into an important parameter matrix, a pruning parameter matrix, and a fine-tuning parameter matrix.

[0105] For the specific implementation steps of this step, reference can be made to the respective embodiments corresponding to the above steps 104 - 106, which involve formulas (3) and (4), and will not be elaborated here.

[0106] Step 4: Starting from the first target network layer among the target network layers, the neural network model to be compressed is processed layer by layer according to the important parameter matrix, the pruning parameter matrix, and the fine-tuning parameter matrix until all target network layers are compressed, and a compressed neural network model is obtained.

[0107] Among them, since the Input Embedding layer is a token vectorization layer and is not compressed, starting from the first Linear layer thereafter, pruning, error minimization, and quantization operations are performed on the model layer by layer. Specifically: Step 4.1: Set the pruning parameter matrix of the th layer of Linear to zero, so that the pruning parameter matrix of this layer is a zero matrix, and unstructured pruning is performed. At this time, there is a pruning error in this layer, and the calculation formula for the pruning error is: .

[0108] Among them, the meanings of the elements in the formula can be referred to the above formulas (6) and (7), and will not be elaborated here. If it is the first layer of Linear currently, that is 0, then is only the pruning error; if is 0, then includes three parts: the pruning error of this layer, the pruning error brought by the previous layer, and the quantization error.

[0109] In this step, the Linear involved in the Transformer layer for processing the Q, K, and V matrices can be calculated in parallel.

[0110] Step 4.2, by minimizing the pruning error, determine the optimal fine-tuning parameter matrix corresponding to the layer Linear.

[0111] For the specific implementation steps of this step, reference can be made to the above steps F1 - F4, which will not be elaborated here. Finally, the weight matrix corresponding to the layer Linear when the error is minimized , , is the important parameter matrix of the layer Linear, is the optimal fine-tuning parameter matrix of the layer Linear.

[0112] Step 4.3, adopt a quantization method that can be quantized layer by layer to perform quantization processing on the weight matrix corresponding to the layer Linear, and obtain the second weight matrix corresponding to the layer Linear.

[0113] Step 4.4, repeat the above steps 4.1 to 4.3 until all Linear layers are compressed to obtain the compressed model.

[0114] It can be understood that after the last Linear layer, there is no next layer to make up for the pruning and quantization losses of this layer. Therefore, compared with the errors of other layers, the error of this layer is relatively large, but after the previous optimization, the overall compression loss is small.

[0115] It should be noted that for the model compression method provided in the embodiments of the present application, the execution subject can be a model compression device, or a control module in the model compression device for executing the model compression method. In the embodiments of the present application, taking the model compression device as the execution subject of the model compression method as an example, the model compression device provided in the embodiments of the present application is described.

[0116] Figure 6 is a schematic structural diagram of a model compression device provided in the embodiments of the present application. As Figure 6 shown, the model compression device includes: an acquisition module 610, an evaluation module 620, a splitting module 630, and a compression processing module 640.

[0117] An acquisition module 610 is configured to acquire a first weight matrix of each target network layer in the neural network model to be compressed; an evaluation module 620 is configured to, for each target network layer, evaluate the importance of each weight parameter in the first weight matrix according to the original input data of the target network layer and the first weight matrix, so as to obtain an evaluation result; a splitting module 630 is configured to split the first weight matrix into an important parameter matrix and an unimportant parameter matrix according to the evaluation result and a preset splitting strategy; a compression processing module 640 is configured to start from the first target network layer among the target network layers, and perform layer-by-layer compression processing on the neural network model to be compressed according to the important parameter matrix and the unimportant parameter matrix until all the target network layers are compressed, so as to obtain a compressed neural network model.

[0118] In one implementation, the unimportant parameter matrix includes a pruning parameter matrix and a fine-tuning parameter matrix; the compression processing module 640 includes: a pruning processing unit, a first determination unit, a second determination unit, and a quantization processing unit.

[0119] The pruning processing unit is configured to, for each target network layer, perform pruning processing on the target network layer according to the pruning parameter matrix, so as to obtain a pruned target network layer; the first determination unit is configured to determine a first pruning error corresponding to the target network layer according to the first output data of the pruned target network layer and the original output data of the target network layer; the second determination unit is configured to determine an optimal fine-tuning parameter matrix corresponding to the target network layer according to the first pruning error; the quantization processing unit is configured to perform quantization processing on the important parameter matrix and the optimal fine-tuning parameter matrix of the target network layer, so as to obtain a second weight matrix corresponding to the target network layer.

[0120] In one implementation, when the target network layer is not the first target network layer, the first determination unit is specifically configured to: determine the output data of the previous target network layer according to the input data of the previous target network layer and its corresponding second weight matrix; determine the first output data of the pruned target network layer according to the output data of the previous target network layer, and the important parameter matrix and the fine-tuning parameter matrix of the target network layer; calculate a loss value between the first output data and the original output data, so as to obtain a first pruning error corresponding to the target network layer.

[0121] In one implementation, when the target network layer is the first target network layer, the first determination unit is specifically configured to: determine the first output data of the pruned target network layer according to the original input data of the target network layer, and the important parameter matrix and the fine-tuning parameter matrix of the target network layer; calculate a loss value between the first output data and the original output data, so as to obtain a first pruning error corresponding to the target network layer.

[0122] In one implementation, the second determination unit is specifically configured to: when the first pruning error is greater than a preset error threshold, iteratively adjust each fine-tuning parameter in the fine-tuning parameter matrix of the target network layer according to a preset parameter adjustment strategy to obtain an adjusted fine-tuning parameter matrix; determine second output data of the pruned target network layer according to the input data of the target network layer, and the important parameter matrix and the adjusted fine-tuning parameter matrix of the target network layer; determine a second pruning error corresponding to the target network layer according to the second output data and the original output data of the target network layer; and when the second pruning error satisfies an iteration termination condition, determine the currently adjusted fine-tuning parameter matrix as the optimal fine-tuning parameter matrix corresponding to the target network layer.

[0123] In one implementation, the evaluation module 620 includes: a calculation unit and a third determination unit.

[0124] The calculation unit is configured to calculate the mean value of the original input data of the target network layer on each column to obtain a mean value matrix corresponding to the original input data; the third determination unit is configured to determine an importance matrix corresponding to the first weight matrix according to the mean value matrix and the first weight matrix; each element in the importance matrix is used to characterize the importance of the weight parameter at the corresponding position in the first weight matrix.

[0125] In one implementation, the acquisition module 610 includes: an acquisition unit, a model processing unit, and a fourth determination unit.

[0126] The acquisition unit is configured to acquire a validation data set; the model processing unit is configured to input the validation data set into a neural network model to be compressed and output the original input data and the original output data of each network layer; the fourth determination unit is configured to, for each target network layer, determine a first weight matrix of the target network layer according to the original input data and the original output data of the target network layer.

[0127] In the embodiment of the present application, by obtaining the first weight matrix of each target network layer in the neural network model to be compressed, for each target network layer, according to the original input data of the target network layer and the first weight matrix, the importance of each weight parameter in the first weight matrix is evaluated to obtain an evaluation result, and according to the evaluation result and the preset splitting strategy, the first weight matrix is split into an important parameter matrix and an unimportant parameter matrix. Furthermore, starting from the first target network layer among the target network layers, according to the important parameter matrix and the unimportant parameter matrix, the neural network model to be compressed is compressed layer by layer until all target network layers are compressed to obtain the compressed neural network model. It can be seen that this technical solution can perform a detailed split of the weight matrix of each target network layer in the model based on the importance of each weight parameter, realizing the refined management of the weight parameters of each target network layer. Moreover, by performing targeted compression processing on each weight parameter matrix obtained by splitting layer by layer, compared with the model compression scheme that adjusts multiple layers or all weight parameters at the same time, the refined processing of each weight parameter matrix is realized, avoiding the situation of damaging the inter-layer relationship, and being able to ensure that the model output has a small loss while increasing the model compression rate, making the model compression effect better and facilitating the deployment and use of the model on devices with weak storage and computing power.

[0128] The model compression device in the embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application does not make specific limitations.

[0129] The model compression device in the embodiment of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiment of the present application does not make specific limitations.

[0130] The model compression device provided in the embodiment of the present application can implement Figures 1 to 5 each process implemented in the method embodiment of , and for the sake of brevity, it will not be repeated here.

[0131] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, which is used to execute the above model compression method. Figure 7 FIG. 4 is a schematic structural diagram of an electronic device for implementing various embodiments of the present application. The electronic device may vary greatly due to configuration or performance differences, and may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call a computer program stored in the memory 730 and executable on the processor 710 to execute the following steps: Obtain the first weight matrix of each target network layer in the neural network model to be compressed; for each target network layer, according to the original input data and the first weight matrix of the target network layer, perform importance evaluation on each weight parameter in the first weight matrix to obtain an evaluation result; according to the evaluation result and a preset splitting strategy, split the first weight matrix into an important parameter matrix and an unimportant parameter matrix; starting from the first target network layer among each target network layer, layer-by-layer compress the neural network model to be compressed according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed to obtain a compressed neural network model.

[0132] In the embodiment of the present application, by obtaining the first weight matrix of each target network layer in the neural network model to be compressed, thus for each target network layer, according to the original input data and the first weight matrix of the target network layer, perform importance evaluation on each weight parameter in the first weight matrix to obtain an evaluation result, and according to the evaluation result and a preset splitting strategy, split the first weight matrix into an important parameter matrix and an unimportant parameter matrix. Furthermore, starting from the first target network layer among each target network layer, layer-by-layer compress the neural network model to be compressed according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed to obtain a compressed neural network model. It can be seen that this technical solution can perform a detailed split on the weight matrix of each target network layer in the model based on the importance of each weight parameter, realizing the refined management of the weight parameters of each target network layer. Moreover, by layer-by-layer performing targeted compression processing on the split weight parameter matrices, compared with the model compression scheme that adjusts multiple layers or all weight parameters simultaneously, it realizes the refined processing of each weight parameter matrix, avoids damaging the inter-layer relationship, can increase the model compression rate while ensuring that the model output has a small loss, making the model compression effect better and facilitating the deployment and use of the model on devices with weak storage and computing power.

[0133] The specific implementation steps can refer to the steps of the above embodiments of the model compression method, and the same technical effects can be achieved. To avoid repetition, they will not be elaborated here.

[0134] It should be noted that the electronic devices in the embodiments of the present application include: servers, terminals, or other devices other than terminals.

[0135] The above structure of the electronic device does not limit the electronic device. The electronic device may include more or fewer components than those shown, or combine some components, or have different component arrangements. For example, the input unit may include a Graphics Processing Unit (GPU) and a microphone, and the display unit may be configured with a display panel in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit includes at least one of a touch panel and other input devices. The touch panel is also called a touch screen. Other input devices may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0136] The memory can be used to store software programs and various data. The memory mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory can include volatile memory or non-volatile memory, or the memory can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synchlink DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM).

[0137] The processor can include one or more processing units; optionally, the processor integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor either.

[0138] The embodiments of the present application also provide a computer-readable storage medium, which is used to store computer-executable instructions. When the computer-executable instructions are executed by a processor, the various processes of the above-mentioned model compression method embodiments are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0139] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0140] The embodiments of the present application further provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements each process of the above-described embodiment of the model compression method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0141] The embodiments of the present application further provide a chip. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-described embodiment of the model compression method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0142] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0143] It should be noted that in this document, the term "including", "comprising", or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0145] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A model compression method, characterized in that, Including: Obtain the first weight matrix of each target network layer in the neural network model to be compressed; For each of the target network layers, perform importance evaluation on each weight parameter in the first weight matrix according to the original input data of the target network layer and the first weight matrix to obtain an evaluation result; According to the evaluation result and a preset splitting strategy, split the first weight matrix into an important parameter matrix and an unimportant parameter matrix; Starting from the first target network layer among the target network layers, compress the neural network model to be compressed layer by layer according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed to obtain a compressed neural network model.

2. The method according to claim 1, wherein The unimportant parameter matrix includes a pruning parameter matrix and a fine-tuning parameter matrix; The step of starting from the first target network layer among the target network layers and compressing the neural network model to be compressed layer by layer according to the important parameter matrix and the unimportant parameter matrix includes: For each of the target network layers, perform pruning processing on the target network layer according to the pruning parameter matrix to obtain a pruned target network layer; Determine the first pruning error corresponding to the target network layer according to the first output data of the pruned target network layer and the original output data of the target network layer; Determine the optimal fine-tuning parameter matrix corresponding to the target network layer according to the first pruning error; Perform quantization processing on the important parameter matrix and the optimal fine-tuning parameter matrix of the target network layer to obtain a second weight matrix corresponding to the target network layer.

3. The method according to claim 2, wherein When the target network layer is not the first target network layer, The step of determining the first pruning error corresponding to the target network layer according to the first output data of the pruned target network layer and the original output data of the target network layer includes: Determine the output data of the previous target network layer according to the input data of the previous target network layer and its corresponding second weight matrix; Determine the first output data of the pruned target network layer according to the output data of the previous target network layer, and the important parameter matrix and the fine-tuning parameter matrix of the target network layer; Calculate the loss value between the first output data and the original output data to obtain the first pruning error corresponding to the target network layer.

4. The method according to claim 2, wherein When the target network layer is the first target network layer, The step of determining the first pruning error corresponding to the target network layer according to the first output data of the pruned target network layer and the original output data of the target network layer includes: Determine the first output data of the pruned target network layer according to the original input data of the target network layer, and the important parameter matrix and the fine-tuning parameter matrix of the target network layer; Calculate the loss value between the first output data and the original output data to obtain the first pruning error corresponding to the target network layer.

5. The method according to claim 2, wherein The step of determining the optimal fine-tuning parameter matrix corresponding to the target network layer according to the first pruning error includes: When the first pruning error is greater than a preset error threshold, each fine-tuning parameter in the fine-tuning parameter matrix of the target network layer is iteratively adjusted according to a preset parameter adjustment strategy to obtain an adjusted fine-tuning parameter matrix; According to the input data of the target network layer, as well as the important parameter matrix and the adjusted fine-tuning parameter matrix of the target network layer, determine the second output data of the target network layer after pruning processing; According to the second output data and the original output data of the target network layer, determine the second pruning error corresponding to the target network layer; When the second pruning error satisfies the iteration termination condition, determine the currently adjusted fine-tuning parameter matrix as the optimal fine-tuning parameter matrix corresponding to the target network layer.

6. The method according to claim 1, wherein The importance evaluation of each weight parameter in the first weight matrix according to the original input data of the target network layer and the first weight matrix, and the obtained evaluation result includes: Calculate the mean value of the original input data of the target network layer on each column to obtain a mean value matrix corresponding to the original input data; According to the mean value matrix and the first weight matrix, determine an importance matrix corresponding to the first weight matrix; each element in the importance matrix is used to characterize the importance of the weight parameter at the corresponding position in the first weight matrix.

7. The method according to claim 1, characterized in that, The obtaining of the first weight matrix of each target network layer in the neural network model to be compressed includes: Obtain a validation data set; Input the validation data set into the neural network model to be compressed, and output the original input data and original output data of each network layer; For each target network layer, determine the first weight matrix of the target network layer according to the original input data and the original output data of the target network layer.

8. A model compression device, characterized in that, It includes: An acquisition module, configured to acquire the first weight matrix of each target network layer in the neural network model to be compressed; An evaluation module, configured to, for each target network layer, perform importance evaluation on each weight parameter in the first weight matrix according to the original input data of the target network layer and the first weight matrix to obtain an evaluation result; A splitting module, configured to split the first weight matrix into an important parameter matrix and an unimportant parameter matrix according to the evaluation result and a preset splitting strategy; A compression processing module, configured to start from the first target network layer among the target network layers, and perform layer-by-layer compression processing on the neural network model to be compressed according to the important parameter matrix and the unimportant parameter matrix until all target network layers are compressed to obtain a compressed neural network model.

9. An electronic device, characterized in that, It includes: A processor; And A memory arranged to store computer-executable instructions, the computer-executable instructions being configured to be executed by the processor, and the computer-executable instructions being executed by the processor to implement the model compression method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the model compression method according to any one of claims 1-7 is implemented.