Model lightweight method based on space-level low-rank decomposition distillation
By performing spatially-level low-rank decomposition and model distillation on the weight matrix of the deep learning model, the application problem of existing deep learning models on resource-constrained devices is solved, and the model is lightweight and performance maintenance is achieved.
Patent Information
- Application Number
- CN202411988200.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
When existing deep learning models are applied on resource-constrained devices, the application in scenarios such as mobile devices and edge computing is limited due to the huge amount of parameters, high computational complexity and slow inference speed.
The model lightweight method based on spatial-level low-rank decomposition is used to decompose the weight matrix of the teacher model through singular value decomposition technology, and the low-rank matrix representation is obtained, and it is distilled as a student model to transfer knowledge to the student model.
It effectively reduces the complexity of the model, maintains the model performance, improves the inference speed, and evaluates the accuracy, inference time and model size of the student model on the test set to achieve excellent performance indicators.
Smart Images

Figure BDA0005223067990000031 
Figure BDA0005223067990000061 
Figure BDA0005223067990000081
Abstract
Description
Technical Field
[0001] The present invention is applied to the fields of machine learning and deep learning, and specifically is a model lightweight method based on spatial-level low-rank decomposition distillation. Background Art
[0002] With the widespread application of deep learning, deep neural networks (DNNs) have achieved remarkable success in areas such as image recognition, natural language processing, and recommendation systems. However, these models usually have problems such as large number of parameters, high computational complexity, and slow inference speed, which limits their application on resource-constrained devices (such as mobile devices and edge computing). Therefore, researchers began to explore model compression and acceleration techniques, among which low-rank decomposition and model distillation are two important methods.
[0003] Low-rank decomposition effectively reduces the number of model parameters and computational complexity by decomposing a high-dimensional weight matrix into the product of multiple low-dimensional matrices. In the inference stage, the amount of multiplication calculations of low-rank matrices is relatively small, thereby improving the inference speed. However, excessive compression may lead to degradation of model performance, especially in complex tasks.
[0004] Model distillation is a method of knowledge transfer that improves the performance of a simple student model by transferring the knowledge of a complex teacher model to the simple student model. It can achieve better performance with fewer parameters by learning the output and intermediate features of the teacher model through the student model. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a model lightweight method based on spatial low-rank decomposition distillation in response to the shortcomings of the prior art.
[0006] In order to solve the above technical problems, the model lightweight method based on spatial low-rank decomposition distillation of the present invention comprises the following steps:
[0007] Select and train a teacher model;
[0008] Performing low-rank decomposition on the weight matrix of the teacher model, and obtaining a low-rank matrix representation by using a singular value decomposition technique;
[0009] The teacher model after low-rank decomposition is used as the student model, and the knowledge of the teacher model before low-rank decomposition is transferred to the student model through distillation technology.
[0010] As a possible implementation, further, the following steps are included:
[0011] Fine-tuning the student model to enhance its generalization ability and accuracy;
[0012] Evaluate the student model’s accuracy, inference time, and model size on the test set.
[0013] As a possible implementation method, further, the step of selecting and training the teacher model specifically includes:
[0014] Select ResNet, VGG, or BERT deep learning model as the teacher model;
[0015] Use standard data sets for training to ensure the accuracy and convergence of the teacher model. This process includes data preprocessing, model architecture definition, loss function selection, and optimization algorithm settings.
[0016] As a possible implementation, further, the weight matrix of the teacher model is subjected to low-rank decomposition, and the steps of obtaining a low-rank matrix representation by singular value decomposition technology specifically include:
[0017] Perform low-rank decomposition on the trained teacher model and use singular value decomposition to decompose the weight matrix to reduce the storage requirements of the model; select the first k singular values and construct a low-rank matrix to form a smaller weight representation;
[0018] The neural network is trained by the SVD form of the neural network, where each network layer is decomposed into two consecutive layers. Specifically, for the weight matrix W∈Rm×n, it is decomposed into U∈Rm×r, V∈Rn×r, where U and V are orthogonal matrices. In the case of full rank, they can be accurately reconstructed as:
[0019] W=Udiag(s)V T
[0020] For a neural network, this amounts to decomposing the weight matrix W into two consecutive layers:
[0021]
[0022] As a possible implementation, further, the weight matrix of the teacher model is subjected to low-rank decomposition, and the step of obtaining a low-rank matrix representation by singular value decomposition technology specifically includes:
[0023] Convolutional layer decomposition, convolution kernel K∈R n×c×h×w Represented as a four-dimensional tensor, where n, c, w, and h represent the number of filters, the number of input channels, and the width and height of the filter, respectively. Using spatial level decomposition, the steps are as follows:
[0024] Reshape the convolution kernel K into a two-dimensional matrix
[0025] Then By SVD we get U∈R n×r ,V∈R cwh×r ,S∈Rr , where U and V are orthogonal matrices, r = min{n,cwh};
[0026] The original convolutional layer is decomposed into two consecutive sub-convolutional layers: K1∈R r×c×w×h and K2∈R n×r×1×1 .
[0027] As a possible implementation, further, the weight matrix of each layer after the low-rank decomposition is converted into two consecutive neural network layers, using U and V as parameters for forward pass and back propagation respectively; the forward pass reduces the computational complexity by inputting the decomposed matrices U and V into two consecutive network layers for calculation, rather than directly using the original weight matrix W.
[0028] As a possible implementation, further, the step of using the teacher model after low-rank decomposition as the student model and transferring the knowledge of the teacher model before low-rank decomposition to the student model through the distillation technology specifically includes:
[0029] The distillation technique includes distilling the difference between the teacher model output before low-rank decomposition and the student model output after low-rank decomposition, and the distillation loss function is designed as:
[0030] L distill =αL output +(1-α)L feature
[0031] Among them, L output is the cross entropy loss between the student model output and the teacher model output, and L feature is the difference between the intermediate layer features of the student model and the teacher model.
[0032] As a possible implementation, further, L output is the cross entropy loss between the student model output and the teacher model output, and L feature is the difference between the intermediate layer features of the student model and the teacher model;
[0033] L output Using cross entropy loss: L feature Using feature loss:
[0034] The present invention adopts the above technical solution and has the following beneficial effects:
[0035] 1. Combination of low-rank decomposition and distillation: The present invention combines low-rank decomposition technology with model distillation to simultaneously reduce model complexity and maintain model performance.
[0036] 2. Based on spatial level decomposition, the pruned model can achieve similar accuracy to the original model.
[0037] 3. Feature matching distillation: The introduction of the distillation strategy of intermediate layer feature matching enables the student model to learn richer representation information from the teacher model, thereby improving the robustness of the model.
[0038] 4. Dynamic distillation weight adjustment: Through experiments, it is found that dynamically adjusting the weight parameter α in the distillation loss function can adaptively adjust according to the performance feedback during the training process and optimize the model training effect. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.
[0040] Example 1
[0041] The present invention provides a model lightweight method based on spatial low-rank decomposition distillation, comprising:
[0042] Select and train a teacher model;
[0043] Performing low-rank decomposition on the weight matrix of the teacher model, and obtaining a low-rank matrix representation by using a singular value decomposition technique;
[0044] The teacher model after low-rank decomposition is used as the student model, and the knowledge of the teacher model before low-rank decomposition is transferred to the student model through distillation technology.
[0045] Fine-tuning the student model to enhance its generalization ability and accuracy;
[0046] Evaluate the student model’s accuracy, inference time, and model size on the test set.
[0047] Among them, the steps of selecting and training the teacher model specifically include:
[0048] Select ResNet, VGG, or BERT deep learning model as the teacher model;
[0049] Use standard data sets for training to ensure the accuracy and convergence of the teacher model. This process includes data preprocessing, model architecture definition, loss function selection, and optimization algorithm settings.
[0050] The weight matrix of the teacher model is decomposed into low rank. The steps of obtaining the low rank matrix representation by singular value decomposition technology include:
[0051] Perform low-rank decomposition on the trained teacher model and use singular value decomposition to decompose the weight matrix to reduce the storage requirements of the model; select the first k singular values and construct a low-rank matrix to form a smaller weight representation;
[0052] The neural network is trained by the SVD form of the neural network, where each network layer is decomposed into two consecutive layers. Specifically, for the weight matrix W∈Rm×n, it is decomposed into U∈Rm×r, V∈Rn×r, where U and V are orthogonal matrices. In the case of full rank, they can be accurately reconstructed as:
[0053] W=Udiag(s)V T
[0054] For a neural network, this amounts to decomposing the weight matrix W into two consecutive layers:
[0055]
[0056] Among them, the weight matrix of the teacher model is decomposed into low rank, and the low rank matrix representation is obtained by singular value decomposition technology, and the specific steps include:
[0057] Convolutional layer decomposition, convolution kernel K∈R n×c×h×w Represented as a four-dimensional tensor, where n, c, w, and h represent the number of filters, the number of input channels, and the width and height of the filter, respectively. Using spatial level decomposition, the steps are as follows:
[0058] Reshape the convolution kernel K into a two-dimensional matrix
[0059] Then By SVD we get U∈R n×r ,V∈R cwh×r ,S∈R r , where U and V are orthogonal matrices, r = min{n,cwh};
[0060] The original convolutional layer is decomposed into two consecutive sub-convolutional layers: K1∈R r×c×w×h and K2∈R n×r×1×1 .
[0061] Among them, the weight matrix of each layer after low-rank decomposition is converted into two consecutive neural network layers, using U and V as parameters for forward pass and back propagation respectively; the forward pass reduces the computational complexity by inputting the decomposed matrices U and V into two consecutive network layers for calculation instead of directly using the original weight matrix W.
[0062] Among them, the teacher model after low-rank decomposition is used as the student model, and the knowledge of the teacher model before low-rank decomposition is transferred to the student model through distillation technology. The specific steps include:
[0063] The distillation technique involves distilling the difference between the teacher model output before low-rank decomposition and the student model output after low-rank decomposition. The distillation loss function is designed as:
[0064] L distill =αL output +(1-α)L feature
[0065] Among them, L output is the cross entropy loss between the student model output and the teacher model output, and L feature is the difference between the intermediate layer features of the student model and the teacher model.
[0066] L output is the cross entropy loss between the student model output and the teacher model output, and L feature is the difference between the intermediate layer features of the student model and the teacher model;
[0067] L output Using cross entropy loss: L feature Using feature loss:
[0068] Example 2
[0069] 1. Selection and training of teacher model:
[0070] Choose a high-performance deep learning model as the teacher model, such as ResNet, VGG, or BERT. These models perform well on corresponding tasks (such as image classification, text processing).
[0071] Use standard datasets (such as CIFAR-10, ImageNet) for training to ensure the accuracy and convergence of the teacher model. This process includes data preprocessing, model architecture definition, loss function selection, and optimization algorithm settings.
[0072] 2. Low rank decomposition processing:
[0073] Perform low-rank decomposition on the trained teacher model and use singular value decomposition (SVD) to decompose the weight matrix to reduce the storage requirements of the model. Select the first k singular values to construct a low-rank matrix to form a smaller weight representation. m×n Perform singular value decomposition and obtain three matrices W = U*S*V T , where U∈R m×r, V∈R n×r , and S is a diagonal matrix of singular values, r is a selected rank, and the low-rank matrix representation helps to reduce storage requirements and maintain accuracy close to the original model. The weight matrix of each layer after the low-rank decomposition is converted into two consecutive neural network layers, using U and V as parameters for forward and back propagation respectively;
[0074] The neural network is trained through the SVD form of the neural network, where each network layer is decomposed into two consecutive layers (there is no additional operation between them). Specifically, for the weight matrix W∈Rm×n, it can be decomposed into three parts: U∈Rm×r and V∈Rn×r, where U and V are orthogonal matrices. In the case of full rank, they can be accurately reconstructed as:
[0075] W=Udiag(s)V T
[0076] For a neural network, this amounts to decomposing the weight matrix W into two consecutive layers:
[0077]
[0078] For the convolution layer, the convolution kernel K∈R n×c×h×w It can be represented as a four-dimensional tensor, where n, c, w, and h represent the number of filters, the number of input channels, and the width and height of the filter, respectively. This paper uses spatial decomposition. This full-rank decomposition model can often achieve similar accuracy to the original model. The steps are as follows:
[0079] First, reshape the convolution kernel K into a two-dimensional matrix
[0080] Then By SVD we get U∈R n×r ,V∈R cwh×r ,S∈R r , where U and V are orthogonal matrices, r = min{n,cwh}.
[0081] At this time, the original convolution layer is decomposed into two consecutive sub-convolution layers: K1∈R r×c×w×h and K2∈R n ×r×1×1 .
[0082] In SVD training, each layer uses the decomposed variables, namely U, s, V, rather than the original convolution kernel K or weight matrix W. The forward pass is to convert U, s, V into two consecutive network layers, and backpropagation and optimization are directly applied to U, s, V. In this way, we can obtain s without performing SVD at each step of training.
[0083] 3. Construction and distillation of student model:
[0084] The model after SVD is used as the student model, and its structure is simplified, and the model before SVD is used as the teacher model. During the distillation process, the knowledge of the teacher model can guide the student model to learn intermediate features, so that the accuracy loss can be well controlled.
[0085] The distillation loss function is designed as:
[0086] L distill =αL output +(1-α)L feature
[0087] Among them, L output is the cross entropy loss between the student model output and the teacher model output, and L feature is the difference between the intermediate layer features of the student model and the teacher model.
[0088] L output Using cross entropy loss: L feature Using feature loss:
[0089] 4. Fine-tuning and optimization of the model:
[0090] The student model is fine-tuned and trained jointly using the original training data and the distilled data to enhance the generalization ability and accuracy of the model.
[0091] 5. Performance evaluation:
[0092] Evaluate the performance of the student model on the test set:
[0093] Accuracy: Required to reach more than 90%.
[0094] Reasoning time: The reasoning time of the student model should be controlled within 60% of the teacher model.
[0095] Model size: Student model parameters should be reduced by more than 30%.
[0096] The above are embodiments of the present invention. For ordinary technicians in this field, according to the teachings of the present invention, all equivalent changes, modifications, substitutions and variations made within the scope of the patent application of the present invention without departing from the principles and spirit of the present invention should fall within the scope of the present invention.
Claims
1. A model lightweight method based on spatial low-rank decomposition distillation, characterized in that: The steps include: Select and train a teacher model; Performing low-rank decomposition on the weight matrix of the teacher model, and obtaining a low-rank matrix representation by using a singular value decomposition technique; The teacher model after low-rank decomposition is used as the student model, and the knowledge of the teacher model before low-rank decomposition is transferred to the student model through distillation technology.
2. The model lightweight method based on spatial low-rank decomposition distillation according to claim 1, characterized in that: The following steps are also included: Fine-tuning the student model to enhance its generalization ability and accuracy; Evaluate the student model’s accuracy, inference time, and model size on the test set.
3. The model lightweight method based on spatial low-rank decomposition distillation according to claim 1, characterized in that: The step of selecting and training the teacher model specifically includes: Select ResNet, VGG, or BERT deep learning model as the teacher model; Use standard data sets for training to ensure the accuracy and convergence of the teacher model. This process includes data preprocessing, model architecture definition, loss function selection, and optimization algorithm settings.
4. The model lightweight method based on spatial low-rank decomposition distillation according to claim 1, characterized in that: The step of performing low-rank decomposition on the weight matrix of the teacher model and obtaining a low-rank matrix representation by using a singular value decomposition technique specifically includes: Perform low-rank decomposition on the trained teacher model and use singular value decomposition to decompose the weight matrix to reduce the storage requirements of the model; select the first k singular values and construct a low-rank matrix to form a smaller weight representation; The neural network is trained by the SVD form of the neural network, where each network layer is decomposed into two consecutive layers. Specifically, for the weight matrix W∈Rm×n, it is decomposed into U∈Rm×r, V∈Rn×r, where U and V are orthogonal matrices. In the case of full rank, they can be accurately reconstructed as: In=Udiag(s)V T For a neural network, this amounts to decomposing the weight matrix W into two consecutive layers:
5. The model lightweight method based on spatial low-rank decomposition distillation according to claim 4, characterized in that: The step of performing low-rank decomposition on the weight matrix of the teacher model and obtaining a low-rank matrix representation through a singular value decomposition technique specifically includes: Convolutional layer decomposition, convolution kernel K∈R n×c×h×w Represented as a four-dimensional tensor, where n, c, w, and h represent the number of filters, the number of input channels, and the width and height of the filter, respectively. Using spatial level decomposition, the steps are as follows: Reshape the convolution kernel K into a two-dimensional matrix Then By SVD we get U∈R n×r ,V∈R cwh×r ,S∈R r , where U and V are orthogonal matrices, r = min{n,cwh}; The original convolutional layer is decomposed into two consecutive sub-convolutional layers: K1∈R r×c×w×h and K2∈R n×r×1×1 .
6. The model lightweight method based on spatial low-rank decomposition distillation according to claim 5, characterized in that: The weight matrix of each layer after the low-rank decomposition is converted into two consecutive neural network layers, using U and V as parameters for forward pass and back propagation respectively; the forward pass reduces the computational complexity by inputting the decomposed matrices U and V into two consecutive network layers for calculation instead of directly using the original weight matrix W.
7. The model lightweight method based on spatial low-rank decomposition distillation according to claim 1, characterized in that: The steps of using the teacher model after low-rank decomposition as the student model and transferring the knowledge of the teacher model before low-rank decomposition to the student model through the distillation technology specifically include: The distillation technique includes distilling the difference between the teacher model output before low-rank decomposition and the student model output after low-rank decomposition, and the distillation loss function is designed as: L distill =αL output +(1-α)L feature Among them, L output is the cross entropy loss between the student model output and the teacher model output, and L feature is the difference between the intermediate layer features of the student model and the teacher model.
8. The model lightweight method based on spatial low-rank decomposition distillation according to claim 7, characterized in that: The L output is the cross entropy loss between the student model output and the teacher model output, and L feature is the difference between the intermediate layer features of the student model and the teacher model; L output Using cross entropy loss: L feature Using feature loss: