Network model lightweight method based on tensor decomposition
By tucker decomposing and quantizing the subnet structure of the pre-trained neural network model, the problem of failure to fully utilize inter-layer commonality in the existing technology is solved, and efficient compression and computational optimization of the neural network model is realized.
Patent Information
- Application Number
- CN202510787074.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing neural networks with subnet structures, the prior art fails to fully utilize the inherent commonality between the layers, resulting in limited compression performance and high computational complexity, making it difficult to effectively deploy on resource-constrained devices.
Using a tensor decomposition method, the subnet structure in the pre-trained neural network model is jointly decomposed through Tucker decomposition, and the shared core tensor and factor matrix are obtained, and combined with quantization technology, parameter compression and calculation complexity are reduced.
While maintaining network accuracy, the compression performance of neural network models is significantly improved, the computational complexity is reduced, the compression rate is improved, and the memory usage and computing costs are reduced.
Smart Images

Figure CN120297337A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network model parameter compression, and specifically to a method for lightweighting a network model based on tensor decomposition. Background Art
[0002] With the continuous expansion of the dataset scale and the improvement of task complexity, deep structured models have become increasingly important in effectively capturing complex patterns and features in data. The rapid development of deep learning has spawned a variety of network architectures, which are usually constructed based on modular design principles, making the network more flexible and scalable. In this design, some network layers may exhibit the same or similar tensor structures, or even share the same parameters, and this same or similar structure is called a subnet structure. These deep neural networks have made remarkable progress in multiple fields and have been widely applied. However, their huge number of parameters and high computational requirements pose many challenges when effectively deployed on resource-constrained devices (such as traditional desktop computers and portable devices). Therefore, reducing the number of parameters and computational complexity is crucial for the practical application of neural networks.
[0003] Based on the technology of low-rank decomposition, the parameters of a neural network can be transformed into smaller matrices or tensors through matrix or tensor decomposition, thereby significantly reducing its storage and computational overhead without significantly affecting the model accuracy. As a high-order extension of a second-order matrix, a tensor can more effectively capture the complex high-order information hidden in the parameters, making the algorithm based on tensor decomposition expected to outperform the traditional matrix decomposition technology in terms of performance.
[0004] For example, the invention patent with the publication number CN116542315A discloses a method and system for compressing parameters of a large-scale neural network based on tensor decomposition. The method includes the following steps: obtaining the weight matrix in the trained large-scale neural network, and setting the row rank and column rank of the weight matrix; according to the row rank and column rank, selecting a set of row number sequences and a set of column number sequences in the weight matrix; performing CUR decomposition on the weight matrix based on the selected row number sequences and column number sequences to obtain a compressed weight matrix, and replacing the weight matrix with the compressed weight matrix; setting a loss function to adjust the compressed weight matrix to obtain an adjusted compressed weight matrix, comparing the sizes of the row rank and column rank, and simplifying the adjusted compressed weight matrix according to the comparison result to achieve the compression of the parameters of the large-scale neural network.
[0005] Combining the above technical solutions, it is found that neural networks with subnet structures exhibit unique advantages in inter-layer correlation and shared features. Although each layer has its independent attributes, there are certain commonalities among them. Current network optimization methods often regard each layer as an independent part and perform separate compression on it. Although this method is effective to some extent, when dealing with neural network models with specific subnet structures, its compression performance is still limited because it fails to fully utilize the inherent commonalities between layers. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides a method for lightweighting a network model based on tensor decomposition, which can effectively solve the problems involved in the above background technology.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for lightweighting a network model based on tensor decomposition, including: obtaining the convolutional kernel parameters in each subnet structure of a pre-trained neural network model, setting the rank of joint decomposition for each layer in each subnet structure of the network model, performing joint decomposition on each layer in the subnet structure of the network model using Tucker decomposition according to the set decomposition rank to obtain compressed convolutional kernels, setting an optimization function to adjust the parameters of the compressed convolutional kernels, and using the adjusted compressed convolutional kernels to replace the original convolutional kernels in the network, performing pseudo-quantization processing on the adjusted compressed convolutional kernels, and updating the parameters according to the gradients in the backpropagation calculation to obtain quantized parameters, replacing the adjusted compressed parameters, and realizing parameter compression of the neural network model. The whole process aims to explore the commonalities of subnet structures in neural networks, and greatly improve the compression performance and reduce the computational complexity of such modular-designed neural network models while maintaining a certain network accuracy.
[0008] As a further method, the process of obtaining the convolutional kernel parameters in each subnet structure of the pre-trained neural network model is as follows: Obtaining the convolutional kernel parameters in each subnet structure of the pre-trained neural network model specifically involves traversing all subnet structures in the pre-trained model (such as residual blocks in ResNet) and extracting the four-dimensional tensors of each convolutional layer , where I and O are the number of input / output channels respectively, K is the convolutional kernel size, and n and m respectively represent the convolutional kernel parameters of the m-th layer in the n-th subnet.
[0009] As a further method, the rank of the set joint decomposition is specifically analyzed as follows: Obtain the Tucker rank of the joint decomposition of each layer in the subnet structure, so as to achieve shared compression. Specifically, it includes separately performing variational Bayesian matrix factorization (VBMF) on the weight tensors of each convolutional layer in the subnet structure to obtain the ranks of the separate decompositions of each layer; after obtaining the separate ranks of each layer, obtain the mean value of the decomposition ranks of each layer within the same subnet; use the mean value of the ranks comprehensively calculated for each layer within the same subnet as the rank of the joint decomposition within the subnet.
[0010] As a further method, the Tucker decomposition is used for joint decomposition. The specific analysis process is as follows: Use the rank of the joint decomposition for each layer within the same subnet to perform Tucker decomposition to obtain the corresponding core tensor and factor matrices; retain the core tensor obtained from the decomposition of the first layer within the same subnet to represent the common properties of the same subnet, and at the same time, the factor matrices obtained from the decomposition of each layer are used as the independent characteristics representing each layer within the subnet; set an optimization function to adjust the above-shared core factor tensor and the parameter of each layer's factor matrix; use the adjusted compressed convolutional kernel to replace the original convolutional kernel in the network.
[0011] As a further method, the further quantization of the compressed parameters includes: Using symmetric quantization to quantize the convolutional layer parameters represented by 32-bit floating-point numbers to obtain the convolutional layer parameters represented by quantized int8; using the convolutional layer parameters represented by quantized int8 for the forward inference of the network, thereby reducing memory usage and computational cost; when performing backpropagation gradient, dequantize the quantized int8 convolutional kernel parameters to correctly calculate the gradient and update the weights.
[0012] Compared with the prior art, the embodiments of the present invention at least have the following advantages or beneficial effects: (1) By providing a method for lightweighting a network model based on tensor decomposition, the present invention obtains the parameters of each convolutional kernel within the subnet structure of a pre-trained neural network model, sets the rank of the joint decomposition of each layer in each subnet structure of the network model, and according to the set decomposition rank, uses Tucker decomposition to perform joint decomposition on each layer in the subnet structure of the network model to obtain compressed convolutional kernels, sets an optimization function to adjust the parameters of the compressed convolutional kernels, and uses the adjusted compressed convolutional kernels to replace the original convolutional kernels in the network, performs pseudo-quantization processing on the adjusted compressed convolutional kernels, and updates the parameters according to the gradients in the reverse calculation to obtain quantized parameters and replace the adjusted compressed parameters, thereby realizing the parameter compression of the neural network model. The whole process aims to explore the commonalities of the subnet structure in the neural network, and significantly improve the compression performance of such modular-designed neural network models and reduce the computational complexity while maintaining a certain network accuracy; (2) The present invention deeply analyzes and researches the subnet structure in the neural network model. To further improve the compression rate of the neural network with such a structure, a sharing algorithm based on Tucker decomposition is developed. This algorithm represents the common components through the core tensor and uses the factor matrices to represent the independent characteristics of each layer within the subnet. This method can effectively extract the common information between layers in each subnet, thus significantly improving the compression rate. (3) By introducing quantization technology, the present invention can further improve the compression effect on the basis of achieving shared compression by Tucker decomposition. By sequentially applying low-rank representation and quantization in two different stages, the orthogonal combination of these two technologies is explored, realizing a double compression process. Brief Description of the Drawings
[0013] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the following drawings. Figure 1 It is a schematic flowchart of the method of the present invention. Figure 2 It is a schematic connection diagram of the system modules of the present invention. Specific Embodiments
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0015] Refer to Figure 1 As shown, the present invention provides a method for lightweighting a network model based on tensor decomposition, including: S1. Obtain the convolutional kernel parameters within the subnet structure in the pre-trained neural network model: Traverse all the subnet structures in the pre-trained model (such as the residual blocks in ResNet), extract the four-dimensional tensors of each convolutional layer, so as to achieve parameter compression according to these convolutional kernels.
[0016] It should be elaborated that a neural network with a subnet structure usually consists of multiple basic blocks, and each block contains multiple layers inside. Assume that the neural network is composed of basic blocks, and similar or identical convolutional kernels are applied in each block to extract features through convolutional operations. For the subnet structure, for the layer, its convolutional kernel parameters are represented as , where I and O are the number of input / output channels respectively, K is the convolution kernel size, and n and m represent the convolution kernel parameters of the m-th layer in the n-th subnet respectively.
[0017] The above-mentioned acquisition of the convolution kernel parameters in the neural network model can be achieved through the interfaces provided by open-source learning frameworks (such as PyTorch) (such as conv_layer.weight.data).
[0018] S2. Set the rank of joint decomposition: Based on variational Bayesian matrix factorization (VBMF), separately obtain the rank of each layer's separate decomposition for the weight tensor of each convolution layer in the subnet structure. After obtaining the separate rank of each layer, obtain the average value of the decomposition ranks of each layer within the same subnet, and use the average value of the ranks obtained by comprehensively calculating each layer within the same subnet as the rank of joint decomposition within the subnet.
[0019] It should be explained that in order to obtain the rank of joint decomposition of each layer within the subnet structure, it is first necessary to perform separate decomposition on each layer. Taking the in the subnet structure convolution kernel parameters of the layer as an example, where respectively represent the spatial dimension, the number of input channels, and the number of output channels. Generally speaking, the value of the spatial dimension is relatively small (such as or ), so it is not processed during parameter compression. That is to say, only the input channels and output channels of the convolution kernel are compressed. Applying VBMF to perform separate decomposition on each layer requires reshaping the convolution kernel parameters.
[0020] First, expand the convolution kernel parameters along the third and fourth dimensions into matrices, which are respectively represented as matrices and with the shape sizes of . Subsequently, perform the following operations on each expanded matrix : Model setting: Assume the expanded matrix , where is the factor matrix, is the noise matrix, and the number of columns of the matrix is the rank to be solved; Prior distribution: Assign an automatic relevance determination (ARD) prior to the elements of , such as a Gaussian distribution, and the precision parameter controls the sparsity of the factors; Variational inference: Optimize the variational lower bound (ELBO) to infer the posterior distribution. During the hyperparameter optimization process, the precisions of unimportant factor columns approach zero and are automatically pruned; Rank determination: The number of columns corresponding to non-zero precision is the rank of the unfolded matrix. ; After the above operations are completed, collect the ranks of all modalities as the ranks for Tucker decomposition.
[0021] It should be explained that the above operations are completed on the matrix unfolded along the third and fourth dimensions of the convolution kernel parameters , so two ranks will be generated. The implementation of the above VBMF algorithm can use ready-made libraries (such as nimfa) in deep learning frameworks or be written by oneself.
[0022] It should be explained that in order to achieve shared compression, it is necessary to ensure that the core tensors in the same sub-network have the same size. This means that their Tucker ranks must be the same. That is to say, after obtaining the individual ranks of each layer, further, take the average of the ranks of all layers within the same sub-network structure to obtain the rank for joint decomposition within this sub-network. Thus, the core tensors in the same sub-network have the same size.
[0023] S3. Perform joint decomposition using Tucker decomposition: Perform joint decomposition on each layer within the same sub-network structure to obtain a shared core factor tensor and factor matrices representing the independent characteristics of each layer, thereby achieving shared compression for each layer within the sub-network structure.
[0024] It should be explained that the above joint decomposition of each layer within the same sub-network structure to obtain a shared core factor tensor and factor matrices representing the independent characteristics of each layer is a crucial step in achieving shared compression based on Tucker decomposition.
[0025] Specifically, the detailed analysis process of the joint decomposition of each layer within the same sub-network structure is as follows: Represent the shared core tensor of the nth sub-network as , and represent the factor matrices obtained from the decomposition of the mth convolution kernel tensor in the nth sub-network as and . To obtain the shared core tensor and the corresponding set of factor matrices and , the network layers under the same sub-network n will be jointly decomposed, and the joint decomposition problem is defined as: ; In the above problem, the initialization of the shared core tensor is obtained by Tucker decomposition using the rank for joint decomposition of the first layer within each sub-network structure, and the corresponding factor matrices and The initialization is obtained by Tucker decomposition of the ranks jointly decomposed by each layer. After initialization, Stochastic Gradient Descent (SGD) is used to update .
[0026] For a neural network model composed of basic blocks, with similar or identical convolutional kernels applied within each block, based on the above joint decomposition, the kernel tensor of each layer can be expressed as: ; where, is the core tensor shared by multiple layers in the th subnet structure, and is a set of factor matrices independent for each layer.
[0027] Before implementing the decomposition, the convolution operation of the m-th convolutional layer in the n-th subnet can be expressed by the following formula: ; where, and represent the input and output tensors respectively, while represent the spatial dimension, input channels, and output channels respectively.
[0028] After implementing the joint decomposition and representing the parameters of the convolutional kernel by the shared core factor tensor and factor matrices, the convolution operation of this convolutional layer is expressed as: ; That is to say, the original convolution operation is divided into three consecutive sub-operations after decomposition, which are respectively expressed as: ; ; ; where , are the input tensor and output tensor, is the intermediate tensor, is the shared core tensor, and is the factor matrix.
[0029] After shared compression, the compression ratio of the th subnet structure is expressed as .
[0030] S4. Further quantization of the compressed parameters: Further quantization is implemented for the shared core factor tensor and factor matrices obtained after jointly compressing each layer within the same subnet structure.
[0031] It should be noted that the shared core factor tensor and factor matrix obtained after the joint compression of each layer within the above-mentioned same subnet structure are stored in the form of 32-bit floating-point numbers, which not only occupy a large amount of memory but also require a large amount of computational cost for operations. The further quantization refers to quantizing the shared core factor tensor and factor matrix obtained after the joint compression of each layer within the same subnet structure represented in the form of 32-bit floating-point numbers into 8-bit integer form using symmetric quantization.
[0032] Specifically, for a 32-bit floating-point parameter , its quantization mathematical formula is expressed as: ; where is the quantization scale factor, is the zero-point offset, and the function rounds the value to the nearest integer within the quantization range. Using symmetric quantization is to process the weight parameters in the neural network. Symmetric quantization keeps the zero points consistent before and after quantization, so there is no need to introduce an offset, that is, .
[0033] In symmetric quantization, the calculation method of the quantization scale factor is: ; where represents the maximum value of the floating-point real number to be quantized, and is 127 in 8-bit integer quantization.
[0034] It should be noted that during the quantization process, the rounding function is usually used to convert floating-point numbers to integers. However, the derivative or gradient of the rounding function is zero at the quantization points. This means that when performing backpropagation to update the weights, the gradient will disappear, hindering effective learning and leading to performance degradation.
[0035] Furthermore, to solve the problem that the gradient disappears when performing backpropagation to update the weights, a method based on quantization-aware training (QAT) is used to quantize the parameters, which involves pseudo-quantization operations. Taking the weight tensor of two convolutional layers as an example, the specific implementation is as follows: First, quantize the floating-point weights of the convolutional layer to obtain the quantized int8 weights . Then, use this to perform the forward inference of the network, thereby reducing memory usage and computational cost. To train , it needs to be de-quantized to obtain the floating-point weights , during which the gradient can be correctly calculated and used for weight update.
[0036] Refer toFigure 2 As shown in Figure 2 , the second aspect of the present invention provides a system for lightweighting a network model based on tensor decomposition, including: a parameter acquisition module, a joint decomposition rank setting module, a Tucker joint decomposition module, and a quantization module.
[0037] The parameter acquisition module is connected to the joint decomposition rank setting module, the joint decomposition rank setting module is connected to the Tucker joint decomposition module, and the Tucker joint decomposition module is connected to the quantization module.
[0038] The parameter acquisition module is used to obtain the convolution kernel parameters in each subnet structure of the pre-trained neural network model, traverse all subnet structures in the pre-trained model (such as the residual blocks in ResNet), and extract the four-dimensional tensors of each convolutional layer , where I and O are the number of input / output channels respectively, K is the convolution kernel size, and n and m respectively represent the convolution kernel parameters of the m-th layer in the n-th subnet.
[0039] The joint decomposition rank setting module is used to obtain the Tucker rank of joint decomposition for each layer in the subnet structure, so as to achieve shared compression. Specifically, it includes separately performing variational Bayesian matrix factorization (VBMF) on the weight tensors of each convolutional layer in the subnet structure to obtain the ranks of individual decompositions for each layer; after obtaining the individual ranks of each layer, obtaining the mean of the decomposition ranks of each layer within the same subnet; and taking the mean of the ranks comprehensively calculated for each layer within the same subnet as the rank of joint decomposition within the subnet.
[0040] The Tucker joint decomposition module is used to perform joint decomposition using Tucker decomposition, perform Tucker decomposition on each layer within the same subnet using the rank of joint decomposition to obtain the corresponding core tensor and factor matrices; retain the core tensor obtained from the decomposition of the first layer within the same subnet to represent the common properties of the same subnet, and at the same time, the factor matrices obtained from the decomposition of each layer are used to represent the independent characteristics of each layer within the subnet; set an optimization function to adjust the above-shared core factor tensor and the factor matrix parameters of each layer; use the adjusted compressed convolution kernel to replace the original convolution kernel in the network.
[0041] The quantization module is used to quantize the convolutional layer parameters represented by 32-bit floating-point numbers using symmetric quantization to obtain the convolutional layer parameters represented by quantized int8; use the convolutional layer parameters represented by quantized int8 for forward inference of the network, thereby reducing memory usage and computational cost; when performing backpropagation gradients, dequantize the quantized int8 convolution kernel parameters to correctly calculate the gradients and update the weights.
Claims
1. A method for lightweighting a network model based on tensor decomposition, characterized in that Including: Obtain the parameters of each convolutional kernel within the subnet structure of the pre-trained neural network model; Set the rank of joint decomposition: According to the convolutional kernel features of each network layer within the neural network subnet structures, set the rank of joint decomposition for each layer; Perform joint decomposition using Tucker decomposition: According to the set decomposition rank, perform Tucker decomposition on each layer in the subnet structure of the network model to obtain compressed convolutional kernels, set an optimization function, adjust the parameters of the compressed convolutional kernels, and replace the original convolutional kernels in the network with the adjusted compressed convolutional kernels; Further quantize the compressed parameters: Perform pseudo-quantization processing on the adjusted compressed convolutional kernels, update the parameters according to the gradients in the backpropagation calculation, obtain the quantized parameters, and replace the adjusted compressed parameters to achieve parameter compression of the neural network model.
2. The lightweight method of a network model based on tensor decomposition according to claim 1, wherein: The pre-trained deep neural network with a subnet structure usually consists of N basic blocks. M similar or identical convolutional kernels are applied within each block to extract features through convolutional operations; Each block contains multiple layers. Taking the block as a unit, obtain the parameters of each convolutional kernel within the subnet structure of the pre-trained neural network model.
3. The lightweight method of a network model based on tensor decomposition according to claim 2, wherein: The specific analysis process for setting the rank of joint decomposition of each layer within the subnet structure of the neural network model is as follows: Perform variational Bayesian matrix factorization (VBMF) based on the obtained parameters of each convolutional kernel within the subnet structure of the neural network model to obtain the rank of individual decomposition for each layer; Analyze and process the ranks of each layer within each subnet to obtain the rank of joint decomposition for each layer within each subnet.
4. The lightweight method for a network model based on tensor decomposition according to claim 3, wherein: According to the set decomposition rank, perform Tucker decomposition on each layer in the subnet structure of the network model to obtain compressed convolutional kernels, set an optimization function, adjust the parameters of the compressed convolutional kernels, and replace the original convolutional kernels in the network with the adjusted compressed convolutional kernels.
5. A lightweight method for a network model based on tensor decomposition according to claim 1, characterized in that: Perform pseudo-quantization processing on the adjusted compressed convolutional kernels, update the parameters according to the gradients in the backpropagation calculation, obtain the quantized parameters, and replace the adjusted compressed parameters to achieve parameter compression of the neural network model.
Citation Information
Patent Citations
Large-scale neural network parameter compression method and system based on tensor decomposition
CN116542315A