Convolutional neural network weight compression method and device
By converting the weight of the convolutional neural network from a four-dimensional tensor to a blocked inverse triangle form, the problem of excessive demand for the computing resources of the convolutional neural network is solved, and the compression of the model and the improvement of the computing efficiency is achieved.
Patent Information
- Application Number
- CN202411012733.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively compress convolutional neural networks, resulting in too large demand for computing resources and difficult to meet performance requirements.
The weight of the convolutional neural network is converted from a four-dimensional tensor to a two-dimensional weight matrix, and converted into a blocked inverse triangle form through improved orthogonal transformation. First-order correction is performed on the symmetric indefinite matrix with large computational volume, and reshape it into a new four-dimensional tensor to achieve compression.
The number of multiplication operations of convolution operations is reduced, the calculation burden is reduced, the stability of numerical calculations is improved, and the model size and run time are optimized.
Smart Images

Figure CN120068976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and device for compressing the weights of a convolutional neural network. Background Art
[0002] With the continuous progress of deep learning technology, the Convolutional Neural Network (CNN) has become an indispensable tool in computer vision, natural language processing, and many other fields. The powerful ability of CNN stems from its hierarchical structure, which enables it to capture complex data features and patterns. However, behind this ability is a huge demand for computing resources, especially when dealing with large-scale datasets and building complex models. With the explosive growth of data volume and the continuous increase in model complexity, the existing processing capabilities and algorithm efficiencies have been difficult to meet the increasingly stringent performance requirements.
[0003] Therefore, how to effectively compress the convolutional neural network has become an urgent problem to be solved in the industry. Summary of the Invention
[0004] The present invention provides a method and device for compressing the weights of a convolutional neural network to solve the problem in the prior art that how to effectively compress the convolutional neural network has become an urgent problem to be solved in the industry.
[0005] The present invention provides a method for compressing the weights of a convolutional neural network, including:
[0006] Converting the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix;
[0007] Converting the two-dimensional weight matrix into a block anti-triangular form through an improved orthogonal transformation, where the improved orthogonal transformation includes at least one of the following: Givens transformation and Householder transformation;
[0008] Calculating the amount of computation for each conversion process based on the number of floating-point operations required by the block anti-triangular algorithm, and performing a first-order correction on the symmetric indefinite matrix whose amount of computation exceeds a preset threshold to obtain a corrected block anti-triangular form;
[0009] Reshaping the lower triangular block matrix of the corrected block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replacing the original four-dimensional tensor with the new four-dimensional tensor to obtain a compressed convolutional neural network.
[0010] The present invention provides a method for compressing the weights of a convolutional neural network, which converts the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix, including:
[0011] Obtain the original four-dimensional tensor of the convolutional layer weights, where the original four-dimensional tensor includes: the number of output filters, the number of input filters, the filter height, and the filter width;
[0012] For each output filter, expand the convolution kernels of all input filters corresponding to the output filter into a one-dimensional vector in depth, and stack the one-dimensional vectors of all output filters to form a two-dimensional weight matrix;
[0013] Among them, the number of rows of the two-dimensional weight matrix is equal to the number of output filters, and the number of columns of the two-dimensional weight matrix is equal to the number of input filters multiplied by the total number of elements of each convolution kernel.
[0014] The present invention provides a method for compressing the weights of a convolutional neural network. Based on the number of floating-point operations required by the block anti-triangular algorithm, calculate the amount of computation for each conversion process, specifically including:
[0015] Obtain the basic data operation categories corresponding to each orthogonal transformation in the block anti-triangular algorithm, and the number of basic data operations corresponding to each orthogonal transformation;
[0016] Based on the first floating-point operation numbers corresponding to each basic data operation, calculate the second floating-point operation number for each orthogonal transformation, and accumulate the second floating-point operation numbers to obtain the third floating-point operation number for the orthogonal transformation in the block anti-triangular algorithm;
[0017] According to the third floating-point operation number, determine the amount of computation for each transformation process in the block anti-triangular algorithm.
[0018] The present invention provides a method for compressing the weights of a convolutional neural network. The specific process of the first-order correction includes:
[0019] Identify the matrix characteristics of a symmetric indefinite matrix whose amount of computation exceeds a preset threshold, and determine the correction target of the symmetric indefinite matrix;
[0020] According to the correction target, perform cyclic correction on the symmetric indefinite matrix until the amount of computation after correction is less than the preset threshold to obtain the corrected symmetric indefinite matrix;
[0021] Convert the corrected symmetric indefinite matrix into a block anti-triangular form to obtain the corrected block anti-triangular form.
[0022] The present invention provides a method for compressing the weights of a convolutional neural network. The steps of reshaping the lower triangular block matrix of the corrected block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replacing the original four-dimensional tensor with the new four-dimensional tensor to obtain a compressed convolutional neural network include:
[0023] Reshape each lower triangular block in the corrected block anti-triangular matrix into the shape of the original convolution kernel one by one, and recombine them to form a new four-dimensional weight tensor;
[0024] Integrate the new four-dimensional weight tensor into the convolutional neural network to replace the original weight tensor of the corresponding layer, so as to obtain a compressed convolutional neural network.
[0025] The present invention also provides a convolutional neural network weight compression device, including:
[0026] A first conversion module for converting the weight of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix;
[0027] A second conversion module for converting the two-dimensional weight matrix into a block anti-triangular form through an improved orthogonal transformation, where the improved orthogonal transformation includes at least one of the following: Givens transformation and Householder transformation;
[0028] A calculation module for calculating the amount of computation of each conversion process based on the number of floating-point operations required by the block anti-triangular algorithm, and performing a first-order correction on the symmetric indefinite matrix whose amount of computation exceeds the preset threshold to obtain a corrected block anti-triangular form;
[0029] A replacement module for reshaping the lower triangular block matrix in the corrected block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replacing the original four-dimensional tensor according to the new four-dimensional tensor to obtain a compressed convolutional neural network.
[0030] The device is also used for:
[0031] Obtain the original four-dimensional tensor of the convolutional layer weight, where the original four-dimensional tensor includes: the number of output filters, the number of input filters, the filter height, and the filter width;
[0032] For each output filter, expand the convolution kernels of all input filters corresponding to the output filter into one-dimensional vectors in depth, and stack the one-dimensional vectors of all output filters to form a two-dimensional weight matrix;
[0033] Wherein, the number of rows of the two-dimensional weight matrix is equal to the number of output filters, and the number of columns of the two-dimensional weight matrix is equal to the number of input filters multiplied by the total number of elements of each convolution kernel.
[0034] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the convolutional neural network weight compression method as described in any one of the above.
[0035] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the convolutional neural network weight compression method described in any one of the above is implemented.
[0036] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the convolutional neural network weight compression method described in any one of the above is implemented.
[0037] For the convolutional neural network weight compression method and device provided by the present invention, by converting the weights from a four-dimensional tensor to a block anti-triangular form, the number of parameters of the network is reduced, because a low-rank matrix can approximate the original weight matrix with fewer parameters. The compression of the weight matrix reduces the number of multiplication operations required for the convolution operation, thereby reducing the computational burden of the model in forward and backward propagation. The first-order correction adjusts the symmetric indefinite matrix with a large amount of computation, improves the condition number of the matrix, and thus improves the stability of numerical calculation. The present application provides various variation methods such as Givens and Householder in the block anti-triangular algorithm and comparison data of computational complexity, which can help select different variation forms for different matrices and reduce the algorithm duration. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart of the convolutional neural network weight compression method provided by the embodiment of the present application;
[0040] Figure 2 It is a schematic structural diagram of the convolutional neural network weight compression device provided by the embodiment of the present application;
[0041] Figure 3 It is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0043] In the embodiments of the present application, the convolutional neural network described may specifically be a neural network applied to image processing or text processing.
[0044] In the embodiments of the present application, a method of block anti-triangular matrix decomposition is proposed to compress the weights in the graph convolutional neural network. The connection layer matrix A is decomposed into A = QMQ T in the form, where Q is an orthogonal matrix and M is a block lower anti-triangular matrix. By replacing the original convolutional layer with the low-rank block matrix of the M matrix, the number of parameters of the model is reduced, thereby reducing the model size, decomposition complexity, and running time.
[0045] Select the weight matrix to analyze the network architecture: Understand each layer of the network and their contributions to the total number of parameters and computational complexity. Generally, the main goal of compression is to reduce the model size and inference time while maintaining the highest possible accuracy. Therefore, the embodiments of the present application focus the decomposition on the layer with the most parameters or the most complex computation.
[0046] Generally, the convolutional layer and the fully connected layer are the main targets of compression because they contain a large number of parameters. The fully connected layer usually has more redundancy that can be exploited, but due to the characteristics of its spatial connection, the effect after the weight matrix decomposition of the convolutional layer may be more complex. Therefore, the embodiments of the present application select the fully connected layer with the most parameters or the most complex computation for compression.
[0047] Figure 1 is a schematic flowchart of the method for compressing the weights of the convolutional neural network provided by the embodiments of the present application. As Figure 1 shown, it includes:
[0048] Step 110, convert the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix;
[0049] Determine the four-dimensional tensor dimensions of the weights of the convolutional layer, that is, the number of output channels Fo, the number of input channels Fi, the height Kh of the convolutional kernel, and the width Kw.
[0050] Map each convolutional kernel from the four-dimensional space to a two-dimensional plane. Usually, by flattening the two-dimensional structure (height and width) of each convolutional kernel into a one-dimensional vector. For each output channel, stack the flattened convolutional kernel vectors corresponding to all input channels to form a row of a two-dimensional matrix. In this way, the weights of each output channel form a matrix row.
[0051] The number of rows of the obtained two-dimensional weight matrix is equal to the number of output channels Fo, and the number of columns is the number of input channels multiplied by the total number of elements of each convolutional kernel Fi×Kh×Kw.
[0052] The converted two-dimensional matrix may need to optimize the storage method to reduce memory occupancy and improve access efficiency. Ensure that the converted two-dimensional weight matrix accurately represents all the weight information in the original four-dimensional tensor without data loss or distortion.
[0053] Use the converted two-dimensional weight matrix for subsequent weight compression algorithms such as singular value decomposition (SVD) or block triangular decomposition.
[0054] Through this process, the weights of the convolutional layer are converted into a more manageable form, laying the foundation for further compression and optimization. This conversion helps reduce the complexity of the model and the number of parameters, thereby improving computational efficiency while maintaining performance.
[0055] The weights of a convolutional layer are usually stored as a four-dimensional tensor in the format (number of output channels, number of input channels, filter height, filter width). To apply matrix decomposition, this four-dimensional tensor needs to be converted into a two-dimensional matrix. Since the weights of the fully connected layer are already two-dimensional matrices, no additional reshaping step is required.
[0056] Step 120, convert the two-dimensional weight matrix into a block triangular form through an improved orthogonal transformation, where the improved orthogonal transformation includes at least one of the following: Givens transformation and Householder transformation;
[0057] In the embodiments of the present application, the Givens transformation or the Householder transformation, or a combination of both, is used to perform the required orthogonal transformation.
[0058] The Givens transformation gradually converts the matrix into an upper triangular or lower triangular form by constructing a series of rotation matrices. In this process, the Givens transformation is used to convert the selected two-dimensional weight matrix into a block triangular form.
[0059] The Householder transformation reduces the elements of certain rows or columns of the matrix to zero by generating a reflection transformation, which helps to further simplify the matrix structure.
[0060] Divide the matrix into multiple smaller sub-blocks, and perform orthogonal transformation on each sub-block independently to facilitate management and reduce the amount of calculation.
[0061] Gradually apply the orthogonal transformation to each sub-block until the entire matrix is converted into a block triangular form. This step may require multiple iterations. Ensure that the applied transformation maintains the orthogonal property of the matrix throughout the transformation process.
[0062] During the transformation process, find ways to reduce the computational load, such as by leveraging the symmetry of matrices or other mathematical properties. Verify whether the transformed block anti-triangular matrix satisfies the expected mathematical characteristics to ensure the correctness of the transformation. Record the details of each step of the transformation, including the type of transformation used, parameters, and intermediate results, for easy review and further analysis.
[0063] Use the obtained block anti-triangular form as the input for subsequent steps, such as matrix compression, decomposition, or other operations.
[0064] Step 130: Calculate the computational workload of each transformation process based on the number of floating-point operations required by the block anti-triangular algorithm. Perform a first-order correction on the symmetric indefinite matrix whose computational workload exceeds the preset threshold to obtain the corrected block anti-triangular form.
[0065] In the embodiments of this application, define the number of floating-point operations (FLOPs) as an indicator to measure the computational workload, including addition, subtraction, multiplication, and division operations. Decompose the block anti-triangular algorithm into multiple steps, including matrix initialization, orthogonal transformation, decomposition, etc. Calculate the required FLOPs for the mathematical operations involved in each step of the algorithm.
[0066] Accumulate the FLOPs of all steps to obtain the total computational workload of the entire algorithm. Determine a preset threshold for the computational workload as the criterion for judging whether a first-order correction is needed.
[0067] Identify the symmetric indefinite matrices whose computational workload exceeds the threshold during the transformation process. For the matrices that exceed the threshold, design a first-order correction scheme, which may include adjusting the diagonal elements of the matrix or changing the matrix structure through orthogonal transformation.
[0068] Apply the first-order correction scheme to the identified matrices to adjust the matrix characteristics to reduce the computational workload. Recalculate the FLOPs for the corrected matrices to verify the correction effect.
[0069] If the computational workload after correction still exceeds the threshold, further optimize the correction scheme. Record the detailed process and results of the first-order correction, including the reduction in FLOPs. Convert the corrected matrix into the block anti-triangular form to prepare for subsequent weight compression and network optimization. Ensure that the corrected matrix meets the requirements in terms of numerical stability and performance.
[0070] Step 140: Reshape the lower triangular block matrix of the corrected block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor. Replace the original four-dimensional tensor with the new four-dimensional tensor to obtain a compressed convolutional neural network.
[0071] In the embodiments of the present application, lower triangular block matrices in a block anti-triangular form are obtained from the first-order correction process. These matrices have a lower computational amount and optimized numerical characteristics. Each lower triangular block matrix is rearranged to restore from the block anti-triangular form to the original convolution kernel shape, and this step involves transposing and reorganizing the matrix.
[0072] The reshaped two-dimensional lower triangular matrices are stacked to construct a new four-dimensional tensor according to the dimensions of the original weight tensor (number of output channels, number of input channels, convolution kernel height, convolution kernel width).
[0073] In the corresponding layer of the convolutional neural network, the original weight tensor is replaced with the new four-dimensional tensor to ensure that the dimensions of the weights are compatible with the network structure. The network structure is adjusted according to the new weight tensor, and if necessary, other parts of the network are modified to adapt to the compression of the weights.
[0074] The replaced convolutional neural network is tested to ensure that the new weight tensor can maintain or be close to the performance of the original network, such as classification accuracy and convergence speed. According to the performance test results, the network parameters are fine-tuned to maximize the compression effect while minimizing the performance loss. The impact of weight compression on model size, computational efficiency, and memory usage is evaluated to ensure that the expected compression target is met.
[0075] After the above steps, the weight compression of the convolutional neural network is completed, and an optimized model with reduced number of parameters and reduced computational cost is obtained.
[0076] In the embodiments of the present application, by converting the weights from a four-dimensional tensor to a block anti-triangular form, the number of parameters of the network is reduced because a low-rank matrix can approximate the original weight matrix with fewer parameters. The compression of the weight matrix reduces the number of multiplication operations required for the convolution operation, thereby reducing the computational burden of the model in forward and backward propagation. The first-order correction adjusts the large computational symmetric indefinite matrix, improves the condition number of the matrix, and thus improves the stability of numerical calculations. The present application provides various transformation methods such as Givens and Householder in the block anti-triangular algorithm and comparison data of computational complexity, which can help select different transformation forms for different matrices and reduce the algorithm duration.
[0077] Optionally, converting the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix includes:
[0078] Obtain the original four-dimensional tensor of the weights of the convolutional layer, where the original four-dimensional tensor includes: the number of output filters, the number of input filters, the filter height, and the filter width;
[0079] For each output filter, the convolution kernels of all the input filters corresponding to the output filter are unfolded into a one-dimensional vector in depth, and the one-dimensional vectors of all the output filters are stacked to form a two-dimensional weight matrix;
[0080] Among them, the number of rows of the two-dimensional weight matrix is equal to the number of output filters, and the number of columns of the two-dimensional weight matrix is equal to the number of input filters multiplied by the total number of elements of each convolution kernel.
[0081] Obtain the original four-dimensional tensor of the weights from the convolutional layer, and the dimensions of this tensor define the number of output filters (Fo), the number of input filters (Fi), the height (Kh) and width (Kw) of each filter.
[0082] Recognize that each input filter corresponds to a convolution kernel, and each convolution kernel has a specific height and width, forming the depth of the filter.
[0083] For each output filter, the convolution kernels of all the input filters corresponding to it are unfolded into a one-dimensional vector in depth. This means converting the two-dimensional structure (Kh×Kw) of each convolution kernel into one dimension.
[0084] Stack all the unfolded one-dimensional vectors in the order of the output filters to form the rows of a two-dimensional matrix.
[0085] Through the above stacking operation, a complete two-dimensional weight matrix is constructed. In this matrix, each row represents the weights of an output filter, and each column is the result of unfolding the convolution kernel corresponding to the input filter.
[0086] The number of rows of the two-dimensional weight matrix is equal to the number of output filters (Fo), and the number of columns is equal to the number of input filters multiplied by the total number of elements of each convolution kernel (Fi)×Kh×Kw.
[0087] Consider optimizing the storage method of the newly constructed two-dimensional weight matrix for subsequent processing and access.
[0088] Ensure the accuracy of the weight information during the conversion process, and verify whether the newly constructed two-dimensional weight matrix correctly represents the weight data in the original four-dimensional tensor.
[0089] In an optional embodiment, the algorithm construction form is recursive. Assume that the embodiment of the present application starts from A (k) =A(1:k,1:k),1≤k<n, i.e., and Q (k) ∈R k×k is an orthogonal matrix, and M (k) ∈R k×k is the corresponding block anti-triangular form. Note that this form is simplest when k = 1. Then the embodiment of the present application shows how to transform A(k+1) = A(1:k+1, 1:k+1) is transformed into the corresponding anti-triangular form by an orthogonal transformation. Let inertia(A (k) ) = (k - , k 0 ,, k + ), k - + k 0 + k + = k, and let k 1 = min(k - , k + ), k 2 = max(k - , k + ) - k 1 . Let M (k) be the corresponding block anti-triangular form of A (k) ,
[0090] i.e.,
[0091] is a non-singular lower anti-triangular matrix, a symmetric indefinite matrix, and are symmetric matrices. Note that some blocks may be zero-dimensional. Without loss of generality, the embodiments of this application assume that k 2 > 0 and X is positive definite and has a Cholesky decomposition X = LL T . Let a = A(1:k, k+1), γ = a k+1,k+1 partition.
[0092]
[0093] Let Because
[0094]
[0095] Now the embodiments of this application illustrate how to improve the orthogonal similarity transformation to reduce it to the corresponding block anti-triangular form, which is divided into the following three cases:
[0096] Case a. Let
[0097]
[0098] be a circulant matrix. is the required block anti-triangular form. In this case, let Then is the corresponding block anti-triangular form. And inertia(A(k +1) ) = (k - , k 0 +1, k + )。
[0099] Cases b.
[0100] This embodiment of the application considers the Householder matrix such that Let
[0101]
[0102] Then is in the corresponding block anti-triangular form,
[0103]
[0104] where
[0105]
[0106] In addition, the inertia (A (k+1) ) = (k - +1, k 0 -1, k + +1).
[0107] Cases c.
[0108] For example, this situation occurs when M (k) is non-singular, i.e., k 0 = 0 and The first step is to eliminate the terms on the main anti-diagonal of the sub-matrix using a Givens matrix of order k such that 1
[0109]
[0110] Y c is a non-singular lower anti-triangular matrix of rank k 1 . Let
[0111]
[0112] and
[0113]
[0114] In addition, let
[0115]
[0116] Then
[0117]
[0118] The embodiments of the present application define a sub-matrix
[0119]
[0120] The next step in this case is to check the correctness of X 1 . Let be generated by a Givens rotation of order k 2 -1 such that Then
[0121]
[0122] where is a sequence of "internal" Givens rotations of order k 2 -1 such that is a non-singular lower triangular matrix. The embodiments of the present application decompose L 1 into
[0123]
[0124] is a lower triangular matrix, i.e., it is proved that
[0125]
[0126] Note that β≠0. Since the Schur complement of L 1 L 1 T in X 2 is X 2 is positive definite and singular, or indefinite if the 2×2 matrix is positive definite and singular or indefinite, respectively. If the embodiments of the present application define
[0127]
[0128] and
[0129] then the embodiments of the present application have the following situation.
[0130] c.1. Matrix T is symmetric and positive definite
[0131]
[0132] is its Cholesky factor. Thus X 2is symmetric positive definite (SPD), and X 1 is the same. Moreover, its Cholesky factor is given by
[0133]
[0134] and then
[0135]
[0136] c.2 Matrix T is non-singular and its eigenvalues are 0 = λ 1 < λ 2 . Let Q 1 be a Givens rotation such that
[0137]
[0138]
[0139] Then
[0140] Let be a Givens rotation of order k 2 -1 such that
[0141]
[0142] and Next
[0143]
[0144] Then and
[0145] Let
[0146] Then
[0147] and
[0148] Let be a Givens rotation of order k 1 such that
[0149]
[0150] where is non-singular lower triangular.
[0151] Let
[0152] Then
[0153] c.3T is an indefinite matrix.
[0154] Let Q 2 be a Givens rotation such that
[0155] Let
[0156] Then
[0157]
[0158] Similar to the previous sub - case, let be a Givens transformation such that
[0159]
[0160] where
[0161] Then
[0162] where Let Let
[0163]
[0164] In addition
[0165] Then
[0166] and
[0167] Note 2. This algorithm is backward - stable because it relies only on Householder transformations and Givens transformations.
[0168] Note 3. The embodiments of this application note that the signs of the anti - diagonal elements of Y and the diagonal elements of L can be chosen arbitrarily.
[0169] Note 4. For the coefficients c and s of the Givens rotations in (2.13) and (2.20) that transform T into an anti - diagonal matrix, the following equations can be calculated:
[0170]
[0171] The parameters c and s can be calculated in a stable way as the absolute value of the root of a quadratic equation, and t = s / c.
[0172] Optionally, based on the number of floating-point operations required by the block inverse triangular algorithm, the computational workload of each transformation process is calculated, specifically including:
[0173] Obtain the basic data operation categories corresponding to each orthogonal transformation in the block inverse triangular algorithm, and the number of basic data operations corresponding to each orthogonal transformation;
[0174] Based on the first floating-point operation quantities corresponding to each basic data operation, calculate the second floating-point operation quantity of each orthogonal transformation, and accumulate the second floating-point operation quantities to obtain the third floating-point operation quantity of the orthogonal transformation in the block inverse triangular algorithm;
[0175] Determine the computational workload of each transformation process in the block inverse triangular algorithm according to the third floating-point operation quantity.
[0176] In the embodiments of the present application, determine the basic data operation categories required for each orthogonal transformation involved in the block inverse triangular algorithm, such as addition, subtraction, multiplication, and division. For each basic data operation, count the number of times it occurs in each orthogonal transformation.
[0177] Allocate floating-point operation quantities (FLOPs) for each basic data operation according to the operation type. Usually, it is assumed that each addition and subtraction is 1 FLOP, and each multiplication and division is also 1 FLOP.
[0178] Multiply the number of times of each type of operation in each orthogonal transformation by the corresponding FLOPs to obtain the total FLOPs of this transformation. Accumulate the FLOPs of all orthogonal transformations to obtain the cumulative FLOPs of the orthogonal transformation in the entire block inverse triangular algorithm.
[0179] Evaluate the computational workload of each transformation process in the block inverse triangular algorithm according to the cumulative FLOPs. If it is found that the FLOPs of some steps are abnormally high, consider algorithm optimization or adjustment to reduce the computational workload. Record the FLOPs calculation results of each step, analyze the algorithm efficiency, and determine whether there is room for further optimization. Determine the threshold of the computational workload according to the algorithm performance requirements and hardware limitations.
[0180] If the FLOPs of the orthogonal transformation exceed the threshold, perform a first-order correction, such as adjusting the transformation parameters or adopting a more efficient algorithm variant. Recalculate the FLOPs for the transformed process after the first-order correction to ensure that the computational workload requirements are met.
[0181] Conduct a final evaluation of the corrected algorithm to ensure that the computational workload is within an acceptable range while maintaining the accuracy and efficiency of the algorithm.
[0182] In an optional embodiment, following the decomposed computational workload given in the embodiments of the present application. Given The algorithm just described in the embodiments of this application can transform A into the corresponding block anti-triangular form through orthogonal transformation. (k+1) Simplify it into the corresponding block anti-triangular form.
[0183] The amount of computation depends on the situation that occurs. The embodiments of this application check the amount of computation for each individual situation. The only identical operation is to generate the 2k floating-point operations required by (2.2). 2 Floating-point operations.
[0184] a. Checking Whether it holds requires O(k) floating-point operations.
[0185] b. The multiplication in (2.3) requires 4k floating-point operations for k. 0 Floating-point operations for k.
[0186] c. The results of (2.4), (2.5), (2.6), (2.7), (2.10), (2.11) and (2.12) are the same for this situation, and the amounts of computation for their floating-point operations are 6k 1 k 2 , (Due to symmetry), 6k 1 k, and 6k 1 k 2 .
[0187] c.1. No additional floating-point operations are required here.
[0188] c.2. The results of (2.14), (2.16), (2.17), (2.18) and (2.19) respectively require 6k 2 k, 6k 1 k 2 , and 6k 1 k floating-point operation groups.
[0189] c.3. The results of (2.21), (2.22), (2.23) and (2.24) The amounts of computation required for their floating-point operations are respectively 6k 1 k 2 and 6kk 2 .
[0190] Therefore, the complexity of the algorithm largely depends on the inertia of the principal submatrix decomposed from the symmetric matrix, and its rank is n 3 .
[0191] The number of multiplications for the Givens rotation operation in this algorithm can be halved if the final rotation is replaced by the fastest Givens transformation.
[0192] Optionally, the specific process of the first-order correction includes:
[0193] Identifying the matrix characteristics of a symmetric indefinite matrix whose amount of computation exceeds a preset threshold, and determining the correction target of the symmetric indefinite matrix;
[0194] Performing cyclic correction on the symmetric indefinite matrix according to the correction target until the amount of computation after correction is less than the preset threshold, to obtain a corrected symmetric indefinite matrix;
[0195] Converting the corrected symmetric indefinite matrix into a block anti-triangular form to obtain a corrected block anti-triangular form.
[0196] In order to reduce the computational complexity, it is important to correct a symmetric indefinite matrix in many applications whose leading | subspaces can be easily detected by the fastest symmetric first-order correction method. The first-order correction of a symmetric anti-diagonal matrix [22, 23] can be simplified by an algorithm, with diagonal elements [4, 7, 11] and a symmetric diagonal matrix plus semi-separable elements [18, 19] in anti-diagonal form. A symmetric matrix A of rank n is decomposed into A = Q T Q T , where T is one of the following matrices and Q is an orthogonal matrix. Although the first-order correction of matrix T can be computed to O(n 2 ), the improved Q-factor requires O(n 3 ).
[0197] Next, the embodiments of the present application will show that a corresponding symmetric block anti-triangular matrix can be improved by a stable O(n 2 ) symmetric first-order correction method.
[0198] Let A = Q M Q T and inertia(A) = (n - , n 0 , n + ), n - + n0 + n+ = n and Q is an orthogonal matrix. Let n 1 = min(n - , n + ), n 2 = max(n - , n + ) - n 1 ,
[0199]
[0200] where is a non-singular lower anti-triangular matrix, X = ε L L T and when ε = 1, n + > n- When ε = -1, n + < n - where L is a lower anti-triangular matrix. The purpose is to transform QMQ T ±yy T = Q(M ± Q T yy T Q)Q T where M is the corresponding symmetric block anti-triangular matrix until a matrix with the same structure as M and orthogonal, eliminating the terms of the vector x = Q T y.
[0201] Let N 1 = n 0 + n 1 and N 2 = n 0 + n 1 + n 2 In the embodiments of the present application, we first consider the case where n 0 , n 1 , n 2 > 0, and the case where n 0 = 0 will be briefly described at the end.
[0202] The embodiments of the present application divide the algorithm into four steps:
[0203] The first step. When n 0 > 1, the first step is to determine a Householder matrix H (1) such that H (1) x(1:n 0 ) can eliminate all terms except the last one. Let then
[0204]
[0205] where Q 1 = H 1 x, The embodiments of the present application observe that M remains unchanged because the first n 0 rows / columns are zero. The embodiments of the present application note that the algorithm reverts to the case where n 0 = 0 if ||x(1:n 0 )|| 2 = 0.
[0206] Due to the special structure of the Householder transformation matrix H 1 , the product requires 4n 0 n, requires 4n 0 . The number of operands can be reduced to 3n 0n and 3n 0 By using the sequence of the fastest Givens transforms instead of a single Householder transform.
[0207] Step 2. When n 1 = min(n - , n + ) > 0, i.e., if M is indefinite, of the n 0 , n 0 + 1,..., N 1 - 1 terms are eliminated by multiplying by a sequence of n 1 Givens rotations G i , i = 1, 2,..., n 1 , such that G i acts on the (n 0 + i + 1)-th and (n 0 + i)-th rows, eliminating the term at position (n 0 + i + 1). Let M (0) = M, the result is to modify the last i terms of the n (i-1) + i + 1 and n 0 + i rows / columns of M 0 , introducing a non-zero term at position (n 0 + i + 1, n + i + 1) and the symmetric position (n + i + 1, n 0 + i + 1).
[0208] For simplicity, the first n 0 - 1 terms of x and the first n 0 - 1 rows and columns of M are not described as they are all zero.
[0209] Due to the anti-triangular structure of M (0) , the computational cost of In addition, the computational cost of 1 is 6n
[0210] Step 3. The purpose of this step is to use a sequence of Givens rotations to eliminate all terms except the central positive definite submatrix of X corresponding to the last term. Since X is in the factored form, i.e., X = εLL T and ε = ±1, other sequences of Givens inner rotations are considered to preserve the factored structure. In particular, the first Givens rotation is applied to modify the (N 1 + 1)-th and N1 +1 row and eliminate the N 1 +2 terms. Then use a similar transformation to calculate The embodiments of the present application observe that this transformation corrects the first two rows / columns of the submatrix of X to produce non-zero terms at the (1,2) position of L. To remove this convexity, a new Givens rotation is applied to the right of L. The embodiments of the present application observe that the last rotation only acts on the matrix L. Without loss of generality, only the central submatrix of X and the corresponding terms of the vector are described in the process described. Let
[0211] Finally the N 1 row (column) is moved to below (to the right of) the positive definite by multiplying by a permutation matrix P. Let and
[0212] Due to the anti-triangular structure the computational cost required is In addition, to preserve the anti-triangular structure and remove the convexity of L requires In addition the computational cost is 6n 2 n.
[0213] The fourth step Add or subtract a first-order matrix
[0214] Because is different from As a submatrix of. Although is positive definite, there is no way to show the positive definiteness of .
[0215] To complete the simplification it is sufficient to simplify the submatrix of to the corresponding block anti-triangular form by applying a symmetric positive definite submatrix that does not start from the second part
[0216] If the submatrix is singular, otherwise it is the second step. Of course, Givens rotations must be applied to the entire matrix
[0217] Let be the matrix obtained after the first step of simplification.
[0218] For clarity, the first two steps described in the embodiments of the present application are divided into three cases and the second step has only one sub - case.
[0219] Case a. Positive definite. Using step c.1 of part 2.1, the final matrix is transformed into the corresponding block anti - triangular form.
[0220] Case b. Singular. Using step c.2 of part 2.1, the final matrix is transformed into the corresponding block anti - triangular form.
[0221] Case c. Indefinite. Using step c.3 of part 2.1, the final matrix is transformed into the corresponding block anti - triangular form. To complete this case and the overall reduction, a similarity transformation based on Givens rotation acts on the N 2 -1 and N 2 rows (columns) to eliminate the (N 2 -1, N 1 )(N 1 , N 2 -1) terms.
[0222] Next, the second step of the embodiments of the present application describes reducing Case a to the corresponding block anti - triangular form ( positive definite). Case c ( indefinite) is also done in a similar way.
[0223] The embodiments of the present application are divided into the following three sub - cases:
[0224] 1.a.1. Positive definite. Using step c.1 of part 2.1, the final matrix is transformed into the corresponding block anti - triangular form.
[0225] 2.a.2. Singular. Then is singular. Using step c.3 of part 2.1, the final matrix is transformed into the corresponding block anti - triangular form. As described in step c.3 of part 2.1, n 1 Givens rotations act on the N 1 -i and N 1 -i + 1 rows (columns) to eliminate the (N 1 -i, n 1 -i + 1) terms, and must be used to reduce the entire matrix to the corresponding block anti - triangular form.
[0226] 3.a.3. Non - positive definite. Using step c.3 in section 2.1, the final matrix is transformed into the corresponding block anti - triangular form.
[0227] Computational complexity According to section 2.1, case c, the computational complexity of the initial sub - steps is mainly due to the need to simplify to the required central part. The amount of work required to transform into the corresponding block anti - triangular form is O(n 2 n 1 +n 2 ). The improved matrix Q requires O(n 2 n), and an additional amount of work O(n 1 n) must be added to sub - case a.2 because of the sequence of products through the decomposed orthogonal matrices.
[0228] 3.5.n 0 = 0 if n 0 = 0, the first step of the above algorithm is skipped. And the second step has the following modification. Let M(0)=M. The sequence of Givens transformations is considered to eliminate the first n 1 - 1 terms of x. Let The matrix is different from the corresponding block anti - triangular matrix for the (i,2n 1 +n 2 - i) and (2n 1 +n 2 - i,i) terms, and is different from 0 when i = 1,...,n 1 - 1. To eliminate the last terms of other sequences, use Givens rotations acting on the rows and columns of 2n 1 +n 2 - i and 2n 1 +n 2 - i + 1 such that
[0229]
[0230] Then the simplification process is similar to the case when n 0 ≠0.
[0231] Optionally, the steps of reshaping the lower triangular block matrix of the modified block anti - triangular form back into a four - dimensional tensor, obtaining a new four - dimensional tensor, and replacing the original four - dimensional tensor with the new four - dimensional tensor to obtain a compressed convolutional neural network include:
[0232] Reshaping each of the lower triangular blocks in the modified block anti - triangular matrix into the shape of the original convolution kernel and recombining them to form a new four - dimensional weight tensor;
[0233] Integrate the new four - dimensional weight tensor into the convolutional neural network, replacing the original weight tensor of the corresponding layer to obtain a compressed convolutional neural network.
[0234] In the embodiment of the present application, extract the lower triangular blocks from the modified block anti - triangular matrix. These blocks are the optimized matrix parts obtained during the modification process. Rearrange each lower triangular block matrix to restore it from a one - dimensional or two - dimensional structure to the original convolutional kernel shape, that is, restore its corresponding height and width dimensions.
[0235] For each output channel, stack the reshaped convolutional kernels in sequence to form a new four - dimensional weight tensor. This step ensures that the new weight tensor has the same four - dimensional structure as the original weight.
[0236] Ensure that the dimensions of the new four - dimensional tensor are consistent with the original weight tensor to ensure that it can be seamlessly integrated into the convolutional neural network. Integrate the new four - dimensional weight tensor into the corresponding layer of the convolutional neural network, replacing the original weight tensor.
[0237] In each layer of the network, replace the original weight with the optimized weight to ensure that the input and output of the layer remain consistent. Adjust the network structure according to the new weight tensor, and if necessary, fine - tune other parts of the network to adapt to the weight compression.
[0238] The convolutional neural network weight compression device provided by the present invention will be described below. The convolutional neural network weight compression device described below can be correspondingly referred to the convolutional neural network weight compression method described above.
[0239] Figure 2 It is a schematic structural diagram of the convolutional neural network weight compression device provided by the embodiment of the present application, as Figure 2 shown, including:
[0240] The first conversion module 210 is used to convert the weight of the convolutional layer of the original convolutional neural network from the original four - dimensional tensor to a two - dimensional weight matrix;
[0241] The second conversion module 220 is used to convert the two - dimensional weight matrix into a block anti - triangular form through an improved orthogonal transformation, where the improved orthogonal transformation includes at least one of the following: Givens transformation and Householder transformation;
[0242] The calculation module 230 is used to calculate the amount of computation for each conversion process based on the number of floating - point operations required by the block anti - triangular algorithm, and perform a first - order correction on the symmetric indefinite matrix whose amount of computation exceeds the preset threshold to obtain a corrected block anti - triangular form;
[0243] The replacement module 240 is used to reshape back to a four-dimensional tensor through the lower triangular block matrix in the corrected block anti-triangular form, obtain a new four-dimensional tensor, and replace the original four-dimensional tensor according to the new four-dimensional tensor to obtain a compressed convolutional neural network.
[0244] Optionally, the device is further configured to:
[0245] Obtain the original four-dimensional tensor of the convolutional layer weights, where the original four-dimensional tensor includes: the number of output filters, the number of input filters, the filter height, and the filter width;
[0246] For each output filter, expand the convolution kernels of all input filters corresponding to the output filter into a one-dimensional vector in the depth direction, and stack the one-dimensional vectors of all output filters to form a two-dimensional weight matrix;
[0247] Wherein, the number of rows of the two-dimensional weight matrix is equal to the number of output filters, and the number of columns of the two-dimensional weight matrix is equal to the number of input filters multiplied by the total number of elements of each convolution kernel.
[0248] In the embodiments of the present application, by converting the weights from a four-dimensional tensor to a block anti-triangular form, the number of parameters of the network is reduced because a low-rank matrix can approximate the original weight matrix with fewer parameters. The compression of the weight matrix reduces the number of multiplication operations required for convolution operations, thereby reducing the computational burden of the model in forward and backward propagation. The first-order correction adjusts the computationally intensive symmetric indefinite matrix, improves the condition number of the matrix, and thus improves the stability of numerical calculations. The present application provides various transformation methods such as Givens and Householder in the block anti-triangular algorithm and comparison data of computational complexity, which can help select different transformation forms for different matrices and reduce the algorithm duration.
[0249] Figure 3 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 3 shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete mutual communication through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the convolutional neural network weight compression method, which includes: converting the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix;
[0250] Convert the two-dimensional weight matrix into a block anti-triangular form through an improved orthogonal transformation, where the improved orthogonal transformation includes at least one of the following: Givens transformation and Householder transformation;
[0251] Calculate the amount of computation for each conversion process based on the number of floating-point operations required by the block anti-triangular algorithm, and perform a first-order correction on the symmetric indefinite matrix whose amount of computation exceeds a preset threshold to obtain a corrected block anti-triangular form;
[0252] Reshape the lower triangular block matrix of the corrected block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replace the original four-dimensional tensor with the new four-dimensional tensor to obtain a compressed convolutional neural network.
[0253] In addition, when the logical instructions in the above-mentioned memory 330 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0254] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the convolutional neural network weight compression method provided by the above-mentioned various methods. The method includes: converting the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor into a two-dimensional weight matrix;
[0255] Convert the two-dimensional weight matrix into a block anti-triangular form through an improved orthogonal transformation, where the improved orthogonal transformation includes at least one of the following: Givens transformation and Householder transformation;
[0256] Calculate the amount of computation for each conversion process based on the number of floating-point operations required by the block anti-triangular algorithm, and perform a first-order correction on the symmetric indefinite matrix whose amount of computation exceeds a preset threshold to obtain a corrected block anti-triangular form;
[0257] Reshape the lower triangular block matrix in the corrected block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replace the original four-dimensional tensor with the new four-dimensional tensor to obtain a compressed convolutional neural network.
[0258] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the convolutional neural network weight compression method provided by the above methods. The method includes: converting the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix;
[0259] Convert the two-dimensional weight matrix into a block anti-triangular form through an improved orthogonal transformation, where the improved orthogonal transformation includes at least one of the following: Givens transformation and Householder transformation;
[0260] Based on the number of floating-point operations required by the block anti-triangular algorithm, calculate the amount of operations for each conversion process, and perform a first-order correction on the symmetric indefinite matrix whose amount of operations exceeds a preset threshold to obtain a corrected block anti-triangular form;
[0261] Reshape the lower triangular block matrix in the corrected block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replace the original four-dimensional tensor with the new four-dimensional tensor to obtain a compressed convolutional neural network.
[0262] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0263] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0264] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A convolutional neural network weight compression method, characterized in that: include: Convert the weights of the convolutional layer of the original convolutional neural network from the original four-dimensional tensor to a two-dimensional weight matrix; The two-dimensional weight matrix is converted into a block inverse triangular form by an improved orthogonal transformation, wherein the improved orthogonal transformation includes at least one of the following: a Givens transformation and a Householder transformation; Based on the number of floating-point operations required by the block anti-triangular algorithm, the amount of operation of each conversion process is calculated, and a first-order correction is performed on the symmetric indefinite matrix whose amount of operation exceeds a preset threshold to obtain a corrected block anti-triangular form; The lower triangular block matrix in the modified block anti-triangular form is reshaped back into a four-dimensional tensor to obtain a new four-dimensional tensor, and the original four-dimensional tensor is replaced according to the new four-dimensional tensor to obtain a compressed convolutional neural network.
2. The convolutional neural network weight compression method according to claim 1, characterized in that: The weights of the convolutional layer of the original convolutional neural network are converted from the original four-dimensional tensor to a two-dimensional weight matrix, including: Obtaining an original four-dimensional tensor of the convolutional layer weights, wherein the original four-dimensional tensor includes: the number of output filters, the number of input filters, the filter height, and the filter width; For each output filter, the convolution kernels of all input filters corresponding to the output filter are expanded in depth into a one-dimensional vector, and the one-dimensional vectors of all output filters are stacked to form a two-dimensional weight matrix; The number of rows of the two-dimensional weight matrix is equal to the number of output filters, and the number of columns of the two-dimensional weight matrix is equal to the number of input filters multiplied by the total number of elements of each convolution kernel.
3. The convolutional neural network weight compression method according to claim 1, characterized in that: Based on the number of floating-point operations required by the block anti-triangulation algorithm, the amount of computation required for each conversion process is calculated, including: Obtaining the basic data operation category corresponding to each orthogonal transformation in the block anti-triangulation algorithm, and the number of basic data operations corresponding to each orthogonal transformation; Based on the first floating-point operation number corresponding to each basic data operation, the second floating-point operation number of each orthogonal transformation is calculated, and each of the second floating-point operation numbers is accumulated to obtain a third floating-point operation number of the orthogonal transformation in the block inverse triangulation algorithm; The amount of computation of each transformation process in the block inverse triangulation algorithm is determined according to the third floating-point operation quantity.
4. The convolutional neural network weight compression method according to claim 1, characterized in that: The specific process of the first-order correction includes: Identifying matrix characteristics of a symmetric indefinite matrix whose computation amount exceeds a preset threshold, and determining a correction target of the symmetric indefinite matrix; The symmetric indefinite matrix is cyclically corrected according to the correction target until the corrected computation amount is less than the preset threshold value, thereby obtaining a corrected symmetric indefinite matrix; The modified symmetric indefinite matrix is converted into a block anti-triangular form to obtain a modified block anti-triangular form.
5. The convolutional neural network weight compression method according to claim 1, characterized in that: The steps of reshaping the lower triangular block matrix in the modified block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replacing the original four-dimensional tensor with the new four-dimensional tensor to obtain a compressed convolutional neural network include: Reshape the lower triangular blocks in the modified block anti-triangular matrix into the shape of the original convolution kernel one by one, and recombine them to form a new four-dimensional weight tensor; The new four-dimensional weight tensor is integrated into the convolutional neural network to replace the original weight tensor of the corresponding layer to obtain a compressed convolutional neural network.
6. A convolutional neural network weight compression device, characterized in that: include: A first conversion module is used to convert the weight of the convolution layer of the original convolutional neural network from an original four-dimensional tensor to a two-dimensional weight matrix; A second conversion module is used to convert the two-dimensional weight matrix into a block inverse triangular form through an improved orthogonal transformation, wherein the improved orthogonal transformation includes at least one of the following: a Givens transformation and a Householder transformation; A calculation module is used to calculate the amount of operation of each conversion process based on the number of floating-point operations required by the block anti-triangular algorithm, and to perform a first-order correction on the symmetric indefinite matrix whose amount of operation exceeds a preset threshold to obtain a corrected block anti-triangular form; A replacement module is used to reshape the lower triangular block matrix in the modified block anti-triangular form back into a four-dimensional tensor to obtain a new four-dimensional tensor, and replace the original four-dimensional tensor according to the new four-dimensional tensor to obtain a compressed convolutional neural network.
7. The convolutional neural network weight compression device according to claim 6, characterized in that: The device is also used for: Obtaining an original four-dimensional tensor of the convolutional layer weights, wherein the original four-dimensional tensor includes: the number of output filters, the number of input filters, the filter height, and the filter width; For each output filter, the convolution kernels of all input filters corresponding to the output filter are expanded in depth into a one-dimensional vector, and the one-dimensional vectors of all output filters are stacked to form a two-dimensional weight matrix; The number of rows of the two-dimensional weight matrix is equal to the number of output filters, and the number of columns of the two-dimensional weight matrix is equal to the number of input filters multiplied by the total number of elements of each convolution kernel.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the convolutional neural network weight compression method as described in any one of claims 1 to 5 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the convolutional neural network weight compression method as described in any one of claims 1 to 5 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the convolutional neural network weight compression method as described in any one of claims 1 to 5 is implemented.