Multi-decomposition compression method of artificial intelligence model based on reinforcement learning
Through the multi-decomposition and compression method of artificial intelligence model based on reinforcement learning, the most suitable decomposition methods and ranks for each network layer of the AI model are automatically selected, which solves the problem of poor compression effect of AI model in the prior art, and achieves efficient model compression and precision maintenance.
Patent Information
- Application Number
- CN202510571921.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When compressing large-scale AI models, it is difficult to achieve optimal compression effects on each network layer of the model, and the compression process is time-consuming and cannot effectively weigh the compression rate and accuracy losses.
Using the multi-decomposition and compression method of artificial intelligence model based on reinforcement learning, the most suitable decomposition method and optimal decomposition rank of each network layer are automatically selected, and by constructing an objective function paradigm of joint model accuracy and compression cost, the accuracy and compression rate of the compression model are optimal.
The model compression efficiency is improved, the optimal trade-off between balancing model compression ratio and model accuracy is achieved, and the direct deployment can be improved to improve practical efficiency.
Smart Images

Figure CN120106151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network model compression, and in particular to a large-scale artificial intelligence model compression multi-decomposition method based on reinforcement learning. Background Art
[0002] As AI model functions continue to strengthen, the scale of model parameters is also showing an increasingly large trend. Common models such as VGG and Resnet have parameters in the millions, which requires a lot of computing resources and time costs to run these models. Therefore, in order to effectively run these large-scale AI models, they must be deployed on resource-rich cloud servers. However, the process of transmitting collected data from the edge to the cloud, processing it in the cloud, and then feeding the results back to the edge has a large transmission delay, resulting in response delays.
[0003] In the current increasing number of intelligent application scenarios, it is urgent to implement edge intelligence at the edge nodes to meet the needs of rapid service response, such as smart home, unmanned driving and other scenarios. To this end, the AI model must be deployed and run on the edge nodes. The advantage of this is that placing the AI model on the edge node can quickly call local data for training without waiting for data to be transmitted to the cloud server. This can not only reduce network latency and transmission costs, but also improve data privacy and security.
[0004] However, the resources of edge nodes are very limited, which makes it difficult to deploy large-scale AI models on edge nodes. Among the existing patents, the invention patent with publication number CN117973476A discloses a lightweight method for vehicle-side models based on matrix weighted low-rank decomposition. The method includes the following steps: collecting and preprocessing sensor data of the test vehicle; training the original model and evaluating the accuracy; performing weighted low-rank decomposition on the model with better accuracy; verifying the accuracy and size of the decomposed model; and finally deploying the lightweight model to the vehicle for verification.
[0005] The invention patent with publication number CN118786439A discloses a method for compressing a deep learning (DL) model by performing low-rank decomposition (LRD) using reinforcement learning (RL). The method includes the following steps: defining the decomposable candidate layers of the DL model; generating compression values and assigning decomposition ranks through an RL agent; training the RL agent based on the updated state; decomposing the candidate layers using LRD, generating and evaluating candidate compression models; and outputting the final compression model after the termination condition is met.
[0006] However, the prior art uses a low-rank matrix decomposition compression model, which is a single decomposition method for the AI model. Although the model is compressed, the structural characteristics of each network layer of the AI model are different, and different low-rank decomposition methods have different effects on parameter matrix compression. Therefore, it is not appropriate to use the same decomposition method for the entire model, and the decomposition method cannot achieve the optimal compression effect on each network layer of the model. In addition, the model compression process of some methods is time-consuming and does not take into account the relationship between compression rate and precision loss, and it is impossible to achieve the optimal trade-off between the two. In order to better solve the above problems, the inventors have summarized two key challenges through analysis: (1) There are many ways to implement low-rank decomposition, such as matrix decomposition, such as singular value decomposition (SVD) and conventional matrix decomposition (MF). Even if the same rank is used, different implementation methods have different effects on AI model compression, which affects the accuracy of the model. Therefore, it is particularly important to select appropriate low-rank decomposition methods for each layer of the AI model to ensure the best compression effect; (2) The performance of low-rank decomposition compression is affected by the rank R of the decomposition matrix. Low-rank decomposition decomposes the original matrix W (m×n) into two low-rank matrices U (m×r) and V (r×n), and the number of parameters after compression is r×(m+n). Selecting a smaller rank r can reduce the number of parameters of the compressed model and improve the compression ratio, but this will also lead to increased model loss and decreased accuracy. Therefore, it becomes very challenging to choose a suitable decomposition rank r to balance the model compression ratio and model accuracy. Summary of the invention
[0007] In view of the above problems existing in the prior art, the present invention provides a multi-decomposition compression method of an artificial intelligence model based on reinforcement learning and an application system of the method. The present invention is based on reinforcement learning, and automatically selects the most suitable decomposition method and optimal decomposition rank of each network layer, thereby improving the model compression efficiency; and constructs an objective function paradigm of joint model accuracy and compression cost to achieve the best compression model accuracy and compression rate; and, using an iterative optimization strategy of decomposition first and then training, directly deploys and improves practical efficiency.
[0008] The technical solution of the present invention is as follows: A multi-decomposition compression method for an artificial intelligence model based on reinforcement learning, comprising the following steps: Step S1, obtaining the weight parameters of the pre-trained model: extracting the weight parameters of each network layer, wherein the weight parameters are expressed in a matrix form; Step S2, optimizing multiple decomposition methods and ranks based on reinforcement learning: constructing a joint application framework of multiple decomposition methods based on reinforcement learning, automatically selecting the most suitable low-rank decomposition method for each network layer and determining the optimal decomposition rank for low-rank decomposition; Step S3, jointly optimizing the model accuracy and compression cost: constructing an objective function paradigm for jointly optimizing the model accuracy and compression cost, and updating the model weight parameters through the function paradigm to achieve the best compression model accuracy and compression rate; Step S4, alternately optimize compression model parameters: adopt an iterative optimization strategy of first decomposing and compressing and then training and learning, train and generate a compression model that can be directly deployed, and optimize the training efficiency of the model based on reinforcement learning.
[0009] Furthermore, the weight parameters in step S1 are in the form of a matrix, which is derived from the multi-decomposition method of the present invention using matrix decomposition. Therefore, the weight parameters of the pre-trained model obtained will be converted into a weight parameter matrix, which will be restored after the training is completed. Preferably, the step S2 comprises: S21. Construct a joint application framework of multiple decomposition methods based on reinforcement learning: After obtaining the weight parameters of the pre-trained model, configure multiple low-rank decomposition methods for each network layer according to the model characteristics, and set a unified benefit value as the selection criterion for each decomposition method based on reinforcement learning; S22, selecting a decomposition method and a decomposition rank: selecting the most suitable low-rank decomposition method for each network layer according to the selection criteria and determining the optimal decomposition rank, and performing low-rank decomposition on the weight parameters of each network layer; Preferably, the low-rank decomposition method in step S21 includes: singular value decomposition SVD and conventional matrix decomposition MF; The singular value decomposition SVD decomposes the weight matrix A into two low-rank matrices U and V, and the number of parameters is compressed from m×n to r×(m+n), where r is much smaller than the minimum value of m and n; The conventional matrix decomposition MF decomposes the weight matrix A into three low-rank matrices U, S and V, and the number of parameters is compressed from m×n to r×(m+r+n), where r is much smaller than the minimum value of m and n.
[0010] Preferably, the specific implementation of configuring multiple low-rank decomposition methods for each network layer in step S21 includes: Based on reinforcement learning, a variety of low-rank decomposition methods are deployed in each neural network layer and the weight parameters of the layer are compressed; The fully connected layer can be directly decomposed; For the tensor parameters in the convolution layer, matrix conversion is required.
[0011] Furthermore, in step S21, for different matrix low-rank decomposition methods, even if the same rank is used, different implementation methods have different effects on AI model compression, which affects the accuracy of the model. Furthermore, in step S21, based on reinforcement learning, different low-rank decomposition methods are used in each neural network layer to compress the weight parameters of the layer. For a convolutional layer, there are n The convolution kernel of , the tensor parameters are expressed as , it can be converted into matrix parameters as , perform SVD decomposition and MF decomposition on the matrix parameters, compress them and then restore them; Preferably, the selection of the most suitable low-rank decomposition method in step S22 is achieved by selecting a function, and the selection criterion is the benefit value.
[0012] Preferably, the steps of selecting the most suitable low-rank decomposition method and determining the optimal decomposition rank in step S22 are as follows: (1) The negative of the sum of the compression accuracy loss and compression rate of the weight parameters of this layer is taken as the benefit value Q. The larger the value, the better the performance. (2) For each decomposition method of each network layer, traverse the possible ranks respectively, analyze the benefit value Q, and the rank with the maximum Q value is the optimal decomposition rank of the network layer, and the maximum benefit value Q is the maximum benefit value of the decomposition method; (3) Select the decomposition method with the maximum benefit value for each network layer by selecting the function, and record the corresponding optimal decomposition rank.
[0013] Furthermore, in step S2, based on the idea of reinforcement learning, multiple low-rank decomposition methods are deployed in each layer of the neural network to compress the weight parameters of the layer, and the accuracy loss and compression rate are taken as the benefit value Q. By selecting the function To realize the selection of decomposition methods for each network layer, the criterion for selecting the function is the size of the benefit value Q obtained by each decomposition method; in addition, the benefit value Also used as a tool to select the best decomposition rank The objective solution formula is: when the rank r selected by traversal makes the benefit value When it reaches the maximum, it is the optimal decomposition rank. Therefore, when the best decomposition method is selected, the best decomposition rank corresponding to the decomposition method is also obtained at the same time.
[0014] Preferably, in step S3, the compression rate in the objective function paradigm is expressed by constructing a compression cost function in the paradigm using a functional relationship of the decomposition rank; The model accuracy in the objective function paradigm is represented by constructing an error cost function; The method for achieving the best of both is: by optimizing and solving the target paradigm, the accuracy and compression rate of the compression model obtained can be comprehensively optimized.
[0015] Furthermore, in step S3, compression of the neural network model will lead to a decrease in the accuracy of the model. The greater the compression rate, the lower the accuracy. Some commonly used neural network compressions only consider the accuracy of the model or the compression rate of the model separately, so it is impossible to optimize both at the same time. Furthermore, in step S3, the reinforcement learning update of the model parameters and the iterative solution of the decomposition and compression of the model parameters are performed until the entire objective function paradigm converges, and finally the target compression model is obtained, and the compression model can be directly deployed without additional retraining or precision adjustment, thereby improving the practicality and deployment efficiency of the compression model; Furthermore, in step S3, according to the objective function paradigm, a stochastic gradient descent algorithm is used to implement the model weight parameters Updates; Preferably, in step S4, the step of first decomposing and compressing and then training the model includes: (1) The weight parameters of each network layer in the model are first decomposed and compressed using multiple decomposition methods deployed; (2) Reinforcement learning is used to select the best decomposition method for each network layer; (3) Then, the objective function paradigm of jointly optimizing accuracy and compression cost is used to train the entire decomposed and compressed model to improve model accuracy; (4) Repeat the iteration until the compression model converges; The compression model is implemented by direct deployment without the need for additional retraining or precision adjustment, thereby improving the practicality and deployment efficiency of the compression model.
[0016] Preferably, in the training step (4) of first decomposing and compressing and then training the model, the convergence of the compression model is achieved by iteratively solving the following two problems: (1) Reinforcement learning update of model parameters; (2) Decomposition and compression of model parameters.
[0017] Furthermore, in step S4, in the framework of reinforcement learning, in order to optimize the efficiency of model training, this paper designs a mode in which the network layer weight parameters are first decomposed and compressed and then trained; Furthermore, in step S4, the weight parameters of each network layer in the model are first decomposed and compressed using the two deployed low-rank decomposition methods, and after reinforcement learning selects the best decomposition method for each network layer, the entire decomposed and compressed model is trained using the objective function paradigm of jointly optimizing accuracy and compression cost to improve model accuracy, and the iteration is repeated until the compression model converges; Furthermore, in step S4, the reinforcement learning compression module and the parameter updating module are performed alternately, and the iterations are repeated until the compression model converges; Furthermore, in step S4, the model compressed by the method can be directly deployed without additional retraining or precision adjustment, thereby improving the practicality and deployment efficiency of the compressed model; The present invention also provides an application system of a multi-decomposition compression method of an artificial intelligence model based on reinforcement learning, comprising the following modules: (1) A pre-trained model parameter acquisition module, used to obtain the weight parameters of each network layer in the pre-trained neural network model; (2) Reinforcement learning compression module, which is used to build a joint application framework of multiple decomposition methods based on reinforcement learning, automatically select the most suitable low-rank decomposition method for each network layer and determine the optimal decomposition rank for low-rank decomposition; (3) Parameter update module, used to update the weight parameters in the model , construct an objective function paradigm for the joint model accuracy and compression cost, and update the model weight parameters through this function paradigm to achieve the best compression model accuracy and compression rate; (4) Iterative optimization module: It adopts an iterative optimization strategy of first decomposing and compressing and then training and learning. It trains and generates a compressed model that can be directly deployed, thereby optimizing the training efficiency of the reinforcement learning-based model.
[0018] Furthermore, the reinforcement learning module configures multiple low-rank decomposition methods for each network layer according to the model characteristics based on the reinforcement learning idea and compresses the weight parameters of the layer, adopts the benefit value as the unified selection standard for each decomposition method, selects the most suitable low-rank decomposition method for each network layer through the selection function and determines the optimal decomposition rank, thereby realizing low-rank decomposition of each network layer; Furthermore, the pre-training model parameter acquisition module is connected to the reinforcement learning compression module, the reinforcement learning compression module is connected to the parameter updating module, and the iterative optimization module is connected to the reinforcement learning compression module and the parameter updating module respectively; Furthermore, this method was tested on five models through comparative experiments, using two data sets, compared with two single compression algorithms, and using seven indicators as comparison criteria. The effectiveness of this method was demonstrated through experimental test results; Furthermore, in the comparative experiment, the five models include: LeNet-300, LeNet-5, ResNet-20, ResNet-32 and Vgg-16. The three algorithms include: LMFBRL-Selected, SVD-Selected and MF-Selected. Among them, LMFBRL-Selected is this method, SVD-Selected means that the entire model only uses a fixed SVF decomposition, and MF-Selected means that the entire model only uses a fixed MF decomposition. The two data sets include: MNIST and CIFAR-10. The seven parameter indicators include: Params, FLOPs, Accuracy (ACC), , , and Among them, Params represents the number of model parameters; FLOP represents the amount of model calculation; Accuracy represents the accuracy of the model, which can also be represented by ACC; Indicates the compression ratio of model parameters; Indicates the model calculation compression ratio; It represents the unit precision loss of the model's parameter compression ratio, reflecting the effectiveness of parameter compression; It represents the unit precision loss of the model's computational compression ratio, reflecting the effectiveness of computational compression.
[0019] Furthermore, the experimental platform of the comparative experiment is equipped with NVIDIA GeForce RTX 4070 GPU, 32GB memory, and the system environment is Windows 11. The development framework uses Python version 3.7 and PyTorch version 1.10.2. In the experiment, the parameters is set to and The algorithm was trained for 40 iterations, initially The value is 1.09. For the SGD algorithm, the learning rate is set to 0.09, the decay rate is set to 0.98, and the momentum is set to 0.9.
[0020] The beneficial technical effects of the present invention are: (1) In order to select appropriate low-rank decomposition methods for each network layer of the AI model, the present invention draws on the idea of reinforcement learning and constructs a low-rank decomposition method selection framework based on reinforcement learning. First, a benefit value with a clearly defined judgment standard is set to select the best decomposition method, and two low-rank decomposition methods, singular value decomposition SVD and conventional matrix decomposition MF, are used in the weight parameters of each layer for decomposition. By comparing the benefit values brought by the models obtained by different decomposition methods, the method that can achieve the maximum benefit value is selected, and then this decomposition method is determined to be the optimal decomposition method for the weight parameters of the network layer; (2) In the framework of reinforcement learning, in order to balance the relationship between the compression ratio and model accuracy of multiple decomposition methods, this invention proposes an objective function paradigm that comprehensively optimizes the accuracy and compression cost as an innovative solution to this problem. According to the functional relationship of the rank r, the spatial cost function in the paradigm is constructed to represent the compression ratio, and the error cost function is constructed to represent the accuracy of the model. By optimizing the spatial cost function and the error cost function in an alternating manner, the optimal decomposition rank r for low-rank matrix decomposition can be obtained; (3) In order to optimize the efficiency of model training under the framework of reinforcement learning, the present invention designs a mode in which the network layer parameter matrix is first decomposed and compressed before training. The weight parameters of each network layer of the model are first decomposed and compressed using various decomposition methods. After reinforcement learning selects the best decomposition method for each network layer, the entire decomposed and compressed model is trained to improve the model accuracy, and then iterated repeatedly until the compression model converges. In addition, the obtained compression model can be directly deployed without additional retraining or accuracy adjustment, thereby improving the practicality and deployment efficiency of the compression model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic flow chart of the method of the present invention; Figure 2 It is a schematic diagram of system module connection of the present invention. DETAILED DESCRIPTION
[0022] The present invention is described in detail below in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] Embodiment 1: This embodiment provides a multi-decomposition compression method based on an artificial intelligence model of reinforcement learning. The specific process is as follows: Figure 1 As shown, the following steps are included: Step S1, obtaining the weight parameters of each network layer in the pre-trained model.
[0024] In this embodiment, the decomposition method adopted is matrix decomposition, so the obtained weight parameters are converted into weight parameter matrices until the training is completed and then restored.
[0025] Step S21, construct a joint application framework of multiple decomposition methods based on reinforcement learning: arrange multiple suitable low-rank decomposition methods for each network layer according to the characteristics of the model, and set a unified benefit value as a selection criterion for each low-rank decomposition method based on the idea of reinforcement learning.
[0026] In this embodiment, the weight parameters obtained are used to construct a joint application framework of multiple decomposition methods based on reinforcement learning. A variety of low-rank decomposition methods include: singular value decomposition SVD and conventional matrix decomposition MF. Reinforcement learning is based on the feedback of the current environment to adjust each step of behavior and ultimately maximize the overall benefit. This embodiment adopts the reinforcement learning Q-learning algorithm based on the greedy strategy, which can be expressed by the following specific formula: ; in is the state subspace, is the set of decomposition methods selected for each network layer, is the action subspace, which represents our comparison operation of each network layer decomposition method. represents the optimal action value function, which is the overall compression model benefit after selecting the decomposition method. Indicates the maximum return of the current state, that is, the benefit value, which is the benefit after selecting the decomposition method for the current network layer. represents the maximum expected profit, is a hyperparameter. The entire model is a state space , select the decomposition method for each network layer in the model as action space A, we need to obtain the benefit value based on each decomposition method To compare, select the decomposition method with the highest value according to the greedy strategy, and finally make the profit of the whole model Reach the maximum.
[0027] For example, for a Lenet5 model with four network layers, when this method is used for compression, it is determined through calculation that the benefit of SVD decomposition in the first layer is greater than that of MF decomposition, but the benefit of MF decomposition in the second to fourth layers is greater than that of SVD decomposition. Therefore, for the compression of the Lenet5 model, the final compression scheme selected is: first layer - SVD decomposition, second layer - MF decomposition, third layer - MF decomposition, fourth layer - MF decomposition.
[0028] Step S22: Select the decomposition method and the decomposition rank: Select the most suitable decomposition method for each network layer from multiple decomposition methods, and obtain the optimal decomposition rank r of the corresponding decomposition method, so as to perform low-rank decomposition on the weight parameters of each network layer. In this embodiment, the selection criterion for the decomposition method takes the negative value of the sum of the accuracy loss and the compression cost, denoted as the benefit value Q. When the sum of the accuracy loss and the compression cost is smaller, it means the greatest benefit. Therefore, the value benefits of different decomposition methods for each layer are expressed by the following formula: ; where represents the benefit value obtained by using the m-th decomposition method in the i-th network layer, represents the accuracy loss caused by decomposition compression, represents the best compression parameter obtained by using the m-th decomposition method for the weight parameters of the i-th network layer. At this time, the corresponding decomposition rank is , represents the compression cost.
[0029] This method uses two low-rank matrix decomposition methods for selection, namely singular value decomposition SVD and conventional matrix decomposition MF. The benefit values corresponding to the two decomposition methods in the i-th network layer are respectively expressed as and , and the cost functions can be respectively expressed as follows: ; ; It should be further explained that: For the conventional matrix decomposition MF, the original weight parameter matrix A(m×n) is decomposed into two low-rank matrices U(m×r) and V(n×r), and the number of parameters is compressed from the original (m×n) to r×(m + n), where r << min(m, n); for the singular value decomposition SVD, the original weight parameter matrix A(m×n) is decomposed into three low-rank matrices U(m×r), S(r×r) and V(r×n), and the number of parameters is compressed from the original (m×n) to r×(m + r + n), where r << min(m, n).
[0030] The cost function is the number of parameters after decomposing the parameters of the i-th network layer. The calculated value corresponds to the selected rank r and the decomposition method, which represents the compression effect of this decomposition method and is part of the benefit value.
[0031] where is the selected rank of decomposition, m and n are the two dimensions of the weight parameter matrix. Therefore, the decomposition method selection function of the model can obtain the following selection function expression according to the benefit value Q: ; Select two sets of parameters in a function and , the corresponding values are two integers: {0,1}, where if the parameter =1, it means The network layer selects full rank decomposition. Otherwise, if the parameter =0, it means The network layer does not select full rank decomposition, and the parameters It indicates whether the SVD decomposition method is selected. Therefore, the constraints of the two are: + .
[0032] In addition, the benefit value Also used as a tool to select the best decomposition rank The objective solution formula is: when the rank r selected by traversal makes the benefit value When it reaches the maximum, it is the optimal decomposition rank. Therefore, when the best decomposition method is selected, the best decomposition rank corresponding to the decomposition method is also obtained at the same time.
[0033] Step S3, jointly optimizing model accuracy and compression cost: constructing an objective function of the joint model accuracy and compression cost, using this objective function to update model parameters, and optimizing compression model accuracy and compression rate at the same time; In this embodiment, for a The neural network model of the layer, the weight parameter matrix of the model is expressed as ,in , Expressed as This method takes into account the mutual influence between the model accuracy and compression cost mentioned above, and constructs the following problem paradigm to weigh the relationship between the two: ; ; In the above formula is the loss function, which is used to express the accuracy of the model, and is the cost function, which is used to represent the number of parameters, thus approximating the compression rate of the model. In this problem paradigm, represents the maximum rank of the parameter matrix of each layer, .in is a hyperparameter used to weigh the importance between compression ratio and model accuracy. ,when The closer it is to 0, the more the model focuses on improving the accuracy of the model. Conversely, the closer it is to 1, the more it focuses on improving the compression rate of the model.
[0034] This method involves the choice of two decomposition methods, namely singular value decomposition SVD and conventional matrix decomposition MF, so the cost function in the problem paradigm is There are many situations that need to be considered. There are two situations that need to be selected for the cost function of each layer, as shown in the following formula: ; ; Combined with the selection function designed previously , and the cost function and The problem paradigm can be expressed as the following objective function paradigm: ; .
[0035] Step S4, alternately optimize compression model parameters: The model training adopts an iterative optimization strategy of first decomposing and compressing and then training and learning. The compression model obtained by training can be directly deployed.
[0036] In this embodiment, the intermediate variable , and introduce constraints , reformulate the objective function as a constrained optimization problem: ; ; This method uses a penalty method to solve the constrained optimization problem mentioned above. Specifically, it uses a quadratic penalty method and implements the augmented Lagrangian method (which works similarly but introduces Lagrange multipliers of the same dimension ). On this basis, the problem in the above formula is rewritten as the driving penalty parameter The objective function paradigm when : ; ; in, is the penalty parameter and increases with the number of model iterations, satisfying . is the Lagrange multiplier vector, whose dimension is the same as the weight matrix To solve the above objective function paradigm, this method designs an alternating optimization method. and , split the objective function paradigm into two sub-problems and solve them alternately.
[0037] The first sub-problem involves learning and updating the model parameters. The parameter matrix of each network layer is The corresponding compression matrix is a known condition. To this end, this method uses the stochastic gradient descent (SGD) algorithm to optimize and solve according to the loss function of the model and the constraints imposed by the compressed matrix on the original matrix. The solution formula is as follows: ; In the above formula, represents the updated parameters, represents the original parameters, is the learning rate, and the fractional term after it represents the parameters This formula realizes the update of model parameters and constitutes the learning and optimization process of the entire model.
[0038] The second sub-problem involves the decomposition and compression of model parameters. The parameter matrix of each network layer is is a known condition. Based on the idea of reinforcement learning, this method designs a multi-decomposition method selection framework, taking the benefit value Q as the standard, and selects the most appropriate decomposition method and decomposition rank r for the weight parameters of each network layer. Therefore, the solution to this problem is to solve the benefit value Q. Since it involves the selection of two decomposition methods, the benefit value solution formulas of the two methods are expressed separately, among which the benefit value solution formula of MF decomposition is expressed as: ; ; ; The formula for solving the benefit value of SVD decomposition is expressed as: ; ; ; According to the above benefit value solving formula, the corresponding benefit value can be obtained and , and simultaneously solve the optimal decomposition rank , combined with the previous selection function , the optimal compression matrix of the i-th network layer can be expressed as: In this way, the corresponding decomposition method, the parameter matrix after decomposition and compression, and the optimal decomposition rank are obtained.
[0039] This method iteratively solves two sub-problems until the entire objective function paradigm converges, and finally obtains the target compression model. The compression model can be directly deployed without additional retraining or accuracy adjustment, thereby improving the practicality and deployment efficiency of the compression model.
[0040] Furthermore, this method was tested on five models using two data sets, compared with two single compression algorithms, and seven parameter indicators were used as comparison criteria. The effectiveness of this method was demonstrated through experimental test results.
[0041] Five models were tested in this experiment: LeNet-300, LeNet-5, ResNet-20, ResNet-32, and Vgg-16.
[0042] Three algorithms are used: LMFBRL-Selected, SVD-Selected and MF-Selected. LMFBRL-Selected is this method, SVD-Selected means that the entire model only uses a fixed SVF decomposition, and MF-Selected means that the entire model only uses a fixed MF decomposition.
[0043] Two datasets were used: MNIST and CIFAR-10.
[0044] The parameter index consists of 7 indicators: Params, FLOPs, Accuracy (ACC), , , and Among them, Params represents the number of model parameters; FLOP represents the amount of model calculation; Accuracy represents the accuracy of the model, which can also be represented by ACC; , represents the model parameter compression ratio, is the number of parameters of the original model, is the number of parameters of the compression model; , represents the model calculation compression ratio, is the computational effort of the original model, is the computational effort of the compression model; , represents the unit precision loss of the model's parameter compression ratio, reflecting the effectiveness of parameter compression. is the original model accuracy, is the compression model accuracy; , which represents the unit precision loss of the model’s computational compression ratio, reflects the effectiveness of computational compression.
[0045] The experimental platform is equipped with NVIDIA GeForce RTX 4070 GPU, 32GB memory, and the system environment is Windows 11. The development framework uses Python version 3.7 and PyTorch version 1.10.2. In the experiment, the parameters is set to and The algorithm was trained for 40 iterations, initially The value is 1.09. For the SGD algorithm, the learning rate is set to 0.09, the decay rate is set to 0.98, and the momentum is set to 0.9.
[0046] Table 1 Comparative analysis of model compression methods
[0047] Comparing the ACC indicator data in the data table, this method LMFBRL-Selected achieved the highest accuracy in 4 out of 5 models, and the difference from the highest accuracy in the ResNet-20 model was only a small 0.1%. This shows that this method can adaptively select layer by layer and has a significant advantage in maintaining model accuracy during compression because it considers the unique structural characteristics of each layer instead of using a fixed method. This experiment proves the effectiveness of the multi-decomposition compression method for artificial intelligence models based on reinforcement learning, which is a reinforcement learning-based method that can adaptively select the optimal decomposition method for each layer in the AI model.
[0048] In addition, this experiment verifies the effectiveness of the proposed joint paradigm, which optimizes the accuracy and compression cost of the model. The method LMFBRL-Selected uses the objective function paradigm of joint model accuracy and compression cost to optimize the solution, while SVD-Selected and MF-Selected only use the method of minimizing the compression cost. and Two indicator data, LMFBRL-Select compresses performance indicators in all models and All of them reached the highest value, reflecting their superior balance between accuracy and compression rate. The results show that the objective function paradigm of the joint model accuracy and compression cost in this method is very effective, achieving the best compression performance of the AI model. However, VD-Selected and MF-Selected only optimize the compression cost and cannot achieve a superior balance between the best accuracy and compression rate.
[0049] Embodiment 2: This embodiment provides an application system of a multi-decomposition compression method based on an artificial intelligence model of reinforcement learning, and the specific structure diagram is as follows: Figure 2 As shown, it includes the following modules: (1) A pre-trained model parameter acquisition module is used to obtain the weight parameters of each network layer in the pre-trained neural network model. This method uses matrix decomposition, so the obtained weight parameters will be converted into a weight parameter matrix until the training is completed and then restored; (2) Reinforcement learning compression module, which is used to construct a joint application framework of multiple decomposition methods based on reinforcement learning to select the decomposition method of each network layer; this method is based on the idea of reinforcement learning Q-learning, and deploys multiple low-rank decomposition methods at each neural network layer to compress the weight parameters of that layer, and takes the accuracy loss and compression rate as the benefit value Q, by selecting the function To realize the selection of decomposition methods for each network layer, the criterion for selecting the function is the size of the benefit value Q obtained by each decomposition method; in addition, the benefit value Also used as a tool to select the best decomposition rank The objective solution formula is: when the rank r selected by traversal makes the benefit value When it reaches the maximum, it is the optimal decomposition rank. , so when the best decomposition method is selected, the best decomposition rank corresponding to the decomposition method is also obtained at the same time; (3) Parameter update module, used to update the weight parameters in the model ; This method jointly optimizes the objective function paradigm of accuracy and compression cost. The compression cost function in the paradigm represents the compression rate, and the error cost function is constructed to represent the accuracy of the model; according to the objective function paradigm, the stochastic gradient descent algorithm is adopted to realize the model weight parameter Updates; (4) An iterative optimization module is used to solve the objective function paradigm and optimize the efficiency of model training. This method designs a mode in which the network layer weight parameters are first decomposed and compressed and then trained and updated. The reinforcement learning compression module and the parameter update module are performed alternately, and the iteration is repeated until the compression model converges. The model compressed by this method can be directly deployed without additional retraining or accuracy adjustment, thereby improving the practicality and deployment efficiency of the compression model.
[0050] The pre-training model parameter acquisition module is connected to the reinforcement learning compression module, the reinforcement learning compression module is connected to the parameter updating module, and the iterative optimization module is connected to the reinforcement learning compression module and the parameter updating module respectively.
[0051] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and implementation modes, and can be fully applicable to various fields suitable for the present invention. For those familiar with the art, for those of ordinary skill in the art, various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to specific details.
Claims
1. A multi-decomposition compression method for an artificial intelligence model based on reinforcement learning, characterized in that: The following steps are involved: Step S1, obtaining the weight parameters of the pre-trained model: extracting the weight parameters of each network layer, wherein the weight parameters are expressed in a matrix form; Step S2, optimizing multiple decomposition methods and ranks based on reinforcement learning: constructing a joint application framework of multiple decomposition methods based on reinforcement learning, automatically selecting the most suitable low-rank decomposition method for each network layer and determining the optimal decomposition rank for low-rank decomposition; Step S3, jointly optimizing the model accuracy and compression cost: constructing an objective function paradigm for jointly optimizing the model accuracy and compression cost, and updating the model weight parameters through the function paradigm to achieve the best compression model accuracy and compression rate; Step S4, alternately optimize compression model parameters: adopt an iterative optimization strategy of first decomposing and compressing and then training and learning, train and generate a compression model that can be directly deployed, and optimize the training efficiency of the model based on reinforcement learning.
2. The multi-decomposition compression method according to claim 1, characterized in that: Step S2 includes: S21. Construct a joint application framework of multiple decomposition methods based on reinforcement learning: After obtaining the weight parameters of the pre-trained model, configure multiple low-rank decomposition methods for each network layer according to the model characteristics, and set a unified benefit value as the selection criterion for each decomposition method based on reinforcement learning; S22. Select decomposition method and decomposition rank: select the most suitable low-rank decomposition method for each network layer according to the selection criteria and determine the optimal decomposition rank, and perform low-rank decomposition on the weight parameters of each network layer.
3. The multi-decomposition compression method according to claim 2, characterized in that: The low-rank decomposition method in step S21 includes: singular value decomposition SVD and conventional matrix decomposition MF; The singular value decomposition SVD decomposes the weight matrix A into two low-rank matrices U and V, and the number of parameters is compressed from m×n to r×(m+n), where r is much smaller than the minimum value of m and n; The conventional matrix decomposition MF decomposes the weight matrix A into three low-rank matrices U, S and V, and the number of parameters is compressed from m×n to r×(m+r+n), where r is much smaller than the minimum value of m and n.
4. The multi-decomposition compression method according to claim 2, characterized in that: The specific implementation of configuring multiple low-rank decomposition methods for each network layer in step S21 includes: Based on reinforcement learning, a variety of low-rank decomposition methods are deployed in each neural network layer and the weight parameters of the layer are compressed; The fully connected layer can be directly decomposed; For the tensor parameters in the convolution layer, matrix conversion is required.
5. The multi-decomposition compression method according to claim 2, characterized in that: The selection of the most suitable low-rank decomposition method in step S22 is achieved by selecting a function, and the selection criterion is the benefit value.
6. The multi-decomposition compression method according to claim 2, characterized in that: The steps of selecting the most suitable low-rank decomposition method and determining the optimal decomposition rank in step S22 are as follows: (1) The negative of the sum of the compression accuracy loss and compression rate of the weight parameters of this layer is taken as the benefit value Q. The larger the value, the better the performance. (2) For each decomposition method of each network layer, traverse the possible ranks respectively, analyze the benefit value Q, and the rank with the maximum Q value is the optimal decomposition rank of the network layer, and the maximum benefit value Q is the maximum benefit value of the decomposition method; (3) Select the decomposition method with the maximum benefit value for each network layer by selecting the function, and record the corresponding optimal decomposition rank.
7. The multi-decomposition compression method according to claim 1, characterized in that: In step S3, the compression rate in the objective function paradigm is expressed by constructing a compression cost function in the paradigm using a functional relationship of the decomposition rank; The model accuracy in the objective function paradigm is represented by constructing an error cost function; The method for achieving the best of both compression model accuracy and compression rate is: by optimizing and solving the target paradigm, the accuracy and compression rate of the compression model obtained can be comprehensively optimized.
8. The multi-decomposition compression method according to claim 1, characterized in that: In step S4, the training step of first decomposing and compressing and then training the model includes: (1) The weight parameters of each network layer in the model are first decomposed and compressed using multiple decomposition methods deployed; (2) Reinforcement learning is used to select the best decomposition method for each network layer; (3) Then, the objective function paradigm of jointly optimizing accuracy and compression cost is used to train the entire decomposed and compressed model to improve model accuracy; (4) Repeat the iteration until the compression model converges; The compression model is implemented by direct deployment without the need for additional retraining or precision adjustment, thereby improving the practicality and deployment efficiency of the compression model.
9. The multi-decomposition compression method according to claim 8, characterized in that: In step (4), the convergence of the compression model is achieved by iteratively solving the following two problems: (1) Reinforcement learning update of model parameters; (2) Decomposition and compression of model parameters.
10. An application system based on the multi-decomposition compression method according to claim 1, characterized in that: Includes the following modules: (1) Pre-trained model parameter acquisition module: used to obtain the weight parameters of each network layer in the pre-trained neural network model; (2) Reinforcement learning compression module: used to build a joint application framework of multiple decomposition methods based on reinforcement learning, automatically select the most suitable low-rank decomposition method for each network layer and determine the optimal decomposition rank for low-rank decomposition; (3) Parameter update module: used to update the weight parameters in the model , construct an objective function paradigm for the joint model accuracy and compression cost, and update the model weight parameters through this function paradigm to achieve the best compression model accuracy and compression rate; (4) Iterative optimization module: It adopts an iterative optimization strategy of first decomposing and compressing and then training and learning to train and generate a compressed model that can be directly deployed, thereby optimizing the training efficiency of the reinforcement learning-based model.
Citation Information
Patent Citations
Vehicle end model lightweight method based on matrix weighted low-rank decomposition
CN117973476A
System and method for compressing deep learning models by low rank decomposition using reinforcement learning
CN118786439A