Image Domain Model Compression Method, Device, Equipment and Medium Based on Rank-Preserving Decomposition and Knowledge Distillation
Through the method of rank-keeping decomposition and knowledge distillation, the image processing model is trained in Cronek decomposition and knowledge distillation, which solves the problem of excessive storage and computing resources demand in lightweight equipment, and achieves efficient compression and performance maintenance of the model.
Patent Information
- Application Number
- CN202510413975.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When existing image processing models are deployed in lightweight devices, there is a problem of excessive demand for storage space and computing resources, and the traditional singular value decomposition method cannot effectively retain the global information of the model, resulting in compression rate and performance losses.
The pre-trained teacher image processing model is decomposed by Cronek decomposition, and the sample images are used for knowledge distillation to obtain the compressed student image processing model to maintain the expression ability and performance of the model.
On the premise of maintaining the performance of the image processing model, greatly compress model parameters, save storage space and computing resources, and realize the stable deployment of lightweight devices.
Smart Images

Figure CN119942275B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to an image domain model compression method, device, equipment and medium based on rank-preserving decomposition and knowledge distillation. Background Art
[0002] In recent years, the application of image processing models has become more and more extensive, which is attributed to the increasingly better effects of image processing models on image processing (such as target tracking, target detection, image classification, etc.). However, the improvement of the current image processing model effect is inseparable from the large-scale model parameters in the image processing model, which means that the equipment deploying the image processing model needs to have large-capacity hardware storage space and / or sufficient computing resources, making it impossible to deploy the image processing model in lightweight devices, limiting the application and development of the image processing field in lightweight devices.
[0003] Currently, most model compression methods based on SVD (Singular Value Decomposition) decompose the weight matrix in the image processing model to reduce the number of model parameters and the amount of calculation, further achieving the purpose of saving storage space and computing resources of the deployed equipment. However, the compression effect of the SVD method depends on the retention strategy of singular values: in practical applications, retaining large singular values will limit the compression ratio of the model, especially when processing complex image processing models, SVD cannot effectively approximate low-rank representations in high-dimensional space. This means that although SVD can reduce the number of model parameters, it often cannot achieve the desired high compression rate and low computational overhead. Although SVD approximates by retaining important singular values, it may introduce performance loss because the SVD method itself does not consider the training objectives of the image processing model. In addition, SVD may not be able to retain the global information of the network, and it is easy to lose global information at high compression rates, thereby affecting the image processing accuracy of the image processing model. As a result, the current method based on low-rank decomposition (i.e., SVD) has the problems of rank limitation and model performance loss, and cannot compress the image processing model parameters to achieve model deployment on lightweight devices while ensuring the performance of the image processing model. Therefore, how to compress the model parameters of the image processing model while maintaining the expressive power of the image processing model and saving the storage space and computing resources of the model deployment device is a technical problem to be solved by the present invention. Summary of the invention
[0004] Based on the above technical problems, the present invention provides an image domain model compression method, device, equipment and medium based on rank-preserving decomposition and knowledge distillation, aiming to effectively maintain the expressiveness of the original image processing model, significantly compress and reduce the storage requirements of the model, realize lightweight deployment of the image processing model, and maintain stable model performance.
[0005] In the first aspect of the present invention, a method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation is provided. The method includes:
[0006] Obtain the device parameter values of the target device for the student image processing model to be deployed and compressed. The device parameter values include at least one of the following: storage space size, operation speed;
[0007] Determine the target number of parameters of the student image processing model according to the device parameter values;
[0008] Perform Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and use the model after Kronecker decomposition as the student image processing model to be trained;
[0009] Use sample images and take the pre-trained teacher image processing model as the learning target to perform knowledge distillation training on the student image processing model to be trained, and obtain the compressed student image processing model. The pre-trained teacher image processing model is used for any one of the following: image classification, image segmentation, object detection;
[0010] Input the image to be processed into the compressed student image processing model to obtain the image processing result. The image processing result is any one of the following: image classification result, image segmentation result, object detection result.
[0011] In the second aspect of the present invention, an apparatus for compressing an image domain model based on rank-preserving decomposition and knowledge distillation is provided. The apparatus includes:
[0012] A parameter acquisition module for obtaining the device parameter values of the target device for the student image processing model to be deployed and compressed. The device parameter values include at least one of the following: storage space size, operation speed;
[0013] A parameter determination module for determining the target number of parameters of the student image processing model according to the device parameter values;
[0014] A model decomposition module for performing Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and using the model after Kronecker decomposition as the student image processing model to be trained;
[0015] A model training module for using sample images and taking the pre-trained teacher image processing model as the learning target to perform knowledge distillation training on the student image processing model to be trained, and obtaining the compressed student image processing model. The pre-trained teacher image processing model is used for any one of the following: image classification, image segmentation, object detection;
[0016] An image processing module, configured to input an image to be processed into the compressed student image processing model, and obtain an image processing result, where the image processing result is any one of the following: an image classification result, an image segmentation result, and an object detection result.
[0017] In a third aspect of the embodiments of the present invention, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the image domain model compression method based on rank-preserving decomposition and knowledge distillation according to the first aspect of the embodiments of the present invention.
[0018] In a fourth aspect of the present invention, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the image domain model compression method based on rank-preserving decomposition and knowledge distillation according to the first aspect of the embodiments of the present invention.
[0019] Through the image domain model compression method based on rank-preserving decomposition and knowledge distillation in this embodiment, after determining the target number of parameters according to the device parameter values of the target device based on the compressed student image processing model to be deployed, perform Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and use the model after performing Kronecker decomposition as the student image processing model to be trained. In this way, in this embodiment, through the decomposition method of Kronecker decomposition, a rank-preserving transformation is performed on the original matrix in the pre-trained teacher image processing model. The decomposed matrix is usually equal to or close to the rank of the original matrix, so that the information in the original matrix can be effectively retained, and the performance loss of the pre-trained teacher image processing model can be avoided. Then, in this embodiment, using the sample image, with the pre-trained teacher image processing model as the learning target, perform knowledge distillation on the student image processing model to be trained after rank-preserving decomposition to obtain the compressed student image processing model, so as to ensure that the compressed student image processing model can maintain a performance similar to that of the teacher model. Finally, while maintaining the expression ability of the teacher image processing model, the model parameters of the student image processing model are compressed, saving the storage space and computing resources of the target device of the compressed student image processing model to be deployed, and realizing the lightweight deployment of the image processing model with stable model performance. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0021] Figure 1It is a flowchart of steps of an image domain model compression method based on rank-preserving decomposition and knowledge distillation shown in an embodiment of the present invention;
[0022] Figure 2 It is a schematic diagram of the operation of a Kronecker product shown in an embodiment of the present invention;
[0023] Figure 3 It is a schematic diagram of the overall process of an image domain model compression method based on rank-preserving decomposition and intermediate layer knowledge distillation shown in an embodiment of the present invention;
[0024] Figure 4 It is a structural block diagram of an image domain model compression device based on rank-preserving decomposition and knowledge distillation provided in an embodiment of the present invention;
[0025] Figure 5 It is a schematic diagram of an electronic device shown in an embodiment of the present invention. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] Please refer to Figure 1 , Figure 1 It is a flowchart of steps of an image domain model compression method based on rank-preserving decomposition and knowledge distillation shown in an embodiment of the present invention. As Figure 1 shown, the image domain model compression method based on rank-preserving decomposition and knowledge distillation provided in this embodiment at least includes the following steps:
[0028] Step S11: Obtain the device parameter values of the target device for deploying the compressed student image processing model, where the device parameter values include at least one of the following: storage space size, operation speed.
[0029] In this embodiment, the target device is used to deploy the compressed student image processing model, and the target device can be any hardware device, such as a server, a computer, a mobile terminal (such as a mobile phone, an IPAD), etc., and there is no limitation thereto. The purpose of this embodiment is to compress the pre-trained teacher image processing model to obtain the compressed student image processing model, so as to deploy the compressed student image processing model on the target device and achieve the lightweight deployment of the compressed student image processing model.
[0030] First, this embodiment can obtain the device parameter values of the target device for deploying the compressed student image processing model. The device parameter values include at least one of the following: storage space size, operation speed. Among them, the device parameter values in this embodiment can refer to the factory-rated device parameter values of the target device, and can also refer to the remaining device parameter values of the target device (such as the available memory capacity), and there is no limitation on this.
[0031] Step S12: Determine the target number of parameters of the student image processing model according to the device parameter values.
[0032] In this embodiment, the target number of parameters of the compressed student image processing model can be determined according to the device parameter values of the target device for deploying the compressed student image processing model. The target number of parameters in this embodiment is the number of model parameters determined according to the device parameter values, so as to meet the device parameter values of the target device when the compressed student image processing model is deployed on the target device subsequently.
[0033] Step S13: Perform Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and use the model after performing Kronecker decomposition as the student image processing model to be trained.
[0034] In this embodiment, after obtaining the target number of parameters, Kronecker decomposition can be performed on the pre-trained teacher image processing model according to the target number of parameters to obtain the model after performing Kronecker decomposition, and the model after performing Kronecker decomposition is used as the student image processing model to be trained. The "teacher image processing model" and "student image processing model" in this embodiment correspond to the "teacher model" and "student model" in knowledge distillation. Among them, the pre-trained teacher image processing model is a pre-trained image processing model, and the pre-trained teacher image processing model is used for any one of the following: image classification, image segmentation, object detection, object tracking. The student image processing model to be trained is a to-be-trained image processing model that learns the knowledge of the pre-trained teacher image processing model, and the student image processing model to be trained is the model after performing Kronecker decomposition on the pre-trained teacher image processing model.
[0035] Step S14: Use the sample images and take the pre-trained teacher image processing model as the learning target to perform knowledge distillation training on the student image processing model to be trained, and obtain the compressed student image processing model. The pre-trained teacher image processing model is used for any one of the following: image classification, image segmentation, object detection.
[0036] After obtaining the student image processing model to be trained in this embodiment, knowledge distillation can be performed on the pre-trained teacher image processing model and the student image processing model to be trained: using the sample images, taking the pre-trained teacher image processing model as the learning target, performing knowledge distillation training on the student image processing model to be trained, and obtaining the compressed student image processing model.
[0037] Step S15: Input the image to be processed into the compressed student image processing model to obtain an image processing result, where the image processing result is any one of the following: an image classification result, an image segmentation result, and an object detection result.
[0038] In this embodiment, after obtaining the compressed student image processing model, the image to be processed can be input into the compressed student image processing model to obtain an image processing result; wherein, the image processing result of this embodiment is any one of the following: an image classification result, an image segmentation result, an object detection result, and an object tracking result. In an alternative embodiment, after obtaining the compressed student image processing model, the compressed student image processing model can be deployed on the target device, and the image to be processed can be input into the compressed student image processing model to obtain an image processing result.
[0039] In this embodiment, through the decomposition method of Kronecker decomposition, a rank-preserving transformation is performed on the original matrix in the pre-trained teacher image processing model. The decomposed matrix usually has the same or similar rank as the original matrix, so that the information in the original matrix can be effectively retained, and the performance loss of the pre-trained teacher image processing model can be avoided. Then, in this embodiment, using the sample images, taking the pre-trained teacher image processing model as the learning target, performing knowledge distillation on the student image processing model to be trained after rank-preserving decomposition, and obtaining the compressed student image processing model, so as to ensure that the compressed student image processing model can maintain a performance similar to that of the teacher model. Finally, while maintaining the expression ability of the teacher image processing model, the model parameters of the student image processing model are compressed, saving the storage space and computing resources of the target device for deploying the compressed student image processing model, and realizing the lightweight deployment of the image processing model with stable model performance.
[0040] Combining the above embodiments, in one embodiment, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, step S12 above may specifically include step S121 and step S122:
[0041] Step S121: When the device parameter values read from the target device are: available memory capacity (unit: MB / GB), the single-computation peak performance of the processor of the target device (e.g., the number of floating-point operations per second that can be processed, FLOPS), and the maximum response time required by the target device (e.g., image processing needs to be completed within 0.5 seconds), calculate the basic number of parameters of the student image processing model based on the device parameter values.
[0042] The basic number of parameters in this embodiment includes: parameter storage amount and upper limit of the number of parameters; among them, the determination method of the parameter storage amount is as follows:
[0043] Determine the parameter storage amount of the student image processing model based on the available memory capacity of the target device:
[0044] Each model parameter defaults to occupying 4 bytes (32-bit floating-point number). If the available memory capacity is M MB, then the parameter storage amount (i.e., the maximum number of parameters) of the student image processing model is: maximum number of parameters = (M×1024²×memory occupancy rate) / 4, where the memory occupancy rate is usually set to 60%-80% to reserve system resources. Example: When the available memory capacity is 4GB (4096MB) and allocated according to a 70% memory occupancy rate, it supports approximately 734 million parameters ((4096×1024²×0.7) / 4 ≈ 734,003,200).
[0045] The determination method of the upper limit of the number of parameters is as follows:
[0046] Determine the operation efficiency constraint of the student image processing model based on the single-computation peak performance of the processor of the target device:
[0047] First, calculate the maximum amount of computation allowed for a single inference (i.e., the tolerable amount of computation) according to the number of floating-point operations per second (P FLOPS) that the processor can process and the maximum response time T seconds required by the target device: tolerable amount of computation = P × T; then, determine the upper limit of the number of parameters according to the number of operations per parameter of the model (e.g., a certain network requires 10 operations per parameter on average): upper limit of the number of parameters = tolerable amount of computation / number of operations per parameter of the model.
[0048] Step S122: Determine the target number of parameters of the student image processing model based on the basic number of parameters of the student image processing model.
[0049] In this embodiment, the minimum value among the basic number of parameters can be determined as the target number of parameters of the student image processing model, that is, the minimum value between the value of the parameter storage amount and the value of the upper limit of the number of parameters is used as the target number of parameters of the student image processing model.
[0050] Combined with the above embodiments, in one implementation manner, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S13 may specifically include step S21 and step S22:
[0051] Step S21: Determine that the first pair of Kronecker matrix sizes is , and the second pair of Kronecker matrix sizes is .
[0052] Step S22: Perform Kronecker decomposition on the parameter matrix of the nth layer of the pre-trained teacher image processing model, obtain a group of Kronecker parameter matrices, and use them as the parameter matrix of the nth layer of the to-be-trained student image processing model.
[0053] In this embodiment, first, according to the target number of parameters, determine the first pair of Kronecker matrix sizes , and, determine the second pair of Kronecker matrix sizes . And, in this embodiment, perform Kronecker decomposition on the parameter matrix of the nth layer of the pre-trained teacher image processing model, obtain a group of Kronecker parameter matrices, and use the group of Kronecker parameter matrices as the parameter matrix of the nth layer of the to-be-trained student image processing model, so as to realize the Kronecker decomposition of the pre-trained teacher image processing model and obtain the model after Kronecker decomposition. Wherein, n is an integer greater than 0, and the group of Kronecker parameter matrices includes: a matrix and that conforms to the first pair of Kronecker matrix sizes, and, a matrix and that conforms to the second pair of Kronecker matrix sizes; .
[0054] In an example, the Kronecker decomposition of a matrix: The Kronecker product is a matrix operation that forms a block matrix from two matrices. Let the matrix , , then the Kronecker product of A and B, denoted as , is a block matrix, where each block is obtained by multiplying by the matrix B. As shown in Figure 2 , Figure 2 is a schematic diagram of the operation of a Kronecker product shown in an embodiment of the present invention. In Figure 2In it, the operation of the Kronecker product on a 2×2 matrix is shown. Decomposing the result of matrix multiplication into the Kronecker product can be considered as replacing the projection of the original linear space with a more constrained linear space, making the overall information expression more compact while retaining the information interaction between each data block. The inverse process of this process is the Kronecker decomposition. The Kronecker theorem states that given , , then any matrix can be expressed as the sum of the Kronecker products of several and .
[0055] In this embodiment, the obtained Kronecker parameter matrix group is expressed as the sum of the Kronecker products of two groups of and , so as to obtain an approximately Kronecker decomposition that is numerically accurate enough. When i takes 1 and 2, the Kronecker parameter matrix group obtained by performing Kronecker decomposition on the parameter matrix is: , where the matrices and the matrix have the size of the first pair of Kronecker matrix sizes , and the matrices and the matrix conform to the second pair of Kronecker matrix sizes .
[0056] The following gives the specific steps for performing Kronecker decomposition in this embodiment:
[0057] (a), Given the scale parameter to be decomposed , where , m and n are the sizes (i.e., the number of rows and columns) of the original matrix (i.e., the matrix to be decomposed).
[0058] (b), Perform a rearrangement operation on the original matrix . Scan the blocks of size in the original matrix one by one, then flatten these blocks into one dimension and splice them to form a matrix of size .
[0059] (c), Perform a rank-1 SVD decomposition on the matrix obtained in the above step.
[0060] (d), Divide the vector into 1 column according to every elements to form a matrix of size , denoted as , for the vector Divide into 1 column according to each element to form a matrix of size , denoted as .
[0061] (e), Let . For , repeatedly execute the operations of steps (b), (c) and (d) to obtain and .
[0062] (f), Obtain the numerical approximation of the Kronecker decomposition of matrix W in this embodiment.
[0063] In this embodiment, through Kronecker decomposition, on the premise of ensuring that the rank of the decomposed matrix is the same as or close to the original and better retaining the information of the original matrix, the number of parameters for storing this matrix is reduced from m×n to 2 ( ), realizing a substantial compression of the image processing model.
[0064] Combining the above embodiments, the present invention also provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the "using the sample image, taking the pre-trained teacher image processing model as the learning target, and performing knowledge distillation training on the to-be-trained student image processing model to obtain the compressed student image processing model" in step S14 above can specifically include step S31 and step S32:
[0065] Step S31: Input the sample image into the pre-trained teacher image processing model, and input the sample image into the to-be-trained student image processing model.
[0066] In this embodiment, the sample image can be input into the pre-trained teacher image processing model, and the sample image can be input into the to-be-trained student image processing model.
[0067] Step S32: Taking the output of the nth layer of the to-be-trained student image processing model learning the output of the nth layer of the pre-trained teacher image processing model as the target, perform knowledge distillation training on the to-be-trained student image processing model to obtain the compressed student image processing model.
[0068] In this embodiment, after the sample image is respectively input into the pre-trained teacher image processing model and the to-be-trained student image processing model, taking the output of the nth layer of the to-be-trained student image processing model learning the output of the nth layer of the pre-trained teacher image processing model as the target, perform knowledge distillation training on the to-be-trained student image processing model to obtain the trained student image processing model as the compressed student image processing model.
[0069] In this embodiment, knowledge distillation is performed based on the outputs of the intermediate layers corresponding to the pre-trained teacher image processing model and the student image processing model to be trained. In this way, even when the teacher model is relatively complex (which means that the distillation process itself may consume a large amount of computing resources) or of low quality, the image domain model compression method based on rank-preserving decomposition and intermediate layer knowledge distillation proposed in this embodiment can perform knowledge distillation only in the intermediate layer of the model: using the pre-trained image processing model as the teacher model, distilling the student image processing model to be trained after rank-preserving decomposition, and optimizing the learning of the student image processing model through the output differences of each layer, so as to ensure that the compressed student image processing model can maintain a performance similar to that of the image processing model, realizing the compression of the image processing model, and greatly reducing the storage requirements of the model while effectively maintaining the expression ability of the original model, thus overcoming the technical defect in the related art that only knowledge distillation is used alone, which is limited by the capacity of the student model and loses key information during the distillation process.
[0070] Combined with the above embodiments, in one implementation, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S21 at least includes step S41, and the above step S22 at least includes step S42:
[0071] Step S41: Determine that the size of the first pair of Kronecker embedding matrices is and the size of the second pair of Kronecker embedding matrices is according to the first target number of parameters.
[0072] In this embodiment, the pre-trained teacher image processing model is a Transformer model, which can be, for example, a DETR object detector, a ViT (Vision Transformer) model, etc., and there is no limitation thereto. The pre-trained teacher image processing model includes an embedding layer, and the embedding layer includes a relatively large lookup table matrix , where is the vocabulary size, d is the embedding dimension, and n is the largest factor of d. The target number of parameters at least includes: the first target number of parameters, which characterizes the target number of parameters of the embedding layer.
[0073] This embodiment can determine that the size of the first pair of Kronecker embedding matrices is and the size of the second pair of Kronecker embedding matrices is according to the first target number of parameters. That is to say, when performing Kronecker decomposition on the embedding layer, the scale parameter determined based on the first target number of parameters is .
[0074] Step S42: Perform Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model to obtain a group of Kronecker embedding matrices, and use it as the embedding layer of the student image processing model to be trained. The group of Kronecker embedding matrices includes: a matrix that conforms to the size of the first pair of Kronecker embedding matrices and a matrix that conforms to the size of the second pair of Kronecker embedding matrices.
[0075] In this embodiment, when performing Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model to obtain a group of Kronecker embedding matrices, and using the group of Kronecker embedding matrices as the parameter matrix of the embedding layer of the student image processing model to be trained, the embedding layer of the student image processing model to be trained is obtained. Among them, the group of Kronecker embedding matrices includes: a matrix that conforms to the size of the first pair of Kronecker embedding matrices and , and, a matrix that conforms to the size of the second pair of Kronecker embedding matrices and ; that is, X = . Specifically, in this embodiment, the specific steps of performing Kronecker decomposition shown in the foregoing embodiment can be followed to perform Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model to obtain a group of Kronecker embedding matrices.
[0076] In this embodiment, the size of the first pair of Kronecker embedding matrices is determined as , and the size of the second pair of Kronecker embedding matrices is determined as , which can make the matrices and degenerate into a row vector in the decomposition. In this way, the embeddings of each word in the lookup table can be separated from each other, and the i-th row in the matrices and independently represents the information related to the embedding of the i-th word.
[0077] Combining the above embodiments, in one implementation, the present invention also provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S21 further includes steps S51 to S53, and, the above step S22 further includes step S54:
[0078] Step S51: In the multi-head self-attention layer of the pre-trained teacher image processing model, splice the weight matrices of the key matrix K, the query matrix Q, and the value matrix V of all self-attention heads respectively to obtain the spliced weight matrices , and .
[0079] In this embodiment, the pre-trained teacher image processing model includes a multi-head self-attention layer (MHA). In the multi-head self-attention layer, the self-attention mechanism maps the input to a key matrix K, a query matrix Q, and a value matrix V, and obtains the final attention matrix through an attention operation and a softmax layer. Among them, the matrices K, Q, and V are obtained by multiplying the input by their respective weight matrices. In the multi-head self-attention (MHA) layer, each attention head has a set of independent weight matrices to allow for a richer data representation.
[0080] In the multi-head self-attention layer of the pre-trained teacher image processing model in this embodiment, the weight matrices corresponding to the key matrix K, the query matrix Q, and the value matrix V of all self-attention heads are respectively concatenated to obtain a concatenated weight matrix , and .
[0081] Step S52: According to the second target number of parameters, determine that the number of rows of the second pair of Kronecker attention weight matrices is the greatest common divisor of the number of rows of the concatenated weight matrix , and , and determine that the number of columns of the second pair of Kronecker attention weight matrices is the greatest common divisor of the number of columns of the concatenated weight matrix , and .
[0082] In this embodiment, the target number of parameters further includes: a second target number of parameters, which represents the target number of parameters of the multi-head self-attention layer. According to the second target number of parameters, it can be determined that the number of rows of the second pair of Kronecker attention weight matrices is the greatest common divisor of the number of rows of the concatenated weight matrix , and , and determine that the number of columns of the second pair of Kronecker attention weight matrices is the greatest common divisor of the number of columns of the concatenated weight matrix , , and . That is to say, when performing Kronecker decomposition on the multi-head self-attention layer, the scale parameter determined based on the second target number of parameters is ( ( , , 's greatest common divisor, 's greatest common divisor).
[0083] Step S53: According to the concatenated weight matrix , and determine the size of the first pair of Kronecker attention weight matrices based on the size of the second pair of Kronecker attention weight matrices and the size of the second pair of Kronecker attention weight matrices.
[0084] In this embodiment, after determining the size of the second pair of Kronecker attention weight matrices (including: the number of rows of the second pair of Kronecker attention weight matrices and the number of columns of the second pair of Kronecker attention weight matrices), the weight matrices after splicing can be used 、 and size ( 、 and The number of rows and columns of are the size of the matrix to be decomposed ), and the size of the second pair of Kronecker attention weight matrices to obtain the size of the first pair of Kronecker attention weight matrices (including: the number of rows of the first pair of Kronecker attention weight matrices and the number of columns of the first pair of Kronecker attention weight matrices). In a specific example, it can be through 、 and The number of rows of is divided by the number of rows of the second pair of Kronecker attention weight matrices to obtain the number of rows of the three first pairs of Kronecker attention weight matrices respectively; by 、 and The number of columns of is divided by the number of columns of the second pair of Kronecker attention weight matrices to obtain the number of columns of the three first pairs of Kronecker attention weight matrices respectively.
[0085] Step S54: Perform Kronecker decomposition on the spliced weight matrices 、 and respectively to obtain a set of Kronecker attention weight matrices 、 and , and use them as the multi-head self-attention layer of the student image processing model to be trained.
[0086] In this embodiment, perform Kronecker decomposition on the spliced weight matrices 、 and of the multi-head self-attention layer of the pre-trained teacher image processing model to obtain a set of Kronecker attention weight matrices 、 and , and the set of Kronecker attention weight matrices 、 and As the parameter matrix of the multi-head self-attention layer of the student image processing model to be trained, the multi-head self-attention layer of the student image processing model to be trained is obtained. Among them, the Kronecker attention weight matrix group , and respectively include: a matrix that conforms to the size of the first pair of Kronecker attention weight matrices, and a matrix that conforms to the size of the second pair of Kronecker attention weight matrices. Specifically, in this embodiment, the Kronecker decomposition can be performed on the concatenated weight matrix , and respectively according to the specific steps of performing Kronecker decomposition shown in the foregoing embodiment, to obtain the Kronecker attention weight matrix group , and .
[0087] Exemplarily, the Kronecker attention weight matrix group includes: a matrix and that conforms to the size of one pair of Kronecker attention weight matrices, and a matrix and that conforms to the size of the second pair of Kronecker embedding matrices; that is, . The Kronecker attention weight matrix group includes: a matrix and that conforms to the size of another pair of Kronecker attention weight matrices, and a matrix and that conforms to the size of the second pair of Kronecker embedding matrices; that is, . The Kronecker attention weight matrix group includes: a matrix and that conforms to the size of another pair of Kronecker attention weight matrices, and a matrix and that conforms to the size of the second pair of Kronecker embedding matrices; that is, . Among them, and , and , and, and may have the same or different sizes; and , and , and, and have the same size.
[0088] In this embodiment, for the weight matrix after splicing in the multi-head self-attention layer , and After performing Kronecker decomposition, the three groups of Kronecker attention weight matrices obtained , and The sizes of the 2 B matrices ( and , and , and, and ) corresponding to them are the same, that is, the sizes of the second pair of Kronecker attention weight matrices are the same, so that the decomposed information has higher consistency.
[0089] Combined with the above embodiments, in one implementation manner, the present invention also provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S21 further includes steps S61 to S62, and the above step S22 further includes step S63:
[0090] Step S61: According to the third target number of parameters, determine that the number of rows of the second pair of Kronecker linear weight matrices is the greatest common divisor of the number of rows of the linear mapping matrix , the first weight matrix and the second weight matrix in the feed-forward neural network layer of the pre-trained teacher image processing model, and determine that the number of columns of the second pair of Kronecker linear weight matrices is the greatest common divisor of the number of columns of the linear mapping matrix , the first weight matrix and the second weight matrix in the feed-forward neural network layer of the pre-trained teacher image processing model.
[0091] In this embodiment, the pre-trained teacher image processing model includes: a feed-forward neural network layer (FFN). The attention matrix finally obtained by the multi-head self-attention layer will be directly passed to the feed-forward neural network layer. In the feed-forward neural network layer, this attention matrix will be passed to a linear mapping matrix and the subsequent two weight matrices: the first weight matrix and the second weight matrix . In this embodiment, the Kronecker decomposition of the parameter matrix of the feed-forward neural network layer is similar to the Kronecker decomposition of the parameter matrix of the multi-head self-attention layer in the foregoing embodiment.
[0092] Specifically, the target parameter quantity in this embodiment further includes: a third target parameter quantity, which represents the target parameter quantity of the feedforward neural network layer. The number of rows of the second pair of Kronecker linear weight matrices can be determined according to the third target parameter quantity as the greatest common divisor of the number of rows of the linear mapping matrix , the first weight matrix and the second weight matrix in the feedforward neural network layer of the pre-trained teacher image processing model, and the number of columns of the second pair of Kronecker linear weight matrices is determined as the greatest common divisor of the number of columns of the linear mapping matrix , the first weight matrix and the second weight matrix in the feedforward neural network layer of the pre-trained teacher image processing model. That is to say, when performing Kronecker decomposition on the feedforward neural network layer, the scale parameter determined based on the third target parameter quantity is ( , , the greatest common divisor of the number of rows of the linear mapping matrix , the first weight matrix and the second weight matrix , the greatest common divisor of the number of columns of the linear mapping matrix , the first weight matrix and the second weight matrix ).
[0093] Step S62: Determine the size of the first pair of Kronecker linear weight matrices according to the sizes of the linear mapping matrix , the first weight matrix and the second weight matrix and the size of the second pair of Kronecker linear weight matrices.
[0094] In this embodiment, after determining the size of the second pair of Kronecker linear weight matrices (including: the number of rows of the second pair of Kronecker linear weight matrices and the number of columns of the second pair of Kronecker linear weight matrices), the size of the first pair of Kronecker linear weight matrices (including: the number of rows of the first pair of Kronecker linear weight matrices and the number of columns of the first pair of Kronecker linear weight matrices) can be obtained according to the sizes of the linear mapping matrix , the first weight matrix and the second weight matrix ( , and the number of rows and columns, that is, the size of the matrix to be decomposed ), and the size of the second pair of Kronecker linear weight matrices. In a specific example, it can be through , and The number of rows of the first pair of Kronecker linear weight matrices is obtained by dividing the number of rows of the second pair of Kronecker linear weight matrices. , and The number of columns of the first pair of Kronecker linear weight matrices is obtained by dividing the number of columns of the second pair of Kronecker linear weight matrices.
[0095] Step S63: Perform Kronecker decomposition on the linear mapping matrix , the first weight matrix and the second weight matrix respectively, to obtain a group of Kronecker linear weight matrices , and , which are used as the feed-forward neural network layer of the to-be-trained student image processing model.
[0096] In this embodiment, Kronecker decomposition is performed on the linear mapping matrix , the first weight matrix and the second weight matrix of the feed-forward neural network layer of the pre-trained teacher image processing model respectively, to obtain a group of Kronecker linear weight matrices , and , and the group of Kronecker linear weight matrices , and are used as the parameter matrices of the feed-forward neural network layer of the to-be-trained student image processing model, to obtain the feed-forward neural network layer of the to-be-trained student image processing model. Among them, the group of Kronecker linear weight matrices , and respectively include: a matrix that conforms to the size of the first pair of Kronecker linear weight matrices, and a matrix that conforms to the size of the second pair of Kronecker linear weight matrices. Specifically, in this embodiment, the Kronecker decomposition can be performed on the linear mapping matrix , the first weight matrix and the second weight matrix respectively according to the specific steps of performing Kronecker decomposition shown in the foregoing embodiment, to obtain a group of Kronecker linear weight matrices , and .
[0097] For example, the group of Kronecker linear weight matrices includes: a matrix and that conform to the size of one of the first pairs of Kronecker linear weight matrices, and, a matrix that conforms to the size of the second pair of Kronecker embedding matrices and ; that is, . The set of Kronecker linear weight matrices includes: a matrix that conforms to the size of another first pair of Kronecker attention weight matrices and , and, a matrix that conforms to the size of the second pair of Kronecker embedding matrices and ; that is, . The set of Kronecker linear weight matrices includes: a matrix that conforms to the size of another first pair of Kronecker attention weight matrices and , and, a matrix that conforms to the size of the second pair of Kronecker embedding matrices and ; that is, . Among them, and , and , and, and can have the same or different sizes; and , and , and, and have the same size.
[0098] In this embodiment, for the linear mapping matrix , the first weight matrix and the second weight matrix in the feedforward neural network layer, after performing Kronecker decomposition, the three sets of Kronecker linear weight matrices , and correspond to 2 B matrices respectively ( and , and , and, and ) have the same scale, that is, the sizes of the second pair of Kronecker linear weight matrices are the same, so that the decomposed information has higher consistency.
[0099] Combining the above embodiments, in one implementation manner, the present invention also provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S32 may specifically include steps S71 to S74:
[0100] Step S71: Obtain the first image embedding representation output by the embedding layer of the pre-trained teacher image processing model, the first attention map output by the multi-head self-attention layer of the pre-trained teacher image processing model, and the first image feature representation output by the feed-forward neural network layer of the pre-trained teacher image processing model after the pre-trained teacher image processing model processes the sample image.
[0101] In this embodiment, knowledge distillation can be performed on the intermediate layer of the student image processing model to be trained after Kronecker decomposition. Specifically, in this embodiment, the same sample image is input into the pre-trained teacher image processing model and the student image processing model to be trained for knowledge distillation. Among them, the sample image is an image used to train the student image processing model.
[0102] In this embodiment, the sample image is input into the pre-trained teacher image processing model, and the pre-trained teacher image processing model processes the sample image to obtain the outputs of the embedding layer, the multi-head self-attention layer, and the feed-forward neural network layer of the pre-trained teacher image processing model respectively; that is, after the pre-trained teacher image processing model processes the sample image, the first image embedding representation output by the embedding layer of the pre-trained teacher image processing model, the first attention map output by the multi-head self-attention layer of the pre-trained teacher image processing model, and the first image feature representation output by the feed-forward neural network layer of the pre-trained teacher image processing model are obtained.
[0103] Step S72: Obtain the second image embedding representation output by the embedding layer of the student image processing model to be trained, the second attention map output by the multi-head self-attention layer of the student image processing model to be trained, and the second image feature representation output by the feed-forward neural network layer of the student image processing model to be trained after the student image processing model to be trained processes the sample image.
[0104] In this embodiment, the sample image is input into the student image processing model to be trained, and the student image processing model to be trained processes the sample image to obtain the outputs of the embedding layer, the multi-head self-attention layer, and the feed-forward neural network layer of the student image processing model to be trained respectively; that is, after the student image processing model to be trained processes the sample image, the second image embedding representation output by the embedding layer of the student image processing model to be trained, the second attention map output by the multi-head self-attention layer of the student image processing model to be trained, and the second image feature representation output by the feed-forward neural network layer of the student image processing model to be trained are obtained.
[0105] In this embodiment, after the input sample image enters the student image processing model to be trained after performing Kronecker decomposition, the input of each layer (embedding layer, multi-head self-attention layer, or feed-forward neural network layer) needs to perform operations in each layer with the elements after performing Kronecker decomposition (i.e., Kronecker embedding matrix group, Kronecker attention weight matrix group, or Kronecker linear weight matrix group) (instead of multiplying with the original corresponding weight matrix). This part of the operation can be simplified by the calculation formula related to the Kronecker product. And most importantly, in this embodiment, compared with the pre-trained teacher image processing model and the student image processing model to be trained, for the number of layers of the model after the operation, the input and output dimensions of each layer remain unchanged, so that it is convenient to align for knowledge distillation.
[0106] Step S73: Determine the distillation loss based on the first image embedding representation, the second image embedding representation, the first attention map, the second attention map, and the first image feature representation and the second image feature representation.
[0107] In this embodiment, when performing knowledge distillation, the difference between the outputs of a specific layer of the pre-trained teacher image processing model and the student image processing model to be trained can be directly obtained without any projection. Specifically, the distillation process of this embodiment only needs to calculate the losses of the outputs of the embedding layer, the multi-head self-attention layer (i.e., the attention matrix), and the feed-forward neural network layer of the two: calculate the distillation loss based on the first image embedding representation, the second image embedding representation, the first attention map, the second attention map, and the first image feature representation and the second image feature representation.
[0108] For example, the embedding loss can be obtained based on the first image embedding representation and the second image embedding representation, the attention loss can be obtained based on the first attention map and the second attention map, and the image feature loss can be obtained based on the first image feature representation and the second image feature representation. Then, the embedding loss, the attention loss, and the image feature loss are added together (such as weighted) as the overall loss function of knowledge distillation to obtain the distillation loss, and a concise knowledge distillation process can be executed.
[0109] Step S74: Update the model parameters of the student image processing model to be trained based on the distillation loss to obtain the compressed student image processing model.
[0110] In this embodiment, the model parameters of the student image processing model to be trained can be updated based on the distillation loss until the distillation loss converges, and then the model parameters of the student image processing model to be trained are fixed to obtain the trained student image processing model. The trained student image processing model obtained after distillation in this embodiment is the compressed student image processing model finally obtained in this embodiment. It effectively maintains the expression ability of the original pre-trained teacher image processing model and greatly compresses and reduces the storage requirements of the model.
[0111] In this embodiment, knowledge distillation is only performed on the output of the embedding layer, the output of the multi-head self-attention layer (attention matrix), and the output of the feed-forward neural network layer related to Kronecker decomposition, and no scale projection is required. Only the simplest knowledge distillation is used locally to complete it, further compressing the number of model parameters, saving the storage space and computing resources of the target device for deploying the compressed student image processing model, and realizing the lightweight deployment of the image processing model with stable model performance.
[0112] In one embodiment, as Figure 3 shown, Figure 3 is the overall process schematic diagram of an image domain model compression method based on rank-preserving decomposition and intermediate layer knowledge distillation shown in an embodiment of the present invention. In Figure 3 , the pre-trained teacher image processing model to be decomposed is a Transformer model, and at least the following are included in the Transformer architecture of the Transformer model: an embedding layer, a multi-head self-attention module (i.e., a multi-head self-attention layer), and a feed-forward neural network layer. Among them, the embedding matrix X is included in the embedding layer, and the attention weights (i.e., the concatenated weight matrix , and ), and the linear weights (i.e., the linear mapping matrix , the first weight matrix and the second weight matrix ) are included in the feed-forward neural network layer.
[0113] Among them, when performing Kronecker decomposition on the pre-trained teacher image processing model, the embedding matrix X in the embedding layer, , and in the multi-head self-attention module, as well as , and in the feed-forward neural network layer are respectively subjected to Kronecker decomposition to obtain Kronecker embeddings (i.e., Kronecker embedding matrix groups: ), Kronecker attention weights (i.e., Kronecker attention weight matrix groups: , and ), and the Kronecker linear weights (i.e., the Kronecker linear weight matrix group: , and ) are used as the parameter matrices of the embedding layer of the student image processing model to be trained, the parameter matrices of the multi-head self-attention layer of the student image processing model to be trained, and the parameter matrices of the feed-forward neural network layer of the student image processing model to be trained, respectively, so as to obtain the embedding layer, multi-head self-attention layer, and feed-forward neural network layer of the student image processing model to be trained, and thus obtain the student image processing model to be trained.
[0114] During knowledge distillation, intermediate layer knowledge distillation is performed on the student image processing model to be trained obtained after decomposition: the entire distillation process only needs to calculate the losses of the outputs of the embedding layer, the outputs of the multi-head self-attention layer (i.e., the attention matrices), and the outputs of the feed-forward neural network layer between the student model (the student image processing model to be trained) and the teacher model (the pre-trained teacher image processing model), so as to obtain the distilled student model, that is, the compressed student image processing model, and thus realize the compression of the Transformer model based on rank-preserving decomposition and intermediate layer knowledge distillation, and greatly reduce the storage requirements of the model while effectively maintaining the expression ability of the original model.
[0115] This embodiment is implemented based on the rank-preserving decomposition (i.e., Kronecker decomposition) of the matrix and its numerical calculation method, and at the same time combines the knowledge distillation of only local elements in the intermediate layer of the model (embedding layer, multi-head self-attention layer, and feed-forward neural network layer). Through Kronecker decomposition, the matrix of the key part of the teacher model will be decomposed into the form of Kronecker product, thereby reducing the model parameters, and this decomposition method is a rank-preserving transformation of the original matrix, and the decomposed matrix is usually equal to or close to the rank of the original matrix, so as to effectively retain the information in the original matrix and avoid performance loss. Immediately afterwards, by performing knowledge distillation only in the intermediate layer of the model, using the pre-trained Transformer model as the teacher model, distilling the model that has undergone rank-preserving decomposition, and optimizing the learning of the student model through the output differences of each layer, so as to ensure that the compressed model can maintain a performance similar to that of the teacher model, and further save the storage space and computing resources of the target device of the student image processing model to be deployed and compressed, and realize the lightweight deployment of the image processing model with stable model performance.
[0116] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0117] Based on the same inventive concept, an embodiment of the present invention provides an image domain model compression device based on rank-preserving decomposition and knowledge distillation. Refer to Figure 4 , Figure 4 which is a structural block diagram of an image domain model compression device based on rank-preserving decomposition and knowledge distillation provided by an embodiment of the present invention. As Figure 4 shown, the image domain model compression device based on rank-preserving decomposition and knowledge distillation of this embodiment may include:
[0118] A parameter acquisition module, configured to acquire the device parameter values of the target device for the student image processing model to be deployed and compressed, where the device parameter values include at least one of the following: storage space size, operation speed;
[0119] A parameter determination module, configured to determine the target number of parameters of the student image processing model according to the device parameter values;
[0120] A model decomposition module, configured to perform Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and use the model after performing Kronecker decomposition as the student image processing model to be trained;
[0121] A model training module, configured to use sample images and take the pre-trained teacher image processing model as the learning target to perform knowledge distillation training on the student image processing model to be trained, so as to obtain the compressed student image processing model, where the pre-trained teacher image processing model is used for any one of the following: image classification, image segmentation, target detection;
[0122] An image processing module, configured to input the image to be processed into the compressed student image processing model to obtain an image processing result, where the image processing result is any one of the following: image classification result, image segmentation result, target detection result.
[0123] Optionally, the model decomposition module includes:
[0124] A size determination module, configured to determine that the first pair of Kronecker matrix sizes is , and the second pair of Kronecker matrix sizes is ;
[0125] A parameter determination module, configured to perform Kronecker decomposition on the parameter matrix of the n-th layer of the pre-trained teacher image processing model, and obtain a group of Kronecker parameter matrices, which are used as the parameter matrix of the n-th layer of the to-be-trained student image processing model. The group of Kronecker parameter matrices includes: a matrix that conforms to the first pair of Kronecker matrix sizes and and, a matrix that conforms to the second pair of Kronecker matrix sizes and and ;
[0126] wherein , and n is an integer greater than 0.
[0127] Optionally, the model training module includes:
[0128] An image input module, configured to input the sample image into the pre-trained teacher image processing model, and input the sample image into the to-be-trained student image processing model;
[0129] A knowledge distillation module, configured to perform knowledge distillation training on the to-be-trained student image processing model with the goal of learning the output of the n-th layer of the pre-trained teacher image processing model by the output of the n-th layer of the to-be-trained student image processing model, so as to obtain the compressed student image processing model.
[0130] Optionally, the pre-trained teacher image processing model is a Transformer model;
[0131] The size determination module at least includes:
[0132] A first size determination module, configured to determine, according to a first target number of parameters, that the first pair of Kronecker embedding matrix sizes is , and the second pair of Kronecker embedding matrix sizes is ;
[0133] The parameter determination module at least includes:
[0134] An embedding layer decomposition module, configured to perform Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model, and obtain a group of Kronecker embedding matrices, which are used as the embedding layer of the to-be-trained student image processing model. The group of Kronecker embedding matrices includes: a matrix that conforms to the first pair of Kronecker embedding matrix sizes and a matrix that conforms to the second pair of Kronecker embedding matrix sizes;
[0135] wherein is the vocabulary size, d is the embedding dimension, and n is the largest factor of d.
[0136] Optionally, the size determination module further includes:
[0137] A matrix splicing module, configured to splice the weight matrices of the key matrix K, the query matrix Q, and the value matrix V of all self-attention heads in the multi-head self-attention layer of the pre-trained teacher image processing model respectively to obtain a spliced weight matrix , and ;
[0138] A first row and column determination module, configured to determine, according to a second target number of parameters, that the number of rows of the second Kronecker attention weight matrix is the greatest common divisor of the number of rows of the spliced weight matrix , and , and determine that the number of columns of the second Kronecker attention weight matrix is the greatest common divisor of the number of columns of the spliced weight matrix , and ;
[0139] A second size determination module, configured to determine the size of the first Kronecker attention weight matrix according to the size of the spliced weight matrix , and and the size of the second Kronecker attention weight matrix;
[0140] The parameter determination module further includes:
[0141] An attention layer decomposition module, configured to perform Kronecker decomposition on the spliced weight matrix , and respectively to obtain a group of Kronecker attention weight matrices , and , and use them as the multi-head self-attention layer of the to-be-trained student image processing model. The group of Kronecker attention weight matrices , and respectively include: a matrix conforming to the size of the first Kronecker attention weight matrix and a matrix conforming to the size of the second Kronecker attention weight matrix.
[0142] Optionally, the size determination module further includes:
[0143] A second row and column determination module, configured to determine, according to a third target number of parameters, that the number of rows of the second Kronecker linear weight matrix is the linear mapping matrix in the feed-forward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix find the greatest common divisor of the number of rows, and determine that the number of columns of the second pair of Kronecker linear weight matrices is the linear mapping matrix in the feedforward neural network layer of the pre-trained teacher image processing model the first weight matrix and the second weight matrix find the greatest common divisor of the number of columns;
[0144] A third dimension determination module, configured to determine the size of the first pair of Kronecker linear weight matrices according to the sizes of the linear mapping matrix the first weight matrix and the second weight matrix and the size of the second pair of Kronecker linear weight matrices;
[0145] The parameter determination module further includes:
[0146] A feedforward neural network layer decomposition module, configured to perform Kronecker decomposition on the linear mapping matrix the first weight matrix and the second weight matrix respectively to obtain a group of Kronecker linear weight matrices , and , and use them as the feedforward neural network layer of the to-be-trained student image processing model. The group of Kronecker linear weight matrices , and respectively include: a matrix conforming to the size of the first pair of Kronecker linear weight matrices, and a matrix conforming to the size of the second pair of Kronecker linear weight matrices.
[0147] Optionally, the knowledge distillation module includes:
[0148] A first acquisition module, configured to acquire, after the pre-trained teacher image processing model processes the sample image, a first image embedding representation output by the embedding layer of the pre-trained teacher image processing model, a first attention map output by the multi-head self-attention layer of the pre-trained teacher image processing model, and a first image feature representation output by the feedforward neural network layer of the pre-trained teacher image processing model;
[0149] A second acquisition module, configured to acquire, after the to-be-trained student image processing model processes the sample image, a second image embedding representation output by the embedding layer of the to-be-trained student image processing model, a second attention map output by the multi-head self-attention layer of the to-be-trained student image processing model, and a second image feature representation output by the feedforward neural network layer of the to-be-trained student image processing model;
[0150] A loss calculation module, configured to determine a distillation loss based on the first image embedding representation, the second image embedding representation, the first attention map, the second attention map, and the first image feature representation and the second image feature representation.
[0151] A parameter update module, configured to update model parameters of the student image processing model to be trained based on the distillation loss, so as to obtain the compressed student image processing model.
[0152] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the image domain model compression method based on rank-preserving decomposition and knowledge distillation as described in any one of the above embodiments of the present invention are implemented.
[0153] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, as Figure 5 shown. Figure 5 is a schematic diagram of an electronic device shown in an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, the steps in the image domain model compression method based on rank-preserving decomposition and knowledge distillation as described in any one of the above embodiments of the present invention are implemented.
[0154] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the related parts, refer to the partial description of the method embodiment.
[0155] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0156] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0157] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0160] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0161] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the said element.
[0162] The above has introduced in detail a method, apparatus, device and medium for compressing an image domain model based on rank-preserving decomposition and knowledge distillation provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An image domain model compression method based on rank-preserving decomposition and knowledge distillation, characterized in that The method includes: Obtaining the device parameter values of the target device for the pre-deployed compressed student image processing model, where the device parameter values include at least one of the following: storage space size, operation speed; Determining the target number of parameters of the student image processing model according to the device parameter values; Performing Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and using the model after performing Kronecker decomposition as the student image processing model to be trained; Using sample images, taking the pre-trained teacher image processing model as the learning target, and performing knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model, where the pre-trained teacher image processing model is used for any one of the following: image classification, image segmentation, object detection; Inputting the image to be processed into the compressed student image processing model to obtain an image processing result, where the image processing result is any one of the following: image classification result, image segmentation result, object detection result; Wherein, the pre-trained teacher image processing model is a Transformer model; performing Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and using the model after performing Kronecker decomposition as the student image processing model to be trained, at least includes: In the multi-head self-attention layer of the pre-trained teacher image processing model, the weight matrices of the key matrix K, the query matrix Q, and the value matrix V of all self-attention heads are respectively concatenated to obtain the concatenated weight matrix , and ; Determine that the number of rows of the second pair of Kronecker attention weight matrices is the greatest common divisor of the numbers of rows of the spliced weight matrix according to the second target parameter quantity 、 and and determine that the number of columns of the second pair of Kronecker attention weight matrices is the greatest common divisor of the numbers of columns of the spliced weight matrix 、 and ; According to the spliced weight matrix , and the size of, and the size of the second pair of Kronecker attention weight matrices, determine the size of the first pair of Kronecker attention weight matrices; For the spliced weight matrix , and perform Kronecker decomposition respectively to obtain a group of Kronecker attention weight matrices , and , and use them as the multi-head self-attention layer of the student image processing model to be trained. The group of Kronecker attention weight matrices , and respectively include: a matrix that conforms to the size of the first pair of Kronecker attention weight matrices, and a matrix that conforms to the size of the second pair of Kronecker attention weight matrices.
2. The method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation according to claim 1, wherein Performing Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and using the model after performing Kronecker decomposition as the student image processing model to be trained, includes: Determine that the first pair of Kronecker matrix dimensions is , and the second pair of Kronecker matrix dimensions is ; The parameter matrix of the nth layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a group of Kronecker parameter matrices, and use them as the parameter matrix of the nth layer of the student image processing model to be trained. The group of Kronecker parameter matrices includes: a matrix that conforms to the first pair of Kronecker matrix sizes and , and, a matrix that conforms to the second pair of Kronecker matrix sizes and ; Among them, , where n is an integer greater than 0.
3. The method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation according to claim 2, wherein Using sample images, taking the pre-trained teacher image processing model as the learning target, and performing knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model, includes: Inputting the sample images into the pre-trained teacher image processing model and inputting the sample images into the student image processing model to be trained; Taking the output of the nth layer of the student image processing model to be trained to learn the output of the nth layer of the pre-trained teacher image processing model as the target, and performing knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model.
4. The method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation according to claim 3, wherein The pre-trained teacher image processing model is a Transformer model; Determine that the first pair of Kronecker matrix sizes is , and the second pair of Kronecker matrix sizes is , including at least: Determine that the size of the first pair of Kronecker embedding matrices is , and the size of the second pair of Kronecker embedding matrices is ; The parameter matrix of the n-th layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a group of Kronecker parameter matrices, and use them as the parameter matrix of the n-th layer of the student image processing model to be trained, including at least: Performing Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model to obtain a group of Kronecker embedding matrices, and using them as the embedding layer of the student image processing model to be trained, where the group of Kronecker embedding matrices includes: a matrix that conforms to the size of the first pair of Kronecker embedding matrices and a matrix that conforms to the size of the second pair of Kronecker embedding matrices; Among them, is the vocabulary size, d is the embedding dimension, and n is the largest factor of d.
5. The method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation according to claim 1, wherein Further includes: Determine that the number of rows of the second pair of Kronecker linear weight matrices is the greatest common divisor of the number of rows of the linear mapping matrix, the first weight matrix, and the second weight matrix in the feed-forward neural network layer of the pre-trained teacher image processing model according to the third target parameter quantity , the first weight matrix and the second weight matrix ; and determine that the number of columns of the second pair of Kronecker linear weight matrices is the greatest common divisor of the number of columns of the linear mapping matrix, the first weight matrix, and the second weight matrix in the feed-forward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix ; According to the linear mapping matrix , the first weight matrix and the second weight matrix , determine the size of the first pair of Kronecker linear weight matrices according to the sizes of the second pair of Kronecker linear weight matrices; For the linear mapping matrix , the first weight matrix and the second weight matrix , perform Kronecker decomposition respectively to obtain a set of Kronecker linear weight matrices , and , and use them as the feedforward neural network layer of the student image processing model to be trained. The set of Kronecker linear weight matrices , and respectively include: a matrix conforming to the size of the first pair of Kronecker linear weight matrices, and a matrix conforming to the size of the second pair of Kronecker linear weight matrices.
6. The method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation according to any one of claims 1 to 5, characterized in that Taking the output of the nth layer of the student image processing model to be trained to learn the output of the nth layer of the pre-trained teacher image processing model as the target, and performing knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model, includes: Obtain the first image embedding representation output by the embedding layer of the pre-trained teacher image processing model, the first attention map output by the multi-head self-attention layer of the pre-trained teacher image processing model, and the first image feature representation output by the feed-forward neural network layer of the pre-trained teacher image processing model after the pre-trained teacher image processing model processes the sample image; Obtain the second image embedding representation output by the embedding layer of the student image processing model to be trained, the second attention map output by the multi-head self-attention layer of the student image processing model to be trained, and the second image feature representation output by the feed-forward neural network layer of the student image processing model to be trained after the student image processing model to be trained processes the sample image; Determine the distillation loss based on the first image embedding representation and the second image embedding representation, the first attention map and the second attention map, and the first image feature representation and the second image feature representation; Update the model parameters of the student image processing model to be trained based on the distillation loss to obtain the compressed student image processing model.
7. An image domain model compression device based on rank-preserving decomposition and knowledge distillation, characterized in that, The device includes: A parameter acquisition module, configured to acquire the device parameter values of the target device for the student image processing model to be deployed and compressed, where the device parameter values include at least one of the following: storage space size, operation speed; A parameter determination module, configured to determine the target number of parameters of the student image processing model according to the device parameter values; A model decomposition module is used to perform Kronecker decomposition on a pre-trained teacher image processing model according to the target number of parameters, and use the model after Kronecker decomposition as the student image processing model to be trained; wherein, the pre-trained teacher image processing model is a Transformer model; performing Kronecker decomposition on the pre-trained teacher image processing model according to the target number of parameters, and using the model after Kronecker decomposition as the student image processing model to be trained includes at least: in the multi-head self-attention layer of the pre-trained teacher image processing model, concatenating the weight matrices of the key matrix K, query matrix Q, and value matrix V of all self-attention heads respectively to obtain a concatenated weight matrix , and ; determining that the number of rows of the second pair of Kronecker attention weight matrices is the greatest common divisor of the number of rows of the concatenated weight matrix , and according to the second target number of parameters, and determining that the number of columns of the second pair of Kronecker attention weight matrices is the greatest common divisor of the number of columns of the concatenated weight matrix , and ; determining the size of the first pair of Kronecker attention weight matrices according to the sizes of the concatenated weight matrix , and and the size of the second pair of Kronecker attention weight matrices; performing Kronecker decomposition on the concatenated weight matrix , and respectively to obtain a group of Kronecker attention weight matrices , and , and using them as the multi-head self-attention layer of the student image processing model to be trained, the group of Kronecker attention weight matrices , and respectively include: a matrix that conforms to the size of the first pair of Kronecker attention weight matrices, and a matrix that conforms to the size of the second pair of Kronecker attention weight matrices; A model training module, configured to use sample images and use the pre-trained teacher image processing model as the learning target to perform knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model, where the pre-trained teacher image processing model is used for any one of the following: image classification, image segmentation, object detection; An image processing module, configured to input the image to be processed into the compressed student image processing model to obtain an image processing result, where the image processing result is any one of the following: image classification result, image segmentation result, object detection result.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the image domain model compression method based on rank-preserving decomposition and knowledge distillation according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image domain model compression method based on rank-preserving decomposition and knowledge distillation according to any one of claims 1 to 6.
Citation Information
Patent Citations
Over-parameterized knowledge distillation method, device, equipment and medium
CN119250176A