Image domain model compression method and device based on rank-preserving decomposition and knowledge distillation, equipment and medium

By adopting the technology of rank-keeping decomposition and knowledge distillation in the image processing model, the problem of difficulty in deploying image processing models on lightweight devices in the prior art is solved, and efficient model compression and performance retention are achieved.

CN119942275AActive Publication Date: 2025-05-06TSINGHUA UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510413975.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

When compressing an image processing model, it is difficult to achieve high compression rate and low computing overhead while maintaining model performance, resulting in the inability to deploy an image processing model on lightweight devices.

Method used

The pre-trained teacher image processing model is compressed through Cronek decomposition to obtain the compressed student image processing model, and the knowledge distillation training ensures that the performance of the student model is close to that of the teacher model.

Benefits of technology

It realizes that while maintaining the expression capabilities of the image processing model, significantly compress model parameters, saves storage space and computing resources, and realizes lightweight deployment of the image processing model, while maintaining the stability of the model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942275A_ABST
    Figure CN119942275A_ABST
Patent Text Reader

Abstract

The invention provides an image domain model compression method and device based on rank-preserving decomposition and knowledge distillation, equipment and a medium, and relates to the field of computers. Comprising the steps of obtaining a device parameter value of a target device of a student image processing model to be deployed and compressed, and determining a target parameter quantity of the student image processing model according to the device parameter value; executing Kronecker decomposition on a pre-trained teacher image processing model according to the target parameter quantity, and taking the decomposed model as a to-be-trained student image processing model; performing knowledge distillation training on a to-be-trained student image processing model by using the sample image and taking a pre-trained teacher image processing model as a learning target to obtain a compressed student image processing model; and inputting the to-be-processed image into the compressed student image processing model to obtain an image processing result, thereby compressing the model parameters of the image processing model under the condition of maintaining the expression capability of the image processing model, and saving the storage space and computing resources of model deployment equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to an image domain model compression method, device, equipment and medium based on rank-preserving decomposition and knowledge distillation. Background Art

[0002] In recent years, the application of image processing models has become more and more extensive, which is attributed to the increasingly better effects of image processing models on image processing (such as target tracking, target detection, image classification, etc.). However, the improvement of the current image processing model effect is inseparable from the large-scale model parameters in the image processing model, which means that the equipment deploying the image processing model needs to have large-capacity hardware storage space and / or sufficient computing resources, making it impossible to deploy the image processing model in lightweight devices, limiting the application and development of the image processing field in lightweight devices.

[0003] Currently, most model compression methods based on SVD (Singular Value Decomposition) decompose the weight matrix in the image processing model to reduce the number of model parameters and the amount of calculation, further achieving the purpose of saving storage space and computing resources of the deployed equipment. However, the compression effect of the SVD method depends on the retention strategy of singular values: in practical applications, retaining large singular values ​​will limit the compression ratio of the model, especially when processing complex image processing models, SVD cannot effectively approximate low-rank representations in high-dimensional space. This means that although SVD can reduce the number of model parameters, it often cannot achieve the desired high compression rate and low computational overhead. Although SVD approximates by retaining important singular values, it may introduce performance loss because the SVD method itself does not consider the training objectives of the image processing model. In addition, SVD may not be able to retain the global information of the network, and it is easy to lose global information at high compression rates, thereby affecting the image processing accuracy of the image processing model. As a result, the current method based on low-rank decomposition (i.e., SVD) has the problems of rank limitation and model performance loss, and cannot compress the image processing model parameters to achieve model deployment on lightweight devices while ensuring the performance of the image processing model. Therefore, how to compress the model parameters of the image processing model while maintaining the expressive power of the image processing model and saving the storage space and computing resources of the model deployment device is a technical problem to be solved by the present invention. Summary of the invention

[0004] Based on the above technical problems, the present invention provides an image domain model compression method, device, equipment and medium based on rank-preserving decomposition and knowledge distillation, aiming to significantly compress and reduce the storage requirements of the model while effectively maintaining the expressiveness of the original image processing model, thereby realizing lightweight deployment of the image processing model and maintaining stable model performance.

[0005] A first aspect of the present invention provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation, the method comprising: Obtaining device parameter values ​​of a target device to which the compressed student image processing model is to be deployed, wherein the device parameter values ​​include at least one of the following: storage space size and computing speed; Determining target parameter values ​​of the student image processing model according to the device parameter values; According to the target parameter amount, performing Kronecker decomposition on the pre-trained teacher image processing model, and using the model after performing the Kronecker decomposition as the student image processing model to be trained; Using sample images and taking the pre-trained teacher image processing model as a learning target, performing knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model, wherein the pre-trained teacher image processing model is used for any of the following: image classification, image segmentation, and object detection; The image to be processed is input into the compressed student image processing model to obtain an image processing result, which is any one of the following: an image classification result, an image segmentation result, and a target detection result.

[0006] A second aspect of the present invention provides an image domain model compression device based on rank-preserving decomposition and knowledge distillation, the device comprising: A parameter acquisition module, used to acquire device parameter values ​​of a target device to be deployed with the compressed student image processing model, wherein the device parameter values ​​include at least one of the following: storage space size and operation speed; A parameter determination module, used for determining a target parameter amount of the student image processing model according to the device parameter value; A model decomposition module, used for performing Kronecker decomposition on the pre-trained teacher image processing model according to the target parameter amount, and using the model after performing Kronecker decomposition as the student image processing model to be trained; A model training module, for performing knowledge distillation training on the student image processing model to be trained using sample images and taking the pre-trained teacher image processing model as a learning target to obtain the compressed student image processing model, wherein the pre-trained teacher image processing model is used for any of the following: image classification, image segmentation, and object detection; The image processing module is used to input the image to be processed into the compressed student image processing model to obtain an image processing result, which is any one of the following: image classification result, image segmentation result, and target detection result.

[0007] A third aspect of an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, an image domain model compression method based on rank-preserving decomposition and knowledge distillation as described in the first aspect of an embodiment of the present invention is implemented.

[0008] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the image domain model compression method based on rank-preserving decomposition and knowledge distillation of the first aspect of the embodiment of the present invention is implemented.

[0009] Through the image domain model compression method based on rank-preserving decomposition and knowledge distillation of this embodiment, after determining the target parameter amount based on the device parameter value of the target device of the compressed student image processing model to be deployed, Kronecker decomposition is performed on the pre-trained teacher image processing model according to the target parameter amount, and the model after Kronecker decomposition is used as the student image processing model to be trained. In this way, this embodiment performs a rank-preserving transformation on the original matrix in the pre-trained teacher image processing model through the decomposition method of Kronecker decomposition. The decomposed matrix is ​​usually equal to or close to the rank of the original matrix, so that the information in the original matrix can be effectively retained, avoiding performance loss of the pre-trained teacher image processing model. Next, this embodiment uses sample images and takes the pre-trained teacher image processing model as the learning target to perform knowledge distillation on the student image processing model to be trained after rank-preserving decomposition, and obtains a compressed student image processing model, thereby ensuring that the compressed student image processing model can maintain similar performance to the teacher model. Ultimately, it achieves the compression of the model parameters of the student image processing model while maintaining the expressiveness of the teacher image processing model, saving the storage space and computing resources of the target device to which the compressed student image processing model is to be deployed, and realizing the lightweight deployment of the image processing model while maintaining stable model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0011] Figure 1 is a flowchart of a method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation according to an embodiment of the present invention; Figure 2 is a schematic diagram of a Kronecker product calculation according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the overall process of an image domain model compression method based on rank-preserving decomposition and intermediate layer knowledge distillation according to an embodiment of the present invention; Figure 4 It is a structural block diagram of an image domain model compression device based on rank-preserving decomposition and knowledge distillation provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of an electronic device shown in an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0013] Please refer to Figure 1 , Figure 1 FIG. 1 is a flowchart of a method for compressing an image domain model based on rank-preserving decomposition and knowledge distillation according to an embodiment of the present invention. Figure 1 As shown, the image domain model compression method based on rank-preserving decomposition and knowledge distillation provided in this embodiment includes at least the following steps: Step S11: obtaining device parameter values ​​of a target device on which the compressed student image processing model is to be deployed, wherein the device parameter values ​​include at least one of the following: storage space size and computing speed.

[0014] In this embodiment, the target device is used to deploy the compressed student image processing model. The target device can be any hardware device, such as a server, a computer, a mobile terminal (such as a mobile phone, an IPAD), etc., without limitation. The purpose of this embodiment is to compress the pre-trained teacher image processing model to obtain a compressed student image processing model, so as to deploy the compressed student image processing model on the target device, thereby realizing lightweight deployment of the compressed student image processing model.

[0015] First, the present embodiment can obtain the device parameter value of the target device on which the compressed student image processing model is to be deployed, and the device parameter value includes at least one of the following: storage space size and operation speed. The device parameter value of the present embodiment can refer to the factory rated device parameter value of the target device, or can refer to the remaining device parameter value of the target device (such as available memory capacity), and there is no limitation on this.

[0016] Step S12: Determine the target parameter value of the student image processing model according to the device parameter value.

[0017] In this embodiment, the target parameter amount of the compressed student image processing model can be determined according to the device parameter value of the target device on which the compressed student image processing model is to be deployed. The target parameter amount of this embodiment is the model parameter amount determined according to the device parameter value, so that when the compressed student image processing model is subsequently deployed on the target device, the device parameter value of the target device is met.

[0018] Step S13: According to the target parameter amount, perform Kronecker decomposition on the pre-trained teacher image processing model, and use the model after the Kronecker decomposition as the student image processing model to be trained.

[0019] In this embodiment, after obtaining the target parameter amount, the pre-trained teacher image processing model can be subjected to Kronecker decomposition according to the target parameter amount to obtain a model after Kronecker decomposition, and the model after Kronecker decomposition is used as the student image processing model to be trained. The "teacher image processing model" and "student image processing model" of this embodiment correspond to the "teacher model" and "student model" in knowledge distillation. Among them, the pre-trained teacher image processing model is a pre-trained image processing model, and the pre-trained teacher image processing model is used for any of the following: image classification, image segmentation, target detection, and target tracking. The student image processing model to be trained is an image processing model to be trained that learns the knowledge of the pre-trained teacher image processing model, and the student image processing model to be trained is a model after Kronecker decomposition is performed on the pre-trained teacher image processing model.

[0020] Step S14: Using sample images and taking the pre-trained teacher image processing model as a learning target, knowledge distillation training is performed on the student image processing model to be trained to obtain the compressed student image processing model, and the pre-trained teacher image processing model is used for any of the following: image classification, image segmentation, and target detection.

[0021] After obtaining the student image processing model to be trained, this embodiment can perform knowledge distillation on the pre-trained teacher image processing model and the student image processing model to be trained: using sample images and taking the pre-trained teacher image processing model as the learning target, knowledge distillation training is performed on the student image processing model to be trained to obtain a compressed student image processing model.

[0022] Step S15: input the image to be processed into the compressed student image processing model to obtain an image processing result, which is any one of the following: image classification result, image segmentation result, and target detection result.

[0023] In this embodiment, after obtaining the compressed student image processing model, the image to be processed can be input into the compressed student image processing model to obtain an image processing result; wherein the image processing result of this embodiment is any one of the following: image classification result, image segmentation result, target detection result, target tracking result. In an optional implementation, after obtaining the compressed student image processing model, the compressed student image processing model can be deployed on a target device, and the image to be processed can be input into the compressed student image processing model to obtain an image processing result.

[0024] In this embodiment, the original matrix in the pre-trained teacher image processing model is transformed to preserve the rank through the decomposition method of Kronecker decomposition. The decomposed matrix is ​​usually equal to or close to the rank of the original matrix, so that the information in the original matrix can be effectively retained, avoiding the performance loss of the pre-trained teacher image processing model. Next, this embodiment uses sample images and takes the pre-trained teacher image processing model as the learning target to perform knowledge distillation on the student image processing model to be trained after the rank-preserving decomposition to obtain a compressed student image processing model, thereby ensuring that the compressed student image processing model can maintain performance similar to that of the teacher model, and finally compressing the model parameters of the student image processing model while maintaining the expressive power of the teacher image processing model, saving the storage space and computing resources of the target device to be deployed with the compressed student image processing model, and realizing the lightweight deployment of the image processing model with stable model performance.

[0025] In combination with the above embodiments, in one implementation, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In the method, the above step S12 may specifically include step S121 and step S122: Step S121: When the device parameter values ​​of the target device read are: available memory capacity (unit: MB / GB), the single calculation peak performance of the processor of the target device (for example, the number of floating-point operations that can be processed per second, FLOPS) and the maximum response time required by the target device (for example, image processing must be completed within 0.5 seconds), calculate the basic parameter quantities of the student image processing model based on the device parameter values.

[0026] The basic parameter amount of this embodiment includes: parameter storage amount and parameter amount upper limit; wherein, the parameter storage amount is determined as follows: Determine the amount of parameter storage for the student image processing model based on the available memory capacity of the target device: Each model parameter occupies 4 bytes (32-bit floating point number) by default. If the available memory capacity is M MB, the parameter storage capacity (i.e. the maximum number of parameters) of the student image processing model is: Maximum number of parameters = (M×1024²×memory occupancy) / 4, where the memory occupancy is usually set to 60%-80% to reserve system resources. For example: If the available memory capacity is 4GB (4096MB), when allocated at 70% memory occupancy, it supports approximately 734 million parameters ((4096×1024²×0.7) / 4 ≈734,003,200).

[0027] The upper limit of the parameter amount is determined as follows: Based on the peak performance of the target device's processor, determine the computational efficiency constraints of the student image processing model: First, calculate the maximum amount of computing allowed for a single inference (i.e., the amount of computing that can be tolerated) based on the number of floating-point operations per second (P FLOPS) that the processor can handle and the maximum response time T seconds required by the target device: Tolerable amount of computing = P × T. Then, based on the number of operations per parameter of the model (for example, a network requires an average of 10 operations per parameter), determine the upper limit of the number of parameters: Upper limit of parameter amount = Tolerable amount of computing / number of operations per parameter of the model.

[0028] Step S122: Determine the target parameter amount of the student image processing model based on the basic parameter amount of the student image processing model.

[0029] In this embodiment, the minimum value among the basic parameter quantities can be determined as the target parameter quantity of the student image processing model, that is, the minimum value between the value of the parameter storage quantity and the value of the parameter quantity upper limit is used as the target parameter quantity of the student image processing model.

[0030] In combination with the above embodiments, in one implementation, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In the method, the above step S13 may specifically include step S21 and step S22: Step S21: According to the target parameter quantity, determine the size of the first pair of Kronecker matrices , the second pair of Kronecker matrix dimensions are .

[0031] Step S22: Parameter matrix of the nth layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a Kronecker parameter matrix group, and use it as the parameter matrix of the nth layer of the student image processing model to be trained.

[0032] In this embodiment, the size of the first pair of Kronecker matrices can be determined according to the target parameter quantity. , and, determine the second pair of Kronecker matrix dimensions And, this embodiment is the parameter matrix of the nth layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a Kronecker parameter matrix group, and use the Kronecker parameter matrix group as the parameter matrix of the nth layer of the student image processing model to be trained, thereby achieving Kronecker decomposition of the pre-trained teacher image processing model and obtaining a model after performing Kronecker decomposition. Where n is an integer greater than 0, and the Kronecker parameter matrix group includes: a matrix that meets the size of the first pair of Kronecker matrices and , and , correspond to the matrices with the dimensions of the second pair of Kronecker matrices and ; .

[0033] In one example, the Kronecker decomposition of a matrix: The Kronecker product is a matrix operation that transforms two matrices into a block matrix. Let the matrix , , then the Kronecker product of A and B is denoted as , is a block matrix where each block Yes Multiply it by matrix B to get. Figure 2 As shown, Figure 2 is a schematic diagram of a Kronecker product operation shown in one embodiment of the present invention. Figure 2 In the figure, the Kronecker product operation for a 2×2 matrix is ​​shown. Decomposing the result of matrix multiplication into a Kronecker product can be considered as replacing the projection of the original linear space with a more constrained linear space, making the overall information expression more compact while retaining the information interaction between each data block. The inverse process of this process is the Kronecker decomposition. Kronecker's theorem shows that given , , then any matrix can be expressed as several and The sum of the Kronecker products of .

[0034] In this embodiment, the obtained Kronecker parameter matrix group is expressed as two groups: and The sum of the Kronecker products of , thus obtaining an approximate Kronecker decomposition that is numerically accurate enough. i takes 1 and 2, and for the parameter matrix The Kronecker parameter matrix group obtained by Kronecker decomposition is: , where the matrix and matrix The dimensions of the first pair of Kronecker matrices are ,matrix and matrix Fit the second pair of Kronecker matrix dimensions .

[0035] The specific steps of performing Kronecker decomposition in this embodiment are given below: (a) Given the scale parameter to be decomposed ,in, , m, n are the original matrices The size (i.e., the number of rows and columns) of the matrix to be decomposed.

[0036] (b) For the original matrix Perform rearrangement operation and scan the original matrix one by one The blocks are then flattened into one dimension and concatenated to form blocks of size The matrix .

[0037] (c) The matrix obtained in the above steps Perform rank 1 SVD decomposition .

[0038] (d) vector According to each elements are divided into 1 column, forming a size of The matrix is ​​denoted as , for the vector According to each elements are divided into 1 column, forming a size of The matrix is ​​denoted as .

[0039] (e) Assume .right Repeat steps (b), (c) and (d) to obtain and .

[0040] (f) Obtain the numerical approximation of the Kronecker decomposition of the matrix W in this embodiment: .

[0041] In this embodiment, by using Kronecker decomposition, the number of parameters storing the matrix is ​​reduced from m×n to 2 ( ), which achieves a significant compression of the image processing model.

[0042] In combination with the above embodiments, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In the method, the above step S14 of "using the sample image, taking the pre-trained teacher image processing model as the learning target, performing knowledge distillation training on the student image processing model to be trained, and obtaining the compressed student image processing model" can specifically include steps S31 and S32: Step S31: inputting the sample image into the pre-trained teacher image processing model, and inputting the sample image into the student image processing model to be trained.

[0043] In this embodiment, the sample image can be input into a pre-trained teacher image processing model, and the sample image can be input into a student image processing model to be trained.

[0044] Step S32: Taking the output of the nth layer of the student image processing model to be trained as the goal to learn the output of the nth layer of the pre-trained teacher image processing model, knowledge distillation training is performed on the student image processing model to be trained to obtain the compressed student image processing model.

[0045] In this embodiment, after the sample images are respectively input into the pre-trained teacher image processing model and the student image processing model to be trained, the output of the nth layer of the student image processing model to be trained is used as the goal to learn the output of the nth layer of the pre-trained teacher image processing model, and the student image processing model to be trained is subjected to knowledge distillation training to obtain the trained student image processing model as the compressed student image processing model.

[0046] In this embodiment, knowledge distillation is performed based on the outputs of the intermediate layers corresponding to the pre-trained teacher image processing model and the student image processing model to be trained. In this way, even when the teacher model is relatively complex (which means that the distillation process itself may consume a lot of computing resources) or the quality is not high, the image domain model compression method based on rank-preserving decomposition and intermediate-layer knowledge distillation proposed in this embodiment can be used to perform knowledge distillation only in the intermediate layer of the model: the pre-trained image processing model is used as the teacher model, and the student image processing model to be trained after the rank-preserving decomposition is distilled, and the learning of the student image processing model is optimized through the output difference of each layer, thereby ensuring that the compressed student image processing model can maintain similar performance to the image processing model, realizing compression of the image processing model, and greatly compressing and reducing the storage requirements of the model while effectively maintaining the expressive power of the original model, thereby overcoming the technical defects of the related technology of only using knowledge distillation alone, being limited by the capacity of the student model, and losing key information in the distillation process.

[0047] In combination with the above embodiments, in one implementation, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S21 at least includes step S41, and the above step S22 at least includes step S42: Step S41: According to the first target parameter quantity, determine the size of the first pair of Kronecker embedding matrices is , the second pair of Kronecker embedding matrix dimensions are .

[0048] In this embodiment, the pre-trained teacher image processing model is a Transformer model, such as a DETR target detector, a ViT (Vision Transformer) model, etc., without limitation. The pre-trained teacher image processing model includes an embedding layer, and the embedding layer includes a large lookup table matrix ,in, is the vocabulary size, d is the embedding dimension, and n is the maximum factor of d. The target parameter quantity at least includes: a first target parameter quantity, which represents the target parameter quantity of the embedding layer.

[0049] In this embodiment, the size of the first pair of Kronecker embedding matrices may be determined according to the first target parameter quantity: , the second pair of Kronecker embedding matrix dimensions are That is, when performing Kronecker decomposition on the embedding layer, the scale parameter determined based on the first target parameter quantity is .

[0050] Step S42: Perform Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model to obtain a Kronecker embedding matrix group, and use it as the embedding layer of the student image processing model to be trained. The Kronecker embedding matrix group includes: a matrix that meets the size of the first pair of Kronecker embedding matrices and a matrix that meets the size of the second pair of Kronecker embedding matrices.

[0051] In this embodiment, the lookup table matrix of the embedding layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a Kronecker embedding matrix group, and use the Kronecker embedding matrix group as a parameter matrix of an embedding layer of a student image processing model to be trained, thereby obtaining an embedding layer of a student image processing model to be trained. The Kronecker embedding matrix group includes: a matrix that meets the size of the first pair of Kronecker embedding matrices and , and , matrices that correspond to the dimensions of the second pair of Kronecker embedding matrices and ; That is, X = Specifically, this embodiment can perform the Kronecker decomposition according to the specific steps of the above-mentioned embodiment, and perform the lookup table matrix of the embedding layer of the pre-trained teacher image processing model. Perform a Kronecker decomposition and obtain the set of Kronecker embedding matrices.

[0052] In this embodiment, the size of the first pair of Kronecker embedding matrices is determined as , the size of the second pair of Kronecker embedding matrices is determined as , which can make the matrix in the decomposition and Degenerates into a row vector, so that the embedding of each word in the lookup table can be separated from each other, in the matrix and The i-th row in represents the information related to the embedding of the i-th word independently.

[0053] In combination with the above embodiments, in one implementation, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S21 also includes steps S51 to S53, and the above step S22 also includes step S54: Step S51: In the multi-head self-attention layer of the pre-trained teacher image processing model, the weight matrices of the key matrix K, the query matrix Q and the value matrix V of all the self-attention heads are concatenated to obtain the concatenated weight matrix , and .

[0054] In this embodiment, the pre-trained teacher image processing model includes a multi-head self-attention layer (MHA), in which the self-attention mechanism maps the input to a key matrix K, a query matrix Q, and a value matrix V, and obtains the final attention matrix through an attention operation and a softmax layer. Among them, K, Q, and matrix V are obtained by multiplying the input by their respective weight matrices. In the multi-head self-attention (MHA) layer, each attention head has an independent set of weight matrices to allow richer data representation.

[0055] In this embodiment, in the multi-head self-attention layer of the pre-trained teacher image processing model, the weight matrices corresponding to the key matrix K, query matrix Q and value matrix V of all self-attention heads are spliced ​​respectively to obtain the spliced ​​weight matrix , and .

[0056] Step S52: According to the second target parameter, determine the number of rows of the second pair of Kronecker attention weight matrices as the concatenated weight matrix , and The greatest common divisor of the number of rows, and determine the number of columns of the second pair of Kronecker attention weight matrices is the concatenated weight matrix , and The greatest common divisor of the number of columns.

[0057] In this embodiment, the target parameter also includes: a second target parameter, which represents the target parameter of the multi-head self-attention layer. According to the second target parameter, the number of rows of the second pair of Kronecker attention weight matrices can be determined as the weight matrix after splicing. , and number of rows The greatest common divisor of , and determine the number of columns of the second pair of Kronecker attention weight matrices is the concatenated weight matrix , and Number of columns That is, when performing Kronecker decomposition on a multi-head self-attention layer, the scale parameter determined based on the second objective parameter is ( , , The greatest common divisor of greatest common divisor of ).

[0058] Step S53: According to the concatenated weight matrix , and The size of the first pair of Kronecker attention weight matrices is determined by the size of the second pair of Kronecker attention weight matrices.

[0059] In this embodiment, after determining the size of the second pair of Kronecker attention weight matrices (including: the number of rows of the second pair of Kronecker attention weight matrices and the number of columns of the second pair of Kronecker attention weight matrices), the weight matrix after splicing can be , and Size ( , and The number of rows and columns is the size of the matrix to be decomposed ), and the size of the second pair of Kronecker attention weight matrices, to obtain the size of the first pair of Kronecker attention weight matrices (including: the number of rows of the first pair of Kronecker attention weight matrices and the number of columns of the first pair of Kronecker attention weight matrices). In a specific example, it can be through , and The number of rows of divided by the number of rows of the second pair of Kronecker attention weight matrices, respectively, to obtain the number of rows of the three first pair of Kronecker attention weight matrices; through , and The number of columns of is divided by the number of columns of the second pair of Kronecker attention weight matrices to obtain the number of columns of the three first pair of Kronecker attention weight matrices.

[0060] Step S54: The weight matrix after splicing , and Perform Kronecker decomposition separately to obtain the Kronecker attention weight matrix group , and , and serves as the multi-head self-attention layer of the student image processing model to be trained.

[0061] In this embodiment, the weight matrix after splicing the multi-head self-attention layer of the pre-trained teacher image processing model is , and Perform Kronecker decomposition to obtain the Kronecker attention weight matrix group , and , and the Kronecker attention weight matrix is , and As the parameter matrix of the multi-head self-attention layer of the student image processing model to be trained, the multi-head self-attention layer of the student image processing model to be trained is obtained. Wherein, the Kronecker attention weight matrix is , and They respectively include: a matrix that conforms to the size of the first pair of Kronecker attention weight matrices, and a matrix that conforms to the size of the second pair of Kronecker attention weight matrices. Specifically, this embodiment can perform the Kronecker decomposition according to the specific steps of the above embodiment to perform the concatenated weight matrix. , and Perform Kronecker decomposition separately to obtain the Kronecker attention weight matrix group , and .

[0062] Example, Kronecker attention weight matrix group Contains: A matrix that matches the size of the first pair of Kronecker attention weight matrices and , and , matrices that correspond to the dimensions of the second pair of Kronecker embedding matrices and ;Right now, Kronecker attention weight matrix group Include: A matrix that matches the dimensions of the first pair of Kronecker attention weight matrices and , and , matrices that correspond to the dimensions of the second pair of Kronecker embedding matrices and ;Right now, Kronecker attention weight matrix group Include: A matrix that matches the dimensions of the first pair of Kronecker attention weight matrices and , and , matrices that correspond to the dimensions of the second pair of Kronecker embedding matrices and ;Right now, .in, and , and ,as well as, and The sizes can be the same or different; and , and ,as well as, and The same size.

[0063] In this implementation, the concatenated weight matrix in the multi-head self-attention layer is , and After performing Kronecker decomposition, the three Kronecker attention weight matrix groups obtained are , and The two corresponding B matrices ( and , and ,as well as, and ), that is, the size of the second pair of Kronecker attention weight matrices is the same, making the decomposed information more consistent.

[0064] In combination with the above embodiments, in one implementation, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S21 also includes steps S61 to S62, and the above step S22 also includes step S63: Step S61: According to the third target parameter, determine the number of rows of the second pair of Kronecker linear weight matrices as the linear mapping matrix in the feedforward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix The greatest common divisor of the number of rows and the number of columns of the second pair of Kronecker linear weight matrices are the linear mapping matrices in the feedforward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix The greatest common divisor of the number of columns.

[0065] In this embodiment, the pre-trained teacher image processing model includes: a feedforward neural network layer (FFN), the attention matrix finally obtained by the multi-head self-attention layer will be immediately passed to the feedforward neural network layer, and in the feedforward neural network layer, the attention matrix will be passed to a linear mapping matrix And the following two weight matrices: The first weight matrix and the second weight matrix In this embodiment, the Kronecker decomposition of the parameter matrix of the feedforward neural network layer is similar to the Kronecker decomposition of the parameter matrix of the multi-head self-attention layer in the aforementioned embodiment.

[0066] Specifically, the target parameter quantity in this embodiment also includes: a third target parameter quantity, which represents the target parameter quantity of the feedforward neural network layer. According to the third target parameter quantity, the number of rows of the second pair of Kronecker linear weight matrices can be determined as the linear mapping matrix in the feedforward neural network layer of the pre-trained teacher image processing model. , the first weight matrix and the second weight matrix The greatest common divisor of the number of rows and the number of columns of the second pair of Kronecker linear weight matrices is the linear mapping matrix in the feedforward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix That is, when performing Kronecker decomposition on a feedforward neural network layer, the scale parameter determined based on the third objective parameter quantity is ( , , linear mapping matrix , the first weight matrix and the second weight matrix The greatest common divisor of the number of rows, the linear mapping matrix , the first weight matrix and the second weight matrix The greatest common divisor of the number of columns).

[0067] Step S62: According to the linear mapping matrix , the first weight matrix and the second weight matrix The dimensions of the first pair of Kronecker linear weight matrices are determined by the dimensions of the second pair of Kronecker linear weight matrices.

[0068] In this embodiment, after determining the size of the second pair of Kronecker linear weight matrices (including: the number of rows of the second pair of Kronecker linear weight matrices and the number of columns of the second pair of Kronecker linear weight matrices), the linear mapping matrix , the first weight matrix and the second weight matrix Size ( , and The number of rows and columns is the size of the matrix to be decomposed ), and the size of the second pair of Kronecker linear weight matrices, to obtain the size of the first pair of Kronecker linear weight matrices (including: the number of rows of the first pair of Kronecker linear weight matrices and the number of columns of the first pair of Kronecker linear weight matrices). In a specific example, it can be through , and The number of rows of divided by the number of rows of the second pair of Kronecker linear weight matrices, respectively, to obtain the number of rows of the three first pair of Kronecker linear weight matrices; by , and The number of columns of is divided by the number of columns of the second pair of Kronecker linear weight matrices to obtain the number of columns of the three first pair of Kronecker linear weight matrices.

[0069] Step S63: The linear mapping matrix , the first weight matrix and the second weight matrix Perform Kronecker decomposition separately to obtain the Kronecker linear weight matrix group , and , and serves as the feedforward neural network layer of the student image processing model to be trained.

[0070] In this embodiment, the linear mapping matrix of the feedforward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix Perform Kronecker decomposition separately to obtain the Kronecker linear weight matrix group , and , and the Kronecker linear weight matrix is , and As the parameter matrix of the feedforward neural network layer of the student image processing model to be trained, the feedforward neural network layer of the student image processing model to be trained is obtained. , and The linear mapping matrix φ(x) is a matrix having a size that matches the first pair of Kronecker linear weight matrices, and the matrix having a size that matches the second pair of Kronecker linear weight matrices. , the first weight matrix and the second weight matrix Perform Kronecker decomposition separately to obtain the Kronecker linear weight matrix group , and .

[0071] Example, Kronecker linear weight matrix group Contains: A matrix that matches the dimensions of a first pair of Kronecker linear weight matrices and , and , matrices that correspond to the dimensions of the second pair of Kronecker embedding matrices and ;Right now, . Kronecker linear weight matrix group Include: A matrix that matches the dimensions of the first pair of Kronecker attention weight matrices and , and , matrices that correspond to the dimensions of the second pair of Kronecker embedding matrices and ;Right now, . Kronecker linear weight matrix group Include: A matrix that matches the dimensions of the first pair of Kronecker attention weight matrices and , and , matrices that correspond to the dimensions of the second pair of Kronecker embedding matrices and ;Right now, .in, and , and ,as well as, and The sizes can be the same or different; and , and ,as well as, and The same size.

[0072] In this implementation, the linear mapping matrix in the feedforward neural network layer is , the first weight matrix and the second weight matrix After performing Kronecker decomposition, the three Kronecker linear weight matrices obtained are , and The two corresponding B matrices ( and , and ,as well as, and ) are consistent in scale, that is, the second pair of Kronecker linear weight matrices have the same size, making the decomposed information more consistent.

[0073] In combination with the above embodiments, in one implementation, the present invention further provides an image domain model compression method based on rank-preserving decomposition and knowledge distillation. In this method, the above step S32 may specifically include steps S71 to S74: Step S71: After the pre-trained teacher image processing model processes the sample image, the first image embedding representation output by the embedding layer of the pre-trained teacher image processing model, the first attention map output by the multi-head self-attention layer of the pre-trained teacher image processing model, and the first image feature representation output by the feedforward neural network layer of the pre-trained teacher image processing model are obtained.

[0074] In this embodiment, knowledge distillation can be performed on the intermediate layer of the student image processing model to be trained after Kronecker decomposition. Specifically, in this embodiment, the same sample image is input into the pre-trained teacher image processing model and the student image processing model to be trained for knowledge distillation. The sample image is an image used to train the student image processing model.

[0075] In this embodiment, a sample image is input into a pre-trained teacher image processing model, and the sample image is processed by the pre-trained teacher image processing model to obtain the outputs of the embedding layer, the multi-head self-attention layer, and the feedforward neural network layer in the pre-trained teacher image processing model, respectively; that is, after the pre-trained teacher image processing model processes the sample image, the first image embedding representation output by the embedding layer of the pre-trained teacher image processing model, the first attention map output by the multi-head self-attention layer of the pre-trained teacher image processing model, and the first image feature representation output by the feedforward neural network layer of the pre-trained teacher image processing model are obtained.

[0076] Step S72: After the student image processing model to be trained processes the sample image, the second image embedding representation output by the embedding layer of the student image processing model to be trained, the second attention map output by the multi-head self-attention layer of the student image processing model to be trained, and the second image feature representation output by the feedforward neural network layer of the student image processing model to be trained are obtained.

[0077] In this embodiment, a sample image is input into the student image processing model to be trained, and the sample image is processed by the student image processing model to be trained to obtain the outputs of the embedding layer, the multi-head self-attention layer and the feedforward neural network layer in the student image processing model to be trained, respectively; that is, after the sample image is processed by the student image processing model to be trained, the second image embedding representation output by the embedding layer of the student image processing model to be trained, the second attention map output by the multi-head self-attention layer of the student image processing model to be trained, and the second image feature representation output by the feedforward neural network layer of the student image processing model to be trained are obtained.

[0078] In this embodiment, after the input sample image enters the student image processing model to be trained after Kronecker decomposition, the input of each layer (embedding layer, multi-head self-attention layer or feedforward neural network layer) needs to be operated with the elements after Kronecker decomposition (i.e., Kronecker embedding matrix group, Kronecker attention weight matrix group or Kronecker linear weight matrix group) in each layer (instead of multiplying with the original corresponding weight matrix). This part of the operation can be simplified by the calculation formula related to the Kronecker product, and most importantly, compared with the pre-trained teacher image processing model and the student image processing model to be trained in this embodiment, the input and output dimensions of each layer remain unchanged for the number of layers of the model after the operation, which can facilitate alignment for knowledge distillation.

[0079] Step S73: Determine a distillation loss based on the first image embedding representation and the second image embedding representation, the first attention map and the second attention map, and the first image feature representation and the second image feature representation.

[0080] In this embodiment, when performing knowledge distillation, the difference between the output of a certain layer of the pre-trained teacher image processing model and the student image processing model to be trained can be directly obtained without any projection. Specifically, the distillation process of this embodiment only needs to calculate the loss of the embedding layer output, the multi-head self-attention layer output (that is, the attention matrix) and the feedforward neural network layer output of the two: based on the first image embedding representation and the second image embedding representation, the first attention map and the second attention map, and the first image feature representation and the second image feature representation, the distillation loss is calculated.

[0081] For example, an embedding loss can be obtained based on the first image embedding representation and the second image embedding representation, an attention loss can be obtained based on the first attention map and the second attention map, and an image feature loss can be obtained based on the first image feature representation and the second image feature representation. Then, the embedding loss, the attention loss, and the image feature loss are added together (e.g., weighted) as the overall loss function of knowledge distillation to obtain the distillation loss, thereby performing a concise knowledge distillation process.

[0082] Step S74: Based on the distillation loss, the model parameters of the student image processing model to be trained are updated to obtain the compressed student image processing model.

[0083] In this embodiment, the model parameters of the student image processing model to be trained can be updated based on the distillation loss until the distillation loss converges, and the model parameters of the student image processing model to be trained are fixed to obtain a trained student image processing model. The trained student image processing model obtained after distillation in this embodiment is the compressed student image processing model finally obtained in this embodiment, which effectively maintains the expressive power of the original pre-trained teacher image processing model and greatly compresses and reduces the storage requirements of the model.

[0084] In this embodiment, knowledge distillation is performed only on the embedding layer output, multi-head self-attention layer output (attention matrix) and feedforward neural network layer output related to Kronecker decomposition, and no scale projection is required. It can be completed by using the simplest knowledge distillation locally, which further compresses the model parameters and saves the storage space and computing resources of the target device where the compressed student image processing model is deployed, thereby achieving lightweight deployment of the image processing model while maintaining stable model performance.

[0085] In one embodiment, if Figure 3 As shown, Figure 3 FIG. 1 is a schematic diagram of the overall process of an image domain model compression method based on rank-preserving decomposition and intermediate layer knowledge distillation according to an embodiment of the present invention. Figure 3 In the example, the pre-trained teacher image processing model to be decomposed is a Transformer model, and the Transformer architecture of the Transformer model includes at least: an embedding layer, a multi-head self-attention module (i.e., a multi-head self-attention layer), and a feedforward neural network layer, wherein the embedding layer includes an embedding matrix X, and the multi-head self-attention module includes an attention weight (i.e., a concatenated weight matrix , and ), and the feedforward neural network layer includes linear weights (i.e., the linear mapping matrix , the first weight matrix and the second weight matrix ).

[0086] Among them, when performing Kronecker decomposition on the pre-trained teacher image processing model, the embedding matrix X in the embedding layer and the multi-head self-attention module are respectively , and and in the feedforward neural network layer , and Perform Kronecker decomposition and obtain Kronecker embeddings (i.e., Kronecker embedding matrix groups: ), Kronecker attention weight (i.e., Kronecker attention weight matrix group: , and ) and Kronecker linear weights (i.e., the Kronecker linear weight matrix group: , and ) are respectively used as the parameter matrix of the embedding layer of the student image processing model to be trained, the parameter matrix of the multi-head self-attention layer of the student image processing model to be trained, and the parameter matrix of the feedforward neural network layer of the student image processing model to be trained, thereby obtaining the embedding layer, multi-head self-attention layer and feedforward neural network layer of the student image processing model to be trained, so as to obtain the student image processing model to be trained.

[0087] When performing knowledge distillation, the intermediate-layer knowledge distillation is performed on the student image processing model to be trained after decomposition: the entire distillation process only needs to calculate the loss of the embedding layer output, the multi-head self-attention layer output (that is, the attention matrix) and the feedforward neural network layer output between the student model (the student image processing model to be trained) and the teacher model (the pre-trained teacher image processing model), so as to obtain the distilled student model, that is, the compressed student image processing model, thereby realizing the compression of the Transformer model based on rank-preserving decomposition and intermediate-layer knowledge distillation, which greatly reduces the storage requirements of the model while effectively maintaining the expressiveness of the original model.

[0088] This embodiment is based on the rank-preserving decomposition of the matrix (i.e., Kronecker decomposition) and its numerical calculation method, and is combined with the knowledge distillation of only the local elements in the middle layers of the model (embedding layer, multi-head self-attention layer, and feedforward neural network layer). Through Kronecker decomposition, the matrix of the key part of the teacher model will be decomposed into the form of Kronecker product, thereby reducing the model parameters. This decomposition method is a rank-preserving transformation of the original matrix. The decomposed matrix is ​​usually equal to or close to the rank of the original matrix, so that the information in the original matrix can be effectively retained to avoid performance loss. Next, by performing knowledge distillation only in the middle layer of the model, using the pre-trained Transformer model as the teacher model, the model that has been rank-preserving decomposition is distilled, and the learning of the student model is optimized through the output difference of each layer, thereby ensuring that the compressed model can maintain performance similar to that of the teacher model, thereby saving the storage space and computing resources of the target device to be deployed with the compressed student image processing model, and realizing the lightweight deployment of the image processing model with stable model performance.

[0089] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0090] Based on the same inventive concept, an embodiment of the present invention provides an image domain model compression device based on rank-preserving decomposition and knowledge distillation. Figure 4 , Figure 4 1 is a structural block diagram of an image domain model compression device based on rank-preserving decomposition and knowledge distillation provided by an embodiment of the present invention. Figure 4 As shown, the image domain model compression device based on rank-preserving decomposition and knowledge distillation in this embodiment may include: A parameter acquisition module, used to acquire device parameter values ​​of a target device to be deployed with the compressed student image processing model, wherein the device parameter values ​​include at least one of the following: storage space size and operation speed; A parameter determination module, used for determining a target parameter amount of the student image processing model according to the device parameter value; A model decomposition module, used for performing Kronecker decomposition on the pre-trained teacher image processing model according to the target parameter amount, and using the model after performing Kronecker decomposition as the student image processing model to be trained; A model training module, for performing knowledge distillation training on the student image processing model to be trained using sample images and taking the pre-trained teacher image processing model as a learning target to obtain the compressed student image processing model, wherein the pre-trained teacher image processing model is used for any of the following: image classification, image segmentation, and object detection; The image processing module is used to input the image to be processed into the compressed student image processing model to obtain an image processing result, which is any one of the following: image classification result, image segmentation result, and target detection result.

[0091] Optionally, the model decomposition module includes: A size determination module is used to determine the size of the first pair of Kronecker matrices according to the target parameter quantity. , the second pair of Kronecker matrix dimensions are ; A parameter determination module for determining the parameter matrix of the nth layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a Kronecker parameter matrix group, and use it as the parameter matrix of the nth layer of the student image processing model to be trained, wherein the Kronecker parameter matrix group includes: a matrix that meets the size of the first pair of Kronecker matrices and , and , correspond to the second pair of Kronecker matrix dimensions of the matrix and ; in, , n is an integer greater than 0.

[0092] Optionally, the model training module includes: An image input module, used for inputting the sample image into the pre-trained teacher image processing model, and inputting the sample image into the student image processing model to be trained; The knowledge distillation module is used to perform knowledge distillation training on the student image processing model to be trained with the output of the nth layer of the student image processing model to be trained as the goal of learning the output of the nth layer of the pre-trained teacher image processing model to obtain the compressed student image processing model.

[0093] Optionally, the pre-trained teacher image processing model is a Transformer model; The size determination module comprises at least: The first size determination module is used to determine the size of the first pair of Kronecker embedding matrices according to the first target parameter quantity. , the second pair of Kronecker embedding matrix dimensions are ; The parameter determination module at least includes: An embedding layer decomposition module, used for performing Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model to obtain a Kronecker embedding matrix group as the embedding layer of the student image processing model to be trained, wherein the Kronecker embedding matrix group includes: a matrix that meets the size of the first pair of Kronecker embedding matrices and a matrix that meets the size of the second pair of Kronecker embedding matrices; in, is the vocabulary size, d is the embedding dimension, and n is the largest factor of d.

[0094] Optionally, the size determination module further includes: A matrix concatenation module is used to concatenate the key matrix K, query matrix Q and value matrix V of all self-attention heads in the multi-head self-attention layer of the pre-trained teacher image processing model to obtain a concatenated weight matrix , and ; The first row and column determination module is used to determine the number of rows of the second pair of Kronecker attention weight matrices as the concatenated weight matrix according to the second target parameter. , and The greatest common divisor of the number of rows, and determine the number of columns of the second pair of Kronecker attention weight matrices is the concatenated weight matrix , and The greatest common divisor of the number of columns; The second size determination module is used to determine the weight matrix after the splicing , and The size of the first pair of Kronecker attention weight matrices is determined by the size of the second pair of Kronecker attention weight matrices; The parameter determination module also includes: The attention layer decomposition module is used to decompose the concatenated weight matrix , and Perform Kronecker decomposition separately to obtain the Kronecker attention weight matrix group , and , and as the multi-head self-attention layer of the student image processing model to be trained, the Kronecker attention weight matrix group , and Respectively include: a matrix that conforms to the size of the first pair of Kronecker attention weight matrices, and a matrix that conforms to the size of the second pair of Kronecker attention weight matrices.

[0095] Optionally, the size determination module further includes: The second row and column determination module is used to determine the number of rows of the second pair of Kronecker linear weight matrices as the linear mapping matrix in the feedforward neural network layer of the pre-trained teacher image processing model according to the third target parameter. , the first weight matrix and the second weight matrix The greatest common divisor of the number of rows and the number of columns of the second pair of Kronecker linear weight matrices are the linear mapping matrices in the feedforward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix The greatest common divisor of the number of columns; The third size determination module is used to determine the size of the linear mapping matrix according to the linear mapping matrix. , the first weight matrix and the second weight matrix The size of the first pair of Kronecker linear weight matrices is determined by the size of the second pair of Kronecker linear weight matrices; The parameter determination module also includes: The feedforward neural network layer decomposition module is used to transform the linear mapping matrix , the first weight matrix and the second weight matrix Perform Kronecker decomposition separately to obtain the Kronecker linear weight matrix group , and , and as the feedforward neural network layer of the student image processing model to be trained, the Kronecker linear weight matrix group , and Respectively include: a matrix that conforms to the size of the first pair of Kronecker linear weight matrices, and a matrix that conforms to the size of the second pair of Kronecker linear weight matrices.

[0096] Optionally, a knowledge distillation module includes: A first acquisition module is used to acquire, after the pre-trained teacher image processing model processes the sample image, a first image embedding representation output by the embedding layer of the pre-trained teacher image processing model, a first attention map output by the multi-head self-attention layer of the pre-trained teacher image processing model, and a first image feature representation output by the feedforward neural network layer of the pre-trained teacher image processing model; A second acquisition module is used to acquire, after the student image processing model to be trained processes the sample image, a second image embedding representation output by the embedding layer of the student image processing model to be trained, a second attention map output by the multi-head self-attention layer of the student image processing model to be trained, and a second image feature representation output by the feedforward neural network layer of the student image processing model to be trained; A loss calculation module, configured to determine a distillation loss based on the first image embedding representation and the second image embedding representation, the first attention map and the second attention map, and the first image feature representation and the second image feature representation; A parameter updating module is used to update the model parameters of the student image processing model to be trained based on the distillation loss to obtain the compressed student image processing model.

[0097] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the image domain model compression method based on rank-preserving decomposition and knowledge distillation as described in any of the above embodiments of the present invention are implemented.

[0098] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, such as Figure 5 shown. Figure 5 1 is a schematic diagram of an electronic device shown in an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes, the steps in the image domain model compression method based on rank-preserving decomposition and knowledge distillation described in any of the above embodiments of the present invention are implemented.

[0099] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0100] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0101] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, embodiments of the present invention may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Furthermore, embodiments of the present invention may take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0102] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0103] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0105] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0106] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0107] The above is a detailed introduction to the image domain model compression method, device, equipment and medium based on rank-preserving decomposition and knowledge distillation provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. An image domain model compression method based on rank-preserving decomposition and knowledge distillation, characterized in that: The method comprises: Obtaining device parameter values ​​of a target device to which the compressed student image processing model is to be deployed, wherein the device parameter values ​​include at least one of the following: storage space size and computing speed; Determining target parameter values ​​of the student image processing model according to the device parameter values; According to the target parameter amount, Kronecker decomposition is performed on the pre-trained teacher image processing model, and the model after the Kronecker decomposition is used as the student image processing model to be trained; Using sample images and taking the pre-trained teacher image processing model as a learning target, performing knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model, wherein the pre-trained teacher image processing model is used for any of the following: image classification, image segmentation, and object detection; The image to be processed is input into the compressed student image processing model to obtain an image processing result, which is any one of the following: an image classification result, an image segmentation result, and a target detection result.

2. The image domain model compression method based on rank-preserving decomposition and knowledge distillation according to claim 1 is characterized in that: According to the target parameter amount, Kronecker decomposition is performed on the pre-trained teacher image processing model, and the model after Kronecker decomposition is used as the student image processing model to be trained, including: According to the target parameter quantity, determine the first pair of Kronecker matrix sizes are , the second pair of Kronecker matrix dimensions are ; The parameter matrix of the nth layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a Kronecker parameter matrix group, and use it as the parameter matrix of the nth layer of the student image processing model to be trained, wherein the Kronecker parameter matrix group includes: a matrix that meets the size of the first pair of Kronecker matrices and , and , correspond to the second pair of Kronecker matrix dimensions of the matrix and ; in, , n is an integer greater than 0.

3. The image domain model compression method based on rank-preserving decomposition and knowledge distillation according to claim 2 is characterized in that: Using the sample image and taking the pre-trained teacher image processing model as the learning target, the student image processing model to be trained is subjected to knowledge distillation training to obtain the compressed student image processing model, including: Inputting the sample image into the pre-trained teacher image processing model, and inputting the sample image into the student image processing model to be trained; With the goal of learning the output of the nth layer of the pre-trained teacher image processing model through the output of the nth layer of the student image processing model to be trained, knowledge distillation training is performed on the student image processing model to be trained to obtain the compressed student image processing model.

4. The image domain model compression method based on rank-preserving decomposition and knowledge distillation according to claim 3 is characterized in that: The pre-trained teacher image processing model is a Transformer model; According to the target parameter quantity, determine the first pair of Kronecker matrix sizes are , the second pair of Kronecker matrix dimensions are , including at least: According to the first target parameter quantity, determine the first pair of Kronecker embedding matrix sizes are , the second pair of Kronecker embedding matrix dimensions are ; The parameter matrix of the nth layer of the pre-trained teacher image processing model Perform Kronecker decomposition to obtain a Kronecker parameter matrix group, and use it as the parameter matrix of the nth layer of the student image processing model to be trained, which at least includes: Performing Kronecker decomposition on the lookup table matrix of the embedding layer of the pre-trained teacher image processing model to obtain a Kronecker embedding matrix group, and using it as the embedding layer of the student image processing model to be trained, wherein the Kronecker embedding matrix group includes: a matrix that meets the size of the first pair of Kronecker embedding matrices and a matrix that meets the size of the second pair of Kronecker embedding matrices; in, is the vocabulary size, d is the embedding dimension, and n is the largest factor of d.

5. The image domain model compression method based on rank-preserving decomposition and knowledge distillation according to claim 4 is characterized in that: According to the target parameter quantity, determine the first pair of Kronecker matrix sizes are , the second pair of Kronecker matrix dimensions are , also includes: In the multi-head self-attention layer of the pre-trained teacher image processing model, the weight matrices of the key matrix K, query matrix Q and value matrix V of all self-attention heads are concatenated to obtain the concatenated weight matrix , and ; According to the second target parameter, the number of rows of the second pair of Kronecker attention weight matrices is determined to be the concatenated weight matrix , and The greatest common divisor of the number of rows, and determine the number of columns of the second pair of Kronecker attention weight matrices is the concatenated weight matrix , and The greatest common divisor of the number of columns; According to the concatenated weight matrix , and The size of the first pair of Kronecker attention weight matrices is determined by the size of the second pair of Kronecker attention weight matrices; The parameter matrix of the nth layer of the pre-trained teacher image processing model Performing Kronecker decomposition to obtain a Kronecker parameter matrix group, and using it as the parameter matrix of the nth layer of the student image processing model to be trained, further comprising: The concatenated weight matrix , and Perform Kronecker decomposition separately to obtain the Kronecker attention weight matrix group , and , and as the multi-head self-attention layer of the student image processing model to be trained, the Kronecker attention weight matrix group , and Respectively include: a matrix that conforms to the size of the first pair of Kronecker attention weight matrices, and a matrix that conforms to the size of the second pair of Kronecker attention weight matrices.

6. The image domain model compression method based on rank-preserving decomposition and knowledge distillation according to claim 5, characterized in that: According to the target parameter quantity, determine the first pair of Kronecker matrix sizes are , the second pair of Kronecker matrix dimensions are , also includes: According to the third target parameter, the number of rows of the second pair of Kronecker linear weight matrices is determined to be the linear mapping matrix in the feedforward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix The greatest common divisor of the number of rows and the number of columns of the second pair of Kronecker linear weight matrices are the linear mapping matrices in the feedforward neural network layer of the pre-trained teacher image processing model , the first weight matrix and the second weight matrix The greatest common divisor of the number of columns; According to the linear mapping matrix , the first weight matrix and the second weight matrix The size of the first pair of Kronecker linear weight matrices is determined by the size of the second pair of Kronecker linear weight matrices; The parameter matrix of the nth layer of the pre-trained teacher image processing model Performing Kronecker decomposition to obtain a Kronecker parameter matrix group, and using it as the parameter matrix of the nth layer of the student image processing model to be trained, further comprising: For the linear mapping matrix , the first weight matrix and the second weight matrix Perform Kronecker decomposition separately to obtain the Kronecker linear weight matrix group , and , and as the feedforward neural network layer of the student image processing model to be trained, the Kronecker linear weight matrix group , and Respectively include: a matrix that conforms to the size of the first pair of Kronecker linear weight matrices, and a matrix that conforms to the size of the second pair of Kronecker linear weight matrices.

7. The image domain model compression method based on rank-preserving decomposition and knowledge distillation according to any one of claims 3 to 6, characterized in that: The method aims to learn the output of the nth layer of the pre-trained teacher image processing model from the output of the nth layer of the student image processing model to be trained, and performs knowledge distillation training on the student image processing model to be trained to obtain the compressed student image processing model, including: After the pre-trained teacher image processing model processes the sample image, a first image embedding representation output by an embedding layer of the pre-trained teacher image processing model, a first attention map output by a multi-head self-attention layer of the pre-trained teacher image processing model, and a first image feature representation output by a feedforward neural network layer of the pre-trained teacher image processing model are obtained; After the student image processing model to be trained processes the sample image, a second image embedding representation output by the embedding layer of the student image processing model to be trained, a second attention map output by the multi-head self-attention layer of the student image processing model to be trained, and a second image feature representation output by the feedforward neural network layer of the student image processing model to be trained are obtained; Determining a distillation loss based on the first image embedding representation and the second image embedding representation, the first attention map and the second attention map, and the first image feature representation and the second image feature representation; Based on the distillation loss, the model parameters of the student image processing model to be trained are updated to obtain the compressed student image processing model.

8. An image domain model compression device based on rank-preserving decomposition and knowledge distillation, characterized in that: The device comprises: A parameter acquisition module, used to acquire device parameter values ​​of a target device to be deployed with the compressed student image processing model, wherein the device parameter values ​​include at least one of the following: storage space size and operation speed; A parameter determination module, used for determining a target parameter amount of the student image processing model according to the device parameter value; A model decomposition module, used for performing Kronecker decomposition on the pre-trained teacher image processing model according to the target parameter amount, and using the model after performing Kronecker decomposition as the student image processing model to be trained; A model training module, for performing knowledge distillation training on the student image processing model to be trained using sample images and taking the pre-trained teacher image processing model as a learning target to obtain the compressed student image processing model, wherein the pre-trained teacher image processing model is used for any of the following: image classification, image segmentation, and object detection; The image processing module is used to input the image to be processed into the compressed student image processing model to obtain an image processing result, which is any one of the following: image classification result, image segmentation result, and target detection result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the image domain model compression method based on rank-preserving decomposition and knowledge distillation as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image domain model compression method based on rank-preserving decomposition and knowledge distillation as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Large-model knowledge distillation low-rank adaptation federated learning method, electronic equipment and readable storage medium

    CN118070876A

  • Over-parameterized knowledge distillation method, device, equipment and medium

    CN119250176A

  • Neural network model compression method, corpus translation method and device

    EP3812969A1

  • Method and apparatus with teacherless student model for classification

    US20240242082A1