Network model lightweight method, device, equipment, medium and program product
By performing random channel pruning and similarity analysis on the original network model, a fusion matrix is generated to determine the post-fusion weight, the problems of knowledge loss and isomorphic model fusion limitations of the network model during edge device deployment are solved, and efficient network model lightweighting and heterogeneous model fusion are achieved.
Patent Information
- Application Number
- CN202510378417.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In the prior art, network models need to be lightweight when deploying edge devices, resulting in knowledge loss and fusion problems that only support isomorphic network models.
By randomly pruning the target layer output channel of the original network model, analyzing the similarity between channels based on the curved wave transformation characteristics and structural similarity, a fusion matrix is generated to determine the weight after fusion, and the fusion of heterogeneous models is achieved.
It avoids the loss of image recognition knowledge, improves the accuracy of image recognition, and supports the fusion of heterogeneous models, improving the ability and effect of network models to lighten the weight.
Smart Images

Figure CN119886233B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural network compression technology, and particularly to a method, device, equipment, medium and program product for lightweighting a network model. Background Art
[0002] In order to enable the deployment of a neural network on edge devices, it is necessary to lightweight the network model, which specifically involves model pruning and model fusion, etc.
[0003] In related technologies, scaling processing based on the kernel set theory is used to scale the original network layer, but this method will cause the loss of a part of image recognition knowledge. When performing model fusion, fusion methods such as ZipIt are used, but this method is only applicable to the fusion of homogeneous network models, that is, each model to be fused has an equal number of channels in the corresponding layer. Summary of the Invention
[0004] This application provides a method for lightweighting a network model to at least solve the problems of network knowledge loss and only supporting the fusion of homogeneous network models in related technologies.
[0005] This application provides a method for lightweighting a network model, including:
[0006] Performing random channel pruning on the output channels of the target layer of the original network model to determine the weights of the target layer of the compressed model corresponding to the original network model;
[0007] Based on the feature maps output by the original network model and the compressed model respectively in the target layer, analyzing the channel similarity metric according to the target similarity metric to determine the channel matching result; the target similarity metric includes structure and wavelet transform features, and the channel matching result represents a set of matching channel pairs;
[0008] Generating a fusion matrix according to the channel matching result, and determining the fused weights of the target layer based on the weights of the target layer of the original network model and the weights of the target layer of the compressed model;
[0009] Obtaining a fused model based on the fused weights corresponding to different network layers, and using the fused model as the final compressed model of the original network model.
[0010] This application also provides a device for lightweighting a network model, including:
[0011] A random channel pruning module, configured to perform random channel pruning on the output channels of the target layer of the original network model to determine the weights of the target layer of the compressed model corresponding to the original network model;
[0012] A channel matching module, configured to analyze the similarity measurement between channels based on the feature maps output by the original network model and the compressed model at the target layer according to the target similarity metric, so as to determine the channel matching result; the target similarity metric includes structural and curvelet transform features, and the channel matching result represents a set of matching channel pairs.
[0013] A fused weight determination module, configured to generate a fusion matrix according to the channel matching result, and determine the fused weight of the target layer based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model.
[0014] A final compressed model determination module, configured to obtain a fused model based on the fused weights corresponding to different network layers, and use the fused model as the final compressed model of the original network model.
[0015] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above network model lightweighting methods when executing the computer program.
[0016] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above network model lightweighting methods are implemented.
[0017] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above network model lightweighting methods are implemented.
[0018] In this application, by analyzing the channel similarity measurement between models based on the curvelet transform feature similarity and structural similarity, and merging the matching channels, the loss of image recognition knowledge can be avoided, and the accuracy of image recognition can be improved; a fusion matrix is generated based on the matching channel pairs, and model fusion is performed through channel similarity measurement, which does not depend on a specific model architecture, so it can be applied to the fusion of heterogeneous models, and the ability and effect of network model lightweighting are improved. Description of the Drawings
[0019] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a flowchart of a network model lightweighting method provided by this application.
[0021] Figure 2A flowchart of a specific network model lightweighting method provided by this application;
[0022] Figure 3 A schematic diagram of a specific channel fusion method under weak compression rate constraint provided by this application;
[0023] Figure 4 A schematic diagram of a specific channel fusion method under strong compression rate constraint provided by this application;
[0024] Figure 5 A schematic diagram of the structure of a network model lightweighting device provided by this application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of this application will be clearly and completely described with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0026] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0027] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be made with reference to the accompanying drawings and specific implementation manners.
[0028] The embodiments of this application provide a network model lightweighting method. Refer to Figure 1 as shown. This method may include the following steps:
[0029] Step S11: Perform random channel pruning on the output channels of the target layer of the original network model to determine the weights of the target layer of the compressed model corresponding to the original network model.
[0030] Pruning is a commonly used method for lightweighting image recognition networks. By removing neurons or parameter groups in the image recognition network, it reduces the number of parameters, storage, and floating-point computational volume during inference in the image recognition network, thereby enabling the deployment and inference acceleration of the image recognition network on resource-constrained edge devices. Specifically, pruning can be performed during training or inference. For example, during inference pruning, heuristic rules are used to select parameter groups with low importance for removal, and the least squares method is directly used to reconstruct the compressed model parameter matrix without relying on retraining to restore the accuracy of the compressed image recognition network, greatly improving the construction efficiency of the model lightweighting process itself. The above original network model is an image recognition network model.
[0031] In some embodiments, performing random channel pruning on the output channels of the target layer of the original network model and determining the weights of the target layer of the compressed model corresponding to the original network model includes: generating a random mask vector for the output channels of the target layer of the original network model, and performing channel pruning on the output channels of the target layer according to the random mask vector; that is, performing random output channel pruning on the target layer of the original network model, which means generating a random mask vector gr for the output channels of the target layer of the original network. According to the feature map output by the target layer of the original network model and the feature map output by the target layer of the compressed model, the least squares method is used to solve for the minimum error to determine the weights of the target layer of the compressed model. For example, the least squares method can be used to solve the optimization function of the feature map reconstruction error to obtain the weights of the target layer of the compressed network model, that is, by minimizing the difference between the feature map output by the target layer of the compressed model and the feature map output by the target layer of the original model to obtain the weights of the compressed network layer by solving , where X represents the input data of the target layer of the original model.
[0032] Step S12: Based on the feature maps output by the target layer of the original network model and the compressed model respectively, analyze the channel similarity metric according to the target similarity metric to determine the channel matching result; the target similarity metric includes structural and wavelet transform features, and the channel matching result represents a set of matching channel pairs.
[0033] In this embodiment, similarity metric learning is used to quantify the similarity between the output feature maps of the compressed model and the original model. By calculating similarity metrics, channels with high similarity in the compressed model and the original model are found, so as to retain the most important feature information in the subsequent fusion process. Specifically, two metrics, namely structural similarity and similarity of curvelet transform features, can be used to quantify the similarity between the output feature maps of the compressed model and the original model, for subsequent channel fusion or other optimization operations. In this way, the most important feature information can be retained during the model compression process, while reducing the computational amount and storage requirements of the model.
[0034] Among them, the inter-channel similarity metric includes cross-model similarity and intra-similarity. The cross-model similarity is to calculate the channel similarity between the feature map of the compressed model and the feature map of the original model, so as to evaluate whether the compressed feature map retains the characteristics of the original feature map. The intra-similarity includes calculating the similarity between channels within the feature map of the compressed model to evaluate the redundancy within the compressed feature map; and calculating the similarity between channels within the feature map of the original model to evaluate the redundancy within the original feature map. That is, similarity metric learning is performed on the output feature maps , of the target layers of the compressed model and the original model, that is, the similarity metrics between the channels of the output feature map of the target layer of the compressed model and the output feature map of the target layer of the original model are calculated respectively. The similarity metrics between the channels of the output feature map within the target layer of the compressed model are calculated, and the similarity metrics between the channels of the output feature map within the target layer of the original model are calculated.
[0035] Curvelet transform can analyze the texture features of an image in multiple scales and directions. In this embodiment, curvelet transform can be embedded into the neural network as a pooling layer for feature extraction. That is, before model fusion, the output feature map of the current network layer is first subjected to curvelet transform, and the low-frequency coefficients of the curvelet transform, that is, the coefficients of the coarse-scale layer, are used as the data for calculating the similarity between channels, while retaining the main information of the feature map while compressing the data volume. Assume that the weight of the target layer (the th layer) of the original network model is a tensor of M×C×h×h, where M is the number of output channels, C is the number of input channels, and h is the convolution kernel size; the weight of the th layer of the corresponding compressed network model is a tensor of N×C×h×h, where N is the number of compressed output channels; the original network model The output feature map of the layer is , and after compression, the network model The output feature map of the layer is , where X represents the input data of the original model of the layer, represents the convolution calculation. The similarity between the output feature maps is used as the basis for measuring the similarity between the weights, and the output feature maps and are subjected to curvelet transform to reduce the amount of data for metric learning.
[0036] Specifically, based on the feature maps output by the original network model and the compressed model at the target layer respectively, the channel - to - channel similarity metric is analyzed according to the target similarity index to determine the channel matching result, including: the output of the target layer of the original network model is the first feature map, and the output of the target layer of the compressed model is the second feature map; calculating the structural similarity between the channels of the first feature map and the channels of the second feature map, calculating the structural similarity between different channels in the first feature map, and the structural similarity between different channels in the second feature map, and generating a structural similarity matrix; calculating the normalized mutual information of the channel feature maps of the first feature map and the channel feature maps of the second feature map on the curvelet transform coarse - scale coefficients, calculating the normalized mutual information of different channel feature maps of the first feature map on the curvelet transform coarse - scale coefficients, and the normalized mutual information of different channel feature maps of the second feature map on the curvelet transform coarse - scale coefficients, and generating a normalized mutual information matrix; generating a mean matrix based on the structural similarity matrix and the normalized mutual information matrix, and determining the channel matching result according to the mean matrix. It can be understood that the feature map output by the target layer is a multi - channel tensor, and the channel feature map is a two - dimensional feature map output by a single channel.
[0037] That is to say, calculate the target similarity index between each channel of and respectively. That is, for each channel bi in and each channel ai in , calculate the structural similarity and the normalized mutual information of the curvelet transform coarse - scale coefficients between them. Inside the target layer of the compressed model, calculate the similarity index between each channel of the output feature map , that is, for any two channels bi and bj in , calculate the similarity between them, and thus the redundancy or correlation between the channels inside the compressed feature map can be evaluated. Inside the target layer of the original model, calculate the similarity index between each channel of the output feature map , that is, for each channel in For any two channels \(a_i\) and \(a_j\) among them, calculate the similarity between them, from which the redundancy or correlation between channels within the original feature map can be evaluated. Based on the structural similarity of each channel, a structural similarity matrix is obtained. ; Based on the normalized mutual information between the curvelet transform coarse-scale coefficients of each channel's feature map, a normalized mutual information matrix is formed. , where \(N\) is the number of output channels of the target layer of the compressed model, and \(M\) is the number of output channels of the target layer of the original network model. Finally, the mean matrix \((SSIM + NMI) / 2\) is obtained. The mean matrix can comprehensively characterize the similarity of different channels in terms of structure and coarse-scale coefficients, and the channel matching results are selected according to the mean matrix.
[0038] By using the normalized mutual information \(NMI\) (Normal mutual information) and the structural similarity \(SSIM\) (Structural Similarity) as the evaluation indicators for the similarity between channels of the feature map, the value ranges of both can be . Among them, the normalized mutual information is equal to the ratio of the average mutual information between two channels to the arithmetic mean of their respective entropies. Considering that different results may occur in the channel similarity matching results obtained based on the two evaluation indicators, in this application, the mean of the two evaluation indicators is used as the comprehensive evaluation indicator, and the maximum value among the comprehensive evaluation indicators of similarity is selected as the similarity matching result within the model and between cross-models, and then a fusion matrix is generated.
[0039] Further, determining the channel matching results according to the mean matrix includes: screening out the maximum value in each row of the mean matrix, and obtaining the channel matching results according to the row and column indices corresponding to the maximum value. That is, from the mean matrix, the maximum similarity is selected row by row to determine the channel matching results \((u, v)\).
[0040] Step S13: Generate a fusion matrix according to the channel matching results, and determine the fused weights of the target layer based on the weights of the target layer of the original network model and the weights of the target layer of the compressed model.
[0041] All channel matching results are represented by generating a fusion matrix, and the fused weights of the target layer are determined according to the fusion matrix, the weights of the target layer of the original network model, and the weights of the target layer of the compressed model.
[0042] In some embodiments, before generating the fusion matrix according to the channel matching result, it may further include: obtaining a network lightweight constraint scenario; the network lightweight constraint scenario includes a strong compression rate constraint and a weak compression rate constraint; the strong compression rate constraint means that the channel compression rate of the model reaches at least N / M, where M is the number of output channels of the target layer of the original network model, and N is the ratio of the number of output channels of the target layer of the compressed model; the weak compression rate constraint means that the channel compression rate of the model reaches at least (N + M) / 2M.
[0043] It can be understood that, for example Figure 2 As shown, in order to meet the model capabilities under different compression requirements, two different processes are provided for subsequent matrix fusion and channel fusion. That is, to improve the flexibility of compression, different model fusion processes are adopted for the two scenarios of weak compression rate constraint and strong compression rate constraint to achieve a better balance between compression rate and accuracy loss. Specifically, the fusion process adopted under the weak compression rate constraint considers the channel similarity between models, the channel similarity within the compressed model, and the channel similarity within the original model when generating the fusion matrix, and updates the channel pruning mask vector according to the channel matching result after network layer fusion. That is, when the channel compression rate permits, more channels participate in the fusion. The fusion process adopted under the strong compression rate constraint omits the calculation of the channel similarity within the original model to ensure that the channel compression rate of the fused weights is equal to the target value.
[0044] Step S14: Obtain the fused model based on the fused weights corresponding to different network layers, and use the fused model as the final compressed model of the original network model.
[0045] Finally, obtain the fused model based on the fused weights corresponding to different network layers respectively, and use this fused model as the final compressed model of the original network model. That is to say, the target layer refers to a certain layer in the network model. By sequentially taking each layer in the network model as the target layer and performing the operations of steps S11 - S13 on each layer, the fused weights corresponding to each layer are obtained. Deploy the above final compressed model on a server or edge device for image recognition.
[0046] In some embodiments, determining the fused weight of the target layer based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model, and obtaining the fused model based on the fused weights corresponding to different network layers includes: determining the current fused weight of the target layer based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model; based on the current fused weight and the historical fused weight generated by historical random channel pruning of the target layer of the original network model, obtaining the target fused weight through moving average processing, so as to obtain the fused model based on the target fused weights corresponding to different network layers. That is, for exampleFigure 2 As shown, for each layer of the network, random pruning, reconstruction, and fusion are performed multiple times, that is, the operations of steps S11 - S13. After obtaining the fused weights each time, moving average processing is performed based on the historical fused weights of this layer until the random pruning, reconstruction, and fusion of this layer reach the preset number of times, and the fused weights obtained by the latest moving average processing are used as the target fused weights. That is, the fused weights:
[0047] ; where r represents the cumulative processing times of the current layer, represents the fused weights obtained by the previous moving average processing.
[0048] In the related art, when pruning the parameter matrix after reconstruction and compression during inference, scaling processing based on the kernel set theory is used to scale the parameter matrix of the original network layer, and the retained channels in the original network model are added to the compressed network model in the channel order; then the compressed parameter matrix is solved by optimizing the feature map reconstruction error. However, this method only scales the retained channels after pruning and does not consider retaining the knowledge contained in the pruned channels, which will cause loss of some image recognition knowledge. In this application, when reconstructing the compressed weight matrix, the channel similarity matching between models is considered, and the paired channels are superimposed and merged, that is, the fusion of the original network model and the compressed network model. Through such fusion, the knowledge learned by the original network model on a large - scale dataset is retained. In the channel space composed of each channel of the feature map of the compressed network layer and each channel of the feature map of the original network layer, similarity metric learning and channel matching are performed, and the paired channels are superimposed and fused to avoid knowledge loss, which is beneficial to the deployment and application of image recognition tasks based on deep learning on edge devices with limited computing and storage resources. In addition, in this application, the similarity of each channel of the activation feature map is also based on the comprehensive similarity metric of curvelet transform and structure.
[0049] The deep model fusion technology can merge the parameters or predictions of multiple network models that solve different tasks and have different structures into one model. In the related art, the graph - based fusion method (ZipIt) is often used, but this method is only applicable to homogeneous network models, that is, each model to be fused has an equal number of channels in the corresponding layer. In this application, through model fusion based on channel similarity, it is not restricted by the model structure and can be applied to heterogeneous model fusion. By model fusion, the accuracy loss caused by network compression is reduced, and it is not necessary to rely on model retraining to restore the accuracy of the network model, optimizing the construction efficiency of the image recognition network compression process itself.
[0050] As can be seen from the above, in this embodiment, pruning is performed layer by layer during the inference of the original network model. Random input channel pruning is sequentially executed, and then the weights of the network layers of the compressed model are solved with the goal of optimizing the reconstruction error of the output feature map. Subsequently, the weights of the original network layer and the weights of the compressed network layer are fused through the model fusion process. Thus, by analyzing the channel similarity metric between models based on the curvelet transform feature similarity and structural similarity, and merging the matching channels, the loss of image recognition knowledge can be avoided. A fusion matrix is generated based on the matching channel pairs, and model fusion is performed through channel similarity metric, which does not depend on a specific model architecture. Therefore, it can be applied to the fusion of heterogeneous models, improving the ability and effect of network model lightweighting.
[0051] An embodiment of the present application discloses a specific method for lightweighting a network model under weak compression rate constraints. The method may include the following steps:
[0052] Step S21: Perform the r-th random channel pruning on the output channels of the layer (i.e., the above-mentioned target layer) of the original network model to determine the weights of the layer of the compressed model corresponding to the original network model.
[0053] Step S22: Based on the feature maps output by the original network model and the compressed model respectively at the layer, analyze the channel similarity metric according to the target similarity index to determine the channel matching result; the target similarity index includes structure and curvelet transform features, and the channel matching result represents a set of matching channel pairs.
[0054] Step S23: Remove duplicate channel matching results from all channel matching results, and generate a fusion matrix based on the de-duplicated channel matching results.
[0055] For example Figure 3 as shown, after calculating the channel similarity, the channel matching result (u, v) is determined by selecting the maximum similarity value row by row from the generated mean matrix (SSIM + NMI) / 2. A fusion matrix T is generated based on the de-duplicated channel matching results. After removing duplicate matching results, a similarity vector S = {si} with a dimension of (N + M) / 2 is obtained, that is, the similarity vector S is used to store the maximum similarity values selected from each row of the mean matrix (SSIM + NMI) / 2. Where , . M is the number of output channels of the layer of the original network model, and N is the number of output channels of the layer of the compressed model. The values of the i-th channel matching result corresponding to the u-th and v-th columns in the i-th row of the fusion matrix T are 1 / 2, and the values of the remaining columns in this row are 0.
[0056] That is to say, the channel matching result includes a first index and a second index. u and v are the indices of two channels, which are used to identify the similar or matching channel pairs found during the channel matching process. The index is an integer from 0 to N+M-1, where N and M represent the sizes of two channel sets. According to the channel matching result, a fusion matrix T is generated. In the fusion matrix T, the elements at the u-th and v-th columns of the i-th row take the value of 1 / 2, indicating that these two channels are matched in the i-th matching result. The elements of the remaining columns in this row take the value of 0, indicating that these channels are not matched in the i-th matching result.
[0057] Step S24: Determine a new pruning mask vector according to the channel matching result, and perform channel pruning on the input channels of the next layer of the original network model layer.
[0058] Since the number of output channels of the compressed network layer with a weak compression rate is relaxed to (M+N) / 2, while the number of input channels of the next layer is M, it is necessary to re-determine the pruning mask vector according to the matching result in the case of a weak compression rate to ensure that the number of output channels of the compressed network layer is equal to the number of input channels of the next layer.
[0059] In some embodiments, determining a new pruning mask vector according to the channel matching result includes: creating a zero vector for the target layer of the original network model; the dimension of the zero vector is half of the sum of N and M; the channel matching result includes a first index and a second index; if the first index is less than N and the second index is greater than or equal to N, then write the difference between the second index and N into the zero vector; if both the first index and the second index are greater than or equal to N, then write the difference between the first index and N, or the difference between the second index and N into the zero vector. If both the first index and the second index are less than N, then write the non-zero value corresponding to the first index in the original mask vector, or the non-zero value corresponding to the second index in the original mask vector into the zero vector to obtain a new pruning mask vector; both the first index and the second index are less than N, indicating that they both belong to the first set and since they are both valid indices, and the corresponding values may both be non-zero, so it is necessary to randomly select one of these two values. Random selection is to break symmetry or avoid preference for a specific index. The original mask vector is the random mask vector used for random channel pruning of the output channels of the original network model layer.
[0060] That is to say, first, generate a zero vector with a dimension of (N+M) / 2 , which will be used to store the values calculated according to the channel matching result. Update the vector according to each channel matching result (u, v) , u and v in the matching results are the indices of channels, and they may belong to two different sets, namely the first set of N channels and the second set of M channels. There are three cases according to the values of u and v: 1. u < N and N ≤ v, that is, u belongs to the first set and v belongs to the second set; at this time, is stored in the vector , and the index of v is adjusted to the relative index of the second set through . 2. u, v ≥ N, the indices are both greater than or equal to N, that is, both u and v belong to the second set; randomly select one value from and and store it in the vector . 3. u, v < N, both u and v belong to the first set (the indices are both less than N), randomly select one of the u-th value or the v-th value of the non-zero elements in the mask vector gr and store it in the vector . Traverse all channel matching results, and update the vector according to the above rules, as the new mask vector.
[0061] Step S25: Divide the fusion matrix into a first fusion matrix corresponding to the original network model and a second fusion matrix corresponding to the compressed model; the number of columns of the first fusion matrix is M, and the number of columns of the second fusion matrix is N; according to the product of the first fusion matrix and the weights of the layer of the original network model, the first product is obtained; according to the product of the second fusion matrix and the weights of the layer of the compressed model, the second product is obtained; according to the sum of the first product and the second product, the layer of the fused weights is obtained.
[0062] Divide the fusion matrix T into the second fusion matrix of the compressed model and the first fusion matrix of the original model respectively according to the number of output channels of the layer of the compressed model and the original model, that is, divide the fusion matrix T into columns as and , , represents concatenation, , . The fused weights , are the weights of the layer of the compressed model, are the weights of the layer of the original model.
[0063] It can be understood that the goal of the weak compression constraint is to maximize the retention of the image recognition knowledge of the original network model, rather than strictly requiring the channel compression rate to reach a predetermined target value. Therefore, a certain flexibility is allowed for the channel compression rate, which can be relaxed to a higher value. At this time, the fusion process considers not only the channel matching between models but also the channel matching within the model. By directly multiplying the weight matrices of the original model and the compressed model with the corresponding fusion matrices through matrix multiplication and then adding them to obtain the fused weight matrix, the matching information within the model and across models can be fully utilized to maximize the retention of the features of the original model.
[0064] Step S26: Obtain the fused model based on the fused weights corresponding to different network layers, and use the fused model as the finally compressed model of the original network model.
[0065] Judge whether r reaches the preset number of times R. If not, continue to perform the (r + 1)-th processing on the layer. If it reaches the preset number of times, use the fused weight as the weight of the new compressed model for the layer. Traverse each network layer of the original network model until the fused weights of each network layer are determined to obtain the final compressed network model.
[0066] Among them, for the specific processes of the above steps S21, S22, and S26, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0067] As can be seen from the above, in this embodiment, duplicate channel matching results are removed from all channel matching results, and a fusion matrix is generated based on the de-duplicated channel matching results; the matching information within the model and across models is fully utilized to maximize the retention of the features of the original model. Determine a new pruning mask vector according to the channel matching result, and perform channel pruning on the input channels of the next layer of the layer of the original network model; ensure that the number of output channels of the compressed network layer is equal to the number of input channels of the next layer.
[0068] The embodiments of the present application disclose a specific method for lightweighting a network model under strong compression rate constraints, which may include the following steps:
[0069] Step S31: Perform random channel pruning on the output channels of the layer (i.e., the above target layer) of the original network model to determine the weights of the layer of the compressed model corresponding to the original network model.
[0070] Step S32: Based on the original network model and the compressed model respectively in The feature maps output by the layer analyze the inter-channel similarity metric according to the target similarity metric to determine the channel matching result; the target similarity metric includes structural and wavelet transform features, and the channel matching result represents a set of matching channel pairs.
[0071] Step S33: Classify the channel matching results. The channel matching results are divided into cross-model matching results, original model internal matching results, and compressed model internal matching results; based on the cross-model matching results and the compressed model internal matching results, a fusion matrix is generated.
[0072] For example Figure 4 As shown, to ensure that the number of channels in the fused network layer is equal to the target compressed channel number N, only the channel matching results between cross-models are used for fusion. For the channels involved in the matching results within the compressed model, they are directly copied to the fused model without fusion processing, and the matching results within the original model are directly ignored.
[0073] According to the mean matrix of the two similarity metrics (SSIM + NMI) / 2, the maximum similarity is selected row by row to determine the channel matching result (u, v), obtaining a similarity vector S={si} of dimension N, and generating a fusion matrix T. Among them , . If the v-th channel matching result is a cross-model match, that is, N≤u, the values in the u-th and v-th columns of the v-th row in the corresponding fusion matrix T are 1 / 2, and the remaining columns in this row are 0. If the v-th channel matching result is a match within the compressed model, that is, u<N, the values in the u-th column of the u-th row and the v-th column of the v-th row in the corresponding fusion matrix T are 1, and the remaining columns in this row are 0. That is, in the case of a strong compression rate, relevant channels are selected from the original model for fusion with the compressed model, and irrelevant channels are directly ignored to avoid interference.
[0074] The following is an example. The following data is only for example and not actual data.
[0075] Let N = 3 (number of channels in the compressed model), M = 4 (number of channels in the original model), and the mean matrix (SSIM + NMI) / 2 be:
[0076] ;
[0077] Determine the channel matching results. The maximum value in the first row is 0.82, and the channel matching result is (u = 0, v = 0); the maximum value in the second row is 0.77, and the channel matching result is (u = 1, v = 1); the maximum value in the third row is 0.79, and the channel matching result is (u = 4, v = 2). For the first row, u = 0 < N, which is an internal match, T[0,0] = T[0,0] = 1, and the values of the remaining columns are 0. For the second row, u = 1 < N, T[1,1] = T[1,1] = 1, and the values of the remaining columns are 0. For the third row, u = 4 > N, T[2,2] = T[2,4] = 0.5, and the values of the remaining columns are 0. The final fusion matrix T is as follows:
[0078] .
[0079] Step S34: Divide the fusion matrix into a first fusion matrix corresponding to the original network model and a second fusion matrix corresponding to the compressed model; the number of columns of the first fusion matrix is M, and the number of columns of the second fusion matrix is N; based on the first fusion matrix, the second fusion matrix, the weights of the layers of the original network model, the weights of the layers of the compressed model, the feature maps output by the layers of the original network model, and the feature maps output by the layers of the compressed model, construct an optimization function representing the feature map reconstruction error; obtain the fusion coefficients by solving the optimization function; based on the fusion coefficients, the first fusion matrix, the weights of the layers of the original network model, the weights of the layers of the compressed model, determine the weights after fusion of the
[0080] First, divide the fusion matrix T into the second fusion matrix of the compressed model and the first fusion matrix of the original model according to the number of output channels of the layers of the compressed model and the original model respectively, that is, divide the fusion matrix by columns into and , , , where
[0081] is a diagonal matrix. Furthermore, solve the fusion coefficients
[0082] by minimizing the feature map reconstruction error, where .
[0083] Fused weights , the weights of the th layer of the compressed model and the weights of the th layer of the original model are converted into fused weights .
[0084] That is, the goal of the strong compression constraint is to strictly ensure that the number of channels in the fused network layer is equal to the target compressed number of channels N. In this case, the fusion process only considers the channel matching results between models. For the matching results within the compressed model, they are directly copied to the fused model without fusion processing; for the matching results within the original model, they are directly ignored. Since the strong constraint conditions limit the structure of the fusion matrix, directly using matrix multiplication may not meet the strict channel number requirements. Therefore, it is necessary to solve the fusion coefficients through an optimization method to minimize the feature map reconstruction error. By optimizing the fusion coefficients, it is ensured that the fused network retains the features of the original model as much as possible while meeting the channel number requirements.
[0085] Step S35: Obtain a fused model based on the fused weights corresponding to different network layers, and use the fused model as the finally compressed model of the original network model.
[0086] Among them, for the specific processes of the above steps S31, S32, and S35, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0087] As can be seen from the above, in this embodiment, a fusion matrix is generated based on the cross-model matching results and the matching results within the compressed model; the compression rate of the fused model meets the standard. By constructing an optimization function representing the feature map reconstruction error; by solving the optimization function, the fusion coefficients are obtained; based on the fusion coefficients, the first fusion matrix, the weights of the th layer of the original network model, the weights of the th layer of the compressed model, the fused weights of the th layer are determined; by optimizing the fusion coefficients, it is ensured that the fused network retains the features of the original model as much as possible while meeting the channel number requirements.
[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform.
[0089] An embodiment of the present application also provides a network model lightweight device. As shown in Figure 5 , the device includes:
[0090] A random channel pruning module 11, which is used to perform random channel pruning on the output channels of the target layer of the original network model, and determine the weights of the target layer of the compressed model corresponding to the original network model;
[0091] A channel matching module 12, which is used to analyze the channel similarity metric based on the feature maps output by the original network model and the compressed model at the target layer according to the target similarity metric to determine the channel matching result; the target similarity metric includes structural and wavelet transform features, and the channel matching result represents a set of matching channel pairs;
[0092] A fused weight determination module 13, which is used to generate a fusion matrix according to the channel matching result, and determine the fused weight of the target layer based on the weights of the target layer of the original network model and the weights of the target layer of the compressed model;
[0093] A final compressed model determination module 14, which is used to obtain a fused model based on the fused weights corresponding to different network layers, and use the fused model as the final compressed model of the original network model.
[0094] As can be seen from the above, in this embodiment, by analyzing the channel similarity metric between models based on the wavelet transform feature similarity and structural similarity, and merging the matching channels, the loss of image recognition knowledge can be avoided; generating a fusion matrix based on the matching channel pairs, and performing model fusion through channel similarity metric, which does not depend on a specific model architecture, so it can be applied to the fusion of heterogeneous models, improving the ability and effect of network model lightweighting.
[0095] In some specific embodiments, the output of the target layer of the original network model is a first feature map, and the output of the target layer of the compressed model is a second feature map; the channel matching module 12 may specifically include:
[0096] A structural similarity calculation unit, which is used to calculate the structural similarity between the channels of the first feature map and the channels of the second feature map, calculate the structural similarity between different channels in the first feature map, and the structural similarity between different channels in the second feature map, and generate a structural similarity matrix;
[0097] A coarse-scale coefficient similarity calculation unit, which is used to calculate the normalized mutual information of the channel feature maps of the first feature map and the channel feature maps of the second feature map on the wavelet transform coarse-scale coefficients, calculate the normalized mutual information of different channel feature maps of the first feature map on the wavelet transform coarse-scale coefficients, and the normalized mutual information of different channel feature maps of the second feature map on the wavelet transform coarse-scale coefficients, and generate a normalized mutual information matrix;
[0098] A mean matrix generation unit, configured to generate a mean matrix based on the structural similarity matrix and the normalized mutual information matrix, and determine a channel matching result according to the mean matrix.
[0099] In some specific embodiments, the mean matrix generation unit may specifically include:
[0100] A channel matching result generation unit, configured to filter out the maximum value in each row of the mean matrix, and obtain the channel matching result according to the row and column indices corresponding to the maximum value.
[0101] In some specific embodiments, the random channel pruning module 11 may specifically include:
[0102] A random mask vector generation unit, configured to generate a random mask vector for the output channels of the target layer of the original network model, and perform channel pruning on the output channels of the target layer according to the random mask vector;
[0103] A weight determination unit, configured to solve the minimization error by using the least squares method according to the feature map output by the target layer of the original network model and the feature map output by the target layer of the compressed model, so as to determine the weight of the target layer of the compressed model.
[0104] In some specific embodiments, the network model lightweight device may specifically include:
[0105] A constraint scenario acquisition unit, configured to acquire a network lightweight constraint scenario before generating a fusion matrix according to the channel matching result; the network lightweight constraint scenario includes a strong compression rate constraint and a weak compression rate constraint; the strong compression rate constraint indicates that the channel compression rate of the model reaches at least N / M, where M is the number of output channels of the target layer of the original network model, and N is the ratio of the number of output channels of the target layer of the compressed model; the weak compression rate constraint indicates that the channel compression rate of the model reaches at least (N + M) / 2M;
[0106] In some specific embodiments, if the network lightweight constraint scenario is the weak compression rate constraint, the fused weight determination module 13 may specifically include:
[0107] A first fusion matrix generation unit, configured to remove duplicate channel matching results from all channel matching results, and generate a fusion matrix based on the deduplicated channel matching results.
[0108] In some specific embodiments, if the network lightweight constraint scenario is the weak compression rate constraint, the fused weight determination module 13 may specifically include:
[0109] A matrix division unit for dividing the fusion matrix into a first fusion matrix corresponding to the original network model and a second fusion matrix corresponding to the compressed model; the number of columns of the first fusion matrix is M, and the number of columns of the second fusion matrix is N;
[0110] A first product determination unit for obtaining a first product according to the product of the first fusion matrix and the weights of the target layer of the original network model;
[0111] A second product determination unit for obtaining a second product according to the product of the second fusion matrix and the weights of the target layer of the compressed model;
[0112] A fused weight determination unit for obtaining the fused weight of the target layer according to the sum of the first product and the second product.
[0113] In some specific embodiments, if the network lightweight constraint scenario is the weak compression rate constraint, the network model lightweight device may specifically include:
[0114] A new pruning mask vector determination unit for determining a new pruning mask vector according to the channel matching result after generating the fusion matrix, and performing channel pruning on the input channels of the next layer of the target layer of the original network model according to the new pruning mask vector.
[0115] In some specific embodiments, the new pruning mask vector determination unit may specifically include:
[0116] A zero vector creation unit for creating a zero vector for the target layer of the original network model; the dimension of the zero vector is half of the sum of N and M; the channel matching result includes a first index and a second index;
[0117] A first writing unit for writing the difference between the second index and N into the zero vector if the first index is less than N and the second index is greater than or equal to N;
[0118] A second writing unit for writing the difference between the first index and N, or the difference between the second index and N into the zero vector if both the first index and the second index are greater than or equal to N;
[0119] A third writing unit for writing the non-zero value corresponding to the first index in the original mask vector, or the non-zero value corresponding to the second index in the original mask vector into the zero vector if both the first index and the second index are less than N, so as to obtain a new pruning mask vector; the original mask vector is a random mask vector used for randomly pruning the output channels of the target layer of the original network model.
[0120] In some specific embodiments, if the network lightweight constraint scenario is the strong compression rate constraint, the fused weight determination module 13 may specifically include:
[0121] A classification unit, configured to classify the channel matching result, and the channel matching result is divided into a cross-model matching result, an in-primitive-model matching result, and an in-compressed-model matching result;
[0122] A second fusion matrix generation unit, configured to generate a fusion matrix based on the cross-model matching result and the in-compressed-model matching result.
[0123] In some specific embodiments, if the network lightweight constraint scenario is the strong compression rate constraint, the fused weight determination module 13 may specifically include:
[0124] A matrix partitioning unit, configured to partition the fusion matrix into a first fusion matrix corresponding to the original network model and a second fusion matrix corresponding to the compressed model; the number of columns of the first fusion matrix is M, and the number of columns of the second fusion matrix is N;
[0125] An optimization function construction unit, configured to construct an optimization function characterizing the feature map reconstruction error based on the first fusion matrix, the second fusion matrix, the weights of the target layer of the original network model, the weights of the target layer of the compressed model, the feature map output by the target layer of the original network model, and the feature map output by the target layer of the compressed model;
[0126] A fusion coefficient solving unit, configured to obtain a fusion coefficient by solving the optimization function;
[0127] A fused weight determination unit, configured to determine the fused weight of the target layer based on the fusion coefficient, the first fusion matrix, the weights of the target layer of the original network model, and the weights of the target layer of the compressed model.
[0128] In some specific embodiments, the fused weight determination module 13 may specifically include:
[0129] A current fused weight determination unit, configured to determine the current fused weight of the target layer based on the weights of the target layer of the original network model and the weights of the target layer of the compressed model;
[0130] A moving average processing unit, configured to obtain a target fused weight through moving average processing based on the current fused weight and the historical fused weight generated by historical random channel pruning of the target layer of the original network model, so as to obtain a fused model based on the target fused weights corresponding to different network layers.
[0131] For the descriptions of the features in the corresponding embodiments of the network model lightweight device, reference can be made to the relevant descriptions in the corresponding embodiments of the network model lightweight method, which will not be elaborated here one by one.
[0132] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the network model lightweight method.
[0133] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned embodiments of the network model lightweight method when running.
[0134] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disk, etc., various media that can store computer programs.
[0135] An embodiment of the present application further provides a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the network model lightweight method.
[0136] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the network model lightweight method.
[0137] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0138] The above has introduced in detail a method for network model lightweighting provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A network model lightweight method, characterized in that: include: Performing random channel pruning on the output channels of the target layer of the original network model to determine the weight of the target layer of the compressed model corresponding to the original network model; Based on the feature maps outputted by the original network model and the compressed model at the target layer respectively, analyzing the similarity metric between channels according to the target similarity index to determine the channel matching result; the target similarity index includes structure and curvelet transform features, and the channel matching result represents a set of matched channel pairs; Generate a fusion matrix according to the channel matching result, and determine the fusion weight of the target layer based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model; A fused model is obtained based on fused weights corresponding to different network layers, and the fused model is used as a final compressed model of the original network model, so as to perform image recognition using the final compressed model; Wherein, based on the feature maps outputted by the original network model and the compressed model at the target layer respectively, the similarity measurement between channels is analyzed according to the target similarity index to determine the channel matching result, including: The output of the target layer of the original network model is a first feature map, and the output of the target layer of the compressed model is a second feature map; Calculating the structural similarity between the channels of the first feature map and the channels of the second feature map, calculating the structural similarity between different channels in the first feature map, and calculating the structural similarity between different channels in the second feature map, and generating a structural similarity matrix; Calculating the normalized mutual information of the channel feature map of the first feature map and the channel feature map of the second feature map on the coarse-scale coefficient of the curvelet transform, calculating the normalized mutual information of different channel feature maps of the first feature map on the coarse-scale coefficient of the curvelet transform, and the normalized mutual information of different channel feature maps of the second feature map on the coarse-scale coefficient of the curvelet transform, and generating a normalized mutual information matrix; A mean matrix is generated based on the structural similarity matrix and the normalized mutual information matrix, and a channel matching result is determined according to the mean matrix.
2. The network model lightweight method according to claim 1, characterized in that: Determining a channel matching result according to the mean matrix includes: The maximum value in each row of the mean matrix is screened out, and the channel matching result is obtained according to the row and column indexes corresponding to the maximum value.
3. The network model lightweight method according to claim 1, characterized in that: Performing random channel pruning on the output channels of the target layer of the original network model to determine the weight of the target layer of the compressed model corresponding to the original network model includes: Generate a random mask vector for the output channel of the target layer of the original network model, and perform channel pruning on the output channel of the target layer according to the random mask vector; According to the feature map output by the target layer of the original network model and the feature map output by the target layer of the compressed model, the least squares method is used to minimize the error to determine the weight of the target layer of the compressed model.
4. The network model lightweight method according to claim 1, characterized in that: Before generating the fusion matrix according to the channel matching result, the method further includes: Obtain a network lightweight constraint scenario; the network lightweight constraint scenario includes a strong compression rate constraint and a weak compression rate constraint; the strong compression rate constraint indicates that the channel compression rate of the model is at least N / M, where M is the number of output channels of the target layer of the original network model, and N is the ratio of the number of output channels of the target layer of the compressed model; the weak compression rate constraint indicates that the channel compression rate of the model is at least (N+M) / 2M; If the network lightweight constraint scenario is the weak compression rate constraint, a fusion matrix is generated according to the channel matching result, including: The duplicate channel matching results are removed from the channel matching results, and a fusion matrix is generated based on the duplicated channel matching results.
5. The network model lightweight method according to claim 4, characterized in that: If the network lightweight constraint scenario is the weak compression rate constraint, the fused weight of the target layer is determined based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model, including: Dividing the fusion matrix into a first fusion matrix corresponding to the original network model and a second fusion matrix corresponding to the compressed model; the number of columns of the first fusion matrix is M, and the number of columns of the second fusion matrix is N; Obtaining a first product according to the product of the first fusion matrix and the weight of the target layer of the original network model; Obtaining a second product according to the product of the second fusion matrix and the weight of the target layer of the compressed model; A fused weight of the target layer is obtained according to the sum of the first product and the second product.
6. The network model lightweight method according to claim 4, characterized in that: If the network lightweight constraint scenario is the weak compression rate constraint, after generating the fusion matrix, the following is also included: A new pruning mask vector is determined according to the channel matching result, and channel pruning is performed on the input channels of the next layer of the target layer of the original network model according to the new pruning mask vector.
7. The network model lightweight method according to claim 6, characterized in that: Determining a new pruning mask vector according to the channel matching result includes: creating a zero vector for the target layer of the original network model; the dimension of the zero vector is half of the sum of N and M; The channel matching result includes a first index and a second index; If the first index is less than N, and the second index is greater than or equal to N, then the difference between the second index and N is written into the zero vector; If both the first index and the second index are greater than or equal to N, the difference between the first index and N, or the difference between the second index and N, is written into the zero vector; If both the first index and the second index are less than N, the non-zero value corresponding to the first index in the original mask vector, or the non-zero value corresponding to the second index in the original mask vector, is written into the zero vector to obtain a new pruning mask vector; the original mask vector is a random mask vector used for random channel pruning of the output channels of the target layer of the original network model.
8. The network model lightweight method according to claim 4, characterized in that: If the network lightweight constraint scenario is the strong compression rate constraint, a fusion matrix is generated according to the channel matching result, including: Classifying the channel matching results, wherein the channel matching results are divided into cross-model matching results, original model internal matching results, and compressed model internal matching results; A fusion matrix is generated based on the cross-model matching results and the compressed model internal matching results.
9. The network model lightweight method according to claim 4, characterized in that: If the network lightweight constraint scenario is the strong compression rate constraint, the fused weight of the target layer is determined based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model, including: Dividing the fusion matrix into a first fusion matrix corresponding to the original network model and a second fusion matrix corresponding to the compressed model; the number of columns of the first fusion matrix is M, and the number of columns of the second fusion matrix is N; Based on the first fusion matrix, the second fusion matrix, the weight of the target layer of the original network model, the weight of the target layer of the compressed model, the feature map output by the target layer of the original network model, and the feature map output by the target layer of the compressed model, an optimization function representing the reconstruction error of the feature map is constructed; Obtaining a fusion coefficient by solving the optimization function; Based on the fusion coefficient, the first fusion matrix, the weight of the target layer of the original network model, and the weight of the target layer of the compressed model, the fused weight of the target layer is determined.
10. The network model lightweight method according to any one of claims 1 to 9, characterized in that: Determining the fused weight of the target layer based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model, and obtaining the fused model based on the fused weights corresponding to different network layers, including: Determine the current fused weight of the target layer based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model; Based on the current fused weights and the historical fused weights generated by historical random channel pruning of the target layer of the original network model, the target fused weights are obtained by moving average processing, so as to obtain a fused model based on the target fused weights corresponding to different network layers.
11. A network model lightweight device, characterized in that: include: A random channel pruning module, used to perform random channel pruning on the output channel of the target layer of the original network model to determine the weight of the target layer of the compressed model corresponding to the original network model; A channel matching module, configured to analyze the inter-channel similarity metric according to a target similarity index based on the feature maps outputted at the target layer by the original network model and the compressed model respectively, to determine a channel matching result; the target similarity index includes structure and curvelet transform features, and the channel matching result represents a set of matched channel pairs; A post-fusion weight determination module, used to generate a fusion matrix according to the channel matching result, and determine the post-fusion weight of the target layer based on the weight of the target layer of the original network model and the weight of the target layer of the compressed model; A final compressed model determination module, used to obtain a fused model based on the fused weights corresponding to different network layers, and use the fused model as the final compressed model of the original network model to perform image recognition using the final compressed model; The output of the target layer of the original network model is the first feature map, and the output of the target layer of the compressed model is the second feature map; The channel matching module is used to calculate the structural similarity between the channels of the first feature map and the channels of the second feature map, calculate the structural similarity between different channels in the first feature map, and calculate the structural similarity between different channels in the second feature map to generate a structural similarity matrix; calculate the normalized mutual information between the channel feature map of the first feature map and the channel feature map of the second feature map on the coarse-scale coefficients of the curvelet transform, calculate the normalized mutual information of different channel feature maps of the first feature map on the coarse-scale coefficients of the curvelet transform, and calculate the normalized mutual information of different channel feature maps of the second feature map on the coarse-scale coefficients of the curvelet transform to generate a normalized mutual information matrix; generate a mean matrix based on the structural similarity matrix and the normalized mutual information matrix, and determine the channel matching result according to the mean matrix.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the network model lightweight method as described in any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the network model lightweight method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the network model lightweight method as described in any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Joint neural network model compression method based on channel pruning and quantitative training
CN111652366A
Flexible deep learning network model compression method based on channel gradient pruning
CN112396179A