Model optimization method and device, image processing method
Patent Information
- Application Number
- CN202211096347.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-08
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-09-08
AI Technical Summary
[0003]相关技术中的模型的剪枝压缩的优化精度不够,导致得到的模型的精度较差
[0011] One embodiment of this disclosure provides a model optimization method that obtains reference feature maps output by each convolutional layer in an image processing model; determines the first weight of each channel of the convolutional layer based on the reference feature maps; determines an intermediate feature map based on the first weights and the reference feature maps, and determines the second weight of each convolutional layer based on the intermediate feature maps; and optimizes the image processing model using the first and second weights to obtain a target image processing model. Compared with existing technologies, this method determines the first weight of each channel in the convolutional layers of the image processing model, considering the importance of each channel in each convolutional layer. It also determines the second weight of each convolutional layer in the entire model, taking into account the importance of the convolutional layers in the model. This allows for more accurate optimization of the image processing model, improving the accuracy of the obtained target image processing model. Furthermore, the obtained target image processing model has a smaller number of parameters, reducing the consumption of computer resources and facilitating its deployment on mobile terminals.
Smart Images

Figure CN117726890B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of model optimization technology, and more specifically, to a model optimization method and apparatus, an image processing method, a computer-readable storage medium, and an electronic device. Background Technology
[0002] With the development of deep learning network models, the technology of deploying deep learning network models on mobile terminals is becoming more and more widely used. However, deep learning network models have a large number of parameters, and pruning and compression of the models are required when deploying them on mobile terminals.
[0003] The optimization accuracy of pruning and compression in related technologies is insufficient, resulting in poor accuracy of the obtained models.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this disclosure is to provide a model optimization method, a model optimization apparatus, an image processing method, a computer-readable medium, and an electronic device, thereby improving the accuracy of model optimization to at least a certain extent.
[0006] According to a first aspect of this disclosure, a model optimization method is provided, comprising: obtaining reference feature maps output by each convolutional layer in an image processing model; determining first weights for each channel of the convolutional layer based on the reference feature maps; determining intermediate feature maps based on the first weights and the reference feature maps, and determining second weights for each of the convolutional layers based on the intermediate feature maps; and optimizing the image processing model using the first weights and the second weights to obtain a target image processing model.
[0007] According to a second aspect of this disclosure, an image processing method is provided, comprising: acquiring an image to be processed and a pre-trained image processing model; optimizing the image processing model using the optimization method of the image processing model to obtain a target image processing model; and inputting the image to be processed into the target image processing model to process the image to be processed.
[0008] According to a third aspect of this disclosure, a model optimization apparatus is provided, comprising: an acquisition module for acquiring reference feature maps output by each convolutional layer in an image processing model; a first determination module for determining first weights for each channel of the convolutional layer based on the reference feature maps; a second determination module for determining intermediate feature maps based on the first weights and the reference feature maps, and determining second weights for each convolutional layer based on the intermediate feature maps; and an optimization module for optimizing the image processing model using the first weights and the second weights to obtain a target image processing model.
[0009] According to a fourth aspect of this disclosure, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described above.
[0010] According to a fifth aspect of this disclosure, an electronic device is provided, characterized in that it includes: one or more processors; and a memory for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method described above.
[0011] One embodiment of this disclosure provides a model optimization method that obtains reference feature maps output by each convolutional layer in an image processing model; determines the first weight of each channel of the convolutional layer based on the reference feature maps; determines an intermediate feature map based on the first weights and the reference feature maps, and determines the second weight of each convolutional layer based on the intermediate feature maps; and optimizes the image processing model using the first and second weights to obtain a target image processing model. Compared with existing technologies, this method determines the first weight of each channel in the convolutional layers of the image processing model, considering the importance of each channel in each convolutional layer. It also determines the second weight of each convolutional layer in the entire model, taking into account the importance of the convolutional layers in the model. This allows for more accurate optimization of the image processing model, improving the accuracy of the obtained target image processing model. Furthermore, the obtained target image processing model has a smaller number of parameters, reducing the consumption of computer resources and facilitating its deployment on mobile terminals.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0014] Figure 1 A schematic diagram of an exemplary system architecture to which embodiments of the present disclosure may be applied is shown;
[0015] Figure 2 This schematically illustrates a flowchart of a model optimization method according to an exemplary embodiment of the present disclosure;
[0016] Figure 3 This schematically illustrates a flowchart of determining a first weight in an exemplary embodiment of the present disclosure;
[0017] Figure 4 The diagram schematically illustrates a framework diagram for obtaining an intermediate feature map in an exemplary embodiment of the present disclosure;
[0018] Figure 5 This schematically illustrates a flowchart of determining a second weight in an exemplary embodiment of the present disclosure;
[0019] Figure 6 This schematically illustrates a flowchart of obtaining a target optimization model in an exemplary embodiment of the present disclosure;
[0020] Figure 7 The illustration schematically shows a comparison diagram of pruning before and after in an exemplary embodiment of the present disclosure;
[0021] Figure 8 This schematically illustrates a flowchart of a method for judging a target image processing model in an exemplary embodiment of the present disclosure;
[0022] Figure 9 A flowchart illustrating another model optimization method in an exemplary embodiment of this disclosure is shown schematically.
[0023] Figure 10 This schematically illustrates a flowchart of another target image processing model making a judgment in an exemplary embodiment of the present disclosure;
[0024] Figure 11 A flowchart illustrating yet another model optimization method in an exemplary embodiment of the present disclosure is shown schematically.
[0025] Figure 12 This schematic diagram illustrates the composition of a model optimization apparatus in an exemplary embodiment of the present disclosure.
[0026] Figure 13 A schematic diagram of an electronic device to which embodiments of the present disclosure may be applied is shown. Detailed Implementation
[0027] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0028] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0029] In related technologies, deep learning network models have a large number of redundant parameters from convolutional layers to fully connected layers, and the activation values of many neurons are close to 0. Removing these neurons can still produce the same model expressive power. This situation is called overparameterization, and the corresponding technique is called model pruning.
[0030] Currently, pruning algorithms are mainly divided into fine-grained weight connection pruning and coarse-grained channel / filter pruning. Coarse-grained pruning has a wider range of applications, as it can produce simplified small models that do not require specialized algorithm support.
[0031] The channel pruning algorithms in related technologies mainly include the following three types:
[0032] The first approach is based on importance factors, which evaluates the effectiveness of a channel and then constrains some channels to make the model structure sparsity, thereby enabling pruning. However, the importance factor-based approach is too subjective, measuring only the importance of a channel locally. It is currently inadequate and cannot comprehensively consider the impact of channels on the network. It's possible that a channel may be unimportant in the current layer but have a significant impact on the model's accuracy across the entire network, leading to excessive accuracy loss.
[0033] The second approach uses reconstruction error to guide pruning, indirectly measuring the impact of a channel on the output. However, channel pruning algorithms based on output reconstruction error only consider the local reconstruction error of a channel at the current layer and cannot assess the importance of the current channel to the entire network.
[0034] The third approach measures channel sensitivity based on changes in the optimization objective. By imposing a global optimization objective on the model itself, the importance of a channel is assessed through changes in the objective function when channels are deleted. While this approach considers the global picture using an optimization objective function, it doesn't take into account local characteristics.
[0035] In summary, the model optimization methods in related technologies tend to focus more on the local impact of the weights during convolution, but ignore the impact on the accuracy of the entire network. They cannot guarantee that the selected channels are globally optimal, and therefore cannot guarantee that the network after coarse-grained pruning is a globally optimal subnetwork.
[0036] Based on the above-mentioned shortcomings, this disclosure provides a new model optimization method. Figure 1 A schematic diagram of a system architecture for implementing the above-described model optimization method is shown. This system architecture 100 may include a terminal 110 and a server 120. The terminal 110 may be a smartphone, tablet, desktop computer, laptop, or other terminal device. The server 120 generally refers to the backend system providing model optimization-related services in this exemplary embodiment, and may be a single server or a cluster of multiple servers. The terminal 110 and the server 120 can be connected via a wired or wireless communication link for data interaction.
[0037] In one implementation, the model optimization method described above can be executed by terminal 110. For example, a user uses terminal 110 to obtain an image processing model, optimizes the image processing model, and outputs a target image processing model.
[0038] In one implementation, the model optimization method described above can be executed by server 120. For example, a user can use terminal 110 to send a reference image to the server. The server inputs the reference image into the image processing model to obtain reference feature maps output by each convolutional layer, optimizes only the image processing model, and returns the target image processing model to terminal 110.
[0039] As can be seen from the above, the execution subject of the model optimization method in this exemplary embodiment can be the aforementioned terminal 110 or server 120, and this disclosure does not limit it in this regard.
[0040] The following is combined with Figure 2 The model optimization method in this exemplary embodiment will be described. Figure 2 An exemplary flow of the model optimization method is shown, which may include:
[0041] Step S210: Obtain the reference feature maps output by each convolutional layer in the image processing model;
[0042] Step S220: Determine the first weight of each channel of the convolutional layer based on the reference feature map;
[0043] Step S230: Determine an intermediate feature map based on the first weight and the reference feature map, and determine the second weight of each convolutional layer based on the intermediate feature map;
[0044] Step S240: Optimize the image processing model using the first weight and the second weight to obtain the target image processing model.
[0045] Based on the above method, the first weight of each channel in the convolutional layer of the image processing model is determined, taking into account the importance of each channel in each convolutional layer. The second weight of each convolutional layer in the whole model is also determined, taking into account the importance of the convolutional layer in the model. When optimizing the model, the image processing model can be optimized more accurately, which can improve the accuracy of the target image processing model.
[0046] The following is about Figure 2 Each step in the process will be explained in detail.
[0047] refer to Figure 2 In step S210, reference feature maps output by each convolutional layer in the image processing model are obtained.
[0048] In one exemplary embodiment of this disclosure, a training dataset and an initial model can be obtained first. The initial model is trained using the training dataset to obtain the aforementioned image processing model, which is a convolutional neural network model. The image processing model includes multiple convolutional layers. Then, any reference image can be obtained and input into the aforementioned image processing model to obtain reference feature maps corresponding to each convolutional layer in the image processing model.
[0049] In this example embodiment, the image processing model can be an image shadow removal model, an image recognition model, an image face recognition model, etc. In this example embodiment, the specific function of the image processing model is not specifically limited.
[0050] In step S220, the first weight of each channel of the convolutional layer is determined based on the reference feature map.
[0051] In one example implementation of this disclosure, determining the first weight of each channel of the convolutional layer based on the reference feature map may include steps S310 to S320.
[0052] In step S310, feature compression is performed on the initial sub-images of each channel of the reference feature map to obtain multiple feature values.
[0053] In this example implementation, refer to Figure 4 As shown, after obtaining the reference feature map 410 of each convolutional layer output, the initial sub-image in each reference feature map can be compressed to obtain multiple feature values 420. These feature values are used to characterize the importance of the corresponding channel of the initial sub-image.
[0054] Specifically, the aforementioned reference feature map may include images from multiple channels. Each channel's image can be defined as an initial sub-image corresponding to that channel. After obtaining the initial sub-images corresponding to each channel, feature compression can be performed on each of the aforementioned initial sub-images. The multiple initial sub-images have the same number of pixels and the same pixel array.
[0055] For example, if the initial sub-image contains H*W pixels, when compressing the features of the initial sub-image, the feature values corresponding to the initial sub-image can be obtained by performing convolution operation on each initial sub-image using an H*W convolution kernel.
[0056] In one example implementation, average pooling can also be performed on each initial sub-image to obtain the aforementioned feature values. Specifically, the feature values corresponding to each initial sub-image can be calculated using the following formula:
[0057]
[0058] Where z represents the aforementioned feature value, i and j represent the row and column identifiers in the initial sub-image, and u(i,j) represents the pixel coordinates.
[0059] It should be noted that there are various ways to perform feature compression on the initial sub-images mentioned above, and no specific limitation is made in this example implementation.
[0060] In step S320, the feature values corresponding to each channel in the reference feature map are normalized to obtain the first weight of each channel.
[0061] In this example implementation, after obtaining the feature values corresponding to the initial sub-images of each channel in the reference feature map, the multiple feature values in each reference feature map can be normalized to obtain the first weight 430 of each initial sub-image, that is, the first weight corresponding to each channel in the reference feature map.
[0062] Specifically, suppose the above reference feature map includes m channels, with feature values z1, z2, ... zn. m When performing normalization, z m The corresponding first weight 430 can be z m / (z1+z2+…+z m ).
[0063] It should be noted that there may be multiple ways to perform normalization operations. The above is an illustrative example and is not specifically limited in this example implementation.
[0064] In step S230, an intermediate feature map is determined based on the first weight and the reference feature map, and a second weight for each of the convolutional layers is determined based on the intermediate feature map.
[0065] In one example embodiment of this disclosure, after obtaining the first weight 430, an intermediate feature image can be obtained by multiplying the initial sub-image with the first weight 430. Specifically, each pixel in the initial sub-image is multiplied by the first weight corresponding to the channel of the initial sub-image to obtain the target sub-image corresponding to that channel.
[0066] After obtaining multiple target sub-images, the multiple target sub-images can be combined according to the channel order of the initial sub-images to obtain the aforementioned intermediate feature map.
[0067] In one exemplary embodiment of this disclosure, after obtaining the aforementioned intermediate feature images, the L1 norm of each target sub-image in each of the intermediate feature images can be calculated. Specifically, the L1 norm of the target sub-image is obtained by summing the absolute values of each pixel in the target sub-image. Then, the second weights of each convolutional layer are calculated based on the aforementioned L1 norm.
[0068] In this example implementation, refer to Figure 5 As shown, calculating the second weight based on the L1 norm may include steps S510 to S520.
[0069] In step S510, the sum of the L1 norms of each target sub-image is used as an importance index of the intermediate feature map.
[0070] In this example implementation, the importance index of the intermediate feature map can be obtained by summing the L1 norms of all target sub-images in the intermediate feature map.
[0071] Specifically, if the number of target sub-images in the intermediate feature map is m, then the L1 norm of the m target sub-images is summed to obtain the above importance index.
[0072] In step S520, the importance index corresponding to each intermediate feature map is normalized to obtain the second weight of each convolutional layer corresponding to each intermediate feature map.
[0073] After obtaining the importance index corresponding to each convolutional layer, the above multiple importance indices can be normalized to obtain the second weight of each convolutional layer corresponding to each intermediate feature map.
[0074] Specifically, assuming the image processing model described above includes n convolutional layers, it can include n intermediate feature maps. The importance index corresponding to the nth convolutional layer can be defined as G. n The second weight corresponding to the nth convolutional layer can be
[0075] It should be noted that there are multiple ways to normalize, and the above is only an example. This example implementation does not limit the normalization method to the above.
[0076] In step S240, the image processing model is optimized using the first weight and the second weight to obtain the target image processing model.
[0077] In this example implementation, refer to Figure 6 As shown, optimizing the image processing model using the first weight and the second weight to obtain the target image processing model may include steps S610 to S630.
[0078] In step S610, the importance ranking of each convolutional layer is determined based on the second weight.
[0079] In this example implementation, after obtaining the second weights corresponding to each of the above convolutional layers, the convolutional layers can be sorted to obtain the importance ranking of each convolutional layer, with the convolutional layer with the largest second weight ranked first, and then the multiple convolutional layers are sorted based on the size of the second weight.
[0080] In step S620, the sparsity of each of the convolutional layers is calculated based on the sorting and the second weight.
[0081] In this example implementation, after obtaining the above importance ranking, the coefficient ratio of each convolutional layer is calculated using the ranking and the second weight. Specifically, assuming there are N convolutional layers, and assuming the second weight corresponding to the nth convolutional layer in the ranking is Q. n Then the sparsity of the nth convolutional layer can be:
[0082]
[0083] Among them, P n Q represents the sparsity of the n-ordered convolutional layers. n The second weight is the weight of the convolutional layer sorted by n.
[0084] In step S630, the image processing model is optimized according to the sparsity rate and the first weight to obtain the target image processing model.
[0085] In this example implementation, after calculating the sparsity of each of the convolutional layers, the channels in each convolutional layer can be ranked according to their importance based on a first weight. Then, the image processing model is optimized according to the ranking to obtain the target image processing model. For example, if a convolutional layer includes m channels and the sparsity is P... n Then the m*P that is sorted last will be... n The intermediate image processing model is obtained by cropping the channels.
[0086] For example, refer to Figure 7 As shown, if the above convolutional layer includes 5 channels, which are sorted as channel 1, channel 2, channel 3, channel 4, and channel 5, and the sparsity is 20%, then channel 5 can be pruned. After pruning each convolutional layer, the target image processing model is obtained.
[0087] In this example implementation, after obtaining the intermediate image processing model, the intermediate image processing model can be trained again using the training data to obtain the target image processing model.
[0088] In one exemplary embodiment of this disclosure, reference is made to Figure 8 As shown, the above model optimization method may further include steps S810 to S830.
[0089] In step S810, the image processing accuracy of the target image processing model is obtained;
[0090] In step S820, in response to the image processing accuracy meeting the preset condition, the image processing model is updated using the target image processing model.
[0091] In this example implementation, after obtaining the target image processing model, the image processing accuracy of the target image processing model can be detected first, and then it can be determined whether the image processing accuracy meets the preset conditions. The preset conditions can be greater than or equal to 90%, or greater than or equal to 95%, or can be customized according to user needs. In this example implementation, no specific limitation is made.
[0092] In this example implementation, when obtaining the target image processing model, the graphics processing accuracy of the target image processing model can be verified using a validation dataset.
[0093] In this example embodiment, if the image processing accuracy of the target image processing model meets the preset conditions, the target image processing model can be used to update the image processing model, and then optimization can be performed again. That is, the target image processing model is used as the image processing model to execute steps S210 to S240 again to obtain a new target image processing model, and the processing accuracy is judged again.
[0094] In step S830, in response to the image processing accuracy not meeting the preset conditions, the image processing model is output.
[0095] In this example implementation, if the target image processing accuracy does not meet the preset condition, it indicates excessive pruning. In this case, the pruning at this point will affect the model's accuracy. Therefore, the pruning result is ignored, and the image processing model is output. Specifically, the image processing model is pruned multiple times until the image processing accuracy of the target image processing model obtained after the Kth pruning does not meet the preset condition. Then, the target image processing model obtained after the K-1th pruning is output. This results in the simplest target image processing model that guarantees image processing accuracy.
[0096] Specifically, refer to Figure 9 As shown, the process can begin with step S910 to obtain the reference feature maps of each convolutional layer output. Then, step S920 is executed to determine the first weights of each channel of the convolutional layer. Next, step S930 is executed to determine the second weights of each convolutional layer. After obtaining the second weights, step S940 is executed to prune the image processing model to obtain an intermediate image processing model. Then, step S950 is executed to train the intermediate image processing model to obtain the target image processing model. Finally, step S960 is executed to determine whether the image processing accuracy of the target image processing model meets the preset conditions. If yes, step S970 is executed to update the image output model using the target image processing model. If no, step S980 is executed to output the image processing model.
[0097] The specific details of each of the above steps have already been explained in detail, so they will not be repeated here.
[0098] In one exemplary embodiment of this disclosure, reference is made to Figure 10 As shown, the above model optimization method may further include steps S1010 to S1030.
[0099] In step S1010, the image processing accuracy of the target image processing model is obtained;
[0100] In step S1020, in response to the image processing accuracy meeting the preset conditions, it is determined whether the target image processing model meets the model optimization conditions.
[0101] In one example embodiment of this disclosure, after obtaining the target image processing model, the image processing accuracy of the target image processing model can be detected first, and then it can be determined whether the image processing accuracy meets the preset conditions. The preset conditions can be greater than or equal to 90%, or greater than or equal to 95%, or can be customized according to user needs. In this example embodiment, no specific limitation is made.
[0102] In this example implementation, when obtaining the target image processing model, the graphics processing accuracy of the target image processing model can be verified using a validation dataset.
[0103] In this example implementation, if the image processing accuracy of the target image processing model meets the preset conditions, it can be calculated that the target image processing model meets the model optimization conditions. The model optimization conditions can be a model pruning rate greater than or equal to 20%, or a model pruning rate greater than or equal to 25%. These conditions can also be customized according to user needs, and are not specifically limited in this example implementation.
[0104] In step S1030, in response to the target image processing model not meeting the model optimization conditions, the image processing model is updated using the target image processing model.
[0105] In this example implementation, if the target image processing model does not meet the above model optimization conditions, the image processing model can be updated using the target image processing model, and then optimized again. That is, the target image processing model is used as the image processing model to execute steps S210 to S240 again to obtain a new target image processing model.
[0106] In this example implementation, if the image processing accuracy of the target image processing model does not meet the preset conditions, the target image processing model is trained again using the training data. If the target image processing module still fails to meet the preset conditions after multiple training iterations, training is stopped, and the image processing model is output.
[0107] In this example implementation, if the target image processing model satisfies the above model optimization conditions, the target processing model is output to complete the image processing.
[0108] Specifically, refer to Figure 11As shown, the process can begin with step S1110 to obtain the reference feature maps of each convolutional layer output. Then, step S1120 is executed to determine the first weights of each channel of the convolutional layer. Next, step S1130 is executed to determine the second weights of each convolutional layer. After obtaining the second weights, step S1140 is executed to prune the image processing model to obtain an intermediate image processing model. Then, step S1150 is executed to train the intermediate image processing model to obtain the target image processing model. Then, step S1160 is executed to determine whether the image processing accuracy of the target image processing model meets the preset conditions. If not, step S1170 is executed to train the target image processing model using the training data. If yes, step S1180 is executed to determine whether the target image processing model meets the model optimization conditions. If yes, step S1190 is executed to output the target image processing model. If not, step S1191 is executed to update the image output model using the target image processing model.
[0109] The specific details of each of the above steps have already been explained in detail, so they will not be repeated here.
[0110] In summary, the first weight of each channel in the convolutional layers of the image processing model was determined, considering the importance of each channel in each convolutional layer. The second weight of each convolutional layer in the entire model was also determined, considering its importance within the model. This allows for more precise optimization of the image processing model, improving the accuracy of the resulting target image processing model. In one embodiment, after pruning the model, multiple pruning cycles are performed until the image processing accuracy of the target image processing model obtained after the Kth pruning cycle does not meet the preset conditions. Then, the target image processing model obtained after the K-1th pruning cycle is output. This results in the simplest target image processing model that guarantees image processing accuracy.
[0111] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0112] Furthermore, this disclosure also provides a new image processing method. Specifically, the image to be processed and a pre-trained image processing model can be obtained first. Then, the image processing model can be optimized using the above-mentioned model optimization method to obtain a target image processing model. The image to be processed is then input into the target image processing model to process the image.
[0113] Further reference Figure 12As shown, this example embodiment also provides a model optimization device 1200, including an acquisition module 1210, a first determination module 1220, a second determination module 1230, and an optimization module 1240. Wherein:
[0114] The acquisition module 1210 can be used to acquire reference feature maps output by each convolutional layer in the image processing model. Specifically, it can first acquire a reference image and a pre-trained image processing model, and then input the reference image into the image processing model to obtain the reference feature maps of each convolutional layer.
[0115] The first determining module 1220 can be used to determine the first weight of each channel of the convolutional layer based on the reference feature map.
[0116] In one example implementation, the first determining module 1220 may be configured to perform feature compression on the initial sub-images of each channel of the reference feature map to obtain multiple feature values; and to normalize the feature values corresponding to each channel in the reference feature map to obtain a first weight for each channel.
[0117] The second determining module 1230 can be used to determine an intermediate feature map based on the first weight and the reference feature map, and to determine the second weight of each of the convolutional layers according to the intermediate feature map.
[0118] In one example implementation, the second determining module 1230 may be configured to multiply the first weight of each channel with the initial sub-image of each channel to obtain the target sub-image corresponding to each channel; determine the intermediate feature map based on each target sub-image, and determine the L1 norm of each target sub-image in the intermediate feature map; and calculate the second weight based on the L1 norm.
[0119] In one example implementation, the second determining module 1230 may be configured to sum the L1 norm of each of the target sub-images as an importance index of the intermediate feature map; and to perform a normalization operation on the importance index corresponding to each of the intermediate feature maps to obtain the second weight of each convolutional layer corresponding to each of the intermediate feature maps.
[0120] The optimization module 1240 can be used to optimize the image processing model using the first weight and the second weight to obtain the target image processing model.
[0121] In one example implementation, the optimization module 1240 may be configured to determine the importance ranking of each of the convolutional layers based on the second weight; calculate the sparsity rate of each of the convolutional layers based on the ranking and the second weight; and optimize the image processing model according to the sparsity rate and the first weight to obtain a target image processing model.
[0122] In one example implementation, the optimization module 1240 may be configured to prune the image processing model according to the sparsity rate and the first weight to obtain an intermediate image processing model; and train the intermediate image processing model using training data to obtain a target image processing model.
[0123] In one exemplary embodiment of this disclosure, the model optimization device can also obtain the image processing accuracy of the target image processing model; update the image processing model using the target image processing model in response to the image processing accuracy meeting a preset condition; and output the image processing model in response to the image processing accuracy not meeting the preset condition.
[0124] In another exemplary embodiment of this disclosure, the model optimization apparatus may further acquire the image processing accuracy of the target image processing model; determine whether the target image processing model meets the model optimization conditions in response to the image processing accuracy meeting a preset condition; update the image processing model using the target image processing model in response to the target image processing model not meeting the model optimization conditions; train the target image processing model using training data in response to the image processing accuracy not meeting the preset condition; and output the target image processing model in response to the target image processing model meeting the model optimization conditions.
[0125] The specific details of each module in the above-mentioned device have been described in detail in the method section of the implementation. For any undisclosed details, please refer to the implementation content of the method section, and therefore will not be repeated here.
[0126] Exemplary embodiments of this disclosure also provide an electronic device for performing the above-described model optimization method. This electronic device may be the terminal 110 or the server 120 described above. Generally, the electronic device may include a processor and a memory, the memory for storing executable instructions of the processor, and the processor configured to perform the above-described model optimization method by executing the executable instructions.
[0127] The following is based on Figure 13 Taking the mobile terminal 1300 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 13 The structure can also be applied to fixed types of equipment.
[0128] like Figure 13As shown, the mobile terminal 1300 may specifically include: a processor 1301, a memory 1302, a bus 1303, a mobile communication module 1304, an antenna 1, a wireless communication module 1305, an antenna 2, a display screen 1306, a camera module 1307, an audio module 1308, a power module 1309, and a sensor module 310.
[0129] Processor 1301 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The model optimization method in this exemplary embodiment can be executed by an AP, GPU, or DSP. When the method involves neural network-related processing, it can be executed by an NPU.
[0130] Mobile terminal 1300 can support one or more encoders and decoders. Thus, mobile terminal 1300 can process images or videos in various encoding formats, such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap) image formats, and MPEG (Moving Picture Experts Group) 1, MPEG2, H.263, H.264, HEVC (High Efficiency Video Coding) video formats.
[0131] The processor 1301 can be connected to the memory 1302 or other components via the bus 1303.
[0132] The memory 1302 can be used to store computer executable program code, which includes instructions. The processor 1301 executes various functional applications and data processing of the mobile terminal 1300 by running the instructions stored in the memory 1302. The memory 1302 can also store application data, such as images, videos, and other files.
[0133] The communication function of mobile terminal 1300 can be implemented through mobile communication module 1304, antenna 1, wireless communication module 1305, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1304 can provide 2G, 3G, 4G, and 5G mobile communication solutions for mobile terminal 1300. Wireless communication module 1305 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1300.
[0134] The display screen 1306 is used to implement display functions, such as displaying the user interface, images, and videos. The camera module 1307 is used to implement shooting functions, such as capturing images and videos. The audio module 1308 is used to implement audio functions, such as playing audio and capturing voice. The power module 1309 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1310 may include a depth sensor 13101, a pressure sensor 13102, a gyroscope sensor 13103, a barometric pressure sensor 13104, etc., to implement corresponding sensing and detection functions.
[0135] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0136] Exemplary embodiments of this disclosure also provide a computer-readable storage medium having a program product stored thereon capable of implementing the methods described above in this specification. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.
[0137] It should be noted that the computer-readable medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0138] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0139] Furthermore, program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0140] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0141] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A model optimization method, characterized in that, include: Obtain reference feature maps from the outputs of each convolutional layer in the image processing model; Determining the first weight of each channel of the convolutional layer based on the reference feature map includes: performing feature compression on the initial sub-image of each channel of the reference feature map to obtain multiple feature values; and normalizing the feature values corresponding to each channel in the reference feature map to obtain the first weight of each channel. Determining an intermediate feature map based on the first weight and the reference feature map includes: multiplying the first weight of each channel with the initial sub-image of each channel to obtain a target sub-image corresponding to each channel; and determining the intermediate feature map based on each target sub-image. Determining the second weights of each convolutional layer based on the intermediate feature map includes: determining the L1 norm of each target sub-image in the intermediate feature map; summing the L1 norms of each target sub-image as an importance index of the intermediate feature map; and normalizing the importance index corresponding to each intermediate feature map to obtain the second weights of each convolutional layer corresponding to each intermediate feature map. The image processing model is optimized using the first weight and the second weight to obtain the target image processing model.
2. The method according to claim 1, characterized in that, The step of optimizing the image processing model using the first weight and the second weight to obtain the target image processing model includes: The importance ranking of each convolutional layer is determined based on the second weight; The sparsity of each of the convolutional layers is calculated based on the sorting and the second weight; The target image processing model is obtained by optimizing the image processing model based on the sparsity rate and the first weight.
3. The method according to claim 2, characterized in that, The step of optimizing the image processing model based on the sparsity rate and the first weight to obtain the target image processing model includes: The image processing model is pruned according to the sparsity rate and the first weight to obtain an intermediate image processing model; The intermediate image processing model is trained using the training data to obtain the target image processing model.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: Obtain the image processing accuracy of the target image processing model; When the image processing accuracy meets the preset condition, the image processing model is updated using the target image processing model. If the image processing accuracy does not meet the preset conditions, the image processing model is output.
5. The method according to any one of claims 1-3, characterized in that, The method further includes: Obtain the image processing accuracy of the target image processing model; In response to the image processing accuracy meeting the preset conditions, determine whether the target image processing model meets the model optimization conditions; If the target image processing model does not meet the model optimization conditions, the image processing model is updated using the target image processing model.
6. The method according to claim 5, characterized in that, The method further includes: In response to the image processing accuracy not meeting the preset conditions, the target image processing model is trained using training data.
7. The method according to claim 5, characterized in that, The method further includes: When the target image processing model satisfies the model optimization conditions, the target image processing model is output.
8. The method according to claim 1, characterized in that, The acquisition of reference feature maps output by each convolutional layer in the image processing model includes: Acquire reference images and pre-trained image processing models; The reference image is input into the image processing model to obtain reference feature maps for each convolutional layer.
9. An image processing method, characterized in that, include: Acquire the image to be processed and the pre-trained image processing model; The image processing model is optimized using the optimization method of the image processing model as described in any one of claims 1 to 8 to obtain the target image processing model; The image to be processed is input into the target image processing model to process the image.
10. A model optimization device, characterized in that, include: The acquisition module is used to acquire reference feature maps output by each convolutional layer in the image processing model; The first determining module is used to determine the first weight of each channel of the convolutional layer based on the reference feature map, including: performing feature compression on the initial sub-image of each channel of the reference feature map to obtain multiple feature values; and performing normalization processing on the feature values corresponding to each channel in the reference feature map to obtain the first weight of each channel. The second determining module is used to determine an intermediate feature map based on the first weight and the reference feature map, and to determine the second weight of each of the convolutional layers based on the intermediate feature map; An optimization module is used to optimize the image processing model using the first weight and the second weight to obtain a target image processing model; The step of determining the intermediate feature map based on the first weight and the reference feature map includes: multiplying the first weight of each channel with the initial sub-image of each channel to obtain the target sub-image corresponding to each channel; and determining the intermediate feature map based on each target sub-image. Determining the second weights of each convolutional layer based on the intermediate feature map includes: determining the L1 norm of each target sub-image in the intermediate feature map; summing the L1 norms of each target sub-image as an importance index of the intermediate feature map; and normalizing the importance index corresponding to each intermediate feature map to obtain the second weights of each convolutional layer corresponding to each intermediate feature map.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 9.
12. An electronic device, characterized in that, include: One or more processors; as well as A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Neural network compression method, apparatus and device, and storage medium
CN111967594A