Pruning methods, devices, servers, and storage media for multimedia classification models
By selecting the feature and weight matrices of the multimedia classification model, performing crossover and mutation processing, and optimizing the Pareto front algorithm, the problems of large search space and low accuracy in multimedia classification model pruning were solved, achieving efficient pruning and improved accuracy.
Patent Information
- Application Number
- CN202210244604.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-03-14
AI Technical Summary
Existing technologies face problems of large search space and low accuracy when pruning multimedia classification models, especially for multimedia classification sub-models with very long coding sequences, resulting in high search difficulty and low accuracy.
By filtering the feature and weight matrices of each layer of the multimedia classification sub-model, and combining crossover mutation processing and Pareto front algorithm, the multimedia classification sub-model is gradually pruned and optimized. Channel number encoding is adopted instead of channel encoding to reduce the length of the encoding sequence and storage space, thereby improving pruning efficiency and accuracy.
High-precision multimedia classification sub-models are selected within a limited time, shortening the pruning time, improving the accuracy and pruning efficiency of the final model, and reducing the storage space of the encoded sequence.
Smart Images

Figure CN114610912B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a pruning method, apparatus, server, and storage medium for a multimedia classification model. Background Technology
[0002] Multimedia classification models are a hot research topic in computer science, applicable to image recognition, object detection, text recognition, image segmentation, and natural language processing. These models are generally large and consume significant computational resources. To improve their performance, pruning is necessary.
[0003] The relevant technology uses a genetic algorithm to prune multimedia classification models. After each pruning, each channel of each layer of the resulting multimedia classification sub-model is encoded one by one to obtain the encoding sequence corresponding to each multimedia classification sub-model. Then, the genetic algorithm searches in the encoding sequence corresponding to each multimedia classification sub-model to obtain the weight matrix of each layer of each multimedia classification sub-model. Based on the weight matrix of each layer of each multimedia classification sub-model, the pruned multimedia classification sub-models are screened. Then, the screened multimedia classification sub-models are subjected to crossover and mutation to finally obtain the optimal multimedia classification sub-model under different pruning rates.
[0004] However, the encoding sequences corresponding to each multimedia classification sub-model in related technologies are very long. For example, the multimedia classification sub-model with the ResNet50 network structure has 24,576 channels excluding the first layer, and the search space of the encoding sequence is 2^34. 24576 Faced with such a huge search space, it is difficult to search using relevant technologies, and the resulting multimedia classification sub-model has low accuracy. Summary of the Invention
[0005] This disclosure provides a pruning method, apparatus, server, and storage medium for a multimedia classification model, which can improve the accuracy of the obtained multimedia classification sub-model. The technical solution is as follows:
[0006] Firstly, a pruning method for a multimedia classification model is provided, the method comprising:
[0007] Based on the feature matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning is determined, wherein I is a positive integer and the value of I is greater than or equal to 2.
[0008] Based on the weight matrix of each layer of the multiple multimedia sub-models obtained by the first pruning and the weight matrix of each layer of the multiple multimedia sub-models obtained by the (I-1)th pruning, the multiple multimedia sub-models obtained by the first and (I-1)th pruning are filtered to obtain multiple multimedia sub-models after filtering.
[0009] The number of channels in each layer of the selected multimedia classification sub-models is cross-mutated to obtain the processed multimedia classification sub-models.
[0010] Based on the number of channels in each layer of the processed multiple multimedia classification sub-models, a pruning operation is performed on the processed multiple multimedia classification sub-models to obtain the multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0011] Based on the multiple multimedia classification sub-models obtained from the (I+1)th pruning, the (I+2)th pruning operation is performed until the number of prunings reaches T. From the multiple multimedia classification sub-models obtained when the number of prunings reaches T, the optimal multimedia classification sub-model under different pruning rates is obtained, where T is a positive integer and the value of T is greater than or equal to I+2.
[0012] In another embodiment of this disclosure, before determining the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning based on the feature matrix of each layer, the method further includes:
[0013] Encode the number of channels in each layer of the multiple multimedia classification sub-models obtained by the I-th pruning to obtain the encoding sequence corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning. The encoding sequence is used to record the number of channels and feature matrix of each layer of the multimedia classification sub-model.
[0014] From the encoding sequences corresponding to the multiple multimedia sub-models obtained by the I-th pruning, obtain the feature matrix of each layer of the multiple multimedia sub-models obtained by the I-th pruning.
[0015] In another embodiment of this disclosure, determining the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning based on the feature matrix of each layer includes:
[0016] Based on the feature matrices corresponding to the input channels and output channels of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning, the weight matrix of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning is determined using the following formula:
[0017] W * =(X T X)-1 XY
[0018] Among them, W * Let X represent the weight matrix of any layer of any multimedia classifier submodel obtained by the I-th pruning, let Y represent the feature matrix corresponding to the input channel of the layer of the multimedia classifier submodel, and let Y represent the feature matrix corresponding to the output channel of the layer of the multimedia classifier submodel.
[0019] In another embodiment of this disclosure, the multiple multimedia classification sub-models obtained from the first and (I-1)th pruning processes are filtered based on the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the first pruning and the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the (I-1)th pruning processes to obtain a plurality of filtered multimedia classification sub-models, including:
[0020] Multiple multimedia classification sub-models obtained from the I-th pruning with weight matrices and multiple multimedia classification sub-models obtained from the (I-1)-th pruning are called to process the multimedia test resource samples, and the classification results of the multimedia resource test samples under each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning are obtained.
[0021] Based on the classification and annotation results of each multimedia classification sub-model obtained by the first and (I-1)th pruning of the multimedia resource test samples, the classification accuracy of each multimedia classification sub-model obtained by the first and (I-1)th pruning is determined.
[0022] Based on the classification accuracy of each multimedia sub-model obtained from the first and (I-1)th pruning, the multiple multimedia sub-models obtained from the first and (I-1)th pruning are filtered to obtain multiple filtered multimedia sub-models.
[0023] In another embodiment of this disclosure, the step of performing cross-mutation processing on the number of channels in each layer of the selected multiple multimedia classification sub-models to obtain processed multiple multimedia classification sub-models includes:
[0024] For some multimedia classification sub-models among the multiple multimedia classification sub-models after screening, the number of channels in at least one layer of any two multimedia classification sub-models in the partial multimedia classification sub-models is cross-exchanged to obtain the cross-exchanged partial multimedia classification sub-models.
[0025] For the remaining multimedia classification sub-models among the multiple filtered multimedia classification sub-models, the number of channels in at least one layer of any two filtered multimedia classification sub-models among the remaining multimedia classification sub-models is mutated to obtain the mutated remaining multimedia classification sub-models.
[0026] The processed multimedia classification sub-models are composed of the partially cross-interchanged multimedia classification sub-models and the mutated remaining multimedia classification sub-models.
[0027] In another embodiment of this disclosure, the step of pruning the processed multimedia classification sub-models based on the number of channels in each layer of the processed multimedia classification sub-models to obtain the (I+1)th pruning of the multimedia classification sub-models includes:
[0028] For each processed multimedia classification sub-model, based on the number of input channels in the next layer of each processed multimedia classification sub-model, the number of output channels in the previous layer is pruned to obtain multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0029] In another embodiment of this disclosure, obtaining the optimal multimedia classification sub-model under different pruning rates from multiple multimedia classification sub-models obtained when the number of prunings reaches T times includes:
[0030] The Pareto front algorithm was used to process multiple multimedia sub-models obtained after pruning T times to obtain the optimal multimedia sub-model under different pruning rates.
[0031] Secondly, a pruning device for a multimedia classification model is provided, the device comprising:
[0032] The determination module is used to determine the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, based on the feature matrix of each layer. Here, I is a positive integer and the value of I is greater than or equal to 2.
[0033] The filtering module is used to filter the multiple multimedia sub-models obtained by the first and (I-1)th pruning based on the weight matrix of each layer of the multiple multimedia sub-models obtained by the first pruning and the weight matrix of each layer of the multiple multimedia sub-models obtained by the (I-1)th pruning, so as to obtain multiple multimedia sub-models after filtering.
[0034] The processing module is used to perform cross-mutation processing on the number of channels in each layer of the multiple multimedia classification sub-models after screening, so as to obtain multiple processed multimedia classification sub-models.
[0035] The pruning module is used to perform pruning operations on the processed multiple multimedia classification sub-models based on the number of channels in each layer of the multiple multimedia classification sub-models, to obtain the multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0036] The acquisition module is used to perform the (I+2)th pruning operation on the multiple multimedia classification sub-models obtained by the (I+1)th pruning, until the number of prunings reaches T. From the multiple multimedia classification sub-models obtained when the number of prunings reaches T, the optimal multimedia classification sub-model under different pruning rates is obtained, where T is a positive integer and the value of T is greater than or equal to I+2.
[0037] In another embodiment of this disclosure, the apparatus further includes:
[0038] The encoding module is used to encode the number of channels in each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, and to obtain the encoding sequence corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning. The encoding sequence is used to record the number of channels and feature matrix of each layer of the multimedia classification sub-model.
[0039] The acquisition module is further configured to acquire the feature matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning from the encoding sequences corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning.
[0040] In another embodiment of this disclosure, the determining module is further configured to determine the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, based on the feature matrices corresponding to the input channels and the feature matrices corresponding to the output channels of each layer, using the following formula:
[0041] W * =(X T X) -1 XY
[0042] Among them, W * Let X represent the weight matrix of any layer of any multimedia classifier submodel obtained by the I-th pruning, let Y represent the feature matrix corresponding to the input channel of the layer of the multimedia classifier submodel, and let Y represent the feature matrix corresponding to the output channel of the layer of the multimedia classifier submodel.
[0043] In another embodiment of this disclosure, the filtering module is used to call multiple multimedia classification sub-models obtained from the I-th pruning with weight matrices and multiple multimedia classification sub-models obtained from the (I-1)-th pruning to process multimedia test resource samples, thereby obtaining classification results of multimedia resource test samples under each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning; based on the classification results and annotation results of the multimedia resource test samples under each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning, the classification accuracy of each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning is determined; based on the classification accuracy of each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning, the multiple multimedia classification sub-models obtained from the I-th and (I-1)-th pruning are filtered to obtain multiple filtered multimedia classification sub-models.
[0044] In another embodiment of this disclosure, the processing module is configured to: cross-interchange the number of channels in at least one layer of any two multimedia classification sub-models among the selected plurality of multimedia classification sub-models to obtain cross-interchangeable partial multimedia classification sub-models; mutate the number of channels in at least one layer of any two selected multimedia classification sub-models among the remaining multimedia classification sub-models to obtain mutated remaining multimedia classification sub-models; and combine the cross-interchangeable partial multimedia classification sub-models and the mutated remaining multimedia classification sub-models to form the processed plurality of multimedia classification sub-models.
[0045] In another embodiment of this disclosure, the pruning module is used to prune the number of output channels of the previous layer based on the number of input channels of the next layer of each processed multimedia classification sub-model, so as to obtain multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0046] In another embodiment of this disclosure, the acquisition module is used to process multiple multimedia sub-models obtained when the number of prunings reaches T times using the Pareto front algorithm to obtain the optimal multimedia sub-model under different pruning rates.
[0047] Thirdly, a server is provided, the server including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the pruning method of the multimedia classification model as described in the first aspect.
[0048] Fourthly, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the storage medium, the at least one piece of program code being loaded and executed by a processor to implement the pruning method of the multimedia classification model described in the first aspect.
[0049] Fifthly, a computer program product is provided, the computer program product including computer program code stored in a computer-readable storage medium, a processor of a server reading the computer program code from the computer-readable storage medium, the processor executing the computer program code, causing the server to perform the pruning method of the multimedia classification model as described in the first aspect.
[0050] The beneficial effects of the technical solutions provided in this disclosure are:
[0051] Based on the feature matrices corresponding to the input and output channels of each layer of the multimedia classification sub-model obtained after each pruning, the weight matrix of each layer can be determined without searching the encoding sequence. This allows for the selection of high-precision multimedia classification sub-models for subsequent evolution within a limited time, improving the accuracy of the final multimedia classification sub-model. Furthermore, this disclosure proposes an encoding method for the multimedia classification sub-model obtained after each pruning. This method encodes the number of channels in each layer of the multimedia classification sub-model instead of individual channels in each layer. Compared to encoding each channel individually, the encoding sequence is shorter, thus reducing storage space. Moreover, the channel-based encoding method eliminates the focus on which channel was pruned in each layer of the intermediate multimedia classification sub-model during subsequent evolution, focusing only on the number of channels retained in each layer. Compared to existing pruning methods, this significantly shortens the pruning time and improves pruning efficiency. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 This is a flowchart of a pruning method for a multimedia classification model provided in an embodiment of this disclosure;
[0054] Figure 2 This is a flowchart of another pruning method for a multimedia classification model provided in this embodiment;
[0055] Figure 3 This is a schematic diagram of the pruning process of a multimedia classification model based on a genetic algorithm provided in an embodiment of this disclosure;
[0056] Figure 4 This is a schematic diagram of a pruning device for a multimedia classification model provided in an embodiment of the present disclosure;
[0057] Figure 5 This is a server for pruning a multimedia classification model, as illustrated in an exemplary embodiment. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this disclosure clearer, the embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings.
[0059] It is understood that the terms "each," "multiple," and "any," etc., used in the embodiments of this disclosure, include "multiple" (two or more), "each" (each of the corresponding multiples), and "any" (any one of the corresponding multiples). For example, multiple words include 10 words, and "each word" refers to each of the 10 words, while "any word" refers to any one of the 10 words.
[0060] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.
[0061] This disclosure provides a pruning method for a multimedia classification model, see [link to relevant documentation]. Figure 1 The method flow provided in this disclosure includes:
[0062] 101. Based on the feature matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, determine the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning.
[0063] Where I is a positive integer, and the value of I is greater than or equal to 2.
[0064] 102. Based on the weight matrices of each layer of the multiple multimedia sub-models obtained by the I-th pruning and the weight matrices of each layer of the multiple multimedia sub-models obtained by the (I-1)-th pruning, the multiple multimedia sub-models obtained by the I-th and (I-1)-th pruning are filtered to obtain the filtered multiple multimedia sub-models.
[0065] 103. Perform cross-mutation on the number of channels in each layer of the selected multimedia classification sub-models to obtain the processed multimedia classification sub-models.
[0066] 104. Based on the number of channels in each layer of the processed multiple multimedia classification sub-models, perform pruning operations on the processed multiple multimedia classification sub-models to obtain the multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0067] 105. Based on the multiple multimedia sub-models obtained from the (I+1)th pruning, perform the (I+2)th pruning operation until the number of prunings reaches T. From the multiple multimedia sub-models obtained when the number of prunings reaches T, obtain the optimal multimedia sub-model under different pruning rates.
[0068] Where T is a positive integer, and the value of T is greater than or equal to I+2.
[0069] The method provided in this disclosure determines the weight matrix of each layer of the multimedia classification sub-model obtained after each pruning based on the feature matrices corresponding to the input and output channels of each layer. This eliminates the need for searching the encoding sequence, allowing for the selection of higher-precision multimedia classification sub-models for subsequent evolution within a limited timeframe, thus improving the accuracy of the final multimedia classification sub-model. Furthermore, this disclosure proposes an encoding method for each pruning multimedia classification sub-model. This method encodes the number of channels in each layer of the multimedia classification sub-model instead of individual channels. Compared to encoding each channel individually, this results in a shorter encoding sequence, reducing storage space. Moreover, the channel-based encoding method eliminates the focus on which channel was pruned in each layer during subsequent evolution, only the number of channels retained in each layer. Compared to existing pruning methods, this significantly shortens the pruning time and improves pruning efficiency.
[0070] In another embodiment of this disclosure, before determining the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, based on the feature matrix of each layer, the method further includes:
[0071] Encode the number of channels in each layer of the multiple multimedia classification sub-models obtained by the I-th pruning to obtain the encoding sequence corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning. The encoding sequence is used to record the number of channels and feature matrix of each layer of the multimedia classification sub-model.
[0072] From the encoding sequences corresponding to the multiple multimedia sub-models obtained by the I-th pruning, obtain the feature matrix of each layer of the multiple multimedia sub-models obtained by the I-th pruning.
[0073] In another embodiment of this disclosure, the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning is determined based on the feature matrix of each layer, including:
[0074] Based on the feature matrices corresponding to the input channels and output channels of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning, the weight matrix of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning is determined using the following formula:
[0075] W * =(X T X) -1 XY
[0076] Among them, W * Let X represent the weight matrix of any layer of any multimedia classifier sub-model obtained by the I-th pruning, let X represent the feature matrix corresponding to the input channel of the multimedia classifier sub-model layer, and let Y represent the feature matrix corresponding to the output channel of the multimedia classifier sub-model layer.
[0077] In another embodiment of this disclosure, based on the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning and the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the (I-1)th pruning, the multiple multimedia classification sub-models obtained from the I-th and (I-1)th pruning are filtered to obtain multiple filtered multimedia classification sub-models, including:
[0078] Multiple multimedia classification sub-models obtained from the I-th pruning with weight matrices and multiple multimedia classification sub-models obtained from the (I-1)-th pruning are called to process the multimedia test resource samples, and the classification results of the multimedia resource test samples under each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning are obtained.
[0079] Based on the classification and annotation results of each multimedia sub-model obtained by the first and second pruning of multimedia resource test samples, the classification accuracy of each multimedia sub-model obtained by the first and second pruning is determined.
[0080] Based on the classification accuracy of each multimedia sub-model obtained from the I-th and I-1-th pruning, the multiple multimedia sub-models obtained from the I-th and I-1-th pruning are filtered to obtain multiple filtered multimedia sub-models.
[0081] In another embodiment of this disclosure, the number of channels in each layer of the selected multiple multimedia classification sub-models is subjected to cross-mutation processing to obtain the processed multiple multimedia classification sub-models, including:
[0082] For a subset of multimedia sub-models among the selected multiple multimedia sub-models, the number of channels in at least one layer of any two multimedia sub-models is cross-interchanged to obtain a cross-interchanged subset of multimedia sub-models.
[0083] For the remaining multimedia sub-models among the multiple selected multimedia sub-models, the number of channels in at least one layer of any two selected multimedia sub-models in the remaining multimedia sub-models is mutated to obtain the mutated remaining multimedia sub-models.
[0084] The cross-interchanged portion of the multimedia sub-models and the mutated remaining multimedia sub-models are combined to form multiple processed multimedia sub-models.
[0085] In another embodiment of this disclosure, based on the number of channels in each layer of the processed multiple multimedia classification sub-models, a pruning operation is performed on the processed multiple multimedia classification sub-models to obtain the multiple multimedia classification sub-models obtained by the (I+1)th pruning, including:
[0086] For each processed multimedia classification sub-model, the number of output channels in the previous layer is pruned based on the number of input channels in the next layer of each processed multimedia classification sub-model, resulting in multiple multimedia classification sub-models obtained from the (I+1)th pruning.
[0087] In another embodiment of this disclosure, the optimal multimedia classification sub-model under different pruning rates is obtained from multiple multimedia classification sub-models obtained when the number of prunings reaches T times, including:
[0088] The Pareto front algorithm was used to process multiple multimedia sub-models obtained after pruning T times to obtain the optimal multimedia sub-model under different pruning rates.
[0089] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0090] This disclosure provides a pruning method for a multimedia classification model. Taking a server executing this disclosure as an example, the server can be a single physical server, or a cluster or distributed system composed of multiple physical servers. This disclosure does not specifically limit the type of server. See also Figure 2 The method flow provided in this disclosure includes:
[0091] 201. The server encodes the number of channels in each layer of the multiple multimedia sub-models obtained from the I-th pruning, and obtains the encoding sequence corresponding to the multiple multimedia sub-models obtained from the I-th pruning.
[0092] The multimedia classification model is trained on a deep learning model based on multimedia resource samples and is used to classify multimedia resources. The training process is as follows: Multiple multimedia resource samples are acquired, each corresponding to a classification label. An initial multimedia classification model is obtained, and its initial parameters are set. Each multimedia resource sample is input into the initial multimedia classification model, and the prediction result for each sample is output. Next, the classification label and prediction result for each multimedia resource sample are input into the target loss function, and the function value of the target loss function is calculated. If the target loss function value does not meet the threshold condition, the model parameters of the initial multimedia classification model are adjusted, and the target loss function value is calculated again until the obtained value meets the threshold condition. The threshold condition can be set according to the processing accuracy. The parameter values of each parameter when the threshold condition is met are obtained, and the initial multimedia classification model corresponding to these parameter values is used as the trained multimedia classification model.
[0093] The functionality of the trained multimedia classification model can be determined by the classification task set during training. For example, if the multimedia resource is an audio resource, and the classification task set during training is to identify whether the music type of the audio resource is country, jazz, or rock, then the multimedia classification model trained according to this classification task can identify the music type of any input audio file. If the multimedia resource is an image resource, and the classification task set during training is to determine whether the image is located in a preset image set, then the multimedia classification model trained according to this classification task can determine whether any input image is located in the preset image set. If the multimedia resource is a text resource, and the classification task set during training is to determine whether the text content type is Chinese, English, or Japanese, then the multimedia classification model trained according to this classification task can identify the text content type of any input text resource. If the multimedia resource is a video resource, and the classification task set during training is to determine whether the video type of the video resource is a comedy, martial arts, or romance film, then the multimedia classification model trained according to this classification task can identify the video type of any input video resource.
[0094] Considering the large size of the trained multimedia classification model, pruning is necessary to reduce the computational load and improve its performance when classifying multimedia resources. Pruning operations include channel pruning and kernel pruning. That is, when pruning the multimedia classification model, either the channels of each layer or the convolutional kernels of each layer can be pruned. Since the convolutional kernels also need to be pruned accordingly after channel pruning, the same pruning effect can be obtained regardless of whether either the channels or convolutional kernels of each layer are pruned. This embodiment of the present disclosure uses channel pruning of each layer of the multimedia classification model as an example for illustration.
[0095] The initial multimedia classification model has an N-layer network structure. A first pruning operation is performed on this model, resulting in P multimedia classification sub-models. Subsequent evolution will be based on these P sub-models obtained from the first pruning. Here, N and P are positive integers, both greater than or equal to 1. It should be noted that, in order to use a genetic algorithm to select the best-structured pruning result from the population (the multimedia classification sub-models obtained from each layer of pruning), this embodiment maintains a constant population size P each time, ensuring that the number of multimedia classification sub-models obtained from each pruning is always P.
[0096] Considering the inherent characteristics of genetic algorithms, after each pruning, the resulting multimedia classification sub-models need to be encoded, and the encoded sequences of each pruned multimedia classification sub-model are stored. Then, based on the stored encoded sequences, the pruned multimedia classification sub-models are evolved. According to modern pruning theory, the performance of a pruned network is strongly correlated with its structure and less correlated with the specific channels in each layer. Encoding each channel in each layer not only has no direct impact on network performance evaluation but also increases the length of the encoded sequence, hindering effective network evolution. Guided by modern pruning theory, this embodiment, when encoding the multiple multimedia classification sub-models obtained after each pruning, does not encode each channel in each layer of the multiple multimedia classification sub-models obtained after each pruning. Instead, it encodes the number of channels in each layer of the multiple multimedia classification sub-models obtained after each pruning, resulting in an encoded sequence corresponding to the multiple multimedia classification sub-models. This encoded sequence is used to record the number of channels in each layer of the multimedia classification sub-model and the feature matrix of each layer. This encoding method significantly reduces the length of the encoded sequence and saves storage space. For example, if a certain layer of a multimedia classification sub-model has 1024 channels, encoding each channel would require 1024 bits, while encoding the number of channels would require 10 bits.
[0097] This disclosure describes a cyclical process for pruning the multimedia classification model. The multimedia classification sub-model output from the previous pruning operation is used as the input for the next pruning operation. To facilitate understanding of the pruning principle of the multimedia classification model pruning method provided in this disclosure, this disclosure uses the (I+1)th pruning based on the multimedia classification sub-model obtained from the I-th pruning as an example. Here, I is a positive integer, and its value is greater than or equal to 2. After obtaining multiple multimedia classification sub-models obtained from the I-th pruning, the server encodes the number of channels in each layer of the multiple multimedia classification sub-models obtained from the I-th pruning, obtaining the encoding sequence corresponding to the multiple multimedia classification sub-models obtained from the I-th pruning.
[0098] 202. The server obtains the feature matrix of each layer of the multiple multimedia sub-models obtained from the I-th pruning from the encoding sequences corresponding to the multiple multimedia sub-models obtained from the I-th pruning.
[0099] Since the encoding sequence of this embodiment records the number of channels and the feature matrix of each layer of the multimedia classification sub-model, the server can obtain the feature matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning from the encoding sequence corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning. Since each layer of the multimedia classification sub-model includes input channels and output channels, and the feature matrix corresponds to the channels, the feature matrix of each layer includes the feature matrix corresponding to the input channels and the feature matrix corresponding to the output channels.
[0100] 203. Based on the feature matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, the server determines the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning.
[0101] Considering that pruning the multimedia classification model or its sub-models inevitably leads to a collapse in network accuracy, a genetic algorithm is needed to "fine-tune" the pruned sub-models to restore accuracy. This fine-tuning process includes using a genetic algorithm to screen the pruned multimedia sub-models (step 204 below) and performing crossover and mutation on the screened sub-models (step 205 below). Before screening the pruned sub-models, the weight matrix of each layer needs to be determined so that the loss of accuracy is limited to the model structure itself. Then, the screening is performed based on the weight matrix of each layer. The genetic algorithm can be NSGA-III (third-generation dominant genetic algorithm), etc. Compared to a "naive" evolutionary algorithm, NSGA-III has two advantages: first, it introduces a hyperplane reference point to ensure population diversity during the evolutionary process, which is crucial for searching for the optimal "network structure"; second, based on Pareto Front, multiple "optimal" sub-network models can be sampled after a single evolution.
[0102] For the multiple multimedia classification sub-models obtained from the I-th pruning, the server can substitute the feature matrices corresponding to the input channels and the feature matrices corresponding to the output channels of each layer of the multiple multimedia classification sub-models into the following formula to determine the weight matrix of each layer of the multiple multimedia classification sub-models:
[0103] W * =(X T X) -1 XY
[0104] Among them, W * Let X represent the weight matrix of any layer of any multimedia classifier sub-model obtained by the I-th pruning, let X represent the feature matrix corresponding to the input channel of the multimedia classifier sub-model layer, and let Y represent the feature matrix corresponding to the output channel of the multimedia classifier sub-model layer.
[0105] 204. Based on the weight matrices of each layer of the multiple multimedia sub-models obtained from the I-th pruning and the weight matrices of each layer of the multiple multimedia sub-models obtained from the (I-1)th pruning, the server filters the multiple multimedia sub-models obtained from the I-th and (I-1)th pruning to obtain the filtered multiple multimedia sub-models.
[0106] The server uses the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning and the (I-1)-th pruning to filter the multiple multimedia classification sub-models obtained from the I-th and (I-1)-th pruning. The following method can be used to obtain the filtered multiple multimedia classification sub-models:
[0107] 2041. The server calls multiple multimedia sub-models obtained from the I-th pruning with weight matrices and multiple multimedia sub-models obtained from the (I-1)-th pruning to process the multimedia test resource samples, and obtains the classification results of the multimedia resource test samples under each multimedia sub-model obtained from the I-th and (I-1)-th pruning.
[0108] Based on the weight matrices of each layer of the multiple multimedia sub-models obtained from the I-th pruning, the server assigns the weight matrices of each layer of each multimedia sub-model obtained from the I-th pruning to each layer of each multimedia sub-model, resulting in multiple multimedia sub-models obtained from the I-th pruning with weight matrices. The server calls the multiple multimedia sub-models obtained from the I-th pruning with weight matrices to process the multimedia test resource samples, obtaining the classification results of the multimedia resource test samples under the multiple multimedia sub-models obtained from the I-th pruning. The server calls the multiple multimedia sub-models obtained from the (I-1)-th pruning with weight matrices to process the multimedia test resource samples, obtaining the classification results of the multimedia resource test samples under the multiple multimedia sub-models obtained from the (I-1)-th pruning. Of course, the server can also choose not to call the multiple multimedia sub-models obtained from the (I-1)-th pruning with weight matrices to process the multimedia test resource samples, but instead directly obtain the classification results of the multimedia resource test samples under the multiple multimedia sub-models obtained from the (I-1)-th pruning in the previous evolution process.
[0109] 2042. Based on the classification and annotation results of each multimedia sub-model obtained by the first and (I-1)th pruning of multimedia resource test samples, the server determines the classification accuracy of each multimedia sub-model obtained by the first and (I-1)th pruning.
[0110] The server evaluates the classification accuracy of multiple multimedia classification sub-models obtained in the I-th pruning based on the classification results and annotation results of the multimedia resource test samples. For example, for any multimedia classification sub-model obtained in the I-th pruning, the server can obtain the correct classification result of the multimedia resource test samples in that sub-model based on the annotation results of the multimedia resource test samples, and divide the number of correct classification results by the total number of classification results to obtain the classification accuracy of that sub-model. Similarly, the server evaluates the classification accuracy of multiple multimedia classification sub-models obtained in the (I-1)-th pruning based on the classification results and annotation results of the multimedia resource test samples. For example, for any multimedia classification sub-model obtained from the (I-1)th pruning, the server can obtain the correct classification result of the multimedia resource test samples in that multimedia classification sub-model based on the annotation results of the multimedia resource test samples, and divide the number of correct classification results by the total number of classification results to obtain the classification accuracy of that multimedia sub-model. Of course, the server can also directly obtain the classification accuracy of each multimedia classification sub-model obtained from the (I-1)th pruning in the previous evolution process.
[0111] 2043. Based on the classification accuracy of each multimedia sub-model obtained from the I-th and I-1-th pruning, the server filters the multiple multimedia sub-models obtained from the I-th and I-1-th pruning to obtain multiple filtered multimedia sub-models.
[0112] Based on the classification accuracy of each multimedia sub-model obtained from the I-th and (I-1)-th pruning processes, the server determines the distribution of multimedia sub-models with different classification accuracies. Then, according to this distribution, it selects multiple multimedia sub-models with different classification accuracies from the multiple models obtained from the I-th and (I-1)-th pruning processes, resulting in P selected multimedia sub-models. When selecting the multiple multimedia sub-models obtained from the I-th and (I-1)-th pruning processes, the server does not select them in descending order of classification accuracy. Instead, it selects at least one multimedia sub-model from each classification accuracy range. This selection method avoids the selected multimedia sub-models having the same pruning rate, which could lead to a lack of diversity in the population.
[0113] 205. The server performs cross-mutation on the number of channels in each layer of the selected multimedia classification sub-models to obtain the processed multimedia classification sub-models.
[0114] When the server performs cross-mutation processing on the number of channels in each layer of multiple selected multimedia classification sub-models, the following method can be used:
[0115] 2051. For a subset of multimedia sub-models among the selected multimedia sub-models, the server cross-interchanges the number of channels in at least one layer of any two multimedia sub-models to obtain a cross-interchange subset of multimedia sub-models.
[0116] For any two multimedia classification sub-models in some multimedia classification sub-models, the number of layers and channels that are crossed are random. For example, the number of channels in the first layer of one multimedia classification sub-model can be exchanged with the number of channels in the second layer of another multimedia classification sub-model, thereby maximizing the diversity of the population.
[0117] It is important to note that after the multimedia classification sub-model performs crossover, the number of channels in each layer needs to be pruned based on the crossover result. To ensure that the pruning process can proceed normally and to avoid the number of channels before pruning being less than the number of channels after pruning, the layers with fewer channels need to be swapped with the layers with more channels during the crossover, and the layers with fewer channels should be replaced with the layers with more channels.
[0118] 2052. For the remaining multimedia sub-models among the multiple selected multimedia sub-models, the server mutates the number of channels in at least one layer of any two selected multimedia sub-models in the remaining multimedia sub-models to obtain the mutated remaining multimedia sub-models.
[0119] For any two multimedia classification sub-models in the remaining multimedia classification sub-models, the number of layers and channels to be mutated are random. For example, the number of channels in the first and second layers of the multimedia classification sub-model can be mutated to ensure the diversity of the population to the maximum extent.
[0120] It is important to note that after the multimedia classification sub-model undergoes mutation, the number of channels in each layer needs to be pruned based on the mutation results. To ensure that the pruning process can proceed normally, the number of channels before pruning should not be less than the number of channels after pruning.
[0121] 2053. The server combines the cross-interchanged portion of the multimedia classification sub-models with the mutated remaining multimedia classification sub-models to form multiple processed multimedia classification sub-models.
[0122] 206. Based on the number of channels in each layer of the processed multiple multimedia classification sub-models, the server performs a pruning operation on the processed multiple multimedia classification sub-models to obtain the multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0123] For each processed multimedia classification sub-model, the server prunes the number of output channels in the previous layer based on the number of input channels in the next layer of each processed multimedia classification sub-model, resulting in multiple multimedia classification sub-models obtained from the (I+1)th pruning. Furthermore, based on the pruning results of the input channels in each layer of the processed multiple multimedia classification sub-models, the server also performs pruning operations on the convolutional kernels of each layer.
[0124] It should be noted that the above explanation uses the (I+1)th pruning operation on the multimedia classification sub-model as an example. Since the value of I is greater than or equal to 2, the above pruning process is applicable to the second to the Tth pruning processes. For the first pruning process, the server obtains multiple multimedia classification sub-models obtained from the first pruning, and then encodes these multiple multimedia classification sub-models to obtain the encoding sequences corresponding to the multiple multimedia classification sub-models obtained from the first pruning. Then, without screening, a genetic algorithm is directly used to perform crossover and mutation processing on the multiple multimedia classification sub-models. Next, based on the number of channels in each layer of the processed multiple multimedia classification sub-models, a pruning operation is performed to obtain the multiple multimedia classification sub-models obtained from the second pruning. Subsequent pruning operations will be performed using the above method.
[0125] 207. The server performs the (I+1)th pruning operation on multiple multimedia sub-models obtained from the (I+1)th pruning, until the number of prunings reaches T. From the multiple multimedia sub-models obtained when the number of prunings reaches T, the optimal multimedia sub-model under different pruning rates is obtained.
[0126] Where T is the pre-set number of pruning steps, also the number of evolutions, and T is a positive integer, with a value greater than or equal to 1+2. When the number of pruning steps reaches T, the server uses the Pareto Front algorithm to process the multiple multimedia classifier sub-models obtained after T pruning steps to obtain the optimal multimedia classifier sub-model under different pruning rates. The Pareto Front algorithm is a multi-objective optimal solution algorithm capable of simultaneously solving for the optimal solutions of multiple objectives. This embodiment uses the NSGA-III genetic algorithm. By employing the Pareto Front algorithm for optimal solution calculation, it can obtain the optimal model under different pruning rates in one step. Compared to NetAdapt / AMC, which can only obtain one optimal model with a fixed pruning rate each time, this improves the efficiency of obtaining the optimal model under different pruning rates and shortens the overall pruning time. Furthermore, the genetic algorithm in this embodiment does not require setting complex hyperparameters, eliminating tedious training during the evolution process, while NetAdapt requires training sub-networks, and AMC requires training reinforcement learning networks.
[0127] Figure 3 A schematic diagram of the pruning process of the multimedia classification model provided in this embodiment is shown. This pruning process can be implemented using the following code:
[0128] Input: Initial network N, pre-trained weight matrix W, population size P, number of mutations M, number of crossovers S, number of evolutions T
[0129] Output: K optimal subnetworks G k
[0130] step:
[0131] 1: G0 = Random(N, P)
[0132] 2: for i = 1: T do
[0133] 3:G metric =Infer(ReParam(G i-1 W)
[0134] 4: G i =NSGA-III.NextGen(G metric M, S)
[0135] 5: end for
[0136] 6:G k =ParetoFront(G T K)
[0137] 7: return G k
[0138] The method provided in this disclosure has the following advantages over other automated pruning methods:
[0139] First, the computational cost of the resulting multimedia classifier submodel after pruning is similar to that of other automated pruning methods, and less than that of other genetic algorithms. See Table 1 for details.
[0140] Table 1
[0141]
[0142] Secondly, existing technologies are typically designed only for convolutional neural networks and have not been improved or validated for attention models (Transformers). Therefore, they cannot simultaneously adapt to both convolutional neural networks and visual attention models. This disclosure is the first method to apply a genetic pruning algorithm to a visual attention model, and the method provided in this disclosure has performance comparable to current, more time-consuming algorithms. See Table 2 for details, where FLOPs (Floating-point Operations Per second) represents the number of floating-point operations performed per second; TOP-1 represents the accuracy of the top-ranked category matching the actual result; and Epochs refers to the number of forward computations and backward propagations completed by all data input into the network.
[0143] Table 2
[0144] Model FLOPs TOP-1 Epochs Scheme Dei-B 17.8G 81.8% 300 Scratch VTP 13.8G 81.3% 100 Finetune EAPruning 13.5G 81.3% 100 Finetune AutoFormer 11.0G 82.4% 500 Supernet EAPruning 11.0G 81.6% 500 Scratch
[0145] Thirdly, the embodiments of this disclosure can significantly improve model inference efficiency without changing the model's accuracy. The method provided in these embodiments achieves speedups of 1.37, 1.34, and 1.4 times on ResNet50, MobileNetV1, and DeiT-Base models, respectively. See Table 3 for details.
[0146] Table 3
[0147]
[0148] The method provided in this disclosure determines the weight matrix of each layer of the multimedia classification sub-model obtained after each pruning based on the feature matrices corresponding to the input and output channels of each layer. This eliminates the need for searching the encoding sequence, allowing for the selection of higher-precision multimedia classification sub-models for subsequent evolution within a limited timeframe, thus improving the accuracy of the final multimedia classification sub-model. Furthermore, this disclosure proposes an encoding method for each pruning multimedia classification sub-model. This method encodes the number of channels in each layer of the multimedia classification sub-model instead of individual channels. Compared to encoding each channel individually, this results in a shorter encoding sequence, reducing storage space. Moreover, the channel-based encoding method eliminates the focus on which channel was pruned in each layer during subsequent evolution, only the number of channels retained in each layer. Compared to existing pruning methods, this significantly shortens the pruning time and improves pruning efficiency.
[0149] See Figure 4 This disclosure provides a pruning device for a multimedia classification model, the device comprising:
[0150] The determination module 401 is used to determine the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning based on the feature matrix of each layer. Here, I is a positive integer and the value of I is greater than or equal to 2.
[0151] The filtering module 402 is used to filter the multiple multimedia sub-models obtained by the first and second pruning based on the weight matrix of each layer of the multiple multimedia sub-models obtained by the first pruning and the weight matrix of each layer of the multiple multimedia sub-models obtained by the (I-1)th pruning, so as to obtain multiple multimedia sub-models after filtering.
[0152] Processing module 403 is used to perform cross-mutation processing on the number of channels in each layer of the multiple multimedia classification sub-models after screening, so as to obtain multiple multimedia classification sub-models after processing.
[0153] The pruning module 404 is used to perform pruning operations on the processed multiple multimedia classification sub-models based on the number of channels in each layer of the processed multiple multimedia classification sub-models, to obtain the multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0154] The acquisition module 405 is used to perform the I+2 pruning operation on multiple multimedia classification sub-models obtained by the I+1th pruning, until the number of prunings reaches T. From the multiple multimedia classification sub-models obtained when the number of prunings reaches T, the optimal multimedia classification sub-model under different pruning rates is obtained, where T is a positive integer and the value of T is greater than or equal to I+2.
[0155] In another embodiment of this disclosure, the device further includes:
[0156] The encoding module is used to encode the number of channels in each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, and to obtain the encoding sequence corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning. The encoding sequence is used to record the number of channels and feature matrix of each layer of the multimedia classification sub-model.
[0157] The acquisition module 405 is also used to acquire the feature matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning from the encoding sequences corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning.
[0158] In another embodiment of this disclosure, the determining module 401 is further configured to determine the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, based on the feature matrices corresponding to the input channels and the feature matrices corresponding to the output channels of each layer, using the following formula:
[0159] W * =(X T X) -1 XY
[0160] Among them, W * Let X represent the weight matrix of any layer of any multimedia classifier sub-model obtained by the I-th pruning, let X represent the feature matrix corresponding to the input channel of the multimedia classifier sub-model layer, and let Y represent the feature matrix corresponding to the output channel of the multimedia classifier sub-model layer.
[0161] In another embodiment of this disclosure, the filtering module 402 is used to call multiple multimedia classification sub-models obtained by the first pruning with weight matrices and multiple multimedia classification sub-models obtained by the (I-1)th pruning to process multimedia test resource samples, thereby obtaining the classification results of multimedia resource test samples under each multimedia classification sub-model obtained by the first and (I-1)th pruning; based on the classification results and annotation results of multimedia resource test samples under each multimedia classification sub-model obtained by the first and (I-1)th pruning, the classification accuracy of each multimedia classification sub-model obtained by the first and (I-1)th pruning is determined; based on the classification accuracy of each multimedia classification sub-model obtained by the first and (I-1)th pruning, the multiple multimedia classification sub-models obtained by the first and (I-1)th pruning are filtered to obtain multiple filtered multimedia classification sub-models.
[0162] In another embodiment of this disclosure, the processing module 403 is configured to: cross-interchange the number of channels in at least one layer of any two multimedia sub-models in the selected multimedia sub-models to obtain a cross-interchangeable partial multimedia sub-model; mutate the number of channels in at least one layer of any two selected multimedia sub-models in the remaining multimedia sub-models to obtain a mutated remaining multimedia sub-model; and combine the cross-interchangeable partial multimedia sub-models and the mutated remaining multimedia sub-models to form the processed multiple multimedia sub-models.
[0163] In another embodiment of this disclosure, the pruning module 404 is used to prune the number of output channels of the previous layer based on the number of input channels of the next layer of each processed multimedia classification sub-model, so as to obtain multiple multimedia classification sub-models obtained by the (I+1)th pruning.
[0164] In another embodiment of this disclosure, the acquisition module 405 is used to process multiple multimedia sub-models obtained when the number of prunings reaches T times using the Pareto front algorithm to obtain the optimal multimedia sub-model under different pruning rates.
[0165] In summary, the apparatus provided in this disclosure can determine the weight matrix of each layer of the multimedia classification sub-model obtained by each pruning based on the feature matrices corresponding to the input channels and output channels of each layer, without needing to search the encoding sequence. This allows for the selection of high-precision multimedia classification sub-models for subsequent evolution within a limited time, improving the accuracy of the final multimedia classification sub-model. Furthermore, this disclosure proposes an encoding method for each pruning multimedia classification sub-model. This method encodes the number of channels in each layer of the multimedia classification sub-model instead of individual channels in each layer. Compared to encoding each channel individually, this results in a shorter encoding sequence, reducing storage space. Moreover, the channel-based encoding method eliminates the focus on which channel was pruned in each layer of the intermediate multimedia classification sub-model during subsequent evolution, focusing only on the number of channels retained in each layer. Compared to existing pruning methods, this significantly shortens the pruning time and improves pruning efficiency.
[0166] Figure 5 This is a server for pruning a multimedia classification model, as illustrated in an exemplary embodiment. (See also...) Figure 5 Server 500 includes processing component 522, which further includes one or more processors, and memory resources represented by memory 532 for storing instructions, such as applications, that can be executed by processing component 522. The applications stored in memory 532 may include one or more modules, each corresponding to a set of instructions. Furthermore, processing component 522 is configured to execute instructions to perform the functions performed by the server in the pruning method of the aforementioned multimedia classification model.
[0167] Server 500 may also include a power supply component 526 configured to perform power management of server 500, a wired or wireless network interface 550 configured to connect server 500 to a network, and an input / output (I / O) interface 558. Server 500 can operate on an operating system, such as Windows Server, stored in memory 532. TM Mac OSX TM Unix TM Linux TM FreeBSD TM Or similar.
[0168] This disclosure provides a computer-readable storage medium storing at least one line of program code, which is loaded and executed by a processor to implement a pruning method for a multimedia classification model. The computer-readable storage medium can be non-transitory. For example, it can be a ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, or optical data storage device.
[0169] This disclosure provides a computer program product including computer program code stored in a computer-readable storage medium. A server's processor reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the server to perform a pruning method for a multimedia classification model.
[0170] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0171] The above description is merely an optional embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A pruning method for a multimedia classification model, characterized in that, The method includes: Based on the feature matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning is determined, wherein I is a positive integer and the value of I is greater than or equal to 2. Based on the weight matrix of each layer of the multiple multimedia sub-models obtained by the first pruning and the weight matrix of each layer of the multiple multimedia sub-models obtained by the (I-1)th pruning, the multiple multimedia sub-models obtained by the first and (I-1)th pruning are filtered to obtain multiple multimedia sub-models after filtering. The number of channels in each layer of the selected multimedia classification sub-models is cross-mutated to obtain the processed multimedia classification sub-models. Based on the number of channels in each layer of the processed multiple multimedia classification sub-models, a pruning operation is performed on the processed multiple multimedia classification sub-models to obtain the multiple multimedia classification sub-models obtained by the (I+1)th pruning. Based on the multiple multimedia classification sub-models obtained from the (I+1)th pruning, a (I+2)th pruning operation is performed until the number of pruning operations reaches T. From the multiple multimedia classification sub-models obtained after the number of pruning operations reaches T, the optimal multimedia classification sub-model under different pruning rates is obtained, where T is a positive integer and the value of T is greater than or equal to I+2. Based on the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning and the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the (I-1)th pruning, the multiple multimedia classification sub-models obtained from the I-th and (I-1)th pruning are filtered to obtain multiple filtered multimedia classification sub-models, including: Multiple multimedia classification sub-models obtained from the I-th pruning with weight matrices and multiple multimedia classification sub-models obtained from the (I-1)-th pruning are called to process the multimedia test resource samples, and the classification results of the multimedia resource test samples under each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning are obtained. Based on the classification and annotation results of each multimedia classification sub-model obtained by the first and (I-1)th pruning of the multimedia resource test samples, the classification accuracy of each multimedia classification sub-model obtained by the first and (I-1)th pruning is determined. Based on the classification accuracy of each multimedia sub-model obtained from the first and (I-1)th pruning, the multiple multimedia sub-models obtained from the first and (I-1)th pruning are filtered to obtain multiple filtered multimedia sub-models.
2. The method according to claim 1, characterized in that, Before determining the weight matrix of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning, the method further includes: Encode the number of channels in each layer of the multiple multimedia classification sub-models obtained by the I-th pruning to obtain the encoding sequence corresponding to the multiple multimedia classification sub-models obtained by the I-th pruning. The encoding sequence is used to record the number of channels and feature matrix of each layer of the multimedia classification sub-model. From the encoding sequences corresponding to the multiple multimedia sub-models obtained by the I-th pruning, obtain the feature matrix of each layer of the multiple multimedia sub-models obtained by the I-th pruning.
3. The method according to claim 1, characterized in that, The feature matrix of each layer of the multiple multimedia classification sub-models obtained based on the I-th pruning is used to determine the weight matrix of each layer of the multiple multimedia classification sub-models obtained based on the I-th pruning, including: Based on the feature matrices corresponding to the input channels and output channels of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning, the weight matrix of each layer of the multiple multimedia classification sub-models obtained from the I-th pruning is determined using the following formula: W*=(XTX)-1XY where W* represents the weight matrix of any layer of any multimedia sub-model obtained by the I-th pruning, X represents the feature matrix corresponding to the input channel of the layer of the multimedia sub-model, and Y represents the feature matrix corresponding to the output channel of the layer of the multimedia sub-model.
4. The method according to claim 1, characterized in that, The method of performing cross-mutation processing on the number of channels in each layer of the selected multiple multimedia classification sub-models to obtain the processed multiple multimedia classification sub-models includes: for some multimedia classification sub-models in the selected multiple multimedia classification sub-models, cross-exchanging the number of channels in at least one layer of any two multimedia classification sub-models in the selected multiple multimedia classification sub-models to obtain the cross-exchanged partial multimedia classification sub-models. For the remaining multimedia classification sub-models among the multiple filtered multimedia classification sub-models, the number of channels in at least one layer of any two filtered multimedia classification sub-models among the remaining multimedia classification sub-models is mutated to obtain the mutated remaining multimedia classification sub-models. The processed multimedia classification sub-models are composed of the partially cross-interchanged multimedia classification sub-models and the mutated remaining multimedia classification sub-models.
5. The method according to claim 1, characterized in that, The process involves pruning the processed multimedia classification sub-models based on the number of channels in each layer, resulting in the (I+1)th pruning of the resulting multimedia classification sub-models, including: For each processed multimedia classification sub-model, based on the number of input channels in the next layer of each processed multimedia classification sub-model, the number of output channels in the previous layer is pruned to obtain multiple multimedia classification sub-models obtained by the (I+1)th pruning.
6. The method according to claim 1, characterized in that, The process of obtaining the optimal multimedia classification sub-model under different pruning rates from multiple multimedia classification sub-models obtained when the number of prunings reaches T times includes: The Pareto front algorithm was used to process multiple multimedia sub-models obtained after pruning T times to obtain the optimal multimedia sub-model under different pruning rates.
7. A pruning device for a multimedia classification model, characterized in that, The device includes: The determination module is used to determine the weight matrix of each layer of the multiple multimedia classification sub-models obtained by the I-th pruning, based on the feature matrix of each layer. Here, I is a positive integer and the value of I is greater than or equal to 2. The filtering module is used to filter the multiple multimedia classification sub-models obtained from the first and (I-1)th pruning processes based on the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the first pruning process and the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the (I-1)th pruning process, to obtain a filtered set of multiple multimedia classification sub-models. The filtering process based on the weight matrices of each layer of the multiple multimedia classification sub-models obtained from the first and (I-1)th pruning processes, to obtain the filtered set of multiple multimedia classification sub-models, includes: Multiple multimedia classification sub-models obtained from the I-th pruning with weight matrices and multiple multimedia classification sub-models obtained from the (I-1)-th pruning are called to process the multimedia test resource samples, and the classification results of the multimedia resource test samples under each multimedia classification sub-model obtained from the I-th and (I-1)-th pruning are obtained. Based on the classification and annotation results of each multimedia classification sub-model obtained by the first and (I-1)th pruning of the multimedia resource test samples, the classification accuracy of each multimedia classification sub-model obtained by the first and (I-1)th pruning is determined. Based on the classification accuracy of each multimedia sub-model obtained from the first and (I-1)th pruning, the multiple multimedia sub-models obtained from the first and (I-1)th pruning are filtered to obtain multiple filtered multimedia sub-models. The processing module is used to perform cross-mutation processing on the number of channels in each layer of the multiple multimedia classification sub-models after screening, so as to obtain multiple processed multimedia classification sub-models. The pruning module is used to prune the processed multimedia sub-models based on the number of channels in each layer, obtaining the (I+1)th pruning sub-model. The acquisition module is used to perform the (I+2)th pruning operation based on the (I+1)th pruning sub-model, until the number of prunings reaches T. From the multimedia sub-models obtained after the number of prunings reaches T, the optimal multimedia sub-model under different pruning rates is acquired, where T is a positive integer and the value of T is greater than or equal to I+2.
8. A server, characterized in that, The server includes a processor and a memory, the memory storing at least one piece of program code, which is loaded and executed by the processor to implement the pruning method of the multimedia classification model as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the pruning method of the multimedia classification model as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes computer program code stored in a computer-readable storage medium. The server's processor reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the server to perform the pruning method of the multimedia classification model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic pruning method and platform for convolutional neural network general compression architecture
CN112396181A