Model pruning method, device and computer equipment

Through the combination of model pruning method and plug-in output module, the problem of excessive resource occupancy of deep learning models is solved, and more efficient resource utilization and parallel operation efficiency are achieved.

CN114819140BActive Publication Date: 2025-05-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210330396.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-05-13
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

The deep learning model is large in scale and occupies high amount of storage and computing resources, making it difficult to efficiently apply to various hardware devices.

Method used

Through the model pruning method, the effective state of the pruning object is determined using the mask information, and the parameter information is iteratively optimized until the end condition is met, and pruning is performed. In addition, add an external output module to the target model, connect it to the unit module, and prune it according to the performance indicators of the external output module.

Benefits of technology

Improve the pruning effect, such as improving the pruning accuracy, reducing the consumption of storage resources and computing resources, and pruning the target model in the depth dimension, improving parallel operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114819140B_ABST
    Figure CN114819140B_ABST
Patent Text Reader

Abstract

The embodiments of this specification disclose a model pruning method, device and computer equipment. The method includes: determining mask information according to pruning parameters; the mask information is used to indicate the effective state of the pruned object in the target model; inputting the sample into the target model after adding the mask information to obtain the first output of the target model; optimizing the parameter information according to the first output; the parameter information includes the model parameters and pruning parameters of the target model; iterating the above steps until the end condition is met; pruning the pruned object according to the mask information. The embodiments of this specification can prune the target model to reduce the occupation of storage resources and computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a model pruning method, device and computer equipment. Background Art

[0002] In recent years, deep learning has achieved great success in the application of artificial intelligence, including computer vision, speech recognition, natural language processing, etc. However, deep learning models are usually large in scale and require high storage and computing resources, making it difficult to efficiently apply deep learning models in various hardware devices.

[0003] Therefore, it is necessary to prune the deep learning model to reduce the usage of storage and computing resources. Summary of the invention

[0004] The embodiments of this specification provide a model pruning method, device and computer equipment to reduce the occupation of storage resources and computing resources. The technical solution of the embodiments of this specification is as follows.

[0005] A first aspect of an embodiment of this specification provides a model pruning method, including:

[0006] Determine mask information according to the pruning parameters; the mask information is used to indicate the valid state of the pruned object in the target model;

[0007] Input the sample into the target model with added mask information to obtain the first output of the target model;

[0008] Optimizing parameter information according to the first output; the parameter information includes model parameters and pruning parameters of the target model;

[0009] Iterate the above steps until the end condition is met;

[0010] According to the mask information, the pruned object is pruned.

[0011] A second aspect of the embodiments of this specification provides a model pruning method, including:

[0012] Input the sample into the target model after adding the external output module, and obtain the first output of the target model and the second output of the external output module; the target model includes a plurality of unit modules with the same structure stacked in sequence, and the external output module is connected to the unit module;

[0013] Optimizing parameter information according to the first output and the second output, the parameter information including model parameters of the target model and model parameters of the plug-in output module;

[0014] Iterate the above steps until the end condition is met;

[0015] According to the performance indicators of the plug-in output module, multiple unit modules are pruned.

[0016] A third aspect of the embodiments of this specification provides a model pruning device, including:

[0017] An iterative unit, configured to iteratively execute the following steps until an end condition is met: determining mask information according to pruning parameters; the mask information is used to indicate a valid state of a pruned object in a target model; inputting a sample into the target model to which the mask information is added, to obtain a first output of the target model; optimizing parameter information according to the first output; the parameter information includes model parameters and pruning parameters of the target model;

[0018] The pruning unit is used to prune the pruning object according to the mask information.

[0019] A fourth aspect of the embodiments of this specification provides a model pruning device, including:

[0020] The iterative unit is used to iteratively execute the following steps until an end condition is met: inputting a sample into a target model after adding an external output module to obtain a first output of the target model and a second output of the external output module; the target model includes a plurality of unit modules with the same structure stacked in sequence, and the external output module is connected to the unit modules; according to the first output and the second output, optimizing parameter information, the parameter information including model parameters of the target model and model parameters of the external output module;

[0021] The pruning unit is used to prune multiple unit modules according to the performance indicators of the plug-in output module.

[0022] According to a fifth aspect of the embodiments of this specification, a computer device is provided, including:

[0023] at least one processor;

[0024] A memory storing program instructions, wherein the program instructions are configured to be suitable for being executed by the at least one processor, and the program instructions include instructions for executing the method as described in the first aspect or the second aspect.

[0025] In the technical solution provided by the embodiment of this specification, the mask information can be determined based on the pruning parameters, and the pruning parameters can be adaptively adjusted during the learning process of the target model. In this way, pruning the pruned object according to the mask information can improve the pruning effect, for example, it can improve the pruning accuracy, thereby reducing the occupation of storage resources and computing resources. In addition, the technical solution provided by the embodiment of this specification adds an external output module to the target model, and the external output module can be trained together with the target model. The performance indicators of the external output module are used to prune multiple unit modules in the target model. Thereby, the target model is pruned in the depth dimension, reducing the occupation of storage resources and computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0027] Figure 1 A schematic diagram of pruning a fully connected neural network in an embodiment of this specification;

[0028] Figure 2 This is a schematic diagram of the structure of the visual converter in the embodiment of this specification;

[0029] Figure 3 Schematic diagram of the model pruning method in the embodiment of this specification;

[0030] Figure 4 is a schematic diagram of first mask information of a linear layer in an embodiment of this specification;

[0031] Figure 5 Schematic diagram of the second mask information of the multi-head self-attention layer in the embodiment of this specification;

[0032] Figure 6 A schematic diagram of pruning a target model in the depth dimension in an embodiment of this specification;

[0033] Figure 7 Schematic diagram of the model pruning method in the embodiment of this specification;

[0034] Figure 8 This is a schematic diagram of the structure of the model pruning device in the embodiment of this specification;

[0035] Fig. 9 This is a schematic diagram of the structure of the model pruning device in the embodiment of this specification;

[0036] Fig.10 Schematic diagram of the structure of a computer device in an embodiment of this specification. DETAILED DESCRIPTION

[0037] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0038] Model pruning can include structured pruning and unstructured pruning. By pruning the model, model redundancy can be reduced to obtain a lightweight model. The lightweight model is smaller in size and occupies less storage resources and computing resources.

[0039] For example, Figure 1 Schematic diagram of pruning a fully connected neural network. Figure 1 As shown, the fully connected neural network may include an L1 layer, an L2 layer, an L3 layer, and an L4 layer. Pruning the fully connected neural network may include: deleting a neuron in the L2 layer, deleting a neuron in the L3 layer, deleting 14 connection edges between the L1 layer and the L2 layer, deleting 8 connection edges between the L2 layer and the L3 layer, and deleting 4 connection edges between the L3 and L4 layers. The connection edges are used to represent weight parameters.

[0040] In the related art, pruning parameters (such as pruning thresholds) can be preset, and the model can be pruned according to the preset pruning parameters. Taking a convolutional neural network model as an example, the convolutional neural network model may include a batch normalization layer (BN). The batch normalization layer has a scaling factor γ (a learnable model parameter). An L1 norm penalty can be applied to the scaling factor γ, and the convolutional neural network model after the L1 norm penalty is applied can be subjected to sparse training. The loss function used in the sparse training can be expressed as L=l(f(x), y)+λ∑|Υ|. l(f(x), y) represents the loss function used in normal training, x represents a sample, f(x) represents the prediction result of the convolutional neural network model, y represents the sample label of the sample, λ represents the sparse coefficient, and ∑|Υ| represents the sum of the absolute values ​​of the scaling factors γ of each batch normalization layer (i.e., the L1 norm of the scaling factor γ). After the sparse training is completed, a scaling factor γ with a smaller value can be selected according to the pruning parameters; the selected scaling factor γ can be used for pruning. For example, the convolutional layer channels corresponding to the smaller scaling factor γ are deleted.

[0041] In the above-mentioned related technologies, pruning parameters need to be pre-set based on manual experience. Such pruning parameters set based on manual experience may not match the actual situation of the model, thereby affecting the pruning effect (eg, low pruning accuracy).

[0042] In addition, the model may include a plurality of unit modules with the same structure stacked in sequence. The depth of the model may refer to the number of the plurality of unit modules. The width of the model may refer to the number of processing units in the unit module. For example, the model may include a convolutional neural network model. The convolutional neural network model may include a plurality of intermediate layers with the same structure stacked in sequence. The intermediate layer may include a convolutional layer, a linear layer, etc. The depth of the convolutional neural network model may refer to the number of intermediate layers. The width of the convolutional neural network model may refer to the number of neurons in the intermediate layer. For another example, the model may include a vision transformer. The vision transformer may include a plurality of encoder blocks with the same structure stacked in sequence. The depth of the vision transformer may refer to the number of encoder blocks. The width of the vision transformer may refer to the number of self-attention heads in a multi-head self-attention layer. In the related art, the model is often pruned in the width dimension, and the model is rarely pruned in the depth dimension. As a result, it is impossible to reduce model redundancy in the depth dimension.

[0043] The embodiment of this specification provides a data processing system.

[0044] The data processing system may include a pruning device and a business device. The pruning device and the business device may be a personal computer, a server, or a server cluster including multiple servers, etc. The pruning device may be used to prune the target model to obtain a lightweight target model; the lightweight target model may be sent to the business device. The business device may receive the lightweight target model; the business data may be input into the lightweight target model to obtain a prediction result. The target model may include an image classification model, a text classification model, an audio classification model, etc. The business data may include image data, text data, audio data, etc. The prediction result may be used to indicate whether the business data is target business data. For example, the prediction result may be used to indicate whether the business data is abnormal business data. The abnormal business data may include business data involving illegal content such as fraud.

[0045] The target model may include a neural network model. The neural network model may include a fully connected neural network model, a neural network model based on a self-attention mechanism, etc. The neural network model based on a self-attention mechanism may include a transformer, etc. The transformer may include a vision transformer.

[0046] The visual converter can be applied to image classification scenarios. The image classification scenarios include but are not limited to commodity classification scenarios (such as recommendation of similar commodities), prohibited items detection scenarios (such as detection of prohibited items), face recognition scenarios (such as liveness detection), urban planning scenarios (such as road network extraction), meteorological scenarios (such as cloud extraction), auto insurance claims scenarios (determining the purpose of images, such as vehicle damage images, property damage images, and document images), etc. The visual converter may include basic visual transformers and their deformations. Please refer to Figure 2 The visual converter may include a vector conversion layer, a plurality of encoder blocks with the same structure stacked in sequence, and a classifier. The vector conversion layer is used to perform vector conversion to generate a vector representation (Embedding). In practical applications, the number of encoder blocks may be P. P may be 8, 9, 12, etc. The encoder block may be used to extract feature vectors. The classifier may be used for prediction.

[0047] Each encoding module may include a normalization layer (Norm), a multi-head self-attention layer (Multi-HeadAttention), and a multi-layer perceptron (MLP), etc. The normalization layer is used for normalization. The multi-head self-attention layer may include a linear layer (Linear, also known as a fully connected layer, hereinafter referred to as the first linear layer) and multiple self-attention heads. The first linear layer is used to summarize the outputs of multiple self-attention heads. In practical applications, the number of self-attention heads may be H. The H may be 3, 4, 6, etc. The multiple self-attention heads have the same structure. Each self-attention head may include a linear layer (Linear, hereinafter referred to as the second linear layer) and a self-attention layer (Self-Attention). The second linear layer is used to perform linear transformation. The self-attention layer is used to perform operations based on the attention mechanism.

[0048] In the image classification scenario, the image data can be segmented to obtain multiple image patches (Patch). The multiple image patches have the same height and width. The image patches can be flattened to obtain a data sequence of the image patches; the data sequence can be input into a vector conversion layer to obtain a vector representation of the image patches (PatchEmbedding). The vector representation of the image patches can be input into multiple encoding modules to obtain a feature vector output by the last encoding module; the feature vector output by the last encoding module can be input into a classifier to obtain the prediction result of the visual converter.

[0049] It is worth noting that in Figure 2 middle, Indicates fusion. For example, in the encoding module, the feature vector can be fused with the output of the multi-head self-attention layer, and the fusion result can be used as the input of the normalization layer.

[0050] The embodiment of this specification provides a model pruning method. The model pruning method can be applied to a pruning device. The model pruning method can be used to prune a target model to reduce model redundancy and obtain a lightweight target model. Figure 3 The model pruning method may include the following steps.

[0051] Step S11: Determine mask information according to pruning parameters.

[0052] In some embodiments, the mask information is used to indicate the effective state of the pruned object in the target model, and the effective state can be understood as the participation state when the target model is predicted, such as whether it participates in the prediction. The pruned object may include a linear layer and / or a multi-head self-attention layer. For ease of description, the mask information of the linear layer is referred to as the first mask information, and the mask information of the multi-head self-attention layer is referred to as the second mask information. The first mask information is used to indicate the effective state of the linear layer. The second mask information is used to indicate the effective state of the multi-head self-attention layer. In some scene examples, the target model may include a visual converter. Thus, the pruned object may include a linear layer in the self-attention head, such as a first linear layer and a second linear layer. Of course, the pruned object may also include other linear layers, such as a linear layer in a multilayer perceptron.

[0053] In some implementations of this embodiment, the weight parameter of the linear layer may include multiple sub-weight parameters. The first mask information may include multiple first sub-mask information. There is a corresponding relationship between the first sub-mask information and the sub-weight parameter. The first sub-mask information is used to indicate the valid state of the sub-weight parameter. For example, the value of the first sub-mask information may be 0 or 1. 0 is used to indicate that the sub-weight parameter is invalid, and 1 is used to indicate that the sub-weight parameter is valid. Of course, 0 or 1 here is only an example, and the first sub-mask information can also be other values. In some scene examples, the weight parameter can be expressed as a weight matrix. The elements in the weight matrix can represent the sub-weight parameters. The first mask information can be expressed as a mask matrix. The elements in the mask matrix can represent the first sub-mask information. The mask matrix and the weight matrix are matrices of the same order. The first sub-mask information with the same two-dimensional coordinates has a corresponding relationship with the sub-weight parameter. The two-dimensional coordinates may include the number of rows and the number of columns.

[0054] In some other implementations of this embodiment, the multi-head self-attention layer may include multiple self-attention heads. The second mask information may include multiple second sub-mask information. There is a corresponding relationship between the second sub-mask information and the self-attention head. The second sub-mask information is used to indicate the valid state of the self-attention head. For example, the value of the second sub-mask information can be 0 or 1. 0 is used to indicate that the self-attention head is invalid, and 1 is used to indicate that the self-attention head is valid. Of course, 0 or 1 here is only an example, and the second sub-mask information can also be other values. In some scenario examples, the number of self-attention heads and the number of second sub-mask information can both be H. Each second sub-mask information can correspond to a self-attention head.

[0055] In some embodiments, the pruning parameters include a pruning threshold and value information of the pruning object. A first pruning rate of the pruning object can be calculated based on the pruning threshold; and mask information of the pruning object can be determined based on the first pruning rate and the value information.

[0056] The pruning threshold may be a critical value for defining a valid state. The initial value of the pruning threshold may be an empirical value, such as 10, 16 or 20. The first pruning rate may be understood as the clipping ratio of the pruned object. The first pruning rate may be calculated using a Sigmoid function, a Tanh function or a ReLU function. For example, the first pruning rate may be calculated according to the formula r=σ(β). r represents the first pruning rate, β represents the pruning threshold, and σ represents the Sigmoid function.

[0057] The value information is used to measure the importance of the pruned object. The value information may include multiple sub-value information. The mask information may include multiple sub-mask information. There may be a corresponding relationship between the sub-value information and the sub-mask information. In this way, a comparison benchmark may be determined based on the first pruning rate and the multiple sub-value information; the sub-value information may be compared with the comparison benchmark; and the sub-mask information corresponding to the sub-value information may be assigned a value based on the comparison result. Specifically, if the sub-value information is less than the comparison benchmark, a first preset value (e.g., a value of 0) may be assigned to the corresponding sub-mask information. If the sub-value information is greater than or equal to the comparison benchmark, a second preset value (e.g., a value of 1) may be assigned to the corresponding sub-mask information.

[0058] For ease of description, the value information of the linear layer is referred to as the first value information, and the value information of the multi-head self-attention layer is referred to as the second value information. The first value information is used to measure the importance of the linear layer. The importance of the linear layer can be characterized by the effective state of the sub-weight parameter in the weight parameter. The second value information is used to measure the importance of the multi-head self-attention layer. The importance of the multi-head self-attention layer can be characterized by the effective state of the self-attention head in the multi-head self-attention layer.

[0059] In some implementations of this embodiment, the first mask information of the linear layer may be determined according to the first pruning rate and the first value information. Specifically, the first value information may include a plurality of first sub-value information. The first mask information may include a plurality of first sub-mask information. Each first sub-value information may correspond to a plurality of first sub-mask information. In this way, a comparison benchmark may be determined according to the first pruning rate and the plurality of first sub-value information; the first sub-value information may be compared with the comparison benchmark, and a value may be assigned to the first sub-mask information corresponding to the first sub-value information according to the comparison result.

[0060] The first pruning rate can be expressed as r, 0≤r≤1. The number of first sub-value information in the first value information can be m. The r×mth smallest first sub-value information can be selected as a comparison benchmark. For example, the m first sub-value information can be arranged in order from small to large; the r×mth first sub-value information can be selected as a comparison benchmark. Of course, in practice, it is not limited to this, and other methods can be used to determine the comparison benchmark.

[0061] If the first sub-value information is less than the comparison benchmark, a first preset value (e.g., a value of 0) may be assigned to the corresponding first sub-mask information, thereby indicating that the sub-weight parameter corresponding to the first sub-mask information is invalid. If the first sub-value information is greater than or equal to the comparison benchmark, a second preset value (e.g., a value of 1) may be assigned to the corresponding first sub-mask information, thereby indicating that the sub-weight parameter corresponding to the first sub-mask information is valid.

[0062] For some example scenarios, see Figure 4 . The weight parameters of the linear layer can be expressed as a weight matrix. The elements in the weight matrix can represent sub-weight parameters. The first mask information can be expressed as a mask matrix. The elements in the mask matrix can represent first sub-mask information. The number of rows of the mask matrix and the weight matrix is ​​m, and the number of columns is n. The first sub-mask information and the sub-weight parameters having the same two-dimensional coordinates have a corresponding relationship. The first value information can include m first sub-value information. Each first sub-value information can correspond to a row of first sub-mask information in the mask matrix.

[0063] If the first sub-value information is less than the comparison benchmark, the corresponding first sub-mask information may be assigned a value of 0, thereby indicating that the sub-weight parameter corresponding to the first sub-mask information is invalid. If the first sub-value information is greater than or equal to the comparison benchmark, the corresponding first sub-mask information may be assigned a value of 1, thereby indicating that the sub-weight parameter corresponding to the first sub-mask information is valid. Figure 4 The gray first sub-value information indicates the first sub-value information that is less than the comparison benchmark, and the white first sub-value information indicates the first sub-value information that is greater than or equal to the comparison benchmark. The gray first sub-mask information indicates the first sub-mask information with a value of 0, and the white first sub-mask information indicates the first sub-mask information with a value of 1. The gray sub-weight parameter indicates an invalid sub-weight parameter, and the white sub-weight parameter indicates a valid sub-weight parameter.

[0064] In some other implementations of this embodiment, the second mask information of the multi-head self-attention layer can be determined based on the first pruning rate and the second value information. Specifically, the second value information may include multiple second sub-value information. The second mask information may include multiple second sub-mask information. Each second sub-value information may correspond to one second sub-mask information. In this way, a comparison benchmark can be determined based on the first pruning rate and the multiple second sub-value information; the second sub-value information can be compared with the comparison benchmark, and the second sub-mask information corresponding to the second sub-value information can be assigned a value based on the comparison result.

[0065] The first pruning rate can be expressed as r, 0≤r≤1. The number of second sub-value information in the second value information can be H. The r×Hth smallest second sub-value information can be selected as a comparison benchmark. For example, the H second sub-value information can be arranged in order from small to large, and the r×Hth second sub-value information can be selected as a comparison benchmark. Of course, in practice, it is not limited to this, and other methods can be used to determine the comparison benchmark.

[0066] If the second sub-value information is less than the comparison benchmark, a first preset value (e.g., a value of 0) may be assigned to the corresponding second sub-mask information, thereby indicating that the self-attention head corresponding to the second sub-mask information is invalid. If the second sub-value information is greater than or equal to the comparison benchmark, a second preset value (e.g., a value of 1) may be assigned to the corresponding second sub-mask information, thereby indicating that the self-attention head corresponding to the second sub-mask information is valid.

[0067] For some example scenarios, see Figure 5 The multi-head self-attention layer may include H self-attention heads. The second mask information may include H second sub-mask information. Each second sub-mask information may correspond to one self-attention head. The second value information may include H second sub-value information. Each second sub-value information may correspond to one second sub-mask information.

[0068] If the second sub-value information is less than the comparison benchmark, the corresponding second sub-mask information may be assigned a value of 0, thereby indicating that the self-attention head corresponding to the second sub-mask information is invalid. If the second sub-value information is greater than or equal to the comparison benchmark, the corresponding second sub-mask information may be assigned a value of 1, thereby indicating that the self-attention head corresponding to the second sub-mask information is valid. Figure 5 As shown. The gray second sub-value information indicates the second sub-value information that is less than the comparison benchmark, and the white second sub-value information indicates the second sub-value information that is greater than or equal to the comparison benchmark. The gray second sub-mask information indicates the second sub-mask information with a value of 0, and the white second sub-mask information indicates the second sub-mask information with a value of 1. The gray self-attention head indicates an invalid self-attention head, and the white self-attention head indicates a valid self-attention head.

[0069] Step S13: Input the sample into the target model with the mask information added, and obtain the first output of the target model.

[0070] In some embodiments, the target model after adding mask information is: the target model after masking the pruned object using the mask information. Through masking, the participation status of the pruned object in the target model prediction can be set.

[0071] In some implementations of this embodiment, the first mask information may be used to perform mask processing on the linear layer.

[0072] For example, the first mask information can be represented as a mask matrix, and the weight parameters of the linear layer can be represented as a weight matrix. The Hadamard Product of the mask matrix and the weight matrix can be calculated. Thus, after the linear layer is masked, the output of the linear layer can contain m elements, of which the jth element can be represented as M represents the mask matrix. The number of rows in the mask matrix M is m, and the number of columns is n. j,k represents the element in the jth row and the kth column of the mask matrix M. W represents the weight matrix, and the number of rows and columns of the weight matrix W is m. j,k represents the jth row and kth column element in the weight matrix W. x represents the input. The x contains n elements. k Represents the kth element in x.

[0073] In some implementations of this embodiment, the second mask information can be used to perform mask processing on the multi-head self-attention layer.

[0074] For example, the second mask information may include multiple second sub-mask information, and the multi-head self-attention layer may include multiple self-attention heads. Each second sub-mask information may correspond to one self-attention head. The second sub-mask information may be multiplied by the output of the corresponding self-attention head. Thus, after masking the multi-layer self-attention heads, the output of the multi-head self-attention layer may be expressed as W proj represents the weight matrix of the first linear layer in the multi-head self-attention layer. H represents the number of self-attention heads. M h Indicates the second sub-mask information corresponding to the h-th self-attention head. h (x) represents the output of the h-th self-attention head. h (x) = α h W h,v x.α h represents the model parameters of the self-attention layer in the h-th self-attention head, W h,v represents the weight matrix of the second linear layer in the h-th self-attention head.

[0075] In some embodiments, the sample may include an image sample, a text sample, an audio text, etc. The sample has a sample label. The sample label may be used to indicate the category of the sample. The target model may include an image classification model, a text classification model, an audio classification model, etc. One or more samples may be input into the target model after adding mask information to obtain a first output of the target model. The first output may be a prediction result of the target model.

[0076] Step S15: Optimize parameter information according to the first output.

[0077] In some embodiments, parameter information can be optimized according to a loss function. The parameter information may include model parameters and pruning parameters of the target model. The loss function may include a first term and a second term. The first term is used to constrain the second pruning rate of the target model so that the target model as a whole approaches or reaches the expected pruning rate. The first term may include an augmented Lagrangian function, such as a two-stage augmented Lagrangian function. Of course, in practice, this is not limited to this, and other functions may be used to constrain the second pruning rate of the target model. Among them, the second pruning rate can be understood as the clipping ratio of the target model, which can be specifically calculated based on the first pruning rate. For example, the second pruning rate can be calculated according to the formula Calculated. L represents the number of pruned objects, r l represents the first pruning rate of the lth pruning object, n l represents the number of parameters of the lth pruned object, and N represents the number of parameters of the target model. The second term is used to represent the deviation between the first output and the sample label. The second term may include a cross-entropy loss function (Cross-Entropy Loss), a mean square error loss function (MSE), etc.

[0078] In some scenario examples, the loss function can be expressed as L = L CE +L p .L p represents the first term. The first term may be a two-stage augmented Lagrangian function. Wherein, λ1 and λ2 represent Lagrange multipliers. λ1 and λ2 can be fixed values. Alternatively, λ1 and λ2 can also be learnable parameters. The initial values ​​of λ1 and λ2 can be empirical values, such as 0, 1, 3, etc. R represents the second pruning rate of the target model, R t represents the expected pruning rate of the target model. Through the two-stage augmented Lagrangian function, the target model as a whole can be guided to approach or reach the expected pruning rate during the learning process. L CE represents the second term. The second term may be a cross entropy loss function.

[0079] In some embodiments, loss information can be calculated using a loss function based on the first output; parameter information can be optimized using a back propagation mechanism based on the loss information. For example, the gradient of the parameter information can be calculated using a back propagation mechanism; and the parameter information can be adjusted based on the gradient of the parameter information.

[0080] For example, the parameter information may include pruning parameters, and the pruning parameters may include pruning thresholds and value information. Compute the gradient of the pruning threshold. represents the gradient of the pruning threshold of the lth pruning object. According to the gradient of the pruning threshold, the larger the number of parameters of the pruning object, the more likely the pruning object is to be pruned. Compute the gradient of the first value information. It represents the gradient of the jth first sub-value information in the first value information. It can be calculated according to the formula Calculate the gradient of the second value information. Represents the gradient of the hth second sub-value information in the second value information.

[0081] Step S17: Pruning the pruned object according to the mask information.

[0082] In some embodiments, steps S11 to S15 may be iteratively executed until an end condition is met. The end condition may be flexibly set according to actual needs. For example, the end condition may be that the number of iterations reaches a preset threshold. For another example, the end condition may also be that the second pruning rate of the target model reaches an expected pruning rate.

[0083] After the iteration process is finished, the pruned objects in the target model can be pruned according to the current mask information. The mask information can be determined according to the pruning parameters. The pruning parameters can be adaptively adjusted during the learning process of the target model. In this way, the pruned objects are pruned according to the mask information, which can improve the pruning effect, for example, the pruning accuracy can be improved. In addition, the pruned objects are pruned according to the mask information, which realizes the pruning of the target model in the width dimension.

[0084] In some embodiments, the mask information may include first mask information. The first mask information may include a plurality of first sub-mask information. The pruning object may include a linear layer. The weight parameter of the linear layer may include a plurality of sub-weight parameters. The sub-weight parameter may be deleted according to the value of the first sub-mask information. For example, the first sub-mask information has a corresponding relationship with the sub-weight parameter. The sub-weight parameter corresponding to the first sub-mask information having a value of the first preset value may be deleted.

[0085] It is worth noting that in some scenario examples, the first mask information can be represented as a mask matrix. The elements in the mask matrix can represent the first sub-mask information. The first value information can include multiple first sub-value information. Each first sub-value information can correspond to a row of first sub-mask information in the mask matrix. So that, for each row of the mask matrix, the values ​​of each first sub-mask information in the row are equal. The weight parameters of the linear layer can be represented as a weight matrix. The elements in the weight matrix can represent sub-weight parameters. So that, by deleting the sub-weight parameters according to the values ​​of the first sub-mask information, the sub-weight parameters in the weight matrix can be deleted in units of rows, thereby reducing the number of rows of the weight matrix. Reducing the number of rows of the weight matrix means reducing the number of elements contained in the output of the linear layer. In other words, this scenario example can prune the output of the linear layer.

[0086] In some embodiments, the mask information may include second mask information. The second mask information may include a plurality of second sub-mask information. The pruning object may include a multi-head self-attention layer. The multi-head self-attention layer may include a plurality of self-attention heads. The self-attention head may be deleted according to the value of the second sub-mask information. For example, the second sub-mask information has a corresponding relationship with the self-attention head. The self-attention head corresponding to the second sub-mask information having a value of the first preset value may be deleted.

[0087] In some embodiments, an output module may be added to the target model. The added output module is referred to as a plug-in output module below. The plug-in output module may be used for prediction. Specifically, the plug-in output module may be a plug-in classification module, such as a plug-in classifier. The plug-in classification module may be used for classification prediction. Of course, the plug-in output model may also be a plug-in regression module. The plug-in regression module may be used for regression prediction.

[0088] The target model may include a plurality of unit modules with the same structure stacked in sequence. Among the plurality of unit modules with the same structure, the output of the previous unit module may be passed as input to the next unit module for processing by the next unit module. The plug-in output module may be connected to the unit module for making predictions based on the output of the unit module. The number of the plug-in output modules may be one or more. Each plug-in output module may be connected to a unit module. Each unit module may be connected to zero or one plug-in output module.

[0089] Of course, the target model itself may also include an output module. In order to distinguish it from an external output module, the output module in the target model is referred to as an internal output module below. The internal output module may be used for prediction, for example, prediction is made based on the output of the last unit module among the multiple unit modules with the same structure. Specifically, the internal output module may be an internal classification module, such as an internal classifier. The internal classification module may be used for classification prediction. Of course, the internal output model may also be an internal regression module. The internal regression module may be used for regression prediction.

[0090] For example, see Figure 6 . The target model may include a visual converter. The multiple modules with the same structure may include a coding module. Multiple plug-in classifiers may be added to the visual converter. Each plug-in classifier may be connected to a coding module. Specifically, for example, the visual converter may include P coding modules with the same structure stacked in sequence. The P may be 8, 9, 12, etc. P-1 plug-in classifiers may be added to the visual converter. The P-1 plug-in classifiers may be connected to the other P-1 coding modules except the last coding module. Each plug-in classifier may be connected to a coding module. The last coding module may be connected to the internal classifier of the visual converter.

[0091] In some embodiments, in step S13, the sample can be input into the target model after adding the mask information and the plug-in output module to obtain the first output of the target model and the second output of the plug-in output module. In step S15, the parameter information can be optimized according to the first output and the second output. After the iterative process is completed, the multiple unit modules can be pruned according to the performance indicators of the plug-in output module. In this way, the target model can be pruned in the width dimension and the depth dimension at the same time. Pruning the target model in the depth dimension can enable the target model to obtain greater parallel operation efficiency.

[0092] Specifically, the loss information can be calculated using a loss function according to the first output and the second output; the parameter information can be optimized using a back propagation mechanism according to the loss information. For example, the gradient of the parameter information can be calculated using a back propagation mechanism; the parameter information can be adjusted according to the gradient of the parameter information. The first output can be a prediction result of the target model. The second output can be a prediction result of the plug-in output module. The parameter information can include model parameters of the target model, model parameters of the plug-in output module, and pruning parameters.

[0093] The loss function may include a first item, a second item and a third item. The first item is used to constrain the second pruning rate of the target model so that the target model as a whole approaches or reaches the expected pruning rate. The second item is used to represent the deviation between the first output and the sample label. The third item is used to represent the deviation between the second output and the sample label. Specifically, the third item may include multiple sub-items. The third item may be the sum of the multiple sub-items. Each sub-item is used to represent the deviation between the second output of an external output module and the sample label. The sub-items may include a cross entropy loss function, a mean square error loss function, and the like.

[0094] In some scenario examples, the loss function can be expressed as L = L CE +L p +L C .L p Indicates the first item. L CE Indicates the second term. C Indicates the third item. P indicates the number of external output modules. ci represents the deviation between the second output of the i-th plug-in output module and the sample label. ci represents the cross entropy loss function.

[0095] Specifically, the performance indicators may include accuracy, recall, precision, F1-score and any combination thereof. The performance of the plug-in output module may be tested using verification data to obtain the performance indicators of the plug-in output module; the target unit module may be determined based on the performance indicators of the plug-in output module; and the unit module after the target unit module may be deleted. For example, the plug-in output module with the best performance may be selected as the target plug-in output module; the unit module connected to the target plug-in output module may be used as the target unit module; and the unit module after the target unit module may be deleted. For another example, the performance of the internal output module may be tested using verification data to obtain the performance indicators of the internal output module. In this way, the output module with the best performance may be selected as the target output module; the unit module connected to the target output module may be used as the target unit module; and the unit module after the target unit module may be deleted. Among them, the output module with the best performance may be an internal output module or a plug-in output module.

[0096] Furthermore, other output modules other than the target output module may be deleted, and the target output module may be used as the output module of the target model itself, and the output of the target output module may be used as the output of the target model. The target output module may be an internal output module or an external output module. Other output modules other than the target output module may include internal output modules and / or external output modules.

[0097] In the model pruning method of the embodiment of this specification, the mask information can be determined based on the pruning parameters, and the pruning parameters can be adaptively adjusted during the learning process of the target model. In this way, pruning the pruning object according to the mask information can improve the pruning effect, for example, the pruning accuracy can be improved.

[0098] The embodiment of this specification also provides another model pruning method. The model pruning method can be applied to a pruning device. The model pruning method is used to prune the target model to reduce model redundancy and obtain a lightweight target model.

[0099] See also Figure 7 The model pruning method may include the following steps.

[0100] Step S21: Input the sample into the target model with the external output module added, and obtain the first output of the target model and the second output of the external output module.

[0101] In some embodiments, an output module may be added to the target model. The added output module is referred to as a plug-in output module below. The plug-in output module may be used for prediction. Specifically, the plug-in output module may be a plug-in classification module, such as a plug-in classifier. The plug-in classification module may be used for classification prediction. Of course, the plug-in output model may also be a plug-in regression module. The plug-in regression module may be used for regression prediction.

[0102] The target model may include a plurality of unit modules with the same structure stacked in sequence. Among the plurality of unit modules with the same structure, the output of the previous unit module may be passed as input to the next unit module for processing by the next unit module. The plug-in output module may be connected to the unit module for making predictions based on the output of the unit module. The number of the plug-in output modules may be one or more. Each plug-in output module may be connected to a unit module. Each unit module may be connected to zero or one plug-in output module.

[0103] Of course, the target model itself may also include an output module. In order to distinguish it from an external output module, the output module in the target model is referred to as an internal output module below. The internal output module may be used for prediction, for example, prediction is made based on the output of the last unit module among the multiple unit modules with the same structure. Specifically, the internal output module may be an internal classification module, such as an internal classifier. The internal classification module may be used for classification prediction. Of course, the internal output model may also be an internal regression module. The internal regression module may be used for regression prediction.

[0104] In some embodiments, the sample may include an image sample, a text sample, an audio text, etc. The sample has a sample label. The sample label may be used to indicate the category of the sample. The target model may include an image classification model, a text classification model, an audio classification model, etc. One or more samples may be input into the target model with the plug-in output module added thereto to obtain a first output of the target model and a second output of the plug-in output module.

[0105] Step S23: Optimize parameter information according to the first output and the second output.

[0106] In some embodiments, loss information can be calculated using a loss function based on the first output and the second output; parameter information can be optimized using a back propagation mechanism based on the loss information. For example, the gradient of parameter information can be calculated using a back propagation mechanism; parameter information can be adjusted based on the gradient of parameter information. The first output can be a prediction result of a target model. The second output can be a prediction result of an external output module. The parameter information can include model parameters of the target model and model parameters of the external output module.

[0107] The loss function may include a second item and a third item. The second item is used to represent the deviation between the first output and the sample label. The third item is used to represent the deviation between the second output and the sample label. Specifically, the third item may include multiple sub-items. The third item may be the sum of the multiple sub-items. Each sub-item is used to represent the deviation between the second output of an external output module and the sample label. The sub-items may include a cross entropy loss function, a mean square error loss function, etc.

[0108] In some scenario examples, the loss function can be expressed as L = L p +L C .L CE Indicates the second term. C Indicates the third item. P indicates the number of external output modules. ci represents the deviation between the second output of the i-th plug-in output module and the sample label. ci represents the cross entropy loss function.

[0109] Step S25: Pruning the multiple unit modules according to the performance indicators of the plug-in output module.

[0110] In some embodiments, steps S21 to S25 may be iteratively executed until an end condition is met. The end condition may be flexibly set according to actual needs. For example, the end condition may be that the number of iterations reaches a preset threshold. After the iterative process is completed, multiple unit modules may be pruned according to the performance indicators of the plug-in output module.

[0111] In some embodiments, the performance indicators may include accuracy, recall, precision, F1-score, and any combination thereof. The performance of the plug-in output module can be tested using verification data to obtain the performance indicators of the plug-in output module; the target unit module can be determined based on the performance indicators of the plug-in output module; and the unit module after the target unit module can be deleted. For example, the plug-in output module with the best performance can be selected as the target output module; the unit module connected to the target output module can be used as the target unit module; and the unit module after the target unit module can be deleted. For another example, the performance of the internal output module can also be tested using verification data to obtain the performance indicators of the internal output module. In this way, the output module with the best performance can be selected as the target output module; the unit module connected to the target output module can be used as the target unit module; and the unit module after the target unit module can be deleted. Among them, the output module with the best performance can be an internal output module or a plug-in output module.

[0112] Furthermore, other output modules other than the target output module may be deleted, and the target output module may be used as the output module of the target model itself, and the output of the target output module may be used as the output of the target model. The target output module may be an internal output module or an external output module. Other output modules other than the target output module may include internal output modules and / or external output modules.

[0113] The model pruning method of the embodiment of this specification adds an external output module to the target model, and the external output module can be trained together with the target model. The performance indicators of the external output module are used to prune multiple unit modules in the target model. This achieves pruning of the target model in the depth dimension, which enables the target model to achieve greater parallel operation efficiency.

[0114] The embodiment of this specification also provides a model pruning device. The model pruning device can be set in a computer device. The computer device can be a personal computer, a server, or a server cluster including multiple servers.

[0115] See also Figure 8 The model pruning device may include the following units.

[0116] The iteration unit 31 is used to iteratively perform the following steps until an end condition is met: determining mask information according to pruning parameters; the mask information is used to indicate the effective state of the pruned object in the target model; inputting the sample into the target model after adding the mask information to obtain a first output of the target model; optimizing parameter information according to the first output; the parameter information includes model parameters and pruning parameters of the target model;

[0117] The pruning unit 33 is used to perform pruning processing on the pruning object according to the mask information.

[0118] The embodiment of this specification also provides another model pruning device. The model pruning device can be set in a computer device. The computer device can be a personal computer, a server, or a server cluster including multiple servers.

[0119] See also Fig. 9 The model pruning device may include the following units.

[0120] The iteration unit 41 is used to iteratively execute the following steps until an end condition is met: inputting a sample into a target model after adding an external output module, obtaining a first output of the target model and a second output of the external output module; the target model includes a plurality of unit modules with the same structure stacked in sequence, and the external output module is connected to the unit modules; optimizing parameter information according to the first output and the second output, wherein the parameter information includes model parameters of the target model and model parameters of the external output module;

[0121] The pruning unit 43 is used to prune multiple unit modules according to the performance indicators of the external output module.

[0122] An embodiment of the computer device of the present specification is introduced below. Fig.10 Schematic diagram of the hardware structure of the computer device in this embodiment. Fig.10 As shown, the computer device may include one or more (only one is shown in the figure) processors, memory and transmission modules. Of course, it can be understood by those skilled in the art that Fig.10 The hardware structure shown is only for illustration and does not limit the hardware structure of the above-mentioned computer device. In practice, the computer device may also include Fig.10 More or fewer component units as shown; or, Fig.10 Different configurations are shown.

[0123] The memory may include a high-speed random access memory; or, it may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory or other non-volatile solid-state memory. Of course, the memory may also include a remotely set network memory. The memory may be used to store program instructions or modules of application software, such as the instructions in this manual. Figure 3 or Figure 7 The program instructions or modules of the corresponding embodiment.

[0124] The processor may be implemented in any suitable manner. For example, the processor may take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. The processor may read and execute program instructions or modules in the memory.

[0125] The transmission module can be used for transmitting data via a network, for example, via a network such as the Internet, an intranet, a local area network, a mobile communication network, etc.

[0126] This specification also provides an embodiment of a computer storage medium. The computer storage medium includes but is not limited to a random access memory (RAM), a read-only memory (ROM), a cache, a hard disk (HDD), a memory card, etc. The computer storage medium stores computer program instructions. When the computer program instructions are executed, the following are achieved: Figure 3 or Figure 7 The program instructions or modules of the corresponding embodiment.

[0127] It should be noted that each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, computer equipment embodiment, and computer storage medium embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In addition, it is understandable that after reading this specification document, those skilled in the art can think of any combination of some or all of the embodiments listed in this specification without creative work, and these combinations are also within the scope of disclosure and protection of this specification.

[0128] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0129] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0130] It can be known from the above description of the implementation mode that the technicians in this field can clearly understand that the present specification can be implemented by means of software plus the necessary general hardware platform. Based on such an understanding, the technical solution of the present specification can be essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present specification or some parts of the embodiments.

[0131] This specification can be used in many general or special computer system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.

[0132] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0133] Although the present specification is described through embodiments, those skilled in the art will appreciate that there are many modifications and changes to the present specification without departing from the spirit of the present specification, and it is intended that the appended claims include these modifications and changes without departing from the spirit of the present specification.

Claims

1. A model pruning method, comprising: Determine the mask information according to the pruning parameters; The mask information is used to indicate a valid state of a pruned object in a target model, wherein the target model includes an image classification model, and the image classification model includes a visual converter; the pruning parameters include a pruning threshold and value information, wherein the pruning threshold is a critical value for defining a valid state; Inputting the sample into the target model after adding the mask information, obtaining a first output of the target model, wherein the sample includes an image sample, and the first output is a prediction result of the target model, wherein the prediction result indicates whether the sample is a target image; Optimizing parameter information according to the first output; the parameter information includes model parameters and pruning parameters of the target model; Iterate the above steps until the end condition is met; According to the mask information, prune the pruning object; Wherein, the determining of mask information includes: calculating a first pruning rate according to a pruning threshold; determining mask information of a pruned object according to the first pruning rate and value information; the first pruning rate represents a pruning ratio of the pruned object; the pruned object includes a linear layer and / or a multi-head self-attention layer, the value information of the linear layer is used to measure the importance of the linear layer, and the importance of the linear layer is characterized by the effective state of a sub-weight parameter in a weight parameter, and the value information of the multi-head self-attention layer is used to measure the importance of the multi-head self-attention layer, and the importance of the multi-head self-attention layer is characterized by the effective state of a self-attention head in the multi-head self-attention layer; The value information includes multiple sub-value information, the mask information includes multiple sub-mask information, and there is a corresponding relationship between the sub-value information and the sub-mask information; the mask information of the pruning object determined according to the first pruning rate and the value information includes: determining a comparison benchmark according to the first pruning rate and the multiple sub-value information; comparing the sub-value information with the comparison benchmark; and assigning a value to the sub-mask information corresponding to the sub-value information according to the comparison result.

2. According to the method of claim 1, the optimization parameter information comprises: According to the loss function, the parameter information is optimized; the loss function includes a first term and a second term, the first term is used to constrain the second pruning rate of the target model, and the second term is used to represent the deviation between the first output and the sample label.

3. According to the method according to claim 1, the mask information includes first mask information and / or second mask information, the first mask information is used to indicate the valid state of the linear layer, and the second mask information is used to indicate the valid state of the multi-head self-attention layer.

4. The method according to claim 3, wherein the first mask information includes a plurality of first sub-mask information, the weight parameter of the linear layer includes a plurality of sub-weight parameters, and the pruning process of the pruned object comprises: The sub-weight parameter is deleted according to the value of the first sub-mask information.

5. According to the method of claim 3, the second mask information includes a plurality of second sub-mask information, the multi-head self-attention layer includes a plurality of self-attention heads, and the pruning process of the pruned object comprises: According to the value of the second sub-mask information, the self-attention head is deleted.

6. The method according to claim 1, wherein inputting the sample into the target model after adding mask information comprises: Input the sample into the target model after adding the mask information and the plug-in output module, and obtain the second output of the plug-in output module; The target model includes a plurality of unit modules with the same structure stacked in sequence, and the external output module is connected to the unit modules; The optimization parameter information includes: Optimizing parameter information according to the first output and the second output; The parameter information includes model parameters of the plug-in output module.

7. The method according to claim 6, wherein the optimization parameter information comprises: According to the loss function, optimize the parameter information; The loss function includes a third term, and the third term is used to represent the deviation between the second output and the sample label.

8. The method according to claim 6, further comprising: According to the performance indicators of the plug-in output module, multiple unit modules are pruned.

9. A model pruning device, comprising: The iteration unit is used to iteratively execute the following steps until an end condition is met: determining mask information according to a pruning parameter; The mask information is used to indicate the effective state of the pruned object in the target model, the target model includes an image classification model, and the image classification model includes a visual converter; the pruning parameters include a pruning threshold and value information, and the pruning threshold is a critical value for defining the effective state; a sample is input into the target model after the mask information is added to obtain a first output of the target model, the sample includes an image sample, and the first output is a prediction result of the target model, and the prediction result indicates whether the sample is a target image; according to the first output, the parameter information is optimized; the parameter information includes model parameters and pruning parameters of the target model; A pruning unit, used for pruning the pruning object according to the mask information; Wherein, the determining of mask information includes: calculating a first pruning rate according to a pruning threshold; determining mask information of a pruned object according to the first pruning rate and value information; the first pruning rate represents a pruning ratio of the pruned object; the pruned object includes a linear layer and / or a multi-head self-attention layer, the value information of the linear layer is used to measure the importance of the linear layer, and the importance of the linear layer is characterized by the effective state of a sub-weight parameter in a weight parameter, and the value information of the multi-head self-attention layer is used to measure the importance of the multi-head self-attention layer, and the importance of the multi-head self-attention layer is characterized by the effective state of a self-attention head in the multi-head self-attention layer; The value information includes multiple sub-value information, the mask information includes multiple sub-mask information, and there is a corresponding relationship between the sub-value information and the sub-mask information; the mask information of the pruning object determined according to the first pruning rate and the value information includes: determining a comparison benchmark according to the first pruning rate and the multiple sub-value information; comparing the sub-value information with the comparison benchmark; and assigning a value to the sub-mask information corresponding to the sub-value information according to the comparison result.

10. A computer device comprising: at least one processor; A memory storing program instructions, wherein the program instructions are configured to be executed by the at least one processor, and the program instructions include instructions for executing the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Convolutional neural network pruning method based on feature map sparsification

    CN110874631A

  • Mask-based depth map convolutional neural network model pruning method and system

    CN111667068A

  • Channel attention guided convolutional neural network dynamic channel pruning method and device

    CN112949840A

  • Model pruning method and device, electronic equipment and storage medium

    CN114037074A