Model processing method and device, equipment, medium and product
By acquiring K clusters of neural network models for quantization and pruning, the problems of model instability and low inference efficiency caused by k-means clustering algorithm are solved, and more efficient neural network model inference is achieved.
Patent Information
- Application Number
- CN202510025305.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-30
AI Technical Summary
During the training of neural network model, the k-means clustering algorithm is used to quantify model parameters, resulting in model instability and inferred inference efficiency.
By obtaining K clusters of the neural network model, quantizing the model, and pruning is performed on this basis, the pruning neural network model is obtained.
This reduces the instability brought by quantitative processing, improves the inference efficiency of neural network models, and makes the model more versatile.
Smart Images

Figure CN120068957A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a model processing method, apparatus, device, medium, and product. Background Art
[0002] During the training process of a neural network model, to reduce resource consumption, related technologies usually use the k-means clustering algorithm to perform quantization processing on model parameters. However, this processing method may cause the neural network model to be unstable and the inference process to be inefficient. Summary of the Invention
[0003] This application mainly provides a model processing method, apparatus, device, medium, and product.
[0004] The technical solution of this application is implemented as follows:
[0005] A model processing method, the method includes:
[0006] Obtain K clustering clusters of a neural network model; the clustering clusters are used to store weight parameters of network nodes in the neural network model; K is a positive integer greater than 1;
[0007] Perform quantization processing on the neural network model based on the K clustering clusters to obtain a quantized neural network model;
[0008] Perform pruning processing on the quantized neural network model to obtain a pruned neural network model.
[0009] In the above solution, the obtaining of the K clustering clusters of the neural network model includes:
[0010] Obtain sample data of a sample image;
[0011] Train the network nodes based on the sample data to obtain the weight parameters;
[0012] Construct a clustering feature tree based on the weight parameters, and cluster leaf nodes of the clustering feature tree to obtain the K clustering clusters.
[0013] In the above solution, the performing quantization processing on the neural network model based on the K clustering clusters to obtain a quantized neural network model includes:
[0014] Perform quantization processing on the weight parameters based on the K clustering clusters to obtain quantized weight parameters;
[0015] Update the weight parameters of the neural network model based on the quantized weight parameters to obtain a quantized neural network model.
[0016] In the above solution, the quantization process of the weight parameters based on the K clustering clusters to obtain the quantized weight parameters includes:
[0017] Performing clustering on the K clustering clusters to obtain the label and clustering center point of each clustering cluster in the K clustering clusters;
[0018] Updating the label based on the weight parameter corresponding to each clustering cluster in the K clustering clusters to obtain the cluster weight of each clustering cluster in the K clustering clusters;
[0019] Restoring the cluster weight based on the label and the clustering center point to obtain the quantized weight parameter.
[0020] In the above solution, the pruning process of the quantized neural network model to obtain the pruned neural network model includes:
[0021] Dividing the convolution kernels of the quantized neural network model into multiple unit convolution kernels;
[0022] Using the pruning condition to prune the multiple unit convolution kernels of the quantized neural network model to obtain the pruned neural network model; the pruning condition includes:
[0023] Pruning the unit convolution kernels with kernel weights less than a preset threshold among the multiple unit convolution kernels.
[0024] In the above solution, the method further includes:
[0025] Obtaining an image to be recognized;
[0026] Inputting the image to be recognized into the pruned neural network model to obtain the target object in the image to be recognized output by the pruned neural network model.
[0027] A model processing device, the device includes:
[0028] An acquisition unit, configured to acquire K clustering clusters of a neural network model; the clustering clusters are used to store the weight parameters of network nodes in the neural network model; K is a positive integer greater than 1;
[0029] A processing unit, configured to perform quantization processing on the neural network model based on the K clustering clusters to obtain a quantized neural network model;
[0030] The processing unit is further configured to perform pruning processing on the quantized neural network model to obtain a pruned neural network model.
[0031] An electronic device, comprising: a processor and a memory for storing a computer program that can run on the processor,
[0032] wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of the above.
[0033] A storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the steps of the method described in any one of the above.
[0034] A computer product, comprising a computer program, characterized in that when the computer program is executed by a processor, it implements the steps of the method described in any one of the above.
[0035] The present invention provides a model processing method, device, equipment, medium and product, which obtains K clustering clusters of a neural network model; the clustering clusters are used to store the weight parameters of network nodes in the neural network model; K is a positive integer greater than 1; based on the K clustering clusters, the neural network model is quantized to obtain a quantized neural network model; the quantized neural network model is pruned to obtain a pruned neural network model. That is to say, in this application, by obtaining K clustering clusters of a neural network model, the clustering clusters are used to store the weight parameters of network nodes in the neural network model, and after quantizing the neural network model based on the K clustering clusters, pruning is performed to obtain a pruned neural network model, so as to reduce the instability brought by the quantization process and improve the efficiency of the inference process of the neural network model, solve the problem that the quantization processing method in the related technology causes the neural network model to be unstable and the inference process to be inefficient, and make the neural network model have better versatility. Description of the Drawings
[0036] Figure 1 It is a schematic architecture diagram for processing the training and inference processes of a neural network model in the related technology;
[0037] Figure 2 It is a schematic flowchart of a model processing method provided by this application;
[0038] Figure 3 It is a schematic flowchart of another model processing method provided by this application;
[0039] Figure 4 It is a schematic flowchart of weight parameter quantization provided by this application;
[0040] Figure 5 It is a schematic diagram of the solution space of a normal penalty function provided by this application;
[0041] Figure 6 It is a schematic diagram of model pruning provided by this application;
[0042] Figure 7 A schematic diagram of the sampling accuracy provided for this application;
[0043] Figure 8 A schematic diagram of the sampling loss provided for this application;
[0044] Figure 9 A schematic diagram of the comparison of model quantization compression provided for this application;
[0045] Figure 10 A schematic diagram of the structure of a model processing device provided for this application;
[0046] Figure 11 A schematic diagram of the structure of an electronic device provided for this application. Detailed implementation manners
[0047] In order to make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be further elaborated in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0048] In the following descriptions, reference is made to "some embodiments", which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0049] The terms "first / second / third" involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0051] In the related art, artificial intelligence (AI) large models generally refer to large-scale language models with tens of billions to hundreds of billions of parameters and very complex deep neural network structures. AI large models have the following obvious characteristics:
[0052] (1) The parameter scale of the model is huge. The number of parameters in large AI models reaches tens of billions or even hundreds of billions, far exceeding that of ordinary neural network models. The increase in the number of parameters has significantly improved the model's learning ability and knowledge capacity.
[0053] (2) The scale of the dataset used is huge. During the training process, large-scale text or image datasets are used, and the dataset scale reaches hundreds of gigabytes (GB) to terabytes (TB).
[0054] (3) The model structure is complex. Large AI models are mainly based on deep neural network structures such as the Transformer attention mechanism.
[0055] (4) The training cost is high. It requires a large cluster to support a long training process. For example, thousands of Graphics Processing Units (GPUs).
[0056] (5) It has strong multitasking capabilities and powerful transfer learning and multitasking learning capabilities, and can be applied to downstream tasks in multiple fields.
[0057] (6) It has high performance and strong universality for different businesses.
[0058] For different scenarios, in actual deployment, large AI models often require specific optimizations. The huge amount of computing resources used in neural network training and inference also brings high computing and storage costs, which virtually restricts the generalization development of neural network models.
[0059] To accelerate the training and inference of neural network models, refer to Figure 1As shown in the figure, in terms of hardware, multi-card parallelism and custom chips can be adopted. Multi-card parallelism uses multiple GPUs or Tensor Processing Units (TPUs) for parallel computing to improve throughput. Custom chips are AI large model chips designed specifically, and can also be combined with Remote Direct Memory Access (RDMA) to reduce the inter-machine loss of multi-machine communication. In terms of software, the model deployment framework can be optimized; in terms of computing, sharded inference and redundant calculation reduction can be adopted; sharded inference only performs inference on a part of the input, thereby reducing the computational amount; redundant calculation reduction is to identify and skip some repeated calculation processes. In terms of the model, pruning, knowledge distillation, quantization compression, etc. can be used to compress the size of the model and reduce the model's computational amount. Pruning means designing an evaluation criterion for network parameters on a pre-trained large model and deleting redundant parameters based on this. Quantization technology processes a large number of parameters through methods such as clustering, quantizing high-precision network parameters to lower bit numbers. Knowledge distillation uses a trained small model to simulate the behavior of a large model. Compared with training a large model, obtaining a small model is more efficient. For models in special scenarios, dynamic adjustment can also be used to dynamically adjust the depth, width, etc. of the model according to specific situations to reduce the computational amount.
[0060] The methods adopted in parameter quantization include clustering quantization, logarithmic quantization, binary quantization, 8-bit quantization, etc. Among them, 8-bit quantization is a relatively mature solution in engineering applications, but it ignores the influence of cluster characteristics compared with clustering quantization. Clustering quantization usually only uses K-means clustering, and the result of K-means clustering is easily affected by the initial clustering center and falls into a local optimal solution rather than a global optimal solution. In addition, K-means clustering is suitable for small data volumes. When the data volume is large, the iteration speed of using the K-means algorithm is too slow.
[0061] Pruning mainly includes structured pruning and unstructured pruning. Among them, unstructured pruning mainly prunes the weight values of individual or entire rows and columns in the weight matrix. The new weight matrix after pruning will become a sparse matrix (the pruned values will be set to 0). If the hardware platform and computing library cannot support efficient sparse matrix calculation, the performance of the pruned model cannot be improved. The basic pruning unit of structured pruning is one or more channels of the filter or weight matrix. Structured pruning mainly includes filter-wise pruning, channel-wise pruning, shape-wise pruning, and stripe-wise pruning; structured pruning does not change the sparsity degree of the weight matrix itself and can well support existing computing frameworks.
[0062] An embodiment of the present application provides a model processing method. Refer to Figure 2 as shown, the method includes the following steps:
[0063] Step S201: Obtain K clustering clusters of the neural network model.
[0064] Among them, the clustering clusters are used to store the weight parameters of the network nodes in the neural network model; K is a positive integer greater than 1.
[0065] It can be understood that the neural network model contains a large amount of data and numerous network nodes. A clustering feature (CF) tree stored in memory can be dynamically established based on the Balanced Iterative Reducing and Clustering using Hierarchies (Birch) algorithm, and the weight parameters of the network nodes in the neural network model are stored in the CF nodes.
[0066] In practical applications, after obtaining the CF tree, the k-means clustering algorithm can be used to cluster the leaf nodes to obtain multiple clustering clusters, and determine the number K of the optimal clustering clusters in the clustering clusters, where K is less than or equal to 2 N .
[0067] Step S202: Perform quantization processing on the neural network model based on the K clustering clusters to obtain a quantized neural network model.
[0068] In practical applications, the birch algorithm can be used to cluster the K optimal clustering clusters to obtain the labels and centroids corresponding to each cluster in the clustering clusters. Based on the number of weight parameters in different clustering clusters, the labels corresponding to each cluster are updated to obtain the cluster weights of the birch clustering clusters. Then, the restoration function is used to restore the weight parameters of the neural network model according to the labels and centroids, update the weight parameters of the neural network model, and obtain a quantized neural network model.
[0069] Step S203: Perform pruning processing on the quantized neural network model to obtain a pruned neural network model.
[0070] In practical applications, the convolution kernels of the quantized neural network model can be divided into multiple unit convolution kernels, and regularization operations are performed on the filter skeletons formed by the filters in each unit convolution kernel. The convolution kernels with weight parameters less than the preset threshold after the operation are trimmed to simplify the neural network model.
[0071] Refer to Figure 3As shown in the figure, after quantizing and pruning the neural network model, the Birch algorithm can ensure that the results follow a normal distribution. When pruning the quantized neural network model, the weight data of the convolutional layer during pruning can be made more centralized, effectively shortening the time required for model inference and improving the inference efficiency of the neural network model.
[0072] As can be seen from the above, in the embodiments of the present application, by obtaining K clustering clusters of the neural network model, where the clustering clusters are used to store the weight parameters of the network nodes in the neural network model, quantizing the neural network model based on the K clustering clusters and then pruning it to obtain the pruned neural network model, to reduce the instability brought by the quantization process and improve the efficiency of the neural network model inference process, solving the problem that the quantization processing method in the related art causes the neural network model to be unstable and the inference process to be inefficient, making the neural network model have better generality.
[0073] In some embodiments of the present application, obtaining K clustering clusters of the neural network model includes:
[0074] Obtaining the sample data of the sample image;
[0075] Training the network nodes based on the sample data to obtain the weight parameters;
[0076] Constructing a clustering feature tree based on the weight parameters and clustering the leaf nodes of the clustering feature tree to obtain K clustering clusters.
[0077] In practical applications, before training the neural network model, a dataset for neural network training can be obtained. In the embodiments of the present application, taking image data as an example, the sample data of the sample image can be obtained.
[0078] The specific steps for constructing the CF tree are as follows:
[0079] 1. Store the weight parameters in the CF node in the form of a triple. Starting from the root node of the CF tree, recursively go down, calculate the distance between the current data point and the target data point to be inserted, find the path with the smallest distance, until the target leaf node closest to the target data point is found;
[0080] 2. Compare whether the calculated distance is less than the threshold T. If it is less, the target leaf node absorbs the current data point; otherwise, continue to the next step.
[0081] 3. Determine whether the number of data points in the leaf node where the target data point is located is less than L. If so, directly insert the data point as part of the target leaf node. Otherwise, the leaf node needs to be split. The splitting principle is to find the two data points with the farthest distance in the leaf node and use these two data points as the starting entries of the two new leaf nodes after splitting. The remaining data points are allocated to these two new leaf nodes according to the principle of the smallest distance. Delete the original leaf node and update the entire CF tree.
[0082] When the data point cannot be inserted, the threshold T needs to be increased at this time and the tree needs to be rebuilt to absorb more leaf node data until all data points are inserted. During the CF tree reconstruction process, a new tree is rebuilt by using the leaf nodes of the old tree. Therefore, the tree reconstruction process does not need to access all points, that is, only access the data once to build the CF tree. If the memory is not enough, increase the threshold and construct a smaller tree based on the original tree.
[0083] In some embodiments of the present application, the neural network model is quantized based on K clustering clusters to obtain the quantized neural network model, including:
[0084] Quantize the weight parameters based on K clustering clusters to obtain the quantized weight parameters;
[0085] Based on the quantized weight parameters, update the weight parameters of the neural network model to obtain the quantized neural network model.
[0086] In practical applications, after obtaining the CF tree, the k-means clustering algorithm can be used to cluster the leaf nodes to obtain multiple clustering clusters, and determine the number K of the optimal clustering clusters in the clustering clusters, where K is less than or equal to 2 N . Perform clustering processing on the K optimal clustering clusters using the birch algorithm to obtain the labels and centriods corresponding to each cluster in the clustering clusters. Based on the number of weight parameters in different clustering clusters, update the labels corresponding to each cluster to obtain the cluster weights of the birch clustering clusters. Then use the restoration function to restore the weight parameters of the neural network model according to the labels and centriods, and update the weight parameters of the neural network model to obtain the quantized neural network model.
[0087] In some embodiments of the present application, quantizing the weight parameters based on K clustering clusters to obtain the quantized weight parameters includes:
[0088] Perform clustering processing on the K clustering clusters to obtain the labels and clustering centers of each clustering cluster in the K clustering clusters;
[0089] Update the labels based on the weight parameters corresponding to each clustering cluster among the K clustering clusters to obtain the cluster weights of each clustering cluster among the K clustering clusters.
[0090] Restore the cluster weights based on the labels and the clustering centroids to obtain the quantized weight parameters.
[0091] In practical applications, referring to Figure 4 As shown, the birch algorithm can be used to perform clustering on the K optimal clustering clusters to obtain the labels and centriods corresponding to each cluster in the clustering clusters. Based on the number of weight parameters in different clustering clusters, update the labels corresponding to each cluster to obtain the cluster weights of the birch clustering clusters. The update method can be as shown in Formula 1:
[0092]
[0093] where w_new i is the updated cluster weight value of the birch clustering cluster i, n represents the number of weight parameters in the current birch clustering cluster, and w_lb i is the label of the birch clustering cluster i.
[0094] Then use the restoration function to restore the weight parameters of the neural network model according to the labels and centriods, update the weight parameters of the neural network model, and obtain the quantized neural network model.
[0095] In some embodiments of the present application, pruning is performed on the quantized neural network model to obtain a pruned neural network model, including:
[0096] Divide the convolution kernels of the quantized neural network model into multiple unit convolution kernels;
[0097] Use the pruning condition to prune the multiple unit convolution kernels of the quantized neural network model to obtain a pruned neural network model; the pruning condition includes:
[0098] Prune the unit convolution kernels with kernel weights less than the preset threshold among the multiple unit convolution kernels.
[0099] It can be understood that the Sliding Window Pruning (SWP) algorithm divides N convolution kernels with dimensions of C×K×K into K×K convolution kernels of 1×1×C, and prunes the convolution kernels in units of 1×1×C. WP prunes the model by pruning the shape of the filter and can be compatible with filter pruning. In the embodiments of the present application, SWP is used as an example to prune the model.
[0100] In practical applications, the filter skeleton (FS) can be compressed by the regularization formula of Formula 2:
[0101]
[0102] where x and y respectively represent the independent variable and the dependent variable of the input data. I is the FS, and W is the weight corresponding to each convolutional layer. α controls the influence degree of the regularization term, and g(I) is a common l1 1 normal penalty function acting on I. g(I) can be written as Formula 3 below, and the matrix form of g(I) can be written as the following function formula for minimizing, that is, Formula 4:
[0103]
[0104] I 1 The normal penalty function is composed of a quadratic function J0 and an absolute value function. Its principle is that on the contour line of the penalty function of I 1 J 0 the function reaches the minimum value on the coordinate axis at (0, W 2 ). As shown in the reference Figure 5 shown, I 1 regularization will make many weights of w equal to 0, and ideal pruning can be performed according to the finally obtained weight parameters.
[0105] FS learns the ideal shape of each filter. To perform ideal pruning, a threshold δ is set here. When the corresponding index of a certain one in FS is lower than the threshold δ, it will not participate in training and will be pruned. During the model inference process, the pruned filter cannot perform convolution operations on the input feature map as a whole. Therefore, it is necessary to perform convolution item by item and then sum up the feature maps. The whole process can be summarized as Formula 5:
[0106]
[0107] As shown in the reference Figure 6 shown, is a point on the (I + 1)-th layer convolution of the feature map. Compared with the standard convolution algorithm, SWP changes the calculation order of the convolution process and does not generate additional computational overhead in the network layer. However, since the data spline positions in the filter are fixed, it is necessary to record the index positions of each data spline before performing item-by-item pruning. This operation reduces the storage overhead compared to recording the entire network parameters. Compared with the N×C×K×K position indexes required to be recorded for single weight parameter pruning, only N×K×K position indexes need to be recorded in the embodiments of the present application.
[0108] In some embodiments of the present application, the method further includes:
[0109] Obtain an image to be recognized;
[0110] The image to be identified is input into the pruned neural network model to obtain the target object in the image to be identified output by the pruned neural network model.
[0111] In practical applications, after obtaining the pruned neural network model, the target object in the image to be identified can be identified through the pruned neural network model. Figure 7 As shown, after the quantization and pruning processing of the embodiment of the present application is adopted, the stability of the neural network model can be improved and the sampling accuracy can be improved.
[0112] refer to Figure 8 As shown in Figure 2, using SWP can achieve higher accuracy and lower training loss faster. Figure 9 As shown in the figure, applying Birch's quantization compression to the model can make the weight data of the convolutional layer more centralized during pruning. Compared with the case of applying only pruning technology, the weight parameter data of the pruned part is quite different from the weight parameter data of the remaining part, thus avoiding the error problem caused by the small difference in weight parameters. Quantizing before pruning also reduces some of the computational complexity of the pruning process.
[0113] Based on the same inventive concept as above, Figure 10 The present invention provides a schematic diagram of a model processing device according to an embodiment of the present invention. The model processing device 1000 includes:
[0114] The acquisition unit 1001 is used to acquire K clusters of the neural network model; the clusters are used to store weight parameters of network nodes in the neural network model; K is a positive integer greater than 1;
[0115] The processing unit 1002 is used to perform quantization processing on the neural network model based on the K clusters to obtain a quantized neural network model;
[0116] The processing unit 1002 is also used to prune the neural network model after quantization to obtain a pruned neural network model.
[0117] In some embodiments of the present application, the acquisition unit 1001 is used to acquire sample data of a sample image;
[0118] Train the network nodes based on sample data to obtain weight parameters;
[0119] A clustering feature tree is constructed based on the weight parameters, and the leaf nodes of the clustering feature tree are clustered to obtain K clusters.
[0120] In some embodiments of the present application, an acquisition unit 1001 is configured to perform quantization processing on weight parameters based on K clustering clusters to obtain the quantized weight parameters;
[0121] Based on the quantized weight parameters, update the weight parameters of the neural network model to obtain the quantized neural network model.
[0122] In some embodiments of the present application, a processing unit 1002 is configured to perform clustering processing on K clustering clusters to obtain the labels and cluster centers of each of the K clustering clusters;
[0123] Update the labels based on the weight parameters corresponding to each of the K clustering clusters to obtain the cluster weights of each of the K clustering clusters;
[0124] Restore the cluster weights based on the labels and cluster centers to obtain the quantized weight parameters.
[0125] In some embodiments of the present application, a processing unit 1002 is configured to divide the convolution kernels of the quantized neural network model into multiple unit convolution kernels;
[0126] Use pruning conditions to prune multiple unit convolution kernels of the quantized neural network model to obtain a pruned neural network model; the pruning conditions include:
[0127] Prune the unit convolution kernels with kernel weights less than a preset threshold among the multiple unit convolution kernels.
[0128] In some embodiments of the present application, a processing unit 1002 is configured to obtain an image to be recognized;
[0129] Input the image to be recognized into the pruned neural network model to obtain the target object in the image to be recognized output by the pruned neural network model.
[0130] Based on the foregoing embodiments, an embodiment of the present application provides an electronic device, Figure 11 FIG. is a schematic hardware structure diagram of an electronic device according to an embodiment of the present invention. The electronic device 1100 includes: at least one processor 1101, a memory 1102. Optionally, the electronic device 1100 may further include at least one communication interface 1103. Each component in the electronic device 1100 is coupled together through a bus system 1104. It can be understood that the bus system 1104 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1104 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 11 all kinds of buses are labeled as the bus system 1104.
[0131] Based on the hardware implementation of the above program modules, the communication interface 1103 can interact with other communication devices for information;
[0132] The processor 1101 is connected to the communication interface 1103 to interact with other communication devices for information and execute the methods provided by the above one or more technical solutions when running a computer program;
[0133] The memory 1102 stores the computer program.
[0134] Specifically, the processor 1101 is used to obtain K clustering clusters of the neural network model; the clustering clusters are used to store the weight parameters of the network nodes in the neural network model; K is a positive integer greater than 1;
[0135] Perform quantization processing on the neural network model based on the K clustering clusters to obtain the quantized neural network model;
[0136] Perform pruning processing on the quantized neural network model to obtain the pruned neural network model.
[0137] In some embodiments of the present application, the processor 1101 is used to obtain the sample data of the sample image;
[0138] Train the network nodes based on the sample data to obtain the weight parameters;
[0139] Construct a clustering feature tree based on the weight parameters and cluster the leaf nodes of the clustering feature tree to obtain K clustering clusters.
[0140] In some embodiments of the present application, the processor 1101 is used to perform quantization processing on the weight parameters based on the K clustering clusters to obtain the quantized weight parameters;
[0141] Update the weight parameters of the neural network model based on the quantized weight parameters to obtain the quantized neural network model.
[0142] In some embodiments of the present application, the processor 1101 is used to perform clustering processing on the K clustering clusters to obtain the labels and clustering center points of each clustering cluster in the K clustering clusters;
[0143] Update the labels based on the weight parameters corresponding to each clustering cluster in the K clustering clusters to obtain the cluster weights of each clustering cluster in the K clustering clusters;
[0144] Restore the cluster weights based on the labels and clustering center points to obtain the quantized weight parameters.
[0145] In some embodiments of the present application, the processor 1101 is configured to divide the convolution kernels of the quantized neural network model into multiple unit convolution kernels;
[0146] Use the pruning condition to prune multiple unit convolution kernels of the quantized neural network model to obtain a pruned neural network model; the pruning condition includes:
[0147] Prune the unit convolution kernels with kernel weights less than a preset threshold among the multiple unit convolution kernels.
[0148] In some embodiments of the present application, the processor 1101 is configured to obtain an image to be recognized;
[0149] Input the image to be recognized into the pruned neural network model to obtain the target object in the image to be recognized output by the pruned neural network model.
[0150] It can be understood that the memory 1102 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM).The memory 1102 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0151] The memory 1102 in the embodiments of the present invention is used to store various types of data to support the operation of the electronic device 1100. Examples of such data include: any computer program for operating on the electronic device 1100, and the program for implementing the method of the embodiments of the present invention may be included in the memory 1102.
[0152] The method disclosed in the above embodiments of the present invention can be applied to the processor 1101 or implemented by the processor 1101. The processor may be an integrated circuit chip with the ability to process signals. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in the form of software. The above processor may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention, it can be directly embodied as being executed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.
[0153] In an exemplary embodiment, the electronic device 1100 can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components for executing the above method.
[0154] Based on the foregoing embodiments, an embodiment of the present application further provides a computer product, including a computer program, and when the computer program is executed by a processor, it implements Figure 1 the steps in the model processing method provided by the corresponding embodiment.
[0155] Based on the foregoing embodiments, an embodiment of the present application further provides a storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are configured to execute Figure 1 the model processing method provided by the corresponding embodiment.
[0156] It should be noted that the above computer storage medium may be a memory such as ROM, PROM, EPROM, EEPROM, FRAM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM; it may also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0157] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0158] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0160] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks.
[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks.
[0163] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A model processing method, characterized in that: The method comprises: Obtain K clusters of the neural network model; the clusters are used to store weight parameters of network nodes in the neural network model; K is a positive integer greater than 1; Quantizing the neural network model based on the K clusters to obtain a quantized neural network model; The quantized neural network model is pruned to obtain a pruned neural network model.
2. The method according to claim 1, characterized in that The step of obtaining K clusters of the neural network model includes: Get sample data of sample image; Training the network nodes based on the sample data to obtain the weight parameters; The cluster feature tree is constructed based on the weight parameters, and the leaf nodes of the cluster feature tree are clustered to obtain the K cluster clusters.
3. The method according to claim 1, characterized in that The quantizing process of the neural network model based on the K clusters to obtain the quantized neural network model includes: quantizing the weight parameter based on the K clusters to obtain the quantized weight parameter; Based on the weight parameters after the quantization processing, the weight parameters of the neural network model are updated to obtain the neural network model after the quantization processing.
4. The method according to claim 3, characterized in that: The step of quantizing the weight parameter based on the K clusters to obtain the quantized weight parameter includes: Performing clustering processing on the K clusters to obtain a label and a cluster center point of each cluster in the K clusters; The label is updated based on the weight parameter corresponding to each cluster in the K clusters to obtain a cluster weight of each cluster in the K clusters; The cluster weight is restored based on the label and the cluster center point to obtain the quantized weight parameter.
5. The method according to claim 1, characterized in that The pruning process is performed on the quantized neural network model to obtain a pruned neural network model, including: Dividing the convolution kernel of the quantized neural network model into a plurality of unit convolution kernels; The plurality of unit convolution kernels of the quantized neural network model are pruned using a pruned condition to obtain the pruned neural network model; the pruned condition includes: The unit convolution kernels whose kernel weights are less than a preset threshold value among the multiple unit convolution kernels are pruned.
6. The method according to claim 1, characterized in that The method further comprises: Obtain an image to be recognized; The image to be identified is input into the pruned neural network model to obtain a target object in the image to be identified output by the pruned neural network model.
7. A model processing device, characterized in that: The device comprises: An acquisition unit, used for acquiring K clusters of the neural network model; the clusters are used for storing weight parameters of network nodes in the neural network model; K is a positive integer greater than 1; A processing unit, configured to perform quantization processing on the neural network model based on the K clusters to obtain a quantized neural network model; The processing unit is also used to prune the neural network model after the quantization process to obtain a pruned neural network model.
8. An electronic device, characterized in that: include: a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of claims 1 to 6.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.