Model pruning method and device, task execution method and device, storage medium and equipment
By acquiring and utilizing the mutual information quantization contribution value of the intermediate layer of the deep learning model for sorting and pruning, the problem of inefficiency in inference of large-scale deep learning models is solved, and the effect of improving efficiency and reducing costs without re-training is achieved.
Patent Information
- Application Number
- CN202510071779.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Large-scale deep learning models consume a lot of computing resources and time during the inference process, resulting in inaccurate inference efficiency and difficult to meet the scenario requirements of high real-time requirements.
By obtaining the mutual information of each intermediate layer, quantizing the contribution value, sorting and pruning the redundant intermediate layer, the processed model is obtained.
Without large-scale retraining, the number of model parameters is reduced, the multiplication and addition of matrix operations is reduced, the inference efficiency is improved, and the adjustment cost is reduced.
Smart Images

Figure CN119990234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model pruning, task execution method, device, storage medium and equipment. Background Art
[0002] With the rapid development of deep learning technology, large-scale deep learning models (such as pre-trained language models, convolutional neural network models, etc.) have been widely used in key fields such as natural language processing and computer vision, and have greatly promoted the progress of artificial intelligence technology. However, since large-scale deep learning models usually have billions or even more parameters, the reasoning process of large-scale deep learning models (that is, the process of predicting or classifying new data through trained models) consumes a lot of computing resources and time, resulting in low reasoning efficiency and difficulty in meeting the needs of scenarios with high real-time requirements.
[0003] In order to improve the inference efficiency of large-scale deep learning models, methods such as model compression (i.e., reducing the number of model parameters or reducing the precision of parameters to reduce the computational complexity) and quantization (i.e., converting the floating-point parameters in the model to low-order integers to reduce the consumption of computing resources and increase the computing speed) can usually be used to perform inference processing on large-scale deep learning models. However, after the large-scale deep learning models are processed by model compression and quantization methods, the models need to be retrained on a large scale to adapt to the model structure after the parameters have changed, resulting in a high cost for adjusting large-scale deep learning models. Summary of the invention
[0004] This specification provides a model pruning, task execution method, device, storage medium and equipment to partially solve the above-mentioned problems existing in the prior art.
[0005] This manual adopts the following technical solutions:
[0006] This manual provides a model pruning method, including:
[0007] Obtaining a mutual information quantization contribution value of each intermediate layer in a process of executing a reasoning task by a to-be-processed model, wherein the mutual information quantization contribution value is used to characterize the proportion of contribution information to reducing uncertainty of an output result of the intermediate layer relative to redundant information when input data input to the intermediate layer is known, wherein the reasoning task includes a visual processing task;
[0008] Sorting the intermediate layers according to the mutual information quantization contribution value of each intermediate layer, and selecting a redundant intermediate layer from the intermediate layers according to the sorting result;
[0009] According to the selection result, the redundant intermediate layers included in the model to be processed are pruned to obtain a processed model.
[0010] Optionally, obtaining the mutual information quantization contribution value of each intermediate layer in the process of executing the inference task of the model to be processed includes:
[0011] For each intermediate layer included in the model to be processed, determining the mutual information amount of the intermediate layer according to the input data and output result of the intermediate layer;
[0012] Determine the redundant information value between the input data and the output result of the intermediate layer according to the mutual information amount, the entropy of the input data of the intermediate layer, and the entropy of the output result of the intermediate layer;
[0013] The mutual information quantization contribution value of the intermediate layer is determined according to the ratio of the mutual information amount to the redundant information value.
[0014] Optionally, determining the mutual information of the intermediate layer according to the input data and the output result of the intermediate layer specifically includes:
[0015] Determine the entropy value of the input data of the intermediate layer and the conditional entropy value of the input data of the intermediate layer when the output result of the intermediate layer is known;
[0016] The mutual information of the intermediate layer is determined according to the entropy value and the conditional entropy value.
[0017] Optionally, obtaining the mutual information quantization contribution value of each intermediate layer in the process of executing the inference task of the model to be processed includes:
[0018] According to the order of the intermediate layers contained in the model to be processed, the input data and output results of each intermediate layer are input into the preset evaluation model in sequence, so that the evaluation model determines the quantitative contribution value of the mutual information of each intermediate layer according to the input data and output results of the intermediate layer and the correlation relationship between the intermediate layers before the intermediate layer.
[0019] Optionally, the method further comprises:
[0020] Inputting the preset first sample data into the processed model to determine the performance parameters of the processed model when performing the reasoning task corresponding to the first sample data;
[0021] Based on the deviation between the performance parameters and the preset original performance parameters, a detection result for the processed model is obtained, and when it is determined that the processed model has an abnormality based on the detection result, the model to be processed is pruned again, and the original performance parameters are the performance parameters of the model to be processed when executing the reasoning task corresponding to the first sample data.
[0022] Optionally, the method further comprises:
[0023] Inputting the preset second sample data into the processed model to obtain output information;
[0024] Determine a target loss value according to a deviation between the output information and actual output information of the second sample data, wherein the target loss value is positively correlated with the deviation;
[0025] Taking minimizing the target loss value as the optimization goal, the model parameters of the processed model are fine-tuned to obtain an adjusted model.
[0026] This specification provides a task execution method, including:
[0027] Acquire the task data to be executed, wherein the task data to be executed includes: at least one of text data, picture data, and video data;
[0028] The task data to be executed is input into the processed model to obtain an output result, and the task is executed according to the output result. The processed model is obtained after pruning by the above-mentioned model pruning method.
[0029] This specification provides a model pruning device, including:
[0030] An acquisition module, used to acquire a mutual information quantization contribution value of each intermediate layer in a process of executing a reasoning task by a to-be-processed model, wherein the mutual information quantization contribution value is used to characterize the proportion of contribution information for reducing the uncertainty of an output result of the intermediate layer relative to redundant information when input data input to the intermediate layer is known, and the reasoning task includes a visual processing task;
[0031] A selection module, used to sort the intermediate layers according to the mutual information quantification contribution value of each intermediate layer, and select a redundant intermediate layer from the intermediate layers according to the sorting result;
[0032] The processing module is used to prune the redundant intermediate layers contained in the model to be processed according to the selection result to obtain a processed model.
[0033] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned model pruning method is implemented.
[0034] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0035] In the model pruning method provided in the present specification, the mutual information quantization contribution value of each intermediate layer is first obtained in the process of the model to be processed performing the reasoning task, wherein the mutual information quantization contribution value is used to characterize the proportion of the contribution information to reducing the uncertainty of the output result of the intermediate layer relative to the redundant information when the input data input to the intermediate layer is known, and the reasoning task includes a visual processing task, and then according to the mutual information quantization contribution value of each intermediate layer, the intermediate layers are sorted, and according to the sorting result, the redundant intermediate layers are selected from the intermediate layers, and according to the selection result, the redundant intermediate layers contained in the model to be processed are pruned to obtain the processed model.
[0036] It can be seen from the above method that the mutual information quantified contribution value of each intermediate layer contained in the model to be processed can be determined according to the input data and output results of each intermediate layer contained in the model to be processed. Therefore, the model to be processed can be pruned to remove the intermediate layers containing more redundant information while ensuring that the performance loss of the model to be processed is within a controllable range. This can reduce the number of parameters of the processed model, reduce the multiplication and addition operations in matrix operations, and avoid large-scale retraining of the processed model, thereby reducing the cost required to adjust large-scale deep learning models. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0038] Figure 1 A flowchart of a model pruning method provided in this specification;
[0039] Figure 2 A schematic diagram of a process for obtaining a mutual information quantization contribution value provided in this specification;
[0040] Figure 3 A flowchart of a task execution method provided in this specification;
[0041] Figure 4 A schematic diagram of a model pruning device provided in this specification;
[0042] Figure 5 A task execution device provided in this manual;
[0043] Figure 6 A method corresponding to the Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0045] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0046] At present, due to the large parameter scale of large-scale deep learning models, the efficiency of performing reasoning tasks through large-scale deep learning models is often low. Therefore, in order to improve the reasoning efficiency of large-scale deep learning models, it is usually necessary to compress and quantize the large-scale deep learning models. After the large-scale deep learning models are compressed and quantized, the large-scale deep learning models need to be retrained on a large scale, and may also cause the performance of the model to deteriorate. Therefore, it is particularly important to improve the reasoning efficiency of large-scale deep learning models without large-scale retraining.
[0047] Figure 1 A flow chart of a model pruning method provided in this specification includes the following steps:
[0048] S101: Obtain the quantitative mutual information contribution value of each intermediate layer in the process of executing the reasoning task of the model to be processed, wherein the quantitative mutual information contribution value is used to characterize the proportion of the contribution information to reducing the uncertainty of the output result of the intermediate layer relative to the redundant information when the input data input to the intermediate layer is known, and the reasoning task includes a visual processing task.
[0049] In this specification, the business platform can obtain the model to be processed, and determine the mutual information quantization contribution value of each intermediate layer based on the ratio of contribution information and redundant information obtained for reducing the uncertainty of the output result of the intermediate layer of each intermediate layer of the model to be processed when the input data input to the intermediate layer is known. Then, according to the mutual information quantization contribution value of each intermediate layer, the model to be processed can be pruned to remove the intermediate layer containing more redundant information, thereby reducing the number of parameters of the processed model and reducing the multiplication and addition operations in matrix operations, thereby reducing the computing resource requirements when performing reasoning tasks through the processed model, and reducing the occupancy of the graphics memory (GPU Memory) when performing reasoning tasks through the processed model, so as to allow the graphics processing unit (GPU) to use a larger batch size (batch size) for data processing, thereby improving the utilization of the GPU.
[0050] Among them, the above-mentioned model to be processed can be a model using a transformer Transformer architecture, such as: a Transformer-based bidirectional encoder representation (Bidirectional Encoder Representations from Transformers, BERT) model, a generative pre-trained Transformer (Generative Pre-trained Transformer, GPT) model, a Qwen model, etc.
[0051] The above intermediate layers may include an attention mechanism layer and a feedforward neural network layer.
[0052] In this specification, the execution entity used to implement the model pruning method can refer to a designated device such as a server set up on the business platform, or it can refer to a terminal device such as a desktop computer, a laptop computer, etc. For the sake of ease of description, the model pruning method provided in this specification is described below using the server as the execution entity as an example.
[0053] In this specification, after obtaining the model to be processed, the server can determine, for each intermediate layer included in the model to be processed, the input data input to the intermediate layer and the output result obtained by the intermediate layer in the process of performing the reasoning task through the model to be processed, and then determine the mutual information quantization contribution value of the intermediate layer according to the input data and output result of the intermediate layer, as shown in the following example: Figure 2 shown.
[0054] Figure 2 A schematic diagram of the process of obtaining the mutual information quantization contribution value provided in this specification.
[0055] Combination Figure 2 It can be seen that the server can first determine the mutual information between the input data and the output result of the intermediate layer. The mutual information here is used to characterize the amount of contribution information for reducing the uncertainty of the output result of the intermediate layer when the input data input to the intermediate layer is known. For details, please refer to the following formula:
[0056] I(X;Y)=H(X)-H(X|Y)
[0057] In the above formula, I(X;Y) is the mutual information between the input data X and the output result Y of the intermediate layer, H(X) is the entropy of the input data of the intermediate layer, and H(X|Y) is the conditional entropy of the input data of the intermediate layer when it is known that the output result of the intermediate layer is Y.
[0058] Furthermore, the server may determine the redundant information value between the input data and the output result according to the mutual information between the input data and the output result of the intermediate layer, the entropy of the input data of the intermediate layer, and the entropy of the output result of the intermediate layer. The redundant information value here is used to characterize the amount of redundant information. Specifically, the following formula may be referred to:
[0059] Redundant information value = H(X)-I(X;Y)-H(Y)
[0060] In the above formula, H(Y) is the entropy of the output result Y of the intermediate layer.
[0061] Furthermore, the server may determine the mutual information quantization contribution value of the intermediate layer according to the ratio between the mutual information amount of the intermediate layer and the redundant information value.
[0062] In actual application scenarios, the server can also input the input data and output results of each intermediate layer of the model to be processed into a preset evaluation model in sequence during the process of executing the reasoning task through the model to be processed, according to the order of the intermediate layers contained in the model to be processed, so that the evaluation model determines the quantitative mutual information contribution value of each intermediate layer according to the input data input into the intermediate layer, the output result obtained by the intermediate layer, and the correlation relationship between the intermediate layers before the intermediate layer.
[0063] The above-mentioned evaluation model may be a long short-term memory neural network model (Long Short-Term Memory network, LSTM).
[0064] In the above content, the association relationship between the intermediate layers located before the intermediate layer may refer to the long-term dependency relationship between the intermediate layers.
[0065] S102: sorting the intermediate layers according to the mutual information quantization contribution value of each intermediate layer, and selecting a redundant intermediate layer from the intermediate layers according to the sorting result.
[0066] S103: According to the selection result, pruning is performed on the redundant intermediate layers included in the model to be processed to obtain a processed model.
[0067] After determining the quantized mutual information contribution value of each intermediate layer contained in the model to be processed, the server can prioritize the intermediate layers according to the quantized mutual information contribution value of each intermediate layer, and select redundant intermediate layers from each intermediate layer according to the order of the intermediate layers after priority sorting, and prune the model to be processed based on the selection result to obtain the processed model.
[0068] In addition, the server can also determine whether to select each intermediate layer included in the model to be processed as a redundant intermediate layer according to whether the mutual information quantization contribution value of the intermediate layer is higher than a preset threshold.
[0069] In this specification, after selecting the redundant intermediate layer, the server can set the weights of the neuron nodes contained in the selected intermediate layer to 0, so as to prune the model to be processed and obtain the processed model. For ease of understanding, the following takes the selection of redundant intermediate layers from each intermediate layer and pruning according to whether the mutual information quantization contribution value of the intermediate layer is higher than a preset threshold as an example to describe the model pruning method in detail. For details, please refer to the following formula:
[0070]
[0071] In the above formula, W i That is, the weight matrix contained in the middle layer of the i-th layer, w i It is the mutual information quantization contribution value of the i-th intermediate layer, and ∈ is the preset threshold of the mutual information quantization contribution value.
[0072] In actual application scenarios, since the model after pruning may have performance loss, in order to reduce the performance loss of the processed model, the server can also obtain the first sample data after obtaining the processed model, and input the first sample data into the processed model to determine the performance parameters of the processed model when performing the reasoning task corresponding to the first sample data. Then, according to the deviation between the performance parameters of the processed model and the preset original performance parameters, the detection result of the processed model can be obtained, and when it is determined that the processed model has an abnormality according to the detection result of the processed model, the model to be processed is pruned again.
[0073] Among them, the above-mentioned performance parameters are used to reflect the performance of the processed model when performing the reasoning task on the first sample data, such as: accuracy, recall rate, average error, memory usage, BLEU score, etc.
[0074] The above-mentioned original performance parameters are performance parameters when the model to be processed performs the reasoning task corresponding to the first sample data.
[0075] In the above content, the method for the server to obtain the detection result for the processed model based on the performance parameters of the processed model can be to determine whether the deviation between the performance parameters of the processed model and the preset original performance parameters exceeds the preset difference threshold. If so, a detection result indicating that the processed model is abnormal can be obtained.
[0076] In addition, the server may also input the performance parameters of the processed model into a preset detection model to obtain a detection result of whether the processed model has an abnormality through the preset detection model.
[0077] In addition, the server can also obtain the second sample data and input the second sample data into the processed model to obtain output information, and then determine the target loss value based on the deviation between the output information and the actual output information corresponding to the second sample data, and use minimizing the target loss value as the optimization goal to fine-tune the model parameters of the processed model to obtain the adjusted model.
[0078] Among them, the above-mentioned target loss value is positively correlated with the deviation between the output information and the actual output information corresponding to the sample data.
[0079] In addition, in order to avoid adding too many parameters or adjusting the values of some model parameters too large in the process of fine-tuning the model parameters of the processed model, so that the task execution through the adjusted model is highly dependent on these model parameters, thereby reducing the generalization ability of the model, the server can also determine the supplementary loss value according to the model parameters of the processed model, and then determine the fusion loss value of the processed model according to the target loss value and the supplementary loss value of the processed model, and take minimizing the fusion loss value of the processed model as the optimization goal, and fine-tune the model parameters of the processed model to obtain the adjusted model.
[0080] Among them, the above-mentioned supplementary loss value increases with the complexity of the processed model. For example, when the more model parameters of the above-mentioned processed model, the larger the above-mentioned supplementary loss value is, and when the sum of squares of the model parameters of the above-mentioned processed model is larger, the above-mentioned supplementary loss value is larger.
[0081] In the above content, the method for the server to determine the fusion loss value of the processed model according to the target loss value and the supplementary loss value of the processed model can refer to the following formula:
[0082]
[0083] In the above formula, is the fusion loss value, is the target loss value, That is, it is the supplementary loss value, λ is the hyperparameter, and θ is the model parameter.
[0084] Among them, the above performance parameters are used to reflect the performance of the processed model when performing reasoning tasks on sample data, such as: accuracy, recall rate, mean error, memory usage, BLEU score, etc.
[0085] In the above content, the method by which the server obtains the detection result for the processed model based on the performance parameters of the processed model can be to determine whether the difference between the performance parameters of the processed model and the performance parameters of the model to be optimized exceeds a preset difference threshold. If so, a detection result indicating that the processed model is abnormal can be obtained.
[0086] In addition, the server can also input the performance parameters of the processed model into a preset detection model to obtain a detection result of whether the processed model has an abnormality through the preset detection model. It should be noted that the above-mentioned reasoning tasks may include visual processing tasks such as image recognition, target detection, and semantic segmentation.
[0087] It should be noted that in this specification, reasoning tasks may include visual processing tasks such as image recognition, object detection, semantic segmentation, etc.
[0088] Among them, when the above-mentioned reasoning task is an image recognition task, the server can input the sample image data into the model to be processed to obtain the quantitative contribution value of the mutual information of each intermediate layer of the model to be processed in the process of executing the reasoning task corresponding to the sample image data, and based on the quantitative contribution value of the mutual information of each intermediate layer of the model to be processed in the process of executing the reasoning task corresponding to the sample image data, prune the model to be processed to obtain the processed model, and then use the processed model to perform image recognition on the image data to be recognized to obtain the recognition result.
[0089] When the above-mentioned reasoning task is a target detection task, the server can input the sample image data into the model to be processed to obtain the quantitative contribution value of the mutual information of each intermediate layer of the model to be processed in the process of executing the reasoning task corresponding to the sample image data, and based on the quantitative contribution value of the mutual information of each intermediate layer of the model to be processed in the process of executing the reasoning task corresponding to the sample image data, prune the model to be processed to obtain a processed model, and then use the processed model to perform target detection on the image data to be detected to determine the category and position of each specific target (for example: road signs, pedestrians, vehicles, etc.) contained in the image data to be detected.
[0090] Of course, the above-mentioned reasoning tasks can also include other deep learning tasks such as natural language processing (for example, machine translation, dialogue generation, sentiment analysis), speech processing (for example, speech recognition, speech synthesis), etc.
[0091] It can be seen from the above method that the server can determine the mutual information quantification contribution value of each intermediate layer contained in the model to be processed based on the input data and output results of each intermediate layer included in the model to be processed, so that the model to be processed can be pruned based on the mutual information quantification contribution value of each intermediate layer while ensuring that the performance loss of the model to be processed is within a controllable range, so as to remove the intermediate layers containing more redundant information. In addition, the number of parameters of the processed model can be reduced to reduce the multiplication and addition operations in matrix operations, while avoiding large-scale retraining of the processed model, so as to reduce the cost required to adjust large-scale deep learning models.
[0092] For ease of understanding, the following describes in detail the process of executing tasks on the processed model obtained by the above model pruning method. Figure 3 shown.
[0093] Figure 3 A flowchart of a task execution method provided in this specification includes the following steps:
[0094] S301: Acquire task data to be executed, where the task data to be executed includes at least one of text data, picture data, and video data.
[0095] S302: Inputting the task data to be executed into the processed model to obtain an output result, and executing the task according to the output result. The processed model is obtained after pruning by the above-mentioned model pruning method.
[0096] In this specification, the server can obtain the task data to be executed, and input the task data to be executed into the processed model to obtain the output result, and execute the task according to the output result.
[0097] The above-mentioned task data to be executed may include: at least one of text data, picture data, video data and other multimodal data.
[0098] The processed model mentioned above can be obtained by pruning using the above model pruning method.
[0099] From the above content, it can be seen that the server can apply the processed model to deep learning tasks such as natural language processing and computer vision without the need for large-scale retraining of the processed model, which can effectively improve reasoning efficiency and reduce hardware resource consumption.
[0100] The above are one or more implementation model pruning and service execution methods of this specification. Based on the same idea, this specification also provides corresponding model pruning and service execution devices, such as Figure 4 , Figure 5 shown.
[0101] Figure 4 A schematic diagram of a model pruning device provided in this specification includes:
[0102] An acquisition module 401 is used to acquire a mutual information quantization contribution value of each intermediate layer in a process of executing a reasoning task by a to-be-processed model, wherein the mutual information quantization contribution value is used to characterize the proportion of contribution information for reducing the uncertainty of an output result of the intermediate layer relative to redundant information when input data input to the intermediate layer is known, and the reasoning task includes a visual processing task;
[0103] A selection module 402 is used to sort the intermediate layers according to the mutual information quantization contribution value of each intermediate layer, and select a redundant intermediate layer from the intermediate layers according to the sorting result;
[0104] The processing module 403 is used to prune the redundant intermediate layers included in the model to be processed according to the selection result to obtain a processed model.
[0105] Optionally, the acquisition module 401 is specifically used to determine, for each intermediate layer included in the model to be processed, the mutual information of the intermediate layer based on the input data and output results of the intermediate layer; determine the redundant information value between the input data and the output result of the intermediate layer based on the mutual information, the entropy of the input data of the intermediate layer, and the entropy of the output result of the intermediate layer; and determine the quantitative contribution value of the mutual information of the intermediate layer based on the ratio of the mutual information and the redundant information value.
[0106] Optionally, the acquisition module 401 is specifically used to determine the entropy value of the input data of the intermediate layer and the conditional entropy value of the input data of the intermediate layer when the output result of the intermediate layer is known; determine the mutual information of the intermediate layer based on the entropy value and the conditional entropy value.
[0107] Optionally, the acquisition module 401 is specifically used to input the input data and output results of each intermediate layer into a preset evaluation model in sequence according to the order of the intermediate layers contained in the model to be processed, so that the evaluation model determines the quantitative mutual information contribution value of each intermediate layer according to the input data and output results of the intermediate layer and the correlation relationship between the intermediate layers before the intermediate layer.
[0108] Optionally, the device further includes: a detection module 404;
[0109] The detection module 404 is used to input the preset first sample data into the processed model to determine the performance parameters of the processed model when executing the reasoning task corresponding to the first sample data; obtain the detection result for the processed model according to the deviation between the performance parameters and the preset original performance parameters, and when it is determined according to the detection result that the processed model has an abnormality, re-prune the model to be processed, and the original performance parameters are the performance parameters of the model to be processed when executing the reasoning task corresponding to the first sample data.
[0110] Optionally, the device further includes: an adjustment module 405;
[0111] The adjustment module 405 is used to input the preset second sample data into the processed model to obtain output information; determine the target loss value according to the deviation between the output information and the actual output information of the second sample data, and the target loss value is positively correlated with the deviation; and fine-tune the model parameters of the processed model with minimizing the target loss value as the optimization goal to obtain an adjusted model.
[0112] Figure 5 A task execution device provided in this specification includes:
[0113] The acquisition module 501 is used to acquire the task data to be executed, wherein the task data to be executed includes: at least one of text data, picture data, and video data;
[0114] The execution module 502 is used to input the task data to be executed into the processed model to obtain an output result, and execute the task according to the output result. The processed model is obtained after pruning by the above-mentioned model pruning method.
[0115] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A model pruning method provided.
[0116] This manual also provides Figure 6 The one shown corresponds to Figure 1 A schematic diagram of the electronic device. Figure 6 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to the software implementation, this specification does not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0117] For the improvement of a technology, it can be clearly distinguished whether it is a hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0118] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0119] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0120] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0121] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0122] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0123] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0125] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0126] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0127] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0128] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0129] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0130] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0131] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0132] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A model pruning method, characterized in that: include: Obtaining a mutual information quantization contribution value of each intermediate layer in a process of executing a reasoning task by a to-be-processed model, wherein the mutual information quantization contribution value is used to characterize the proportion of contribution information to reducing uncertainty of an output result of the intermediate layer relative to redundant information when input data input to the intermediate layer is known, wherein the reasoning task includes a visual processing task; Sorting the intermediate layers according to the mutual information quantization contribution value of each intermediate layer, and selecting a redundant intermediate layer from the intermediate layers according to the sorting result; According to the selection result, the redundant intermediate layers included in the model to be processed are pruned to obtain a processed model.
2. The method according to claim 1, characterized in that Obtain the mutual information quantification contribution value of each intermediate layer in the process of executing the inference task of the model to be processed, including: For each intermediate layer included in the model to be processed, determining the mutual information amount of the intermediate layer according to the input data and output result of the intermediate layer; Determine the redundant information value between the input data and the output result of the intermediate layer according to the mutual information amount, the entropy of the input data of the intermediate layer, and the entropy of the output result of the intermediate layer; The mutual information quantization contribution value of the intermediate layer is determined according to the ratio of the mutual information amount to the redundant information value.
3. The method according to claim 2, characterized in that According to the input data and output result of the intermediate layer, the mutual information of the intermediate layer is determined, specifically including: Determine the entropy value of the input data of the intermediate layer and the conditional entropy value of the input data of the intermediate layer when the output result of the intermediate layer is known; The mutual information of the intermediate layer is determined according to the entropy value and the conditional entropy value.
4. The method according to claim 1, characterized in that Obtain the mutual information quantification contribution value of each intermediate layer in the process of executing the inference task of the model to be processed, including: According to the order of the intermediate layers contained in the model to be processed, the input data and output results of each intermediate layer are input into the preset evaluation model in sequence, so that the evaluation model determines the quantitative contribution value of the mutual information of each intermediate layer according to the input data and output results of the intermediate layer and the correlation relationship between the intermediate layers before the intermediate layer.
5. The method according to claim 1, characterized in that The method further comprises: Inputting the preset first sample data into the processed model to determine the performance parameters of the processed model when performing the reasoning task corresponding to the first sample data; Based on the deviation between the performance parameters and the preset original performance parameters, a detection result for the processed model is obtained, and when it is determined that the processed model has an abnormality based on the detection result, the model to be processed is pruned again, and the original performance parameters are the performance parameters of the model to be processed when executing the reasoning task corresponding to the first sample data.
6. The method according to claim 1, characterized in that The method further comprises: Inputting the preset second sample data into the processed model to obtain output information; Determine a target loss value according to a deviation between the output information and actual output information of the second sample data, wherein the target loss value is positively correlated with the deviation; Taking minimizing the target loss value as the optimization goal, the model parameters of the processed model are fine-tuned to obtain an adjusted model.
7. A task execution method, characterized in that: include: Acquire the task data to be executed, wherein the task data to be executed includes: at least one of text data, picture data, and video data; The task data to be executed is input into a processed model to obtain an output result, and the task is executed according to the output result. The processed model is obtained by pruning according to the method described in any one of claims 1 to 6.
8. A model pruning device, characterized in that: include: An acquisition module, used to acquire a mutual information quantized contribution value of each intermediate layer in the process of executing a reasoning task by a to-be-processed model, wherein the mutual information quantized contribution value is used to determine the proportion of contribution information relative to redundant information for reducing uncertainty of an output result of the intermediate layer when input data input to the intermediate layer is known, and the reasoning task includes a visual processing task; A selection module, used to sort the intermediate layers according to the mutual information quantification contribution value of each intermediate layer, and select a redundant intermediate layer from the intermediate layers according to the sorting result; The processing module is used to prune the redundant intermediate layers contained in the model to be processed according to the selection result to obtain a processed model.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Convolutional neural network lightweight pruning method and system based on multi-view features
CN117217281A
Joint pruning and quantization scheme for deep neural networks
US20210089922A1
Model processing method, federated learning method, and related device
WO2023279975A1
Cited By
Traffic flow prediction method and device based on model lightweight, and medium
CN121438577A