A network pruning method, device, terminal, and computer-readable storage medium

By generating preprocessing neural networks and pruning based on the loss value, the problem of poor performance of deep neural networks after pruning is solved, and an algorithmic indicator that maintains excellent on small models is realized, suitable for image, text and speech processing.

CN118821893BActive Publication Date: 2025-07-04ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411307364.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-07-04
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

In the prior art, the deep neural network after pruning has poor performance and it is difficult to maintain excellent algorithmic indicators on small models.

Method used

By obtaining the data information and processing module of the initial neural network, disconnecting the nodes to generate a preprocessing neural network, determining the loss value of the operating node, and pruning based on the greedy strategy and preset pruning rate to generate the target neural network.

Benefits of technology

The impact of pruning processing on the performance of the target neural network is reduced, and the algorithmic indicators of the small model are improved, which are suitable for image processing, text processing and speech processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118821893B_ABST
    Figure CN118821893B_ABST
Patent Text Reader

Abstract

The present invention provides a network pruning method, apparatus, terminal, and computer-readable storage medium. The network pruning method includes: obtaining data information and an initial neural network; the initial neural network includes a plurality of processing modules, and the processing modules have nodes; disconnecting the connections between the operation nodes and the nodes in other processing modules to obtain a preprocessing neural network corresponding to the operation nodes; inputting the data information into the initial neural network and the preprocessing neural network corresponding to the operation nodes respectively to determine the loss value of the operation nodes; based on the loss values of the nodes in the current processing module, performing pruning processing on the nodes in the current processing module to obtain a target neural network. This application determines the contribution degree of the operation nodes to the initial neural network according to the difference between the processing results corresponding to the two neural networks, so as to prune the initial neural network according to the contribution degree of the operation nodes to the initial neural network, and reduce the impact of the pruning processing on the performance of the target neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural networks, and in particular, to a network pruning method, device, terminal, and computer-readable storage medium. Background Art

[0002] Deep neural networks have been applied to various fields of life and industry. When using a deep neural network model, generally, a larger model can obtain better algorithm metrics, but it often requires higher inference costs or higher requirements for the inference hardware platform. In order to obtain a model more suitable for application deployment, a smaller model structure is usually used. If a small model is directly trained, the algorithm metrics will drop significantly. In order to obtain relatively excellent algorithm metrics on a small model, one of the current mainstream processing methods is to perform pruning and fine-tuning based on a large-parameter model. However, the performance of the deep neural network after pruning by the current network pruning method is poor. Summary of the Invention

[0003] The main technical problem to be solved by the present invention is to provide a network pruning method, device, terminal, and computer-readable storage medium, so as to solve the problem of poor performance of the pruned model in the prior art.

[0004] To solve the above technical problem, the first technical solution adopted by the present invention is: to provide a network pruning method, which is applicable to any one of the application scenarios of image processing, text processing, and speech processing. The network pruning method includes:

[0005] Obtain data information and an initial neural network; the initial neural network includes a plurality of processing modules, and the processing module has nodes;

[0006] Take each processing module in the initial neural network as the current processing module in turn, take each node in the current processing module as an operation node, and disconnect the connection between the operation node and each node in other processing modules to obtain a preprocessing neural network corresponding to the operation node;

[0007] Input the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively, and determine the loss value of the operation node;

[0008] Based on the loss values of the nodes in the current processing module, perform pruning processing on the nodes in the current processing module to obtain a target neural network.

[0009] Among them, the network pruning method includes:

[0010] Start from the first processing module of the initial neural network, take each processing module as the current processing module in turn in the order from front to back, and perform pruning on the processing module through a channel selection method based on a greedy strategy until the nodes in the last processing module are pruned.

[0011] Among them, the network pruning method includes:

[0012] Using the training sample set to train the pruned target neural network to optimize the parameters of the pruned target neural network.

[0013] Among them, disconnecting the connection between the operation node and each node in other processing modules to obtain the preprocessing neural network corresponding to the operation node includes:

[0014] Setting the connection parameters between the operation node in the current processing module and each node in the adjacent connected processing module to zero to obtain the preprocessing neural network corresponding to the operation node; the processing modules before the current processing module in the preprocessing neural network have completed pruning processing.

[0015] Among them, inputting the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively to determine the loss value of the operation node includes:

[0016] Processing the data information through the initial neural network to obtain the first processing result of the data information;

[0017] Processing the data information through the preprocessing neural network to obtain the second processing result of the data information;

[0018] Taking the difference between the first processing result and the second processing result corresponding to the data information as the loss value of the operation node.

[0019] Among them, the initial neural network has a preset pruning rate;

[0020] Based on the loss values of the nodes in the current processing module, pruning the nodes in the current processing module to obtain the target neural network includes:

[0021] Sorting the nodes included in the current processing module in ascending order according to the loss values;

[0022] According to the preset pruning rate of the initial neural network, selecting a preset number of nodes with the top rankings as the nodes to be pruned;

[0023] Disconnecting the connections between the nodes to be pruned in the current processing module and other nodes, and pruning the nodes to be pruned to obtain the target neural network.

[0024] Among them, based on the loss values of the nodes in the current processing module, pruning the nodes in the current processing module to obtain the target neural network includes:

[0025] Disconnecting the connection between the node corresponding to the minimum loss value in the current processing module and the nodes in other processing modules, and pruning the node corresponding to the minimum loss value to obtain the target neural network.

[0026] Among them, the network pruning method further includes:

[0027] Pre-train the initial neural network; during the training process, each node in the initial neural network is connected with a Dropout layer; after the initial neural network is trained, remove the Dropout layer.

[0028] To solve the above technical problems, the second technical solution adopted by the present invention is: to provide a network pruning device, which is applicable to any application scenario in image processing, text processing, and speech processing. The network pruning device includes:

[0029] An acquisition module, configured to acquire data information and an initial neural network; the initial neural network includes a plurality of processing modules, and the processing modules have nodes;

[0030] A processing module, configured to sequentially use each processing module in the initial neural network as the current processing module, and use each node in the current processing module as an operation node, and disconnect the connection between the operation node and each node in other processing modules to obtain a preprocessing neural network corresponding to the operation node;

[0031] A calculation module, configured to input the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively, and determine the loss value of the operation node;

[0032] A pruning module, configured to perform pruning processing on the nodes in the current processing module based on the loss values of the nodes in the current processing module to obtain a target neural network.

[0033] To solve the above technical problems, the third technical solution adopted by the present invention is: to provide a terminal, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. The processor is configured to execute program data to implement the steps in the network pruning method as described above.

[0034] To solve the above technical problems, the fourth technical solution adopted by the present invention is: to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the network pruning method as described above.

[0035] The beneficial effects of the present invention are as follows: Different from the prior art, a network pruning method, apparatus, terminal, and computer-readable storage medium are provided. The network pruning method includes: obtaining data information and an initial neural network; the initial neural network includes multiple processing modules, and each processing module has nodes; taking each processing module in the initial neural network as the current processing module in turn, taking each node in the current processing module as an operation node, and disconnecting the connection between the operation node and each node in other processing modules to obtain a preprocessing neural network corresponding to the operation node; inputting the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively, and determining the loss value of the operation node; based on the loss values of the nodes in the current processing module, pruning the nodes in the current processing module to obtain a target neural network. In this application, a preprocessing neural network corresponding to the operation node is generated by preprocessing the initial neural network, and then the same data information is processed based on the preprocessing neural network and the initial neural network respectively. The contribution degree of the operation node to the initial neural network is determined according to the difference between the processing results corresponding to the two neural networks, so as to prune the initial neural network according to the contribution degree of the operation node to the initial neural network, and reduce the impact of the pruning process on the performance of the target neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 is a flowchart of the network pruning method provided by the present invention;

[0038] Figure 2 is a structural diagram of a specific embodiment of the initial neural network provided by the present invention;

[0039] Figure 3 is a connection structure diagram between the nodes in the current processing module and the nodes in other processing modules provided by the present invention;

[0040] Figure 4 is a framework diagram of an embodiment of the network pruning apparatus provided by the present invention;

[0041] Figure 5 is a framework diagram of an embodiment of the terminal provided by the present invention;

[0042] Figure 6 is a framework diagram of an embodiment of the computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will combine with the accompanying drawings of the specification to elaborate in detail on the solutions of the embodiments of the present application.

[0044] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0045] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. Furthermore, "multiple" in this article means two or more than two.

[0046] To enable those skilled in the art to better understand the technical solutions of the present invention, the following will further elaborate in detail on a network pruning method provided by the present invention in combination with the drawings and specific implementation manners.

[0047] Unless otherwise defined, all technical and scientific terms used in this article have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in this article are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0048] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0049] Greedy strategy: Always make locally optimal choices. For a specific optimization problem, the greedy strategy guarantees that it will definitely converge to the global optimal solution.

[0050] Dropout layer: Refers to that during the training process of a deep learning network, for neural network units, they are temporarily discarded from the network with a certain probability.

[0051] The network pruning method provided by the embodiments of the present application can be implemented independently by a server or a terminal, or jointly implemented by the server and the terminal. In some embodiments, the terminal or the server can implement the network pruning method provided by the embodiments of the present application by running a computer program. For example, the computer program can be a native program or a software module in an operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a client supporting a virtual scenario, such as a game APP; it can also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module or plug-in.

[0052] The following takes the server implementation as an example to illustrate the network pruning method provided by the embodiments of the present application.

[0053] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the network pruning method provided by the present invention.

[0054] A network pruning method is provided in this embodiment. The network pruning method is applicable to any application scenario in image processing, text processing, and speech processing. The network pruning method includes the following steps.

[0055] S1: Obtain data information and an initial neural network; the initial neural network includes multiple processing modules, and the processing modules have nodes.

[0056] S2: Take each processing module in the initial neural network as the current processing module in turn, take each node in the current processing module as an operation node, and disconnect the connection between the operation node and each node in other processing modules to obtain a preprocessing neural network corresponding to the operation node.

[0057] S3: Input the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively, and determine the loss value of the operation node.

[0058] S4: Based on the loss values of the nodes in the current processing module, perform pruning processing on the nodes in the current processing module to obtain a target neural network.

[0059] In the network pruning method of this embodiment, a preprocessing neural network corresponding to the operation node is generated by preprocessing the initial neural network, and then the same data information is processed based on the preprocessing neural network and the initial neural network respectively. The contribution degree of the operation node to the initial neural network is determined according to the difference between the processing results corresponding to the two neural networks, so as to prune the initial neural network according to the contribution degree of the operation node to the initial neural network, and reduce the impact of the pruning processing on the performance of the target neural network.

[0060] Specifically, the specific implementation of obtaining data information and the initial neural network in step S1 is as follows.

[0061] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a specific embodiment of the initial neural network provided by the present invention.

[0062] In one embodiment, the initial neural network includes multiple processing modules, and the processing modules are specifically network layers. The initial neural network can be a convolutional layer, a normalization layer, an activation layer, a pooling layer, a fully connected layer, etc. Among them, the processing modules can be a convolutional layer, a normalization layer, and a fully connected layer, and the convolutional layer, the normalization layer, and the fully connected layer all contain training parameters. In this embodiment, by reducing the number of parameters in the convolutional layer, the normalization layer, and / or the fully connected layer, pruning of the initial neural network is realized.

[0063] In one embodiment, the initial neural network can be a Transformer model, and the processing modules can be a multi-layer perceptron module (MLP), a self-attention module, etc. Among them, the parameters in the Transformer model include MLP parameters, linear transformation parameters of Q, K, V, etc. The connection between the multi-layer perceptron module and the self-attention module can be the connection between fully connected layers.

[0064] Among them, the fully connected layer, the normalization layer, and the convolutional layer each include multiple channels, and each channel serves as a node. In this embodiment, pruning of the initial neural network is realized by pruning the channels.

[0065] In a specific embodiment, the initial neural network can be an object detection model, an object classification model, a language processing model, etc.

[0066] In a specific embodiment, when the initial neural network is an image processing model, the data information can be image information; when the initial neural network is a text processing model, the data information can be text information; when the initial neural network is a language processing model, the data information can be speech information. In other embodiments, the data information is input information that can be processed by the initial neural network.

[0067] The initial neural network provided in this embodiment is a basic network model with a large number of parameters.

[0068] In one embodiment, the initial neural network is pre-trained; during the training process, each node in the initial neural network is connected to a Dropout layer; after the initial neural network is trained, the Dropout layer is removed.

[0069] In a specific embodiment, a training data set is obtained. The training data set includes multiple training samples. The initial neural network is trained with the training samples to optimize the connection parameters between each node and other nodes in each processing module of the initial neural network, thereby improving the processing effect of the initial neural network. During the training process, each node is connected to a Dropout layer. By means of the Dropout layer, overfitting of the initial neural network can be avoided, so as to enhance the robustness of the initial neural network, and to a certain extent, the reduction of the model metrics after pruning the initial neural network can also be reduced.

[0070] Specifically, the specific implementation manner of disconnecting the connection between the operation node and each node in other processing modules to obtain the preprocessing neural network corresponding to the operation node in step S2 is as follows.

[0071] In one embodiment, each processing module in the preprocessing neural network is traversed, and each processing module in the initial neural network is sequentially used as the current processing module. Each node in the current processing module is traversed, and each node in the current processing module is respectively used as the operation node.

[0072] In a specific embodiment, the network pruning method further includes starting from the first processing module of the initial neural network, sequentially using each processing module as the current processing module in the order from front to back, and pruning the processing module by a channel selection method based on a greedy strategy until the nodes in the last processing module are pruned completely.

[0073] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the connection structure between the nodes in the current processing module and the nodes in other processing modules provided by the present invention.

[0074] In a specific embodiment, the connection parameters between the operation node in the current processing module m and the nodes in the previous adjacent processing module m - 1 and the connection parameters between the nodes in the next adjacent processing module m + 1 are both set to zero to obtain the preprocessing neural network corresponding to the operation node; each processing module before the current processing module in the preprocessing neural network has completed the pruning process.

[0075] In one embodiment, the network pruning method further includes starting from the last processing module of the initial neural network, sequentially using each processing module as the current processing module in the order from back to front, and pruning the processing module by a channel selection method based on a greedy strategy until the nodes in the first processing module are pruned completely.

[0076] In a specific embodiment, the connection parameters between the operation nodes in the current processing module m and the nodes of the adjacent subsequent processing module m + 1, and the connection parameters between the nodes of the adjacent previous processing module m - 1 are all set to zero, obtaining a preprocessing neural network corresponding to the operation nodes.

[0077] In an embodiment, the subsequent processing module and the previous processing module after the current processing module in the preprocessing neural network are not pruned.

[0078] In an embodiment, the subsequent processing module after the current processing module in the preprocessing neural network has completed pruning.

[0079] Among them, in the preprocessing neural network corresponding to the operation nodes, the connection parameters between the operation nodes and the nodes in the adjacent previous processing module, and the connection parameters between the operation nodes and the nodes in the adjacent subsequent processing module are all zero.

[0080] Specifically, in step S3, the data information is respectively input into the initial neural network and the preprocessing neural network corresponding to the operation nodes, and the specific implementation manner of determining the loss value of the operation nodes is as follows.

[0081] In an embodiment, the data information is processed by the initial neural network to obtain a first processing result of the data information; the data information is processed by the preprocessing neural network to obtain a second processing result of the data information; the difference between the first processing result and the second processing result corresponding to the data information is used as the loss value of the operation nodes.

[0082] In an embodiment, multiple data information is processed by the initial neural network to obtain first processing results of the respective data information; multiple data information is processed by the preprocessing neural network to obtain second processing results of the respective data information, and the differences between the first processing results and the second processing results corresponding to the respective data information are summed and averaged as the loss value of the operation nodes.

[0083] Through the above steps, the loss values corresponding to the respective nodes in the current processing module can be obtained.

[0084] Specifically, in step S4, based on the loss values of the nodes in the current processing module, the specific implementation manner of pruning the nodes in the current processing module to obtain the target neural network is as follows.

[0085] Among them, the initial neural network has a preset pruning rate.

[0086] In one embodiment, the connection between the node corresponding to the minimum loss value in the current processing module and the nodes in other processing modules is disconnected, and the node corresponding to the minimum loss value is used as a node to be pruned. The node to be pruned is pruned to obtain a target neural network.

[0087] In a specific embodiment, the nodes to be pruned are screened out according to the loss values ​​corresponding to the nodes in the current processing module.

[0088] (Formula 1)

[0089] Where: n∈S m ∪S,m i Represents the i-th node in the m-th processing module.

[0090] In one embodiment, the nodes contained in the current processing module are sorted in order from small to large according to the loss value; according to the preset pruning rate of the initial neural network, a preset number of nodes with a high ranking are selected as nodes to be pruned; the connection between the nodes to be pruned and other nodes in the current processing module is disconnected, and the nodes to be pruned are pruned to obtain the target neural network. For example, the preset pruning rate is 20%, and when the current processing module contains 10 nodes, 2 nodes in the current processing module are removed. When the preset pruning rate of the initial neural network is 20%, the nodes in each processing module can be pruned by 20%.

[0091] Among them, the smaller the loss value corresponding to the node in the processing module, the smaller the contribution of the node to the initial neural network, and pruning the node will have a smaller impact on the performance of the neural network.

[0092] Through the above steps, the nodes to be pruned in each processing module can be determined, and the nodes to be pruned can be pruned and the connections between the nodes to be pruned and other nodes can be disconnected to obtain the target neural network.

[0093] In one embodiment, the pruned target neural network is trained using the training sample set to optimize the parameters of the pruned target neural network. Specifically, during the training process, each node in the initial neural network of the pruned target neural network is connected with a Dropout layer; after the initial neural network training is completed, the Dropout layer is removed.

[0094] The network pruning method provided in this embodiment generates a preprocessed neural network corresponding to an operation node by preprocessing an initial neural network, then processes the same data information based on the preprocessed neural network and the initial neural network respectively, and determines the contribution degree of the operation node to the initial neural network according to the difference between the processing results corresponding to the two neural networks, so as to prune the initial neural network according to the contribution degree of the operation node to the initial neural network and reduce the impact of the pruning process on the performance of the target neural network.

[0095] Please refer to Figure 4 , Figure 4 which is a schematic framework diagram of an embodiment of the network pruning device provided by the present invention.

[0096] This embodiment provides a network pruning device 60. The network pruning device 60 is applicable to any application scenario among image processing, text processing, and speech processing. The network pruning device 60 includes an acquisition module 61, a processing module 62, a calculation module 63, and a pruning module 64.

[0097] The acquisition module 61 is used to acquire data information and an initial neural network; the initial neural network includes a plurality of processing modules, and the processing module has nodes.

[0098] The processing module 62 is used to sequentially use each processing module in the initial neural network as the current processing module, use each node in the current processing module as an operation node respectively, and disconnect the connection between the operation node and each node in other processing modules to obtain a preprocessed neural network corresponding to the operation node.

[0099] The calculation module 63 is used to input the data information into the initial neural network and the preprocessed neural network corresponding to the operation node respectively, and determine the loss value of the operation node.

[0100] The pruning module 64 is used to prune the nodes in the current processing module based on the loss values of the nodes in the current processing module to obtain a target neural network.

[0101] The network pruning device provided in this embodiment generates a preprocessed neural network corresponding to an operation node by preprocessing an initial neural network, then processes the same data information based on the preprocessed neural network and the initial neural network respectively, and determines the contribution degree of the operation node to the initial neural network according to the difference between the processing results corresponding to the two neural networks, so as to prune the initial neural network according to the contribution degree of the operation node to the initial neural network and reduce the impact of the pruning process on the performance of the target neural network.

[0102] Please refer to Figure 5 , Figure 5 which is a schematic framework diagram of an embodiment of the terminal provided by the present invention.

[0103] The terminal 80 provided in this embodiment includes a memory 81 and a processor 82 which are coupled to each other. The processor 82 is configured to execute program instructions stored in the memory 81 to implement the steps of any of the above-described network pruning method embodiments. In a specific implementation scenario, the terminal 80 may include, but is not limited to, a microcomputer, a server. In addition, the terminal 80 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.

[0104] Specifically, the processor 82 is configured to control itself and the memory 81 to implement the steps of any of the above-described network pruning method embodiments. The processor 82 may also be referred to as a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip having signal processing capabilities. The processor 82 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 82 may be implemented jointly by integrated circuit chips.

[0105] In the above solution, the network pruning method is applicable to any of the application scenarios of image processing, text processing, and speech processing. The network pruning method includes: obtaining data information and an initial neural network; the initial neural network includes a plurality of processing modules, and the processing module has nodes; sequentially taking each processing module in the initial neural network as the current processing module, taking each node in the current processing module as an operation node, and disconnecting the connection between the operation node and each node in other processing modules to obtain a preprocessing neural network corresponding to the operation node; inputting the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively, and determining the loss value of the operation node; based on the loss values of the nodes in the current processing module, pruning the nodes in the current processing module to obtain a target neural network.

[0106] Please refer to Figure 6 , Figure 6 which is a framework schematic diagram of an embodiment of a computer-readable storage medium provided by the present invention.

[0107] The computer-readable storage medium 90 provided in this embodiment stores program instructions 901 that can be run by a processor. The program instructions 901 are configured to implement the steps of any of the above-described network pruning method embodiments.

[0108] In the above solution, the network pruning method is applicable to any application scenario in image processing, text processing, and speech processing. The network pruning method includes: obtaining data information and an initial neural network; the initial neural network includes multiple processing modules, and the processing modules have nodes; taking each processing module in the initial neural network as the current processing module in turn, taking each node in the current processing module as an operation node, and disconnecting the connection between the operation node and each node in other processing modules to obtain a preprocessing neural network corresponding to the operation node; inputting the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively to determine the loss value of the operation node; based on the loss values of the nodes in the current processing module, pruning the nodes in the current processing module to obtain a target neural network.

[0109] In some embodiments, the functions or modules included in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the method embodiments above. The specific implementation can refer to the description of the method embodiments above. For the sake of brevity, it will not be repeated here.

[0110] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0111] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0112] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0113] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0114] The above are only the embodiments of the present invention, and do not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A network pruning method, characterized in that, The network pruning method is applicable to any application scenario in image processing, text processing, and speech processing, and the network pruning method includes: Acquire data information and an initial neural network; the initial neural network includes a plurality of processing modules, and the processing modules have nodes; Using each of the processing modules in the initial neural network as a current processing module in turn, using each of the nodes in the current processing module as an operation node, disconnecting the operation node from each of the nodes in other processing modules to obtain a preprocessing neural network corresponding to the operation node; Inputting the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively, and determining the loss value of the operation node; Based on the loss value of each node in the current processing module, pruning the nodes in the current processing module to obtain a target neural network; The disconnecting of the operation node from the nodes in other processing modules to obtain a preprocessing neural network corresponding to the operation node includes: The connection parameters between the operation node in the current processing module and each of the nodes of the adjacently connected processing modules are set to zero to obtain the preprocessing neural network corresponding to the operation node.

2. The network pruning method according to claim 1, wherein The network pruning method comprises: Starting from the first processing module of the initial neural network, each processing module is used as the current processing module in order from front to back, and the processing modules are pruned through a channel selection method based on a greedy strategy until the nodes in the last processing module are pruned.

3. The network pruning method according to claim 1, characterized in that The network pruning method comprises: The pruned target neural network is trained using the training sample set to optimize the parameters of the pruned target neural network.

4. The network pruning method according to any one of claims 1 to 3, characterized in that, The processing module before the current processing module in the preprocessing neural network has completed pruning processing.

5. The network pruning method according to claim 4, characterized in that: The step of inputting the data information into the initial neural network and the preprocessing neural network corresponding to the operation node, respectively, and determining the loss value of the operation node comprises: Processing the data information through the initial neural network to obtain a first processing result of the data information; Processing the data information by the preprocessing neural network to obtain a second processing result of the data information; The difference between the first processing result and the second processing result corresponding to the data information is used as the loss value of the operation node.

6. The network pruning method according to claim 1, wherein The initial neural network has a preset pruning rate; The step of pruning the nodes in the current processing module based on the loss value of each node in the current processing module to obtain a target neural network includes: Sort the nodes included in the current processing module in ascending order according to the loss values; According to a preset pruning rate of the initial neural network, selecting a preset number of the nodes ranked at the top as nodes to be pruned; Disconnect the connection between the node to be pruned and other nodes in the current processing module, and perform pruning processing on the node to be pruned to obtain the target neural network.

7. The network pruning method according to claim 1, wherein: The pruning of the nodes in the current processing module based on the loss values of the nodes in the current processing module to obtain a target neural network includes: Disconnect the connection between the node corresponding to the minimum loss value in the current processing module and the nodes in other processing modules, and perform pruning processing on the node corresponding to the minimum loss value to obtain the target neural network.

8. The network pruning method according to claim 1, wherein The network pruning method further includes: Pre-training the initial neural network; during the training process, each node in the initial neural network is connected with a Dropout layer; after the initial neural network is trained, remove the Dropout layer.

9. A network pruning device, characterized in that, The network pruning device is applicable to any one of the application scenarios of image processing, text processing, and speech processing. The network pruning device includes: An acquisition module, configured to acquire data information and an initial neural network; the initial neural network includes a plurality of processing modules, and the processing module has nodes; A processing module, configured to sequentially use each of the processing modules in the initial neural network as the current processing module, and use each of the nodes in the current processing module as an operation node, and disconnect the connection between the operation node and each of the nodes in other processing modules to obtain a preprocessing neural network corresponding to the operation node; and is further configured to set the connection parameters between the operation node in the current processing module and each of the nodes in the adjacent connected processing module to zero to obtain the preprocessing neural network corresponding to the operation node; A calculation module, configured to input the data information into the initial neural network and the preprocessing neural network corresponding to the operation node respectively, and determine the loss value of the operation node; A pruning module, configured to perform pruning processing on the nodes in the current processing module based on the loss values of the nodes in the current processing module to obtain a target neural network.

10. A terminal, characterized in that, The terminal includes a memory, a processor, and a computer program stored in the memory and running on the processor. The processor is configured to execute program data to implement the steps in the network pruning method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps in the network pruning method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Data processing method and device and storage medium

    CN111062477A

  • Neural network automatic pruning method and device and electronic equipment

    CN111967591A