Method and apparatus for training a neural network for image recognition
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2021-06-15
- Publication Date
- 2026-08-07
AI Technical Summary
随着人工神经网络的学习量增大,构成人工神经网络的连接性可能变得复杂
Smart Images

Figure CN114358274B_ABST
Abstract
Description
[0001] This application claims the benefit of Korean Patent Application No. 10-2020-0132151, filed on October 13, 2020, with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical Field
[0002] The following description relates to methods and apparatus for training neural networks for image recognition. Background Technology
[0003] Artificial neural networks can be computational architectures. Using neural network devices, input data can be analyzed and useful information can be extracted.
[0004] Artificial neural network devices can utilize large amounts of computation to process complex input data. However, as the learning load of an artificial neural network increases, the connectivity that constitutes it can become more complex. Furthermore, while the accuracy of old learning data increases, the confidence of estimates for new data may decrease. In other words, overfitting may occur. Additionally, the increased complexity of the artificial neural network may lead to an excessively large memory allocation, potentially causing problems with miniaturization and commercialization. Summary of the Invention
[0005] This summary is provided to introduce, in a simplified form, the selection of concepts further described in the following detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.
[0006] In one general aspect, a method for training a neural network for image recognition includes: receiving an input image set; executing a pre-trained neural network as input to the input image set to obtain result values inferred using the pre-trained neural network, and obtaining a first task accuracy of the pre-trained neural network based on predetermined result values corresponding to the input image set and the result values inferred using the pre-trained neural network; pruning the neural network based on channel units by adjusting the weights between nodes of multiple channels based on preset learning weights and based on channel-wise pruning parameters corresponding to channels of each of multiple layers of the pre-trained neural network; executing the pruned neural network as input to the input image set to obtain result values inferred using the pruned neural network, and obtaining a task accuracy of the pruned neural network based on predetermined result values and the result values inferred using the pruned neural network; updating the learning weights based on the first task accuracy and the task accuracy of the pruned neural network; updating the channel-wise pruning parameters based on the updated learning weights and the task accuracy of the pruned neural network; and re-pruning the pruned neural network based on channel units, based on the updated learning weights and based on the updated channel-wise pruning parameters.
[0007] In one general aspect, an apparatus for training a neural network for image recognition includes: a processor configured to: receive an input image set; execute a pre-trained neural network as input to the input image set to obtain result values inferred using the pre-trained neural network; obtain a first task accuracy of the pre-trained neural network based on predetermined result values corresponding to the input image set and the result values inferred using the pre-trained neural network; prune the neural network based on channel units by adjusting the weights between nodes of multiple channels based on preset learning weights and based on channel-wise pruning parameters corresponding to channels of each of multiple layers of the pre-trained neural network; execute the pruned neural network as input to the input image set to obtain result values inferred using the pruned neural network; obtain a task accuracy of the pruned neural network based on predetermined result values and the result values inferred using the pruned neural network; update learning weights based on the first task accuracy and the task accuracy of the pruned neural network; update channel-wise pruning parameters based on the updated learning weights and the task accuracy of the pruned neural network; and re-prune the pruned neural network based on channel units, based on the updated learning weights and based on the updated channel-wise pruning parameters.
[0008] In one general aspect, a method for training a neural network for image recognition includes: receiving an input image set; executing a pre-trained neural network on the input image set as input to obtain result values inferred using the pre-trained neural network, and obtaining a first accuracy of the pre-trained neural network based on predetermined result values corresponding to the input image set and the result values inferred using the pre-trained neural network; for each channel of the pre-trained neural network, pruning the weights of the channel based on pruning parameters and learning weights of the channel; executing a pruned neural network on the input image set as input to obtain result values inferred using the pruned neural network, and obtaining a second accuracy of the pruned neural network based on predetermined result values and the result values inferred using the pruned neural network; updating the learning weights based on a comparison between the first accuracy of the pre-trained neural network and the second accuracy of the pruned neural network; for each channel, updating the pruning parameters of the channel based on the updated learning weights and the second accuracy; and for each channel, re-pruning the weights of the channel based on the updated pruning parameters and the updated learning weights of the channel.
[0009] In one general aspect, a method for training a neural network includes: storing a pre-trained neural network in a memory; reading the pre-trained neural network from the memory by a processor; pruning the pre-trained neural network by the processor; and storing the pruned neural network in the memory, wherein the step of pruning the pre-trained neural network by the processor includes: obtaining a first task accuracy of an inference task processed by the pre-trained neural network; pruning the neural network based on channel units by adjusting the weights between nodes of multiple channels based on preset learning weights and based on channel-by-channel pruning parameters corresponding to channels of each of multiple layers of the pre-trained neural network; updating the learning weights based on the first task accuracy and the task accuracy of the pruned neural network; updating the channel-by-channel pruning parameters based on the updated learning weights and the task accuracy of the pruned neural network; and re-pruning the pruned neural network based on channel units, based on the updated learning weights, and based on the updated channel-by-channel pruning parameters.
[0010] In one general aspect, a method for training a neural network includes: storing a pre-trained neural network in a memory; reading the pre-trained neural network from the memory by a processor; pruning the pre-trained neural network by the processor; and storing the pruned neural network in the memory, wherein the step of pruning the pre-trained neural network by the processor includes: for each channel of the pre-trained neural network, pruning the weights of the channel based on pruning parameters and learning weights of the channel; updating the learning weights based on a comparison of a first accuracy of inference performed using the pre-trained neural network and a second accuracy of inference performed using the pruned neural network; for each channel, updating the pruning parameters of the channel based on the updated learning weights and the second accuracy; and for each channel, re-pruning the weights of the channel based on the updated pruning parameters and the updated learning weights of the channel.
[0011] In one general aspect, a neural network device includes: a processor configured to: acquire a first task accuracy of an inference task processed by a pre-trained neural network; prune the neural network based on channel units by adjusting weights between nodes of a plurality of channels based on preset learning weights and based on channel-wise pruning parameters corresponding to channels of each of a plurality of layers of the pre-trained neural network; update the learning weights based on the first task accuracy and the task accuracy of the pruned neural network; update the channel-wise pruning parameters based on the updated learning weights and the task accuracy of the pruned neural network; and re-prune the pruned neural network based on channel units, based on the updated learning weights and based on the updated channel-wise pruning parameters.
[0012] In one general aspect, a neural network pruning method includes: obtaining a first task accuracy of an inference task processed by a pre-trained neural network; pruning the neural network based on channel units by adjusting the weights between nodes of multiple channels based on preset learning weights and on channel-wise pruning parameters corresponding to channels of each of multiple layers of the pre-trained neural network; updating the learning weights based on the first task accuracy and the task accuracy of the pruned neural network; updating the channel-wise pruning parameters based on the updated learning weights and the task accuracy of the pruned neural network; and re-pruning the pruned neural network based on channel units, based on the updated learning weights and on the updated channel-wise pruning parameters.
[0013] Channel-by-channel trimming parameters may include a first parameter used to determine the trimming threshold.
[0014] The trimming step may include: trimming the channel elements in the plurality of channels to occupy 0 at a threshold ratio or a greater ratio.
[0015] The steps for updating the learning weights may include: in response to the task accuracy of the pruned neural network being less than the first task accuracy, updating the learning weights to increase the task accuracy of the pruned neural network.
[0016] The method may include repeatedly performing a pruning-evaluation operation to determine the task accuracy of the repruned neural network and the learning weights of the repruned neural network.
[0017] The method may include determining whether to perform an additional pruning-evaluation operation based on a preset epoch and the task accuracy of the re-pruned neural network.
[0018] The method may include: in response to repeatedly performing a pruning-evaluation operation, comparing the determined learning weights with a lower limit threshold of the learning weights; and based on the comparison result, determining whether to terminate the current pruning session and initiate a subsequent pruning session in which the learning weights are set to an initial reference value.
[0019] A non-transitory computer-readable storage medium stores instructions that, when executed by a processor, configure the processor to perform the method.
[0020] In another general aspect, a neural network pruning device includes: a processor configured to: acquire a first task accuracy of an inference task processed by a pre-trained neural network; prune the neural network based on channel units by adjusting weights between nodes of a plurality of channels based on preset learning weights and on channel-wise pruning parameters corresponding to channels of each of a plurality of layers of the pre-trained neural network; update the learning weights based on the first task accuracy and the task accuracy of the pruned neural network; update the channel-wise pruning parameters based on the updated learning weights and the task accuracy of the pruned neural network; and re-prune the pruned neural network based on channel units, based on the updated learning weights and on the updated channel-wise pruning parameters.
[0021] Channel-by-channel trimming parameters may include a first parameter used to determine the trimming threshold.
[0022] For pruning, the processor can be configured to prune, among the plurality of channels, the channel elements included in the channel occupying 0 at a threshold ratio or a greater ratio.
[0023] To update the learning weights, the processor can be configured to update the learning weights in response to the task accuracy of the pruned neural network being less than the first task accuracy, thereby increasing the task accuracy of the pruned neural network.
[0024] The processor can be configured to repeatedly perform a pruning-evaluation operation to determine the task accuracy of the repruned neural network and the learning weights of the repruned neural network.
[0025] The processor can be configured to determine whether to perform additional pruning-evaluation operations based on the task accuracy of the neural network after a preset number of rounds and the re-pruning process.
[0026] The processor can be configured to: in response to repeatedly performing the pruning-evaluation operation to compare the determined learning weights with a lower bound threshold of the learning weights, and based on the result of the comparison, determine whether to terminate the current pruning session and initiate a subsequent pruning session in which the learning weights are set to the initial reference value.
[0027] The device may include: a memory storing instructions that, when executed by a processor, configure the processor to perform steps such as obtaining the accuracy of a first task, pruning the neural network, updating the learning weights, updating the channel-by-channel pruning parameters, and re-pruning the pruned neural network.
[0028] In another general aspect, a neural network pruning method includes: for each channel of a pre-trained neural network, pruning the weights of the channel based on pruning parameters and learning weights of the channel; updating the learning weights based on a comparison of a first accuracy of inference performed using the pre-trained neural network and a second accuracy of inference performed using the pruned neural network; for each channel, updating the pruning parameters of the channel based on the updated learning weights and the second accuracy; and for each channel, re-pruning the weights of the channel based on the updated pruning parameters and the updated learning weights of the channel.
[0029] The steps for updating learning weights may include updating the learning weights based on whether the second accuracy is greater than the first accuracy.
[0030] The steps for updating the learning weights may include: decreasing the learning weights in response to a second accuracy being greater than a first accuracy; and increasing the learning weights in response to a second accuracy being less than or equal to a first accuracy.
[0031] The steps of pruning weights may include: determining the value of the weight transformation function based on the pruning parameters; and pruning the weights based on the value of the weight transformation function.
[0032] The step of pruning weights may include pruning a larger number of weights in response to the first value, compared to pruning a second value in response to the first value being less than the first value.
[0033] Other features and aspects will become clear from the following detailed description, the accompanying drawings, and the claims. Attached Figure Description
[0034] Figure 1 This shows an example of an operation performed in an artificial neural network.
[0035] Figure 2 An example of pruning is shown.
[0036] Figure 3 An example of a device used for pruning artificial neural networks is shown.
[0037] Figure 4 An example of an electronic system is shown.
[0038] Figure 5 This is a flowchart illustrating an example of a pruning algorithm performed by a device used to prune an artificial neural network.
[0039] Figure 6 This is a flowchart illustrating an example of an operation performed by a device for pruning artificial neural networks to update pruning-evaluation operation variables.
[0040] Figure 7 An example of a system including an artificial neural network and a device for pruning the artificial neural network is shown.
[0041] Throughout the accompanying drawings and detailed embodiments, unless otherwise described or provided, the same reference numerals will be understood to denote the same elements, features, and structures. The drawings may not be to scale, and for clarity, illustration, and convenience, the relative dimensions, scale, and depiction of elements in the drawings may be exaggerated. Detailed Implementation
[0042] The following detailed embodiments are provided to aid the reader in gaining a comprehensive understanding of the methods, apparatus, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but may be changed as will become clear upon understanding this disclosure, except for operations that must occur in a specific order. Furthermore, for clarity and brevity, descriptions of features known in the art upon understanding this disclosure may be omitted.
[0043] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein have been provided only to illustrate some of the many feasible ways of implementing the methods, apparatus, and / or systems described herein that will become clear upon understanding the disclosure of this application.
[0044] The structural or functional descriptions of the examples disclosed in this disclosure are intended for illustrative purposes only, and the examples may be implemented in various forms. The examples are not intended to be limiting, but rather to cover various modifications, equivalents, and substitutions within the scope of the claims.
[0045] Although the terms “first” or “second” are used to describe various components, assemblies, regions, layers, or parts, the components, assemblies, regions, layers, or parts are not limited by the terms. These terms should only be used to distinguish one component, assembly, region, layer, or part from another component, assembly, region, layer, or part. For example, within the scope of the claims according to the concept of this disclosure, the “first” component, “first” assembly, “first” region, “first” layer, or “first” part referred to in the examples described herein may be referred to as a “second” component, “second” assembly, “second” region, “second” layer, or “second” part, or similarly, a “second” component, “second” assembly, “second” region, “second” layer, or “second” part may be referred to as a “first” component, “first” assembly, “first” region, “first” layer, or “first” part.
[0046] Throughout this specification, it will be understood that when a component or element is referred to as being "on," "connected to," or "joined to" another component or element, the component or element may be directly on, directly connected to, or joined to the other component or element, or there may be one or more intermediate components or elements in between. Conversely, when a component or element is referred to as being "directly on," "directly connected to," or "directly joined to" another component or element, there are no intermediate components or elements. Similarly, expressions such as "between" and "immediately between," and "adjacent to" and "closely adjacent to" can also be interpreted as described above.
[0047] The terminology used herein is for the purpose of describing particular examples only and is not intended to limit this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. As used herein, the term “and / or” includes any one and any combination of any two or more of the associated listed items. As used herein, the terms “comprising” and / or “including” as used in this specification indicate the presence of the described features, integrals, steps, operations, elements, components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. The use of the term “may” herein with respect to examples or embodiments (e.g., regarding what an example or embodiment may include or implement) indicates the presence of at least one example or embodiment that includes or implements such a feature, while all examples are not limited thereto.
[0048] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the examples pertain. It will also be understood that, unless expressly defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and in this disclosure, and shall not be interpreted in an idealized or overly formalized sense.
[0049] In the following description, examples will be described in detail with reference to the accompanying drawings, and the same reference numerals in the drawings always denote the same elements.
[0050] Examples can be implemented in various types of products (such as data centers, servers, personal computers (PCs), laptops, tablets, smartphones, televisions, smart electronic devices, smart vehicles, self-service kiosks, and wearable devices). The examples are described below with reference to the accompanying drawings. Although the same reference numerals are shown in different drawings, the same reference numerals always denote the same elements.
[0051] One or more embodiments of the method and apparatus can perform compression while maintaining the performance of the neural network and reducing the system cost of implementing the neural network.
[0052] Figure 1 This shows an example of an operation performed in an artificial neural network.
[0053] Artificial neural networks can be computational systems that implement information processing methods. While neural networks may be called "artificial" neural networks, this term is not intended to assign any relation to how the neural network architecture computationally maps or thereby intuitively identifies information and how human nodes operate. In other words, the term "artificial neural network" is simply a specialized term referring to the hardware implementation of the neural network.
[0054] As one method for implementing artificial neural networks, deep neural networks (DNNs) can include multiple layers. For example, a DNN can include an input layer, an output layer, and multiple hidden layers. The input layer is configured to receive input data, the output layer is configured to output the predicted value based on the input data, and the multiple hidden layers are positioned between the input layer and the output layer.
[0055] Based on the algorithms used to process information, DNNs can be classified into Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), etc.
[0056] The method of training artificial neural networks is called deep learning. Various algorithms (e.g., CNN schemes and RNN schemes) can be used for deep learning.
[0057] Here, training an artificial neural network can represent determining and updating the weights and biases between layers or between multiple nodes in different layers that belong to adjacent layers.
[0058] For example, the multiple hierarchical structures and the weights and biases between multiple layers or nodes can be collectively referred to as the connectivity of an artificial neural network. Therefore, training an artificial neural network can be represented as constructing and learning connectivity.
[0059] Reference Figure 1 The artificial neural network 100 may include an input layer, a hidden layer and an output layer. The artificial neural network 100 may perform operations based on received input data (e.g., I1 and I2) and generate output data (e.g., O1 and O2) based on the results of the operations.
[0060] The artificial neural network 100 can be a DNN or an n-layer artificial neural network including one or more hidden layers. For example, refer to... Figure 1 The artificial neural network 100 can be a DNN comprising an input layer (layer 1), one or more hidden layers (layers 2 and 3), and an output layer (layer 4). The DNN can include CNNs, RNNs, deep belief networks (DBNs), restricted Boltzmann machines (RBMs), etc. However, this is provided only as an example.
[0061] When the artificial neural network 100 is implemented as a DNN architecture, it can include more layers capable of processing effective information. Therefore, compared to an artificial neural network comprising a single layer, the artificial neural network 100 can handle complex datasets. Although Figure 1 The artificial neural network 100 shown includes four layers, but this is provided only as an example. The artificial neural network 100 may include more or fewer layers, or more or fewer channels. The artificial neural network 100 may include... Figure 1 Layers with different structures.
[0062] Each layer included in the artificial neural network 100 may include multiple channels. Multiple channels may correspond to multiple nodes, processing elements (PEs), or similar terms. For example, refer to... Figure 1 Layer 1 may include two channels (nodes), and each of layers 2 and 3 may include three channels (e.g., CH1 to CH3). However, this is provided only as an example. Each layer included in the artificial neural network 100 may include a variety of numbers of channels (nodes).
[0063] Channels included in each layer of the artificial neural network 100 can be interconnected and process data. For example, a single channel can receive data from other channels and perform operations, and can output the results of the operations to other channels.
[0064] The input and output of each channel can be referred to as input activation and output activation, respectively. That is, activation can be a parameter corresponding to the output of a single channel and simultaneously to the input of channels included in subsequent layers. The activation of each channel can be determined based on the activations, weights, and biases received from channels included in the previous layer. Weights can be parameters used to calculate the output activation of each channel, and can also be values assigned to the connectivity relationships between channels.
[0065] Each channel can be processed by a computing unit or processing element, which is configured to receive input and output activation. The input and output of each channel can be mapped. For example, when σ represents the activation function, This represents the weights from the k-th node included in layer (i-1) to the j-th node included in layer i. This represents the bias value of the j-th node included in the i-th layer, and When the j-th node in the i-th layer is activated, the activation... It can be represented as Equation 1 below.
[0066] Equation 1:
[0067]
[0068] Reference Figure 1 The activation of the first channel (CH1) in the second layer (layer 2) can be represented as also, It can have according to Equation 1 The value of . Equation 1 above is provided only as an example to describe the activations, weights, and biases used by the artificial neural network 100 to process data. Activations can be values obtained by passing a weighted average of activations received from the previous layer to an activation function (e.g., a sigmoid function and a rectified linear unit (ReLU) function).
[0069] Figure 2 An example of pruning is shown.
[0070] exist Figure 2 In this context, artificial neural network 210 may correspond to a portion of a layer of an artificial neural network that has been pre-trained (or previously trained) before pruning, and artificial neural network 220 may correspond to a portion of a corresponding layer of a pruned artificial neural network.
[0071] Reference Figure 2In the artificial neural network 210, connections can be formed between all node combinations including two node combinations in three channels contained in each of two adjacent different layers. For example, each node in a channel layer can be connected to each node in a subsequent layer of the channel.
[0072] exist Figure 2 In this context, the pre-trained artificial neural network 210 can be fully connected, and the weights representing the connection strength between two nodes belonging to different layers included in the artificial neural network 210 can be values greater than 0. When there are connections between nodes in all adjacent layers, the overall complexity of the artificial neural network 210 increases. Furthermore, due to the overfitting problem, the prediction results of the artificial neural network 210 may have reduced accuracy and confidence.
[0073] To reduce complexity and / or overfitting, pruning of the artificial neural network 210 can be performed. For example, if there are weights with a preset threshold or smaller among multiple weights in the artificial neural network 210 before pruning, weakening or removal of the corresponding weights can be performed.
[0074] To determine which parts of the artificial neural network 210 can be pruned, the artificial neural network 210 can be explored. Here, the artificial neural network 210 can be pruned by removing or reducing portions of the parameters or channels of the layers of the artificial neural network 210 that do not substantially damage or reduce the accuracy of the artificial neural network 210.
[0075] Pruning can be performed on each channel of a layer of the artificial neural network 210 that does not substantially affect the output of the artificial neural network 210. For example, pruning can be performed on one or more input feature maps of each channel of a layer that does not substantially affect the output feature map generated by each channel of the layer.
[0076] It can identify and retrieve connections between nodes whose weights are less than a specified threshold. Connections corresponding to all weights identified or determined to have values less than the threshold can be removed, zeroed out, or ignored.
[0077] When the weight is small (e.g., when the weight is less than a specified lower threshold), the corresponding channel of the artificial neural network 210 can be detected. In this case, the detected channel can be selected as a candidate to be removed from the artificial neural network 210.
[0078] Pruning may include reducing the precision of the numerical form of at least one value of the artificial neural network 210. For example, it can be determined whether the precision of the numerical form used for the weights can be reduced by analyzing at least one weight in the channels of each of the multiple layers of the artificial neural network 210. Due to the reduced precision of the numerical form used, lower precision arithmetic hardware can be used in turn. Low-precision arithmetic hardware can be more power-efficient and densely embedded than high-precision arithmetic hardware. Compared to a typical artificial neural network that uses more bits than required, one or more embodiments of the artificial neural network can achieve relatively high performance (e.g., fast runtime and / or low power consumption) by using the minimum number of bits required to indicate the precision and range of the parameters.
[0079] Furthermore, the pruning methods of one or more embodiments can prune specific channels among multiple channels with a slight loss of accuracy. In one example, channels among multiple channels whose channel elements occupy 0 at a threshold ratio or greater (e.g., channels whose number of channel elements being 0 is greater than or equal to the total number of channel elements) can be pruned. For example, when most values of channel 5 in layer 2 converge to 0 and the amount of data is small, the corresponding channel can be pruned.
[0080] For example, the weights of the artificial neural network 210 can be transformed by applying a weight transformation function to the artificial neural network 210 before pruning, as shown in Equation 2 below.
[0081] Equation 2:
[0082] w′ n,c,m =g(w n,c,m α n,c ,β n,c )·W n,c,m
[0083] In equation 2, w n,c,m w' represents the weights of the artificial neural network (e.g., artificial neural network 210) before pruning. n,c,m This represents the weights of the pruned artificial neural network (e.g., artificial neural network 220), where n represents the corresponding layer index, c represents the corresponding channel index, m represents the corresponding weight index, and g represents the weight transformation function.
[0084] For example, the weighting transformation function g can be expressed as Equation 3 below.
[0085] Equation 3:
[0086]
[0087] In equation 3, β n,c α represents the threshold-determining variable. n,cThis represents the variable used to determine the gradient of the weight transformation function around a threshold. In one example, β n,c This can be a parameter used to determine the pruning threshold. Here, α n,c and β n,c These can be referred to as trimming parameters, which can be determined for each channel of the corresponding layer. For example, α n,c and β n,c It can be the trimming parameter corresponding to the c-th channel of the n-th layer.
[0088] As α increases, the value of the weight transformation function can gradually change based on the threshold. As β increases, the region of weights that are transformed into values close to 0 can decrease. When the value of the weight transformation function is less than or equal to the expected value, the weight value can be transformed into 0.
[0089] For example, in In the weight region where the value is less than 0.5, the weight transformation function g(w) n,c,m α n,c ,β n,c ) = 0.
[0090] The weight transformation function g can take various forms, and it is understood that, upon understanding this disclosure, other applicable functions are also within the scope of this disclosure.
[0091] Channel-by-channel pruning parameters can be determined through learning. The artificial neural network can be trained to determine the model parameters that minimize the loss function. The loss function can be determined as an index to determine the optimal model parameters during the learning process of the artificial neural network. In one example, the artificial neural network can be trained based on the loss function shown in Equation 4 below.
[0092] Equation 4:
[0093]
[0094] Where 0 < λ < 1, 0 < η < ∞.
[0095] In Equation 4, L' represents the loss function, L represents the task accuracy, λ represents the learning weight, N represents the number of layers, and C represents the number of channels.
[0096] Referring to Equation 4, the ratio between task accuracy and pruning amount can be determined based on the learning weight λ. For example, as the learning weight λ gets closer to 1, the pruning amount can increase and the task speed can improve, which may lead to a decrease in task accuracy. Conversely, as the learning weight λ gets closer to 0, task accuracy can be improved and the pruning amount can be reduced, which may lead to a decrease in task speed. In the following, as a non-limiting example, refer to... Figures 3 to 6This paper describes a method for determining learning weights and pruning parameters, and pruning artificial neural networks based on the determined learning weights and pruning parameters.
[0097] Figure 3 An example of a device (or neural network device) for pruning an artificial neural network (e.g., device 300) is shown.
[0098] Reference Figure 3 The device 300 may include a memory 310 (e.g., one or more memories) and a controller 320 (e.g., one or more processors).
[0099] Device 300 may include memory 310 and controller 320, the controller 320 being connected to memory 310 via a system bus or other suitable circuitry.
[0100] Device 300 can store instructions in memory 310. Controller 320 can process the operation of pruning the artificial neural network by executing instructions called from memory 310 via the system bus.
[0101] Memory 310 may include local memory or at least one physical memory device (such as at least one mass storage device). Here, local memory may include random access memory (RAM) or other volatile memory devices typically used during the actual execution of instructions. Mass storage devices may be implemented as hard disk drives (HDDs), solid-state drives (SSDs), or other non-volatile memory devices. Furthermore, device 300 may include at least one cache that provides temporary storage for at least a portion of the instructions to reduce the number of times the mass storage device searches for instructions during pruning operations. In one example, the memory may store a pre-trained artificial neural network and a pruned artificial neural network.
[0102] In response to the execution of executable instructions stored in memory 310 by device 300, controller 320 may perform various operations disclosed herein. For example, memory 310 may store instructions that enable controller 320 to execute. Figure 1 , Figure 2 and Figures 4 to 7 At least one operation described in the document.
[0103] Depending on the specific type of device to be implemented, device 300 may include more than Figure 3 The number of components shown may be a small number of components, or may include a small number of components. Figure 3 Additional components not shown. Furthermore, at least one component may be included in or constitute part of another component.
[0104] The controller 320 can acquire an initial value for the task accuracy of the inference task processed by the pre-trained artificial neural network. Hereinafter, the initial value for the task accuracy of the inference task processed by the pre-trained artificial neural network may be referred to as the first task accuracy.
[0105] Artificial neural networks can learn target tasks and build inference models. Furthermore, based on the constructed inference models, artificial neural networks can output inferences about external input values.
[0106] Related to tasks performed using artificial neural networks, artificial neural networks can be applied to facial recognition modules or software for smartphones, recognition / classification operations (such as object recognition, speech recognition, and image classification), medical and diagnostic devices, and unmanned systems. Furthermore, artificial neural networks can be implemented as dedicated processing devices configured to extract meaningful information by processing video data.
[0107] The controller 320 can obtain an initial value for task accuracy by evaluating the pre-trained artificial neural network. Task accuracy may include the mean squared error (MSE), which represents the error between the expected result value and the inferred result value by the artificial neural network. Here, the smaller the MSE, which corresponds to the task accuracy, the better the performance of the artificial neural network.
[0108] For example, to measure task accuracy, controller 320 may input a dataset for measuring the performance of the artificial neural network. Controller 320 may calculate the Mean Squared Error (MSE) between the expected result value (e.g., a predetermined result value corresponding to the input dataset) and the result value inferred using the artificial neural network based on the input dataset, and may determine task accuracy as the MSE, determine task accuracy including the MSE, or determine task accuracy based on the MSE. Here, the dataset may be determined differently based on the domain of the task the artificial neural network expects. For example, datasets such as CIFAR-10 and CIFAR-100 may be used in the domain of image classification. In one example, controller 320 may execute the artificial neural network with an input dataset (e.g., an input image set) as input to obtain the result value inferred using the artificial neural network, and obtain the task accuracy of the artificial neural network based on the predetermined result value corresponding to the input dataset and the result value inferred using the artificial neural network. In one example, device 300 may be a device for training a neural network for image recognition; however, the example is not limited to this.
[0109] The controller 320 can obtain the task accuracy of the pre-trained artificial neural network as an initial value for task accuracy before performing pruning. For example, task accuracy can be calculated by receiving a training dataset and performing an evaluation based on the results predicted by the artificial neural network using the training dataset. Task accuracy can be the prediction loss; the smaller the prediction loss, the more accurate the inference result of the artificial neural network. In addition, various methods can be used to measure task accuracy. The controller 320 can evaluate the artificial neural network multiple times and obtain the task accuracy based on the average of multiple evaluation results.
[0110] The controller 320 can prune the artificial neural network based on channel units (e.g., on a channel-by-channel basis) by adjusting the weights between nodes belonging to a channel according to preset learning weights based on channel-by-channel pruning parameters corresponding to channels in a plurality of layers constituting the pre-trained artificial neural network. For example, the controller 320 can prune the artificial neural network by adjusting at least a portion of the connections between nodes belonging to a channel used for sending and receiving channel information of each of the plurality of layers in the pre-trained artificial neural network.
[0111] Here, the controller 320 can acquire or determine information about the connections between channels included in the multiple layers constituting the pre-trained artificial neural network. That is, the controller 320 can acquire or determine information about the channels included in each of the multiple layers of the multi-layer artificial neural network and information about the connections between the channels included in the multiple layers. Furthermore, the information about the connections may include information about the weights of the connections between adjacent layers among the multiple layers.
[0112] The controller 320 can prune the artificial neural network by adjusting at least a portion of the connections between channels included in the layers of the artificial neural network. For example, the controller 320 can compress the artificial neural network by adjusting the pruning operation of the weights of the connections. (See above for reference.) Figure 2 The descriptions related to pruning can be applied here.
[0113] The controller 320 can determine (or update) the learning weights of the artificial neural network based on an initial value of task accuracy and the task accuracy of the pruned artificial neural network. Here, when the task accuracy of the pruned artificial neural network is less than the initial value, the controller 320 can determine learning weights to increase the task accuracy. That is, when the inference task accuracy determined based on the task accuracy of the pruned artificial neural network is less than the inference task accuracy of the artificial neural network before pruning, determined based on the initial value of task accuracy, the learning weights can be reduced to prevent a decrease in inference task performance. For example, when the value of task accuracy is proportional to the performance of the inference task, the controller 320 can determine that the value of the task accuracy of the pruned artificial neural network is less than the value of the task accuracy before pruning, and therefore can reduce the learning weights. Conversely, when the value of task accuracy is inversely proportional to the performance of the inference task, the controller 320 can determine that the value of the task accuracy of the pruned artificial neural network is less than the value of the task accuracy before pruning, and therefore can increase the learning weights.
[0114] The learning weights can represent the degree of pruning performed compared to the time required to perform a single pruning operation. The amount of time used to perform a single pruning operation can be the same for each device performing the pruning. Therefore, as the learning weights increase, the amount of information about the pruning process in each pruning stage can increase.
[0115] The controller 320 can update the channel-wise pruning parameters based on the determined learned weights and the task accuracy of the pruned artificial neural network. For example, the controller 320 can update the parameters based on a variable (e.g., the threshold determination variable β in Equation 3) that is used to determine the weights as criteria for performing pruning. n,c The loss function is determined by the established learning weights and the task accuracy of the pruned artificial neural network, and the threshold determination variable can be updated to reduce the loss function.
[0116] The controller 320 can re-prune the pruned artificial neural network based on the updated learning weights according to the updated channel-by-channel pruning parameters. The controller 320 can repeatedly perform the pruning-evaluation operation to re-prune the artificial neural network and determine the task accuracy and learning weights of the re-pruned artificial neural network. Here, the controller 320 can determine whether to perform an additional pruning-evaluation operation based on a preset number of epochs and the task accuracy of the re-pruned artificial neural network. The pruning-evaluation operation can be a unit that measures a single pruning and the task accuracy.
[0117] Controller 320 can compare determined learning weights with a lower bound threshold of the learning weights in response to repeatedly performing pruning-evaluation operations. Based on the comparison result, controller 320 can determine whether to terminate the current pruning session and initiate a subsequent pruning session where the learning weights are set to initial reference values. For example, when it is determined that the determined learning weights are less than or equal to the lower bound threshold of the learning weights, controller 320 can terminate the current pruning session, set the learning weights to the initial reference values, and initiate a subsequent pruning session. Here, by terminating the current pruning session, controller 320 can store information about the artificial neural network pruned during the current pruning session. By terminating each of multiple pruning sessions, controller 320 can determine the optimal pruning method by comparing the stored performance of the artificial neural network. A session can be a unit that continues until the learning weights are updated to initial values according to multiple pruning-evaluation operations.
[0118] The controller 320 can perform pruning-evaluation operations, thus acquiring the task accuracy of the pruned or re-pruned artificial neural network. For example, whenever the controller 320 performs each of multiple pruning-evaluation operations during a single session, the controller 320 can acquire and / or store the task accuracy and calculate the average task accuracy for each single session. The average task accuracy calculated in a session can be used as a benchmark for determining the learning weights in subsequent sessions.
[0119] Figure 4 An example of an electronic system is shown.
[0120] Reference Figure 4 The electronic system 400 can analyze input data in real time based on an artificial neural network, extract effective information, and determine the situation and / or control components of the electronic device to which the electronic system 400 is installed based on the extracted information. For example, the electronic system 400 can be a drone, robotic equipment (such as an advanced driver assistance system (ADAS)), a smart TV, a smartphone, a medical device, a mobile device, a video display device, a measuring device, and an Internet of Things (IoT) device, or the electronic system 400 can be included in a drone, robotic equipment (such as an advanced driver assistance system (ADAS)), a smart TV, a smartphone, a medical device, a mobile device, a video display device, a measuring device, and an Internet of Things (IoT) device. In addition, the electronic system 400 can be installed in at least one of various types of electronic devices.
[0121] Reference Figure 4The electronic system 400 may include a controller 410 (e.g., one or more processors), RAM 420, neuromorphic device 430, memory 440 (e.g., one or more memories), and communication module 450. A portion of the hardware components of the electronic system 400 may be mounted on at least one semiconductor chip.
[0122] The controller 410 controls the overall operation of the electronic system 400. The controller 410 may include a single core or multiple cores. The controller 410 can process or execute programs and / or data stored in the memory 440. The controller 410 can control the functions of the neuromorphic device 430 by executing programs stored in the memory 440. Furthermore, the controller 410 can perform pruning to reduce the amount of weight information used by the neuromorphic device 430. The controller 410 may be implemented as a central processing unit (CPU), graphics processing unit (GPU), access point (AP), etc. Figure 4 The controller 410 and memory 440 can respectively correspond to Figure 3 The controller 320 and the memory 310.
[0123] RAM 420 may temporarily store programs, data, or instructions. For example, programs and / or data stored in memory 440 may be temporarily stored in RAM 420 according to control or start codes of controller 410. RAM 420 may be implemented as a memory (such as dynamic RAM (DRAM) or static RAM (SRAM)).
[0124] The neuromorphic device 430 can perform operations based on received input data and generate information signals based on the results of the operations. The neuromorphic device 430 may correspond to a hardware accelerator dedicated to artificial neural networks or a device including a hardware accelerator.
[0125] The information signal may include one of various types of recognition signals (e.g., any one of voice recognition signals, object recognition signals, video recognition signals, and biometric recognition signals). For example, the neuromorphic device 430 may receive frame data included in a video stream as input data and may generate a recognition signal about an object included in an image represented by the frame data from the frame data. However, this is provided only as an example. The neuromorphic device 430 may receive various types of input data and may generate recognition signals based on the type or function of the electronic device to which the electronic system 400 is installed.
[0126] Memory 440 may be a storage device configured to store data and to store the OS, various types of programs, and various types of data. According to an example, memory 440 may store intermediate results generated during the operation execution processing of the neuromorphic device 430 or weights used during the operation execution processing.
[0127] Memory 440 may be DRAM, but this is provided by way of example only. Memory 440 may include any one or any combination of volatile memory and non-volatile memory. Examples of non-volatile memory include ROM, random access programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, PRAM, MRAM, RRAM, FRAM, etc. Examples of non-volatile memory include DRAM, SRAM, SDRAM, PRAM, MRAM, RRAM, FeRAM, etc. According to the example, memory 440 may include any one or any combination of HDD, SSD, CF, SD, micro-SD, mini-SD, xD, and Memory Stick.
[0128] The communication module 450 may include various types of wired or wireless interfaces capable of communicating with external devices. For example, the communication module 450 may include wired local area network (LAN), wireless local area network (WLAN) (such as Wi-Fi), wireless personal area network (WPAN) (such as Bluetooth), wireless universal serial bus (USB), ZigBee, near field communication (NFC), radio frequency identification (RFID), power line communication (PLC), and mobile cellular network accessible communication interfaces (such as 3G, 4G, and LTE).
[0129] Figure 5 This is a flowchart illustrating an example of a pruning algorithm (or, a method for training a neural network) performed by a device used to prune an artificial neural network.
[0130] Figure 5 The processing can be done by Figure 3 Device 300 for pruning artificial neural networks and Figure 4 The electronic system 400 is used to perform the operation. In one example, the device 300 and the electronic system 400 are used to perform the operation. Figure 5 The processing can be a method for training neural networks for image recognition; however, examples are not limited to this.
[0131] In the following text, Figure 5 and Figure 6 The numerical values, ranges of variables, and equations used can be modified to the extent readily available after understanding this disclosure, and the modified numerical values are within the scope of the disclosure.
[0132] Reference Figure 5In operation 501, the device can acquire information about a pre-trained artificial neural network. As described above, the information about the artificial neural network may be, or may include, the weights between multiple nodes belonging to various channels in adjacent layers included in multiple layers.
[0133] In operation 502, the device can initiate an initial session. Here, the index n used to identify the session can be set to 1.
[0134] In operation 503, the device can set initial values for the trim-evaluation operation variables. For example, the device can set s to 0, where s is an index representing the number of trim-evaluation operations repeated in each session.
[0135] In addition, the device can represent the λ of the learning weights. s λ is set as the initial reference weight value. init .
[0136] For example, λ init λ can represent the maximum value within the available range of learning weights. As another example, λ init It can be a random value within the available range of learning weights. For example, the learning weight λ s It can have values from 0 to 1. Here, λ init It can be 1. However, this is only provided as an example.
[0137] In addition, the device can set the average task accuracy T for each session. n-1 Here, the subscript is not n but n-1, indicating that the average task accuracy of the previous session is used as the comparison standard with the task accuracy of the current session. However, this is provided only as an example. Various methods can be used to determine the task accuracy used as the comparison standard.
[0138] For example, the device can repeatedly perform evaluations based on a pre-determined training dataset for a pre-trained artificial neural network before pruning, and then the average of the obtained evaluation results can be determined as T. init Task accuracy can be determined by the device. As another example, task accuracy can be determined by an external device that implements the artificial neural network or another external device, from which the device can then obtain the learning accuracy.
[0139] In addition, the device can obtain the lower bound threshold λ of the learning weights as the termination criterion for each session. min For example, the lower threshold λ min This can be a value close to the lower bound determined based on the available range and lower bound of the learning weights. For example, the learning weight λ can have values from 0 to 1. Here, λ min It can be 10 -20 However, this is provided only as an example.
[0140] Furthermore, as an example of pruning, the device can determine the threshold variable β. s Set to the initial value β init To determine the thresholds of weights used as the criteria for performing pruning. Here, β s The magnitude of the weight threshold can be determined as a variable in the weight transformation function that transforms existing weights. For example, with β... s As the threshold increases, there is a probability that more weights will be pruned, converted to 0, or converted to smaller weight values.
[0141] In operation 504, the device can perform a pruning operation on the artificial neural network based on the pruning-evaluation operation variables.
[0142] For example, the device can acquire information about the values of all or at least a portion of the weights in each adjacent layer of an artificial neural network, and can do so based on predetermined learning weights λ. s and pruning parameters (e.g., pruning parameters include threshold determination variable β) s Use this to perform pruning.
[0143] In operation 504, when the weight value of the pre-trained artificial neural network or the pruned artificial neural network in the previous pruning-evaluation operation is less than the pruning reference threshold, the device can either convert the corresponding weight value to 0 or reduce the weight value to a smaller value. Conversely, in operation 504, when the weight value of the pre-trained artificial neural network or the pruned artificial neural network in the previous pruning-evaluation operation is greater than the pruning reference threshold, the device can either convert the corresponding weight value to 1 or increase the corresponding weight value to a larger value.
[0144] In operation 505, the device can update the pruning-evaluation operation variables (e.g., the learned weights λ). s As a non-limiting example, see [reference]. Figure 6 The method for updating pruning-evaluation operational variables is further described.
[0145] In operation 506, the device can determine the updated learning weights λ. s Is it less than the lower bound threshold λ of the learning weight? min .
[0146] When the updated learning weight λ is determined in operation 506 s The lower bound threshold λ of the learning weights is greater than or equal to the threshold value. min If this happens, the device can return to operation 504 and execute subsequent procedures.
[0147] When the updated learning weight λ is determined in operation 506 s Less than the lower bound threshold λ of the learning weight minAt this point, the device can execute a session termination procedure. The pruning process may slow down due to the decreased task accuracy of the pruned artificial neural network. Therefore, the device can terminate the current session n and increase the learning weights λ. s Then proceed with the next n+1 sessions.
[0148] In operation 507, the device may update the average inference task accuracy of the current session, which is a benchmark used to evaluate the task accuracy of the pruned artificial neural network in the subsequent session n+1. For example, the device may use the average task accuracy obtained when performing the pruning-evaluation operation a preset number of times in the current session n to update the inference task accuracy in session n. For example, the device may update the inference task accuracy according to Equation 5 below.
[0149] Equation 5:
[0150] T n =AVG(T s )
[0151] In equation 5, T n This represents the inference task accuracy in session n, where AVG represents the operator used to calculate the average, and T... s This represents the sum of the task accuracy values of the pruned artificial neural network obtained in session n. Upon understanding this disclosure, it will be understood that any method for calculating the inference task accuracy (the benchmark for comparing task accuracy in subsequent sessions n=1) in various ways, other than simple averaging, is within the scope of this disclosure.
[0152] In operation 508, the device can determine whether to perform an additional pruning-evaluation operation based on the task accuracy of the re-pruned artificial neural network and a preset number of rounds.
[0153] For example, in operation 508, the device may terminate the pruning operation when at least a preset percentage (e.g., 70%) of the entire artificial neural network has been pruned, or when the artificial neural network has been trained through a preset number of rounds. Here, the degree of pruning can be determined based on the task accuracy of the re-pruned artificial neural network. For example, the higher the task accuracy of the re-pruned artificial neural network, the lower the degree of pruning. Conversely, the lower the task accuracy of the re-pruned artificial neural network, the higher the degree of pruning.
[0154] In operation 509, the device can store information about the pruned artificial neural network in session n before updating the session. Furthermore, in operation 509, the device can increase the decreased learning weights λ. s For example, the device can transmit λ s λ is set as the initial reference weight value. init .
[0155] Figure 6 This is a flowchart illustrating an example of an operation performed by a device for pruning artificial neural networks to update pruning-evaluation operation variables.
[0156] Figure 6 The processing can be done by Figure 3 Device 300 for pruning artificial neural networks and Figure 4 The electronic system 400 is used to perform this action. In a non-limiting example, Figure 6 The operation can correspond to Figure 5 Operation 505.
[0157] Reference Figure 6 In operation 601, the device can initiate an operation to update the trim-evaluation operation variable.
[0158] In operation 602, the device can update s with s+1, representing the number of times the pruning-evaluation operation is repeated in session n. In the following text, s represents the updated value.
[0159] In operation 603, the device can determine the mission accuracy L. s-1 For example, L can be determined based on an evaluation performed on a pruned artificial neural network. s-1 As mentioned above, as the inference becomes more accurate, the task accuracy L increases. s-1 The value can be reduced.
[0160] Task accuracy L s-1 This could be, for example, predicting loss. However, this is only provided as an example.
[0161] In operation 604, the device can add the task accuracy history from the current session n. For example, the device can store the accuracy data used to determine... Figure 5 The average inference task accuracy T for each session described in the text n L s-1 .
[0162] In operation 605, the device can determine the task accuracy L. s-1 Is the value greater than the average inference task accuracy T in the previous session n-1? n-1 The value of .
[0163] When the task accuracy of the pruned artificial neural network (e.g., task accuracy L) s-1 The inference task accuracy is less than the initial value or the accuracy of the previous session (e.g., the average inference task accuracy T). n-1 When this happens, the device can determine learning weights to increase task accuracy.
[0164] When determining the task accuracy L in operation 605 s-1The value is greater than the average inference task accuracy T in the previous session n-1. n-1 When the value is given, for example in operation 606, the device can update the learning weight λ according to the following equation 6. s .
[0165] Equation 6:
[0166]
[0167] When determining the task accuracy L in operation 605 s-1 The value is less than or equal to the average inference task accuracy T in the previous session n-1. n-1 When the value is given, for example in operation 607, the device can update the learning weight λ according to the following equation 7. s .
[0168] Equation 7:
[0169]
[0170] Here, it can be based on the learned weight λ s And task accuracy L s-1 To determine the threshold and the variable β n,c and the variable α used to determine the gradient of the weight transformation function around a threshold (e.g., a pruning parameter). n,c .
[0171] In operation 608, the device can terminate the update of the trim-evaluation operation variables.
[0172] Generally, as the pruning rate of an artificial neural network increases, its accuracy may decrease. Figures 4 to 6 For example, the variables required to perform pruning can be determined based on data and equations acquired by the device. That is, the optimal pruning is determined by an algorithm that minimizes pruning costs, without the need for fine-tuning or manually setting key parameters with relatively high sensitivity. Therefore, the device of one or more embodiments can reduce the amount of time and cost spent on pruning, which can lead to improved pruning efficiency. Furthermore, by pruning the artificial neural network, the complexity of the artificial neural network can be reduced, and the amount of memory allocated accordingly can be reduced.
[0173] Figure 7 An example of a system is shown that includes an artificial neural network (e.g., artificial neural network 750) and a device for pruning the artificial neural network (e.g., device 700).
[0174] For example, artificial neural network 750 may represent a device included in an external server and / or database of device 700.
[0175] Reference Figure 7The device 700 may include a memory 710 (e.g., one or more memories), a controller 720 (e.g., one or more processors), and a communicator 730.
[0176] Without leaving Figure 7 In the case of the range of examples, with Figure 3 The description relating to memory 310 and controller 320 can be applied to Figure 7 The memory 710 and the controller 720.
[0177] The device 700 can form a communication network with the artificial neural network 750 through the communicator 730.
[0178] Device 700 can obtain information about artificial neural network 750 from artificial neural network 750 via communicator 730. During pruning operations, device 700 can access information about artificial neural network 750 via communicator 730. Therefore, it is not necessary to store information about artificial neural network 750 in memory 710.
[0179] Furthermore, device 700 can be executed in various ways. For example, device 700 can be executed in a user terminal and the pruned artificial neural network can be obtained by accessing an external artificial neural network. As another example, artificial neural network 750 and device 700 can be implemented integratedly in the user terminal. As yet another example, device 700 and artificial neural network 750 can be implemented separately from the user terminal, and the user terminal can obtain only the pruned artificial neural network 750 through device 700.
[0180] Regarding Figures 1 to 7The described devices, memories, controllers, electronic systems, RAM, neuromorphic devices, communication modules, communicators, device 300, memory 310, controller 320, electronic system 400, controller 410, RAM 420, neuromorphic device 430, memory 440, communication module 450, device 700, memory 710, controller 720, communicator 730, and other devices, units, modules, apparatuses, and components are implemented or represent hardware components by hardware components. Examples of hardware components that can be used to perform the operations described in this application suitably include: controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components performing the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer may be implemented using one or more processing elements, such as logic gate arrays, controllers and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field-programmable gate arrays, programmable logic arrays, microprocessors, or any other means or combination of means configured to respond to and execute instructions in a defined manner to achieve a desired result. In one example, the processor or computer includes or is connected to one or more memories storing instructions or software executed by the processor or computer. Hardware components implemented by the processor or computer may execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) for performing the operations described herein. The hardware components may also access, manipulate, process, create, and store data in response to the execution of instructions or software. For simplicity, the singular terms “processor” or “computer” may be used in the description of the examples described herein, but in other examples, multiple processors or computers may be used, or a processor or computer may include multiple processing elements or multiple types of processing elements or both. For example, a single hardware component or two or more hardware components may be implemented using a single processor, or two or more processors, or a processor and a controller. One or more hardware components may be implemented by one or more processors, or processors and controllers, and one or more other hardware components may be implemented by one or more other processors, or additional processors and additional controllers. One or more processors, or processors and controllers, may implement a single hardware component or two or more hardware components.The hardware components can have any one or more different processing configurations, examples of which include: a single processor, a discrete processor, a parallel processor, a single instruction single data (SISD) multiprocessing, a single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.
[0181] Figures 1 to 7 The methods for performing the operations described in this application, as shown, are executed by computing hardware (e.g., by one or more processors or a computer), which is implemented as described above to execute instructions or software to perform the operations performed by the methods described in this application. For example, a single operation or two or more operations may be executed by a single processor, or two or more processors, or a processor and a controller. One or more operations may be executed by one or more processors, or a processor and a controller, and one or more other operations may be executed by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, may execute a single operation or two or more operations.
[0182] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement hardware components and perform the methods described above can be written as computer programs, code segments, instructions, or any combination thereof to individually or collectively instruct or configure one or more processors or computers, such as machines or special-purpose computers, to perform operations performed by the hardware components and methods described above. In one example, the instructions or software include machine code (such as machine code generated by a compiler) that is directly executed by one or more processors or computers. In another example, the instructions or software include high-level code that is executed by one or more processors or computers using an interpreter. The instructions or software can be written using any programming language based on the block diagrams and flowcharts shown in the accompanying drawings and the corresponding descriptions in the specification, which disclose algorithms for performing operations performed by the hardware components and methods described above.
[0183] Instructions or software used to control computing hardware (e.g., one or more processors or computers) to implement hardware components and perform the methods described above, along with any associated data, data files, and data structures, may be recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards or microcards (e.g., Secure Digital (SD) or Extreme Digital (XD))), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and to provide the instructions or software and any associated data, data files, and data structures to one or more processors or computers, such that one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system, such that the instructions and software and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0184] While this disclosure includes specific examples, it will be clear upon understanding this disclosure that various changes in form and detail may be made to these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein are to be considered descriptive only and not for limiting purposes. The description of features or aspects in each example is to be considered applicable to similar features or aspects in other examples. Suitable results may be achieved if the described techniques are performed in a different order, and / or if components in the described system, architecture, apparatus, or circuit are combined in a different manner, and / or replaced or supplemented by other components or their equivalents. Therefore, the scope of the disclosure is not limited by the specific embodiments but by the claims and their equivalents, and all variations within the scope of the claims and their equivalents should be construed as included in the disclosure.
Claims
1. A method for training a neural network for image recognition, the method comprising: Receive input image set; A pre-trained neural network is executed with an input image set as input to obtain the result value inferred by the pre-trained neural network, and the first task accuracy of the pre-trained neural network is obtained based on the predetermined result value corresponding to the input image set and the result value inferred by the pre-trained neural network. The neural network is pruned on a channel-by-channel basis by adjusting the weights between nodes in multiple channels based on preset learning weights and channel-by-channel pruning parameters corresponding to the channels in multiple layers of the pre-trained neural network to determine the pruning threshold. The pruned neural network is executed with the input image set as input to obtain the result value inferred by the pruned neural network, and the task accuracy of the pruned neural network is obtained based on the predetermined result value and the result value inferred by the pruned neural network. The learning weights are updated based on the accuracy of the first task and the task accuracy of the pruned neural network. The channel-wise pruning parameters are updated based on the updated learning weights and the task accuracy of the pruned neural network. and The pruned neural network is re-pruned on a channel-by-channel basis based on the updated learning weights and the updated channel-by-channel pruning parameters.
2. The method according to claim 1, wherein, The channel-by-channel trimming parameters include a first parameter used to determine the trimming threshold.
3. The method according to claim 1, wherein, The pruning steps include: Trim the channel elements among the multiple channels to occupy 0 channels at a threshold ratio or a greater ratio.
4. The method according to claim 1, wherein, The steps for updating the learning weights include: in response to the task accuracy of the pruned neural network being less than the accuracy of the first task, updating the learning weights to increase the task accuracy of the pruned neural network.
5. The method according to any one of claims 1 to 4, further comprising: Repeatedly perform the pruning-evaluation operation to determine the task accuracy of the repruned neural network and the learning weights of the repruned neural network.
6. The method according to claim 5, further comprising: Whether to perform additional pruning-evaluation operations is determined based on the task accuracy of the neural network after a preset number of rounds and the pruning process.
7. The method according to claim 5, further comprising: In response to the repeated execution of the pruning-evaluation operation, the determined learning weights are compared with the lower bound threshold of the learning weights; and Based on the comparison results, determine whether to terminate the current pruning session and initiate a subsequent pruning session with the learning weights set to the initial reference values.
8. An apparatus for training a neural network for image recognition, the apparatus comprising: The processor is configured as follows: Receive input image set, A pre-trained neural network is executed with an input image set as input to obtain the inference results using the pre-trained neural network. The first task accuracy of the pre-trained neural network is then obtained based on the predetermined result values corresponding to the input image set and the inference results using the pre-trained neural network. The neural network is pruned on a channel-by-channel basis by adjusting the weights between nodes across multiple channels based on preset learning weights and channel-by-channel pruning parameters corresponding to the channels of each layer in the pre-trained neural network to determine the pruning threshold. The pruned neural network is executed with the input image set to obtain the inference results using the pruned neural network. The task accuracy of the pruned neural network is then obtained based on the predetermined result values and the inference results using the pruned neural network. The learned weights are updated based on the accuracy of the first task and the task accuracy of the pruned neural network. The channel-wise pruning parameters are updated based on the updated learned weights and the task accuracy of the pruned neural network. The pruned neural network is re-pruned on a channel-by-channel basis based on the updated learning weights and the updated channel-by-channel pruning parameters.
9. The device according to claim 8, wherein, The channel-by-channel trimming parameters include a first parameter used to determine the trimming threshold.
10. The device according to claim 8, wherein, For pruning, the processor is configured to prune channel elements included in the channels among the plurality of channels by a threshold ratio or a greater ratio that occupies 0 channels.
11. The device according to claim 8, wherein, To update the learning weights, the processor is configured to update the learning weights in response to the task accuracy of the pruned neural network being less than that of the first task, thereby increasing the task accuracy of the pruned neural network.
12. The device according to any one of claims 8 to 11, wherein, The processor is configured to repeatedly perform a pruning-evaluation operation to determine the task accuracy of the repruned neural network and the learning weights of the repruned neural network.
13. The device according to claim 12, wherein, The processor is configured to determine whether to perform an additional pruning-evaluation operation based on the task accuracy of the neural network after a preset number of rounds and the re-pruning process.
14. The device according to claim 12, wherein, The processor is configured to: in response to repeatedly performing the pruning-evaluation operation to compare the determined learning weights with a lower bound threshold of the learning weights, and based on the result of the comparison, determine whether to terminate the current pruning session and initiate a subsequent pruning session in which the learning weights are set to the initial reference value.
15. The device according to any one of claims 8 to 11, further comprising: The memory stores instructions that, when executed by a processor, configure the processor to perform the steps of receiving an input image set, obtaining a first task accuracy, pruning a neural network, obtaining the task accuracy of the pruned neural network, updating the learning weights, updating the channel-by-channel pruning parameters, and re-pruning the pruned neural network.
16. A method for training a neural network for image recognition, the method comprising: Receive input image set; A pre-trained neural network is executed with an input image set as input to obtain the result value inferred by the pre-trained neural network, and the first accuracy of the pre-trained neural network is obtained based on the predetermined result value corresponding to the input image set and the result value inferred by the pre-trained neural network. For each channel of the pre-trained neural network, the weights of the channel are pruned based on the pruning parameters used to determine the pruning threshold and the learned weights of the channel; A pruned neural network is executed with an input image set as input to obtain the result value inferred by the pruned neural network, and a second accuracy of the pruned neural network is obtained based on the predetermined result value and the result value inferred by the pruned neural network. The learned weights are updated based on a comparison between the first accuracy of the pre-trained neural network and the second accuracy of the pruned neural network. For each channel, the pruning parameters of the channel are updated based on the updated learning weights and the second accuracy; and For each channel, the channel weights are re-pruned based on the updated pruning parameters and updated learned weights of that channel.
17. The method according to claim 16, wherein, The steps for updating the learning weights include: updating the learning weights based on whether the second accuracy is greater than the first accuracy.
18. The method according to claim 17, wherein, The steps to update learning weights include: In response to the second accuracy being greater than the first accuracy, the learning weights are reduced; and In response to the second accuracy being less than or equal to the first accuracy, the learning weight is increased.
19. The method according to any one of claims 16 to 18, wherein, The steps for pruning weights include: The value of the weight transformation function is determined based on the trimming parameters; and The weights are trimmed based on the value of the weight transformation function.
20. The method according to any one of claims 16 to 18, wherein, The steps for pruning weights include pruning a larger number of weights in response to the first value, compared to pruning a second value in response to the first value being less than the first value.
21. A method for training a neural network for image recognition, the training method comprising: Store the pre-trained neural network in memory; The processor reads the pre-trained neural network from memory; The pre-trained neural network is pruned by the processor; The pruned neural network is stored in memory. The steps involved in pruning the pre-trained neural network using a processor include: Receive input image set; A pre-trained neural network is executed with an input image set as input to obtain the inference results using the pre-trained neural network. The first task accuracy of the pre-trained neural network is then obtained based on the predetermined result values corresponding to the input image set and the inference results using the pre-trained neural network. The neural network is pruned on a channel-by-channel basis by adjusting the weights between nodes across multiple channels based on preset learning weights and channel-by-channel pruning parameters corresponding to the channels of each layer in the pre-trained neural network to determine the pruning threshold. The pruned neural network is executed with the input image set as input to obtain the result value inferred by the pruned neural network, and the task accuracy of the pruned neural network is obtained based on the predetermined result value and the result value inferred by the pruned neural network. The learned weights are updated based on the accuracy of the first task and the task accuracy of the pruned neural network. The channel-wise pruning parameters are updated based on the updated learned weights and the task accuracy of the pruned neural network. The pruned neural network is re-pruned on a channel-by-channel basis based on the updated learning weights and the updated channel-by-channel pruning parameters.
22. A method for training a neural network for image recognition, the training method comprising: Store the pre-trained neural network in memory; The processor reads the pre-trained neural network from memory; The pre-trained neural network is pruned by the processor; The pruned neural network is stored in memory. The steps involved in pruning the pre-trained neural network using a processor include: Receive input image set; A pre-trained neural network is executed with an input image set as input to obtain the result value inferred by the pre-trained neural network, and the first accuracy of the pre-trained neural network is obtained based on the predetermined result value corresponding to the input image set and the result value inferred by the pre-trained neural network. For each channel of the pre-trained neural network, the weights of the channel are pruned based on the pruning parameters used to determine the pruning threshold and the learned weights of the channel; A pruned neural network is executed with an input image set as input to obtain the result value inferred by the pruned neural network, and a second accuracy of the pruned neural network is obtained based on the predetermined result value and the result value inferred by the pruned neural network. The learned weights are updated based on a comparison between the first accuracy of inference performed using a pre-trained neural network and the second accuracy of inference performed using a pruned neural network. For each channel, the pruning parameters of that channel are updated based on the updated learning weights and the second accuracy; and For each channel, the channel weights are re-pruned based on the updated pruning parameters and updated learned weights of that channel.
23. A neural network device, the neural network device comprising: The processor is configured as follows: Receive input image set, A pre-trained neural network for image recognition is executed with an input image set as input to obtain result values inferred using the pre-trained neural network. A first task accuracy of the pre-trained neural network is obtained based on predetermined result values corresponding to the input image set and the result values inferred using the pre-trained neural network. The neural network is pruned on a channel-by-channel basis by adjusting the weights between nodes across multiple channels based on preset learning weights and channel-by-channel pruning parameters corresponding to the channels of each layer in the pre-trained neural network to determine the pruning threshold. The pruned neural network is executed with the input image set to obtain the inference results using the pruned neural network. The task accuracy of the pruned neural network is then obtained based on the predetermined result values and the inference results using the pruned neural network. The learned weights are updated based on the accuracy of the first task and the task accuracy of the pruned neural network. The channel-wise pruning parameters are updated based on the updated learned weights and the task accuracy of the pruned neural network. The pruned neural network is re-pruned on a channel-by-channel basis based on the updated learning weights and the updated channel-by-channel pruning parameters.
24. A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, configure the processor to perform the method according to any one of claims 1 to 7 and 16 to 22.
Citation Information
Patent Citations
Practice method for surgery using microscope and Practice device for the same
KR1020200132151A
Apparatus and method of compressing neural network
US20200184333A1