Efficient second-order pruning of computer-implemented neural networks

JP2023016023A5Pending Publication Date: 2025-07-29ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022114585
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-20
Filing Date
2022-07-19
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing pruning methods for neural networks often degrade the overall performance or accuracy of the network, and fail to predict the impact of pruning multiple network structures accurately, making it difficult to simplify neural networks without compromising their functionality.

Method used

A method involving the use of pruning vectors and Hessian matrix derivatives to determine the impact of pruning neural network structures, allowing for the generation of a simplified neural network with minimal performance loss by identifying non-essential structures.

Benefits of technology

The method generates a simplified neural network with fewer neurons and layers, suitable for devices with limited resources, maintaining high accuracy and computation speed, while avoiding coarse approximations that overlook structural correlations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for generating a simplified computer-implemented neural network.SOLUTION: A neural network 1a includes a plurality of neural network structures, each neural network structure being assigned a pruning vector which describes a change in a plurality of weights as a result of pruning the neural network structure. The method further includes calculating a product of a matrix and a structure vector, the matrix including second order partial derivatives of a loss function with respect to the plurality of weights, where the weights correspond to components of the structure vector. The method further includes determining two or more changes in the loss function with respect to the predefined neural network, and determining the two or more changes in the loss function using the calculated product, the respective pruning vector, and a current plurality of weights of the predefined neural network. The method finally includes pruning the neural network structure based on the determined two or more changes in the loss function.SELECTED DRAWING: Figure 2a
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Technical Field The present invention relates to techniques for generating a simplified computer-implemented neural network from a given neural network. Corresponding aspects relate to a computer program product and a computer-implemented system. [Background technology]

[0002] background Computer-implemented neural networks are increasingly used in a variety of technical devices. In this regard, neural networks for many technical applications may have a complex structure (e.g., with a large number of neurons, layers, and corresponding links). This may impose (excessively) high demands on the hardware required to apply the neural network. Therefore, a compromise must be found between the performance of a computer-implemented neural network and its complexity.

[0003] To address the above-mentioned problems, so-called pruning methods have been developed. These pruning methods aim, on the one hand, to reduce the size of neural networks and simplify their overall structure, while, on the other hand, to maintain good overall performance of the simplified neural networks (or to minimize the overall performance loss as much as possible). Therefore, neural networks simplified using these methods can be used, for example, for small technical devices with limited hardware resources (e.g., power tools, gardening equipment, or household appliances). In other examples, it may be necessary to reduce the evaluation time of a computer-implemented neural network to ensure a sufficiently fast response of a technical device (e.g., an autonomous robot). For this reason, simplifying a computer-implemented neural network is advantageous. Summary of the Invention [Problem to be solved by the invention]

[0004] However, some prior art pruning methods suffer from the problem that the approximations used result in a degradation of the overall performance or accuracy of the neural network generated by the pruning method compared to the original neural network. Furthermore, some prior art pruning methods are unable to predict how pruning multiple network structures will affect the overall performance of the remaining network structures of the neural network due to possible correlations between these network structures. Therefore, it is often a significant effort to prune multiple network structures in the original neural network without degrading the overall performance or accuracy.

[0005] Therefore, there is a need to develop new and efficient techniques for generating simplified computer-implemented neural networks for devices that can solve some or all of the above problems. [Means for solving the problem]

[0006] Summary of the Invention A first general aspect of the present disclosure relates to a method for generating a simplified computer-implemented neural network. The method includes receiving a given neural network, the given neural network including a plurality of neural network structures and described by a plurality of weights. A pruning vector is associated with each neural network structure from the plurality of neural network structures of the first aspect, the pruning vector describing changes to the plurality of weights resulting from pruning each neural network structure. The method further includes calculating a product of a matrix and the structure vector, the matrix including second-order partial derivatives of a loss function with respect to the plurality of weights, and each weight from the plurality of weights belonging to two or more neural network structures to be pruned from the plurality of neural network structures corresponds to a respective component of the structure vector. In a next step, the method of the first aspect includes determining two or more changes to the loss function for the given neural network, each change from the two or more changes resulting from pruning a corresponding neural network structure from the two or more neural network structures to be pruned. Further, determining the two or more changes in the loss function is performed using the calculated product, the respective pruning vectors, and the current weights of the given neural network. Finally, the method includes pruning at least one neural network structure from the plurality of neural network structures based on the determined two or more changes in the loss function to generate a simplified neural network.

[0007] A second general aspect of the present disclosure relates to a computer program configured to perform a computer-implemented method according to the first general aspect of the present disclosure.

[0008] A third general aspect of the present disclosure relates to a computer-implemented system for generating and / or applying a computer-implemented neural network for an apparatus, the system being configured to perform a method according to the first general aspect of the present disclosure. Additionally or alternatively, the computer-implemented system of the third general aspect is configured to execute a computer program according to the second general aspect of the present disclosure.

[0009] The techniques of the first through third general aspects may have one or more of the following advantages.

[0010] First, the present technique can enable the generation of a simplified neural network of a smaller size (e.g., having a smaller number of neurons and / or links and / or layers) compared to the original given neural network. In this case, the generated simplified neural network should not suffer too much loss in overall performance or accuracy (ideally, no loss in overall performance or accuracy should occur). Such simplified (pruned) computer-implemented neural networks of the present technique may be suitable for technical devices with relatively reasonable hardware resources (e.g., portable electrical devices or devices without permanent network connections) or in technical environments where relatively high calculation and evaluation speeds are useful (e.g., in at least partially autonomous vehicles). Therefore, such pruned neural networks can be particularly adapted for these resource-saving technical environments and / or for technical environments with higher requirements for calculation and evaluation speeds.

[0011] Second, the techniques of the present disclosure provide a means, particularly for complex neural networks, to better and more quickly estimate which neural network structures make a small or even negligible contribution to the overall performance of a given original neural network than some prior art techniques.

[0012] Third, the present technique does not use coarse approximations (e.g., using diagonal or near-diagonal matrices for each Hessian matrix), as is the case in some known prior art methods. As a result, the present technique can enable more efficient and more accurate determination of the network structure to be pruned compared to some prior art methods.

[0013] In this disclosure, several terms are used as follows:

[0014] The term "neural network" is understood to mean any artificial neural network that may have a particular topology and a number of neurons with corresponding links (see also the following description). According to some embodiments, the neural network may be a convolutional neural network ("convolutional neural network" in English, or "CNN" for short), defined, for example, by the number of filters, filter size, step size, etc. For example, a convolutional neural network may be used for image classification purposes, and may perform one or more transformations on a digital image, for example, based on convolution (e.g., using fully connected layers), nonlinearities (ReLU), pooling, or classification algorithms. A neural network may also be configured as a multi-layer feedforward or recurrent network, as a neural network with direct or indirect feedback, or as a multi-layer perceptron. These neural networks can be used in a vehicle computer or other components of a vehicle, or in at least partially autonomous robots (e.g., to assess the operating state of a vehicle or robot and / or to control the functions of a vehicle or robot based on state data and / or environmental data of the vehicle or robot as input data). This list of examples is not exhaustive (further examples are provided below).

[0015] Thus, the term "network structure" (hereinafter sometimes referred to as "structure" for short) includes any subset of elements of a neural network, such as neurons with their respective weights and / or links, that may be arranged in the form of layers of a neural network, as described below. A network structure may also include an entire layer of a neural network.

[0016] The term "pruning ratio" of a neural network is understood to be any variable that can characterize the extent to which a pruned neural network has been changed compared to the original neural network (e.g., with respect to the number of neural network structures pruned or remaining, or with respect to the desired overall performance of the pruned neural network). The pruning ratio of a neural network can be defined, for example, as the ratio between the number of neural network structures pruned from the neural network and the (original) number of neural network structures in the neural network.

[0017] The term "randomly initialized neural network" means that the initial (i.e., initial) weights for a neural network are selected as random numbers or initialized with different values ​​(e.g., to break up the symmetric distribution of the initial weights of the neural network).

[0018] Thus, a "trained neural network" is a neural network that has been trained, for example, by performing a stochastic gradient descent or other method using one dataset (also referred to in the present disclosure as a training dataset) or multiple datasets (e.g., relating to sensor data) so that a corresponding loss function is minimized by training (within a predetermined numerical precision and / or until a predetermined termination criterion is reached). In this context, the term "incompletely trained neural network" is understood to be, for example, a neural network whose corresponding loss function has not yet been minimized by training (within a predetermined numerical precision and / or until a predetermined termination criterion is reached), as will be explained in more detail below. For example, a "randomly initialized neural network" can be trained to become a "(incompletely) trained neural network."

[0019] Hereinafter, the term "device" refers to any device to which a computer-implemented neural network can be applied / used to control and / or monitor the device, such as a vehicle (e.g., an automobile, such as a car, that is operated / assisted at least partially autonomously, or a ship, train, aircraft, or spacecraft), a vehicle computer, a partially autonomous or autonomous robot (e.g., an industrial robot or machine), or a group thereof (e.g., factory equipment), a tool, a home or garden appliance, or a monitoring device (further examples are also provided below). For example, a computer-implemented neural network can be used in these devices to classify or otherwise evaluate status or environmental data (e.g., image data) collected for these devices (e.g., by corresponding sensors). The classification or evaluation results can be used to control and / or monitor the device. [Brief explanation of the drawings]

[0020] [Figure 1a] 1 is a flowchart illustrating an example of a method for generating a simplified computer-implemented neural network according to a first embodiment. [Figure 1b] 4 is a flow chart showing further possible method steps according to the first aspect; [Figure 1c] 4 is a flow chart showing further possible method steps according to the first aspect; [Figure 2a] 1 shows a schematic representation of a given neural network before and after pruning (simplification), with neurons 2 and their links 3 shown as nodes (circles) or edges (arrows). [Figure 2b] 1 shows a schematic diagram of a pruning vector 8;δp associated with node p of a neural network, describing the change in current weights due to pruning the links of node p (all components of the pruning vector are set to zero except for components 7a, 7b corresponding to the weights belonging to node p to be pruned). [Figure 3a]Figure 10 shows a schematic of four different pruning scenarios in which the present technique can be implemented: (a) First, a simplified randomly initialized neural network is generated by pruning from a randomly initialized neural network, and then trained 10a. [Figure 3b] 10A-10C are schematic diagrams illustrating four different pruning scenarios in which the present technology can be implemented. (a) First, a randomly initialized neural network is trained, and then a simplified trained neural network is generated by pruning from the trained neural network. Finally, the generated simplified trained neural network is (further) trained. [Figure 3c] 10A-10C are diagrams illustrating four different pruning scenarios that can be implemented by the present technology. (c) First, a randomly initialized neural network is trained, and then two or more neural network structures to be pruned from the trained neural network are determined. Next, a randomly initialized neural network is pruned based on the determined two or more neural network structures to be pruned from the trained neural network. Finally, the generated simplified randomly initialized neural network is trained. [Figure 3d]10(d) is a diagram schematically illustrating four different pruning scenarios possible with the present technology. (d) First, two or more neural network structures to be pruned of a randomly initialized neural network are determined, and the randomly initialized neural network is trained. Then, based on the determined two or more neural network structures to be pruned of the randomly initialized neural network, the corresponding trained neural network is pruned to generate a simplified trained neural network. Finally, the generated simplified trained neural network is (further) trained 10(d). The notation "NN" refers to the corresponding neural network. The notation "Mask" symbolically indicates two steps of the present method, by identifying a network structure to be pruned for a neural network (arrows leading into the "Mask" block) and applying it to the corresponding neural network (arrows leading out of the "Mask" block). See the description below. [Figure 4] Figure 4 shows the mean classification accuracy 20, 22 for the pruning methods used in Figures 3a and 3b, and the error of this mean (vertical bars; three trials were used), based on different pruning ratio values. The results of the present technique are compared with the results of a random pruning method 21, 23, which randomly selects (using a uniform distribution) a structure to be pruned from multiple neural network structures until the desired pruning ratio (x-axis in Figure 4) is achieved. The method of the first embodiment was applied to the convolutional neural network "DenseNet-40-BC" using the "Cifar10" dataset. [Figure 5]FIG. 12 shows the distribution of layer pruning ratios at initialization 12a (corresponding to the scenario shown in FIG. 3a) and after training 12b (corresponding to the scenario shown in FIG. 3b) as a function of layer index of the convolutional neural network “DenseNet-40-BC” for the trials of FIG. 4. (The layer index numbers the layers of the convolutional neural network “DenseNet-40-BC,” which includes multiple layers.) The pruning ratio of the multiple neural network structures of the convolutional neural network “DenseNet-40-BC” is 50% in the example of FIG. 4 (which corresponds to the value 0.5 on the x-axis of FIG. 4). DETAILED DESCRIPTION OF THE INVENTION

[0021] Detailed Description First, a technique for generating a simplified computer-implemented neural network is described based on Figures 1a-1c. Then, an example given neural network before and after pruning and the pruning vector δ are shown. p This will be explained based on Figures 2a and 2b. Next, Figures 3a to 3d show four example pruning scenarios of the present technology in a simplified manner. Finally, Figures 4 and 5 illustrate further aspects of the pruning method of the present disclosure.

[0022] 1a-1c, a first general aspect relates to a method for generating a simplified computer-implemented neural network. As described in more detail below, a simplified neural network is a neural network of smaller size (e.g., having fewer neurons and / or links and / or layers) compared to the original neural network from which it was generated. For example, the simplified neural network may include 90% or less, 70% or less, 50% or less, or 30% or less of the number of neurons and / or links and / or layers of the original neural network.

[0023] FIG. 2a (schematically illustrates an exemplary neural network 1). A neural network can be formed from multiple neurons (an exemplary neuron 2 is highlighted in FIG. 2a), which form nodes of the neural network 1 and are interconnected via edges 3. The neurons of the computer-implemented neural network of the present disclosure are arranged in layers (e.g., the third layer 4 of FIG. 2a includes three neurons). In the present disclosure, the incoming edges or links of a neuron (or node) are considered to be part of the respective layer (i.e., the incoming link and the node are located in the same layer). A computer-implemented neural network can include two or more, three or more, or five or more layers. Neurons and their links can have different structures and can be represented as nodes or edges using a graph. As noted above, any subset of elements of a neural network is referred to as a "neural network structure." In some examples, a neural network structure may include one or more edges, one or more nodes, or a combination of one or more nodes and edges (e.g., a node with an edge leading to the node and an edge leading away from the node). Networks different from the network illustrated in FIG. 2a may include additional elements (e.g., feedback or memory). Such elements may also be part of or be part of the neural network structure. In other examples, the elements of a neural network may be parameters for describing the neural network (this aspect is described in more detail below).

[0024] The output of a given neuron j may depend on the applied inputs of one or more neurons i. In particular, first, for neuron j, we can form a weighted sum of the applied inputs, and then, for all neurons, we can calculate the θ ij We can define weights of the form (θij (Note that θ = 0 can mean that neuron j has no link with neuron i.) Finally, the output of neuron j can be calculated after applying the activation functions defined for each neuron to the pre-computed sum. Thus, in some examples, the topology of the neural network and / or the weights θ for all neurons can be calculated. ij A neural network can be defined by specifying the weights of the neural network. These weights can therefore also be elements of the neural network structure in the sense of the present disclosure. That is, the neural network structure can include one or more weights of the neural network (these weights can correspond to one or more edges in the graphical description of the neural network). According to the terminology explained above, all weights of links or edges coming into a node or neuron of a given layer belong to this layer. Weights of edges or links going out of this given layer will belong to other layers.

[0025] A computer-implemented neural network can be created and trained (e.g., fully or partially trainable) for use in a device to process data (e.g., sensor data) generated by the device and calculate output data to be used, for example, for monitoring and / or controlling the device (applications of computer-implemented neural networks are described below). Thus, the characteristics of the device, or its response to a given event, can ultimately be determined by the topology and weights θ of the neural network. ij In another example, the neural network can be "hidden" in the weights θ ij In the present disclosure, the weight θ ij are discussed as exemplary parameters.

[0026] Some methods used in the prior art for pruning the network structure from a trained neural network are implemented as follows: After neural networks with different topologies are generated depending on the task, the weights θ ij can be selected accordingly. In some examples, this selection of weights can be referred to as training or learning the neural network, starting from a random initialization of the weights. This step is performed on a computer system. During "supervised learning," multiple input data x k (e.g., sensor data) and a corresponding plurality of desired output data y k (For example, the state of the technical device or the state of the environment of the technical device, or the control variable) can be used (i.e., the input data and the output data each form one pair). k ,y k ), k=1,...,N is called the training dataset. Training a neural network can be formulated as an optimization problem in which, given an input x k The output generated by the neural network for

number

number

number

[0027] Minimization is done by over all weights θ ij As a result of this minimization, the trained weights

number

number

[0028] The trained neural network has a complex topology and a large number of neurons and links (e.g., 10 4 More than 10 5 , or more than one link), thus resulting in undesirably high computational hardware requirements. Furthermore, as already mentioned above, this complex neural network is first simplified by pruning so that it can be used on a corresponding device. The right side of FIG. 2a shows a simplified (pruned) neural network 1a. For example, between the top layer and the layer below it, 5, some edges (links) have been pruned. Furthermore, in the penultimate layer, 6, one node has been pruned together with its associated edges.

[0029] This simplification is achieved by pruning one or more neural network structures to reduce the loss function L D , and analyzing the changes in each of the neurons. Pruning the structure may involve deleting one or more links (i.e., edges) between neurons and / or completely deleting one or more neurons along with their incoming and outgoing links. In other cases (or other diagrams), pruning may involve deleting or zeroing one or more weights (which may be an alternative description of deleting one or more links or edges). In other cases, the neural network may include elements (e.g., feedback or memory) beyond the structure shown in FIG. 2a. Such network structures may also be pruned by the methods of the present disclosure.

[0030] In some examples, the change in the loss function due to pruning the structure is calculated by the weights θ ij Loss function L for D can be approximated to a given order by a Taylor expansion of, for example, an expansion of the form:

number

[0031] where δθ is the weight of a given neural network (e.g.,

number

number

number

number

number

number

number

number

[0032] Loss function L D The formula for the Taylor expansion of is shown only as an example, and other formulas can be adopted depending on the normalization selected for each vector (e.g., δθ and δθ T (One can also assume a factor of 1 / 2 in the second term of the Taylor expansion shown above in

[0033] When multiple neural network structures are pruned, the change in the loss function can be expressed as a sum over multiple elements of a matrix that describes the change in the loss function due to pruning one or more neural network structures from the multiple neural network structures. For example, the change in the loss function can be expressed in the following form: δL D (θ)≒1 / 2Σ beschn.pq Q pq can be determined in the form: pq teeth,

number

[0034] The formula for the change in the loss function seems simple, but δL D There can be some difficulties in calculating (θ). First, the Hessian matrix H(θ) (and thus the matrix Q) pq ) is usually very large, P×P, where P denotes the total number of links in the neural network. For this reason, the Hessian matrix is ​​approximated by a diagonal matrix (or nearly so) in some previously known methods in the prior art. However, this approximation ignores possible correlations between network structures, e.g., between an individual network structure and all other network structures from a set of neural network structures (this aspect will be explained in more detail below). Moreover, this can lead to (partially significant) inaccuracies when assessing which network structures have an impact on the performance of the neural network. Second, δL D The calculation of (θ) involves a number of computational steps on a computer-implemented system, which depend on the number of training data sets N and the pruning vector δ p With the dimension P of 2) Furthermore, the number of computational steps is completely independent of the fact that the respective dimension L of the network structure being considered for pruning may be much smaller than the total number of links in the neural network P. As a result, the above evaluation for complex structures with large P may be computationally difficult to perform. Third, an additional problem may arise in that training neural networks (especially relatively large and complex neural networks) may be a computationally very expensive and therefore time-consuming task. These problems can be addressed by the techniques of the present disclosure in some implementations.

[0035] The first step of the method for generating a simplified computer-implemented neural network of the present disclosure involves receiving 100 a given neural network 1, where the given neural network includes multiple neural network structures S and is described by multiple weights 7a and 7b. For example, the given neural network 1 may include multiple neural network structures S in the form of one or more layers 4-6. In some examples, multiple neural network structures form a given neural network. In other examples, multiple neural network structures form part of a given neural network (e.g., one or more neurons with their respective links, an input layer or an output layer, one or more hidden layers, or a combination thereof may belong to multiple neural network structures). In the present technology, each layer may be provided by, for example, multiple neurons and corresponding incoming links. In this case, the weighted output of a neuron located in one layer may be the input of another neuron located in another layer. For example, the neural network 1 of FIG. 2a includes four layers. For a given neural network, the current weights are given by the corresponding numerical values.

number

[0036] In the technique of the present disclosure, a pruning vector δ is assigned to each of the neural network structures from among the plurality of neural network structures. p are associated with each other, and the pruning vector δ p describes the weight changes due to pruning each neural network structure. In other words, the pruning vector δ p considers that pruning a corresponding structure p changes the architecture of a given network, for example, because the output of a neuron (i.e., node) is changed by pruning one or more links (i.e., edges) coming into this neuron. Furthermore, one or more changed outputs of a neuron may be the corresponding inputs of one or more other neurons, so that the pruning vector δ pIt may be possible to consider that such a change in a node may propagate further within the neural network (e.g., to other neurons in the same layer or other layers). In some examples, each pruning vector for a corresponding neural network structure to be pruned may include weights from multiple weights belonging to the corresponding network structure to be pruned as components 7a, 7b. Furthermore, corresponding other components 7c of the pruning vector, which correspond to weights that do not belong to the corresponding network structure to be pruned, may be set to zero. This situation is shown, for example, in FIG. 2b.

[0037] In a further step of the method according to the invention, a matrix (for example the Hessian matrix H(θ) as explained above) and a structure vector δ struc and the product (H(θ)δ struc ) is calculated, and the matrix is ​​a loss function L D Furthermore, each weight from the plurality of weights belonging to two or more neural network structures to be pruned from the plurality of neural network structures is included in the structure vector δ struc (e.g., each component of the structure vector can be proportional to a respective weight). Furthermore, the present technique can be applied to two or more variations of the loss function λ for a given neural network. p , where each change from the two or more changes results from pruning a corresponding neural network structure from the two or more neural network structures to be pruned. In other words, the two or more changes λ pis associated with two or more corresponding network structures to be pruned. Further, in some examples, two or more neural network structures to be pruned may exist within different layers of a given neural network. Additionally or alternatively, two or more neural network structures from multiple neural network structures can be arranged within one (single) layer of a given neural network. In some examples, the number M of two or more neural network structures to be pruned can be made less than the number S of neural network structures in the multiple neural network structures, i.e., M < S can be achieved (e.g., the corresponding M / S ratio can be a value of 0.9 or less, 0.5 or less, 0.1 or less). In other examples, the number M of two or more neural network structures to be pruned can be selected to be equal to the number S of neural network structures in the multiple neural network structures S, i.e., M = S can be selected.

[0038] Furthermore, determining two or more changes of the loss function can be implemented using the calculated product, each pruning vector, and the current multiple weights of a given neural network [Number] (e.g., the corresponding numerical values as described above). Further, the method may, in some cases, include receiving a dataset (X) that describes the behavior of the device, where in this case, the dataset (X) is composed of a plurality of pairs (X = {(x i , y i ) | i ≤ n}). Each pair can be formed, for example, from input data and respective output data, and a given neural network generates the respective output data for the input data of each pair. In some examples, the loss function min(L DAs can be concluded from the example equation above for ), this step of the method may be necessary to determine more than one change in the loss function. For example, the data set can ultimately be substituted into equations for each of the more than one change in the loss function.

[0039] In this context, the dataset (x k ,y k ) can contain different kinds of data, and in each pair, the input data (x k ) and output data (y k ) are paired (k=1,...,N). For example, the input data and output data may each be a scalar (e.g., a scalar measurement), a vector of any length (i.e., a length greater than or equal to 1), or a matrix. The input data may represent an environmental influence or an internal operating state of the device. In one example, the input data may include sensor data. Alternatively or additionally, the input data may include image data and / or audio data. The output data may be a state of the device or the environment, or a recognized event of the device or the environment (e.g., a "battery almost empty" state or a "raining" state for an electrical device). In a further example, the output variable may be a control variable (e.g., for an actuator) or may otherwise identify a response of the device. In either case, the output data may be usable for controlling the device (e.g., selecting an operating parameter or operating mode of the device) and / or monitoring the device.

[0040] The techniques of the present disclosure ultimately include pruning 400 at least one neural network structure from the plurality of neural network structures based on the determined two or more changes in the loss function to generate a simplified neural network 1a. In some examples, the values ​​of the determined two or more changes in the loss function can be used to identify which neural network structures make a small or even negligible contribution to the overall performance of the original given neural network (e.g., the overall performance does not degrade beyond a predetermined measure). For example, in this regard, after pruning, the loss function δL D Only network structures that do not cause an increase in (θ) or that cause an increase that does not exceed a predetermined measure can be classified as network structures to be pruned. Therefore, such classified network structures can be pruned from a given neural network to generate a simplified neural network for a device. The resulting simplified neural network can provide data for the device more quickly and / or require fewer hardware resources.

[0041] In this technique, the matrix is ​​a loss function (L D The Hessian matrix (H(θ)) contains the second partial derivatives of the vector (δ(θ)). In this case, the product is the Hessian matrix and the structure vector (δ struc ) may be a Hessian vector product with the pruning vector δ . In the present disclosure, determining each change from two or more changes in the loss function resulting from pruning each neural network structure from two or more neural network structures to be pruned may be performed by using a pruning vector δ . pand the calculated product (e.g., the Hessian vector product described above). As explained above, each component of the structure vector may be proportional to a respective weight, and each weight from the plurality of weights belonging to the two or more neural network structures to be pruned may be included in the structure vector δ. struc Thus, the structure vector may in some instances be a weighted sum δ of pruning vectors belonging to two or more neural network structures to be pruned. struc =Σ q∈M μ q δ q where μ q are the corresponding weight coefficients, and M is the number of neural network structures to be pruned. In some cases, all weight coefficients can be the same, e.g., μ q = 1. In this case, the structure vector is the sum δ of pruning vectors belonging to two or more neural network structures to be pruned. struc =Σ q∈M δ q Furthermore, in some cases, for example, if the ratio between the number of two or more neural network structures to be pruned and the number of neural network structures in the plurality of neural network structures, i.e., M / S, is above a predefined threshold corresponding to a high pruning ratio (e.g., M / S is 0.1 or more, 0.5 or more, or 0.8 or more), the structure vector can be expressed as the sum of all pruning vectors of the plurality of neural network structures (S).

number

[0042] The first contribution, as defined in this disclosure, is the respective change in the loss function λ p Weight θ for ij Regarding the loss function L D Furthermore, as noted above, the matrix described above may in some cases be taken to be the Hessian matrix H(θ), and therefore the product is the Hessian matrix and the structure vector δ struc Therefore, in some examples, the respective change in the loss function λ due to pruning the network structure p p The first contribution to is the pruning vector δ associated with each neural network structure. p and the Hessian vector product

number

number

number

number

[0043] Furthermore, the step 300 of "determining" two or more changes in the loss function may further include determining two or more respective pruning vectors δ p and the current weights of a given neural network

number

number

number

number

number

number

[0044] In some examples of the present technology, each change λ from two or more changes in the loss function due to pruning the network structure p pmay be proportional to or equal to the respective first contributions. p may be proportional to or equal to the absolute value of the first contribution. As explained above, this means that the gradient ∂L D This may be the case when (θ) / ∂θ is negligibly small, and thus when each second contribution can be ignored relative to each first contribution. In yet another example, determining each change from two or more changes in the loss function may include calculating 330 a first product by multiplying each first contribution by a first weighting factor. In a next step, the "determining" step may include calculating 340 a second product by multiplying each second contribution by a second weighting factor. Furthermore, the "determining" step may include summing 350 the absolute value of the first product and the absolute value of the second product to calculate the respective change. In this case, for example, the respective change in the loss function due to pruning the network structure p may be calculated using the following formula:

number

number

[0045] In the present disclosure, the step of "pruning" based on the determined two or more changes in the loss function further includes pruning 410 a neural network structure corresponding to the smallest change from the two or more changes in the loss function. In some examples, for this purpose, the two or more changes can first be ordered (e.g., in ascending or descending order). Next, the method according to the present invention includes iteratively pruning 420 two or more corresponding neural network structures from the two or more neural network structures to be pruned, each subsequent neural network structure to be pruned corresponding to the next largest value of the two or more changes (e.g., the ordered two or more changes). Generally, the iterative pruning can be performed until the size of the simplified neural network is below a desired size. In one example, the desired size can be given by the minimum number of neurons in the simplified neural network or in one layer of the simplified neural network. In other examples, the desired size can be defined by a minimum number of links between neurons in the simplified neural network or by a minimum number of links between neurons in one layer of the simplified neural network. The desired size can also be given, for example, as a minimum number of unpruned layers or structures of the simplified neural network. In other examples, the pruning method is performed until the overall performance of the simplified neural network falls below a predefined threshold. For example, the overall performance can be estimated using accuracy (e.g., the accuracy of the classification results) (see, for example, FIG. 4), which itself can be calculated based on a loss function. In one example, the predefined threshold can be defined as the ratio between the overall performance of the simplified neural network and the overall performance of a given neural network. In other examples, the predefined threshold can correspond to a selected number. In some examples, the overall change in the loss function can be defined as the sum of two or more changes in the loss function corresponding to the neural network structures to be pruned. For example, if M network structures are pruned, the overall change is λ Ges =Σq∈M λ q It can be written as:

[0046] The present disclosure may further include providing a randomly initialized neural network. In some examples, weights of the randomly initialized neural network may be initialized to random numbers randomly distributed over a given interval. In some cases, a random number generator that provides evenly or unevenly distributed values ​​over a given interval is used for this purpose. In other examples, weights of the randomly initialized neural network may be initialized with different (not necessarily random) values ​​(e.g., to break up the symmetric distribution of the initial weights).

[0047] In some examples, as described above, the given neural network may be a randomly initialized neural network. In this case, a simplified randomly initialized neural network may be generated from the randomly initialized neural network, where two or more neural network structures of the randomly initialized neural network are pruned based on the determined two or more changes in the loss function for the randomly initialized neural network. The generated simplified randomly initialized neural network may then be trained (e.g., using the received dataset (X)). This scenario 10a can be seen, for example, in FIG. 3a. In some cases, the resulting neural network may be retrained using a dataset different from the received dataset (X), e.g., to achieve the same number of overall training steps as the method summarized in FIGS. 3b-3d (there are two overall training steps shown in FIGS. 3b-3d; see subsequent discussion for further details). Figure 4 (left part 11a) shows the mean classification accuracy 20 and the error of this mean for the pruning method used in Figure 3a (vertical bars; three trials were used), each based on different pruning ratio values. A pruning ratio of zero corresponds to a randomly initialized neural network that is not pruned. The results of our technique are compared with the results of each random pruning method 21, which randomly selects (using a uniform distribution) from multiple neural network structures of a given neural network until the desired pruning ratio (x-axis in Figure 4) is achieved. It can be seen that the classification accuracy achieved by using our technique is better than that obtained based on the random pruning method for pruning ratio values ​​greater than 0.2 and less than 0.7.Furthermore, Figure 5 (left part 12a) shows an exemplary distribution of layer pruning ratios for a trained simplified randomly initialized neural network (corresponding to the scenario shown in Figure 3a) as a function of the layer index of the neural network for the trial in Figure 4 (left part 11a). The present technique has been applied to the convolutional neural network "DenseNet-40-BC" (see the following link https: / / arxiv.org / abs / 1608.06993v5) using the "Cifar10" dataset (see also "Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700-4708, 2017") (see the following link https: / / www.cs.toronto.edu / ~kriz / learning-features-2009-TR.pdf).

[0048] In an alternative method, a randomly initialized neural network is first trained to generate a trained neural network. In this case, the generated trained neural network is selected as the given neural network. In some cases, the received data set (X) described above for training the randomly initialized neural network can be used for this purpose. Alternatively or additionally, other data sets can be used in this regard. Next, the method can include generating a simplified trained neural network from the generated trained neural network, where two or more neural network structures of the generated trained neural network are pruned based on the determined two or more changes in the loss function for the generated trained neural network. Finally, the generated simplified trained neural network can be trained. In some cases, training can be performed using the received data set (X). This scenario 10b is summarized, for example, in FIG. 3b. Figure 4 (right-hand portion 11b) shows the mean classification accuracy 22 and the error of this mean for the pruning method used in Figure 3b (vertical bars; three trials were used), respectively, based on different pruning ratio values. The results of our technique are compared with the results of a random pruning method 23, which randomly selects (using a uniform distribution) a structure to be pruned from multiple neural network structures of a given neural network until the desired pruning ratio (x-axis in Figure 4) is achieved. It can be seen that the classification accuracy achieved using our technique is better than that obtained based on the random pruning method for pruning ratio values ​​greater than 0.1 and less than 0.7 (i.e., the entire interval shown). Furthermore, the classification accuracy using our technique is better when the given neural network corresponds to the generated trained neural network (curve 22 is significantly higher than curve 20 in Figure 4).In other words, in some cases, it may be necessary to first train and then prune a randomly initialized neural network to achieve a predetermined accuracy (e.g., 85% or higher or 90% or higher) and corresponding overall performance. Furthermore, Figure 5 (right-hand portion 12b) shows an exemplary distribution of layer pruning ratios for a trained simplified neural network (corresponding to the scenario shown in Figure 3b) as a function of neural network layer index for the trial of Figure 4 (right-hand portion 11b).

[0049] In a further alternative method, a randomly initialized neural network is first selected as the given neural network. In a next step, the randomly initialized neural network is trained to generate a trained neural network. In some cases, the received dataset (X) described above for training the randomly initialized neural network can be used for this purpose. Alternatively or additionally, other datasets can be used in this regard. Thereafter, the method can include determining two or more neural network structures to be pruned from the generated trained neural network based on the determined two or more changes in the loss function for the generated trained neural network. Next, the method can include pruning two or more corresponding neural network structures of the randomly initialized neural network based on the determined two or more neural network structures to be pruned from the generated trained neural network to generate a simplified randomly initialized neural network. Finally, the generated simplified randomly initialized neural network can be trained. In some cases, training can be performed using the received dataset (X). This scenario 10c is shown, for example, in FIG. 3c.

[0050] A further alternative method may include first determining two or more neural network structures to be pruned of the randomly initialized neural network based on the determined two or more changes in the loss function for the randomly initialized neural network. The technique may then include training the randomly initialized neural network to generate a trained neural network, where the generated trained neural network is a given neural network. In some cases, the received dataset (X) described above for training the randomly initialized neural network may be used for this purpose. Alternatively or additionally, other datasets may be used in this regard. Next, the method may include pruning two or more corresponding neural network structures of the generated trained neural network based on the determined two or more neural network structures to be pruned of the randomly initialized neural network to generate a simplified trained neural network. Finally, the generated simplified trained neural network may be trained. In some cases, training may be performed using the received dataset (X). This scenario 10d is shown, for example, in FIG. 3d.

[0051] In the present disclosure, in some cases, an incompletely trained neural network (the term "incompletely trained" can be understood in the sense explained above) can be used instead of a randomly initialized neural network. Alternatively or additionally, an incompletely trained neural network can be selected as a given neural network. In some examples, the trained neural networks disclosed in connection with the above methods (see FIGS. 3a-3d) can correspond to the respective incompletely trained neural networks.

[0052] As previously mentioned, the simplified computer-implemented neural networks of the present disclosure can be used in a variety of devices. In general, the present disclosure also relates to methods that include pruning a computer-implemented neural network and then using the computer-implemented neural network in a device. This use can include, for example, controlling (or regulating) the device with the simplified neural network, recognizing an operational state (e.g., a malfunction) of the device or a state of the device's environment with the simplified neural network, or evaluating the operational state of the device or a state of the device's environment with the simplified neural network. In these cases, the input data can include state data (e.g., at least in part, sensor data) related to the internal state of the device. Additionally or alternatively, the input data can include state data (e.g., at least in part, sensor data) related to the device's environment. The output data of the simplified neural network can characterize the operational state or other internal state of the device (e.g., whether an error, an abnormality, or a critical operational state exists). The output data can be used to control the device in response to the characterized operational state or other internal state. Alternatively or additionally, the output data may include control data for the device.

[0053] In some cases, the device may be a vehicle (e.g., an automobile, such as a car that is operated / assisted at least partially autonomously, as described above, or a ship, train, aircraft, or spacecraft). In other cases, the device may be a component of a vehicle (e.g., a vehicle computer). In still other cases, the device may be an electrical appliance (e.g., a tool, a home appliance, or a garden appliance). In yet other examples, the device may be a device in the form of the Internet of Things. Alternatively or additionally, the device may be a battery-operated device (e.g., with a power consumption of less than 5 kW maximum power). As described above, simplified computer-implemented neural networks are advantageous in these environments because they can be constructed in a relatively resource-efficient manner.

[0054] The simplified computer-implemented neural network can be used for classifying time series and / or for classifying image data (i.e., the device is an image classifier). The image data can be, for example, camera image data, lidar image data, radar image data, ultrasound image data, or thermal image data (e.g., generated by a corresponding sensor). The image data can include individual images or video data. In some examples, the computer-implemented neural network can be configured for or usable in a monitoring device (e.g., for manufacturing processes and / or quality assurance) or a medical imaging system (e.g., for retrieving diagnostic data). The image classifier can be configured to receive and classify image data into multiple classes. In some examples, the method includes a one-dimensional (R n ) input vector, and a two-dimensional (R m The image classification may include mapping the input vector components to output data in the form of an output vector. For example, the components of the input vector may represent multiple received image data. Each component of the output vector may represent a result of image classification computed based on a simplified computer-implemented neural network. In some examples, the image classification may include semantic segmentation of the image (e.g., region-wise and / or pixel-wise classification of the image). The image classification may be, for example, object classification. For example, the presence of one or more objects in the image data may be detected (e.g., in a driver assistance system for automatically recognizing traffic signs or lanes).

[0055] In other examples (or additionally), the computer-implemented neural network may be configurable or usable to monitor the operating state and / or environment of the at least partially autonomous robot. In some examples, the at least partially autonomous robot may be an industrial robot. In other examples, the device may be a machine or group of machines (e.g., factory equipment) whose operating state and / or environment is monitored. For example, the operating state of a machine tool may be monitored. In these examples, the input data x may include state data of the at least partially autonomous robot, the machine or group of machines, and / or their environment, and the output data y may include information regarding the operating state and / or environment of the respective device.

[0056] In further examples, the system to be monitored may be a communications network. In some examples, the network may be a telecommunications network (e.g., a 5G network). In these examples, the input data x may include utilization data at nodes of the network, and the output data y may include information regarding resource allocation (e.g., channels, bandwidth of channels, or other resources of the network). In other examples, network malfunctions may be recognized.

[0057] In other examples (or additionally), the computer-implemented neural network may be configurable or usable to control (or regulate) a technical device. The device itself may be one of the devices discussed above (or below) (e.g., an at least partially autonomous robot or machine). In these examples, the input data x may include state data of the technical device and / or its environment, and the output data y may include control variables of the respective technical system.

[0058] In yet other examples (or additionally), the computer-implemented neural network may be configurable or usable to filter the signal (e.g., to remove interfering components and / or noise). In some cases, the signal may be an audio signal or a video signal. In these examples, the output data y may include the filtered signal.

[0059] A second general aspect of the present disclosure relates to a computer program configured to perform the computer-implemented method according to the first general aspect of the present disclosure. The present disclosure also relates to a computer-readable medium (e.g., a machine-readable storage medium such as an optical storage medium or a non-volatile memory, e.g., a flash memory) and a signal storing or encoding the computer program of the present disclosure.

[0060] A third general aspect of the present disclosure relates to a computer-implemented system for generating and / or applying a computer-implemented neural network configured to perform a method according to the first general aspect of the present disclosure. Additionally or alternatively, the computer-implemented system of the third general aspect is configured to execute a computer program according to the second general aspect of the present disclosure. The computer-implemented system may have at least one processor, at least one memory (which may contain a program that, when executed, performs the method of the present disclosure), and at least one interface for input and output. The computer-implemented system may be a "standalone" system or may be a distributed system that communicates via a network (e.g., the Internet).

Claims

1. A method for generating a simplified computer-implemented neural network, The method includes a step (100) of receiving a given neural network (1), the given neural network (1) including a plurality of neural network structures (S) and being described by a plurality of weights (7a, 7b; θ), and a pruning vector (δ p ) being associated with each neural network structure from the plurality of neural network structures, the pruning vector (δ p ) describing a change in the plurality of weights by pruning each of the neural network structures, When the method includes a step (200) of calculating a product (H(θ)δ struc ) of a matrix (H(θ)) and a structure vector (δ struc ), the matrix includes a second partial derivative of a loss function (L D (θ)) with respect to the plurality of weights, and each weight from the plurality of weights belonging to two or more neural network structures to be pruned from the plurality of neural network structures corresponds to a respective component of the structure vector (δ struc ). When the method includes a step (300) of determining two or more changes (λ p ) of the loss function with respect to the given neural network, each change from the two or more changes results from pruning a corresponding neural network structure from the two or more neural network structures to be pruned, and determining the two or more changes of the loss function includes the calculated product, each of the pruning vectors, and the current plurality of weights of the given neural network 【Number 1】 implemented using, the method includes a step (400) of pruning at least one neural network structure from the plurality of neural network structures based on the determined two or more changes of the loss function to generate a simplified neural network (1a). A method for generating a simplified computer-implemented neural network.

2. Each pruning vector (δ) of the corresponding neural network structure to be pruned p includes, as components (7a, 7b), weights from the plurality of weights (θ) belonging to the corresponding network structure to be pruned, Corresponding other components (7c) of the pruning vector corresponding to weights not belonging to the corresponding network structure to be pruned are set to zero, The method according to claim 1.

3. The matrix is a Hessian matrix (H(θ)) that includes the second-order partial derivative function of the loss function (L D (θ)) with respect to the plurality of weights, The product (H(θ)δ struc ) is the Hessian vector product of the Hessian matrix and the structure vector (δ struc ). The method according to claim 1.

4. Determining each change from two or more changes of the loss function resulting from pruning each neural network structure from two or more neural network structures to be pruned is the pruning vector (δ associated with each of the neural network structures p ), and the scalar product of the calculated product (H(θ)δ struc ) 【Number 2】 including calculating each first contribution (310) by calculating, the current plurality of weights of the given neural network are then substituted into the first contribution, The method according to claim 1.

5. Determining two or more changes of the loss function includes, for each of the two or more pruning vectors (δ p ), the current plurality of weights of the given neural network [Number 3] with the gradient (∂L D (θ) / ∂θ) is carried out, The gradient (∂L D (θ) / ∂θ) includes the first derivative of the loss function according to the weights, The method according to claim 4.

6. Determining each change from two or more changes of the loss function that results as a result of pruning each neural network structure from two or more neural network structures to be pruned is the pruning vector (δ p ), which is associated with each of the neural network structures, and the gradient (∂L D (θ) / ∂θ) includes calculating each second contribution by calculating a scalar product therebetween (320). the current plurality of weights of the given neural network are then substituted into the second contribution, optionally, determining each change from two or more changes of the loss function includes calculating each second contribution, The method according to claim 5.

7. determining each change from two or more changes of the loss function includes the following steps, namely, a step (330) of calculating a first product by multiplying each of the first contributions by a first weight coefficient; a step (340) of calculating a second product by multiplying each of the second contributions by a second weight coefficient; a step (350) of summing the absolute value of the first product and the absolute value of the second product, or a step of summing the first product and the second product, to calculate each of the changes; including, optionally, the same first weight coefficient and the same second weight coefficient are used to determine each change from two or more changes of the loss function, The method according to claim 6.

8. pruning based on the determined two or more changes of the loss function further includes the following steps, namely, a step (410) of pruning the neural network structure corresponding to the smallest change from the two or more changes of the loss function Repeatedly pruning two or more corresponding neural network structures from the two or more neural network structures to be pruned (420); further comprising; each subsequent neural network structure to be pruned corresponds to the next largest value among two or more changes, optionally, the repeatedly pruning step is performed until the size of the simplified neural network is below a desired size and / or until the overall performance of the simplified neural network is below a predefined threshold. The method according to claim 1.

9. The method further comprises providing a randomly initialized neural network, optionally, a plurality of weights of the randomly initialized neural network are initialized to random numbers randomly distributed within a given interval. The method according to claim 1.

10. The given neural network is a randomly initialized neural network, the method further comprises generating a simplified randomly initialized neural network from the randomly initialized neural network, and two or more neural network structures of the randomly initialized neural network are pruned based on the determined two or more changes of the loss function with respect to the randomly initialized neural network. The method further comprises training the generated simplified randomly initialized neural network. The method according to claim 9.

11. The method further comprises training the randomly initialized neural network to generate a trained neural network, and the generated trained neural network is the given neural network. The method further comprises generating a simplified trained neural network from the generated trained neural network, and two or more neural network structures of the generated trained neural network are pruned based on the determined two or more changes of the loss function with respect to the generated trained neural network. The method further includes training the generated simplified trained neural network. The method according to claim 9. **Claim 12** The given neural network is an image classifier, The image classifier is configured to receive input data in the form of image data and, optionally, classify the image data into one or more classes based on semantic segmentation of the image data. The method according to claim 1. **Claim 13** A computer program configured to perform all steps of the method according to any one of claims 1 to 12. **Claim 14** A computer-implemented system for generating and / or applying a computer-implemented neural network for an apparatus configured to perform the method according to any one of claims 1 to 12 or configured to execute the computer program according to claim 13.