Methods for compressing neural networks
By implementing neural networks locally among queue participants and selecting components that should be pruned, transmitting them to the central server for summary and pruning, the problem of low neural network compression efficiency in the prior art is solved, and efficient neural network compression and optimization of computing resources are achieved.
Patent Information
- Application Number
- CN202080062080.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-04
- Filing Date
- 2020-08-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-08-04
AI Technical Summary
In the prior art, when compressing convolutional neural networks, it is difficult to effectively select components that should be pruned, resulting in low network compression efficiency and wasted computing resources.
Transfer to the central server for summary and trimming by implementing the neural network locally among the queue participants of the transport vehicle queue and determining the selection of components to be pruned during the inference phase.
The accuracy and efficiency of component selection that should be pruned is improved, and through collective selection and summary, efficient compression of neural networks is achieved, reducing waste of computing resources.
Smart Images

Figure CN114287008B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for compressing a neural network. The invention also relates to a queue participant, a central server and a system. Background Art
[0002] Modern driver assistance systems and driving functions for automated driving increasingly use machine learning, in particular to recognize the vehicle's surroundings, including other traffic participants (e.g. pedestrians and other vehicles), and to describe their behavior. In this case, input data (input) from different sources (e.g. camera, radar, lidar) are evaluated by deep neural networks, which in particular perform pixel-by-pixel classification (semantic segmentation) or generate closed boxes (bounding boxes) of recognized objects.
[0003] In both cases, convolutional networks (CNNs) are usually used, which parameterize the weights of so-called filters based on the input during training. In this case, the convolutional networks used increasingly use a large number of filters and layers, so that the time and computing effort required for processing (inference) from the input data to the output (output) increases. Since the use of neural networks in the field of autonomous driving is severely limited in terms of the required computing time due to the dynamic environment, and at the same time the hardware (computing power) available in the vehicle cannot be expanded arbitrarily, the size of the neural network is a limiting factor in terms of the usability in such systems.
[0004] So-called pruning attempts to reduce the size of a neural network by removing individual elements (i.e. neurons, parameters or entire filters). The choice of neurons or filters to be removed is important here. Different filters can affect the output of the network to different degrees. Therefore, it is suitable to select such filters by the selected strategy that their removal has the least effect on the output (quality) and at the same time prune as many filters as possible in order to achieve a significant reduction in the size of the network and thus achieve the least possible inference time and training time.
[0005] A system and method for pruning a convolutional network (CNN) is known from US2018 / 0336468A1. The method comprises extracting convolutional layers from a trained CNN, wherein each convolutional layer comprises: a kernel matrix having at least one filter constructed in a corresponding output channel of the kernel matrix; a feature map group having a feature map corresponding to each filter. An absolute kernel weight is determined for each kernel, and each filter is summed to determine the strength of each filter. The strength of each filter is compared to a threshold, and if the determined strength is below the threshold, the filter is removed. The feature map corresponding to each removed filter is removed to prune the CNN. The CNN is retrained to produce a pruned CNN with fewer convolutional layers.
[0006] A method, a computer-readable medium and a system for pruning a neural network are known from US2018 / 0114114A1. The method comprises the following steps: receiving a first-order gradient of a cost function relative to a layer parameter for a trained neural network, and calculating a pruning criterion for each layer parameter based on the first-order gradient corresponding to the layer parameter, wherein the pruning criterion indicates the importance of each neuron contained in the trained neural network and assigned to the layer parameter. The method comprises the additional step of identifying at least one neuron with the lowest importance and removing the at least one neuron from the trained neural network to produce a pruned neural network. Summary of the invention
[0007] The invention is based on the object of improving a method for compressing a neural network, in particular with regard to the selection of elements of the neural network that are to be pruned.
[0008] According to the invention, this object is achieved by the method according to the invention, the queue participants according to the invention, the central server according to the invention and the system according to the invention. Advantageous embodiments of the invention result from the invention.
[0009] In a first aspect of the invention, in particular, a method for compressing a neural network is provided, wherein queue participants of a transport queue locally implement the neural network and respectively determine a selection of elements of the neural network to be pruned during at least one inference phase, wherein the queue participants transmit the respectively determined selections to a central server, wherein the central server aggregates the respectively transmitted selections and generates an aggregated selection, and wherein the central server prunes the neural network based on the aggregated selections.
[0010] In a second aspect of the invention, in particular, a queue participant for a queue of transport vehicles is created, which includes a computing device, wherein the computing device is designed in such a way that a neural network is implemented locally and a selection of elements of the neural network to be pruned is determined during at least one reasoning phase, and the determined selection is transmitted to a central server.
[0011] In a third aspect of the invention, in particular, a central server is created, which includes a computing device, wherein the computing device is designed in such a way that the selections of elements of the neural network transmitted respectively by the queue participants are aggregated and an aggregated selection is generated, and the neural network is pruned based on the aggregated selection.
[0012] In a fourth aspect of the invention, in particular a system is created, comprising at least one queue participant according to the second aspect of the invention and a central server according to the third aspect of the invention. In particular, the system implements the method according to the first aspect of the invention.
[0013] The method and the system enable a neural network to be compressed in an efficient manner. This is achieved by the queue participants that are dependent on the transport queue. The queue participants implement the neural network with the aid of a computing device. During at least one inference phase, the elements of the neural network that are to be pruned are determined. The determined elements are each transmitted in the form of a selection to a central server. The central server collects the corresponding selections of the queue participants and generates a compiled selection therefrom. The neural network is then pruned based on the compiled selection, in particular with the aid of the central server.
[0014] One advantage of the present invention is that the selection of the elements to be pruned is performed based on an increased database, since a large number of queue participants select elements of the neural network taking into account different situations. The more queue participants are considered when determining the selection of the elements of the neural network to be pruned, the more situations can be taken into account. This improves the selection of the elements to be pruned.
[0015] The computing devices of the queue participants each have, in particular, a storage device or can access such a storage device. The computing device can be designed as a combination of hardware and software, for example as a program code executed on a microcontroller or microprocessor. The computing device in particular runs the neural network, i.e. the computing device performs the computational operations required for running the neural network on the input data provided, thereby inferring and providing activations or values in the form of output data at the output of the neural network. For this purpose, a local copy of the neural network, i.e. the structure of the neural network and the associated weights and parameters, is stored in a respectively associated storage device.
[0016] The computing device of the central server is designed accordingly. A central copy of the neural network is located in a storage device of the central server. The pruning is performed on the central copy of the neural network.
[0017] The input data are in particular sensor data, in particular a sensor data stream of sensor data detected over time. In particular, the sensor data are detected by a sensor and supplied to an input layer of the neural network, for example via an input interface designed for this purpose. The sensor is in particular a camera, a lidar sensor or a radar sensor. In principle, however, fused sensor data can also be used.
[0018] The train participants are in particular motor vehicles. However, in principle, the train participants can also be other land vehicles, water vehicles, air vehicles or space vehicles.
[0019] The transmission of the selection from the queue participant to the central server is carried out in particular via a communication interface provided for this purpose of the queue participant and the central server. Here, the communication is carried out in particular wirelessly.
[0020] The neural network is in particular a deep neural network, in particular a convolutional network (convolutional neural network, CNN). In particular, it is provided that the neural network performs a perception function. For example, the neural network can be used to classify objects in detected sensor data (e.g. camera data). In addition, regions of the sensor data in which objects can be found (bounding boxes) can also be identified.
[0021] The elements are in particular neurons of a neural network. If the neural network is a convolutional network, the elements are in particular filters of the convolutional network.
[0022] Pruning of a neural network is to be understood to mean, in particular, that the structure of the neural network is changed, in particular pruned or reduced. This is achieved by removing elements and / or parts of elements (e.g. parameters or input channels etc.) from the neural network. Due to the changed structure, in particular the pruned structure, a pruned neural network with less computing power can be used for the input data. The pruned neural network is then compressed in terms of its structure.
[0023] In particular, it can be provided that the method is repeated cyclically. For example, the method can be repeated until a termination criterion is met. For example, the termination criterion can be a functional quality that is lower than that of the (previously) pruned neural network.
[0024] In particular, it can be provided that following the pruning, a retraining of the pruned neural network is subsequently carried out in order to restore the functional quality (performance) of the (pruned) neural network after the pruning.
[0025] During pruning, it is especially provided that the neural network is pruned in a uniform manner. Uniform here should mean that on average all regions of the neural network are pruned to the same extent. This can prevent regions or individual layers of the neural network from being pruned excessively and thus having a negative impact on the function or function quality of the neural network.
[0026] The selection of the elements of the neural network can be carried out in various ways. In a simple embodiment, for example, the elements of the neural network are selected by the queue participants for pruning, i.e., the elements have the smallest influence on the output result of the neural network. Furthermore, it can also be provided that the elements are selected at whose output the activation is always below a predefined threshold.
[0027] For example, the selected elements are collected in the form of a list, table or database and arranged into the selection. In the list, table or database, for example, the univocal identification of the respectively selected element and possibly other characteristics or values, such as the maximum, minimum and / or average activation of the observed element are indicated. In particular, the list, table or database includes the respectively used selection criteria of the respective element or values associated with the respectively used selection criteria. The selection criteria define the conditions for the selection of the element. The list, table or database is transmitted to the central server, in particular in the form of a digital data packet via a communication interface.
[0028] The compilation of the lists / sorts / tables can be performed in various ways. In particular, an average value can be calculated and / or other averaging methods can be used, such as an arithmetic mean, a geometric mean, a moment-based averaging method, a group-based averaging method, a geographical averaging method related to the driving route and / or the participants in the queue and / or a safety-oriented or situation-related averaging method.
[0029] In particular, the aggregating may include determining the elements most frequently selected by the queue participants, wherein the aggregated selections include the most frequently selected elements.
[0030] It can be provided that after pruning the pruned neural network is output, for example in the form of a numerical data set describing the structure and weights or parameters of the neural network.
[0031] It can be provided that only selected queue participants of a queue of transport vehicles carry out the method. The other queue participants can only use the neural network and receive the pruned neural network from a central server, for example after pruning the neural network. The selection of the elements to be pruned is not carried out or transmitted by the other queue participants. This has the advantage that the selected queue participants can be better equipped technically to carry out the method, for example with regard to the storage capacity of the sensor device and / or the storage device.
[0032] In one embodiment, it is provided that the selection is transmitted to the central server if at least one transmission criterion is met. In this way, unnecessary communication between the queue participants and the central server can be avoided. For example, the transmission criterion can be a certain number of locally collected elements of the neural network. The transmission criterion can also require that the selection no longer changes after a number of passes or inference phases.
[0033] In one embodiment, it is provided that the queue participants each create a ranking of the selected elements and transmit the selection in the form of the created ranking to a central server, wherein the central server creates a summarized ranking based on the transmitted rankings for aggregation, and wherein the neural network is pruned based on the summarized ranking. As a result, those elements can be pruned in a targeted manner which achieve the highest value according to the criterion according to which the ranking was created.
[0034] In one embodiment, it is provided that input data that vary over time are supplied to the neural network for determining the corresponding selection, wherein the time activation difference of the elements of the neural network is determined for the input data that are adjacent in time, and wherein the selection of the elements of the neural network is carried out according to the determined time activation difference. This achieves: compressing the neural network while taking into account the stability-oriented criteria. This is achieved in the following way: the neural network is supplied with input data that vary over time. Since the input data vary over time, the activation or value at the corresponding output of each element of the neural network also changes. The time change of the activation or value at the output is then mapped by the time activation difference of the elements of the neural network. This embodiment is based on the following consideration: because the input data provided, especially based on the detected sensor data, usually only changes slightly when the time changes slightly, the activation difference determined at the output of the element for this time change should also only change slightly. Therefore, a large activation difference indicates an unstable element in the neural network. Unstable elements can be identified by determining the time activation difference. Once the unstable element is identified, the unstable element can be removed from the structure of the neural network in a centralized server later by pruning.
[0035] An activation difference is in particular a difference which is determined by an activation or a value of an output which is inferred or calculated by an element of the neural network at adjacent, in particular successive, points in time. If the input data are camera data, for example, the input data can correspond to two temporally successive individual images of the camera data. Thus, temporal fluctuations of the activation at the output of an element of the neural network in relation to the temporally varying input data are mapped by means of the activation difference.
[0036] The input data are in particular sensor data, in particular a sensor data stream of sensor data detected over time. In particular, the sensor data are detected by a sensor and supplied to an input layer of the neural network, for example via an input interface designed for this purpose.
[0037] If the neural network is a convolutional network, the activation differences are determined for the filters of the convolutional network respectively.
[0038] If the method is repeated cyclically, the repetition of the method can be terminated, for example, when the activation difference falls below a predefined threshold value, ie, when a predefined degree of stability has been reached.
[0039] In particular, it is provided that an average value is calculated element by element from the determined activation differences, wherein pruning is performed as a function of the average value determined. As a result, peaks that occur briefly in the activation differences can be taken into account or weakened. Such peaks alone do not then lead to the relevant element of the neural network being identified as unstable. The relevant element is only selected for pruning if the average value determined for the element from a plurality of activation differences exceeds a threshold value, for example. The average value can be determined as an arithmetic mean, a time mean or a geometric mean. The averaging can be performed both by the queue participants and by a central server.
[0040] The aggregation of the selections in the central server (which may also be referred to as aggregation) may in particular include the calculation of the arithmetic mean, the geometric mean method (for example, the center of gravity of the vector formed by the individual sortings may be formed), the moment-based averaging method, the group-based averaging method, the geographic averaging method and / or the safety-focused averaging method.
[0041] The group-based averaging method can be designed in particular as a local filter, i.e. subsets are selected from the selections of the individual queue participants and the averaging is performed only within these subsets. The subsets can be formed, for example, based on geographical location or region, travel route, time point (clock time, time of day, day of the week, month, season, etc.) or other characteristics. The subsets are always formed based on the commonality of the queue participants who provided the selection or the circumstances in which the selection was created.
[0042] The safety-focused averaging method takes into account in particular the situation in which the elements of the neural network are selected for pruning. This situation can then have an influence on the determined ranking. For example, a safety-critical situation (e.g. when children are around the motor vehicle) can lead to the selected elements being moved further up in the ranking, i.e. in the direction of a higher order. In particular, if the elements are selected using temporal activation differences, the "penalty" of the elements or filters can thereby be increased (or decreased), so that the situation-dependent environment has an influence on the ranking.
[0043] In one specific embodiment, it is provided that at least a portion of the determined activation differences is determined as a function of at least one influencing parameter. This makes it possible, for example, for the activation differences to be influenced over time or depending on the situation. For example, the speed (provided, for example, by a GPS sensor), the current weather (provided, for example, by a rain sensor) and / or the steering angle (provided, for example, by a steering angle sensor) can be used to strengthen or weaken the determined time activation differences depending on the situation. In particular, situation-dependent sensor properties can be taken into account.
[0044] In one embodiment, it is provided that a ranking of the elements is created based on the determined time activation differences, wherein pruning is performed according to the created ranking. In particular, the ranking is created based on the corresponding average values of the determined activation differences. Based on the created ranking, it can be provided, for example, that a preset number of rankings are taken into account during pruning in the central server later, for example 10, 100 or 1000 elements of a neural network (e.g. 10, 100 or 1000 filters of a convolutional network) having the maximum (average) activation difference. The creation of the ranking makes it possible to compress the neural network and to selectively select and prune the most unstable elements. The rankings created by the queue participants and transmitted to the central server as corresponding selections are aggregated by the central server into an aggregated ranking.
[0045] In an improved embodiment, it is provided in particular that the ranking is determined in such a way that the elements of the neural network with the largest time activation differences are pruned, thereby removing the most unstable elements of the neural network from the neural network.
[0046] A ranking function for determining the (in)stability of an element of a neural network can be defined in a simple case by the temporal activation differences between the activations of the elements with respect to (temporal) changes in the input data (eg changes in temporally adjacent individual images of a video).
[0047] If the input data is a single video image, such as a camera image of a surrounding camera, a structural similarity index (Structural Similarity Index, SSIM) between single video images at different time points can be used in the ranking function to determine the difference between temporally adjacent single video images.
[0048] For convolutional neural networks (CNN), in an improvement of the ranking function for the observed filters (also called filter kernels), the temporal activation differences in the previous convolutional layers of the neural network are also taken into account in the convolutional layers of the CNN. As a result, the influence of the previous convolutional layers can be taken into account or removed when determining or calculating the temporal activation differences. The idea behind this is that the activation differences can propagate through the neural network, because the temporal activation differences of the convolutional layers are passed to the subsequent convolutional layers. By taking into account the temporal activation differences of the respective previous convolutional layers, the activation differences calculated for the individual filters can be compared over the entire neural network or across layers. As a result, the ranking of the elements of the neural network can be better determined.
[0049] Improvedly, it can be arranged that averaging is performed over multiple time steps. In addition, it can be arranged that, for example, averaging is performed over multiple sets of input data, for example, averaging is performed over multiple video sequences each consisting of a single video image.
[0050] In an improved embodiment, it is provided that the aggregated ranking is created taking into account at least one target variable. Thus, an extended ranking can be generated. In particular, in addition to the stability of the neural network, further target settings can also be taken into account by taking into account the at least one target variable. For example, these target settings can relate to the robustness of the neural network. For example, if the neural network is used to recognize objects in camera data, robustness with respect to brightness changes can be sought as a target variable. Here, when selecting the time difference by the input data, it is taken into account that the time difference only appears as a difference in the at least one target variable. Then the ranking is created in a similar manner. In addition, filters with a greater influence in the neural network can still be retained despite large activation differences. Filters with the same function as other filters (for example, filters that filter camera data by convolution with a filter function) can also be moved to a higher order in the ranking, i.e., become preferred when pruning or removing. Filters that are not similar to other filters can be moved to a lower order in the ranking, i.e., pruned less preferentially. Another target variable can be majority-oriented, i.e., responsible for the recognition or filtering out of many different features in the input data. Accordingly, the determined ranking is adapted in such a way that a large number of different features or a minimum number of different features can still be detected or filtered out. A further target variable can also be the efficiency (also called performance) of the neural network. In short, the aggregated ranking can be determined and / or adapted in such a way that all predefined target variables are taken into account when selecting the elements or filters of the neural network that are to be pruned.
[0051] In one embodiment, it is provided that the determination of the selection, in particular the determination of the selection by activation differences, and the pruning are restricted to selected layers of the neural network. Thus, for example, layers used for feature extraction may not be involved in pruning. In particular, for neural networks that perform different tasks (multi-task learning) or whose task performance is broken down into different sub-steps, pruning can thus be focused on specific sub-tasks. An example of this is a region proposal network that first identifies relevant image parts for object recognition and then classifies and evaluates the image parts. In this case, pruning can be focused on classification in a target-oriented manner to prevent the neglect of relevant image regions due to pruning triggers.
[0052] In one embodiment, it is provided that the neural network is retrained after pruning. This can improve the functional quality of the pruned or compressed neural network. In this case, it is not necessary to carry out a complete training. Instead, it can be provided that the pruned or compressed neural network is retrained using only a portion of the training data originally used for training. The retraining is carried out in particular with the aid of a central server.
[0053] In one embodiment, it is provided that the elements used for pruning are at least temporarily deactivated. For this purpose, for example, the parameters of the elements to be pruned are set to zero, so that the elements in the neural network no longer have an influence on the results in the subsequent layers or output layers of the neural network. This has the advantage that the deactivation can be canceled again more easily than if an element of the neural network was removed. For example, if after deactivating an element it is found that the functional quality of the neural network is too strongly impeded, the element can be activated again. In particular, the reactivated element can then be marked so that it is not deactivated and / or pruned again in the subsequent through-running of the method. In other cases, the deactivated elements can be removed in a later step, i.e. the structure of the neural network is adapted. This is achieved in particular when the functional quality or other target parameters are achieved despite the deactivation of the element.
[0054] In one embodiment, it is provided that the pruning is only performed when at least one trigger criterion is met. As a result, pruning can always be performed when a certain state is reached. In particular, a permanent or too frequent pruning and / or a too drastically changing selection of elements of the neural network that are to be pruned can be prevented. For example, the trigger criterion can be a preset number of elements in the sequence. Furthermore, the trigger criterion can also additionally or alternatively be the convergence of the sequence. Convergence of the sequence means that the elements in the determined sequence no longer change at least for a preset number of sequences over a preset number of passes of the method or a preset time period. In particular, the presence of the trigger criterion is checked with the aid of a central server.
[0055] In one specific embodiment, it is provided that the pruned neural network is then transmitted to at least one queue participant. This is done, for example, in the form of a digital data packet, which is transmitted via a communication interface and includes the structure and weights and parameters of the pruned neural network. The queue participant receives the pruned neural network and can then replace the neural network stored in the corresponding storage device with the received pruned neural network. The method can then be re-implemented on the pruned neural network.
[0056] In one embodiment, it is provided that the determination of the ranking is designed as an iterative process. In this case, the transmission and compilation (or aggregation step) on the central server is followed by a redistribution of the ranking (for example as a table) to the queue participants of the transport means queue. The queue participants then continue the ranking or update the ranking until a relevant deviation from the last assigned ranking occurs and the next aggregation step is triggered. The process ends when the ranking has not changed or has only changed slightly.
[0057] Further features for the configuration of the queue participants and the central server and the system are derived from the description of the configuration of the method. The advantages of the queue participants and the central server and the system are respectively the same as in the configuration of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The present invention will be explained in more detail below with reference to the accompanying drawings using preferred embodiments.
[0059] Figure 1 A schematic diagram showing one embodiment of the system;
[0060] Figure 2 A schematic flow chart shows an embodiment of a method for compressing a neural network. DETAILED DESCRIPTION
[0061] exist Figure 1 A schematic diagram of one embodiment of a system 1 is shown in . The system 1 comprises a plurality of queue participants 2 in the form of motor vehicles 50 and a central server 20 .
[0062] Each queue participant 2 comprises a computing device 3, a storage device 4 and a communication interface 5. Each queue participant 2 also comprises a sensor device 6, for example in the form of a camera 51, which provides camera data as input data 7 via an input interface 8, and the input data is supplied to the computing device 3 from the input interface 8. A neural network 9 is respectively stored in the storage device 4, i.e. the structure, weights and parameters of the neural network 9 are described univocally. The computing device 3 can carry out computing operations in the storage device 4 and in particular run the neural network 9.
[0063] The central server 20 includes a computing device 21 , a storage device 22 , and a communication interface 23 .
[0064] The computing device 3 of the queue participant 2 is designed such that the neural network 9 stored in the memory device 4 is locally executed on the input data 7 and a selection 10 of the elements of the neural network 9 to be pruned is determined in each case during at least one inference. If a selection 10 is made, the computing device 3 transmits the determined selection 10 to the central server 20 via the communication interface 5.
[0065] The computing device 21 of the central server 20 is designed in such a way that it receives the selections 10 of elements of the neural network 9 transmitted by the queue participants 2 in each case via the communication interfaces 5, 23, combines them, and generates a combined selection 11. The computing device 21 then prunes the neural network 9 based on the combined selection 11, thereby generating a pruned neural network 12.
[0066] It can be provided that the pruned neural network 12 is then transmitted to the queue participant 2 via the communication interface 5 , 23 . The computing device 3 of the queue participant 2 can then replace the neural network 9 in the storage device 4 with the pruned neural network 12 .
[0067] It can be provided that the selection 10 is transmitted to the central server 20 if at least one transmission criterion 13 is met. For example, the transmission criterion 13 can be a certain number of locally collected or selected elements of the neural network 9. The transmission criterion 13 can also be that the selection does not change after a number of local passes or inference phases.
[0068] It can be provided that the queue participants 2 each create a ranking 14 of the selected elements and the selection 10 is transmitted to the central server 20 in the form of the created ranking 14, wherein the central server 20 creates a summarized ranking 15 for summarization based on the transmitted ranking 14. The neural network 9 is then pruned based on the summarized ranking 15.
[0069] It can be provided that the neural network 9 is respectively supplied with input data 7 which vary over time in order to determine a corresponding selection 10, wherein time activation differences of elements of the neural network 9 are determined for temporally adjacent input data 7, and wherein the selection of the elements of the neural network 9 is carried out as a function of the determined time activation differences. In particular, it can be provided that for the selection, a ranking 14 of the elements of the neural network 9 is created based on the time activation differences determined for the elements in each case. The selection 10 then includes the ranking 14 created in this way.
[0070] It can be provided that the neural network 9 is retrained in the central server 2 following the pruning.
[0071] It can be provided that the elements used for pruning are at least temporarily deactivated. For this purpose, for example, the parameters of the elements are set to zero, so that the elements in the neural network 9 no longer have any influence on the results in the subsequent layers or output layers of the neural network 9. If, after deactivating an element, it is found, for example, that the functional quality of the neural network 9 is too strongly impeded, the element can be activated again. In particular, the reactivated element can then be marked so that it is not deactivated and / or pruned again in a subsequent pass-through operation of the method. In a later step, in particular after retraining the neural network 9, the deactivated elements can be removed, i.e. the structure of the neural network 9 is adapted, so that a pruned neural network 12 is generated. This occurs in particular when the functional quality or other target parameters are achieved despite the deactivation of the element.
[0072] Furthermore, it can be provided that pruning is only carried out when at least one trigger criterion 30 is met. For example, the trigger criterion 30 can be the reaching of a predefined number of elements in the aggregated order 15. Furthermore, the trigger criterion 30 can also be, in addition or as an alternative, the convergence of the aggregated order 15. The convergence of the aggregated order 15 means that the elements in the aggregated order 15 no longer change over a predefined number of runs of the method or over a predefined period of time at least for a predefined number of sequences. In particular, the presence of the trigger criterion 30 is checked by means of the central server 20.
[0073] exist Figure 2 A schematic flow chart for illustrating one specific embodiment of a method for compressing a neural network 9 is shown in FIG. The method is implemented with the aid of a system 1, which is, for example, Figure 1 . In this case, one part of the method is implemented in each queue participant 2 and another part is implemented in the central server 20. The flowchart is explained using input data 7 as an example, which is provided in the form of a video 40, which consists of individual video images 41 (frames). The individual video images 41 are respectively associated with the time points t i Corresponding.
[0074] In method step 100, the video individual images 41 are fed to the neural network 9. This is shown for two adjacent time steps, namely for the time point t i The corresponding video single image 41 and the subsequent time point t i+1The corresponding single video image 41 is shown. The neural network 9 is applied to the single video image 41 and a result is inferred at the output of the output layer of the neural network 9. The result can, for example, include object recognition or object classification and / or the creation of a bounding box for the recognized object. During the inference, the value of the activation 43 for the element of the neural network 9 is detected or read out. If the neural network 9 is designed as a convolutional network, the activation 43 corresponds to the corresponding value at the output of the filter of the convolutional network. For example, the results are respectively provided as a list in which, for each element of the neural network 9, a list of values for the observed time point t is stored. i and t i+1 Related activation43.
[0075] In method step 101 , the values of activation 43 for the individual elements are used to determine the values for the two observed points in time t i and t i+1 This takes place element by element for all elements of the neural network 9. In this case, in particular, the differences between the values of the activation 43 of the individual elements are taken into account. For each element of the neural network 9, then, with respect to two time points t i and t i+1 The time activation differences 44 are available. For example, the result is provided as a list in which the time difference t i and t i+1 The time activation difference is 44.
[0076] It may be provided that the temporal activation differences 44 are averaged and that the subsequent method steps 102 - 106 are carried out based on the averaged activation differences 44 .
[0077] It can be provided that at least a part of the determined time activation difference 44 is determined as a function of at least one influencing parameter 45. For example, the speed (provided, for example, by a GPS sensor), the current weather (provided, for example, by a rain sensor) and / or the steering angle (provided, for example, by a steering angle sensor) can be used to strengthen or weaken the determined time activation difference 44 depending on the situation.
[0078] In method step 102 , the determined temporal activation differences 44 are sorted according to magnitude, resulting in a ranking 14 in which the elements of neural network 9 having the largest temporal activation differences 44 are ranked high.
[0079] In method step 103 , the created ranking 14 of the queue participants 2 is transmitted to the central server 20 via a communication interface.
[0080] In method step 104, a summarized ranking 15 is generated in the central server 20 based on the transmitted rankings 14. The summarization (which can also be referred to as aggregation) can include, in particular, the calculation of the arithmetic mean, the geometric mean method (for example, the center of gravity of the vector formed by the individual rankings 14 can be formed), the moment-based averaging method, the group-based averaging method, the geographical averaging method and / or the safety-focused averaging method.
[0081] It can be provided that the aggregated ranking 15 is created taking into account at least one target variable 46. For example, the at least one target variable 46 can relate to the robustness of the neural network 9. In addition, elements or filters that have a greater influence within the neural network 9 can be retained despite large time activation differences 44. Elements or filters that have the same function as other elements or filters (for example, filters that filter camera data by convolution with a filter function) can also be moved forward in the aggregated ranking 15, i.e. become preferred when pruning or removing. Elements or filters that are not similar to other elements or filters can be moved backward in the aggregated ranking 15, i.e. be deleted with less priority. A further target variable 46 can be majority-oriented, i.e. be responsible for the recognition or filtering out of many different features in the input data 7. Accordingly, the aggregated ranking 15 is adapted in such a way that many different features or a minimum number of different features are still recognized or filtered out. A further target variable 46 can also be the efficiency (performance) of the neural network 9.
[0082] It can be provided that the elements for pruning are at least temporarily deactivated. In a subsequent method step, for example after successful retraining, the neural network 9 can be pruned.
[0083] In method step 105 , neural network 9 is pruned by removing the elements of neural network 9 having the largest temporal activation differences 44 from the structure of neural network 9 according to aggregated ranking 15 . As a result, pruned neural network 12 is provided.
[0084] It can be provided that in method step 200 a preliminary check is carried out to determine whether at least one triggering criterion 30 is met. Only if triggering criterion 30 is met is the pruning of neural network 9 carried out.
[0085] In method step 106 , in order to improve the functional quality of pruned neural network 12 , provision may be made to retrain pruned neural network 12 .
[0086] The described method steps 100-106 are for a further point in time t i+x In particular, it is provided that the trimming is performed in method step 105 based on the aggregated ranking 15, which is created for the average value of the time activation differences 44. In this case, in particular at a plurality of time points ti Find the average value above.
[0087] In particular, it is provided that method steps 100-106 are repeated cyclically, wherein the current input data 7 are respectively used. In particular, it is provided that method steps 100-106 are repeated using a (retrained) pruned neural network 12, wherein the neural network 9 is replaced by a corresponding (retrained) pruned neural network 12 for this purpose.
[0088] The described embodiment of the method makes it possible to compress neural network 9 , wherein at the same time the stability of neural network 9 is increased since unstable elements of neural network 9 are removed or deactivated.
[0089] The determination of the sorting based on the time activation difference is explained below with the help of a mathematical example. Here, it is assumed that the input data consists of a video sequence, which consists of individual video images. Depending on the number of individual video images per time unit (e.g., frames per second), only slight changes occur in adjacent individual video images. The method takes advantage of this and uses it for stability-based pruning of neural networks. The method is particularly useful for well-trained neural networks, particularly convolutional neural networks (CNNs). The filters of CNNs are particularly considered to be elements of neural networks. Here, the following filters (also referred to as filter kernels) are considered to be unstable, the activation of which shows large changes in activation in adjacent individual video images, i.e., in input data that changes over time (i.e., where the time activation difference is large). In an embodiment of the method, such filters are evaluated higher in the sorting. As input data, only unlabeled (without labels) input data is required, such as individual video images of a video sequence, which is detected by a camera for detecting the surroundings of a motor vehicle.
[0090] A (convolutional) layer comprises in particular a plurality of filters (also called filter kernels), wherein each filter in particular receives the entire output of the preceding layer and each filter provides an associated feature map (English: feature map) as output.
[0091] For a single image of a video in a sequential dataset (video sequence) with height H, width W, channels C, and time point t, where x t ∈G H×W×C Defined as a single image of a video of dataset χ, where Video single image x t It is a neural network The input of (i.e., corresponding to the input data at the input end), where θ is the parameter of the neural network. Include (convolutional) layer, whose output is With height H l , width W l and a certain number of feature maps k for l∈{1,...,L} l .
[0092] The j-th feature map of the output of layer l is denoted by Ψ l,j , i.e., the feature map associated with filter j of layer l, where j∈{1,...,k i}. It can be defined as the group of all feature maps in the neural network at time point t.
[0093] Stability is defined in particular as t The change of output Ψ l,j The change in (activation difference), that is, for filter j in layer l and time point t, produces a ranking function ranking rank:
[0094]
[0095] in,
[0096]
[0097] In short, for the filter under consideration, the greater the resulting value of the ranking function, the greater the instability. The ranking function is determined in particular for each element of the neural network, ie in the present example for each filter.
[0098] Output l,j (x t ), i.e., the activation of the filter, may have two causes. First, the change at the output may be due to the single video image x at the input. t Secondly, the change may also be caused by the change in the previous layer Ψ l-1 (x t ) is caused by a change in activation at the output terminal.
[0099] To calculate the difference between successive individual video images (ie temporally adjacent input data), for example, a structural similarity index known per se (Structural Similarity Index, SSIM) can be used. SSIM is used to measure the similarity between two images.
[0100] Δx t =1-SSIM(x t , x t-1 ) (2)
[0101] Because l,j The stability of t-1The stability of the output (activation) of , and thus the filters in layer l should be prevented from being pruned due to instability in layer l-1. Therefore, the contribution of the previous layer is taken into account in the ranking function. is the normalized output of layer l at time t to the subsequent layer. This is calculated in the following formula. In order to be able to compare the changes in the output (activation) of the filters throughout the neural network, the output (activation) is multiplied by the height H of layer l. l , Width W l and the number of channels (feature maps, English: feature map) k l To standardize.
[0102] To calculate the stability-based rank for each filter, equations (2) and (3) are combined with equation (1). The rank is defined here as in is a video sequence. Formula (4) defines the order as the (temporal) activation difference averaged over all time points t (i.e. over the number T of individual images of the video) relative to the change at the filter input (i.e. relative to the (temporal) change of the input data):
[0103]
[0104] Here, λ is a weighting factor, by means of which the influence of the previous layer can be set. In this case, the influence of the previous layer is determined by means of the size H of the respectively considered layer l. l ×W l To add weight.
[0105] This determines the average value, that is, the order Aggregate and take the (arithmetic) average over multiple unlabeled (without labels) video sequences A, as shown in formula (5).
[0106] R l,j A larger value of means that the observed filter (the jth filter in the lth layer) has a greater instability. It is also possible to use other methods for averaging, such as a moment-based averaging method or an averaging method in which the individual activation differences are taken into account in a weighted manner.
[0107] The ranking created in this way is transmitted to a central server where it is aggregated with the rankings transmitted by other queue participants.
[0108] The neural network is then pruned based on the summarized ranking, for example, filters in the higher (5, 10, 20, ..., etc.) orders of the summarized ranking are removed from the neural network because these filters are the most unstable filters or elements of the neural network.
[0109] The mathematical examples involve video sequences or individual video images as input data. However, in principle, the approach is similar for other types of input data.
[0110] Reference numerals list
[0111] 1 System
[0112] 2 Queue Participants
[0113] 3 Computing devices (queue participants)
[0114] 4 Storage Devices (Queue Participants)
[0115] 5 Communication interface (queue participants)
[0116] 6 Sensor device
[0117] 7 Input Data
[0118] 8 Input Interface
[0119] 9 Neural Networks
[0120] 10 Choice
[0121] 11 Aggregated selection
[0122] 12 Pruned Neural Network
[0123] 13 Transmission Standards
[0124] 14 Sorting
[0125] 15 Summarized ranking
[0126] 20 Central Server
[0127] 21 Computing device (central server)
[0128] 22 Storage device (central server)
[0129] 23 Communication interface (central server)
[0130] 30 Trigger criteria
[0131] 40 Videos
[0132] 41 Video single image
[0133] 43 Activation
[0134] 44 Time activation difference
[0135] 45 Influencing parameters
[0136] 46 Target Parameters
[0137] 50 Motor Vehicles
[0138] 100-106 Methods and Steps
[0139] 200 Methods and Steps
[0140] t i Time
[0141] t i+1 Time
[0142] t i+x Further time point.
Claims
1. Method for compressing neural networks (9), in, The train participants (2) of the transport train locally implement the neural network (9) and in each case determine during at least one inference phase a selection (10) of elements of the neural network (9) that are to be pruned, Therein, the queue participants (2) transmit the respectively determined selections (10) to a central server (20), wherein the central server (20) aggregates the separately transmitted selections (10) and generates an aggregated selection (11), wherein the central server (20) prunes the neural network (9) based on the aggregated selections (11), Characterized in that the pruned neural network (9) is subsequently transmitted to at least one queue participant (2), wherein the at least one queue participant (2) receives the pruned neural network (9) from the central server (20) and replaces the neural network (9) stored locally in the storage device with the received pruned neural network (9).
2. The method according to claim 1, characterized in that If at least one transmission criterion (13) is met, the selection (10) is transmitted to the central server (20).
3. The method according to claim 1 or 2, characterized in that: The queue participants (2) each create a ranking (14) of the selected elements and transmit the selection (10) in the form of the created ranking (14) to the central server (20), wherein the central server (20) creates a summarized ranking (15) based on the transmitted ranking (14) for aggregation, and wherein the neural network (9) is pruned based on the summarized ranking (15).
4. The method according to claim 1 or 2, characterized in that: Input data (7) that vary over time are respectively supplied to the neural network (9) for determining the corresponding selection (10), wherein time activation differences (44) of elements of the neural network (9) are determined for temporally adjacent input data (7), and wherein the selection of elements of the neural network (9) is carried out as a function of the determined time activation differences (44).
5. The method according to claim 1 or 2, characterized in that: The neural network (9) is retrained following the pruning.
6. The method according to claim 1 or 2, characterized in that: Deactivate the component used for trimming at least temporarily.
7. The method according to claim 1 or 2, characterized in that: The pruning is performed only when at least one triggering criterion (30) is met.
8. Queue participants (2) for a transport queue, including: Computing device (3), The method is characterized in that the computing device (3) is designed to locally execute the neural network (9) and to determine a selection (10) of elements of the neural network (9) that are to be pruned during at least one inference phase, and transmitting the determined selection (10) to a central server (20); In addition, a pruned neural network (9) is received from the central server (20), and the neural network (9) stored locally in the storage device is replaced by the received pruned neural network (9).
9. Central server (20), comprising: Computing device (3), The computing device (3) is characterized in that the computing device (3) is designed to aggregate the selections (10) of elements of the neural network (9) respectively transmitted by the queue participants (2) according to claim 8 and generate an aggregated selection (11), and to prune the neural network (9) based on the aggregated selection (11), and to subsequently transmit the pruned neural network (9) to at least one queue participant (2).
10. A system for pruning a neural network (1), characterized in that It includes: At least one queue participant (2) according to claim 8 and a central server (20) according to claim 9.
Citation Information
Patent Citations
Systems and methods for pruning neural networks for resource efficient inference
US20180114114A1
Pruning filters for efficient convolutional neural networks for image recognition in surveillance applications
US20180336468A1
Local learning system in artificial intelligence device
US20190114543A1