Method and system for training neural network for multi-core data processing system
Through improved training methods and systems, the execution of neural networks on multi-core processor systems is optimized, and the problems of resource and routing constraints are solved, achieving more efficient computing performance and resource utilization.
Patent Information
- Application Number
- CN202380078773.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-14
- Filing Date
- 2023-11-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to optimize the execution of neural networks on multi-core data processing systems while taking into account specific resources and routing constraints, resulting in uneven allocation of computing load and storage requirements, affecting system performance.
An improved method and system is provided for training a neural network to perform data processing tasks on a multi-core processor system. The method includes providing a neural network architecture, obtaining training data, updating neural network weights through backpropagation, and mapping adjustments based on computational performance losses to optimize the execution of neural networks on multi-core processor systems.
By optimizing the mapping and weight update of neural networks, computational performance losses, such as delay and power consumption, improve system resource utilization and execution efficiency.
Smart Images

Figure CN120226018A_ABST
Abstract
Description
Background Art
[0001] The present disclosure relates to a method for training a neural network for a multi-core data processing system.
[0002] The present disclosure also relates to a system for training a neural network for a multi-core data processing system.
[0003] In the past few decades, increasingly powerful neural networks have been developed and applied to an increasing number of applications, such as image processing, data analysis, device control, etc. A neural network includes a plurality of neural network elements, and the plurality of neural network elements are interconnected with other neural network elements through interconnections with corresponding weights. Training data pairs are used to determine the weights during the training phase. Each training data pair includes corresponding input data to be provided at the input of the neural network to be trained and corresponding true output data for comparison with the output data calculated by the neural network. A loss indicating the difference between the output data calculated by the neural network and the true output data is calculated. Although the neural network elements of a neural network can potentially be linked in any possible way, neural networks are typically organized in neural network layers. Although a neural network can in principle be executed by a single data processor, for practical purposes this typically involves too low a throughput and too high a latency. Therefore, it is necessary to map a neural network onto a multi-core processor. This means that the computational load and storage requirements involved in neural network operations are distributed in a distributed manner to the processor cores of the multi-core processor. As an example, in a neural network with multiple layers, a corresponding number of processor cores with corresponding processing capabilities and corresponding storage spaces can be assigned to each of the layers, such that the corresponding processor cores perform all the calculations for the neural elements in the layer, and the storage space is used to store the feature maps of the layer. In practice, the mapping that can optimally utilize the resources of the multi-core processor strongly depends on the number of available cores and the nature of the layers.
[0004] In theory, balanced load can be achieved by allocating more computing power to computationally intensive layers and less to other layers. In this regard, it should be noted that Shen discussed the problem of optimizing global resource allocation in "HALP: hardware-aware latency pruning" arXiv:2110.10811v1 [cs.CV] 20 Oct 2021, aiming to maximize accuracy while limiting latency under a predefined budget. The methods disclosed therein include computing a lookup table (once) for estimating the latency of each layer based on the static properties of the layer, namely the number of input channels and output channels. However, in practice, the possibilities of resource allocation are restricted by specific resource and routing constraints. That is, the mapping of a neural network also involves the allocation of communication channels, such as using a network-on-chip (NoC) in a multi-core processor. In particular, in a sparse neural network, the computational load distribution strongly depends on the weights assigned to the neural network interconnect during the training of the neural network for a specific purpose. Summary of the Invention
[0005] A first object of the present disclosure is to provide an improved method for training a neural network, which aims to optimize the execution of the neural network on a multi-core data processing system while considering specific resource and routing constraints.
[0006] A second object of the present disclosure is to provide an improved training system for training a neural network, which aims to optimize the execution of the neural network on a multi-core data processing system while considering specific resource and routing constraints.
[0007] According to the first object, an improved method is provided herein for training a neural network to perform a data processing task when executed by a processor system including a plurality of processor cores. The specific multi-core processor system for which the neural network is trained to perform the data processing task is also denoted as the target processor system. In one example, the target processor system is incorporated in a mobile phone. In another example, the target processor system is provided by a GPU on a laptop computer.
[0008] An improved method for training includes providing a neural network architecture for a neural network. The neural network architecture specifies the components of the network, such as convolutional layers, pooling layers, and their interrelationships. The network architecture can be a known network architecture, such as a version of ResNet, MobileNet, or VGG, or others. The network provided herein can be pre-trained or not trained at all. In any case, it is assumed that it is not yet (optimally) suitable for performing data processing tasks on the target processor system. Thus, the improved method includes obtaining training data, which will be used to train the neural network to perform data processing tasks. The training data includes input data and ground truth output data.
[0009] The improved method provides a mapping of the neural network on a multi-core processor system. This mapping in particular determines the assignment of the cores of the multi-core processor system to the layers of the neural network. This mapping can be provided by conventional mapping methods. An example of this is described by Dai et al. in "ChamNet: Towards Efficient Network Design through Platform-Aware Model Adaptation", arXiv:1812.08934v1 [cs.CV] 21 Dec 2018.
[0010] Then the following sequence of training steps is repeated. The repetition can be a predetermined number of times or can stop when a convergence criterion is detected.
[0011] In the first training step in the sequence, a sample of the input data of the training data is provided to the training processor system that executes the neural network being trained. In this regard, it should be noted that the training processor system is not necessarily the target processor. Generally, the training processor system has higher processing power than the target processor system.
[0012] In the next training step in the sequence, in response to the input data, the training processor system executes the neural network being trained to generate output data.
[0013] The ground truth data corresponding to the sample of the input data is provided, and a task loss indicating the deviation between the generated output data and the ground truth output data is calculated. The task loss is, for example, the cross-entropy loss for classification, the L1 norm for regression, or a combination thereof.
[0014] In addition, a computational performance loss is calculated, which indicates the cost of executing a neural network if the neural network is executed by a target processor system according to the provided mapping. The computational performance loss is, for example, a measure of latency (e.g., per layer or total latency), throughput, required supply power, or a combination thereof. The computational performance loss can, for example, be proportional to the latency, but alternatively, it can be proportional to the first difference between the latency and the minimum value of the latency that can be theoretically achieved. As another example, the computational performance loss can be inversely proportional to the throughput, but alternatively, it can be proportional to the second difference between the inverse of the throughput and the minimum value of this metric that can be theoretically achieved. As another example, the computational performance loss can be proportional to the power consumption, but alternatively, it can be proportional to the third difference between the power consumption and the minimum value of the power consumption that can be theoretically achieved. In yet another example, the computational performance loss is a weighted sum of two or more of the first difference, the second difference, and the third difference. In particular, the total latency and the total throughput are very important computational performance metrics.
[0015] Then, a total loss that is positively correlated with the task loss and the computational performance loss is calculated. As an example, the total loss is calculated as a weighted sum of the task loss and the computational performance loss. However, other implementations of the loss weighting unit LWU are possible as long as the operations it performs are differentiable.
[0016] In subsequent operations of the sequence, the weights of the neural network being trained are updated by backpropagation of the total loss.
[0017] Due to the fact that the total loss is partially determined by the latency loss, the backpropagation not only leads to the adaptation of network parameters that helps improve accuracy, but also leads to a reduction in the computational loss, for example, by reducing the latency expected when the neural network being trained is mapped onto the target hardware. Thus, contrary to the method disclosed by Shen et al., in which a fixed computational cost for each set of feature maps of the target hardware is assumed, the improved method calculates the cost of each set of neurons according to the mapping of the network to the multi-core architecture, and this mapping changes as the weight sparsity and accuracy change during the entire training process.
[0018] In one embodiment, the initially created mapping is used to actually map the trained neural network onto the target processor system. Even in the case where this initial mapping is not optimal, the improved training method still helps to achieve a low computational performance loss because it tends to train the neural network such that the computational load occurring in the neural network layers that dominate the overall computational performance loss of the neural network is reduced. In an example, the above sequence of training steps is repeated a predetermined number of times. In another example, the sequence of training steps is repeated until a convergence criterion is met. This applies, for example, to the case where the total loss reaches a value less than a specified threshold.
[0019] In this regard, it should be noted that activation sparsity also plays a role in that, depending on the weights selected, a neural network layer can have a lower or higher output activity. This in turn affects the requirements of subsequent neural network layers.
[0020] In some embodiments, the method includes performing a neural network remapping step each time a predetermined number of training sequences have been executed. In the neural network remapping step, the currently provided mapping and weight distribution are replaced with a different mapping and weight distribution taking into account the computational load associated with the activation sparsity occurring in the neural network layers identified during the predetermined number of training sequences. The role of the weight distribution is twofold. First, increasing the number of zero weights affects the mapping not only by reducing the computational load but also by reducing the requirements for memory resources. Second, weight adjustment can result in different activation rates of neurons. After the remapping and updated weight distribution, the training sequences are again executed a predetermined number of times to minimize the total loss of the neural network according to the new mapping.
[0021] In these or other embodiments, the method for training can further include pruning the neural network each time a predetermined number of training sequences have been executed.
[0022] In these or other embodiments, the method for training can further include adaptively adjusting the quantization level of a number of neural network computations each time a predetermined number of training sequences have been executed.
[0023] In these or other embodiments, the method for training also includes adaptively adjusting the activation sparsity each time a predetermined number of training sequences have been executed.
[0024] In an embodiment, the computational performance loss is determined based on the amount of non-zero values in the input of a layer and its kernel size. Thereby, computational performance parameters such as the latency per kernel per layer can be estimated. This information can also be used to manipulate one or more of pruning, quantization, activation regularization, gating, or dropout more aggressively on layers with a higher latency per kernel by adjusting their coefficients per layer. Contrary to known methods, these optimizations are performed as part of a mapping-aware training process. Thereby, it is achieved that the optimizations do contribute to mapping the neural network on the target hardware.
[0025] According to a method applicable in the absence of computational performance information (such as throughput information, power consumption information, and / or latency information for each layer), training methods such as evolutionary training, reinforcement learning, Bayesian optimization, etc. can be used as learning heuristic search strategies to find the set of weights that gives the lowest computational performance loss. In this method, the weights of the neural network elements are randomly mutated, and the "computational performance loss space" is evaluated. Then, the set of weights with the smallest amount of computational performance loss is selected again for reproduction. When the loss of the set of weights is evaluated and the smallest fitting set is eliminated, the loop starts again. Similarly, if the latency information for each layer is not present, other methods (such as reinforcement learning and Bayesian optimization) are applicable.
[0026] In an embodiment, then by calculating the computational performance loss as the accumulation of the latency of each core of all layers, where Kx l , Ky l , Kz l are the core sizes of layer l in the x, y, and z dimensions respectively, C l is the number of cores assigned to layer l, and where, is an indicator of the ratio of non-zero weights in the core, where n is a hyperparameter.
[0027] The computational performance loss is summed with the task loss in the backpropagation phase, allowing weight updates in the direction of the minimum computational performance loss and the minimum task loss. Similarly, the mapping algorithm can be re-evaluated after n training steps to update C l and ensure its effectiveness.
[0028] In the case where the loss function is differentiable with respect to the output / weight, the gradient of the loss can be calculated with respect to the weight of the last layer using the chain rule, and therefore the gradient of the loss is calculated with respect to the weight of the previous layer. The weight is then changed by the negative value of the gradient, which is the direction of the minimum total loss value [Goodfellow, Ian; Bengio, Yoshua; Courville, Aaron (2016). "6.5 Back-Propagation and Other Differentiation Algorithms (6.5 Back-Propagation and Other Differentiation Algorithms)". Deep Learning. MIT Press, pp. 200-220. ISBN 9780262035613], the improved method disclosed herein identifies computationally intensive layers and bottlenecks by continuously mapping each layer to different cores of a large-scale multi-core system, and implements sparsity in these identified elements of the processing chain. Therefore, the improved method results in a mapping neural network with reduced computational performance loss, for example, with reduced inference latency and enhanced throughput compared to what can be achieved using the global optimization techniques discussed above. Furthermore, in contrast to prior art methods, the improved method maximizes resource utilization in multi-core systems.
[0029] According to a second object, an improved training system for training a neural network is provided. The improved training system is configured to train a neural network to perform a data processing task when executed by a multi-core processor system including a plurality of processor cores.
[0030] The training system is also configured to receive a specification of a neural network architecture for the neural network, and training data for determining neural network parameter values by training the neural network.
[0031] The training data consists of multiple pairs of input data and associated true output data.
[0032] The training system is configured to provide a mapping of a neural network for a multi-core processor system and execute the neural network for training. To this end, the training system is configured to repeat a training sequence that includes the following operations: updating the weights of the neural network by backpropagation of the total loss calculated during the execution of the neural network using training data by the training system. More specifically, in response to a sample of input data, the training system executes the neural network being trained to generate output data. The training system calculates a task loss that indicates the deviation between the generated output data and the true output data corresponding to this sample of the input data. The training system calculates a computational performance loss that indicates the cost of executing the neural network if the neural network is executed by the multi-core processor system according to the provided mapping. Subsequently, the training system calculates a total loss that is positively correlated with both the task loss and the computational performance loss. Then, the training system updates the weights of the neural network being trained by backpropagation of the total loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] These and other aspects of the present invention will be described in more detail with reference to the accompanying drawings. Among them,
[0034] Figure 1 FIG. schematically shows a first embodiment of an improved training method;
[0035] Figure 2 FIG. schematically shows a second embodiment of an improved training method;
[0036] Figure 3 FIG. schematically shows a first embodiment of an improved training system;
[0037] Figure 4 FIG. schematically shows an aspect of a second embodiment of an improved training system;
[0038] Figure 5 FIG. schematically shows an aspect of a third embodiment of an improved training system. DETAILED DESCRIPTION
[0039] In the description, unless otherwise specified, the same reference numerals in the various drawings indicate the same elements.
[0040] Figure 1 FIG. schematically shows a first embodiment of a method for computationally-aware training of a neural network NN that is to be executed by a multi-core processor system MPS (i.e., a processor system including a plurality of processor cores PC) to perform a data processing task.
[0041] Among them, the reference sign S1 represents the process of providing a neural network architecture NNA for a neural network. The neural network architecture NNA specifies the components of the network, such as convolutional layers, pooling layers, and their interrelationships. The network architecture can be a known network architecture, such as ResNet version, MobileNet version, or VGG or other versions. The network specified in process S1 still needs to be trained with training data in order to determine the network parameters that can be used to perform data processing tasks. An appropriate mapping of the trained neural network on the multi-core processor system MPS must also be found.
[0042] In processes S3 and S5, training data TD{D1, GT1; D2, GT2;...D2n, GTn} is provided, which includes multiple pairs of input data Di and corresponding true output data GTi. The training data TD includes, for example, image data and true data specifying the object category and its position in the image data.
[0043] During the training process, a sequence with the following processes is repeated.
[0044] In process S2, a tentative mapping MP of the neural network NN on the multi-core processor system MPS is provided. The mapping specifies which parts of the multi-core processor will be assigned to which parts of the neural network.
[0045] The mapping assigns, for example, a corresponding set of one or more cores to each neural network layer of the neural network. The multi-core processor MPS can include a communication network, for example, an on-chip communication network (NoC) that enables the processor cores to communicate with each other. In one example, the network capacity is allocated to the processor cores in a predetermined manner. In another example, the allocation of the network capacity can be part of the mapping.
[0046] The mapping component for performing the mapping process can be pre-initialized through the initialization process S20, where an initial mapping MP0 is provided based on the specifications HWS of the target hardware MPS. The tentative mapping is used to enable the estimation of computational performance parameters such as the latency occurring in the neural network NN, the throughput of the neural network, and the power consumption of the neural network if the tentative mapping is actually implemented on the target hardware.
[0047] In process S4, in response to the input data Di provided in process S3, the training processor TP executes the neural network NN in its current training state to generate output data Oi. In this case, the training processor TP that executes the neural network NN is denoted as TP(NN). In this regard, it should be noted that in this process, the neural network NN does not have to be actually executed by the multi-core processor MPS for which it is trained to be used. It can be executed by any type of processor that can execute the training process within a reasonable amount of time.
[0048] In process S6, the total loss Lti is calculated. The calculated total loss indicates the process of training the neural network NN for the purpose of performing a data processing task when the target multi-core processor MPS executes the data processing task according to the current mapping. The total loss is calculated as follows in sub-processes S6A, S6B, and S6C of process S6.
[0049] In sub-process S6A, a task loss is calculated, which is a function of the output data Oi generated in process S4 and the true output data GTi provided in process S5, as it indicates the deviation between the generated output data Oi and the true output data GTi. The task loss may include one or more of a regression loss and a classification loss.
[0050] In sub-process S6B, a computational performance loss is calculated, which indicates an estimate of the cost of executing the neural network when the neural network is mapped onto the target multi-core processor MPS. In an example, the computational performance loss indicates an estimate of the total latency expected to occur in the multi-core processor MPS, i.e., the length of the time interval between when the neural network receives input data and provides output data when executed by the multi-core processor MPS according to the mapping.
[0051] In sub-process S6C, the total loss is calculated based on the task loss and the computational performance loss.
[0052] In process S7, the weights of the neural network NN are updated by backpropagation of the total loss.
[0053] In process S8, it is determined whether the neural network NN meets the performance requirements when mapped onto the target hardware MPS according to the current mapping MP. In one example, if the total loss does not exceed a predetermined maximum value, the performance requirements are met. In another example, if the task loss and the computational performance loss each do not exceed their respective predetermined maximum values, the performance requirements are met. In an alternative embodiment, it is assumed that the performance requirements are met if the training sequence is applied a predetermined number of times.
[0054] Figure 2 A second embodiment of the method is shown, which differs from the method shown in Figure 1 in that a neural network optimization process S10 is further applied. In the example shown, the neural network optimization process S10 is applied whenever it is determined in block S9 that the training sequence has been executed a predetermined number of times. The training and optimization processes continue until it is determined in process S8 that the performance requirements are met.
[0055] In one example, the neural network is optimized by pruning in process S10. As a result of the pruning, the number of non-zero parameters of the neural network is reduced, which tends to contribute to more efficient execution with shorter latency. Hao Li et al. specify an exemplary pruning method in PRUNING FILTERS FOR EFFICIENT CONVNETS, which was published as a conference paper at ICLR 2017 and is available via arXiv:1608.08710v3 [cs.CV] 10 Mar 2017.
[0056] In another example, the neural network is optimized by weight quantization in process S10. A reduction in the number of quantization levels can enable more efficient execution on the target hardware. In one example, the weights of the network or even the data being processed are quantized to binary variables. This method is discussed by Rastegari et al. in XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks, available via arXiv:1603.05279v4 [cs.CV] 2 Aug 2016.
[0057] Optimization of the neural network, for example by pruning and / or quantization, may also lead to a reduction in accuracy. The loss of accuracy is mitigated by subsequently performing a further training sequence.
[0058] Another optional optimization process that can be combined with pruning and / or quantization in a spiking neural network is sparsity control. In a spiking neural network, a neural element emits an output message if the activation threshold is exceeded. During the sparsity control process, the activation threshold is increased, such that the number of output messages is reduced. Consequently, the computational load on the recipient neural elements is also reduced. Additionally, the message network load is decreased. These effects all contribute to more efficient execution with shorter latency. Furthermore, for this type of optimization, the potential loss of accuracy is mitigated by subsequently performing a further training sequence.
[0059] Kurtz et al. discuss sparsity control in Inducing and Exploiting Activation Sparsity for Fast Neural Network Inference. See Proceedings of the 37th International Conference on Machine Learning, Vienna, Austria, PMLR 119, 2020.
[0060] In this regard, further refer to "ARTS: an adaptive regularization training schedule for activation sparsity exploration" by Zeqi Zhu et al., see https: / / pure.tue.nl / ws / portalfiles / portal / 215820255 / ARTS_DSD_2022_Camera_Ready.pdf.
[0061] The authors proposed a training schedule therein, which provides joint weight and activation sparsification in the model by adaptively changing the regularization coefficient through training. The present invention provides further improvements in the following aspects: using mapping information to specifically optimize the use of activation sparsity of the neural network mapped on the target hardware.
[0062] Figure 3 A first embodiment of the improved training system is schematically shown. Unless otherwise specified, features corresponding to those Figure 1 and Figure 2 presented have the same reference numerals.
[0063] The improved training system is configured to train a neural network NN, which will be executed by a multi-core processor system MPS including a plurality of processor cores PC to perform data processing tasks.
[0064] As Figure 3 shown, the training system is configured to receive a specification of the neural network architecture NNA of the neural network. It is also configured to receive training data TD including input data Di and ground truth output data GTi for determining neural network parameter values by training the neural network.
[0065] The training system includes a mapping unit MAP and is thus configured to provide a mapping of the neural network NN for the multi-core processor system MPS. In an example, the mapping data Cl specifies how the cores of the processor are allocated to components (e.g., neural network layers) of the neural network NN.
[0066] The training system also has the processing ability to perform the training of a neural network. In the example shown, the processing ability is indicated as a general neural network processor NNP, which is configured to execute the neural network specified by the neural network architecture NNA such that it can be trained by repeating a training sequence. As described above, the hardware used by the training system is generally different from the hardware used for training because the trained neural network can be applied by a large number of end users, each with its appropriate hardware, such as a mobile phone or a laptop computer.
[0067] As Figure 3 shown, the training system includes a task loss calculation unit CAL, a computational performance loss calculation unit CLL, a total loss calculation unit LWU, and a backpropagation unit BP1. In this regard, it should be noted that these units can be provided in the system as corresponding dedicated hardware units, but various other implementations are possible. For example, these units can be software modules executed by a common data processing facility.
[0068] As described above, a neural network with the neural network architecture NNA will be executed such that it can be trained by repeating a training sequence. The training sequence will now be described in more detail.
[0069] The general neural network processor NNP or other processing facility of the training system executes the neural network NN being trained while providing a sample Di of input data to the input of the neural network. As a result of the execution, the neural network NN generates output data Oi in response to the sample of the input data Di.
[0070] The task loss calculation unit CAL calculates a task loss Lli, which indicates the deviation between the generated output data Oi and the true output data GTi corresponding to this sample of the input data Di. For example, the task loss calculation unit CAL calculates a classification loss indicating the amount of misclassification and / or a regression loss indicating the degree to which the position estimate deviates from the position indicated by the true output data GTi.
[0071] The computational performance loss calculation unit CLL calculates a computational performance loss L2i indicating the cost of executing the neural network if it were executed by a multi-core processor system MPS, based on the mapping provided by the mapping unit MAP.
[0072] In an embodiment, the computational performance loss calculation unit CLL calculates the computational performance loss as follows.
[0073]
[0074] where Kx l , Ky l , Kz l are the kernel sizes of layer l in the x, y, and z dimensions, respectively, and Cl is the number of kernels assigned to layer l as indicated by the mapper mapping. Further, where ||X l || n is the Ln norm of the layer, which is an indication of the number of non-zero elements in the layer, and which is defined by defined.
[0075] Next, the total loss calculation unit LWU calculates a total loss Lti that is positively correlated with both the task loss Lli and the computational performance loss L2i. In the example shown, the total loss Lti is calculated as a weighted sum of the task loss Lli and the computational performance loss L2i, i.e., Lti = cl * Lli + c2 * L2i.
[0076] However, other implementations of the loss weighting unit LWU are possible as long as the operations it performs are differentiable. In another example, the total loss is calculated as follows:
[0077] Lti = (Lli + al) * (L2i + a2).
[0078] Thus, the total loss allows gradient descent training, where the weights of each layer are adjusted by backpropagation towards the minimum loss.
[0079] Based on the total loss Lti, the first backpropagation component BP1 applies a backpropagation operation to the neural network being trained in order to update its weights. Due to the fact that the total loss is partially determined by the latency loss, backpropagation not only results in an adaptive adjustment of the network parameters, which helps to improve accuracy, but also results in a reduction in the latency expected for the neural network being trained when mapped onto the target hardware.
[0080] As Figure 3 shown, as a first option, a mapping control component BP2 is provided, which is configured to control the mapper MAP with a control signal ΔCl in order to update the mapping taking into account the total loss Lti.
[0081] Other optional components are a pruning component PR, a quantization control component QTZ, and a sparsity control component SPR.
[0082] The optional pruning component PR is configured to apply a pruning operation to the current set of neural network parameters. Thus, it reduces the number of non-zero parameters as indicated by the Ln norm ||X l || n for each neural network layer. It is also possible to reduce the number of non-zero parameters by reducing the size of one or more of the kernel sizes Kx l , Ky l , Kz lPruning is applied. The pruning operation affects both the task loss Lli and the computational performance loss L2i. Generally, pruning will help reduce latency, i.e., reduce the computational performance loss, but also tends to increase the task loss by reducing the accuracy. The relationship between pruning operations and their impact on the task loss Lli and the computational performance loss L2i cannot be expressed as a differentiable function. Therefore, pruning operations are generally not applied during the gradient descent process, but different training processes (such as evolutionary training, reinforcement learning, Bayesian optimization, etc.) are used to perform them.
[0083] The optional quantization control component QTZ determines the quantization level n for performing the neural network computations of the layer in the target processor MPS. q The number of levels. Generally, the number of quantization levels is expressed as a power of 2. In addition, a change in the number of quantization levels of the quantization control component QTZ affects both the task loss Lli and the computational performance loss L2i. Generally, a reduction in the number of quantization levels will help reduce the computational performance loss, for example, by reducing latency or by increasing throughput or reducing the required supply power. In the extreme case, the number of quantization levels is reduced to 2. In this case, the neural network operations of the layer are simplified to Boolean operations that only require minimal computational time. On the other hand, a reduction in the number of quantization levels also tends to reduce the accuracy, i.e., increase the task loss. In addition, the relationship between quantization control operations and their impact on the task loss Lli and the computational performance loss L2i cannot be expressed as a differentiable function. Therefore, quantization control operations are generally not applied during the gradient descent process, but different training processes (such as evolutionary training, reinforcement learning, Bayesian optimization, etc.) are used to perform them.
[0084] The optional sparsity control component SPR determines the runtime behavior of the neural network NN using the parameter s. This parameter determines the sparsity of the output data provided by the neural network element to its receivers. This can be the case where the parameter s defines an activation threshold that the state value of the neural element needs to exceed to spike to the receiving neural network element. A higher activation threshold results in increased sparsity. Generally, increased sparsity will result in reduced latency because the receiving neural network elements receive a lower number of input data and need to perform a lower number of operations per unit time. In addition, the reduced network load due to increased sparsity can help reduce latency, reduce computational power, and / or increase throughput. That is, increased sparsity generally results in a reduction in the computational performance loss. However, in addition, increased sparsity tends to increase the task loss.
[0085] In addition, the relationship between sparsity control operations and their impact on the task loss Lli and the computational performance loss L2i cannot be expressed as a differentiable function. Therefore, sparsity control operations are generally not applied during the gradient descent process, but different training processes (such as evolutionary training, reinforcement learning, Bayesian optimization, etc.) are used to perform them.
[0086] Optionally, it may include each or a combination of components BP2, PR, QTZ, and SPR to help reduce the total loss.
[0087] Figure 4 Details of another embodiment of the training system are shown, which includes a multi-layer perceptron MLP configured to predict the latency per core per layer. Here, the input of the multi-layer perceptron is the matrix MNN, which has the identifiers of the neural network elements at one axis and their features, such as Ln-norm, kernel size, etc. at the other axis. As Figure 4 shown in the lower part of, in this method, the multi-layer perceptron MLP is pre-trained using training data TDL including a plurality of training examples. Each training example ATRj specifies a set of values for the attributes that can configure the layer, for example, its kernel size, its number of inputs, its stride, and the associated ground truth GTLj indicating the expected latency in the case where the layer is executed by a single processor core.
[0088] Therefore, during inference, when training the neural network NN, the multi-layer perceptron MLP does not need to have explicit information about the mapping. However, given the neural network elements and other features such as Ln-norm and kernel size, it can approximately process the load of each element (e.g., layer) per core LLC l,i in the training data Di. The computing performance loss calculation unit component CLL calculates the computing performance loss L2i as follows.
[0089]
[0090] Figure 5 Another example is shown. In this case, the specification Cl of the current mapping of the neural network NN on the target hardware HW is also provided at the input of the multi-layer perceptron. Therefore, here, the input of the multi-layer perceptron is the matrix M NN , which NN has the identifiers of the neural network elements at one axis and their features, such as Ln-norm, kernel size, etc. at the other axis. In this example, the number of cores assigned to each network layer is also included as a feature. Using this input data, the multi-layer perceptron MLP is configured to predict the total latency of the neural network in its current training state when the neural network is mapped to the target hardware MPS. This embodiment requires re-evaluating the mapping after n training steps to ensure its effectiveness for the updated weights and network structure.
[0091] It will be apparent to those skilled in the art that the elements listed in the system claims are intended to include any hardware (e.g., discrete or integrated circuits or electronic components) or software (e.g., programs or portions of programs) that reproduce or are designed to reproduce the specified function in operation, whether alone or in combination with other functions, whether in isolation or in cooperation with other elements. For example, in the embodiments of the training system discussed with reference to Figure 3 In the embodiments of the training system discussed, for the sake of clarity of the drawings, several functions are shown as discrete blocks. These include the general neural network processor NNP, the task loss calculation unit CAL, the computational performance loss calculation unit CLL, the total loss calculation unit LWU, the mapping unit MAP, etc. Two or more of these functional blocks can be implemented as software modules to be executed by a general-purpose processor. A single functional block (e.g., the general neural network processor NNP) can also be implemented as a multi-core processor. The present invention can be implemented by hardware including several different elements and by a properly programmed computer. In an apparatus claim listing several means, several of these means can be implemented by the same item of hardware
[0092] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single component or other unit can implement the functions of several items recited in the claims. The fact that certain means are recited in mutually different claims does not indicate that a combination of these means cannot be used advantageously. Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A method for training a neural network (NN), which will be executed by a processor system (MPS) including a plurality of processor cores (PC) to perform data processing tasks, the method comprising: Providing (S1) a neural network architecture (NNA) for the neural network, the neural network having a plurality of parameters to be trained using training data, the training data including input data (Di) and ground truth output data (GTi); Providing (S2) a mapping (MP) of the neural network (NN) for the multi-core processor system (MPS); Repeating the following sequence of training steps (S3 to S7): Providing (S3) the input data (Di) of the training data to a training processor system (TPS) that executes the neural network (NN) being trained; Wherein, in response to the input data (Di), the training processor system (TPS) executes the neural network (NN) being trained to generate (S4) output data (Oi); Providing (S5) the ground truth data (GTi) corresponding to the input data (Di); Calculating (S6, S6A) a task loss (Lli), the task loss indicating the deviation between the generated output data (Oi) and the ground truth output data (GTi); Calculating (S6, S6B) a computational performance loss (L2i), the computational performance loss indicating the cost of executing the neural network if the neural network is executed by the multi-core processor system (MPS) according to the provided mapping (MP); Calculating (S6, S6C) a total loss (Lti), the total loss being positively correlated with both the task loss (Lli) and the computational performance loss (L2i); Updating the weights of the neural network (NN) being trained by backpropagation (S7) of the total loss (Lti).
2. The training method according to claim 1, further comprising performing a neural network remapping step each time after a predetermined number of training sequences have been executed.
3. The training method according to claim 1 or 2, further comprising pruning the neural network each time after a predetermined number of training sequences have been executed.
4. The training method according to claim 1, 2 or 3, further comprising adaptively adjusting the number of quantization levels of neural network computations each time after a predetermined number of training sequences have been executed.
5. The training method according to one or more of the preceding claims, further comprising adaptively adjusting activation sparsity each time after a predetermined number of training sequences have been executed.
6. The training method according to one or more of the preceding claims, further configured to execute a multi-layer perceptron when calculating the computational performance loss.
7. The method for training according to claim 6, wherein, The multi-layer perceptron is provided at its input with a matrix that has the identification of neural network elements on one axis and their characteristics, such as the l_n norm, kernel size, on the other axis, where the multi-layer perceptron estimates the latency per kernel per layer, and where the training system further estimates the latency per layer by dividing the latency per kernel per layer by the number of kernels assigned to each layer, and estimates the total latency by accumulating the latency estimated for each layer.
8. The method for training according to claim 6, wherein, The multi-layer perceptron is provided at its input with a matrix that has the identification of neural network elements on one axis, and the number of assigned processing kernels and their characteristics, such as the l_n norm, kernel size, on the other axis, where the multi-layer perceptron provides an estimate of the total latency at its output.
9. A training system for training a neural network (NN) that is to be executed by a multi-core processor system (MPS) including a plurality of processor cores (PC) to perform a data processing task, the training system being configured to receive a specification of a neural network architecture (NNA) for the neural network, and training data including input data (Di) and ground truth output data (GTi) for determining neural network parameter values by training the neural network, the training system being configured to provide a mapping (MP) of the neural network (NN) for the multi-core processor system (MPS) and execute the neural network for training, the training system being configured to repeat a training sequence including the following operations: Execute the neural network (NN) being trained in response to a sample of the input data (Di) to generate output data (Oi); Compute a task loss (Lli) that indicates a deviation between the generated output data (Oi) and the ground truth output data (GTi) corresponding to the sample of the input data (Di); Compute a computational performance loss (L2i) that indicates the cost of executing the neural network if the neural network is executed by the multi-core processor system (MPS) according to the provided mapping (MP); Compute a total loss (Lti) that is positively correlated with both the task loss (Lli) and the computational performance loss (L2i); Update the weights of the neural network (NN) being trained by backpropagation (S7) of the total loss (Lti).
10. The training system according to claim 9, further configured to perform a neural network remapping step each time a predetermined number of training sequences have been executed.
11. The training system according to claim 9 or 10, further configured to prune the neural network each time a predetermined number of training sequences have been executed.
12. The training system according to claim 9, 10 or 11, further configured to adaptively adjust the number of quantization levels of neural network computations each time a predetermined number of training sequences have been executed.
13. The training system according to claim 9, 10, 11 or 12 is further configured to adaptively adjust the activation sparsity each time after a predetermined number of training sequences have been executed.
14. The training system according to one or more of claims 9 to 13 is further configured to execute a multi-layer perceptron when calculating the computational performance loss.