A method for training a deep neural network to classify data
By incorporating a clustering-based regularization process that promotes sparsity and clustering of filters/neurons, the method enhances the interpretability of deep neural networks, addressing the challenges of complex decision boundaries and overfitting in classification models.
Patent Information
- Application Number
- JP2021140892
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-07
- Filing Date
- 2021-08-31
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-08-31
AI Technical Summary
Existing deep learning classification models face challenges in interpretability due to complex decision boundaries and high risk of overfitting, especially in applications like healthcare and autonomous driving where explainability is crucial.
A clustering-based regularization process is implemented during the training of deep neural networks, which adds a regularization activation penalty to the loss function. This penalty encourages neuron activation to converge to a previous probability distribution, promoting sparsity and clustering of filters/neurons, thereby enhancing interpretability without requiring object part annotations.
The proposed method leads to more interpretable filters/neurons with small, compact activation regions, improving the model's ability to explain decisions and reducing redundancy in representations, thus facilitating better human understanding and trust in AI systems.
Smart Images

Figure 0007694265000046 
Figure 0007694265000047 
Figure 0007694265000048
Abstract
Description
Technical Field
[0001] Embodiments relate to a method of training a deep neural network (DNN) to classify data.
Background Art
[0002] State-of-the-art deep learning classification models contain millions, sometimes billions, of parameters, which result in very complex decision boundaries. The decision boundaries of deep learning models form high-dimensional manifolds that are not visualizable. Furthermore, having many parameters increases the risk of overfitting. Overfitting is typically detected by inspection in the train-test error / accuracy, but simply seeking model selection accuracy is not always appropriate. In some cases, for a machine learning model, it is important to be a faithful approximation of human-like reasoning even if it is less accurate than a conventional model trained for the same task.
[0003]
[0004] The sparsity of neuron activation is assumed to be desirable and appropriate for self-explainable models as it can lead to more interpretable models. In conventional neural networks, due to sparsity constraints on filters, the filters are forced to mimic the mammalian visual cortex areas V1 and V2. Furthermore, sparsity can induce fewer rules without sacrificing accuracy, so sparsity can improve the performance and interpretability of rule extraction logic programs. The fewer the number of rules, the easier it is for humans to interpret them, and thus they can explain the decisions made by neural networks in a better and more lightweight way.
[0005] Interpretable machine learning models are desirable in many real-life scenarios including important fields such as healthcare, autonomous driving, and finance. Here, the models should explain their decision-making processes; otherwise, they will not be trusted. For example, explaining the decision-making process of a neural network can assist doctors in making better diagnoses of patient conditions and reduce human errors.
[0006] In typical convolutional neural networks (CNNs), filters may fire in response to multiple object parts within the input image, and often, the important regions of activation are very large. This makes it difficult to evaluate the cause of filter activation and hinders interpretability. Furthermore, in typical CNNs, images are often associated with high activation by many filters, and the lack of sparsity in this activation makes it more difficult to explain the CNN's decision-making process based on filter activation. Therefore, linking neurons (or clusters of neurons) to specific object parts is considered a desirable step towards explaining the decisions made by neural networks based on their neuron activation.
[0007] One way to train a more interpretable model is to perform some kind of clustering within the filter / neuron space. The main idea is to encourage filters / neurons to form groups in response to common object parts or patterns that exist in a particular class or are shared between classes. Each neuron / filter can then be associated with a particular object part or topic, after a labeling process that may be manual or automatic. Neurons with high activation regions are more important for explaining the decisions made by the model. These activations can be used by a rule extraction program to explain the specific decisions made by a complex model, improving the interpretability of the learned representation while maintaining fidelity.
[0008] Many supervised approaches have been proposed that use object part annotation to associate filters with specific object parts. However, such detailed data is very expensive to obtain because it requires a great deal of effort to label, and most data does not have such annotations. Therefore, it is very useful to train the model in an unsupervised way (without object part annotations) and teach that those filters are interpretable by representing specific object parts.
[0009] One previous proposal associates filters with specific object parts by introducing an additional penalty called "filter loss" into an objective function that assigns each filter f to a category c having an image that most activates the filter f. These losses are expressed in terms of the mutual information between the feature map and several templates, causing each filter f to represent a specific object of category c and remain silent about other categories. That is, each filter is associated with one class. This results in redundant representations, for example, instead of having different filters that fire for "dog tails", "cat tails", "bird tails", and having a unified filter that represents "tails" and can be activated in multiple classes simultaneously. Clearly, this method succeeds in unraveling the representation and linking the filters to objects of a specific class, but does not promote a moderate representation (sparsity) that might help reduce the redundancy of the representation.
[0010] Some regularization approaches have been proposed to achieve sparse activations, but none of them simultaneously achieve clustering in the filter space (e.g., filters representing object parts or topics) and a sparse representation. SUMMARY OF THE INVENTION
[0011] An embodiment of a first aspect is a method implemented by a computer for training a deep neural network (DNN) to classify data, where the data may be, for example, an image or in tabular form, and the method includes: For a batch of N training data X i where i = 1 to N, and c iis the class of the training data, and during the training of the DNN, in at least one layer I of the DNN having neuron j, performing a clustering-based regularization process, in which process a regularization activation penalty is added to the loss function of the batch of training data to be optimized during training, whereby the regularization activation penalty includes components associated with each neuron within the layer that depends on each class of the training data, step including.
[0012] The clustering-based regularization process may include, for each class, obtaining a probability distribution before neuron activation, prior to adding the regularization activation penalty. The regularization activation penalty may be structured to induce neuron activation to converge to the previous probability distribution.
[0013] The previous probability distribution may be a coarse distribution in which only a low percentage of neurons within layer I are activated for the class.
[0014] The previous probability distributions of at least some classes may intersect.
[0015] Embodiments provide a clustering-based regularization technique to train more interpretable filters in a convolutional neural network (CNN) or more interpretable neurons in a generally feed forward neural network (FFNN), while achieving a desired degree of sparsity in simultaneous activations. Thus, the DNN can learn faster. Further, it enables pruning of unimportant neurons, thereby reducing the memory requirements of the DNN.
[0016] The proposed method encourages each filter in the convolutional layer to represent object parts or concepts without requiring object part annotations. This is achieved by imposing a penalty on the activation of the filter using the ground truth label of each image / sample as a supervisory signal. After being trained by the proposed method, the activation regions of the filters become small and compact. Thus, after the labeling process (which can be manual or automatic), each filter / neuron may be associated with a specific object part or concept. This results in more interpretable filter / neuron activations. This is an important step towards explainability in artificial intelligence.
[0017] The proposed method may also be used for transfer learning where a machine learning model is trained in one domain and then applied to another domain with little or no additional training. By utilizing the learned interpretable representation, less data is required to train the model, thus reducing the cost of obtaining big data for businesses.
[0018] In one embodiment, the clustering-based regularization process may include, for each neuron, clustering the components of the regularization activation penalty associated with the neuron, where the amount of the component is the probability p that the neuron is activated according to the previous probability distribution. jci The components of the regularization activity penalty can be calculated using the following formula:
Equation
[0019] The regularization activation penalty R(W 1:l ) can be calculated using the following formula:
Number
[0020] In another embodiment, the clustering-based regularization process may include, in each iteration of the process, determining the previous probability distribution for each class before adding the regularization activation penalty.
[0021] The step of determining the previous probability distribution for each class may include determining the probability distribution using the neuron activations of the previous class from the previous iteration.
[0022] The clustering-based regularization process may further include identifying a group of neurons whose number of neuron activations of the class satisfies a predetermined criterion using the determined previous probability distribution.
[0023] The predetermined criterion is whether the neuron is ranked within the top K neurons when the neurons are ranked according to the number of neuron activations of the class from the previous probability distribution, where K is an integer, whether the number of neuron activations of the class from the previous probability distribution exceeds a predetermined activation threshold, and may be at least one of them.
[0024] The regularization activation penalty may include penalty components calculated for each neuron not included in the group, and there are no penalty components for the neurons within the group.
[0025] Alternatively, the regularization activation penalty may include penalty components calculated for each neuron in the layer, and the penalty components of neurons not included in the group are greater than those of neurons in the group. In the regularization process based on clustering, neurons may be ranked according to the number of activations of neurons in the class from the previous probability distribution. The penalty component of each neuron may be inversely proportional to the ranking of the neuron.
[0026] An embodiment of the method further includes determining the saliency of neurons in the layer and discarding at least one neuron in the layer that is less salient than others in the layer. That is, as described above, unimportant neurons may be pruned.
[0027] An embodiment of the method may further include applying a weight regularization technique to the layer after performing a clustering-based regularization process.
[0028] To obtain rules that explain the activation of neurons, after training is completed, rule extraction techniques may be applied to the DNN. That is, the proposed method may be combined with a post-hoc rule extraction program (e.g., but not limited to, the one proposed in EP3291146) to achieve better and more interpretable rules. Since fewer filters / neurons fire for a particular image / sample, the sparsity of activation may improve interpretability. Using sparse activation, the rule extraction program may generate fewer but more interpretable rules while maintaining high fidelity.
[0029] As described above, a manual or automatic neuron labeling process may be applied to the DNN after training is completed to associate neurons with specific object parts or concepts.
[0030] In certain implementations, the method according to the embodiments may be used to train a DNN for use in controlling a semi-autonomous vehicle. For example, in an example of transfer learning, a CNN trained to recognize traffic signals using a dataset including images of traffic signals from one country using the method according to the embodiments is trained faster to recognize traffic signals from another country than a CNN trained in a different manner.
[0031] Embodiments of the second aspect provide a computer program or a computer program product including instructions that, when executed by a computer, cause the computer to execute any of the methods / method steps described herein, and a non-transitory computer-readable storage medium including instructions that, when executed by a computer, cause the computer to execute any of the methods / method steps described herein.
[0032] Embodiments of the third aspect are devices for training a deep neural network (DNN) to classify data, where the data may be, for example, in the form of images or tables. The device includes at least one processor and at least one memory storing the DNN, the data to be classified, and instructions. The instructions cause the processor to For a batch of N training data X i where i = 1 to N, and c i is the class of the training data, in at least one layer I of the DNN having neurons j, perform a clustering-based regularization process in which a regularization activation penalty is added to the loss function of the batch of training data optimized during training, whereby the regularization activation penalty includes components associated with each neuron in the layer that depends on each class of the training data.
Brief Description of the Drawings
[0033] By way of example, reference is made to the following accompanying drawings.
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 14C
Figure 14D
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21A
Figure 21B
Figure 22
Figure 23
Figure 24
Figure 25
Best Mode for Carrying Out the Invention
[0034] <Summary> This proposal aims to address the aforementioned inefficiencies in a unified approach. The goal is to train more interpretable filters / neurons by introducing sparsity in activation, and cluster them to encourage firing in response to small semantically meaningful regions. As a result, these regions may be associated with specific object parts in a subsequent labeling process. This clustering is achieved without using specific object part annotations, but instead using only the ground truth labels of each sample as the teacher signal, making the method widely applicable.
[0035] As described above, previous proposals associate filters with specific object parts through an appropriate "filter loss". However, the proposed loss can introduce redundancy in the training of different filters for each class for the same concept in the representation, especially when the model has high capacity for the problem to be solved (e.g., instead of learning the general concept of "tail", they learn "cat's tail", "dog's tail", etc.). The method proposed in this application aims to address this inefficiency by introducing sparsity in the representation. By encouraging a concise representation, the filters are induced to capture the most discriminative information and hopefully avoid the problem of redundancy.
[0036] FIG. 1 is a flowchart of a method implemented by a computer for training a DNN to classify (image or tabular) data according to an embodiment. Step S1 includes selecting which layer or layers L m (m = m1,..., m n ) of the DNN should be regularized in the clustering-based regularization process, and which hyperparameter λm should be used (these may be input by the user). Step S2 is N training data X iFor the batch, i = 1 to N, c i is the training data X i and includes the step of performing clustering-based regularization processing in the selected layer I of the DNN during the training of the DNN. In the regularization processing, the regularization activation penalty is added to the loss function of the batch of training data to be optimized during training. The regularization activation penalty includes components associated with each neuron in the layer that depends on each class of the training data.
[0037] For example, in the forward pass, the penalty R m (W 1:m ) is calculated for each layer according to the algorithm to be used for the layer (described later). After reaching the classification head, the loss to be minimized is as follows:
Equation
[0038] In the present application, with the first algorithm, hereinafter referred to as "Algorithm I" or "Elite-BackProp with prior distribution", for each class, a prior distribution regarding filter activation is imposed, thereby resulting in a semantically meaningful clustering of the filter space. Allowing different classes' prior distributions to intersect to doThis allows for the modeling of concepts shared between classes (e.g., intuitively learning only one filter for the concept of head instead of one filter per class). Furthermore, by imposing a sparse prior distribution as shown in Figure 21B, a desirable level of sparsity in filter activation can be achieved to reduce the redundancy of the representation (i.e., the same abstract concept, e.g., "head", is represented by different filters).
[0039] The "Elite - BackProp with prior" algorithm is supervised in the sense that a prior distribution regarding filter activation needs to be manually defined for each class. Simple prior distributions such as uniform or Gaussian are easy to configure, but defining more complex prior distributions can be more difficult as the number of classes increases, as the possibilities increase combinatorially. The question of "how many filters should be activated for different classes" is difficult to answer empirically. Furthermore, an "incorrect" prior distribution may still result in redundancy in filter activation. For example, when the model has high capabilities for the problem at hand, a sparse prior distribution is appropriate, but if a dense prior distribution as shown in Figure 21A is defined, redundancy in the representation is expected. height To address the aforementioned problems, a second algorithm, hereinafter referred to as "Algorithm II" or "Elite - BackProp topK", is proposed. Here, the topK activations per class, called "Elite", are rewarded, and all other filters outside the "Elite" of the class are penalized. For each class, the "Elite" is defined during training in a fully supervised manner from the activation history. Basically, the filters are ranked according to their activations
[0040] and the topK are selected as the "Elite". re, any filter outside of "Elite" is penalized according to its ranking. The lower the rank, the greater the penalty. The proposed method is not limited to defining "Elite" for each class, and different approaches using thresholds (instead of ranking) may be utilized. Furthermore, it is also possible to impose a penalty on all filters inversely proportional to their ranking (i.e., impose a small penalty on "Elite" as well). Thus, each class is associated with the distribution of penalties for each filter. This is considered equivalent to imposing a prior distribution on filter activation. The "Elite-BackProp topK" algorithm addresses the redundancy problem by not requiring manual definition of the prior distribution and encouraging simple representations in a completely unsupervised manner.
[0041] After training with the proposed regularization method, the filters / neurons become more interpretable as they have high activation regions and sparse activations on meaningful parts / objects of the input image. The selective step may be to prune and discard unimportant neurons (low importance) for speed and memory benefits. Subsequently, a labeling process may follow to associate all filters with specific words that describe their activation (this may be manual or automatic). This is done either by visualizing each field of high activation of the filter across different filters or in an unsupervised manner using few-shot learning techniques.
[0042] Furthermore, a rule extraction logic program (although not limited to, proposed in EP3291146A) may be used after being trained by the proposed regularization to extract the knowledge of an interpretable neural model and explain its decision-making. Any other existing or future method of mapping filters / neurons to documents and generating a classification regarding those documents may also be used. Most rule extraction programs take as input the activations of filters from a subset of layers and measure the association with the target output. After associating the filters with documents, rules are generated to explain a particular decision-making, raising the interpretability of the underlying representation. This is always beneficial in some fields such as healthcare where doctors need to know not only the output classification but also the decision-making process of the model. For example, when detecting tumors or other diseases from images, it is beneficial to have access to a neural network where the filters fire in semantically meaningful image regions that assist in the diagnosis when the disease is present.
[0043] In summary, this proposal aims to train a neural network by introducing sparsity of activations through a regularization method based on two clusterings that use the ground truth label of each sample as a signal, such that the filters / neurons represent semantically meaningful object parts or concepts that are not necessarily associated with one specific class. The proposed method encourages simple representations, makes the activation regions of each filter small and compact, and facilitates associating its activation with object parts in another labeling process. Finally and importantly, the sparsity that introduces the characteristics of the proposed method may be combined with pruning for speed and memory benefits, making it even easier to implement deep neural networks on mobile devices. teacher The method described in this application is implemented in TensorFlow (trademark), although any other deep learning framework such as PyTorch (trademark) or Caffe (trademark) may be used.
[0044]
[0045] This proposal presents two regularization methods to achieve sparsity in clustering and activation. The proposed methods are described for convolutional neural networks as shown in Figure 15, but theoretically the same rationale applies to any architecture by replacing "filters" with "neurons". That is, when the data to be classified is in the form of images, CNNs are appropriate, but data in a general tabular form classified by a Feed Forward Neural Network (FFNN) may also be used with the proposed regularization approach.
[0046] Before explaining the proposed methods, some necessary preliminary knowledge is discussed and the notation for the remainder of this proposal is explained (see also the glossary at the end of this document).
[0047] The proposed clustering-based regularization method should be provided in layer l of a CNN consisting of the following filters:
Number
[0048] The following batch of images is given:
Number
Number
Number
Number
Number
[0049] 1. Algorithm I: "Elite-BackProp with Activation of Prior Distribution" 1.1. Detailed Explanation of the Method In this algorithm, the previous probability distribution regarding the filter activation in the layer is specified separately for each class. Next, a penalty is introduced into the loss function to encourage the activation of the filter so as to converge to the specified prior distribution. The objective is that a set of filters "fires" only for a specific class, differentiating this class from other classes, and that another set of filters may "fire" for multiple classes, representing the object parts shared among classes. Furthermore, by specifying a sparse prior distribution regarding the filter activation, the redundancy of the representation can be controlled. The filters / neurons are induced to localize meaningful objects for each class, which results in small activation regions, raises the interpretability of each filter / neuron, and consequently makes the model interpretable.
[0050] The method proposed to achieve this has imposing a penalty on filter activations that have a low probability of being activated for a particular class. Here, the low probability is measured in terms of the selected prior distribution. For example, if filter f i has the following probability of being activated for class c:
Equation
Equation
[0051] Generally, for image X iis of class \(c_i\in\{1,\ldots,1C\}\), \(p\) jci is the previous probability distribution specified for class \(c\) i with respect to filter activation, the penalty added to the loss for this image is as follows:
Number
Number
Number
Number
[0052] Note that the activation \(A_{ij}^{(l)}\) is a function of all weights \(W\) 1:l . Therefore, the proposed method regularizes all weights up to layer \(l\) and encourages the filters to cluster and have an incentive to represent specific object parts as specified by the prior distribution \(p\) jci .
[0053] The loss function that we optimize during training takes the following form:
Number
[0054]
Number
[0055] It is possible to specify any desired prior distribution, and during training, the filters are encouraged to have activations similar to the prior distribution. Intuitively, as described above, by clustering according to the prior distribution, it is to encourage the filters to represent specific object parts. Therefore, they are clustered to distinguish categories or to represent common topics shared among them.
[0056] The method of "instructing" the filters to have activations according to the prior distribution is discussed below and starts by first looking for the distribution of filter activations for the first class. During training, if the images of class 1 have high activations for filters 25 or 37 in layer l, a large penalty should be imposed. This is because according to the prior distribution, filters 25 and 37 should not be activated for that class. On the other hand, according to the prior distribution, filter 12 has a very high probability of being activated for class 1, so no penalty is imposed on the activation of filter 12. Filter 16 should be slightly penalized because the probability of its activation for class 1 is not 1.0. Therefore, the filter activation is penalized with a penalty inversely proportional to the specified prior distribution of filter activations. Figure 3B shows the prior distribution of the penalties applied to the filters during training according to the prior distribution of filter activations for class 1, class 2, and class 3.
[0057] There are no restrictions on what prior distribution is specified. For example, a uniform prior distribution of filter activation, as shown in FIG. 6, may be specified. Note that this prior distribution does not enforce representing common topics between different classes. This does not imply that some topics will not be learned during training. For example, if class 1 and class 2 share a common topic / object, but the "bucket" of intersecting filters is not specified in the prior distribution, these common objects may still be learned from some filters. The filters activated for class 1 may learn this topic, and the filters activated for class 2 may learn this topic. However, although it is very common in neural networks for different filters to learn the same topic, this introduces some redundancy in the representation.
[0058] 1.3. Weight regularization after the Elite - BackProp layer To limit the weights to a small Euclidean sphere, it is desirable to apply weight regularization (e.g., Ridge) to the layer following the application of the Elite - BackProp algorithm. The reason is that the proposed method imposes a penalty on the activation according to the prior distribution (e.g., "Elite" in Algorithm II as described later). If no limit is imposed, the model can freely learn arbitrarily large weights to cancel out the regularization effect. This problem is shown in FIG. 4. Assume that the Elite - BackProp algorithm is applied to layer l and the penalty imposed on the neuron activation for class A is as follows: All neuron activations except the first and second are penalized. That is,
Equation
Number
Number
Number
[0059] In a post - processing step, pruning techniques may be applied to remove unimportant filters. The importance (also known as saliency) of each filter / weight in the CNN / FFNN may be determined in terms of a metric (e.g., Lp, Lpq + p, Lp,q norms, and group sparsity), and the filters may be sorted according to this metric. Later, the least important filters / weights may be discarded by setting their effects to zero, and the pruned network may be fine - tuned (re - trained) to converge to a simpler function with a minimal loss in accuracy. This process may be performed multiple times in an iterative manner.
[0060] As described above, it is very difficult to define a good prior distribution for filter activation for each class. In particular, defining the appropriate number of filters to be activated for different classes gives rise to a separate combinatorial problem that is very time-consuming to solve. Furthermore, a poor choice of prior distribution can still cause redundancy in the representation, especially when the model has high capacity for the problem at hand.
[0061] To address these inefficiencies, a non - supervised method is proposed below that naturally defines a "prior distribution" and achieves a parsimonious representation and clustering of filters into semantically meaningful concepts.
[0062] 2. Algorithm II: "Elite - BackProp top - K activations" This chapter describes a non - supervised method to address the limitations of the aforementioned "Elite - BackProp by prior distribution activation". The main idea is to define a more natural "prior distribution" for the activation of filters for each class for each iteration of the algorithm. This prior distribution is not constant during training and is updated by examining the history of filter activations from all previous iterations. Filters that had high activations in the past for a class are rewarded, and filters that had low activations are penalized. to do Thus, this procedure constructs, for each class at each iteration, a histogram of activities that can be regarded as the prior distribution at that iteration. One can choose to reward a subset of activations per class by defining a top - K approach (i.e., ranking the filters and rewarding the highest activations or equivalently penalizing the minimum of J - K) or by defining a threshold, but the proposed method is not limited to these approaches. The goal is to encourage a parsimonious representation by rewarding a subset of filters (referred to as "Elite") in a non - supervised way in order to reduce redundancy and provide an incentive to focus on the most discriminative information.
[0063] The proposed second algorithm uses the ground truth labels of each class as the monitoring signal to achieve activity sparsity and clustering in the filter space. In this algorithm, the prior distribution of filter activation for each class is not specified. Instead, only the number K is specified, which is used to pick the "elite" of the filters with high activation for a specific class without any monitoring during the training process. This algorithm is described using the topK approach, but the proposed method is not limited to this.
[0064] The 'Elite' of each class is composed as follows. In each forward propagation path, the activation of the top K filters of each class is found and their activations are dynamically accumulated (for each filter and each class). Penalties are applied to all filters that do not belong to the "Elite" of the corresponding class in the backpropagation path. In this way, only the "Elite" of the filter neurons are activated in each class after training, which induces the desired degree of sparsity controlled by K (the higher K, the less sparse it means) and cluster formation in the filter neuron space. After training, there are also some filters that belong to the "Elite" of many classes. to do This is a situation that occurs naturally when the topics represented by the filters are shared among those classes.
[0065] The purpose of this method is to induce sparsity depending only on the "Elite" of the filter neurons of each class. In this way, the "Elite" filters represent only the top K most important objects / topics of each class to achieve good classification performance and have the incentive to freely represent the common objects shared among the classes or objects that identify the class. Finally, the filters that do not belong to any of the "Elite" are removed later to gain speed and memory advantages as a post-processing step, and the remaining network is fine-tuned.
[0066] The pruning technique aims to first find the "importance" (also known as saliency) of each filter / weight in the CNN / FFNN from the perspective of measurement criteria (such as Lp, Lp,q norms or those related to group sparsity), and classify them according to that measurement criterion. Then, by setting their effects to zero, the least important filters / weights are discarded. Finally, fine-tuning (retraining) of the pruned network is performed to converge to a simpler function while minimizing the loss of accuracy. This process can be executed iteratively multiple times.
[0067] 2.1. Detailed Explanation of the Method The Elite-BackProptopK method is as follows: In each forward propagation path, for each image X in the batch i the activation from the layer where the normalization layer is applied is found. c i is the i class of image X Grand Truce Let D be a dictionary with the target class as the key, and a vector that cumulatively stores the activation of the filter for each value. The dictionary D is initialized as a zero C l dimensional vector for each class. Here, the number of filters in the layer is such that in each iteration, after the layer reaches during the forward propagation path, the activation of the filter is calculated and the memory is updated dynamically:
Number
Number
[0068] When the dictionary D is updated, the filters are ranked according to their activation, and the "Elite" E(c i) is defined. That is, for the top K filters with the highest activation of the class, the activation of "Elite" is not penalized, but filters that do not belong to "Elite" are penalized in the backpropagation path, and the penalty is inversely proportional to the rank of the filter. The lower the rank, the higher the penalty. Figure 5, Figure 19, Figure 20, and the algorithms shown below explain the proposed method in detail.
Number
Number
Number
[0069] Just as in the case of Algorithm I, for the normalization of the "Elite - BackProp top - K activations" layer, weight regularization should desirably continue (see Section 1.3).
[0070] Working Example Experiments using the proposed regularization approach and a uniform prior or topK, as shown in Figure 16, were conducted on two datasets, namely, the road dataset obtained from Placesse365 (Zhou, B.; Lapedirza, A.; Khosla, A.; Oliva, A.;; Toralba, A.A.A.A.S.Pattern Anal. Mach. Intell. 2018, 40, 1452-1464) and the CUB200-2011 bird dataset (C. Wah, S. Branson, P. Welinder, P. Perona, S. Belongie: Caltech-ucsd birds-200-2011 dataset, 2011). The data may be in the form of images (thus, CNNs are more appropriate) or generally in tabular form (thus, FFNNs may also be used for our regularization approach).
[0071] A. Road Dataset This dataset contains three categories ("forest road", "highway", "street") from the Placesse365 dataset and aims to classify road scenes. The scenes may be described through sub-objects and the topics present in them, and thus this toy dataset was selected for testing the proposed method. The selected train-validation-test split is 10,445 - 1,500 - 3,055, with 500 images per class for validation and approximately 1018 images per class for testing.
[0072] To standardize the data, for each item, the mean value per channel was subtracted and divided by the standard deviation per channel of the images in the loaded dataset. Furthermore, to augment the data, the following transformations were performed on each image. The probability of horizontal flip was 50%, the probability of changing brightness was 30%, the probability of Gaussian blur was 20%, the probability of smoothing was 35%, the probability of converting the image to black and white was 20%, and the probability of adding salt and pepper noise was 30%.
[0073] A toy architecture for training the ROAD dataset is shown in Figure 22. The following parameters were used with the TensorFlow™ deep learning framework (although other DL frameworks such as PyTorch™ or Caffe™ could be used instead). · Optimizer: Adam algorithm with a learning rate of 0.00005, β1 = 0.9, β2 = 0.999, and ε = 1e-08. · The learning rate had a decay rate of 0.5 and a patience of 5 epochs. · Training was performed for 60 epochs. · The regularization layer was added after the GAP layer, and a uniform prior distribution was defined over 150 filter activations as shown in Figure 6. · As described above, it was also necessary to apply L2 weight regularization to the layers following the Elite-BackProp activation regularizer (see Section 1.3). In this case, L2 regularization with reg_val = 0.01 was applied.
[0074] A.1 Results after training with Elite-BackProp with prior distribution In this section, quantitative and qualitative results regarding the sparsity of filter activations and visualization of activation regions after training with Elite-BackProp with a uniform and sparse prior distribution are reported.
[0075] A.1.1 Training from scratch with a uniform prior distribution The architecture shown in Figure 22 was trained on the road dataset using Elite-BackProp with the uniform prior distribution defined in Figure 6. The test accuracy was 86.57% and the validation accuracy was 86%.
[0076] The average filter activation per class was calculated as follows: for each class in the test dataset, all images belonging to that class were passed through the CNN, and the spatial average of each filter after the last convolutional layer to which regularization was applied was recorded. That is,
Equation
Number
[0077] A simple global threshold t = 0.3 for activating the filter can be manually specified by visual inspection of the activation in Figure 7. If the average activation of the i-th filter f of the i-th image X i of i is below that threshold, that is,
Number
Number
[0078] Filters with high activation have the greatest impact on the classification scores for the linear layer, followed by softmax after the GAP layer.
[0079] Qualitative Analysis of Training from Scratch with a Uniform Prior Distribution The purpose of this qualitative analysis is to evaluate whether the filters are clustered after training by Elite-BackProp according to the semantically meaningful interpretable regions of the input images (according to the specified prior distribution). Figure 8 shows the high activation of filters trained to highly activate classes 1, 2, and 3 respectively by a specific uniform prior distribution in Figure 6. It shows the top activation of each filter for the road dataset (which contains 3 classes from the Places365 dataset as described above). In this figure, some examples of active filters that emit light for each class detecting trees (class 1), traffic signs (class 2), and buildings - sky (class 3) can be seen. Filters 1 - 50 fire for objects in class 1 (first row), filters 51 - 100 for class 2 (second row), and filters 101 - 150 for class 3 (third row).
[0080] The top 10 activation regions of the filters are calculated as follows:
Equation
[0081] The proposed algorithm Elite-BackProp can also be used to fine-tune a pre-trained model with a sparse prior distribution to impose sparser activation. In this case, the sparse prior Elite-Backprop defines the "Elite" of the filters for each class, where Elite is calculated from the activation of the pre-trained model. The "Elite" for each class represents the most effective filters for that class, and all filters outside the "Elite" are penalized during training.
[0082] The sparse prior Elite-BackProp can be regarded as a combination of the techniques of "Elite-BackProp with Prior Distribution" (Algorithm I) and "Elite-BackProptopK" (Algorithm II) and can be utilized to effectively fine-tune an existing model while inducing sparsity.
[0083] Fine-tuning by Elite-BackProp As described above, the elements of Algorithms I and II may be combined to fine-tune an existing model using Elite-BackProp with a sparse prior distribution. This may be achieved as follows. a) Preprocessing step (before training): For each class c of the training dataset: · Loop through all images of that class and pass them through the trained model.
[0084] · Extract the activation from the l-th convolutional layer where Elite-BackProp will be applied in the future.
[0085] · Calculate the average activation of each filter across all images of the current class.
[0086] · Rank the activations and select the top K activations (where the user specifies K) to create the "Elite" of the filters for the current class.
[0087] · Construct the following sparse prior distribution: For each filter f, if the filter belongs to the "Elite" of its class, assign probability p i with p jc = 1.0, and in other cases, assign probability p jc = 0.0. b) Fine-tuning by Elite-BackProp with a sparse prior distribution In this step, the pre-defined sparse prior distribution p jc is used, and a regularization layer is added to the l-th convolutional layer of the architecture, Quantitative analysis after fine-tuning with a sparse prior distribution In this chapter, the quantitative results after training using Elite-BackProp and a sparse prior distribution are presented. The tables in Figures 9 and 10 show the average filter activations before and after fine-tuning the VGG16 architecture shown in Figure 22 using Elite-BackProp after the GAP layer.
[0088] To construct the sparse prior distribution, the steps outlined in the previous chapter are continued. That is, for each class in the road dataset, loop through all images of that class to find on average the top 20 filter activations. Then, as described above, a sparse prior distribution is constructed that assigns probability 1.0 to the top 20 filters for each class and probability 0.0 to all others.
[0089] After fine-tuning with Elite-BackProp and a sparse prior distribution, perform the same process as described in the "Training from Scratch with Uniform Prior Distribution" section to evaluate the sparsity of the activations. For each class ("forest road", "highway", "street"), loop through all images belonging to that class and calculate all filter activations from the last convolutional layer. Similarly, for each image, a vector of 150 filter activations is obtained.
[0090] A filter is considered activated if its activation exceeds the threshold of that filter (as described above, for example, a global threshold, or
Number
[0091] Thus, in images of the "forest road" class, on average 30.674 filters are activated before applying the Elite-BackProp algorithm, and on average 4.52 filters are activated after training with "Elite-BackProp". to do This means that for each image, only a few filters are highly activated, making it easier to explain the classification decision.
[0092] In the table of Figure 10, sparsity is evaluated without using the Global Average Pooling (GAP) layer. This shows that significant benefits are obtained from using the proposed Elite‐BackProp with a sparse prior distribution, indicating that the sparsity-inducing property of the proposed method does not depend on the GAP layer.
[0093] A.1.3. Quantitative Analysis of Rule Extraction The rule extraction framework proposed in EP3291146A was used to extract knowledge from the trained CNN, and the number of rules was measured along with the classification accuracy of the extracted rules. The results are shown in the tables presented in Figures 11 and 12, where it can be seen that the use of the Elite-BackProp algorithm is associated with a reduction in the number of rules without sacrificing accuracy (same fidelity). Therefore, the use of the Elite-BackProp algorithm may lead to a more compact representation and interpretability. This is because a smaller number of rules may be more interpretable by humans.
[0094] For the architecture shown in Figure 22 related to the sparsity level shown in Figure 9, the rule extraction analysis is shown in the table of Figure 11.
[0095] For the architecture described in Figure 22 related to the sparsity level shown in Figure 10 without using the GAP layer, the rule extraction analysis is shown in the table of Figure 12.
[0096] From the results so far, in the case of the road dataset, global average pooling is considered useful for using rule extraction in combination with regularization. However, even without using the GAP layer, the proposed regularization leads to a more unique meaning (literal) equivalent to a simpler representation with little sacrifice in accuracy. B. CUB200-2011 Dataset The CUB200-2011 dataset contains 11.8K images of 200 bird species. Each category contains 12 to 33 images (an average of 22.4 images per category). Since this dataset is very small, extensive augmentation was applied as described above for the road dataset. The tracked training-validation-test split was 5696, 1600, and 4493. To standardize the data, for each item, the mean value per channel was subtracted and divided by the standard deviation per channel of the images within CUB200.
[0097] The architecture shown in Figure 23 was trained on the CUB200-2011 dataset using the TensorFlow (trademark) deep learning framework (other DL frameworks such as PyTorch (trademark) or Caffe (trademark) can be used instead) and the following parameters: Trained in the normal way. · Optimizer: Adam algorithm with learning rate 0.00005, β1 = 0.9, β2 = 0.999, ε = 1e-08. · The learning rate has a decay rate of fraction 0.3 and a patience of 5 epochs. · Trained for 100 epochs. · The regularization layer is connected after global average pooling, and the following uniform prior distribution is defined over 1000 filter activations: {class1:[1,2,3,4,5], class2:[6,7,8,9,10],…,class200:[996,997,998,999,1000]} Thus, for each class, 5 separate filter activations are specified in the list. · As described above, following Elite-BackProp (see Section 1.3), it was also necessary to apply L2 weight regularization to the layer. In this case, L2 regularization with reg_val = 0.01 was applied.
[0098] B.1. Results after training using Elite-BackProp topK activation In this chapter, we report quantitative and qualitative results regarding the sparsity of activation, and visualize the activation regions of filters after 100 epochs of training using Elite-BackProptopK activation with K = 20. The average accuracy of test and validation at the 88th epoch of the architecture in Figure 23 was 46.75% and 48.25% respectively.
[0099] Quantitative analysis We measured the average filter activation of all images in a given category (as described above: loop through all images of a specific class to obtain activations and then take the average), and plotted those values as shown in Figure 13.
[0100] It can be seen that Elite-BackProptopK introduces large spikes to filters that are on average very active for images of class 1. After visualizing their highly activated receptive fields within the image, filters with high activation spikes can be associated with objects of class 1.
[0101] In class 31, it is also revealed that large activation spikes occur in some filters that form the "Elite" of this class. Note that some filters overlap with class 1.
[0102] Qualitative analysis of training from scratch with Elite-BackProptopK Experiments were conducted using two different visualization methods. For the first visualization approach, in the previous chapter "Qualitative analysis of training from scratch with uniform prior distribution", the threshold μ j +σ j for the road dataset was described. The second visualization approach was proposed by D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba. Network Analysis: Quantifying the Interpretability of Deep Visual Representations, CVPR 2017. Both approaches yield similar results in terms of visualizing important regions of activation.
[0103] For completeness, the second visualization approach of Bau et al. is described below: For each filter f i the feature map F ij (and maxpool if present in the architecture) after ReLu operation is computed on different input images X i on the l-th layer where the proposed regularization is applied. Next, the distribution of activation scores at all positions of all feature maps is computed. Then, an activation threshold t fj is set as follows to retain the top activations from all spatial positions (r, s) of all feature maps: [Number] Finally, after thresholding the feature map to obtain a binary mask, it is scaled up to match the resolution of the input image, and the input image is masked and visualized.
[0104] Visualizations of high activations in some filters within the CUB200-2011 dataset are shown in FIGS. 14A - 14D. In the image of FIG. 14A, the filter detects "head", in the image of FIG. 14B, the filter detects "body", in the image of FIG. 14C, the filter detects "wing", and in the image of FIG. 14D, the filter detects "tree branch". From the visualization of the top activation regions in FIGS. 14A - 14D, it is clear that the filters have learned specific object parts or environmental concepts without the need for annotation of object parts during training.
[0105] In summary, the proposed method is a clustering-based regularization process with the following characteristics: (1) Filters are clustered and activated on specific object parts present in one or more classes, and the activation regions are small, compact, and semantically meaningful. (2) Filters are clustered considering the ground Truce truth labels for each image to guide the regularization. Filters are related to the class to which the image belongsGrand Truce Different penalties are imposed on each image in the batch according to the class. (3) One embodiment (Algorithm I) clusters the filters according to a specified prior distribution for the filter activations for each class. Each class may be associated with a prior distribution on the filter activations. The filters are trained to converge to that distribution. This clustering encourages the filters to fire in response to compact and semantically meaningful regions of the input image and associate with object parts of a particular class. Further, the activation regions are small and compact. (4) Another embodiment (Algorithm II) ranks the filters within a layer according to the activations accumulated for each class during training. Each class is associated with an "Elite" group of filters. A penalty is applied to all filters outside the 'Elite' group during backpropagation. As a result, a sparse representation is obtained, and filters that do not belong to any "Elite" group may be pruned for efficiency. The "Elite" group can be constructed using, for example, the top-K approach, a threshold, etc. (5) The functional form of the loss penalty by regularization activation changes from one iteration to the next. Penalties for Algorithms I and II:
Number
Number
Number
[0106] The proposed method has the following advantages: · It enhances the interpretability of the machine learning model by prompting neurons to form clusters that "fire" in response to specific target parts / concepts. · Since the activation regions of neurons are small and compact, it is easier to relate to the target site. · The sparsity of activation is introduced to address the redundancy problem of the prior art. For example, instead of having different filters for "cat's tail", "dog's tail", etc., a frugal representation is encouraged to help train one filter representing the general concept of "tail".
[0107] Sparsity also has the following advantages: · This can be combined with pruning of unimportant neurons (neurons with low amplitudes) for speed improvement and memory reduction with minimal loss of accuracy. This helps to provide deep learning algorithms to resource - limited mobile devices. · Annotating filters after training and linking them to specific object parts is made easier by sparsity (fewer filters require annotation due to sparsity). · It improves the performance of the rule extraction logic program. This is due to the sparsity and clustering of neurons. · Fewer rules are generated that can capture semantic information more compactly without sacrificing accuracy or faithfulness. · Fewer rules are more interpretable by humans.
[0108] Embodiments are applicable to any area where an interpretable neuron-filter having self-explainable model is required or sparsity of representation is desirable.
[0109] After achieving the desired level of sparsity, unnecessary filters / neurons can be pruned to increase speed and reduce memory requirements. This results in a more compact and lightweight DNN model, facilitating its incorporation into resource-constrained portable devices (e.g., mobile devices).
[0110] After training using the proposed regularization method, the filters / neurons become more interpretable and fire (i.e., have high activation regions) towards meaningful parts / objects of the input image. Subsequently, a labeling process may follow for each filter to associate all filters with specific words that describe their activation (this may be manual or automatic). Essentially, by visualizing the receptive fields of high activation of filters across different images, each filter can be associated with a word that describes its activation. Further, a rule extraction logic program may be used after training with the proposed regularization to extract knowledge of the interpretable neural model, not limited to, and explain its decision-making. Such a rule extraction program takes as input the activation of filters from a subset of layers and measures the relationship with the target output. After setting the activation thresholds for each filter / neuron, each of them may be either active or inactive, and for example, by creating a decision tree or graph, rules for explaining a particular decision are created, enhancing the interpretability of the underlying representation. This is highly beneficial in areas such as: · Healthcare: Assist doctors in diagnosing diseases from tabular data or images. In many cases, doctors need to know not only the output classification but also the decision-making process of the model. For example, when detecting tumors or other diseases from images, it is beneficial to have access to a neural network where the filter fires in semantically meaningful image regions that assist in the diagnosis if a disease is present. For example, when detecting a tumor, the filter can be made to fire only towards abnormal morphological objects associated with the presence of a particular type of tumor. Further, during training, ground truth polygon annotations for the location and shape of the tumor are not required, enabling the proposed method to be easily applied without supervision and without the need to obtain segmentation data. For example, using the proposed regularization term, a binary classifier on tumor / non-tumor images can be trained to give the filter an incentive to form clusters towards small and compact regions that identify the class representing a particular target part. After labeling the filter (either manually or automatically) and quantizing its activation, a rule extraction logic program can be used to generate rules for each input image of the patient that explain the decision made by the neural model. For example, there is a rule such as "Since Filter A is activated, Filter B is activated, and Filter C is activated, there is a malignant tumor in the patient with a probability of X%." Thus, when the filter activation exceeds a threshold (as described in the proposed method), the filter has detected the presence of a particular shape-color-object in the image. The probability is easily generated by having a softmax layer in the CNN output, and the uncertainty in the estimation can be measured by various methods such as MCMC dropout. The important part is to first make the CNN filters more interpretable and train them to link to specific object parts. · Autonomous driving: Based on input images from the environment, an autonomous vehicle determines maneuvers such as turning, accelerating, braking, and stopping. To increase confidence in the decision-making by such a system, it is beneficial to explain the decision-making process, auditing, semi-autonomous vehicle assistance, and debugging assistance. This can be done by training more interpretable filters within a CNN where each filter (or cluster of filters) can represent (detect) topics such as road object parts or white stripes, pedestrian or animal crosswalks, traffic signals, etc. As described above, when the filter activity exceeds a threshold, a particular object part / topic is present in the image. Subsequently, a rule extraction program can be used to distill knowledge using the filter activity as input and target the decision of the CNN. The rule extraction program can generate compact rules due to the induced sparsity in the proposed method of explaining the decisions made by the classifier. An example of a rule is "Since filter A is activated, the vehicle has stopped", which could be translated to braking due to detecting a red light (if filter A represents a red light and emits when a red light is present).
[0111] Explanation techniques can be applied, for example, to audits where an insurance agent would be interested in knowing why a vehicle took an incorrect action, what caused an accident, and who should be held responsible. Additionally, semi-autonomous vehicles can benefit from being more robust (improving generalization ability) to unseen environments with more interpretable filters. · Transfer learning: Filters that represent semantically meaningful concepts or object parts can be used in transfer learning scenarios where a machine learning model is trained in one domain and needs to be applied to another domain. As an example, a model trained on European traffic signs with Elite - BackProp can have filters that, in an unsupervised way, encourage the detection of primitive shapes and objects like circles and triangles within traffic signs. Later, this knowledge can be applied to a new domain, such as traffic signs in other regions, to detect and interpret traffic signs in the new domain. As already mentioned, the property that induces sparsity in the proposed method means that fewer filters require annotations, and as a result, less annotated data is needed, leading to speed - up and low cost for the business.
[0112] Transferable Road Sign Recognition - Usage Example Next, the application of the embodiment to transferable traffic sign recognition will be described with reference to FIG. 24.
[0113] There are road signs in every country. Individual signs vary from country to country, but may share common purposes, such as speed limits, restrictions on direction of movement, and warnings to drivers. Therefore, once one gets used to traffic signs in one country, one can understand the meaning of other signs with similar meanings in other countries. On the other hand, since traffic signs appear in various scenes including angles, lighting, and occlusion, existing image recognition technologies need to be trained on a large number of traffic sign images from each country. Despite these difficulties, using an example of the proposed method, an image recognition system trained in a human - like way on a set of traffic signs for one country can recognize traffic signs in different countries without zero - shot training.
[0114] Assume that a neural network is being trained using images of Japanese traffic signs. Figures C and D in Figure 24 show some of these traffic signs. All signs are target classes to be learned. Note that among the traffic signs, there are some with the same purpose. For example, the symbol in Figure C indicates something is prohibited, and the symbol in Figure D indicates a right turn is permitted.
[0115] Figure E in Figure 24 shows traffic signs prohibiting a right turn in the UK (left) or Japan (right). In the sense that the prohibited direction is shown in the UK and the permitted direction is shown in Japan, as can be seen in the image, the signs are contrasted with each other. However, they serve the same purpose, that is, they restrict the direction of traffic when they appear at intersections. There are no signs in Japan like those used in the UK, but traffic signs contain the same concepts such as prohibition and instruction. By combining these concepts, the proposed method makes it possible to construct rules for an image recognition system trained on Japanese traffic signs that can recognize UK traffic signs.
[0116] For example, the EliteBack-Prop algorithm can capture the concept of "prohibition" with both signs in Figure C using a prior distribution or top-K. This is because it has the ability to train kernels activated by common concepts among classes. Similarly, the "right arrow" can be captured from the symbol in Figure D. For example, using rule extraction techniques such as those proposed in EP3291146A, a rule set for recognizing Japanese signs can be extracted. That is, when provided with the images in Figures C and D, one of the rules matches and each is correctly classified. For example, the rule "X ∧ Y → U-turn prohibited" classifies the left image in Figure C, where kernel X represents "prohibition" and Y represents "U-turn". The rule "U ∧ V → only right turn" classifies the left image in Figure D, where kernel U represents "right arrow" and V represents "blue background". Since Japanese traffic signs do not represent "right turn prohibited", it does not become "X ∧ U" → right turn prohibited".
[0117] If the user wants the system to recognize "right turn prohibited" in the UK without using additional training images, the rule "X ∧ U → right turn prohibited" can be manually added to the rule set. However, this rule can be automatically generated by training the system on a relatively small image set that contains far fewer images than are normally used for neural network training and applying the aforementioned rule extraction technique.
[0118] FIG. 25 is a block diagram of a computing device, such as a data storage server, that embodies the present invention, performs some or all of the operations of a method of embodying the present invention, and can be used to perform some or all of the tasks of the apparatus of the embodiments. For example, the computing device of FIG. 25 can be used only to execute all the tasks of FIGS. 18 and 20, execute all the operations of the method shown in FIG. 1, or execute one or more of the processes described with reference to FIGS. 2, 5, 17, and 19.
[0119] The computing device includes a processor 993 and a memory 994, which may be configured to perform tasks of, for example, a deep neural network. Optionally, the computing device also includes a network interface 997 for communicating with other such computing devices, for example, other computing devices of embodiments of the present invention.
[0120] For example, one embodiment may be composed of a network of such computing devices. Optionally, the computing device also includes one or more input mechanisms, such as a keyboard and a mouse 996, and a display unit, such as one or more monitors 995. The components can be connected to each other via a bus 992.
[0121] Memory 994 may include a computer-readable medium, and this term may refer to a single medium or multiple media (e.g., a centralized or distributed database and / or associated cache and server) configured to carry computer-executable instructions or store data structures such as a road and Cub dataset. The computer-executable instructions may be accessible, for example, by a general-purpose computer, a special-purpose computer, or a special-purpose processing device (e.g., one or more processors), and may include instructions and data for causing one or more functions or operations to be performed. For example, the computer-executable instructions may include instructions for performing all tasks or functions to be executed by each or all of FIGS. 18 and 20, or instructions for performing all operations of the method of FIG. 1, or instructions for performing one or more processes described with reference to FIGS. 2, 5, 17, and 19. Such instructions may be executed by one or more processors 993. The term "computer-readable storage medium" can store, encode, or hold a set of instructions for machine execution, and may include any medium that can cause a machine to execute any one or more of the methods of the present disclosure. Thus, the term "computer-readable storage medium" includes, but is not limited to, solid-state memory, optical media, and magnetic media. For example, without limitation, the term "computer-readable storage medium" may include solid-state memory, optical media, and magnetic media. By way of example, and not limitation, such computer-readable media may include non-transitory computer-readable media including random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disk storage devices, magnetic disk storage devices, or other magnetic storage devices, flash memory elements (e.g., individual memory devices).
[0122] Processor 993 controls the computing device and is configured to perform processing operations such as executing computer program code stored in memory 994 to implement the methods described with reference to FIGS. 1, 2, 5, 17, and 19 and defined in the claims. Memory 994 stores data read and written by processor 993. As mentioned herein, the processor may include one or more general-purpose processing devices such as a microprocessor, a central processing unit, etc. The processor may include a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets or processors that implement a combination of instruction sets. Further, the processor may include one or more special-purpose processing devices such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. In one or more embodiments, the processor is configured to execute the operations and instructions discussed herein.
[0123] Display unit 995 can display a representation of data stored by the computing device and can also display a cursor, dialog boxes, and screens that enable interaction between the user and the programs and data stored in the computing device. Input mechanism 996 can enable the user to input data and instructions into the computing device.
[0124] Network interface (Network I / F) 997 can connect to a network such as the Internet and can connect to other such computing devices via the network. Network I / F 997 can control input / output data with other devices via the network.
[0125] Other peripheral devices such as microphones, speakers, printers, power supply units, fans, cases, scanners, trackballs, etc. may be included in the computing device.
[0126] The method of embodying the present invention can be executed on a computing device as shown in FIG. 25. Such a computing device does not necessarily have all the components shown in FIG. 25 and can be composed of a subset of these components. The method of embodying the present invention can be executed by a single computing device that communicates with one or more data storage servers via a network. The computing device may be a data storage device that stores at least a part of the data.
[0127] The method of embodying the present invention can be executed by a plurality of computing devices that cooperate with each other. One or more of the plurality of computing devices may be a data storage server that stores at least a part of the data.
[0128] The above-described embodiments of the present invention can advantageously be used independently of any other embodiments or in any feasible combination with one or more other embodiments. Brief Explanation of Technical Terms - Glossary · Post-hoc method = Post hoc method = A method that attempts to explain, locally or globally, the decisions made by a trained neural network by approximating the underlying complex model with a simpler and more interpretable alternative model. · FFNN (Feed Forward Neural Network) = Feed forward neural network · CNN (Convolutional Neural Network) = Convolutional neural network · GAP (Global Average Pooling layer) = Global average pooling layer · Kernel = A K1×K2 matrix that synthesizes the input image or feature map (varies by layer) · The convolution operator is represented by *. · W ij (l) represents the weight connecting the i-th neuron and the j-th neuron in layer l. · W j (l) represents the set of all weights (as a matrix) in layer l.
Number
Number
Number
Equation
Equation
[0129] In addition to the above embodiments, the following supplementary notes are further disclosed. (Supplementary Note 1) A method implemented by a computer for training a deep neural network (DNN) to classify data, comprising: For a batch of N training data X i where i = 1 to N, and c i is the class of the training data, in at least one layer I of the DNN having neurons j, performing a clustering-based regularization process, in which a regularization activation penalty is added to the loss function of the batch of training data to be optimized during training, whereby the regularization activation penalty includes components associated with each neuron within the layer that depends on each class of the training data. A method comprising the above steps. (Supplementary Note 2) The regularization process based on clustering includes the step of obtaining the previous probability distribution regarding neuron activation for each class before adding the regularization activation penalty. The regularization activation penalty is structured to include the activation of neurons that should converge to the previous probability distribution. The method according to Appendix 1. (Appendix 3) The previous probability distribution is a rough distribution in which only a low proportion of neurons in layer I are activated for the class, the method according to Appendix 2. (Appendix 4) The previous probability distribution overlaps with at least some classes, the method according to Appendix 2 or 3. (Appendix 5) The regularization process based on clustering is the step of clustering the components of the regularization activation penalty associated with each neuron, where the amount of the component is the probability p that the neuron is activated according to the previous probability distribution. jci The method according to any one of Appendices 2 to 4, further including the step. (Appendix 6) The component of the regularization activation penalty is calculated using the following formula: [Formula] Here, A ij (l) is the activation of neuron j in layer I for training data X i The method according to Appendix 5. (Appendix 7) The regularization activation penalty R(W 1:l ) is calculated using the following formula: [Formula] Here, W 1:l represents the set of weights from layer 1 to l, the method according to Appendix 6. (Appendix 8) The regularization process based on the clustering further includes, in each iteration of the process, determining the prior probability distribution for each class before adding the regularization activation penalty. The method according to any one of Appendices 2 to 4. (Appendix 9) The step of determining the prior probability distribution for each class includes determining a probability distribution using the neuron activation of the class from the previous iteration, according to the method described in Appendix 8. (Appendix 10) The regularization process based on the clustering further includes identifying a group of neurons whose number of activations of the neurons of the class satisfies a predetermined criterion using the determined prior probability distribution, according to the method described in Appendix 8 or 9. (Appendix 11) The predetermined criterion is when the neurons are ranked according to the number of activations of the neurons of the class from the prior probability distribution, whether the neurons are ranked within the top K neurons, where K is an integer, whether the number of activations of the neurons of the class from the prior probability distribution exceeds a predetermined activation threshold, at least one of which, according to the method described in Appendix 10. (Appendix 12) The regularization activation penalty includes a penalty component calculated for each neuron not included in the group, and there is no penalty component for the neurons within the group, according to the method described in Appendix 10 or 11. (Appendix 13) The regularization activation penalty includes a penalty component calculated for each neuron in the layer, and the penalty component for the neurons not included in the group is greater than that for the neurons within the group, according to the method described in Appendix 10 or 11. (Appendix 14) In the regularization process based on the clustering, the neurons are ranked according to the number of activations of the neurons of the class from the prior probability distribution, and the penalty component of each neuron is inversely proportional to the ranking of the neuron, the method according to Supplementary Note 13. (Supplementary Note 15) The method according to any one of Supplementary Notes 1 to 14, further comprising the step of determining the saliency of neurons in the layer and discarding at least one neuron in the layer that is less salient than others in the layer.
Explanation of Signs
[0130] 993 Processor 994 Memory 995 Display 996 Input 997 Network I / F
Claims
1. A computer-implemented method for training a deep neural network (DNN) to classify data, comprising: For a batch of N training data X, where i = 1 to N, and c i is the class of the training data, in at least one layer I of the DNN having neurons j, performing a clustering-based regularization process, wherein the clustering represents that the training data is classified into classes in the next layer I+1 by activation of the neurons in layer I, and in the clustering-based regularization process, a regularization activation penalty is added to the loss function of the batch of training data to be optimized during training, i comprising: The regularization activation penalty includes components associated with each neuron within the layer, The lower the regularization activation penalty associated with a neuron for a class, the higher the probability of activation of the neuron for that class, The clustering-based regularization process Before adding the regularization activation penalty, in each iteration of the clustering-based regularization process, for each class, determining a prior probability distribution regarding neuron activation using the neuron activation of the class from the previous iteration; Identifying a group of neurons for which the number of activations of the neurons of the class meets a predetermined criterion using the determined prior probability distribution; comprising: The regularization activation penalty for each neuron is determined based on whether each neuron is included in the group.
2. The clustering-based regularization process includes obtaining the prior probability distribution before adding the regularization activation penalty, The regularization activation penalty is structured to include the activation of neurons that should converge to the prior probability distribution. The method according to claim 1. **Claim 3** The prior probability distribution is a distribution in which, within layer I, one or more neurons are activated for one class, but one neuron is not activated for multiple classes. The method according to claim 2. **Claim 4** The prior probability distribution overlaps with at least some classes. The method according to claim 2. **Claim 5** The regularization process based on clustering is a step of clustering, for each neuron, the component of the regularization activation penalty associated with the neuron, where the amount of the component is the probability p that the neuron is activated according to the prior probability distribution. jci The method according to any one of claims 2 to 4, further including the step. **Claim 6** The component of the regularization activation penalty is calculated using the following formula: **Equation 1** where A ij (l) is the activation of neuron j in layer I for training data X i The method according to claim 5. **Claim 7** The regularization activation penalty R(W 1:l ) is calculated using the following formula: **Equation 2** where W 1:l represents the set of weights from layer 1 to l. The method according to claim 6. **Claim 8** The predetermined criterion is When the neurons are ranked according to the number of activations of the neurons of the class from the prior probability distribution, whether the neurons are ranked within the top K neurons, where K is an integer whether the number of activations of the neurons of the class from the prior probability distribution exceeds a predetermined activation threshold The method according to any one of claims 1 to 7, which is at least one of the above.
9. The regularization activation penalty includes a penalty component calculated for each neuron not included in the group, and there is no penalty component for the neurons within the group. The method according to any one of claims 1 to 7.
10. The regularization activation penalty includes a penalty component calculated for each neuron in the layer, and the penalty component of the neurons not included in the group is greater than that of the neurons within the group. The method according to any one of claims 1 to 7.
11. In the regularization process based on the clustering, the neurons are ranked according to the number of activations of the neurons of the class from the prior probability distribution, and the penalty component of each neuron is inversely proportional to the ranking of the neuron. The method according to claim 10.
12. The method according to any one of claims 1 to 11, further comprising the step of determining the saliency of the neurons in the layer and discarding at least one neuron in the layer that is less salient than the others in the layer.
Citation Information
Patent Citations
Neural network acceleration and embedding compression system and method using activity sparsification
JP2021528796A
Neural network acceleration and embedding compression systems and methods with activation sparsification
US20190392323A1