Incorporation of decision tree in neural network

Replacing neural network layers with decision trees and quantizing inputs/outputs creates a hybrid model that addresses resource inefficiencies, reducing computational cost and memory usage while maintaining accuracy, suitable for devices with limited resources.

JP2025111557AActive Publication Date: 2025-07-30GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025068106
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-30
Estimated Expiration
2041-05-10

AI Technical Summary

Technical Problem

Large-scale neural networks consume significant computing resources and have high latency due to their complexity, requiring large memory and processor cycles, and existing techniques like sparse matrices and fused weight matrices introduce inefficiencies and memory access issues.

Method used

Replace groups of neural network layers with decision trees, using quantized inputs and outputs, and train the remaining layers with quantized training data to create a hybrid model that reduces computational cost and memory requirements.

Benefits of technology

The hybrid model reduces computational cost and memory usage, enabling efficient inference calculations on devices with limited resources by utilizing less expensive hardware, such as multiplexers, and maintains accuracy with minimal error from quantization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111557000001_ABST
    Figure 2025111557000001_ABST
Patent Text Reader

Abstract

To provide methods, systems, and apparatus, including computer programs encoded on computer storage media, for scheduling operations represented on a computation graph.SOLUTION: One of methods comprises: receiving data representing a neural network comprising a plurality of layers arranged in a sequence; selecting one or more groups of layers each comprising one or more layers adjacent to each other in the sequence; and, for each group of layers, using respective decision trees, replacing the group of layers. Each of the decision trees receives, as input, a quantized version of the inputs to respective first layers in the group and generates, as output, a quantized version of the outputs of respective last layers in the group. A tree depth of each decision tree is based at least in part on the number of layers of the group.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field This specification relates to the incorporation of decision trees in large-scale neural networks.

Background Art

[0002] Background A neural network is a machine learning model that applies one or more layers of non-linear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current values of a set of respective parameters.

[0003] More specifically, each neural network layer includes a plurality of nodes, and each layer represents a series of operations defined by the neural network. Generally, these operations can be arithmetic operations that include linear operations such as addition and multiplication, and non-linear operations such as non-linear activation functions such as the "Relu" or "Sigmoid" functions. The linear operation combines the layer input and the layer weights. The linear operation of each layer can be implemented using tensor operations in which the layer weights are represented in the form of a matrix or a tensor and the layer input of the layer is represented in the form of a vector.

[0004] Large neural networks, i.e., neural networks having many layers and a large number of parameters, have shown excellent performance in various machine learning tasks. However, these large neural networks have high latency and may consume a large amount of computing resources. For example, they require a large amount of memory to make predictions and consume a significant number of processor cycles. One of the conventional techniques for improving the computational efficiency of neural networks is to operate on a part of the weight matrix to make it a sparse matrix. A sparse matrix is a matrix in which a large number of terms are zero. SUMMARY OF THE INVENTION

[0005] Summary This specification describes a technique for incorporating a decision tree into a large neural network to generate a new machine learning model.

[0006] Generally, one innovative aspect of the subject matter described in this specification can be embodied in a method that includes the following operations. The method includes receiving data representing a neural network having a plurality of layers arranged in a sequence, and selecting one or more groups of layers from the plurality of layers, where each group of layers includes one or more layers adjacent to each other in this sequence. The method further includes generating a new machine learning model corresponding to the neural network.

[0007] Generating the new machine learning model includes, for each group of layers, selecting a respective decision tree to replace the group of layers. Each decision tree receives, as input, a quantized version of the input to the first layer in the group and generates, as output, a quantized version of the output of the last layer in the group. The tree depth of each decision tree is at least partially based on the number of layers in the group.

[0008] Other embodiments of this aspect include a corresponding computer system, an apparatus, and a computer program recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0009] Each of the foregoing embodiments and other embodiments can optionally include one or more of the following features, either alone or in combination. In particular, one embodiment includes a combination of all of the following features.

[0010] The method can further include an operation of training a new machine learning model by training at least a portion of the layers in the neural network that were not replaced by respective decision trees based on the training data of the neural network.

[0011] As discussed above, the operation of selecting each of one or more groups of layers can further include selecting each of the respective initial layers in the neural network, generating each of a plurality of candidate groups, each having the respective initial layer as the first layer in the candidate group, determining, for each of the plurality of candidate groups, each performance metric of the candidate group by measuring the performance of the corresponding new machine learning model having layers in the candidate group to be replaced by the respective decision tree, and for each of the plurality of candidate groups, selecting one of the candidate groups as a group based on each performance metric.

[0012] The selection of each initial layer can be performed by a random process or based on a sequence of neural networks. The quantized version of the input to each first layer in the group and the quantized version of the output of each last layer in the group can be generated using binary quantization or ternary quantization. Each decision tree layer that replaces a group of layers can include a GradientBoost decision tree or an AdaBoost decision tree.

[0013] The method may further include outputting a new machine learning model to a system configured to implement the new machine learning model, the system comprising one or more computing units for implementing a decision tree by one or more functions selected from an additional function, a selection function, or a switching function. That is, a new machine learning model (e.g., being trained) can be output to a system that includes one or more computing units (such as a multiplexer or an arithmetic logic unit) for implementing a decision tree without requiring more expensive computing units such as a multiply-accumulator unit (MAC) to perform multiplication.

[0014] The subject matter described herein can be implemented in certain embodiments to realize one or more of the following advantages.

[0015] The described system that implements the techniques described below can reduce the computational cost and improve the efficiency of performing inference calculations for large neural networks.

[0016] First, the described technique for replacing one or more network layers of a large neural network with a decision tree can reduce the amount of computation when the system performs inference calculations for the neural network. For example, the neural net A decision tree for replacing one or more layers of a workpiece can have only one or some layers (e.g., a tree stamp or a shallow tree). The computational cost for performing operations on tree stamps and shallow trees is much less than the computational cost required by calculating the neural network layers of a large neural network. As another example, if the decision tree is an Adaboost tree, since there is no need to perform multiplication operations for the Adaboost tree, the system can perform fewer operations than calculating a conventional neural network layer that requires both multiplication and addition, thus improving efficiency.

[0017] Second, the described technique can reduce the total size of a neural network by quantizing at least the inputs and outputs of the network layers replaced by one or more decision trees. By quantizing these inputs and outputs to reduce the number of significant digits involved in the calculation, it is possible to reduce the computational cost, especially for the inserted decision tree and at least the neural network layers adjacent to the decision tree (i.e., the layers preceding or following the decision tree). Since quantization reduces the size of the neural network, quantization can also reduce the total memory / storage requirements of the computing system. Considering this, the described technique enables devices with less memory and computing power (e.g., smartphones, tablets) to efficiently perform inference calculations on the modified neural network. In some situations, one or more hardware accelerators in a device can be customized to perform inference calculations on a specific modified neural network (i.e., one or more layers are replaced by one or more decision trees, and the inputs and outputs of one or more layers are quantized), thereby reducing the memory usage of the device, reducing power consumption, and enabling inference calculations to be performed more efficiently and quickly.

[0018] In addition, when performing inference calculations of a neural network, in many cases, the high precision provided by non-quantized inputs and outputs is not necessary to accurately detect and represent the presence or absence of important features. That is, the error introduced by quantizing the inputs and outputs to the layers replaced by decision trees is minimal compared to the improvement in efficiency.

[0019] In fact, quantized data is being used to train and calculate the neural network, but it is important for the techniques described below to quantize the inputs and outputs to be suitable for each decision tree and the remaining neural network layers. More specifically, since the system can quantize floating-point numbers and reduce the number of digits representing the floating-point numbers, the system can use fewer digits for the sign, exponent, and mantissa of the floating-point numbers. In the case of binary quantization and ternary quantization, for example, the system can map floating-point numbers to integers such as {1, -1} or {1, 0, -1}.

[0020] Also, the techniques described can efficiently train the modified neural network (i.e., the new neural network where layers are replaced by decision trees). The system can simply fine-tune the parameters in the modified neural network using at least a portion of the same training examples used to train the original neural network. Therefore, the time period required to train the modified neural network can be significantly shortened compared to training the original neural network.

[0021] Furthermore, the techniques described can reduce costs by using low-cost programmable hardware to perform the operations of the decision tree. For example, the computing system The tem can use a multiplexer (MUX) unit to calculate the operations in the decision tree instead of a multiply-accumulator unit (MAC). Those skilled in the art know that the MUX unit consumes less power and space than the MAC unit. Therefore, the computing system can include only programmable hardware units suitable for the modified neural network, rather than expensive hardware accelerators such as GPUs or TPUs. Therefore, the total cost of constructing a hardware system for performing inference calculations of the modified neural network is much less than that of the original neural network.

[0022] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 7

[0024] Detailed Description Like reference numerals and designations in the various drawings indicate like elements.

[0025] One conventional approach for enhancing the efficiency in the training and calculation of large-scale neural networks is to construct respective sparse matrices for the activation inputs and weights of each layer. However, the approach of constructing sparse matrices can give rise to new problems. For example, using sparse matrices can cause the computing system to be unable to access the same memory address during a time period, which is also called the lack of inference locality. More specifically, when constructing a sparse matrix, the computing system stores the non-zero terms of the original matrix at different memory addresses that can be physically far apart from each other. In some situations, the memory address of each term in the sparse matrix may change dynamically during the calculation. For example, a zero term may become non-zero after being added to another non-zero term. Therefore, when the computing system accesses the data stored in the memory, it may be troubled by memory latency, cache thrashing, and even cache pollution, and ultimately reduce the arithmetic efficiency of training and input processing using large-scale neural networks.

[0026] Alternatively, another prior art approach is to fuse the weight matrix with the layer logic. For example, a system adopting this fused weight technique Since the ro term can be determined, multiplication and accumulation operations are not performed on this term. Considering this, the fusion weighting technique can reduce the computational cost of large-scale neural networks by not performing calculations that result in zero output because one of the inputs is zero. However, fusion weighting comes at a cost. First, the fusion weighting technique must not be changed when the neural network is deployed on a hardware accelerator, i.e., the zero terms in the weight matrix must remain zero during calculation. However, in practice, it may be necessary to fine-tune the weight matrix of the deployed neural network using new training data. Second, when zero terms in a sparse weight matrix are shared by multiple operations, such as addition after multiplication and then another multiplication, the system still needs to perform operations for the zero terms.

[0027] The techniques described below in the specification can address the above problems. More specifically, the described techniques can efficiently perform one or more inference calculations for large-scale neural networks using quantization and decision trees. Generally, the described techniques relate to replacing one or more layers of a neural network with respective decision trees, where the input to the decision tree and the output from the decision tree are the quantized versions of the input and output of the corresponding layer.

[0028] This specification describes a system implemented as a computer program on one or more computers in one or more locations that generates a new machine learning model by replacing one or more groups of network layers of a neural network with one or more decision trees to reduce the computational cost and improve the efficiency of performing inference calculations using economical hardware.

[0029] FIG. 1 shows an exemplary neural network deployment system 100 including an exemplary neural network modification engine 120.

[0030] Generally, the neural network deployment system 100 receives, as input, data 110 representing a neural network and outputs a new trained machine learning model 180. The neural network deployment system 100 includes a neural network modification engine 120 for generating a new machine learning model 130 for the input neural network model. The new machine learning model 130 is a hybrid of the original input neural network model and one or more neural network layers that are replaced by one or more decision trees. Details of the generation of the new machine learning model 130 are described below. The neural network deployment system 100 is also configured with a training engine 140 for training the new machine learning model using training data 150, and a memory 160 configured to store and provide data for the training engine 140 (e.g., training and output data for the machine learning model, and data defining the machine learning model).

[0031] More specifically, the data 110 representing the neural network received by the deployment system 100 can include information defining the neural network, such as the operations of each layer of the neural network and the weights of each network layer.

[0032] The data 110 can also represent other aspects of the neural network. For example, the data 110 can include data representing the number of network layers in the neural network, the number of respective nodes in each layer, one or more types of inter-layer connections, such as element-wise connections or full connections, and data representing the type of each layer in the neural network, such as a pooling layer, a fully connected layer, or a SoftMax layer, etc. and so on.

[0033] Generally, the data 110 can represent a trained neural network trained on a plurality of training data 150, or a neural network that has not yet been trained.

[0034] The neural network deployment system 100 can provide the received data 110 to the neural network modification engine 120. The modification engine 120 can select one or more groups of layers of the neural network and replace each group of layers with respective decision trees to output a new machine learning model 130. The selection of one or more groups of layers will be described in more detail below.

[0035] Each decision tree for replacing each group of layers can be stored in the memory 160 and is accessible for the modification engine 120. More specifically, the data representing the decision tree and stored in the memory 160 can include data specifying the total number of nodes (e.g., the root and a plurality of leaves), the connectivity between nodes (e.g., how a leaf is connected to one or more other leaves), and one or more node operations (e.g., logical comparisons for one or more nodes).

[0036] The modification engine 120 can automatically determine each decision tree for replacing a group of layers. Alternatively, the type of decision tree for replacing each group of layers can be determined in advance by the user or by a computer program implemented by one or more computers external to the modification engine 120. The decision tree can be a GradientBoost tree or an AdaBoost tree.

[0037] GradientBoost (or gradient boost) trees can be obtained via gradient boosting, which is a machine learning method that creates a model in the form of a combination (e.g., weighted sum) of one or more simple prediction models (e.g., decision trees). AdaBoost (or Adaptive Boosting) adaptively combines one or more simple prediction models (e.g., decision trees) so that during training, large weights are assigned to simple prediction models with low performance (e.g., measures of incorrect classification, etc.), and the trained model generated by AdaBoost is more likely to accurately generate predictions when given a specific input.

[0038] The modification engine 120 can also quantize at least a portion of the neural network represented by the data 110. For example, the modification engine 120 can quantize the input to the first layer of a group of layers to be replaced by decision trees and the output from the last layer of the group of layers. Alternatively, the modification engine 120 can quantize the entire neural network such that the input and output to each layer are quantized. Further, the modification engine 120 can quantize at least a portion of the training data 150 and use the quantized training data to train each decision tree, or at least a portion of the new machine learning model 130, or both.

[0039] Quantization is the process of mapping a large set of input values to a smaller set of output values and is typically used for rounding and truncation. More specifically, quantization can be used to reduce the precision of numerical values. For example, quantization can reduce the precision of floating - point numbers from 32 bits to 8 bits. In the context of neural networks, the modification en The engine 120 can quantize the activation tensor, weight tensor, or layer output of the neural network layer with a precision of 8 bits to 4 bits, and further to 1 bit (for example, 1 bit for the mantissa). Details of quantization (for example, binary quantization and ternary quantization) are described in connection with FIGS. 2A and 2B.

[0040] The change engine 120 can provide a new machine learning model 130 to the training engine 140. Then, the training engine 140 can train at least a part of the new machine learning model 130 based on the training data 150 used to train the original neural network 110, and output the trained new machine learning model 180 as the output of the system 100. More specifically, the training engine 140 can use the training data 150 to train the remaining layers of the neural network that have not been replaced by the decision tree during a time period. Alternatively, or in addition, the training engine 140 can train the entire new machine learning model 130 based on the quantized training data, assuming the gradient of each decision tree.

[0041] The time period for training the new machine learning model 130 may be several minutes, several hours, or several days. Alternatively, or in addition, the time period can be based on the size of the training data 150 used to train at least a part of the new machine learning model 130. For example, the time period can be determined by the time required to train (for example, fine-tune) the new machine learning model 130 using 100 mini-batches or 1000 mini-batches of the training data 150.

[0042] The training engine 140 can train a decision tree that replaces one or more layers of the original neural network.

[0043] More specifically, the training engine 140 trains the decision tree before replacing the network layers of the original neural network with the decision tree. The training engine 140 can also fine-tune the decision tree in the new machine learning model 130 as described above.

[0044] To train the decision tree, the training engine 140 can use the corresponding portion of the training data 150, which is the same but quantized, for training the original neural network as the training data. Specifically, the training engine 140 trains the decision tree using each training sample. Each training example includes a quantized version of the layer input to the first layer of the group of layers and a quantized version of the layer output from the last layer of the group of layers. The group of layers is replaced by the decision tree. The quantized versions of the layer input and layer output are associated with the layer input and layer output of the corresponding training data 150 that were used to train the group of layers in the original neural network.

[0045] When training the decision tree, the training engine 140 can define a loss and adjust the node operations to reduce the loss until the loss falls below a specific criterion. The loss can be a hinge loss indicating label error, or a logarithmic loss representing information gain based on entropy theory, or any other loss suitable for training the decision tree. The training engine 140 can adjust node operations such as the respective threshold values of the logical comparison operations at one or more nodes. In some situations The training engine 140 can reduce overfitting by pruning the decision tree during training, i.e., by removing one or more leaves and all the branches and child leaves associated with the one or more leaves. The training engine 140 can leave a portion of the original training data set as a validation set for detecting and improving overfitting.

[0046] The training engine 140 used to train at least a portion of the new machine learning model 130 can include a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), or any other computing unit suitable for performing neural network operations. In particular, the new machine learning model can still have non-replaced layers, including linear operations (e.g., mainly tensor operations), and thus the training engine  140 can include more TPUs than CPUs or GPUs to facilitate the training process.

[0047] The newly trained machine learning model 180 can be used to efficiently generate inferences when an input is provided. The newly trained machine learning model 180 can be used to generate inferences using less computing power. For example, since one or more layers of the neural network layer are replaced by a shallow decision tree with fewer nodes, the newly trained machine learning model 180 requires only a memory size for storage that is less than the memory size of the original neural network. As another example, since the newly trained machine learning model is compatible with quantized inputs and outputs represented using fewer bits, it reduces the system memory bandwidth during computation. Further, since a computing unit such as a MUX unit or an arithmetic logic (ALC) unit can be used to execute operations in the decision tree, a neural network inference engine or system for executing the inference calculation of the decision tree of the new machine learning model 180 can replace a large and expensive computing unit such as a TPU or GPU with a smaller and less expensive programmable core having one or more MUX units or ALC units. Therefore, the total cost and total size of the device for generating inferences of the new machine learning model can be reduced.

[0048] FIG. 2A illustrates a portion of an exemplary new machine learning model 295 having a decision tree with a binary quantization output 225.

[0049] As shown in FIG. 2A, a portion of the original neural network 200 represented by input data 110 includes a plurality of network layers. The plurality of network layers can include a group 290 of network layers determined by the system 100 to be replaced by a corresponding decision tree, a first network layer 210 preceding the first layer of the group of network layers 290, and a second network layer 230 following the last layer of the group of network layers.

[0050] Each layer of the plurality of layers in a part of the neural network 200 has a plurality of nodes each representing a linear operation and a non - linear operation. For example, the network layer 210 includes nodes 210a - f. As another example, the network layer 230 includes nodes 230a - 230f. Also, each layer of the group 290 of network layers includes a respective number of nodes (not shown).

[0051] Some of the network layers of the neural network 200 are arranged in a sequence such that for each input to the neural network, the preceding layer generates a layer output and provides that output as a layer activation input to the subsequent layer. For example, the first layer of the group 290 of network layers receives from the network layer 210 as a layer activation input 213. As another example, the last layer of the group 290 of network layers provides the layer output 217 to the subsequent layer 230. The inputs and outputs can have respective precisions according to the computational requirements. For example, the inputs and outputs can have a floating - point format with 32 - bit precision. As another example, the inputs to the first few layers of the neural network and the outputs from the last few layers may have a higher precision than the intermediate layers.

[0052] In some embodiments, the system 100 can use different quantization methods to quantize the inputs and outputs of one or more layers of the neural network, reducing the precision. For example, the activation input 213 and the layer output 217 can be quantized using binary quantization or ternary quantization and can also have 8 - bit or even 1 - bit precision. Alternatively, the system 100 can quantize only a part of the weight input or activation input for the network layer.

[0053]

[0054] ​As described above, quantization is a process that reduces the precision of numerical values (e.g., reducing the number of bits representing the sign, exponent, and mantissa of a numerical value). Binary quantization and ternary quantization are branches of the quantization process.

[0055] Regarding binary quantization, in relation to the neural network deployment system 100 and FIG. 2A, the system 100 can quantize floating-point numbers into the binary set {1, -1}. Binary quantization can be considered as a specific quantization process in which the system 100 uses only 1 bit for the sign, 0 bits for the exponent power, and 1 bit for the significant digits of the floating-point number. More specifically, to give a few examples, using binary quantization, the floating-point weight 0.744 can be quantized to 1, and the floating-point activation input -0.21 can be quantized to -1.

[0056] Similarly, ternary quantization is an alternative to binary quantization. Although the model size becomes larger, it can improve the accuracy. In relation to FIG. 2B, the system 100 can quantize floating-point numbers into the ternary set {1, 0, -1}. More specifically, a normalized weight in floating-point format greater than 0.66 can be quantized to 1, another normalized weight smaller than 0.66 and greater than -0.66 can be quantized to 0, and another normalized weight smaller than -0.66 can be quantized to -1.

[0057] The system 100 can apply each scaling coefficient to each quantized input and output by multiplying each quantized input or output by its respective scaling coefficient. Given a specific input, the system 100 can obtain each scaling coefficient based on a measure of the approximation (or similarity) between the quantized neural network and the original non-quantized neural network. After training the modified neural network, the system 100 can determine the scaling coefficient and set the scaling coefficient as a constant when performing inference calculations for the modified neural network.

[0058] Returning to the reference of FIG. 2A, the neural network modification engine 120 can generate a new machine learning model 295 (or a new portion thereof) by replacing a group 290 of network layers of an original portion of the neural network 200 with a decision tree 220 having a tree depth based on at least the number of layers in the group 290 of network layers. The decision tree can be of any type as long as it is suitable for replacing the group of network layers. For example, the decision tree can be a GradientBoost decision tree or an AdaBoost decision tree. As long as it is suitable, it can be of any type. For example, the decision tree can be a GradientBoost decision tree or an AdaBoost decision tree.

[0059] The system 100 can perform binary quantization on at least the input 213 to the first layer of the group 290 of network layers and the output 217 from the last layer of the group 290 of network layers. Alternatively, the system 100 can perform binary quantization on the entire neural network.

[0060] The decision tree 200 can receive the binary quantization input 215 from the preceding layer 210 and output the quantization output 225 to the subsequent layer 230. More specifically, the quantization input 215 includes binary inputs (either 1 or -1) from the nodes 210a - f of the preceding layer 210. Similarly, for illustrative purposes only, the quantization output 225 can be either an output 225a representing {1} or an output 225b representing {-1}. Each output from the decision tree is quantized as either 1 or -1 and provided to the subsequent layer 230.

[0061] Before replacing the group 290 of network layers with the decision tree 220, the system 100 can train the decision tree 220 using the quantized version of the corresponding part of the training examples. More specifically, the system 100 can obtain the quantized version of the input to the first layer of the group 290 of layers and the quantized version of the output from the last layer of the group 290 of layers when the training examples of the original neural network are provided. The system 100 can set the quantized version of the input as the input to the decision tree 220, set the quantized version of the output as the output from the decision tree 220, and use the quantized versions of the input and output to train the decision 220.

[0062] FIG. 2B shows a portion of another exemplary new machine learning model 255 having a decision tree with a ternary quantization output 275.

[0063] Similar to FIG. 2A, the system 100 can replace different groups 285 of layers in a portion of the neural network 250 with different decision trees 270 and use ternary quantization to quantize the inputs and outputs of the decision trees. The quantized input 265 and the quantized output 275 include one of the value sets {-1, 0, 1}. For purposes of illustration only, the quantized output 275 can be one of an output 275a representing {1}, an output 275b representing {0}, and an output 275c representing {-1}.

[0064] For ease of illustration, in FIG. 2A the number of groups of network layers is 3 and in FIG. 2B the number is 5, but the number of groups of network layers replaced by decision trees can be any suitable value determined by the system 100. For example, the number can be 1, 10, or 50.

[0065] FIG. 3 is a flowchart of an exemplary process 300 for generating a new trained machine learning model. For convenience, process 300 is described as being executed by one or more computer systems located in one or more locations. For example, neural network deployment system 100, such as shown in FIG. 1 and appropriately programmed in accordance with this specification, can execute process 300.

[0066] The system receives (310) data representing a neural network having a plurality of layers arranged in a sequence. The data received can include, among other examples, the operations performed by each of the network layers of the neural network and the weights of each network layer. The network layers are arranged in sequence such that the layer output from a preceding layer is provided as a layer input to a subsequent layer. The system selects one or more groups of layers from the plurality of layers, where each group of layers comprises one or more layers adjacent to each other in the sequence (320). For example, the system can select three groups of layers, where the first group of layers includes only one layer, the second group of layers includes three layers, and the last group of layers includes five layers having the second-to-last layer of the neural network.

[0067] The system generates a new machine learning model corresponding to the neural network by replacing each of the one or more selected groups of layers with respective decision trees (330).

[0068]

[0069] ​Each decision tree for replacing each group of layers can have a tree depth based at least on the number of layers in the group of layers. For example, the tree depth of the decision tree can be equal to the number of layers in the group of layers. As another example, the tree depth of the decision tree is 3, and the group of layers replaced by the decision tree has five network layers.

[0070] The decision tree can include any suitable tree type. For example, the decision tree can be a GradientBoost tree or an AdaBoost tree.

[0071] The system can at least quantize the input to the first layer of the group of layers and the output from the last layer of the group of layers, and provide the quantized input and output to the respective decision trees. More specifically, for each of one or more groups of layers, each decision tree receives, as input, a quantized version of the input to each first layer in the group, and generates, as output, a quantized version of the output of each last layer in the group.

[0072] In some embodiments, the system can obtain the quantized versions of the input and output to each decision tree based on respective scaling factors. More specifically, the system can generate the quantized version of the input to the decision tree by multiplying the quantized input by the respective scaling factors. As described above, the scaling factors are obtained based at least on a measure of similarity between the quantized layer and the original layer.

[0073] The system trains a new machine learning model based on the training data of the original neural network (340). The system trains at least a portion of the layers in the original neural network that were not replaced by respective decision trees. In some embodiments, the system trains the layers of the neural network that follow one or more groups of layers of the neural network in the sequence.

[0074] In some situations, the system can train the entire new machine learning model for the original neural network using the same but quantized training samples. In some embodiments, the system can use the quantized inputs and outputs for forward propagation during training and floating - point inputs and outputs for updating weights during backpropagation. Alternatively, the system can also calculate data representing respective gradients for each decision tree in the new machine learning model. The system can select an appropriate algorithm to select a group of layers and replace the group of layers with respective decision trees. In some embodiments, the system can repeatedly select a group of layers. More specifically, one exemplary algorithm for selecting a group of layers is described below.

[0075] Assuming the neural network includes N network layers, the system indexes each layer of the neural network according to a sequence, with layer i where i ∈ [0, 1, 2, ···, N - 1].

[0076] The system sets the total number of layers (all layers) from which to select a group of layers equal to the size of the neural network (N layers).

[0077]

[0078] ​The system randomly selects layer L as the first layer of the group of layers from among all the layers. In some embodiments, the system can select the first layer of the group of layers according to a layer sequence. For example, the system can start from layer 0 as the first layer of the group of layers.

[0079] For each layer in the sequence from layer L to the last layer N, the system first checks whether the current layer has already been replaced by or belongs to the decision tree.

[0080] If it is determined that the current layer has not been replaced by or does not belong to the decision tree, the system adds the current layer to the group of layers.

[0081] The system performs quantization on the input to the initial layer L of the group of layers and the output from the current layer, and temporarily replaces the layers from layer L to the current layer with their respective decision trees. As described above, the system can perform binary quantization or ternary quantization on the input and output of the neural network layer. In some embodiments, the system quantizes all the outputs of each layer in the neural network to limit the size of the accumulator to a specific number of decision trees.

[0082] The system updates the layer information of the group of layers and measures the performance of the modified network (e.g., inference accuracy). The details of the determination and performance measurement of the group of layers are described in more detail in relation to FIG. 4.

[0083] After determining the group of layers, the system replaces the group of layers with their respective decision trees.

[0084] The system trains the remaining layers from the last layer of the group to the last layer of the neural network. More specifically, the system fine-tunes the weights of the remaining layers using the corresponding portions of the same training examples that were used to train the original neural network.

[0085] FIG. 4 is a flow diagram of an exemplary process 400 for selecting one or more groups of neural network layers. For convenience, process 400 is described as being executed by a system of one or more computers located in one or more locations. For example, a neural network deployment system 100, such as system 100 shown in FIG. 1 and appropriately programmed in accordance with this specification, can execute process 400. For example, a neural network deployment system 100, such as system 100 shown in FIG. 1 and appropriately programmed in accordance with this specification, can execute process 400.

[0086] The system selects each initial layer in the neural network (410). As described above, the system can select the initial layer of a group of layers randomly or based on a sequence of layers.

[0087] The system generates a plurality of candidate groups, each having a respective initial layer as the first layer in the candidate group (420). More specifically, the system tentatively sets the number of layers for each candidate group having the same initial layer. For example, the system generates a first candidate group having only two layers, a second candidate group having four layers, and a third candidate group having six layers. In some embodiments, the candidate groups can have a consecutive number of layers. For example, the first candidate group includes a single layer, the second candidate group includes two layers, and the third candidate group includes three layers.

[0088] For each of the plurality of candidate groups, the system determines (430) a respective performance measure of the candidate group that measures the performance of the corresponding new machine learning model in which the layers in the candidate group are replaced by respective decision trees. More specifically, the system generates a plurality of new machine learning models, each of the plurality of new machine learning models including a respective candidate group of layers that are replaced by respective decision trees. The system performs inference calculations for each of the new machine learning models using the same input data and obtains a respective performance score for each of the new machine learning models. Each performance score can be obtained based on inference accuracy.

[0089] The system selects (440), for each of the plurality of candidate groups, one of the candidate groups of layers as the group of layers to be replaced by the decision tree based on the respective performance measures. In some embodiments, the system determines the maximum performance measure among the respective performance measures and selects, as the group, the candidate group associated with the maximum performance measure from among the plurality of candidate groups. Alternatively, the system selects the candidate group that has a relatively high performance score but for which the execution of the inference calculation is the fastest.

[0090] FIG. 5 shows an example of a decision tree 500. As shown in FIG. 5, the decision tree 500 can include a plurality of nodes having a particular tree depth. The tree depth is determined based on the total number of tree layers. Each layer of the decision can include nodes representing respective node operations, such as logical comparisons or other suitable criteria.

[0091] For example, as shown in FIG. 5, the decision tree 500 can have a plurality of nodes including a root node 510 and respective non-root nodes 520, 530 in different tree layers. The root node 510 is the starting point of the decision tree 500, has no parent node, and the non-root nodes 520 and 530 each have a parent node and are also called child nodes. Generally, each node except the leaf nodes in the tree layer can have one or more branches (arrows in FIG. 5) connecting the node to respective child nodes in the next tree layer. A node having a child node is also called a non-leaf node, such as nodes 510, 520a, and 520b.

[0092] For the leaf nodes (e.g., nodes 530a, 530b, 530c, and 530d) in the deepest or last tree layer, each can represent the inference output of the decision tree 500 and has no branches connecting to other child nodes. The inference output can be a prediction, such as the probability P1 of leaf node 530a.

[0093] Referring to FIG. 5, the system (e.g., system 100 in FIG. 1) when performing an inference operation on the decision tree using input data representing coefficients n f1 , n f2 , and n f3 , each non-leaf node except the leaf nodes in the deepest tree layer has a respective node operation (e.g., a logical comparison as shown in FIG. 5). The system performs a node operation on the root node 510, obtains a logical result (e.g., true or false) from the node operation, and based on the result, can approach the corresponding child node along each branch. For example, if the result is false, the system approaches node 520a. In some embodiments, the system can use the integer 0 and the integer 1 to represent false and true. Finally, the system approaches a specific leaf node in the deepest layer of the decision tree and returns the inference represented by the specific leaf node.

[0094] FIG. 6 shows an exemplary embodiment of a decision tree using fixed function hardware 600.

[0095] As shown in FIG. 6, a system (e.g., the system 100 shown in FIG. 1) can perform inference calculations of a decision tree using fixed function hardware 600. More specifically, the system can receive an input x610 and return a final inference output 620 for the exemplary decision tree shown in FIG. 5. The decision tree 600 can include a plurality of computing units 630 and 640 to obtain inference results from different nodes and generate the final inference output 620 represented by a specific leaf node in the decision tree.

[0096] A system using fixed function hardware 600 can assign each input data to a corresponding tree node and perform respective functions at each node using respective comparators and multiplexers. As shown in FIG. 6, the system assigns the input data x ∈ [h 1, l1] to the top node (corresponding to the root node 510 in FIG. 5), assigns the input data x ∈ [h 2, l2] to the left node 630a (corresponding to the left non-leaf node 520a in the second tree layer), and assigns the input data x ∈ [h 3, l3] to the right node 630b (corresponding to the right non-leaf node 520b in the second tree layer). The system can perform operations at each node using respective comparators 640 and multiplexers 630. For example, the system performs operations at the top node using comparator 630c and multiplexer 630c. More specifically, the system receives an input x, which can be a vector of real values, and the range x ∈ [h 1,Select a part of the input having l1] and provide it to the topmost root node. The system uses a comparator 640c to compare a part of the input with a reference a1, which can represent a real-valued scalar. If the input data x is greater than or equal to a1, the system outputs a result representing "true" using a multiplexer 630c assigned to the node. For example, the multiplexer can output an integer 1 representing "true". The system can continue to perform operations at the child nodes linked along the corresponding branch of the tree at the current node. After processing the operations for the non-leaf nodes in the second-to-last tree layer, the system can output a final result 620 based on the value (e.g., probability) represented by the corresponding leaf node in the decision tree.

[0097] FIG. 7 illustrates an exemplary programmable core 700 for performing inference calculations of a decision tree. illustrates.

[0098] As shown in FIG. 7, a system (e.g., system 100 shown in FIG. 1) can use a programmable core 700 to perform inference calculations for a decision tree. The programmable core 700 can receive a tree input 710 and generate an inference output 720 for the tree input. The programmable core 700 can include a plurality of computational components, such as a MUX unit, an ALU unit, and a static random access memory (SRAM). The programmable core 700 can perform node operations simply by adding, selecting, and switching the functions configured in the computational components. Therefore, the programmable core 700 does not need to include a MAC unit for performing multiplication and addition as node operations for generating the inference output of the decision tree.

[0099] Referring to FIG. 7, to perform the operations at the root node (e.g., root node 510 in FIG. 5), the system receives an input 710 and the received input 710 is combined by the combining unit with an array (x, a, l, h, i 0 、i 1 ) that is specific to the root node and has been previously stored in the queue and modified. As described above, a represents the node reference, l and h represent the numerical range of the input x, and the indices i 0 、i 1 can each represent the tree branches connecting the current node to its respective child nodes and can be assigned integer values to represent the results of operations performed on the node operations.

[0100] The system can select the inputs combined using any unit and perform node operations at the root node. For example, the system can use each comparator to perform a comparison between the input x ∈ [h, l] and the reference a1. Accordingly, the system can assign the indices i 0 、i 1 integer values to represent the result of the comparison and direct the calculation of the decision tree to the corresponding child nodes. For example, if the system determines that x ∈ [h, l] is greater than the reference a1, the system can assign i 0 = 0, i 1 = 1, so that the system can execute an operation on the corresponding next child node along the tree branch represented by i 1 .

[0101] To perform an operation on the next child node, the system first applies a switching unit to determine whether the next child node is a non-leaf node or a leaf node based on the numbering criterion N. To make this determination, the system can number each node using respective tags K (e.g., integers) and determine the type of the next child node based on a predetermined numbering criterion. For example, in relation to FIG. 5, the system can tag nodes 510, 520a, and 520b as nodes 0, node 1, and node 2, and can tag leaf nodes 530a, 530b, 530c, and 530d as nodes 3, node 4, node 5, and node 6. The system can set the numbering criterion N = 3 such that nodes with tag K < 3 are non-leaf nodes and nodes with tag K >= 3 are leaf nodes.

[0102] In response to a determination that the next child node is a non-leaf node, the system can update each array (x, a, l, h, i 0 , i 1 ) of the next child node from the data stored in the non-leaf SRAM. Similarly, in response to a determination that the next child node is a leaf node, the system can provide the final output associated with the leaf node stored in the leaf SRAM.

[0103] For parallel computing, the system can employ a fork function to store a portion of the input data of the nodes in the queue and respective arrays identifying the operations and types of the nodes in the SRAM. During parallel computing, the system can automatically determine when to modify the input data using each array, using one or more combining units based on the respective latencies observed for each computing unit.

[0104] The embodiments referred to in this specification provide an improved method for training a neural network. The neural network can be configured to receive any kind of digital data input and generate any kind of score, classification, or regression output based on that input. The input data items can comprise image data (including video data here), audio data, or text data, such as words or word parts of natural language (or their representations, such as embeddings). The input data items can comprise sequential data, such as a sequence of data samples representing digitized audio, or an image represented as a sequence of pixels, or a video represented as a sequence of images, or a sequence representing a sequence of words in natural language. As used herein, "image" includes, for example, LIDAR images.

[0105] In some embodiments, the neural network output can comprise a feature representation, which can then be further processed to generate a system output. For example, the system output can comprise a classification output for classifying the input data item into one of a plurality of categories, such as categories of images, videos, or audio (e.g., data representing the likelihood that the input data item or an object / element of the input data item belongs to the category), or a segmentation output for segmenting regions of the input data item into objects or actions represented, for example, in an image or video. Or, the system output can be an action selection output in a reinforcement learning system.

[0106] In some other embodiments, the network output may comprise another data item of the same type or a different type. For example, the input data item may be an image, audio, or text, and the output data item may be a modified version of an image, audio, or text, e.g., changing the style, content, properties, pose, etc. of the input data item, or one or more objects or elements within the input data item, or filling in (missing) parts of the input data item, or predicting another version of the data item, or an extension of a video or audio data item, or providing an upsampled (or downsampled) version of the input data item. For example, the input data item may be a textual representation in a first language, and the output data item may be a translation of the text into another language, or a score for a translation of the text into another language. In another example, an input image may be converted into a video, or a wireframe model, or a CAD model, or a 2D input image may be converted into 3D or vice versa. Or, the input data item may comprise features derived from an oral utterance or a sequence of oral utterances, and the network system output may comprise a score for each of a series of text segments, where each score represents the estimated likelihood that a part of the text is the correct replacement based on the features. In another example, the input data item may be an image, audio, or text, and the output data item may be a representation of the input data item in a different format. For example, a neural network may convert text to speech (in the case of speech recognition) or vice versa, or convert an image (or video) to text (e.g., for captions). When generating an output comprising continuous data, the neural network may include one or more convolutions, e.g., dilated convolutional layers.

[0107] In some other embodiments, the network output may comprise an output for selecting actions to be performed by an agent, such as a robot or other mechanical agent, in an environment such as a real-world environment or a simulation of a real-world environment.

[0108] In some embodiments, a neural network is configured to receive input data items and process the input data items to generate a feature representation of the input data items according to network parameters. Generally, a feature representation of a data item is an ordered set of numerical values, such as a vector, that represents the data item as a point in a multi-dimensional feature space. In other words, each feature representation may include numerical values for each of a plurality of features of the input data item. As described above, a neural network can be configured to receive any kind of digital data input and generate a feature representation from that input. For example, the input data item, which may also be referred to as a network input, can be an image, a portion of a document, a text sequence, audio data, medical data, and the like.

[0109] Once trained, the feature representation can be provided as input to another system for use, for example, in performing a machine learning task on the network input. Exemplary tasks can include feature-based search, clustering, near-duplicate detection, verification, feature matching, domain adaptation, video-based semi-supervised learning, and in the case of video, can include, for example, object tracking over video frames, gesture recognition of gestures being performed by entities depicted in the video.

[0110] When the input to the neural network is an image or features extracted from an image, the output generated by the neural network for a given image can be scores for each of a series of object categories, where each score represents the estimated likelihood that the image contains an image of an object belonging to the category. More specifically, each of the input image or the features extracted from the input image can include one or more pixels each having its respective intensity value. The neural network is configured to process each intensity value of the input image or the features extracted from the image to generate predictions such as, for example, image classification, image recognition, or image segmentation.

[0111] As another example, when the input to the neural network is an Internet resource (e.g., a web page), a document, or a part of a document, or features extracted from an Internet resource, a document, or a part of a document, the output generated by the neural network for a given Internet resource, document, or part of a document can be scores for each of a series of topics, where each score represents the estimated likelihood that the Internet resource, document, or part of a document pertains to that topic.

[0112] As another example, when the input to the neural network is features of the impression context of a particular advertisement, the output generated by the neural network can be a score representing the estimated likelihood that the particular advertisement will be clicked.

[0113] As another example, when the input to the neural network is features of recommendations personalized for a user, such as features characterizing the context of the recommendations, such as features characterizing previous actions taken by the user, the output generated by the neural network can be scores for each of a series of content items, where each score represents the estimated likelihood that the user will respond positively to the recommended content item.

[0114] As another example, if the input to the neural network is a sequence of text in one language, the output generated by the neural network can be scores for each of a series of text segments in another language, where each score represents the estimated likelihood that a portion of the text in the other language is an appropriate translation of the input text into the other language.

[0115] As another example, if the input to the neural network is a sequence representing an oral utterance, the output generated by the neural network can be scores for each of a series of text segments, where each score represents the estimated likelihood that a portion of the text is a correct transcription of the utterance.

[0116] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of their combinations. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random access memory device, or a serial access memory device, or one or more of their combinations. Alternatively, or in addition, the program instructions may be encoded in an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiver device for execution by a data processing apparatus.

[0117] The term "data processing apparatus" refers to data processing hardware and includes, by way of example, any kind of device, apparatus, and machine for processing data, including a programmable processor, a computer, or multiple processors or computers. The apparatus may also or further include dedicated logic circuitry such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The apparatus may optionally include, in addition to the hardware, code for creating an execution environment for computer programs, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0118] A computer program, which may also or may be described as a program, software, software application, app, module, software module, script, or code, can be described in any form of programming language including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, such as as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program can be stored, for example, in one or more scripts stored in a markup language document, as part of a file holding other programs or data, in a single file dedicated to the program in question, or in multiple coordinated files, such as files storing one or more modules, subprograms, or portions of code. The computer program can be deployed to be executed on one computer or on computers arranged at one site or distributed across multiple sites and interconnected by a data communication network.

[0119] As used herein, the term "database" is used broadly to refer to any collection of data, where the data need not be structured in any particular way, or structured at all, and can be stored in a memory device in one or more locations. Thus, for example, an index database can include multiple collections of data, each of which can be organized and accessed in a different manner.

[0120] Similarly, as used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components and installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and executed on the same one or more computers.

[0121] The processes and logical flows described herein can be executed by one or more programmable computers that perform operations on input data and generate output, by executing one or more computer programs. The processes and logical flows can also be executed by, for example, a dedicated logic circuit configuration such as an FPGA or ASIC, or by a combination of a dedicated logic circuit configuration and one or more programmed computers.

[0122] A computer suitable for the execution of a computer program can be based on a general-purpose or special-purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit receives instructions and data from a read-only memory, or a random access memory, or both. The essential elements of a computer are a central processing unit for executing or implementing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be complemented by, or incorporated in, a dedicated logic circuit configuration. Generally, a computer also includes, or is operably coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, for receiving data from them, or for transferring data to them, or both. However, a computer need not have such devices. Further, a computer can be incorporated in another device, such as, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as, for example, a universal serial bus (USB) flash drive.

[0123] Computer-readable media suitable for storing computer program instructions and data include, by way of example, any form of non-volatile memory, media, and memory devices, such as semiconductor memory devices, for example, EPROM, EEPROM, and flash memory devices, magnetic disks, for example, internal hard disks or removable disks, magneto-optical disks, and CD ROM disks and DVD-ROM disks.

[0124] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can be used to provide interaction with the user, for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user can be in any form, including acoustic, voice, or tactile input. Additionally, the computer can interact with the user by sending documents to, and receiving documents from, the device used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser, and the computer can also interact with the user by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user as a reply. By sending documents to, and receiving documents from, the device, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser, or by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user as a reply, the computer can interact with the user.

[0125] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing general and computationally intensive parts of machine learning training or production, i.e., inference, workload.

[0126] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.

[0127] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes, for example, as a data server, back-end components, or includes middleware components, such as, for example, an application server, or includes front-end components, such as, for example, a graphical user interface, a web browser, or a client computer having an app through which a user can interact with embodiments of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as, for example, a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as, for example, the Internet.

[0128] The computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between a client and a server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, the server transmits data, such as an HTML page, to a user device for the purpose of displaying data to a user who interacts with a device acting as a client and receiving user input from the user. Data generated at the user device, such as the result of a user interaction, can be received at the server from the device.

[0129] This specification includes details of many specific embodiments, but these should not be construed as limitations on the scope of any invention or what can be claimed, but rather as descriptions of features that may be particular to specific embodiments of a particular invention. In this specification, specific features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented individually or in any suitable partial combination in multiple embodiments. Further, features may be described above as operating in a particular combination and may initially be claimed as such, but in some cases, one or more features from the claimed combination may be excluded, and the claimed combination may be directed to a partial combination or a variation of a partial combination.

[0130] Similarly, operations are shown in the drawings and recited in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed, for the desired results to be achieved. Multitasking and parallel processing may be advantageous in certain situations. Further, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the program components and systems described are generally understood to be able to be integrated together into a single software product or packaged into multiple software products.

[0131] Specific embodiments of the subject matter are described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims can be performed in a different order and still achieve the desired result. As one example, the processes illustrated in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method implemented by one or more computers, the method comprising: receiving data representing a neural network comprising a plurality of layers arranged in a sequence; selecting one or more groups of layers from the plurality of layers, each group of layers comprising one or more layers adjacent to each other in the sequence; the method further comprising: generating a new machine learning model corresponding to the neural network, the generating of the new machine learning model comprising: for each group of layers, selecting a respective decision tree to replace the group of layers, each respective decision tree receiving, as input, a quantized version of an input to a first layer in the group and generating, as output, a quantized version of an output of a last layer in the group, and a tree depth of each respective decision tree being at least partially based on a number of layers in the group.

2. The method of claim 1, further comprising training the new machine learning model by training at least a portion of the layers in the neural network that were not replaced by respective decision trees based on training data of the neural network.

3. Training at least a portion of the layers in the neural network that were not replaced by respective decision trees comprises: The method of claim 2, comprising training the layers of the neural network subsequent to the one or more groups of layers of the neural network according to the sequence.

4. Selecting each of the one or more groups of layers comprises: selecting each initial layer in the neural network; generating a plurality of respective candidate groups, each having the respective initial layer as a first layer in a candidate group; for each of the plurality of respective candidate groups, determining a respective performance metric of the candidate group by measuring a performance of a corresponding new machine learning model having the layers in the candidate group that are to be replaced by respective decision trees. The method according to any one of claims 1 to 3, comprising, for each of the plurality of candidate groups, selecting one of the candidate groups as the group based on each performance metric.

5. In order to generate each group of layers, selecting each of the initial layers in the neural network, The method according to claim 4, comprising selecting each of the initial layers by a random process or based on the sequence of the neural network.

6. Selecting one of the candidate groups as the group based on each performance metric, Determining the maximum performance metric from among each of the performance metrics, From each of the plurality of candidate groups, the candidate group associated with the maximum performance metric The method according to claim 4 or claim 5, comprising selecting as the group.

7. The quantized version of the input to each first layer in the group and the quantized version of the output of each last layer in the group are generated using binary quantization or ternary quantization. The method according to any one of claims 1 to 6.

8. Each decision tree layer for replacing the group of layers comprises a GradientBoost decision tree or an AdaBoost decision tree. The method according to any one of claims 1 to 7.

9. Each layer comprises a respective set of weights, and the method The method according to any one of claims 1 to 8, further comprising quantizing at least a portion of the weights associated with the layer for each layer in the neural network that is not in the one or more groups of layers.

10. The tree depth of each of the decision trees is equal to the number of layers in the group. The method according to any one of claims 1 to 9.

11. The quantization version of the input to each first layer in the group, or the quantization version of the output of each last layer in the group, is generated by one or more scaling factors, according to the method of any one of claims 1 to 10.

12. The neural network represented by the received data is a neural network that has been first trained by a training dataset, each decision tree for replacing each group of layers is trained based on a respective part of the quantization version of the training dataset, each training sample of the respective part of the quantization version of the training dataset comprises (i) the quantization version of the layer input to the first layer of the group and (ii) the quantization version of the layer output from the last layer of the group, according to the method of any one of claims 1 to 11.

13. The method further includes outputting the new machine learning model to a system configured to implement the new machine learning model, the system comprising one or more computing units for implementing the decision tree by one or more functions selected from an additional function, a selection function, or a switching function, according to the method of any one of claims 1 to 12.

14. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 13.

15. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Computation graph processing

    JP2018538607A

  • Machine learning classification on hardware accelerators with stacked memory

    US20160379137A1

  • Methods, processing engines, and microprocessors for classifying data according to decision trees

    US20190005396A1