Optimizing algorithms for hardware devices
The optimized tensor decomposition through neural network generation solves the problem that multilinear algorithms in the prior art are difficult to adapt to specific hardware devices, and optimizes hardware performance, including reducing computing complexity and running time.
Patent Information
- Application Number
- CN202380071037.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-03
- Filing Date
- 2023-10-02
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has difficulty in automatically optimizing multilinear algorithms to adapt to the architecture of specific hardware devices, resulting in the inadequate utilization of hardware performance.
By generating optimized tensor decomposition using neural networks, reparameterize target tensors to generate more efficient multilinear algorithms. The neural network processes the tensor state according to network parameters, generates a strategy for modifying the tensor, and determines the target network output through a tree search to optimize the execution of the algorithm on the target hardware device.
Optimization of multilinear algorithms executed on specific hardware devices is realized, and performance indicators of hardware devices such as computational complexity, runtime, power consumption, and cache performance are improved.
Smart Images

Figure CN120112901A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Application No. 17 / 959,210 filed on October 3, 2022. The disclosure of the prior application is considered part of the disclosure of the present application and is incorporated by reference into the disclosure of the present application. Background Art
[0002] This specification relates to optimizing multilinear algorithms for hardware devices using neural networks.
[0003] Multilinear mapping, especially bilinear mapping (such as matrix multiplication), is a fundamental computing task performed by various hardware devices, such as central processing units (CPUs), graphics processing units (GPUs), tensor processing units (TPUs), application-specific integrated circuits (ASICs), etc.
[0004] A neural network is a machine learning model that uses one or more layers of nonlinear units to predict the output of a received input. In addition to the output layer, some neural networks also include one or more hidden layers. The output of each hidden layer is used as the input to the next layer in the network (i.e., the next hidden layer or output layer). Each layer of the network generates an output from the received input based on the current value input of a corresponding set of parameters. Summary of the invention
[0005] This specification describes a method executed by one or more computers for obtaining an optimized algorithm that (i) is functionally equivalent to a target algorithm and (ii) optimizes one or more target properties when executed on a target set of one or more hardware devices.
[0006] The method includes: initializing a target tensor representing a target algorithm; generating a tensor decomposition of the target tensor parameterized for a candidate algorithm using a neural network having a plurality of network parameters, wherein the neural network is configured to receive a state of the tensor as an input and process the input according to the network parameters to generate a network output including a strategy for applying a modification to the tensor, wherein generating the tensor decomposition includes: for each step in a sequence of steps: obtaining a current state of the target tensor; determining a target network output for the current state by performing a tree search of a state tree having nodes representing the state of the target tensor starting from a root node representing the current state, wherein the tree search is guided by the neural network according to the network parameters; applying the modification to the target tensor using the target network output for the current state; and determining whether to terminate the sequence based at least in part on whether the target tensor is equal to a zero tensor after applying the modification; and generating the tensor decomposition based on the modification applied to the target tensor at each step in the sequence of steps; generating a target attribute value for each of the target attributes when the candidate algorithm is executed on a target set of hardware devices; determining a benchmark score for the tensor decomposition based on the target attribute values of the candidate algorithm score); generating training examples according to the tensor decomposition and the benchmark score; and storing the training examples in a training data repository for use in updating network parameters of the neural network.
[0007] The method may include selecting a particular candidate algorithm as the optimized algorithm based on a benchmark score for the particular candidate algorithm generated by using the neural network.
[0008] In some implementations, a policy may define a probability distribution over possible rank-one terms to be subtracted from a tensor.
[0009] In some implementations, the network output may include a return output defining an estimated return resulting from the tensor being in that state. The estimated return may be an estimate of an expected base score for the tensor decomposition. The estimated return may be an estimate of an expected rank for the tensor decomposition.
[0010] In some implementations, performing a tree search may include: based on action scores assigned to edges connecting nodes of a state tree, traversing the edges until a leaf node is reached, wherein the edge represents a possible modification to be applied to a target tensor; processing the state of the target tensor represented by the leaf node using a neural network according to network parameters to generate a network output for the leaf node; expanding the state tree at the leaf node using a strategy for the leaf node; and for each edge of the state tree that has been traversed: incrementing a visit count of the edge; and updating the action score for the edge based on a value constructed based on a return output for the leaf node.
[0011] In some further implementations, performing a tree search may include: storing in a permutation table one or more nodes encountered during the tree search; while traversing the edges, determining that a newly encountered node represents the same state of the target tensor as a previously encountered node stored in the permutation table; and in response, replacing the newly encountered node with the previously encountered node.
[0012] In some implementations, determining the target network output according to the tree search may include: if the total access count of all edges of the root node is greater than the maximum total access count, smoothing the access counts of the edges of the root node using an adaptive temperature scheme.
[0013] In some implementations, determining the target network output according to the tree search can include ignoring edges of the root node that have action scores lower than the action score of an edge of the root node that has a highest visit count.
[0014] In some implementations, initializing the target tensor may include performing a change of basis on the target tensor, and generating the tensor decomposition may include performing an inverse basis change on the tensor decomposition.
[0015] The method may include: generating a set of tensor decompositions of one or more synthetic tensors, wherein the synthetic tensors are randomly initialized tensors; generating a set of synthetic training examples based on the set of tensor decompositions of the synthetic tensors; and storing the set of synthetic training examples in a training data repository for use in updating network parameters of a neural network.
[0016] The method may include: retrieving a training state of a tensor associated with a training target from a training data repository; processing the training state using a neural network to generate a training network output based on network parameters; determining a gradient of an objective function with respect to the network parameters that promotes the training network output to satisfy the training target for the training state; and updating the network parameters based on the gradient.
[0017] In some implementations, the optimized algorithm can be recursively executed on the target set of hardware devices.
[0018] In some implementations, the target algorithm may compute a bilinear mapping. The bilinear mapping may be a matrix multiplication.
[0019] In some implementations, the target attributes may include at least one of: the computational complexity of the optimized algorithm, the runtime of the target group hardware device when executing the optimized algorithm, the cache performance of the target group hardware device when executing the optimized algorithm, the locality of reference when the optimized algorithm is executed on the target group hardware device, or the power consumption of the target group hardware device when executing the optimized algorithm.
[0020] In some implementations, the target attributes may include one or more of: runtime of the target group hardware device when executing the optimized algorithm, cache performance of the target group hardware device when executing the optimized algorithm, locality of reference of the optimized algorithm when executed on the target group hardware device, or power consumption of the target group hardware device when executing the optimized algorithm (but not including computational complexity of the optimized algorithm).
[0021] In some implementations, the target set of hardware devices can be a simulation of a set of hardware devices.
[0022] In some implementations, the target group hardware device may include at least one of: a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), or an application specific integrated circuit (ASIC).
[0023] The method may include: receiving a new input; and executing a target algorithm on the new input by executing an optimized algorithm on a target set of hardware devices.
[0024] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0025] The systems and methods disclosed in this specification can use neural networks to automatically obtain optimized multilinear algorithms, such as algorithms optimized for a particular set of hardware devices. The result is superior performance of hardware devices over existing algorithms, as measured by any combination of performance metrics, such as reduced computational complexity, reduced runtime, reduced power consumption, improved cache metrics, increased reference locality, etc. Automatically generating efficient algorithms using machine learning techniques disclosed herein can go beyond the scope of human intuition and outperform human-designed algorithms.
[0026] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 An example optimization system is shown.
[0028] Figure 2 is a flowchart of an example process for generating a tensor decomposition.
[0029] Figure 3 is a flow chart of an example process for performing a tree search.
[0030] Figure 4is a flowchart of an example process for generating training examples from a tensor decomposition.
[0031] Figure 5 is a flow chart of an example process for updating network parameter values of a neural network.
[0032] Figure 6 is an example of a matrix multiplication algorithm parameterized by tensor decomposition.
[0033] 7A to 7C is experimental data showing matrix multiplication algorithms optimized for various hardware devices.
[0034] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION
[0035] Figure 1 An example optimization system 100 is shown that can automatically generate optimized algorithms for a set of hardware devices. The optimization system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
[0036] The system 100 receives a target algorithm to be optimized for a target set of one or more hardware devices, such as one or more of the following: a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), an application-specific integrated circuit (ASIC), etc. The target algorithm can describe any set of operations to be performed by the hardware device, where the only stipulation is that the target algorithm can be represented as a tensor. That is, the target algorithm can compute any multilinear mapping ,in and is a finite-dimensional vector space, and is a function that implements the target algorithm.
[0037] Multilinear mapping, especially bilinear mapping, is a basic computing task that can manipulate large amounts of data. Optimizing such ubiquitous tasks allows hardware devices to fully utilize their computing resources. For example, GPUs and TPUs can efficiently perform bilinear mappings, such as structured matrix multiplication, polynomial multiplication, convolution, custom machine learning operations, etc., due to their parallel processing architecture. However, determining the appropriate algorithm for their specific architecture usually relies on human intuition and is therefore suboptimal. In contrast, system 100 can automatically use a neural network to generate the best algorithm, i.e., an algorithm optimized for a specific hardware architecture, which is trained via reinforcement learning on one or more performance targets (e.g., as a weighted sum) of a hardware device.
[0038] When executing the optimized algorithm on a hardware device, the performance goal can be characterized by any number of target attributes. The target attribute can include an intrinsic measure of the computational complexity of the optimized algorithm. Alternatively or in addition, the target attribute can include one or more parameters that are specific to the set of hardware devices (e.g., parameters that depend on the architecture of the hardware devices) and indicate the performance of the optimized algorithm when implemented on the hardware devices, such as running time, cache hit indicators, reference locality, power consumption, etc. As will be described in more detail, the system 100 can benchmark these attributes on the fly to train the neural network towards generating an algorithm that optimizes any combination of the target attributes in a manner specific to the set of hardware devices.
[0039] In addition, the system 100 can generate an optimized algorithm that is functionally equivalent to the target algorithm. In this case, "functionally equivalent" means that for every possible input, the optimized algorithm calculates a corresponding output that is exactly the same as the target algorithm. In other words, the optimized algorithm is provably correct and is not close to the target algorithm. This can be a particularly advantageous quality when considering basic computing tasks because the optimized algorithm will not accumulate errors when executed on a hardware device. Note that the term "identical" is used to refer to the specific number of significant digits (e.g., the number of significant digits as the number of digits in the multiplied value) being exactly the same, because in hardware implementations of multiplication operations, rounding errors may occur when two floating point values are multiplied. If the target algorithm includes a truncation step in order to generate a function that is provably correct, the optimized algorithm can be used to generate a function that is provably correct. If the output is encoded to a specific number of significant digits, an optimized algorithm can do so as well.
[0040] For clarity, a brief review of multilinear mappings is outlined below to demonstrate the target tensor How the target algorithm can be represented. Multilinear functions The vector space The variables in are mapped to the vector space In addition, the function In each variable has linear properties, so
[0041] Here, and is a scalar. A multilinear map of a single variable is a linear map. A multilinear map of two variables is a bilinear map (e.g., matrix multiplication), and so on.
[0042] To implement the target algorithm, the hardware device executes a function on a vector basis This base specifies the hardware device in the implementation How to index variables when Although any vector basis can be implemented, in most cases vector spaces are represented by linear arrays because this is often how computing devices index data structures. In particular, each vector space The variables can be represented as column vectors ,in is a scalar element of the variable. However, the element itself can be associated with matrices or, more generally, tensors. For example, It can contain elements of matrices in either row-major order or column-major order. This basis provides a very convenient and powerful framework to implement arbitrary multilinear algorithms and is therefore adopted in the description of the algorithms in this paper.
[0043] Therefore, for each vector space Choose a "canonical" basis of orthogonal unit vectors , and for Select Base . A unit vector can be represented as a column vector with a single entry of 1 and the rest of the entries being 0,
[0044] Therefore, the variables of each vector space can be expanded with respect to the unit vector as and . Note that the number of unit vectors in each basis depends on the dimensionality of their corresponding vector space. Each vector space can have any number of dimensions. Using the aforementioned linear properties, the function can be parameterized in this basis as,
[0045] In this case, parameterization means that and Execute function . Tensor components The target tensor is defined by a collection of The target tensor fully specify a multilinear function in a chosen basis In particular, the amount Independent of And with Related.
[0046] Here, Represents the dot product. The target tensor Usually with The order corresponding to different indexes ,in is the number of variables entering the multilinear map. The target tensor Size Depending on each vector space Dimension ,as well as Dimension .
[0047] Output vector Elements It is succinctly expressed as
[0048] Each is computed as the tensor components Scaling elements The hardware device can be configured to sum the products of Execution throughout The target algorithm is implemented by using nested loops to compute each product and then the corresponding sum of the products. All such computations are compactly represented in the target tensor Therefore, the target tensor is the tensor representation of the target algorithm.
[0049] As a concrete example of a bilinear map, consider a map with size The matrix and Matrix multiplication to produce matrices of the same size To implement an algorithm for matrix multiplication, the elements of each matrix can be stored in a linear array, for example, in row-major or column-major order. , and .in this case, Is to form a dimension A set of orthogonal bases of the vector space unit vectors. Therefore, the element , and The matrices corresponding to them , and Each element in corresponds to a bilinear function. By and As input and The corresponding results are calculated in to implement the algorithm for square matrix multiplication. Using the method described above, each can be expressed as and A linear combination of the multiplications between .
[0050] Tensor Components The collection has , and the definition has size of order The "matrix multiplication tensor" By parameterizing, The weight is independent of And with the bilinear function Related.
[0051] Therefore, the matrix multiplication tensor is the tensor representation of the matrix multiplication algorithm. A similar procedure applies to any bilinear map or, in general, to multilinear maps.
[0052] However, a particular algorithm represented by a particular tensor may be suboptimal. For example, matrix multiplication of a tensor Has cubic computational complexity because it involves and between the elements This is an unnecessary amount of multiplication. For example, system 100 can construct an optimized algorithm (eg, Strassen's algorithm) that involves only 7 multiplications instead of 8. Reducing the number of multiplications in an algorithm can be an effective means of improving the performance of a set of hardware devices because multiplications are typically more computationally expensive than additions.
[0053] The system 100 decomposes the target tensor by means of tensor decomposition. And thus reparameterize the target algorithm to generate an optimized algorithm. System 100 generates an optimized algorithm by reparameterizing the target tensor Factoring 1 rank item The tensor decomposition is obtained by linear combination of
[0054] As the name implies, the rank of each rank 1 term is , which imposes constraints on their permissible forms. A tensor decomposition with rank 1 terms is called a rank Therefore, the target tensor The rank of is called the most or Note that the rank of a tensor should not be related to its order is confusing, because rank and order are two different but related quantities. The rank of a tensor depends on its size and refers to the number of linearly independent directions it can represent in the tensor product space. In other words, rank is a measure of the "non-degenerate nature" of a tensor. For example, a tensor with size of A tensor (i.e., matrix) of order is represented by Defines the maximum rank. With size of The tensor of order is given by The maximum rank defined, and so on.
[0055] Rank 1 represents a single linearly independent direction in the tensor product space, and can therefore always be expressed as The outer (tensor) product of vector factors ,
[0056] Each factor and express and A single direction in the vector space of . The combination is the tensor product space A single linearly independent direction in .
[0057] System 100 relative to the set factor in the tensor decomposition to parameterize the optimized algorithm. That is, the output vector Each element of can be calculated as
[0058] in
[0059] The hardware device can be Cycle through Index to calculate each product And then calculate the corresponding sum of the products to implement the optimized algorithm. In this case, by The loop replacement is done by , thereby achieving a more computationally efficient algorithm. As can be seen in the above formula, the rank of the decomposition Managing computational complexity , that is, the product , which means that low-rank decomposition is usually desirable. The specific form of controls the density of the decomposition, that is, the number of additions. The target tensor with a low-rank decomposition Meaning that the target algorithm has many redundant operations (it is highly degenerate). System 100 aims to eliminate such redundancy.
[0060] Figure 6 An example of an optimized matrix multiplication algorithm that can be obtained by the optimization system 100 is shown. Here, the matrix multiplication tensor By factor , and is reparameterized. Figure 6 As can be seen, the number of multiplications between the elements of the input matrices A and B is given by Therefore, The decomposition of corresponds to a square matrix multiplication algorithm with reduced computational complexity. Figure 6 The algorithm of the form shown in can be used to multiply block matrices, for example, Matrix Submatrices of size can be used The algorithm multiplies. In addition, Figure 6 The algorithm can be executed recursively on a set of hardware devices to multiply matrices of arbitrary size, that is, A matrix of size can be computed with complexity Therefore, it is possible to efficiently multiply much larger matrices by leveraging a low-rank decomposition of a smaller-sized matrix multiplication tensor.
[0061] Figure 1 shows how the system 100 uses the decomposition engine 100A to generate The system 100 provides an autonomous means of overcoming the tensor decomposition problem and optimizing a general multilinear algorithm for any set of hardware devices.
[0062] After finding an optimized algorithm for a target group of hardware devices, system 100 can replace the target algorithm with the optimized algorithm. That is, all new inputs related to the operations performed by the target algorithm received by the group of hardware devices can be performed by the optimized algorithm instead. The result is the superior performance of the device, as measured by any combination of one or more target attributes (that is, a combination defined by the weighted sum of weights), such as reduced complexity, reduced running time, reduced power consumption, improved cache indicators, etc. Since multilinear mapping (e.g., einsum operation) becomes a bottleneck for many systems (e.g., training large language models), these improvements that are superior to existing algorithms can cause significant impacts. As the number of specialized hardware devices surges (as expected in the foreseeable future), such customized algorithms will become more and more common. If both algorithm and hardware design are jointly optimized, system 100 can provide even higher gains.
[0063] Similarly, the system 100 may also specify data (e.g., a set of parameterized factors) of the optimized algorithm. ) to a different set of hardware devices to implement the optimized algorithm in place of the target algorithm. For example, if system 100 obtains an optimized matrix multiplication algorithm for a GPU in a GPU cluster, system 100 can send data specifying the optimized algorithm to all GPUs in the cluster. The performance enhancement achieved on one device is scalable because the optimized algorithm can be implemented on any other device with the same (or similar) processing architecture.
[0064] Reference is now made to the neural network employed by system 100. By a set of network parameters 120 parameterized neural networks are configured to receive tensors takes the state of as input and processes it to generate the network output The network output includes a method for applying modifications (i.e. actions) to to produce the modified tensor Strategy The policy defines a set of actions that can be applied as a modification to In order to produce Action In general, this strategy Provides possible rank 1 terms that can be combined of Possible factors Therefore, the action and The choice of factors and Subtracting the resulting rank 1 term corresponds to .
[0065] Network output can also include return output . Generally speaking, the output returned is Provides a state for the tensor However, the return output may alternatively be a regressed value, for example, an expected benchmark score corresponding to the mean of the distribution. Having said that, modeling the distribution generally provides improved performance and flexibility for the system 100 because a high degree of variability in the benchmark scores can be captured.
[0066] The benchmark score 110 represents a performance target when executing the candidate algorithm 108 parameterized by the tensor decomposition 106 on a target set of hardware devices. As will be discussed in more detail below, the benchmark score 110 is a numerical value that assigns a relative score to the tensor decomposition 106.
[0067] In some implementations, the output is returned The distribution over the expected rank of the tensor decomposition 106 is provided. In this case, the expected rank serves as a proxy metric for the benchmark score 110 without actually executing the candidate algorithm 108 on a hardware device. This is because the rank of the tensor decomposition 106 is the computational complexity of candidate algorithm 108 Thus, when optimizing only complexity, system 100 does not need to refer to any particular set of hardware devices and will tend to converge to a low-rank tensor decomposition.
[0068] In general, the neural network included in the optimization system 100 can have any suitable architecture to achieve its desired functionality. In particular, the neural network can include any suitable neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) in any suitable number (e.g., 5, 10, or 100 layers) and arranged in any suitable configuration (e.g., as a linear sequence of layers).
[0069] A high-level overview of the optimization system 100 is outlined below. Figures 2 to 5 Detail the details of each process involved.
[0070] Decomposition Engine (DE) 100A converts tensors The initial state is set to the target tensor , then executes the sequence of steps to generate the tensor decomposition 106. Each step The tensor then passes through the state 102a-102n. At each step in the sequence, DE 100A selects an action And the current status according to Update. Action With factor The choice corresponds to that of , and the rank 1 term is obtained by subtracting To modify the current state To determine the appropriate action DE 100A performs a tree search 104a-104n of the state tree at each step, such as Monte Carlo Tree Search (MCTS). The neural network is based on the network parameters 120 Guided tree search to determine the appropriate action for this step .
[0071] DE 100A at every step Apply actions , until the tensor is at the final step The zero tensor is reached After that, all factors selected by DE 100A is incorporated into the tensor decomposition 106 that parameterizes the candidate algorithm 108. The sequence The total number of steps in is equal to the rank of the decomposition 106 and the computational complexity of the candidate algorithm 108 In addition, it can be seen that the sequence of factors selected satisfies , which guarantees the correctness of the candidate algorithm 108. To avoid a potentially infinite sequence, the system 100 can limit the number of steps to a maximum, after which the sequence is terminated. For example, the maximum value can be about the target tensor In some cases, the system 100 assigns a penalty score to incomplete decompositions in order to train the neural network away from such action sequences.
[0072] The system 100 benchmarks the candidate algorithm 108 on the fly by potentially executing the candidate algorithm 108 multiple times on the set of hardware devices. For example, the system 100 may execute the candidate algorithm 108 multiple times with randomly initialized inputs and then average the performance metrics. Alternatively or in addition, the system 100 may execute the candidate algorithm 108 on a simulation of a hardware device that may perform multiple executions in parallel.
[0073] The system 100 then obtains a benchmark score 110 that characterizes a target attribute of the hardware device executing the candidate algorithm 108, such as computational complexity, reference locality, runtime, cache hit index, power consumption, etc. The benchmark score 110 is a numerical value that appropriately weights (i.e., according to the weight of the weighted sum) a combination of target attribute values of the target attribute being optimized. For example, the target attribute value may include a rank , the amount of local memory referenced, the average runtime, the average cache hit rate, the average energy consumed, etc. The system 100 is flexible because it supports complex random and non-differentiable benchmark scores 110. In some implementations, the benchmark score 110 strictly includes the rank of the tensor decomposition 106 , in which case the candidate algorithm 108 need not be executed on any particular set of hardware devices. The system 100 is then strictly optimized for low-rank decomposition, i.e., minimizing computational complexity. More preferably, the benchmark score is based on (additionally or alternatively) one or more of the target attributes that are specific to the set of hardware devices and indicative of the corresponding performance of the candidate algorithm 108 performed by the set of hardware devices.
[0074] The system 100 generates training examples 112a based on the tensor decompositions 106 and the benchmark scores 110 stored in the training data repository 114. The benchmark scores 110 can reinforce tensor decompositions 106 that have achieved relatively good scores (e.g., low rank and / or running time). Alternatively or in addition, the system 100 can generate synthetic training examples 112b and add them to the data repository 114. The synthetic training examples 112b are compared with the randomly initialized synthetic tensors. Supervised learning can train the network to mimic the decomposition of a synthetic dataset.
[0075] Sequentially or in parallel with DE 100A, training engine (TE) 100B randomly samples training tensor states from training data repository 114 118. TE 100B uses training status 118 to update the network parameters of the neural network 120. Training status 118 can be derived from training examples 112a or synthetic training examples 112b. In the case of being derived from training examples, the neural network is trained via reinforcement learning. In the case of being derived from synthetic training examples, the neural network is trained via supervised learning. And the composite tensor The hybrid training strategy for training neural networks on the tensor decomposition of can significantly outperform each training strategy individually. With Very different properties. TE 100B can continuously sample the training state 118, to continuously update network parameters 120.
[0076] The updated neural network may be utilized by DE 100A to obtain an improved tensor decomposition 106, which is then benchmarked and used as training examples 112a for TE 100B. This process may be repeated until the system 100 converges on a candidate algorithm 108 with the best achievable benchmark score 110 (relative to the baseline benchmark score of the target algorithm) on the optimizer, in which case the candidate algorithm 108 is optimized.
[0077] In other words, the system can continue to update the neural network while generating candidate algorithms until the termination criteria are met. Once the termination criteria are met, the system can select the most recently generated candidate or the candidate with the best benchmark score as the optimized algorithm. For example, the criteria can be met when the benchmark score 110 meets a certain threshold, the benchmark score 110 changes negligibly relative to a previous score stored in the data repository 114, a threshold number of candidates have been generated, a threshold amount of wall clock time has passed, etc. Note that the system 100 can obtain multiple optimized algorithms that perform equally well on the set of hardware devices, i.e., they have equivalent benchmark scores. The system 100 or a user can select any of these optimized algorithms to implement the target algorithm.
[0078] Figure 2 Is used to generate the target tensor Flowchart of an example process 200 for tensor decomposition 106 of FIG. 106. For convenience, process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a decomposition engine appropriately programmed according to the present specification, such as Figure 1 The decomposition engine 100A can perform process 200.
[0079] The Decomposition Engine (DE) converts tensors Initialize to the target tensor (202). In some implementations, DE initializes Before performing a basis transformation This can be achieved using the following transformations on the tensor components.
[0080] in and It is to define a new basis reversible matrices. The resulting decomposition can be converted back to the original (canonical) basis by performing an inverse basis transformation on the recovered set of factors. This procedure injects diversity into the system 100, which can help the neural network both in decomposing tensors and in learning from tensor decompositions. DE can randomly sample bases and perform decompositions in parallel on all such bases. For numerical stability, all basis matrices can be of type The unimodulus of the determinant.
[0081] Integer index is set to The initial value of .
[0082] DE gets the current state of the tensor (204). In the case of , this is equivalent to obtaining the generated in step 202 .
[0083] The DE performs a tree search of the state tree to determine the The target network output (206) is as follows. Figure 3 Describes the details of tree search. A tree search guided by a neural network returns a target network output that includes actions The improved strategy provides a possible rank 1 term that can be combined into Possible factors in Improved distribution on .
[0084] DE uses the target network output for the current state to adjust the current tensor state Apply the modification (208). DE is then Sampling to determine appropriate action .action With factor Corresponding to the choice of , this factor is obtained by subtracting the rank 1 term From the current state of the tensor Modify a tensor.
[0085] DE determines whether the modified tensor is equal to the zero tensor (210). If the tensor is not equal to the zero tensor , then process 200 repeats from step 204. If the tensor is equal to the zero tensor , then process 200 proceeds to step 212.
[0086] For the final iteration Determine if a tensor is equal to the zero tensor Afterwards, DE generates a tensor decomposition 106 (212) based on all modifications applied to the tensor. That is, DE merges the tensors generated by the actions All selected factors , to obtain the tensor break down.
[0087] Figure 3is a flow chart of an example process 300 for performing a tree search of a state tree to determine a target network output (e.g., Monte Carlo Tree Search (MCTS)). For convenience, process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a decomposition engine appropriately programmed according to the present specification, such as Figure 1 The decomposition engine 100A may perform the process 300 , for example, as part of executing step 206 of the method 200 .
[0088] The state tree includes representation tensors The state of the action is represented by Each state-action pair Stores a set of edge statistics corresponding to visit counts, action values, and prior probabilities, respectively , and .
[0089] The detailed steps used to perform the tree search are provided by T. Hubert, J. Schrittwieser, I. Antonoglou, M. Barekatain, S. Schmitt, and D. Silver, “Learning and planning in complex action spaces,” in International Conference on Machine Learning (ICML), 2021, which outlines a similar sampling-based MCTS used for Sampled AlphaZero.
[0090] DE identifies the representation tensor The root node (302) of the current state.
[0091] Starting from the root node, DE traverses the state tree until it reaches a leaf state The DE can traverse the state tree based on the action scores assigned to the edges, for example by maximizing on a probabilistic upper confidence tree (PUCT) boundary.
[0092] DE evaluation indicates leaf status The leaf node (306) of According to network parameters Handling Leaf Status , to generate network output for leaf nodes. The network output includes the strategy for leaf nodes , and can also include return output for leaf nodes .
[0093] DE extension indicates leaf state The network output can be used to extend the state tree with new (child) nodes connected to the leaf nodes. The new nodes represent tensor states. Specifically, we can use the The strategy determines a set of actions . Every action With factor The selection and resulting rank 1 term Associated. From the leaf state Subtract the rank 1 term to generate a new node Note that since, for example, is greater than for most cases of interest Huge action space , this group of actions Sampling is usually done according to a policy and is not fully enumerated. Therefore, the system 100 samples a fixed number of actions according to the policy for the leaf nodes. ,in .
[0094] In some implementations, the return output is Build values for leaf nodes This value could be the mean of the returned outputs, i.e. the expected benchmark scores, but more generally it could be constructed to facilitate some behavior of the tree search. For example, the value It can be a risk-seeking value that encourages exploration in order to find the best trajectory of the state tree. To achieve this, the neural network can generate a Corresponding The return output is composed of the outputs. In this way, the return output predicts the distribution of returns for the state in the form of values predicted for the quantiles mentioned above. In order to construct the risk preference value , DE can use the average of the predicted values above the 75th percentile. A review of quantile regression learning is provided by: W. Dabney, M. Rowland, M. Bellemare, and R. Munos, "Distributional reinforcement learning with quantile regression", Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, 2018.
[0095] The DE updates edge statistics for all traversed edges (310). While passing backward up the state tree, the DE increments the visit counts and action values (e.g., using the value ).
[0096] The DE executes consecutive trajectories starting from the root node until one or more criteria are met, such as maximum elapsed time, maximum number of leaf nodes evaluated, etc. In addition, if different trajectories reach nodes representing exactly the same state of a tensor, the DE can use a permutation table to reassemble the different trajectories. This can happen frequently because actions are usually commutative, i.e., changing the order of actions does not change the resulting state. In this case, the permutation table is a cache of frequently encountered nodes. If a node representing a particular state recurs via a different sequence of actions, the DE can permute this node with a previously encountered node (representing the same state) from the permutation table, thereby avoiding re-evaluating the subtree under this node. This generally increases the quality of the information collected from the state tree.
[0097] From the root node After trajectories, DE determines the current state The target network output (312) includes the action The improved strategy based on sampling can be based on the root node The normalized visit count of the edge is determined by is the visit count of the edge of the root node. Then, DE can use the improved strategy Select Action .
[0098] In some implementations, the selected action The subtree under (i.e. the selected edge at the root node) is targeted in the next iteration Reuse in tree searches.
[0099] In a further implementation, the tree search uses an adaptive temperature scheme to smooth the normalized access counts (i.e., suppress ) because some states can accumulate an order of magnitude more visits than other states due to permutation table and subtree reuse. For example, the normalized counts can be calculated using the function Smoothing is performed, where if , then the function is , otherwise the function is 1. Here, is a hyperparameter representing the maximum total visit count of all edges of the root node.
[0100] In yet a further implementation, when returning an improved policy to be used for action selection by the DE, the tree search ignores edges of the root node that have action scores lower than the action scores of the most visited edges.
[0101] Figure 4 is a flow chart of an example process 400 for generating training examples based on tensor decomposition. For convenience, process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, an optimization system appropriately programmed according to the present specification, such as Figure 1 The optimization system 100 can perform process 400.
[0102] The system obtains a candidate algorithm parameterized by tensor decomposition (402).
[0103] The system executes the candidate algorithm potentially multiple times on the target set of hardware devices (404). In some implementations, the candidate algorithm is executed on simulations of the set of hardware devices, in which case many simulations can be run in parallel.
[0104] The system generates target attribute values for the target attributes of the hardware device (406). The target attributes represent any number of performance goals of the hardware device, such as complexity, running time, etc. The target attribute values assign numerical values to these attributes. For example, the target attribute values may include the rank of the tensor decomposition and the average running time of the candidate algorithms executed on the hardware device .
[0105] The system determines a benchmark score for the tensor decomposition based on the target attribute values (408). The benchmark score can weight the target attribute values in a linear or nonlinear combination to emphasize certain target attributes. For example, the benchmark score Can include candidate algorithms Rank and average running time ,in is a user-specified coefficient that controls the relative importance of computational complexity versus computational speed. In this case, a high benchmark score corresponds to a combination of low complexity and fast runtime.
[0106] The system generates training examples based on the tensor decomposition and the benchmark scores (410). The training examples typically include a sequence of The tensor at each step in The training examples also include The target network output determined by tree search, i.e., the improved strategy .
[0107] The system stores the training examples in a training data repository for use in training the neural network via reinforcement learning (412). The training examples can be sorted in the data repository relative to their baseline scores to reinforce the tensor decompositions with the best relative scores. In some cases, this means discarding training examples with poor scores and retaining training examples with acceptable scores. The system can extract additional training examples from the same tensor decomposition by reordering the rank 1 terms (because the sums are commutative). In particular, the system can randomly swap two actions to generate additional training examples. This helps the system explore actions that it previously discovered only later in the decomposition sequence.
[0108] Alternatively or additionally, the system can incorporate synthetic training examples into the data repository for use in training a neural network via supervised learning. The synthetic training examples include random tensors The system can randomly sample factors According to the corresponding rank 1 term Build Therefore, this factor by definition consists of Tensor decomposition of . Since according to the rank 1 term Get a random tensor The forward process of is fundamental, so the system can generate an arbitrarily large set of synthetic training examples.
[0109] By repeatedly performing process 400 while training the neural network on training examples sampled from the training data repository, the neural network will generate candidate algorithms with increasing benchmark scores. Once a termination criterion for terminating training has been met, the system can select one of the candidate algorithms that have been generated using process 400 as the optimized algorithm, for example, by selecting the most recently generated candidate or by identifying the candidate algorithm with the best benchmark score among the already generated candidates.
[0110] Figure 5 is the network parameter value used to update the neural network Flowchart of an example process 500. For convenience, process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training engine appropriately programmed according to the present specification, such as Figure 1 The training engine 100B can execute process 500.
[0111] The training engine (TE) randomly samples training states from the training data repository (502).
[0112] TE according to network parameters Using Neural Networks to Process Training Status , to generate the training network output (504).
[0113] Each training state Associated with the training target. If the training state From the training examples, the training goal is to maximize the training strategy and for status To achieve this, the objective function can include a measure and In some implementations, the term for the divergence loss between The training objective also aims to maximize the output returned by the training compared to the actual benchmark score. The similarity between the estimated benchmark scores defined. For example, the objective function can include a term that measures the loss of the quantile regression distribution.
[0114] On the other hand, if the training state From synthetic training examples, the training goal is to maximize the training strategy and next action (i.e., the next rank 1 term ). That is, from the training state Start, next action Is used from the state Subtract the next rank 1 term Factor This corresponds to a certain step in the decomposition of the composite tensor. is constructed from known factors and rank-1 terms, so the next (best) action is known a priori. In this case, the objective function can include the state of the decomposition given a trained neural network to mimic the synthetic training data A term that measures the likelihood of an action in the case of .
[0115] TE determines the objective function with respect to network parameters The gradient of (506).
[0116] TE updates the network parameter values according to the gradient of the objective function (508) For example, TE can use a stochastic gradient descent method, such as RMSprop, Adam with decoupled weight decay, etc., to update the network parameter values.
[0117] Fig. 7A and Figure 7BThe optimized runtime for NVidia V100 GPU and TPU v2 are shown respectively. That is, the system 100 obtains a matrix multiplication algorithm for using Candidate algorithms for matrix multiplication and for computing The candidate algorithms are benchmarked by the running time (and complexity) of the block matrix, where each block (submatrix) is The optimized matrix multiplication algorithm is based on Figure 6 The algorithm is parameterized and compared to Strassen's algorithm. The speedup is measured relative to standard (e.g., cuBLAS for GPUs) matrix multiplication on the same hardware. The speedup is reported for various matrix sizes by recursively executing the optimized algorithm (although the algorithm is optimized on only one matrix size). The median speedup is reported over 200 runs, with 100% speedup on runs with 0. The standard deviation of .
[0118] like Fig. 7A and Figure 7B As can be seen, system 100 can customize algorithms for specific hardware devices such as GPUs and TPUs. However, as mentioned above, due to the different processing architectures of different hardwired devices, an optimized algorithm for a specific hardware device may not perform optimally on a different hardwired device. Figure 7C The results for the two devices are shown. The difference in speedup between two algorithms (tailored for the GPU and the TPU) is benchmarked on matrix sizes of 100,000 and 100,000, respectively. For example, when the algorithm optimized for the TPU is run on the GPU, the speedup relative to standard matrix multiplication is only 4.4%.
[0119] The target algorithm may be selected as a multilinear mapping that is a component of a data processing task, such as any known data processing task that performs and / or is performed on input data obtained from the real world (e.g., sensor data, such as data derived from a still camera, video camera, LIDAR sensor, or microphone) to generate data for generating a still or moving image of one or more objects in the real world or similar to such objects (e.g., the still or moving image is displayed on a screen), a sound data signal (e.g., the sound data signal is used as an input to a speaker to generate a sound signal), or control data for any form of mechatronic agent (e.g., a robot) configured to perform navigation and / or manipulation tasks in a real world environment. For example, the data processing task may be a classification task that classifies input data into one or more categories corresponding to the content of the input data by generating output data indicating corresponding one or more categories in the category; or the data processing task may be a task of generating a sound or image representing the semantic content of the data input; or the data processing task may be a task of generating control data based on sensor data describing the environment. Alternatively, the data processing task may be a task of converting data input as an encoding of a natural language (e.g., text in a first natural language) into data output as another encoding of a natural language having the same semantic content (e.g., text in another different natural language). Thus, the examples make it possible to obtain improved algorithms specific to a set of one or more hardware devices for performing components of a data processing task.
[0120] This specification uses the term "configured" in connection with system and computer program components. For a system of one or more computers to be configured to perform a particular operation or action, it is meant that the system has installed thereon software, firmware, hardware, or a combination thereof that, in operation, causes the system to perform the operation or action. For one or more computer programs to be configured to perform a particular operation or action, it is meant that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0121] Embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, program instructions may be encoded on an artificially generated propagation signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by a data processing device.
[0122] The term "data processing equipment" refers to data processing hardware, and includes all kinds of equipment, devices and machines for processing data, including, for example, a programmable processor, a computer or a plurality of processors or computers. The equipment may also be or further include a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the equipment may also optionally include code for creating an execution environment for a computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0123] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script or code, may be written in any form of programming language (including compiled or interpreted languages or declarative or procedural languages); and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines or code portions). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed over multiple sites and interconnected by a data communications network.
[0124] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines can be installed and run on the same one or more computers.
[0125] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by a dedicated logic circuit system (e.g., an FPGA or ASIC) or by a combination of a dedicated logic circuit system and one or more programmed computers.
[0126] A computer suitable for executing a computer program can be based on a general or special microprocessor or both, or any other kind of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for fulfilling or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by a special logic circuit system or incorporated into the special logic circuit system. Typically, the computer will also include one or more large-capacity storage devices for storing data, such as a disk, a magneto-optical disk, or an optical disk, or operatively coupled to receive data from it or transfer data to it or both. However, the computer does not have to have such a device. In addition, the computer can be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.
[0127] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0128] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser in response to a request received from a web browser on the user's device. In addition, the computer may interact with the user by sending a text message or other form of message to a personal device (e.g., a smart phone running a messaging application), and receiving a response message from the user in return.
[0129] A data processing device used to implement a machine learning model may also include, for example, dedicated hardware accelerator units for processing general-purpose and computationally intensive parts of machine learning training or production (i.e., inference, workloads).
[0130] Machine learning models can be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework).
[0131] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0132] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from it. Data generated at the user device, for example, the result of a user interaction, can be received from the device at the server.
[0133] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on the scope of any invention or the scope that may be claimed, but should be interpreted as descriptions of features that may be peculiar to a particular embodiment of a particular invention. Certain features described in the context of separate embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. In addition, although features may be described above as working in certain combinations and even initially claimed as such, one or more features from the claimed combination may be deleted from the combination in some cases, and the claimed combination may involve a sub-combination or a variant of a sub-combination.
[0134] Similarly, although operations are depicted in the drawings and described in the claims in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the operations shown be performed, to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0135] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for obtaining an optimized algorithm executed by one or more computers, wherein the optimized algorithm (i) is functionally equivalent to a target algorithm, and (ii) optimizes one or more target properties when executed on a target set of one or more hardware devices, wherein the method include: Initialize a target tensor representing the target algorithm; generating a tensor decomposition of the target tensor that parameterizes the candidate algorithm using a neural network having a plurality of network parameters, wherein the neural network is configured to receive a state of the tensor as an input and process the input according to the network parameters to generate a network output including a strategy for applying a modification to the tensor, and wherein generating the tensor decomposition comprises: For each step in the sequence of steps: Get the current state of the target tensor; determining a target network output for the current state by performing a tree search of a state tree having nodes representing states of the target tensor starting from a root node representing the current state, wherein the tree search is guided by the neural network according to the network parameters; applying a modification to the target tensor using the target network output for the current state; and determining whether to terminate the sequence based at least in part on whether the target tensor is equal to a zero tensor after applying the modification; and generating the tensor decomposition according to the modifications applied to the target tensor at each step in the sequence of steps; generating a target attribute value for each of the target attributes when executing the candidate algorithm on the target set of hardware devices; determining a benchmark score for the tensor decomposition based on the target property value of the candidate algorithm; generating training examples according to the tensor decomposition and the benchmark score; and The training examples are stored in a training data repository for use in updating the network parameters of the neural network.
2. The method of claim 1, further comprising: include: A specific candidate algorithm is selected as the optimized algorithm based on a benchmark score for the specific candidate algorithm generated by using the neural network.
3. A method as claimed in any preceding claim, in, The strategy defines a probability distribution over the possible rank-1 terms to be subtracted from the tensor.
4. A method as claimed in any preceding claim, in, The network outputs further include a return output defining an estimated return resulting from the tensor being in the state.
5. The method according to claim 4, in, The estimated return is an estimate of the expected benchmark score for the tensor decomposition.
6. The method according to claim 4, in, The estimated return is an estimate of the expected rank of the tensor decomposition.
7. The method according to any one of claims 4 to 6, in, Performing the tree search includes: Based on action scores assigned to edges connecting nodes of the state tree, traversing the edges until a leaf node is reached, wherein the edges represent possible modifications to be applied to the target tensor; Processing the state of the target tensor represented by the leaf node using the neural network according to the network parameters to generate a network output for the leaf node; extending the state tree at the leaf node using a policy for the leaf node; and For each edge of the state tree that has been traversed: Incrementing the visit count of the edge; and The action score for the edge is updated based on a value constructed from the return output for the leaf node.
8. The method according to claim 7, in, Performing the tree search further comprises: storing in a permutation table one or more nodes encountered during the tree search; While traversing the edges, determining that a newly encountered node represents the same state of the target tensor as a previously encountered node stored in the permutation table; and In response, the newly encountered node is replaced with the previously encountered node.
9. The method according to any one of claims 7 to 8, in, Determining the target network output according to the tree search further comprises: If the total access counts of all edges of the root node are greater than the maximum total access count, an adaptive temperature scheme is used to smooth the access counts of the edges of the root node.
10. The method according to any one of claims 7 to 9, in, Determining the target network output according to the tree search further comprises: Edges of the root node that have action scores lower than the action score of the edge of the root node with the highest visit count are ignored.
11. A method as claimed in any preceding claim, in, Initializing the target tensor includes performing a basis transformation on the target tensor, and wherein generating the tensor decomposition includes performing an inverse basis transformation on the tensor decomposition.
12. The method according to any preceding claim, further comprising: include: generating a set of tensor decompositions of one or more composite tensors, wherein the composite tensor is a randomly initialized tensor; generating a set of synthetic training examples according to the set of tensor decompositions of the synthetic tensor; and The combined training examples are stored in the training data repository for use in updating the network parameters of the neural network.
13. A method as claimed in any preceding claim, further comprising: include: retrieving a training state of a tensor associated with a training target from the training data repository; processing the training state using the neural network according to the network parameters to generate a training network output; determining a gradient of an objective function with respect to the network parameters, the objective function promoting the training network output to satisfy the training goal for the training state; and The network parameters are updated according to the gradient.
14. A method as claimed in any preceding claim, in, The optimized algorithm is recursively executed on the target set of hardware devices.
15. A method as claimed in any preceding claim, in, The target algorithm computes a bilinear map.
16. The method of claim 15, in, The bilinear map is a matrix multiplication.
17. A method as claimed in any preceding claim, in, The target attributes include at least one of the following: The computational complexity of the optimized algorithm, the runtime of the target set of hardware devices when executing the optimized algorithm, cache performance of the target set of hardware devices when executing the optimized algorithm, locality of reference of the optimized algorithm when executed on the target set of hardware devices, or The target set of hardware devices consumes power when executing the optimized algorithm.
18. A method as claimed in any preceding claim, in, The target set of hardware devices is a simulation of a set of hardware devices.
19. A method as claimed in any preceding claim, in, The target group of hardware devices includes at least one of the following: a central processing unit CPU, a graphics processing unit GPU, a tensor processing unit TPU or an application-specific integrated circuit ASIC.
20. The method of any preceding claim, further comprising: include: Receive new input; as well as The target algorithm is performed on the new input by executing the optimized algorithm on the target set of hardware devices.