Techniques for removing masks from trimmed neural networks
By generating unmasked outputs and replacing masked outputs, leveraging dense versions of tensors and decentralized operations, the problem of poor performance of traditional neural networks in real-time applications is solved, achieving faster inference speeds and smaller memory footprints.
Patent Information
- Application Number
- CN202510130277.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-22
- Filing Date
- 2020-01-20
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional neural networks perform poorly in real-time and low-latency applications, especially in areas such as automatic vehicle control. Although pruning neural networks reduce size, they fail to significantly accelerate inference time.
By generating unmasked output and replacing masked output, replacing the original tensor with dense versions of tensors, and inserting scatter operations in subsequent nodes to extend the tensor dimensions, enabling quick execution of operations.
Optimized neural networks perform inference operations faster than the original pruned neural networks, are suitable for real-time applications, and have a small memory footprint, saving memory resources.
Smart Images

Figure CN120218158A_ABST
Abstract
Description
[0001] Relevant information of divisional application
[0002] This application is a divisional application of a Chinese patent application with the application number "202010065865.1" and the invention title "Techniques for Removing Masks from Pruned Neural Networks", which was filed on January 20, 2020. Technical Field
[0003] The present disclosure generally relates to neural networks, and more particularly to techniques for removing masks from pruned neural networks. Background Art
[0004] A "tensor" is a mathematical construct commonly used in linear algebra applications such as machine learning and artificial intelligence. Scalars, vectors, and matrices are examples of tensors. A neural network typically includes one or more tensors, which are processed during the execution of the neural network to perform one or more operations. The values of a given tensor included in the neural network are modified through a training process to make one or more current outputs of the neural network close to one or more target outputs. After training is completed, some or all of the tensors included in the neural network may be large. Operations associated with large tensors generally cannot be executed quickly. Therefore, traditional neural networks are generally not suitable for real-time, low-latency applications such as autonomous vehicle control. Summary of the Invention
[0005] In one aspect, a computer-implemented method is described. The method includes: causing an unmasked output of the first neural network portion to be generated at least partially based on a masked output of the first neural network portion, wherein the dimension of the unmasked output is less than the dimension of the masked output; causing the unmasked output to replace the masked output; and causing a scatter operation to be performed to expand the dimension of the unmasked output to a dimension corresponding to the masked output.
[0006] In another aspect, a non-transitory computer-readable medium is described. The non-transitory computer-readable medium stores program instructions that, when executed by at least one processor, cause the at least one processor to at least: cause an unmasked output of the first neural network layer to be generated at least partially based on a masked output of the first neural network layer, wherein the dimension of the unmasked output is different from the dimension of the masked output; cause the unmasked output to replace the masked output; and cause a first operation to be performed to scale the dimension of the unmasked output to a dimension corresponding to the masked output.
[0007] In another aspect, a system is described. The system includes: a memory that stores one or more instructions; and a processor that executes the instructions to at least: cause an unmasked output of a first neural network layer to be generated based at least in part on a masked output of the first neural network layer, wherein a dimension of the unmasked output is less than a dimension of the masked output, cause the unmasked output to replace the masked output, and cause a scatter operation to be performed to expand the dimension of the unmasked output to a dimension corresponding to the masked output. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to understand the manner in which the above-recited features of the various embodiments can be obtained, the inventive concepts briefly summarized above may be described in more detail by reference to the various embodiments, in which some embodiments are illustrated in the accompanying drawings. It should be noted, however, that the drawings illustrate only typical embodiments of the inventive concepts and are therefore not to be considered limiting in any way, and that there are other equivalent embodiments.
[0009] Figure 1 A system configured to implement one or more aspects of the various embodiments is shown.
[0010] Figure 2 A graphical representation of a neural network according to various embodiments is shown.
[0011] Figure 3 An example of how a node evaluates a function based on an input tensor to generate an output tensor according to various embodiments is shown.
[0012] Figure 4 An example of how a node evaluates a function based on a masked input tensor to generate an output tensor according to various embodiments is shown.
[0013] Figure 5 An example of how a node evaluates a function based on a dense version of an input tensor according to various embodiments is shown.
[0014] Figure 6 Adjacent nodes on which a scatter operation can propagate according to various embodiments are shown.
[0015] Figure 7 An illustration of how a scatter operation propagates between Figure 6 adjacent nodes according to various embodiments is shown.
[0016] Figure 8 A flowchart of method steps for removing a mask from a neural network according to various embodiments.
[0017] Figure 9 A block diagram of a computer system configured to implement one or more aspects of the various embodiments is shown.
[0018] Figure 10 is a block diagram of a parallel processing unit (PPU) included in a parallel processing subsystem of Figure 9 .
[0019] Figure 11 is a block diagram of a general processing cluster (GPC) included in a parallel processing unit (PPU) of Figure 10 . DETAILED DESCRIPTION
[0020] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.
[0021] As described above, a neural network may include one or more tensors that are processed during the execution of an artificial neural network to perform one or more operations. Tensors can be very large and thus, in some cases, cannot be processed quickly. To address this issue, an artificial neural network can be "pruned" to increase the speed of processing the corresponding tensors.
[0022] Pruning a neural network generally involves generating one or more masks that zero out elements of one or more tensors included in the neural network. Although pruning reduces the overall size of the neural network, pruning a neural network does not significantly and / or reliably accelerate the inference time of the neural network. In particular, operations involving tensors of the pruned neural network (at least partially zeroed out by the mask) are still performed even if the results of these operations do not affect the final inference output of the neural network.
[0023] Various embodiments disclosed herein include a technique for removing a mask from a pruned neural network, where the pruned neural network is represented by a graph of nodes. A given node in the graph includes at least one tensor (W) and a corresponding mask (M). The operation M·W zeros out the elements of W. A first function associated with the given node is evaluated based on M·W to produce an output tensor X of the node. The output tensor X is provided as an input to one or more subsequent nodes in the graph.
[0024] In one embodiment, to remove a mask M applied to a given node, the dense version of a tensor W denoted by w is used to replace M·W. The dimension of the tensor w is smaller than that of W, and thus, operations involving w are performed much faster than those involving the tensor W. In particular, the evaluation of a first function based on w is faster than the evaluation of a first function based on M·W. The evaluation of the first function based on the tensor w generates a smaller output tensor, denoted as x. Since subsequent nodes in the neural network expect the given node to provide an output of a tensor X with a larger dimension, a scatter operation is inserted in the subsequent nodes to add zeros to the tensor x so as to expand the tensor x to produce the tensor X (or a tensor with an equivalent dimension). Since less data can be processed, the operations associated with the given node can be performed in a fast manner. The scatter operation can be propagated forward through the graph, towards the output, to accelerate other functions. The scatter operation can also be merged with other scatter operations and / or absorbed into subsequent nodes. By the disclosed techniques, a pruned neural network including one or more masks can be modified to produce an optimized neural network.
[0025] At least one technical advantage of the techniques described herein is that the optimized neural network performs inference operations faster than the original pruned neural network. Thus, the optimized neural network is well-suited for use in real-time applications such as autonomous vehicles. Another advantage of the techniques described herein is that, compared with the pruned neural network, the optimized neural network can have a smaller memory footprint, thus saving memory resources. These technical advantages represent multiple technological advancements over prior art methods.
[0026] System Overview
[0027] Figure 1 A system configured to implement one or more aspects of various embodiments is shown. As shown, in one embodiment, a neural network optimization pipeline 100 includes a training engine 110, a pruning engine 120, and a demasking engine 130.
[0028] In one embodiment, the training engine 110 generates and trains an initial neural network 102 to generate a trained neural network 112. During training, the training engine 110 iteratively adjusts one or more tensors included in the initial neural network 102 based on training data to cause the output of the initial neural network 102 to more closely match a target output. When training is complete, the trained neural network 112 includes one or more tensors derived from one or more tensors included in the initial neural network 102. The training engine 110 may implement any technically feasible training process to generate the trained neural network 112, including backpropagation and / or gradient descent, etc. The trained neural network 112 is generally capable of performing inference operations to generate output data based on input data not included in the training data.
[0029] In one embodiment, the pruning engine 120 prunes the trained neural network 112 to generate a masked neural network 122. During pruning, the pruning engine 120 identifies redundant elements within one or more tensors included in the trained neural network 112, and then generates one or more masks that cause those specific elements to be multiplied by zero (set to zero). When processing a tensor, elements of the tensor that do not contribute to the output can be considered "redundant". Since the identified elements are redundant, those elements can be set to zero without adversely affecting the functional characteristics of the masked neural network 122. Setting the redundant elements to zero in this way can reduce the computational load associated with processing some or all of the tensors included in the masked neural network 122.
[0030] In one embodiment, the training engine 110 and the pruning engine 120 interoperate with each other to simultaneously train the initial neural network 102 and prune the initial neural network 102. For example, the training engine 110 may perform a first training to modify a portion of the initial neural network 102, and then the pruning engine 120 may perform pruning to insert one or more masks into at least a partially trained version of the initial neural network 102. In this way, the training engine 110 and the pruning engine 120 can coordinate their operations to generate the masked neural network 122.
[0031] In one embodiment, the unmasking engine 130 performs an unmasking process using the masked neural network 122 to generate an optimized neural network 132. During the unmasking process, the unmasking engine 130 removes one or more masks previously introduced by the pruning engine 120 from the masked neural network 122 and, as described above, applies various other modifications to the masked neural network 122 to produce the optimized neural network 132. The optimized neural network 132 has the same or similar functional characteristics as the masked neural network 122. However, the optimized neural network 132 can perform various processing operations, including inference operations and the like, faster than the masked neural network 122 can perform those processing operations.
[0032] In one embodiment, any one of the initial neural network 102, the trained neural network 112, the masked neural network 122, and the optimized neural network 132 can be represented by a node graph coupled together by a set of edges. Each node can be associated with one or more tensors and one or more functions evaluated based on one or more tensors. Figure 2 illustrates a node graph that can represent Figure 1 any of the neural networks shown in
[0033] Example graphical representation of a neural network
[0034] Figure 2 Illustrates graphical representations of neural networks according to various embodiments. As shown, in one embodiment, the graphical representation 200 includes an input 210, a set of nodes N0, N1, and N2, and an output 220. Node N0 processes the input 210 to generate an output provided to node N1. Node N1 processes the received input to generate an output provided to node N2. Node N2 processes the received input to generate the output 220. Here, the graphical representation 200 is merely an example of a node graph that can represent a neural network.
[0035] In one embodiment, each node included in the graphical representation 200 corresponds to a neural network-oriented function and one or more tensors. For example, node N0 can correspond to a convolution function, a concatenation function, a matrix multiplication function, an activation function, or a rectified linear unit (ReLU) function, etc. Additionally, the function associated with node N0 can be evaluated based on one or more input tensors to generate one or more output tensors. In one embodiment, a given input tensor can be associated with an inbound edge of a given node, and a given output tensor can be associated with an outbound edge of a given node.
[0036] In one embodiment, the graphical representation 200 represents the initial neural network 102, and one or more nodes of the graphical representation 200 produce output tensors that do not match the target output. In the above combinationFigure 1 During the discussed training process, the training engine 110 iteratively adjusts one or more input tensors associated with one or more nodes to make the output tensor more closely match the target output.
[0037] In one embodiment, the graphical representation 200 represents a trained neural network 112, and one or more nodes of the graphical representation correspond to at least partially redundant tensors. As mentioned herein, "redundant" elements of a tensor are elements that do not contribute to the output of a function evaluated based on the tensor. During the pruning process discussed above in conjunction with Figure 1 the pruning engine 120 incorporates masks into one or more nodes to zero out the redundant elements of the associated tensors, thereby generating a masked neural network 122. Figures 3 - 4 An example of how the pruning engine 120 incorporates masks into nodes is shown.
[0038] In one embodiment, as described above, the graphical representation 200 represents the masked neural network 122, and one or more nodes of the graphical representation 200 include masks that zero out elements of the input tensors associated with those nodes. During the demasking process discussed above in conjunction with Figure 1 the demasking engine 130 removes these masks and makes various other modifications to some or all of the nodes to generate an optimized neural network 132. The optimized neural network 132 includes tensors that are smaller in size than the corresponding tensors included in the masked neural network 122. Thus, the optimized neural network 132 can be executed more quickly compared to the masked neural network 122. Figures 5 - 7 An example of how the demasking engine 130 demasks nodes and performs various other optimization operations is shown.
[0039] Examples of pruning and demasking processes
[0040] Figure 3 Examples of how nodes evaluate functions based on input tensors to generate output tensors according to various embodiments are shown. As shown, in one embodiment, node 300 includes tensor W, function f1, and tensor X. Node 300 can be included in the graphical representation of the trained neural network 112, etc. During the training process discussed above in conjunction with Figures 1 - 2 the training engine 110 generates node 300. When node 300 is executed, function f1 is evaluated based on tensor W to produce tensor X. Function f1 can be any technically feasible function configured to operate on one or more input tensors to generate one or more output tensors.
[0041] In one embodiment, the pruning engine 120 can modify node 300 to simplify the evaluation of function f1, as described below in conjunction withFigure 4 which is described in more detail.
[0042] Figure 4 illustrates an example of how a node according to various embodiments evaluates a function based on a masked input tensor to generate an output tensor. As shown, in one embodiment, node 400 includes tensor W, mask M, function f1, and tensor X. Node 400 may be included in the graphical representation of the trained neural network 112. During the pruning process discussed above in conjunction with Figures 1 - 2 the pruning engine 120 generates node 400 based on Figure 3 node 300 as shown. In particular, the pruning engine 120 identifies elements of W that do not contribute to tensor X and then generates a mask M to zero out these elements, thereby reducing the computational burden of evaluating function f1. When node 400 is executed, function f1 is evaluated based on tensor W and mask M to generate tensor X.
[0043] In one embodiment, evaluating f1(W·M) may be computationally more efficient than evaluating f1(W). However, in practice, implementing mask M to zero out redundant elements of tensor W may not significantly improve computational efficiency because these zeroed-out elements still have to be processed during the evaluation of f1(W·M). To address this issue, the demasking engine 130 may perform techniques to remove the mask and perform other optimizations, as described in more detail below in conjunction with Figures 5 - 7 which is described in more detail.
[0044] Figure 5 illustrates an example of how a node according to various embodiments evaluates a function based on a dense version of an input tensor. As shown, in one embodiment, node 500 includes tensor w, function f1, tensor x, scatter operation S1, and tensor X. During the demasking process discussed above in conjunction with Figures 1 - 2 the demasking engine 130 generates node 500 based on Figure 4 node 400. In particular, the demasking engine 130 replaces tensor W with tensor w. Tensor w is a denser version of tensor W and has a smaller dimension than tensor W. To generate tensor w, the demasking engine 130 identifies the portions of tensor W that are zeroed out by Figure 4 the mask M and then removes these portions of tensor W to produce tensor w. The remaining portions of tensor w are complementary to the zeroed-out portions of tensor W and may be referred to as corresponding to them. When node 500 is executed, function f1 is evaluated based on tensor w instead of tensor W. Since the dimension of tensor w is less than that of tensor W, the evaluation of f1(w) is significantly faster than f1(W).
[0045] In one embodiment, evaluating f1(w) produces tensor x, whose dimension is less than Figures 3 - 4The dimension of the tensor X. In some cases, nodes located after node 500 may expect the input tensor to have the same dimension as tensor X. To address this issue, the demasking engine 130 generates a scattering operation S1 to expand tensor x to generate tensor X (or a tensor of equivalent dimension). The scattering operation S1 inserts zeros into tensor x corresponding to those elements in tensor W that were previously zeroed out by mask M and then removed, thereby restoring the dimension of tensor x to the dimension associated with tensor X. Thus, nodes located after node 500 receive a tensor with the expected dimension of tensor X as input.
[0046] In one embodiment, instead of or in addition to inserting zeros, the scattering operation S1 can insert any technically feasible value into a given tensor. For example, assume that function f1 is a sigmoid function that maps the zeros contained in tensor W to the zeros contained in tensor X. The scattering operation S1 can insert 1 into tensor x to compensate for the zeros removed from tensor W.
[0047] In one embodiment, the demasking engine 130 performs the above technique on each node included in the graphical representation 200, thereby incorporating one or more scattering operations into the graphical representation. Then, as described in more detail below in conjunction with Figures 6 - 7 The demasking engine 130 iteratively propagates one or more of these scattering operations to the output of the graphical representation 200.
[0048] Figure 6 Adjacent nodes are shown on which scattering operations can be propagated according to various embodiments. As shown, Figure 5 Node 500 is adjacent to a subsequent node 600. Node 600 includes function f2 and tensor Y. Function f2 is evaluated based on tensor X to produce tensor Y. The demasking engine 130 propagates the scattering operation S1 from node 500 to node 600 to alleviate the computational burden of evaluating function f2(X), as described in more detail below in conjunction with Figure 7 To be described in more detail.
[0049] Figure 7 Shows how scattering operations can propagate between Figure 6 Adjacent nodes according to various embodiments. As shown, node 700 includes tensor w, function f1, and tensor x. Node 700 includes the same elements as node 500, except that the scattering operation S1 and tensor X are omitted. Node 710 includes function f2, tensor y, scattering operation S2, and tensor Y. The demasking engine 130 generates the scattering operation S2 by propagating the scattering operation S1 forward to node 710. Thus, function f2 can be evaluated based on tensor x instead of the larger tensor X. Since the dimension of tensor x is less than the dimension of the output tensor X, f2(x) can be evaluated faster than f2(X).
[0050] In one embodiment, evaluating f2(x) produces an output tensor y that has a smaller dimension than Figure 6 the output tensor Y. However, subsequent nodes may expect an input having the dimension of Y. To address this issue, the scatter operation S2 expands the tensor y into the output tensor Y (or a tensor of equivalent dimension). The scatter operation S2 inserts zeros into the tensor y corresponding to any elements of W and / or X that are zeroed out by the mask. Thus, the nodes located after node 710 receive an input tensor having the expected dimension associated with Y.
[0051] In one embodiment, the demasking engine 130 propagates the Figure 6 scatter operation S1 of forward to node 710 by combining the scatter operation S1 with any scatter operations previously associated with node 710. For example, node 710 may include a scatter operation introduced by the demasking engine 130 in the manner described above in connection with Figure 5 . The demasking engine 130 combines the scatter operation S1 with any such pre-existing scatter operations associated with node 710. Generally, any two or more scatter operations can be combined when the zero values inserted by those scatter operations are aligned along the same dimension. For example, two scatter operations that insert zeros along different rows can be combined to form a scatter operation that inserts zeros along both rows.
[0052] In one embodiment, the demasking engine 130 propagates the Figure 6 scatter operation S1 of forward to node 710 and then stacks the scatter operation S1 together with any scatter operations previously associated with node 710. When these two scatter operations insert zeros along different dimensions, the demasking engine 130 can propagate the given scatter operation and then stack it adjacent to the other scatter operation. For example, the scatter operation S1 can be propagated forward to insert a row of zeros into the output tensor y, and the demasking engine 130 can stack the scatter operation S1 adjacent to another scatter operation that inserts a column of zeros into the output tensor y.
[0053] In one embodiment, the demasking engine 130 propagates the scatter operation S1 forward and causes node 710 to absorb the scatter operation. For example, referring to Figure 6 , assume that the scatter operation S1 inserts a column of zeros into the tensor x to generate the tensor X. Also assume that the function f2 is a matrix multiplication operation that multiplies the tensor X by an input tensor. Since the column of zeros inserted into X multiplies the corresponding rows of the input tensor, the scatter operation S1 can be removed as long as the corresponding rows of the input tensor are also removed. This method does not change the dimension of the output tensor but eliminates the need for the scatter operation S1.
[0054] In one embodiment, in addition to, or instead of, generating a scatter operation, the unmasking engine 130 generates a gather operation. For example, the unmasking engine 130 can generate a gather operation that resides after a node and selects a subset of the output tensor of that node to pass to a subsequent node. Because the subsequent node only receives a subset of the output tensor, computations involving that subset can be performed more quickly compared to equivalent computations performed using the entire output tensor. The unmasking engine 130 can also propagate the gather operation towards the input of the graphical representation 200 and combine, stack, and / or absorb the gather operations similar to how the unmasking engine 130 combines, stacks, and / or absorbs scatter operations.
[0055] Generally referring Figures 3 - 7 , in various embodiments, the techniques described in conjunction with these figures can be advantageously applied to generate an optimized neural network 132. The disclosed techniques can be applied to any technically feasible neural network and / or any of its technically feasible graphical representations. The disclosed techniques can also be applied to any part of a neural network, including one or more layers, components, or elements, etc. The optimized neural network 132 can perform inference operations significantly faster than the masked neural network 122 while retaining the functional characteristics of the masked neural network 122. Thus, the disclosed techniques represent a significant advancement over prior art that cannot provide a similar improvement in inference speed.
[0056] The process of unmasking a masked neural network
[0057] Figure 8 is a flowchart of method steps for removing a mask from a neural network according to various embodiments. Although the method steps are described in the context of Figures 1 - 7 the system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of this embodiment.
[0058] As shown, method 800 begins at step 802, in which the unmasking engine 130 identifies a first node included in the graphical representation of the masked neural network. In one embodiment, as Figure 4 shown, the graphical representation of the masked neural network can be the graphical representation 200 configured with a set of masked tensors. Figure 2
[0059] In step 804, the de-masking engine 130 determines a first tensor, a first mask, and a first function included in the first node. In one embodiment, the de-masking engine 130 traverses a graphical representation of neural network nodes node by node and iteratively processes each node. The de-masking engine 130 can analyze the first node by parsing the program code associated with the first node to extract the first tensor, the first mask, and the first function.
[0060] In step 806, the de-masking engine 130 removes the first mask from the first node. In one embodiment, the first mask zeros out elements of the first tensor that do not contribute to the output of the first function. These elements can be considered redundant and can be safely removed without adversely affecting the output of the first function. Figure 1 The pruning engine 120 can generate the first mask via the pruning process described above.
[0061] In step 808, the de-masking engine 130 replaces the first tensor with a densified version of the first tensor. As used herein, the term "densified" refers to a more dense form of the tensor. In one embodiment, the de-masking engine 130 generates a densified version of a given tensor by analyzing the mask associated with the tensor and identifying the portions of the tensor that are zeroed out by the mask. The de-masking engine 130 can then remove these portions from the tensor to produce a smaller, more dense version of the tensor. A given function evaluated based on the smaller, more dense version of the tensor can be evaluated more quickly than the given function evaluated based on the original, larger tensor.
[0062] In step 810, the de-masking engine 130 adds a first scatter operation near the first node after the first function. Since the first function receives the densified version of the tensor as input, the first function may produce a smaller, more dense output when evaluated compared to the output produced when the first function is evaluated based on the original, larger tensor. In various embodiments, the first scatter operation expands that smaller, more dense output to have dimensions associated with the previous output of the first function. Thus, input data having that dimension is provided to a downstream node expecting an input having a particular dimension.
[0063] In step 812, the unmasking engine 130 propagates the first scatter operation to the output of the graphical representation. The unmasking engine 130 can propagate the first scatter operation via one or more different techniques. In one embodiment, the unmasking engine 130 removes the first scatter operation from a position after the first node and produces a second scatter operation at a position after the second node. The second node then receives a smaller, denser output of the first function as input and can thus be evaluated more quickly. The second scatter operation expands the output of the second function to be consistent with the expected dimensions associated with subsequent nodes. In another embodiment, the unmasking engine 130 combines the first scatter operation with at least one other scatter operation when those scatter operations insert zeros along the same axis. In another embodiment, the unmasking engine 130 stacks the first scatter operation with at least one other scatter operation when those scatter operations insert zeros along different axes. In yet another embodiment, the unmasking engine 130 causes subsequent nodes to absorb the scatter operation by modifying the input tensors processed by the subsequent nodes.
[0064] Generally referring Figures 1 - 8 , those skilled in the art will understand that the disclosed techniques can be implemented by any technically feasible combination of computer hardware and / or software. The following describes in conjunction with Figures 9 - 11 an example computer system configured to execute the neural network optimization pipeline 100 and / or the optimized neural network 132 in more detail.
[0065] Example hardware architecture
[0066] Figure 9 is a block diagram showing a computer system 900 configured to implement one or more aspects of various embodiments. In some embodiments, the computer system 900 is a server machine operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network. In one embodiment, various elements of the computer system 900 execute Figure 1 the neural network optimization pipeline 100 and / or the optimized neural network 132.
[0067] In various embodiments, the computer system 900 includes, but is not limited to, a central processing unit (CPU) 902 and a system memory 904, which is coupled to a parallel processing subsystem 912 via a memory bridge 905 and a communication path 913. The memory bridge 905 is also coupled to an I / O (input / output) bridge 907 via a communication path 906, and the I / O bridge 907 is in turn coupled to a switch 916.
[0068] In one embodiment, the I / O bridge 907 is configured to receive user input information from an optional input device 908, such as a keyboard or a mouse, and forward the input information to the CPU 902 via the communication path 906 and the memory bridge 905 for processing. In some embodiments, the computer system 900 may be a server machine in a cloud computing environment. In such an embodiment, the computer system 900 may not have the input device 908. Instead, the computer system 900 may receive equivalent input information by receiving commands in the form of messages sent over a network and received via the network adapter 918. In one embodiment, the switch 916 is configured to provide connections between the I / O bridge 907 and other components of the computer system 900, such as the network adapter 918 and various plug-in cards 920 and 921.
[0069] In one embodiment, the I / O bridge 907 is coupled to the system disk 914, which may be configured to store content as well as applications and data for use by the CPU 902 and the parallel processing subsystem 912. In one embodiment, the system disk 914 provides non-volatile storage for applications and data and may include a fixed or removable hard disk drive, a flash device, and a CD-ROM (compact disc read-only memory), a DVD-ROM (digital versatile disc-ROM), a Blu-ray, an HD-DVD (high definition DVD), or other magnetic, optical, or solid-state storage devices. In various embodiments, other components, such as a universal serial bus or other port connection, an optical disc drive, a digital versatile disk drive, a video recording device, etc., may also be connected to the I / O bridge 907.
[0070] In various embodiments, the memory bridge 905 may be a north bridge chip, and the I / O bridge 907 may be a south bridge chip. Additionally, the communication paths 906 and 913 and other communication paths within the computer system 900 may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0071] In some embodiments, the parallel processing subsystem 912 includes a graphics subsystem that delivers pixels to an optional display device 910, which may be any conventional cathode ray tube, liquid crystal display, light emitting diode display, etc. In such an embodiment, the parallel processing subsystem 912 includes circuitry optimized for graphics and video processing, including, for example, video output circuitry. As described below in connection with Figure 8 and Figure 9More specifically described, such circuitry may incorporate across one or more parallel processing units (PPUs) included within parallel processing subsystem 912, also referred to herein as parallel processors. In other embodiments, parallel processing subsystem 912 includes circuitry optimized for general-purpose and / or compute processing. Again, such circuitry may incorporate across one or more PPUs included within parallel processing subsystem 912, which are configured to perform such general-purpose and / or compute operations. In other embodiments, one or more PPUs included within parallel processing subsystem 912 may be configured to perform graphics processing, general-purpose processing, and compute processing operations. System memory 904 includes at least one device driver configured to manage the processing operations of one or more PPUs within parallel processing subsystem 912.
[0072] In various embodiments, parallel processing subsystem 912 may be integrated with Figure 9 one or more other elements to form a single system. For example, parallel processing subsystem 912 may be integrated with CPU 902 and other connection circuitry on a single chip to form a system-on-chip (SoC).
[0073] In one embodiment, CPU 902 is the main processor of computer system 900, which controls and coordinates the operation of other system components. In one embodiment, CPU 902 issues commands that control the operation of the PPU. In some embodiments, communication path 913 is a serial bus (PCI Express) link, as is well known in the art, where dedicated channels are allocated to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture. The PPU may be equipped with any number of local parallel processing memories (PP memories).
[0074] It will be appreciated that the systems shown herein are illustrative and may be subject to variations and modifications. The connection topology may be modified as needed, including the number and arrangement of bridges, the number of CPUs 902, and the number of parallel processing subsystems 912. For example, in some embodiments, system memory 904 may connect directly to CPU 902 rather than through memory bridge 905, and other devices will communicate with system memory 904 via memory bridge 905 and CPU 902. In other embodiments, parallel processing subsystem 912 may be connected to I / O bridge 907 or directly to CPU 902 rather than to memory bridge 905. In other embodiments, I / O bridge 907 and memory bridge 905 may be integrated into a single chip rather than existing as one or more discrete devices. Finally, in certain embodiments, there may not be Figure 9One or more components as shown. For example, switch 916 can be omitted, and network adapter 918 and insertion cards 920, 921 will be directly connected to I / O bridge 907.
[0075] Figure 10 is a block diagram of a parallel processing unit (PPU) 1002 included in Figure 9 parallel processing subsystem 912 according to various embodiments. Although Figure 10 illustrates one PPU 1002, as described above, parallel processing subsystem 912 can include any number of PPU 1002. As shown, PPU 1002 is coupled to local parallel processing (PP) memory 1004. PPU 1002 and PP memory 1004 can be implemented using one or more integrated circuit devices (such as programmable processors, application specific integrated circuits (ASICs), or storage devices) or in any other technically feasible manner.
[0076] In some embodiments, PPU 1002 includes a graphics processing unit (GPU), which can be configured to implement a graphics rendering pipeline to perform various operations related to generating pixel data based on graphics data provided by CPU 902 and / or system memory 904. When processing graphics data, PP memory 1004 can be used as a graphics memory, which stores one or more conventional frame buffers and, if needed, one or more other rendering targets. Among them, PP memory 1004 can be used to store and update pixel data, and transfer the final pixel data or display frame to an optional display device 910 for display. In some embodiments, PPU 1002 can also be configured for general processing operations and computing operations. In some embodiments, computer system 900 can be a server machine in a cloud computing environment. In such an embodiment, computer system 900 may not have a display device 910. Instead, computer system 900 can generate equivalent output information by sending commands in the form of messages over the network via network adapter 918.
[0077] In some embodiments, CPU 902 is the main processor of computer system 900, which controls and coordinates the operations of other system components. In one embodiment, CPU 902 issues commands to control the operation of PPU 1002. In some embodiments, CPU 902 writes a command stream for PPU 1002 to a data structure (in Figure 9 or Figure 10(not explicitly shown in the figure), the data structure can be located in the system memory 904, the PP memory 1004, or another storage location accessible to both the CPU 902 and the PPU 1002. A pointer to the data structure is written into the command queue, also referred to herein as the push buffer, to initiate the processing of the command data stream in the data structure. In one embodiment, the PPU 1002 reads the command stream from the command queue and then executes the commands asynchronously relative to the operations of the CPU 902. In an embodiment where multiple push buffers are generated, the application can specify an execution priority for each push buffer through the device driver to control the scheduling of different push buffers.
[0078] In one embodiment, the PPU 1002 includes an I / O (input / output) unit 1005 that communicates with the rest of the computer system 900 via the communication path 913 and the memory bridge 905. In one embodiment, the I / O unit 1005 generates packets (or other signals) for transmission on the communication path 913 and also receives all incoming packets (or other signals) from the communication path 913, directing the incoming packets to the appropriate components of the PPU 1002. For example, commands related to processing tasks can be directed to the host interface 1006, while commands related to memory operations (e.g., read from or written to the PP memory 1004) can be directed to the crossbar unit 1010. In one embodiment, the host interface 1006 reads each command queue and sends the command stream stored in the command queue to the front end 1012.
[0079] As described above in connection with Figure 9 the connection of the PPU 1002 to the rest of the computer system 900 can vary. In some embodiments, the parallel processing subsystem 912 including at least one PPU 1002 is implemented as an insert card that can be inserted into an expansion slot of the computer system 900. In other embodiments, the PPU 1002 can be integrated with a bus bridge (e.g., the memory bridge 905 or the I / O bridge 907) on a single chip. Again, in other embodiments, some or all of the elements of the PPU 1002 can be included with the CPU 902 in a single integrated circuit or system-on-chip (SoC).
[0080] In one embodiment, the front end 1012 sends the processing tasks received from the host interface 1006 to a work distribution unit (not shown) within the task / work unit 1007. In one embodiment, the work distribution unit receives a pointer to the encoded processing task as task metadata (TMD) and stores it in memory. The pointer to the TMD is included in a command stream that is stored as a command queue and received by the front end unit 1012 from the host interface 1006. The processing tasks that can be encoded as TMD include an index associated with the data to be processed and status parameters and commands that define how to process the data. For example, the status parameters and commands can define a program to be executed on the data. Also for example, the TMD can specify the number and configuration of a set of CTAs. Generally, each TMD corresponds to one task. The task / work unit 1007 receives the tasks from the front end 1012 and ensures that the GPC 1008 is configured to an active state before starting the processing tasks specified by each TMD. A priority can be specified for each TMD used to schedule the execution of the processing tasks. Processing tasks can also be received from the processing cluster array 1030. Optionally, the TMD can include a parameter that controls whether to add the TMD to the head or tail (or a list of pointers to processing tasks) of the processing task list, thus providing another level of control over the execution priority.
[0081] In one embodiment, the PPU 1002 implements a highly parallel processing architecture based on the processing cluster array 1030, which includes a set of C general processing clusters (GPCs) 1008, where C ≥ 1. Each GPC 1008 is capable of executing a large number of threads (e.g., hundreds or thousands) in parallel, where each thread is an instance of a program. In various applications, different GPCs 1008 can be assigned to process different types of programs or perform different types of computations. The assignment of GPCs 1008 can vary according to the workload generated by each type of program or computation.
[0082] In one embodiment, the memory interface 1014 includes a set of D partition units 1015, where D ≥ 1. Each partition unit 1015 is coupled to one or more dynamic random access memories (DRAMs) 1020 residing within the PPM memory 1004. In some embodiments, the number of partition units 1015 is equal to the number of DRAMs 1020, and each partition unit 1015 is coupled to a different DRAM 1020. In other embodiments, the number of partition units 1015 may be different from the number of DRAMs 1020. Those of ordinary skill in the art will understand that the DRAMs 1020 may be replaced with any other technically suitable storage device. In operation, various render targets, such as texture maps and frame buffers, may be stored on the DRAMs 1020, allowing the partition units 1015 to write portions of each render target in parallel to effectively utilize the available bandwidth of the PP memory 1004.
[0083] In one embodiment, a given GPC 1008 may process data to be written to any DRAM 1020 within the PPM memory 1004. In one embodiment, the crossbar unit 1010 is configured to route the output of each GPC 1008 to the input of any partition unit 1015 or to any other GPC 1008 for further processing. The GPC 1008 communicates with the memory interface 1014 via the crossbar unit 1010 to read from or write to the various DRAMs 1020. In some embodiments, the crossbar unit 1010 has a connection to the I / O unit 1005 in addition to its connection to the PPM memory 1004 via the memory interface 1014, enabling the processing cores within different GPCs 1008 to communicate with the system memory 904 or other memories that are not local to the PPU 1002. In Figure 10 an embodiment, the crossbar unit 1010 is directly connected to the I / O unit 1005. In various embodiments, the crossbar unit 1010 may use virtual channels to separate the traffic flow between the GPCs 1008 and the partition units 1015.
[0084] In one embodiment, GPC 1008 can be programmed to perform processing tasks related to a wide variety of applications, including but not limited to linear and non-linear data transformations, filtering of video and / or audio data, modeling operations (e.g., applying physical laws to determine the position, velocity, and other properties of an object), image rendering operations (e.g., tessellation shaders, vertex shaders, geometry shaders, and / or pixel / fragment shader programs), general computing operations, and the like. In operation, PPU 1002 is configured to transfer data from system memory 904 and / or PP memory 1004 to one or more on-chip memory units, process the data, and write the resulting data back to system memory 904 and / or PP memory 1004. The data can then be accessed by other system components, including CPU 902, another PPU 1002 in parallel processing subsystem 912, or another parallel processing subsystem 912 in computer system 900.
[0085] In one embodiment, any number of PPUs 1002 can be included in parallel processing subsystem 912. For example, multiple PPUs 1002 can be provided on a single plug-in card, or multiple plug-in cards can be connected to communication path 913, or one or more PPUs 1002 can be integrated into a bridge chip. The PPUs 1002 in a multi-PPU system can be the same as or different from each other. For example, different PPUs 1002 may have different numbers of processing cores and / or different numbers of PP memories 1004. In embodiments where there are multiple PPUs 1002, those PPUs can run in parallel to process data with a higher throughput (than may be possible with a single PPU 1002). Systems that include one or more PPUs 1002 can be implemented in a variety of configurations and forms, including but not limited to desktop computers, laptop computers, handheld personal computers or other handheld devices, servers, workstations, gaming consoles, embedded systems, and the like.
[0086] Figure 11 is a block diagram of a general processing cluster (GPC) 1008 included in a Figure 10 parallel processing unit (PPU) 1002 according to various embodiments. As shown, GPC 1008 includes, but is not limited to, a pipeline manager 1105, one or more texture units 1115, a pre-raster (preROP) unit 1125, a work distribution crossbar 1130, and an L1.5 cache 1135.
[0087] In one embodiment, the GPC 1008 may be configured to execute a large number of threads in parallel to perform graphics, general processing, and / or computing operations. As used herein, a "thread" refers to an instance of a particular program executed for a particular set of input data. In some embodiments, single-instruction multiple-data (SIMD) instruction issue techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In other embodiments, using a common instruction unit configured to issue instructions to a set of processing engines within the GPC 1008, single-instruction multiple-thread (SIMT) techniques are used to support the parallel execution of a large number of generally synchronized threads. Different from the SIMD execution mechanism in which all processing engines typically execute the same instruction, SIMT execution allows different threads to more easily follow different execution paths through a given program. Those of ordinary skill in the art will understand that the SIMD processing mechanism represents a functional subset of the SIMT processing mechanism.
[0088] In one embodiment, the operation of the GPC 1008 is controlled via a pipeline manager 1105, which assigns processing tasks received from a work assignment unit (not shown) within the task / work unit 1007 to one or more streaming multi-processors (SMs) 1110. The pipeline manager 1105 may also be configured to control the work assignment crossbar 1130 by specifying the destination of the processed data output by the SM 1110.
[0089] In various embodiments, the GPC 1008 includes a set of M SMs 1110, where M ≥ 1. Moreover, each SM 1110 includes a set of functional execution units (not shown), such as execution units and load / store units. The processing operations specific to any functional execution unit can be pipelined, which enables new instructions for execution to be issued before the previous instruction has completed execution. Any combination of the functional execution units within a given SM 1110 may be provided. In various embodiments, the functional execution units may be configured to support a variety of different operations, including integer and floating-point arithmetic (e.g., addition and multiplication), comparison operations, Boolean operations (AND, OR, XOR), shifts, and operations on various algebraic functions (e.g., planar interpolation and trigonometric, exponential, and logarithmic functions, etc.). Advantageously, the same functional execution unit may be configured to perform different operations.
[0090] In various embodiments, each SM 1110 includes a plurality of processing cores. In one embodiment, the SM 1110 includes a large number (e.g., 128, etc.) of different processing cores. Each core may include fully pipelined, single-precision, double-precision, and / or mixed-precision processing units, which include floating-point arithmetic logic units and integer arithmetic logic units. In one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0091] In one embodiment, the tensor cores are configured to perform matrix operations, and in one embodiment, one or more tensor cores are included among the cores. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In one embodiment, each tensor core operates on a 4×4 matrix and performs matrix multiplication and accumulation operations D = A×B + C, where A, B, C, and D are 4×4 matrices.
[0092] In one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, while the accumulation matrices C and D can be 16-bit floating-point or 32-bit floating-point matrices. The tensor cores operate on 16-bit floating-point input data using 32-bit floating-point accumulation. 16-bit floating-point multiplication requires 64 operations and produces a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition for 4×4×4 matrix multiplication. In practice, the tensor cores are used to perform larger two-dimensional or higher-dimensional matrix operations, which are composed of these smaller elements. APIs such as the CUDA 9 C++ API expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to effectively use the tensor cores from CUDA-C++ programs. At the CUDA level, the warp-level interface assumes that a 16×16 size matrix spans all 32 threads of a warp.
[0093] Neural networks rely heavily on matrix mathematical operations, and complex multi-layer networks require a large amount of floating-point performance and bandwidth to improve efficiency and speed. In various embodiments, the SM 1110 has thousands of processing cores, optimized for matrix mathematical operations, and provides performance in the tens to hundreds of TFLOPS, thus providing a computing platform capable of providing the performance required for artificial intelligence and machine learning applications based on deep neural networks.
[0094] In various embodiments, each SM 1110 may also include a plurality of special function units (SFUs) that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, the SFU may include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, the SFU may include a texture unit configured to perform texture map filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texture pixels) from memory and sample the texture map to produce a sampled texture value for use in a shader program executed by the SM. In various embodiments, each SM 1110 also includes a plurality of load / store units (LSUs) that implement load and store operations between the shared memory / L1 cache and the register file within the SM 1110.
[0095] In one embodiment, each SM 1110 is configured to process one or more thread groups. As used herein, a "thread group" or "warp" refers to a group of threads that execute the same program on different input data, with one thread in the group being assigned to a different execution unit within the SM 1110. A thread group may include fewer threads than the number of execution units within the SM 1110, in which case some executions may be idle during the cycles in which the thread group is being processed. A thread group may also include more threads than the number of execution units within the SM 1110, in which case the processing may occur in consecutive clock cycles. Since each SM 1110 can support up to G thread groups simultaneously, up to G*M thread groups can be executed in the GPC 1008 at any given time.
[0096] Additionally, in one embodiment, multiple related thread groups may be simultaneously active (in different stages of execution) within the SM 1110. Such a collection of thread groups is referred to herein as a "cooperative thread array" ("CTA") or "thread array". The size of a particular CTA is equal to m*k, where k is the number of threads that execute simultaneously within the thread group, typically an integer multiple of the number of execution units within the SM 1110, and m is the number of thread groups that are simultaneously active within the SM 1110. In some embodiments, a single SM 1110 can support multiple CTAs simultaneously, where such CTAs are at the granularity of work assignment to the SM 1110.
[0097] In one embodiment, each SM 1110 includes a level 1 (L1) cache, or uses space in a corresponding L1 cache external to the SM 1110 to support (among other things) load and store operations performed by the execution units. Each SM 1110 may also access a level 2 (L2) cache (not shown) shared among all GPCs 1008 in the PPU 1002. The L2 cache may be used to transfer data between threads. Finally, the SM 1110 may also access off-chip "global" memory, which may include the PP memory 1004 and / or the system memory 904. It should be understood that any memory external to the PPU 1002 may be used as global memory. Additionally, as Figure 11 shown, a one-and-a-half level (L1.5) cache 1135 may be included in the GPC 1008 and is configured to receive and store data requested by the SM 1110 from memory via the memory interface 1014. Such data may include, but is not limited to, instructions, unified data, and constant data. In embodiments having multiple SM 1110s within a GPC 1008, the multiple SM 1110s may beneficially share common instructions and data cached in the L1.5 cache 1135.
[0098] In one embodiment, each GPC 1008 may have an associated memory management unit (MMU) 1120, which is configured to map virtual addresses to physical addresses. In various embodiments, the MMU 1120 may reside within the GPC 1008 or within the memory interface 1014. The MMU 1120 includes a set of page table entries (PTEs) that are used to map virtual addresses to the physical addresses of tiles or memory pages and optionally to cache line indices. The MMU 1120 may include an address translation lookaside buffer (TLB) or cache that may reside within the SM 1110, within one or more L1 caches, or within the GPC 1008.
[0099] In one embodiment, in graphics and compute applications, the GPC 1008 may be configured such that each SM 1110 is coupled to a texture unit 1115 for performing texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data.
[0100] In one embodiment, each SM 1110 sends the processed tasks to the work distribution crossbar 1130 to provide the processed tasks to another GPC 1008 for further processing, or store the processed tasks in the L2 cache (not shown), the parallel processing memory 1004, or the system memory 904 via the crossbar unit 1010. Additionally, the pre-raster operation (preROP) unit 1125 is configured to receive data from the SM 1110, direct the data to one or more raster operation (ROP) units within the partition unit 1015, perform color mixing optimization, organize pixel color data, and perform address translation.
[0101] It will be appreciated that the architectures described herein are illustrative and can be varied and modified. Among other things, any number of processing units can be included within the GPC 1008, such as the SM 1110, the texture unit 1115, or the preROP unit 1125. Additionally, as described above in connection with Figure 6 the PPU 1002 can include any number of GPCs 1008 configured to be functionally similar to each other such that the execution behavior does not depend on which GPC 1008 receives a particular processing task. Additionally, each GPC 1008 operates independently of the other GPCs 1008 in the PPU 1002 to perform tasks for one or more applications.
[0102] In summary, the unmasking engine removes masks from the pruned neural network represented by the node graph. The unmasking engine analyzes the tensors and masks associated with a given node in the node graph to determine the portions of the tensor that are zeroed out by the mask. The unmasking engine then removes these portions from the tensor to generate a dense tensor with dimensions smaller than the original tensor. Based on the dense tensor, the function associated with the node can be evaluated more quickly compared to the original tensor. The unmasking engine adds a scatter operation after the node to scale the dimensions of the dense tensor to the dimensions associated with the original tensor.
[0103] At least one technical advantage of the techniques described herein is that the optimized neural network performs inference operations faster than the original pruned neural network. Thus, the optimized neural network is well-suited for use in real-time applications such as autonomous vehicles. Another advantage of the techniques described herein is that the optimized neural network can have a smaller memory footprint compared to the pruned neural network, thereby saving memory resources. These technical advantages represent multiple technological advancements over prior art methods.
[0104] 1. Some embodiments include a computer-implemented method that includes: causing an unmasked output of a first neural network portion to be generated based at least in part on a masked output of the first neural network portion, wherein the unmasked output has a dimension less than a dimension of the masked output, causing the unmasked output to replace the masked output, and causing a scatter operation to be performed to expand the dimension of the unmasked output to a dimension corresponding to the masked output.
[0105] 2. The computer-implemented method according to clause 1, wherein the unmasked output is associated with a first tensor and the masked output is associated with a second tensor.
[0106] 3. The computer-implemented method according to any one of clauses 1-2, wherein causing the unmasked output to be generated includes: determining a first portion of the first tensor, the first portion of the first tensor corresponding to one or more zeros included in a first mask, wherein the masked output is derived based on the first tensor and the first mask, generating the second tensor based on the first portion of the first tensor, and evaluating a first function based on the second tensor to generate the unmasked output.
[0107] 4. The computer-implemented method according to any one of clauses 1-3, wherein the first mask zeros out the first portion of the first tensor, and wherein the first function is evaluated based on the first tensor to produce a first result independent of the first portion of the first tensor.
[0108] 5. The computer-implemented method according to any one of clauses 1-4, wherein the second tensor includes only a second portion of the first tensor.
[0109] 6. The computer-implemented method according to any one of clauses 1-5, wherein a processor evaluates the first function based on the second tensor faster than the processor evaluates the first function based on the first tensor.
[0110] 7. The computer-implemented method according to any one of clauses 1-6, wherein causing the unmasked output to replace the masked output includes: replacing the first tensor with the second tensor, wherein the second tensor has a dimension less than a dimension of the first tensor.
[0111] 8. The computer-implemented method according to any one of clauses 1-7, wherein causing the scatter operation to be performed includes: inserting one or more zeros into the unmasked output.
[0112] 9. The computer-implemented method according to any one of clauses 1-8 further includes: combining the dispersion operation with one or more additional dispersion operations associated with one or more neural network layers.
[0113] 10. The computer-implemented method according to any one of clauses 1-9 further includes: absorbing the dispersion operation into a second neural network portion located after the first neural network portion.
[0114] 11. Some embodiments include a non-transitory computer-readable medium that stores program instructions that, when executed by at least one processor, cause the at least one processor to at least: cause an unmasked output of the first neural network layer to be generated based at least in part on a masked output of the first neural network layer, where the dimension of the unmasked output is different from the dimension of the masked output, cause the unmasked output to replace the masked output, and cause a first operation to be performed to scale the dimension of the unmasked output to correspond to the dimension of the masked output.
[0115] 12. The non-transitory computer-readable medium according to clause 11, wherein the first operation includes a first dispersion operation that is performed to expand the dimension of the unmasked output to correspond to the dimension of the masked output.
[0116] 13. The non-transitory computer-readable medium according to any one of clauses 11-12 further includes combining the first dispersion operation with a second dispersion operation associated with a second neural network layer that is located after the first neural network layer in a sequence of neural network layers.
[0117] 14. The non-transitory computer-readable medium according to any one of clauses 11-13, wherein the first operation includes a first aggregation operation that is performed to reduce the dimension of the unmasked output to correspond to the dimension of the masked output.
[0118] 15. The non-transitory computer-readable medium according to any one of clauses 11-14 further includes: combining the first aggregation operation with a second aggregation operation associated with a second neural network layer that is located before the first neural network layer in a sequence of neural network layers.
[0119] 16. The non-transitory computer-readable medium according to any one of clauses 11-15, wherein causing the generation of the unmasked output includes: determining a first portion of a first tensor, the first portion of the first tensor corresponding to one or more zeros included in a first mask, wherein the masked output is derived based on the first tensor and the first mask, generating a second tensor based on the first portion of the first tensor, and evaluating a first function based on the second tensor to generate the unmasked output.
[0120] 17. The non-transitory computer-readable medium according to any one of clauses 11-16, wherein the second tensor includes only a second portion of the first tensor and does not include the first portion of the first tensor, and wherein evaluating the first function based on the second tensor is faster than evaluating the first function based on the first tensor.
[0121] 18. Some embodiments include a system that includes: a memory that stores one or more instructions, and a processor that executes the instructions to at least: cause an unmasked output of a first neural network layer to be generated at least partially based on a masked output of the first neural network layer, wherein a dimension of the unmasked output is less than a dimension of the masked output, cause the unmasked output to replace the masked output, and cause a scatter operation to be performed to expand the dimension of the unmasked output to a dimension corresponding to the masked output.
[0122] 19. The system according to clause 18, wherein the processor further executes the instructions to: combine the scatter operation with one or more scatter operations, wherein the one or more scatter operations include at least one dimension aligned with a corresponding dimension associated with the scatter operation.
[0123] 20. The system according to any one of clauses 18-19, wherein the processor further executes the instructions to: stack the scatter operation adjacent to one or more scatter operations, wherein the one or more scatter operations include at least one dimension not aligned with a corresponding dimension associated with the scatter operation.
[0124] In any way, any and all combinations of any claim elements recited in any claim and / or any elements described in this application fall within the intended scope of the present disclosure and protection.
[0125] The descriptions of the various embodiments have been given for purposes of illustration, but these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to a person of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0126] Aspects of the present embodiment may be embodied as a system, a method, or a computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects that may generally be referred to herein as a "module" or "system". In addition, aspects of the present disclosure may take the form of a computer program product embodying one or more computer-readable media having computer-readable program code embodied thereon.
[0127] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0128] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, enable the functions / acts specified in the flowchart and / or block diagram block to be implemented. Such a processor may be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0130] Although the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, and the scope of the present disclosure is determined by the appended claims.
Claims
1. A computer-implemented method, comprising: Generating an unmasked output of the first neural network portion at least partially based on a masked output of the first neural network portion, wherein a dimension of the unmasked output is less than a dimension of the masked output; Causing the unmasked output to replace the masked output; Causing a scatter operation to be performed to expand the dimension of the unmasked output to a dimension corresponding to the masked output.
2. The computer-implemented method according to claim 1, wherein the unmasked output is associated with a first tensor, and the masked output is associated with a second tensor.
3. The computer-implemented method according to claim 2, wherein causing the unmasked output to be generated comprises: Determining a first portion of the first tensor, the first portion of the first tensor corresponding to one or more zeros included in a first mask, wherein the masked output is derived based on the first tensor and the first mask; Generating the second tensor based on the first portion of the first tensor; And Evaluating a first function based on the second tensor to generate the unmasked output.
4. The computer-implemented method according to claim 3, wherein the first mask zeros out the first portion of the first tensor, and wherein the first function is evaluated based on the first tensor to produce a first result independent of the first portion of the first tensor.
5. The computer-implemented method according to claim 3, wherein the second tensor only includes a second portion of the first tensor.
6. The computer-implemented method according to claim 3, wherein a processor evaluates the first function based on the second tensor faster than the processor evaluates the first function based on the first tensor.
7. The computer-implemented method according to claim 2, wherein, Causing the unmasked output to replace the masked output comprises: replacing the first tensor with the second tensor, wherein a dimension of the second tensor is less than a dimension of the first tensor.
8. The computer-implemented method according to claim 1, wherein causing the decentralized operation to be performed comprises: Inserting one or more zeros into the unmasked output.
9. A non-transitory computer-readable medium storing program instructions that, when executed by at least one processor, cause the at least one processor to at least: Generate an unmasked output of the first neural network layer at least partially based on a masked output of the first neural network layer, wherein a dimension of the unmasked output is different from a dimension of the masked output; Cause the unmasked output to replace the masked output; Cause a first operation to be performed to scale the dimension of the unmasked output to a dimension corresponding to the masked output.
10. A system, comprising: A memory storing one or more instructions; And A processor that executes the instructions to at least: Generate an unmasked output of the first neural network layer at least partially based on a masked output of the first neural network layer, wherein a dimension of the unmasked output is less than a dimension of the masked output, Cause the unmasked output to replace the masked output, and Cause a scatter operation to be performed to expand the dimension of the unmasked output to a dimension corresponding to the masked output.