Analysis Techniques for Improved Supertiling Machine Learning Processing
Supertiling the layers of machine learning networks based on memory usage changes optimizes memory access and processing efficiency on devices with limited resources, addressing bottlenecks in executing complex ML techniques.
Patent Information
- Application Number
- JP2022578583
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-24
- Filing Date
- 2021-06-07
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2041-06-07
AI Technical Summary
Efficient execution of complex machine learning techniques, such as neural networks, on devices with limited computing and memory resources is challenging due to varying memory requirements across layers, leading to bottlenecks in processing time and efficiency.
The technique involves supertiling the layers of a machine learning network, grouping them based on memory usage changes, and optimizing memory access by processing tensor portions across multiple layers before moving to the next, utilizing on-chip memory efficiently.
This approach reduces the need for external memory access, enhances processing efficiency, and optimizes resource utilization on devices with limited resources.
Smart Images

Figure 0007698375000001 
Figure 0007698375000002 
Figure 0007698375000003
Abstract
Description
Technical Field
[0001] Machine learning (ML) has become an increasingly important part of the computing environment. Machine learning is a type of artificial intelligence (AI), and ML helps enable software systems to learn to recognize patterns from data without being directly programmed. A neural network (NN) is a type of ML that utilizes a set of linked and hierarchical functions (e.g., nodes, neurons, etc.) that are weighted to evaluate input data. Some NNs, sometimes called convolutional neural networks (CNNs), can perform a convolution operation in NN layers based on the received input and weights. The convolution operation is a mathematical transformation applied to two functions to produce a third function that represents how the shape of one function is modified by a second function. Examples of CNNs include inverse convolutional neural networks, pooling neural networks, upsampling neural networks, deep neural networks, etc., and CNNs are commonly used in a wide range of applications for recognition and classification, such as image recognition and classification, prediction and recommendation systems, speech and language recognition and translation.
[0002] As ML becomes increasingly useful, it is desirable to efficiently execute complex ML techniques, such as NNs and CNNs, on devices with relatively limited computing and memory resources, such as embedded or other low-power devices. To help efficiently execute a given ML model on target hardware resources, the ML model can be analyzed and optimized to be executed using supertiling to adjust the ML model for the target hardware resources to be used.
Summary of the Invention
[0003] The present disclosure relates to techniques for improving ML model execution. The techniques include determining an amount of memory used to process layers of a machine learning network having a plurality of layers, smoothing the amount of memory used to process the layers of the machine learning network based on the number of layers, identifying a change layer in which the smoothed amount of memory used changes more than a memory change threshold amount, grouping the layers of the machine learning network into a first layer grouping based on the identified change layer, and outputting the first layer grouping.
[0004] Another aspect of the present disclosure relates to a non-transitory program storage device storing instructions for causing one or more processors to perform the following acts. These acts include determining an amount of memory used to process layers of a machine learning network having a plurality of layers, smoothing the amount of memory used to process the layers of the machine learning network based on the number of layers, identifying a change layer in which the smoothed amount of memory used changes more than a memory change threshold amount, grouping the layers of the machine learning network into a first layer grouping based on the identified change layer, and outputting the first layer grouping.
[0005] Another aspect of the present disclosure relates to a device, the device including a memory and one or more processors operably coupled to the memory, the one or more processors being configured to execute non-transitory instructions for causing the one or more processors to perform the following acts. These acts include determining an amount of memory used to process layers of a machine learning network having a plurality of layers, smoothing the amount of memory used to process the layers of the machine learning network based on the number of layers, identifying a change layer in which the smoothed amount of memory used changes more than a memory change threshold amount, grouping the layers of the machine learning network into a first layer grouping based on the identified change layer, and outputting the first layer grouping.
Brief Description of the Drawings
[0006] For a detailed description of various examples, reference is now made to the accompanying drawings.
[0007]
Figure 1
[0008]
Figure 2
[0009]
Figure 3A
[0010]
Figure 3B
[0011]
Figure 4
[0012]
Figure 5A
Figure 5B
[0013]
Figure 6A
[0014]
Figure 6B
[0015]
Figure 7A
Figure 7B
[0016]
Figure 8
[0017]
Figure 9
Mode for Carrying Out the Invention
[0018] Figure 1 illustrates a data flow through an exemplary CNN 100 according to an aspect of the present disclosure. The CNN 100 shown here includes two layers, namely a first layer 102 and a second layer 104. It should be understood that although the CNN in this example includes two layers, other CNNs can include any number of layers. A layer represents a mathematical function applied to an input tensor and results in an output tensor. Examples of mathematical functions include convolution / transposed convolution functions, pooling, element-wise addition, concatenation, and the like. A tensor is a multi-dimensional generalized matrix that includes one or more nodes containing values. As an example, in the case of an image, a node can describe a pixel and can include the values of the x and y coordinates of the pixel, as well as the values of the R, G, and B channels that describe the color of the pixel. A tensor can have a height axis represented by HI, H2, H3 here corresponding to the dimensions of the image, a width axis W1, W2, and W3, and a channel axis represented by C1, C2, and C3 corresponding to color channel information (RGB information). In this example, a first tensor 106 is input to the first layer 102 along with a set of operation parameters 108 to generate a second tensor 110. Similarly, the second tensor 110 can be input to the second layer 104, processed based on operation parameters 112, and output a third tensor 114. The operation parameters 108 and 112 can include, for example, weights for applying to the processing of a given layer. Generally, an initial tensor such as the first tensor 106 is an input to the CNN 100, and the last tensor, here the third tensor 114, is an output from the CNN 100. A tensor between the input tensor and the output tensor, here the second tensor 110, may be referred to as an intermediate tensor.
[0019] In some cases, as shown by tensor 200 in FIG. 2, a tensor can be divided into tiles for processing, and the tiles can be sized, for example, based on the pipeline design of the processor. For example, a tile can include one or more nodes based on several parallel pipelines available on the processor. Note that hereafter, the tensor is shown as a two-dimensional structure for clarity. In a general implementation, all tiles of a given tensor are processed by a particular layer before processing starts on the next tensor and layer. For example, referring back to FIG. 1, the processing of the first tensor 106 in the first layer 102 can be completed for the entire first tensor 106, and can be output to the second tensor 110 before the processing of the second tensor 110 in the second layer 104.
[0020] Generally, it is advantageous to store as much information as possible required to execute a CNN in memory as close as possible to the processor to improve performance. Generally, memory close to the processor is sometimes called on-chip memory, while memory relatively far from the processor is sometimes called system memory, main memory, or random access memory (RAM), and even further memory is sometimes called storage, disk, or hard disk drive. Examples of on-chip memory include static random access memory (SRAM) and cache memory. Cache memory can be further divided into levels such as level 1 (L1), level 2 (L2), and level 3 (L3), where a larger number generally indicates that the cache is further away from the processor (e.g., slower access). As an example of processing an intermediate input tensor in the corresponding layer, the input tensor is stored in level 3 (L3) memory cache, while the weights, CNN model, as well as the input tile and output information are stored in level 2 (L2) cache. When a part of the tensor is processed, the output can be temporarily stored in the L2 cache, and then when the input tensor is processed, it can be output to another intermediate tensor, e.g., the L3 cache. When the next tensor is output to the L3 cache, the system is ready to process the next layer. In some cases, the initial input tensor and the final output can be stored in system memory. Storing and accessing intermediate tensors completely in the cache can take several clock cycles (e.g., processing cycles) when the processor may need to stall while waiting for data, reducing processing efficiency, and helps reduce the need to access external memory such as system memory like double data rate (DDR) memory.
[0021] The size of the memory can be fixed, but the size required by the intermediate tensors can vary. For example, a CNN can have an input tensor of half megabyte (MB) size and can be associated with two intermediate tensors of 5MB and 12MB respectively. For example, if the near-processor memory such as the L3 cache is only 8MB, the 12MB intermediate tensor cannot fit entirely within the L3 cache, and a portion of the 12MB intermediate tensor is likely to be stored in the system memory. Since memory access to the system memory is substantially longer than access to the cache memory, in this case, the processing time of the 12MB intermediate tensor is bottlenecked by the memory input / output time.
[0022] Figure 3A is a block diagram illustrating supertile processing 300 in accordance with an aspect of the present disclosure. Instead of processing the entire tensor through a layer before processing the next tensor and layer, a portion of the tensor can be processed as a supertile across multiple layers before the next supertile is processed. For example, as shown in FIG. 3, a first tensor 302 can be divided into three parts or supertiles, supertiles 304, 306, and 308. Supertile 304 can be processed in a first layer 310 to output supertile 304 which is a part of a second tensor 312. Similarly, supertile 304 of the second tensor 312 can then be processed in a second layer 314 to output supertile 304 of a third tensor 316. Thus, supertile 304 is processed across multiple layers before supertile 306 is processed. In this example, the supertiling is performed across the height axis or dimension. In other cases, supertiling can be performed on other axes such as the horizontal or vertical axis by removing values from one dimension of the tensor. After supertile 304 is processed by a set of layers, supertile 306 is processed by the set of layers. After the processing of supertile 306 is completed, supertile 308 is then processed by the set of layers.
[0023] In some cases, a portion of the input signal tensor is overwritten by the corresponding output that processes that portion of the input signal tensor. FIG. 3B is a block diagram illustrating a supertile processing resource usage 320 in accordance with an aspect of the present disclosure. This example illustrates an on-chip memory 322, a processor 324, and another memory 326. In this example, the memory 322 includes a first portion 328 of a first tensor. The first portion 328 can be, in this example, an intermediate tensor output from a previous layer (not shown). The first portion 328 can be processed in a first layer 330 with first ML network information 332 having model and / or weight information to generate a first layer output 334. The first output 334 is written back to the on-chip memory 322, overwriting the portion of the on-chip memory 322 that stored the first portion 328 to obtain a second portion 336 of a second tensor. In some cases, the second portion 336 may be a different size than the first portion 328. When the second portion 336 is smaller in size compared to the first portion 328, the remaining portion 338 of the first portion 328 can be discarded. In some cases, the output from the first layer 332 can be dynamically written on top of the corresponding portion of the first portion 328 within the on-chip 322 as the output is generated. Once generated, the second portion 336 is processed in a second layer 340 with second ML network information 342 to generate a second layer output 344 that is written back to the on-chip memory 322, overwriting the portion of the on-chip memory 322 that stored the second portion 336 to obtain a third portion 346 of a third tensor.
[0024] FIG. 4 illustrates supertile processing for a plurality of supertile paths 400 in accordance with an aspect of the present disclosure. This example includes a group of layers having at least four intermediate tensors, a first tensor 402A-402D, a second tensor 404A-404D, a third tensor 406A-406D, and a fourth tensor 408A-408D, which are shown here in a single dimension having 20 tiles, with other dimensions omitted for clarity. Layers are also omitted in this example. Note that since tensors 402-408 in this example are intermediate tensors, the first tensor 402 is the output tensor from a separate input tensor (not shown) and the corresponding layer. As described above, the first tensor 402 is input to the first layer to generate the second tensor 404, which is input to the second layer to generate the third tensor 406, which is input to the third layer to generate the fourth tensor 408. Four supertile paths are used to generate the complete fourth tensor 408, which can be input to another layer, e.g., another layer outside of this group of layers.
[0025] Each layer discussed in this example is a 3×3 convolutional layer. In a 3×3 convolutional layer, each tile is processed with one adjacent tile in each dimension of the layer. Each tensor includes two zero paddings represented by -1 and 20 entries. These zero paddings can be used as adjacent tiles when processing tiles on the edge of a given tensor. Here, at the end of each supertile path, the fourth tensor 408 has five completed tiles 410. Since each layer is a 3×3 convolutional layer, tile 5 of the third tensor 406A is used to generate tile 4 of the fourth tensor 408A. Similarly, tile 6 of the second tensor 404A is used to generate tile 5 of the third tensor 406A, and so on. After the first supertile path is completed, the second supertile path is implemented. Similar to the first supertile path, five completed tiles 412 are generated after the second supertile path is completed. As described in relation to FIG. 4, there may be overlapping regions between supertile paths. For example, tiles 4 and 5 for the third tensor 406B can be used to generate the five completed tiles 412 of the fourth tensor 408B. Tiles 4 and 5 of the third tensor 406B were previously calculated and stored in the first supertile path. When generating the third tensor 406B, tiles 4 and 5 of the third tensor 406B are reloaded rather than recalculated. Similarly, tiles 5 and 6 of the second tensor 404B and tiles 6 and 7 of the first tensor 402B can also be reloaded. In some cases, some tiles included within a supertile can vary across supertile paths. For example, in the case of the fourth supertile path, the first tensor 402D may have two tiles instead of eight tiles as in the case of other supertile paths. When the size of the tensor varies across layer groups, the size of the largest tensor can be used as part of determining the size of the supertile. In this example, since each preceding layer needs to calculate more tiles than the next one, the size and thus the memory space required to calculate the tiles of the first tensor 402A for the first path is a limiting factor for the size of the entire supertile.That is, the size of the supertile (e.g., tile height) can be selected to enable the calculations required for the first tensor 402A to fit into a memory such as an L3 cache in the first pass.
[0026] FIGS. 5A and 5B illustrate a supertile process 500 for multiple supertile paths across multiple supertile groups, in accordance with aspects of the present disclosure. Generally, a CNN can have any number of layers, and in some cases, a particular CNN may have more layers than can actually be executed as a single supertile. For example, a CNN having a relatively large input tensor and a relatively small output tensor may benefit from executing the layers of the CNN within multiple supertiles rather than a single supertile. In some cases, the layers of the CNN can be grouped into one or more layers grouped into each supertile group 502, and can be grouped into supertile groups 502A and 502B (collectively 502).
[0027] Each supertile group can be associated with certain supertile group characteristics. These supertile group characteristics can include characteristics such as the number of layers within the supertile group, the tile height associated with the layers, and context memory. In this example, the number of layers within the first supertile group 502A includes four layers 504, which are here layers 1, 2, 3, and 4. In this example, the second supertile group 502B also includes four layers 518, which are here layers 5, 6, 7, and 8. It can be understood that each supertile group can have a different number of layers. Each layer can be associated with multiple tile heights. In some cases, each layer can be associated with a first tile height, a normal tile height, and a final tile height. The first tile height can indicate the number of tiles for each layer during a first execution. In some cases, the first execution can be a virtual or preheat supertile path, which is here labeled as pass 0 506. The virtual supertile path may not generate tiles that are complete at the last tensor of the layer group. Instead, the virtual supertile path calculates a set of tiles that overlap with the tiles of the next normal supertile path and stores these (e.g., backed-up) calculated tiles for the next path. In this example, for the first layer, the first tile height is 3, the second layer is 2, the third layer is 1, and the fourth layer is 0.
[0028] The normal tile height can indicate the number of tiles for each layer during the steady-state execution of the supertile path, which is here labeled as pass 1 508, pass 2 510, and pass 3 512. In this example, the normal tile height for all layers is 5. It will be understood that the normal tile height for each layer can be different. The final tile height indicates the number of tiles for each layer for the last path of the supertile execution, which is here pass 4 514. In this example, the final tile height of the first layer is 2, the second layer is 3, the third layer is 4, and the fourth layer is 5.
[0029] The context memory supertile group characteristics refer to the stored or backed-up tiles 516 for a path. In this example, the context memory size is 6 tiles.
[0030] The supertile group and related supertile group characteristics can be defined for a CNN to help regulate the execution of the CNN for certain hardware resources. Each CNN can have a unique combination of the number of layers, the tensor dimensions of each layer, and what each layer is doing. For example, certain layers such as those implementing a pooling function, a convolution function, etc., can be associated with downsampling characteristics, where the layer takes an input tensor of a certain dimension and outputs a tensor with a reduced dimension. Other layers such as those implementing a resize function, a transposed convolution function, etc., can be associated with upsampling characteristics, where the layer takes an input tensor of a certain dimension and outputs a tensor with an increased dimension.
[0031] To help regulate the execution of a CNN for a given hardware resource, the CNN can be modeled to determine the total amount of memory (e.g., amount of memory) required for each layer of the CNN. This total amount of memory can include all the memory required to execute the layers of the CNN, including the input tensor, the output tensor, the backed-up tiles, the operation parameters required for the layer, etc. The supertile group can be defined based on this total amount of memory.
[0032] FIG. 6A is a line graph 600 plotting the total amount of memory used for each layer of a CNN according to an aspect of the present disclosure. In FIG. 6A, 64 layers 602 of the CNN are shown on the X-axis, and on the Y-axis, the total amount of memory 604 used per layer is shown in megabytes. In this example, the total amount of memory used by the layers of the CNN can vary quite significantly between layers. According to an aspect of the present disclosure, this local noise can be addressed by smoothing the sum value of the memory used across the layers within a window.
[0033] FIG. 6B is a line graph 650 plotting the windowed total amount of memory for a layer of a CNN according to an aspect of the present disclosure. Windowing is performed across the layers of the CNN to generate the windowed total amount data shown by plot 652. In some cases, the windowed total value for layer i can be the maximum total amount from layer i to layer i+W, where W is the window size. For example, in FIG. 650, the window size can be set to 8, and thus the windowed total amount for layer 1 is the maximum total value of layers 1-9. Referring back to line graph 600, layer 5 has the maximum total value of layers 1-9 at 25 MB, and thus the windowed total amount for layer 1 is 25 MB. As another example, for layer 6, the windowed total amount for layer 6 is the maximum total value of layers 6-14, or approximately 9 MB based on layers 8, 9, and 12. In some cases, W can be a predetermined value. For example, W can be a coded default value, received from a user, etc. In some cases, W can be determined based on one or more factors, e.g., as a function of the total number of layers in the CNN, the type of layer (e.g., convolution, deconvolution, pooling, etc.), as a function of some specific types of layers, based on the ordering of the layers, cost function, and modeling.
[0034] Based on the windowed total amount data, points where the total amount has changed by a certain amount, which can be called the volume change coefficient, can be identified. These identified points can be used to determine the initial boundaries of the supertiling groups. In the exemplary line graph 650, points can be identified between layers 5 and 6, layers 12 and 13, layers 24 and 35, and layers 49 and 50. In this example, there is a total amount change between layers 33 and 34 and between layers 54 and 55, but the total amount change at these points can be less than the volume change rate, and thus these points are not identified. Thus, five supertiling groups can be defined as including layers [1:5], [6:12], [13:24], [25:49], and [50:64]. When a relatively small volume change rate is used, additional supertiling groups such as [1:5], [6:12], [13:24], [25:49], [50:54], [55:64], or [1:5], [6:12], [13:24], [25:33], [34:49], [50:54], [55:64] can be defined. In some cases, the volume change rate can be pre-determined, for example, as a default value received from the user. In other cases, the volume change coefficient can be determined based on one or more factors, for example, based on cache or memory size, the maximum total amount across all layers, the ratio of the maximum total value to the minimum total value, etc. The volume change rate can be selected to balance noise reduction and the number of identified points. In some cases, multiple volume change rates can be used to determine multiple sets of supertiling groups for comparison, for example, via performance simulations (e.g., modeling).
[0035] After a super tiling group is identified, the super tiling group can be improved. In some cases, the super tiling group can be improved based on cost minimization implemented across super tiling group variations. For example, an initial super tiling group variation can be a super tiling group identified based on a total quantity change. A cost factor can be determined and associated with this initial super tiling group variation. This cost factor can be determined based on a performance simulation (e.g., modeling) of a CNN being executed using the initial super tiling group variation. The performance simulation can consider the memory access latency, processing speed, and power consumption of target hardware resources (e.g., for which the CNN execution on the hardware resources is optimized). The cost factor is then associated with the initial super tiling group variation. Variations of the super tiling group are determined by moving one or more group boundaries of the super tiling group within an improvement range N of the initial group boundaries. In some cases, the improvement range can be both positive and negative and can be relatively small. As an example, an initial group boundary 654 can be identified between layer 24 and layer 25 between the initial super tiling groups [13:24] and [25:33], with a range of N = 1. Two determined variations of the initial group boundary can then be [13, 23], [24, 33], and [13, 25], [26, 33], and these determined variations can then be evaluated via a performance simulation and associated with a cost factor. The variation with the relatively lowest cost factor can be selected as the final super tiling group configuration. In some cases, each group boundary of the initial group boundary can be improved. In some cases, one group boundary with a total quantity change above or below a certain threshold constant size can be improved. In some cases, such as when two super tiling groups are within each other's improvement ranges, the two super tiling groups can be merged.In some cases, different step sizes for the improvement range can be used, for example, adjusting the group boundary by two layers instead of one layer.
[0036] In accordance with aspects of the present disclosure, the tile height and the number of tiles can be configured for the supertiling group. In some cases, this determination can be based on backpropagation from the tile height of the last layer of the supertiling group, such as layer 4 in the example shown in FIG. 5. To determine the tile height via backpropagation, the volume of memory required for each layer can be determined. Based on the volume of memory required for each layer and the amount of memory available on the target hardware resources, the minimum number of tiles (e.g., passes) required to process the layer while maintaining the tile memory usage within the amount of memory available on the target hardware resources can be determined. Once the minimum number of tiles is determined for each layer, the maximum number of the minimum number of tiles for that layer is identified. In some cases, the number of tiles for the group of layers can be constant, except for the first and last passes. Based on this maximum number of the minimum number of tiles, the tile height of the last layer can be determined for the first pass, passes, and normal passes. Based on the tile height of the last layer, the tile height of the layer before the last layer can be determined. This process is then repeated until the tile height of the first layer is determined.
[0037] FIGS. 7A and 7B are flowcharts illustrating group boundary determination in accordance with aspects of the present disclosure. At block 702, the window size is determined. In some cases, the window size is determined They can be retrieved from, for example, memory. In some cases, the window size can be determined based on one or more factors such as the total number of layers of the CNN, the cost function, etc. In block 704, the total windowed amount of the layers of the CNN can be determined based on the size of the window. For example, a layer can have a total windowed amount based on the maximum sum of other layers within the number of windows of that layer. In block 706, the change in the total windowed amount between a layer and the next layer is compared to the volume change rate. In block 708, if the total windowed amount change is less than the volume change rate, in block 706, the next layer and the layers after the next layer are evaluated. If the total windowed amount change is greater than the volume change rate, in block 710, the boundary between the layers is marked as the initial supertile boundary. In block 712, if there are additional layers, the additional layers are looped through. In block 714, if there is an additional volume change rate to be considered, the layers of the CNN are looped through again using the additional volume change rate. In block 716, one or more sets of the marked initial supertile group boundaries can be output.
[0038] In block 718, if there is a set of unimproved supertile groups, in block 720, a CNN can be modeled to determine the cost coefficients of the supertile group boundaries within the improvement range. For example, the CNN can be modeled by executing the CNN with simulated inputs and using the supertile grouping to be modeled. The modeling can use simulated target hardware, such as by using a virtual machine, and record operation information such as memory usage, latency of the memory being used, processor usage, and power consumption. In some cases, each variant of the supertile group boundary within the improvement range can be simulated, and a cost factor can be associated with that variant. In block 722, the variant with the lowest cost factor among the variants of the supertile group boundary within the improvement range can be selected as the supertile group boundary. In block 724, if there are additional supertile group boundaries to be evaluated, the execution returns to 720 to evaluate those additional supertile group boundaries. If there are no more supertile group boundaries to be evaluated, the execution returns to 718. If there is no additional set of supertile groups to be evaluated in block 718, in block 726, if there are multiple sets of improved supertile groups, the cost factors across the multiple sets of improved supertile groups are compared, and in block 728, the set of improved supertile groups with the lowest cost factor is selected. Otherwise, in block 730, the improved supertile groups are output.
[0039] FIG. 8 is a flowchart showing a technique 800 for determining layer grouping according to an aspect of the present disclosure. In block 802, the amount of memory used to process the layers of a machine learning network having a plurality of layers is determined. For example, a CNN can be executed using simulated inputs to determine the memory usage by the layers of the CNN. In block 804, the amount of memory used to process the layers of the machine learning network can be smoothed based on the number of layers. For example, the amount of memory used to process the layers of a CNN can be smoothed using a window. The window can have a window size indicating the number of layers included in the window. In some cases, the smoothed amount of memory can be based on the maximum amount of memory used by any layer within the rolling window. In block 806, a layer in which the smoothed amount of memory used changes more than a memory change threshold amount is identified. For example, a point at which the smoothed amount of memory used changes more than a volume change coefficient can be identified as a boundary. In block 808, the layers of the machine learning network can be grouped into a first layer grouping based on the identified layers. For example, a super tiling group can be defined based on the identified boundary. In block 810, the first layer grouping is output.
[0040] As shown in FIG. 9, device 900 includes a processing element such as a processor 905 that includes one or more hardware processors, and each hardware processor can have a single or multiple processor cores. Examples of processors include, but are not limited to, a central processing unit (CPU) or a microprocessor. Although not shown in FIG. 9, the processing elements that make up processor 905 can also include one or more other types of hardware processing components such as a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), and / or a digital signal processor (DSP). In some cases, processor 905 can be configured to perform the tasks described in connection with FIGS. 7-8.
[0041] Processor 905 is operably and communicably coupled to on-chip memory 925 such as cache memory, SRAM, registers, etc. Regarding the cache memory, the cache memory may include one or more L1 caches, one or more L2 caches, and one or more L3 caches. The L1 cache may be integrated into the package together with the processor 905. The L2 and / or L3 caches may be integrated within the processor package or may be in a package separate from the processor package. In some cases, the L2 and / or L3 caches, or a portion thereof, may be integrated with a memory controller that helps manage memory traffic to the processor 905.
[0042] FIG. 9 illustrates that the memory 910 can be operably and communicably coupled to the processor 905. The memory 910 can be a non-transitory computer-readable storage medium (e.g., a non-transitory program storage device) configured to store various types of data. For example, the memory 910 may include one or more volatile devices such as random access memory (RAM). In some cases, SRAM and circuitry as described in FIGS. 4-8 may be part of the memory 910. The non-volatile storage device 920 (e.g., a non-transitory program storage device) may include one or more disk drives, optical drives, solid state drives (SSDs), tape drives, flash memories, electrically erasable programmable read-only memories (EEPROMs), and / or any other type of memory designed to maintain data over a period of time after a power loss or shutdown operation. The non-volatile storage device 920 may also be used to store programs that are loaded into RAM when such programs are executed.
[0043] Those skilled in the art will recognize that software programs can be developed, coded, and compiled in various computing languages for various software platforms and / or operating systems, and then loaded and executed by the processor 905. In one example, the compilation process of a software program can convert program code written in one programming language into another computer language so that the processor 905 can execute the programming code. For example, the compilation process of a software program may generate an executable program that operates the ML network.
[0044] After compilation, the encoded instructions can be loaded from storage 920, from memory 910, into processor 905 as computer-executable instructions or processing steps, and / or can be embedded within processor 905 (e.g., via a cache or on-board ROM). Processor 905 can be configured to execute the stored instructions or processing steps to perform instructions or processing steps for converting a computing device into a non-generic, and in particular, a specially programmed machine or device. Stored data, e.g., data stored by storage device 920, can be accessed by processor 905 during the execution of computer-executable instructions or processing steps to instruct one or more components within computing device 900. Storage 920 can be partitioned or divided into multiple sections that can be accessed by different software programs. For example, storage 920 can include sections designated for specific purposes, such as storing program instructions or data for updating the software of computing device 900. In one example, the software to be updated includes the ROM or firmware of the computing device. In some cases, computing device 900 includes multiple operating systems. For example, computing device 900 can include a general-purpose operating system for normal operation. Computing device 900 can also include another operating system, such as a boot loader, to perform specific tasks such as upgrading and recovering the general-purpose operating system and to enable access to computing device 900 at a level not generally available via the general-purpose operating system. Both the general-purpose operating system and the other operating system can access sections of storage 920 designated for specific purposes.
[0045] One or more communication interfaces may include a wireless communication interface for interfacing with one or more wireless communication devices. In some cases, elements coupled to the processor may be included on hardware shared with the processor. For example, communication interface 925, storage 920, and memory 910 may be included in a single chip or package, such as a system-on-chip (SOC), along with other elements such as digital radio. The computing device may also include input and / or output devices not shown, examples of which include devices for human input such as sensors, cameras, mice, keyboards, touchscreens, monitors, display screens, tactile or motion generators, speakers, lights, etc.
[0046] As used herein, the term "coupled" may encompass a connection, communication, or signal path that enables a functional relationship consistent with this specification. For example, if device A generates a signal for controlling device B to perform an action, (A) in a first example, device A is coupled to device B by a direct connection, or (b) in a second example, when intervening component C does not change the functional relationship between device A and device B, device A is coupled to device B via intervening component C, and device B is controlled by device A via a control signal generated by device A.
[0047] Within the scope of the claims of the present invention, modifications may be made to the illustrated exemplary embodiments, and other embodiments are possible.
Claims
1. A method comprising: determining, by a processing element, an amount of memory used to process a layer of a machine learning network having a plurality of layers; smoothing, by the processing element, the amount of memory used to process the layers of the machine learning network based on the number of layers; identifying, by the processing element, a changing layer, wherein the smoothed amount of memory used changes by more than a memory change threshold amount; grouping, by the processing element, the layers of the machine learning network into a first layer grouping based on the identified changing layer; outputting, by the processing element, the first layer grouping; A method, comprising the above steps.
2. The method according to claim 1, further comprising: modeling, by the processing element, the machine learning network based on the first layer grouping; associating, by the processing element, a first cost with the first layer grouping; generating, by the processing element, a second layer grouping by adjusting a group boundary of the first layer grouping; modeling, by the processing element, the machine learning network based on the second layer grouping; associating, by the processing element, a second cost with the second layer grouping; outputting, by the processing element, a lower cost layer grouping based on a comparison between the first cost and the second cost; A method, further comprising the above steps.
3. The method according to claim 2, wherein the first and second costs are based on at least one of an expected number of memory accesses or processing cycles.
4. The method according to claim 2, wherein the group boundary is adjusted within a range of predetermined values around the group boundary.
5. The method according to claim 1, wherein the first layer grouping includes a first set of layers and a second set of layers.
6. The method according to claim 5, wherein a number of first layers in the first set of layers is different from a number of second layers in the second set of layers.
7. The method according to claim 1, further comprising: determining, by the processing element, a minimum number of tiles for the layers of the first layer grouping based on an amount of memory used by the layers. The processing element determines the number of tiles for the last layer of the first layer grouping based on the minimum number of tiles of the tile; The processing element determines the number of tiles of other layers of the first layer grouping based on the number of tiles for the last layer; The method further includes.
8. A non-transitory program storage device, to one or more processors, determining the amount of memory used to process the layers of a machine learning network having a plurality of layers; smoothing the amount of memory used to process the layers of the machine learning network based on the number of layers; identifying a changing layer in which the smoothed amount of memory used changes more than a memory change threshold amount; grouping the layers of the machine learning network into a first layer grouping based on the identified changing layer; outputting the first layer grouping; A non-transitory program storage device including stored instructions for causing the above.
9. The non-transitory program storage device according to claim 8, wherein the instructions cause the one or more processors to model the machine learning network based on the first layer grouping; associating a first cost with the first layer grouping; generating a second layer grouping by adjusting the group boundary of the first layer grouping; modeling the machine learning network based on the second layer grouping; associating a second cost with the second layer grouping; outputting a lower cost layer grouping based on a comparison between the first cost and the second cost; The non-transitory program storage device further causing the above.
10. The non-transitory program storage device according to claim 9, wherein the first and second costs are based on at least one of an expected number of memory accesses or processing cycles.
11. The non-transitory program storage device according to claim 9, wherein the group boundary is adjusted within a predetermined value range around the group boundary.
12. The non-transitory program storage device according to claim 8, A non-transitory program storage device in which the first layer grouping includes a set of first layers and a set of second layers. **Claim 13** The non-transitory program storage device according to claim 12, wherein the number of first layers in the set of first layers is different from the number of second layers in the set of second layers. **Claim 14** The non-transitory program storage device according to claim 8, wherein the instructions cause the one or more processors to determine a minimum number of tiles for the layers of the first layer grouping based on an amount of memory used by the layer, determine a number of tiles for the last layer of the first layer grouping based on the minimum number of tiles, determine a number of tiles for other layers of the first layer grouping based on the number of tiles for the last layer, and further cause the one or more processors to perform the above operations. **Claim 15** A device comprising: a memory; one or more processors operably coupled to the memory, the one or more processors being configured to determine an amount of memory used to process layers of a machine learning network having a plurality of layers, smooth the amount of memory used to process the layers of the machine learning network based on the number of layers, identify a changing layer in which the smoothed amount of memory used changes more than a memory change threshold amount, group the layers of the machine learning network into a first layer grouping based on the identified changing layer, output the first layer grouping, and execute non-transitory instructions to cause the one or more processors to perform the above operations. The device further includes the one or more processors. **Claim 16** The device according to claim 15, wherein the instructions cause the one or more processors to model the machine learning network based on the first layer grouping, associate a first cost with the first layer grouping, generate a second layer grouping by adjusting a group boundary of the first layer grouping, model the machine learning network based on the second layer grouping, associate a second cost with the second layer grouping, Outputting a lower-cost layer group based on a comparison between the first cost and the second cost; A device that further causes this.
17. The device according to claim 16, wherein the first and second costs are based on at least one of the expected number of memory accesses or processing cycles.
18. The device according to claim 16, wherein the group boundary is adjusted within a range of a predetermined value around the group boundary.
19. The device according to claim 15, wherein the first layer grouping includes a first set of layers and a second set of layers.
20. The device according to claim 15, wherein the instructions cause the one or more processors to determine a minimum number of tiles for the layers of the first layer grouping based on the amount of memory used by the layers, determine the number of tiles for the last layer of the first layer grouping based on the minimum number of tiles, and determine the number of tiles for the other layers of the first layer grouping based on the number of tiles for the last layer. A device that further causes this.
Citation Information
Patent Citations
Methods and arrangements to manage memory in cascaded neural networks
US20190042925A1
Scheduling neural network processing
WO2018212799A1