Method and system for generating sub-network distribution of deep neural network
By dividing a deep neural network into sub-networks in an embedded system and optimizing performance parameters using the hardware system's topology and runtime statistics, the problem of resource constraints in embedded systems is solved, and efficient machine learning inference and computational performance optimization are achieved.
Patent Information
- Application Number
- CN202480015574.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-08
- Filing Date
- 2024-03-07
- Publication Date
- 2025-10-31
AI Technical Summary
Embedded systems have limited resources, making it difficult to directly execute complex deep neural network calculations. Furthermore, the latency and memory limitations caused by distributing to cloud computing restrict the computing power of embedded devices.
By receiving topology and performance information of the hardware system, the deep neural network is divided into multiple sub-networks using runtime statistics, and the sub-networks are adjusted on the computing nodes of the hardware system to optimize performance parameters such as latency, energy consumption, and memory usage.
It enables efficient execution of machine learning inference in embedded systems, dynamically adapts to hardware requirements, optimizes computing performance and resource utilization, and reduces latency and energy consumption.
Smart Images

Figure CN120883218A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for generating subnetwork distributions of a deep neural network. Specifically, this invention relates to a method for dividing a deep neural network into subnetwork distributions based on runtime statistics of the hardware system and system requirements. Background Technology
[0002] Machine learning inference is defined as the process of taking an input and producing an output through a machine learning model. The input can be, for example, an image, and the output can be a model prediction of what the image depicts. The model can be provided, for example, in the form of a deep neural network (DNN).
[0003] There is growing interest in using general-purpose machine learning models, and in particular, DNNs, in a variety of applications such as embedded systems. However, due to the limited resources of embedded systems, many DNNs are too demanding to be executed directly on embedded devices or systems, and inference is instead performed in the cloud where computation can be carried out in powerful computing clusters.
[0004] Embedded systems (such as mobile phones and smart camera devices) are typically limited by a number of resources—battery, computing speed, and memory. Driven by advancements in areas such as autonomous driving, embedded devices are becoming increasingly powerful—some even featuring dedicated computing units for DNN inference. Therefore, it is sometimes no longer necessary to distribute computing to the cloud.
[0005] Currently, performing DNN computations locally may be beneficial, as the additional latency of sending data to the cloud may outweigh any benefits of performing computations remotely. Memory limitations in embedded systems can be overcome by stacking multiple devices together, with each device hosting a portion of the computation. This again negates the benefits of distributing computation to the cloud. This defines a new era in which embedded devices begin to rival the capabilities of central computers.
[0006] However, many challenges remain to be overcome in order to efficiently perform machine learning inference in embedded systems, and further development is required. Therefore, further development of methods for running machine learning models in systems that include multiple hardware resources, such as embedded systems, is desired. Summary of the Invention
[0007] In view of the above-mentioned and other disadvantages of the prior art, the object of the present invention is to provide an improved method for generating subnetwork distributions of deep neural networks for hardware systems including multiple computing nodes.
[0008] According to a first aspect of the invention, a computer-implemented method is provided for generating sub-network distributions of a deep neural network for a hardware system comprising multiple computing nodes. The method includes: in a processor device: receiving a deep neural network; receiving topology and performance information of the hardware system; receiving runtime statistics of the hardware system; receiving at least one target performance parameter for the execution of the deep neural network on the hardware system; for each target performance parameter, dividing the deep neural network into at least one sub-network distribution based on the runtime statistics of the hardware system and based on the target performance parameter, wherein each sub-network in the sub-network distribution is adapted to execute on a computing node of the hardware system; and for each sub-network distribution, determining a performance metric for the execution of the deep neural network on the hardware system for the at least one target performance parameter.
[0009] This invention is based on the understanding that by partitioning deep neural networks intended to execute on a hardware system based on known attributes of hardware resources and runtime statistics of the hardware system, the inference of the deep neural network can be optimized for selected performance parameters and for a given hardware system. Therefore, runtime statistics are hardware-dependent, while performance parameters are related to the execution of a particular deep neural network on a specific hardware system.
[0010] Furthermore, the described method allows for local optimization for multiple different performance parameters, enabling the dynamic selection of appropriate subnetwork distributions based on the time-point requirements and runtime performance of the deep neural network when it is executed on the hardware system.
[0011] In this context, a subnetwork distribution is described as multiple subnetworks that together form a deep neural network. This method involves dividing a deep neural network into multiple subnetworks, each configured to: receive input; and provide output, i.e., inference or partial inference. Input can be received from another subnetwork or from the functionality of the hardware system in which the deep neural network is executed. It should be noted that, depending on the requirements, a subnetwork distribution may not include the functionality of the complete deep neural network but may still provide inference. In other words, a subnetwork distribution does not need to include all the features, functions, and / or computations of a complete deep neural network.
[0012] Furthermore, each unique partition (distribution) of a deep neural network is considered a separate subnetwork distribution. Therefore, even if multiple subnetworks in two subnetwork distributions are the same, the distribution is considered a unique subnetwork distribution if the content of the subnetworks is different.
[0013] The inclusion of multiple computing nodes in this hardware system should be understood as implying the presence of multiple hardware resources capable of executing subnetworks of a deep neural network. A computing node can be, for example, a single computer, one of several processors / accelerators on a PCB board, a digital signal processor (DSP) in a chip, a neural processing unit in a field-programmable array (FPGA), a coprocessor, an accelerator, etc. Furthermore, a computing node can also be defined as a time slot on a processor device, such that a single processor device can be described as multiple computing nodes based on the available time slots, or where each time slot is considered a computing node; that is, a computing node can be referred to as a time slice. A computing node can also be viewed as a process running on a processing device, such that a single processing device can include multiple computing nodes that can be utilized in parallel or at least independently. A computing node can also be one or more cores in a multi-core processing device.
[0014] The topology of a hardware system describes which hardware resources are available within the system, and further, it describes the connections between these resources. Therefore, deep neural networks can be partitioned not only based on the properties of the computing nodes themselves, but also on the properties of the connections between them, allowing for adjustments to the communication between subnetworks based on a known topology. Furthermore, the performance information of the hardware system can be considered an inherent property of the hardware system.
[0015] Runtime statistics can be generated through measurements of the hardware system's components and compute nodes during the initial benchmarking step of executing a deep neural network by the hardware system, or some or all of the runtime statistics can be derived through simulation or other methods of estimating hardware performance (e.g., using model-based prediction). Runtime statistics can be collected and stored as lookup tables in a database, and / or runtime statistics can be embedded into simulation or estimation models. Advantageously collecting runtime statistics from different configurations of the deep neural network allows them to be used to estimate how modifications to the deep neural network affect the network's execution parameters, i.e., performance parameters. Based on the runtime statistics, the optimization engine knows how changes to the deep neural network will affect the target performance parameters described in the requirements. The optimization engine can be described as and / or include a performance predictor that can predict how modifications to the subnetwork distribution will affect the performance parameters.
[0016] Adjusting a subnetwork to execute on a hardware system's compute nodes means modifying the deep neural network to improve its performance on those nodes. Such adjustments can include, for example, modifying network operational parameters such as the number of filters in convolution operators. Adjustments can also include, for example, reducing the kernel size of convolution operations, removing one or more operations from the network, sparsifying the network by inserting zeros into the network weight tensors, or otherwise altering the architecture to improve hardware performance. Thus, adjusting a subnetwork means that the split points between layers in the network can be changed, and the layers themselves can be reconfigured.
[0017] Performance metrics describe the measured, simulated, or predicted performance of a deep neural network on a hardware system, allowing comparisons of different sub-network distributions for different performance parameters. A basic performance metric could be, for example, the total execution time of a sub-network distribution. Other performance metrics can describe energy consumption, memory usage, or the accuracy of deep neural network inference. Furthermore, performance metrics can be defined for the entire hardware system or for one or more individual compute nodes.
[0018] According to an example implementation, the performance information of the hardware system includes at least one of the following: memory size of computing nodes, memory availability of computing nodes, time availability of computing nodes, computing resources of computing nodes, latency of computing nodes, and latency between computing nodes. Memory availability and computing resources of computing nodes can describe the amount of memory and computing resources of a specific computing node that can be used for execution and / or storage in the subnetwork, and similarly, time availability of computing nodes can describe the size / duration of a computing node's time slot or other means of resource sharing. Furthermore, the latency of computing nodes describes the time spent performing certain operations, and latency between computing nodes can include both node response time and the time required to transfer the required amount of data between nodes. Performance information may also include computer node latency for different types of operations. Furthermore, it should be noted that the topology and performance information received by the hardware system preferably only describes the resources available for execution of the deep neural network, and thus the performance information received by the hardware system does not necessarily include all resources or components of the hardware system. Therefore, the performance information of the hardware system describes the capabilities of the hardware system based on parameters directly linked to the hardware system. Even if some performance parameters can be determined or estimated based on runtime statistics, these are different because runtime statistics are related to the execution of the deep neural network on the hardware system.
[0019] According to an example implementation, the target performance parameter is at least one of latency, energy consumption, runtime memory usage, data size when stored on compute nodes, number of computations used for inference (e.g., FLOPs), throughput, and accuracy. Thus, the deep neural network is partitioned considering at least one target performance parameter in conjunction with the hardware system topology. The objective could be, for example, to execute the deep neural network's inference as quickly as possible; in this case, partitioning into subnetworks aims to optimize the latency of the subnetwork distribution, i.e., the total execution time. Objectives in forming the subnetwork distribution could also include optimizing the subnetwork distribution for several performance parameters, or finding solutions that achieve favorable trade-offs between different performance parameters. Performance parameters can be described at the system level, for individual compute nodes, or a combination of both. Here, throughput can be described, for example, as the number of operations per unit of time that can be improved by implementing parallel operations.
[0020] According to an example implementation, dividing a deep neural network into multiple subnetworks includes: receiving a previous performance metric of a target performance parameter for a first subnetwork distribution executed on a hardware system; and dividing the deep neural network into a second subnetwork distribution based on runtime statistics describing the performance of the hardware system, thereby improving the previous performance metric. Therefore, the described method is an iterative process in which the first subnetwork distribution can be formed and tested, for example, through simulation, to determine the performance metric. Based on known runtime statistics, how changes to the subnetwork distribution will affect the performance metric can be anticipated, and therefore, changes are not performed randomly but based on knowledge of the hardware system. An iterative optimization process can be performed until no further improvement to the performance metric can be found, or any improvement is minor and / or detrimental to other performance parameters.
[0021] According to an example implementation, the operation of dividing a deep neural network into at least one sub-network distribution includes: accessing a lookup table that includes runtime statistics; and, for a target performance parameter, determining whether a change in the sub-network distribution improves a performance metric for the performance parameter. The use of a lookup table provides an efficient way to determine how a change in the sub-network distribution affects the performance parameter. The lookup table may, for example, include information about the performance of network components. For example, the network may have convolutions using a specific number of filters, and by querying the lookup table, it can be determined whether using a different number of filters makes the convolutions faster. Alternatives to the lookup table are measuring the target performance parameter on a hardware system using a hardware simulator, or predicting the target performance using a performance predictor model. Measuring the performance parameter can be more accurate, but is generally more time-consuming. Predictions may be effective depending on the complexity of the model used, but are also time-consuming and / or have lower accuracy.
[0022] According to an example implementation, partitioning a deep neural network includes adjusting the size of the subnetworks to not exceed the available storage space of the corresponding compute nodes. Furthermore, partitioning the deep neural network includes adjusting the operations of the subnetworks based on the runtime environment of the corresponding compute nodes. Different runtime environments have different properties and thus can perform certain operations in different ways and with different latency and efficiency. The runtime environment can also set limits on the amount of available runtime memory, which is then taken into account when partitioning the deep neural network. Therefore, it is advantageous to form the subnetworks using knowledge of the runtime environment of the compute nodes on which the subnetworks will be executed. Examples of runtime environments include NVIDIA TensorRT, ONNXRuntime, Tensorflow Lite, and Arm NN.
[0023] According to an exemplary embodiment of the present invention, dividing a deep neural network into at least one subnetwork distribution includes pruning, performing Neural Architecture Search (NAS), and / or performing quantization of the deep neural network and / or its subnetworks. The foregoing actions describe different ways of modifying a deep neural network, but other methods may also be used.
[0024] According to an example implementation, the method further includes: providing a plurality of subnetworks to a scheduling node by a first processor device; receiving a target performance parameter by the scheduling node; comparing the target performance parameter with corresponding performance metrics of the target performance parameters of the plurality of subnetwork distributions by the scheduling node; and selecting, by the scheduling node, the subnetwork distribution with the best performance metric for the received target performance parameter from the plurality of subnetwork distributions. The best performance metric is selected based on one or more target performance parameters. For example, if the target performance parameter is latency, the subnetwork with the lowest latency can be considered the subnetwork distribution with the best performance metric. Similarly, if the target performance parameter is throughput, the subnetwork with the highest throughput can be considered the subnetwork distribution with the best performance metric. Furthermore, the best performance metric can also be the performance metric that is closest to a specific threshold above or below a given target performance parameter. Therefore, an additional condition when selecting a subnetwork distribution could be that the selected performance metric is not allowed to be above or below a given threshold. The subnetwork distribution can also be optimized based on a combination of multiple target performance parameters, such as when inference must be performed within a specified time period without consuming more than a given amount of energy.
[0025] Scheduling nodes are typically components or compute node parts of a hardware system, or components or compute node parts arranged to connect to the hardware system. Scheduling nodes receive multiple sub-network distributions and performance metrics of different performance parameters for each sub-network distribution, enabling them to select the sub-network distribution that best matches the current needs of the hardware system when a request to execute a deep neural network originates from the hardware system.
[0026] According to a second aspect of the invention, a computer system including a processing circuitry is also provided, the processing circuitry being configured to: receive a deep neural network; receive topology and performance information of a hardware system; receive runtime statistics of the hardware system; receive at least one target performance parameter for the execution of the deep neural network on the hardware system; for each target performance parameter, divide the deep neural network into at least one subnetwork distribution based on the runtime statistics of the hardware system and based on the target performance parameter, wherein each subnetwork in the subnetwork distribution is adapted to execute on a computing node of the hardware system; and for each subnetwork distribution, determine a performance metric for the execution of the deep neural network on the hardware system for at least one target performance parameter.
[0027] The additional effects and features of the second aspect of the invention are largely similar to those described above in conjunction with the first aspect of the invention.
[0028] Other features and advantages of the invention will become apparent when examined in conjunction with the appended claims and the following description. Those skilled in the art will recognize that different features of the invention can be combined to create embodiments other than those described below, without departing from the scope of the invention. Attached Figure Description
[0029] These and other aspects of the invention will now be described in more detail with reference to the accompanying drawings, which illustrate exemplary embodiments of the invention, in which:
[0030] Figure 1 This is a flowchart outlining the steps of a method according to an example implementation;
[0031] Figure 2 The features of the system and method according to embodiments of the present invention are illustrated schematically;
[0032] Figure 3 The features of the system and method according to embodiments of the present invention are illustrated schematically;
[0033] Figure 4 The features of the system and method according to embodiments of the present invention are illustrated schematically;
[0034] Figure 5 An example implementation of the system and method according to embodiments of the present invention is illustrated schematically;
[0035] Figure 6 An example implementation of the system and method according to embodiments of the present invention is illustrated schematically;
[0036] Figure 7An example implementation of the system and method according to embodiments of the present invention is illustrated schematically. Detailed Implementation
[0037] In this detailed description, various embodiments of the methods and systems according to the present invention will be described.
[0038] Figure 1 This is a flowchart outlining the steps of a method for generating a subnetwork distribution of a deep neural network according to an example implementation, and the method will be further referred to Figure 2 To describe, the Figure 2 Features of the system and method according to embodiments of the present invention are illustrated schematically. Furthermore, Figure 3 A hardware system 300 comprising multiple computing nodes 302a to 302h is schematically illustrated, on which a deep neural network will execute. Each computing node represents a set of computing resources (CPU, RAM, HDD) associated with actual physical hardware. The computing node also includes the temporal availability of resources. Thus, a computing node can be defined by a specific set of hardware resources available during a specific time period.
[0039] A computer-implemented method for generating subnetwork distributions of a deep neural network includes, in a processor device 200 (CPU): receiving a deep neural network 202; receiving topology 204 and performance information 206 of a hardware system 300; receiving runtime statistics 208 of the hardware system 300; receiving at least one target performance parameter 210 for the execution of the deep neural network on the hardware system 300; for each target performance parameter 210, dividing the deep neural network into at least one subnetwork distribution 212a to 212c based on the runtime statistics 208 of the hardware system 300 and based on the target performance parameter 210, wherein each subnetwork in the subnetwork distribution is adapted to execute on a computing node of the hardware system; and for each subnetwork distribution, determining performance metrics 214a to 214c for the execution of at least one target performance parameter of the deep neural network 202 on the hardware system 300.
[0040] Processor device 200 may include a microprocessor, a microcontroller, a programmable digital signal processor, or another programmable device. The processor device may also, or alternatively, include an application-specific integrated circuit (ASIC), a programmable gate array (FPGA) or programmable array logic, a programmable logic device, or a digital signal processor. Where the processor device includes a programmable device such as a microprocessor, microcontroller, or programmable digital signal processor as mentioned above, the processor device may also include computer-executable code for controlling the operation of the programmable device.
[0041] Processor device 200 may also include or be coupled to nontransitory computer-readable storage media such as storage devices, which may include, for example, internal or external hard disk drives (HDDs) (e.g., Enhanced Integrated Drive Electronic Devices (EIDE) or Serial Advanced Technology Accessories (SATA)), HDDs for storage (e.g., EIDE or SATA), flash memory, etc. Storage devices and other drives associated with computer-readable media and computer-usable media can provide non-volatile storage of data, data structures, computer-executable instructions, etc.
[0042] The processor device can receive the described information from any device or service that communicates with it (such as internal, external, or remote data storage, cloud computing environments, etc.). Those skilled in the art will readily recognize that the described method can be implemented in any suitable software environment, where the steps of the method can be initiated and executed using manual or automated input.
[0043] like Figure 2 As schematically shown, each subnetwork distribution 212a to 212c is a unique partition of the deep neural network. Although Figure 2 Each subnetwork distribution 212a to 212c has a unique topology, but it should be noted that even if two unique subnetwork distributions have the same partitioning / topology, they can be distinguished by the type of operation performed in a given subnetwork.
[0044] Furthermore, each subnetwork distribution 212a to 212c includes at least one corresponding performance metric 214a to 214c, such that, based on the target performance parameters, a distribution with the optimal performance metric for the selected performance parameters can be selected. The performance information of the hardware system 300 includes at least one of the following: memory size of computing nodes, memory availability of computing nodes, time availability of computing nodes, latency of computing nodes, and latency between computing nodes. Computing nodes may, for example, include processor devices with known memory size and known processing speed, wherein the memory size can be either static memory size or dynamic memory availability. The performance information, together with the topology of the hardware system 300 describing, for example, the available connections between nodes, provides the information needed to optimize the subnetworks of the respective hardware nodes for one or more target performance parameters (including latency, energy consumption, memory requirements, and accuracy). In practice, target performance parameters can be defined in many different ways and are defined as combinations of target performance parameters. Target performance parameters can, for example, refer to objectives for the execution of the deep neural network as a whole, such as total energy efficiency or total execution time, but it is also possible to specify target performance parameters for individual computing nodes.
[0045] According to the example, dividing a deep neural network into multiple subnetworks includes: receiving a prior performance metric of the target performance parameters of a first subnetwork distribution executed on a hardware system; and dividing the deep neural network into a second subnetwork distribution based on runtime statistics describing the hardware system performance, thereby improving the prior performance metric. Thus, the method relates to an iterative approach for optimizing the subnetwork distribution for one or more performance parameters. The optimization process may also include specific desired trade-offs between different performance parameters or between the weights of different performance parameters.
[0046] Dividing a deep neural network into at least one sub-network distribution may also include: accessing a lookup table that includes runtime statistics; and determining, for a target performance parameter, whether the change in the sub-network distribution improves the performance metric. Thus, the lookup table can provide a quick and direct way to determine whether a change in the sub-network distribution will result in a performance improvement, without having to simulate or perform the execution of the deep neural network.
[0047] Partitioning a deep neural network can also include resizing the subnetworks to a size that does not exceed the available storage space of the corresponding compute nodes. Resizing can include operations such as pruning to remove parts of the deep neural network that may not be needed for a given application, and also includes regular compression, which can be performed as part of generating the subnetworks, but may also be performed by the compute nodes themselves before the data packets are transmitted to subsequent compute nodes. The operations on the subnetworks can also be modified based on the runtime environment of the corresponding compute nodes.
[0048] Furthermore, dividing a deep neural network into at least one sub-network distribution may include performing Neural Architecture Search (NAS) and / or performing quantization of the deep neural network. The sub-network distribution may also be generated based on sequential execution, parallel execution, or a combination thereof.
[0049] According to an example implementation, the method further includes: providing a plurality of subnetworks to a scheduling node by a first processor device; receiving target performance parameters by the scheduling node; comparing the target performance parameters with corresponding performance metrics of the target performance parameters of the plurality of subnetwork distributions by the scheduling node; and selecting, by the scheduling node, the subnetwork distribution that has the best performance metric for the received target performance parameters from the plurality of subnetwork distributions. Therefore, the scheduling node can be any one of nodes 302a to 302h of the hardware system 300. Therefore, the scheduling node can also be a separate computing node separate from the hardware system.
[0050] In an exemplary scenario, the target performance parameter could be total execution time. The scheduling node then receives a request to execute the deep neural network, and the execution that produces inference should be completed within the requested total execution time. The scheduling node then moves to the performance metric "total execution time" of different sub-network distributions and selects the sub-network distribution with the lowest total execution time. The selected sub-network distribution is then launched on the compute nodes of the hardware system.
[0051] The request to the scheduling node can also include several target performance parameters such as execution time and memory consumption, where different performance parameters can be assigned different weights or priorities. Therefore, the scheduling node can consider several different performance metrics when selecting the sub-network distribution. Depending on the selected target performance parameters, the scheduling node can also select different computing nodes. The first computing node 1 may have a lower processing speed and a larger memory size, while the second computing node 2 may have a higher processing speed and a smaller memory size, thus influencing the selection of the sub-network distribution by the scheduling node.
[0052] Figure 4 An example subnetwork distribution 400 is schematically illustrated, comprising three subnetworks 402a to 402c, wherein each subnetwork 402a to 402c subsequently includes multiple operations to be performed by the respective subnetwork 402a to 402c. The subnetwork distribution 400 includes: an input portion 404 in which the required input for inference is provided to the deep neural network; an output portion 406 in which the inference generated by the deep neural network is provided; and a hidden portion 408 including hidden layers in which operations not directly visible or usable by the end user are performed.
[0053] The following text describes four example scenarios 1 through 4, in which different parameters affect the division of a deep neural network into subnetwork distributions.
[0054] Scene 1
[0055] Figure 5A hardware system 500 comprising two computing nodes 502a to 502b communicating with each other via 5G is schematically illustrated. Complete inference has latency as a target performance parameter. A subnetwork distribution can be generated in the anticipated scenario, where the first node 502a is responsible for taking actions based on the output from the inference, but the first node 502a alone cannot meet the latency requirements. Therefore, the deep neural network is divided into two subnetworks, where the output of the first subnetwork is transmitted from the first node 502a to the second node 502b and used as the input to the second subnetwork. Based on the latency imposed by transmitting data from the first node 502a to the second node 502b, it can be determined that inference does not meet the latency requirements without modification of the initial subnetwork distribution. The latency requirements can be met, for example, by modifying the architecture of the two subnetworks through pruning, taking into account information about the hardware located at the nodes and the chosen runtime environment. The inference latency can be further reduced by modifying the first network to reduce the size of the data that must be transmitted between the two nodes 502a to 502b. When dividing a deep neural network into sub-network distributions, several sub-network distributions and optimizations can be considered for latency in order to find the best trade-off between performance and latency.
[0056] Scene 2
[0057] Figure 6 A hardware system 600 is schematically illustrated, comprising a central computing node 601 connected to multiple computing nodes 602a to 602e, wherein the computing nodes 602a to 602e are at least capable of serial communication. The illustrated hardware system 600 can be viewed as representing a hardware system comprising several memory-constrained computing nodes, wherein none of the nodes can independently host the entire deep neural network. The deep neural network is then partitioned into a distribution of subnetworks, wherein each subnetwork satisfies the memory constraints of each corresponding node 602a to 602e. Each subnetwork can be further optimized to address any additional latency constraints associated with the full inference process.
[0058] Scene 3
[0059] Figure 7 A hardware system 700 is schematically shown, comprising a central computing node 701 connected to multiple computing nodes 702a to 702e. Figure 7 This describes a hardware system where a complete inference of a deep neural network cannot be completed within a single available time slot, where computation nodes 702a to 702e represent time slots on a single processing device. The deep neural network is then partitioned into subnetworks, where the inference latency of each subnetwork does not exceed the corresponding available time slots 702a to 702e, which can be referred to as time slices. Each subnetwork can be further optimized to reduce the overall latency of the complete inference.
[0060] Although the present invention has been described with reference to specific exemplary embodiments, many different changes, modifications, etc., will become apparent to those skilled in the art. Furthermore, it should be noted that parts of the methods and systems may be omitted, interchanged, or arranged in various ways, and the methods and systems will still be able to perform the functions of the present invention.
[0061] Furthermore, based on a study of the drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude multiple. The fact that certain measures are enumerated in mutually different dependent claims does not indicate that combinations of these measures cannot be used advantageously.
Claims
1. A computer-implemented method for generating at least one sub-network distribution and determining a performance metric for a deep neural network for a hardware system (300) comprising multiple computing nodes (302a to 302h), said method comprising, in a processor device (200): Receive (100) deep neural network (202); Receive (102) the topology (204) and performance information (206) of the hardware system; Receive (104) runtime statistics (208) of the hardware system; Receive (106) at least one target performance parameter (210) for the execution of the deep neural network on the hardware system; For each target performance parameter, the deep neural network is divided (108) into at least one sub-network distribution (212a to 212c) based on the runtime statistics of the hardware system and the target performance parameter, wherein each sub-network in the sub-network distribution is adapted to execute on the computing nodes of the hardware system, and For each sub-network distribution, determine (110) a performance metric (214a to 214c) for at least one target performance parameter of the execution of the deep neural network on the hardware system.
2. The computer-implemented method according to claim 1, wherein, The performance information of the hardware system includes at least one of the following: memory size of the computing nodes, memory availability of the computing nodes, time availability of the computing nodes, computing resources of the computing nodes, latency of the computing nodes, and latency between computing nodes.
3. The method according to claim 1 or 2, wherein, The target performance parameter is at least one of latency, energy consumption, number of computations used for inference, memory requirements, throughput, and accuracy.
4. The method according to any one of the preceding claims, wherein, Dividing the deep neural network into multiple sub-networks includes: Receive a previous performance metric for the target performance parameter of the first sub-network distribution executed on the hardware system; and The deep neural network is divided into a second sub-network distribution based on runtime statistics describing the performance of the hardware system, thereby improving the previous performance metric.
5. The method according to any one of the preceding claims, wherein, The operation of dividing the deep neural network into at least one sub-network distribution includes: Access the lookup table that includes the runtime statistics; For the target performance parameter, determine whether changes in the subnetwork distribution improve the performance metric of the performance parameter.
6. The method according to any one of the preceding claims, wherein, Dividing the deep neural network includes adjusting the size of the subnetworks to not exceed the available storage space of the corresponding computing nodes.
7. The method according to any one of the preceding claims, wherein, Dividing the deep neural network includes adjusting the operation of the sub-network based on the runtime environment of the corresponding computing node.
8. The method according to any one of the preceding claims, wherein, Dividing the deep neural network into at least one subnetwork distribution includes pruning, performing neural architecture search (NAS), performing quantization of the deep neural network and / or quantization of the subnetwork.
9. The method according to any one of the preceding claims further comprises: The first processor device provides the plurality of sub-networks to the scheduling node; The target performance parameters are received by the scheduling node; The scheduling node compares the target performance parameters with the corresponding performance metrics of the target performance parameters distributed across the multiple sub-networks; as well as The scheduling node selects from the plurality of sub-network distributions the sub-network distribution that has the "best" performance metric for the received target performance parameters.
10. A computer program product comprising program code for performing the method according to any one of claims 1 to 9 when executed by a processing device.
11. A computer system including a processing circuitry (200), said processing circuitry (200) being configured to: Receive deep neural network (202); Receive the topology (204) and performance information (206) of the hardware system; Receive runtime statistics data (208) of the hardware system; Receive at least one target performance parameter (210) for the execution of the deep neural network on the hardware system; For each target performance parameter, the deep neural network is divided into at least one sub-network distribution based on the runtime statistics of the hardware system and the target performance parameter, wherein each sub-network in the sub-network distribution is adapted to execute on the computing nodes of the hardware system, and For each subnetwork distribution, a performance metric is determined for the execution of the deep neural network on the hardware system based on at least one target performance parameter.
12. The computer system according to claim 11, wherein, Dividing the deep neural network into multiple sub-networks includes: Receive a previous performance metric of the target performance parameter of the first sub-network distribution executed on the hardware system; and The deep neural network is divided into a second sub-network distribution based on runtime statistics describing the performance of the hardware system, thereby improving the previous performance metric.
13. The computer system according to claim 11 or 12, wherein, The operation of dividing the deep neural network into at least one sub-network distribution includes: Access the lookup table that includes the runtime statistics; For the target performance parameter, determine whether changes in the subnetwork distribution improve the performance metric of the performance parameter.
14. The computer system according to any one of claims 11 to 13, wherein, Dividing the deep neural network includes adjusting the size of the subnetworks to not exceed the available storage space of the corresponding computing nodes.
15. The computer system according to any one of claims 11 to 14, further comprising a scheduling device, the scheduling device being configured to: Receive the plurality of sub-networks from the processor device; Receive target performance parameters; Compare the target performance parameters with the corresponding performance metrics of the target performance parameters distributed across the multiple sub-networks; as well as Select the subnetwork distribution that has the best performance metric for the received target performance parameters from the plurality of subnetwork distributions.