Split network acceleration architecture

By splitting the large neural network into multiple AI inference accelerators and using the switch to directly route the intermediate inference request results, the problems of insufficient memory capacity and large host bandwidth consumption are solved, and efficient neural network inference is achieved.

CN113366501BActive Publication Date: 2025-06-10QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080011992.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-05
Filing Date
2020-02-06
Publication Date
2025-06-10
Estimated Expiration
2040-02-06

AI Technical Summary

Technical Problem

When the prior art accelerates the inference of large-scale neural networks, it faces the problems of insufficient memory capacity and large host bandwidth consumption, resulting in inefficiency.

Method used

Split the large neural network into multiple independent AI inference accelerators, each accelerator stores a part of the network, and directly routes the intermediate inference request results through the switch without involving the host processor.

Benefits of technology

Reduce dependence on host memory bandwidth and CPU cycles, improve the efficiency and speed of inference calculations, and avoid memory bottlenecks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113366501B_ABST
    Figure CN113366501B_ABST
Patent Text Reader

Abstract

Describes a method for accelerating machine learning on a computing device. The method includes: hosting a neural network in a first inference accelerator and a second inference accelerator. The neural network is split between the first inference accelerator and the second inference accelerator. The method further includes: directly routing intermediate inference request results between the first inference accelerator and the second inference accelerator. The method further includes: generating a final inference request result from these intermediate inference request results.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Patent Application No. 16 / 783,047, entitled "SPLIT NETWORK ACCELERATING ARCHITECTURE", filed on February 5, 2020, which in turn claims the benefit of U.S. Provisional Patent Application No. 62 / 802,150, entitled "SPLIT NETWORK ACCELERATING ARCHITECTURE", filed on February 6, 2019. The disclosures of these applications are hereby incorporated by reference in their entireties. Technical Field

[0003] Certain aspects of the present invention generally relate to artificial neural networks, and more particularly to hardware accelerators for splitting artificial neural networks. Background Art

[0004] An artificial neural network, which may include a group of interconnected artificial neurons, can be a computing device or can represent a method to be executed by a computing device. The artificial neural network can have corresponding structures and / or functions in biological neural networks. However, artificial neural networks can provide useful computing techniques for certain applications where traditional computing techniques may be cumbersome, impractical, or incompetent. Since artificial neural networks can infer a function from observations, such networks can be useful in applications where the complexity of the task and / or data makes it cumbersome to design the function using conventional techniques.

[0005] In computing, hardware acceleration is the use of computer hardware to execute some functions more efficiently than is possible with software running on a more general - purpose central processing unit (CPU). The hardware that performs the acceleration can be referred to as a hardware accelerator. A hardware accelerator can improve the execution of a particular algorithm by allowing greater concurrency, having specific data paths for temporary variables in the algorithm, and potentially reducing the overhead of instruction control.

[0006] Overview

[0007] A method for accelerating machine learning on a computing device is described. The method includes: hosting a neural network in a first inference accelerator and a second inference accelerator. The neural network is split between the first inference accelerator and the second inference accelerator. The method further includes: directly routing intermediate inference request results between the first inference accelerator and the second inference accelerator. The method further includes: generating a final inference request result from these intermediate inference request results.

[0008] A system for accelerating machine learning is described. The system includes: a neural network, which is hosted in a first inference accelerator and a second inference accelerator. The neural network is split between the first inference accelerator and the second inference accelerator. The system further includes: a switch for directly routing intermediate inference request results between the first inference accelerator and the second inference accelerator. The system still further includes: a host device for receiving final inference request results generated from these intermediate inference request results.

[0009] A system for accelerating machine learning is described. The system includes: a neural network, which is hosted in a first inference accelerator and a second inference accelerator. The neural network is split between the first inference accelerator and the second inference accelerator. The system further includes: means for directly routing intermediate inference request results between the first inference accelerator and the second inference accelerator. The system still further includes: a host device for receiving final inference request results generated from these intermediate inference request results.

[0010] The features and technical advantages of the present disclosure have been outlined more broadly so that the following detailed description may be better understood. Additional features and advantages of the present disclosure will be described below. Those skilled in the art should appreciate that the present disclosure can be readily used as a basis for modifying or designing other structures for implementing the same purpose as the present disclosure. Those skilled in the art should also recognize that such equivalent constructs do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features that are considered to be characteristics of the present disclosure, both in its organization and method of operation, together with further objects and advantages, will be better understood when the following description is considered in conjunction with the accompanying drawings. However, it should be clearly understood that each drawing is provided for purposes of illustration and description only, and is not intended as a definition of the limitation of the present disclosure. Brief Description of the Drawings

[0012] The features, nature, and advantages of the present disclosure will become more apparent when the detailed description set forth below is understood in conjunction with the accompanying drawings, in which like reference numerals always denote corresponding parts.

[0013] Figure 1 An example implementation of designing a neural network using a system on a chip (SOC) including a general-purpose processor in accordance with certain aspects of the present disclosure is illustrated.

[0014] Figure 2A 、 2B and 2C are diagrams illustrating a neural network in accordance with aspects of the present disclosure.

[0015] Figure 2D is a diagram illustrating a neural network in accordance with aspects of the present disclosure.

[0016] Figure 3 is a block diagram illustrating a neural network acceleration architecture according to aspects of the present disclosure.

[0017] Figure 4 is a block diagram illustrating control flow and data flow in a neural network acceleration architecture according to aspects of the present disclosure.

[0018] Figure 5 is a block diagram illustrating control flow and data flow in a neural network acceleration architecture according to aspects of the present disclosure.

[0019] Figure 6 is a block diagram illustrating control flow and data flow in a neural network acceleration architecture according to further aspects of the present disclosure.

[0020] Figure 7 Illustrates a method for accelerating machine learning on a computing device according to aspects of the present disclosure.

[0021] Detailed description

[0022] The following detailed description, presented in conjunction with the accompanying drawings, is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.

[0023] Based on this teaching, those skilled in the art should appreciate that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether implemented independently of or in combination with any other aspect of the present disclosure. For example, any number of the aspects described may be used to implement an apparatus or practice a method. Additionally, the scope of the present disclosure is intended to cover such apparatus or methods practiced using other structures, functionality, or a combination of structures and functionality that supplement or are different from the various aspects of the present disclosure as described. It should be understood that any aspect of the present disclosure disclosed may be implemented by one or more elements of the claims.

[0024] Although specific aspects are described herein, numerous variations and permutations of these aspects fall within the scope of the present disclosure. Although some benefits and advantages of the preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses, or objectives. Instead, the aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the drawings and the following description of the preferred aspects. The detailed description and the drawings merely illustrate the present disclosure and do not limit the present disclosure, the scope of the present disclosure being defined by the appended claims and their equivalent technical solutions.

[0025] Examples of accelerated deep learning use a deep learning accelerator (e.g., an artificial intelligence inference accelerator) to train a neural network. Another example of accelerated deep learning uses a deep learning accelerator to operate a trained neural network to perform inference. Yet another example of accelerated deep learning is to use a deep learning accelerator to train a neural network and then use the trained neural network to perform inference. Information from the trained neural network and / or variants of the trained neural network are also used to perform inference.

[0026] As mentioned, an artificial intelligence accelerator can be used to train a neural network. Training of a neural network generally involves determining one or more weights associated with the neural network. For example, the weights associated with the neural network are determined through hardware acceleration using a deep learning accelerator. Once the weights associated with the neural network are determined, the trained neural network can be used to perform inference, where the trained neural network calculates a result (e.g., an activation) by processing input data based on the weights associated with the trained neural network.

[0027] However, in practice, a deep learning accelerator has a fixed amount of memory (e.g., 128 megabytes (MB) of static random access memory (SRAM)). As a result, the capacity of a deep learning accelerator is sometimes not large enough to accommodate and store a single network. For example, some networks have weights that are larger in size than the fixed amount of memory available from the deep learning accelerator. One solution to accommodate large networks is to split the weights into separate storage devices (e.g., dynamic random access memory (DRAM)). These weights are then read from the DRAM during each inference. However, this implementation uses more power and may result in a memory bottleneck.

[0028] Another solution to accommodate large networks is to split the network into multiple slices and pass intermediate results from one accelerator to another through a host. Unfortunately, passing intermediate inference request results through the host consumes host bandwidth. For example, using a host interface (e.g., a peripheral component interconnect express (PCIe) interface) to pass intermediate inference request results consumes host memory bandwidth. Additionally, passing intermediate inference request results through the host (e.g., a host processor) consumes the central processing unit cycles of the host processor and lengthens the overall inference computation latency.

[0029] One aspect of the present disclosure splits a large neural network across multiple separate artificial intelligence (AI) inference accelerators (AIIAs). Each of the separate AI inference accelerators may be implemented in a separate system-on-chip (SoC). For example, each AI inference accelerator is assigned and stores a portion of the weights or other parameters of the neural network. Intermediate inference request results are passed from one AI inference accelerator to another AI inference accelerator independently of the host processor. Thus, the host processor is not involved in the passing of the intermediate inference request results.

[0030] In some aspects, a designated AI inference accelerator among these AI inference accelerators communicates directly with the host (e.g., the host processor). For example, non-designated AI inference accelerators communicate with the host processor through the designated AI inference accelerator. The separate AI inference accelerators are also coupled to a peripheral interface or a server. The separate AI inference accelerators may be implemented in the same or different modules (or cards) or in the same or different packages.

[0031] In one aspect, a switch device (e.g., a Peripheral Component Interconnect Express (PCIe) switch or other similar peripheral interconnect switch) is coupled to the separate AI inference accelerators to pass intermediate inference request results between the separate AI inference accelerators without involving the host processor. Thus, each of the separate artificial intelligence inference accelerators together performs inference for the neural network split across the separate artificial intelligence inference accelerators.

[0032] Figure 1 An example implementation of a system-on-chip (SoC) 100 that may include a central processing unit (CPU) 102 or a multi-core CPU is illustrated, such as split network acceleration. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., a neural network with weights), latencies, frequency slot information, and task information may be stored in a memory block associated with the neural processing unit (NPU) 108, in a memory block associated with the CPU 102, in a memory block associated with the graphics processing unit (GPU) 104, in a memory block associated with the digital signal processor (DSP) 106, in memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from memory block 118.

[0033] SoC 100 may also include additional processing blocks customized for specific functions (such as connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, unlicensed Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.)) and a multimedia processor 112 that can, for example, detect and recognize gestures. In one implementation, the NPU is implemented in the CPU, DSP, and / or GPU. SoC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which may include a global positioning system).

[0034] Deep learning architectures can perform object recognition tasks by learning to represent the input at successively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses the main bottlenecks of conventional machine learning. Before the advent of deep learning, machine learning approaches for object recognition problems might have relied heavily on human-engineered features, perhaps combined with shallow classifiers. Shallow classifiers can be two-class linear classifiers, for example, where the weighted sum of the feature vector components is compared to a threshold to predict which class the input belongs to. Human-engineered features can be templates or kernels customized by engineers with domain expertise for a specific problem domain. In contrast, deep learning architectures can learn to represent features similar to those that human engineers might design, but it learns through training. Additionally, deep networks can learn to represent and recognize new types of features that humans might not have considered.

[0035] Deep learning architectures can learn a hierarchy of features. For example, if visual data is presented to the first layer, the first layer can learn to recognize relatively simple features in the input stream (such as edges). In another example, if auditory data is presented to the first layer, the first layer can learn to recognize spectral power in specific frequencies. The second layer, which takes the output of the first layer as input, can learn to recognize combinations of features, such as recognizing simple shapes for visual data or combinations of sounds for auditory data. For example, higher layers can learn to represent complex shapes in visual data or spoken phrases in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.

[0036] Deep learning architectures may perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motor vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher levels to recognize cars, trucks, and airplanes.

[0037] Neural networks can be designed with a wide variety of connectivity patterns. In a feedforward network, information is passed from lower layers to higher layers, where each neuron in a given layer communicates to neurons in a higher layer. As described above, hierarchical representations can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one chunk of input data presented sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be beneficial when the identification of high-level concepts can assist in discerning specific low-level features of the input.

[0038] The connections between the layers of a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, a neuron in the first layer can communicate its output to every neuron in the second layer, such that each neuron in the second layer receives input from every neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, a neuron in the first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that each neuron in a layer will have the same or similar connectivity pattern, although the connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern can result in spatially distinct receptive fields in higher layers, since neurons in higher layers in a given region can receive input that has been tuned through training to the properties of a restricted portion of the total input to the network.

[0039] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 can be configured such that the connection strengths associated with the input to each neuron in the second layer are shared (e.g., 208). Convolutional neural networks can be well-suited for problems where the spatial location of the input is meaningful.

[0040] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2DA detailed example of the DCN 200 designed to recognize visual features input from an image capture device 230 (such as an in-vehicle camera) is explained. The DCN 200 of the current example can be trained to identify traffic signs and the numbers provided on the traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic signals.

[0041] The DCN 200 can be trained using supervised learning. During training, an image (such as the image 226 of a speed limit sign) can be presented to the DCN 200, and then a "forward pass" can be computed to generate an output 222. The DCN 200 can include a feature extraction section and a classification section. Upon receiving the image 226, the convolutional layer 232 can apply a convolutional kernel (not shown) to the image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 can be a 5x5 kernel that generates 28x28 feature maps. In this example, since four different convolutional kernels are applied to the image 226 at the convolutional layer 232, four different feature maps are generated in the first set of feature maps 218. The convolutional kernel can also be referred to as a filter or a convolutional filter.

[0042] The first set of feature maps 218 can be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (such as 14x14) is smaller than the size of the first set of feature maps 218 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate subsequent sets of feature maps (not shown).

[0043] In Figure 2D the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. In addition, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 can include a number corresponding to a possible feature of the image 226 (such as "sign", "60", and "100"). A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.

[0044] In this example, the probabilities of "flag" and "60" in output 222 are higher than the probabilities of other outputs in output 222 (such as "30", "40", "50", "70", "80", "90", and "100"). Before training, the output 222 generated by DCN 200 is likely to be incorrect. Thus, the error between output 222 and the target output can be calculated. The target output is the ground truth of image 226 (e.g., "flag" and "60"). The weights of DCN 200 can then be adjusted to align the output 222 of DCN 200 more closely with the target output.

[0045] Figure 3 is a block diagram illustrating a neural network acceleration architecture 300 according to aspects of the present disclosure. The neural network acceleration architecture 300 includes a first AI inference accelerator (e.g., first AIIA 330), a second AI inference accelerator (e.g., second AIIA 340), and a host processor 310. In this configuration, the host processor 310 includes a host runtime block 312 for executing a host application 320 to operate a neural network. In this example, the neural network of the host application 320 exceeds the fixed amount of memory provided by a single AI inference accelerator (AIIA).

[0046] According to this aspect of the present disclosure, the neural network of the host application 320 is split across the first AIIA 330 and the second AIIA 340. That is, the neural network of the host application 320 is hosted in the first AIIA 330 and the second AIIA 340. However, splitting the neural network of the host application 320 between the first AIIA 330 and the second AIIA 340 involves routing intermediate inference request results 308 generated by the first AIIA 330 and / or the second AIIA 340. In this example, the intermediate inference request results 308 are routed between the first AIIA 330 and the second AIIA 340 independently of the host processor 310. It should be appreciated that the intermediate inference request results 308 are intermediate results for generating the inference request results 306.

[0047] The inference request results 306 are provided to the host processor 310 through the inference accelerators specified in the first AIIA 330 and the second AIIA 340. In this configuration, the non-specified inference accelerator (e.g., second AIIA 340) communicates with the host processor 310 through the specified inference accelerator (e.g., first AIIA 330). For example, the inference request results 306 can be generated by the first AIIA 330, where the inference request results 306 are based on the intermediate inference request results 308 received from the second AIIA 340.

[0048] In one aspect of the present disclosure, routing of the intermediate inference request result 308 is implemented based on addresses imposed by the switch device 302, including the peripheral interface 304 (e.g., x8 peripheral interface). In this configuration, the switch device 302 (e.g., a Peripheral Component Interconnect Express (PCIe) switch) is coupled to the first AIIA 330 and the second AIIA 340 via a peripheral interconnect 309. The switch device 302 is coupled to the host processor 310 via the peripheral interface 304 (e.g., a PCIe x8 interface).

[0049] According to this aspect of the present disclosure, the switch device 302 supports peer-to-peer communication between the first AIIA 330 and the second AIIA 340 to enable direct memory access (DMA) transfers without intervention from the host processor 310. In one configuration, the switch device 302 (e.g., a PCIe switch) is configured to route based on an address range. For example, a peer address range is set for which data is transferred peer-to-peer by the switch device 302. In this configuration, the incoming address translation units (iATUs) in the first AIIA 330 and the second AIIA 340 are set to map local addresses to the peer address space (e.g., the PCIe address space of the switch device 302). For example, this switching is transparent to the endpoints in the first AIIA 330 and / or the second AIIA 340. The DMA transfer is set using appropriate addresses targeted at locations in the peer devices. Additionally, the base address registers (BARs) in the first AIIA 330 and / or the second AIIA 340 may provide access windows to the peer AIIAs (such as the first AIIA 330 and / or the second AIIA 340).

[0050] For example, the switch device 302 supports mapping of the local addresses of the first AIIA 330 and the second AIIA 340 to exchange the intermediate inference request result 308 without intervention from the host processor 310. In some aspects, data from the host processor 310 may be accessed by the first AIIA 330 based on the DMA implementation mentioned. Similarly, data from the first AIIA 330 may be accessed by the second AIIA 340 based on the DMA implementation mentioned, as further illustrated in Figure 4 which is further described in

[0051] Figure 4 is a block diagram illustrating control flow and data flow in a neural network acceleration architecture according to aspects of the present disclosure. In this configuration, the neural network acceleration architecture 400 may be associated with Figure 3is similar to the neural network acceleration architecture 300 shown. In this example, data accessed by the first AIIA 430 can be through pointers in a request queue (RQ) 416 stored in the memory 414 of the host 410 (e.g., host processor). Inference inputs (e.g., pointers) are queued in the RQ 416. When it is determined that a pointer to data exists in the RQ 416, the data is passed and received at a virtual channel (e.g., VC 432) of the first AIIA 430, as shown by the first control flow 460. Additionally, inference request results are provided from the first AIIA 430 to the completion queue (QC) 418 of the host 410, as shown by the second control flow 462.

[0052] As further illustrated in Figure 4 the second AIIA 440 also includes a virtual channel (e.g., VC 442) through which the second AIIA 440 receives data (as shown by the third control flow 464) and transmits data (as shown by the fourth control flow 466). One way to communicate inference request results from the first AIIA 430 to the host 410 is through the global synchronization manager (GSM) 450 of the first AIIA 430. The first AIIA 430 may also include a request queue (RQ) 436 and a completion queue (CQ) 438 in the memory 434 of the first AIIA 430.

[0053] According to this configuration, the host 410 is the driver of the control flow and data flow for the first AIIA 430 using the RQ 416 and CQ 418 in the memory 414 of the host 410. Similarly, the first AIIA 430 is the driver of the control flow and data flow for the second AIIA 440 using the RQ 436 and CQ 438 in the memory 434 of the first AIIA 430. For example, the network signal processor (NSP) (not shown) of the first AIIA 430 generates RQ elements for the DMA of inputs going to the second AIIA 440 in the RQ 436 within the memory 434 of the first AIIA 430.

[0054] In this example, the NSP of the first AIIA 430 generates additional RQ elements for the DMA of results from the second AIIA 440 (see data flow 470) in the RQ 436. This additional RQ element may have a final DMA (e.g., doorbell) which can be written to the synchronization manager (e.g., GSM 450) in the first AIIA 430. Writing to the GSM 450 can be performed by the second AIIA 440 using an interconnect protocol (e.g., Peripheral Component Interconnect (PCI) protocol), as shown by the fourth control flow 466. The RQ tail (T) pointer is incremented in the VC 442 of the second AIIA 440 using, for example, the third control flow 464.

[0055] In this configuration, the operation of the second AIIA 440 assumes communication with the host 410; however, the host address for this communication is mapped to the first AIIA 430. Once the intermediate result is completed by the second AIIA 440, a doorbell to the GSM 450 in the first AIIA 430 is written across the switch device 302, as Figure 3 shown. The DMA channel established in the first AIIA 430 can be assigned a pre-synchronization condition to DMA the inference request result back to the host 410, as shown by the data stream 472.

[0056] Figure 5 is a block diagram illustrating the control flow and data flow in a neural network acceleration architecture according to aspects of the present disclosure. In this configuration, the network acceleration architecture 500 can be similar to the neural network acceleration architecture 400 shown in Figure 4 and includes a host 410 and a first AIIA 430. In this configuration, the second AIIA 540 includes virtual channels (e.g., VC 542) and a memory 544 (including a request queue (e.g., RQ 546) and a command queue (CQ) 548). The network acceleration architecture 500 further includes a third AIIA 580 and a fourth AIIA 590.

[0057] In operation, the data accessed by the second AIIA 540 can come from the RQ436 in the memory 434 of the first AIIA 430. The inference inputs are queued in the RQ 436. When it is determined that data exists in the RQ 436, the data is passed and received at the VC 542 of the second AIIA 540, as shown by the control flow 560. Additionally, the data accessed by the third AIIA 580 can come from the RQ 546 in the memory 544 of the second AIIA 540. The inference inputs are queued in the RQ 546. When it is determined that data exists in the RQ546, the data is passed and received at the third AIIA580, as shown by the control flow 562. When the intermediate inference request result is stored in the completion queue CQ 548 of the second AIIA 540, the intermediate inference request result is passed from the second AIIA 540 to the third AIIA 580, as shown by the data stream 570.

[0058] Once the intermediate result is completed by the third AIIA 580, these intermediate inference request results are written to the fourth AIIA590. Once the intermediate result is completed by the fourth AIIA 590, the fourth AIIA 590 can use the doorbell mechanism for the GSM 450 in the first AIIA 430 so that these intermediate results cross the switch device 302 ( Figure 3) are written, as shown by data stream 574. The DMA channels established in the first AIIA 430 can be assigned pre-synchronization conditions to DMA the inference request results back to the host 410, as shown by data stream 472. The configuration of the network acceleration architecture 500 includes additional computing capabilities provided by adding a third AIIA 580 and a fourth AIIA 590. Unfortunately, the circular data movement that provides the inference request results to the first AIIA 430, shown by data streams 570, 572, and 574, involves additional data transfers, as shown by data stream 574.

[0059] Figure 6 is a block diagram illustrating control flow and data flow in a neural network acceleration architecture according to aspects of the present disclosure. In one configuration, the network acceleration architecture 600 can be similar to the Figure 4 neural network acceleration architecture 400 shown therein. In one configuration, the network acceleration architecture 600 includes a host 610, a first AIIA 630, and a second AIIA 640. In this configuration, the host 610 controls both the first AIIA 630 and the second AIIA 640, relative to the Figure 4 neural network acceleration architecture 400 shown therein, which eliminates additional data transfers. However, the network acceleration architecture 600 involves additional costs due to having a host interface with the first AIIA 630 and a host interface with the second AIIA 640.

[0060] In this aspect of the present disclosure, the first AIIA 630 includes a first request queue (RQ1) 616 and a first completion queue (CQ1) 618 in the first host memory 614. In this configuration, the host 610 is the driver for the input data transfer 670 (In) from the host 610 to the first AIIA 630 (see control flow 660). Additionally, the first AIIA 630 includes a virtual channel (VC) 632 for the input data transfer 670, which includes a head (H) pointer and a tail (T) pointer to organize the input data transfer 670 to generate intermediate inference request results for the first part of the neural network.

[0061] In this configuration, the second AIIA 640 includes a second request queue (RQ2) 622 and a second completion queue (CQ2) 624 in the second host memory 620 for incoming data transfer 672 from the first AIIA 630 and outgoing data transfer 674 to the host 610. Additionally, the second AIIA 640 includes virtual channels (VCs) 642 for incoming data transfer 672 and outgoing data transfer 674, which include head (H) and tail (T) pointers to organize the incoming data transfer 672 and outgoing data transfer 674 to generate intermediate inference request results for a second part of the neural network. In this configuration, the host 610 is the driver for the incoming data transfer 672, where the data movement is from the first AIIA 630 to the second AIIA 640. Additionally, the host 610 can be configured to complete the final inference request result based on the intermediate inference request results calculated by the first AIIA 630 for this part of the neural network and the intermediate inference request results calculated by the second AIIA 640 for the second part of the neural network.

[0062] In operation, for a single inference request, the host 610 is designated to place incoming requests for these inputs in the first request queue RQl 616 to compute the single inference request. Additionally, the host 610 is further configured to place incoming requests in the second request queue RQ2 622 for the transfer of intermediate inference request results from the first AIIA 630 to the second AIIA 640. The host 610 is also configured to place outgoing requests for the inference request results in the second request RQ2 622. After the first AIIA 630 completes the inference for the first half of the neural network, the transfer of the intermediate inference request results is performed.

[0063] In this aspect of the disclosure, once the inference of the first half of the neural network is complete, the first AIIA 630 can write a value to a signal lamp in the second AIIA 640 (e.g., one (l) greater than the stored value); however, the first AIIA 630 may not know the value of the signal lamp in the second AIIA 640. The first AIIA 630 can increment the signal lamp in the second AIIA 640 by writing to the GSM 650 in the second AIIA 640 through a cross-switch device (not shown) (as shown by the control flow 668). In this configuration, when the signal lamp reaches a predetermined value, the second AIIA 640 performs a pre-synchronization condition. For example, the second AIIA 640 initiates a DMA transfer to move data according to the request queue element (RQE) in the second request RQ2 622, as shown by the outgoing data transfer 674. In this configuration, the second AIIA 640 writes to the doorbell register of the network signal processor (NSP) memory (e.g., according to the RQE) to notify the NSP that it can start processing the intermediate results in the second half of the network. Notifying the host 610 that the final inference request result is complete is performed by placing an entry in the second command queue (CQ2) 624 and incrementing the CQ2 tail pointer, as shown by the control flows 662 and 664.

[0064] Figure 7 A method for accelerating machine learning on a computing device is illustrated. Method 700 begins at block 702, where a neural network is hosted in a first inference accelerator and a second inference accelerator. For example, as Figure 4 shown, the neural network is split between the first AIIA 430 and the second AIIA 440. The method continues at block 704,

[0065] where intermediate inference request results are routed between the first inference accelerator and the second inference accelerator. For example, as Figure 4 shown, the data stream between the first AIIA 430 and the second AIIA 440 is routed independently of the host 410 (e.g., the host processor of the computing device). At block 706, a final inference request result is generated from these intermediate inference request results. For example, in Figure 4 shown, the final inference request result is generated by the second AIIA 440 based on the intermediate inference request results from the first AIIA 430. In some aspects, method 700 can be performed by a system-on-chip (SoC) that includes two separate inference accelerators.

[0066] In some aspects, method 700 can be performed by the SOC 100 ( Figure 1) is performed. That is, by way of example and not limitation, each element of method 700 can be performed by SOC 100 or one or more processors (e.g., CPU 102 and / or NPU 108) and / or other components included therein.

[0067] A system for accelerating machine learning includes means for directly routing intermediate inference request results between a first inference accelerator and a second inference accelerator. In one aspect, the routing means can be a switch device 302 configured to perform the recited functions. In another configuration, the foregoing means can be any module or any device configured to perform the functions described by the foregoing means.

[0068] The various operations of the methods described above can be performed by any suitable means capable of performing the corresponding functions. These means can include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the figures, those operations can have corresponding paired means plus function components with similar numbers.

[0069] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" can include computing, calculating, processing, deriving, researching, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, and the like. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and similar actions. Further, "determine" can include parsing, selecting, choosing, establishing, and similar actions.

[0070] As used herein, the phrase "at least one of" recited in a list of items refers to any combination of those items, including a single member. By way of example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c.

[0071] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure can be implemented or performed with a general purpose processor, digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the processor can be any commercially available processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0072] The steps of the methods or algorithms described in connection with the present disclosure may be implemented directly in hardware, in software modules executed by a processor, or in a combination of the two. The software modules may reside in any form of storage medium known in the art. Some examples of storage media that may be used include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software modules may include a single instruction or many instructions and may be distributed over several different code segments, distributed among different programs, and across multiple storage media. The storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In an alternative, the storage medium may be integrated into the processor.

[0073] The methods disclosed herein include one or more steps or acts for implementing the described methods. These method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of the steps or acts is specified, the order and / or use of the specific steps and / or acts may be altered without departing from the scope of the claims.

[0074] The described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented with a bus architecture. Depending on the particular application and overall design constraints of the processing system, the bus may include any number of interconnecting buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect, among other things, a network adapter to the processing system via the bus. The network adapter may be used to implement signal processing functions. For certain aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link together various other circuits such as a timing source, peripherals, voltage regulators, power management circuits, and similar circuits, which are well known in the art and will not be described further.

[0075] The processor may be responsible for managing the bus and general processing, including the execution of software stored on a machine-readable medium. The processor may be implemented with one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry capable of executing software. Software should be construed broadly to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. By way of example, the machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard disk drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging material.

[0076] In a hardware implementation, the machine-readable medium may be a part of the processing system separate from the processor. However, as will be readily appreciated by those skilled in the art, the machine-readable medium or any part thereof may be external to the processing system. By way of example, the machine-readable medium may include transmission lines, carrier waves modulated with data, and / or computer products separate from the device, all of which may be accessed by the processor via a bus interface. Alternatively or additionally, the machine-readable medium or any part thereof may be integrated into the processor, such as may be the case with a cache and / or a general register file. Although the various components discussed may be described as having a specific location, such as local components, they may also be configured in various ways, such as some components being configured as part of a distributed computing system.

[0077] The processing system can be configured as a general-purpose processing system that has one or more microprocessors providing processor functionality, and an external memory providing at least a portion of the machine-readable medium, all linked together by an external bus architecture to other support circuitry. Alternatively, the processing system can include one or more neuromorphic processors for implementing the neuron models and nervous system models described herein. As another alternative, the processing system can be implemented with an application specific integrated circuit (ASIC) with a processor, bus interface, user interface, support circuitry, and at least a portion of the machine-readable medium integrated on a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuits capable of performing the various functions described throughout this disclosure. Depending on the particular application and the overall design constraints imposed on the system, those of ordinary skill in the art will recognize how best to implement the functionality described with respect to the processing system.

[0078] The machine-readable medium can include several software modules. These software modules include instructions that, when executed by the processor, cause the processing system to perform various functions. These software modules can include a transmission module and a reception module. Each software module can reside in a single storage device or be distributed across multiple storage devices. As an example, when a triggering event occurs, the software module can be loaded from a hard drive into RAM. During the execution of the software module, the processor can load some of the instructions into a cache to improve access speed. One or more cache lines can then be loaded into the general register file for execution by the processor. When referring to the functionality of the software modules below, it will be understood that such functionality is implemented by the processor when the processor executes instructions from the software module. Further, it should be appreciated that aspects of the present disclosure result in improvements to the capabilities of a processor, computer, machine, or other system implementing such aspects.

[0079] If implemented in software, each function can be stored on or transmitted via a computer-readable medium as one or more instructions or code. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium accessible by a computer. By way of example and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL), or wireless technology such as infrared (IR), radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, radio, and microwave is included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and disc, where disk often magnetically reproduces data, while disc optically reproduces data with a laser. Thus, in some aspects, computer-readable media can include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media can include transitory computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.

[0080] Accordingly, some aspects can include a computer program product for performing the operations given herein. For example, such a computer program product can include a computer-readable medium having instructions stored (and / or encoded) thereon that can be executed by one or more processors to perform the operations described herein. For some aspects, the computer program product can include packaging materials.

[0081] Furthermore, it should be appreciated that modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by a user terminal and / or base station where applicable. For example, such devices can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage device (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or floppy disk, etc.) such that once the storage device is coupled to or provided to the user terminal and / or base station, the device can obtain the various methods. Additionally, any other suitable technology can be utilized that is adapted to provide the methods and techniques described herein to a device.

[0082] It will be understood that the claims are not limited to the exact configurations and components illustrated above. Various modifications, substitutions, and variations can be made in the layout, operation, and details of the methods and apparatuses described above without departing from the scope of the claims.

Claims

1. A method for accelerating machine learning on a computing device, comprising: hosting a neural network in a first inference accelerator and a second inference accelerator, the neural network being split between the first inference accelerator and the second inference accelerator; implementing a request queue for the second inference accelerator in a memory in either the first inference accelerator or a host processor; routing intermediate inference request results directly between the first inference accelerator and the second inference accelerator; and generating a final inference request result from the intermediate inference request results.

2. The method according to claim 1, wherein generating the final inference request result comprises: generating the final inference request result by the second inference accelerator in response to the intermediate inference request result from the first inference accelerator.

3. The method according to claim 2, further comprising: transmitting the final inference request result directly from the second inference accelerator to a host processor of the computing device.

4. The method according to claim 2, further comprising: transmitting the final inference request result from the second inference accelerator to a host processor of the computing device via the first inference accelerator.

5. The method according to claim 1, wherein routing the intermediate inference request results directly between the first inference accelerator and the second inference accelerator is performed independently of a host processor of the computing device.

6. The method according to claim 1, wherein generating the final inference request result further comprises: after the final inference request result from the second inference accelerator is transferred by direct memory access (DMA) to the first inference accelerator, writing the final inference request result from the second inference accelerator to a global synchronization manager (GSM) of the first inference accelerator across a switch device.

7. The method according to claim 6, wherein the DMA transfer of the final inference request result is based on a request queue element in a request queue of the second inference accelerator stored in a memory of the first inference accelerator.

8. The method according to claim 1, wherein routing the intermediate inference request results comprises: after the intermediate inference request result from the first inference accelerator is transferred by direct memory access (DMA) to the second inference accelerator, writing the intermediate inference request result from the first inference accelerator to a global synchronization manager (GSM) of the second inference accelerator across a switch device.

9. A system for accelerating machine learning, comprising: a neural network, the neural network being hosted in a first inference accelerator and a second inference accelerator, the neural network being split between the first inference accelerator and the second inference accelerator, wherein a request queue for the second inference accelerator is implemented in a memory in either the first inference accelerator or a host processor; a switch for directly routing intermediate inference request results between the first inference accelerator and the second inference accelerator; and A host device, the host device being configured to receive a final inference request result generated from the intermediate inference request result.

10. The system of claim 9, wherein the final inference request result is generated by the second inference accelerator in response to the intermediate inference request result from the first inference accelerator.

11. The system of claim 10, wherein the final inference request result is directly transmitted from the second inference accelerator to the host device.

12. The system of claim 10, wherein the final inference request result is transmitted from the second inference accelerator to the host device via the first inference accelerator.

13. The system of claim 9, wherein the first inference accelerator includes a global synchronization manager (GSM) configured to notify the first inference accelerator that the final inference request result from the second inference accelerator is transferred to the first inference accelerator by direct memory access (DMA).

14. The system of claim 9, wherein the second inference accelerator includes a global synchronization manager (GSM) configured to notify the second inference accelerator after the intermediate inference request result from the first inference accelerator is transferred to the second inference accelerator by direct memory access (DMA).

15. A system for accelerating machine learning, comprising: a neural network, the neural network being hosted in a first inference accelerator and a second inference accelerator, the neural network being split between the first inference accelerator and the second inference accelerator, wherein a request queue for the second inference accelerator is implemented in a memory in one of the first inference accelerator and a host processor; means for directly routing an intermediate inference request result between the first inference accelerator and the second inference accelerator; and a host device, the host device being configured to receive a final inference request result generated from the intermediate inference request result.

16. The system of claim 15, wherein the final inference request result is generated by the second inference accelerator in response to the intermediate inference request result from the first inference accelerator.

17. The system of claim 16, wherein the final inference request result is directly transmitted from the second inference accelerator to the host device.

18. The system of claim 16, wherein the final inference request result is transmitted from the second inference accelerator to the host device via the first inference accelerator.

Citation Information

Patent Citations

  • Data processing device and server

    CN106776461A

  • Processing circuit and neural network operating method thereof

    CN108470009A

  • Field-Programmable Gate Array Based Accelerator System

    US20100076915A1