A method for deploying a deep learning model to an acceleration unit

By dividing the deep learning model into engine computing segments that adapt to the memory size of the acceleration unit core, and deploying it to the acceleration unit core based on storage space and dependencies, the frequent data exchange caused by insufficient on-chip storage is solved, and execution performance and efficiency are improved.

CN113743567BActive Publication Date: 2025-07-08PINGTOU GE (HANGZHOU) SEMICON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010467422.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-28
Publication Date
2025-07-08
Estimated Expiration
2040-05-28

AI Technical Summary

Technical Problem

Due to the limited on-chip storage resources of the acceleration unit, on-chip and off-chip data exchange is frequently performed when processing large-scale deep learning networks, affecting execution performance.

Method used

The deep learning model is divided into multiple engine computing fragments. The storage requirements of each fragment do not exceed the on-chip memory size of the acceleration unit core, and are deployed to the corresponding acceleration unit core for execution according to the storage space and dependencies.

Benefits of technology

It reduces data exchange on-chip and off-chip, improves the execution performance and utilization efficiency of acceleration units, and improves the inference speed of deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113743567B_ABST
    Figure CN113743567B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for deploying a deep learning model to an acceleration unit. In this method, according to the on-chip memory size of the acceleration unit cores in the acceleration unit, the deep learning model is divided into multiple engine operation segments. Each engine operation segment includes a part of multiple processing nodes in the deep learning model, and the storage space used when executed in the acceleration unit cores does not exceed the on-chip memory size. Then, in this method, according to the storage space size and / or dependency relationship to be used by each engine operation segment, each engine operation segment is deployed to the corresponding acceleration unit core for execution. The present invention also discloses an artificial intelligence device, a data center, and a computing device that adopt this method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and more particularly to the field of deep learning and the execution of deep learning models in an acceleration unit. Background Art

[0002] With the rapid development of artificial intelligence in recent years, artificial intelligence, especially deep learning algorithms, has been widely applied in fields such as visual images, speech recognition, and natural language processing.

[0003] Artificial intelligence requires more and more computing power, resulting in large-scale deep learning model computing scenarios. These computing scenarios require a large amount of computing resources. For this reason, acceleration units such as neural network processors (NPUs) have emerged to accelerate the computing of artificial intelligence, especially deep learning models.

[0004] When the processor is a CPU (central processing unit) or a GPU (graphics processing unit), they generally have sufficient memory resources to store the parameters required for the entire deep learning model network to run. Therefore, existing deep learning frameworks will directly hand over the entire deep learning model graph or sub-graph to the processor for processing without considering the storage space problem. However, due to design and cost reasons, acceleration units such as NPUs generally have much less on-chip storage resources than CPUs / GPUs. When using an acceleration unit to process a large-scale deep learning network, there will be a situation where the on-chip storage resources are not enough to load the entire network model, resulting in frequent data exchange between on-chip and off-chip, which affects the execution performance of the acceleration unit.

[0005] Therefore, a new deep learning model deployment scheme is needed, which can divide the deep learning model to adapt to the size of the on-chip memory of the acceleration unit, reduce the data exchange between on-chip and off-chip memories, and thus improve the execution performance of the acceleration unit. Summary of the Invention

[0006] For this purpose, the present invention provides a method and a computing system for deploying a deep learning model to an acceleration unit, in an attempt to solve or at least alleviate at least one of the above problems.

[0007] According to one aspect of the present invention, there is provided a method for deploying a deep learning model to an acceleration unit, including the steps of: dividing the deep learning model into multiple engine operation segments according to the size of the on-chip memory of the acceleration unit cores in the acceleration unit, where the storage space used by each engine operation segment when executed in the acceleration unit core does not exceed the size of the on-chip memory and includes a part of multiple processing nodes in the deep learning model; and deploying each engine operation segment to the corresponding acceleration unit core for execution according to the size of the storage space to be used by each engine operation segment and / or the dependency relationship.

[0008] Optionally, in the method according to the present invention, the step of dividing the deep learning model into multiple engine operation segments includes: constructing the engine operation segments by using multiple processing nodes that are sequentially executed in order, so that the storage space required at each running moment when the engine operation segments are executed in the acceleration unit cores does not exceed the size of the on-chip memory.

[0009] Optionally, in the method according to the present invention, the storage space required at the running moment includes the sum of the storage space required by the currently executed processing node and the storage space required by the dependent nodes. The dependent nodes include the processing nodes that have been executed and whose processing results are required by the currently executed processing node.

[0010] Optionally, in the method according to the present invention, the step of dividing the deep learning model into multiple engine operation segments includes: constructing sequentially executed engine operation segments according to the execution order of the processing nodes.

[0011] Optionally, in the method according to the present invention, constructing sequentially executed engine operation segments according to the execution order of the processing nodes includes: at the first running moment, determining the currently executed processing node at this running moment and the storage space required at this running moment; at a plurality of consecutive running moments after the first running moment, determining the currently executed processing node and the storage space required at each running moment until the second running moment, where the difference between the storage space required at the second running moment and the size of the on-chip memory is within a predetermined threshold; and dividing the processing nodes executed between the first running moment and the second running moment into engine operation segments.

[0012] Optionally, in the method according to the present invention, constructing sequentially executed engine operation segments according to the execution order of the processing nodes further includes: after determining the engine operation segments, determining the subsequent running moment as the new first running moment to determine the new second running moment, and dividing the processing nodes executed between the new first running moment and the second running moment into new engine operation segments.

[0013] Optionally, in the method according to the present invention, the step of dividing the deep learning model into multiple engine operation segments includes: constructing parallelly executed engine operation segments according to the storage space required at each running moment.

[0014] Optionally, in the method according to the present invention, the storage space required by the processing node includes the storage space occupied by one or more of the following items: input variables, weight values, activation parameters, and output variables.

[0015] Optionally, in the method according to the present invention, the storage space of the engine operation fragment is set to the maximum storage space required by the multiple runtimes included in the engine operation fragment; and the step of deploying each engine operation fragment to the corresponding acceleration unit core for execution includes: deploying the engine operation fragment with the largest storage space and the engine operation fragment with the smallest storage space to the same acceleration unit core for execution.

[0016] Optionally, in the method according to the present invention, dividing the deep learning model into multiple engine operation fragments further includes the steps of: determining the running time of each engine operation fragment according to the multiple runtimes included in each engine operation fragment; and determining the dependency relationship between each engine operation fragment according to the processing nodes included in each engine operation fragment.

[0017] Optionally, in the method according to the present invention, the step of deploying each engine operation fragment to the corresponding acceleration unit core for execution includes: deploying the engine operation fragments without dependency relationships to different acceleration unit cores for execution.

[0018] Optionally, in the method according to the present invention, the step of deploying each engine operation fragment to the corresponding acceleration unit core for execution includes: deploying the engine operation fragments with similar running times to different acceleration unit cores for execution.

[0019] Optionally, in the method according to the present invention, the deep learning model includes multiple model layers, and each processing node corresponds to a model layer or a part of a model layer.

[0020] According to another aspect of the present invention, there is provided an artificial intelligence device, including an acceleration unit and a processor. The acceleration unit includes one or more acceleration unit cores, and each acceleration unit core has a corresponding on-chip memory. The processor is adapted to execute the method according to the present invention so as to deploy the deep learning model to the acceleration unit for execution.

[0021] According to still another aspect of the present invention, there is provided a data center, which includes the artificial intelligence device according to the present invention.

[0022] According to still another aspect of the present invention, there is provided a data center, which includes multiple acceleration units for deploying the deep learning model according to the method of the present invention.

[0023] According to yet another aspect of the present invention, there is provided a computing device, including: at least one processor; and a memory storing program instructions, wherein the program instructions are configured to be adapted to be executed by at least one processor, and the program instructions include instructions for executing any of the above methods.

[0024] According to the solution of the present invention, considering the size of the on-chip memory of the acceleration unit core, the deep learning model is divided into multiple engine operation segments. Each engine operation segment includes a part of the deep learning model that is continuously executed, and the storage space required for each engine operation segment to execute does not exceed the size of the on-chip memory of the acceleration unit core. In this way, each engine operation segment can be executed in a single acceleration unit core without data exchange inside and outside the unit core, thereby significantly improving the inference speed of the deep learning model in the acceleration unit.

[0025] In addition, according to the solution of the present invention, after the deep learning model is divided into each engine operation segment, the engine operation segments can be deployed to the corresponding acceleration unit cores in an appropriate manner according to the storage space required by each engine operation segment, the mutual relationship (whether there is data dependence), and the execution timing, etc., thereby improving the utilization efficiency of the acceleration unit core and further improving the inference speed of the deep learning model in the acceleration unit.

[0026] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are hereinafter specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To achieve the above and related purposes, certain illustrative aspects are described herein in conjunction with the following description and the accompanying drawings, which indicate various ways in which the principles disclosed herein can be practiced, and all aspects and their equivalent aspects are intended to fall within the scope of the claimed subject matter. By reading the following detailed description in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present disclosure will become more apparent. Throughout the present disclosure, like reference numerals generally refer to like components or elements.

[0028] Figure 1A FIG. shows a schematic diagram of an exemplary computer system according to an embodiment of the present invention;

[0029] Figure 1B FIG. shows a schematic diagram of a deep neural network as a machine learning model according to an embodiment of the present invention;

[0030] Figure 2A FIG. shows an internal structure diagram of a computing device according to an embodiment of the present invention;

[0031] Figure 2B FIG. shows a connection relationship diagram of a scheduling unit and an acceleration unit inside a computing device according to an embodiment of the present invention;

[0032] Figure 3Shows the internal structure diagram of an acceleration unit core according to an embodiment of the present invention;

[0033] Figure 4 Shows a schematic diagram of a method for deploying a deep learning model into an acceleration unit according to an embodiment of the present invention;

[0034] Figure 5 Shows a schematic diagram of a method for determining an engine operation segment according to an embodiment of the present invention;

[0035] Figure 6A and 6B respectively show schematic diagrams of processing nodes at different running times according to an embodiment of the present invention; and

[0036] Figure 7A and 7B respectively show schematic diagrams of deploying an engine operation segment to different acceleration unit cores for execution according to an embodiment of the present invention. Detailed implementation manners

[0037] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0038] Figure 1A Depicts a block diagram of an example computing system 9100 according to an example embodiment of the present disclosure. The system 9100 includes a user computing device 9110, a server computing system 9130, and a training computing system 9150 communicatively coupled via a network 9180.

[0039] The user computing device 9110 can be any type of computing device, including but not limited to, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (a smart phone or a tablet), a game console or controller, a wearable computing device, an embedded computing device, an edge computing device, or any other type of computing device. The user computing device 9110 can be deployed as an edge intelligent device at a user site and interact with the user to process user input.

[0040] The user computing device 9110 may store or include one or more machine learning models 9120. The machine learning models 9120 may be designed to perform various tasks, such as image classification, object detection, speech recognition, machine translation, content filtering, and so on. The machine learning models 9120 may be machine learning models of other types such as neural networks (e.g., deep neural networks) or including non-linear models and / or linear models. Examples of the machine learning models 9120 include but are not limited to various types of deep neural networks (DNNs), such as feedforward neural networks, recurrent neural networks (RNNs, e.g., long short-term memory recurrent neural networks (LSTMs), Transformer neural networks (Transformers) with or without an attention mechanism), convolutional neural networks (CNNs), or other forms of neural networks. The machine learning models 9120 may include one machine learning model or may be a combination of multiple machine learning models.

[0041] Figure 1B Shown therein is a deep learning model, i.e., a neural network, as a machine learning model 9120 according to some embodiments. The neural network has a hierarchical architecture, and each network layer has one or more processing nodes (referred to as neurons or filters) for processing. In a deep neural network, the output after processing by the previous layer is the input of the next layer, where the first layer in the architecture receives the network input for processing, and the output of the last layer is provided as the network output. As Figure 1B shown, the machine learning model 9120 includes network layers 9122, 9124, 9126, etc., where the network layer 9124 receives the network input and the network layer 9126 provides the network output.

[0042] In a deep neural network, the main processing operations within the network are interleaved linear and non-linear transformations. These processes are distributed among the respective processing nodes. Figure 1B An enlarged view of a node 9121 in the model 9120 is also shown. The node 9121 receives a plurality of input values a1, a2, a3, etc., and processes the input values based on corresponding processing parameters (such as weights w1, w2, w3, etc.) to generate an output z. The node 171 may be designed to process the input using an activation function, which can be expressed as:

[0043] z = σ(w T α + b) (1)

[0044] where α ∈ R N represents the input vector of the node 9121 (which includes elements a1, a2, a3, etc.); w ∈ R NRepresents the weight vector (including elements w1, w2, w3, etc.) in the processing parameters used by node 9121, where each weight is used to weight the corresponding input; N represents the number of input values; b ∈ R N Represents the bias vector (including elements b1, b2, b3, etc.) in the processing parameters used by node 9121, where each bias is used to bias the corresponding input and the weighted result; σ() represents the activation function used by node 9121, and the activation function can be a linear function or a non-linear function. Commonly used activation functions in neural networks include the sigmoid function, ReLu function, tanh function, maxout function, etc. The output of node 9121 can also be referred to as the activation value. Depending on the network design, the output (i.e., the activation value) of each network layer can be provided as input to one, multiple, or all nodes of the next layer.

[0045] Each network layer in the machine learning model 9120 can include one or more nodes 9121. When viewing the processing in the machine learning model 9121 in terms of network layers, the processing of each network layer can also be similarly represented in the form of formula (1) or formula (2). At this time, a represents the input vector of the network layer, and w represents the weight of the network layer.

[0046] It should be understood that Figure 1B The architecture of the illustrated machine learning model, as well as the number of network layers and processing nodes therein, are all schematic. In different applications, according to needs, the machine learning model can be designed to have other architectures.

[0047] Continuing to refer to Figure 1A , in some implementations, the user computing device 9110 can receive the machine learning model 9120 from the server computing system 130 through the network 9180, store it in the memory of the user computing device, and use or implement it by an application in the user computing device.

[0048] In other implementations, the user computing device 9110 can call the machine learning module 9140 stored and implemented in the server computing system 9130. For example, the machine learning model 9140 can be implemented by the server computing system 9130 as part of a web service, so that the user computing device 9110 can call the machine learning model 9140 implemented as a web service, for example, through the network 9180 and according to the client-server relationship. Therefore, the machine learning modules that can be used at the user computing device 102 include the machine learning model 9120 stored and implemented at the user computing device 9110 and / or the machine learning model 9140 stored and implemented at the server computing system 9130.

[0049] The server computing system 9130 can include one or more server computing devices. In cases where the server computing system 9130 includes multiple server computing devices, these server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0050] As described above, the server computing system 9130 can store or include one or more machine learning models 9140. Similar to the machine learning model 9120, the machine learning model 9140 can be designed to perform various tasks, such as image classification, object detection, speech recognition, machine translation, content filtering, and so on. The model 9140 can include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.

[0051] The user computing device 9110 and / or the server computing system 9130 can train the model 9120 and / or 9140 via interaction with a training computing system 9150 communicatively coupled through a network 9180. The training computing system 9150 can be separate from the server computing system 9130 or can be a part of the server computing system 9130.

[0052] Similar to the server computing system 9130, the training computing system 9150 can include one or more server computing devices or otherwise be implemented by one or more server computing devices.

[0053] The training computing system 9150 can include a model trainer 9160 that trains the machine learning models 9120 and / or 9140 stored at the user computing device 9110 and / or the server computing system 9130 using various training or learning techniques such as, for example, backpropagation of error. In some implementations, performing backpropagation of error can include performing truncated backpropagation through time. The model trainer 9160 can perform a variety of generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0054] Specifically, the model trainer 9160 can train the machine learning models 9120 and / or 9140 based on a collection of training data 9162. The training data 9162 can include multiple different training data sets, each of which, for example, respectively helps train the machine learning models 9120 and / or 9140 to perform multiple different tasks. For example, the training data sets include data sets that help the machine learning models 9120 and / or 9140 perform object detection, object recognition, object segmentation, image classification, and / or other tasks.

[0055] In addition, in some implementations, the model trainer 9160 may modify the machine learning model 9140 in the server computing system 9130 to obtain a machine learning model 9120 suitable for use in the user computing device 9110. These modifications may include, for example, reducing the number of various parameters in the model, storing parameter values with a smaller precision, etc., so that the trained machine learning model 9120 and / or 9140 is suitable for running considering the different processing performances of the server computing system 9130 and the user computing device 9110.

[0056] The model trainer 9160 includes computer logic for providing the desired functionality. The model trainer 9160 may be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 9160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 9160 includes a set of one or more computer-executable instructions stored in a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium. In some implementations, the model trainer 9160 may be replicated and / or distributed across multiple different devices.

[0057] The network 9180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. Generally, communication over the network 9180 may be carried via any type of wired and / or wireless connection, using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML, and JSON), and / or security schemes (e.g., VPN, HTTPS, SSL).

[0058] Figure 1A The user computing device 9110, the server computing system 9130, and the training computing system 9150 in the illustrated example computing system 9100 may all be implemented by the computing device 9200 as described below. Figure 2A A schematic diagram of a computing device 200 according to an embodiment of the present invention is shown.

[0059] Figure 2AA structural block diagram inside a computing device 200 according to an embodiment of the present disclosure is shown. The computing device 200 includes a memory 210, a scheduling unit cluster 270, and an acceleration unit cluster 280 connected by a bus. The scheduling unit cluster 270 includes a plurality of scheduling units 220. The acceleration unit cluster 280 includes a plurality of acceleration units 230. In the embodiments of the present disclosure, the acceleration units are specialized processing units mainly designed to accelerate the operation processing speed of neural network models, and can be embodied as processing units (NPUs) specifically designed for neural network operation processing, graphics processing units (GPUs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. The scheduling units are processing units that schedule the acceleration units and allocate the sequence of instructions to be executed to each acceleration unit, and can take various forms such as a central processing unit (CPU), an application-specific integrated circuit (ASIC), and a field-programmable gate array (FPGA).

[0060] In the architecture design of traditional central processing units, a large part of the space in the architecture is occupied by the control unit and the storage unit, while the space occupied by the computing unit is insufficient. Therefore, it is very effective in logical control, but inefficient in large-scale parallel computing. Therefore, various specialized acceleration units have been developed to perform more effective processing to improve the operation speed for different functions and different fields of computing. The acceleration units proposed in the present invention are processing units dedicated to accelerating the operation processing speed of neural network models. It is an architecture that uses data-driven parallel computing and is a processing unit for processing a large number of operations (such as convolution, pooling, etc.) of each neural network node. Since the data and intermediate results in a large number of operations (such as convolution, pooling, etc.) of each neural network node are closely related throughout the computing process and are frequently used, with the existing central processing unit architecture, due to the small memory capacity inside the core of the central processing unit, a large amount of frequent access to the off-core memory is required, resulting in inefficient processing. By using such acceleration units dedicated to accelerating the operation processing speed of neural network models, since each core has on-chip (on-chip) memory with a storage capacity suitable for neural network computing, frequent access to the off-core memory is avoided, and thus the processing efficiency can be greatly improved and the computing performance can be enhanced.

[0061] The acceleration units 230 are to be scheduled by the scheduling units 220. As Figure 2A shown, various neural network models are stored in the memory 210, including the nodes of these models and the weight data of the nodes, etc. These neural network models are Figure 2AOne scheduling unit 220 is deployed to an acceleration unit 230. That is, the scheduling unit 220 can send the addresses of the parameters (such as the weights of each node) in the model in the form of instructions to the acceleration unit 230 in the memory 210. When the acceleration unit 230 actually uses the neural network model for calculation, it will directly address these parameters in the memory 210 according to the addresses of these parameters (such as weights) in the memory 210 and cache them in its on-chip memory. When the acceleration unit 230 actually uses the neural network model for calculation, the scheduling unit 220 will also send the input parameters of the model to the acceleration unit 230 in the form of instructions and cache them in the on-chip memory of the acceleration unit 230. In this way, the acceleration unit 230 can perform inference calculations based on these input parameters and the parameters (such as weights) in the model.

[0062] Internal Structures of Scheduling Unit and Acceleration Unit

[0063] The following combines Figure 2B the internal structure diagrams of the scheduling unit 220 and the acceleration unit 230 to specifically illustrate how the scheduling unit 220 schedules the acceleration unit 230 to work.

[0064] As Figure 2B shown, the scheduling unit 220 includes multiple processor cores 222 and a cache 221 shared by the multiple processor cores 222. Each processor core 222 includes an instruction fetch unit 203, an instruction decoding unit 224, an instruction issuing unit 225, and an instruction execution unit 226.

[0065] The instruction fetch unit 223 is used to transfer the instructions to be executed from the memory 210 to the instruction register (which can be Figure 3 one of the register files 229 shown for storing instructions) and receive the next instruction fetch address or calculate the next instruction fetch address according to the instruction fetch algorithm. The instruction fetch algorithm includes, for example: incrementing or decrementing the address according to the instruction length.

[0066] After fetching the instructions, the scheduling unit 220 enters the instruction decoding stage. The instruction decoding unit 224 decodes the fetched instructions according to a predetermined instruction format to obtain the operand acquisition information required by the fetched instructions, so as to prepare for the operations of the instruction execution unit 225. The operand acquisition information points to, for example, immediate numbers, registers, or other software / hardware that can provide source operands.

[0067] The instruction issuing unit 225 is located between the instruction decoding unit 224 and the instruction execution unit 226 and is used for instruction scheduling and control to efficiently allocate each instruction to different instruction execution units 226, making it possible to perform parallel operations on multiple instructions.

[0068] After the instruction emission unit 225 emits an instruction to the instruction execution unit 226, the instruction execution unit 226 starts to execute the instruction. However, if the instruction execution unit 226 determines that the instruction should be executed by the acceleration unit, it forwards the instruction to the corresponding acceleration unit for execution. For example, if the instruction is an instruction for neural network inference, the instruction execution unit 226 no longer executes the instruction, but sends the instruction to the acceleration unit 230 via the bus, and the acceleration unit 230 executes it.

[0069] Inside the acceleration unit 30, there are multiple acceleration unit cores 236 (4 cores are shown in Figure 2, but those skilled in the art should understand that the acceleration unit 230 may also include other numbers of cores 236), a command processor 237, a direct memory access mechanism 235, and a bus channel 231.

[0070] The bus channel 231 is the channel for instructions to enter and exit the acceleration unit 230 from the bus.

[0071] The direct memory access (DMA) mechanism 235 is a function provided by some computer bus architectures. It enables data to be directly written from an attached device to the memory on the computer motherboard. This method greatly improves the data access efficiency compared to the method where all data transfers between devices have to go through the scheduling unit. Because of such a mechanism, the cores of the acceleration unit 230 can directly access the memory 210 to read parameters (such as the weights of each node) in the neural network model, etc., greatly improving the data access efficiency.

[0072] The command processor 237 distributes the instructions sent from the scheduling unit 220 to the acceleration unit 230 to the cores 236 for execution. The instruction execution unit 226 sends the sequence of instructions to be executed that need to be executed by the acceleration unit 230 to the acceleration unit 230. After the sequence of instructions to be executed enters through the bus channel 231, it is cached in the command processor 237. The command processor 237 selects a core 236 and distributes the instruction sequence to it for execution. In addition, the command processor 237 is also responsible for the synchronization operation between the cores 236.

[0073] Acceleration Unit Core

[0074] Figure 3 It is the internal structure diagram of the acceleration unit core 236 according to an embodiment of the present disclosure.

[0075] In one embodiment, as Figure 3 shown, the acceleration unit core 236 includes a tensor engine 310, a pooling engine 320, a memory copy engine 330, an sequencer 350, an instruction buffer 340, an on-chip memory 360 (also referred to as on-chip memory), and a constant buffer 370.

[0076] The instruction sequence assigned by the command processor 237 to the acceleration unit core 236 first enters the instruction cache 340 for caching. Then, the sequencer 350 fetches instructions from the instruction cache 340 in a first-in-first-out order and assigns them to the tensor engine 310, the pooling engine 320, or the memory copy engine 330 for execution according to the nature of the instructions. The tensor engine 310 is responsible for processing operations such as convolution and matrix multiplication in the neural network model. The pooling engine 320 is responsible for processing the pooling operations in the neural network model. The memory copy engine 330 is responsible for copying the operands stored in the on-chip memory 360 within the core 236 to the memory shared among the cores 236 or to the on-chip memory 360 within other cores 236. The sequencer 350 determines whether to assign the instruction to the tensor engine 310, the pooling engine 320, or the memory copy engine 330 according to the nature of the operation of the fetched instruction, such as convolution, matrix multiplication, pooling, or operand copy.

[0077] The on-chip memory 360 is the in-core memory that stores the weight parameters in the neural network model, as well as the input parameters and various intermediate results when the neural network model is actually used. The constant buffer 370 is a buffer that stores other constant parameters (e.g., hyperparameters in the neural network model) in the neural network model except for the weight parameters. As described above, during the process of the scheduling unit 220 pre-configuring the neural network model in the acceleration unit 230, the scheduling unit 220 sends the addresses of the parameters in the model in the memory 210 to the acceleration unit 230 in the form of instructions. These parameters include the weights of the nodes and other parameters (e.g., hyperparameters). For the weights, during the actual operation of the neural network model, the acceleration unit 230 fetches them from the corresponding positions in the memory 210 and places them in the on-chip memory 360. For other parameters, during the actual operation of the neural network model, the acceleration unit 230 fetches them from the corresponding positions in the memory 210 and places them in the constant buffer 370. Additionally, when the instruction for actual inference is assigned by the command processor 237 to the core 236 for execution, the input parameters (inputs to the neural network model) in the instruction are also stored in the on-chip memory 360. Additionally, when the tensor engine 310 and the pooling engine 320 perform convolution or pooling operations, the various intermediate results obtained are also stored in the on-chip memory 360.

[0078] Figure 4 FIG. 400 is a schematic diagram of a method for deploying a deep learning model to an acceleration unit according to an embodiment of the present invention. The method 400 is adapted to be executed in the computing device 200 shown above with reference to Figure 2A so as to deploy the deep learning model to an acceleration unit (such as the acceleration unit 230 described with reference to Figure 2A - 2B ). As described above with reference to Figure 1BAn example of a deep learning model is given. The deep learning model includes multiple model layers, and each model layer includes one or more processing nodes. According to an embodiment of the present invention, the processing node corresponds to all the processing nodes in a model layer, that is, the processing node corresponds to the model layer. According to another embodiment, the processing node corresponds to a part of the processing nodes in a model layer, that is, the processing node corresponds to a part of the model layer. According to still another embodiment, the processing node may correspond to multiple model layers. As long as the deep learning model can be characterized as multiple processing nodes executed in a certain order, such a partitioning method is within the protection scope of the present invention.

[0079] Method 400 starts at step S410. In step S410, according to the spatial size of the on-chip memory 360 of the acceleration unit core 236 in the acceleration unit 230, the deep learning model to be deployed is divided into multiple engine operation segments. Each engine operation segment includes a part of the processing nodes in the deep learning model to be deployed, and the storage space used when executed in the acceleration unit core 236 does not exceed the spatial size of the on-chip memory 360.

[0080] As referred to above Figure 3 When the deep learning model is to be deployed to the acceleration unit core 236 for execution, the input data per batch required for model calculation, the weight values of the processing nodes, the activation thresholds, and the output data generated after processing, etc. are stored in the on-chip memory 360 of the acceleration unit core 236. To avoid data exchange between on-chip and off-chip during model calculation due to the excessive size of the data required for deep learning model calculation and inability to store all of it in the on-chip memory 360, it is necessary to split the deep learning model in step S410. And the storage space used when each split engine operation segment is executed does not exceed the storage space size of the on-chip memory 360, thus avoiding performance degradation caused by frequent data interaction.

[0081] According to one embodiment, in step S410, the running order of each processing node in the deep learning model is simulated. Considering the computing characteristics of the deep learning model, various methods can be used to simulate the computing process of the deep learning model. For example, the above simulation can be performed using a software application developed for the acceleration unit 230.

[0082] During the simulation, various running times can be determined. The running time is some characteristic time points during the execution of the model. For example, the characteristic time points are the moments when some processing nodes complete execution, the moments when some other processing nodes start execution, etc. At each characteristic time point, a change in the storage space required for the calculation occurs. During the simulation, the processing nodes that perform calculations at each running time can be determined, and further, the storage space required by these processing nodes can be determined. The sum of the storage spaces required by these processing nodes is determined as the storage space required for the deep learning model calculation at this running time.

[0083] According to one embodiment, the storage space required for a processing node to execute includes the storage space size occupied by the data to be loaded into the on-chip memory 360 when the processing node is deployed to the acceleration unit core 236 for execution. These include but are not limited to the input data per batch, the weight values of the processing node, the activation threshold, and the output data generated after processing described above.

[0084] According to one embodiment, in step S410, through the above simulation process, an engine operation segment can be constructed using multiple processing nodes that are sequentially executed in order. The storage space required by this engine operation segment at each running time when executed in the acceleration unit core does not exceed the space size of the on-chip memory.

[0085] According to one embodiment, each currently executing processing node has associated dependent nodes. The dependent nodes include the processing nodes that have been executed and whose processing results are required by the currently executing processing node. The storage space required by the engine operation segment at each running time includes the sum of the storage space required by the currently executing processing node in this engine operation segment and the storage space required by the dependent nodes.

[0086] There can be various ways to construct the engine operation segment. According to one way, when the size of each model layer of the deep learning model is appropriate and the association between the model layers is large, that is, when the execution of the processing nodes in one model layer requires the output results of the processing nodes in multiple previous model layers, the processing nodes can be set to correspond to the entire model layer, and the deep learning model can be split into multiple sequentially executed engine operation segments.

[0087] According to another way, when the size of each model layer in the deep learning model is large and the association between the model layers is small, that is, when the execution of the processing nodes in one model layer requires the output results of the processing nodes in a few previous model layers, the processing nodes can be set to a part of the model layer, and the deep learning model can be split into multiple parallelly executed engine operation segments according to the storage space limit required by the processing nodes.

[0088] It is also possible to combine the above two methods, that is, the deep learning model can be split into multiple engine operation segments, some of which are executed sequentially and some are executed in parallel. All such construction methods are within the protection scope of the present invention.

[0089] Figure 5 FIG. shows a schematic diagram of a method 500 for splitting a deep learning model into multiple engine operation segments that are sequentially executed in step S410 according to an embodiment of the present invention. Figure 6A and Figure 6B FIG. shows Figure 5 a schematic diagram of a processing node at different running times during the execution of the method 500 shown.

[0090] The method 500 starts at step S510, in which the processing node currently being executed at the start time T1 of splitting is determined, and the storage space required at this running time is determined. As described above, the storage space required at each running time is the sum of the storage spaces required by the processing node and the dependent nodes being executed at this running time.

[0091] As shown in reference Figure 6A FIG., it is assumed that at the running time T1, the processing nodes currently being processed in the deep learning model are 2, 3, and 4. Therefore, the processing nodes 2, 3, and 4 can be assigned to the running pool. The dependent nodes, that is, the processing nodes that have been executed but whose calculation results still need to be used by the processing nodes, are 0 and 1. Therefore, the processing nodes 0 and 1 can be assigned to the survival pool. The processing nodes that have not been executed are 5, 6, 7, etc., and these processing nodes are assigned to the to-be-processed pool. Therefore, at the running time T1, the on-chip storage space required by the acceleration unit core is the storage space required by each processing node (i.e., processing nodes 0 - processing node 4) in the survival pool and the running pool, denoted as mem[T1].

[0092] Subsequently, in step S520, for a continuous plurality of running times after the first running time T1, the processing node currently being executed at each running time is determined and the storage space required at this running time is calculated. Similarly, Figure 6BShows the next runtime T2 after runtime T1. The processing nodes currently being processed in the deep learning model are 5, 6, and 7. Therefore, processing nodes 5, 6, and 7 can be assigned to the running pool. At this time, the dependent nodes, that is, the processing nodes that have been executed but whose calculation results still need to be used by the processing nodes, are 3 and 4. Therefore, processing nodes 3 and 4 can be assigned to the survival pool. The processing nodes that have not been executed are 8, 9, 10, etc., and these processing nodes are assigned to the pending pool. Therefore, at runtime T2, the on-chip storage space required by the acceleration unit core is the storage space required by each processing node (i.e., processing nodes 3 - processing node 7) in the survival pool and the running pool, denoted as mem[T2].

[0093] Then in step S530, it is determined whether the difference between the storage space required at the runtime calculated in step S520 and the space size of the on-chip memory 360 is within a predetermined threshold (i.e., whether it is relatively close to the space size of the on-chip memory 360) to determine whether this runtime is the end time of splitting. If this runtime is not the end time of splitting, the storage space required at this runtime will be recorded, such as mem[T2], mem[T3]…, mem[Tn], and so on. And further in step S560, it is determined whether this runtime is the last runtime of the deep learning model, that is, when the splitting of this deep learning template ends, and at this time this moment is the last runtime. Then in step S570, all the processing nodes executed between the last runtime and the start time of splitting are divided into the last engine operation segment, and the execution of method 500 ends.

[0094] On the contrary, if it is determined in step S560 that this runtime is not the last runtime of the deep learning model, then method 500 returns to step S520 to continue calculating the storage space required for the next runtime.

[0095] If it is determined in step S530 that this runtime is the end time of splitting, then in step S540, the processing nodes executed between the start time of splitting and the end time of splitting are divided into engine operation segments. Subsequently, in step S550, the next runtime is set as the new start time of splitting, so as to return to step S510 to start determining new engine operation segments.

[0096] Using method 500, a deep learning model can be split into multiple engine operation segments that can be executed sequentially.

[0097] Return reference Figure 4According to the description, after splitting the deep learning model into multiple engine operation segments in step S410, in step S420, according to the storage space size and / or dependency relationship required by each engine operation segment, each engine operation segment is deployed to the corresponding acceleration unit core for execution.

[0098] According to one embodiment, the deployment can be performed according to the storage space size of the engine operation segment. The storage space of the engine operation segment is set to the maximum storage space required by multiple runtimes included in the engine operation segment. As referenced above Figure 5 As described, for the storage spaces required by multiple runtimes in an engine operation unit, mem[T1], mem[T2], mem[T3]…, mem[Tn], the maximum value among these storage spaces is set as the storage space of the engine operation segment.

[0099] Subsequently, the engine operation segment with the largest storage space and the engine operation segment with the smallest storage space can be deployed to the same acceleration unit core for execution.

[0100] Figure 7A Such a deployment method is shown. As Figure 7A shown, for four engine operation segments e0, e1, e2, and e3, the storage space of e0 is the largest (0x500) while the storage space of e1 is the smallest (0x100), and the sum of the storage spaces of the two engine operation segments does not exceed the on-chip storage space (0x800) of the acceleration unit core. Therefore, e0 and e1 can be deployed to the same acceleration unit core (such as core0), and e2 and e3 can be deployed to another acceleration unit core (such as core1).

[0101] According to another embodiment, the deployment can be performed according to the execution timing and dependency relationship of the engine operation segments. When constructing the engine motion unit in step S410, through software simulation, the multiple runtimes included in each engine operation segment can be determined, thereby further determining the running time of each engine operation segment. In addition, according to the association relationship between the running nodes and the dependency nodes, the dependency relationship between each engine operation segment can be determined by using the processing nodes included in each engine operation segment.

[0102] After determining the running times of each engine operation segment, the execution timing of the engine operation segment can be determined. In step S420, the engine operation segments with close running times can be deployed to different acceleration unit cores for execution. In addition, after determining the dependency relationship between each engine operation segment, in step S420, the engine operation segments without dependency relationship can be deployed to different acceleration unit cores for execution.

[0103] Figure 7BAnother deployment method is given. As Figure 7B shown, there is an ordered data dependency between engine processing units op0, op1 / op2, and op3. There is no data dependency between engine processing units op1 and op2, and their running time periods are quite similar (the length of the blocks in the figure represents the relationship of running time). Using Figure 7B the deployment method provided, deploy engine processing units op1 and op2 on different acceleration unit cores core1 / core2, which can perform parallel computing to improve performance. Moreover, the execution times of engine processing units op1 and op2 are quite similar, and there will be no waiting phenomenon for the execution of op3 deployed on acceleration unit core3. Through such a deployment scheme, according to the storage space size, mutual relationship (whether there is data dependency), and execution timing required by each engine operation segment, etc., the engine operation segments can be deployed to the corresponding acceleration unit cores in an appropriate manner to execute, thereby improving the utilization efficiency of the acceleration unit cores and the inference speed of the deep learning model in the acceleration unit.

[0104] The various technologies described herein can be implemented in combination with hardware or software, or a combination of them. Thus, the method and device of the present invention, or certain aspects or parts of the method and device of the present invention, can take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, USB flash drive, floppy disk, CD-ROM, or any other machine-readable storage medium, where when the program is loaded into a machine such as a computer and executed by the machine, the machine becomes a device for practicing the present invention.

[0105] In the case where the program code is executed on a programmable computer, the computing device generally includes a processor, a processor-readable storage medium (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device. Among them, the memory is configured to store the program code; the processor is configured to execute the method of the present invention according to the instructions in the program code stored in the memory.

[0106] By way of example and not limitation, the readable medium includes a readable storage medium and a communication medium. The readable storage medium stores information such as computer-readable instructions, data structures, program modules, or other data. The communication medium generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery medium. A combination of any of the above is also included within the scope of the readable medium.

[0107] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the examples of the present invention. Based on the above description, the structure required to construct such systems will be apparent. Additionally, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of a particular language above is for the purpose of disclosing example embodiments of the present invention.

[0108] In the specification provided herein, a number of specific details are set forth. However, it will be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure an understanding of the present specification.

[0109] Similarly, it should be understood that, in order to streamline the present disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present invention.

[0110] Those skilled in the art should understand that the modules, or units, or components of the devices in the examples disclosed herein may be arranged in the devices as described in the embodiments, or alternatively may be located in one or more devices different from those in the examples. The modules in the foregoing examples may be combined into one module or further divided into multiple sub-modules.

[0111] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device thus disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0112] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0113] In addition, some of the embodiments are described herein as a combination of methods or method elements that can be implemented by a processor of a computer system or by other devices performing the functions. Therefore, a processor having the necessary instructions for implementing the method or method element forms a device for implementing the method or method element. In addition, the elements described herein in the device embodiments are examples of such devices: the device is used to implement the functions performed by the elements for the purpose of implementing the invention.

[0114] As used herein, unless otherwise specified, the use of ordinal numbers "first", "second", "third", etc. to describe ordinary objects only indicates different instances of similar objects, and does not intend to imply that the objects so described must have a given order in terms of time, space, sorting, or in any other way.

[0115] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art, having the benefit of the foregoing description, will appreciate that other embodiments can be conceived within the scope of the invention as thus described. In addition, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and not to limit or circumscribe the inventive subject matter. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure herein is illustrative and not restrictive, and the scope of the invention is defined by the appended claims.

Claims

1. A method for deploying a deep learning model to an acceleration unit, the deep learning model including a plurality of processing nodes, the method comprising the steps: Divide the deep learning model into multiple engine operation segments according to the on-chip memory size of the acceleration unit cores in the acceleration unit, where the storage space used by each engine operation segment when executed in the acceleration unit cores does not exceed the on-chip memory size, and includes a part of multiple processing nodes in the deep learning model; and Deploying each engine operation segment to a corresponding acceleration unit core for execution according to the storage space size and / or dependency relationship to be used by each engine operation segment; wherein, The step of dividing the deep learning model into a plurality of engine operation segments includes: At a first runtime, determining the processing node currently being executed at the first runtime and the storage space required at the first runtime; At a plurality of consecutive runtimes after the first runtime, determining the processing node currently being executed at each runtime and the storage space required at that runtime until a second runtime, wherein the difference between the storage space required at the second runtime and the size of the on-chip memory is within a predetermined threshold; and Dividing all the processing nodes executed between the first runtime and the second runtime into engine operation segments.

2. The method according to claim 1, wherein the step of dividing the deep learning model into a plurality of engine operation segments includes: Constructing engine operation segments using a plurality of processing nodes that are sequentially executed in order, so that the storage space required at each runtime when the engine operation segments are executed in the acceleration unit core does not exceed the size of the on-chip memory.

3. The method according to claim 2, wherein the storage space required at the runtime includes the sum of the storage space required by the currently executed processing node and the storage space required by the dependent nodes, and the dependent nodes include the processing nodes that have been executed and whose processing results are required by the currently executed processing node.

4. The method according to claim 1, wherein constructing sequentially executed engine operation segments according to the execution order of the processing nodes further includes: After determining the engine operation segments, setting the subsequent runtime as a new first runtime to determine a new second runtime, and dividing all the processing nodes executed between the new first runtime and the second runtime into new engine operation segments.

5. The method according to claim 2 or 3, wherein the step of dividing the deep learning model into a plurality of engine operation segments includes: Constructing parallelly executed engine operation segments according to the storage space required at each runtime.

6. The method according to claim 2 or 3, wherein the storage space required by the processing node includes the storage space occupied by one or more of the following items: Input variables, weight values, activation parameters, and output variables.

7. The method according to claim 2 or 3, wherein setting the storage space of the engine operation segment to the maximum storage space required by the plurality of runtimes included in the engine operation segment; and The step of deploying each engine operation segment to a corresponding acceleration unit core for execution includes: Deploying the engine operation segment with the largest storage space and the engine operation segment with the smallest storage space to the same acceleration unit core for execution.

8. The method according to claim 2 or 3, the step of dividing the deep learning model into a plurality of engine operation segments further includes the steps: Determine the running time of each engine operation segment according to the multiple run times included in each engine operation segment; and Determine the dependency relationship between each engine operation segment according to the processing nodes included in each engine operation segment.

9. The method according to claim 8, wherein the step of deploying each engine operation segment to a corresponding acceleration unit core for execution includes: Deploy the engine operation segments without dependency relationship to different acceleration unit cores for execution.

10. The method according to claim 8, wherein the step of deploying each engine operation segment to a corresponding acceleration unit core for execution includes: Deploy the engine operation segments with close running times to different acceleration unit cores for execution.

11. The method according to any one of claims 2 or 3, wherein the deep learning model includes a plurality of model layers, and each processing node corresponds to a model layer or a part of a model layer.

12. An artificial intelligence device, comprising: An acceleration unit, including one or more acceleration unit cores, and each acceleration unit core has a corresponding on-chip memory; And A processor, adapted to execute the method according to any one of claims 1-11, so as to deploy the deep learning model to the acceleration unit for execution.

13. A data center, comprising: A plurality of artificial intelligence devices according to claim 12.

14. A data center, comprising: A plurality of acceleration units, each acceleration unit includes one or more acceleration unit cores, and each acceleration unit core has a corresponding on-chip memory, wherein the acceleration unit uses the method according to any one of claims 1-11 to deploy the deep learning model.

15. A computing device, comprising: At least one processor; And A memory storing program instructions, wherein the program instructions are configured to be adapted to be executed by the at least one processor, and the program instructions include instructions for executing the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Software virtual machine for acceleration of transactional data processing

    CN103930875A

  • Method for generating instruction sequence and method and device for executing neural network operation

    CN109919311A