Method and device for optimizing a deep learning computational graph

By merging and dividing operators in DNN computation graphs based on output properties and platform capacity, the method optimizes DNN processing efficiency and reduces cache pressure, addressing the inefficiencies of existing fusion methods.

DE112022007835T5Pending Publication Date: 2025-08-07INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE112022007835
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Deep neural network (DNN) optimization is computationally expensive and time-consuming, particularly for complex DNNs with large cache pressures due to inefficient operator fusion methods like fixed pattern and polyhedral-based loop fusion.

Method used

A method involving merging memory-intensive operators into computational-intensive operators, dividing the resulting graph into partial graphs, and merging operators within each partial graph to optimize the computation graph, using heuristic rules based on output properties and platform capacity.

Benefits of technology

Improves optimization efficiency and reduces cache pressure by effectively restructuring the computation graph, enhancing performance in DNN processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are apparatus and methods for optimizing a deep learning computational graph. The method includes obtaining a deep learning computational graph including computationally intensive operators and memory-intensive operators; merging the memory-intensive operators into the computationally intensive operators to generate a new computational graph; splitting the new computational graph into subcomputational graphs; and merging computationally intensive operators in each of the subcomputational graphs to generate an optimized computational graph. Other embodiments may also be disclosed and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

field of technology

[0001] Embodiments described herein generally relate to deep learning (DL) networks and, more particularly, relate to a method and apparatus for optimizing deep learning computational graphs. Technical background

[0002] Deep neural network (DNN) models have become deeper and more complex these days, with hundreds or even more layers. To obtain a suitable DNN, a deep learning computation graph generated from its corresponding intermediate representation (IR) should be optimized. However, for such a complex DNN, optimization is generally computationally expensive and time-consuming, and can lead to significant cache pressure. Summary

[0003] One aspect of the disclosure provides a method for optimizing a deep learning computation graph, comprising: obtaining a deep learning computation graph comprising computationally intensive operators and memory-intensive operators; merging the memory-intensive operators into the computationally intensive operators to generate a new computation graph; splitting the new computation graph; and merging computationally intensive operators in each of the sub-computation graphs to generate an optimized computation graph.

[0004] Another aspect of the disclosure provides an apparatus for optimizing a deep learning computation graph, comprising: interface circuitry; and processor circuitry coupled to the interface circuitry and configured to: obtain a deep learning computation graph comprising computationally intensive operators and memory-intensive operators; merge the memory-intensive operators into the computationally intensive operators to generate a new computation graph; split the new computation graph; and merge computationally intensive operators in each of the partial computation graphs to generate an optimized computation graph.

[0005] Another aspect of the disclosure provides a computer-readable medium having instructions stored thereon, the instructions, when executed by a processor, causing the processor to: obtain a deep learning computational graph comprising computationally intensive operators and memory-intensive operators; merge the memory-intensive operators into the computationally intensive operators to generate a new computational graph; split the new computational graph; and merge computationally intensive operators to generate an optimized computational graph. Brief description of the drawings

[0006] Embodiments of the disclosure are illustrated by way of example and not limitation in conjunction with the figures of the accompanying drawings, in which like reference numerals refer to similar elements and wherein: Fig. 1 illustrates a flowchart of an example method for optimizing a deep learning computation graph according to an embodiment of the present application; Fig. 2 illustrates a flowchart of another example of a method for optimizing a deep learning computation graph according to an embodiment of the present application; Fig. 3 illustrates a flowchart of another example of a method for optimizing a deep learning computation graph according to an embodiment of the present application; Fig. 4 illustrates a flowchart of another example of a method for optimizing a deep learning computation graph according to an embodiment of the present application; Fig. 5 illustrates a schematic diagram of an example of splitting for batch and scan according to an embodiment of the present application; Fig. 6 illustrates a schematic diagram of another example of dividing for a scan according to an embodiment of the present application; Fig. 7 illustrates a block diagram of an example of an apparatus for optimizing a deep learning computation graph according to an embodiment of the present application; Fig. 8 is a block diagram illustrating components capable of reading instructions from a machine-readable or computer-readable medium and performing any one or more of the methodologies discussed herein, according to some example embodiments; and Fig. 9 is a block diagram of an example processor platform in accordance with some embodiments of the disclosure. Detailed description of embodiments

[0007] Various aspects of the exemplary embodiments are described using terms commonly employed by those skilled in the art to convey the substance of the disclosure to others skilled in the art. However, it will be apparent to those skilled in the art that many alternative embodiments may be practiced using portions of the described aspects. For purposes of explanation, specific numbers, materials, and configurations are set forth in order to provide a thorough understanding of the illustrative embodiments. However, it will be apparent to those skilled in the art that alternative embodiments may be practiced without the specific details. In other instances, well-known features may be omitted or simplified to avoid misunderstanding from the illustrative embodiments.

[0008] Furthermore, various operations are described sequentially as multiple discrete operations in a manner that is helpful for understanding the example embodiments; however, the order of description should not be construed to imply that these operations are necessarily order-dependent. In particular, it is not necessary that these operations be performed in the order of presentation.

[0009] The terms "in one embodiment," "in an embodiment," "in some embodiments," and "in some embodiments" are used repeatedly herein. The term generally does not refer to the same embodiment; but it can refer to both. The terms "comprising," "having," and "including" are synonymous unless the context dictates otherwise. The terms "A or B" and "A / B" mean "(A), (B), or (A and B)."

[0010] To obtain a suitable deep neural network (DNN) from a deep learning framework, the following processing stages can generally be performed: a) Graph construction stage, in which a computational graph is constructed via its intermediate representation (IR) according to the information from the deep learning framework; b) compilation stage, in which the computation graph is transformed (e.g., optimized) while the IR is optimized and reduced to the hardware-specific IR; and c) Code generation stage, in which the binary code or equivalent representation is generated based on the optimized IR.

[0011] During the compilation stage, the computation graph is optimized, for example, by operator fusion.

[0012] Traditionally, operator fusion is usually performed using a fixed-pattern approach or polyhedral-based loop fusion. However, the fixed-pattern approach is limited by the fixed special operators it contains and cannot be universally applied. Polyhedral-based loop fusion may miss potential fusion opportunities due to the lack of operator-level information.

[0013] Furthermore, for a complex DNN with a larger number of layers, the efficiency of the above approaches may be very low and they may cause large cache pressure.

[0014] Fig. 1 illustrates a flowchart of an example method for optimizing a deep learning computation graph according to an embodiment of the present application.

[0015] The method herein may be performed by any suitable device, such as a deep learning compiler or a processor.

[0016] The deep learning computation graph herein may be a deep learning computation graph for any DNN model (such as a convolutional neural network (CNN) model with a large batch size) for inference (e.g., RN50 throughput in MLPerf) or training (e.g., with a deep learning recommendation model (DLRM)). The deep learning computation graph may be a computation graph created by a deep learning compiler, as described above.

[0017] With reference to Fig. 1, a deep learning computation graph is obtained in block S110, which includes computation-intensive operators and memory-intensive operators.

[0018] For example, the computationally intensive operator of the deep learning computational graph can be a convolution or a Matmul. The memory-intensive operator can be element-wise, binary, or memory movement. In general, for a deep learning computational graph for a CNN model, the computationally intensive operator can be followed by one or more memory-intensive operators.

[0019] In block S120, the memory-intensive operators are merged into the computation-intensive operators to generate a new computation graph.

[0020] In one embodiment, one or more sequential memory-intensive operators can be merged into a preceding or following computation-intensive operator. In this way, a new computation graph can be generated. The new computation graph can include multiple layers, and each of the layers can include multiple computation-intensive operators. Thus, the new computation graph contains only computation-intensive operators.

[0021] In block S130, the new calculation graph is divided into sub-calculation graphs.

[0022] For example, each subcomputational graph may contain one or more layers of the new computational graph.

[0023] In one embodiment, the new computation graph may be divided into the sub-computation graphs based on an output property of each layer of the new computation graph and a platform capacity of a platform on which the new computation graph is to be executed.

[0024] For example, the output property of each layer can be an output property indicating the buffer capacity that should be allocated for operator fusion for that layer.

[0025] In one embodiment, each of the computationally intensive operators of the new computational graph may output an output activation, and the output activations of the computationally intensive operators of each layer form an output stack. The output property may include a size of an output stack for each layer and / or a size of an output activation of each computationally intensive operator of the new computational graph.

[0026] Regarding the platform capacity, in one embodiment, the new computation graph may be executed on a central processing unit (CPU), and then the platform capacity may include a data cache unit (DCU) size and a mid-level cell (MLC) size of the CPU.

[0027] The embodiments for splitting partial computation graphs based on an output property and the platform capacity in block S130 are further described with reference to the following Fig. 2 and Fig. 3 described.

[0028] In block S140, the computationally intensive operators in each of the sub-computation graphs are merged to generate an optimized computation graph.

[0029] For example, after merging the computationally intensive operators in each of the subcomputation graphs, all computationally intensive operators form the optimized computation graph.

[0030] By dividing the computation graph into subcomputation graphs and merging the operators in each subcomputation graph, the efficiency and cache pressure can be improved to optimize the deep learning computation graph.

[0031] Fig. 2 illustrates a flowchart of another example of a method for optimizing a deep learning computation graph according to an embodiment of the present application.

[0032] In Fig. 2, blocks S110, S120 and S140 are similar to blocks S110, S120 and S140 in Fig. 1, and the difference between Fig. 2 and Fig. 1 is that the block S130 in Fig. 1 specifically by blocks S131, S132 and S133 in Fig. 2 is illustrated.

[0033] With reference to Fig. 2, in block S131, a division parameter is obtained by means of a heuristic rule based on the output characteristic of each layer.

[0034] In one embodiment, the splitting parameter may include a stack split number (x) and a spatial split number (y). The stack split number may correspond to a number of sub-stacks into which an output stack is to be split, and the spatial split number may correspond to a number of sub-activations into which an output activation is to be split. Splitting for the sub-stacks and sub-activations is further described below.

[0035] In block S132, a buffer size (a i ) for each layer to be assigned for an output batch and a weight for the layer based on the output property of the layer.

[0036] The buffer size used for each layer can be estimated using any estimation method.

[0037] In block S133, the layers are sequentially divided into the partial computation graphs in a topology order of the new computation graph based on the division parameter, the buffer size, and the platform capacity.

[0038] An embodiment for dividing the layers in block S133 is described with reference to the following Fig. 3 described.

[0039] Fig. 3 illustrates a flowchart of another example of a method for optimizing a deep learning computation graph according to an embodiment of the present application.

[0040] In Fig. 3, blocks S110, S120, S130 and S140 are similar to blocks S110, S120, S130 and S140 in Fig. 2, and the difference between Fig. 3 and Fig. 2 is that the block S133 in Fig. 2 specifically by blocks S133-1 and S133-2 in Fig. 3 is illustrated.

[0041] With reference to Fig. 3, in block S133-1, a reduced buffer size ( a R_i ) for each layer based on the splitting parameter (x and y) and the buffer size (a i ) for the shift.

[0042] In one embodiment, the reduced buffer size can be expressed by the following equation (1): aR_i=wi+ai×1x×y

[0043] In equation (1) a R_i the reduced buffer size for an i-th layer of the new computation graph in the topology order, W i shows the weight for the i-th layer, a i indicates the buffer size for the i-th layer, x indicates the stack split number, y indicates the spatial split number, and each of a R_i , W i , ai , x and y is greater than 0.

[0044] In block S133-2, one or more sequential layers that satisfy a predetermined condition are divided into a partial computation graph.

[0045] In one embodiment, in block S133-2, the one or more sequential layers for which a sum of the reduced buffer sizes for the one or more sequential layers is less than or equal to the platform capacity and a sum of the reduced buffer sizes for the one or more sequential layers and a subsequent layer is greater than the platform capacity are divided into a partial computation graph.

[0046] In other words, the reduced buffer sizes can be accumulated layer by layer from the first layer of the new computation graph, and once the above predetermined condition is met, the accumulation is performed again from the next layer.

[0047] In one embodiment, for each sub-computation graph, the following equation (2) is satisfied: ∑i=Ni=MaR_i≤T×(L1+L2)<∑i=Ni=M+1aR_i

[0048] In equation (2), N indicates a starting layer of the partial computation graph, M indicates a last layer of the partial computation graph, T indicates a threshold set based on the CPU, L1 indicates the size of the DCU, L2 indicates the size of the MLC, N, M, L1 and L2 are greater than 0, and T is greater than 0 and less than or equal to 1 (e.g., T ranges from 0.9 to 0.95).

[0049] After the splitting of the partial computation graphs is completed, the fusion of the computationally intensive operators can be performed. One embodiment of the fusion is shown in the following Fig. 4 shown.

[0050] Fig. 4 illustrates a flowchart of another example of a method for optimizing a deep learning computation graph according to an embodiment of the present application.

[0051] In Fig. 4, blocks S110, S120 and S130 are similar to blocks S110, S120 and S130 in Fig. 2, and the difference between Fig. 4 and Fig. 2 is that the block S140 in Fig. 2 specifically through blocks S141 and S142 in Fig. 4 is illustrated.

[0052] With reference to Fig. 4, in block S141, the output stack for each layer of the partial calculation graph is divided into the sub-stacks by the stack division number.

[0053] For example, if the stack split number is set to 2, the output stack is divided into 2 sub-stacks for each layer.

[0054] In block S142, the output activation of each computation-intensive operator of the partial computation graph is divided into the partial activations by the spatial division number.

[0055] For example, if the spatial division number is determined as 3, the output activation of each computationally intensive operator is divided into 3 partial activations.

[0056] In one embodiment, each output activation may include one or more output samples (and each sample may include one or more channels).

[0057] In this condition, the division in block S142 may be performed by dividing each output sample of the output activation along a height direction of the sample into sub-samples by the spatial division number.

[0058] The embodiments for splitting for partial stacks and partial activations are further described with reference to the following Fig. 5 and Fig. 6 described.

[0059] In block S143, the computationally intensive operators in the partial computation graph are merged based on the partial stacks and the partial activations.

[0060] The fusion of the computationally intensive operators in block S143 can be implemented using any suitable fusion solution.

[0061] The following Table 1 provides an example machine-readable language for implementing the fusion in block S140. Table 1 0uter Loop{ / / batch axis Inner loop { / / height axis(optional) partial compute-intensive op1 partial compute-intensive op2 } } 0uter Loop{ / / batch axis Inner loop { / / height axis(optional) partial compute-intensive op3 ... partial compute-intensive opN } }

[0062] It is understood that the merging in block S140 may be implemented by any other machine-readable language that can realize the merging as described above.

[0063] Fig. 5 illustrates a schematic diagram of an example of splitting for batch and scan according to an embodiment of the present application.

[0064] Fig. Figure 5 shows an output stack containing four samples: Sample 1, Sample 2, Sample 3, and Sample 4. Each sample contains four channels. For example, if the computation graph is used to process an image with four channels (e.g., red, green, blue, and write channels), each sample can contain four corresponding channels.

[0065] In Fig. 5, the X-axis and Y-axis directions indicated by the arrows are a stacking direction for dividing the stack and a height direction for dividing the samples, respectively.

[0066] At the Fig. In the embodiment shown in Figure 5, the stack pitch number x is 2 and the spatial pitch number y is 3. The output stack is divided into 2 sub-stacks along the stacking direction X, as indicated by line L3, one of the sub-stacks containing samples 1 and 2 and the other sub-stack containing samples 3 and 4. Each of samples 1, 2, 3, and 4 is divided into 3 sub-samples along the height direction Y, as indicated by lines L1 and L2.

[0067] This diving for sampling may be suitable in the case where the corresponding computationally intensive operators of two sequential layers have kernels of the same dimension, e.g., 1x1 convolution kernels (e.g., with stride = 1).

[0068] In the case where the corresponding computationally intensive operators of two sequential layers have kernels with different dimensions, the division for the samples of a current layer may depend on the output samples of a subsequent layer. An embodiment of this case is described in the following Fig. 6 shown.

[0069] Fig. 6 illustrates a schematic diagram of another example of dividing for a scan according to an embodiment of the present application.

[0070] At the Fig. In the embodiment shown in Figure 6, samples c1 and c2 are samples output from a current layer, and samples f1 and f2 are samples output from a subsequent layer. The spatial division number y is 2. The kernel of the computationally intensive operator corresponding to samples c1 or c2 is a 1x1 convolution kernel (e.g., with stride = 1), and the kernel of the computationally intensive operator corresponding to samples f1 or f2 is a 3x3 convolution kernel (e.g., with stride = 1).

[0071] In this case, to obtain two sub-scans for sample f1 or f2 (as indicated by line L4), each sub-scan divided by the corresponding sample c1 or c2 should include four lines of the lines shown in sample c1 or c2. That is, for sample c1 or c2, the first four lines can be divided into one sub-scan (as indicated by line L5), and the last four lines can be divided into another sub-scan (as indicated by line L6), meaning that the middle two lines are reused across the two sub-scans.

[0072] It is understood that the above division for the output stack and sampling and the above dimensions of kernels of the computation-intensive operators are provided only as examples, the stack and sampling may be divided in any other way, and the kernels of the computation-intensive operators may have other dimensions according to actual requirements.

[0073] According to the method for optimizing the deep learning computation graph of the embodiments of the present application, optimization efficiency and cache pressure can be improved by dividing the computation graph into the sub-computation graphs and merging the operators in each sub-computation graph.

[0074] Fig. 7 illustrates a block diagram of an example of an apparatus 700 for optimizing a deep learning computation graph according to an embodiment of the present application.

[0075] With reference to Fig. 7, the apparatus 700 for optimizing a deep learning computation graph according to an embodiment of the present application includes processor circuitry 710 and interface circuitry 720 coupled to each other.

[0076] Processor circuitry 710 may be a deep learning compiler or any other processor. Processor circuitry 710 is configured to: obtain a deep learning computational graph including computationally intensive operators and memory-intensive operators; merge the memory-intensive operators into the computationally intensive operators to generate a new computational graph; split the new computational graph into subcomputational graphs; and merge computationally intensive operators in each of the subcomputational graphs to generate an optimized computational graph.

[0077] In one embodiment, the processor circuitry 710 may be configured to merge the memory-intensive operators into the computation-intensive operators by: merging one or more sequential memory-intensive operators into a previous or a following computation-intensive operator to generate the new computation graph.

[0078] In one embodiment, the new computational graph may include multiple layers, and each of the layers may include multiple computationally intensive operators.

[0079] In one embodiment, the new computation graph may be divided into the sub-computation graphs based on an output property of each layer of the new computation graph and a platform capacity of a platform on which the new computation graph is to be executed.

[0080] In one embodiment, each of the computationally intensive operators of the new computational graph outputs an output activation, and the output activations of the computationally intensive operators of each layer form an output stack. The output property may include a size of an output stack for each layer and / or a size of an output activation of each computationally intensive operator of the new computational graph.

[0081] In one embodiment, the platform capacity may include a data cache unit (DCU) size and a mid-level cell (MLC) size of a central processing unit (CPU) on which the new computation graph is to be executed.

[0082] In one embodiment, processor circuitry 710 may be configured to divide the new computation graph into the sub-computation graphs by: obtaining a division parameter using a heuristic rule based on the output property of each layer; obtaining a buffer size for each layer to be allocated for an output stack and a weight for the layer based on the output property of the layer; and sequentially dividing the layers, in a topology order of the new computation graph, into the sub-computation graphs based on the division parameter, the buffer size, and the platform capacity.

[0083] In one embodiment, the split parameter may include a stack split number and a spatial split number. The stack split number may correspond to a number of sub-stacks into which an output stack is to be split, and the spatial split number may correspond to a number of sub-activations into which an output activation is to be split.

[0084] In one embodiment, processor circuitry 710 may be configured to sequentially divide the layers into the division graphs in the topology order of the new computation graph based on the division parameter, the buffer size, and the platform capacity by: obtaining a reduced buffer size for each layer based on the division parameter and the buffer size for the layer; and dividing one or more sequential layers for which a sum of the reduced buffer sizes for the one or more sequential layers is less than or equal to the platform capacity and a sum of the reduced buffer sizes for the one or more sequential layers and a subsequent layer is greater than the platform capacity, into a partial computation graph.

[0085] In one embodiment, the reduced buffer size may be expressed by equation (1) above, and each shared partial computation graph may satisfy equation (2) above.

[0086] In one embodiment, processor circuitry 710 may be configured to merge the computationally intensive operators in each of the subcomputational graphs by: for each subcomputational graph, dividing the output stack for each layer of the subcomputational graph into the sub-stacks by the stack split number; dividing the output activation of each computationally intensive operator of the subcomputational graph into the sub-activations by the spatial split number; and merging the computationally intensive operators in the subcomputational graph based on the sub-stacks and the sub-activations.

[0087] In one embodiment, the processor circuitry 710 may be configured to divide the output activation of each computationally intensive operator of the sub-computation graph into sub-activations by the spatial division number by: dividing each output sample of the output activation into sub-samples along a height direction of the sample by the spatial division number.

[0088] The details of the operations performed by the processor circuitry 710 of the means 700 for optimizing the deep learning computation graph may refer to the above, in Fig. 1 to Fig. 6, which are not repeated here.

[0089] According to the means for optimizing the deep learning computation graph of the embodiments of the present application, optimization efficiency and cache pressure can be improved by dividing the computation graph into the sub-computation graphs and merging the operators in each sub-computation graph.

[0090] Further, a computer-readable medium is provided. The computer-readable medium is stored with instructions. The instructions, when executed by a processor, cause the processor to: obtain a deep learning computational graph including computationally intensive operators and memory-intensive operators; merge the memory-intensive operators into the computationally intensive operators to generate a new computational graph; split the new computational graph into subcomputational graphs; and merge computationally intensive operators in each of the subcomputational graphs to generate an optimized computational graph.

[0091] For example, the instructions, when executed by a processor, may cause the processor to perform the operations described above subject to Fig. 1 to Fig. 6, which are not repeated here.

[0092] Fig. Figure 8 is a block diagram illustrating components capable of reading instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and performing any one or more of the methodologies discussed herein, according to some example embodiments. In particular, Fig. 8 is a diagrammatic representation of hardware resources 800, including one or more processors (or processor cores) 810, one or more memory / storage devices 820, and one or more communication resources 830, each of which may be communicatively coupled via a bus 840. For embodiments employing node virtualization (e.g., NFV), a hypervisor 802 may be executed to provide an execution environment for one or more network slices / sub-slices for utilizing the hardware resources 800.

[0093] For example, the processors 810 may include a processor 812 and a processor 814, which may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a visual processing unit (VPU), a field-programmable gate array (FPGA), or any suitable combination thereof.

[0094] The memory / storage devices 820 may include main memory, disk storage, or any suitable combination thereof. The memory / storage devices 820 may include, but are not limited to, any type of volatile or non-volatile memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state memory, etc.

[0095] The communication resources 830 may include interconnect or network interface components or other suitable devices for communicating with one or more peripheral devices 804 or one or more databases 806 via a network 808. For example, the communication resources 830 may include wired communication components (e.g., for coupling via a Universal Serial Bus (USB)), cellular communication components, NFC components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components.

[0096] Instructions 850 may comprise software, a program, an application, an applet, an app, or other executable code to cause at least any of the processors 810 to perform one or more of the methodologies discussed herein. The instructions 850 may reside, in whole or in part, in at least one of the processors 810 (e.g., in the processor's cache memory), in the memory / storage devices 820, or in any suitable combination thereof. Furthermore, any portion of the instructions 850 may be transferred to the hardware resources 800 from any combination of the peripherals 804 or the databases 806. Accordingly, the memory of the processors 810, the memory / storage devices 820, the peripherals 804, and the databases 806 are examples of computer-readable and machine-readable media.

[0097] Fig.9 is a block diagram of an example processor platform in accordance with some embodiments of the disclosure. Processor platform 900 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cellular phone, a smartphone, a tablet such as an iPad™), a personal digital assistant (PDA), an internet device, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset, or other wearable device, or any other type of computing device.

[0098] The processor platform 900 of the illustrated example includes a processor 912. The processor 912 of the illustrated example is hardware. The processor 912 may be implemented, for example, by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers of any desired family or manufacturer. The hardware processor may be a semiconductor-based (e.g., silicon-based) device. In some embodiments, the processor implements one or more of the previously described methods or processes.

[0099] The processor 912 of the illustrated example includes a local memory 913 (e.g., a cache). The processor 912 of the illustrated example is in communication with a main memory, which includes a volatile memory 914 and a non-volatile memory 916, via a bus 918. The volatile memory 914 may be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The non-volatile memory 916 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 914, 916 is controlled by a memory controller.

[0100] The processor platform 900 of the illustrated example also includes interface circuitry 920. The interface circuitry 920 may be implemented by any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth® interface, a Near Field Communication (NFC) interface, and / or a PCI Express interface.

[0101] In the illustrated example, one or more input devices 922 are connected to the interface circuitry 920. The input device(s) 922 enable a user to input data and / or commands into the processor 912. The input device(s) may be implemented, for example, by an audio sensor, a microphone, a camera (still or video), a keyboard, a button, a mouse, a touchscreen, a trackpad, a trackball, and / or a speech recognition system.

[0102] Also connected to the interface circuitry 920 of the illustrated example is one or more output devices 924. The output devices 924 may be implemented, for example, by display devices (e.g., a light-emitting diode (LED), an organic light-emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube (CRT) display, an IPS (In-Place Switching) display, a touchscreen, etc.), a tactile output device, a printer, and / or speakers. The interface circuitry 920 of the illustrated example thus typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.

[0103] The interface circuitry 920 of the illustrated example also includes a communication device, such as a transmitter, a receiver, a transceiver, a modem, a home gateway, a wireless access point, and / or a network interface, to enable the exchange of data with external machines (e.g., computing devices of any type) over a network 926. Communication may occur, for example, via an Ethernet connection, a DSL (Digital Subscriber Line) connection, a telephone line connection, a coaxial cable system, a satellite system, a wireless line-of-site system, a cellular telephone system, etc.

[0104] For example, the interface circuitry 920 may include a training data set input through the input device(s) 922 or retrieved from the network 926.

[0105] The processor platform 900 of the illustrated example also includes one or more mass storage devices 928 for storing software and / or data. Examples of such mass storage devices 928 include floppy disk drives, hard disks, CD drives, Blu-ray disk drives, RAID (Redundant Array of Independent Disks) systems, and DVD (Digital Versatile Disk) drives.

[0106] Machine-executable instructions 932 may be stored in the mass storage device 928, in the volatile memory 914, in the non-volatile memory 916, and / or on a removable non-volatile computer-readable storage medium, such as a CD or DVD.

[0107] The following sections describe examples of different embodiments.

[0108] Example 1 includes a method for optimizing a deep learning computation graph, comprising: obtaining a deep learning computation graph comprising computationally intensive operators and memory-intensive operators; merging the memory-intensive operators into the computationally intensive operators to generate a new computation graph; splitting the new computation graph; and merging computationally intensive operators in each of the sub-computation graphs to generate an optimized computation graph.

[0109] Example 2 includes the method of Example 1, wherein merging the memory-intensive operators into the compute-intensive operators comprises: merging one or more sequential memory-intensive operators into a previous or a following compute-intensive operator.

[0110] Example 3 includes the method of Example 1 or 2, wherein the new computation graph comprises a plurality of layers and each of the layers comprises a plurality of computationally intensive operators, and wherein the new computation graph is divided into the sub-computation graphs based on an output property of each layer of the new computation graph and a platform capacity of a platform on which the new computation graph is to be executed.

[0111] Example 4 includes the method of any one of Examples 1-3, wherein splitting the new computation graph into the sub-computation graphs comprises: obtaining a splitting parameter using a heuristic rule based on the output property of each layer; obtaining a buffer size for each layer to be allocated for an output batch and a weight for the layer based on the output property of the layer; and sequentially splitting the layers, in a topology order of the new computation graph, into the sub-computation graphs based on the splitting parameter, the buffer size, and the platform capacity.

[0112] Example 5 includes the method of any of Examples 1-4, wherein sequentially splitting the layers in the topology order of the new computation graph into the sub-computation graphs based on the splitting parameter, the buffer size, and the platform capacity comprises: obtaining a reduced buffer size for each layer based on the splitting parameter and the buffer size for the layer; and splitting one or more sequential layers for which a sum of the reduced buffer sizes for the one or more sequential layers is less than or equal to the platform capacity and a sum of the reduced buffer sizes for the one or more sequential layers and a subsequent layer is greater than the platform capacity, into a sub-computation graph.

[0113] Example 6 includes the method of any of Examples 1-5, wherein each of the computationally intensive operators of the new computational graph outputs an output activation and the output activations of the computationally intensive operators of each layer form an output stack, and wherein the output property comprises a size of an output stack for each layer and / or a size of an output activation of each computationally intensive operator of the new computational graph.

[0114] Example 7 includes the method of any of Examples 1-6, wherein the splitting parameter comprises a stack splitting number and a spatial splitting number, wherein the stack splitting number corresponds to a number of sub-stacks into which an output stack is to be split, and the spatial splitting number corresponds to a number of sub-activations into which an output activation is to be split.

[0115] Example 8 includes the method of any of Examples 1-7, where the reduced buffer size is expressed by: aR_i=wi+ai×1x×y, where a R_i indicates the reduced buffer size for an i-th layer of the new computation graph in the topology order, W i the weighting for the i-th layer, a i indicates the buffer size for the i-th layer, x indicates the stack split number, y indicates the spatial split number, and each of a R_ i, W i , a i , x and y is greater than 0.

[0116] Example 9 includes the method of any of Examples 1-8, wherein the platform capacity comprises a data cache unit (DCU) size and a mid-level cell (MLC) size of a central processing unit (CPU) on which the new computation graph is to be executed.

[0117] Example 10 includes the method of any of Examples 1-9, wherein for each sub-computation graph: ∑i=Ni=MaR_i≤T×(L1+L2)<∑i=Ni=M+1aR_i, where N indicates a starting layer of the sub-computation graph, M indicates a last layer of the sub-computation graph, T indicates a threshold set based on the CPU, L1 indicates the size of the DCU, L2 indicates the size of the MLC, N, M, L1 and L2 are greater than 0 and T is greater than 0 and less than or equal to 1.

[0118] Example 11 includes the method of any of Examples 1-10, wherein merging the computationally intensive operators in each of the subcomputational graphs comprises: for each subcomputational graph, dividing the output stack for each layer of the subcomputational graph into the substacks by the stack split number; dividing the output activation of each computationally intensive operator of the subcomputational graph into the subactivations by the spatial split number; and merging the computationally intensive operators in the subcomputational graph based on the substacks and the subactivations.

[0119] Example 12 includes the method of any of Examples 1-11, wherein each output activation comprises one or more output samples, and wherein dividing the output activation of each computationally intensive operator of the sub-computation graph into the sub-activations by the spatial division number comprises: dividing each output sample of the output activation into sub-samples along a height direction of the sample by the spatial division number.

[0120] Example 13 includes an apparatus for optimizing a deep learning computation graph, comprising: interface circuitry; and processor circuitry coupled to the interface circuitry and configured to: obtain a deep learning computation graph comprising computationally intensive operators and memory-intensive operators; merging the memory-intensive operators into the computationally intensive operators to generate a new computation graph; splitting the new computation graph; and merging computationally intensive operators in each of the partial computation graphs to generate an optimized computation graph.

[0121] Example 14 includes the apparatus of Example 13, wherein the processor circuitry is configured to merge the memory-intensive operators into the compute-intensive operators by: merging one or more sequential memory-intensive operators into a previous or a following compute-intensive operator.

[0122] Example 15 includes the apparatus of example 13 or 14, wherein the new computation graph comprises a plurality of layers, each of the layers comprising a plurality of computationally intensive operators, and wherein the new computation graph is divided into the sub-computational graphs based on an output property of each layer of the new computational graph and a platform capacity of a platform on which the new computational graph is to be executed.

[0123] Example 16 includes the apparatus of any one of Examples 13-15, wherein the processor circuitry is configured to divide the new computation graph into the sub-computation graphs by: obtaining a division parameter using a heuristic rule based on the output characteristic of each layer; obtaining a buffer size for each layer to be allocated for an output stack and a weight for the layer based on the output characteristic of the layer; and sequentially dividing the layers, in a topology order of the new computation graph, into the sub-computation graphs based on the division parameter, the buffer size, and the platform capacity.

[0124] Example 17 includes the apparatus of any of Examples 13-16, wherein the processor circuitry is configured to sequentially split the layers in the topology order of the new computation graph based on the split parameter, the buffer size, and the platform capacity into the split graphs by: obtaining a reduced buffer size for each layer based on the split parameter and the buffer size for the layer; and splitting one or more sequential layers for which a sum of the reduced buffer sizes for the one or more sequential layers is less than or equal to the platform capacity and a sum of the reduced buffer sizes for the one or more sequential layers and a subsequent layer is greater than the platform capacity, into a partial computation graph.

[0125] Example 18 includes the apparatus of any of Examples 13-17, wherein each of the computationally intensive operators of the new computational graph outputs an output activation and the output activations of the computationally intensive operators of each layer form an output stack, and wherein the output property comprises a size of an output stack for each layer and / or a size of an output activation of each computationally intensive operator of the new computational graph.

[0126] Example 19 includes the apparatus of any of Examples 13-18, wherein the split parameter comprises a stack split number and a spatial split number, wherein the stack split number corresponds to a number of sub-stacks into which an output stack is to be split, and the spatial split number corresponds to a number of sub-activations into which an output activation is to be split.

[0127] Example 20 includes the facility of any of Examples 13-19, wherein the reduced buffer size is expressed by: aR_i=wi+ai×1x×y, where a R_i indicates the reduced buffer size for an i-th layer of the new computation graph in the topology order, W i the weighting for the i-th layer, a i indicates the buffer size for the i-th layer, x indicates the stack split number, y indicates the spatial split number, and each of a R_i , i, W i , a i , x and y is greater than 0.

[0128] Example 21 includes the apparatus of any of Examples 13-20, wherein the platform capacity comprises a data cache unit (DCU) size and a mid-level cell (MLC) size of a central processing unit (CPU) on which the new computation graph is to be executed.

[0129] Example 22 includes the facility of any of Examples 13-21, wherein for each sub-computation graph: ∑i=Ni=MaR_i≤T×(L1+L2)<∑i=Ni=M+1aR_i, where N indicates a starting layer of the sub-computation graph, M indicates a last layer of the sub-computation graph, T indicates a threshold set based on the CPU, L1 indicates the size of the DCU, L2 indicates the size of the MLC, N, M, L1 and L2 are greater than 0 and T is greater than 0 and less than or equal to 1.

[0130] Example 23 includes the apparatus of any of Examples 13-22, wherein the processor circuitry is configured to merge the computationally intensive operators in each of the subcomputational graphs by: for each subcomputational graph, dividing the output stack for each layer of the subcomputational graph into the substacks by the stack split number; dividing the output activation of each computationally intensive operator of the subcomputational graph into the subactivations by the spatial split number; and merging the computationally intensive operators in the subcomputational graph based on the substacks and the subactivations.

[0131] Example 24 includes the apparatus of any of Examples 13-23, wherein each output activation comprises one or more output samples, and wherein the processor circuitry is configured to divide the output activation of each computationally intensive operator of the sub-computation graph into the sub-activations by the spatial division number by: dividing each output sample of the output activation into sub-samples along a height direction of the sample by the spatial division number.

[0132] Example 25 includes an apparatus for optimizing a deep learning computation graph, comprising: means for obtaining a deep learning computation graph comprising computationally intensive operators and memory-intensive operators; means for merging the memory-intensive operators into the computationally intensive operators to generate a new computation graph; means for splitting the new computation graph; and means for merging computationally intensive operators in each of the partial computation graphs to generate an optimized computation graph.

[0133] Example 26 includes the apparatus of Example 25, wherein the means for merging the memory-intensive operators into the computation-intensive operators comprises: means for merging one or more sequential memory-intensive operators into a previous or a following computation-intensive operator.

[0134] Example 27 includes the apparatus of example 25 or 26, wherein the new computation graph comprises a plurality of layers and each of the layers comprises a plurality of computationally intensive operators, and wherein the new computation graph is divided into the sub-computational graphs based on an output property of each layer of the new computational graph and a platform capacity of a platform on which the new computational graph is to be executed.

[0135] Example 28 includes the apparatus of any one of Examples 25-27, wherein the means for dividing the new computation graph into the sub-computation graphs comprises: means for obtaining a division parameter using a heuristic rule based on the output property of each layer; means for obtaining a buffer size for each layer to be allocated for an output batch and a weight for the layer based on the output property of the layer; and means for sequentially dividing the layers, in a topology order of the new computation graph, into the sub-computation graphs based on the division parameter, the buffer size, and the platform capacity.

[0136] Example 29 includes the apparatus of any of Examples 25-28, wherein the means for sequentially splitting the layers in the topology order of the new computation graph into the sub-computation graphs based on the splitting parameter, the buffer size, and the platform capacity comprises: means for obtaining a reduced buffer size for each layer based on the splitting parameter and the buffer size for the layer; and means for splitting one or more sequential layers for which a sum of the reduced buffer sizes for the one or more sequential layers is less than or equal to the platform capacity and a sum of the reduced buffer sizes for the one or more sequential layers and a subsequent layer is greater than the platform capacity, into a sub-computation graph.

[0137] Example 30 includes the apparatus of any of Examples 25-29, wherein each of the computationally intensive operators of the new computational graph outputs an output activation, and the output activations of the computationally intensive operators of each layer form an output stack, and wherein the output property comprises a size of an output stack for each layer and / or a size of an output activation of each computationally intensive operator of the new computational graph.

[0138] Example 31 includes the apparatus of any of Examples 25-30, wherein the split parameter comprises a stack split number and a spatial split number, wherein the stack split number corresponds to a number of sub-stacks into which an output stack is to be split, and the spatial split number corresponds to a number of sub-activations into which an output activation is to be split.

[0139] Example 32 includes the facility of any of Examples 25-31, where the reduced buffer size is expressed by: aR_i=wi+ai×1x×y, where a R_i indicates the reduced buffer size for an i-th layer of the new computation graph in the topology order, W i the weighting for the i-th layer, a i indicates the buffer size for the i-th layer, x indicates the stack split number, y indicates the spatial split number, and each of a R_i , i, W i , a i , x and y is greater than 0.

[0140] Example 33 includes the apparatus of any of Examples 25-32, wherein the platform capacity comprises a data cache unit (DCU) size and a mid-level cell (MLC) size of a central processing unit (CPU) on which the new computation graph is to be executed.

[0141] Example 34 includes the facility of any of Examples 25-33, where for each sub-computation graph: ∑i=Ni=MaR_i≤T×(L1+L2)<∑i=Ni=M+1aR_i, where N indicates a starting layer of the sub-computation graph, M indicates a last layer of the sub-computation graph, T indicates a threshold set based on the CPU, L1 indicates the size of the DCU, L2 indicates the size of the MLC, N, M, L1 and L2 are greater than 0 and T is greater than 0 and less than or equal to 1.

[0142] Example 35 includes the apparatus of any of Examples 25-34, wherein the means for merging the computationally intensive operators in each of the subcomputational graphs comprises: for each subcomputational graph, means for dividing the output stack for each layer of the subcomputational graph into the substacks by the stack split number; means for dividing the output activation of each computationally intensive operator of the subcomputational graph into the subactivations by the spatial split number; and means for merging the computationally intensive operators in the subcomputational graph based on the substacks and the subactivations.

[0143] Example 36 includes the apparatus of any of Examples 25-35, wherein each output activation comprises one or more output samples, and wherein the means for dividing the output activation of each computationally intensive operator of the sub-computation graph into the sub-activations by the spatial division number comprises: means for dividing each output sample of the output activation into sub-samples along a height direction of the sample by the spatial division number.

[0144] Although specific embodiments have been shown and described herein for descriptive purposes, a wide variety of alternative and / or equivalent embodiments or calculated implementations may be implemented in place of the embodiments shown and described without departing from the scope of the present disclosure. This application is intended to cover any adaptations or variations of the embodiments discussed herein. Therefore, it is expressly intended that the embodiments described herein be limited only by the appended claims and the equivalents thereof.

Claims

[1] A method for optimizing a deep learning computation graph, comprising: Obtaining a deep learning computation graph that includes computation-intensive operators and memory-intensive operators; Merging the memory-intensive operators into the computation-intensive operators to generate a new computation graph; Dividing the new computation graph into sub-computation graphs; and Merging computationally intensive operators in each of the subcomputation graphs to produce an optimized computation graph. [2] The method of claim 1, wherein merging the memory-intensive operators into the computation-intensive operators comprises: Merging one or more sequential memory-intensive operators into a previous or following computation-intensive operator. [3] The method of claim 1, wherein the new computational graph comprises multiple layers and each of the layers comprises multiple computationally intensive operators, and wherein the new computation graph is divided into the sub-computation graphs based on an output property of each layer of the new computation graph and a platform capacity of a platform on which the new computation graph is to be executed. [4] The method of claim 3, wherein dividing the new computation graph into the sub-computation graphs comprises: Obtaining a division parameter using a heuristic rule based on the output property of each layer; Obtaining a buffer size for each layer to be allocated for an output batch and a weight for the layer based on the output property of the layer; and sequentially splitting the layers in a topology order of the new computation graph into the sub-computation graphs based on the splitting parameter, buffer size, and platform capacity. [5] The method of claim 4, wherein sequentially dividing the layers in the topology order of the new computation graph into the partial computation graphs based on the dividing parameter, the buffer size, and the platform capacity comprises: Obtaining a reduced buffer size for each layer based on the splitting parameter and the buffer size for the layer; and Dividing one or more sequential layers for which a sum of the reduced buffer sizes for the one or more sequential layers is less than or equal to the platform capacity and a sum of the reduced buffer sizes for the one or more sequential layers and a subsequent layer is greater than the platform capacity into a partial computation graph. [6] The method of claim 5, wherein each of the computationally intensive operators of the new computational graph outputs an output activation, and the output activations of the computationally intensive operators of each layer form an output stack, and wherein the output property comprises a size of an output stack for each layer and / or a size of an output activation of each computationally intensive operator of the new computational graph. [7] The method of claim 6, wherein the pitch parameter comprises a stack pitch number and a spatial pitch number, where the stack division number corresponds to a number of sub-stacks into which an output stack is to be divided, and the spatial division number corresponds to a number of partial activations into which an output activation is to be divided. [8] The method of claim 7, wherein the reduced buffer size is expressed by: aR_i=wi+ai×1x×y, where a R_i indicates the reduced buffer size for an i-th layer of the new computation graph in the topology order, w i indicates the weight for the i-th layer, a i indicates the buffer size for the i-th layer, x indicates the stack split number, y indicates the spatial split number and each of a R_i , i, w i , a i , x and y is greater than 0. [9] The method of claim 8, wherein the platform capacity comprises a size of a data cache unit (DCU) and a size of a mid-level cell (MLC) of a central processing unit (CPU) on which the new computation graph is to be executed. [10] Method according to claim 9, wherein for each partial computation graph: ∑i=Ni=MaR_i≤T×(L1+L2)<∑i=Ni=M+1aR_i, where N indicates a starting layer of the sub-computation graph, M indicates a last layer of the sub-computation graph, T indicates a threshold set based on the CPU, L1 indicates the size of the DCU, L2 indicates the size of the MLC, N, M, L1 and L2 are greater than 0, and T is greater than 0 and less than or equal to 1. [11] The method of any of claims 7-10, wherein merging the computationally intensive operators in each of the subcomputation graphs comprises: for each subcomputation graph, Dividing the output stack for each layer of the subcomputation graph into substacks by the stack division number; Dividing the output activation of each computationally intensive operator of the partial computation graph into the partial activations by the spatial division number; and Merging the computationally intensive operators in the partial computation graph based on the partial stacks and partial activations. [12] The method of claim 11, wherein each output activation comprises one or more output samples, and wherein dividing the output activation of each computationally intensive operator of the partial computation graph into the partial activations by the spatial division number comprises: Dividing each output sample of the output activation along a height direction of the sample by the spatial division number into sub-samples. [13] Apparatus for optimizing a deep learning computation graph, comprising: an interface circuit; and a processor circuit coupled to the interface circuit and configured to: Obtaining a deep learning computation graph that includes computation-intensive operators and memory-intensive operators; Merging the memory-intensive operators into the computation-intensive operators to generate a new computation graph; Dividing the new computation graph into sub-computation graphs; and Merging computationally intensive operators in each of the subcomputation graphs to produce an optimized computation graph. [14] The device of claim 13, wherein the processor circuitry is configured to merge the memory-intensive operators into the computation-intensive operators by: Merging one or more sequential memory-intensive operators into a previous or following computation-intensive operator. [15] The apparatus of claim 13, wherein the new computational graph comprises multiple layers, each of the layers comprising multiple computationally intensive operators, and wherein the new computation graph is divided into the sub-computation graphs based on an output property of each layer of the new computation graph and a platform capacity of a platform on which the new computation graph is to be executed. [16] Device according to claim 15, wherein the processor circuitry is designed to divide the new calculation graph into the partial calculation graphs by: Obtaining a division parameter using a heuristic rule based on the output property of each layer; Obtaining a buffer size for each layer to be allocated for an output batch and a weight for the layer based on the output property of the layer; and sequentially splitting the layers in a topology order of the new computation graph into the sub-computation graphs based on the splitting parameter, buffer size, and platform capacity. [17] The device of claim 16, wherein the processor circuitry is configured to sequentially divide the layers in the topology order of the new computation graph based on the division parameter, the buffer size, and the platform capacity by: Obtaining a reduced buffer size for each layer based on the splitting parameter and the buffer size for the layer; and Dividing one or more sequential layers for which a sum of the reduced buffer sizes for the one or more sequential layers is less than or equal to the platform capacity and a sum of the reduced buffer sizes for the one or more sequential layers and a subsequent layer is greater than the platform capacity into a partial computation graph. [18] The apparatus of claim 17, wherein each of the computationally intensive operators of the new computation graph outputs an output activation, and the output activations of the computationally intensive operators of each layer form an output stack, and wherein the output property comprises a size of an output stack for each layer and / or a size of an output activation of each computationally intensive operator of the new computational graph. [19] The device of claim 18, wherein the pitch parameter comprises a stack pitch number and a spatial pitch number, where the stack division number corresponds to a number of sub-stacks into which an output stack is to be divided, and the spatial division number corresponds to a number of partial activations into which an output activation is to be divided. [20] The device of claim 19, wherein the reduced buffer size is expressed by: aR_i=wi+ai×1x×y, where a R_i indicates the reduced buffer size for an i-th layer of the new computation graph in the topology order, w i indicates the weight for the i-th layer, a i indicates the buffer size for the i-th layer, x is the stack split number indicates, y indicates the spatial division number and each of a R_i , i, w i , a i , x and y is greater than 0. [21] The device of claim 20, wherein the platform capacity comprises a size of a data cache unit (DCU) and a size of a mid-level cell (MLC) of a central processing unit (CPU) on which the new computation graph is to be executed. [22] Device according to claim 21, wherein for each partial computation graph: ∑i=Ni=MaR_i×(L1+L2)<∑i=Ni=M+1aR_i, where N indicates a starting layer of the sub-computation graph, M indicates a last layer of the sub-computation graph, T indicates a threshold set based on the CPU, L1 indicates the size of the DCU, L2 indicates the size of the MLC, N, M, L1 and L2 are greater than 0, and T is greater than 0 and less than or equal to 1. [23] Device according to one of claims 19-22, wherein the processor circuitry is arranged to merge the computationally intensive operators in each of the sub-computation graphs by: for each sub-computation graph, Dividing the output stack for each layer of the subcomputation graph into substacks by the stack division number; Dividing the output activation of each computationally intensive operator of the partial computation graph into the partial activations by the spatial division number; and Merging the computationally intensive operators in the partial computation graph based on the partial stacks and partial activations. [24] A computer-readable medium having instructions stored thereon, the instructions, when executed by a processor, causing the processor to: Obtaining a deep learning computation graph that includes computation-intensive operators and memory-intensive operators; Merging the memory-intensive operators into the computation-intensive operators to generate a new computation graph; Dividing the new computation graph into sub-computation graphs; and Merging computationally intensive operators in each of the subcomputation graphs to produce an optimized computation graph. [25] The computer-readable medium of claim 24, wherein the new computational graph comprises a plurality of layers, each of the layers comprising a plurality of computationally intensive operators, and wherein the new computation graph is divided into the sub-computation graphs based on an output property of each layer of the new computation graph and a platform capacity of a platform on which the new computation graph is to be executed.