Compilation Method, Device and Related Products of Deep Learning Algorithm

By receiving and judging the operation data of the deep learning programming library and using NCLAPI for flexible compilation, the problem of performance optimization of deep learning algorithms on different hardware platforms is solved, and more efficient compilation and processing performance is achieved.

CN112183712BActive Publication Date: 2025-07-22ANHUI CAMBRICON INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201910596132.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-07-03
Publication Date
2025-07-22
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

Performance optimization of deep learning algorithms on different hardware platforms is difficult to meet the problems of variability and high coupling, and the existing technology is difficult to improve flexibility and efficiency through early compilation optimization.

Method used

By receiving operation data transmitted from the deep learning programming library interface, judging the operation instruction type and executing corresponding compilation operations, generating binary code for deep learning algorithms, and using the neural computing programming library interface (NCLAPI) for flexible compilation, supporting operation fusion, customization, specialization and offline optimization.

Benefits of technology

It improves the flexibility and efficiency of the compilation process, improves the performance optimization effect of deep learning algorithms on the hardware platform, and enhances the processing performance of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112183712B_ABST
    Figure CN112183712B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a compilation method, apparatus and related products for deep learning algorithms. The product includes a controller unit, and the controller unit includes: an instruction cache unit, an instruction processing unit and a storage queue unit; the instruction cache unit is used to store calculation instructions associated with the artificial neural network operation; the instruction processing unit is used to parse the calculation instructions to obtain a plurality of operation instructions; the storage queue unit is used to store an instruction queue, and the instruction queue includes: a plurality of operation instructions or calculation instructions to be executed in the front-to-back order of the queue. Through the above method, the present disclosure can improve the operation efficiency of related products when performing operations on neural network models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning, and in particular, to a method, apparatus, and related products for compiling a deep learning algorithm. Background Art

[0002] In the field of artificial intelligence technology, neural network algorithms are a very popular machine learning algorithm recently, and have achieved very good results in various fields, such as image recognition, speech recognition, natural language processing, etc. With the development of neural network algorithms, the complexity of the algorithms is also getting higher and higher. In order to improve the recognition rate, the scale of the model is also gradually increasing. Summary of the Invention

[0003] In view of this, the present disclosure provides a method, apparatus, and related products for compiling a deep learning algorithm, which can improve the performance optimization effect of the deep learning algorithm for the corresponding hardware platform.

[0004] According to a first aspect of the present disclosure, there is provided a method for compiling a deep learning algorithm, the method comprising: receiving operation data transmitted by a deep learning programming library interface; obtaining an operation instruction included in the operation data; determining an instruction type of the operation instruction, and performing a compilation operation corresponding to the instruction type according to a determination result to obtain a binary code of the deep learning algorithm.

[0005] According to a second aspect of the present disclosure, there is provided a device for compiling a deep learning algorithm, comprising: an operation data receiving module, configured to receive operation data transmitted by a deep learning programming library interface; an operation instruction obtaining module, configured to obtain an operation instruction included in the operation data; and a compilation module, configured to determine an instruction type of the operation instruction, and perform a compilation operation corresponding to the instruction type according to a determination result to obtain a binary code of the deep learning algorithm.

[0006] According to a third aspect of the present disclosure, there is provided a deep learning operation device, the deep learning operation device comprising the device for compiling a deep learning algorithm as described in the second aspect above, and the deep learning operation device is configured to complete a set deep learning operation.

[0007] According to a fourth aspect of the present disclosure, there is provided a combined operation device, the combined operation device comprising the deep learning operation device as described in the third aspect above, a general-purpose interconnection interface, and other processing devices; the deep learning operation device interacts with the other processing devices to jointly complete a calculation operation specified by a user.

[0008] According to a fifth aspect of the present disclosure, a deep learning chip is provided. The deep learning chip includes: a compilation device for a deep learning algorithm as described in the second aspect above; or, a deep learning operation device as described in the third aspect above; or, a combined operation device as described in the fourth aspect above.

[0009] According to a sixth aspect of the present disclosure, an electronic device is provided. The electronic device includes: a compilation device for a deep learning algorithm as described in the second aspect above; or, a deep learning operation device as described in the third aspect above; or, a combined operation device as described in the fourth aspect above; or, a deep learning chip as described in the fifth aspect above.

[0010] By receiving operation data transmitted through a deep learning programming library interface, and according to the instruction type of the operation instruction in the operation data, a compilation operation corresponding to the instruction type is performed to obtain a binary code of the deep learning algorithm. The compilation method, device, and related products of the deep learning algorithm according to various embodiments of the present disclosure can enable the compilation process to adaptively change according to different types of operation instructions, thereby greatly improving the flexibility and efficiency of compilation, effectively improving the performance optimization effect of the deep learning algorithm for the corresponding hardware platform, and then improving the processing performance of the deep learning processor.

[0011] Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings included in and constituting a part of this specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0013] Figure 1 A flowchart showing a compilation method of a deep learning algorithm according to an embodiment of the present disclosure.

[0014] Figure 2 A schematic diagram showing the overall architecture of a neural calculus programming library interface according to an embodiment of the present disclosure.

[0015] Figure 3 A diagram showing the correspondence between the attributes, classifications, and meanings of tensor data according to an embodiment of the present disclosure.

[0016] Figure 4 A diagram showing the correspondence between tensor data and shape symbols according to an embodiment of the present disclosure.

[0017] Figure 5 A diagram showing the build-in operation instructions supported by NCLAPI according to an embodiment of the present disclosure.

[0018] Figure 6 Schematic diagram showing the implementation of automatic optimization of accuracy according to an embodiment of the present disclosure.

[0019] Figure 7 Schematic diagram showing a calculation model according to an embodiment of the present disclosure.

[0020] Figure 8 Schematic diagram showing the generation principle of customized operation instructions according to an embodiment of the present disclosure.

[0021] Figure 9 Schematic diagram showing the operation fusion result according to an embodiment of the present disclosure.

[0022] Figure 10 Schematic diagram showing the operation fusion result according to an embodiment of the present disclosure.

[0023] Figure 11 Schematic diagram showing the relevant programming interfaces of operation fusion according to an embodiment of the present disclosure.

[0024] Figure 12 Schematic diagram showing the process of creating a fusion operation according to an embodiment of the present disclosure.

[0025] Figure 13 Schematic diagram showing the data flow of a three-layer calculation model according to an embodiment of the present disclosure.

[0026] Figure 14 Schematic block diagram showing the implementation of a hybrid programming model according to an embodiment of the present disclosure.

[0027] Figure 15 Schematic diagram showing the difference between the offline mode and the online mode according to an embodiment of the present disclosure.

[0028] Figure 16 Schematic diagram showing the offline interface according to an embodiment of the present disclosure.

[0029] Figure 17 Schematic diagram showing the architecture of TensorFlow according to an embodiment of the present disclosure.

[0030] Figure 18 Schematic diagram showing the comparison between the NCLAPI and the interfaces of mainstream deep learning programming libraries according to an embodiment of the present disclosure.

[0031] Figure 19 Schematic diagram showing the overall architecture of NCLA according to an embodiment of the present disclosure.

[0032] Figure 20 Schematic diagram showing the implementation of a static operation pool according to an embodiment of the present disclosure.

[0033] Figure 21Shows a schematic architecture diagram of CDUCA according to an embodiment of the present disclosure.

[0034] Figure 22 Shows the form of the original computational graph according to an embodiment of the present disclosure.

[0035] Figure 23 Shows the flowchart of the operation of the computational graph engine according to an embodiment of the present disclosure.

[0036] Figure 24 Shows a schematic diagram of the structure including sub-structures within an image classification network according to an embodiment of the present disclosure.

[0037] Figure 25 Shows the flowchart of the compilation method of the deep learning algorithm according to an embodiment of the present disclosure.

[0038] Figure 26 Shows a schematic diagram of instruction pipelining according to an embodiment of the present disclosure.

[0039] Figure 27 Shows an implementation diagram of optimizing model data according to an embodiment of the present disclosure.

[0040] Figure 28 Shows a schematic diagram of optimization after data splitting according to an embodiment of the present disclosure.

[0041] Figure 29 Shows a schematic diagram of the modules and functions of the runtime system according to an embodiment of the present disclosure.

[0042] Figure 30 Shows a block diagram of the compilation device for the deep learning algorithm according to an embodiment of the present disclosure.

[0043] Figure 31 Shows a block diagram of the combined processing device according to an embodiment of the present disclosure. Detailed Description of the Embodiment

[0044] The following will describe various exemplary embodiments, features, and aspects of the present disclosure in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0045] The special term "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not necessarily have to be construed as superior to or better than other embodiments.

[0046] In addition, to better illustrate the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0047] To alleviate the deteriorating memory wall problem, deep learning processors usually design on-chip memory near the computing units. The latency of accessing on-chip memory is much lower than that of accessing off-chip memory. Therefore, reasonably utilizing on-chip memory is the key to unleashing the performance of deep learning processors. However, on-chip memory does not have the functions of cache such as data prefetching, data replacement, and handling data conflicts. These tasks must be completed by programmers writing instructions. On the other hand, the capacity of on-chip memory is very limited. Programmers must split both the operations and the data simultaneously, which results in a tight coupling between the operations and the data. Taking a certain deep learning processor as an example, the capacity of its on-chip neuron storage (NBin) is only 2KB, while a 1080P RGB image represented in half precision requires 12MB of space. Therefore, only the input data needs to be split into 6K times for loading. The data splitting causes the operations to be split accordingly. For example, the number of nested loops will increase, and the number of iterations of the inner loop will decrease.

[0048] The characteristics of on-chip memory (display management, limited capacity) and the characteristics of deep learning algorithms (complex and variable processing layers, diverse operation data types, processing high-dimensional tensors, etc.) jointly lead to the program optimization in the deep learning field being extremely sensitive to algorithm and hardware changes. Taking the 2D convolution operation as an example, it contains at least 9 shape parameters (N, CI, HI, WI, CO, Kh, Kw, Sh, Sw), 7 nested loops, and multiple operation precisions (half, fxm.b, Qts.b, intx, etc.). A change in the combination of the above parameters or the capacity of on-chip memory will affect the optimal splitting strategy of data and operations, resulting in different generated binary codes.

[0049] From the above reasons, due to the particularity of deep learning algorithms and the particularity of deep learning processor architectures, there are two important characteristics in the program optimization in the deep learning field. One is being extremely sensitive to algorithm and hardware changes, and the other is a very high coupling degree between operations and data. Since the program optimization in the deep learning field is extremely sensitive to algorithm and hardware changes, it is difficult for us to meet the performance requirements under different algorithms and different hardware platforms through ahead-of-time compilation optimization (AOT). Therefore, how to propose a compilation method for deep learning algorithms based on the particularity of deep learning algorithms and the particularity of deep learning processor architectures has become an urgent problem to be solved.

[0050] Figure 1 A flowchart showing a method for compiling a deep learning algorithm according to an embodiment of the present disclosure. As shown in the figure, the method may include:

[0051] Step S11: Receive operation data passed by a deep learning programming library interface.

[0052] Step S12: Obtain the operation instructions included in the operation data.

[0053] Step S13: Determine the instruction type of the operation instruction, and perform a compilation operation corresponding to the instruction type according to the determination result to obtain the binary code of the deep learning algorithm.

[0054] In the above-mentioned disclosed embodiment, the binary code is a hardware instruction for guiding a hardware device to execute a deep learning algorithm. Specifically, which hardware devices are guided and the specific content of the hardware instruction are not limited in the embodiments of the present disclosure and can be flexibly selected according to actual situations.

[0055] By receiving the operation data passed by the deep learning programming library interface, performing a compilation operation corresponding to the instruction type according to the instruction type of the operation instruction in the operation data to obtain the binary code of the deep learning algorithm, the method, device and related products for compiling the deep learning algorithm according to various embodiments of the present disclosure can make the compilation process adaptively change according to the type of the operation instruction, thereby greatly improving the flexibility and efficiency of compilation, effectively improving the performance optimization effect of the deep learning algorithm for the corresponding hardware platform, and then improving the processing performance of the deep learning processor.

[0056] The specific implementation source of the operation data passed by the deep learning programming library interface received in step S11 is not limited. In one possible implementation, the operation data can be created or called according to a user instruction received by the deep learning programming library interface.

[0057] By creating or calling operation data according to a user instruction received by the deep learning programming library interface, and the operation data can be used to obtain the binary code of the deep learning algorithm subsequently. Through the above process, a general programming interface can be provided for the user to realize the effective conversion between the user instruction and the machine instruction.

[0058] In the above steps, the implementation manner of the deep learning programming library interface is not limited and can be flexibly selected according to actual situations. In one possible implementation, the deep learning programming library interface can be a neurological calculus library interface (NCLAPI, neurological calculus Library API), and the specific implementation manner of this interface can be determined according to actual situations and is not limited to the following disclosed embodiments. Figure 2Shows a schematic diagram of the overall architecture of a neural calculus programming library interface according to an embodiment of the present disclosure. As shown in the figure, in one example, the implementation of the NCLAPI interface can be as follows: By simulating neural calculus, the NCLAPI interface is enabled to have good deep learning modeling capabilities; By designing reshapable operations and corresponding operation rules, various performance optimizations are flexibly supported; By designing a hybrid programming model, the flexibility of the programming model is improved; The design of data structures and interfaces is simplified, hiding hardware details inside the data structures and interfaces.

[0059] In the above-mentioned disclosed embodiment, neural calculus is a functional deep learning modeling method. In one example, neural calculus can represent the input data, output data, and model parameters of the input layer with tensors, and represent the deep learning processing layer with functions. The functions can be combined according to certain rules to construct various deep learning calculation models. Since functions themselves have combinability and reusability, neural calculus can well express the combinability and reusability of deep learning algorithms. The neural calculus designed according to the above idea has powerful deep learning modeling capabilities. Currently known deep learning frameworks such as Tensor and Mxnet use directed graphs to model deep learning calculation models. Through experiments, it can be proven that any directed graph can be mapped into a function combination, and any function combination can be mapped into a directed acyclic graph. Therefore, neural calculus has the same deep learning modeling capabilities as directed graphs. Since in one example, the NCLAPI interface can be enabled to have good deep learning modeling capabilities by simulating neural calculus, therefore, under an appropriate simulation method, NCLAPI can have the same learning modeling capabilities as directed graphs.

[0060] By Figure 2 It can be seen that in one example, there can be two data structures in NCLAPI, namely tensor (nclTensor) and malleable operation (nclOperator). nclTensor is used to describe the input data, output data, and model parameters of the deep learning processing layer, and nclOperator is used to describe the deep learning processing layer. These two data structures are simulated from the tensors and functions of neural calculus and have been modified and extended according to the actual programming model.

[0061] As can be seen from step S11, for the compilation method of the deep learning algorithm proposed in the embodiments of the present disclosure, it is first necessary to receive the operation data passed by the deep learning programming library interface. In the above-mentioned disclosed embodiments, it is also proposed that in one example, the deep learning programming library interface can be NCLAPI, and there can be two data structures, namely tensors and plastic operations, in NCLAPI. Therefore, in a possible implementation manner, the operation data passed by the deep learning programming library interface can be tensor data corresponding to the nclTensor data structure, or operation instructions corresponding to the nclOperator data structure, or can include both tensor data and operation instructions at the same time.

[0062] In a possible implementation manner, the tensor data in NCLAPI is an abstract representation of multi-dimensional data and can be used to represent the input data, output data, and model parameters of the deep learning processing layer. In one example, the input, output, and weight data of the convolutional layer can all be represented as tensor data. Therefore, in a possible implementation manner, the tensor data can have the following characteristics: including multiple attributes; being able to describe multi-dimensional data such as scalars, vectors, matrices, and tensors; following certain naming rules; and being able to describe multiple data types.

[0063] It is proposed in the above-mentioned disclosed embodiments that the tensor data follows certain naming rules, and these naming rules can be flexibly set according to the actual situation and are not limited to the following disclosed embodiments. In a possible implementation manner, the naming rules followed by the tensor data can be: it can only be composed of letters, numbers, and underscores; the first character must be an English letter; and it cannot contain punctuation marks and type specifiers.

[0064] Since deep learning computing models usually process data of a fixed size, taking the image classification model AlexNet as an example, the shapes of the input and output data of each of its processing layers are fixed, while the values of the data change frequently with the input. Therefore, the attributes of the data and the values of the data have completely different update frequencies. From the perspectives of data structure reuse and programming flexibility, the attributes of the data and the values of the data should be decoupled. Therefore, in a possible implementation manner, the tensor data in the embodiments of the present disclosure is only used to describe the attributes of the data, and the values of the data can be described by pointers pointing to memory regions, and the neural calculus tensor is completely mapped through the combination of tensor data and pointers.

[0065] It has been proposed in the above-mentioned disclosed embodiments that the characteristics of the tensor data are that it can include multiple attributes, and specifically which attributes are included can be flexibly set and selected according to the actual situation. In a possible implementation manner, the tensor data can include a shape attribute (shape), a logical data type attribute (dtype), a physical data type attribute (pdtype), and a physical layout attribute (layout).

[0066] By setting four attributes, namely the shape attribute, logical data type attribute, physical data type attribute, and physical layout attribute, for tensor data, various data in deep learning algorithms can be described sufficiently and well, enabling the compilation method to better adapt to various data situations in deep learning algorithms and enhancing the generality of the compilation method.

[0067] During the application process, the above attributes can be further classified based on the usage of tensor data, and the specific classification method can also be set according to the actual situation. Figure 3 Show a correspondence diagram between the attributes, classification, and meanings of tensor data according to an embodiment of the present disclosure. As shown in the figure, in one example, the above four attributes can be divided into two categories: visible attributes and invisible attributes. Among them, the shape attribute and logical data type attribute can be classified as visible attributes. During actual use, the visible attributes can be set through the tensor assignment interface; the physical data type attribute and physical layout attribute can be classified as invisible attributes, which can be maintained, modified, and used inside the programming library, thereby shielding hardware details from the outside world and reducing the programming complexity.

[0068] Through Figure 3 It can be seen that the physical data type attribute can be used to indicate the precision of data stored in the memory of the hardware device, while the logical data type attribute can be used to indicate the precision of data stored in the host memory. Therefore, the precision represented by the physical data type attribute and the logical data type attribute can be the same or different. In one possible implementation, the physical data type data and the logical data type attribute can be different. In this case, the compilation process can be set to implement the automatic precision optimization function, that is, during the compilation process, the data type with the fastest running speed can be automatically selected for operation, and this process can be transparent to the user. The specific implementation process of the automatic precision optimization function can be determined according to the actual situation and will be specifically described in the subsequent disclosed embodiments.

[0069] The above-mentioned disclosed embodiments also propose that tensor data can describe multiple data types, and which specific data types can be described can be flexibly determined according to the actual situation. In one possible implementation, tensor data can describe data types such as low bit-width and quantization. To enable tensor data to support low bit-width and quantization, the present disclosure embodiments design different data types (including logical data types and physical data types) for tensor data, including:

[0070] Double precision floating point: double; single precision floating point: float; half precision floating point: half; fixed point: fxm.b (m represents the number of integer bits, b represents the total number of bits); quantization: Qts.b (s represents the scale factor of the tensor, b represents the bias of the tensor); integer type: intx; unsigned integer type: uintx.

[0071] Deep learning algorithms can support channel-wise quantization of images, where the scale and bias of each channel can be different. Although channel-wise quantization cannot be described by Qts.b, it can be implemented using the scale and add operations provided by NCLAPI. Therefore, the expressive power of NCLAPI for quantization is still complete. In addition, considering that other data types may appear in the future, NCLAPI can also support the extension of the data types described by tensor data.

[0072] As proposed in the above disclosed embodiments, nclTensor is used to describe the input data, output data, and model parameters of deep learning processing layers. Since tensor data corresponds to nclTensor, and the most common deep learning processing layers are convolution, pooling, and RNN, their input data, output data, and model parameters are all high-dimensional data. Figure 4 A diagram showing the correspondence between tensor data and shape symbols according to an embodiment of the present disclosure is shown. As shown, in one possible implementation, the shapes of tensor data in deep learning operations can be defined according to the illustrated correspondence.

[0073] As proposed in the above disclosed embodiments, operation data can be created or called according to user instructions received by the deep learning programming library interface. Since tensor data is a possible implementation of operation data, it can be created or called. Tensor data must be created before use, and the specific creation and call processes can be flexibly set according to the actual situation. In one possible implementation, the call process can be to assign a value to the tensor data. In one possible implementation, the creation process can be to initialize its visible attributes when creating the tensor data. In one possible implementation, since deep learning frameworks such as Tensor decouple the creation of data objects and the setting of attributes, in order to be integrated into the deep learning framework without destroying the code structure of the deep learning framework, the creation process can be to first create an uninitialized tensor data, and then call the tensor assignment interface (nclSet-TensorAttr) for attribute assignment.

[0074] In a possible implementation, the operation instructions in NCLAPI are an abstract representation of transformations. It can be used to represent deep learning processing layers or general computing. In one example, the operation instructions can be used to represent deep learning processing layers, such as convolution, pooling, fully connected, etc. In the embodiments of the present disclosure, the operations performed by the operation instructions can be collectively referred to as malleable operations.

[0075] In a possible implementation, the operation instructions can consist of three parts: input parameters (inputparams), output parameters (output params), and operation type (OpType). Among them, the input parameters correspond to the set of input tensors of the transformation, that is, they can be the nclTensors and pointers corresponding to all input data. The output parameters correspond to the set of output tensors of the transformation, that is, they can be the nclTensors and pointers corresponding to all output data. The implementation methods of the input parameters and output parameters are not limited. In one example, it can be stipulated that an operation instruction allows zero or more (tensor data, pointer) as input parameters and one or more (tensor data, pointer) as output parameters.

[0076] The operation type can be used to specify what kind of data transformation the operation instruction performs. Users can specify different operation types when creating operation instructions. These operation types can represent three types of data transformations: value transformation, attribute transformation, and null transformation. Therefore, the operation instructions can not only describe deep learning processing layers but also describe general computing such as data segmentation, data splicing, and size scaling.

[0077] For deep learning programming library interfaces, they can provide a series of predefined self-owned operation instructions. In the embodiments of the present disclosure, for the NCLAPI interface, these operation instructions can be called build-in operation instructions (build-in operators). Figure 5 Show a diagram of the build-in operation instructions supported by NCLAPI according to an embodiment of the present disclosure. It can be seen from the figure that the operation instructions supporting in-place algorithms are the build-in operation instructions supported by NCLAPI.

[0078] The nature of the operation instructions can be flexibly set according to the actual situation. In a possible implementation, in order to make the behavior of the program easier to analyze and predict, except for the build-in operation instructions, the remaining operation instructions can all have unidirectionality, non-intersection, and idempotency. Among them, unidirectionality means that the operation instructions do not change the input parameters (including tensor data and the data pointed to by the pointer), non-intersection means that the input parameters and output parameters of the operation instructions cannot have the same name, and idempotency means that the result of the operation instruction call only depends on the input parameters and is not affected by the number of calls.

[0079] As proposed in the above disclosed embodiments, the operation data can be created or invoked according to user instructions received through a deep learning programming library interface. Since the operation instruction is a possible implementation of the operation data, it can be created or invoked. The specific processes of creation and invocation can be flexibly set according to the actual situation. In one possible implementation, the operation instruction needs to be created first and then invoked. The invocation of the operation instruction means mapping an operation instruction to a deep learning processor for execution. The operation instruction natively supports runtime variability of operation parameters. Specifically, the creation of the operation instruction only needs to be performed once, while the invocation of the operation instruction can be repeated, and different input and output parameters can be specified each time the operation instruction is invoked.

[0080] In one possible implementation, to optimize program performance, for the NCLAPI interface, two important functions are supported during the execution of the operation instruction invocation, namely asynchronous execution and automatic precision optimization.

[0081] In an example, asynchronous execution means that after the operation instruction invocation function is called by the host side, it will immediately return, and the CPU can perform other operations while the deep learning processor is performing calculations, thereby improving the overall utilization rate of the system and program performance. To ensure the completion of the asynchronous call execution, NCLAPI provides the device synchronization interface nclSyncDevice, which blocks the execution of the CPU until the device finishes the operation.

[0082] Figure 6 FIG. shows a schematic diagram of the implementation of automatic precision optimization according to an embodiment of the present disclosure. As shown, in an example, automatic precision optimization means that before the operation instruction is executed on the device, the programming library will automatically select the data type with the shortest execution time, convert the original data into the optimal format, and then perform the operation. Automatic precision optimization will jointly consider the time overhead of both data format conversion and operation instruction execution to ensure the shortest overall execution time of the operation. In addition, to meet the unidirectionality and idempotency of the operation, the programming library will apply for temporary space to complete the data format conversion to ensure that the original input data is not overwritten.

[0083] In one possible implementation, the operation instruction can also support operation connection. Operation connection means using the output parameter of an operation instruction A as the input parameter of another operation instruction B, and B processes the output data of A after A finishes the calculation. The necessary and sufficient condition for two operation instructions A and B to be connectable is that they each have at least one output tensor T1 and one input tensor T2, and the attributes of T1 and T2 are exactly the same. Operation connection has a directionality, and the direction of connection is from the operation providing the data to the operation using the data.

[0084] In a possible implementation, a deep learning computation model can be represented as a function composition. A function composition is a sequence of functions with only unidirectional function connections (in a function composition, the direction of function connections can only be from left to right). Since operation instructions can be mapped from functions, a deep learning computation model can be represented as a sequence of operation instructions with only unidirectional operation connections (in the sequence of operation instructions, the direction of operation instruction connections can only be from left to right). In the embodiments of the present disclosure, such a sequence of operation instructions is referred to as a unidirectional operation sequence. A directed graph can be converted into a unidirectional operation sequence according to a certain algorithm. In addition, in order to avoid the in-place algorithm from destroying the unidirectionality of operations, tensor alias technology can be used to eliminate in-place operations.

[0085] In one example, the algorithm for converting a directed graph into a sequence of unidirectional operation instructions can be: first, convert the directed graph g into a directed acyclic graph g'. Then, perform a topological sort on the directed acyclic graph g' to obtain:

[0086] g”(V{vertex1,vertex2,...,vertex n},E)

[0087] Map the vertices in the graph g” into a sequence of operation instructions in order; map the edges in g” into tensors, add the outgoing directed edges of a vertex to the output parameters of the vertex, and add the incoming edges of a vertex to the input parameters of the vertex.

[0088] In one example, the implementation form of tensor alias technology can be:

[0089]

[0090] Figure 7 A schematic diagram of a computation model according to an embodiment of the present disclosure is shown. As shown in the figure, the computation model can be represented as two unidirectional operation instruction sequences: (conv, pool, bn, relu, add) and (conv, pool, relu, bn, add). By calling the operation instructions in the order in which they appear in the unidirectional operation instruction sequence, one execution of the computation model can be completed. It should be noted that since there are operation connections from right to left in the relu and add operations, the above computation model cannot be represented as an operation instruction sequence of (conv, pool, add, bn, relu).

[0091] In a possible implementation, since deep learning algorithms usually process data of a fixed size, performance optimization can be achieved by fixing the parameters of the operation instructions. The operation instructions default to support all parameters being variable at runtime. To utilize fixed parameters to optimize operation performance, the embodiments of the present disclosure design parameter binding and operator specialization functions for the operation instructions. Parameter binding means fixing some or all of the input parameters of an operation instruction; operator specialization means converting an operation instruction with parameter binding into a new operation instruction, and this new operation instruction can be called a specialized operation instruction. The specialized operation instruction still satisfies the definition and properties of the operation instruction and supports all functions of the operation instruction (supporting specialization, fusion, etc.). The classification method of the specialized instructions can be flexibly set according to the actual situation. In a possible implementation, the specialized operation instructions can be divided according to the number of parameter bindings. In an example, the specialized operation instructions include fully specialized operation instructions, partially specialized operation instructions, and pseudo-specialized operation instructions; among them,

[0092] The fully specialized operation instructions include the operation instructions obtained by binding all the input parameters of an operation instruction and then converting;

[0093] The partially specialized operation instructions include the operation instructions obtained by binding N input parameters of an operation instruction and then converting, where N is a positive integer less than the number of input parameters of the operation instruction;

[0094] The pseudo-specialized operation instructions include the operation instructions obtained by directly converting an operation instruction without binding its input parameters.

[0095] It can be seen from the above embodiments of the present disclosure that in the examples of the present disclosure, the specialized operation instructions are divided into three categories: binding all the input parameters of an operation is a fully specialized operation instruction, binding some input parameters is a partially specialized operation instruction, and not binding any input parameters is a pseudo-specialized operation instruction. The bound parameters can be deleted from the input parameters of the operation instruction. When the user calls the operation instruction with bound parameters, there is no need to specify the bound parameters. Therefore, the fully specialized operation instruction can be called without specifying any input parameters.

[0096] By means of parameter binding and specialization operation instructions, during the compilation process, partial evaluation optimization of the program can be performed to reduce the running time of operation instructions. The specific implementation methods of parameter binding and specialization operation instructions are not limited. In one possible implementation, the nclSpecializeOperator interface can be responsible for implementing the specialization operation instructions. It can perform just-in-time compilation optimization on operations and return specialization operation instructions with shorter execution time on the hardware device. In an example, there may be a certain convolution operation. When the attributes and values of the input, output, and weights of this convolution operation are completely determined, parameter binding and operation specialization can be used to generate a faster-running convolution operation. Parameter binding and specialization operation instructions can be widely applied to real computing models. In an example, deep learning computing models usually process data of a fixed size. Therefore, the shape of the tensor can be parameter-bound, and then operation specialization can be performed to obtain specialization operation instructions, thereby optimizing the performance of the program. In an example, for the inference application scenario, the weights are pre-trained constants. Therefore, the weight data of the operation can be parameter-bound, and then operation specialization can be performed to obtain specialization operation instructions, thereby optimizing the program performance.

[0097] In one possible implementation, deep learning programming libraries usually only support processing layers with high usage frequency and long execution time (such as convolution, fully connected, RNN, pooling, activation), resulting in the programming library being unable to well support end-to-end execution. To solve the above problems, the embodiments of the present disclosure design an operation customization function for operation instructions to make them have customizable characteristics. Operation customization means writing an operation in a domain-specific programming language and then inserting it into the programming library in the form of binary code. In the embodiments of the present disclosure, such an operation is called a customized operator. The customized operator still satisfies the definition and properties of the operation instruction and supports all functions of the operation instruction (supporting specialization, fusion, etc.).

[0098] Since the customized operator is inserted into the programming library in the form of binary code, it indicates that the customized operator needs to be pre-compiled to generate binary code. The generation process of the binary code corresponding to the customized operator can be flexibly determined according to the actual situation of the NCLAPI and the deep learning programming library where it is located. Figure 8 The schematic diagram of the generation of the customized operator according to an embodiment of the present disclosure is shown. As shown in the figure, in one possible implementation, the generation process of the binary code corresponding to the customized operator may include:

[0099] According to the interface and data structure definition of the operation instruction, encapsulate the user instruction corresponding to the operation instruction to obtain the encapsulated user instruction;

[0100] Compile the encapsulated user instructions to obtain a compilation result;

[0101] Insert the compilation result into the static operation pool in the form of dynamic linking or static linking to obtain the binary code corresponding to the customized operation instruction.

[0102] The static operation pool in the above-mentioned disclosed embodiments is a storage area in the deep learning programming library, and its specific implementation manner will be specifically described in the subsequent disclosed embodiments.

[0103] By compiling the encapsulated user instructions to obtain a compilation result, and inserting the compilation result into the static operation pool of the deep learning programming library in the form of dynamic linking or static linking to obtain the binary code corresponding to the customized operation instruction, through this process, pre-compiled customized operation instructions can be generated, and the compilation results of the customized operation instructions can be saved. Non-library-built-in operation instructions that appear repeatedly multiple times can be converted into a packaged customized operation instruction. Thus, when implementing a certain deep learning algorithm, the operations to be executed can be directly implemented by calling the customized operation instruction, avoiding repeated and useless instruction editing. Moreover, since the compilation result of the customized operation instruction has been saved in the deep learning programming library, the binary code corresponding to the customized operation instruction can be directly called during compilation without repeated compilation multiple times, effectively improving the compilation efficiency and shortening the compilation time.

[0104] In an example, according to the generation process of the binary code corresponding to the customized operation instruction proposed in the above-mentioned disclosed embodiments, the specific process of operation customization can be as follows: Implement a customized transformation using a programming language to obtain the code to be inserted (insert code); Package the code to be inserted according to the interface and data structure definition of the operation instruction, and complete work such as data format conversion; Compile the code to be inserted and insert it into the deep learning programming library in the form of dynamic linking or static linking to complete operation customization; Use the customized operation instruction normally like an operation instruction. It should be noted that the name of the customized operation instruction is specified by the user, but it cannot conflict with the operation name of the build-in operation.

[0105] In a possible implementation manner, the operation instruction proposed in the embodiments of the present disclosure can also support an operation fusion function. Operation fusion (operator fusion) refers to combining multiple malleable operations in the call order into a new malleable operation, and this new operation can be called a fusion operation instruction (fusion operator). The fusion operation instruction still satisfies the definition and properties of the operation instruction and supports all functions of the operation instruction (supporting specialization, fusion, etc.). In an example, the formal representation of operation fusion is as follows:

[0106] op fused= Fuse(op1, op2,..., op n )

[0107] In the embodiments of the present disclosure, operation fusion satisfies weak transformation equivalence: the calculation results of the fused operation instructions and the calculation results of the original operation instruction sequence can be regarded as equal within the allowable error range, and it is formally expressed as error < epsilon, where epsilon is determined by the sensitivity of the application process to precision. In one example, the fused operation instructions can participate in operation fusion again, which is called high-order fusion, and the output obtained by high-order fusion is still a malleable operation, which is expressed as follows:

[0108] op fused2 = Fuse(op1, op2,..., op fused ,..., op n )

[0109] Operation fusion can bring two benefits: optimizing performance and simplifying programming. In terms of optimizing performance, the programming library can perform compilation optimization at the computational graph level inside the fused operation (for example, reducing the overall amount of computation and memory access through optimization techniques such as linear transformation and constant folding), thereby reducing the execution time of the operation on the device; in terms of simplifying programming, a single fused operation instruction can be used to represent common functional blocks (such as the residual block in ResNet) or even the entire computational model in deep learning algorithms. These highly abstract components can be reused repeatedly, thereby improving the efficiency of program development.

[0110] In one possible implementation, operation fusion needs to meet certain conditions. In one example, this condition can be that the operation instructions to be fused can be represented as a continuous subsequence in a unidirectional operation instruction sequence. As Figure 7 shown in the schematic diagram of the computational model, as proposed in the above-mentioned embodiments of the present disclosure, in one example, the computational model can be represented as the following two unidirectional operation instruction sequences, namely seq1: (conv, pool, bn, relu, add) and seq2: (conv, pool, relu, bn, add). Any subsequence in these two operation instruction sequences can perform operation fusion, Figure 9 shown in the schematic diagram of the operation fusion result according to an embodiment of the present disclosure. As shown in the figure, in one example, the three operation instructions conv, pool, and bn in seq1 can be fused, and at this time, the computational model (fusion, relu, add) can be obtained; Figure 10Shows a schematic diagram of the operation fusion result according to an embodiment of the present disclosure. As shown in the figure, in one example, the two operation instructions bn and add in seq2 can be fused. At this time, the calculation model (conv, pool, relu, fusion) can be obtained. In one example, the three operation instructions pool, relu, and add cannot be fused because they are neither continuous subsequences of seq1 nor seq2. If they are forcibly fused into the operation fusion, no matter which position in the sequence (conv, bn) the fusion is inserted, there will be a cyclic data dependency, that is, bn depends on the output result of fusion, and fusion also depends on the output result of bn. Therefore, operation fusion cannot be performed at this time.

[0111] According to the above principle of operation fusion, in a possible implementation, the creation process of the fusion operation instruction may include:

[0112] Create the name of the fusion operation instruction.

[0113] Determine the fusion operation sub-instructions according to the operation instructions to be fused.

[0114] Determine the operation connection relationship between the fusion operation sub-instructions according to the call order of the operation instructions to be fused.

[0115] Connect the fusion operation sub-instructions according to the operation connection relationship to obtain the connection result.

[0116] Set the input parameters and output parameters of the fusion operation instruction according to the user instruction corresponding to the fusion operation instruction.

[0117] Package the name, connection result, input parameters, and output parameters to obtain the fusion operation instruction.

[0118] Through the above creation process of the fusion operation instruction, multiple operation instructions can be conveniently fused into one fusion operation instruction, thereby effectively optimizing the compilation performance, reducing the execution time of the operation instructions on the device, and at the same time, common functional blocks or even the entire calculation model in the deep learning algorithm can be represented by a single fusion operation instruction. These abstract instructions can be reused repeatedly, thereby improving the program development efficiency.

[0119] The specific programming interface of the operation fusion can be flexibly set according to the actual situation. Figure 11 Shows a schematic diagram of the relevant programming interface of the operation fusion according to an embodiment of the present disclosure. In one example, based on the programming interface shown in the figure and combined with the above creation process of the fusion operation instruction, the process of creating a fusion operation can be obtained. Figure 12A schematic process diagram showing the creation of a fusion operation according to an embodiment of the present disclosure. As shown in the figure, in one example, the steps of creating a fusion operation may include:

[0120] Create a fusion operation and specify the operation type name of the fusion operation (which must not conflict with the build-in operation types).

[0121] Call the nclAddFusionOperator / nclSetFusionOperators interface to specify all sub-operations to be fused.

[0122] Call the nclLinkOperator interface to specify the operation connection relationships between the sub-operations.

[0123] Call the nclAddFusionInput and nclAddFusionOutput interfaces to set the input and output parameters of the fusion operation.

[0124] Call the nclFuseOperator interface to complete the operation fusion.

[0125] The purpose of function connection for sub-operations in the above steps is to construct a computational graph, and the nclFuseOperator interface will perform timely compilation and optimization on the computational graph, thereby accelerating the operation execution.

[0126] It can be seen from the above disclosed embodiments that the operation instructions proposed in the embodiments of the present disclosure may include operation fusion instructions, or may also include other types of operation instructions such as build-in operation instructions and specialization operation instructions. For different types of operation instructions, the implemented programming models are also different. In the embodiments of the present disclosure, NCLAPI adopts a hybrid programming model, that is, it supports both imperative programming and declarative programming. In one possible implementation manner, a hybrid programming model can be designed based on operation fusion, that is, the programming model without using fusion operations is an imperative programming model, and the programming model using fusion operations is a declarative programming model, and the two programming models can be used in combination.

[0127] The implementation of the programming model can be flexibly set according to the actual situation. In one possible implementation, the programming model can be designed based on three factors: data flow, execution flow, and control flow. In terms of data flow, in order to complete the data transfer between the host and the device, the embodiments of the present disclosure design a data copy interface nclMemcpy for NCLAPI. In terms of control flow, in order to control the device execution and synchronize between the host and the device, the embodiments of the present disclosure design an operation call interface nclInvokeOperator and a device synchronization interface nclSyncDevice for NCLAPI. In terms of execution flow, the embodiments of the present disclosure classify the execution methods of the computing model into three categories: layer-by-layer call: call all operations in the computing model one by one; fusion call: perform operation fusion on the entire computing model and then call the fused operation; segmented fusion call: perform segmented operation fusion on the computing model and then call the fused operation segment by segment. And these three execution methods are used to distinguish the NCLAPI programming model: the execution method of layer-by-layer call corresponds to the imperative programming model; the execution method of fusion call corresponds to the declarative programming model; the execution method of segmented fusion call corresponds to the hybrid programming. The above-mentioned embodiments of the present disclosure have proposed that in one possible implementation, the operation data is created or called according to the user instructions received by the deep learning programming library interface. In one example, based on the three execution methods of the computing model of the above-mentioned embodiments of the present disclosure, it can be known that triggering the call of the operation instructions in the deep learning programming library interface according to the user instructions may include: according to the user instructions, calling all corresponding operation instructions one by one; or, according to the user instructions, fusing all corresponding operation instructions to obtain a fused operation instruction and calling the fused operation instruction; or, according to the user instructions, segmenting all corresponding operation instructions to obtain segmented results, fusing each segmented result respectively to obtain corresponding segmented fusion operation instructions, and calling the segmented fusion operation instructions in sequence.

[0128] In one possible implementation, these three call methods can be used to distinguish the NCLAPI programming model. The execution method of layer-by-layer call can correspond to the imperative programming model; the execution method of fusion call can correspond to the declarative programming model; the execution method of segmented fusion call can correspond to the hybrid programming.

[0129] Figure 13A data flow schematic diagram of a three-layer computing model according to an embodiment of the present disclosure is shown. As shown, in one example, based on the programming model in the figure, the process of calling operation instructions through different calling methods can be as follows: Layer-by-layer calling: Call the nclInvokeOperator interface three times to execute the conv, pool, and fc operations respectively; Fusion calling: First fuse the conv, pool, and fc operations into a single operation, and then call the nclInvokeOperator interface once to execute the fused operation; Segmented fusion calling: Fuse the conv and pool operations, and then call the nclInvokeOperator interface twice to execute the fused operation and the fc operation respectively.

[0130] Figure 14 An implementation block diagram of a hybrid programming model according to an embodiment of the present disclosure is shown. As shown, in one example, the hybrid programming model is jointly composed of an imperative programming model and a declarative programming model. In one example, the complete programming process of the imperative programming model can be as follows: Initialization: Initialize the device and the running environment; Create an operation: Create a single operation, and selectively perform parameter binding and operation specialization; Call the operation: Prepare the operation parameters (including creating a Tensor and allocating device addresses), copy the host-side input data to the device memory, perform the operation call, synchronize the device, and read the output result from the device memory; Release the operation resources: Destroy the resources that are no longer used in the previous two steps, including tensors, operations, memory, etc.; Repeat creating, calling, and releasing operations until all operations in the computing model are executed; Exit: Shut down the device and destroy the running environment. In one example, the complete programming process of the declarative programming model can be as follows: Initialization: Initialize the device and the running environment; Create sub-operations: Create all sub-operations that need to participate in the fusion, and selectively perform parameter binding on the sub-operations; Create a fused operation: Create a fused operation, add the sub-operations to be fused, specify the operation connection relationship between the sub-operations, set the input and output parameters of the fused operation, perform operation fusion, and selectively perform operation specialization; Call the fused operation: Prepare the operation parameters (including creating a Tensor and allocating device addresses), copy the host-side input data to the device memory, perform the operation call, synchronize the device, and read the output result from the device memory; Release the operation resources: Release the sub-operation resources and release the fused operation resources; Exit: Shut down the device and destroy the running environment.

[0131] As proposed in the above disclosed embodiments, specializing operation instructions and fusing operation instructions can optimize the compilation performance. However, in one possible implementation, the compilation of both of these operation instructions needs to be implemented through just-in-time compilation, which may lead to an increase in the running time on the host side and a longer total running time of the program. Therefore, in one possible implementation, the compilation time can be further optimized and saved, and the compilation efficiency can be improved through an offline mode.

[0132] Since deep learning algorithms have a high degree of reusability, the optimized computational model can be repeatedly used for inference. Therefore, for the same computational model, the specialized operation instructions and fused operation instructions usually only need to be created once, while the operation instruction calls can be repeated. The more times the operation is called, the higher the benefits of specialized operation and operation fusion. Therefore, in a possible implementation, the embodiments of the present disclosure propose an offline mode for NCLAPI to eliminate the secondary compilation overhead of operation specialization and operation fusion on the host side. The implementation of the offline mode can be flexibly set according to the actual situation. In one example, operation specialization or operation fusion can be used in a separate program to optimize the operation instructions in advance and directly used in another program. In a possible implementation, the operation instructions optimized in advance can be referred to as offline operators. Corresponding to the offline mode is the online mode. Figure 15 FIG. shows a schematic diagram of the difference between the offline mode and the online mode according to an embodiment of the present disclosure. As shown in the figure, in one example, in the online mode, operation specialization or operation fusion can be first performed in the same program, and then the specialized operation instructions or fused operation instructions can be called.

[0133] It can be seen from the above disclosed embodiments that in a possible implementation, the offline operation instructions and their compiled binary codes need to be saved in advance for subsequent direct use. Therefore, in the embodiments of the present disclosure, an offline cache is also designed for NCLAPI. The implementation of the offline cache can be flexibly determined according to the actual situation. In a possible implementation, the offline cache includes an offline file and an index table; wherein, the offline file is used to save the pre-compiled results of the offline operation instructions; the index table is used to indicate the position of the pre-compiled results of the offline operation instructions in the offline file. The specific implementation of the offline file and the index table can also be flexibly selected according to the actual situation. In one example, the offline operations are saved to the offline file, and the position in the offline file is indexed through the index table. The index table is implemented by key-value pairs (Key, Value). The Key is the name of the offline operation, and the Value is a pointer used to point to the binary code corresponding to the offline operation in the offline file. The specific interface of the offline operation instructions can be set according to the actual situation. Figure 16A schematic diagram showing an offline interface according to an embodiment of the present disclosure is shown. As shown, in one example, the interface implementation method of the offline operation instruction can be: the nclSaveOperator interface saves the specified operation instruction to the offline cache, and uses the string specified by op_type as the name and index Key of the operation. The usage method of the offline operation instruction is exactly the same as that of the build-in operation instruction, except for the operation type. Therefore, the user can still use the nclCreateOperator interface to create an offline operation instruction. The NCLAPI can preferentially match the build-in operation instruction according to the given operation name. If it fails to match, it will then search in the offline cache to see if there is a corresponding offline operation instruction.

[0134] As can be seen from the above disclosed embodiments, the NCLAPI can transfer operation data and has good adaptability to deep learning algorithms. Therefore, in one possible implementation, the NCLAPI can be integrated into the deep learning framework. In one possible implementation, the deep learning framework can be class-extended to encapsulate tensor data and operation instructions within the data of the deep learning framework, realizing the integration of the deep learning framework and the deep learning programming library interface. In practical applications, with the different specific implementation methods of the deep learning framework, the implementation method of integrating the NCLAPI into the deep learning framework may also change accordingly.

[0135] In one example, the deep learning framework can be Caffe. Since Caffe contains three key data structures: Blob, Layer, and Net. Blob is mainly used to store data, complete data copying between the host and the device, and provide data access interfaces. Layer is used to represent operations (such as convolution, pooling, etc.), and it takes Blob as input and output. Caffe designs an inheritance system for Layer, and different operations can be implemented by writing subclasses of Layer. Layer has three key methods: Setup, Forward, and Backward, which are responsible for operation initialization, forward calculation, and backward calculation respectively. To support different devices, the same Layer subclass can contain multiple Forward and Backward methods.

[0136] Net saves all Blobs and Layers. It represents the complete computational model with a directed acyclic graph composed of Layers. Net has three key methods: Init, Forward, and Backward. The Init method converts the computational model defined by NetParameter (converted from prototxt) into Blobs and Layers, and calls the Setup method to initialize all Layers. The Forward method performs forward inference on the entire computational model, and the Backward method performs backward training on the computational model. Caffe uses prototxt to model the deep learning computational model. Users describe the processing layers, data, and the connection relationships between processing layers according to the syntax of prototxt. Caffe receives the prototxt file, converts it into Blobs, Layers, and Net, and then executes.

[0137] According to the composition of Caffe, in one example, the embodiments of the present disclosure integrate NCLAPI into Caffe in the following manner:

[0138] Extend the Blob class and encapsulate the nclTensor data structure and related interfaces (such as nclCreateTensor, nclSetTensorAttr, nclMemcpy) into the Blob.

[0139] Extend the Layer class and encapsulate the nclOperator data structure of NCLAPI and related interfaces into the Layer. Specifically, encapsulate interfaces such as nclCreateOperator, nclSpecializeOperator, and nclBlindOutputTensor into the Setup method of the Layer, and encapsulate nclInvokeOperator into the Forward and Backward methods of the Layer. In order to support new devices and operators without damaging the existing Layers, the embodiments of the present disclosure adopt the implementation method of adding a new subclass to the Layer.

[0140] Extend the Net class and encapsulate the operation fusion interface of NCLAPI into the Net. Since all Layers can be obtained in the Net, the Net is most suitable as the carrier for operation fusion. The embodiments of the present disclosure add an operation fusion module to the Net, which can perform segmented fusion or complete fusion on the computational model.

[0141] In one example, the deep learning framework can be TensorFlow. Figure 17The architecture diagram of TensorFlow according to an embodiment of the present disclosure is shown. As shown in the figure, in one example, the scalability of the architecture was well considered in the design of TensorFlow. It reserved an operator addition and device registration mechanism and provided a detailed official documentation. Therefore, it is relatively easy to integrate third-party deep learning programming libraries and deep learning processors into TensorFlow itself. Since the Distributed master of TensorFlow is responsible for the division of computational subgraphs and task allocation, in one example, the operations of NCLAPI can be integrated into the Distributed master module to perform operation fusion on subgraphs.

[0142] In one example, the embodiments of the present disclosure integrate NCLAPI into TensorFlow in the following manner:

[0143] Expand the Tensor class and encapsulate the nclTensor data structure and related interfaces (such as nclCreateTensor, nclSetTensorAttr, nclMemcpy) into Tensor;

[0144] Register new devices (deep learning processors) and NCLAPI operators according to the official documentation of TensorFlow;

[0145] Integrate the operation fusion function into the Distributed master module.

[0146] It can be seen from the above various disclosed embodiments that NCLAPI uses tensors to represent multi-dimensional data such as scalars, vectors, and matrices, and uses operation instructions to represent deep learning processing layers. The operation instructions support operation fusion, operation customization, operation specialization, runtime-variable operation parameters, offline optimization, and a hybrid programming model (imperative + declarative), thus well solving the problems of performance optimization and programming flexibility. In one example, a specific implementation of NCLAPI can be deployed on the DaDianNao deep learning processor platform, and it has also been integrated into mainstream deep learning frameworks such as Caffe and TensorFlow. Practice has proved that NCLAPI can run mainstream deep learning algorithms including image classification, object detection, and natural language processing, and it has strong generality and flexibility. In addition, the design of the data structure and interfaces of NCLAPI simulates neural calculus, thus proving that neural calculus can be used as a theoretical basis for guiding the design of deep learning programming libraries.

[0147] Figure 18A comparison schematic diagram of the NCLAPI and the interfaces of mainstream deep learning programming libraries according to an embodiment of the present disclosure is shown. It can be seen from the figure that compared with the interfaces of mainstream deep learning programming libraries, the NCLAPI can support a hybrid programming model, which can simultaneously meet the requirements of performance optimization and programming flexibility; it can support operation customization, has strong operation extensibility, and can better support end-to-end execution performance optimization; it can support operation fusion, operation specialization, and offline mode, and can optimize the performance of the program from multiple perspectives. At the same time, TensorFlow integrated with the NCLAPI can support a large number of deep learning computing models end-to-end (without using any CPU operations).

[0148] As can be seen from the above-mentioned various disclosed embodiments, in a possible implementation manner, the compilation method of the deep learning algorithm proposed in the embodiments of the present disclosure can be implemented based on the NCLAPI interface. After receiving the operation data transmitted by the NCLAPI, it can determine the instruction type according to the operation instructions included in the operation data, and perform the compilation operation corresponding to the instruction type according to the determination result to obtain the binary code of the deep learning algorithm. Therefore, based on this compilation method, the embodiments of the present disclosure also propose a deep learning programming library architecture adapted to this compilation method - the neurological calculus library architecture (NCLA). Figure 19 A schematic diagram of the overall architecture of the NCLA according to an embodiment of the present disclosure is shown. As shown in the figure, in a possible implementation manner, the NCLA can be composed of a just-in-time compilation system (NCLCS), a static operator pool (NCLSOPP), and a runtime system (NCLRT). Moreover, the NCLA can be integrated with the NCLAPI and perform compilation according to the operation data transmitted by the NCLAPI. The just-in-time compilation system can perform arithmetic and data collaborative compilation optimization on any operation instruction during runtime to generate efficient binary code; the static operator pool can be used to store the already optimized binary code, thereby eliminating the overhead of secondary compilation; the runtime system can provide basic functions such as device management, memory management, operation execution, and device synchronization, so that the operation instructions can be deployed end-to-end to the deep learning processor for execution.

[0149] As proposed in the above-mentioned disclosed embodiments, program optimization in the field of deep learning is extremely sensitive to algorithm and hardware changes. Therefore, in the implementation process, it is difficult to meet the performance requirements under different algorithms and different hardware platforms through ahead-of-time compilation optimization (AOT). Therefore, in a possible implementation manner, the disclosed embodiments of the present disclosure may adopt just-in-time compilation optimization (JIT) to design NCLA. Just-in-time compilation optimization can dynamically adjust the optimization strategy for different algorithms and different hardware platforms at runtime to achieve universal performance optimization. However, just-in-time compilation will introduce additional runtime overhead. A common means to alleviate this problem is to introduce a just-in-time compilation cache. Deep learning algorithms have a high degree of reusability, so the cache can play a huge role. In a possible implementation manner, the disclosed embodiments of the present disclosure use a static operation pool as the cache of the just-in-time compilation system. Since the coupling degree between operations and data is extremely high, optimizing operations or data alone cannot fully exert the performance of the deep learning processor. Therefore, in a possible implementation manner, a compilation framework for co-optimization of operations and data can be adopted to design just-in-time compilation.

[0150] Based on the above principle, the operation instructions included in the operation data transmitted by NCLAPI can be split. The specific splitting method can be flexibly selected according to the actual situation. In a possible implementation manner, the operation instructions can be divided into static operation instructions and dynamic operation instructions. The dynamic operation instructions can trigger NCLA to perform just-in-time compilation, while the static operation instructions can trigger NCLA to perform a lookup operation. Which specific instructions are included in the static operation instructions and the dynamic operation instructions can be determined according to the actual situation of the operation instructions transmitted by the deep programming interface, and are not limited to the following disclosed embodiments.

[0151] In a possible implementation manner, the static operation instructions may include one or more of custom operation instructions, build-in operation instructions, and offline operation instructions; among them,

[0152] The custom operation instructions include operation instructions with custom functions implemented according to the encoding method and encapsulation form of the operation instructions;

[0153] The build-in operation instructions include the self-owned operation instructions included in the deep learning programming library interface;

[0154] The offline operation instructions include pre-compiled dynamic operation instructions, where the pre-compiled results are stored in the offline cache.

[0155] The specific implementation manners of the custom operation instructions, build-in operation instructions, and offline operation instructions have been described in the above-mentioned disclosed embodiments, and will not be repeated here.

[0156] In a possible implementation, the dynamic operation instructions may include specialized operation instructions and / or fused operation instructions; wherein,

[0157] The specialized operation instructions include the operation instructions obtained by binding input parameters to the operation instructions and then converting them;

[0158] The fused operation instructions include the operation instructions obtained by combining multiple operation instructions according to the call order.

[0159] The specific implementation processes of the specialized operation instructions and the fused operation instructions have also been described in the above-mentioned disclosed embodiments, and will not be elaborated here.

[0160] Therefore, in a possible implementation, step S13 may include:

[0161] Step S131, determining the instruction type of the operation instruction.

[0162] Step S132, when the instruction type is a static operation instruction, looking up the corresponding binary code in the static operation pool according to the name of the static operation instruction, and using it as the binary code of the deep learning algorithm.

[0163] Based on the principle proposed in the above-mentioned disclosed embodiments, it can be seen that by looking up the corresponding binary code in the static operation pool according to the name of the static operation instruction when the instruction type is a static operation instruction, and using it as the binary code of the deep learning algorithm, it is possible to directly look up the corresponding binary code for the reusable operation instructions, avoid repeated compilations, eliminate the overhead of secondary compilation, and improve the compilation efficiency.

[0164] Furthermore, in a possible implementation, step S132 may include:

[0165] Looking up the binary code corresponding to the name in the static operation pool according to the name of the static operation instruction.

[0166] When the lookup result is successful, returning the binary code as the binary code of the deep learning algorithm.

[0167] In a possible implementation, step S132 may further include:

[0168] When the lookup result is failed, regarding the static operation instruction as a dynamic operation instruction and performing real-time compilation.

[0169] In the above-described disclosed embodiments, in a possible implementation manner, the customized operation instruction can be written by a user using a programming language in the field of deep learning. First, it is pre-compiled to generate binary code, and then inserted into the static operation pool in a dynamic linking or static linking manner. The binary code of the offline operation instruction can be generated by an just-in-time compilation system, and the user inserts it into the static operation pool by calling the nclSaveOperator interface. The build-in operation instruction is an operation provided by the NCLAPI. Considering that the program optimization in the field of deep learning is extremely sensitive to algorithms and hardware, in order to reduce the development cost, the disclosed embodiments do not adopt the method of manual optimization to implement the build-in operation instruction. Instead, the just-in-time compilation system is used in advance to perform pseudo-specialization (specialization without binding any input parameters) on the operation instruction to generate the binary code corresponding to the build-in operation instruction, and then insert it into the static operation pool.

[0170] As can be seen from the above-described disclosed embodiments, the static operation pool can be used to store the binary code corresponding to the static operation instruction, and its specific implementation manner can be flexibly set according to the actual situation, not limited to the following disclosed embodiments. In a possible implementation manner, the static operation pool may include a static operation source file, and the static operation source file includes a static code segment, a static data segment, a dynamic code segment, and a dynamic data segment; wherein, the static code segment is used to save the binary code corresponding to the build-in operation instruction; the dynamic code segment is used to save the binary code corresponding to the customized operation instruction; the static data segment is used to save the tensor data corresponding to the build-in operation instruction; and the dynamic data segment is used to save the tensor data corresponding to the customized operation instruction.

[0171] It was also proposed in the above-described disclosed embodiments that the binary code corresponding to the offline operation instruction can be stored in the offline cache. Therefore, in a possible implementation manner, the static operation pool may further include an offline cache.

[0172] Based on the above-described disclosed embodiments, Figure 20A schematic diagram showing the implementation of a static operation pool according to an embodiment of the present disclosure. As shown in the figure, in one example, four types of segments can be added to the source file (.so) of the deep learning programming library: a static code segment, a static data segment, a dynamic code segment, and a dynamic data segment. The static code segment and the static data segment are used to store the binary codes corresponding to the build-in operation instructions and the corresponding constant data, and the dynamic code segment and the dynamic data segment are used to store the binary codes corresponding to the custom operation instructions and the corresponding constant data. The reasons for distinguishing between the dynamic segment and the static segment are as follows: The custom operation instructions are written by the user and will continue to expand. Therefore, the embodiment of the present disclosure designs a dynamic segment with a variable size for it; while the build-in operation instructions are provided by the deep learning programming library itself and will not change. Therefore, the embodiment of the present disclosure designs a static segment with a fixed size for it. In addition, the embodiment of the present disclosure does not embed the binary codes corresponding to the offline operation instructions into the source file (.so) of the deep learning programming library because the offline operation instructions are usually heavyweight computational models (such as AlexNet, ResNet, etc.) and they occupy a huge storage space. Therefore, the embodiment of the present disclosure designs a separate offline cache for the offline operations and uses the file system to save the offline operation instructions. The offline cache consists of an index table (indextable) and offline files. The index table is implemented using key-value pairs (Key, Value). The Key is the type name of the offline operation instruction, and the Value is the binary code reference pointer corresponding to the offline operation instruction.

[0173] Based on the implementation of the static operation pool proposed in the above-mentioned disclosed embodiment, in a possible implementation, according to the name of the static operation instruction, searching for the binary code corresponding to the name in the static operation pool may include:

[0174] According to the name specified when the static operation instruction is created, in the static operation pool, sequentially search for the binary code corresponding to the custom operation, the binary code corresponding to the build-in operation, and the binary code corresponding to the offline operation to obtain the binary code corresponding to the name.

[0175] In one example, the user can specify the name of the operation when calling the operation creation interface (nclCreateOperator). NCLA uses the name of the operation as an index to search for the corresponding binary code in the static operation pool. The search order is sequentially the custom operation instruction, the build-in operation instruction, and the offline operation instruction. If the search hits, the corresponding binary code reference is returned; otherwise, just-in-time compilation is triggered.

[0176] The above-described disclosed embodiments illustrate specific compilation methods when the operation instruction is a static operation instruction. From the above-described disclosed embodiments, it can also be learned that the operation instruction can also be a dynamic operation instruction, which can trigger the NCLA to perform just-in-time compilation. Therefore, in a possible implementation manner, step S13 may further include step S133: when the instruction type is a dynamic operation instruction, perform real-time compilation on the dynamic operation instruction to obtain a real-time compilation result, which serves as the binary code of the deep learning algorithm.

[0177] As can be seen from the above process, if the NCLAPI interface passes dynamic operation instructions such as fused operation instructions or specialized operation instructions, just-in-time compilation (also known as real-time compilation) will be triggered when calling the operation fusion or operation specialization interfaces (nclFuseOperator, nclSpecializeOperator). The just-in-time compilation system generates highly optimized binary code for the operation, which is handed over to the runtime system for execution when the operation is called. In a possible implementation manner, the user can also call the nclSaveOperator interface to save the optimized operation instruction (such as a fused operation instruction), then the binary code corresponding to the operation instruction can be saved to the offline cache and indexed by its operation name. Otherwise, to ensure that the size of the programming library does not expand rapidly, the binary code corresponding to the unsaved operation instruction will be discarded after the program exits.

[0178] The specific process of performing real-time compilation on the dynamic operation instruction is not limited and can be flexibly selected according to the actual situation. In the disclosed embodiments of the present disclosure, a deep learning compilation framework for coordinated optimization of computation and data (computation and data unified compilation architecture, CDUCA) is involved to implement real-time compilation of the dynamic operation instruction. Figure 21 The architecture diagram of the CDUCA according to an embodiment of the present disclosure is shown. As shown in the figure, in a possible implementation manner, the CDUCA includes three components: a computation graph engine, a code generator, and a data optimizer. It performs coordinated optimization of computation and data at multiple levels to generate efficient binary code. The specific methods of the three components are not unique. In a possible implementation manner, the functions that the three components can implement are as follows:

[0179] Computation graph engine: Using optimization techniques such as linear transformation and constant folding, perform high-level algorithm-oriented optimization on the original computation graph and constant data to generate an optimized computation graph and constant data;

[0180] Code Generator: Adopt a heuristic search strategy based on a cost model to perform operations and data co-compilation optimization on the computational graph and constant data, and generate efficient target platform code and data descriptors;

[0181] Data Optimizer: Parse the data descriptors, perform optimizations such as splitting, reordering, and precision conversion on the constant data for the target platform, and then package the optimized constant data and target platform code (address relocation, etc.) to generate the final binary code.

[0182] Based on the architecture of the above disclosed embodiments, in a possible implementation, step S133 may include:

[0183] Step S1331, obtain the original computational graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction.

[0184] Step S1332, perform collaborative processing for the deep learning algorithm based on the original computational graph and original model data to obtain the first computational graph and the first model data.

[0185] Step S1333, generate hardware instructions and data descriptors according to the first computational graph.

[0186] Step S1334, perform processing for the hardware platform on the first model data according to the data descriptors to obtain the second model data.

[0187] Step S1335, obtain the binary code of the deep learning algorithm according to the hardware instructions and the second model data.

[0188] It can be seen from the above disclosed embodiments that the inputs of CDUCA are the computational graph and constant data. Therefore, in a possible implementation, it is necessary to obtain the original computational graph and original model data corresponding to the dynamic instruction through step S1331. The implementation manner of step S1331 is not limited. In a possible implementation, step S1331 may include:

[0189] Obtain the original computational graph corresponding to the dynamic operation instruction by parsing the dynamic operation instruction.

[0190] Obtain the original model data according to the parameters of the dynamic operation instruction.

[0191] Based on the above disclosed embodiments, in an example, the way to obtain the original computational graph and original model data corresponding to the dynamic instruction may be: directly obtain the original model data according to the parameters of the dynamic operation instruction, and the original computational graph can be generated by NCLA parsing the dynamic operation instruction. The specific parsing process can be flexibly determined according to the actual situation and is not limited here. Figure 22Shows the form of the original computational graph according to an embodiment of the present disclosure. As shown in the figure, in one example, the original computational graph contains two types of graph nodes, Tensor and Operator, corresponding to the input and output data of the operation instruction and the data transformation respectively.

[0192] It can be seen from the above disclosed embodiments that step S1332 can correspond to the computational graph engine component in CDUCA, and the implementation manner of the computational graph engine is not limited and can be flexibly determined according to the actual situation.

[0193] The nodes in the computational graph can be used to represent the operations performed in the deep learning process. Common operations in deep learning algorithms can include convolution, fully connected, activation, pooling, batch normalization, scaling, etc. These operations can be classified into two categories: linear transformation operations and non-linear transformation operations according to their specific implementation forms. Among them, all linear transformation operations can be expressed in the form of multiplying and adding vectors or matrices. Therefore, the general expression form of linear transformation operations can be:

[0194] Y = X * W + B

[0195] Where X and Y are variables, and W and B are model data constants.

[0196] Any linear transformation operation can be expressed in the above general expression form. Therefore, operations that cannot be expressed in the above general expression form are non-linear transformation operations. In one example, among the common operations of the above-listed deep learning algorithms, convolution, fully connected, batch normalization, and scaling operations are linear transformation operations, while pooling and activation are non-linear transformation operations. In one example, when expressing the fully connected operation in the above general expression form, X and Y represent the input neuron matrix and output neuron matrix of the fully connected respectively, W represents the weight matrix, and B represents the bias matrix. The specific expression methods of other linear transformation operations through the general expression form will not be elaborated here.

[0197] For linear transformation operations, if there are two consecutive linear transformation operations, which are respectively:

[0198] Y1 = X1 * W1 + B1

[0199] Y2 = Y1 * W2 + B2

[0200] Since linear transformation operations satisfy the distributive law and associative law, and both W and B in the general expression form of linear transformation operations are constants, the two linear transformation operations can be respectively equivalent linear transformed as follows:

[0201] Y2 = (X1 * W1 + B1) * W2 + B2

[0202] Y2 = X1 * W1 * W2 + B1 * W2 + B2

[0203] W' = W1 * W2, B' = B1 * W2 + B2

[0204] Y2 = X1 * W' + B'

[0205] Through the above equivalent linear transformation, the original linear transformation operation can be optimized as a whole through linear transformation optimization means and constant folding optimization means, and finally simplified to a single linear transformation operation. By optimizing through the above linear transformation optimization means and constant folding optimization means, while reducing the amount of computation, the model data can be compressed. On the one hand, it can reduce the storage overhead of the model data, and on the other hand, it can reduce the memory access volume during operation.

[0206] Therefore, in a possible implementation, the specific implementation form for the collaborative processing of deep learning algorithms can be linear transformation and constant folding.

[0207] Figure 23 The working flow chart of the computing graph engine according to an embodiment of the present disclosure is shown. As shown in the figure, in a possible implementation, step S1332 may include:

[0208] Step S13321, read the original computing graph and the original model data.

[0209] Step S13322, identify the continuous linear transformation operation nodes in the original computing graph.

[0210] Step S13323, process the continuous linear transformation operation nodes through linear transformation and constant folding to obtain the first computing graph and the first model data.

[0211] As can be seen from the above, 2 consecutive linear transformation operations can be simplified to 1 linear transformation operation through linear transformation and constant folding. When there are 3 consecutive linear transformation operations, the first 2 consecutive linear transformation operations can be simplified to 1 linear transformation operation through linear transformation and constant folding, and then this simplified linear transformation operation and the remaining another linear transformation operation can be simplified to 1 linear transformation operation again through linear transformation and constant folding. By analogy, when there are more consecutive linear transformation operations, through linear transformation and constant folding, these linear transformation operations can be merged and simplified to at least 1 linear transformation operation.

[0212] Since each linear transformation operation corresponds to a corresponding linear transformation operation node in the computational graph, in one possible implementation, consecutive linear transformation operation nodes may include: at least 2 consecutive linear transformation operation nodes. In this way, consecutive linear transformation operation nodes can correspond to at least 2 consecutive linear transformation operations, and the number of consecutive linear transformation operation nodes is not limited and can be determined according to the actual situation of the computational graph.

[0213] In one possible implementation, the specific process of step S13323 may include:

[0214] Perform linear transformation and constant folding on consecutive linear transformation operation nodes in the original computational graph, and merge the consecutive linear transformation operation nodes to obtain a first computational graph.

[0215] Merge the model data corresponding to the consecutive linear transformation operation nodes to obtain first model data.

[0216] In one example, the consecutive linear transformation operation nodes in the original computational graph may be nodes 1, 2, 3, and 4 connected in sequence, and the corresponding model data combination may be model data group 1. Through linear transformation and constant folding, nodes 1, 2, 3, and 4 can be merged into 1 node, and finally node 5 is obtained. In this process, since constant folding occurs in model data group 1, the model data contained therein may be merged, and finally a merged model data group is obtained, which can be called model data group 2. In one example, the consecutive linear transformation operation nodes in the original computational graph may be nodes 1, 2, 3, and 4 connected in sequence, and the corresponding model data combination may be model data group 1. Through linear transformation and constant folding, nodes 1, 2, and 3 can be merged into 1 node 6, and node 6 and node 4 are not merged. At this time, nodes 4 and 6 can be finally obtained. In this process, since constant folding occurs in nodes 1, 2, and 3, the corresponding model data combination in model data group 1 may be merged, and finally a merged model data group is obtained, which can be called model data group 3. Since the model data corresponding to the original node 4 does not undergo constant folding, combining model data group 3 and the model data corresponding to node 4 can obtain model data group 4 corresponding to the current overall linear transformation operation. By analogy from the above two examples, when the number of consecutive linear transformation operation nodes in the original computational graph changes, the specific process of step S13323 can also change accordingly, which will not be listed one by one here. In one example, the neural network corresponding to the deep learning algorithm may be the classic image classification network Resnet. Figure 24A schematic structural diagram showing a sub-structure included in an image classification network according to an embodiment of the present disclosure. As can be seen from the figure, in the Resnet network, a sub-structure of convolution + batch normalization + scaling can be included. Since convolution, batch normalization, and scaling are all linear transformation operations, in the computational graph corresponding to the Resnet network, the 3 operation nodes corresponding to this sub-structure can be linearly transformed and constant folded in the manner of the linear transformation operation in the above content, and finally merged into 1 operation node, and the model data involved in these 3 operation nodes is also merged accordingly.

[0217] Through the process in any of the above forms, the operation process of the deep learning algorithm can be optimized to obtain the first computational graph and the first model data after collaborative processing for the deep learning algorithm, thereby reducing the memory access amount when running this deep learning algorithm and also reducing the storage overhead when storing the model data.

[0218] Based on the first computational graph and the first model data, hardware instructions and data descriptors can be generated through step S1333. The specific implementation form of step S1333 is not limited, and any process that can generate hardware instructions based on the computational graph can be used as the implementation form of step S1333.

[0219] The main purpose of step S1333 is to generate hardware instructions readable by the corresponding hardware platform based on the first computational graph optimized in step S1332. To optimize the memory access performance of the hardware platform, on-chip memory is often designed in the part of the hardware platform close to the computing position. In one example, this part can be the position close to the computing unit of the deep learning processor. When accessing the hardware platform, the speed of accessing the on-chip memory is often faster than that of accessing other positions. In one example, other positions can be off-chip double data rate synchronous dynamic random access memory (DDR). Affected by the position and function of the on-chip memory, the capacity of the on-chip memory is limited. Therefore, the utilization rate of the on-chip memory can directly affect the performance optimization effect of the deep learning algorithm for the corresponding hardware platform. However, through the collaborative processing method for the deep learning algorithm in step S1332, the optimization of the operation can be achieved, but the utilization rate of the on-chip memory cannot be improved. In a possible implementation manner, the utilization rate of the on-chip memory can be improved by optimizing the process of generating hardware instructions, and then the performance optimization effect of the deep learning algorithm for the corresponding hardware platform can be enhanced. Therefore, in a possible implementation manner, step S1333 can include: processing the first computational graph according to a cost model and combining a heuristic search strategy to obtain hardware instructions and data descriptors.

[0220] The specific process of processing the first computational graph according to the cost model and combining the heuristic search strategy to obtain the hardware instructions and data descriptors can be flexibly selected according to the actual situation. Figure 25 The flowchart showing the compilation method of the deep learning algorithm according to an embodiment of the present disclosure is shown. As shown in the figure, in a possible implementation manner, step S1333 may include:

[0221] Step S13331, modeling the first computational graph through a cost model to generate a search space and an objective function.

[0222] Step S13332, performing a search within the search space through a heuristic search strategy. When the objective function reaches a threshold, generate hardware instructions and data descriptors for the hardware platform.

[0223] There can be multiple implementation forms for generating hardware instructions based on the computational graph. For example, the computational graph can be used as input, passed through a model that can generate hardware instructions, and optimization can be performed on the output result of the model to obtain the final hardware instructions for the hardware platform. Which model that can generate hardware instructions is specifically applied is not limited. In a possible implementation manner, a cost model can be used to model the first computational graph. The cost model estimates the total time for operations to be executed on the deep learning processor (including overhead such as data format conversion). The main factors it considers are the memory access time, operation time, and the overlap rate between the two. These three factors directly determine the performance of the program. Specifically, the deep learning processor includes an independent operation unit and a memory access unit, and the compute instructions and memory access instructions can be executed in a pipelined and overlapping manner. Figure 26 The schematic diagram of instruction pipelining according to an embodiment of the present disclosure is shown. As shown in the figure. To minimize the total running time of the program, it is necessary to reduce the compute time, memory access time, and increase the overlap rate of their execution. Therefore, after modeling the first computational graph through the cost model, the generated hardware instructions are actually a search space composed of multiple possible hardware instructions. The hardware instructions included in the search space can indicate various selection methods of the hardware platform during operation: In one example, the hardware instructions can indicate that the hardware platform divides a complete piece of data into several times of loading and places it in the on-chip cache; In one example, the hardware instructions can indicate that the hardware platform divides a complete piece of data into several times of loading and swaps it in and out between the on-chip cache and off-chip DDR according to requirements; In one example, the hardware instructions can indicate how much data needs to be processed when performing a vector operation once. Since there are multiple hardware instructions in the search space, which specific application instructions are ultimately applied requires searching within the search space to find the optimal instruction combination as the final obtained hardware instructions.

[0224] As can be seen from the above content, the cost model can give the estimated running time of hardware instructions. Therefore, after passing the first computational graph through the cost model, a corresponding objective function can also be generated. The specific form of the objective function is not limited here and can be flexibly set according to the actual situation. This objective function can indicate the time consumption of the finally generated hardware instructions during operation when searching within the search space. Therefore, when the objective function reaches the threshold, it indicates that the generated hardware instructions meet the running time requirements, and further indicates that the generated hardware instructions can improve the performance of the hardware instructions during operation on the hardware platform. Since the specific implementation form of the objective function is not limited, the threshold of the corresponding objective function is also not limited and can be flexibly set according to the actual situation.

[0225] There is also no limitation on how to search within the generated search space. In one possible implementation, a brute-force search method can be directly used for searching. However, such a search process takes too long and may prolong the compilation process, and may further reduce the performance optimization effect of the deep learning algorithm for the corresponding hardware platform. In one possible implementation, a heuristic search strategy can be used to search within the search space.

[0226] The purpose of the heuristic search strategy is to improve the search efficiency within the search space. Therefore, it is not limited to a specific search method and can be flexibly selected according to the actual situation. In one possible implementation, the heuristic search strategy can include: a search strategy for improving the utilization rate of on-chip caches within the hardware platform; or, a search strategy for reducing the operation granularity and memory access granularity on the basis of ensuring the utilization rates of the arithmetic units and access units within the hardware platform. In one example, the purpose of the heuristic search strategy can be to try to fully utilize the on-chip cache. Therefore, the search strategy can be set as a search strategy for improving the utilization rate of on-chip caches within the hardware platform at this time. In one example, the purpose of the heuristic search strategy can be to select an instruction combination with relatively small operation and access granularities as much as possible on the premise of ensuring the utilization rates of the arithmetic units and memory access units in the hardware platform, so as to improve the coverage of calculation and memory access. Therefore, the search strategy can be set as a search strategy for reducing the operation granularity and memory access granularity on the basis of ensuring the utilization rates of the arithmetic units and access units within the hardware platform at this time. In one example, the heuristic search strategy can also be a balanced strategy of the above two strategies, that is, a search strategy that makes the above two strategies reach the comprehensive optimal situation.

[0227] Through any of the above search strategies, hardware instructions for a hardware platform can be generated after searching within the search space. Since the hardware instructions included in the search space can indicate various selection methods during the operation of the hardware platform, when searching within the search space, in addition to considering the running performance of the above hardware instructions, it is also possible to consider using the generated hardware instructions to guide the hardware platform to optimize the model data, so as to further improve the operation memory access speed and memory access efficiency of the hardware platform. Therefore, based on this consideration, while generating hardware instructions in step S13332, a data descriptor can also be generated to guide how to further optimize the model data.

[0228] Based on the data descriptor generated simultaneously in step S1333, the first model data can be further optimized to obtain the second model data. The specific implementation form of how the data descriptor optimizes the model data is not limited. Figure 27 Fig. shows an implementation diagram of optimizing model data according to an embodiment of the present disclosure. As shown in the figure, in a possible implementation, step S1334 may include: splitting the first model data according to the data descriptor and the arithmetic requirements within the hardware platform; aligning the first model data according to the data descriptor and the arithmetic requirements within the hardware platform; or, performing dimensional transformation on the first model data according to the data descriptor and the arithmetic requirements within the hardware platform; or, selecting the precision of the first model data according to the data descriptor and the arithmetic requirements within the hardware platform.

[0229] In the above disclosed embodiments, in a possible implementation, for data splitting (tiling), loops, on-chip memory management, and instruction pipelining can be optimized; for data alignment (align), deep learning processors support vector and matrix operations, so there are certain requirements for data alignment. Aligned access can accelerate memory access speed and reduce DDR bank conflicts; for data reordering (reorder), the input data of operations such as convolution is a multi-dimensional array, and the arrangement order of the multi-dimensional array will affect the number of memory access jumps. Reordering the data can improve memory access locality and reduce the MMU miss rate; for precision conversion (type convert), deep learning processors usually support low-precision operations, such as half-precision floating-point and qubit quantization. Different precision operations have different performances. Performing appropriate precision conversion on the data can improve the overall performance of the program; precision conversion is a specific implementation of the NCLAPI precision automatic optimization function.

[0230] The specific implementation of data splitting is not limited. Figure 28A schematic diagram showing the optimization after data splitting according to an embodiment of the present disclosure is shown. As shown in the figure, in one example, constant data with 2 channels, a height of 4, and a width of 4 can be optimized. Assume that the data descriptor requires the following transformation of the data: both the height and width are split by 2; the arrangement order of the data is adjusted from HWC to CHW; the data operation precision is adjusted from float32 to half. The data optimizer performs physical transformations such as splitting, rearrangement, precision conversion, and alignment on the data from the DDR perspective according to the above description, and finally obtains the optimized data.

[0231] In one example, during the operation according to the hardware instructions, the hardware platform may have certain requirements for data alignment. If this process is performed within the hardware platform, it will greatly reduce the performance of the hardware platform during operation. Therefore, according to the data descriptor, the data to be aligned can be aligned in advance during the compilation process, so as to accelerate the memory access speed when the hardware platform is working. The specific alignment standards and methods are not limited here and can be determined according to the actual requirements of the hardware platform; in one example, during the operation according to the hardware instructions, some algorithms, such as the convolution algorithm, may need to interpret the data as a multi-dimensional array, and the arrangement order of the multi-dimensional array will affect the number of memory access jumps within the hardware platform, thus affecting the performance of the hardware platform during operation. Therefore, according to the data descriptor, the dimensional transformation of the model data can be performed in advance according to the operation requirements during the compilation process to improve the memory access locality, so as to minimize the memory access jumps of the hardware platform. The specific dimensional transformation and the transformation method are not limited here and can be determined according to the actual requirements of the hardware platform; in one example, since the operation speeds corresponding to different operation precisions are different, and the hardware platform may also support different operation precision requirements. In one example, the hardware platform may support low-precision operations more. If high-precision operations are used, it may reduce the running speed of the hardware platform. Therefore, according to the data descriptor, the preferred model data precision can be selected in advance according to the precision requirements of the hardware platform during the compilation process. The specific precision of the data to be selected is not limited here again and can be determined according to the requirements of the hardware platform. In one example, the preferred precision can be 16-bit quantization precision, and in one example, the preferred precision can be 8-bit quantization precision.

[0232] The process of optimizing the first model data into the second model data as described above can be any one of the above four methods, or any combination form of the above four methods, or may also include other methods that are more conducive to improving the speed of the hardware platform, which are not listed here. By completing the optimization of the model data in advance during the compilation process, the performance of the hardware platform during operation can be greatly improved, such as improving the memory access speed and memory access efficiency.

[0233] Based on the above process, the binary code of the deep learning algorithm can be obtained. The content included in this binary code can exist in various forms. In one possible implementation, step S1335 may include: packing the hardware instructions and the second model data to obtain the binary code of the deep learning algorithm.

[0234] In one example, the hardware instructions may be the hardware instructions generated according to the first computational graph, and the second model data may be the model data obtained by respectively performing co-processing for the deep learning algorithm and processing for the hardware platform on the original model data. That is, the finally obtained binary code of the deep learning algorithm is the content obtained by successively going through step S1332, step S1333, and step 1334. In one example, the hardware instructions may be the hardware instructions directly generated according to the original computational graph, and the second model data may be the model data obtained by only undergoing processing for the hardware platform on the original model data. That is, the finally obtained executable file of the deep learning algorithm is the content obtained by only successively going through step S1333 and step S1334. In one example, the hardware instructions may be the hardware instructions generated according to the first computational graph, and the second model data may be the first model data. That is, the finally obtained executable file of the deep learning algorithm is the content obtained by only successively going through step S1332 and step S1333. In one example, the hardware instructions may be the hardware instructions directly generated according to the original computational graph, and the second model data may be the original model data. That is, the finally obtained executable file of the deep learning algorithm may be the content obtained by only going through step S1333. As can be seen from the above examples, step S1332, step S1333, and step S1334 may not exist simultaneously and can be flexibly combined according to the actual situation.

[0235] The CDUCA implemented through the above disclosed embodiments can integrate the foregoing compilation optimization technologies such as memory reuse, operation fusion, latency hiding, linear algebra transformation, common sub-expression elimination, constant propagation, dead code elimination, and data parallelism. In addition, the hierarchical design structure of CDUCA has strong scalability, and developers can integrate various compilation optimization technologies in each module of CDUCA. For example, the operation aggregation technology can be integrated in the computational graph engine module, and the polyhedral compilation optimization technology can be integrated in the code generator module. By performing real-time compilation on dynamic operation instructions through CDUCA, the compilation efficiency can be effectively improved, thereby improving the running speed of the hardware device.

[0236] As can be seen from the above disclosed embodiments, the binary code of the deep learning algorithm can be generated through two methods: just-in-time compilation and static lookup. However, it can also be seen from the above disclosed embodiments that in the overall architecture of NCLA, there is also a runtime system. Therefore, in one possible implementation, the method proposed in the embodiments of the present disclosure further includes: through the runtime system, executing the binary code of the deep learning algorithm in the deep learning processor.

[0237] The implementation manner of the runtime system is not limited. Figure 29 The module and function schematic diagram of the runtime system according to an embodiment of the present disclosure are shown. As shown in the figure, the runtime system can be responsible for the interaction between the host and the deep learning processor. It can encapsulate the device driver interface and provide functions such as device management, memory management, operation execution, and device synchronization for the upper layer.

[0238] Through the above disclosed embodiments, it is possible to achieve the compilation of the deep learning algorithm through NCLA. Through experiments, it can be confirmed that a specific implementation of NCLA can currently be deployed on the deep learning processor platform and can support mainstream deep learning algorithms including image classification, object detection, and natural language processing. The embodiments of the present disclosure have conducted experiments on several common deep learning applications using TensorFlow. On average, the performance of the binary code generated by the NCLA just-in-time compilation system can reach 83.24% of the performance of the manually optimized code. In the best case, NCLA can at least exert 72.61% of the hardware peak performance. In addition, the successful cases of NCLA further confirm the generality of neural calculus and the NCLAPI mentioned in the above disclosed embodiments.

[0239] Figure 30 The block diagram of the compilation device of the deep learning algorithm according to an embodiment of the present disclosure is shown. As shown in the figure, the device 20 includes: an operation data receiving module 21, configured to receive operation data transmitted by the deep learning programming library interface; an operation instruction obtaining module 22, configured to obtain the operation instruction included in the operation data; and a compilation module 23, configured to determine the instruction type of the operation instruction and perform a compilation operation corresponding to the instruction type according to the determination result to obtain the binary code of the deep learning algorithm.

[0240] In one possible implementation, the operation instruction is created or called according to the user instruction received by the deep learning programming library interface.

[0241] In a possible implementation, the compilation module includes: a judgment unit for judging the instruction type of the operation instruction; a static lookup unit for, when the instruction type is a static operation instruction, looking up a corresponding binary code in a static operation pool according to the name of the static operation instruction as the binary code of the deep learning algorithm.

[0242] In a possible implementation, the static lookup unit is configured to: look up a binary code corresponding to the name in the static operation pool according to the name of the static operation instruction; and when the lookup result is successful, return the binary code as the binary code of the deep learning algorithm.

[0243] In a possible implementation, the static lookup unit is further configured to: when the lookup result is failed, regard the static operation instruction as a dynamic operation instruction for real-time compilation.

[0244] In a possible implementation, the static operation instructions include one or more of customized operation instructions, build-in operation instructions, and offline operation instructions; wherein, the customized operation instructions include operation instructions with custom functions implemented according to the encoding method and encapsulation form of the operation instruction; the build-in operation instructions include the native operation instructions included in the deep learning programming library interface; the offline operation instructions include pre-compiled dynamic operation instructions, and the result of the pre-compilation is stored in an offline cache.

[0245] In a possible implementation, the process of generating the binary code corresponding to the customized operation instruction includes: encapsulating the user instruction corresponding to the operation instruction according to the interface and data structure definition of the operation instruction to obtain an encapsulated user instruction; compiling the encapsulated user instruction to obtain a compilation result; and inserting the compilation result into the static operation pool in a dynamic link or static link manner to obtain the binary code corresponding to the customized operation instruction.

[0246] In a possible implementation, the static operation pool includes a static operation source file, and the static operation source file includes a static code segment, a static data segment, a dynamic code segment, and a dynamic data segment; wherein, the static code segment is used to store the binary code corresponding to the build-in operation instruction; the dynamic code segment is used to store the binary code corresponding to the customized operation instruction; the static data segment is used to store the tensor data corresponding to the build-in operation instruction; and the dynamic data segment is used to store the tensor data corresponding to the customized operation instruction.

[0247] In a possible implementation, the static operation pool further includes an offline cache, and the offline cache includes an offline file and an index table; wherein, the offline file is used to save the pre-compiled result of the offline operation instruction; the index table is used to indicate the position of the pre-compiled result of the offline operation instruction in the offline file.

[0248] In a possible implementation, the static lookup unit is further configured to: sequentially look up the binary code corresponding to the customized operation, the binary code corresponding to the build-in operation, and the binary code corresponding to the offline operation in the static operation pool according to the name specified when the static operation instruction is created, so as to obtain the binary code corresponding to the name.

[0249] In a possible implementation, the compilation module further includes a dynamic compilation unit, which is configured to: when the instruction type is a dynamic operation instruction, perform real-time compilation on the dynamic operation instruction to obtain a real-time compilation result as the binary code of the deep learning algorithm.

[0250] In a possible implementation, the dynamic compilation unit includes: an original data acquisition subunit, which is configured to obtain an original computation graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction; a collaborative processing subunit, which is configured to perform collaborative processing for the deep learning algorithm according to the original computation graph and the original model data to obtain a first computation graph and a first model data; a hardware instruction generation subunit, which is configured to generate hardware instructions and data descriptors according to the first computation graph; a model data processing subunit, which is configured to process the first model data for the hardware platform according to the data descriptor to obtain a second model data; a binary code generation subunit, which is configured to obtain the binary code of the deep learning algorithm according to the hardware instructions and the second model data.

[0251] In a possible implementation, the original data acquisition subunit is configured to: obtain an original computation graph corresponding to the dynamic operation instruction by parsing the dynamic operation instruction; obtain the original model data according to the parameters of the dynamic operation instruction.

[0252] In a possible implementation, the collaborative processing subunit is configured to: read the original computation graph and the original model data; identify continuous linear transformation operation nodes in the original computation graph; process the continuous linear transformation operation nodes through linear transformation and constant folding to obtain a first computation graph and a first model data.

[0253] In a possible implementation, the hardware instruction generation subunit is configured to: process the first computation graph according to a cost model and combine a heuristic search strategy to obtain hardware instructions and data descriptors.

[0254] In a possible implementation, the model data processing subunit is configured to: align the first model data according to the data descriptor and the operation requirements within the hardware platform; or transform the dimensions of the first model data according to the data descriptor and the operation requirements within the hardware platform; or select the precision of the first model data according to the data descriptor and the operation requirements within the hardware platform.

[0255] In a possible implementation, the binary code generation subunit is configured to: package the hardware instructions and the second model data to obtain the binary code of the deep learning algorithm.

[0256] In a possible implementation, the dynamic operation instructions include specialization operation instructions and / or fusion operation instructions; wherein, the specialization operation instructions include the operation instructions obtained by binding input parameters to the operation instructions and then performing conversion; the fusion operation instructions include the operation instructions obtained by combining multiple operation instructions according to the call order.

[0257] In a possible implementation, the specialization operation instructions include complete specialization operation instructions, partial specialization operation instructions, and pseudo-specialization operation instructions; wherein, the complete specialization operation instructions include the operation instructions obtained by binding all input parameters to the operation instructions and then performing conversion; the partial specialization operation instructions include the operation instructions obtained by binding N input parameters to the operation instructions and then performing conversion, where N is a positive integer less than the number of input parameters of the operation instructions; the pseudo-specialization operation instructions include the operation instructions obtained by directly performing conversion on the operation instructions without binding input parameters.

[0258] In a possible implementation, the creation process of the fusion operation instructions includes: creating the name of the fusion operation instructions; determining the fusion operation sub-instructions according to the operation instructions to be fused; determining the operation connection relationship between the fusion operation sub-instructions according to the call order of the operation instructions to be fused; connecting the fusion operation sub-instructions according to the operation connection relationship to obtain a connection result; setting the input parameters and output parameters of the fusion operation instructions according to the user instructions corresponding to the fusion operation instructions; and packaging the name, connection result, input parameters, and output parameters to obtain the fusion operation instructions.

[0259] In a possible implementation, the device further includes an execution module, configured to: execute the binary code of the deep learning algorithm in the deep learning processor through the runtime system.

[0260] In a possible implementation, the operation data further includes tensor data, where the tensor data includes a shape attribute, a logical data type attribute, a physical data type attribute, and a physical layout attribute.

[0261] In a possible implementation, the device is further configured to: perform class extension on the deep learning framework, encapsulate the tensor data and the operation instruction within the data of the deep learning framework, and implement the integration of the deep learning framework and the deep learning programming library interface.

[0262] In a possible implementation, the present disclosure also proposes a deep learning computing device, which includes the compilation device of any one of the above possible deep learning algorithms, and the deep learning computing device is used to complete the set deep learning operations.

[0263] Figure 31 A block diagram of a combined processing device according to an embodiment of the present disclosure is shown. As shown in the figure, the combined processing device includes the above deep learning computing device, a general interconnect interface, and other processing devices.

[0264] The deep learning computing device interacts with other processing devices to jointly complete the operations specified by the user. The other processing devices include one or more types of general / special purpose processors such as a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processor. The number of processors included in the other processing devices is not limited. The other processing devices serve as an interface between the deep learning computing device and external data and control, including data transfer, and complete basic controls such as starting and stopping the deep learning computing device; the other processing devices can also cooperate with the deep learning computing device to jointly complete the computing tasks. The general interconnect interface is used to transfer data and control instructions between the deep learning computing device and other processing devices. The deep learning computing device obtains the required input data from other processing devices and writes it into the storage device on the chip of the deep learning computing device; it can obtain control instructions from other processing devices and write them into the control cache on the chip of the deep learning computing device; it can also read the data in the storage module of the deep learning computing device and transfer it to other processing devices.

[0265] The combined processing device may further include a storage device, which is respectively connected to the deep learning computing device and the other processing devices. The storage device is used to store the data in the deep learning computing device and the other processing devices, and is particularly suitable for data that cannot be fully stored in the internal storage of the deep learning computing device or other processing devices for the required operations.

[0266] The combined processing device can be used as a system-on-chip (SOC) for devices such as mobile phones, robots, drones, and video surveillance devices, effectively reducing the core area of the control part, improving the processing speed, and reducing the overall power consumption. In this case, the general interconnect interface of the combined processing device is connected to certain components of the device. Some components include, for example, cameras, displays, mice, keyboards, network cards, and Wi-Fi interfaces.

[0267] In a possible implementation, the present disclosure also provides a deep learning chip, which includes the above-mentioned deep learning computing device or combined processing device.

[0268] In a possible implementation, the present disclosure also provides a chip packaging structure, which includes the above-mentioned chip.

[0269] In a possible implementation, the present disclosure also provides a board, which includes the above-mentioned chip packaging structure.

[0270] In a possible implementation, the present disclosure also provides an electronic device, which includes the above-mentioned board.

[0271] The electronic device includes a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a driving recorder, a navigator, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, a headset, a mobile storage device, a wearable device, a vehicle, a household appliance, and / or a medical device.

[0272] The vehicle includes an airplane, a ship, and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, a rice cooker, a humidifier, a washing machine, a light, a gas stove, and a range hood; the medical device includes a nuclear magnetic resonance instrument, a B-ultrasound instrument, and / or an electrocardiogram instrument.

[0273] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.

[0274] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0275] In several embodiments provided by the present disclosure, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0276] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0277] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software program module.

[0278] If the above integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The aforementioned memory includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.

[0279] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memory (abbreviation: ROM), random access memory (abbreviation: RAM), magnetic disks, or optical discs, etc.

[0280] The above has introduced the embodiments of the present disclosure in detail. Specific examples are used herein to elaborate on the principles and implementation manners of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.

[0281] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0282] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0283] The above has described the embodiments of the present disclosure. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical applications, or the improvements to the technology in the market, or to enable other ordinary skill in the art in this technical field to understand the disclosed embodiments.

[0284] The foregoing can be better understood in accordance with the following clauses:

[0285] Clause A1. A method for compiling a deep learning algorithm, the method comprising:

[0286] Receiving operation data transmitted by a deep learning programming library interface;

[0287] Obtain the operation instructions included in the operation data;

[0288] Judge the instruction type of the operation instruction, and perform a compilation operation corresponding to the instruction type according to the judgment result to obtain the binary code of the deep learning algorithm.

[0289] Clause A2. According to the method described in Clause A1, the operation data is created or called according to the user instruction received by the deep learning programming library interface.

[0290] Clause A3. According to the method described in Clause A2, the judging the instruction type of the operation instruction and performing a compilation operation corresponding to the instruction type according to the judgment result to obtain the binary code of the deep learning algorithm includes:

[0291] Judge the instruction type of the operation instruction;

[0292] When the instruction type is a static operation instruction, look up the corresponding binary code in the static operation pool according to the name of the static operation instruction, and use it as the binary code of the deep learning algorithm.

[0293] Clause A4. According to the method described in Clause A3, the looking up the corresponding binary code in the static operation pool according to the name of the static operation instruction and using it as the binary code of the deep learning algorithm when the instruction type is a static operation instruction includes:

[0294] Look up the binary code corresponding to the name in the static operation pool according to the name of the static operation instruction;

[0295] When the lookup result is successful, return the binary code as the binary code of the deep learning algorithm.

[0296] Clause A5. According to the method described in Clause A4, the looking up the corresponding binary code in the static operation pool according to the name of the static operation instruction and using it as the binary code of the deep learning algorithm when the instruction type is a static operation instruction further includes:

[0297] When the lookup result is failed, regard the static operation instruction as a dynamic operation instruction and perform real-time compilation.

[0298] Clause A6. According to the method described in Clause A3, the static operation instruction includes one or more of a custom operation instruction, a build-in operation instruction, and an offline operation instruction; among them,

[0299] The custom operation instruction includes an operation instruction with a custom function implemented according to the encoding method and encapsulation form of the operation instruction;

[0300] The described build-in operation instructions include the proprietary operation instructions included in the deep learning programming library interface;

[0301] The described offline operation instructions include pre-compiled dynamic operation instructions, where the pre-compilation result is saved in the offline cache.

[0302] Clause A7. According to the method described in Clause A6, the process of generating the binary code corresponding to the customized operation instructions includes:

[0303] According to the interface and data structure definitions of the operation instructions, encapsulate the user instructions corresponding to the operation instructions to obtain the encapsulated user instructions;

[0304] Compile the encapsulated user instructions to obtain a compilation result;

[0305] Insert the compilation result into the static operation pool in a dynamic link or static link manner to obtain the binary code corresponding to the customized operation instructions.

[0306] Clause A8. According to the method described in Clause A6, the static operation pool includes static operation source files, and the static operation source files include a static code segment, a static data segment, a dynamic code segment, and a dynamic data segment; where

[0307] The static code segment is used to save the binary code corresponding to the build-in operation instructions;

[0308] The dynamic code segment is used to save the binary code corresponding to the customized operation instructions;

[0309] The static data segment is used to save the tensor data corresponding to the build-in operation instructions;

[0310] The dynamic data segment is used to save the tensor data corresponding to the customized operation instructions.

[0311] Clause A9. According to the method described in Clause A8, the static operation pool also includes an offline cache, and the offline cache includes an offline file and an index table; where

[0312] The offline file is used to save the pre-compilation result of the offline operation instructions;

[0313] The index table is used to indicate the position of the pre-compilation result of the offline operation instructions in the offline file.

[0314] Clause A10. According to the method described in Clause A6, searching for the binary code corresponding to the name in the static operation pool according to the name of the static operation instructions includes:

[0315] In the static operation pool, sequentially search for the binary code corresponding to the custom operation, the binary code corresponding to the build-in operation, and the binary code corresponding to the offline operation according to the name specified when the static operation instruction is created, to obtain the binary code corresponding to the name.

[0316] Clause A11. According to the method described in Clause A3, for determining the instruction type of the operation instruction and performing a compilation operation corresponding to the instruction type according to the determination result to obtain the binary code of the deep learning algorithm, it further includes:

[0317] When the instruction type is a dynamic operation instruction, perform real-time compilation on the dynamic operation instruction to obtain a real-time compilation result as the binary code of the deep learning algorithm.

[0318] Clause A12. According to the method described in any one of Clauses A1 to A11, for performing real-time compilation on the dynamic operation instruction when the operation instruction includes a dynamic operation instruction to obtain a real-time compilation result as the binary code of the deep learning algorithm, it includes:

[0319] Obtain the original computation graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction;

[0320] Perform collaborative processing for the deep learning algorithm according to the original computation graph and original model data to obtain a first computation graph and first model data;

[0321] Generate hardware instructions and data descriptors according to the first computation graph;

[0322] Process the first model data for the hardware platform according to the data descriptor to obtain second model data;

[0323] Obtain the binary code of the deep learning algorithm according to the hardware instructions and the second model data.

[0324] Clause A13. According to the method described in Clause A12, for obtaining the original computation graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction, it includes:

[0325] Obtain the original computation graph corresponding to the dynamic operation instruction by parsing the dynamic operation instruction;

[0326] Obtain the original model data according to the parameters of the dynamic operation instruction.

[0327] Clause A14. For the method described in Clause A12, for the original computational graph and original model data according to the deep learning algorithm, collaborative processing for the deep learning algorithm is performed to obtain a first computational graph and first model data, including:

[0328] Read the original computational graph and the original model data;

[0329] Identify consecutive linear transformation operation nodes within the original computational graph;

[0330] Process the consecutive linear transformation operation nodes through linear transformation and constant folding to obtain a first computational graph and first model data.

[0331] Clause A15. For the method described in Clause A12, for generating hardware instructions and data descriptors according to the first computational graph, including:

[0332] Process the first computational graph according to a cost model, and combine with a heuristic search strategy to obtain hardware instructions and data descriptors.

[0333] Clause A16. For the method described in Clause A12, for processing the first model data for a hardware platform according to the data descriptor to obtain second model data, including:

[0334] According to the data descriptor, perform data alignment on the first model data according to the arithmetic requirements within the hardware platform; or,

[0335] According to the data descriptor, perform dimensional transformation on the first model data according to the arithmetic requirements within the hardware platform; or,

[0336] According to the data descriptor, perform precision selection on the first model data according to the arithmetic requirements within the hardware platform.

[0337] Clause A17. For the method described in Clause A12, for obtaining the binary code of the deep learning algorithm according to the hardware instructions and the second model data, including:

[0338] Package the hardware instructions and the second model data to obtain the binary code of the deep learning algorithm.

[0339] Clause A18. For the method described in Clause A11, the dynamic operation instructions include specialized operation instructions and / or fused operation instructions; where

[0340] The specialized operation instructions include operation instructions obtained by binding input parameters to the operation instructions and then converting;

[0341] The fusion operation instruction includes an operation instruction obtained by combining a plurality of the operation instructions according to the call order.

[0342] Clause A19. According to the method described in Clause A18, the specialization operation instruction includes a full specialization operation instruction, a partial specialization operation instruction, and a pseudo-specialization operation instruction; wherein,

[0343] The full specialization operation instruction includes an operation instruction obtained by binding all input parameters to the operation instruction and then converting it;

[0344] The partial specialization operation instruction includes an operation instruction obtained by binding N input parameters to the operation instruction and then converting it, where N is a positive integer less than the number of input parameters of the operation instruction;

[0345] The pseudo-specialization operation instruction includes an operation instruction obtained by directly converting the operation instruction without binding input parameters.

[0346] Clause A20. According to the method described in Clause A18, the creation process of the fusion operation instruction includes:

[0347] Create the name of the fusion operation instruction;

[0348] Determine the fusion operation sub-instructions according to the operation instructions to be fused;

[0349] Determine the operation connection relationship between the fusion operation sub-instructions according to the call order of the operation instructions to be fused;

[0350] Connect the fusion operation sub-instructions according to the operation connection relationship to obtain a connection result;

[0351] Set the input parameters and output parameters of the fusion operation instruction according to the user instruction corresponding to the fusion operation instruction;

[0352] Package the name, connection result, input parameters, and output parameters to obtain the fusion operation instruction.

[0353] Clause A21. According to the method described in Clause A1, the method further includes: through a runtime system, executing the binary code of the deep learning algorithm in a deep learning processor.

[0354] Clause A22. According to the method described in Clause A1, the operation data further includes tensor data, wherein the tensor data includes a shape attribute, a logical data type attribute, a physical data type attribute, and a physical layout attribute.

[0355] Clause A23. According to the method described in Clause A22, the method further includes:

[0356] Perform class extension on the deep learning framework, encapsulate the tensor data and the operation instructions within the data of the deep learning framework, and implement the integration of the deep learning framework and the interface of the deep learning programming library.

[0357] Clause A24. A compilation device for a deep learning algorithm, comprising:

[0358] An operation data receiving module, configured to receive operation data transmitted by an interface of a deep learning programming library;

[0359] An operation instruction obtaining module, configured to obtain the operation instructions included in the operation data;

[0360] A compilation module, configured to determine the instruction type of the operation instructions, and perform a compilation operation corresponding to the instruction type according to the determination result to obtain the binary code of the deep learning algorithm.

[0361] Clause A25. The device according to Clause A24, wherein the operation instructions are created or called according to user instructions received by the interface of the deep learning programming library.

[0362] Clause A26. The device according to Clause A25, wherein the compilation module comprises:

[0363] A determination unit, configured to determine the instruction type of the operation instructions;

[0364] A static search unit, configured to, when the instruction type is a static operation instruction, search for a corresponding binary code in a static operation pool according to the name of the static operation instruction, and use the binary code as the binary code of the deep learning algorithm.

[0365] Clause A27. The device according to Clause A26, wherein the static search unit is configured to:

[0366] Search for a binary code corresponding to the name in the static operation pool according to the name of the static operation instruction;

[0367] When the search result is successful, return the binary code as the binary code of the deep learning algorithm.

[0368] Clause A28. The device according to Clause A27, wherein the static search unit is further configured to:

[0369] When the search result is failed, regard the static operation instruction as a dynamic operation instruction and perform real-time compilation.

[0370] Clause A29. The device according to Clause A26, wherein the static operation instructions include one or more of customized operation instructions, build-in operation instructions, and offline operation instructions; wherein,

[0371] The customized operation instruction includes an operation instruction with a custom function implemented according to the encoding method and encapsulation form of the operation instruction.

[0372] The build-in operation instruction includes the proprietary operation instructions included in the deep learning programming library interface.

[0373] The offline operation instruction includes pre-compiled dynamic operation instructions, where the pre-compilation result is saved in the offline cache.

[0374] Article A30. For the device according to Article A29, the process of generating the binary code corresponding to the customized operation instruction includes:

[0375] Encapsulate the user instruction corresponding to the operation instruction according to the interface and data structure definition of the operation instruction to obtain the encapsulated user instruction.

[0376] Compile the encapsulated user instruction to obtain the compilation result.

[0377] Insert the compilation result into the static operation pool in a dynamic link or static link manner to obtain the binary code corresponding to the customized operation instruction.

[0378] Article A31. For the device according to Article A29, the static operation pool includes static operation source files, and the static operation source files include a static code segment, a static data segment, a dynamic code segment, and a dynamic data segment; where

[0379] The static code segment is used to save the binary code corresponding to the build-in operation instruction.

[0380] The dynamic code segment is used to save the binary code corresponding to the customized operation instruction.

[0381] The static data segment is used to save the tensor data corresponding to the build-in operation instruction.

[0382] The dynamic data segment is used to save the tensor data corresponding to the customized operation instruction.

[0383] Article A32. For the device according to Article A31, the static operation pool further includes an offline cache, and the offline cache includes an offline file and an index table; where

[0384] The offline file is used to save the pre-compilation result of the offline operation instruction.

[0385] The index table is used to indicate the position of the pre-compilation result of the offline operation instruction in the offline file.

[0386] Clause A33. For the device according to Clause A29, the static lookup unit is further configured to:

[0387] According to the name specified when the static operation instruction is created, sequentially look up the binary code corresponding to the custom operation, the binary code corresponding to the build-in operation, and the binary code corresponding to the offline operation in the static operation pool to obtain the binary code corresponding to the name.

[0388] Clause A34. For the device according to Clause A26, the compilation module further includes a dynamic compilation unit, which is configured to:

[0389] When the instruction type is a dynamic operation instruction, perform real-time compilation on the dynamic operation instruction to obtain a real-time compilation result, which serves as the binary code of the deep learning algorithm.

[0390] Clause A35. For the device according to any one of Clauses A24 to A34, the dynamic compilation unit includes:

[0391] An original data acquisition subunit, which is configured to obtain an original computational graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction;

[0392] A collaborative processing subunit, which is configured to perform collaborative processing for the deep learning algorithm according to the original computational graph and the original model data to obtain a first computational graph and first model data;

[0393] A hardware instruction generation subunit, which is configured to generate hardware instructions and data descriptors according to the first computational graph;

[0394] A model data processing subunit, which is configured to process the first model data for the hardware platform according to the data descriptor to obtain second model data;

[0395] A binary code generation subunit, which is configured to obtain the binary code of the deep learning algorithm according to the hardware instructions and the second model data.

[0396] Clause A36. For the device according to Clause A35, the original data acquisition subunit is configured to:

[0397] By parsing the dynamic operation instruction, obtain the original computational graph corresponding to the dynamic operation instruction;

[0398] According to the parameters of the dynamic operation instruction, obtain the original model data.

[0399] Clause A37. For the device according to Clause A35, the co - processing subunit is used for:

[0400] Read the original computation graph and the original model data;

[0401] Identify the consecutive linear transformation operation nodes within the original computation graph;

[0402] Process the consecutive linear transformation operation nodes through linear transformation and constant folding to obtain a first computation graph and first model data.

[0403] Clause A38. For the device according to Clause A35, the hardware instruction generation subunit is used for:

[0404] Process the first computation graph according to a cost model, and combine with a heuristic search strategy to obtain hardware instructions and data descriptors.

[0405] Clause A39. For the device according to Clause A35, the model data processing subunit is used for:

[0406] Align the first model data according to the operation requirements within the hardware platform according to the data descriptor; or,

[0407] Perform dimensional transformation on the first model data according to the operation requirements within the hardware platform according to the data descriptor; or,

[0408] Select the precision of the first model data according to the operation requirements within the hardware platform according to the data descriptor.

[0409] Clause A40. For the device according to Clause A35, the binary code generation subunit is used for:

[0410] Package the hardware instructions and the second model data to obtain the binary code of the deep - learning algorithm.

[0411] Clause A41. For the device according to Clause A34, the dynamic operation instructions include specialized operation instructions and / or fused operation instructions; wherein,

[0412] The specialized operation instructions include the operation instructions obtained by converting the operation instructions after binding input parameters;

[0413] The fused operation instructions include the operation instructions obtained by combining multiple operation instructions according to the call order.

[0414] Clause A42. For the device according to Clause A41, the specialized operation instructions include full - specialization operation instructions, partial - specialization operation instructions, and pseudo - specialization operation instructions; wherein,

[0415] The fully specialized operation instruction includes the operation instruction obtained by binding all input parameters to the operation instruction and then converting it;

[0416] The partially specialized operation instruction includes the operation instruction obtained by binding N input parameters to the operation instruction and then converting it, where N is a positive integer less than the number of input parameters of the operation instruction;

[0417] The pseudo-specialized operation instruction includes the operation instruction obtained by directly converting the operation instruction without binding input parameters.

[0418] Clause A43. For the device described in Clause A41, the creation process of the fused operation instruction includes:

[0419] Create the name of the fused operation instruction;

[0420] Determine the fused operation sub-instructions according to the operation instructions to be fused;

[0421] Determine the operation connection relationship between the fused operation sub-instructions according to the call order of the operation instructions to be fused;

[0422] Connect the fused operation sub-instructions according to the operation connection relationship to obtain a connection result;

[0423] Set the input parameters and output parameters of the fused operation instruction according to the user instruction corresponding to the fused operation instruction;

[0424] Package the name, connection result, input parameters and output parameters to obtain the fused operation instruction.

[0425] Clause A44. For the device described in Clause A24, the device further includes an execution module for: executing the binary code of the deep learning algorithm in the deep learning processor through the runtime system.

[0426] Clause A45. For the device described in Clause A24, the operation data further includes tensor data, where the tensor data includes a shape attribute, a logical data type attribute, a physical data type attribute and a physical layout attribute.

[0427] Clause A46. For the device described in Clause A45, the device is further used for:

[0428] Perform class extension on the deep learning framework, encapsulate the tensor data and the operation instruction in the data of the deep learning framework, and implement the integration of the deep learning framework and the deep learning programming library interface.

[0429] Clause A47. A deep learning computing device, the deep learning computing device includes one or more compilation devices for deep learning algorithms as described in any one of Clauses A24 - A46, and the deep learning computing device is used to complete set deep learning computations.

[0430] Clause A48. A combined computing device, the combined computing device includes one or more deep learning computing devices as described in any one of Clause A47, a general interconnection interface, and other processing devices;

[0431] The deep learning computing device interacts with the other processing devices to jointly complete the calculation operations specified by the user.

[0432] Clause A49. A deep learning chip, the deep learning chip includes:

[0433] A compilation device for deep learning algorithms as described in any one of Clauses A24 - 46; or,

[0434] A deep learning computing device as described in Clause A47; or,

[0435] A combined computing device as described in Clause A48.

[0436] Clause A50. An electronic device, the electronic device includes:

[0437] A compilation device for deep learning algorithms as described in any one of Clauses A24 - A46; or,

[0438] A deep learning computing device as described in Clause A47; or,

[0439] A combined computing device as described in Clause A48; or,

[0440] A deep learning chip as described in Clause A49.

[0441] The above has introduced the embodiments of the present disclosure in detail. Specific examples are used in this article to elaborate on the principles and implementation manners of the present disclosure. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present disclosure. At the same time, those skilled in the art, based on the idea of the present disclosure, the changes or deformations made in terms of the specific implementation manners and application scope of the present disclosure all fall within the protection scope of the present disclosure. In summary, the content of this specification should not be construed as a limitation to the present disclosure.

Claims

1. A compilation method for a deep learning algorithm, characterized in that, The method includes: Receiving operation data passed by the deep learning programming library interface; Obtaining the operation instructions included in the operation data; Judging the instruction type of the operation instruction, and performing a compilation operation corresponding to the instruction type according to the judgment result to obtain the binary code of the deep learning algorithm; Among them, the judging the instruction type of the operation instruction, and performing a compilation operation corresponding to the instruction type according to the judgment result to obtain the binary code of the deep learning algorithm includes: Judging the instruction type of the operation instruction; When the instruction type is a static operation instruction, looking up the corresponding binary code in the static operation pool according to the name of the static operation instruction as the binary code of the deep learning algorithm; Among them, the judging the instruction type of the operation instruction, and performing a compilation operation corresponding to the instruction type according to the judgment result to obtain the binary code of the deep learning algorithm further includes: When the instruction type is a dynamic operation instruction, performing real-time compilation on the dynamic operation instruction to obtain a real-time compilation result as the binary code of the deep learning algorithm; Among them, when the operation instruction includes a dynamic operation instruction, performing real-time compilation on the dynamic operation instruction to obtain a real-time compilation result as the binary code of the deep learning algorithm includes: Obtaining the original computation graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction; Performing collaborative processing for the deep learning algorithm according to the original computation graph and original model data to obtain a first computation graph and first model data; Generating hardware instructions and data descriptors according to the first computation graph; Processing the first model data for the hardware platform according to the data descriptor to obtain a second model data; Obtaining the binary code of the deep learning algorithm according to the hardware instructions and the second model data.

2. The method according to claim 1, wherein The operation data is created or called according to the user instruction received by the deep learning programming library interface.

3. The method according to claim 1, characterized in that The looking up the corresponding binary code in the static operation pool according to the name of the static operation instruction as the binary code of the deep learning algorithm when the instruction type is a static operation instruction includes: Looking up the binary code corresponding to the name in the static operation pool according to the name of the static operation instruction; When the lookup result is successful, returning the binary code as the binary code of the deep learning algorithm.

4. The method according to claim 3, wherein The looking up the corresponding binary code in the static operation pool according to the name of the static operation instruction as the binary code of the deep learning algorithm when the instruction type is a static operation instruction further includes: When the lookup result is failed, regarding the static operation instruction as a dynamic operation instruction and performing real-time compilation.

5. The method according to claim 1, characterized in that, The static operation instruction includes one or more of a customized operation instruction, a build-in operation instruction, and an offline operation instruction; among them, The customized operation instruction includes an operation instruction with a custom function implemented according to the encoding method and encapsulation form of the operation instruction. The described build-in operation instructions include the proprietary operation instructions included in the deep learning programming library interface; The described offline operation instructions include the pre-compiled dynamic operation instructions, where the pre-compiled result is stored in the offline cache.

6. The method according to claim 5, wherein The generation process of the binary code corresponding to the described customized operation instructions includes: According to the interface and data structure definition of the operation instructions, encapsulate the user instructions corresponding to the operation instructions to obtain the encapsulated user instructions; Compile the encapsulated user instructions to obtain a compilation result; Insert the compilation result into the static operation pool in a dynamic link or static link manner to obtain the binary code corresponding to the customized operation instructions.

7. The method according to claim 5, wherein The static operation pool includes static operation source files, and the static operation source files include a static code segment, a static data segment, a dynamic code segment, and a dynamic data segment; where The static code segment is used to store the binary code corresponding to the build-in operation instructions; The dynamic code segment is used to store the binary code corresponding to the customized operation instructions; The static data segment is used to store the tensor data corresponding to the build-in operation instructions; The dynamic data segment is used to store the tensor data corresponding to the customized operation instructions.

8. The method according to claim 7, wherein The static operation pool also includes an offline cache, and the offline cache includes an offline file and an index table; where The offline file is used to store the pre-compiled result of the offline operation instructions; The index table is used to indicate the position of the pre-compiled result of the offline operation instructions in the offline file.

9. The method according to claim 5, wherein The searching for the binary code corresponding to the name in the static operation pool according to the name of the static operation instructions includes: According to the name specified when the static operation instruction is created, sequentially search for the binary code corresponding to the customized operation, the binary code corresponding to the build-in operation, and the binary code corresponding to the offline operation in the static operation pool to obtain the binary code corresponding to the name.

10. The method according to claim 1, characterized in that, The obtaining of the original computation graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction includes: By parsing the dynamic operation instruction, obtain the original computation graph corresponding to the dynamic operation instruction; According to the parameters of the dynamic operation instruction, obtain the original model data.

11. The method according to claim 1, characterized in that, The collaborative processing for the deep learning algorithm according to the original computation graph and original model data of the deep learning algorithm to obtain a first computation graph and first model data includes: Read the original computation graph and the original model data; Identify the continuous linear transformation operation nodes in the original computation graph; Process the continuous linear transformation operation nodes through linear transformation and constant folding to obtain a first computation graph and first model data.

12. The method according to claim 1, characterized in that, The generation of hardware instructions and data descriptors according to the first computation graph includes: Process the first computation graph according to the cost model, and combine with the heuristic search strategy to obtain hardware instructions and data descriptors.

13. The method according to claim 1, characterized in that Performing processing of the first model data for the hardware platform according to the data descriptor, to obtain second model data, includes: According to the data descriptor, aligning the first model data according to the operation requirements within the hardware platform; or, According to the data descriptor, performing dimensional transformation on the first model data according to the operation requirements within the hardware platform; or, According to the data descriptor, selecting the precision of the first model data according to the operation requirements within the hardware platform.

14. The method according to claim 1, wherein Obtaining the binary code of the deep learning algorithm according to the hardware instruction and the second model data, includes: Packaging the hardware instruction and the second model data to obtain the binary code of the deep learning algorithm.

15. The method according to claim 1, wherein The dynamic operation instruction includes a specialization operation instruction and / or a fusion operation instruction; wherein, The specialization operation instruction includes an operation instruction obtained by binding input parameters to the operation instruction and then converting; The fusion operation instruction includes an operation instruction obtained by combining multiple operation instructions according to the call order.

16. The method according to claim 15, characterized in that, The specialization operation instruction includes a full specialization operation instruction, a partial specialization operation instruction, and a pseudo-specialization operation instruction; wherein, The full specialization operation instruction includes an operation instruction obtained by binding all input parameters to the operation instruction and then converting; The partial specialization operation instruction includes an operation instruction obtained by binding N input parameters to the operation instruction and then converting, where N is a positive integer less than the number of input parameters of the operation instruction; The pseudo-specialization operation instruction includes an operation instruction obtained by directly converting the operation instruction without binding input parameters.

17. The method according to claim 15, wherein The creation process of the fusion operation instruction includes: Creating the name of the fusion operation instruction; Determining the fusion operation sub-instructions according to the operation instructions to be fused; Determining the operation connection relationship between the fusion operation sub-instructions according to the call order of the operation instructions to be fused; Connecting the fusion operation sub-instructions according to the operation connection relationship to obtain a connection result; Setting the input parameters and output parameters of the fusion operation instruction according to the user instruction corresponding to the fusion operation instruction; Packaging the name, connection result, input parameters, and output parameters to obtain the fusion operation instruction.

18. The method according to claim 1, wherein The method further includes: executing the binary code of the deep learning algorithm in the deep learning processor through a runtime system.

19. The method according to claim 1, wherein The operation data further includes tensor data, wherein the tensor data includes a shape attribute, a logical data type attribute, a physical data type attribute, and a physical layout attribute.

20. The method according to claim 19, wherein The method further includes: Performing class extension on the deep learning framework, encapsulating the tensor data and the operation instruction within the data of the deep learning framework, to implement the integration of the deep learning framework and the deep learning programming library interface.

21. A compilation device for a deep learning algorithm, characterized in that, Includes: An operation data receiving module, configured to receive operation data transmitted by the deep learning programming library interface; An operation instruction obtaining module, configured to obtain the operation instruction included in the operation data; A compilation module, which is used to determine the instruction type of the operation instruction, and perform a compilation operation corresponding to the instruction type according to the determination result to obtain the binary code of the deep learning algorithm; Among them, the compilation module includes: A judgment unit, which is used to judge the instruction type of the operation instruction; A static lookup unit, which is used to, when the instruction type is a static operation instruction, look up the corresponding binary code in the static operation pool according to the name of the static operation instruction as the binary code of the deep learning algorithm; Among them, the compilation module further includes a dynamic compilation unit, which is used to: When the instruction type is a dynamic operation instruction, perform real-time compilation on the dynamic operation instruction to obtain a real-time compilation result as the binary code of the deep learning algorithm; Among them, the dynamic compilation unit includes: An original data acquisition subunit, which is used to obtain an original computational graph and original model data corresponding to the dynamic operation instruction according to the dynamic operation instruction; A collaborative processing subunit, which is used to perform collaborative processing for the deep learning algorithm according to the original computational graph and the original model data to obtain a first computational graph and first model data; A hardware instruction generation subunit, which is used to generate hardware instructions and data descriptors according to the first computational graph; A model data processing subunit, which is used to process the first model data for the hardware platform according to the data descriptor to obtain second model data; A binary code generation subunit, which is used to obtain the binary code of the deep learning algorithm according to the hardware instructions and the second model data.

22. The device according to claim 21, characterized in that, The operation instruction is created or called according to the user instruction received by the deep learning programming library interface.

23. The device according to claim 21, wherein The static lookup unit is used to: Look up the binary code corresponding to the name in the static operation pool according to the name of the static operation instruction; When the lookup result is successful, return the binary code as the binary code of the deep learning algorithm.

24. The device according to claim 23, wherein, The static lookup unit is further used to: When the lookup result is failed, take the static operation instruction as a dynamic operation instruction and perform real-time compilation.

25. The device according to claim 21, characterized in that, The static operation instruction includes one or more of a customized operation instruction, a build-in operation instruction, and an offline operation instruction; among them, The customized operation instruction includes an operation instruction with a custom function implemented according to the encoding method and encapsulation form of the operation instruction; The build-in operation instruction includes the own operation instructions included in the deep learning programming library interface; The offline operation instruction includes a dynamically compiled operation instruction that has been pre-compiled, and the pre-compiled result is stored in the offline cache.

26. The device according to claim 25, characterized in that, The generation process of the binary code corresponding to the customized operation instruction includes: Encapsulate the user instruction corresponding to the operation instruction according to the interface and data structure definition of the operation instruction to obtain the encapsulated user instruction; Compile the encapsulated user instruction to obtain a compilation result; Insert the compilation result into the static operation pool in a dynamic link or static link manner to obtain the binary code corresponding to the customized operation instruction.

27. The device according to claim 25, characterized in that, The static operation pool includes a static operation source file, which includes a static code segment, a static data segment, a dynamic code segment, and a dynamic data segment; wherein, The static code segment is used to store the binary code corresponding to the build-in operation instructions; The dynamic code segment is used to store the binary code corresponding to the customized operation instructions; The static data segment is used to store the tensor data corresponding to the build-in operation instructions; The dynamic data segment is used to store the tensor data corresponding to the customized operation instructions.

28. The device according to claim 27, characterized in that, The static operation pool further includes an offline cache, which includes an offline file and an index table; wherein, The offline file is used to store the pre-compiled results of the offline operation instructions; The index table is used to indicate the location of the pre-compiled results of the offline operation instructions in the offline file.

29. The device according to claim 25, characterized in that The static lookup unit is further used for: According to the name specified when the static operation instruction is created, sequentially look up the binary code corresponding to the customized operation, the binary code corresponding to the build-in operation, and the binary code corresponding to the offline operation in the static operation pool to obtain the binary code corresponding to the name.

30. The device according to claim 29, wherein, The raw data acquisition subunit is used for: By parsing the dynamic operation instruction, obtain the raw computation graph corresponding to the dynamic operation instruction; According to the parameters of the dynamic operation instruction, obtain the raw model data.

31. The device according to claim 29, wherein The collaborative processing subunit is used for: Read the raw computation graph and the raw model data; Identify the continuous linear transformation operation nodes in the raw computation graph; Process the continuous linear transformation operation nodes through linear transformation and constant folding to obtain a first computation graph and first model data.

32. The device according to claim 29, characterized in that, The hardware instruction generation subunit is used for: Process the first computation graph according to the cost model, and combine with the heuristic search strategy to obtain hardware instructions and data descriptors.

33. The apparatus according to claim 29, wherein, The model data processing subunit is used for: According to the data descriptor, align the first model data according to the operation requirements within the hardware platform; or, According to the data descriptor, perform dimensional transformation on the first model data according to the operation requirements within the hardware platform; or, According to the data descriptor, perform precision selection on the first model data according to the operation requirements within the hardware platform.

34. The device according to claim 29, characterized in that, The binary code generation subunit is used for: Package the hardware instructions and the second model data to obtain the binary code of the deep learning algorithm.

35. The device according to claim 28, wherein The dynamic operation instructions include specialized operation instructions and / or fused operation instructions; wherein, The specialized operation instructions include the operation instructions obtained by converting after binding input parameters to the operation instructions; The fused operation instructions include the operation instructions obtained by combining multiple operation instructions according to the call order.

36. The device according to claim 35, wherein The specialized operation instructions include fully specialized operation instructions, partially specialized operation instructions, and pseudo-specialized operation instructions; wherein, The fully specialized operation instructions include the operation instructions obtained by converting after binding all input parameters to the operation instructions; The partial specialization operation instructions include the operation instructions obtained by binding N input parameters to the operation instructions and then converting them, where N is a positive integer less than the number of input parameters of the operation instructions; The pseudo-specialization operation instructions include the operation instructions directly obtained by not binding input parameters to the operation instructions and then converting them.

37. The device according to claim 36, characterized in that, The creation process of the fusion operation instructions includes: Creating the name of the fusion operation instructions; Determining the fusion operation sub-instructions according to the operation instructions to be fused; Determining the operation connection relationship between the fusion operation sub-instructions according to the call order of the operation instructions to be fused; Connecting the fusion operation sub-instructions according to the operation connection relationship to obtain a connection result; Setting the input parameters and output parameters of the fusion operation instructions according to the user instructions corresponding to the fusion operation instructions; Packaging the name, connection result, input parameters and output parameters to obtain the fusion operation instructions.

38. The device according to claim 21, wherein The device further includes an execution module for executing the binary code of the deep learning algorithm in the deep learning processor through a runtime system.

39. The device according to claim 21, characterized in that, The operation data further includes tensor data, where the tensor data includes a shape attribute, a logical data type attribute, a physical data type attribute, and a physical layout attribute.

40. The device according to claim 39, characterized in that, The device is further used for: Performing class extension on the deep learning framework, encapsulating the tensor data and the operation instructions in the data of the deep learning framework, and implementing the integration of the deep learning framework and the deep learning programming library interface.

41. A deep learning computing device, characterized in that, The deep learning operation device includes one or more compilation devices of the deep learning algorithms according to any one of claims 21-40, and the deep learning operation device is used to complete the set deep learning operations.

42. A combined operation device, characterized in that, The combined operation device includes one or more deep learning operation devices according to any one of claim 41, a general interconnection interface, and other processing devices; The deep learning operation device interacts with the other processing devices to jointly complete the calculation operations specified by the user.

43. A deep learning chip, characterized in that, The deep learning chip includes: The compilation device of the deep learning algorithm according to any one of claims 21-40; or, The deep learning operation device according to claim 41; or, The combined operation device according to claim 42.

44. An electronic device, characterized in that, The electronic device includes: The compilation device of the deep learning algorithm according to any one of claims 21-40; or, The deep learning operation device according to claim 41; or, The combined operation device according to claim 42; or, The deep learning chip according to claim 43.

Citation Information

Patent Citations

  • Method and system for instruction-set architecture simulation using just in time compilation

    US20030217248A1