Deep learning machine instruction generation method and device supporting multiple types of backend computing hardware

By optimizing the compilation process of the deep learning framework, segmenting the calculation graph and generating corresponding operator codes, the cumbersome model adaptation problems caused by insufficient hardware support capabilities in the existing technology are solved, and better compatibility and execution efficiency are achieved on a variety of back-end hardware.

WO2025123654A1PCT designated stage expired Publication Date: 2025-06-19SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Patent Information

Application Number
PCT/CN2024/103553
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-07-04
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing deep learning framework assumes that the support capabilities of NVIDIA hardware are the most complete at the compilation level, resulting in the adaptation and change of deep learning models in the case of insufficient hardware support capabilities, and it is impossible to effectively support multiple back-end computing hardware.

Method used

By obtaining the deep learning model program and converting it into a calculation chart, comparing the calculation chart with the target hardware support information, and determining whether there is an unsupported operator. If there is, the calculation graph is segmented, a secondary subgraph is generated, and the corresponding operator code is generated according to the subgraph type, and the final connection is made to generate a complete set of machine instructions.

Benefits of technology

It achieves better compatibility and execution efficiency under the support of multiple back-end hardware, and is suitable for a variety of application scenarios, avoiding the cumbersome model adaptation caused by insufficient hardware support capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024103553_19062025_PF_FP_ABST
    Figure CN2024103553_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a deep learning machine instruction generation method and device supporting multiple types of backend computing hardware. The method comprises the following steps: obtaining a deep learning model program, and converting the deep learning model program into a series of computational graphs; comparing each computational graph with target hardware support information, determining whether an operator which is not supported by the target hardware exists in the computational graph, if yes, partitioning the current computational graph to obtain second-level sub-graphs, and if not, directly using the current computational graph as a second-level sub-graph, wherein a corresponding sub-graph type is marked on each second-level sub-graph; generating a corresponding operator code on the basis of the sub-graph type marked on each second-level sub-graph; and connecting the operator codes to generate a complete machine instruction set. Compared with the prior art, the present invention has the advantages of being compatible with various hardware scenarios and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Deep learning machine instruction generation method and device supporting multiple back-end computing hardware Technical Field

[0001] The present invention relates to the field of deep learning framework processing technology, and in particular to a method and device for generating deep learning machine instructions that supports multiple back-end computing hardware. Background Art

[0002] The compilation process of mainstream deep learning frameworks such as PyTorch and TensorFlow is primarily divided into two phases: graph capture and operator generation. Specifically, user-written deep learning programs are converted and captured into a computation graph, where nodes represent basic deep learning operators, such as multiplication and convolution, that perform specific computational functions. Further optimizations and transformations are then performed on the computation graph, and the operator execution components are embedded into the computation graph, enabling the deep learning model to run on the hardware. Typically, the operator execution components are implemented or supported by the hardware vendor. Hardware vendors such as NVIDIA and Suiyuan provide a pre-implemented operator library for direct user access, as well as hardware-specific compilers that compile user-written operator implementation code into hardware-executable programs. However, the operators included in the operator libraries provided by different hardware vendors vary, and some vendors' compilers may even be unable to successfully compile user-written operator programs. Operator support by these hardware vendors is typically grouped together as hardware support capabilities, encompassing both operator library support and hardware compiler support.

[0003] Because NVIDIA provides the most comprehensive hardware support, covering all current operators, all mainstream deep learning frameworks in the industry use NVIDIA as the default backend. This largely disregards insufficient hardware support, assuming that operators in the captured computational graph are supported by the hardware. However, this assumption doesn't always hold true in reality. Many deep learning computing hardware produced by small and medium-sized manufacturers lacks comprehensive support, often failing to support certain operators. This often requires custom adaptation and modification of these mainstream deep learning frameworks, a tedious task. Therefore, a holistic design and solution is needed to effectively address this situation.

[0004] Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a deep learning machine instruction generation method and device that is compatible with multiple hardware scenarios and supports multiple back-end computing hardware.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] A method for generating deep learning machine instructions supporting multiple back-end computing hardware includes the following steps:

[0008] Obtain a deep learning model program, and convert the deep learning model program into a series of computational graphs;

[0009] Compare each of the computation graphs with the target hardware support information to determine whether there are operators in the computation graph that are not supported by the target hardware. If so, split the current computation graph to obtain a second-level subgraph. If not, directly use the current computation graph as a second-level subgraph, and mark each second-level subgraph with a corresponding subgraph type.

[0010] Generate corresponding operator codes based on the subgraph types marked on each secondary subgraph;

[0011] Connect the operator codes to generate a complete set of machine instructions.

[0012] Furthermore, the deep learning model program is converted into the computational graph in a dynamic, static, or mixed dynamic and static form.

[0013] Furthermore, the target hardware support information includes a list for describing operators that are not supported by the target hardware.

[0014] Furthermore, the generation process of the secondary subgraph includes the following steps:

[0015] Traverse each node in the current computation graph, match the operator used by the node with the target hardware support information, determine whether the operator of the current node is an operator not supported by the target hardware, and split the consecutive nodes supported by the same hardware into the same secondary subgraph.

[0016] Furthermore, the process of generating the corresponding operator code includes:

[0017] Generate a corresponding processing path based on the subgraph type marked on each secondary subgraph;

[0018] Each secondary subgraph is distributed based on the processing path, and the corresponding backend hardware or CPU operator library is called to generate the operator code.

[0019] The present invention also provides a computer-readable storage medium comprising one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs include instructions for executing the deep learning machine instruction generation method that supports multiple back-end computing hardware as described above.

[0020] The present invention also provides a deep learning machine instruction generation device supporting multiple back-end computing hardware, comprising:

[0021] A graph splitter is used to obtain a series of computational graphs converted by the deep learning model program, compare each of the computational graphs with the target hardware support information, and determine whether there are operators in the computational graph that are not supported by the target hardware. If so, the current computational graph is split to obtain second-level subgraphs. If not, the current computational graph is directly used as a second-level subgraph, and each second-level subgraph is marked with the corresponding subgraph type;

[0022] The code generation scheduler is used to generate corresponding operator codes based on the subgraph type marked on each secondary subgraph, and connect the operator codes to generate a complete set of machine instructions.

[0023] Furthermore, the target hardware support information includes a list for describing operators that are not supported by the target hardware.

[0024] Furthermore, the generation process of the secondary subgraph includes the following steps:

[0025] Traverse each node in the current computation graph, match the operator used by the node with the target hardware support information, determine whether the operator of the current node is an operator not supported by the target hardware, and split the consecutive nodes supported by the same hardware into the same secondary subgraph.

[0026] Furthermore, the process of generating the corresponding operator code includes:

[0027] Generate a corresponding processing path based on the subgraph type marked on each secondary subgraph;

[0028] Each secondary subgraph is distributed based on the processing path, and the corresponding backend hardware or CPU operator library is called to generate the operator code.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] The present invention optimizes and improves the graph capture mechanism of the deep learning framework compilation process, takes into account the reality that the back-end hardware support capabilities vary, performs secondary segmentation of the computational graph according to the different back-end hardware capabilities, and distributes and schedules it to different hardware for execution, so that it can have better compatibility and execution efficiency with the support of multiple back-end hardware, and is suitable for a variety of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] FIG1 is a schematic diagram of the overall technical solution of the present invention;

[0032] FIG2 is a schematic diagram of the conversion from a high-level computational graph to a segmented secondary subgraph according to the present invention. DETAILED DESCRIPTION

[0033] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0034] The current deep learning framework compilation process first analyzes the user-written model program and converts it into a computational graph, either dynamically or statically. These computational graphs use operators as their basic units. Operator execution programs are then generated by directly calling the operator library or relying on the hardware compiler provided by the hardware manufacturer. In this process, some operators may not be supported by the relevant hardware manufacturers, requiring additional processing for these unsupported operators. The present invention is intended to address this technical problem.

[0035] In this embodiment, a method for generating deep learning machine instructions supporting multiple back-end computing hardware is provided, as shown in FIG1 , including the following steps:

[0036] S1. Obtain a deep learning model program and convert the deep learning model program into a series of computational graphs.

[0037] In specific implementations, the deep learning model program can be converted into the computation graph in a dynamic, static, or hybrid form. More specifically, translation optimization and conversion can be performed using a Just-In-Time (JIT) or AOT approach. The nodes in the computation graph are basic deep learning operators, such as multiplication and convolution, that perform specific computational functions.

[0038] S2. Compare each computation graph with the target hardware support information to determine whether there are any operators in the computation graph that are not supported by the target hardware. If so, split the current computation graph to obtain second-level subgraphs. If not, the current computation graph is directly used as a second-level subgraph. Each second-level subgraph is marked with a corresponding subgraph type. The subgraph type indicates whether the subgraph can be supported by the backend hardware or needs to be executed by an operator library that supports more operators.

[0039] In a specific embodiment, the target hardware support information includes a list describing operators that are not supported by the target hardware. This list can be considered a blacklist. When the graph partitioner performs a query comparison, it uses the operator name of the current node operator to match the operator names in this list. If a match is found, it indicates that the operator is not supported. This can concisely and clearly describe the limitations of a backend hardware. In this embodiment, the target hardware can be multiple hardware types, enabling multiple hardware types to work together.

[0040] In a specific implementation, the generation of the secondary subgraph is as follows: by traversing each node in the current computation graph, matching the operator used by the node with the target hardware support information, determining whether the operator of the current node is an operator not supported by the target hardware, and dividing the continuous nodes supported by the same hardware into the same secondary subgraph.

[0041] In this embodiment, the operator library of the default CPU version can support all operators, and the calculation graph is divided according to the hardware support capabilities of the CPU and the target hardware to achieve a balance between task completion and processing efficiency. Specifically, by traversing the nodes of the obtained calculation graph, the operators used by each node are checked and compared with the information of the target hardware support capabilities. Once it is found that the operator of the current node is not supported, the calculation graph is divided here. First, check whether the previous subgraph belongs to the same supported hardware. If the same, collect this node into the previous adjacent subgraph, which is marked with supported back-end hardware. If different, this unsupported node is separately divided into a secondary subgraph and marked as about to be distributed to the CPU for execution. Similarly, continuous nodes that cannot be supported by the back-end hardware are also divided into the same secondary subgraph, and then continue to process the remaining nodes backward. Repeat this process until all nodes are traversed.

[0042] Figure 2 shows the conversion from a high-level computational graph to a segmented second-level subgraph. In the figure, A, B, C, D, E, and F represent different operators. Operators B and C are not supported by the backend hardware and are marked as subgraph types processed by the CPU.

[0043] S3. Generate corresponding operator codes based on the subgraph types marked on each secondary subgraph.

[0044] Each secondary subgraph contains several operators, each of which can be implemented by the operator library of the supported hardware, or the entire secondary subgraph can be implemented by the corresponding hardware compiler to generate corresponding machine instructions for these operators in batches.

[0045] In a specific implementation, different subgraphs are assigned different processing paths based on their subgraph types. For example, subgraphs not supported by the hardware are directly processed as machine code snippets that call the CPU version of the operator library, while subgraphs supported by the backend are processed as code snippets that call the backend hardware version of the operator library or directly call the hardware compiler to dynamically generate operator code and call it. These different code snippets correspond to different secondary subgraphs at the machine code level.

[0046] S4. Connect the operator codes to generate a complete set of machine instructions.

[0047] In a specific implementation, the code snippets generated in step S3 are connected, and the execution flows of the machine codes corresponding to these different secondary subgraphs are connected in series, thereby finally generating a complete set of machine instructions corresponding to the computation graph.

[0048] If the above method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0049] In another embodiment, a deep learning machine instruction generation device that supports multiple back-end computing hardware is provided, including a graph splitter and a code generation scheduler, wherein the graph splitter related to the hardware capability is used to obtain a series of computational graphs converted by the deep learning model program, and the deep learning model program is converted into sub-graphs by a graph capturer that is unrelated to the hardware capability, and each of the computational graphs is compared with the target hardware support information to determine whether there are operators in the computational graph that are not supported by the target hardware. If so, the current computational graph is split to obtain a second-level sub-graph. If not, the current computational graph is directly used as a second-level sub-graph, and each second-level sub-graph is marked with a corresponding sub-graph type; the code generation scheduler is used to generate corresponding operator codes based on the sub-graph type marked on each second-level sub-graph, and connect each operator code to generate a complete set of machine instructions. The overall implementation equation of the above-mentioned device includes: dividing the received original calculation graph into two-level sub-graphs through a graph divider, each second-level sub-graph contains several operators, each operator can be implemented by the operator library of the supported hardware or the entire second-level sub-graph can generate corresponding machine instructions for these operators in batches by the corresponding hardware compiler; and distributing each second-level sub-graph to the corresponding hardware for corresponding processing based on the marking information on each second-level sub-graph through a code generation scheduler.

[0050] Based on the deep learning framework compilation process, the above-mentioned device organically embeds the hardware capability-related graph segmenter, hardware support information data and code generation scheduler into it, so as to achieve more efficient and reliable processing of deep learning model programs.

[0051] It will be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0052] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0053] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0054] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.

Claims

1. A method for generating deep learning machine instructions supporting multiple backend computing hardware, characterized in that: The following steps are involved: Obtain a deep learning model program, and convert the deep learning model program into a series of computational graphs; Compare each of the computation graphs with the target hardware support information to determine whether there are operators in the computation graph that are not supported by the target hardware. If so, segment the current computation graph to obtain a secondary subgraph. If not, the current computation graph is directly used as a secondary subgraph, and each secondary subgraph is marked with a corresponding subgraph type. Generate corresponding operator codes based on the subgraph types marked on each secondary subgraph; Connect the operator codes to generate a complete set of machine instructions.

2. The method for generating deep learning machine instructions supporting multiple backend computing hardware according to claim 1, characterized in that: The deep learning model program is converted into the computational graph in a dynamic, static or mixed dynamic and static form.

3. The method for generating deep learning machine instructions supporting multiple backend computing hardware according to claim 1, characterized in that: The target hardware support information includes a list for describing operators that are not supported by the target hardware.

4. The method for generating deep learning machine instructions supporting multiple backend computing hardware according to claim 1, characterized in that: The generation process of the secondary subgraph includes the following steps: Traverse each node in the current computation graph, match the operator used by the node with the target hardware support information, determine whether the operator of the current node is an operator not supported by the target hardware, and split the continuous nodes supported by the same hardware into the same secondary subgraph.

5. The method for generating deep learning machine instructions supporting multiple backend computing hardware according to claim 1, characterized in that: The process of generating the corresponding operator code includes: Generate a corresponding processing path based on the subgraph type marked on each secondary subgraph; Each secondary subgraph is distributed based on the processing path, and the corresponding backend hardware or CPU operator library is called to generate the operator code.

6. A computer-readable storage medium, characterized in that: It includes one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing a deep learning machine instruction generation method that supports multiple back-end computing hardware as described in any one of claims 1-5.

7. A deep learning machine instruction generation device supporting multiple back-end computing hardware, characterized in that: include: A graph splitter is used to obtain a series of computational graphs converted by a deep learning model program, compare each of the computational graphs with the target hardware support information, and determine whether there are operators in the computational graph that are not supported by the target hardware. If so, the current computational graph is split to obtain a second-level subgraph. If not, the current computational graph is directly used as a second-level subgraph, and each second-level subgraph is marked with a corresponding subgraph type; The code generation scheduler is used to generate corresponding operator codes based on the subgraph type marked on each secondary subgraph, and connect each operator code to generate a complete set of machine instructions.

8. The deep learning machine instruction generation device supporting multiple back-end computing hardware according to claim 7, characterized in that: The target hardware support information includes a list for describing operators that are not supported by the target hardware.

9. The deep learning machine instruction generation device supporting multiple back-end computing hardware according to claim 7, characterized in that: The generation process of the secondary subgraph includes the following steps: Traverse each node in the current computation graph, match the operator used by the node with the target hardware support information, determine whether the operator of the current node is an operator not supported by the target hardware, and split the continuous nodes supported by the same hardware into the same secondary subgraph.

10. The deep learning machine instruction generation device supporting multiple back-end computing hardware according to claim 7, characterized in that: The process of generating the corresponding operator code includes: Generate a corresponding processing path based on the subgraph type marked on each secondary subgraph; Each secondary subgraph is distributed based on the processing path, and the corresponding backend hardware or CPU operator library is called to generate the operator code.

Citation Information

Patent Citations

  • Online reasoning optimization method and device in deep learning and computer storage medium

    CN112825154A

  • Method, system and equipment for generating computational graph of deep learning compiler and medium

    CN113885845A

  • Generation method and device of operator, equipment and storage medium

    CN114911465A

  • Platform framework extension method and device and storage medium

    CN116360712A

  • Deep learning machine instruction generation method and device supporting multiple back-end computing hardware

    CN117892836A

Cited By

  • Deep learning model-oriented graph level compilation optimization method and deep learning model-oriented graph level compilation optimization equipment

    CN121209886A