Neural network model deployment method, device, electronic device and storage medium
By converting the calculation graph of the original neural network model into an intermediate calculation graph composed of the operator set supported by the target hardware device, and further transforming according to the hardware constraints, the problem that the neural network model cannot be directly deployed on the new hardware device is solved, and an efficient and unified deployment method is realized, improving the availability and ease of use of hardware devices.
Patent Information
- Application Number
- CN202210744352.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-06-27
AI Technical Summary
Existing neural network models for deep learning and brain-like computing cannot be directly deployed on new hardware devices because they use different algorithmic frameworks than those supported by hardware, resulting in the availability and ease of use of hardware devices.
A neural network model deployment method is proposed. By obtaining the calculation diagram of the original neural network model, it is converted into an intermediate calculation diagram composed of the operators supported by the target hardware device, and further converted into a target calculation diagram adapted to the hardware device according to the constraints of the hardware device, thereby generating a hardware executable file.
It realizes efficient and unified conversion of neural network models under different algorithm frameworks into hardware executable files, meeting the data accuracy, resource and operator set constraints of hardware devices, and improving the availability and ease of use of hardware devices.
Smart Images

Figure CN115099399B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a neural network model deployment method, device, electronic device and storage medium. Background Art
[0002] With the development of artificial intelligence technology and non-von Neumann hardware architecture, many new hardware devices have emerged in the academic and industrial fields, such as multi-core chips, many-core chips, deep learning accelerators, neuromorphic chips, brain-like computing chips, etc.
[0003] However, the algorithm framework of the neural network model created based on current deep learning, brain-like computing and other technologies may be different from the algorithm framework supported by the above-mentioned new hardware devices. The neural network models under different algorithm frameworks may not be directly deployed and executed in the above-mentioned hardware devices, which affects the availability and ease of use of the above-mentioned hardware devices. Summary of the invention
[0004] In view of this, the present disclosure proposes a neural network model deployment method, device, electronic device and storage medium to efficiently and uniformly convert original neural network models under different algorithm frameworks into hardware executable files that can be deployed on the above-mentioned hardware devices.
[0005] According to one aspect of the present disclosure, a neural network model deployment method is provided, including: obtaining an original computational graph corresponding to an original neural network model to be deployed on a hardware device; converting the original computational graph into an intermediate computational graph based on a target operator set, wherein the target operator set is an operator set supported by the hardware device; converting the intermediate computational graph into a target computational graph adapted to the hardware device according to hardware constraints corresponding to the hardware device, wherein the hardware constraints are determined based on data accuracy, hardware resources and at least one of the target operator set corresponding to the hardware device; and determining a hardware executable file based on the target computational graph, wherein the hardware executable file is used to be deployed on the hardware device.
[0006] In one possible implementation, the converting of the original computation graph into an intermediate computation graph based on a target operator set includes: replacing the original computation nodes in the original computation graph with the target operators in the target operator set to obtain intermediate computation nodes equivalent to the original computation nodes; generating intermediate storage nodes associated with the intermediate computation nodes according to data storage rules corresponding to the hardware device, wherein the data storage rules are used to indicate data storage methods supported by the hardware device; establishing directed edges between the intermediate computation nodes and the intermediate storage nodes according to data dependencies between the nodes in the original computation graph to obtain the intermediate computation graph, wherein the data dependencies characterize the association relationships between the nodes in the original computation graph.
[0007] In one possible implementation, the converting of the original computation graph into an intermediate computation graph based on a target operator set includes: replacing the original computation nodes in the original computation graph with the target operators in the target operator set to obtain intermediate computation nodes equivalent to the original computation nodes; generating intermediate storage nodes associated with the intermediate computation nodes according to data storage rules corresponding to the hardware device, wherein the data storage rules are used to indicate data storage methods supported by the hardware device; establishing directed edges between the intermediate computation nodes and the intermediate storage nodes according to data dependencies between the nodes in the original computation graph to obtain the intermediate computation graph, wherein the data dependencies characterize the association relationships between the nodes in the original computation graph.
[0008] In a possible implementation, the method of replacing the original computing node in the original computing graph with the target operator in the target operator set to obtain an intermediate computing node equivalent to the original computing node includes: determining a first original computing node in the original computing graph that can be equivalently replaced by the target operator in the target operator set, and determining a second original computing node in the original computing graph that cannot be equivalently replaced by the target operator in the target operator set; using the target operator that can equivalently replace the first original computing node as a computing node equivalent to the first original computing node; converting the second original computing node that cannot be equivalently replaced into a placeholder node, and recording the operator information of the second original computing node corresponding to the placeholder node; wherein the intermediate computing node includes: a computing node equivalent to the first original computing node and a placeholder node converted from the second original computing node.
[0009] In a possible implementation, the original computing graph includes multiple first original computing nodes, and the target operator that can equivalently replace the first original computing node, as a computing node equivalent to the first original computing node, includes: for any first original computing node in the original computing graph, when there are at least two target operators in the target operator set that can equivalently replace the first original computing node, determining the computing node equivalent to the first original computing node according to the computing efficiency and / or storage efficiency corresponding to each of the at least two target operators.
[0010] In one possible implementation, after establishing a directed edge between the intermediate computing node and the intermediate storage node, the conversion of the original computing graph into an intermediate computing graph based on a target operator set also includes: in a case where the original computing graph includes a third original computing node for performing a data rearrangement operation, according to the node position of the third original computing node in the original computing graph, the third original computing node is converted into data rearrangement information, and the data rearrangement information is recorded on the directed edge corresponding to the node position; wherein the data rearrangement operation includes an operation for changing the data arrangement without generating any calculation, and the data rearrangement information is used to indicate the data rearrangement operation.
[0011] In one possible implementation, after obtaining the data rearrangement information, converting the original computational graph into an intermediate computational graph based on a target operator set also includes: eliminating redundant data rearrangement information according to the data storage rules corresponding to the hardware device, and / or merging at least two data rearrangement information on the same directed edge into a single data rearrangement information.
[0012] In one possible implementation, the intermediate computing node includes: a computing node equivalent to the first original computing node, the first original computing node being a node in the original computing graph that can be equivalently replaced by the target operator in the target operator set; wherein, after establishing a directed edge between the intermediate computing node and the intermediate storage node, the original computing graph is converted into an intermediate computing graph based on the target operator set, and further includes: according to the operator computing rules corresponding to the hardware device, at least two computing nodes that satisfy the operator computing rules and have a data dependency relationship are merged into a single computing node, wherein the hardware computing rules are used to indicate that the hardware device can perform at least two target operators for a single calculation.
[0013] In a possible implementation, the hardware constraint includes: at least one of a data accuracy requirement of the hardware device, a hardware resource restriction, and a completeness requirement of the target operator set, wherein the data accuracy requirement represents the hardware device's requirement for the data accuracy of the target computation graph, the hardware resource restriction represents the hardware device's restriction on the hardware resources required for the target computation graph, and the completeness requirement represents that the target operator in the target operator set cannot equivalently replace all the original computing nodes in the intermediate computation graph; wherein, converting the intermediate computation graph into a target computation graph adapted to the hardware device according to the hardware constraint corresponding to the hardware device includes: when the hardware constraint includes the data accuracy requirement, converting the accuracy of the intermediate computation graph to obtain a target computation graph that meets the data accuracy requirement; and / or, when the hardware constraint includes the hardware resource restriction, optimizing the structure of the intermediate computation graph to obtain a target computation graph that meets the specified execution efficiency requirement; and / or, when the hardware constraint includes the completeness requirement of the target operator set, generating a target computation graph that is approximately equivalent to the intermediate computation graph according to the target operator set.
[0014] In one possible implementation, the intermediate computation graph includes placeholder nodes converted from second original computation nodes that cannot be equivalently replaced by target operators in the target operator set, and the placeholder nodes correspond to operator information of the second original computation nodes; wherein, when the hardware constraint condition includes the completeness requirement of the target operator set, generating a target computation graph approximately equivalent to the intermediate computation graph according to the target operator set includes: performing knowledge distillation on the intermediate neural network model corresponding to the intermediate computation graph to obtain a lightweight neural network model approximately equivalent to the intermediate neural network model, the lightweight neural network model is constructed based on the target operator set, and the target computation graph includes the computation graph corresponding to the lightweight neural network model; and / or, determining a target operator approximately equivalent to the second original computation node from the target operator set according to the operator information corresponding to the placeholder node, and replacing the placeholder node with a target operator approximately equivalent to the second original computation node to obtain the target computation graph.
[0015] In one possible implementation, the data accuracy requirement indicates the target data accuracy required by the hardware device, wherein, when the hardware constraint includes the data accuracy requirement, the accuracy of the intermediate calculation graph is converted to obtain a target calculation graph that meets the data accuracy requirement, including: quantizing the intermediate calculation graph according to the target data accuracy to obtain a target calculation graph with the target data accuracy; and / or, when the intermediate calculation node of the intermediate calculation graph contains a nonlinear function and the target data accuracy limits the input data of the nonlinear function to a finite discrete value, by calculating the output data corresponding to each input data of the nonlinear function, generating a lookup table corresponding to the nonlinear function, and storing the lookup table in an intermediate storage node connected to the intermediate calculation node of the nonlinear function to obtain the target calculation graph.
[0016] In one possible implementation, the optimization of the structure of the intermediate computation graph includes at least one of the following processing: deleting some intermediate computation nodes and / or some intermediate storage nodes in the intermediate computation graph; deleting some directed edges in the intermediate computation graph; decomposing some intermediate computation nodes in the intermediate computation graph into multiple sub-computation nodes; merging some node combinations with branch structures in the intermediate computation graph into node combinations without branch structures; modifying some node combinations with branch structures in the intermediate computation graph whose branch structures are unbalanced in load into node combinations with branch structures that are load-balanced.
[0017] According to another aspect of the present disclosure, a neural network model deployment device is provided, including: an acquisition module, used to acquire an original calculation graph corresponding to an original neural network model to be deployed on a hardware device; an equivalent conversion module, used to convert the original calculation graph into an intermediate calculation graph based on a target operator set, wherein the target operator set is an operator set supported by the hardware device; an approximate conversion module, used to convert the intermediate calculation graph into a target calculation graph adapted to the hardware device according to hardware constraints corresponding to the hardware device, wherein the hardware constraints are determined based on data accuracy, hardware resources and at least one of the target operator set corresponding to the hardware device; a determination module, used to determine a hardware executable file based on the target calculation graph, wherein the hardware executable file is used to be deployed on the hardware device.
[0018] In one possible implementation, the equivalent conversion module includes: a replacement submodule, used to replace the original computing node in the original computing graph with the target operator in the target operator set, so as to obtain an intermediate computing node equivalent to the original computing node; a generation submodule, used to generate an intermediate storage node associated with the intermediate computing node according to the data storage rule corresponding to the hardware device, wherein the data storage rule is used to indicate the data storage method supported by the hardware device; an establishment submodule, used to establish a directed edge between the intermediate computing node and the intermediate storage node according to the data dependency relationship between the nodes in the original computing graph, so as to obtain the intermediate computing graph, wherein the data dependency relationship represents the association relationship between the nodes in the original computing graph.
[0019] In a possible implementation, the method of replacing the original computing node in the original computing graph with the target operator in the target operator set to obtain an intermediate computing node equivalent to the original computing node includes: determining a first original computing node in the original computing graph that can be equivalently replaced by the target operator in the target operator set, and determining a second original computing node in the original computing graph that cannot be equivalently replaced by the target operator in the target operator set; using the target operator that can equivalently replace the first original computing node as a computing node equivalent to the first original computing node; converting the second original computing node that cannot be equivalently replaced into a placeholder node, and recording the operator information of the second original computing node corresponding to the placeholder node; wherein the intermediate computing node includes: a computing node equivalent to the first original computing node and a placeholder node converted from the second original computing node.
[0020] In a possible implementation, the original computing graph includes multiple first original computing nodes, and the target operator that can equivalently replace the first original computing node, as a computing node equivalent to the first original computing node, includes: for any first original computing node in the original computing graph, when there are at least two target operators in the target operator set that can equivalently replace the first original computing node, determining the computing node equivalent to the first original computing node according to the computing efficiency and / or storage efficiency corresponding to each of the at least two target operators.
[0021] In one possible implementation, after establishing the directed edge between the intermediate computing node and the intermediate storage node, the equivalent conversion module further includes: a rearrangement operation conversion submodule, which is used to, when the original computing graph includes a third original computing node for performing a data rearrangement operation, convert the third original computing node into data rearrangement information according to the node position of the third original computing node in the original computing graph, and record the data rearrangement information on the directed edge corresponding to the node position; wherein the data rearrangement operation includes an operation for changing the data arrangement and not generating calculations, and the data rearrangement information is used to indicate the data rearrangement operation.
[0022] In one possible implementation, after obtaining the data rearrangement information, the equivalent conversion module also includes: a rearrangement information optimization submodule, which is used to eliminate redundant data rearrangement information according to the data storage rules corresponding to the hardware device, and / or merge at least two data rearrangement information on the same directed edge into a single data rearrangement information.
[0023] In one possible implementation, the intermediate computing node includes: a computing node equivalent to the first original computing node, the first original computing node is a node in the original computing graph that can be equivalently replaced by the target operator in the target operator set; wherein, after establishing a directed edge between the intermediate computing node and the intermediate storage node, the equivalent conversion module also includes: a node fusion submodule, which is used to fuse at least two computing nodes that satisfy the operator calculation rules and have data dependencies into a single computing node according to the operator calculation rules corresponding to the hardware device, wherein the hardware calculation rules are used to indicate that the hardware device can perform at least two target operators for a single calculation.
[0024] In a possible implementation, the hardware constraint includes: at least one of: a data accuracy requirement of the hardware device, a hardware resource limitation, and a completeness requirement of the target operator set, wherein the data accuracy requirement represents the hardware device's requirement for the data accuracy of the target computation graph, the hardware resource limitation represents the hardware device's limitation on the hardware resources required for the target computation graph, and the completeness requirement represents that the target operator in the target operator set cannot equivalently replace all the original computing nodes in the intermediate computation graph; wherein the approximate conversion module includes: an accuracy conversion submodule, which is used to convert the accuracy of the intermediate computation graph to obtain a target computation graph that meets the data accuracy requirement when the hardware constraint includes the data accuracy requirement; and / or a structure optimization submodule, which is used to optimize the structure of the intermediate computation graph to obtain a target computation graph that meets the specified execution efficiency requirement when the hardware constraint includes the hardware resource limitation; and / or a completeness conversion submodule, which is used to generate a target computation graph that is approximately equivalent to the intermediate computation graph according to the target operator set when the hardware constraint includes the completeness of the target operator set.
[0025] In one possible implementation, the intermediate computation graph includes placeholder nodes converted from second original computation nodes that cannot be equivalently replaced by target operators in the target operator set, and the placeholder nodes correspond to operator information of the second original computation nodes; wherein, when the hardware constraint condition includes the completeness requirement of the target operator set, generating a target computation graph approximately equivalent to the intermediate computation graph according to the target operator set includes: performing knowledge distillation on the intermediate neural network model corresponding to the intermediate computation graph to obtain a lightweight neural network model approximately equivalent to the intermediate neural network model, the lightweight neural network model is constructed based on the target operator set, and the target computation graph includes the computation graph corresponding to the lightweight neural network model; and / or, determining a target operator approximately equivalent to the second original computation node from the target operator set according to the operator information corresponding to the placeholder node, and replacing the placeholder node with a target operator approximately equivalent to the second original computation node to obtain the target computation graph.
[0026] In one possible implementation, the data accuracy requirement indicates the target data accuracy required by the hardware device, wherein, when the hardware constraint includes the data accuracy requirement, the accuracy of the intermediate calculation graph is converted to obtain a target calculation graph that meets the data accuracy requirement, including: quantizing the intermediate calculation graph according to the target data accuracy to obtain a target calculation graph with the target data accuracy; and / or, when the intermediate calculation node of the intermediate calculation graph contains a nonlinear function and the target data accuracy limits the input data of the nonlinear function to a finite discrete value, by calculating the output data corresponding to each input data of the nonlinear function, generating a lookup table corresponding to the nonlinear function, and storing the lookup table in an intermediate storage node connected to the intermediate calculation node of the nonlinear function to obtain the target calculation graph.
[0027] In one possible implementation, the optimization of the structure of the intermediate computation graph includes at least one of the following processing: deleting some intermediate computation nodes and / or some intermediate storage nodes in the intermediate computation graph; deleting some directed edges in the intermediate computation graph; decomposing some intermediate computation nodes in the intermediate computation graph into multiple sub-computation nodes; merging some node combinations with branch structures in the intermediate computation graph into node combinations without branch structures; modifying some node combinations with branch structures in the intermediate computation graph whose branch structures are unbalanced in load into node combinations with branch structures that are load-balanced.
[0028] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0029] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0030] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0031] According to an embodiment of the present disclosure, by obtaining the original computational graph corresponding to the original neural network model to be deployed on the hardware device, the original computational graph is converted into an intermediate computational graph consisting of a target operator set supported by the hardware device, which is equivalent to converting the original computational graph defined by one original operator set into an intermediate computational graph defined by another operator set supported by the hardware device, and then converting the intermediate computational graph into a target computational graph adapted to the hardware device according to the hardware constraints corresponding to the hardware device. The hardware executable file determined based on the target computational graph can not only conform to the target operator set supported by the hardware device, but also meet the constraints of the hardware device on data accuracy, storage rules, hardware resources and target operator set, thereby realizing efficient and unified conversion of various original neural network models into hardware executable files that can be deployed on hardware devices.
[0032] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0034] Figure 1 A flowchart of a neural network model deployment method according to an embodiment of the present disclosure is shown.
[0035] Figure 2 A schematic diagram of a computation graph according to an embodiment of the present disclosure is shown.
[0036] Figure 3 A schematic diagram showing a neural network model deployment method according to an embodiment of the present disclosure.
[0037] Figure 4 A schematic diagram showing the conversion of an original computation graph into an intermediate computation graph according to an embodiment of the present disclosure.
[0038] Figure 5 A schematic diagram showing the conversion of an original computation graph into an intermediate computation graph according to an embodiment of the present disclosure.
[0039] Figure 6 A schematic diagram showing a conversion data rearrangement operation according to an embodiment of the present disclosure is shown.
[0040] Figure 7 A schematic diagram of a conversion intermediate computation graph according to an embodiment of the present disclosure is shown.
[0041] Figure 8 A block diagram of a neural network model deployment device according to an embodiment of the present disclosure is shown.
[0042] Fig. 9 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0043] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0044] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0045] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.
[0046] With the in-depth research on artificial intelligence algorithms and non-von Neumann hardware architectures, many new hardware devices have emerged in the academic field and industry, such as multi-core chips, many-core chips, deep learning accelerators, neuromorphic chips, brain-like computing chips, etc. The emergence of these new hardware devices has also put forward new requirements for the supporting software tool chain. The supporting software tool chain is an important guarantee for the availability and ease of use of hardware devices. Its purpose is to complete the entire process from new algorithms such as deep learning and brain-like computing to deployment and execution. Since the current deep learning and brain-like computing algorithms can generally be expressed in the form of neural network calculation graphs, the corresponding software tool chain is generally based on the compiler based on the calculation graph. Since there are currently a variety of deep learning and brain-like computing algorithm frameworks, the main task of the compiler is generally to realize the conversion from the original calculation graph described by different deep learning and brain-like computing algorithm frameworks to a unified target calculation graph that meets hardware constraints, that is, to complete the conversion from a calculation graph defined by an original operator set to a calculation graph defined by another target operator set.
[0047] In order to enable the compiler to efficiently and uniformly convert computation graphs, that is, to ensure that the target computation graph can better meet the constraints given by the hardware device, such as the constraints of the hardware operator set, the constraints of the data accuracy, the constraints of the hardware resources, etc., the present disclosure proposes a neural network model deployment method, which can realize the efficient and uniform conversion of various original neural network models represented by computation graphs into hardware executable files that can be deployed on hardware devices. The hardware executable file can be understood as a file that can be executed on the hardware device. The embodiments of the present disclosure do not limit the type of hardware executable files.
[0048] Figure 1 A flowchart of a neural network model deployment method according to an embodiment of the present disclosure is shown. The neural network model deployment method can be executed by an electronic device such as a terminal device or a server, and the method can be implemented by a processor of the terminal device calling a computer-readable instruction stored in a memory, or the method can be executed by a server.
[0049] like Figure 1 As shown, the neural network model deployment method includes:
[0050] In step S11, the original computational graph corresponding to the original neural network model to be deployed on the hardware device is obtained.
[0051] Among them, the hardware devices may include but are not limited to: many-core neural network acceleration chips, neuromorphic chips, brain-like many-core chips, many-core graphics processing chips, many-core chips with vector or matrix acceleration units, deep learning accelerators, etc. The original neural network model may include but is not limited to artificial neural network (ANN), spiking neural network (SNN), hybrid neural network (HNN), dynamic neural network and multi-neural network model, etc. It should be understood that the types of hardware devices and the types of original neural network models corresponding to the embodiments of the present disclosure are not limited.
[0052] It is known that the computational graph of a neural network model is an intermediate representation (IR) that contains the structure and function of a neural network, and is widely used in neural network programming, compilation, execution and other systems. The computational graph is composed of nodes and directed edges connecting the nodes. In a possible implementation, the nodes include computing nodes and storage nodes. The computing nodes represent the operators to be executed by the neural network model, and the storage nodes represent the input data required by the computing nodes or the output data calculated by the computing nodes. The computing nodes and storage nodes are connected by directed edges, and the directed edges represent the data dependency between the nodes. Figure 2 A schematic diagram of a computational graph according to an embodiment of the present disclosure is shown, such as Figure 2 As shown, the computational graph includes computational nodes and storage nodes, as well as directed edges represented by arrows between the nodes.
[0053] It should be understood that those skilled in the art can use a computational graph generation method known in the art to determine the original computational graph corresponding to the original neural network model, and the embodiments of the present disclosure are not limited to this.
[0054] In step S12, the original computation graph is converted into an intermediate computation graph based on a target operator set, where the target operator set is an operator set supported by the hardware device.
[0055] Among them, the operator set is the set composed of operators, and the target operator set is the set composed of operators supported by the hardware device. The operators supported by the hardware device can be understood as operators that can perform calculations on the hardware device. The operator sets under different algorithm frameworks are different. The source operator set that constitutes the original calculation graph is usually different from the target operator set. Therefore, the original calculation graph based on the source operator set can be converted into an intermediate calculation graph based on the target operator set.
[0056] Among them, converting the original computation graph into an intermediate computation graph based on the target operator set can be understood as equivalently converting the original computation graph into the intermediate computation graph. Equivalent conversion refers to a conversion operation that does not change the semantics of the computation graph, that is, for the same input, the computation graphs before and after the equivalent conversion should obtain the same output, or, using the intermediate computation graph and the original computation graph to perform operations on the same input data, respectively, will obtain the same output data.
[0057] Figure 3 A schematic diagram showing a neural network model deployment method according to an embodiment of the present disclosure is shown as follows: Figure 3 As shown, in a possible implementation, the original computation graph is converted into an intermediate computation graph based on a target operator set, which can be understood as equivalently converting the original computation graph into an intermediate computation graph, wherein the equivalent conversion may include operator conversion and graph conversion, the operator conversion is used to equivalently convert the operators in the original computation graph into target operators in the target operator set, that is, at least one operator in the original computation graph can be equivalently replaced with at least one target operator in the target operator set, which is equivalent to generating each node in the intermediate computation graph; the graph conversion is used to connect the converted target operators into a computation graph and optimize the structure of the connected computation graph, which is equivalent to generating directed edges between each node in the intermediate computation graph.
[0058] In step S13, the intermediate computation graph is converted into a target computation graph adapted to the hardware device according to the hardware constraints corresponding to the hardware device.
[0059] Among them, the hardware constraint condition is determined based on at least one of the data accuracy, hardware resources and target operator set corresponding to the hardware device. Among them, data accuracy can be understood as the data accuracy required for the hardware device to perform calculations or storage, and hardware resources can include computing resources and storage resources. It should be understood that the hardware resources that the hardware device can provide are limited, and the hardware resources required for the neural network model should be less than the hardware resources that the hardware device itself can provide; as above, the target operator set supported by the hardware device is different from the source operator set corresponding to the original calculation graph, then there may be a situation where the target operator cannot completely replace the operator in the original calculation graph. The intermediate calculation graph can be approximately converted based on the target operator set supported by the hardware device to obtain a target calculation graph that can be fully supported by the hardware device.
[0060] Among them, according to the hardware constraints corresponding to the hardware device, the intermediate calculation graph is converted into a target calculation graph adapted to the hardware device. It can be understood that according to the hardware constraints corresponding to the hardware device, the intermediate calculation graph is approximately converted into a target calculation graph adapted to the hardware device; approximate conversion refers to a conversion operation that may change the semantics of the original calculation graph to a certain extent, that is, for the same input, the calculation graphs before and after the conversion may have different outputs, but the intermediate calculation graph before the conversion and the target calculation graph after the conversion are approximately equivalent, or, for specific network processing tasks (such as target detection, image recognition, etc.), the difference between the accuracy of the intermediate calculation graph and the target calculation graph (or other indicators describing the completion effect of the neural network) is still within the allowable range.
[0061] like Figure 3 As shown, based on the above hardware constraints, the intermediate calculation graph is approximately converted, which may include at least one of precision conversion, resource conversion, structural conversion and completeness conversion. When the precision supported by the hardware device is different from the calculation or storage precision of the intermediate calculation graph, the intermediate calculation graph needs to be converted in precision; wherein, precision conversion refers to the conversion of data precision of the intermediate calculation graph, for example, it may include changing the precision of the data stored in the storage node, the precision of the input data of the input calculation node and the precision of the output data obtained by the calculation node, so as to meet the precision constraint of the hardware device on the target calculation graph, or it can save computing resources and storage resources by reducing the precision, or improve the performance of the target calculation graph on the implemented network processing task by improving the precision; resource conversion refers to the operation of adding or reducing directed edges and / or adding or reducing nodes to the intermediate calculation graph, so that the hardware resources required to be consumed by the target calculation graph are less than the hardware resources required to be consumed by the intermediate calculation graph, that is, the target calculation graph can save more hardware resources; structural conversion refers to changing part of the structure of the intermediate calculation graph to obtain the target calculation graph, so that the network structure corresponding to the target calculation graph can be more suitable for the execution calculation of the hardware device, and at the same time, it can also save the hardware resources required to be consumed by the target calculation graph;
[0062] Completeness conversion refers to the approximate approximation capability of the neural network itself. When it is impossible to directly replace the operators in the original computation graph with the target operators in the target operator set, one or more operators in the target operator set can be used to approximate one or more operators in the original computation graph. At the same time, the target computation graph and the intermediate computation graph obtained before and after the completeness conversion can still maintain approximate equivalence. In some possible implementations, for example, the intermediate computation graph can be approximated by a complete target computation graph based on the target operator set. In this implementation, the similarity between the target computation graph and the intermediate computation graph may not be restricted, that is, the data accuracy, required hardware resources and network structure of the two may be different. When the completeness of the target operator set is not enough to completely replace the operators in the original computation graph, the completeness conversion is equivalent to expanding the completeness of the target operator set, so that the target operator set after the completeness conversion has the ability to support the complete replacement of the original computation graph.
[0063] In step S14, a hardware executable file is determined based on the target computation graph, and the hardware executable file is used to be deployed on a hardware device.
[0064] As described above, the computation graph of the neural network model is an intermediate representation that includes the structure and function of the neural network. The computation graph cannot usually be executed directly on the hardware device. In order to deploy the neural network model represented by the target computation graph on the hardware device, after obtaining the target computation graph, the target computation graph can be converted into a hardware executable file. As described above, the hardware executable file can be understood as a file that can be executed on the hardware device, such as a configuration file, a binary file, etc. The embodiment of the present disclosure does not limit the file type of the hardware executable file.
[0065] It should be understood that those skilled in the art may use compilation techniques known in the art to convert the target computation graph into a hardware executable file, and this is not limited to the embodiments of the present disclosure.
[0066] In the deployment process of the above-mentioned original neural network model, equivalent transformation and approximate transformation can be combined. The intermediate computational graph obtained by equivalent transformation can retain all the semantics of the original computational graph and meet the algorithmic constraints of the target operator set. Approximate transformation can make the target computational graph meet a series of other hardware constraints and at the same time play the role of expanding the target operator set.
[0067] The neural network model deployment method of the disclosed embodiment can be applied to the conversion of a neural network computation graph (i.e., the original computation graph) to a hardware primitive graph (i.e., the target computation graph), wherein the equivalent conversion can make the intermediate computation graph conform to a given target operator set and hardware storage, and the approximate conversion can meet hardware constraints, optimize hardware resource utilization, reduce hardware resource requirements, and expand the completeness of the hardware operator set.
[0068] The neural network model deployment method of the disclosed embodiment can be applied to the programming system, compilation system and related hardware execution system of the neural network or brain-like algorithm, so that the neural network model or brain-like algorithm can be efficiently deployed on the above-mentioned hardware devices, and can also expand the completeness of the above-mentioned system, reducing the completeness requirements for the hardware equipment; it can also be applied to neural network-like tasks, such as multi-agent simulation, chip simulation, brain simulation, graphics, scientific computing, parallel search algorithms, etc., which are not limited to the disclosed embodiments.
[0069] According to an embodiment of the present disclosure, by obtaining the original computational graph corresponding to the original neural network model to be deployed on the hardware device, the original computational graph is converted into an intermediate computational graph consisting of a target operator set supported by the hardware device, which is equivalent to converting the original computational graph defined by one original operator set into an intermediate computational graph defined by another operator set supported by the hardware device, and then converting the intermediate computational graph into a target computational graph adapted to the hardware device according to the hardware constraints corresponding to the hardware device. The hardware executable file determined based on the target computational graph can not only conform to the target operator set supported by the hardware device, but also meet the constraints of the hardware device on data accuracy, storage rules, hardware resources and target operator set, thereby realizing efficient and unified conversion of various original neural network models into hardware executable files that can be deployed on hardware devices.
[0070] As described above, converting the original computation graph into an intermediate computation graph based on the target operator set may include steps such as operator conversion and graph conversion. Figure 4 A schematic diagram showing how an original computation graph is converted into an intermediate computation graph according to an embodiment of the present disclosure is shown. Figure 4 As shown, operator conversion may include generating intermediate computing nodes according to a target operator set, and generating intermediate storage nodes according to a data storage rule; graph conversion may include generating directed edges between intermediate computing nodes and intermediate storage nodes.
[0071] Based on the above Figure 4 In the equivalent conversion process shown in FIG. 1 , in a possible implementation, in step S12, the original computation graph is converted into an intermediate computation graph based on the target operator set, including:
[0072] Step S121: Use the target operator in the target operator set to replace the original computing node in the original computing graph, and obtain an intermediate computing node equivalent to the original computing node. (Equivalent to the above-mentioned generation of computing nodes according to the target operator set)
[0073] Among them, based on the principle of equivalent replacement, that is, for the same input, the outputs of the two calculation graphs before and after the equivalent replacement are the same, for each original calculation node in the original calculation graph, a target operator equivalent to the original calculation node is determined, and the original calculation node in the original calculation graph is replaced with the target operator equivalent to the original calculation node to obtain an intermediate calculation node.
[0074] Among them, the equivalent replacements that may occur may include at least one of the following: replacing a single original computing node in the original computing graph with a single target operator, replacing a single original computing node in the original computing graph with an operator combination consisting of multiple target operators, replacing multiple original computing nodes in the original computing graph with a single target operator, and replacing multiple original computing nodes in the original computing graph with multiple target operators.
[0075] Step S122: Generate an intermediate storage node associated with the intermediate computing node according to the data storage rule corresponding to the hardware device, where the data storage rule is used to indicate the data storage method supported by the hardware device. (Equivalent to generating a storage node according to the data storage rule described above)
[0076] It should be understood that the data storage methods of different hardware devices may be different and known, and the embodiments of the present disclosure do not limit the specific content of the data storage rules. According to the data storage rules corresponding to the hardware devices to be deployed, the intermediate storage nodes associated with each intermediate computing node can be generated, that is, the input storage nodes required by each intermediate computing node and the output storage nodes required by each intermediate computing node can be generated. If an intermediate storage node is both the output of an intermediate computing node and the input of another intermediate computing node, it can be represented once in the intermediate computing graph, such as Figure 2 The storage node 3 in the computation graph shown is both an output storage node of the computation node 1 and an input storage node of the computation node 2 .
[0077] Step S123: According to the data dependency relationship between each node in the original calculation graph, a directed edge between the intermediate calculation node and the intermediate storage node is established to obtain an intermediate calculation graph, wherein the data dependency relationship represents the association relationship between each node in the original calculation graph. (Equivalent to the above-mentioned generation of directed edges between the calculation node and the storage node)
[0078] It should be understood that the directed edges between the nodes in the original computation graph can indicate the data dependency relationship between the nodes, and the directed edges between the intermediate computation nodes and the intermediate storage nodes can be established based on the data dependency relationship, and the directed edges can indicate the data dependency relationship between the nodes in the intermediate computation graph. It should be understood that the embodiments of the present disclosure do not limit the way of establishing directed edges.
[0079] In the embodiments of the present disclosure, the original computation graph can be effectively converted into an intermediate computation graph based on the target operator set and data storage rules supported by the hardware device.
[0080] As described above, there may be a situation where the target operator set cannot completely replace the operators (that is, the original computing nodes) in the original computing graph. This situation can be understood as the different completeness of the target operator set and the source operator set. In this case, the computing graph conversion may not be completed, which will cause the entire compilation process to fail. Therefore, the intermediate computing graph can be approximately converted to reduce the failure of the compilation process. In order to facilitate the approximate conversion of the intermediate computing graph in this case in step S13, in a possible implementation method, in step S121, the original computing nodes in the original computing graph are replaced with the target operators in the target operator set to obtain intermediate computing nodes equivalent to the original computing nodes, including:
[0081] Determine a first original computing node in the original computing graph that can be equivalently replaced by a target operator in the target operator set, and determine a second original computing node in the original computing graph that cannot be equivalently replaced by a target operator in the target operator set; use the target operator that can equivalently replace the first original computing node as a computing node equivalent to the first original computing node; convert the second original computing node that cannot be equivalently replaced into a placeholder node, and record the operator information of the second original computing node corresponding to the placeholder node; wherein the intermediate computing node includes: a computing node equivalent to the first original computing node and a placeholder node converted from the second original computing node.
[0082] Among them, for example, based on the principle of equivalent replacement mentioned in the above-mentioned embodiments of the present disclosure, the first original computing node in the original computing graph that can be equivalently replaced by the target operator in the target operator set can be first determined, then the original computing nodes in the original computing graph other than the first original computing node can be the second original computing nodes that cannot be equivalently replaced by the target operator in the target operator set.
[0083] The operator information of the second original computing node includes, for example, at least: information describing the operator in the second original computing node, such as the type, precision, shape, parameters, input data, and output data of the operator. The operator information can be recorded in a placeholder node, and of course can also be recorded in other storage spaces, which is not limited in the embodiments of the present disclosure. The placeholder node is used to characterize that the occupied node is an original computing node that cannot be equivalently replaced. The placeholder node and the operator information can facilitate the subsequent approximate conversion of the intermediate computing graph.
[0084] It should be understood that if the target operator set can completely replace the operators (that is, the original computing nodes) in the original computing graph, that is, all the original computing nodes in the original computing graph can be equivalently replaced by the target operators in the target operator set, then after determining the first original computing node in the original computing graph that can be equivalently replaced by the target operator in the target operator set, the target operator that can equivalently replace the first original computing node can be directly used as a computing node equivalent to the first original computing node.
[0085] In the embodiments of the present disclosure, in the case where the target operator set cannot completely replace the operators in the original computation graph, placeholder nodes can be effectively used to generate an intermediate computation graph to be subjected to approximate conversion processing, which is helpful to reduce compilation failures.
[0086] Figure 5 A schematic diagram showing how an original computation graph is converted into an intermediate computation graph according to an embodiment of the present disclosure is shown. Figure 5 As shown, operator conversion may also include replacing generated computing nodes according to differences in computing efficiency of target operators; specifically, for some original computing nodes on the original computing graph, if there are multiple target operators in the target operator set that can equivalently replace a certain original computing node, but the computing efficiencies of the multiple target operators are different, then the target operator with the highest computing efficiency among the multiple target operators may be selected to replace the computing node equivalent to the original computing node, that is, a target operator with better computing efficiency may be selected to replace the computing node generated by the above step S121.
[0087] In a possible implementation, the original computation graph includes a plurality of first original computation nodes. In step S12, a target operator that can equivalently replace the first original computation node is used as a computation node equivalent to the first original computation node, including:
[0088] Step S124: For any first original computing node in the original computing graph, if there are at least two target operators in the target operator set that can equivalently replace the first original computing node, determine a computing node equivalent to the first original computing node according to the computing efficiency and / or storage efficiency of each of the at least two target operators. In this way, the operators in the converted intermediate computing graph can meet the requirements of computing efficiency and / or storage efficiency.
[0089] Among them, the computational efficiency and storage efficiency corresponding to the target operator can be determined in advance through experiments, and the storage efficiency corresponding to the target operator can be understood as the storage efficiency corresponding to the input data and output data of the target operator; the computational efficiency and storage efficiency corresponding to different operators may be different or the same. It should be understood that if the computational efficiency and storage efficiency of at least two target operators that can equivalently replace the first original computing node are different, the target operator with the highest computational efficiency among the at least two target operators can be selected as the computing node equivalent to the first original computing node; or the target operator with the highest storage efficiency can also be selected as the computing node equivalent to the first original computing node; or the target operator with the highest comprehensive computational efficiency and storage efficiency can also be selected as the computing node equivalent to the first original computing node, and this is not limited to the embodiments of the present disclosure.
[0090] If the computing efficiency of at least two target operators that can equivalently replace the first original computing node is the same, one of the target operators can be randomly selected, or the default target operator can still be selected as the computing node equivalent to the first original computing node; or the target operator with the highest storage efficiency can be further selected as the computing node equivalent to the first original computing node based on the storage efficiency of the at least two target operators. If the storage efficiency corresponding to the at least two target operators is also the same, one of the target operators can still be randomly selected, or the default target operator can be selected, and this is not limited to the embodiments of the present disclosure.
[0091] In one possible implementation, when there are at least two target operators in the target operator set that can equivalently replace the first original computing node, the above-mentioned step S121 can directly adopt the default target operator equivalent to the original computing node, and step S124 can be used as an independent step to replace the default target operator used in step S121, that is, a target operator with higher computational efficiency than the default target operator used in step S121 can be selected through step S124 to replace the computing node generated in step S121; of course, step S124 can also be used as a sub-step of step S121, so that the target operator equivalent to the first original computing node can be directly selected in step S121 according to the computational efficiency, and the embodiments of the present disclosure are not limited to this.
[0092] It should be understood that, for any first original computing node in the original computing graph, if there is a target operator in the target operator set that can equivalently replace the first original computing node, then the target operator can be directly used as a computing node equivalent to the first original computing node; step S124 can be an optional step in step S12.
[0093] like Figure 5 As shown, the graph conversion may also include conversion data reordering operations; wherein, the data reordering operations include operations for changing the data arrangement and not generating calculations, such as transposition, changing data types, and the like. Since these data reordering operations do not actually perform calculations on hardware devices, these data reordering operations may be directly converted into data reordering information, and the data reordering information is used to indicate the data reordering operation, or the data reordering information is used to describe the data reordering operator, that is, conversion data reordering operations may be understood as converting the original computing nodes in the original computing graph for performing the data reordering operation into data reordering information.
[0094] In a possible implementation, after establishing the directed edge between the intermediate computing node and the intermediate storage node, in step S12, the original computing graph is converted into an intermediate computing graph based on the target operator set, further comprising:
[0095] Step S125: When the original computation graph includes a third original computation node for performing a data rearrangement operation, the third original computation node is converted into data rearrangement information according to the node position of the third original computation node in the original computation graph, and the data rearrangement information is recorded on the directed edge corresponding to the node position. In this way, the intermediate computation graph can be made more suitable for the execution operation of the hardware device.
[0096] The node position may represent the position of the third original computing node in the original computing graph, and based on the node position, the data rearrangement information may be recorded on the directed edge corresponding to the node position. Figure 6 A schematic diagram showing a conversion data rearrangement operation according to an embodiment of the present disclosure is shown as follows: Figure 6 As shown, the node a2 between the node a1 and the node a3 in the original computation graph for performing the transposition operation can be converted into the data reordering information recorded on the directed edge between the node b1 and the node b2 in the intermediate computation graph, wherein the node b1 can be a node equivalent to the node a1, the node b2 can be a node equivalent to the node a3, and the directed edge between the node b1 and the node b2 can be a directed edge corresponding to the node position of the node a2.
[0097] Considering that some data rearrangement information may be redundant and useless for the data storage rules of the hardware device, for example, the data storage rules indicate that the data storage form of the hardware device is one-dimensional data, then if the data rearrangement information indicates that the high-dimensional data is expanded into one-dimensional data, this is redundant data rearrangement information for the hardware device; for another example, if multiple data rearrangement information indicates multiple changes in data shape, such as there are two data rearrangement information respectively indicating "change from shape A to shape B, and then from shape B to shape C", then the multiple data rearrangement information can be fused into one data rearrangement information to indicate a change in data shape, for example, the fused data rearrangement information can directly indicate the change from shape A to shape C. Figure 5 As shown, the graph conversion may also include optimizing data rearrangement information, which is used to eliminate or merge some redundant and useless data rearrangement information according to the data storage rules corresponding to the hardware device.
[0098] In a possible implementation, after obtaining the data rearrangement information, in step S12, the original computation graph is converted into an intermediate computation graph based on the target operator set, further comprising:
[0099] Step S126: Eliminate redundant data rearrangement information according to the data storage rules corresponding to the hardware device, and / or merge at least two pieces of data rearrangement information on the same directed edge into a single piece of data rearrangement information. In this way, redundant and useless data rearrangement information on the intermediate computation graph can be eliminated or merged.
[0100] Among them, redundant data rearrangement information can be understood as information that is useless for data storage or data calculation; at least two data rearrangement information on the same directed edge can indicate multiple consecutive data rearrangement operations, so at least two data rearrangement information on the same directed edge can be directly merged to make the intermediate calculation graph more concise and effective. It should be understood that step S126 can be an optional step.
[0101] As described above, the intermediate computing nodes include: computing nodes equivalent to the first original computing nodes, the first original computing nodes are nodes in the original computing graph that can be equivalently replaced by the target operator in the target operator set; considering that some node combinations composed of multiple computing nodes can be implemented by a hardware device performing a single calculation, for example, the maximum pooling operation and the relu activation function can be synthesized into an operator to perform a single calculation in the hardware device, and the convolution operation and the normalization operation can also be synthesized into an operator to perform a single calculation in the hardware device, such as Figure 5 As shown, the graph conversion may also include a fusion computing node, which is used to fuse at least two computing nodes in the intermediate computing graph into a single computing node according to the operator computing rules corresponding to the hardware device, thereby reducing the computing and storage costs of the intermediate computing graph.
[0102] In a possible implementation, after establishing the directed edge between the intermediate computing node and the intermediate storage node, in step S12, the original computing graph is converted into an intermediate computing graph based on the target operator set, further comprising:
[0103] Step S127: According to the operator calculation rules corresponding to the hardware device, at least two computing nodes that meet the operator calculation rules and have data dependencies are merged into a single computing node, wherein the hardware calculation rules are used to indicate at least two target operators that the hardware device can perform calculations on at a single time. In this way, the consumption cost of the intermediate calculation graph in terms of calculation and storage can be reduced.
[0104] It should be understood that those skilled in the art can set the operator calculation rules according to actual needs, and the embodiments of the present disclosure do not limit the specific content of the operator calculation rules. Among them, at least two computing nodes that meet the operator calculation rules and have data dependencies can be understood as at least two computing nodes that can perform calculations at a time on the hardware device and are related to each other in the original calculation graph. Step S127 can be an optional step.
[0105] It should be noted that those skilled in the art can customize the specific execution order of the above steps S121 to S127 according to actual needs. The embodiment of the present disclosure does not limit the execution order of the above steps S121 to S127. For example, S121 to S122 can be executed before S123 and S125, and S126 and S127 can be executed after S123 and S125.
[0106] As described above, the hardware constraints are determined based on at least one of the data precision, hardware resources and target operator set corresponding to the hardware device, and the approximate conversion intermediate calculation graph may include at least one of precision conversion, resource conversion, structure conversion and completeness conversion. Figure 7 A schematic diagram of a conversion intermediate calculation graph according to an embodiment of the present disclosure is shown, as shown in Figure 7 As shown, based on at least one of the above-mentioned hardware constraints, the precision conversion may include at least one of quantization and generation of lookup tables, the resource conversion may include at least one of node deletion, directed edge deletion, and node decomposition, the structural conversion may include at least one of branch merging and branch balancing, and the completeness conversion may include at least one of knowledge distillation and approximate approximation.
[0107] Among them, quantization is used to convert the data precision of the intermediate calculation graph according to the data precision requirement. Among them, when the precision requirement of the hardware device for the target calculation graph is different from the data precision of the intermediate calculation graph, the intermediate calculation graph is quantized. If the precision requirement for the target calculation graph is the same as the data precision of the intermediate calculation graph, no quantization is required. It should be understood that those skilled in the art can use data precision conversion technology known in the art to achieve data precision conversion of the intermediate calculation graph. The embodiments of the present disclosure do not limit the data precision conversion method of the intermediate calculation graph.
[0108] The lookup table is referred to as LUT, which is used to generate a lookup table for some nonlinear functions (such as e x , lnx, etc.), the output data corresponding to all discrete input data within the input precision range indicated by the data precision requirement is recorded as a lookup table and stored in the computational graph, thereby efficiently realizing the approximate calculation of some nonlinear functions that cannot be natively supported in the target operator set; if the precision conversion of the nonlinear function is realized by generating a lookup table, the input data can be a finite discrete value, which is equivalent to the precision conversion of the continuous nonlinear function. In one possible implementation, since the nonlinear function can also be realized by Taylor expansion and other methods, when the nonlinear function is realized by Taylor expansion, it is not necessary to generate the lookup table corresponding to the nonlinear function, and generating the lookup table can be an optional step.
[0109] Among them, deleting nodes is used to delete certain nodes in the intermediate calculation graph according to hardware resource limitations; deleting directed edges is used to delete certain directed edges in the intermediate calculation graph according to hardware resource limitations; after deleting nodes and / or directed edges, retraining and fine-tuning are generally required, that is, retraining the intermediate neural network model corresponding to the intermediate calculation graph, so that the target calculation graph can be approximately equivalent to the intermediate calculation graph, while reducing the computing resources and storage resources required for the target calculation graph, thereby improving the execution efficiency of the target calculation graph; node decomposition (which can also be tensor decomposition) is used to decompose the computing nodes in a certain network layer in the neural network model into computing nodes of multiple sub-network layers according to hardware resource limitations, thereby reducing the original amount of calculation and parameter amount of the neural network in this layer. The embodiments of the present disclosure do not limit the specific implementation methods of deleting nodes, deleting directed edges, and decomposing nodes.
[0110] Among them, the merging branch is used to merge some node combinations with branch structures in the intermediate calculation graph into node combinations without branch structures according to the hardware resource limitations, so as to reduce the dynamic storage and computing amount of the target calculation graph during the execution process, so that the target calculation graph itself is more suitable for execution by certain hardware devices; the balancing branch is used to modify some node combinations with multi-branch structures in the intermediate calculation graph according to the hardware resource limitations, so that the different branches in the node combination are load-balanced as much as possible, which is conducive to improving the utilization rate of hardware resources; wherein, load balancing can be understood as the computing speeds of the various branches in the node combination are similar, the storage differences are small, etc. The embodiments of the present disclosure do not limit the specific implementation methods of merging branches and balancing branches. Figure 2 The node combination consisting of computing node 2, storage node 4, computing node 3, computing node 4, and storage node 5 in the computing graph shown can be a node combination with a branch structure, computing node 4 and storage node 5 can be a branch in the node combination, and computing node 2 and storage node 4 can be another branch in the node combination.
[0111] Among them, knowledge distillation refers to using the output results of a neural network model with strong learning ability and large scale to train a lightweight neural network model with weak learning ability and small scale, so as to transfer the knowledge learned in the neural network model to the lightweight neural network model, that is, to achieve the role of compressing the neural network model; by performing knowledge distillation on the intermediate neural network model corresponding to the intermediate calculation graph, a lightweight neural network model composed of target operators in the target operator set is obtained, and the calculation graph corresponding to the lightweight neural network model can be the target calculation graph, and the lightweight neural network model can be different from the intermediate neural network model in structure, type, parameters, etc., which can not only reduce the hardware resources required for the target calculation graph, but also play the role of processing the placeholder nodes in the intermediate calculation graph, that is, it plays the role of processing the operators in the original calculation graph that the hardware device cannot support.
[0112] Among them, approximate approximation refers to using the target operator in the target operator set to approximate the computing nodes that cannot be equivalently replaced in the original computing graph, that is, replacing the placeholder nodes in the intermediate computing graph with one or more target operators (or sub-neural networks) that can be supported by the hardware device. It can not only play the role of processing the placeholder nodes in the intermediate computing graph, but also expand the completeness of the target operator set.
[0113] It should be noted that the embodiments of the present disclosure do not limit the specific execution order of the above-mentioned precision conversion, resource conversion, structure conversion and completeness conversion. There is no limitation on the specific execution order of quantization, generation of lookup table, deletion of nodes, deletion of directed edges, decomposition of nodes, merging branches, balancing branches, knowledge distillation and approximation.
[0114] As described above, the hardware constraint condition is determined based on at least one of the data accuracy, hardware resources and target operator set corresponding to the hardware device. In a possible implementation, the hardware constraint condition includes: at least one of the data accuracy requirement of the hardware device, the hardware resource limitation and the completeness requirement of the target operator set, wherein the data accuracy requirement represents the hardware device's requirement for the data accuracy of the target computation graph, the hardware resource limitation represents the hardware device's limitation on the hardware resources required for the target computation graph, and the completeness requirement represents that the target operator in the target operator set cannot equivalently replace all the original computation nodes in the intermediate computation graph.
[0115] In the embodiments of the present disclosure, the conversion of an original computation graph to another target computation graph can be achieved by combining equivalent conversion and approximate conversion. The equivalent conversion can completely preserve the semantics of the original computation graph, and the approximate conversion can make the intermediate computation graph obtained after the equivalent conversion further meet other hardware constraints, while also expanding the completeness of the target operator set corresponding to the target computation graph.
[0116] In the embodiment of the present disclosure, the conversion method that combines equivalent conversion and approximate conversion can be applied to the conversion of a neural network computation graph (equivalent to the original computation graph) to a hardware primitive computation graph (equivalent to the target computation graph). The intermediate computation graph after the equivalent conversion can satisfy the target operator set and hardware storage rules supported by the hardware device while retaining the semantics of the original computation graph. On this basis, the approximate conversion can enable the hardware primitive computation graph to meet the data accuracy requirements, improve the utilization of hardware resources, reduce the various hardware resources required to execute hardware executable files, and expand the completeness of the target operator set.
[0117] Based on the above hardware constraints, in a possible implementation, in step S13, according to the hardware constraints corresponding to the hardware device, the intermediate computation graph is converted into a target computation graph adapted to the hardware device, including:
[0118] Step S131: when the hardware constraint condition includes data precision requirement, convert the precision of the intermediate computation graph to obtain a target computation graph that meets the data precision requirement; and / or,
[0119] Step S132: When the hardware constraints include hardware resource limitations, optimizing the structure of the intermediate computation graph to obtain a target computation graph that meets the specified execution efficiency requirements; and / or,
[0120] Step S133: When the hardware constraint condition includes the completeness requirement of the target operator set, a target computation graph approximately equivalent to the intermediate computation graph is generated according to the target operator set.
[0121] As described above, the data accuracy requirement indicates the target data accuracy required by the hardware device. In a possible implementation, in step S131, when the hardware constraint condition includes the data accuracy requirement, the accuracy of the intermediate calculation graph is converted to obtain a target calculation graph that meets the data accuracy requirement, including: quantizing the intermediate calculation graph according to the target data accuracy to obtain a target calculation graph with the target data accuracy; and / or, when the intermediate calculation node of the intermediate calculation graph contains a nonlinear function and the target data accuracy limits the input data of the nonlinear function to a finite discrete value, by calculating the output data corresponding to each input data of the nonlinear function, generating a lookup table corresponding to the nonlinear function, and storing the lookup table in an intermediate storage node connected to the intermediate calculation node of the nonlinear function, to obtain the target calculation graph. Among them, the intermediate storage node connected to the intermediate calculation node of the nonlinear function can be determined according to the data dependency relationship (i.e., directed edge) between each node in the intermediate calculation graph. In this way, the intermediate calculation graph can be effectively converted into a target calculation graph that meets the data accuracy requirement.
[0122] The embodiment of the present disclosure does not limit the quantization method of the intermediate calculation graph and the generation method of the lookup table. The output data corresponding to each input data of the nonlinear function is calculated, that is, the output value corresponding to the discrete value indicated by the target data precision is calculated by using the nonlinear function. For example, assuming that the nonlinear function is e x , the target data accuracy indicates that the input data of the nonlinear function is 0 and 1, then the output data corresponding to 0 can be 1, and the output data corresponding to 1 can be 2.718.
[0123] In step S132, specifying the execution efficiency requirement may include improving the execution efficiency or reducing the execution efficiency; the target computation graph that meets the specified execution efficiency requirement may be understood as a target computation graph with higher execution efficiency or a target computation graph with lower execution efficiency, that is, the execution efficiency of the target computation graph obtained after optimizing the structure of the intermediate computation graph may be lower or higher than that of the intermediate computation graph; specifically, the execution efficiency of the intermediate computation graph may be adjusted by modifying the structure of the intermediate computation graph according to the hardware resource limitation of the hardware device. For example, if the hardware resources of the hardware device can support an intermediate computation graph with higher execution efficiency, the intermediate computation graph may be converted into a target computation graph with higher execution efficiency; if the hardware resources of the hardware device cannot support the current execution efficiency of the intermediate computation graph, the intermediate computation graph may be converted into a target computation graph with lower execution efficiency.
[0124] In one possible implementation, in step S132, the structure of the intermediate computation graph is optimized, including at least one of the following processing: deleting some intermediate computation nodes and / or some intermediate storage nodes in the intermediate computation graph (equivalent to deleting nodes); deleting some directed edges in the intermediate computation graph (equivalent to deleting directed edges as described above); decomposing some intermediate computation nodes in the intermediate computation graph into multiple sub-computation nodes (equivalent to the decomposition nodes as described above); merging some node combinations with branch structures in the intermediate computation graph into node combinations without branch structures (equivalent to the merged nodes as described above); modifying some node combinations with branch structures in the intermediate computation graph whose branch structures are unbalanced in load into node combinations with branch structures that are load-balanced (equivalent to the balancing nodes as described above).
[0125] As described above, the embodiments of the present disclosure do not limit the specific implementation methods of deleting nodes, deleting directed edges, decomposing nodes, merging branches, and balancing branches. Load imbalance can be understood as the large difference in computing speed and storage between the branches in the node combination. Correspondingly, load balancing can be understood as the similar computing speed and small storage difference between the branches in the node combination. By optimizing the structure of the intermediate calculation graph, the hardware equipment can support the computing resources and storage resources required by the target calculation graph, adjust the execution efficiency of the target calculation graph, reduce the amount of calculation and parameters of a single network layer, reduce the dynamic storage and calculation amount of the target calculation graph during the calculation process, make the target calculation graph itself more suitable for hardware execution, and make the different branches in the node combination as load-balanced as possible, which is conducive to improving the utilization of hardware resources.
[0126] As described above, the intermediate computation graph includes placeholder nodes converted from second original computation nodes that cannot be equivalently replaced by target operators in the target operator set, and the placeholder nodes correspond to operator information of the second original computation nodes; in a possible implementation, in step S133, when the hardware constraint condition includes the completeness requirement of the target operator set, a target computation graph approximately equivalent to the intermediate computation graph is generated according to the target operator set, including:
[0127] Perform knowledge distillation on the intermediate neural network model corresponding to the intermediate computational graph to obtain a lightweight neural network model that is approximately equivalent to the intermediate neural network model, wherein the lightweight neural network model is constructed based on a target operator set, and the target computational graph includes a computational graph corresponding to the lightweight neural network model; and / or, based on operator information corresponding to the placeholder node, determine a target operator that is approximately equivalent to the second original computational node from the target operator set, and replace the placeholder node with a target operator that is approximately equivalent to the second original computational node to obtain a target computational graph.
[0128] Among them, knowledge distillation of the intermediate neural network model can be understood as using the output results of the intermediate neural network model with strong learning ability and large scale to train the lightweight neural network model with weak learning ability and small scale, so as to realize the transfer of knowledge learned in the intermediate neural network model to the lightweight neural network model. The lightweight neural network model can be a network model built based on the target operator set, that is, the lightweight neural network model can be a network model supported by the hardware device, or a network model that can be deployed and executed in the hardware device. In this way, not only can the hardware resources required for the target computation graph be reduced, but it can also play the role of processing the placeholder nodes in the intermediate computation graph, that is, it plays the role of processing the operators in the original computation graph that the hardware device cannot support.
[0129] Among them, according to the operator information corresponding to the placeholder node, the target operator approximately equivalent to the second original computing node is determined from the target operator set, which can be understood as determining the target operator that can be supported by the hardware device and is approximately equivalent to the second original computing node; approximately equivalent can be understood as for the same input, the output results between the target operator and the second original computing node are different but similar, or the difference between the output results of the target operator and the second original computing node for the same input data is within an allowable range. As described above, the operator information can describe the second original computing node, so based on the operator information, a target operator approximately equivalent to the second original computing node can be found.
[0130] It should be understood that the target operator approximately equivalent to the second original computing node can be a single target operator or an operator combination consisting of multiple target operators, that is, a placeholder node can be replaced by a single target operator or an operator combination consisting of multiple target operators. It can also be that a single target operator replaces multiple placeholder nodes or multiple target operators replace multiple placeholder nodes. In this way, not only can the placeholder nodes in the intermediate computing graph be processed, but also the completeness of the target operator set can be expanded.
[0131] It should be noted that the embodiment of the present disclosure does not limit the specific execution order of the above steps S131, S132, and S133. In the embodiment of the present disclosure, the intermediate calculation graph is converted into a target calculation graph through at least one of the above steps S131, S132, and S133, which can adapt to the hardware device and meet various constraints of the hardware device, thereby realizing efficient and unified conversion of the original neural network models under different algorithm frameworks into hardware executable files that can be deployed on the above hardware devices.
[0132] Figure 8 A block diagram of a neural network model deployment device according to an embodiment of the present disclosure is shown as follows: Figure 8 As shown, the neural network model deployment device includes:
[0133] An acquisition module 101 is used to acquire an original computational graph corresponding to an original neural network model to be deployed on a hardware device;
[0134] An equivalent conversion module 102, used to convert the original calculation graph into an intermediate calculation graph based on a target operator set, where the target operator set is an operator set supported by the hardware device;
[0135] An approximate conversion module 103, configured to convert the intermediate computation graph into a target computation graph adapted to the hardware device according to a hardware constraint condition corresponding to the hardware device, wherein the hardware constraint condition is determined based on at least one of data accuracy, hardware resources, and the target operator set corresponding to the hardware device;
[0136] The determination module 104 is used to determine a hardware executable file based on the target computation graph, and the hardware executable file is used to be deployed on the hardware device.
[0137] In one possible implementation, the equivalent conversion module 102 includes: a replacement submodule, used to replace the original computing node in the original computing graph with the target operator in the target operator set, so as to obtain an intermediate computing node equivalent to the original computing node; a generation submodule, used to generate an intermediate storage node associated with the intermediate computing node according to the data storage rule corresponding to the hardware device, wherein the data storage rule is used to indicate the data storage method supported by the hardware device; an establishment submodule, used to establish a directed edge between the intermediate computing node and the intermediate storage node according to the data dependency relationship between the nodes in the original computing graph, so as to obtain the intermediate computing graph, wherein the data dependency relationship represents the association relationship between the nodes in the original computing graph.
[0138] In a possible implementation, the method of replacing the original computing node in the original computing graph with the target operator in the target operator set to obtain an intermediate computing node equivalent to the original computing node includes: determining a first original computing node in the original computing graph that can be equivalently replaced by the target operator in the target operator set, and determining a second original computing node in the original computing graph that cannot be equivalently replaced by the target operator in the target operator set; using the target operator that can equivalently replace the first original computing node as a computing node equivalent to the first original computing node; converting the second original computing node that cannot be equivalently replaced into a placeholder node, and recording the operator information of the second original computing node corresponding to the placeholder node; wherein the intermediate computing node includes: a computing node equivalent to the first original computing node and a placeholder node converted from the second original computing node.
[0139] In a possible implementation, the original computing graph includes multiple first original computing nodes, and the target operator that can equivalently replace the first original computing node, as a computing node equivalent to the first original computing node, includes: for any first original computing node in the original computing graph, when there are at least two target operators in the target operator set that can equivalently replace the first original computing node, determining the computing node equivalent to the first original computing node according to the computing efficiency and / or storage efficiency corresponding to each of the at least two target operators.
[0140] In one possible implementation, after establishing the directed edge between the intermediate computing node and the intermediate storage node, the equivalent conversion module 102 further includes: a rearrangement operation conversion submodule, which is used to, when the original computing graph includes a third original computing node for performing a data rearrangement operation, convert the third original computing node into data rearrangement information according to the node position of the third original computing node in the original computing graph, and record the data rearrangement information on the directed edge corresponding to the node position; wherein the data rearrangement operation includes an operation for changing the data arrangement and not generating any calculation, and the data rearrangement information is used to indicate the data rearrangement operation.
[0141] In one possible implementation, after obtaining the data rearrangement information, the equivalent conversion module 102 further includes: a rearrangement information optimization submodule, which is used to eliminate redundant data rearrangement information according to the data storage rules corresponding to the hardware device, and / or merge at least two data rearrangement information on the same directed edge into a single data rearrangement information.
[0142] In one possible implementation, the intermediate computing node includes: a computing node equivalent to the first original computing node, the first original computing node being a node in the original computing graph that can be equivalently replaced by the target operator in the target operator set; wherein, after establishing a directed edge between the intermediate computing node and the intermediate storage node, the equivalent conversion module 102 further includes: a node fusion submodule, which is used to fuse at least two computing nodes that satisfy the operator computing rules and have data dependencies into a single computing node according to the operator computing rules corresponding to the hardware device, wherein the hardware computing rules are used to indicate that the hardware device can perform at least two target operators for a single calculation.
[0143] In a possible implementation, the hardware constraint includes: at least one of: a data accuracy requirement of the hardware device, a hardware resource limitation, and a completeness requirement of the target operator set, wherein the data accuracy requirement represents the hardware device's requirement for the data accuracy of the target computation graph, the hardware resource limitation represents the hardware device's limitation on the hardware resources required for the target computation graph, and the completeness requirement represents that the target operator in the target operator set cannot equivalently replace all the original computing nodes in the intermediate computation graph; wherein the approximate conversion module 103 includes: an accuracy conversion submodule, which is used to convert the accuracy of the intermediate computation graph to obtain a target computation graph that meets the data accuracy requirement when the hardware constraint includes the data accuracy requirement; and / or a structure optimization submodule, which is used to optimize the structure of the intermediate computation graph to obtain a target computation graph that meets the specified execution efficiency requirement when the hardware constraint includes the hardware resource limitation; and / or a completeness conversion submodule, which is used to generate a target computation graph that is approximately equivalent to the intermediate computation graph according to the target operator set when the hardware constraint includes the completeness requirement of the target operator set.
[0144] In one possible implementation, the intermediate computation graph includes placeholder nodes converted from second original computation nodes that cannot be equivalently replaced by target operators in the target operator set, and the placeholder nodes correspond to operator information of the second original computation nodes; wherein, when the hardware constraint condition includes the completeness requirement of the target operator set, generating a target computation graph approximately equivalent to the intermediate computation graph according to the target operator set includes: performing knowledge distillation on the intermediate neural network model corresponding to the intermediate computation graph to obtain a lightweight neural network model approximately equivalent to the intermediate neural network model, the lightweight neural network model is constructed based on the target operator set, and the target computation graph includes the computation graph corresponding to the lightweight neural network model; and / or, determining a target operator approximately equivalent to the second original computation node from the target operator set according to the operator information corresponding to the placeholder node, and replacing the placeholder node with a target operator approximately equivalent to the second original computation node to obtain the target computation graph.
[0145] In one possible implementation, the data accuracy requirement indicates the target data accuracy required by the hardware device, wherein, when the hardware constraint includes the data accuracy requirement, the accuracy of the intermediate calculation graph is converted to obtain a target calculation graph that meets the data accuracy requirement, including: quantizing the intermediate calculation graph according to the target data accuracy to obtain a target calculation graph with the target data accuracy; and / or, when the intermediate calculation node of the intermediate calculation graph contains a nonlinear function and the target data accuracy limits the input data of the nonlinear function to a finite discrete value, by calculating the output data corresponding to each input data of the nonlinear function, generating a lookup table corresponding to the nonlinear function, and storing the lookup table in an intermediate storage node connected to the intermediate calculation node of the nonlinear function to obtain the target calculation graph.
[0146] In one possible implementation, the optimization of the structure of the intermediate computation graph includes at least one of the following processing: deleting some intermediate computation nodes and / or some intermediate storage nodes in the intermediate computation graph; deleting some directed edges in the intermediate computation graph; decomposing some intermediate computation nodes in the intermediate computation graph into multiple sub-computation nodes; merging some node combinations with branch structures in the intermediate computation graph into node combinations without branch structures; modifying some node combinations with branch structures in the intermediate computation graph whose branch structures are unbalanced in load into node combinations with branch structures that are load-balanced.
[0147] According to an embodiment of the present disclosure, by obtaining the original computational graph corresponding to the original neural network model to be deployed on the hardware device, the original computational graph is converted into an intermediate computational graph consisting of a target operator set supported by the hardware device, which is equivalent to converting the original computational graph defined by one original operator set into an intermediate computational graph defined by another operator set supported by the hardware device, and then converting the intermediate computational graph into a target computational graph adapted to the hardware device according to the hardware constraints corresponding to the hardware device. The hardware executable file determined based on the target computational graph can not only conform to the target operator set supported by the hardware device, but also meet the constraints of the hardware device on data accuracy, storage rules, hardware resources and target operator set, thereby realizing efficient and unified conversion of various original neural network models into hardware executable files that can be deployed on hardware devices.
[0148] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0149] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0150] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0151] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0152] Fig. 9 1 is a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server or a terminal device. Fig. 9 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0153] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.
[0154] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0155] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0156] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0157] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0158] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0159] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0160] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0161] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0162] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0163] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A neural network model deployment method, It is characterized in that include: Obtain the original computational graph corresponding to the original neural network model to be deployed on the hardware device; Converting the original computation graph into an intermediate computation graph based on a target operator set, where the target operator set is an operator set supported by the hardware device; According to a hardware constraint condition corresponding to the hardware device, converting the intermediate computation graph into a target computation graph adapted to the hardware device, wherein the hardware constraint condition is determined based on at least one of data precision, hardware resources, and the target operator set corresponding to the hardware device; Based on the target computation graph, determine a hardware executable file, where the hardware executable file is used to be deployed on the hardware device; The converting of the original computation graph into an intermediate computation graph based on a target operator set includes: Using the target operator in the target operator set to replace the original computing node in the original computing graph, to obtain an intermediate computing node equivalent to the original computing node; generating an intermediate storage node associated with the intermediate computing node according to a data storage rule corresponding to the hardware device, wherein the data storage rule is used to indicate a data storage method supported by the hardware device; According to the data dependency relationship between the nodes in the original calculation graph, a directed edge between the intermediate calculation node and the intermediate storage node is established to obtain the intermediate calculation graph, wherein the data dependency relationship represents the association relationship between the nodes in the original calculation graph; The hardware constraint condition includes: at least one of the data accuracy requirement of the hardware device, the hardware resource limitation and the completeness requirement of the target operator set, wherein the data accuracy requirement represents the requirement of the hardware device for the data accuracy of the target computation graph, the hardware resource limitation represents the limitation of the hardware device on the hardware resources required for the target computation graph, and the completeness requirement represents that the target operator in the target operator set cannot equivalently replace all the original computing nodes in the intermediate computation graph; The converting of the intermediate computation graph into a target computation graph adapted to the hardware device according to the hardware constraint conditions corresponding to the hardware device comprises: In the case where the hardware constraint condition includes the data precision requirement, converting the precision of the intermediate computation graph to obtain a target computation graph that meets the data precision requirement; and / or, In the case where the hardware constraint condition includes the hardware resource limitation, optimizing the structure of the intermediate computation graph to obtain a target computation graph that meets the specified execution efficiency requirement; and / or, In a case where the hardware constraint condition includes a completeness requirement of the target operator set, a target computation graph approximately equivalent to the intermediate computation graph is generated according to the target operator set.
2. The method according to claim 1, It is characterized in that The step of replacing the original computing nodes in the original computing graph with the target operator in the target operator set to obtain the intermediate computing nodes equivalent to the original computing nodes includes: Determine a first original computing node in the original computing graph that can be equivalently replaced by a target operator in the target operator set, and determine a second original computing node in the original computing graph that cannot be equivalently replaced by a target operator in the target operator set; Using a target operator that can equivalently replace the first original computing node as a computing node equivalent to the first original computing node; Convert the second original computing node that cannot be equivalently replaced into a placeholder node, and record operator information of the second original computing node corresponding to the placeholder node; The intermediate computing nodes include: computing nodes equivalent to the first original computing nodes and placeholder nodes converted from the second original computing nodes.
3. The method according to claim 2, It is characterized in that The original computation graph includes a plurality of first original computation nodes, and the target operator that can equivalently replace the first original computation node, as a computation node equivalent to the first original computation node, includes: For any first original computing node in the original computing graph, if there are at least two target operators in the target operator set that can equivalently replace the first original computing node, determine a computing node equivalent to the first original computing node based on the computing efficiency and / or storage efficiency corresponding to each of the at least two target operators.
4. The method according to claim 1 or 2, It is characterized in that After establishing the directed edge between the intermediate computing node and the intermediate storage node, converting the original computing graph into an intermediate computing graph based on the target operator set further includes: In a case where the original computation graph includes a third original computation node for performing a data rearrangement operation, converting the third original computation node into data rearrangement information according to a node position of the third original computation node in the original computation graph, and recording the data rearrangement information on a directed edge corresponding to the node position; The data rearrangement operation includes an operation for changing data arrangement without generating any calculation, and the data rearrangement information is used to indicate the data rearrangement operation.
5. The method according to claim 4, It is characterized in that After obtaining the data rearrangement information, converting the original computation graph into an intermediate computation graph based on the target operator set further includes: According to the data storage rule corresponding to the hardware device, redundant data rearrangement information is eliminated, and / or at least two pieces of data rearrangement information on the same directed edge are merged into a single piece of data rearrangement information.
6. The method according to claim 1 or 2, It is characterized in that The intermediate computing nodes include: computing nodes equivalent to the first original computing nodes, the first original computing nodes being nodes in the original computing graph that can be equivalently replaced by the target operator in the target operator set; After establishing the directed edge between the intermediate computing node and the intermediate storage node, converting the original computing graph into an intermediate computing graph based on the target operator set further includes: According to the operator calculation rule corresponding to the hardware device, at least two computing nodes that satisfy the operator calculation rule and have a data dependency relationship are merged into a single computing node, wherein the hardware calculation rule is used to indicate at least two target operators that the hardware device can perform calculations on at a single time.
7. The method according to claim 1, It is characterized in that The intermediate computation graph includes a placeholder node converted from a second original computation node that cannot be equivalently replaced by a target operator in the target operator set, and the placeholder node corresponds to operator information of the second original computation node; Wherein, when the hardware constraint condition includes the completeness requirement of the target operator set, generating a target computation graph approximately equivalent to the intermediate computation graph according to the target operator set includes: performing knowledge distillation on the intermediate neural network model corresponding to the intermediate computation graph to obtain a lightweight neural network model that is approximately equivalent to the intermediate neural network model, wherein the lightweight neural network model is constructed based on the target operator set, and the target computation graph includes the computation graph corresponding to the lightweight neural network model; and / or, According to the operator information corresponding to the placeholder node, a target operator approximately equivalent to the second original computing node is determined from the target operator set, and the placeholder node is replaced with a target operator approximately equivalent to the second original computing node to obtain the target computing graph.
8. The method according to claim 1, It is characterized in that The data precision requirement indicates the target data precision required by the hardware device, wherein, when the hardware constraint condition includes the data precision requirement, converting the precision of the intermediate calculation graph to obtain a target calculation graph that meets the data precision requirement includes: quantizing the intermediate computation graph according to the target data accuracy to obtain a target computation graph having the target data accuracy; and / or, When a nonlinear function is included in the intermediate calculation node of the intermediate calculation graph and the target data precision limits the input data of the nonlinear function to a finite discrete value, a lookup table corresponding to the nonlinear function is generated by calculating the output data corresponding to each input data of the nonlinear function, and the lookup table is stored in an intermediate storage node connected to the intermediate calculation node of the nonlinear function to obtain the target calculation graph.
9. The method according to claim 1, It is characterized in that The optimizing the structure of the intermediate computation graph includes at least one of the following processes: Deleting some intermediate computing nodes and / or some intermediate storage nodes in the intermediate computing graph; Deleting some directed edges in the intermediate computation graph; Decomposing some intermediate computing nodes in the intermediate computing graph into multiple sub-computing nodes; Merging some node combinations with branch structures in the intermediate calculation graph into node combinations without branch structures; Modify some node combinations in the intermediate calculation graph that have a branch structure and unbalanced loads of the branch structure into a node combination with a balanced load of the branch structure.
10. A neural network model deployment device, It is characterized in that include: An acquisition module is used to acquire the original computational graph corresponding to the original neural network model to be deployed on the hardware device; An equivalent conversion module, used to convert the original calculation graph into an intermediate calculation graph based on a target operator set, where the target operator set is an operator set supported by the hardware device; An approximate conversion module, configured to convert the intermediate computation graph into a target computation graph adapted to the hardware device according to a hardware constraint condition corresponding to the hardware device, wherein the hardware constraint condition is determined based on at least one of data accuracy, hardware resources and the target operator set corresponding to the hardware device; A determination module, used to determine a hardware executable file based on the target computation graph, wherein the hardware executable file is used to be deployed on the hardware device; Wherein, the equivalent conversion module includes: a replacement submodule, used to replace the original computing node in the original computing graph with the target operator in the target operator set, so as to obtain an intermediate computing node equivalent to the original computing node; a generation submodule, used to generate an intermediate storage node associated with the intermediate computing node according to the data storage rule corresponding to the hardware device, wherein the data storage rule is used to indicate the data storage method supported by the hardware device; an establishment submodule, used to establish a directed edge between the intermediate computing node and the intermediate storage node according to the data dependency relationship between the nodes in the original computing graph, so as to obtain the intermediate computing graph, wherein the data dependency relationship represents the association relationship between the nodes in the original computing graph; Wherein, the hardware constraint conditions include: at least one of the data accuracy requirements, hardware resource restrictions and completeness requirements of the target operator set of the hardware device, wherein the data accuracy requirements represent the requirements of the hardware device for the data accuracy of the target calculation graph, the hardware resource restrictions represent the restrictions of the hardware device on the hardware resources required for the target calculation graph, and the completeness requirements represent that the target operators in the target operator set cannot equivalently replace all the original computing nodes in the intermediate calculation graph; wherein, the approximate conversion module includes: an accuracy conversion submodule, which is used to convert the accuracy of the intermediate calculation graph to obtain a target calculation graph that meets the data accuracy requirements when the hardware constraint conditions include the data accuracy requirements; and / or, a structure optimization submodule, which is used to optimize the structure of the intermediate calculation graph to obtain a target calculation graph that meets the specified execution efficiency requirements when the hardware constraint conditions include the hardware resource restrictions; and / or, a completeness conversion submodule, which is used to generate a target calculation graph that is approximately equivalent to the intermediate calculation graph according to the target operator set when the hardware constraint conditions include the completeness of the target operator set.
11. An electronic device, It is characterized in that include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 9 when executing the instructions stored in the memory.
12. A computer-readable storage medium having computer program instructions stored thereon, It is characterized in that When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Neural network processing method and device, computer device and storage medium
CN110674936A
Neural network model deployment method and device, electronic equipment and storage medium
CN111222637A