An operator monitoring method and system based on a computational graph
Through the operator monitoring method based on the computational graph, the static graph of the CPU and artificial intelligence processor is captured and compared, and the problem of incomplete corresponding implementation of deep learning networks on different devices is solved, and the operators in the network are quickly positioned and adapted to improve efficiency.
Patent Information
- Application Number
- CN202210147619.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-02-17
AI Technical Summary
In the deep learning programming framework, due to the different programming languages, the implementation of the same deep learning network on the CPU and the artificial intelligence processor is not completely corresponding, and it is difficult to quickly locate the operators that need to be implemented, especially when non-CPU devices adapt to the network.
Using an operator monitoring method based on the computing graph, by obtaining the computing graph of the network to be monitored, the static graph compiled by the CPU and the artificial intelligence processor on the network is captured, and the network operators corresponding to each node in the computing graph are compared, and operators that are not adapted and implemented in the artificial intelligence processor are determined.
It quickly locates all operators in the network and operators that need to be adapted on the artificial intelligence processor, and improves the efficiency of statistical operators of non-CPU devices and the monitoring efficiency of unadapted operators.
Smart Images

Figure CN114489604B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence chips, and particularly relates to an operator monitoring method and system based on a computational graph. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Due to deep learning programming frameworks, such as Pytorch, in implementation, due to different programming languages, the same deep learning programming is not exactly the same in the end. For example, Python and C++ are two different programming languages, and the implementations on the Python side and the C++ side do not exactly correspond, and the specific implementation of the central processing unit (CPU) does not exactly correspond to the specific implementation of the artificial intelligence processor. Therefore, it is impossible to simply obtain the specific operators to be implemented through the Python side, which makes it difficult to quickly locate the operators to be implemented when adapting the network to non-CPU devices. Summary of the Invention
[0004] To solve the technical problems existing in the above background art, the present invention provides an operator monitoring method and system based on a computational graph, which can quickly locate all the operators in a network and the operators that need to be adapted on the artificial intelligence processor.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The first aspect of the present invention provides an operator monitoring method based on a computational graph, which includes:
[0007] Obtain the computational graph of the network to be monitored;
[0008] Capture the first static graph returned after the central processing unit compiles the network to be monitored;
[0009] Capture the second static graph returned after the artificial intelligence processor compiles the network to be monitored;
[0010] Compare the computational graph with the first static graph to obtain the network operators in the first static graph corresponding to each node in the computational graph; compare the first static graph with the second static graph to obtain the network operators that are not adapted and implemented on the artificial intelligence processor.
[0011] Further, each node in the first static graph is a network operator supported by the central processing unit.
[0012] Further, each node in the second static graph is a network operator supported by the artificial intelligence processor or the central processing unit.
[0013] Further, each network operator is configured with an exception capture mechanism.
[0014] Further, for the network operators not implemented by the artificial intelligence processor, if the exception capture mechanism works properly, a static graph mixed with central processor nodes and artificial intelligence processor nodes can be generated, and this mixed static graph is used as the second static graph.
[0015] The second aspect of the present invention provides an operator monitoring system based on a computational graph, which includes:
[0016] A computational graph acquisition module, which is configured to: acquire the computational graph of the network to be monitored;
[0017] A first static graph capture module, which is configured to: capture the first static graph returned by the central processor after compiling the network to be monitored;
[0018] A second static graph capture module, which is configured to: capture the second static graph returned by the artificial intelligence processor after compiling the network to be monitored;
[0019] A comparison module, which is configured to: compare the computational graph and the first static graph to obtain the network operators in the first static graph corresponding to each node in the computational graph; compare the first static graph and the second static graph to obtain the network operators not adapted and implemented by the artificial intelligence processor.
[0020] Further, each network operator is configured with an exception capture mechanism.
[0021] Further, the second static graph capture module is further configured to: for the network operators not implemented by the artificial intelligence processor, if the exception capture mechanism works properly, a static graph mixed with central processor nodes and artificial intelligence processor nodes can be generated, and this mixed static graph is used as the second static graph.
[0022] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in an operator monitoring method based on a computational graph as described above are implemented.
[0023] The fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in an operator monitoring method based on a computational graph as described above are implemented.
[0024] Compared with the prior art, the beneficial effects of the present invention are:
[0025] The present invention provides an operator monitoring method based on a computational graph, which obtains the CPU operators corresponding to each module in the computational graph by comparing the computational graph of the network to be monitored with a first static graph; and obtains the operators that are not adapted and implemented on the artificial intelligence processor by comparing the first static graph with a second static graph, avoiding the need for manual confirmation of all operators in a network and the operators that need to be adapted on the artificial intelligence processor, improving the efficiency of counting operators for non-CPU devices and the monitoring efficiency of operators that are not adapted and implemented in non-CPU devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which form a part of this specification, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0027] Figure 1 is a flowchart of an operator monitoring method based on a computational graph according to Embodiment 1 of the present invention;
[0028] Figure 2 is a structural diagram of an operator monitoring system based on a computational graph according to Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0030] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0031] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0032] Embodiment 1
[0033] This embodiment provides an operator monitoring method based on a computational graph, as Figure 1 shown, including the following steps:
[0034] Step 1, obtain the computational graph of the network to be monitored.
[0035] Specifically, the computational graph includes a static graph and a dynamic graph. The computational graph includes nodes and edges. Among them, the nodes represent operator nodes, and the edges represent data flow directions. For example, the network to be monitored can be a neural network to be monitored. The neural network includes various layers, such as convolutional layers, pooling layers, etc. Each layer in the neural network corresponds to each operator node in the computational graph.
[0036] Specifically, print the computational graph of the network to be monitored at the Python end, so that at the top level, which layers exist in the network to be monitored or which modules exist in the network to be monitored can be obtained according to the nodes in the computational graph. The nodes in the computational graph and the network layers of the network to be detected have a one-to-one correspondence relationship. For example, the convolutional operator node in the computational graph corresponds to the convolutional layer of the network to be monitored.
[0037] Step 2: Capture the first static graph returned by the central processing unit after compiling the network to be monitored.
[0038] Perform Just In Time Compilation (jit.trace) with CPU-type input, capture the first static graph obtained by the CPU. By traversing the network nodes in the first static graph, the name of the operator or the string representation of the operator directly output can be obtained, and all the names of the operators that will actually run in all CPU modes can be obtained through processing.
[0039] The operator that can pass the compilation indicates that the operator can run with CPU-type input, that is, the operator can run in the CPU mode. The network nodes of the first static graph (computational graph) correspond to operators. By traversing the network nodes in the first static graph, the running operators can be obtained. Optionally, the network node can directly correspond to the operator name, or the network node can correspond to the string of the output operator, and this string has a one-to-one correspondence relationship with the operator. This application does not make any restrictions on this.
[0040] Among them, Just In Time Compilation (jit) is a mechanism in PyTorch and a method of program optimization. Using torch.jit.trace, the existing modules or Python functions of the network to be monitored can be obtained, providing sample inputs, and then running the function to record the operations performed on all tensors.
[0041] Pytorch is a deep learning programming framework applicable to programming languages such as Python and C++. It is used to achieve efficient parallel computing on GPUs / AI processors and the construction of deep learning networks, and has advantages such as easy expansion, fast implementation, and strong stability in production deployment. The main feature of PyTorch is tensor computing similar to Numpy. Writing new neural network modules or the interface design of the Tensor API in PyTorch is simple and straightforward with a minimum of abstractions. The operator modules commonly used in building neural network modules in PyTorch include torch.ops, torch.Tensor.ops, torch.nn.Modules, and torch.nn.functional.ops. The operator module interfaces used in writing neural network modules in PyTorch correspond to the same function in C++ implementation, and the Python interfaces define multiple calling forms, facilitating the construction of neural network modules. However, a function is just a combination of statements that perform a task, and only includes a function name, parameters, a function body, and a return type. The function call execution method is to complete a corresponding calculation once it is executed, and the return is executed according to the input passed during execution. The same is true for the functions corresponding to the PyTorch calculation module. Compared with the implementation method of static graphs, the operators in each calculation module are instances of a class. A class is object-oriented, and a class can also save some attribute states and can encapsulate multiple functions.
[0042] Among them, performing a certain operation on any function can be regarded as an operator, such as an addition operator, a convolution operator, etc.
[0043] Each node in the first static graph is a CPU node, that is, a network operator supported by the CPU (a CPU operator).
[0044] Step 3: Capture the second static graph returned after the AI processor compiles the network to be monitored.
[0045] Specifically, the AI processor can be understood as other processors except the Central Processing Unit (CPU), and can include but are not limited to general or special processors such as Graphics Processing Unit (GPU), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0046] Each node in the second static graph is a network operator supported by an AI processor or a central processing unit. That is, the second static graph includes nodes supported by the CPU and nodes supported by the AI processor, or the second static graph only includes nodes supported by the AI processor. Specifically, if no operator that is not implemented on the AI processor is encountered, then the second static graph only includes nodes supported by the AI processor, that is, each node in the second static graph is an AI processor node, that is, a network operator supported by the AI processor (an AI processor operator); if an operator that is not implemented on the AI processor is encountered, the second static graph will have CPU nodes, that is, it includes both CPU and AI processor nodes. Further, if an operator that is not implemented on the AI processor is encountered and the second static graph can be successfully generated, then it includes both CPU and AI processor nodes. If the second static graph cannot be generated, locate the operator that causes the failure, and this operator is a CPU node, and the subsequent ones continue to be judged according to the above method.
[0047] Perform just-in-time compilation (jit.trace) with the input of the AI processor type to capture the second static graph obtained by the AI processor.
[0048] Each network operator is configured with an exception capture mechanism (try-catch mechanism). For the exception capture mechanism of each AI processor operator, the static graph can be normally generated when some AI processor operator exceptions are encountered.
[0049] If an operator that is not implemented on the AI processor device is encountered, if the exception capture mechanism works properly, a static graph mixed with CPU nodes and AI processor nodes can be generated, and this static graph mixed with CPU nodes and AI processor nodes is used as the second static graph.
[0050] Among them, both the first static graph and the second static graph are composed of edges and nodes. The edges represent the data flow direction, and the nodes represent operators. Deep learning requires a large number of neural network construction and operation modules. Based on this requirement, there are currently many deep learning frameworks, such as Caffe, MXNet, and TensorFlow in addition to PyTorch. Comparing the design ideas of other frameworks, almost all of them are based on computational graphs, and computational graphs can be divided into static computational graphs and dynamic computational graphs. Static graphs mean defining first and then running, defining once and running multiple times, while dynamic computational graphs are defined during the running process, constructed during running, and can be constructed and run multiple times. Based on such a functional design idea, the operation module corresponds to a node in the graph, and the graph contains multiple operation nodes. Each operator is regarded as a graph node, the data flow direction is regarded as a directed edge, and the entire static acyclic directed computational graph composed of these points and edges is traversed (the C++ representation can be obtained at the underlying level in Pytorch) to obtain the running device information of each node.
[0051] Step 4: Compare the computational graph of the network to be monitored with the first static graph to obtain the network operators in the first static graph corresponding to each node in the computational graph; compare the first static graph with the second static graph to obtain the network operators that have not been adapted and implemented on the artificial intelligence processor. That is, by comparing the computational graph of the network to be monitored with the first static graph, it can be obtained which specific operators a module in Pytorch is implemented by; through the first static graph and the second static graph, it can be obtained which operators have not been adapted and implemented on the artificial intelligence processor device.
[0052] Specifically, the operators corresponding to the nodes in the first static graph are the operators running in the CPU mode, that is, the operators adapted and implemented on the CPU device, while the operators corresponding to the nodes in the second static graph include the operators adapted by the artificial intelligence processor or the mixed operators of the operators adapted by the intelligent processor and the operators adapted by the CPU device. By comparing the first static graph and the second static graph, the operators adapted by the artificial intelligence processor can be obtained, that is, which operators have not been implemented on the artificial intelligence processor device can be obtained.
[0053] As an implementation method, each node in the computational graph is compared one by one with each operator in the first static graph to obtain the operator in the first static graph corresponding to each node in the computational graph; each operator in the first static graph is compared one by one with each operator in the second static graph to obtain the operators missing in the second static graph, which are the operators that have not been adapted and implemented on the artificial intelligence processor.
[0054] In the worst case, if the hybrid static graph cannot be generated in step 3, the unimplemented operator can be located through the layer with network exception, and transferring it to the CPU for running can further count the implementation status of the remaining network operators.
[0055] The method of this application does not require manual confirmation of all operators in a network and the operators that need to be adapted on the artificial intelligence processor device, improving the efficiency of counting operators for non-CPU devices.
[0056] Embodiment 2
[0057] This embodiment provides an operator monitoring system based on a computational graph, as Figure 2 shown, which specifically includes the following modules:
[0058] A computational graph acquisition module, which is configured to: acquire the computational graph of the network to be monitored;
[0059] A first static graph capture module, which is configured to: capture the first static graph returned by the central processing unit after compiling the network to be monitored;
[0060] A second static graph capture module, which is configured to: capture the second static graph returned by the artificial intelligence processor after compiling the network to be monitored;
[0061] A comparison module, which is configured to: compare the computational graph with the first static graph to obtain the network operators in the first static graph corresponding to each node in the computational graph; compare the first static graph with the second static graph to obtain the network operators that are not adapted and implemented on the artificial intelligence processor.
[0062] Among them, the central processing unit is configured to acquire the network to be monitored, receive a compilation instruction, obtain the compiled model, and return the first static graph corresponding to the model; the artificial intelligence processor is configured to acquire the network to be monitored, receive a compilation instruction, obtain the compiled model, and return the second static graph corresponding to the model.
[0063] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1 one by one, and its specific implementation process is the same, so it will not be repeated here.
[0064] Embodiment 3
[0065] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a method for monitoring operators based on a computational graph as described in Embodiment 1 above.
[0066] Embodiment 4
[0067] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in an operator monitoring method based on a computational graph as described in the above-mentioned Embodiment 1.
[0068] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.
[0069] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0072] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0073] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An operator monitoring method based on a computational graph, characterized in that Including: Obtain the computational graph of the network to be monitored; Capture the first static graph returned after the central processing unit compiles the network to be monitored; Capture the second static graph returned after the artificial intelligence processor compiles the network to be monitored; Compare the computational graph with the first static graph to obtain the network operators in the first static graph corresponding to each node in the computational graph; Compare the first static graph with the second static graph to obtain the network operators not adapted and implemented by the artificial intelligence processor.
2. The operator monitoring method based on a computational graph according to claim 1, characterized in that Each node in the first static graph is a network operator supported by the central processing unit.
3. The operator monitoring method based on a computational graph according to claim 1, characterized in that Each node in the second static graph is a network operator supported by the artificial intelligence processor or the central processing unit.
4. The operator monitoring method based on a computational graph according to claim 1, characterized in that Each network operator is configured with an exception capture mechanism.
5. The operator monitoring method based on a computational graph according to claim 4, characterized in that For the network operators not implemented by the artificial intelligence processor, if the exception capture mechanism works properly, a static graph mixed with central processing unit nodes and artificial intelligence processor nodes can be generated, and this mixed static graph is used as the second static graph.
6. An operator monitoring system based on a computational graph, characterized in that Including: A computational graph acquisition module configured to: obtain the computational graph of the network to be monitored; A first static graph capture module configured to: capture the first static graph returned after the central processing unit compiles the network to be monitored; A second static graph capture module configured to: capture the second static graph returned after the artificial intelligence processor compiles the network to be monitored; A comparison module configured to: compare the computational graph with the first static graph to obtain the network operators in the first static graph corresponding to each node in the computational graph; compare the first static graph with the second static graph to obtain the network operators not adapted and implemented by the artificial intelligence processor.
7. The operator monitoring system based on a computational graph according to claim 6, characterized in that Each network operator is configured with an exception capture mechanism.
8. The operator monitoring system based on a computational graph according to claim 7, characterized in that The second static graph capture module is further configured to: for the network operators not implemented by the artificial intelligence processor, if the exception capture mechanism works properly, a static graph mixed with central processing unit nodes and artificial intelligence processor nodes can be generated, and this mixed static graph is used as the second static graph.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that When the program is executed by a processor, it implements the steps in an operator monitoring method based on a computational graph as described in any one of claims 1-5.
10. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the steps in an operator monitoring method based on a computational graph as described in any one of claims 1-5.
Citation Information
Patent Citations
Neural network calculation graph processing method, computer storage medium and electronic equipment
CN111723935A
Optimization method and apparatus for computation graph, computer device, and storage medium
WO2021114757A1