Computational Graph Visualization Method, Device, Equipment and Medium Applied to AI Chip

By building a bidirectional node graph structure and visually displaying it, the problems of high complexity of algorithm model calculation graphs and diversity of model frameworks in AI chips are solved, and compatibility of diverse network model architectures and cost reduction of algorithm model analysis and optimization are achieved.

CN116227566BActive Publication Date: 2025-06-24SHANGHAI SUIYUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310221781.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-09
Publication Date
2025-06-24
Estimated Expiration
2043-03-09

AI Technical Summary

Technical Problem

When existing AI chips perform complex multi-node structure calculations, the algorithm model calculation graph is complex and there are many neural network model frameworks, which makes the existing visualization methods unable to be compatible with diverse network model architectures, increasing the cost of algorithm model analysis and optimization.

Method used

By obtaining the initial calculation diagram generated by any deep learning framework corresponding to the algorithm model under the specified AI chip, a two-way node graph structure matching the initial calculation diagram is built, the calculation diagram structure diagram, memory usage peak diagram, and node running time evaluation diagram are obtained, and visually displayed.

Benefits of technology

Compatibility with different deep learning frameworks is achieved, the cost of algorithm model analysis optimization is reduced, and the comprehensibility and optimization efficiency of computational graphs are improved through visual presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116227566B_ABST
    Figure CN116227566B_ABST
Patent Text Reader

Abstract

The present invention discloses a computational graph visualization method, device, equipment and medium applied to an AI chip. The method includes: obtaining an initial computational graph generated by any deep learning framework corresponding to an algorithm model under a specified AI chip; obtaining AI chip information, and constructing a first bidirectional node graph structure body matching the initial computational graph based on the AI chip information; obtaining a computational graph structure diagram, a peak memory usage value and a node operation evaluation graph according to the first bidirectional node graph structure body. Fast search is realized on the basis of the bidirectional node graph structure, hierarchical visualization of large-scale deep learning models is carried out, and the computational graph structure diagram, the peak memory usage value graph and the node operation evaluation graph are visually displayed. While adding various optimization information, the visualization of the computational graph is realized, compatibility with different deep learning frameworks can be achieved, and the cost of optimizing and analyzing algorithm models on the AI chip is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to artificial intelligence chip technology, and in particular, to a computational graph visualization method, device, equipment and medium applied to an AI chip. Background Art

[0002] Existing artificial intelligence (AI) chips are usually used for calculations of complex multi-node structures such as neural network models, thereby improving the computing efficiency of algorithm models corresponding to the multi-node structures.

[0003] However, when applying AI chips to perform calculations on complex multi-node structures, visualization technology is needed to support model analysis and optimization. However, due to the high complexity of the algorithm model calculation graph and the numerous neural network model frameworks, existing visualization methods are not compatible with diverse network model architectures, thereby increasing the cost of algorithm model analysis and optimization. Summary of the invention

[0004] Embodiments of the present invention provide a method, device, equipment and medium for visualizing a computational graph applied to an AI chip to achieve a visual display of the computational graph.

[0005] In a first aspect, an embodiment of the present invention provides a method for visualizing a computational graph applied to an AI chip, comprising: obtaining an initial computational graph generated by any deep learning framework corresponding to an algorithm model under a specified AI chip, wherein the initial computational graph generated by any deep learning framework includes nodes and node descriptions of the algorithm model;

[0006] Acquire AI chip information, and construct a first bidirectional node graph structure matching the initial calculation graph based on the AI ​​chip information;

[0007] According to the first bidirectional node graph structure, a calculation graph structure diagram, a memory usage peak diagram and a node running time evaluation diagram are obtained, and the calculation graph structure diagram, the memory usage peak diagram and the node running time evaluation diagram are visualized and displayed.

[0008] In a second aspect, an embodiment of the present invention provides a computational graph visualization device applied to an AI chip, comprising: an initial computational graph acquisition module, for acquiring an initial computational graph generated by any deep learning framework corresponding to an algorithm model under a specified AI chip, wherein the initial computational graph generated by any deep learning framework includes nodes of the algorithm model;

[0009] A first bidirectional node graph structure construction module, used to obtain AI chip information, and construct a first bidirectional node graph structure matching the initial calculation graph based on the AI ​​chip information;

[0010] A visualization display module, configured to obtain a computation graph structure diagram, a peak memory usage graph, and a node running time evaluation graph according to the first bidirectional node graph structure, and perform visualization display on the computation graph structure diagram, the peak memory usage graph, and the node running time evaluation graph.

[0011] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-described method is implemented.

[0012] In a fourth aspect, an embodiment of the present invention provides a storage medium storing computer-executable instructions, on which a computer program is stored. When the program is executed by a processor, the above-described method is implemented. Description of the Drawings

[0013] Figure 1 is a flowchart of a computation graph visualization method applied to an AI chip according to Embodiment 1 of the present invention;

[0014] Figure 2 is a schematic diagram of a computation graph structure diagram provided by Embodiment 1 of the present invention;

[0015] Figure 3 is a schematic diagram of a peak memory usage graph provided by Embodiment 1 of the present invention;

[0016] Figure 4 is a schematic diagram of a node running time evaluation graph provided by Embodiment 1 of the present invention;

[0017] Figure 5 is a flowchart of a computation graph visualization method applied to an AI chip according to Embodiment 2 of the present invention;

[0018] Figure 6 is a schematic structural diagram of a computation graph visualization device applied to an AI chip according to Embodiment 3 of the present invention;

[0019] Figure 7 is a schematic structural diagram of a computer device according to Embodiment 4 of the present invention. Detailed Description of the Embodiments

[0020] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, rather than all the structures.

[0021] Embodiment 1

[0022] Figure 1The following is a flowchart of a computational graph visualization method provided in the first embodiment of the present invention. This embodiment is applicable to the situation of visualizing a computational graph. The method can be executed by a computational graph visualization device applied to an AI chip, and the device can be implemented in software and / or hardware. The method includes:

[0023] Step S101, obtain an initial computational graph generated by any deep learning framework corresponding to an algorithm model under a specified AI chip.

[0024] Among them, the algorithm model can be a neural network model or a Bayesian model. The specific type of the algorithm model is not limited in this embodiment. And under the specified AI chip, an initial computational graph corresponding to the algorithm model can be generated through any deep learning framework. The initial computational graph generated through any deep learning framework can include nodes and node descriptions of the algorithm model. The initial computational graph obtained at this time can specifically be in the form of an intermediate representation source code. For example: %22 = "gather"(%arg60, %19) {dimension_numbers = collapsed_slice_dims = dense<0>:tensor<1xi64>, index_vector_dim = 1:i64, offset_dims = dense<1>:tensor<1xi64>, start_index_map = dense<0>:tensor<1xi64>, indices_are_sorted = false, name = "gather.278", op_id = 7:i64, slice_size = dense<[1,1024]>:tensor<2xi64>, tensor_split = dense<[4,1]>:tensor<1x2xi32>, unique_name = "common20_gather"}. Among them, gather is the name of the node, and the content contained in {} is the node description. Of course, only an example is given in this embodiment, and the number of nodes included in the initial computational graph and the content of the node description are not specifically limited.

[0025] Step S102, obtain AI chip information, and construct a first bidirectional node graph structure body that matches the initial computational graph based on the AI chip information.

[0026] Optionally, construct a first bidirectional node graph structure that matches the initial computation graph based on the AI chip information, including: performing backward traversal on each node in the initial computation graph to obtain node dependencies, and constructing a second bidirectional node graph structure based on the nodes, node descriptions, and node dependencies; obtaining the first bidirectional node graph structure that matches the initial computation graph according to the second bidirectional node graph structure and the AI chip information.

[0027] Optionally, obtaining the first bidirectional node graph structure that matches the initial computation graph according to the second bidirectional node graph structure and the AI chip information includes: simplifying the second bidirectional node graph structure according to a specified rule to obtain a simplified third bidirectional node graph structure, where the third bidirectional node graph structure includes simplified nodes, simplified node descriptions, and node dependencies; determining the hardware parameter information included in the AI chip information and the underlying hardware operator information based on the AI chip; obtaining the first bidirectional node graph structure that matches the initial computation graph according to the third bidirectional node graph structure, the hardware parameter information, and the underlying hardware operator information.

[0028] Optionally, simplifying the second bidirectional node graph structure according to a specified rule to obtain a simplified third bidirectional node graph structure includes: deleting nodes of a specified type in the second bidirectional node graph structure to obtain simplified nodes, where the specified type includes non-computation nodes; deleting node descriptions of a specified form in the second bidirectional node graph structure to obtain simplified node descriptions, where the specified form includes shape corresponding arrangements and information owned for downward mapping to the hardware; constructing a third bidirectional node graph structure based on the simplified nodes, simplified node descriptions, and node dependencies.

[0029] Optionally, obtaining the first bidirectional node graph structure that matches the initial computation graph according to the third bidirectional node graph structure, the hardware parameter information, and the underlying hardware operator information includes:

[0030] Performing node traversal calculation on the third bidirectional node graph structure according to the hardware parameter information and the underlying hardware operator information to obtain the node running information of each node, where the node running information includes estimated running time, storage space usage, computing power, and bandwidth; constructing the first bidirectional node graph structure based on the simplified nodes, simplified node descriptions, node dependencies, and node running information.

[0031] Specifically, in this embodiment, after obtaining the initial computational graph as shown above, when constructing the first bidirectional node graph structure that matches the initial computational graph, first, a backward traversal is performed on each node in the initial computational graph to obtain the node dependency relationship. For example, in the initial computational graph, there are four nodes, namely node a, node b, node c, and node d, and the node descriptions of the four nodes. By performing a backward traversal on each node, the node dependency relationship can be obtained. For example, if it is determined that the input of node a is the output of node b, then it is determined that node a depends on node b. The operation of node d requires the outputs of node b and node c, so node d depends on node b and node c. Thus, based on the traversal results, the dependency relationship of each node can be obtained, and based on the known nodes, node descriptions, and the node dependency relationship obtained through traversal, a second bidirectional node graph structure is constructed.

[0032] Among them, after obtaining the second bidirectional node graph structure, when constructing the first bidirectional node graph structure based on the second bidirectional node graph structure, first, the obtained second bidirectional node graph structure is simplified according to a specified rule. The specified rule can be to delete the nodes of a specified type and the node descriptions in a specified form in the second bidirectional node graph structure. For example, for the second bidirectional node graph structure that includes four nodes: node a, node b, node c, and node d, the node descriptions of the four nodes, and the node dependency relationship, nodes a, b, and c are computational nodes, while node d is a non-computational node. Then, according to the specified rule, the non-computational node can be deleted to obtain the simplified nodes: node a, node b, and node c. At the same time, the node descriptions in the second bidirectional node graph structure can also be deleted according to the specified rule to obtain the simplified node descriptions. For example, redundant node descriptions such as the corresponding arrangement of the shape shape, and the node descriptions of the internal operations of special nodes, such as the part to be deleted is for mapping the information owned by the underlying hardware. Of course, in this embodiment, only examples are given, and the specific content of the specified rule is not limited. Thus, based on the simplified nodes, the simplified node descriptions, and the node dependency relationship, a third bidirectional node graph structure is constructed.

[0033] It should be noted that since the hardware parameter information of the AI chip and the underlying hardware operator information based on the AI chip have been obtained previously, after obtaining the simplified third bidirectional node graph structure according to the second bidirectional node graph structure, the simplified third bidirectional node graph structure can be traversed and calculated based on the hardware parameter information and the underlying hardware operator information to calculate the node operation parameters of each node. The node operation information may include estimated operation time, storage space usage, computing power, bandwidth, etc. Since the hardware parameter information and the underlying hardware operator information indicate the operation environment information of each node, the operation environment information of the above-mentioned nodes can be obtained when the operation environment information is known. Thus, the first bidirectional node graph structure is constructed based on the known simplified nodes, simplified node descriptions, and node dependencies, as well as the node operation information obtained through the third bidirectional node graph structure.

[0034] Step S103, obtain a computation graph structure diagram, a peak memory usage graph, and a node operation time evaluation graph according to the first bidirectional node graph structure, and visually display the computation graph structure diagram, the peak memory usage graph, and the node operation time evaluation graph.

[0035] Optionally, obtaining a computation graph structure diagram, a peak memory usage graph, and a node operation time evaluation graph according to the first bidirectional node graph structure includes: sequentially connecting the simplified nodes according to the node dependencies to obtain an initial structure diagram; adding the simplified node descriptions and node operation information to the matching nodes in the initial structure diagram to construct a computation graph structure diagram; obtaining the adjacent nodes of each specified node according to the node dependencies, and calculating the peak memory usage of each specified node according to the storage space usage of the adjacent nodes and the storage space usage of the specified node, and constructing a peak memory usage graph according to the peak memory usage of each specified node; determining the operation time ratio of nodes with the same name according to the estimated operation time of each node, and constructing a node operation time evaluation graph according to the operation time ratio of nodes with the same name.

[0036] Specifically, in this embodiment, after obtaining the first bidirectional node graph structure, the simplified nodes will be sequentially connected according to the node dependencies to obtain an initial structure diagram. For example, when it is determined that the simplified nodes include: node a, node b, node c, node d, node e, and node f, the six nodes will be sequentially connected according to the obtained dependencies to obtain an initial structure diagram. Of course, only six nodes are used as an example in this embodiment, and in actual applications, the computation graph structure diagram usually contains a large number of nodes. At the same time, the simplified node descriptions and node operation information are added to the matching nodes in the initial structure diagram to construct a computation graph structure diagram, as Figure 2 shown in the schematic diagram of the obtained computation graph structure diagram.

[0037] Among them, in this embodiment, adjacent nodes of each specified node are obtained according to the node dependency relationship. For example, the adjacent node of node b is node a. At the same time, the storage space usage of node a, which is 10, and the storage space usage of node b, which is 8, are obtained according to the running information of each node. Since node a runs before node b, when calculating the peak memory usage of node b, specifically, the storage space usage of node b and the storage space usage of node a are added together to calculate that the peak memory usage of node b is 18. Of course, in this embodiment, only the calculation of the storage space usage of node b is used as an example for illustration. The method for solving the peak memory usage of other nodes is roughly the same, and will not be elaborated in this embodiment. And a peak memory usage graph of the computational graph structure is constructed according to the peak memory usage of each obtained node, as Figure 3 shown in the schematic diagram of the obtained peak memory usage graph. Of course, in this embodiment, it is only an example for illustration, and the specific numerical values of the peak memory usage of each node in the peak memory usage graph are not limited.

[0038] Among them, in this embodiment, a node running status list is obtained according to the running information of each node, as shown in Table 1 below:

[0039] Table 1

[0040]

[0041]

[0042] Among them, due to space limitations, only three node names are used as examples in Table 1, and in the Figure 2 shown computational graph structure, there will be multiple nodes with the same name. In this embodiment, the number of nodes with the same name included in the computational graph structure is not limited. For example, when it is determined that there are 83 convolutions, the running time ratio in the last column specifically refers to the total running time ratio of 83 convolution nodes. And a node running time evaluation graph as shown in Figure 4 can be constructed according to the running time ratio of nodes with the same name.

[0043] Optionally, visualize the computational graph structure diagram, the peak memory usage diagram, and the node running time evaluation diagram, including: receiving a splitting instruction for the computational graph structure diagram, where the splitting instruction includes the number of nodes; sequentially splitting the computational graph structure diagram according to the number of nodes to obtain split computational graph structure diagrams, where each split computational graph structure diagram is marked with a different identifier; receiving a visualization display instruction for the split computational graph structure diagram, where the visualization display instruction includes a specified identifier; and visualizing the split computational graph structure, the peak memory usage diagram, and the node running time evaluation diagram that match the specified identifier according to the visualization display instruction.

[0044] It is worth mentioning that in this embodiment, the obtained computational graph structure diagram and the peak memory usage diagram can be directly displayed to realize the visualization of the computational graph. In addition, in this embodiment, slicing technology can also be used to perform hidden display on the computational graph structure diagram.

[0045] In a specific implementation, after obtaining the computational graph structure diagram as shown in Figure 2 , receive a splitting instruction for the computational graph structure diagram. The splitting instruction includes the number of nodes, such as 2. Then the terminal will sequentially split the computational graph structure diagram as shown in Figure 2 to obtain split computational graph structure diagrams. For example, split computational graph structure diagram A including nodes a and b, split computational graph structure diagram B including nodes c and d, and split computational graph structure diagram C including nodes e and f. Therefore, when the received visualization display instruction for the split computational graph structure diagram includes the identifiers: A and B, then only split computational graph structure diagram A and split computational graph structure diagram B can be displayed according to the instruction, that is, only nodes a, b, c, and d are displayed, while nodes d and e are hidden, thus improving the flexibility of computational graph visualization.

[0046] This application obtains the initial computational graph generated by any deep learning framework corresponding to a specified AI chip by the algorithm model, and obtains the visual computational graph structure diagram and the peak memory usage diagram according to the structure body constructed by the initial computational graph. Thus, when visualizing the computational graph, it can be compatible with different deep learning frameworks, reducing the cost of algorithm model analysis and optimization.

[0047] Embodiment 2

[0048] Figure 5The flowchart of a computational graph visualization method for an AI chip provided in the second embodiment of this application. This embodiment is based on the above embodiment. After visually displaying the computational graph structure diagram, the peak memory usage graph, and the node running time evaluation graph, it further includes: receiving a search instruction from the user, searching the computational graph structure diagram according to the search instruction to obtain the target position of the key node in the computational graph structure diagram, and marking the key node at the target position.

[0049] Step S201, obtain the initial computational graph generated by any deep learning framework corresponding to the algorithm model under the specified AI chip.

[0050] Step S202, obtain the AI chip information, and construct a first bidirectional node graph structure that matches the initial computational graph based on the AI chip information.

[0051] Optionally, constructing a first bidirectional node graph structure that matches the initial computational graph based on the AI chip information includes: performing a backward traversal on each node in the initial computational graph to obtain the node dependency relationship, and constructing a second bidirectional node graph structure according to the node, the node description, and the node dependency relationship; obtaining the first bidirectional node graph structure that matches the initial computational graph according to the second bidirectional node graph structure and the AI chip information.

[0052] Optionally, obtaining the first bidirectional node graph structure that matches the initial computational graph according to the second bidirectional node graph structure and the AI chip information includes: simplifying the second bidirectional node graph structure according to a specified rule to obtain a simplified third bidirectional node graph structure, where the third bidirectional node graph structure includes the simplified nodes, the simplified node descriptions, and the node dependency relationship; determining the hardware parameter information included in the AI chip information and the underlying hardware operator information based on the AI chip; obtaining the first bidirectional node graph structure that matches the initial computational graph according to the third bidirectional node graph structure, the hardware parameter information, and the underlying hardware operator information.

[0053] Step S203, obtain the computational graph structure diagram, the peak memory usage graph, and the node running time evaluation graph according to the first bidirectional node graph structure, and visually display the computational graph structure diagram, the peak memory usage graph, and the node running time evaluation graph.

[0054] Optionally, obtain the computational graph structure diagram, the peak memory usage diagram, and the node running time evaluation diagram according to the first bidirectional node graph structure, including: sequentially connect the simplified nodes according to the node dependencies to obtain the initial structure diagram; add the descriptions of the simplified nodes and the node running information to the corresponding nodes in the initial structure diagram to construct the computational graph structure diagram; obtain the adjacent nodes of each specified node according to the node dependencies, and calculate the peak memory usage of each specified node according to the storage space usage of the adjacent nodes and the storage space usage of the specified node, and construct the peak memory usage diagram according to the peak memory usage of each specified node; determine the running time ratio of the nodes with the same name according to the estimated running time of each node, and construct the node running time evaluation diagram according to the running time ratio of the nodes with the same name.

[0055] Optionally, visually display the computational graph structure diagram and the peak memory usage diagram, including: receive a splitting instruction for the computational graph structure diagram, where the splitting instruction includes the number of nodes; sequentially split the computational graph structure diagram according to the number of nodes to obtain the split computational graph structure diagrams, where each split computational graph structure diagram is marked with a different identifier; receive a visual display instruction for the split computational graph structure diagrams, where the visual display instruction includes a specified identifier; visually display the split computational graph structure, the peak memory usage diagram, and the node running time evaluation diagram that match the specified identifier according to the visual display instruction.

[0056] Step S204, receive the user's search instruction, search the computational graph structure diagram according to the search instruction to obtain the target position of the key node in the computational graph structure diagram, and mark the key node at the target position.

[0057] Specifically, in this embodiment, after visually displaying the computational graph structure diagram, since each node name is included in the computational graph structure diagram, when receiving a search instruction including the key node name, where the search instruction includes the key node name. For example, if the search instruction includes node a, the computational graph structure diagram can be searched according to the search instruction to obtain the target position of node a in the computational graph structure diagram, and node a is marked at the target position. For example, it can be marked with a circle or a star. In this embodiment, the specific marking method for the key node is not limited, as long as the key node to be searched can be located, it is within the protection scope of this application, and it is not limited in this embodiment. By automatically positioning the key node in the computational graph structure diagram, it is convenient for the user to quickly find the nodes that need to be focused on according to the actual needs from the computational graph structure diagram with complex structure and numerous nodes, thereby improving the user's search efficiency.

[0058] This application obtains the initial computation graph generated by any deep learning framework corresponding to an algorithm model on a specified AI chip, and obtains a visual computation graph structure diagram and a peak memory usage graph based on the structure constructed from the initial computation graph. Thus, when visualizing the computation graph, compatibility with different deep learning frameworks can be achieved, reducing the cost of algorithm model analysis and optimization. By automatically locating key nodes in the computation graph structure diagram, it is convenient for users to quickly find the nodes that need to be focused on according to actual requirements from the computation graph structure diagram with complex structures and numerous nodes, thereby improving the user's search efficiency.

[0059] Embodiment III

[0060] Figure 6 FIG. 7 is a schematic structural diagram of a computation graph visualization device applied to an AI chip provided in Embodiment III of the present invention. This device can execute the computation graph visualization method applied to an AI chip involved in the above embodiments. This device can be implemented in software and / or hardware manners, such as Figure 6 As shown, the computation graph visualization device applied to an AI chip includes: an initial computation graph acquisition module 310, a first bidirectional node graph structure construction module 320, and a visualization display module 330.

[0061] The initial computation graph acquisition module 310 is configured to obtain the initial computation graph generated by any deep learning framework corresponding to an algorithm model on a specified AI chip, where the initial computation graph generated by any deep learning framework includes nodes of the algorithm model;

[0062] The first bidirectional node graph structure construction module 320 is configured to obtain AI chip information and construct a first bidirectional node graph structure matching the initial computation graph based on the AI chip information;

[0063] The visualization display module 330 is configured to obtain a computation graph structure diagram, a peak memory usage graph, and a node running time evaluation graph according to the first bidirectional node graph structure, and visually display the computation graph structure diagram, the peak memory usage graph, and the node running time evaluation graph.

[0064] Optionally, the first bidirectional node graph structure construction module includes:

[0065] The second bidirectional node graph structure construction sub-module is configured to perform backward traversal on each node in the initial computation graph to obtain node dependency relationships, and construct a second bidirectional node graph structure according to the nodes, node descriptions, and node dependency relationships;

[0066] The first bidirectional node graph structure acquisition sub-module is configured to obtain a first bidirectional node graph structure matching the initial computation graph according to the second bidirectional node graph structure and the AI chip information.

[0067] Optionally, the first dual - node graph structure acquisition sub - module includes:

[0068] The third dual - node graph structure acquisition subunit is used to simplify the second dual - node graph structure according to specified rules to obtain a simplified third dual - node graph structure, where the third dual - node graph structure includes simplified nodes, simplified node descriptions, and node dependency relationships;

[0069] The hardware information determination subunit is used to determine the hardware parameter information included in the AI chip information and the underlying hardware operator information based on the AI chip;

[0070] The first dual - node graph structure acquisition subunit is used to obtain a first dual - node graph structure that matches the initial computational graph according to the third dual - node graph structure, hardware parameter information, and underlying hardware operator information.

[0071] Optionally, the third dual - node graph structure acquisition subunit is used to delete nodes of a specified type in the second dual - node graph structure to obtain simplified nodes, where the specified type includes non - computational nodes;

[0072] Delete node descriptions in the second dual - node graph structure in a specified form to obtain simplified node descriptions, where the specified form includes shape - corresponding arrangements and information owned for downward mapping to the hardware;

[0073] Construct a third dual - node graph structure according to the simplified nodes, simplified node descriptions, and node dependency relationships.

[0074] Optionally, the first dual - node graph structure acquisition subunit is used to perform node traversal calculations on the third dual - node graph structure according to the hardware parameter information and the underlying hardware operator information to obtain node operation information for each node, where the node operation information includes estimated running time, storage space usage, computing power, and bandwidth;

[0075] Construct a first dual - node graph structure according to the simplified nodes, simplified node descriptions, node dependency relationships, and node operation information.

[0076] Optionally, the visualization display module is used to sequentially connect the simplified nodes according to the node dependency relationships to obtain an initial structure diagram;

[0077] Add the simplified node descriptions and node operation information to the matching nodes in the initial structure diagram to construct a computational graph structure diagram;

[0078] Obtain the adjacent nodes of each specified node according to the node dependency relationship, and calculate the peak memory usage of each specified node based on the storage space usage of the adjacent nodes and the storage space usage of the specified node. Construct a peak memory usage graph based on the peak memory usage of each specified node;

[0079] Determine the running time ratio of nodes with the same name according to the estimated running time of each node, and construct a node running time evaluation graph based on the running time ratio of nodes with the same name.

[0080] Optionally, the visualization display module is further configured to receive a splitting instruction for the computational graph structure diagram, where the splitting instruction includes the number of nodes;

[0081] Sequentially split the computational graph structure diagram according to the number of nodes to obtain a split computational graph structure diagram, where each split computational graph structure diagram is marked with a different identifier;

[0082] Receive a visualization display instruction for the split computational graph structure diagram, where the visualization display instruction includes a specified identifier;

[0083] Visualize the split computational graph structure that matches the specified identifier and the peak memory usage graph according to the visualization display instruction.

[0084] Optionally, the device further includes a positioning module, configured to receive a search instruction from a user, where the search instruction includes the name of a key node;

[0085] Search the computational graph structure diagram according to the search instruction to obtain the target position of the key node in the computational graph structure diagram, and mark the key node at the target position.

[0086] Embodiment 4

[0087] Figure 7 The structural schematic diagram of a computer device provided in Embodiment 4 of the present invention is shown as Figure 7 shown. The computer device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the computer device can be one or more, Figure 7 taking one processor 610 as an example; the processor 610, the memory 620, the input device 630, and the output device 640 in the computer device can be connected through a bus or other means, Figure 7 taking connection through a bus as an example.

[0088] The memory 620, being a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the computational graph visualization method applied to the AI chip in the embodiments of the present invention. The processor 610 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 620, that is, to implement the above-mentioned computational graph visualization applied to the AI chip.

[0089] A computational graph visualization method applied to an AI chip, comprising:

[0090] Obtain an initial computational graph generated by any deep learning framework corresponding to an algorithm model under a specified AI chip, wherein the initial computational graph generated by any deep learning framework includes nodes and node descriptions of the algorithm model;

[0091] Obtain AI chip information, and construct a first bidirectional node graph structure that matches the initial computational graph based on the AI chip information;

[0092] Obtain a computational graph structure diagram, a peak memory usage diagram, and a node running time evaluation diagram according to the first bidirectional node graph structure, and visually display the computational graph structure diagram, the peak memory usage diagram, and the node running time evaluation diagram.

[0093] The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 620 may further include a memory remotely set relative to the processor 610, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0094] The input device 630 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device. The output device 640 may include a display device such as a display screen.

[0095] Embodiment Five

[0096] Embodiment Five of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a computational graph visualization method applied to an AI chip when executed by a computer processor;

[0097] Obtain the initial computation graph generated by any deep learning framework corresponding to the algorithm model under the specified AI chip, where the initial computation graph generated by any deep learning framework includes the nodes and node descriptions of the algorithm model;

[0098] Obtain AI chip information, and construct a first bidirectional node graph structure that matches the initial computation graph based on the AI chip information;

[0099] Obtain the computation graph structure diagram, the peak memory usage graph, and the node running time evaluation graph according to the first bidirectional node graph structure, and visually display the computation graph structure diagram, the peak memory usage graph, and the node running time evaluation graph.

[0100] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the above method operations, and can also execute related operations in the computation graph visualization method applied to the AI chip provided by any embodiment of the present invention.

[0101] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.

[0102] It should be noted that in the above embodiments, the included units and modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0103] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A computational graph visualization method applied to an AI chip, characterized in that, Including: Obtain an initial computation graph generated by any deep learning framework corresponding to an algorithm model under a specified AI chip, where the initial computation graph generated by the any deep learning framework includes nodes and node descriptions of the algorithm model; Obtain AI chip information, and construct a first bidirectional node graph structure matching the initial computation graph based on the AI chip information; Obtain a computation graph structure diagram, a memory usage peak diagram, and a node running time evaluation diagram according to the first bidirectional node graph structure, and visually display the computation graph structure diagram, the memory usage peak diagram, and the node running time evaluation diagram; The constructing the first bidirectional node graph structure matching the initial computation graph includes: performing backward traversal on each node in the initial computation graph to obtain node dependencies, and constructing a second bidirectional node graph structure according to the nodes, the node descriptions, and the node dependencies; Obtain the first bidirectional node graph structure matching the initial computation graph according to the second bidirectional node graph structure and the AI chip information; The obtaining the first bidirectional node graph structure matching the initial computation graph according to the second bidirectional node graph structure and the AI chip information includes: simplifying the second bidirectional node graph structure according to a specified rule to obtain a simplified third bidirectional node graph structure, where the third bidirectional node graph structure includes simplified nodes, simplified node descriptions, and node dependencies; Determine the hardware parameter information included in the AI chip information and the underlying hardware operator information based on the AI chip; Obtain the first bidirectional node graph structure matching the initial computation graph according to the third bidirectional node graph structure, the hardware parameter information, and the underlying hardware operator information.

2. The method according to claim 1, characterized in that, The simplifying the second bidirectional node graph structure according to a specified rule to obtain a simplified third bidirectional node graph structure includes: Delete nodes of a specified type in the second bidirectional node graph structure to obtain the simplified nodes, where the specified type includes non-computation nodes; Delete node descriptions of a specified form in the second bidirectional node graph structure to obtain the simplified node descriptions, where the specified form includes shape corresponding arrangements and information owned for downward mapping of the hardware; Construct the third bidirectional node graph structure according to the simplified nodes, the simplified node descriptions, and the node dependencies.

3. The method according to claim 1, characterized in that The obtaining the first bidirectional node graph structure matching the initial computation graph according to the third bidirectional node graph structure, the hardware parameter information, and the underlying hardware operator information includes: Perform node traversal calculation on the third bidirectional node graph structure according to the hardware parameter information and the underlying hardware operator information to obtain node running information of each node, where the node running information includes estimated running time, storage space usage, computing power, and bandwidth; Construct the first bidirectional node graph structure according to the simplified nodes, the simplified node descriptions, the node dependency relationships, and the node running information.

4. The method according to claim 3, wherein Obtaining the computational graph structure diagram, the peak memory usage graph, and the node running time evaluation graph according to the first bidirectional node graph structure includes: Sequentially connect the simplified nodes according to the node dependency relationships to obtain an initial structure diagram; Add the simplified node descriptions and the node running information to the corresponding nodes in the initial structure diagram to construct the computational graph structure diagram; Obtain the adjacent nodes of each specified node according to the node dependency relationships, and calculate the peak memory usage of each specified node according to the storage space usage of the adjacent nodes and the storage space usage of the specified node. Construct the peak memory usage graph according to the peak memory usage of each specified node; Determine the running time ratio of nodes with the same name according to the estimated running time of each node, and construct the node running time evaluation graph according to the running time ratio of nodes with the same name.

5. The method according to claim 1, wherein Visualizing the computational graph structure diagram, the peak memory usage graph, and the node running time evaluation graph includes: Receive a splitting instruction for the computational graph structure diagram, where the splitting instruction includes the number of nodes; Sequentially split the computational graph structure diagram according to the number of nodes to obtain split computational graph structure diagrams, and each of the split computational graph structure diagrams is marked with a different identifier; Receive a visualization display instruction for the split computational graph structure diagram, where the visualization display instruction includes a specified identifier; Visualize the split computational graph structure, the peak memory usage graph, and the node running time evaluation graph that match the specified identifier according to the visualization display instruction.

6. The method according to claim 1, characterized in that, After obtaining the computational graph structure diagram, the peak memory usage graph, and the node running time evaluation graph according to the first bidirectional node graph structure, it further includes: Receive a search instruction from the user, where the search instruction includes the name of a key node; Search the computational graph structure diagram according to the search instruction to obtain the target position of the key node in the computational graph structure diagram, and mark the key node at the target position.

7. A computational graph visualization device applied to an AI chip, characterized in that, It includes: An initial computational graph acquisition module for acquiring an initial computational graph generated by any deep learning framework corresponding to an algorithm model under a specified AI chip, where the initial computational graph generated by the any deep learning framework includes the nodes of the algorithm model; A first bidirectional node graph structure construction module for acquiring AI chip information and constructing a first bidirectional node graph structure matching the initial computational graph based on the AI chip information; A visualization display module for obtaining a computational graph structure diagram, a peak memory usage graph, and a node running time evaluation graph according to the first bidirectional node graph structure, and visualizing the computational graph structure diagram, the peak memory usage graph, and the node running time evaluation graph. The first bidirectional node graph structure construction module is used to perform backward traversal on each node in the initial computation graph to obtain node dependency relationships, and construct a second bidirectional node graph structure according to the node, the node description, and the node dependency relationships; Obtain a first bidirectional node graph structure that matches the initial computation graph according to the second bidirectional node graph structure and the AI chip information; The first bidirectional node graph structure construction module is further used to obtain a first bidirectional node graph structure that matches the initial computation graph according to the second bidirectional node graph structure and the AI chip information, including: simplifying the second bidirectional node graph structure according to a specified rule to obtain a simplified third bidirectional node graph structure, where the third bidirectional node graph structure includes simplified nodes, simplified node descriptions, and node dependency relationships; Determine the hardware parameter information included in the AI chip information and the underlying hardware operator information based on the AI chip; Obtain a first bidirectional node graph structure that matches the initial computation graph according to the third bidirectional node graph structure, the hardware parameter information, and the underlying hardware operator information.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1-6 is implemented.

9. A storage medium storing computer-executable instructions, on which a computer program is stored, characterized in that, When the program is executed by the processor, the method described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Deep learning calculation method and device, chip and medium

    CN113326137A

  • Optimization method for executing deep learning tasks in distributed mode and distributed system

    CN115543639A