Computation graph representation method and device suitable for hardware design space search

By introducing multiple types of edges and separate computing nodes and data nodes, the existing computing graph representation solutions are solved, and the problems of insufficient expression capabilities and cumbersome debugging in hardware design space search are achieved, and more efficient hardware design space search and flexible resource management are achieved.

CN120373425APending Publication Date: 2025-07-25BEIJING YIXIN YIYU MICROELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510361790.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing computing graph representation schemes have problems such as insufficient edge type function, lack of support for complex computing graph structures, lack of dynamic trace support, and coupling of computing and data access in hardware design space search, resulting in limited expression capabilities, cumbersome debugging process, insufficient flexibility and resource management difficulties.

Method used

A variety of types of edges are introduced, including data edges, index edges and mask access edges, which support dynamic trace recording, and separate computing nodes from data nodes, providing dynamic trace support and flexible access modes to achieve decoupling of computing and data access.

Benefits of technology

It improves the expression ability and resource utilization of computing graphs, simplifies the debugging process, improves the efficiency and flexibility of hardware design space search, and is suitable for hardware design space search for static and dynamic neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373425A_ABST
    Figure CN120373425A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a computational graph representation method and device suitable for hardware design space search. The method comprises the steps of constructing a computational graph based on a deep learning program; and performing hardware design space search based on the calculation graph to determine a hardware design scheme meeting a preset performance index. Wherein various types of edges are introduced into the computational graph, and computational nodes and data nodes are separately designed. And the computational graph supports dynamic traces. The method has the advantages that the expression ability is obviously enhanced, the debugging and hardware design space search efficiency is greatly improved, the flexibility advantage is outstanding, and the application scene is wide. Moreover, the method has been successfully applied to a project of a special accelerator automatic generation technology, and is specifically applied to an LLMCompler module and a DSE module in the project.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and hardware design space search, and particularly relates to a computational graph representation method and device suitable for hardware design space search. Background Art

[0002] In the fields of deep learning and hardware design space search, the representation of computational graphs is crucial for model construction, optimization, and hardware adaptation. Currently, there are various computational graph representation schemes, among which the more typical ones are ONNX (Open Neural Network Exchange), PyTorch TorchScript, etc.

[0003] 1. ONNX (Open Neural Network Exchange), Open Neural Network Exchange

[0004] ONNX is an open standard format designed to promote model interoperability between different deep learning frameworks. It describes the computational flow of the model by defining a series of operators and tensors. The computational graph of ONNX consists of nodes and edges, where nodes represent operators and edges represent data streams. Users can export models from frameworks that support ONNX (such as PyTorch, TensorFlow) to the ONNX format. When constructing the computational graph, it is necessary to manually specify each operator and its dependencies, and build a complete computational graph description by defining the attributes of the operator and the shape and data type of the tensors.

[0005] 2. PyTorch TorchScript

[0006] TorchScript is an intermediate representation form of PyTorch for converting a dynamic computational graph into a static computational graph. It supports generating a static computational graph through tracing or scripting. In the tracing mode, the model tracing module automatically generates a static computational graph by tracing the forward propagation process of the model; the scripting mode requires users to write TorchScript code to explicitly define the computational logic of the model. The generated computational graph can be exported as a TorchScript file and deployed on different platforms.

[0007] In addition, there are also some lightweight solutions, such as MLIR (Multi-Level Intermediate Representation). MLIR is a general intermediate representation framework that supports multiple levels of abstraction. It can be used to represent various types of computational graphs, including deep learning models, traditional compiler optimizations, etc. Users need to convert models from different sources (such as TensorFlow, PyTorch) into the MLIR intermediate representation. During the conversion process, it is necessary to deeply understand the internal mechanism of MLIR and re-describe and convert the model according to the defined syntax and semantic rules.

[0008] Although the above existing computational graph representation schemes have achieved certain success in the representation and optimization of deep learning models, there are still some significant drawbacks in practical applications, especially in the task of hardware design space search:

[0009] 1. Insufficient functionality of edge types

[0010] The edges in the existing computational graph representation schemes usually only represent data flow and lack support for different types of edges. For example, when dealing with complex data access patterns such as indirect access, masking operations, reduction operations, etc., the existing edge types cannot accurately describe the specific behaviors of these operations. This leads to limited expressiveness of the computational graph, especially in dealing with complex scenarios such as irregular data structures and sparse matrices, where it is difficult to achieve efficient memory access and resource management.

[0011] 2. Lack of support for complex computational graph structures

[0012] When dealing with complex computational graph structures (such as conditional branches, loops, indirect access, etc.), the existing computational graph representation schemes usually adopt relatively simple linear representation methods and cannot effectively express complex control flows and data dependencies. This not only limits the expressiveness of the computational graph but also causes difficulties in performance optimization, especially in cases involving dynamic control flows.

[0013] 3. Lack of support for dynamic traces

[0014] When the existing computational graph representation schemes are designed, they focus more on the representation of static computational graphs and ignore the importance of dynamic traces. For complex computational graphs that need to be frequently debugged and optimized, the lack of dynamic trace support will make the debugging process extremely cumbersome. Especially in cases involving dynamic control flows such as conditional branches and indirect access, it is difficult for developers to monitor the execution path of the computational graph in real time, discover potential performance bottlenecks, and perform targeted optimizations. When conducting hardware design space search, due to the lack of dynamic execution information, it is impossible to accurately evaluate the impact of different hardware configurations on the execution performance of the computational graph.

[0015] 4. Computational and data access coupling

[0016] In existing computational graph representation schemes, computations and data accesses are usually mixed together, resulting in a lack of modularity in the design of computational graphs. This coupled design leads to insufficient flexibility. When dealing with complex computational graph structures, it is difficult to flexibly adjust the computational logic or data access patterns. In terms of resource management, due to the lack of clear separation between computations and data accesses, it is difficult to efficiently manage computational resources and memory resources, resulting in low resource utilization. When debugging and optimizing computational graphs, it is also difficult to clearly distinguish issues related to computations and data accesses. Summary of the Invention

[0017] Aiming at the technical deficiencies mentioned in the background art, the purpose of the embodiments of the present invention is to provide a computational graph representation method and device applicable to hardware design space search.

[0018] To achieve the above objective, in a first aspect, the embodiments of the present invention provide a computational graph representation method applicable to hardware design space search, including:

[0019] Constructing a computational graph based on a deep learning program; the computational graph includes nodes and edges; the edges include data edges, index edges, and mask access edges; the data edges are used to represent the data flow between computational nodes; the index edges are indices for indirect access, which can carry index information to indicate the specific access location or method of data; the mask access edges are used to selectively process part of the data according to mask conditions;

[0020] Performing a hardware design space search based on the computational graph to determine a hardware design solution that meets preset performance indicators.

[0021] As a specific implementation manner of the present application, the nodes include IO nodes, data nodes, computational nodes, constant nodes, and branch nodes; based on the nodes, the computational graph can provide dynamic trace support, including fine-grained dynamic trace recording, branch jump information recording, and indirect access address recording.

[0022] As a specific implementation manner of the present application, the data nodes and computational nodes adopt a separated design; the computational nodes are limited to computations; the data nodes are only used for data storage or access, and the provided access modes include: affine access, indirect access, slice access, merge access, and mask access.

[0023] As a specific implementation manner of the present application, the fine-grained dynamic trace recording is specifically: when each operator is executed, record the input, output, and execution timestamp;

[0024] The specific recording of branch jump information is as follows: For a computational graph containing conditional branches, branch nodes are used to record the specific situation of each branch jump, including the conditional result, the selected path, and the execution timestamp.

[0025] The specific recording of indirect access addresses is as follows: For operations involving indirect access, the actual accessed address and mask value are recorded.

[0026] In a second aspect, an embodiment of the present application further provides a computational graph representation device applicable to hardware design space search, including:

[0027] A construction unit for constructing a computational graph based on a deep learning program; the computational graph includes nodes and edges; the edges include data edges, index edges, and mask access edges; the data edges are used to represent the data flow between computational nodes; the index edges are indexes for indirect access and can carry index information to indicate the specific access location or method of data; the mask access edges are used to selectively process part of the data according to mask conditions.

[0028] A search unit for performing hardware design space search based on the computational graph to determine a hardware design scheme that meets preset performance indicators.

[0029] Implementing the embodiments of the present invention has the following advantages:

[0030] First, stronger expressive ability. By introducing INDEX edges, the present invention can flexibly handle complex memory access patterns, improving computational efficiency and resource utilization. Especially for complex scenarios such as sparse matrix operations, dynamic data streams, and multi-input operations, the advantages of the present invention are particularly obvious.

[0031] Second, more efficient search and more convenient debugging. The support for dynamic traces simplifies the debugging process. Developers can monitor the execution path of the computational graph in real time during runtime, discover potential performance bottlenecks, and perform targeted optimizations. In the hardware design space search task, the rich information provided by dynamic traces also plays a key role. Traditional search methods often blindly explore in a large number of invalid design spaces due to the lack of in-depth understanding of the dynamic execution of the computational graph, resulting in low efficiency. However, the present invention can accurately analyze the performance of different computational operations and data access patterns in hardware execution based on the information recorded by dynamic traces.

[0032] For example, by analyzing the address distribution and frequency of indirect access operations and combining with the cache characteristics of the hardware, the search algorithm can be guided to more specifically screen out hardware configurations that can optimize cache utilization. For a computational graph with frequent conditional branches, according to the recorded branch jump situation, predict the resource requirements of the hardware under different execution paths, so as to quickly find a suitable hardware architecture, greatly reducing the search time and improving the search efficiency.

[0033] III. Higher flexibility for adapting to hardware design space search. In the field of hardware design space search, the separate design of computing nodes and data nodes brings significant advantages. This separation makes the computing graph highly modular. In traditional solutions, computing and data access are closely related. When conducting hardware design space search, optimizing the computing is likely to cause data access problems, and vice versa, greatly limiting the flexibility of the search. However, the present invention decouples the two, enabling developers to optimize the computing logic and data access patterns independently. When searching for a computing graph adapted to specific hardware, such as when conducting design space search for a chip with a specific cache architecture, developers can flexibly adjust the access pattern of data nodes according to the cache characteristics of the chip, such as adopting a block loading strategy more suitable for the cache, and independently optimize the algorithms of computing nodes, such as selecting more efficient computing operators, so as to more efficiently search for the computing graph configuration that meets the hardware performance requirements.

[0034] IV. Wide range of application scenarios. The present invention is not only applicable to representing static neural networks, but also to dynamic neural networks, with obvious advantages. At the same time, the present invention is not only applicable to tasks such as hardware design space search, but also to the inference tasks of traditional deep learning networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art.

[0036] Figure 1 is the flowchart of the computing graph representation method applicable to hardware design space search provided by the embodiment of the present invention;

[0037] Figure 2 is the overall structure diagram of the present invention;

[0038] Figure 3 is the schematic diagram of constructing a computing graph and conducting hardware design space search;

[0039] Figure 4 is the structure diagram of the computing graph representation device applicable to hardware design space search provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0042] Term Explanation:

[0043] 1. ONNX: Open Neural Network Exchange, an open neural network exchange.

[0044] 2. PyTorch: An open-source deep learning framework developed by the artificial intelligence research team of Facebook.

[0045] 3. TensorFlow: An open-source deep learning framework developed by Google.

[0046] 4. MLIR: Multi-Level Intermediate Representation, a multi-level intermediate representation.

[0047] 5. LLMCompiler: A compilation module based on large language models.

[0048] 6. DSE: Design Space Exploration, a design space search.

[0049] The inventive points involved in the present invention can be briefly described as follows:

[0050] Inventive Point 1: Enhanced Edge Type Support

[0051] In view of the lack of support for different types of edges in the existing computational graph representation scheme, the present invention introduces multiple types of edges, which can accurately describe different types of data access patterns and control flows, and solves the problem of insufficient edge type functions in the existing computational graph representation scheme. Specifically, the present invention supports the following two enhanced edge types:

[0052] DATA edge, used to represent the data flow between computational nodes. Each DATA edge connects two nodes, indicating that the data output from one node flows to another node as input, which is the most common edge type in the computational graph. It is applicable to all scenarios involving data transfer, such as data transfer between layers in a neural network, input and output in matrix multiplication, etc.

[0053] INDEX edge, an index for indirect access, which can carry additional index information to indicate the specific access location or method of data. Indirect access performs data access through an index array and is applicable to scenarios such as sparse matrices and irregular data structures.

[0054] MASK edge, for masked access. Masked access selectively processes part of the data according to the mask condition, and is applicable to fields such as image processing and natural language processing.

[0055] The INDEX edge and the MASK edge are used in conjunction with the indirect access mode and the masked access mode of the DataNode, greatly enhancing the expressive power of the computational graph, being able to flexibly handle complex memory access patterns, and improving the computational efficiency and resource utilization rate. It is more suitable for scenarios that need to process complex data access patterns, such as sparse matrix operations and dynamic data streams.

[0056] That is, the present invention introduces multiple types of edges, which can more accurately describe different types of data access patterns and control flows, and solves the problem of the single function of edge types in existing computational graph representation schemes. Specifically, in addition to supporting standard data flow edges, the present invention also supports INDEX edges and MASK edges. The INDEX edge can carry additional index information indicating the specific access location or method of the data. The masking operation selectively processes part of the data according to the mask condition.

[0057] By introducing additional types of edges, the present invention can flexibly handle more complex memory access patterns. Especially when dealing with complex scenarios such as sparse matrix operations and dynamic data streams, the advantages of the present invention are particularly obvious, thereby improving the expressive power of the computational graph.

[0058] Innovation Point Two: Support for Dynamic Traces

[0059] Aiming at the problem of ignoring the importance of dynamic traces in existing computational graph representation schemes, the present invention provides comprehensive support for dynamic traces, allowing users to record and analyze the execution path of the computational graph at runtime, helping developers better understand and optimize the computational graph. Specifically, the present invention can record information such as its input, output, and execution timestamp when each operator is executed. These dynamic traces can not only be used for debugging but also for tasks such as hardware design space search. For a computational graph containing conditional branches, the present invention can record the specific situation of each branch jump, including the condition result, the selected path, and the execution timestamp. This enables developers to clearly understand the behavior of the computational graph under different inputs. And for operations involving indirect access, the present invention can record the actual accessed address and mask value, helping developers analyze the memory access pattern and optimize the cache utilization rate and memory bandwidth.

[0060] Innovation Point Three: Separation of Computational Nodes and Data Nodes

[0061] In view of the problem that the mixing of computation and data access in the existing computational graph representation scheme leads to a lack of modularity in the design of the computational graph, the present invention separates computational nodes and data nodes. This design brings multiple advantages, significantly enhancing the flexibility and optimization potential of the computational graph. Specifically, the advantages include the following four points:

[0062] 1. Clear role division. The computational node (ComputeNode) is specifically responsible for performing specific computational operations, such as addition, multiplication, activation functions, etc. The responsibility of the computational node is limited to computation and does not involve the storage or access method of data. The data node (DataNode) is specifically responsible for describing the access behavior of data, such as input shape, output shape, data type, access mode, etc. The responsibility of the data node is to ensure the correct transfer and access of data and does not involve specific computational logic.

[0063] 2. Decoupling of computation and data access. By separating computational nodes and data nodes, the present invention achieves complete decoupling of computation and data access. This makes the design of the computational graph more modular, and developers can independently optimize computational logic and data access patterns without affecting each other. For example, when optimizing computational logic, developers can focus on improving the efficiency of the algorithm without worrying about changes in the data access pattern; similarly, when optimizing the data access pattern, developers can focus on improving memory access efficiency without worrying about the complexity of the computational logic.

[0064] 3. Support for flexible access patterns. The data node supports multiple access patterns, such as affine access (AFFINE), indirect access (INDIRECT), slice access (SLICE), merge access (MERGE), mask access (MASK), etc. This flexibility enables the computational graph to adapt to various complex data access requirements, especially in scenarios such as processing irregular data structures, sparse matrices, dynamic data streams, etc. For example, when processing image padding, the data node can easily achieve boundary expansion through the affine access pattern; when processing matrix transpose, the data node can efficiently rearrange data through the multi-dimensional access pattern.

[0065] 4. Efficient resource management. Due to the separation of computational nodes and data nodes, the present invention can manage computational resources and memory resources more efficiently. The computational node can dynamically allocate computational resources as needed, while the data node can optimize the memory layout and access path according to the data access pattern. For example, when processing large-scale data, the data node can adopt a block loading method to reduce memory occupancy; when processing small-scale data, the data node can adopt a sequential access pattern to improve the cache hit rate.

[0066] Please refer to Figure 1, an embodiment of the present invention provides a computational graph representation method applicable to hardware design space search, including:

[0067] S1, constructing a computational graph based on a deep learning program.

[0068] Please refer to Figure 2 , the computational graph mainly includes five types of nodes and three types of edges, which are introduced separately below:

[0069] I. IO node (IONode), used to represent input and output data.

[0070]

[0071] Among them, the data types supported by data_type include DataType.INT8 / 16 / 32 / 64, DataType.UINT8 / 16 / 32 / 64, DataType.FLOAT, DataType.DOUBLE, and DataType.BOOL.

[0072] II. Data node (DataNode), used to describe the access behavior of data.

[0073]

[0074] The access patterns (AccessPattern) of the data node include AFFINE (affine access), INDIRECT (indirect access), SLICE (slice access), MERGE (merge access), and MASK (mask access). The following is a detailed explanation of each access pattern:

[0075] 1. AFFINE (affine access) is used to represent a regular array access pattern and supports index offset (index_offset). Example: A[i+1][j], A[2*i][j-1].

[0076] 2. INDIRECT (indirect access) is accessed through an index array, and an additional index array needs to be provided. Example: A[idx[i]].

[0077] 3. SLICE (slice access) accesses a continuous data segment and needs to specify the start position (start) and end position (end). Example: A[start:end].

[0078] 4. MERGE (merge access) merges multiple inputs into one output and needs to specify the shape and offset of each input. Example: Connecting multiple arrays into a large array.

[0079] 5. MASK (Mask Access) selectively processes part of the data according to the mask conditions. It is applicable to scenarios where data filtering or selective processing is required, such as selecting specific area pixels in image processing or ignoring certain words in natural language processing.

[0080] The access expression of the data node is used to describe the specific pattern of data access, mainly for the pattern_expr parameter of the DataNode. The AffineExpr of the access expression is defined as follows:

[0081]

[0082] The access expression describes how to calculate the actual access position from the input index, supporting linear combinations (linear combinations of multiple dimensions), constant offsets (fixed offsets), and symbolic variables (variables that can only be determined at runtime).

[0083] It should be noted that there are the following three precautions for the use of data nodes:

[0084] First, the length of the coefficient list must match the number of input dimensions. For one-dimensional access, use a single-element list.

[0085] Second, the index range needs to ensure that the generated index is within the valid range, considering the possible out-of-bounds caused by offsets.

[0086] Third, the values of all used symbolic variables must be provided at runtime. Symbolic variables are usually used to represent dynamic sizes or offsets.

[0087] III. Compute Node, which is used to represent a computing operation.

[0088]

[0089] Supported operation types include logical operations, arithmetic operations, scalar mathematical operations, special operations, and selection operations, etc. Logical operations include bitwise AND, bitwise OR, and equality, etc. Arithmetic operations include addition, multiplication, modulo, and left shift, etc. Scalar mathematical operations include maximum value, ceiling, and rounding, etc. Special operations include matrix multiplication and dropout operations, etc. Selection operations include MUX multiplexing, etc.

[0090] In addition, the computing node can also implement tensor reduction operations, which allow data to be reduced according to the specified grouping method. The reduction types support sum reduction, maximum reduction, and minimum reduction. The grouping method of reduction is specified by a mask array (mask) to indicate which group each element belongs to. mask[i] represents which group the i-th input element belongs to, where the mask value ranges from [0, num_groups - 1], and a negative mask value indicates that the element does not participate in the reduction. According to the above definition, operations such as row reduction and irregular grouping reduction can be implemented.

[0091] It should be noted that in the existing computational graph representation schemes, computation and data access are usually mixed together, resulting in a non-modular design of the computational graph. This mixed design brings the following four main problems:

[0092] 1. Increased development difficulty. Since computation and data access are tightly coupled, when searching the hardware design space, changes in the data access pattern need to be considered simultaneously when considering the computation logic, and vice versa. This increases the complexity and cost of the search.

[0093] 2. Lack of flexibility. Existing computational graph representation schemes are difficult to flexibly adjust the computation logic or data access pattern when dealing with complex computational graph structures (such as conditional branches, loops, indirect access, etc.).

[0094] 3. Difficult resource management. Since computation and data access are not clearly separated, existing computational graph representation schemes are difficult to efficiently manage computational resources and memory resources when dealing with large-scale data. For example, developers cannot independently optimize the allocation of computational resources or the memory layout, resulting in low resource utilization.

[0095] 4. Inconvenient debugging and optimization. Since computation and data access are mixed together, it is difficult to clearly distinguish problems in the computation logic and data access pattern when debugging and optimizing the computational graph. This makes the debugging process more complex, especially in the case of dynamic control flow, where it is difficult for developers to accurately trace the root cause of the problem.

[0096] To solve the above problems, the present invention introduces a separated design of computing nodes and data nodes. By completely decoupling computation and data access, the present invention significantly improves the flexibility and optimization potential of the computational graph.

[0097] Specifically, this separated design brings the following advantages:

[0098] 1. Clear role division. The ComputeNode is specifically responsible for performing specific computing operations, such as addition, multiplication, activation functions, etc. The responsibility of the ComputeNode is limited to computing and does not involve data storage or access methods. This clear role division makes the computing more focused on the implementation of algorithms and reduces unnecessary complexity. The DataNode is specifically responsible for describing data access behaviors, such as input shape, output shape, data type, access mode, etc. The responsibility of the DataNode is to ensure the correct transfer and access of data without involving specific computing logic. This separation enables the data access mode to be optimized independently of the computing logic, further improving the flexibility of the computation graph.

[0099] 2. Decouple computing and data access. By separating the ComputeNode and the DataNode, the present invention achieves a complete decoupling of computing and data access, making the design of the computation graph more modular and also making the computation graph more scalable. Developers can flexibly select suitable combinations of ComputeNodes and DataNodes according to different application scenarios, thereby constructing more complex computation graph structures. For example, when dealing with sparse matrix operations, developers can choose an appropriate index access mode (INDEX). This flexibility enables the computation graph to adapt to various complex data access requirements.

[0100] 3. Support for flexible access modes. The DataNode supports multiple access modes, such as affine access (AFFINE), indirect access (INDIRECT), slice access (SLICE), merge access (MERGE), etc. These access modes enable the computation graph to adapt to various complex data access requirements, especially performing well in scenarios such as dealing with irregular data structures, sparse matrices, and dynamic data streams.

[0101] 4. Efficient resource management. Due to the separation of the ComputeNode and the DataNode, the present invention can manage computing resources and memory resources more efficiently. The ComputeNode can dynamically allocate computing resources as needed, while the DataNode can optimize the memory layout and access path according to the data access mode. For example, when dealing with large-scale data, the DataNode can adopt a block loading method to reduce memory occupancy; when dealing with small-scale data, the DataNode can adopt a sequential access mode to improve the cache hit rate.

[0102] IV. ConstantNode, which is used to represent constant values in the program.

[0103]

[0104] V. BranchNode, which is used to represent a conditional branch structure.

[0105]

[0106] Among them, BranchTrace is used to record branch jump information. Each branch node can record the results of multiple executions. There is a list of branch execution records in BranchTrace. The objects in this list are BranchExecution. BranchExecution has three parameters, namely the condition result, the selected path (True or False), and the execution timestamp. The execution timestamp is used to identify the execution order.

[0107] VI. DATA EDGE

[0108] The DATA edge is used to represent the data flow between computing nodes. Each DATA edge connects two nodes, indicating that the data output from one node flows to another node as input. It is the most common edge type in the computation graph. It is applicable to all scenarios involving data transfer, such as data transfer between layers in a neural network, input and output in matrix multiplication, etc.

[0109] VII. INDEX EDGE

[0110] The INDEX edge is used for the index of indirect access and can carry additional index information to indicate the specific access location or method of the data. Indirect access performs data access through an index array and is applicable to scenarios such as sparse matrices and irregular data structures.

[0111] VIII. MASK EDGE

[0112] The MASK edge is used for masked access. Masked access selectively processes part of the data according to the mask condition and is applicable to fields such as image processing and natural language processing.

[0113] The INDEX edge and the MASK edge are used in conjunction with the indirect access mode and the masked access mode of the DataNode, greatly enhancing the expressive power of the computation graph, being able to flexibly handle complex memory access patterns, and improving the computing efficiency and resource utilization rate. It is more applicable to scenarios that require processing complex data access patterns, such as sparse matrix operations and dynamic data streams.

[0114] It should be noted that the embodiments of the present invention introduce multiple types of edges, which can more accurately describe different types of data access patterns and control flows, and solve the problem of the single function of edge types in the existing computation graph representation scheme. Specifically, in addition to supporting standard data flow edges, the present invention also supports INDEX edges and MASK edges. The INDEX edge can carry additional index information to indicate the specific access location or method of the data. The masking operation selectively processes part of the data according to the mask condition.

[0115] By introducing additional types of edges, the present invention can flexibly handle more complex memory access patterns. Especially when dealing with complex scenarios such as sparse matrix operations and dynamic data streams, the advantages of the present invention are particularly obvious, thereby improving the expressive power of the computational graph.

[0116] S2. Perform a hardware design space search based on the computational graph to determine a hardware design solution that meets the preset performance indicators.

[0117] Furthermore, as Figure 3 shown, the present invention can start from a deep learning program, construct a computational graph, and perform a hardware design space search. The pre-built operator template library is a template library of commonly used operators in deep learning predefined based on the basic components of the present invention, which can be directly selected and used. At the same time, this library can be expanded at any time according to specific applications.

[0118] Implementing the computational graph representation method applicable to hardware design space search provided by the embodiments of the present invention has the following advantages:

[0119] First, the expressive power is significantly enhanced.

[0120] In the existing computational graph representation schemes, the edge types are single, making it difficult to accurately describe complex data access patterns, and the efficiency is low when dealing with scenarios such as irregular data structures and sparse matrices. The present invention introduces INDEX edges and MASK edges, greatly enhancing the expressive power of the computational graph. For example, in sparse matrix operations, INDEX edges can accurately depict indirect access patterns, accurately indicating the specific access locations or methods of data, enabling the computational graph to manage memory access more efficiently, and effectively solving the problem of insufficient expressive power of the computational graph in complex scenarios.

[0121] Second, the debugging and hardware design space search efficiency are greatly improved.

[0122] Traditional computational graph representation schemes lack support for dynamic traces, resulting in a cumbersome debugging process and being unable to accurately evaluate the impact of different hardware configurations on the execution performance of the computational graph during hardware design space search. The present invention provides comprehensive support for dynamic traces, recording information such as input, output, and execution timestamps when each operator is executed, and can also record relevant information in detail for operations such as conditional branches and indirect accesses. When debugging complex computational graphs, developers can use this dynamic trace information to quickly locate potential performance bottlenecks, and the average debugging time is shortened by about twice compared with traditional schemes. In the hardware design space search, based on the information recorded by the dynamic traces, the performance of different computational operations and data access patterns during hardware execution can be accurately analyzed, greatly reducing the search time and improving the search accuracy.

[0123] Third, the flexibility advantage is prominent.

[0124] In the existing computational graph representation schemes, the coupling between computation and data access leads to problems such as insufficient flexibility, difficult resource management, and inconvenience in debugging and optimization. The present invention separates computational nodes from data nodes, achieving complete decoupling of the two, and can flexibly represent various combinations of data access and computation. At the same time, when searching the hardware design space, it can effectively accelerate the search.

[0125] IV. Widely applicable to a variety of application scenarios.

[0126] The prior art has limitations in dealing with static and dynamic neural networks and different types of deep learning tasks. The computational graph representation method of the present invention is not only applicable to static neural networks, but also performs excellently in dynamic neural networks, and has good applicability in aspects such as hardware design space search and traditional deep learning network inference tasks.

[0127] In addition, the present invention has been successfully applied to the project of the automatic generation technology of dedicated accelerators, specifically applied to the LLMCompiler module and DSE module in this project. LLMCompiler is an automated program performance analysis and optimization tool, aiming to identify performance bottlenecks from high-level language programs and generate a Task Graph for Design Space Exploration (DSE). The Task Graph therein uses the present invention to describe and define the computational graph. In this module, by introducing enhanced edge type support, dynamic trace support, and the separated design of computational nodes and data nodes, the generated computational graph well adapts to the computational logic description requirements in high-level language programs. The separated design of computational nodes and data nodes is also very suitable for the DSE module to perform the design space search task of dedicated accelerators, facilitating better allocation of computational and bandwidth resources.

[0128] Based on the same inventive concept, an embodiment of the present invention further provides a computational graph representation device suitable for hardware design space search, as Figure 4 shown, including:

[0129] A construction unit, configured to construct a computational graph based on a deep learning program; the computational graph includes nodes and edges; the edges include data edges, index edges, and mask access edges; the data edges are used to represent the data flow between computational nodes; the index edges are indexes for indirect access, and can carry index information to indicate the specific access location or method of data; the mask access edges are used to selectively process part of the data according to mask conditions;

[0130] A search unit, configured to perform a hardware design space search based on the computational graph to determine a hardware design solution that meets preset performance indicators.

[0131] Further, the nodes include IO nodes, data nodes, computing nodes, constant nodes, and branch nodes; based on these nodes, the computational graph can provide dynamic tracing support, including fine-grained dynamic tracing records, branch jump information records, and indirect access address records.

[0132] Among them, the fine-grained dynamic tracing record is specifically: when each operator is executed, record the input, output, and execution timestamp;

[0133] The branch jump information record is specifically: for a computational graph containing conditional branches, use branch nodes to record the specific situation of each branch jump, including the condition result, the selected path, and the execution timestamp;

[0134] The indirect access address record is specifically: for operations involving indirect access, record the actual accessed address and the mask value.

[0135] Further, the data nodes and computing nodes adopt a separated design; the computing nodes are limited to computing; the data nodes are only used for data storage or access, and the provided access modes include: affine access, indirect access, slice access, merge access, and mask access. The computing nodes support logical operations, arithmetic operations, scalar data operations, special operations, selection operations, and tensor reduction.

[0136] It should be noted that for the specific working process of this embodiment, please refer to the method embodiment part mentioned above, and details will not be elaborated here.

[0137] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A computational graph representation method applicable to hardware design space search, characterized in that, Including: Constructing a computational graph based on a deep learning program; the computational graph includes nodes and edges; the edges include data edges, index edges, and mask access edges; The data edges are used to represent the data flow between computational nodes; the index edges are indices for indirect access, which can carry index information to indicate the specific access location or method of data; the mask access edges are used to selectively process part of the data according to mask conditions; Performing a hardware design space search based on the computational graph to determine a hardware design solution that meets the preset performance metrics.

2. The computational graph representation method according to claim 1, characterized in that The data edges are applicable to all scenarios involving data transfer; the scenarios applicable to the index edges include sparse matrices, input and output in matrix multiplication; the scenarios applicable to the mask access edges include image processing and natural language processing.

3. The computational graph representation method according to claim 1, characterized in that The nodes include IO nodes, data nodes, computational nodes, constant nodes, and branch nodes; based on the nodes, the computational graph can provide dynamic trace support, including fine-grained dynamic trace recording, branch jump information recording, and indirect access address recording.

4. The computational graph representation method according to claim 3, wherein The data nodes and computational nodes adopt a separated design; the computational nodes are limited to computing; The data nodes are only used for data storage or access, and the available access modes include: affine access, indirect access, slice access, merge access, and mask access.

5. The computational graph representation method according to claim 3, wherein The computational nodes support logical operations, arithmetic operations, scalar data operations, special operations, selection operations, and tensor reduction.

6. The computational graph representation method according to claim 3, wherein, The fine-grained dynamic trace recording is specifically: when each operator is executed, record the input, output, and execution timestamp; The branch jump information recording is specifically: for a computational graph containing conditional branches, use branch nodes to record the specific situation of each branch jump, including the conditional result, the selected path, and the execution timestamp; The indirect access address recording is specifically: for operations involving indirect access, record the actual accessed address and mask value.

7. A computational graph representation device applicable to hardware design space search, characterized in that, Including: A construction unit for constructing a computational graph based on a deep learning program; the computational graph includes nodes and edges; the edges include data edges, index edges, and mask access edges; the data edges are used to represent the data flow between computational nodes; the index edges are indices for indirect access, which can carry index information to indicate the specific access location or method of data; the mask access edges are used to selectively process part of the data according to mask conditions; A search unit for performing a hardware design space search based on the computational graph to determine a hardware design solution that meets the preset performance metrics.

8. The computational graph representation device according to claim 7, wherein The nodes include IO nodes, data nodes, computational nodes, constant nodes, and branch nodes; based on the nodes, the computational graph can provide dynamic trace support, including fine-grained dynamic trace recording, branch jump information recording, and indirect access address recording.

9. The computational graph representation device according to claim 8, wherein The data nodes and computational nodes adopt a separated design; the computational nodes are limited to computing; The data nodes are only used for data storage or access, and the available access modes include: affine access, indirect access, slice access, merge access, and mask access.

10. The computational graph representation device according to claim 8, wherein The fine-grained dynamic trace recording is specifically: when each operator is executed, record the input, output, and execution timestamp; The branch jump information record is specifically as follows: for a computational graph containing conditional branches, branch nodes are used to record the specific situation of each branch jump, including the conditional result, the selected path, and the execution timestamp; The indirect access address record is specifically as follows: for operations involving indirect access, the actual accessed address and the mask value are recorded.