An operator library generation method and application for coarse-grained reconfigurable AI array
By converting the CNN model into a directed acyclic graph and performing frequent subgraph mining, identifying high-frequency computing patterns, and building an operator library, the problems of low computational efficiency of CGRA arrays and insufficient reuse of common structures in existing technologies are solved, and efficient convolutional neural network reasoning is achieved.
Patent Information
- Application Number
- CN202511028158.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing compilers and deep learning frameworks are unable to effectively utilize the coarse-grained computing characteristics of CGRA, resulting in wasted computing power and poor performance. They are also difficult to adapt to changes in different network structures and hardware configurations, and lack the potential for reuse of common structures across models.
The CNN neural network inference model is converted into a directed acyclic graph. The frequent subgraph mining algorithm is used to identify high-frequency computing patterns. The frequent subgraphs are extracted as coarse-grained operators, which are partitioned and mapped on the CGRA array to build an operator library and optimize data transmission and storage utilization.
It improves processing unit utilization, significantly reduces execution time, achieves a 12.5% increase in inference speed, and reduces the workload of operator development and the difficulty of software ecosystem construction.
Smart Images

Figure CN120525013B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to extracting frequent / common operators suitable for coarse-grained reconfigurable AI (artificial intelligence) arrays from computational graphs of various convolutional neural network inference models, and in particular to a method and application for generating an operator library for coarse-grained reconfigurable AI arrays. Background Art
[0002] With the development of artificial intelligence (AI), the scale of CNNs (convolutional neural networks) continues to expand, and the computational complexity of these models has reached gigabytes. The large number of multiplication-accumulation calculations and model parameters poses severe challenges to the performance and real-time nature of computing devices. CGRAs (coarse-grained reconfigurable arrays) offer excellent configuration flexibility and higher performance when accelerating high-parameter and high-throughput computations, making them an ideal platform for handling compute-intensive operations and deploying CNN models.
[0003] Existing optimization tools for reconfigurable arrays are primarily targeted at traditional large-scale commercial reconfigurable array processors, and their mapping algorithms are typically designed for specific underlying hardware. This makes it difficult for existing open-source optimization tools to be directly applied to specific scales, such as multi-level storage CGRA processors with weight broadcasting. Furthermore, existing deep learning compilers are unable to achieve efficient hardware-dependent graph partitioning. Specifically, existing technologies suffer from the following drawbacks:
[0004] (1) Existing compilers (such as the Tensor Virtual Machine (TVM), the Tensor Algebra Super Optimizer (TASO), and Partial Evaluation and Tabulation (PET)) have limited optimization capabilities in solving hardware-related graph-level optimizations and computational graph partitioning schemes, and are unable to fully utilize the coarse-grained computing characteristics of CGRA, resulting in wasted computing power and poor performance, reducing actual computing efficiency.
[0005] (2) Some studies have also proposed data partitioning and scheduling strategies for different storage levels to address the frequent on-chip / off-chip data movement caused by the large number of parameters and intermediate feature maps in CNNs. However, these methods are usually based on fixed algorithm templates or heuristic rules, making it difficult to adapt to changes in different network structures and hardware configurations, and the optimization space is relatively limited. In particular, for the multi-level storage CGRA architecture with weight broadcast, it is impossible to flexibly adjust the optimization strategy to adapt to hardware constraints.
[0006] (3) In the design of existing deep learning frameworks and operator libraries, people often only focus on the high-performance implementation of underlying atomic operators or simple operator fusion, but ignore the reuse potential across models and tasks and the analysis of high-frequency subgraph patterns. This leads to the lack of a composite operator priority strategy based on subgraph frequency, and the inability to capture the common structure across models at the operator library level. Summary of the Invention
[0007] The embodiments of the present application provide an operator library generation method and application for a coarse-grained reconfigurable AI array, which is used to solve the problem that the optimization tools in the prior art are not suitable for multi-level storage CGRA with weight broadcast.
[0008] In one aspect, an embodiment of the present application provides a method for generating an operator library for a coarse-grained reconfigurable AI array, comprising:
[0009] Convert the CNN neural network inference model into a directed acyclic graph according to different mapping strategies , a directed acyclic graph contains multiple nodes;
[0010] The different mapping strategies are defined as:
[0011]
[0012] in, Represents a directed acyclic graph of the CNN neural network inference model in the ONNX (Open Neural Network Exchange) model structure, S Indicates the mapping strategy number. Express Use mapping strategy for conversion, Representation mapping strategy S What's included, v Represents a node, Indicates the node type, represents the kernel size of the convolution / pooling layer, Indicates the number of steps of the convolution / pooling layer, The shape representing the node input / output;
[0013] Extract common operators from the directed acyclic graph, and divide the directed acyclic graph into multiple subgraphs using a frequent subgraph mining algorithm based on the Apriori algorithm according to the common operators and the connection order of the nodes in the directed acyclic graph.
[0014] Analyze the frequency of subgraphs. According to the frequency, the subgraphs whose occurrence times are greater than the support threshold are classified as frequent subgraphs, otherwise they are non-frequent subgraphs. Frequent subgraphs are called coarse-grained operators.
[0015] For multi-level storage CGRA arrays with weight broadcast, when the computation space required by the coarse-grained operator is larger than the on-chip storage space of the CGRA array C When dividing the convolution operation in the coarse-grained operator, the feature map of the input CGRA array is first divided into multiple data blocks, and then the convolution operation in the coarse-grained operator is divided according to the formula:
[0016]
[0017]
[0018] in, , h Indicates the height of the largest data block, w Indicates the width of the maximum data block, C in Indicates the number of channels of the largest data block, W w and W h Respectively represent the convolution kernel width and height of the fine-grained operator, T rep represents the number of repetitions of the convolution kernel, I c Indicates the number of channels of the feature map;
[0019] Using the machine instructions of the CGRA array, coarse-grained operators and fine-grained operators are manually mapped respectively to build an operator library for the CGRA array.
[0020] In one possible implementation, two (k-1)-subgraphs are connected. If their first (k-2) nodes are the same, they are connected into a k-subgraph. Multiple k-subgraphs form a candidate set, which is pruned to obtain a frequent subgraph.
[0021] In one possible implementation, the method for pruning the candidate set is:
[0022] Calculate the level difference of all nodes in each k-subgraph in the candidate set. If the level difference of a k-subgraph is greater than 2, delete the k-subgraph from the candidate set.
[0023] In one possible implementation, a method for dividing a feature map into multiple data blocks includes:
[0024] First, maximize the input dimension of the data block. Then, make the data block large enough in height to contain the size required for four convolution steps. Then, make the width as large as possible under the constraints of the formula. If the width exceeds the width of the feature map, continue to increase the height.
[0025] On the other hand, an embodiment of the present application further provides a method for image classification using an operator library generation method for a coarse-grained reconfigurable AI array, the method comprising the following steps:
[0026] Get the ONNX model file of the convolutional neural network model;
[0027] The front-end conversion module parses the ONNX model file and extracts the network structure composition information, generating a JSON (JavaScript Object Notation) configuration file containing detailed information of each operation node and data node;
[0028] Traverse all nodes in the JSON configuration file, analyze the parameter information of each node, match the node parameter information with the constructed operator library, and reuse the operator if there is an operator in the operator library that completely matches the parameter information of the current node.
[0029] Complete the node connection of the entire network based on the matched operators, determine the corresponding data block information based on the parameter information of each node, extract the model parameters, and perform quantization and format conversion according to the storage and computing requirements of the CGRA array;
[0030] The above operator library is called and configuration parameters for controlling array behavior are added on this basis to generate executable code deployed on the weight-broadcast multi-level storage CGRA array. The CGRA array executes the image classification inference process according to the generated code and outputs the classification results.
[0031] The operator library generation method and application for a coarse-grained reconfigurable AI array in this application have the following advantages:
[0032] (1) This application is specifically designed for the weight broadcast multi-level storage CGRA array architecture. First, the network model is converted into a high-level computational graph representation. Then, the frequent subgraph mining algorithm is used to iteratively identify the effective computing node combinations and frequently occurring common substructures in the network structure, extracting coarse-grained operators. Finally, considering the storage hierarchy and capacity constraints of CGRA, the slicing method is used to decompose the calculation of large blocks of data that cannot be deployed in CGRA into multiple fine-grained small blocks of data. The communication cost between on-chip storage and off-chip storage is reduced by optimizing the data exchange mechanism. This method can effectively reduce the data transmission overhead between frequent operators and achieve efficient utilization of hardware computing resources. By jointly constructing a convolutional neural network operator library for the weight broadcast multi-level storage CGRA array with two operators of different granularities, the mapping difficulty of the convolutional neural network is effectively reduced, and the workload of operator development is reduced.
[0033] (2) Analyze and identify high-frequency neural network structures through the frequency of subgraphs in neural networks, thereby building an operator library, giving priority to the development of composite operators with high reuse potential, and building a general operator library by capturing common computing patterns across models, reducing the workload of CNN operator development and lowering the difficulty of building an AI chip application software ecosystem.
[0034] (3) Experiments show that deploying the CNN model on a weight-broadcast multi-level storage CGRA array using the method proposed in this application increases the processing unit (PE) utilization by 10%, significantly reduces the execution time of operators on the array, and achieves a 12.5% increase in inference speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0036] Figure 1 Schematic diagram of the top-level structure of the PE array provided in an embodiment of the present application.
[0037] Figure 2 A schematic diagram of the common operator extraction results provided in an embodiment of the present application.
[0038] Figure 3 Schematic diagram of the difference between the coarse-grained operator before and after division provided in an embodiment of the present application.
[0039] Figure 4 This is a graph representation after subgraph mining provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0041] The method of this application is applicable to a multi-level storage CGRA array structure of weight broadcast, Figure 1 This cluster array structure is shown, which contains 36 PEs, cache units for storing input / output data, and cache units for storing weight data. The outer layer of the cluster is a set of dual-mode PEs, which can be set to instruction flow mode or data flow mode, while the inner layer is the data flow PE. In instruction flow mode, the dual-mode PE can access the peripheral cache units to read and write data through memory access instructions. The entire cluster is composed of 36 PE units arranged in a 6×6 layout, and the operating mode is controlled by configuration information. After configuration is completed, the data flow PE receives input data, performs calculations, and transmits the results to the specified path. These data flow PEs mainly perform basic calculations such as multiplication and accumulation operations, maximum pooling, and activation functions.
[0042] The optimization method of this application targets the overall structural characteristics of the CGRA array and uses frequent subgraph mining technology to identify high-frequency computing patterns in different CNN models (such as structural combinations such as Conv+BN). By capturing common computing requirements across models, it provides a decision-making basis for the operator priority development and hardware resource allocation of the CGRA array. On this basis, the extracted operators (including coarse-grained operators and fine-grained operators) are used to build a complete operator library, thereby realizing operator-based algorithm mapping. In the process of building an operator library for the CGRA array, operators are divided into two categories: coarse-grained and fine-grained. This application defines these two types of operators as follows:
[0043] Coarse-grained operators are hardware-independent, abstract computational units that describe computational logic at a high level and are independent of specific hardware architectures. These operators focus on the functionality and semantics of the algorithm, abstracting complex computational tasks into independent functional modules. For example, image feature extraction can be implemented as a coarse-grained operator based on this definition, regardless of whether it is implemented on an FPGA (field programmable gate array), an ASIC (application-specific integrated circuit), or other hardware platforms. This hardware independence gives coarse-grained operators strong versatility and portability, enabling cross-platform reuse and effectively reducing duplication of development work for different hardware platforms.
[0044] Fine-grained operators contain instruction codes that can be executed directly on hardware and are closely associated with specific hardware architectures and instruction sets. Taking the CGRA array as an example, fine-grained operators make full use of its resources such as PEs, lookup tables, and registers to convert algorithm implementations into instructions that can be directly executed by the hardware. In the implementation of the convolution algorithm, fine-grained operators finely decompose and optimize the calculation process based on the hardware characteristics of the CGRA array, generating efficient assembly code. This enables fine-grained operators to fully utilize the array's hardware computing resources and achieve efficient computing task execution through instruction-level optimization.
[0045] The following describes in detail a method for generating an operator library for a coarse-grained reconfigurable AI array provided by an embodiment of the present application. The method includes the following steps:
[0046] S100, converts the CNN neural network inference model into a directed acyclic graph according to different mapping strategies , a directed acyclic graph contains multiple nodes;
[0047] The different mapping strategies are defined as:
[0048]
[0049] in, Represents a directed acyclic graph (DAG) of the CNN neural network inference model in the ONNX model structure. This DAG trains the convolutional neural network inference model based on the PyTorch framework and exports the model to the ONNX format through the model conversion tool provided by PyTorch. In the ONNX representation, the computational graph structure of the neural network is modeled as a DAG. S Indicates the mapping strategy number. Express Use mapping strategy for conversion, Representation mapping strategy S What's included, v Represents a node, Indicates the node type, represents the kernel size of the convolution / pooling layer, Indicates the number of steps of the convolution / pooling layer, The shape representing the input / output of a node.
[0050] For example, the graph representation of CNN is a directed acyclic graph. By analyzing the frequency of occurrence of subgraphs in the neural network model, high-frequency network structures can be identified, composite operators with high reuse potential can be developed, and computing efficiency and hardware resource adaptability can be optimized. For a single neural network model, it can be represented as a directed acyclic graph. G ( V , E ),in V Represents a set of nodes, each node v i ∈ V represents a neuron, E represents an edge set, each edge ( v i , v j )∈ E Represents neurons v i to neurons v j For data flow or dependency between M neural networks, which can be combined into a set { G 1, G 2,⋯, G t ,⋯, G M},in G t =( V t , E t ) indicates the t A neural network model.
[0051] S110, extracting common operators from the directed acyclic graph, and dividing the directed acyclic graph into multiple subgraphs using a frequent subgraph mining algorithm based on the Apriori algorithm according to the common operators and the connection order of the nodes in the directed acyclic graph.
[0052] For example, in the subgraph mining process, each node can be regarded as a 1-subgraph, and a subgraph containing multiple nodes is called a k-subgraph, where k represents the number of nodes in the subgraph. Figure 2 As shown in the figure, if a subgraph contains three nodes: convolution Conv, normalization BN, and activation ReLu, and this type of subgraph appears multiple times in different network structures, then {Conv, BN, ReLu} is a frequent subgraph S1.
[0053] S120, analyzing the frequency of the subgraphs, and classifying the subgraphs whose occurrence times are greater than the support threshold as frequent subgraphs according to the frequency, and otherwise as infrequent subgraphs. Frequent subgraphs are called coarse-grained operators.
[0054] For example, the set of nodes in a subgraph identified through frequent subgraph mining is defined as a coarse-grained operator in the convolutional neural network operator library. The subgraph length is the number of nodes in the subgraph. The subgraph S1 mentioned above has a length of 3 and is called a 3-subgraph. A subgraph containing only one node is called a 1-subgraph. Different subgraphs appear different numbers of times in each neural network model. Some subgraphs may appear only once, while some frequent subgraphs may appear in multiple neural network models, or even multiple times in the same neural network model.
[0055] CNN contains a variety of computing nodes, many of which have large computational loads and strong data dependencies between nodes, so there is a lot of room for optimization. In order to seize these optimization opportunities, this application designs a frequent subgraph mining algorithm G-Growth, which is used to identify frequently appearing subgraphs and lay the foundation for subsequent operator extraction. The G-Growth algorithm specifically includes: scanning each subgraph, treating each node in the subgraph as a 1-subgraph, counting the frequency of each 1-subgraph, and screening to obtain frequent 1-subgraphs, recorded as L 1. From the frequent 1-subgraph L Starting from 1, generate a candidate set of k-subgraphs through the connection operation C k , for the candidate set C k Perform pruning to obtain all frequent subgraphs.
[0056] Furthermore, the frequency of the subgraph is expressed as follows:
[0057]
[0058] in, Representing a subgraph X The frequency of Indicates that it contains subgraphs X The number of directed acyclic graphs, N Represents the total number of directed acyclic graphs; a subgraph whose frequency is greater than or equal to the support threshold is determined to be a frequent subgraph.
[0059] Furthermore, a candidate set is generated by the connection operation C k The method is: connect two (k-1)-subgraphs. If their first (k-2) nodes are the same, connect them into a k-subgraph, and multiple k-subgraphs form a candidate set.
[0060] Furthermore, for the candidate set C k Pruning can improve the efficiency of frequent subgraph mining. In the embodiment of the present application, there are two methods for pruning the candidate set. One is to narrow the search space based on the hierarchical structure of the nodes in the subgraph, calculate the hierarchical difference of all nodes in each k-subgraph in the candidate set, and if the hierarchical difference of a k-subgraph is greater than 2, remove the k-subgraph from the candidate set. C k The second is to detect the frequency of the subgraph in advance, that is, in the pruning process, follow the principle: if a subgraph is frequent, then all its subsets must be frequent; conversely, if a subset of a subgraph is not frequent, then the subgraph cannot be frequent. Based on this principle, the pruning method of this application is as follows: if the candidate set C k If a (k-1)-subgraph of a k-subgraph is not a frequent subgraph, then the k-subgraph is removed from the candidate set C k Delete in.
[0061] After pruning the candidate set, each subgraph is scanned again, the frequency of each subgraph is calculated, and a new candidate set is generated. The new candidate set is pruned to obtain a new frequent subgraph. The subgraph scanning, candidate set generation and pruning process are repeated until no new frequent subgraph can be generated.
[0062] Based on the results of frequent subgraph mining, this application performs coarse-grained partitioning of identified frequent subgraphs based on the logical unit characteristics of the CGRA array to improve hardware efficiency and performance. This strategy specifically focuses on the combination patterns of core operators such as convolution (Conv) and activation (ReLu). In traditional methods, the results of the convolution operator are typically stored in an on-chip cache before the activation or pooling (Maxpool) operator is calculated. This not only increases data transmission latency but also increases memory access pressure. To address this issue, different operators with adjacent computation relationships can be combined into a single, coarse-grained operator. For example, the convolution operator and activation operator identified by frequent subgraph mining can be combined into a single coarse-grained operator. This operator extraction strategy significantly reduces the communication overhead of convolution and activation calculations, and significantly reduces the cost of accessing off-chip memory. This operator extraction and coarse-grained partitioning based on frequent subgraph mining significantly reduces the data transmission overhead between hardware units and improves computational efficiency. Figure 3 The results of graph partitioning based on the extraction of coarse-grained operators are shown. The left side shows the graph structure before coarse-grained partitioning, and the right side shows the graph structure after coarse-grained partitioning. Compared with the Apriori algorithm, the G-Growth algorithm proposed in this application reduces the subgraph mining time by an average of 8.7% under different mapping strategies.
[0063] S130, for the multi-level storage CGRA array of weight broadcast, when the computation space required by the coarse-grained operator is larger than the on-chip storage space of the CGRA array C When dividing the convolution operation in the coarse-grained operator, the convolution operation in the coarse-grained operator is divided to obtain multiple sub-graphs with smaller computation and storage requirements, which are called fine-grained operators. When dividing the convolution operation in the coarse-grained operator, the feature map of the input CGRA array is first divided into multiple data blocks, and then divided according to the formula:
[0064]
[0065]
[0066] in, , h Indicates the height of the largest data block, w Indicates the width of the maximum data block, C in Indicates the number of channels of the largest data block, W w and W h Respectively represent the convolution kernel width and height of the fine-grained operator, T rep represents the number of repetitions of the convolution kernel, I cIndicates the number of channels of the feature map.
[0067] Exemplarily, the above-mentioned G-Growth algorithm can extract coarse-grained operators based on the relationship and frequency of occurrence of nodes in the neural network model. This strategy pays special attention to the combination pattern of core operators such as convolution operators and activation operators. Below, a fine-grained operator partitioning method is designed based on the storage constraints of the multi-level storage CGRA array structure of weight broadcasting to divide convolution operators with large computational complexity and storage requirements into finer-grained operators. On the one hand, this division alleviates the storage pressure of the computing device, and on the other hand, fine-grained operators can expose more optimization space and parallel space. At the same time, computing operations with different input and output sizes can divide more computing operations with the same input and output, thereby reducing the number of operator types and the workload of operator developers.
[0068] When the convolution operator has an input feature map of a larger size, a convolution operator with a fixed convolution kernel size is applied to it. The input feature map can be divided into multiple smaller data blocks in the spatial dimension, and then a fine-grained operator with the same size convolution kernel is applied to each data block for convolution operation. For example, for a larger two-dimensional image, it is divided into multiple small image blocks, each small image block is used as an independent input, and the same convolution kernel is used to perform convolution calculations on these small image blocks respectively. Finally, the convolution results of these small image blocks are combined, which is equivalent to performing a convolution operation on the entire large feature map. Based on this method, the present application proposes an M-Base algorithm that can split the convolution kernel in different spatial dimensions, reducing the storage requirements of a single operation without losing the accuracy of the calculation.
[0069] Depend on Figure 1 From the top-level PE array structure, you can see that the CGRA array has a multi-level storage setting. Therefore, when designing fine-grained operators that can be executed in the CGRA, you must not only consider the parallel computing capabilities of the array, but also pay attention to the constraints of the storage module. In the convolution calculation, the input feature map and weight data of the CGRA array have different storage constraints. The CGRA array can only calculate integer data. The bit width of the input feature map is 16 bits, and the bit width of the weight data is 8 bits. The input data and weight data in a single task cannot exceed 1Mb. The weight data will be repeatedly organized in the memory according to the number of convolution steps and the number of computing units involved in the array. This is to reduce data readback and enable faster data transmission and calculation. The weight data is determined according to the storage space occupied by different situations. The number of repetitions of the convolution kernel T rep It needs to be determined according to the number of convolution steps, and its calculation method is shown in the following formula:
[0070]
[0071] in, Indicates rounding down. N H is the number of PEs in the height direction, T conv_h and T conv_w They represent the number of convolution steps in the height and width directions of the feature map respectively. The calculation method is shown in the following formula: I w and I h Represent the size of the feature map in the width and height directions respectively, W w and W h Respectively represent the convolution kernel width and height of the fine-grained operator:
[0072]
[0073]
[0074] in, s It indicates the sliding stride size of the convolution kernel. It can be seen that the convolution kernel changes continuously as the size of the convolution feature map grows in each dimension.
[0075] In order to maximize the utilization of the array by a single task, the following splitting scheme is designed:
[0076] set up T conv_w and T conv_h are the number of convolution steps of the feature map in the width and height directions respectively, and the constraints are satisfied when dividing the feature map: W h * W w * I c * T conv_w *( T conv_h / 4+1) is less than 8192, where I c The maximum is 8. When dividing the feature map, first maximize the input dimension of the data block, then make the data block contain the size required for four convolution steps in the height direction, and then make the width as large as possible under the above constraints. If the width exceeds the width of the feature map, continue to increase it in the height direction, and finally split the part that cannot be contained in the large block into small blocks.
[0077] S140 , using machine instructions of the CGRA array, manually mapping the coarse-grained operators and the fine-grained operators respectively, to construct an operator library for the CGRA array.
[0078] The present application also provides a method for image processing using the above-mentioned operator library generation method for a coarse-grained reconfigurable AI array. The method uses image classification tasks as a specific application scenario, and the method includes:
[0079] Get the ONNX model file of a given convolutional neural network model (such as ResNet, VGG, etc.);
[0080] The front-end conversion module parses the ONNX model file and extracts the network structure composition information, generating a JSON configuration file containing detailed information of each operation node and data node;
[0081] Traverse all nodes in the JSON configuration file, analyze the parameter information of each node (including operation type, input and output dimensions, convolution kernel size, step size, etc.), and match the node parameter information with the constructed operator library. If there is an operator in the operator library that completely matches the parameter information of the current node, it will be directly reused. If an operator node is not in the operator library, it will be manually mapped using CGRA instructions and added to the operator library. Because the frequent subgraph mining algorithm covers a wide range, the operator library contains multiple operator combinations in common convolutional neural networks, so it can match the main operators in most network structures.
[0082] Based on the matched operators, the node connections of the entire network are completed. The corresponding data block information, including input data block, weight data block, output data block and their sizes, is determined according to the parameter information of each node. The model parameters of all or specific layers are extracted and quantized and format converted according to the storage and computing requirements of the CGRA array. At the same time, the multi-level storage structure of the array is considered to optimize the data scheduling strategy between different storage levels.
[0083] The constructed operator library is called and configuration parameters for controlling array behavior are added on this basis to generate executable code that can be deployed on the multi-level storage CGRA array with weight broadcasting. The CGRA array executes the image classification inference process according to the generated code and outputs the classification results.
[0084] In view of the fact that existing optimization technologies are multi-faceted in traditional reconfigurable array design and difficult to be directly applied to the multi-level storage specific CGRA array problem of weight broadcast, this application proposes a CNN operator extraction and optimization method based on frequent subgraph mining technology. This method is based on the mined frequent subgraph and storage constraint block, and the optimized graph is divided into G ( V , E) is mapped to the computing resources of the CGRA array and the operation is completed. First, based on the structural scale of the array, the G-Growth algorithm is used to extract frequent subgraphs of different networks and count their frequency of occurrence to capture common computing patterns across models, providing a decision basis for the priority development of operator categories. Because the CGRA array has strict control over the parameters of executable operators, it is necessary to divide larger-scale operators into smaller-scale operators (for example, dividing large-scale convolution operators into small convolution operators) to adapt to the array resources and relieve the pressure on the storage module of the CGRA array.
[0085] When extracting frequent subgraphs between different CNNs, the frequent subgraph mining algorithm designed by this application is applied. However, in the subgraph mining algorithm, the judgment process of graph isomorphism is computationally time-consuming, and the degree of information redundancy carried by the graph is closely related to the judgment efficiency. In order to improve the judgment efficiency as much as possible, the information contained in the graph should be kept highly concise. This application uses Figure 4 The format shown is used to represent different neural network models. The t line represents the number of the graph, the v line defines the vertices, and the e line defines the edges. As can be seen from the figure, the node labels are represented by integers, but different nodes in the neural network model contain rich information. In order to achieve more efficient graph isomorphism detection, the experiment maps the nodes in the neural network model so that they correspond to unique integer identifiers. By establishing a mapping relationship between nodes and integers, a mapping table is constructed. This mapping table remains consistent when processing different graphs, ensuring that the mapping rules of nodes are stable and unified throughout the experiment, thereby providing a standardized and efficient data foundation for subsequent experiments based on frequent subgraph mining, so as to reduce the additional computational overhead caused by inconsistent node representation, thereby accelerating the process of graph isomorphism detection.
[0086] A node in a neural network model contains a lot of information, such as its type, data input and output size, and, for convolutions, kernel size and stride. Therefore, determining whether two nodes are identical requires different criteria. Because different hardware backends handle different computations at different granularities, this application designs four different mapping strategies, as shown in Table 1.
[0087] Table 1 Different mapping strategies
[0088]
[0089] Strategy 1 determines whether two nodes are identical solely based on node types such as convolution, pooling, and activation. Strategy 2 builds on Strategy 1 by also considering the kernel size of convolution and pooling operations, treating 3x3 convolutions and 5x5 convolutions as distinct nodes. Strategy 3 builds on Strategy 2 by also considering the number of steps, for example, treating a 3x3 convolution with a step of 1 and a 3x3 convolution with a step of 2 as distinct nodes. Strategy 4 builds on Strategy 3 by also considering the input and output shapes. Using mapping strategies with varying degrees of laxity, the criteria for determining whether two nodes are identical vary, allowing the selection of appropriate isomorphism criteria based on how hardware processes different nodes in a neural network model.
[0090] To analyze the ability of frequent subgraph mining algorithms to identify common structures and their frequency across different neural network models, experiments were conducted using five classic convolutional neural networks (AlexNet, LeNetv5, MobileNet, ResNet18, and Vgg16). Table 2 shows that when using Strategy 1 (considering only node types), 32 common subgraphs with a frequency of 2 were found, 18 with a frequency of 3, 6 with a frequency of 4, and 2 with a frequency of 5. These experimental data provide important guidance for prioritizing operator development, allowing development resources to be focused on frequently used operators and reducing duplication of development work.
[0091] Table 2 Common subgraphs of different networks
[0092]
[0093] The M-Base algorithm effectively addresses the storage constraints of the CGRA array. The algorithm divides large convolution operations into smaller operations that meet storage constraints, alleviating storage pressure while increasing operator reuse. Table 3 demonstrates that after operator splitting, using Strategy 1, the number of common subgraphs with a frequency of 2 increases from 32 to 53, those with a frequency of 3 increase from 18 to 27, those with a frequency of 4 increase from 6 to 10, and those with a frequency of 5 increase from 2 to 4. This result demonstrates that operator splitting can significantly improve the recognition rate of common subgraphs, further reducing the workload of operator development.
[0094] Table 3 Common subgraphs of different networks after operator splitting
[0095]
[0096] The application of different node mapping strategies enables the method of this application to flexibly adjust the node matching granularity based on the characteristics of the hardware backend. Experiments have found that as the strictness of the judgment conditions increases, the number of common subgraphs gradually decreases. Specific analysis shows that when using strategy 4 (considering node type, convolution / pooling window, stride, and input and output shape), no common subgraphs can be found before operator splitting; after operator splitting, the number of common subgraphs with a frequency of 2 reaches 28, with a frequency of 3 reaching 10, and a frequency of 4 reaching 2. These results show that the method proposed in this application can effectively discover common subgraphs and improve the accuracy and practicality of frequent subgraph mining.
[0097] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0098] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A method for generating an operator library for a coarse-grained reconfigurable AI array, characterized in that: include: Convert the CNN neural network inference model into a directed acyclic graph according to different mapping strategies , the directed acyclic graph includes a plurality of nodes; The different mapping strategies are defined as follows: in, Represents the directed acyclic graph represented by the CNN neural network inference model in the ONNX model structure, S represents the number of the mapping strategy, Express Use mapping strategy for conversion, Representation mapping strategy S Included content, v Represents a node, Indicates the node type, represents the kernel size of the convolution / pooling layer, Indicates the number of steps of the convolution / pooling layer, The shape representing the node input / output; Extracting common operators from the directed acyclic graph, and dividing the directed acyclic graph into multiple subgraphs using a frequent subgraph mining algorithm based on the Apriori algorithm according to the common operators and the connection order of the nodes in the directed acyclic graph; Analyze the frequency of the subgraphs, and classify the subgraphs whose occurrence times are greater than the support threshold as frequent subgraphs according to the frequency, otherwise as infrequent subgraphs. The frequent subgraphs are called coarse-grained operators; For the multi-level storage CGRA array of weight broadcast, when the computation space required by the coarse-grained operator is larger than the on-chip storage space of the CGRA array C When dividing the convolution operation in the coarse-grained operator, the feature map input to the CGRA array is first divided into multiple data blocks, and then the convolution operation in the coarse-grained operator is divided according to the formula: in, , h Indicates the height of the largest data block, w Indicates the width of the maximum data block, C in Indicates the number of channels of the largest data block, W w and W h Respectively represent the convolution kernel width and height of the fine-grained operator, T rep represents the number of repetitions of the convolution kernel, I c Represents the number of channels of the feature map; The coarse-grained operators and the fine-grained operators are manually mapped respectively using machine instructions of the CGRA array to construct an operator library for the CGRA array.
2. The method for generating an operator library for a coarse-grained reconfigurable AI array according to claim 1, characterized in that: Two (k-1)-subgraphs are connected. If their first (k-2) nodes are the same, they are connected into a k-subgraph. Multiple k-subgraphs form a candidate set, and the candidate set is pruned to obtain the frequent subgraph.
3. The method for generating an operator library for a coarse-grained reconfigurable AI array according to claim 2, characterized in that: The method for pruning the candidate set is: The level differences of all nodes in each k-subgraph in the candidate set are calculated. If the level difference of a k-subgraph is greater than 2, the k-subgraph is deleted from the candidate set.
4. The method for generating an operator library for a coarse-grained reconfigurable AI array according to claim 1, characterized in that: The method of dividing the feature map into a plurality of data blocks comprises: First, the input dimension of the data block is maximized, and then the data block is made to contain the size required for four convolution steps in the height direction. Then, under the constraints of the formula, the width direction is made as large as possible. If the width direction exceeds the width of the feature map, the height direction is further increased.
5. A method for image classification using the operator library generation method for a coarse-grained reconfigurable AI array according to any one of claims 1 to 4, characterized in that: include: Get the ONNX model file of the convolutional neural network model; The front-end conversion module parses the ONNX model file and extracts the network structure composition information, generating a JSON configuration file containing detailed information of each operation node and data node; Traverse all nodes in the JSON configuration file, analyze the parameter information of each node, match the node parameter information with the constructed operator library, and reuse the operator if there is an operator in the operator library that completely matches the parameter information of the current node; Complete the node connection of the entire network based on the matched operators, determine the corresponding data block information based on the parameter information of each node, extract the model parameters, and perform quantization and format conversion according to the storage and computing requirements of the CGRA array; Calling the operator library constructed according to any one of claims 1 to 4 and adding configuration parameters for controlling array behavior on this basis, generating executable code deployed on a weight-broadcast multi-level storage CGRA array, and the CGRA array executes the image classification inference process according to the generated code and outputs the classification result.
Citation Information
Patent Citations
Operator scheduling method, device and system based on directed acyclic graph
CN115729648A
Coarse-grained reconfigurable array operator design method and system for deep learning
CN116301892A