Generation Method of Hardware Description Language Supporting Machine Learning and Compilation Toolchain

By generating a hardware description language that supports machine learning, the problem of long-term deployment of neural network models on FPGAs is solved, and the effect of rapid deployment and cost reduction is achieved.

CN116661793BActive Publication Date: 2025-07-01TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210152028.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2025-07-01
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

In the prior art, the deployment of neural network models in FPGAs takes a long time, resulting in extended development cycles and high learning costs.

Method used

A method for generating a hardware description language that supports machine learning is provided, including obtaining a neural network model built in a first high-level language, generating a first computing graph for describing the neural network model, generating a second high-level language based on the first computing graph and preset language generation rules, and inputting it into an intelligent scheduling model to generate a hardware description language.

Benefits of technology

The deployment of neural network model is achieved quickly in FPGA, reducing labor costs and conversion time and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116661793B_ABST
    Figure CN116661793B_ABST
Patent Text Reader

Abstract

The present invention provides a method for generating a hardware description language supporting machine learning and a compilation tool chain. The method includes: obtaining a neural network model constructed using a first high-level language; generating a first computational graph for describing the neural network model; generating a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent; inputting the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, and the intelligent scheduling model is trained through second high-level language samples and hardware description language samples. The present invention is used to solve the defect in the prior art that the deployment of a neural network model in an FPGA takes a long time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code generation, and in particular to a method for generating a hardware description language supporting machine learning and a compilation tool chain. Background Art

[0002] In recent years, neural network models have achieved rapid development. Due to their excellent performance in fields such as image and speech recognition and natural language processing, they have become increasingly familiar and accepted by people and are applied in their respective fields. However, with the development of model construction technology, neural networks have gradually evolved from the initially simple networks with only a few layers (LeNet) to deep convolutional neural networks (VGG-Net) and residual networks (ResNet) with a huge number of parameters and complex network connection relationships. The increasing complexity of neural network models poses a great challenge to underlying computing devices.

[0003] Due to the advantages of rich wiring resources, reprogrammability, high integration, and high programming flexibility of Field Programmable Gate Array (FPGA for short), it has gradually become a preferred choice for neural network model deployment. However, FPGA neural network model deployment also faces many challenges.

[0004] For example, the additional overhead brought by fine-grained reconfigurability: The reconfigurability of FPGA has reached the bit level, which means that it takes a long time to reconfigure the neural network model on the chip each time, greatly extending the development cycle of the application.

[0005] Another example is the programming complexity: For the deployment of neural network models on FPGA, researchers need to be exposed to complex hardware programming, which brings extremely high learning costs to users and also extends the development cycle of the application.

[0006] Therefore, how to quickly complete the deployment of neural networks is an important issue that needs to be solved urgently in the industry. Summary of the Invention

[0007] The present invention provides a method for generating a hardware description language supporting machine learning and a compilation tool chain to solve the defect of long time consumption in deploying neural network models in FPGA in the prior art, and to achieve the rapid deployment of neural network models in FPGA.

[0008] The present invention provides a method for generating a hardware description language supporting machine learning, including:

[0009] Obtaining a neural network model constructed using a first high-level language;

[0010] Generating a first computational graph for describing the neural network model;

[0011] Generate a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent;

[0012] Input the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, where the intelligent scheduling model is trained with second high-level language samples and hardware description language samples.

[0013] According to a method for generating a hardware description language supporting machine learning provided by the present invention, the generating a second high-level language based on the first computational graph and a preset language generation rule includes:

[0014] Extract each first operator node in the first computational graph and the association relationship between the first operator nodes;

[0015] Perform the following processing procedure on any one of the first operator nodes:

[0016] Determine the type and specification information corresponding to the first operator node; determine a language conversion template corresponding to the type, and modify the template specification information in the language conversion template based on the specification information to obtain a sub-second high-level language corresponding to the first operator node;

[0017] Based on the sub-second high-level language, the association relationship, and the preset language generation rule, obtain the second high-level language.

[0018] According to a method for generating a hardware description language supporting machine learning provided by the present invention, the obtaining the second high-level language based on the sub-second high-level language, the association relationship, and the preset language generation rule includes:

[0019] Based on the association relationship, perform a topological sort on the first operator nodes to obtain a topological sequence;

[0020] Based on the topological sequence, number the first operator nodes to obtain a target topological sequence, where the target topological sequence is used to indicate the operation order between the first operator nodes;

[0021] Based on the target topological sequence, determine the execution order of each sub-second high-level language and obtain an initial second high-level language;

[0022] Based on the preset language generation rule, optimize the initial second high-level language to obtain the second high-level language.

[0023] A method for generating a hardware description language supporting machine learning according to the present invention, where inputting the second high-level language into an intelligent scheduling model to obtain the hardware description language output by the intelligent scheduling model includes:

[0024] Input the second high-level language into the intelligent scheduling model to obtain a second computational graph;

[0025] Extract each second operator node in the second computational graph;

[0026] Schedule each second operator node in the second computational graph based on a preset scheduling strategy until each second operator node is scheduled to completion to obtain a scheduling result;

[0027] Generate the hardware description language based on the scheduling result.

[0028] According to a method for generating a hardware description language supporting machine learning provided by the present invention, the second computational graph is used to indicate relevant information that the i-th second operator node is scheduled to the j-th time slice, the scheduling range of the i-th second operator node, and the scheduling time difference corresponding to the i-th second operator node, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1.

[0029] According to a method for generating a hardware description language supporting machine learning provided by the present invention, before scheduling each second operator node in the second computational graph based on a preset scheduling strategy until each second operator node is scheduled to completion to obtain a scheduling result, it further includes: determining the second operator node to be scheduled corresponding to the current time slice and the second operator node that has been scheduled;

[0030] When the second operator node to be scheduled and the second operator node that has been scheduled do not match, determine a target second operator node from the second operator nodes that have been scheduled;

[0031] The scheduling each second operator node in the second computational graph based on a preset scheduling strategy until each second operator node is scheduled to completion to obtain a scheduling result includes:

[0032] Schedule the target second operator node based on the preset scheduling strategy, the relevant information, the scheduling range, and the scheduling time difference until each second operator node is scheduled to completion to obtain a scheduling result.

[0033] The present invention also provides a compilation toolchain, including:

[0034] An acquisition module for acquiring a neural network model constructed using a first high-level language;

[0035] A generation module, configured to generate a first computational graph for describing the neural network model;

[0036] A first high-level synthesis module, configured to generate a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent;

[0037] A second high-level synthesis module, configured to input the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, where the intelligent scheduling model is trained by second high-level language samples and hardware description language samples.

[0038] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for generating a hardware description language supporting machine learning as described in any one of the above is implemented.

[0039] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for generating a hardware description language supporting machine learning as described in any one of the above is implemented.

[0040] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for generating a hardware description language supporting machine learning as described in any one of the above is implemented.

[0041] The method for generating a hardware description language supporting machine learning and a compilation tool chain provided by the present invention are applied to a compilation tool chain, and the compilation tool chain is applied to an FPGA chip. The method obtains a neural network model constructed using a first high-level language through the compilation tool chain; generates a first computational graph for describing the neural network model; generates a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent; inputs the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model. It can be seen that based on this compilation tool chain, the present invention can directly obtain the hardware description language corresponding to the neural network model. This process does not require human participation and does not require human research on hardware programming, saving labor costs, improving the conversion rate of the neural network model to the hardware description language, saving conversion time, and improving the user experience. Furthermore, the hardware description language can be deployed in the FPGA. By converting the neural network model into a hardware description language, the present invention provides an effective conversion basis for deploying the neural network model in the FPGA, solves the defect of long time consumption for deploying the neural network model in the FPGA, and realizes the rapid completion of the deployment operation of the neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 is one of the flow schematic diagrams of the method for generating a hardware description language supporting machine learning provided by the present invention;

[0044] Figure 2 is the second of the flow schematic diagrams of the method for generating a hardware description language supporting machine learning provided by the present invention;

[0045] Figure 3 is the third of the flow schematic diagrams of the method for generating a hardware description language supporting machine learning provided by the present invention;

[0046] Figure 4 is one of the structural schematic diagrams of the compilation toolchain provided by the present invention;

[0047] Figure 5 is the second of the structural schematic diagrams of the compilation toolchain provided by the present invention;

[0048] Figure 6 is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0050] The following will be combined with Figures 1 - 3 to describe the method for generating a hardware description language supporting machine learning of the present invention. This method is applied to a compilation toolchain, and the compilation toolchain is applied to an FPGA.

[0051] Next, in conjunction with the above compilation toolchain, the specific implementation of the method for generating a hardware description language supporting machine learning will be described, as Figure 1 shown:

[0052] Step 101, obtain a neural network model constructed using a first high-level language.

[0053] Among them, the first high-level language includes the Python language.

[0054] Specifically, the compilation toolchain of the present invention includes middleware (TVM). Among them, the present invention can obtain an externally trained neural network model, or can also use the Python front-end extended by TVM to build a neural network model.

[0055] Step 102, generate a first computation graph for describing the neural network model.

[0056] Specifically, after obtaining the neural network model, a third-party library will be called to convert the neural network model into a general neural network exchange format, for example, the Open Neural Network Exchange (ONNX for short); furthermore, the obtained ONNX-format neural network model will be read and converted into a Relay intermediate representation, and the Relay intermediate representation will be used as the first computation graph. Among them, Relay is a versatile programming language used for the intermediate representation of machine learning system expressions. In addition, the Relay intermediate representation is a set of intermediate expressions defined by TVM itself, which is convenient and concise to use, supports subgraphs to be called as functions, and TVM has also implemented a series of computation graph optimization Passes on the basis of Relay. Among them, Pass is a base class that can be selected by users. The use of Relay isolates the neural network model from the generation of backend code, facilitating platform-independent optimization and platform adaptation for developers.

[0057] Step 103, generate a second high-level language based on the first computation graph and a preset language generation rule.

[0058] Among them, the first high-level language and the second high-level language are inconsistent.

[0059] Among them, the second high-level language includes C language.

[0060] In a specific embodiment, after obtaining the first computation graph, traverse the first computation graph, extract each first operator node in the first computation graph, and the association relationship between each first operator node; perform the following processing process on any one first operator node: determine the type and specification information corresponding to the first operator node; determine the language conversion template corresponding to the type, and modify the template specification information in the language conversion template based on the specification information to obtain the sub-second high-level language corresponding to the first operator node; based on the sub-second high-level language, the association relationship, and the preset language generation rule, obtain the second high-level language.

[0061] Among them, the types of the first operator nodes include: addition operation, subtraction operation, multiplication operation, division operation, etc. Different types correspond to different language conversion templates, and the difference between the first operator nodes of the same type is only the specification information participating in the operation. Among them, the type is determined based on the operator.

[0062] Among them, the specification information is a tensor. That is, the difference between the first operator nodes of the same type lies in the size of the tensors participating in the operation. A tensor is a data container that includes data of various dimensions. For example, a scalar can be called a zero-dimensional tensor, and a matrix can be called a two-dimensional tensor, etc. It is the most basic data format in machine learning. In a neural network computation graph, tensors are the inputs and outputs of operators. For example, a convolution operator generally takes a three-dimensional tensor as input and output, and a fully connected operator generally takes a one-dimensional tensor as input and output. However, the neural network computation graph is a relatively abstract high-level representation and does not explain how the underlying operations should be carried out. In fact, the operations between tensors are, at the bottom layer, the operations between each element of the tensors. For example, when defining the addition of two one-dimensional tensors, the underlying operation is to add the elements at the corresponding positions of the two tensors pairwise to form a new tensor. This computational pattern of traversing each element position of the tensor is essentially the same as the loops and nested loops in the C language. Therefore, the compilation toolchain of the present invention uses a multi-layer loop method to express each tensor operation, which not only conforms to the principle of tensor operations but also is convenient to express in the C language. Specifically, since the difference between the first operator nodes of the same type lies in the size of the tensors participating in the operation, and the element-level operations are the same, this feature is also the reason why the language conversion template of the present invention is effective. Therefore, for the first operator nodes of the same type, the corresponding C language can be obtained by using the pre-set language conversion template, and the tensor can be modified for specific operations.

[0063] For example, mapping the first operator node corresponding to a fully connected layer to two nested loops, and the two nested loops are the scales of the input and output tensors. The same is true for the first operator node corresponding to the convolutional layer. Load the pre-set language conversion template and modify the convolutional kernel size, convolutional stride, and the number of channels and the length and width of the image to be convolved for specific convolutional operations.

[0064] Among them, an example code of the language conversion template is as follows:

[0065]

[0066] Among them, i and j represent loop control variables, FM0_SIZE represents the dimension size of the input tensor of the fully connected layer, FM1_SIZE represents the dimension size of the output tensor of the fully connected layer, fm0_ptr represents the input tensor of the fully connected layer, fm1_ptr represents the output tensor of the fully connected layer, and fclweight represents the weight of the fully connected layer.

[0067] Among them, for the implementation schematic diagram of the above example code, see Figure 2 .

[0068] Among them, the preset language generation rules are used to indicate the generation of C language code for high-level synthesis. Among them, the preset language generation rules can also be defined as the C language code rules for high-level synthesis. When generating C language based on the first computational graph, the preset language generation rules must be followed. Otherwise, the generated C language cannot obtain the hardware description language. The preset language generation rules include:

[0069] (1) There should be no C language operations involving system calls in the program, that is, there should be no functions designed for system calls, such as printing characters, reading and writing files, etc.;

[0070] (2) There should be no dynamic memory allocation and release operations in the program, that is, there should be no functions designed for dynamic memory allocation and release, such as the malloc function and the free function. Since dynamic memory allocation is a runtime operation, high-level synthesis cannot obtain the size of the required memory, and dynamic memory allocation requires system library support;

[0071] (3) There should be no pure software-defined recursive operations in the program, that is, there should be no recursive function calls, because the FPGA hardware backend does not support stacks;

[0072] (4) The standard template library cannot be used in the program, that is, there should be no functions in the STL standard template library. Since there are many operations in the standard template library that are not supported by high-level synthesis, such as dynamic memory allocation and recursion;

[0073] (5) There should be no forced pointer type conversion for custom pointers in the program, that is, there should be no explicit or implicit pointer type conversion. Since the FPGA hardware memory is relatively fixed, a pointer can only point to a fixed block of memory.

[0074] Next, Table 1 is used to illustrate the preset language generation rules by way of example:

[0075]

[0076] Table 1 Preset Language Generation Rules

[0077] Under the restrictions of the above rules, the compilation toolchain can generate C language that conforms to high-level synthesis. However, for C code with the same function, different coding styles and coding methods have a great impact on the results. For example, due to differences in coding styles and coding methods, problems such as excessive area of high-level synthesis results and excessive use of flip-flops occur. To solve the above problems, the present invention summarizes a series of C code coding styles that are not friendly to high-level synthesis and forms a set of C code coding styles that are friendly to high-level synthesis on this basis.

[0078] Specifically, the C language code generated by the compilation toolchain cannot have functions with complex definitions. Here, the complexity of a function is reflected in the number of parameters and the type of return value. Under traditional CPU architectures, function parameters and return values are passed through registers, but in FPGAs, function parameters and return values will be synthesized into the input ports of function modules. Therefore, a larger number of parameters and more complex return values will cause the function module to occupy more port resources and have a larger area. Secondly, the C code cannot perform overly frequent type conversion operations, as type conversion requires additional connection overhead. Therefore, a larger number of type conversions will result in too high a latency and low operating efficiency. Also, the C code cannot define too many redundant variables. Compared with CPUs, resources on FPGAs are very scarce, and redundant variables will cause the synthesis result to require more flip-flop resources, which is actually a waste of resources.

[0079] Therefore, certain restrictions are imposed on the coding style of the C language, as follows:

[0080] (1) Complex functions should be avoided in the program. Complex means that the function has a large number of parameters and return values and a complex format;

[0081] (2) Uncertain data types and frequent type conversions should be avoided in the program;

[0082] (3) Too many redundant variables should be avoided in the program.

[0083] Next, Table 2 is used to illustrate the C language coding style by way of example:

[0084]

[0085] Table 2 C Language Coding Style

[0086] Among them, / / / / / indicates that there is no example.

[0087] In view of the characteristics of high-level synthesis programs and FPGA hardware, the present invention example proposes and summarizes C code specifications that are friendly to high-level synthesis. These C code specifications that are friendly to high-level synthesis require the intermediate code generation module to generate code under certain restrictions. Experimental data shows that compared with the code that does not comply with the specifications, the intermediate C code that complies with the high-level synthesis-friendly specifications not only greatly saves the use of hardware resources, but also significantly improves the operating efficiency of the program deployed on the FPGA.

[0088] In a specific embodiment, based on the association relationship, topological sorting is performed on each first operator node to obtain a topological sequence; based on the topological sequence, each first operator node is numbered to obtain a target topological sequence, and the target topological sequence is used to indicate the operation order between each first operator node; based on the target topological sequence, the execution order of each sub-second high-level language is determined, and an initial second high-level language is obtained; based on the preset language generation rules, the initial second high-level language is optimized to obtain a second high-level language.

[0089] Among them, when generating C language based on the language conversion module, at least one operator is scheduled, and the multiple operators scheduled in this process are defined as an operator set.

[0090] Through the above embodiments, the present invention solves the problem of generating C language for each first operator node through technical means such as a language conversion module and language generation rules. Then, how to solve the communication problem between each first operator node? The present invention adopts the technical means of global numbering to build a bridge for data communication between each first operator node to solve the communication problem between each first operator node.

[0091] Among them, a first computational graph represents a complete computational process, including: input of data, transfer of data between first operator nodes, and output of data. The specific implementation of the technical means of global numbering is as follows:

[0092] Number each first operator node according to the topological sequence. If the first calculation Figure 1 There are a total of N first operator nodes. The initial first operator node of the first computational graph is defined as node No. 1, the terminating first operator node is located as node No. N, and the input tensor is defined as tensor No. 0, where N is an integer greater than 1.

[0093] Through the above technical means of global numbering, the present invention numbers and names the first operator nodes and tensors according to the topological sequence when traversing the first computational graph. In this way, when generating code, if the code content involves data interaction between two first operator nodes, the object of data operation can be directly determined through the number.

[0094] Step 104, input the second high-level language into the intelligent scheduling model to obtain the hardware description language output by the intelligent scheduling model.

[0095] Among them, the intelligent scheduling model is trained through second high-level language samples and hardware description language samples.

[0096] Among them, the initial model framework of the intelligent scheduling model is the framework bambu, which is embedded into bambu based on a pre-designed intelligent scheduling algorithm. Furthermore, it is trained based on second high-level language samples and hardware description language samples to obtain the final intelligent scheduling model.

[0097] Among them, the specific implementation code of the intelligent scheduling algorithm is as follows:

[0098]

[0099] Among them, bb represents the basic block in the second high-level language, N represents the set of operators that can be scheduled next, E represents the set of dependencies between the second operator nodes in the second computational graph, C represents the set of operators with resource conflicts, G represents the second computational graph, M represents the scheduling result table corresponding to the scheduling result, o i represents the operator that has been scheduled, c i represents o i the time slice to which o is scheduled.

[0100] In a specific embodiment, the specific implementation of obtaining the hardware description language corresponding to C language by using the trained intelligent scheduling model is as follows: input the second high-level language into the intelligent scheduling model to obtain the second computational graph; extract each second operator node in the second computational graph; schedule each second operator node in the second computational graph based on the preset scheduling strategy until each second operator node is scheduled to obtain the scheduling result; generate the hardware description language based on the scheduling result.

[0101] Among them, the second computational graph is a fine-grained directed acyclic graph.

[0102] Among them, the preset scheduling strategy is an execution strategy for converting C language into a hardware description language based on a directed acyclic graph designed in advance.

[0103] In a specific embodiment, the second computational graph is used to indicate the relevant information of scheduling the i-th second operator node to the j-th time slice, the scheduling range of the i-th second operator node, and the scheduling time difference corresponding to the i-th second operator node, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1.

[0104] Among them, the scheduling range of the i-th second operator node is the time range in which the i-th second operator node can move; the scheduling time difference corresponding to the i-th second operator node is the difference between the time point corresponding to the i-th second operator node in the earliest best scheduling algorithm and the time point corresponding to the i-th second operator node in the latest best scheduling algorithm. Among them, the earliest best scheduling algorithm is used to indicate the earliest time corresponding to the i-th second operator node when all second operator nodes are scheduled, and the latest best scheduling algorithm is used to indicate the latest time corresponding to the i-th second operator node when all second operator nodes are scheduled.

[0105] Specifically, as Figure 3 shown, taking 4 second operator nodes as an example for illustration:

[0106] Among them, the 4 second operator nodes are respectively represented by c1, c2, c3, and c4. In Figure 3 I, II, and III represent numbers, that is, the execution order. In Figure 3 in the coordinate axes, c-step represents the second operator node, op represents the execution order, Currentschedule represents the specific scheduling implementation of the scheduling policy, Current possible movements represents the second operator nodes that can be moved at the current moment, and All possible movements represents all the second operator nodes that can be moved. Among them, 1 means schedulable, that is, movable, and 0 means non-schedulable, that is, immovable.

[0107] Next, Table 3 and Table 4 are used to illustrate the multiple layers of the intelligent scheduling model and the input and output of each layer through examples.

[0108] Layer name Input Output Conv2d 3*50*50 64*50*50 ReLU 64*50*50 64*50*50 Conv2d 64*50*50 64*50*50 ReLU 64*50*50 64*50*50 MaxPool2d 64*50*50 64*25*25 Conv2d 64*25*50 128*25*25 ReLU 128*25*25 128*25*25 MaxPool2d 128*25*25 128*12*12 Conv2d 128*12*12 256*12*12 ReLU 256*12*12 256*12*12 Conv2d 256*12*12 256*12*12 ReLU 256*12*12 256*12*12 MaxPool2d 256*12*12 256*6*6

[0109] Table 3 Multiple layers of the intelligent scheduling model and the input and output of each layer

[0110]

[0111]

[0112] Table 4 represents the multiple layers of the intelligent scheduling model and the input and output of each layer

[0113] Among them, the layer exemplified in Table 3 is used for feature extraction, and the layer exemplified in Table 4 is used for classification, and is used to output the second computer nodes that can be moved at the current time point.

[0114] Specifically, when training the intelligent scheduling model, the initial hyperparameters are established based on the best results of supervised learning. After this result, the model is trained according to a large number of second high-level language samples and hardware description language samples. During the training process, each time a legal second operator node is selected, that is, a second operator node that will not affect the execution order after moving, and then move it down.

[0115] Among them, the training code of the intelligent scheduling model is as follows:

[0116]

[0117] Among them, N represents the number of Monte Carlo search times, T represents the time step of Monte Carlo simulation, represents the state in reinforcement learning, represents the action in reinforcement learning, ρ refers to the neural network parameters, α refers to the learning rate, and Δρ is the gradient used to update the model.

[0118] In a specific embodiment, before scheduling each second operator node in the second computation graph based on a preset scheduling policy until the scheduling of each second operator node is completed to obtain a scheduling result, determine the second operator node to be scheduled corresponding to the current time slice and the scheduled second operator node that has undergone scheduling; when the second operator node to be scheduled does not match the scheduled second operator node, determine the target second operator node from the scheduled second operator nodes; based on the preset scheduling policy, relevant information, scheduling range, and scheduling time difference, schedule the target second operator node until the scheduling of each second operator node is completed to obtain a scheduling result.

[0119] Among them, there is a corresponding relationship between time points and time slices.

[0120] Specifically, when the intelligent scheduling model performs scheduling, determine whether there are second operator nodes with resource conflicts at the current time point. Among them, when it is determined that the second operator node to be scheduled does not match the scheduled second operator node, it is determined that there are second operator nodes with resource conflicts. At this time, among the scheduled second operator nodes, the scheduled second operator node corresponding to the shortest time when the scheduling is completed is used as the target second operator node, and scheduling is performed downward from the target second operator node until all second operator nodes meet the requirements of resource constraints. Among them, it is necessary to judge whether its successor is "adjacent" during scheduling. After scheduling these adjacent corresponding second operator nodes, schedule the target second operator node.

[0121] Among them, when scheduling any second operator node, due to the introduction of resource conflicts, it is necessary to update the local and global resource distributions. On this basis, each conflicting second operator node is scheduled downward at each step until the hardware constraints are met. The specific implementation code is as follows:

[0122]

[0123]

[0124] Among them, P represents the probability vector of different second operator nodes being scheduled, and RL refers to reinforcement learning.

[0125] Among them, during the scheduling process, an operation of adding time constraints is performed. By modifying the time slice of the second operator node, the schedulability when resource scheduling conflicts cannot be resolved is increased. The specific implementation code is as follows:

[0126]

[0127] Among them, op represents the second operator node to be scheduled.

[0128] Among them, bambu calls the intelligent scheduling algorithm to enter the main program through the parameter class, calls the comprehensive step management class through the main program, and then calls the comprehensive step factory class, and then calls the algorithm class to perform the scheduling of the second operator node based on the preset scheduling strategy.

[0129] Finally, the generated hardware description language is logically synthesized and physically designed and finally deployed on the FPGA. Specifically, the specific implementation process of the present invention is summarized as follows: Based on the computational graph describing neural network calculations, a corresponding intermediate C code program is generated. This C code program is then input into a high-level synthesis program to generate a hardware description language that supports FPGA programming. The user defines the neural network model in the form of a computational graph, and the intermediate code generation module generates the corresponding intermediate C code program, which is used to generate a hardware description language that supports FPGA programming through the high-level synthesis program. This compilation toolchain supports a variety of general mathematical calculations as well as the training and inference tasks of multi-layer perceptrons and various convolutional neural networks.

[0130] The example of the present invention aims at the problem of optimizing hardware resource scheduling in high-level synthesis, proposes a scheduling algorithm based on reinforcement learning, realizes a high-level synthesis solution based on the existing semi-automatic high-level synthesis framework bambu, and proposes and summarizes the transformation method of the scheduling problem and the reinforcement learning scheduling model, including but not limited to data preprocessing, training methods and related parameters. The scheduling problem is transformed through definition to generate the model input, and based on the scheduling model, a cycle performance that meets the resource constraints and is close to the theoretical optimum is given.

[0131] The present invention provides a method for generating a hardware description language supporting machine learning and a compilation toolchain. The method is applied to the compilation toolchain, and the compilation toolchain is applied to the FPGA chip. The method obtains a neural network model constructed using a first high-level language through the compilation toolchain; generates a first computational graph for describing the neural network model; based on the first computational graph and a preset language generation rule, generates a second high-level language, where the first high-level language and the second high-level language are inconsistent; inputs the second high-level language into an intelligent scheduling model to obtain the hardware description language output by the intelligent scheduling model. It can be seen that based on this compilation toolchain, the present invention can directly obtain the hardware description language corresponding to the neural network model. This process does not require human participation and does not require human research on hardware programming, saving labor costs, improving the conversion rate of the neural network model to the hardware description language, saving the conversion time, improving the user experience. Furthermore, the hardware description language can be deployed in the FPGA. The present invention provides an effective conversion basis for deploying the neural network model in the FPGA by converting the neural network model into a hardware description language, solves the defect of the long time consumption for deploying the neural network model in the FPGA, and realizes the rapid completion of the deployment operation of the neural network model.

[0132] The compilation toolchain provided by the present invention will be described below. The compilation toolchain described below can be mutually referred to corresponding to the method for generating a hardware description language supporting machine learning described above, and repeated parts will not be elaborated. For example, Figure 4 As shown, the compilation toolchain includes:

[0133] An acquisition module 401, configured to acquire a neural network model constructed in a first high-level language;

[0134] A generation module 402, configured to generate a first computational graph for describing the neural network model;

[0135] A first high-level synthesis module 402, configured to generate a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent;

[0136] A second high-level synthesis module 403, configured to input the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, and the intelligent scheduling model is trained by a second high-level language sample and a hardware description language sample.

[0137] In a specific embodiment, the first high-level synthesis module 402 is specifically configured to extract each first operator node in the first computational graph, as well as the association relationship between each first operator node; perform the following processing process on any one first operator node: determine the type and specification information corresponding to the first operator node; determine a language conversion template corresponding to the type, and modify the template specification information in the language conversion template based on the specification information to obtain a sub-second high-level language corresponding to the first operator node; based on the sub-second high-level language, the association relationship, and a preset language generation rule, obtain the second high-level language.

[0138] In a specific embodiment, the first high-level synthesis module 402 is specifically configured to perform a topological sorting on each first operator node based on the association relationship to obtain a topological sequence; number each first operator node based on the topological sequence to obtain a target topological sequence, where the target topological sequence is used to indicate the operation order between each first operator node; determine the execution order of each sub-second high-level language based on the target topological sequence, and obtain an initial second high-level language; optimize the initial second high-level language based on a preset language generation rule to obtain the second high-level language.

[0139] In a specific embodiment, the second high-level synthesis module 403 is specifically configured to input the second high-level language into an intelligent scheduling model to obtain a second computational graph; extract each second operator node in the second computational graph; schedule each second operator node in the second computational graph based on a preset scheduling strategy until each second operator node is scheduled to completion to obtain a scheduling result; generate a hardware description language based on the scheduling result.

[0140] In a specific embodiment, the second computation graph is used to indicate the relevant information of scheduling the i-th second operator node to the j-th time slice, the scheduling range of the i-th second operator node, and the scheduling time difference corresponding to the i-th second operator node, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1.

[0141] In a specific embodiment, the second high-level synthesis module 403 is further configured to determine the second operator node to be scheduled and the scheduled second operator node corresponding to the current time slice; when the second operator node to be scheduled does not match the scheduled second operator node, determine the target second operator node from the scheduled second operator nodes; specifically, the second high-level synthesis module 403 is configured to schedule the target second operator node based on a preset scheduling policy, relevant information, scheduling range, and scheduling time difference until the scheduling of each second operator node is completed to obtain a scheduling result.

[0142] Specifically, the compilation tool chain according to the generation process of the hardware description language can also be divided into a front end, a middle end, and a back end. For details, see Figure 5 , and the compilation tool chain includes a compilation framework front end 501, a compilation framework middle end 502, and a compilation framework back end 503.

[0143] Specifically, the embodiment of the present invention uses an existing open-source neural network compilation framework (TVM) as the compilation framework front end 501, uses high-level synthesis technology (High Level Synthesis, abbreviated as HLS) to design the coding framework middle end 502, and uses high-level synthesis software and FPGA development software as the compilation framework back end 503. Each module of the compilation framework front end 501, the compilation framework middle end 502, and the compilation framework back end 503 corresponds to input data. Therefore, it is necessary to design the output data of each module so that each module can be organically synthesized into a whole, that is, the compilation tool chain.

[0144] Among them, the compilation framework front end 501 is configured to obtain a neural network model, convert it into a neural network interchange format, and then obtain a Relay intermediate representation.

[0145] The compilation framework middle end 502 is configured to traverse the first computation graph and generate C language.

[0146] The compilation framework back end 503 is configured to generate a second computation graph, schedule the second operator nodes of the second computation graph, and obtain a hardware description language.

[0147] The compilation tool chain according to the generation process of the hardware description language further includes a deployment back end for deploying the hardware description language in an FPGA.

[0148] Figure 6Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 6 shown. The electronic device may include: a processor 601, a communications interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communications interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 can call the logical instructions in the memory 603 to execute a method for generating a hardware description language for supporting machine learning. The method includes: obtaining a neural network model constructed using a first high-level language; generating a first computational graph for describing the neural network model; generating a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent; inputting the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, and the intelligent scheduling model is trained by second high-level language samples and hardware description language samples.

[0149] In addition, when the logical instructions in the above-mentioned memory 603 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0150] The present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating a hardware description language for supporting machine learning provided by the above-mentioned various methods. The method includes: obtaining a neural network model constructed using a first high-level language; generating a first computational graph for describing the neural network model; generating a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent; inputting the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, and the intelligent scheduling model is trained by second high-level language samples and hardware description language samples.

[0151] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for generating a hardware description language supporting machine learning provided by the above-mentioned various methods. The method includes: obtaining a neural network model constructed using a first high-level language; generating a first computational graph for describing the neural network model; generating a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent; inputting the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, and the intelligent scheduling model is trained using second high-level language samples and hardware description language samples.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a hardware description language supporting machine learning, characterized in that, The method includes: Obtaining a neural network model constructed using a first high-level language; Generating a first computational graph for describing the neural network model; Generating a second high-level language based on the first computational graph and a preset language generation rule, where the first high-level language and the second high-level language are inconsistent; Inputting the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, where the intelligent scheduling model is trained using second high-level language samples and hardware description language samples; The inputting the second high-level language into the intelligent scheduling model to obtain the hardware description language output by the intelligent scheduling model includes: Inputting the second high-level language into the intelligent scheduling model to obtain a second computational graph, where the second computational graph is a fine-grained directed acyclic graph; Extracting each second operator node in the second computational graph; Scheduling each second operator node in the second computational graph based on a preset scheduling policy until each second operator node is scheduled to completion to obtain a scheduling result; Generating the hardware description language based on the scheduling result; The second computational graph is used to indicate the relevant information of the i-th second operator node scheduled to the j-th time slice, the scheduling range of the i-th second operator node, and the scheduling time difference corresponding to the i-th second operator node, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1. The scheduling range of the i-th second operator node is the time range within which the i-th second operator node can be moved; the scheduling time difference corresponding to the i-th second operator node is the difference between the time point corresponding to the i-th second operator node in the earliest best scheduling algorithm and the time point corresponding to the i-th second operator node in the latest best scheduling algorithm. The earliest best scheduling algorithm is used to indicate the earliest time corresponding to the i-th second operator node when all second operator nodes are scheduled to completion, and the latest best scheduling algorithm is used to indicate the latest time corresponding to the i-th second operator node when all second operator nodes are scheduled to completion.

2. The method for generating a hardware description language supporting machine learning according to claim 1, wherein The generating the second high-level language based on the first computational graph and a preset language generation rule includes: Extracting each first operator node in the first computational graph and the association relationship between the first operator nodes; Performing the following processing procedure on any one of the first operator nodes: Determining the type and specification information corresponding to the first operator node; determining a language conversion template corresponding to the type, and modifying the template specification information in the language conversion template based on the specification information to obtain a sub-second high-level language corresponding to the first operator node; Obtaining the second high-level language based on the sub-second high-level language, the association relationship, and the preset language generation rule.

3. The method for generating a hardware description language supporting machine learning according to claim 2, characterized in that, The obtaining the second high-level language based on the sub-second high-level language, the association relationship, and the preset language generation rule includes: Performing a topological sort on the first operator nodes based on the association relationship to obtain a topological sequence; Number each of the first operator nodes based on the topological sequence to obtain a target topological sequence, where the target topological sequence is used to indicate the operation order among the first operator nodes; Determine the execution order of each of the sub-second high-level languages based on the target topological sequence, and obtain an initial second high-level language; Optimize the initial second high-level language based on the preset language generation rules to obtain the second high-level language.

4. The method for generating a hardware description language supporting machine learning according to any one of claims 1-3, characterized in that, Before scheduling each second operator node in the second computational graph based on the preset scheduling strategy until each second operator node is scheduled to obtain a scheduling result, it further includes: determining a second operator node to be scheduled corresponding to the current time slice and a scheduled second operator node that has undergone scheduling; When the second operator node to be scheduled does not match the scheduled second operator node, determine a target second operator node from the scheduled second operator nodes; Scheduling each second operator node in the second computational graph based on the preset scheduling strategy until each second operator node is scheduled to obtain a scheduling result, includes: Scheduling the target second operator node based on the preset scheduling strategy, the relevant information, the scheduling range, and the scheduling time difference until each second operator node is scheduled to obtain a scheduling result.

5. A compilation toolchain, characterized in that, Includes: An acquisition module for acquiring a neural network model constructed using a first high-level language; A generation module for generating a first computational graph for describing the neural network model; A first high-level synthesis module for generating a second high-level language based on the first computational graph and preset language generation rules, where the first high-level language and the second high-level language are inconsistent; A second high-level synthesis module for inputting the second high-level language into an intelligent scheduling model to obtain a hardware description language output by the intelligent scheduling model, where the intelligent scheduling model is trained using second high-level language samples and hardware description language samples; Inputting the second high-level language into the intelligent scheduling model to obtain the hardware description language output by the intelligent scheduling model, includes: Inputting the second high-level language into the intelligent scheduling model to obtain a second computational graph, where the second computational graph is a fine-grained directed acyclic graph; Extract each second operator node in the second computational graph; Schedule each second operator node in the second computational graph based on a preset scheduling strategy until each second operator node is scheduled to obtain a scheduling result; Generate the hardware description language based on the scheduling result; The second computation graph is used to indicate the relevant information of scheduling the i-th second operator node to the j-th time slice, the scheduling range of the i-th second operator node and the scheduling time difference corresponding to the i-th second operator node, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1. The scheduling range of the i-th second operator node is the time range within which the i-th second operator node can be moved; the scheduling time difference corresponding to the i-th second operator node is the difference between the time point corresponding to the i-th second operator node in the earliest best scheduling algorithm and the time point corresponding to the i-th second operator node in the latest best scheduling algorithm; the earliest best scheduling algorithm is used to indicate the earliest time corresponding to the i-th second operator node when all the second operator nodes are scheduled, and the latest best scheduling algorithm is used to indicate the latest time corresponding to the i-th second operator node when all the second operator nodes are scheduled.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for generating a hardware description language supporting machine learning according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for generating a hardware description language supporting machine learning according to any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for generating a hardware description language supporting machine learning according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method, electronic device and computer program product for processing machine learning model

    US20210034582A1

  • Hardware platform specific operator fusion in machine learning

    WO2021114530A1