A Circuit Logic Synthesis Method for FPGA Mapping Based on Reinforcement Learning

Through a logical synthesis method based on reinforcement learning, using graph attention network to analyze circuit structure and statistical features, the problems of difficulty in selecting circuit optimization strategies and insufficient applicability of models in traditional methods are solved, and efficient optimization of new circuits is achieved.

CN120068743BActive Publication Date: 2025-07-11BOYA XINKE (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510537106.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-11
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Traditional logic synthesis methods are difficult to choose the optimal optimization strategy when facing different circuit structures, and the existing reinforcement learning models cannot be directly applied to new circuits after training, which takes a long time to train, resulting in insufficient model optimization capabilities.

Method used

A logical synthesis method based on reinforcement learning is constructed, and the circuit structure is analyzed through the graph attention network, combined with statistical features and comprehensive sequence features, and a model that can directly optimize the new circuit is trained, and a graph embedded representation and enhanced state are used for logical optimization.

Benefits of technology

It realizes precise optimization of different circuits, improves the efficiency and performance of logic synthesis, avoids the time consumption caused by retraining, and provides better optimization effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068743B_ABST
    Figure CN120068743B_ABST
Patent Text Reader

Abstract

The present invention provides a circuit logic synthesis method for FPGA mapping based on reinforcement learning to solve the problem of insufficient ability of model optimization for unseen circuits in traditional netlist optimization methods. First, a dataset for logic synthesis is constructed, and then the dataset is preprocessed to obtain the statistical features and graph structure of the circuit; subsequently, a logic synthesis model for FPGA mapping based on reinforcement learning is constructed and pre-trained using the preprocessed dataset, and finally, the pre-trained logic synthesis model for FPGA mapping based on reinforcement learning is used to perform logic optimization on the circuit. The circuit feature extraction module of the present invention uses a GNN module to analyze the logic circuit and netlist structure to obtain graph embeddings, and analyzes the statistical features of the logic circuit and netlist through an FCN module, and splices the two parts of the features to obtain an enhanced representation state of the precise circuit, improving the understanding ability of the reinforcement learning model between different circuits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electronic design automation, and particularly relates to a circuit logic synthesis method for FPGA mapping based on reinforcement learning. Background Art

[0002] With the continuous expansion of the scale of modern very large scale integrated circuit design, logic synthesis, as a key step in the electronic design automation (EDA) process, has an important impact on chip performance and resource utilization. The main goal of logic synthesis is to convert the register transfer level (RTL) description of a circuit into a gate-level netlist and optimize its area, power consumption, and timing performance, which is a key link in the design of field programmable gate arrays (FPGAs). Logic synthesis usually includes three stages: translation, logic optimization, and mapping. Among them, the logic optimization stage simplifies the logical expression of the circuit to reduce the circuit scale and optimize its performance.

[0003] In the process of logic optimization, optimization operators in the open-source tool ABC are often used to form a synthesis sequence to achieve the purpose of optimizing the logic circuit. However, the effects of different optimization sequences composed of the same optimization operators vary greatly. Traditional methods use heuristic scripts (such as resyn2 in ABC) to optimize the circuit, but the applicability of such methods is limited, and it is difficult to select the optimal optimization strategy for different circuit structures. Therefore, it is of great significance to find a dedicated optimization sequence for the circuit. However, due to the extremely large search space of logic optimization, it is computationally infeasible to enumerate all possible optimization sequences.

[0004] In recent years, reinforcement learning (RL) has been introduced into logic synthesis optimization to automatically search for and select the optimal optimization sequence. Designers can use the reinforcement learning model to avoid the exhaustive process of the circuit synthesis flow, obtain a dedicated optimization sequence for the circuit, and improve the performance of the chip. However, most of the relevant reinforcement learning models are discarded after being trained on a circuit, and new circuits need to be retrained. The training process is time-consuming, and this extremely long time consumption is unacceptable when dealing with large circuits. Therefore, a suitable method is needed to directly optimize the circuit when facing a new circuit and achieve better performance than traditional methods (such as resyn2).

[0005] One of the solutions to this problem is to continuously train the circuit on the training set so that it can directly optimize the new circuit without retraining for each circuit. Currently, the general model for netlist optimization uses the statistical features of logic circuits and netlists as the features for reinforcement learning. However, for two circuits with the same statistical features, their structures may be completely different. The defect brought by the method that only uses statistical features makes it difficult for the model to accurately optimize unseen circuits. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present invention provides a logic synthesis method for FPGA mapping based on reinforcement learning to solve the problem of insufficient ability of the model to optimize unseen circuits in traditional netlist optimization methods. Specifically, the present invention constructs the structures of logic circuits and netlists into directed acyclic graphs, and obtains the graph embedding representation by parsing the correlation between different circuit structures through a graph attention network (GAT). The graph embedding representation, statistical features, and synthesis sequence features are used as the states enhanced by reinforcement learning. The reinforcement learning model is trained on the training set using the enhanced states to obtain a model that can directly optimize new circuits without retraining.

[0007] The present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides a logic synthesis method for FPGA mapping based on reinforcement learning, and the method includes:

[0009] Step S1: Construct a data set for logic synthesis; the data set contains input data, and the input data is a circuit described by RTL verilog;

[0010] Step S2: Preprocess the data set;

[0011] Step S3: Construct a logic synthesis model for FPGA mapping based on reinforcement learning, and perform pre-training using the preprocessed data set;

[0012] The logic synthesis model for FPGA mapping based on reinforcement learning includes a circuit feature extraction module and a reinforcement learning module;

[0013] The circuit feature extraction module is used to learn the statistical features, synthesis sequence features, and graph structure of the circuit; the graph structure includes the structure of the logic circuit and the structure of the netlist;

[0014] The reinforcement learning module is used to splice the learned statistical features, synthesis sequence features, and graph structure of the circuit to obtain the state of reinforcement learning, and select a suitable optimization algorithm for the circuit according to the state;

[0015] Preferably, in the logic synthesis model for FPGA mapping based on reinforcement learning, the circuit feature extraction module includes an FCN (Fully Connected Network) module and a GNN (Graph Neural Network) module;

[0016] Preferably, the input of the FCN module is the statistical features and synthesis sequence features of the circuit, specifically including a first linear layer, a first ReLU layer, a second linear layer, a second ReLU layer, and a third linear layer;

[0017] Preferably, the input of the GNN module is the structure of the logic circuit and the structure of the netlist, specifically including a first GAT (Graph Attention Network) layer, a first ReLU layer, a second GAT layer, a second ReLU layer, a third GAT layer, and a pooling layer;

[0018] The first GAT layer, the second GAT layer, and the third GAT layer all use the attention mechanism to learn the weight distribution of different neighbors, aggregate the neighbor node features, and update their own node features; the differences in the structures of the three GAT layers lie in the different input and output dimensions. At the same time, the first GAT layer uses 4 - head attention, allowing the model to learn node relationships from different angles, capture richer information, and enhance the expression ability for the graph structure. The mathematical description of the aggregation is as follows:

[0019]

[0020] Among them, is a non - linear activation function, is the attention coefficient, is the feature transformation matrix, is the node 's feature, is the neighbor node set of node ;

[0021]

[0022] Among them is the weight vector of the attention mechanism, concat represents the concatenation operation, and LeakyReLU is the activation function, is the feature of node 's, is the index variable, is the feature of node ;

[0023] The pooling layer refers to performing max - pooling and average - pooling operations on the processed graph structure and then performing a concatenation operation;

[0024] Preferably, the reinforcement learning module includes a state, an Actor network, and a Critic network;

[0025] where the state is obtained by concatenating vectors obtained from the FCN module and the GNN module;

[0026] The Actor network processes the state using a neural network, analyzes the current circuit state, and obtains a probability vector for the corresponding action, and selects the corresponding circuit optimization action according to this vector , and this optimization action is an optimization algorithm of the logic synthesis tool ABC; using this action to optimize the circuit in ABC, the difference before and after the circuit is obtained as the reward for this action ; specifically, it includes a first linear layer, a first ReLU layer, and a second linear layer;

[0027] The reward for the action refers to the difference in the number of 6-input lookup tables in the circuit netlist before and after applying the optimization action, and the mathematical description is as follows:

[0028]

[0029] where is the number of 6-input lookup tables after optimization, is used to normalize the obtained reward to facilitate the stable update of the network;

[0030] The Critic network processes the state using a neural network, analyzes the current circuit state, and obtains an evaluation of the state , provides an advantage function for evaluating the pros and cons of the currently selected action, specifically including a first linear layer, a first ReLU layer, and a second linear layer. The reinforcement learning model is updated according to the pros and cons of this action, and the advantage function refers to the Generalized Advantage Estimation (GAE), and the mathematical description is as follows:

[0031]

[0032] where represents the discount factor, represents the evaluation of the circuit state by the Critic network;

[0033] This advantage function is used to update the entire reinforcement learning model, and the mathematical representation is as follows:

[0034]

[0035] where, represents the probability difference between the decisions of the new and old models , is the clipping rate of the clipping function, clamps within Between them, the update amplitude of the control network ensures smooth update. This formula is used to update the entire logic synthesis model including the feature extraction module and the reinforcement learning module.

[0036] Step S4: Use the pre-trained logic synthesis model for FPGA mapping based on reinforcement learning to perform logic optimization on the new circuit to verify the cross-circuit logic optimization ability of the model.

[0037] Preferably, the step S2 includes the following steps:

[0038] S201: Convert the circuit described by RTL verilog into a circuit in AIG format, and obtain the initial statistical features and the initial graph structure of the AIG circuit according to the initial AIG format circuit. The node types in the AIG format circuit include logic AND gate nodes, standard input nodes, and standard output nodes; Map the circuit described in AIG format to a circuit in the initial netlist format containing 6-input lookup tables through FPGA, and obtain the statistical features and the initial graph structure of the netlist according to the netlist format circuit. The node types in the netlist format circuit include 6-input lookup table nodes, standard input nodes, and standard output nodes;

[0039] The statistical features of the AIG circuit refer to the number of logic AND gate nodes, the maximum number of layers, the average number of layers, the minimum number of layers, the average value and variance of the fan-outs of all inputs, and the average value and variance of the fan-outs of all logic AND gate nodes in the current logic circuit; The statistical features of the netlist refer to the number of 6-input lookup tables, the number of layers, and the number of edges in the current netlist;

[0040] Construct a directed acyclic graph according to the AIG format circuit As the graph structure of the AIG circuit; where is the set of nodes, and each node has corresponding node attributes, including node type, the number of fan-ins, and the number of NOT gates; is the set of edges, where each edge represents a directed connection between the corresponding two nodes; Construct a directed acyclic graph according to the circuit in netlist format As the graph structure of the netlist; where is the set of nodes, and each node has corresponding node attributes, including node type, the number of fan-ins, and the truth table, is the set of edges, where each edge represents a directed connection between the corresponding two nodes;

[0041] The node type is 0, 1, or 2, where 0 represents a standard input node, 1 represents a logical AND gate node and a 6-input lookup table node, and 2 represents a standard output node; the fan-in number refers to the number of inputs of the node; the number of NOT gates refers to the number of NOT gates in the node inputs; the truth table represents the logical function corresponding to the 6-input lookup table node; the reinforcement learning model understands the differences and connections between circuits from multiple perspectives by learning the statistical features and graph structures of the logic circuits and netlists, and based on these commonalities and differences, takes appropriate optimization actions for different circuits and different states of the same circuit, combines multiple optimization actions into a comprehensive sequence, and performs logical optimization on the circuit to achieve better results than commonly used heuristic scripts.

[0042] S202: Normalize the statistical features;

[0043] Preferably, the pre-training in step S3 is specifically as follows: Sequentially select the circuits in the dataset for training. After each training is completed, retain the model parameters and use them as the initialization parameters for the training of the next circuit. By gradually training on the entire dataset, finally obtain a pre-trained logic synthesis model for FPGA mapping based on reinforcement learning.

[0044] The pre-training process is as follows: ① The reinforcement learning model selects the corresponding optimization action according to the current circuit state, optimizes the circuit to obtain a reward and a new circuit state, forming {state, action, new state, reward}, which is called a trajectory (Trajectory). Loop this process until a trajectory of a specified length is obtained, and the model is updated using the trajectory of the specified length; ② Continuously repeat process ① on the same circuit until the model converges on this circuit; ③ Retain the model parameters and use them as the initialization parameters, continue to read the next circuit and repeat processes ①②③ until the pre-training is completed for all circuits.

[0045] In a second aspect, the present invention provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method.

[0046] In a third aspect, the present invention provides a machine-readable storage medium, characterized in that the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the method.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] The circuit feature extraction module of the present invention uses a GNN module to analyze the graph structure of the logic circuit and the netlist, obtaining a graph embedding. It analyzes the statistical features of the logic circuit and the netlist through an FCN module, and splices the two parts of features to obtain an enhanced representation state of the precise circuit, improving the understanding ability of the reinforcement learning model for different circuits. The reinforcement learning utilizes this enhanced state to learn the differences and connections between different circuits from the statistical features, the graph structure of AIG, and the graph structure of the netlist. By continuously training and accumulating experience, it can give an exclusive and excellent optimization sequence for different circuits when facing them, and better optimize the netlist. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flowchart of a logic synthesis method for FPGA mapping based on reinforcement learning provided by the present invention.

[0050] Figure 2 is a complete model framework diagram provided by the present invention.

[0051] Figure 3 is a diagram of the circuit feature extraction module provided by the present invention.

[0052] Figure 4 is a diagram of the reinforcement learning module provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The following drawings further describe in detail the principles and features of the logic synthesis method for FPGA mapping based on reinforcement learning proposed by the present invention. The examples given are only used to explain the present invention, rather than limiting the scope of the present invention. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for the purpose of conveniently and clearly assisting in explaining the embodiments of the present invention. In addition, the structures shown in the drawings are often part of the actual structure. In particular, the focus to be shown in each drawing is different, and sometimes different scales are used.

[0054] The present invention provides a logic synthesis method for FPGA mapping based on reinforcement learning, and its flowchart is as Figure 1 shown. First, a dataset for logic synthesis is constructed, and the test circuits are pre-divided into two groups, a training set and a test set. The dataset contains input data, and the input data is a circuit described by RTL verilog. The input data is preprocessed. A logic synthesis model for FPGA mapping based on reinforcement learning is constructed, a loss function is constructed, and the model is pre-trained on the preprocessed training set. The test set is input into the pre-trained model for netlist optimization of logic synthesis, and the model is evaluated using evaluation metrics.

[0055] Therefore, a logic synthesis method for FPGA mapping based on reinforcement learning specifically includes:

[0056] S1: Construct a logically synthesized dataset, and pre-divide the test circuits into two groups, a training set and a test set, with 8 circuits in each group. The circuits in the training set are used to train the model, and the circuits in the test set simulate new circuits to verify the effect of cross-circuit experiments; the dataset contains input data; the input data is a circuit described in RTL verilog;

[0057] S2: Construct a logical synthesis environment and preprocess the input data;

[0058] S201: Use the open-source logic synthesis tool ABC to construct a logical synthesis environment, and convert the circuit described in verilog into an AIG-format circuit; read the circuit described in AIG in the logic synthesis tool ABC to obtain the statistical characteristics and graph structure of the AIG circuit;

[0059] In the present invention, the logical circuit and the AIG circuit refer to the same;

[0060] The statistical characteristics of the logical circuit refer to the number of logical AND gate nodes, the maximum number of layers, the average number of layers, the minimum number of layers, the average value and variance of the fan-outs of all inputs, and the average value and variance of the fan-outs of all logical AND gate nodes in the current logical circuit;

[0061] Construct a directed acyclic graph according to the AIG-format circuit As the graph structure of the AIG circuit; where is a set of nodes, and each node has corresponding node attributes, including node type, number of fan-ins, and number of NOT gates; is a set of edges, where each edge represents a directed connection between two corresponding nodes;

[0062] The node type is 0, 1, or 2, where 0 represents a standard input node, 1 represents a logical AND gate node, and 2 represents a standard output node; the number of fan-ins refers to the number of inputs of the node; the number of NOT gates refers to the number of NOT gates in the inputs of the node;

[0063] Use the if-a-K 6 command related to FPGA mapping in the logic synthesis tool ABC to map the logical circuit AIG into a netlist containing 6-input lookup tables, and obtain the statistical characteristics and graph structure of the netlist;

[0064] The statistical characteristics of the netlist refer to the number of 6-input lookup tables, the number of layers, and the number of edges in the current netlist;

[0065] Construct a directed acyclic graph according to the circuit in netlist format As the graph structure of the netlist; where is a set of nodes, each node having corresponding node attributes, including node type, number of fan - ins, and truth table, is a set of edges, where each edge represents a directed connection between two corresponding nodes;

[0066] The node type is 0, 1, or 2, where 0 represents a standard input node, 1 represents a 6 - input lookup table node, and 2 represents a standard output node; the number of fan - ins refers to the number of inputs of the node; the truth table represents the logical function corresponding to the 6 - input lookup table node;

[0067] S202: Normalize the statistical features. Specifically, it means compressing the statistical features within the interval [-1, 1] to ensure the stable learning of the subsequent model;

[0068] S3: Build a logic synthesis model for FPGA mapping based on reinforcement learning and pre - train the model on the pre - processed training set;

[0069] The logic synthesis model for FPGA mapping based on reinforcement learning is as Figure 2 , including a circuit feature extraction module and a reinforcement learning module;

[0070] S301: Build the action space of reinforcement learning; the action space is seven optimization algorithms in the heuristic script resyn2 commonly used in the synthesis tool ABC, including balance, rewrite, refactor, resub, rewrite - z, refactor - z, resub - z; mark each action as 0 - 6, representing these seven optimization algorithms respectively;

[0071] S302: Build the reward function of reinforcement learning: Encourage the reinforcement learning model to minimize the number of 6 - input lookup tables after FPGA mapping as much as possible. At the same time, standardize the rewards to ensure stable training between different circuits. The design is as follows:

[0072]

[0073] where represents the number of 6 - input lookup tables in the netlist at the th step;

[0074] S303: Build a circuit feature extraction module to learn the features of the circuit, including an FCN module and a GNN module;

[0075] The input of the FCN module is the statistical features and synthesis sequence features of the circuit, as Figure 3As shown in the figure, it specifically includes a first linear layer, a first ReLU layer, a second linear layer, a second ReLU layer, and a third linear layer. Among them, the activation function ReLU is used after the first two linear layers to increase the non-linearity of the network.

[0076] The comprehensive sequence feature consists of the numbers of the 5 most recently selected actions (optimization algorithms) and the current comprehensive sequence length (the total number of executed actions). Among them, the action numbers are initially default set to all 7s, and the current comprehensive sequence length is set to 0, indicating that no optimization actions have been selected yet. The comprehensive sequence feature will change continuously during the interaction between the model and the circuit to help the model distinguish different circuit states. Specifically, the initial value of the comprehensive sequence feature is 777770. After selecting action 2, it becomes 277771, and after selecting action 5, it becomes 257772, and so on.

[0077] The input of the GNN module is the structure of the logic circuit and the structure of the netlist, which specifically includes a first GAT layer, a first ReLU layer, a second GAT layer, a second ReLU layer, a third GAT layer, and a pooling layer.

[0078] The input of the GNN module is the logic circuit structure and the netlist structure. Through three GAT layers, each node can learn node embeddings from its neighbor nodes. There is a ReLU activation function after the first two GAT layers to increase non-linear features. After the last GAT layer, there is a pooling layer to reduce the scale of the graph and generate a higher-level feature representation.

[0079] The first GAT layer, the second GAT layer, and the third GAT layer all use the attention mechanism to learn the weight distribution of different neighbors, aggregate the neighbor node features, and update their own node features. The differences in the structures of the three GAT layers lie in the input and output dimensions. At the same time, the first GAT layer uses a 4-head attention mechanism, allowing the model to learn node relationships from different angles, capture richer information, and improve the expression ability of the graph structure. The mathematical description of the aggregation is as follows:

[0080]

[0081] Among them, is a non-linear activation function, is the attention coefficient, is the feature transformation matrix, is the node 's feature, is the set of neighbor nodes of node i;

[0082]

[0083] Among them is the weight vector of the attention mechanism, concat represents the concatenation operation, and LeakyReLU is the activation function. is the node feature, is the index variable, is the node feature; the input of the first GAT layer uses the node features, the input of the second GAT layer uses the result output by the first ReLU layer, and the third GAT layer uses the result output by the second ReLU layer.

[0084] Among them, the pooling layer consists of max-pooling and average-pooling, which is used to map the graph to a graph with a unified number of nodes. The mathematical description is as follows:

[0085]

[0086] Among them represents the graph embedding, represents the node embedding.

[0087] S304: Build a reinforcement learning module; Figure 4 is the reinforcement learning module, which adopts the Actor-Critic framework and is modeled with the PPO algorithm.

[0088] The reinforcement learning module concatenates the statistical features and structure of the circuit to obtain the state of reinforcement learning, and selects a suitable optimization algorithm for the circuit according to the state; it includes the state, Actor (policy) network, and Critic (value) network;

[0089] Among them, the state is obtained by concatenating vectors obtained by the FCN module and the GNN module;

[0090] The Actor network uses a neural network to process the state, analyzes the current circuit state, and obtains a probability vector of the corresponding action, and selects the corresponding circuit optimization action according to this vector and this optimization action is an optimization algorithm of the logic synthesis tool ABC; using this action to optimize the circuit in ABC, the difference before and after the circuit is obtained as the reward of this action ; specifically includes the first linear layer, the first ReLU layer, and the second linear layer;

[0091] The Critic network is used to analyze the current circuit state and obtain the evaluation of the state and provides an advantage function for evaluating the pros and cons of the currently selected action. Specifically, it includes the first linear layer, the first ReLU layer, and the second linear layer. The reinforcement learning model updates according to the pros and cons of this action. The advantage function refers to the Generalized Advantage Estimation (GAE), and the mathematical description is as follows:

[0092]

[0093] wherein represents the discount factor, representing the evaluation of the action by the Critic network;

[0094] The evaluation of this action is used to update the network, and the mathematical representation is as follows:

[0095]

[0096] wherein, represents the probability difference between the decisions of the new and old models , is the clipping rate of the clipping function, clips to be between to control the update amplitude of the network to ensure smooth update. This formula is used to update the entire logic synthesis model including the feature extraction module and the reinforcement learning module.

[0097] The reinforcement learning module is used to splice the statistical features and structure of the circuit to obtain the state of reinforcement learning, and select a suitable optimization algorithm for the circuit according to the state;

[0098] S305: Iterate the above process, collect trajectories of length and update the entire network; continuously repeat this process until the model converges on this circuit or reaches the maximum number of iterations , then the training of the model on this circuit is completed;

[0099] S306: Train one circuit on the test set each time. After the training is completed, inherit the model parameters and continue to train a new circuit. Repeat the above process until all the circuits in the training set are trained, and a pre-trained model is obtained.

[0100] S4: Use the pre-trained logic synthesis model for FPGA mapping based on reinforcement learning to directly perform logic optimization on the new circuit without retraining.

[0101] The experimental environment of the logic synthesis method for FPGA mapping based on reinforcement learning proposed by the present invention is shown in Table 1 below:

[0102] Table 1

[0103] Category Configuration Operating System Linux CPU Intel(R) Xeon(R) Gold 6348 CPU @ 2.6GHz GPU NVIDIA GeForce RTX 4090 RAM 500GB Pytorch 2.2.1 Pytorch Geometric 2.5.3

[0104] The optimization effects of the trained model and the heuristic script resyn2 method on 8 circuits in the test set are shown in Table 2 below:

[0105] Table 2

[0106] Test Circuit resyn2 Method in This Paper adder 0 2.01% square 1.98% 2.75% sqrt 36.48% 43.72% voter 33.31% 31.41% int2float 2.13% 14.89% priority 16.29% 46.21% mem 1.23% 11.48% div 65.79% 77.92% MEAN 19.65% 28.80%

[0107] The present invention provides an electronic device, comprising a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method.

[0108] The memory configuration involved in the present invention may include fast random access memory (RAM) and non-volatile storage media, such as disk storage devices like hard disks. The system establishes a communication connection with at least one external network element through at least one communication interface (which may be wired or wireless), and supports data exchange in various network environments such as the Internet, wide area network, local area network, or metropolitan area network.

[0109] The bus can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0110] Among them, the memory is used to store program codes. When the processor receives instructions to execute these programs, it will run the programs stored in the memory. The method flow defined in any embodiment described in the present invention can be integrated into the operation of the processor or implemented by the processor. In short, after receiving the execution command, the processor will execute the program in the memory to implement the method disclosed in the present invention.

[0111] The processor can be an integrated circuit chip with signal processing functions. When implementing the method of the present invention, each step can be realized by the hardware logic circuit inside the processor or software instructions. The processor may be a general-purpose processor, such as a central processing unit (CPU), a network processor (NP), etc., or may be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate and transistor logic devices, discrete hardware components. These processors can implement or execute various methods, steps, and logic flows disclosed in the present invention.

[0112] The general-purpose processor can be a microprocessor or any standard processor. The method steps of the present invention can be directly executed by the hardware decoding processor or executed by a combination of hardware and software modules in the decoding processor. The software module can be stored in mature storage media such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), or electrically erasable programmable memory (EEPROM). These storage media are located in the memory, and the processor reads information from them and combines its hardware to complete the steps of the above method.

[0113] An embodiment of the present invention provides a computer program product stored in a readable storage medium and including a series of program codes. The instructions included in these codes can implement the processes described in the foregoing method embodiments. The specific implementation details have been described in detail in the previous method embodiments and will not be repeated here. In short, such a computer program product enables the storage medium to be used to execute the method disclosed by the present invention.

Claims

1. A circuit logic synthesis method for FPGA mapping based on reinforcement learning, characterized in that The method includes the following steps: Construct a dataset for logic synthesis; the dataset contains circuits described by RTL verilog; Preprocess the dataset; Construct a logic synthesis model for FPGA mapping based on reinforcement learning, and use the preprocessed dataset for pre-training; wherein, the logic synthesis model for FPGA mapping based on reinforcement learning includes a circuit feature extraction module and a reinforcement learning module; the circuit feature extraction module is used to learn the statistical features, graph structure and synthesis sequence features of the circuit; the reinforcement learning module is used to splice the learned statistical features, graph structure and synthesis sequence features of the circuit to obtain the state of reinforcement learning, and select a suitable optimization algorithm for the circuit according to the state; the synthesis sequence feature is composed of the numbers of the last 5 executed actions and the total number of executed actions; the actions are from the action space of reinforcement learning; Use the pre-trained logic synthesis model for FPGA mapping based on reinforcement learning to perform logic optimization on the circuit.

2. The method according to claim 1, wherein The preprocessing specifically is: Convert the circuit described by RTL verilog into a logic circuit in AIG format; Obtain the statistical features of the logic circuit according to the logic circuit in AIG format; the statistical features include the number of logic AND gate nodes, the maximum number of layers, the average number of layers, the minimum number of layers, the average value and variance of the fan-outs of all inputs, and the average value and variance of the fan-outs of all logic AND gate nodes; Construct a directed acyclic graph according to the logic circuit in AIG format as the graph structure of the logic circuit; where is a set of nodes, and each node has corresponding node attributes, including node type, the number of fan-ins, and the number of logic NOT gates; is a set of edges, where each edge represents a directed connection between two corresponding nodes; Map the logic circuit in AIG format through FPGA into a circuit in netlist format containing 6-input lookup tables, and obtain the statistical features of the netlist; the statistical features of the netlist refer to the number, number of layers and number of edges of 6-input lookup tables in the current netlist; Construct a directed acyclic graph according to the circuit in netlist format A graph structure as a netlist; where is a set of nodes, each node having corresponding node attributes, including node type, number of fan-ins, and truth table; is a set of edges, where each edge represents a directed connection between two corresponding nodes; Normalize the statistical features of the logic circuit and the netlist.

3. The method according to claim 2, wherein The node type is 0, 1 or 2, where 0 represents a standard input node, 1 represents a logic AND gate node, and 2 represents a standard output node; the fan-in number refers to the number of inputs of the node; the number of logic NOT gates refers to the number of NOT gates in the inputs of the node.

4. The method according to claim 2, wherein The reinforcement learning specifically is: Construct an action space for reinforcement learning, and the action space is composed of multiple circuit optimization algorithms; Construct a reward function for reinforcement learning, with the goal of reducing the number of 6-input lookup tables after FPGA mapping, and standardize the reward; Construct a circuit feature extraction module of the logic synthesis model for FPGA mapping based on reinforcement learning to learn the statistical features, graph structure and synthesis sequence features of the circuit; Construct a reinforcement learning module of the logic synthesis model for FPGA mapping based on reinforcement learning, splice the learned statistical features, graph structure and synthesis sequence features of the circuit to obtain the state of reinforcement learning, and select a suitable optimization algorithm for the circuit according to the state; Use the preprocessed dataset to pre-train the logic synthesis model for FPGA mapping based on reinforcement learning, and during the pre-training process, use the reward corresponding to each selected action to guide the model.

5. The method according to claim 4, wherein The circuit feature extraction module includes a fully connected network module and a graph neural network module; The input of the fully connected network module is the statistical features and comprehensive sequence features of the circuit, specifically including a first linear layer, a first ReLU layer, a second linear layer, a second ReLU layer, and a third linear layer; The input of the graph neural network module is the graph structure, specifically including a first graph attention layer, a first ReLU layer, a second graph attention layer, a second ReLU layer, a third graph attention layer, and a pooling layer.

6. The method according to claim 5, wherein The first graph attention layer, the second graph attention layer, and the third graph attention layer all learn the weight distribution of different neighbors through the attention mechanism, aggregate the neighbor node features, and update their own node features; the mathematical description of the aggregation is as follows: ; Among them, is a non-linear activation function, is an attention coefficient, is a feature transformation matrix, is a node feature, is the set of neighbor nodes of node i; ; Among them is the weight vector of the attention mechanism, concat represents the concatenation operation, and LeakyReLU is the activation function, is the node 's feature, is the index variable, is the node 's feature; The pooling layer performs max-pooling and average-pooling operations on the processed graph structure and performs a concatenation operation.

7. The method according to claim 6, characterized in that, The reinforcement learning module includes a state, a policy network, and a value network; among them, The state is obtained by concatenating the vectors obtained from the fully connected network module and the graph neural network module; The policy network is used to analyze the current circuit state and obtain a probability vector corresponding to the action, and select the corresponding optimized action according to the probability vector , obtain the difference before and after the circuit, and use it as the reward ; specifically, it includes a first linear layer, a first ReLU layer, and a second linear layer; The value network is used to analyze the current circuit state and obtain an evaluation of the state , and provide an advantage function for evaluating the currently selected action Advantages and disadvantages. The model is updated according to the advantages and disadvantages of the action 8. The method according to claim 4, characterized in that, The pre-training is specifically as follows: Circuits in the dataset are sequentially selected for training. After each training is completed, the model parameters are retained and used as the initialization parameters for the training of the next circuit. By gradually training on the entire dataset, a pre-trained logic synthesis model for FPGA mapping based on reinforcement learning is finally obtained.

9. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1-8.

10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • FPGA automatic parameter adjustment optimization method and system based on machine learning

    CN111241778A

  • Logic optimization command sequence combination method based on game reinforcement learning

    CN119514440A