An FPGA Placement Method Based on Simulated Annealing Algorithm
By optimizing the set of instances to be optimized in the simulated annealing algorithm, selecting component instances with poor performance indicators for layout optimization, the problem of low FPGA layout efficiency in the prior art is solved, and fast convergence and efficient layout results are achieved.
Patent Information
- Application Number
- CN202210736841.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-27
AI Technical Summary
The existing simulated annealing algorithm has a long execution time and low layout efficiency during the FPGA layout optimization process, and requires a large number of random exchanges and mobile component instances to achieve global optimal layout results.
By initializing the set of instances to be optimized in the simulated annealing algorithm, selecting component instances with poor performance indicators for layout optimization, and determining performance indicators based on the delay parameters and regional congestion of the component instances, and iteratively updates the set of instances to be optimized until the layout goal is reached.
The layout efficiency of the simulated annealing algorithm is improved, and the layout results that quickly converge to meet the optimization goals are met, the number of random exchanges is reduced, and the selection effectiveness and layout efficiency of component instances are improved.
Smart Images

Figure CN115017853B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of FPGA, and in particular to an FPGA placement method based on the simulated annealing algorithm. Background Art
[0002] A Field-Programmable Gate Array (FPGA) is a chip widely used in household appliances, large machinery, and even aerospace. With the expansion of the scale of FPGA chips, the placement of the chips becomes increasingly crucial and important, directly affecting the performance of the chips such as area and frequency. Therefore, when placing the chips, it is necessary to comprehensively consider various costs. Under the condition of meeting multiple constraints, how to optimize the chip placement to ensure the performance and routability of the chips has become the key to ensuring the chip quality.
[0003] The simulated annealing algorithm is a commonly used algorithm in the process of optimizing the placement of FPGAs at present. The simulated annealing algorithm is a heuristic iterative search algorithm, which was proposed by Metropolis in 1953, and its design principle comes from the real physical annealing process.
[0004] The process of using the simulated annealing algorithm to solve the FPGA placement problem can be referred to Figure 1The flowchart shown first randomly forms a global initial layout result on the FPGA. The global initial layout result includes the initial layout positions of all component instances, and initializes an annealing initial temperature, an annealing exit temperature, and a reachable global state change space. The simulated annealing algorithm mainly includes an outer loop and an inner loop. The outer loop is used to control the annealing temperature and the state change space, and the inner loop is used to exchange and move the layout positions of component instances. The outer loop starts from the annealing initial temperature and the reachable global state change space. Enter the inner loop under the current outer loop. Randomly select two component instances as the starting point and the target point respectively within the current state change space, and exchange the layout positions of the starting point and the target point to change the layout result. Each layout result change operation will be accepted with a certain probability, and this probability is related to the current annealing temperature and the impact generated by the layout result change operation. The impact generated by the layout result change operation is generally determined by calculating the cost function. The higher the annealing temperature, the greater the acceptance probability, and the more obvious the optimization effect of this layout result change operation on the layout result and the greater the acceptance probability. This process simulates the movement of solid particles. After the inner loop ends, the annealing temperature will be lowered and the state change space will be reduced to enter the next outer loop, simulating the process of the energy of the solid gradually decreasing and the activity of the solid particles gradually decreasing during the cooling process of the solid. Exit the outer loop until the annealing temperature reaches the exit temperature. The initial temperature is generally relatively high, which can make the layout result change operation of component instances easily accepted, simulating the high-temperature state of the solid. The exit temperature is generally relatively low, so that the layout result change operation of component instances becomes difficult to accept, simulating the low-temperature state of the solid. After the outer loop exits, set the annealing temperature to 0, and then perform several exchanges of the layout positions of component instances to change their states. At this time, only the layout result change operation that optimizes the layout result can be accepted, and finally complete the simulated annealing algorithm to obtain the global optimal layout result on the FPGA.
[0005] As described above, when the simulated annealing algorithm is used to solve the FPGA layout problem, it is necessary to exchange and move the layout results of component instances through multiple inner and outer loops, and often a large number of random exchanges and moves are required to ensure convergence to the global optimal layout result. Therefore, the algorithm execution time is usually long and the layout efficiency is low. Summary of the Invention
[0006] In view of the above problems and technical requirements, the present application proposes an FPGA layout method based on the simulated annealing algorithm. The technical solution of the present application is as follows:
[0007] An FPGA layout method based on the simulated annealing algorithm, the method includes:
[0008] Perform a global random initial layout on the FPGA according to the user netlist to obtain a layout result that has not reached the layout target;
[0009] Initialize the set of instances to be optimized according to the layout result obtained from the initial layout. The set of instances to be optimized includes several instances to be optimized, and an instance to be optimized is a component instance in the user netlist whose performance index is worse than a preset threshold under the current layout result.
[0010] Based on the layout result obtained from the initial layout, use the simulated annealing algorithm to iteratively select instances to be optimized from the set of instances to be optimized to optimize the layout result, and iteratively update the set of instances to be optimized until a layout result that meets the layout goal globally for the FPGA is obtained.
[0011] A further technical solution thereof is that the performance index of a component instance in the user netlist is related to the delay parameter and / or the congestion degree of the area where the component instance is located under the current layout result. The greater the delay of the component instance under the current layout result, the higher the congestion degree of the area where it is located, and the worse the performance index of the component instance.
[0012] A further technical solution thereof is that the set of instances to be optimized includes at least two types of instances to be optimized with different layout criticalities, and the layout criticality of an instance to be optimized is related to the performance index of the instance to be optimized; when iteratively selecting instances to be optimized from the set of instances to be optimized at an annealing temperature, the selection frequency of each instance to be optimized in the set of instances to be optimized corresponds to the layout criticality of the instance to be optimized.
[0013] A further technical solution thereof is that the worse the performance index of an instance to be optimized, the higher the layout criticality, and the higher the corresponding selection frequency.
[0014] A further technical solution thereof is that the set of instances to be optimized is divided into several non - overlapping subsets according to the layout criticality. Each subset corresponds to a layout criticality and includes the instances to be optimized in the set of instances to be optimized that have the layout criticality. Moreover, the higher the layout criticality, the fewer the total number of instances to be optimized included in the corresponding subset; when iteratively selecting instances to be optimized from the set of instances to be optimized at an annealing temperature, randomly select instances to be optimized from the corresponding subset according to the selection probability corresponding to each layout criticality.
[0015] A further technical solution thereof is that the selection probability of the subset with the lowest corresponding layout criticality and the largest total number of instances to be optimized included is not lower than the probability threshold.
[0016] A further technical solution thereof is that the method of using the simulated annealing algorithm to optimize the layout result and iteratively update the set of instances to be optimized includes:
[0017] The annealing temperature starts from the initial temperature. At the current annealing temperature, iteratively select instances to be optimized from the current set of instances to be optimized to optimize the layout result, and obtain the optimized layout result at the current annealing temperature.
[0018] Update the set of instances to be optimized with the optimized layout result, reduce the annealing temperature and the swap radius, and execute again the step of using the simulated annealing algorithm to select an instance to be optimized from the set of instances to be optimized to optimize the layout result until the annealing temperature reaches the exit temperature.
[0019] A further technical solution thereof is to iteratively select an instance to be optimized from the current set of instances to be optimized at the current annealing temperature to optimize the layout result, including:
[0020] Select an instance to be optimized from the current set of instances to be optimized, and optimize the layout position of the selected instance to be optimized within the swap radius according to the current annealing temperature and the acceptance rate corresponding to the preset cost function.
[0021] If the iteration termination condition is not reached, then re-execute the step of selecting an instance to be optimized from the current set of instances to be optimized.
[0022] If the iteration termination condition is reached, then obtain the optimized layout result at the current annealing temperature.
[0023] The beneficial technical effects of the present application are:
[0024] The present application discloses an FPGA layout method based on a simulated annealing algorithm. This method optimizes the process of using the conventional simulated annealing algorithm to solve the FPGA layout problem. In the algorithm loop exchange, instead of randomly exchanging the layout positions of component instances, it finds the key instances to be optimized with poor performance indicators, and more selects these instances to be optimized for layout optimization, improving the effectiveness of selecting component instances, enabling the simulated annealing process to converge more quickly and obtain a layout result that meets the optimization goal, and improving the layout efficiency.
[0025] When selecting these instances to be optimized for layout optimization, the criticality between these instances to be optimized can be further distinguished, so as to select the instances to be optimized with different performance indicators according to different selection frequencies, thereby more selecting the instances to be optimized with worse performance indicators for layout optimization, and further improving the effectiveness of selecting component instances and the layout efficiency.
[0026] When selecting the key instances to be optimized with poor performance indicators, this method considers the delay parameter of the component instance in the current layout result and / or the congestion degree of the area where it is located, so that different performance indicators can be focused on for optimization according to actual needs, with high flexibility and can meet different layout goals. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of the existing simulated annealing algorithm for FPGA layout optimization.
[0028] Figure 2 It is a flowchart of a method for implementing FPGA placement based on the simulated annealing algorithm provided by an embodiment of the present application.
[0029] Figure 3 It is a flowchart of a method for implementing FPGA placement based on the simulated annealing algorithm provided by another embodiment of the present application. Detailed implementation manners
[0030] The following further describes the detailed implementation manners of the present application with reference to the accompanying drawings.
[0031] The present application discloses an FPGA placement method based on the simulated annealing algorithm. The method includes the following steps. Please refer to Figure 2 :
[0032] Step 210: Perform a global random initial placement on the FPGA according to the user netlist to obtain a placement result that does not reach the placement target. The method of random initial placement can be the same as the random initial placement in the initialization process of the conventional simulated annealing algorithm, and the present application does not elaborate on this.
[0033] Step 220: Initialize the set of instances to be optimized according to the placement result obtained from the initial placement. The set of instances to be optimized includes several instances to be optimized. The instance to be optimized is a component instance in the user netlist whose performance index is worse than the preset threshold under the current placement result.
[0034] Step 230: On the basis of the placement result obtained from the initial placement, use the simulated annealing algorithm to iteratively select instances to be optimized from the set of instances to be optimized to optimize the placement result, and iteratively update the set of instances to be optimized until a placement result that reaches the placement target globally for the FPGA is obtained.
[0035] As recorded in the background art section, when the conventional simulated annealing algorithm is executed, two component instances are randomly selected from the placement result obtained from the random initial placement to exchange their placement positions to change the placement result. To solve the problem of low placement efficiency caused by a large number of random exchanges, the present application optimizes the process of the simulated annealing algorithm. Instead of randomly selecting component instances to exchange, it finds the component instances with poor performance indexes. These component instances are the key instances that need to be optimized during the placement process. By more frequently selecting these instances to be optimized to exchange the placement positions during the execution of the simulated annealing algorithm, the effectiveness of the optimization is improved, enabling the simulated annealing algorithm to converge more quickly and obtain a placement result that meets the placement target.
[0036] In the above step 220, a set of instances to be optimized under the current layout result is determined according to the performance metrics of each component instance under the current layout result. The performance metrics of the component instances in the user netlist are related to the delay parameters of the component instances under the current layout result and / or the congestion degree of the regions where they are located. The greater the delay of the component instance under the current layout result, the higher the congestion degree of the region where it is located, and the worse the performance metrics of the component instance. In one embodiment, the regions where the component instances are located are divided according to the domains of the FPGA architecture. Or in another embodiment, the regions where the component instances are located are divided according to custom rules, and the included region ranges can be set customarily. The congestion degree of the regions where the component instances are located can be calculated and determined using existing congestion estimation algorithms.
[0037] When determining the performance metrics based on the delay parameters of the component instances under the current layout result and / or the congestion degree of the regions where they are located, there are multiple different implementation methods:
[0038] In one embodiment, the performance metrics are determined only according to the delay parameters of the component instances under the current layout result, and the corresponding relationship between different delay parameters and performance metrics is pre-configured.
[0039] In another embodiment, the performance metrics are determined only according to the congestion degree of the regions where the component instances are located under the current layout result, and the corresponding relationship between different congestion degrees and performance metrics is pre-configured.
[0040] In another embodiment, the performance metrics are determined according to both the delay parameters of the component instances under the current layout result and the congestion degree of the regions where they are located. Then, the relationship between different delay parameters and delay metric coefficients is pre-configured, and the corresponding relationship between the congestion degrees of different regions where they are located and congestion metric coefficients is pre-configured. Then, the performance metrics of the component instances can be determined by performing weighted calculation on the delay metric coefficients and congestion metric coefficients of the component instances. When performing weighted calculation, the weights of the delay metric coefficients and congestion metric coefficients can be equal, or the weights of the two can be configured to be unequal according to actual needs.
[0041] In another embodiment, considering that congestion does not occur in all test cases, especially in some test cases with less resource usage, when the congestion degree of the region where the component instance is located does not reach the congestion threshold, the performance metrics are determined only according to the delay parameters of the component instance under the current layout result. When the congestion degree of the region where the component instance is located reaches the congestion threshold, the performance metrics are determined according to both the delay parameters of the component instance under the current layout result and the congestion degree of the region where it is located.
[0042] In the above step 230, the method of using the simulated annealing algorithm to optimize the layout result and iteratively update the set of instances to be optimized is as follows. Please refer to Figure 3 the flowchart shown:
[0043] Step 310: Initialize the system parameters of the simulated annealing algorithm, mainly including the initial temperature, the exit temperature, and the state change space. The state change space can generally be defined by the swap radius, and the initial swap radius can generally reach the entire FPGA chip. The annealing temperature of the simulated annealing algorithm starts from the initial temperature. In actual applications, there is no specific order between the step of initializing the system parameters in this step 310 and steps 210 and 220.
[0044] Step 320: If the current annealing temperature has not reached the exit temperature, then at the current annealing temperature, iteratively select an instance to be optimized from the current set of instances to be optimized to optimize the layout result, and obtain the optimized layout result at the current annealing temperature. At the initial temperature, the current set of instances to be optimized is the set of instances to be optimized initialized based on the layout result obtained from the initial layout.
[0045] The process of iteratively selecting an instance to be optimized from the current set of instances to be optimized at the current annealing temperature to perform an inner loop on the layout result includes the following steps:
[0046] Step 321: Select an instance to be optimized from the current set of instances to be optimized.
[0047] In one embodiment, each time a random instance to be optimized is selected from the set of instances to be optimized, that is, the selection frequencies of all instances to be optimized in the set of instances to be optimized are the same.
[0048] In another embodiment, in order to further improve the effectiveness of the selected instance to be optimized and accelerate convergence, the set of instances to be optimized includes at least two types of instances to be optimized with different layout criticalities, and the layout criticality of the instance to be optimized is related to the performance index of the instance to be optimized. Each instance to be optimized has and only has one corresponding layout criticality, and each layout criticality has at least one instance to be optimized. For example, the instances to be optimized configured within the same performance index range all belong to the same layout criticality, and the performance index ranges corresponding to different layout criticalities do not overlap. Then, when iteratively selecting an instance to be optimized from the set of instances to be optimized at an annealing temperature, the selection frequency of each instance to be optimized in the set of instances to be optimized corresponds to the layout criticality of the instance to be optimized. That is, in this embodiment, the instances to be optimized are not completely randomly selected from the set of instances to be optimized, but different instances to be optimized with different layout criticalities are selected according to different selection frequencies.
[0049] In order to speed up optimization and algorithm convergence, the worse the performance index of the instance to be optimized, the higher the layout criticality, and the higher the corresponding selection frequency. That is, the instances to be optimized with poor performance indicators are selected more frequently to exchange layout positions, so as to focus on optimizing these instances to be optimized with poor performance indicators to achieve the desired optimization purpose.
[0050] In order to realize the function of selecting instances to be optimized with different layout criticalities according to different selection frequencies, this embodiment provides an implementation method: the set of instances to be optimized is divided into several subsets according to the layout criticality, any two subsets do not intersect each other, that is, the intersection is empty, and the union of all subsets is the set of instances to be optimized. Each subset corresponds to a layout criticality and contains instances to be optimized with layout criticality in the set of instances to be optimized, and the higher the layout criticality, the fewer the total number of instances to be optimized included in the corresponding subset. For example, in one example, the set of instances to be optimized, which includes a total of N instances to be optimized, is divided into subset 1, subset 2, and subset 3. Subset 1 contains 10%*N instances to be optimized, and these instances to be optimized in subset 1 all correspond to layout criticality A. Subset 2 contains 20%*N instances to be optimized, and these instances to be optimized in subset 2 all correspond to layout criticality B. Subset 3 contains 70%*N instances to be optimized, and these instances to be optimized in subset 3 all correspond to layout criticality C. The layout criticality of layout criticality A, layout criticality B, and layout criticality C decreases in sequence.
[0051] When iteratively selecting instances to be optimized from the set of instances to be optimized at an annealing temperature, instances to be optimized are randomly selected from the corresponding subsets according to the selection probability corresponding to each layout criticality. The selection frequency of each instance to be optimized is affected by the number of instances to be optimized contained in the subset to which it belongs and the selection probability of the subset. For example, when the number of instances to be optimized contained in two subsets is equal, if the selection probabilities of the two subsets are not equal, the selection frequencies of the instances to be optimized in the two subsets are also different, and the selection frequency of the instances to be optimized in the subset with a larger selection probability is relatively higher. For another example, when the selection probabilities of the two subsets are equal, if the number of instances to be optimized contained in the two subsets is not equal, the selection frequencies of the instances to be optimized in the two subsets are also different, and the selection frequency of the instances to be optimized in the subset containing fewer instances to be optimized is relatively higher. The selection probabilities corresponding to various layout criticalities can be customized and adjusted, and the specific setting values are not limited. It only needs to ensure that the higher the layout criticality of the instance to be optimized, the higher the corresponding selection frequency.
[0052] For example, in the above example, assume that the selection probability corresponding to layout criticality A is 40%, the selection probability corresponding to layout criticality B is 20%, and the selection probability corresponding to layout criticality C is 40%. Then, when iteratively selecting instances to be optimized from the set of instances to be optimized at an annealing temperature, randomly select instances to be optimized from subset 1 with a selection probability of 40%, randomly select instances to be optimized from subset 2 with a selection probability of 20%, and randomly select instances to be optimized from subset 3 with a selection probability of 40%. Although the selection probabilities of layout criticality A and layout criticality C are equal, since the number in subset 3 is more than that in subset 1, the selection frequency of the instances to be optimized in subset 1 is higher than that of the instances to be optimized in subset 3. Similarly, although the selection probability of layout criticality C is greater than that of layout criticality B, since the number in subset 3 is more than that in subset 2, the selection frequency of the instances to be optimized in subset 2 is still higher than that of the instances to be optimized in subset 3.
[0053] Through the above method, it is possible to more conveniently iteratively select instances to be optimized from the set of instances to be optimized according to the rule that the higher the layout criticality, the higher the corresponding selection frequency. Through the above method, it is possible to optimize the instances to be optimized with worse performance indicators more frequently, but the selection probability of the subset with the lowest layout criticality and the largest total number of instances to be optimized included is not lower than the probability threshold, and this probability threshold is a preset value. That is, although it is necessary to try to select instances to be optimized with higher layout criticality as much as possible, it is also necessary to select a certain number of instances to be optimized with low layout criticality, so as to take into account the scale of the solution space and ensure that the algorithm finally converges to the global optimal solution.
[0054] Step 322: Optimize the layout position of the selected instance to be optimized within the current exchange radius in the current layout result according to the current annealing temperature and the acceptance rate corresponding to the preset cost function. In the first outer loop when the annealing temperature is the initial temperature, the current exchange radius is the exchange radius at initialization.
[0055] Similar to the conventional simulated annealing algorithm, it is necessary to select a starting point and a target point in the current state change space for exchanging layout positions. It is possible to only select an instance to be optimized from the set of instances to be optimized as the starting point according to the method provided in this application, while the target point is still randomly selected, or the target point can also be selected from the set of instances to be optimized according to the method provided in this application. Each layout result change operation will be accepted with a certain probability, which is similar to the conventional simulated annealing algorithm and will not be elaborated in this application.
[0056] Step 323: If the iteration termination condition is met, the optimized layout result at the current annealing temperature is obtained. If the iteration termination condition is not met, the step of selecting an instance to be optimized from the current set of instances to be optimized is re-executed. This iteration termination condition can be customized. Generally, it can be limited by the number of inner loop iterations. The number of inner loop iterations when the iteration termination condition is met can be set during the initialization of system parameters. The iteration termination conditions at different annealing temperatures can be the same or different, and can actually be customized according to needs.
[0057] Step 330: Update the set of instances to be optimized using the optimized layout result at the current annealing temperature, and decrease the annealing temperature and the swap radius. That is, re-determine the set of instances to be optimized according to the optimized layout result, which is similar to the method of determining the set of instances to be optimized during initialization.
[0058] Step 340: Decrease the annealing temperature and reduce the swap radius. If the current annealing temperature after the decrease still has not reached the exit temperature, re-execute Steps 320 and 330, and optimize the layout result and update the set of instances to be optimized at the current annealing temperature after the decrease. If the current annealing temperature after the decrease reaches the exit temperature, end the loop of the simulated annealing algorithm, and finally obtain the layout result of the FPGA global reaching the layout target. After ending the loop of the simulated annealing algorithm, additional swapping operations of the layout positions of component instances can also be performed as described in the background art part. This application does not elaborate and limit the process after exiting the loop of the simulated annealing algorithm.
Claims
1. An FPGA placement method based on the simulated annealing algorithm, characterized in that, The method includes: Performing a global random initial placement on the FPGA according to the user netlist to obtain a placement result that does not meet the placement target; Initializing a set of instances to be optimized based on the placement result obtained from the initial placement. The set of instances to be optimized includes several instances to be optimized, and an instance to be optimized is a component instance in the user netlist whose performance index is worse than a preset threshold under the current placement result; On the basis of the placement result obtained from the initial placement, using the simulated annealing algorithm to iteratively select an instance to be optimized from the set of instances to be optimized to optimize the placement result, and iteratively updating the set of instances to be optimized until a placement result that meets the placement target globally for the FPGA is obtained; The method of using the simulated annealing algorithm to optimize the placement result and iteratively update the set of instances to be optimized includes: starting from the initial temperature, selecting an instance to be optimized from the current set of instances to be optimized, and optimizing the placement position of the selected instance to be optimized within the exchange radius according to the current annealing temperature and the acceptance rate corresponding to the preset cost function; if the iteration termination condition is not reached, then re-execute the step of selecting an instance to be optimized from the current set of instances to be optimized; if the iteration termination condition is reached, then obtain the optimized placement result at the current annealing temperature; updating the set of instances to be optimized using the optimized placement result, and reducing the annealing temperature and the exchange radius, and then re-execute the step of using the simulated annealing algorithm to select an instance to be optimized from the set of instances to be optimized to optimize the placement result until the annealing temperature reaches the exit temperature; The set of instances to be optimized includes at least two types of instances to be optimized with different placement criticalities, and the placement criticality of an instance to be optimized is related to the performance index of the instance to be optimized; when iteratively selecting an instance to be optimized from the set of instances to be optimized at an annealing temperature, the selection frequency of each instance to be optimized in the set of instances to be optimized corresponds to the placement criticality of the instance to be optimized, and the worse the performance index of the instance to be optimized and the higher the placement criticality, the higher the corresponding selection frequency.
2. The method according to claim 1, wherein The performance index of a component instance in the user netlist is related to the delay parameter of the component instance under the current placement result and / or the congestion degree of the area where it is located. The greater the delay of the component instance under the current placement result and the higher the congestion degree of the area where it is located, the worse the performance index of the component instance.
3. The method according to claim 1, wherein The set of instances to be optimized is divided into several non-overlapping subsets according to the placement criticality. Each subset corresponds to a placement criticality and includes the instances to be optimized in the set of instances to be optimized that have the placement criticality. Moreover, the higher the placement criticality, the fewer the total number of instances to be optimized included in the corresponding subset; when iteratively selecting an instance to be optimized from the set of instances to be optimized at an annealing temperature, randomly select an instance to be optimized from the corresponding subset according to the selection probability corresponding to each placement criticality.
4. The method according to claim 3, wherein The selection probability of the subset with the lowest corresponding placement criticality and the largest total number of instances to be optimized included is not less than the probability threshold.
Citation Information
Patent Citations
Field programmable gate array (FPGA) layout method for realizing layout legalization by utilizing netlist local re-integration
CN113408224A