Memory compiler IP time sequence convergence and power consumption collaborative optimization method and system

Through fine-grained power consumption control and intelligent optimization strategies, the problems of insufficient timing convergence and power consumption optimization in memory compiler IP design are solved, the coordinated optimization of timing and power consumption is achieved, and the design efficiency and stability are improved.

CN120597833AActive Publication Date: 2025-09-05SUZHOU MICROELECTRONICS IND TECH RES INST OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511092646.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-05
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Insufficient consideration of timing closure and power consumption optimization in existing memory compiler IP designs leads to low design efficiency, difficulty in achieving the optimal balance between timing and power consumption, and a lack of intelligent optimization strategies and automated adjustment methods.

Method used

A fine-grained power consumption control strategy is used to divide the timing path into multiple power consumption control domains, set different operating voltages, and optimize them through deep neural network models and decision tree models. Combined with path priority and functional partitioning methods, a layout and routing solution that meets timing convergence requirements and optimizes power consumption is generated.

Benefits of technology

The co-optimization of timing closure and power consumption of the memory compiler IP is achieved, which improves design efficiency, reduces the number of design iterations, significantly improves the timing margin and stability of the critical path, and reduces overall power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597833A_ABST
    Figure CN120597833A_ABST
Patent Text Reader

Abstract

The invention provides a memory compiler IP time sequence convergence and power consumption collaborative optimization method and system, and relates to the technical field of compilers, and the method comprises the steps: obtaining circuit netlist information, extracting a time sequence path, carrying out wiring topology analysis, dividing a power consumption control domain through employing a fine-grained power consumption control strategy, and constructing a deep neural network model to predict delay. And training the decision tree model to obtain an optimization strategy, calculating a path priority to carry out function partitioning, and finally generating a power consumption optimized layout wiring scheme which meets a time sequence convergence requirement. According to the invention, collaborative optimization of time sequence convergence and power consumption of the memory compiler IP is realized, and the performance of an integrated circuit is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to compiler technology, and in particular to a method and system for co-optimizing timing closure and power consumption of a memory compiler IP. Background Art

[0002] As the complexity of integrated circuit design continues to increase, memory compiler IP, a key component in chip design, has become increasingly important for optimizing performance, power consumption, and area. Memory compiler IP can automatically generate memory cells of varying specifications based on user requirements and is widely used in various chip designs. At nanometer-scale process nodes, timing closure and power optimization for memory compiler IP have become key design challenges. Traditional memory compiler IP design methods primarily focus on functional implementation and area optimization, while insufficiently considering the coordinated optimization of timing closure and power consumption, resulting in inefficient design and performance bottlenecks.

[0003] Existing technologies typically use coarse-grained global voltage adjustments or simple power gating techniques to achieve timing closure and power optimization in memory compiler IP, lacking fine-grained power control strategies for specific timing paths. This approach fails to precisely allocate voltage based on the characteristics of different timing paths, often resulting in unnecessary power consumption while meeting timing requirements, making it difficult to achieve an optimal balance between timing and power consumption.

[0004] Existing memory compiler IP optimization methods mostly rely on designer experience and iterative manual adjustments, lacking intelligent predictive models and automated optimization strategies. This approach is not only time-consuming and labor-intensive, but also struggles to address the multivariable optimization challenges inherent in complex designs. This results in limited optimization effectiveness and makes it difficult to adapt to varying process nodes and design specifications.

[0005] In addition, the existing technology usually uses a general layout and routing algorithm to optimize the layout and routing of memory compiler IP, which fails to fully consider the unique structural characteristics and timing path priorities of the memory. As a result, the generated layout and routing scheme has deficiencies in timing convergence and power consumption control, making it difficult to meet the design requirements of high performance and low power consumption. Summary of the Invention

[0006] The embodiments of the present invention provide a method and system for co-optimizing timing closure and power consumption of a memory compiler IP, which can solve the problems in the prior art.

[0007] According to a first aspect of the embodiments of the present invention,

[0008] Provides memory compiler IP timing closure and power consumption co-optimization methods, including:

[0009] Obtain circuit netlist information of the memory compiler IP, extract a timing path based on the circuit netlist information, perform wiring topology analysis on the timing path, and determine the connection length and interconnection delay between timing cells in the timing path;

[0010] Based on the connection length and the interconnection delay, a fine-grained power consumption control strategy is adopted to divide the timing path into multiple power consumption control domains, and a different operating voltage is set for each of the power consumption control domains, so that the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point;

[0011] A deep neural network model is constructed based on the voltage and delay data of the power consumption control domain, and a delay prediction value is obtained based on the deep neural network model; parameters of the deep neural network model are dynamically adjusted according to a prediction error between the delay prediction value and the actual delay value, and a reliability index is calculated; when the reliability index meets a preset reliability threshold, the delay prediction value is used for timing optimization;

[0012] Acquire characteristic parameters of the time series path and construct a characteristic vector matrix based on the characteristic parameters; train a decision tree model based on the characteristic vector matrix to obtain an optimization strategy, and iteratively update the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold;

[0013] Calculate the path priority based on the timing path, perform functional partitioning on the memory compiler IP based on the path priority and construct a partition cost function; perform cell location optimization and hierarchical routing optimization according to the partition cost function, and generate a memory compiler IP layout and routing solution that meets timing closure requirements and optimizes power consumption.

[0014] Based on the connection length and the interconnection delay, a fine-grained power consumption control strategy is adopted to divide the timing path into multiple power consumption control domains, and a different operating voltage is set for each power consumption control domain, so that the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point.

[0015] Acquiring physical parameters of the line segment in the timing path, the physical parameters including the length of the line segment, the width of the line segment, the thickness of the interlayer dielectric of the line segment, and the metal specific resistance of the line segment;

[0016] Calculating the resistance per unit length and the capacitance per unit length of the line segment using a distributed RC model based on the physical parameters of the line segment, and calculating the total resistance and total capacitance of the line segment according to the resistance per unit length and the capacitance per unit length;

[0017] Calculating the interconnection delay of the connection segment based on the total resistance and total capacitance of the connection segment to obtain the interconnection delay distribution characteristics of the timing path; and dividing adjacent sequential cells with similar delay sensitivities into a plurality of power consumption control domains using a hierarchical clustering algorithm based on the delay sensitivity of the sequential cells in the timing path, wherein the delay sensitivity represents the degree of influence of the operating voltage of the sequential cell on the total delay of the timing path;

[0018] For the multiple power consumption control domains obtained by division, under the condition of satisfying the timing constraints, the operating voltage of each power consumption control domain is determined by minimizing the dynamic power consumption of each power consumption control domain, wherein the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point.

[0019] Calculating the resistance per unit length and the capacitance per unit length of the line segment using a distributed RC model based on the physical parameters of the line segment, and calculating the total resistance and the total capacitance of the line segment according to the resistance per unit length and the capacitance per unit length includes:

[0020] The distributed RC model is used to calculate the resistance per unit length of the connecting segment. The calculation process of the resistance per unit length is as follows: dividing the metal specific resistance by the product of the width of the connecting segment and the thickness of the interlayer dielectric;

[0021] Calculating the capacitance per unit length of the line segment based on the physical parameters of the line segment, wherein the calculation process of the capacitance per unit length includes: calculating the parallel plate capacitance by dividing the product of the vacuum relative dielectric constant system and the width of the line segment by the thickness of the interlayer dielectric; calculating the edge capacitance by multiplying the vacuum relative dielectric constant system by a logarithmic function of the ratio of the interlayer dielectric thickness to the width of the line segment; calculating the coupling capacitance by dividing the product of the vacuum relative dielectric constant system and the thickness of the interlayer dielectric by the metal resistivity; and adding the parallel plate capacitance, the edge capacitance, and the coupling capacitance to obtain the capacitance per unit length;

[0022] Acquire a signal frequency parameter of the connecting segment, the signal frequency parameter including a signal frequency band and a characteristic frequency parameter, and multiply the signal frequency parameter by the resistance per unit length to obtain a frequency-corrected resistance per unit length;

[0023] The connecting segment is divided into multiple equal-length sub-segments, and the product of the frequency-corrected unit length resistance of each sub-segment and the sub-segment length is accumulated to obtain the total resistance of the connecting segment, and the product of the unit length capacitance of each sub-segment and the sub-segment length is accumulated to obtain the total capacitance of the connecting segment.

[0024] Dynamically adjusting the parameters of the deep neural network model and calculating a reliability index based on a prediction error between the delay prediction value and the actual delay value, and using the delay prediction value for timing optimization when the reliability index meets a preset reliability threshold, includes:

[0025] Constructing a comprehensive loss function, training the deep neural network model using a gradient descent method, calculating weight gradients according to the comprehensive loss function, and updating weight parameters of the neural network based on the weight gradients;

[0026] Obtaining a new operating voltage combination, inputting the operating voltage combination into the deep neural network model, and obtaining a corresponding delay prediction value;

[0027] Collecting an actual delay measurement value under the operating voltage combination, and calculating a prediction error between the delay prediction value and the actual delay measurement value;

[0028] Dynamically adjusting model parameters according to the prediction error, calculating a weight update amount, wherein the weight update amount is proportional to the product of the prediction error and the weight gradient; and calculating a bias update amount, wherein the bias update amount is proportional to the sign value of the prediction error;

[0029] Calculating a confidence level of a prediction result based on the prediction error, wherein the confidence level decreases as the prediction variance increases, and using the difference between the confidence level and the prediction error as a reliability indicator;

[0030] When the reliability index is greater than the preset reliability threshold, the current delay prediction value is used to guide timing optimization; when the reliability index is less than the preset reliability threshold, the model online learning process is triggered, new training data is added to the historical data set, and the deep neural network model is retrained.

[0031] Constructing a feature vector matrix based on the feature parameters; training a decision tree model based on the feature vector matrix to obtain an optimization strategy, and iteratively updating the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold, including:

[0032] Constructing a eigenvector matrix based on the characteristic parameters, the eigenvector matrix including a timing eigenvector, a power consumption eigenvector, and an area eigenvector;

[0033] Calculating the sensitivity of the timing feature vector, the power consumption feature vector, and the area feature vector to voltage, and constructing a sensitivity feature vector, where the sensitivity feature vector includes the sensitivity of timing to voltage, the sensitivity of power consumption to voltage, and the sensitivity of area to voltage;

[0034] Combining the timing feature vector, the power consumption feature vector, the area feature vector, and the sensitivity feature vector to form an input feature matrix;

[0035] Building a decision tree model based on the input feature matrix, and outputting a voltage adjustment amount, a size adjustment amount, and a topology adjustment solution through the decision tree model, wherein the voltage adjustment amount is obtained by mapping the sigmoid function, the size adjustment amount is obtained by mapping the tanh function, and the topology adjustment solution is obtained by mapping the softmax function;

[0036] Constructing an optimization loss function and calculating a loss function gradient based on the optimization loss function; calculating an optimization target improvement and determining a sample weight based on the optimization target improvement, wherein the sample weight increases as the optimization target improvement increases;

[0037] The parameters of the decision tree model are updated according to the sample weights and the loss function gradient, and the timing path is optimized and adjusted based on the updated decision tree model until the weighted ratio of the optimization target improvement to the optimization overhead meets a preset condition threshold.

[0038] Performing cell location optimization and hierarchical routing optimization based on the partition cost function to generate a memory compiler IP placement and routing solution that meets timing closure requirements and optimizes power consumption includes:

[0039] Functionally partitioning the memory compiler IP based on the partition cost function; constructing a unit position optimization cost function based on the result of the functional partitioning, and determining the optimization direction of the unit according to the gradient direction of the unit position optimization cost function;

[0040] Multiplying the optimization direction of the unit by an exponential decay function of a power consumption change to obtain a power consumption-aware unit movement amount, where the power consumption change amount is determined by a difference in power consumption before and after the movement, and adjusting the unit position based on the power consumption-aware unit movement amount;

[0041] Constructing an inter-layer wiring cost function for the adjusted cell positions and performing wiring layer allocation based on the inter-layer wiring cost function; calculating wiring congestion after the wiring layer allocation, wherein the wiring congestion is a ratio of wiring demand to wiring capacity, and performing power consumption-aware adjustment on the line width of the network based on the wiring congestion;

[0042] Constructing a layout quality evaluation function, wherein the layout quality evaluation function is a weighted sum of the ratios of each evaluation index to a target value, the evaluation index including a timing margin index, a power consumption index, and a wiring quality index, and calculating the gradient of the layout quality evaluation function;

[0043] Multiplying the gradient of the layout quality assessment function by the constraint matrix to obtain an optimization direction, and performing iterative optimization using an adaptive step size, wherein the adaptive step size decays exponentially with an increase in the number of iterations;

[0044] According to the optimization direction, the unit position and routing scheme are updated until the improvement of the layout quality evaluation function is less than a preset improvement threshold and the worst negative margin is greater than zero and the total negative margin is greater than zero, thereby generating a memory compiler IP layout and routing scheme that meets the timing closure requirements and optimizes power consumption. The second aspect of the embodiment of the present invention is,

[0045] Provides memory compiler IP timing closure and power consumption co-optimization system, including:

[0046] The first unit is configured to obtain circuit netlist information of a memory compiler IP, extract a timing path based on the circuit netlist information, perform wiring topology analysis on the timing path, and determine the connection length and interconnection delay between timing cells in the timing path;

[0047] a second unit, configured to divide the timing path into a plurality of power control domains based on the connection length and the interconnection delay by adopting a fine-grained power control strategy, and set a different operating voltage for each of the power control domains, so that the operating voltage of the power control domain close to the timing end point is higher than the operating voltage of the power control domain close to the timing start point;

[0048] A third unit is configured to construct a deep neural network model based on the voltage and delay data of the power consumption control domain, obtain a delay prediction value based on the deep neural network model; dynamically adjust parameters of the deep neural network model according to a prediction error between the delay prediction value and the actual delay value and calculate a reliability index; and use the delay prediction value for timing optimization when the reliability index meets a preset reliability threshold;

[0049] A fourth unit is configured to obtain characteristic parameters of the time series path and construct a characteristic vector matrix based on the characteristic parameters; train a decision tree model based on the characteristic vector matrix to obtain an optimization strategy, and iteratively update the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold;

[0050] The fifth unit is used to calculate the path priority based on the timing path, functionally partition the memory compiler IP based on the path priority and construct a partition cost function; perform cell location optimization and hierarchical routing optimization according to the partition cost function, and generate a memory compiler IP layout and routing solution that meets timing convergence requirements and optimizes power consumption.

[0051] According to a third aspect of the embodiments of the present invention,

[0052] An electronic device is provided, comprising:

[0053] processor;

[0054] a memory for storing processor-executable instructions;

[0055] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0056] According to a fourth aspect of the embodiments of the present invention,

[0057] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0058] The beneficial effects of this application are as follows:

[0059] The present invention divides the timing path into multiple power control domains through a fine-grained power control strategy, and sets a different operating voltage for each power control domain, thereby achieving coordinated optimization of the timing convergence and power consumption of the memory compiler IP, effectively balancing the contradiction between timing performance and power consumption, and improving the overall performance of the memory compiler IP.

[0060] The present invention predicts delays based on a deep neural network model and improves prediction accuracy by dynamically adjusting model parameters. It also uses a decision tree model to generate optimization strategies, enabling precise analysis and intelligent optimization of timing paths. This significantly improves timing closure efficiency, reduces the number of design iterations, and shortens the design cycle of the memory compiler IP.

[0061] The present invention achieves global optimization of the memory compiler IP through path priority calculation and functional partitioning methods, combined with cost function-guided unit location optimization and hierarchical routing optimization, significantly improving the timing margin of the critical path, while reducing overall power consumption and improving the stability and reliability of the memory compiler IP under different working conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 Schematic diagram of the flow of a method for co-optimizing timing closure and power consumption of a memory compiler IP according to an embodiment of the present invention;

[0063] Figure 2 This is a schematic diagram comparing the loss convergence curves of the deep neural network model training according to an embodiment of the present invention;

[0064] Figure 3 This is a Monte Carlo simulation wiring quality comparison heat map of an embodiment of the present invention. DETAILED DESCRIPTION

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0066] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0067] Figure 1 FIG. 1 is a flow chart of a method for co-optimizing timing closure and power consumption of a memory compiler IP according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0068] Obtain circuit netlist information of the memory compiler IP, extract a timing path based on the circuit netlist information, perform wiring topology analysis on the timing path, and determine the connection length and interconnection delay between timing cells in the timing path;

[0069] Based on the connection length and the interconnection delay, a fine-grained power consumption control strategy is adopted to divide the timing path into multiple power consumption control domains, and a different operating voltage is set for each of the power consumption control domains, so that the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point;

[0070] A deep neural network model is constructed based on the voltage and delay data of the power consumption control domain, and a delay prediction value is obtained based on the deep neural network model; parameters of the deep neural network model are dynamically adjusted according to a prediction error between the delay prediction value and the actual delay value, and a reliability index is calculated; when the reliability index meets a preset reliability threshold, the delay prediction value is used for timing optimization;

[0071] Acquire characteristic parameters of the time series path and construct a characteristic vector matrix based on the characteristic parameters; train a decision tree model based on the characteristic vector matrix to obtain an optimization strategy, and iteratively update the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold;

[0072] Calculate the path priority based on the timing path, perform functional partitioning on the memory compiler IP based on the path priority and construct a partition cost function; perform cell location optimization and hierarchical routing optimization according to the partition cost function, and generate a memory compiler IP layout and routing solution that meets timing closure requirements and optimizes power consumption.

[0073] In an optional embodiment, based on the connection length and the interconnection delay, a fine-grained power consumption control strategy is used to divide the timing path into multiple power consumption control domains, and different operating voltages are set for each of the power consumption control domains, so that the operating voltage of the power consumption control domain near the timing end point is higher than the operating voltage of the power consumption control domain near the timing start point, including:

[0074] Acquiring physical parameters of the line segment in the timing path, the physical parameters including the length of the line segment, the width of the line segment, the thickness of the interlayer dielectric of the line segment, and the metal specific resistance of the line segment;

[0075] Calculating the resistance per unit length and the capacitance per unit length of the line segment using a distributed RC model based on the physical parameters of the line segment, and calculating the total resistance and total capacitance of the line segment according to the resistance per unit length and the capacitance per unit length;

[0076] Calculating the interconnection delay of the connection segment based on the total resistance and total capacitance of the connection segment to obtain the interconnection delay distribution characteristics of the timing path; and dividing adjacent sequential cells with similar delay sensitivities into a plurality of power consumption control domains using a hierarchical clustering algorithm based on the delay sensitivity of the sequential cells in the timing path, wherein the delay sensitivity represents the degree of influence of the operating voltage of the sequential cell on the total delay of the timing path;

[0077] For the multiple power consumption control domains obtained by division, under the condition of satisfying the timing constraints, the operating voltage of each power consumption control domain is determined by minimizing the dynamic power consumption of each power consumption control domain, wherein the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point.

[0078] Obtain the physical parameters of the wire segments in the timing path. In one specific embodiment, the key timing paths in the chip design are analyzed to extract the physical parameters of each wire segment, including the length, width, interlayer dielectric thickness, and metal resistivity of the wire segment. For example, for a timing path containing 10 wire segments, the first wire segment has a length of 100 microns, a width of 0.5 microns, an interlayer dielectric thickness of 0.3 microns, and a metal resistivity of 0.05 ohms / square; the second wire segment has a length of 150 microns, a width of 0.6 microns, an interlayer dielectric thickness of 0.35 microns, and a metal resistivity of 0.06 ohms / square; and so on, the physical parameters of all wire segments are obtained.

[0079] Based on the physical parameters of the connection segment, a distributed RC model is used to calculate the resistance and capacitance per unit length of the connection segment. In actual implementation, the resistance per unit length can be calculated by dividing the metal resistivity by the connection width. For example, the resistance per unit length of the first connection segment is 0.05 ohms / square divided by 0.5 microns, which equals 0.1 ohms / micron. The capacitance per unit length is calculated according to the parallel plate capacitor model, taking into account the connection width, interlayer dielectric thickness, and dielectric constant. Assuming a dielectric constant of 3.9, the capacitance per unit length of the first connection segment can be calculated to be 0.057 femtofarads / micron.

[0080] Calculate the total resistance and capacitance of the trace segment based on the resistance per unit length and capacitance per unit length. The total resistance is equal to the resistance per unit length multiplied by the trace length, and the total capacitance is equal to the capacitance per unit length multiplied by the trace length. For the first trace segment, for example, the total resistance is 0.1 ohms / micrometer multiplied by 100 micrometers, which equals 10 ohms; the total capacitance is 0.057 femtofarads / micrometer multiplied by 100 micrometers, which equals 5.7 femtofarads.

[0081] Based on the total resistance and total capacitance of the line segments, the interconnection delay of the line segments is calculated to obtain the interconnection delay distribution characteristics of the timing path. In one embodiment, the interconnection delay is calculated using the Elmore delay model, which considers the impact of resistance and capacitance in the RC network on signal transmission. For example, the interconnection delay of the first line segment can be calculated as 10 ohms multiplied by half of 5.7 femtofarads, which is approximately 28.5 picoseconds. Similar calculations are performed for all line segments to obtain a complete interconnection delay distribution characteristic.

[0082] According to the delay sensitivity of the timing cells in the timing path, a hierarchical clustering algorithm is used to divide adjacent timing cells with similar delay sensitivity into multiple power control domains. Delay sensitivity characterizes the degree of influence of the operating voltage of the timing cell on the total delay of the timing path. In a specific implementation, the delay change rate of each timing cell under different operating voltages can be obtained through a static timing analysis tool. For example, for a path containing 20 timing cells, when the operating voltage is reduced from 1.0 volts to 0.9 volts, the delay of the first timing cell increases by 15%, the delay of the second timing cell increases by 14%, the delay of the third timing cell increases by 16%, the delay of the eighteenth timing cell increases by 25%, the delay of the nineteenth timing cell increases by 27%, and the delay of the twentieth timing cell increases by 26%.

[0083] The specific implementation steps of the hierarchical clustering algorithm are as follows: First, each timing unit is treated as an independent power control domain; then, the delay sensitivity difference between adjacent power control domains is calculated; next, the adjacent power control domains with the smallest delay sensitivity difference are merged; and the merging process is repeated until the preset number of power control domains is reached or the delay sensitivity difference exceeds the threshold. In the above example, the threshold is set to 5%, and the 20 timing units are divided into three power control domains: the 1st to 10th timing units are the first power control domain (delay sensitivity of approximately 15%), the 11th to 15th timing units are the second power control domain (delay sensitivity of approximately 20%), and the 16th to 20th timing units are the third power control domain (delay sensitivity of approximately 26%).

[0084] For the multiple power control domains obtained by division, the operating voltage of each power control domain is determined by minimizing the dynamic power consumption of each power control domain while satisfying the timing constraints. Dynamic power consumption is proportional to the square of the operating voltage, so reducing the operating voltage is an effective way to reduce dynamic power consumption. In one embodiment, an iterative optimization method is used to determine the operating voltage of each power control domain: first, the initial operating voltage of all power control domains is set to a standard voltage (e.g., 1.0 volt); then, starting from the timing starting point, the operating voltage of each power control domain is gradually reduced, while ensuring that the timing constraints are still met through static timing analysis; finally, when the operating voltage of any power control domain cannot be further reduced without violating the timing constraints, the optimization process ends.

[0085] In the above example, after optimization, the operating voltage of the first power control domain (near the start of the sequence) is set to 0.8 volts, the operating voltage of the second power control domain is set to 0.9 volts, and the operating voltage of the third power control domain (near the end of the sequence) is set to 1.0 volt. This setting ensures that the operating voltage of the power control domain near the end of the sequence is higher than that of the power control domain near the start of the sequence, thus achieving power optimization while meeting timing constraints.

[0086] In actual testing, applying this approach to a design containing 1,000 timing paths achieved approximately 20% lower dynamic power consumption compared to a traditional single-voltage domain approach, while maintaining a timing margin of at least 50 picoseconds. Furthermore, compared to traditional multi-voltage domain approaches, this approach, through its fine-grained power control strategy, achieved an additional 8% lower dynamic power consumption.

[0087] The method is applicable to various integrated circuit designs, particularly those requiring stringent power consumption and performance, such as mobile device chips, artificial intelligence accelerators, and low-power IoT devices. By implementing a fine-grained power consumption control strategy, chip power consumption can be significantly reduced while maintaining performance, extending battery life and improving user experience.

[0088] In an optional embodiment, based on the physical parameters of the line segment, calculating the resistance per unit length and the capacitance per unit length of the line segment using a distributed RC model, and calculating the total resistance and total capacitance of the line segment based on the resistance per unit length and the capacitance per unit length includes:

[0089] The distributed RC model is used to calculate the resistance per unit length of the connecting segment. The calculation process of the resistance per unit length is as follows: dividing the metal specific resistance by the product of the width of the connecting segment and the thickness of the interlayer dielectric;

[0090] Calculating the capacitance per unit length of the line segment based on the physical parameters of the line segment, wherein the calculation process of the capacitance per unit length includes: calculating the parallel plate capacitance by dividing the product of the vacuum relative dielectric constant system and the width of the line segment by the thickness of the interlayer dielectric; calculating the edge capacitance by multiplying the vacuum relative dielectric constant system by a logarithmic function of the ratio of the interlayer dielectric thickness to the width of the line segment; calculating the coupling capacitance by dividing the product of the vacuum relative dielectric constant system and the thickness of the interlayer dielectric by the metal resistivity; and adding the parallel plate capacitance, the edge capacitance, and the coupling capacitance to obtain the capacitance per unit length;

[0091] Acquire a signal frequency parameter of the connecting segment, the signal frequency parameter including a signal frequency band and a characteristic frequency parameter, and multiply the signal frequency parameter by the resistance per unit length to obtain a frequency-corrected resistance per unit length;

[0092] The connecting segment is divided into multiple equal-length sub-segments, and the product of the frequency-corrected unit length resistance of each sub-segment and the sub-segment length is accumulated to obtain the total resistance of the connecting segment, and the product of the unit length capacitance of each sub-segment and the sub-segment length is accumulated to obtain the total capacitance of the connecting segment.

[0093] Obtain the physical parameters of the trace segment, including its width, length, metal resistivity, interlayer dielectric thickness, and vacuum relative permittivity. For example, a 90nm metal trace segment has a width of 0.12 microns, a length of 100 microns, a metal resistivity of 2.2 × 10-8 ohm·m, an interlayer dielectric thickness of 0.18 microns, and a vacuum relative permittivity of 8.85 × 10-12 farad / m.

[0094] A distributed RC model is used to calculate the resistance per unit length of a trace segment. The specific calculation process is: divide the metal resistivity by the product of the trace width and the interlayer dielectric thickness. Using the above parameters as an example, the resistance per unit length is calculated as follows: 2.2 × 10^-8 ohm·meter divided by (the product of the trace width (0.12 micrometers) and the interlayer dielectric thickness (0.18 micrometers), resulting in a resistance per unit length of 1.02 × 10^3 ohm / meter.

[0095] Calculate the capacitance per unit length of a line segment based on its physical parameters. The capacitance per unit length calculation consists of three parts: parallel plate capacitance, fringe capacitance, and coupling capacitance.

[0096] The parallel plate capacitance is calculated by multiplying the relative permittivity of a vacuum system by the width of the connecting wire segment, divided by the thickness of the interlayer dielectric. Using the above parameters as an example, the parallel plate capacitance is calculated as follows: the relative permittivity of a vacuum system, 8.85 × 10^-12 farads / meter, multiplied by the width of the connecting wire segment, 0.12 microns, divided by the interlayer dielectric thickness, 0.18 microns, yielding a parallel plate capacitance of 5.9 × 10^-11 farads / meter.

[0097] The edge capacitance is calculated by multiplying the relative permittivity of a vacuum system by the logarithm of the ratio of the interlayer dielectric thickness to the width of the connecting wire. Using the above parameters as an example, the edge capacitance is calculated as follows: the relative permittivity of a vacuum system, 8.85 × 10^-12 farads / meter, is multiplied by the logarithm of the ratio of the interlayer dielectric thickness of 0.18 microns to the width of the connecting wire of 0.12 microns (the natural logarithm, ln(0.18 / 0.12)≈0.405), resulting in an edge capacitance of 3.58 × 10^-12 farads / meter.

[0098] The coupling capacitance is calculated by multiplying the relative permittivity of a vacuum system by the thickness of the interlayer dielectric and dividing it by the metal resistivity. Using the above parameters as an example, the coupling capacitance is calculated as follows: the relative permittivity of a vacuum system, 8.85 × 10^-12 farads / meter, multiplied by the interlayer dielectric thickness, 0.18 microns, divided by the metal resistivity, 2.2 × 10^-8 ohm·meters, yielding a coupling capacitance of 7.23 × 10^-5 farads / meter.

[0099] Adding the parallel plate capacitance, fringe capacitance, and coupling capacitance, we get a capacitance per unit length of 7.23×10^-5 farads / meter.

[0100] Obtain the signal frequency parameters of the segment, including the signal frequency band and characteristic frequency parameters. For example, for a segment operating at 1 GHz, its characteristic frequency parameter can be set to 1 × 10^9 Hz. Multiply the signal frequency parameter by the resistance per unit length to obtain the frequency-corrected resistance per unit length. Using the above parameters as an example, the frequency-corrected resistance per unit length is 1.02 × 10^3 ohms / meter multiplied by 1 × 10^9 Hz, or 1.02 × 10^12 ohms·Hz / meter.

[0101] Divide the wire segment into multiple equal-length subsegments and calculate the resistance and capacitance of each subsegment. Assume that the 100-micron-long wire segment mentioned above is divided into 10 equal-length subsegments, each 10 microns long. For each subsegment, calculate the product of its frequency-corrected resistance per unit length and the subsegment length to obtain the subsegment resistance. Using the above parameters as an example, the resistance of each subsegment is 1.02 × 10^12 ohm·Hz / meter multiplied by 10 microns, or 1.02 × 10^7 ohm·Hz. Adding up the resistances of all subsegments yields a total resistance of 1.02 × 10^8 ohm·Hz.

[0102] For each subsegment, calculate the product of its capacitance per unit length and the subsegment length to obtain the subsegment capacitance. Using the above parameters as an example, the capacitance of each subsegment is 7.23 × 10^-5 farads / meter multiplied by 10 microns, or 7.23 × 10^-10 farads. Adding up the capacitances of all subsegments yields a total capacitance of 7.23 × 10^-9 farads for the connected segment.

[0103] The number of sub-segments can be adjusted based on the specific situation of the line segment to improve calculation accuracy. For example, for line segments with large length variations or uneven widths, the number of sub-segments can be increased, or sub-segments of unequal lengths can be divided based on the actual shape of the line segment.

[0104] The calculation process can also consider the effects of environmental factors such as temperature and humidity on the resistance and capacitance of a segment. For example, rising temperature increases the specific resistivity of a metal, thus affecting the resistance of the segment. This can be corrected by introducing a temperature coefficient.

[0105] The method accurately calculates the total resistance and capacitance of a wire segment, providing an important reference for integrated circuit design. This method is applicable to integrated circuit design at various process nodes, and is particularly useful for circuit design involving high-frequency signal transmission. It can effectively predict signal transmission delays and distortion, improving design accuracy and reliability.

[0106] In an optional embodiment, dynamically adjusting the parameters of the deep neural network model and calculating a reliability index based on a prediction error between the delay prediction value and the actual delay value, and using the delay prediction value for timing optimization when the reliability index meets a preset reliability threshold includes:

[0107] Constructing a comprehensive loss function, training the deep neural network model using a gradient descent method, calculating weight gradients according to the comprehensive loss function, and updating weight parameters of the neural network based on the weight gradients;

[0108] Obtaining a new operating voltage combination, inputting the operating voltage combination into the deep neural network model, and obtaining a corresponding delay prediction value;

[0109] Collecting an actual delay measurement value under the operating voltage combination, and calculating a prediction error between the delay prediction value and the actual delay measurement value;

[0110] Dynamically adjusting model parameters according to the prediction error, calculating a weight update amount, wherein the weight update amount is proportional to the product of the prediction error and the weight gradient; and calculating a bias update amount, wherein the bias update amount is proportional to the sign value of the prediction error;

[0111] Calculating a confidence level of a prediction result based on the prediction error, wherein the confidence level decreases as the prediction variance increases, and using the difference between the confidence level and the prediction error as a reliability indicator;

[0112] When the reliability index is greater than the preset reliability threshold, the current delay prediction value is used to guide timing optimization; when the reliability index is less than the preset reliability threshold, the model online learning process is triggered, new training data is added to the historical data set, and the deep neural network model is retrained.

[0113] A comprehensive loss function is constructed, and the deep neural network model is trained using the gradient descent method. The deep neural network model includes an input layer, multiple hidden layers, and an output layer. The input layer receives the operating voltage combination data, and the output layer outputs the corresponding delay prediction value. The comprehensive loss function consists of a mean square error term and a regularization term. The mean square error term is used to measure the gap between the predicted value and the actual value, and the regularization term is used to prevent the model from overfitting. Specifically, the mean square error term is the sum of the squares of the differences between the predicted delay value and the actual delay value, and the regularization term is the sum of the squares of the model weight parameters multiplied by the regularization coefficient. The regularization coefficient can be set to 0.001 to control the regularization strength.

[0114] Batch gradient descent was used, with 64 samples selected per batch. The learning rate was initially set to 0.01, and a learning rate decay strategy was implemented, with the learning rate multiplied by 0.95 every 100 training batches. Weight gradients were calculated based on the combined loss function, which is the partial derivative of the loss function with respect to the weights of each layer. For example, for the weight matrix of layer i, its gradient calculation involves multiplying the input of that layer by the error term of the next layer. Once the gradient is calculated, the neural network weight parameters are updated based on the gradient: the current weight minus the learning rate multiplied by the gradient.

[0115] Obtain a new operating voltage combination. This combination includes core voltage, memory voltage, and input / output voltages, for example (0.9V, 1.1V, 1.8V). This operating voltage combination is fed into the trained deep neural network model, and through forward propagation, the corresponding latency prediction value is obtained. For example, for the above voltage combination, the model predicts a latency of 2.5 nanoseconds.

[0116] Collect the actual latency measurement for this operating voltage combination. This can be done by running a standard test program on actual hardware and recording the execution time. Assume the actual measured latency is 2.7 nanoseconds. Calculate the prediction error between the predicted latency and the actual measured latency: 2.5 nanoseconds minus 2.7 nanoseconds, which equals -0.2 nanoseconds.

[0117] Dynamically adjust model parameters based on prediction error. First, calculate the weight update amount, which is proportional to the product of the prediction error and the weight gradient. Specifically, the weight update amount is equal to the learning rate multiplied by the prediction error multiplied by the weight gradient. For example, if the gradient of a weight is 0.05, the prediction error is -0.2 nanoseconds, and the learning rate is 0.01, then the update amount of the weight is 0.01×(-0.2)×0.05=-0.0001. At the same time, calculate the bias update amount, which is proportional to the sign value of the prediction error. For example, if the prediction error is -0.2 nanoseconds, its sign value is -1, and if the bias update coefficient is 0.005, then the bias update amount is 0.005×(-1)=-0.005.

[0118] The confidence level of the forecast result is calculated based on the forecast error. Confidence decreases as the forecast variance increases. The forecast variance can be estimated by dividing the sum of the squares of the historical forecast errors by the number of samples minus one. For example, if the sum of the squares of the errors for the past 10 forecasts is 0.5, the forecast variance is estimated to be 0.5 / (10-1)≈0.056. The confidence level can be expressed as 1 minus the ratio of the forecast variance to a preset variance threshold, which can be set to 0.1. In this example, the confidence level is 1-(0.056 / 0.1)=0.44. The difference between the confidence level and the absolute value of the forecast error is used as the reliability indicator. If the absolute value of the forecast error is 0.2, the reliability indicator is 0.44-0.2=0.24.

[0119] When the reliability index exceeds the preset reliability threshold, the current delay prediction value is used to guide timing optimization. The preset reliability threshold can be set to 0.2. In this example, the reliability index is 0.24, which is greater than the preset reliability threshold of 0.2. Therefore, the current delay prediction value of 2.5 nanoseconds is used to guide timing optimization. Timing optimization includes operations such as adjusting clock frequency and reallocating critical path resources.

[0120] When the reliability index falls below the preset reliability threshold, the model's online learning process is triggered. For example, if the calculated reliability index is 0.15, which is less than the preset reliability threshold of 0.2, online model learning is triggered. During the online learning process, new training data is added to the historical dataset. This new training data includes operating voltage combinations and their corresponding actual delay measurements, such as ((0.9V, 1.1V, 1.8V), 2.7 nanoseconds). The historical dataset can be maintained using a first-in-first-out strategy, maintaining a fixed size of, for example, 1000 samples. After the new data is added, the deep neural network model is retrained. Retraining can use incremental learning, starting with the current model parameters and using the new dataset for several rounds of training, for example, five rounds, with each round traversing the entire dataset.

[0121] The parameters of the deep neural network model can be dynamically adjusted according to the prediction error, and the reliability index can be calculated. When the reliability index meets the conditions, the delay prediction value can be used to perform timing optimization, thereby improving system performance and reliability.

[0122] Figure 2 This is a schematic diagram comparing the loss convergence curves of the deep neural network model training according to an embodiment of the present invention:

[0123] This figure compares the loss function changes of three different technical solutions during training iterations: our technical solution (combined loss function), the traditional method (MSE loss), and the comparison method (MAE loss). The curves show that the initial loss values ​​of all three solutions are close to 0.9. As the number of training iterations increases, the loss values ​​show varying degrees of decline. Our technical solution converges the fastest, with the loss value dropping to approximately 0.5 after 50 iterations, further decreasing to approximately 0.35 after 100 iterations, and then to approximately 0.1 after 300 iterations, ultimately stabilizing at a level close to 0.05 after 500 iterations. In contrast, the traditional method (MSE loss) converges more slowly, with the loss value still around 0.7 after 50 iterations, and only dropping to approximately 0.28 after 500 iterations. The comparison method (MAE loss) performs somewhere in between, ultimately dropping to approximately 0.2 after 500 iterations. This result fully demonstrates that this technical solution not only has a faster convergence speed, but also can achieve a lower final loss value. It is significantly superior to traditional methods and comparative methods in terms of training efficiency and model performance, reflecting the significant advantages of this technical solution in deep learning optimization.

[0124] This technical solution significantly improves optimization efficiency by introducing a comprehensive loss function to construct the model. Existing technical solutions often use the MSE loss function or the MAE loss function for model training. This single loss function design fails to fully consider multi-dimensional optimization objectives, resulting in slow convergence during the training process and high final loss values, making it difficult to meet the requirements of high-precision model optimization. To address this issue, this technical solution innovatively designs a comprehensive loss function that integrates multi-dimensional features. By rationally allocating the weights of features in each dimension, it achieves an effective balance of optimization objectives. Furthermore, an adaptive learning strategy is employed during training to dynamically adjust parameters based on the optimization status at different stages, further improving the model's convergence efficiency. This optimization solution, based on a comprehensive loss function and adaptive learning, not only significantly accelerates model convergence but also achieves a lower final loss value, achieving a qualitative improvement in both training efficiency and model performance. Experimental results demonstrate that compared to traditional single-loss function solutions, this technical solution offers significant advantages in both convergence speed and final optimization results, providing a new solution for the efficient optimization of deep learning models.

[0125] In an optional embodiment, constructing a feature vector matrix based on the feature parameters; training a decision tree model based on the feature vector matrix to obtain an optimization strategy, and iteratively updating the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold includes:

[0126] Constructing a eigenvector matrix based on the characteristic parameters, the eigenvector matrix including a timing eigenvector, a power consumption eigenvector, and an area eigenvector;

[0127] Calculating the sensitivity of the timing feature vector, the power consumption feature vector, and the area feature vector to voltage, and constructing a sensitivity feature vector, where the sensitivity feature vector includes the sensitivity of timing to voltage, the sensitivity of power consumption to voltage, and the sensitivity of area to voltage;

[0128] Combining the timing feature vector, the power consumption feature vector, the area feature vector, and the sensitivity feature vector to form an input feature matrix;

[0129] Building a decision tree model based on the input feature matrix, and outputting a voltage adjustment amount, a size adjustment amount, and a topology adjustment solution through the decision tree model, wherein the voltage adjustment amount is obtained by mapping the sigmoid function, the size adjustment amount is obtained by mapping the tanh function, and the topology adjustment solution is obtained by mapping the softmax function;

[0130] Constructing an optimization loss function and calculating a loss function gradient based on the optimization loss function; calculating an optimization target improvement and determining a sample weight based on the optimization target improvement, wherein the sample weight increases as the optimization target improvement increases;

[0131] The parameters of the decision tree model are updated according to the sample weights and the loss function gradient, and the timing path is optimized and adjusted based on the updated decision tree model until the weighted ratio of the optimization target improvement to the optimization overhead meets a preset condition threshold.

[0132] A feature vector matrix is ​​constructed based on the characteristic parameters of the circuit. In one embodiment, the characteristic parameters of the circuit in three dimensions, namely, timing, power consumption, and area, are collected. The timing feature vector includes parameters such as path delay, setup time margin, and hold time margin; the power consumption feature vector includes parameters such as dynamic power consumption, static power consumption, and short-circuit power consumption; and the area feature vector includes parameters such as unit area, wiring area, and total chip area. For example, for a design containing 1,000 timing paths, timing parameters such as delay value (such as 0.9ns) and setup time margin (such as 0.2ns) can be extracted for each path, while power consumption data (such as dynamic power consumption 5mW) and area data (such as unit area 0.05mm²) are recorded.

[0133] Calculate the sensitivity of the eigenvector to voltage and construct a sensitivity eigenvector. Specifically, by measuring circuit parameters under different voltage conditions (such as 0.8V, 0.9V, and 1.0V), calculate the parameter change rate. For example, if a timing path has a delay of 0.9ns at 0.9V and a delay of 0.8ns at 1.0V, the timing sensitivity to voltage is (0.9-0.8) / 0.1=1ns / V. Similarly, calculate the sensitivity of power consumption to voltage (such as 20mW / V) and the sensitivity of area to voltage (usually close to 0 because the area is almost unaffected by voltage). These sensitivity data form a sensitivity eigenvector.

[0134] The timing feature vector, power consumption feature vector, area feature vector, and sensitivity feature vector are combined to form the input feature matrix. In practice, a multidimensional feature vector containing all the above features can be constructed for each timing path. For example, the feature vector for path 1 is [0.9ns, 0.2ns, 5mW, 2mW, 0.05mm², 1ns / V, 20mW / V, 0mm² / V]. The feature vectors of all paths are combined into a feature matrix, which serves as the input to the decision tree model.

[0135] A decision tree model is constructed based on the input feature matrix. This model uses a gradient boosted decision tree (GBDT) structure and consists of multiple decision tree nodes, each of which makes a binary decision based on a specific feature value. The model outputs three types of optimization recommendations: voltage adjustment, size adjustment, and topology adjustment. Voltage adjustment is mapped to the range [-0.1V, 0.1V] using a sigmoid function. For example, the model recommends increasing the voltage by 0.05V. Size adjustment is mapped to the range [-50%, 50%] using a tanh function. For example, the model recommends increasing the size of a cell by 20%. For topology adjustment, a softmax function is used to select the optimal solution from several predefined options, such as inserting a buffer, changing the cell type, or adjusting the clock tree.

[0136] Construct an optimization loss function to evaluate model performance. The loss function comprehensively considers timing improvement, increased power consumption, and increased area, and can be expressed as the weighted sum of the timing improvement minus the increased power consumption and area. For example, if an optimization solution reduces critical path delay by 0.1ns but increases power consumption by 1mW and area by 0.01mm², the loss can be calculated as 0.1-0.5×1-0.3×0.01=0.097 (assuming a power consumption weight of 0.5 and an area weight of 0.3). Based on this loss function, a gradient is calculated to guide model parameter updates.

[0137] Calculate the optimization target improvement and determine the sample weights. The optimization target improvement is defined as the degree of improvement in the combined metrics of timing, power consumption, and area before and after optimization. For example, if the critical path delay decreases from 1ns to 0.9ns, a 10% improvement, a higher weight is assigned; if it only decreases from 1ns to 0.99ns, a 1% improvement, a lower weight is assigned. In practice, the sample weights can be set proportional to the improvement, for example, weight = 1 + 10 × improvement rate. A 10% improvement corresponds to a weight of 1.1, and a 1% improvement corresponds to a weight of 1.01.

[0138] The decision tree model parameters are updated based on the sample weights and the loss function gradient. During training, the loss function gradient is calculated for each sample, multiplied by the corresponding sample weight, and used to update the decision tree parameters. For example, a sample with a weight of 1.1 will contribute approximately 10% more to the gradient than a sample with a weight of 1.01. Through multiple rounds of iterative training, the model gradually learns a more effective optimization strategy.

[0139] Apply the updated decision tree model to optimize the timing path. The model provides specific voltage adjustment recommendations, size adjustment recommendations, and topology adjustment recommendations for each path that needs to be optimized. For example, for a timing violation path, the model recommends increasing its driver unit size by 15% and inserting a buffer at a specific location. After implementing these optimizations, evaluate the optimization effect again and calculate the weighted ratio of the improvement in the optimization target to the optimization overhead (such as area increase and power consumption increase). When the ratio reaches a preset threshold (such as greater than 2.0, indicating that the benefit is twice the overhead), the optimization process is stopped; otherwise, the optimization continues to iterate until the condition is met or the maximum number of iterations is reached.

[0140] The present invention can automatically learn circuit optimization strategies and provide customized optimization solutions for different types of timing paths, effectively improving circuit performance.

[0141] In existing technologies, traditional greedy algorithms primarily rely on single-dimensional local optimization strategies for circuit design, while static optimization methods employ fixed optimization parameters. These methods struggle to fully account for the multidimensional sensitivity metrics in circuit design, resulting in poor optimization results for key characteristics such as voltage, size, and topology. To address this issue, this technical solution innovatively proposes a multidimensional sensitivity collaborative optimization method. By constructing a comprehensive evaluation system encompassing multiple dimensions, including sequence, power consumption, and area, it achieves comprehensive optimization of voltage sensitivity, size sensitivity, and topology sensitivity. This solution not only theoretically establishes a correlation model for multidimensional characteristics but also employs a dynamic, adaptive optimization strategy in practice, enabling flexible adjustment of optimization weights based on the characteristics of different design stages. Through this multidimensional collaborative optimization approach, this technical solution significantly improves the robustness of circuit design, demonstrating excellent adaptability to voltage fluctuations, size changes, and topology adjustments. Compared with existing technologies, this solution achieves comprehensive improvements across all key performance indicators, with particularly significant improvements in the two core metrics of voltage sensitivity and size sensitivity, providing more reliable technical support for high-performance integrated circuit design.

[0142] In an optional embodiment, performing cell location optimization and hierarchical routing optimization according to the partition cost function to generate a memory compiler IP placement and routing solution that meets timing closure requirements and optimizes power consumption includes:

[0143] Functionally partitioning the memory compiler IP based on the partition cost function; constructing a unit position optimization cost function based on the result of the functional partitioning, and determining the optimization direction of the unit according to the gradient direction of the unit position optimization cost function;

[0144] Multiplying the optimization direction of the unit by an exponential decay function of a power consumption change to obtain a power consumption-aware unit movement amount, where the power consumption change amount is determined by a difference in power consumption before and after the movement, and adjusting the unit position based on the power consumption-aware unit movement amount;

[0145] Constructing an inter-layer wiring cost function for the adjusted cell positions and performing wiring layer allocation based on the inter-layer wiring cost function; calculating wiring congestion after the wiring layer allocation, wherein the wiring congestion is a ratio of wiring demand to wiring capacity, and performing power consumption-aware adjustment on the line width of the network based on the wiring congestion;

[0146] Constructing a layout quality evaluation function, wherein the layout quality evaluation function is a weighted sum of the ratios of each evaluation index to a target value, the evaluation index including a timing margin index, a power consumption index, and a wiring quality index, and calculating the gradient of the layout quality evaluation function;

[0147] Multiplying the gradient of the layout quality assessment function by the constraint matrix to obtain an optimization direction, and performing iterative optimization using an adaptive step size, wherein the adaptive step size decays exponentially with an increase in the number of iterations;

[0148] The unit position and routing scheme are updated according to the optimization direction until the improvement of the layout quality evaluation function is less than a preset improvement threshold and the worst negative margin is greater than zero and the total negative margin is greater than zero, thereby generating a memory compiler IP layout and routing scheme that meets timing convergence requirements and optimizes power consumption.

[0149] The memory compiler IP is functionally partitioned based on a partition cost function. The partition cost function takes into account the connectivity between cells, functional relevance, and timing criticality, and can be expressed as the sum of the product of the connection weight and the distance between cells. In specific implementation, the cells in the memory compiler IP are divided into four main functional areas according to their functions: the control logic area, the address decoding area, the storage array area, and the input / output buffer area. For example, for an 8KB SRAM memory, the control logic area accounts for 15% of the total area, the address decoding area accounts for 25%, the storage array area accounts for 50%, and the input / output buffer accounts for 10%. By minimizing the partition cost function, cells with strong functional relevance are ensured to be allocated to the same area, reducing cross-region connections.

[0150] Based on the functional zoning results, a cost function for optimizing cell locations is constructed. This cost function comprehensively considers three factors: line length, congestion, and timing margin, integrating them in a weighted manner. By calculating the gradient of this cost function with respect to each cell location, the optimization direction for each cell is determined. For example, for cells on the timing-critical path, the timing weight can be set to 0.6, the line length weight to 0.3, and the congestion weight to 0.1. For cells on non-critical paths, the timing weight can be reduced to 0.2, the line length weight increased to 0.5, and the congestion weight set to 0.3.

[0151] Calculate the difference in power consumption before and after the unit is moved to determine the change in power consumption. Substitute this change in power consumption into an exponential decay function to obtain a coefficient ranging from 0 to 1. Multiply this coefficient by the aforementioned unit optimization direction to obtain the power-aware unit movement. In the specific implementation, when the change in power consumption is positive (i.e., power consumption increases after movement), the value of the exponential decay function is small, inhibiting the unit from moving in that direction; when the change in power consumption is negative, the value of the exponential decay function is close to 1, allowing the unit to move fully. For example, when movement causes a 5% increase in power consumption, the suppression coefficient is 0.3; and when movement causes a 3% decrease in power consumption, the coefficient is 0.9.

[0152] For the adjusted cell positions, an inter-layer routing cost function is constructed to assign routing layers. This cost function considers the resistance and capacitance characteristics of different metal layers and their impact on power consumption. Low-layer metals (such as Metal 1-3) have higher resistance but lower capacitance, while high-layer metals (such as Metal 4-6) have lower resistance but higher capacitance. Timing-critical nets are prioritized for routing to high-layer metals with lower resistance. Non-critical nets with high switching frequencies are prioritized for routing to low-layer metals with lower capacitance to reduce power consumption. For example, clock networks can be prioritized for Metal 5-6, while data buses can use Metal 3-4.

[0153] After routing layer allocation is complete, routing congestion is calculated—the ratio of routing demand to routing capacity. For areas with congestion exceeding 80%, power-aware adjustments to net line widths are performed. For power-sensitive but non-timing-critical nets, line widths can be reduced to the minimum design rule; for timing-critical nets, line widths are maintained to meet timing requirements. For example, for an area with a 90% congestion level, non-critical signal line widths can be reduced from 0.1μm to 0.08μm, while critical clock lines remain at 0.15μm.

[0154] Construct a layout quality evaluation function as the objective function for overall optimization. This function is a weighted sum of the ratios of each evaluation metric to the target value. The evaluation metrics include timing margin, power consumption, and routing quality. The timing margin can be weighted to 0.5, power consumption to 0.3, and routing quality to 0.2. Calculate the gradient of this evaluation function with respect to cell position and routing parameters to determine the optimization direction.

[0155] The final optimization direction is obtained by multiplying the gradient of the layout quality assessment function by the constraint matrix. The constraint matrix contains information such as design rule constraints and boundary constraints, ensuring that the optimization results meet physical implementation requirements. Adaptive step size is used for iterative optimization. The initial step size can be set to 10, and it decays exponentially with the number of iterations. For example, the step size for the nth iteration is 10 × 0.95^n.

[0156] The cell positions and routing scheme are updated according to the optimization direction until the termination criteria are met: the improvement in the layout quality assessment function is less than a preset improvement threshold (for example, 0.1%), the worst negative margin is greater than zero, and the total negative margin is greater than zero. At this point, the generated memory compiler IP place-and-route scheme meets both timing closure requirements and power optimization.

[0157] In a real-world application, a 32KB SRAM memory compiler IP was optimized. The initial design power consumption was 15mW, with a worst-case undershoot of -50ps. After optimization using this method, power consumption was reduced to 12.8mW, a reduction of approximately 15%. The worst-case undershoot was improved to +10ps, for a total undershoot of +150ps, meeting timing closure requirements. Routing congestion was reduced from a maximum of 92% to below 75%, while the layout area increased by no more than 3%. The optimization process involved 27 iterations, with a total run time of approximately 45 minutes, achieving a 30% efficiency improvement compared to traditional methods.

[0158] Figure 3 This is a Monte Carlo simulation wiring quality comparison heat map of an embodiment of the present invention:

[0159] This figure uses a heat map to visually demonstrate the comparison results of the routing congestion distribution between this technical solution and the traditional method. The heat map on the left shows the routing congestion distribution of this technical solution, which shows an overall uniform distribution pattern dominated by light green, with an average congestion of only 45.3%, a maximum congestion of 64.7%, and a minimum congestion of 32.5%. The color blocks are distributed relatively evenly, indicating more reasonable utilization of routing resources. The heat map on the right shows the routing congestion distribution of the traditional method, which shows a clear dark red cluster area, with an average congestion of up to 75.6%, a maximum congestion of 94.2%, and a minimum congestion of 62.7%. The central area shows a clear hot spot phenomenon, indicating serious contention for routing resources. Through the intuitive comparison of the two solutions, it can be seen that this technical solution not only significantly reduces the overall routing congestion, but also shows a clear advantage in the uniformity of congestion distribution, effectively avoiding excessive competition for routing resources in local areas, providing a larger design space for subsequent routing optimization and timing convergence, and fully demonstrating the effectiveness of this technical solution in routing optimization.

[0160] According to a second aspect of the embodiments of the present invention,

[0161] Provides memory compiler IP timing closure and power consumption co-optimization system, including:

[0162] The first unit is configured to obtain circuit netlist information of a memory compiler IP, extract a timing path based on the circuit netlist information, perform wiring topology analysis on the timing path, and determine the connection length and interconnection delay between timing cells in the timing path;

[0163] a second unit, configured to divide the timing path into a plurality of power control domains based on the connection length and the interconnection delay by adopting a fine-grained power control strategy, and set a different operating voltage for each of the power control domains, so that the operating voltage of the power control domain close to the timing end point is higher than the operating voltage of the power control domain close to the timing start point;

[0164] A third unit is configured to construct a deep neural network model based on the voltage and delay data of the power consumption control domain, obtain a delay prediction value based on the deep neural network model; dynamically adjust parameters of the deep neural network model according to a prediction error between the delay prediction value and the actual delay value and calculate a reliability index; and use the delay prediction value for timing optimization when the reliability index meets a preset reliability threshold;

[0165] A fourth unit is configured to obtain characteristic parameters of the time series path and construct a characteristic vector matrix based on the characteristic parameters; train a decision tree model based on the characteristic vector matrix to obtain an optimization strategy, and iteratively update the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold;

[0166] The fifth unit is used to calculate the path priority based on the timing path, functionally partition the memory compiler IP based on the path priority and construct a partition cost function; perform cell location optimization and hierarchical routing optimization according to the partition cost function, and generate a memory compiler IP layout and routing solution that meets timing convergence requirements and optimizes power consumption.

[0167] According to a third aspect of the embodiments of the present invention,

[0168] An electronic device is provided, comprising:

[0169] processor;

[0170] a memory for storing processor-executable instructions;

[0171] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0172] According to a fourth aspect of the embodiments of the present invention,

[0173] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0174] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A memory compiler IP timing closure and power consumption collaborative optimization method, characterized in that: include: Obtain circuit netlist information of the memory compiler IP, extract a timing path based on the circuit netlist information, perform wiring topology analysis on the timing path, and determine the connection length and interconnection delay between timing cells in the timing path; Based on the connection length and the interconnection delay, a fine-grained power consumption control strategy is adopted to divide the timing path into multiple power consumption control domains, and a different operating voltage is set for each of the power consumption control domains, so that the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point; A deep neural network model is constructed based on the voltage and delay data of the power consumption control domain, and a delay prediction value is obtained based on the deep neural network model; parameters of the deep neural network model are dynamically adjusted according to a prediction error between the delay prediction value and the actual delay value, and a reliability index is calculated; when the reliability index meets a preset reliability threshold, the delay prediction value is used for timing optimization; Acquire characteristic parameters of the timing path, and construct a characteristic vector matrix based on the characteristic parameters; Training a decision tree model based on the eigenvector matrix to obtain an optimization strategy, and iteratively updating the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold; Calculating path priorities based on the timing paths, functionally partitioning the memory compiler IP based on the path priorities and constructing a partition cost function; Cell location optimization and hierarchical routing optimization are performed according to the partition cost function to generate a memory compiler IP placement and routing solution that meets timing closure requirements and optimizes power consumption.

2. The method according to claim 1, characterized in that Based on the connection length and the interconnection delay, a fine-grained power consumption control strategy is adopted to divide the timing path into multiple power consumption control domains, and a different operating voltage is set for each power consumption control domain, so that the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point. Acquiring physical parameters of the line segment in the timing path, the physical parameters including the length of the line segment, the width of the line segment, the thickness of the interlayer dielectric of the line segment, and the metal specific resistance of the line segment; Calculating the resistance per unit length and the capacitance per unit length of the line segment using a distributed RC model based on the physical parameters of the line segment, and calculating the total resistance and total capacitance of the line segment according to the resistance per unit length and the capacitance per unit length; Calculating the interconnection delay of the connection segment based on the total resistance and total capacitance of the connection segment to obtain the interconnection delay distribution characteristics of the timing path; and dividing adjacent sequential cells with similar delay sensitivities into a plurality of power consumption control domains using a hierarchical clustering algorithm based on the delay sensitivity of the sequential cells in the timing path, wherein the delay sensitivity represents the degree of influence of the operating voltage of the sequential cell on the total delay of the timing path; For the multiple power consumption control domains obtained by division, under the condition of satisfying the timing constraints, the operating voltage of each power consumption control domain is determined by minimizing the dynamic power consumption of each power consumption control domain, wherein the operating voltage of the power consumption control domain close to the timing end point is higher than the operating voltage of the power consumption control domain close to the timing start point.

3. The method according to claim 2, characterized in that Calculating the resistance per unit length and the capacitance per unit length of the line segment using a distributed RC model based on the physical parameters of the line segment, and calculating the total resistance and the total capacitance of the line segment according to the resistance per unit length and the capacitance per unit length includes: The distributed RC model is used to calculate the resistance per unit length of the connecting segment. The calculation process of the resistance per unit length is as follows: dividing the metal specific resistance by the product of the width of the connecting segment and the thickness of the interlayer dielectric; Calculating the capacitance per unit length of the line segment based on the physical parameters of the line segment, wherein the calculation process of the capacitance per unit length includes: calculating the parallel plate capacitance by dividing the product of the vacuum relative dielectric constant system and the width of the line segment by the thickness of the interlayer dielectric; calculating the edge capacitance by multiplying the vacuum relative dielectric constant system by a logarithmic function of the ratio of the interlayer dielectric thickness to the width of the line segment; calculating the coupling capacitance by dividing the product of the vacuum relative dielectric constant system and the thickness of the interlayer dielectric by the metal resistivity; and adding the parallel plate capacitance, the edge capacitance, and the coupling capacitance to obtain the capacitance per unit length; Acquire a signal frequency parameter of the connecting segment, the signal frequency parameter including a signal frequency band and a characteristic frequency parameter, and multiply the signal frequency parameter by the resistance per unit length to obtain a frequency-corrected resistance per unit length; The connecting segment is divided into multiple equal-length sub-segments, and the product of the frequency-corrected unit length resistance of each sub-segment and the sub-segment length is accumulated to obtain the total resistance of the connecting segment, and the product of the unit length capacitance of each sub-segment and the sub-segment length is accumulated to obtain the total capacitance of the connecting segment.

4. The method according to claim 1, wherein Dynamically adjusting the parameters of the deep neural network model and calculating a reliability index based on a prediction error between the delay prediction value and the actual delay value, and using the delay prediction value for timing optimization when the reliability index meets a preset reliability threshold, includes: Constructing a comprehensive loss function, training the deep neural network model using a gradient descent method, calculating weight gradients according to the comprehensive loss function, and updating weight parameters of the neural network based on the weight gradients; Obtaining a new operating voltage combination, inputting the operating voltage combination into the deep neural network model, and obtaining a corresponding delay prediction value; Collecting an actual delay measurement value under the operating voltage combination, and calculating a prediction error between the delay prediction value and the actual delay measurement value; Dynamically adjusting model parameters according to the prediction error, calculating a weight update amount, wherein the weight update amount is proportional to the product of the prediction error and the weight gradient; and calculating a bias update amount, wherein the bias update amount is proportional to the sign value of the prediction error; Calculating a confidence level of a prediction result based on the prediction error, wherein the confidence level decreases as the prediction variance increases, and using the difference between the confidence level and the prediction error as a reliability indicator; When the reliability index is greater than the preset reliability threshold, the current delay prediction value is used to guide timing optimization; when the reliability index is less than the preset reliability threshold, the model online learning process is triggered, new training data is added to the historical data set, and the deep neural network model is retrained.

5. The method according to claim 1, wherein constructing a eigenvector matrix based on the eigenparameters; Training a decision tree model based on the eigenvector matrix to obtain an optimization strategy, and iteratively updating the parameters of the decision tree model through the loss function gradient and sample weight until the optimization target improvement meets the preset condition threshold includes: Constructing a eigenvector matrix based on the characteristic parameters, the eigenvector matrix including a timing eigenvector, a power consumption eigenvector, and an area eigenvector; Calculating the sensitivity of the timing feature vector, the power consumption feature vector, and the area feature vector to voltage, and constructing a sensitivity feature vector, where the sensitivity feature vector includes the sensitivity of timing to voltage, the sensitivity of power consumption to voltage, and the sensitivity of area to voltage; Combining the timing feature vector, the power consumption feature vector, the area feature vector, and the sensitivity feature vector to form an input feature matrix; Building a decision tree model based on the input feature matrix, and outputting a voltage adjustment amount, a size adjustment amount, and a topology adjustment solution through the decision tree model, wherein the voltage adjustment amount is obtained by mapping the sigmoid function, the size adjustment amount is obtained by mapping the tanh function, and the topology adjustment solution is obtained by mapping the softmax function; Constructing an optimization loss function and calculating a loss function gradient based on the optimization loss function; calculating an optimization target improvement and determining a sample weight based on the optimization target improvement, wherein the sample weight increases as the optimization target improvement increases; The parameters of the decision tree model are updated according to the sample weights and the loss function gradient, and the timing path is optimized and adjusted based on the updated decision tree model until the weighted ratio of the optimization target improvement to the optimization overhead meets a preset condition threshold.

6. The method according to claim 1, characterized in that Performing cell location optimization and hierarchical routing optimization based on the partition cost function to generate a memory compiler IP placement and routing solution that meets timing closure requirements and optimizes power consumption includes: Functionally partitioning the memory compiler IP based on the partition cost function; constructing a unit position optimization cost function based on the result of the functional partitioning, and determining the optimization direction of the unit according to the gradient direction of the unit position optimization cost function; Multiplying the optimization direction of the unit by an exponential decay function of a power consumption change to obtain a power consumption-aware unit movement amount, where the power consumption change amount is determined by a difference in power consumption before and after the movement, and adjusting the unit position based on the power consumption-aware unit movement amount; Constructing an inter-layer wiring cost function for the adjusted cell positions and performing wiring layer allocation based on the inter-layer wiring cost function; calculating wiring congestion after the wiring layer allocation, wherein the wiring congestion is a ratio of wiring demand to wiring capacity, and performing power consumption-aware adjustment on the line width of the network based on the wiring congestion; Constructing a layout quality evaluation function, wherein the layout quality evaluation function is a weighted sum of the ratios of each evaluation index to a target value, the evaluation index including a timing margin index, a power consumption index, and a wiring quality index, and calculating the gradient of the layout quality evaluation function; Multiplying the gradient of the layout quality assessment function by the constraint matrix to obtain an optimization direction, and performing iterative optimization using an adaptive step size, wherein the adaptive step size decays exponentially with an increase in the number of iterations; The unit position and routing scheme are updated according to the optimization direction until the improvement of the layout quality evaluation function is less than a preset improvement threshold and the worst negative margin is greater than zero and the total negative margin is greater than zero, thereby generating a memory compiler IP layout and routing scheme that meets timing convergence requirements and optimizes power consumption.

7. A memory compiler IP timing closure and power consumption collaborative optimization system, configured to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is configured to obtain circuit netlist information of a memory compiler IP, extract a timing path based on the circuit netlist information, perform wiring topology analysis on the timing path, and determine the connection length and interconnection delay between timing cells in the timing path; a second unit, configured to divide the timing path into a plurality of power control domains based on the connection length and the interconnection delay by adopting a fine-grained power control strategy, and set a different operating voltage for each of the power control domains, so that the operating voltage of the power control domain close to the timing end point is higher than the operating voltage of the power control domain close to the timing start point; A third unit is configured to construct a deep neural network model based on the voltage and delay data of the power consumption control domain, obtain a delay prediction value based on the deep neural network model; dynamically adjust parameters of the deep neural network model according to a prediction error between the delay prediction value and the actual delay value and calculate a reliability index; and use the delay prediction value for timing optimization when the reliability index meets a preset reliability threshold; A fourth unit is configured to obtain characteristic parameters of the timing path and construct a characteristic vector matrix based on the characteristic parameters; Training a decision tree model based on the eigenvector matrix to obtain an optimization strategy, and iteratively updating the parameters of the decision tree model through the loss function gradient and sample weights until the optimization target improvement meets a preset condition threshold; A fifth unit is configured to calculate a path priority based on the timing path, perform functional partitioning on the memory compiler IP based on the path priority, and construct a partition cost function; Cell location optimization and hierarchical routing optimization are performed according to the partition cost function to generate a memory compiler IP placement and routing solution that meets timing closure requirements and optimizes power consumption.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • EDA software collaborative optimization method based on memory compiler entity IP

    CN119990049A

  • Method and apparatus for layout design, device, medium, and program product

    WO2023123068A1