EDA software collaborative optimization method based on memory compiler entity IP
By establishing a timing netlist model and adjusting the design of the data driver and decoder, and dynamically adjusting the operating voltage and frequency of the memory cell, the timing and power consumption optimization problems in the memory compiler IP are solved, achieving higher performance and reliability.
Patent Information
- Application Number
- CN202510479248.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the existing memory compiler IP design, data drivers with fixed drive strength cannot be dynamically adjusted, resulting in timing violations and waste of power consumption; decoder design lacks systematic optimization, resulting in uneven allocation of timing margins; traditional hotspot management methods are passive and cannot effectively deal with local hotspot problems.
By obtaining the layout information of the memory compiler IP, establishing a timing netlist model, adjusting the driving strength of the data driver and adopting a progressive driving strength allocation strategy, optimizing the hierarchical structure of the decoder, and inserting low-power cells into the memory cell array to dynamically adjust the working voltage and frequency to achieve dynamic balance of local power consumption.
The coordinated optimization of timing and power consumption is achieved, the overall performance and stability of memory compiler IP is improved, power consumption is reduced, chip service life is extended, and adaptability and reliability are improved.
Smart Images

Figure CN119990049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to memory technology, and in particular to an EDA software collaborative optimization method based on a memory compiler entity IP. Background Art
[0002] Memory compiler IP is a key component in integrated circuit design. It can automatically generate memory cells that meet specific specifications according to design requirements. With the continuous development of integrated circuit technology and the increase in design complexity, memory compiler IP has played an increasingly important role in chip design. Memory compiler IP usually contains key components such as memory cell array area, data driver, decoder and clock tree. The performance and power consumption of these parts directly affect the performance and reliability of the entire chip.
[0003] At the current semiconductor process nodes, chip design faces severe power consumption and timing challenges. Electronic design automation (EDA) tools play an indispensable role in the design and optimization of memory compiler IP. Traditional EDA tools usually independently optimize layout, power distribution and timing paths, but with the continuous advancement of technology, this isolated optimization method has been difficult to meet the design requirements of high performance and low power consumption. Especially at advanced process nodes, the impact of temperature on timing path delay is more significant, and the hot spot problem caused by the uneven power distribution is more prominent, which puts higher requirements on the design of memory compiler IP.
[0004] Traditional memory compiler IPs are usually designed with fixed drive strength data drivers, which cannot dynamically adjust the drive capability according to changes in signal transmission distance. This may lead to timing violations during long-distance signal transmission and power consumption waste during short-distance signal transmission, making it difficult to achieve global optimization of timing and power consumption. In addition, the fixed drive strength strategy makes the relationship between signal transmission delay and distance nonlinear, increasing the difficulty of timing convergence.
[0005] Existing decoder designs often use a simple multi-level structure, lacking systematic optimization of the fan-out number of each level. This design method makes it difficult to balance the delay distribution between the decoder levels, resulting in uneven distribution of timing margins on the critical path, which ultimately affects the overall performance. At the same time, the buffering strategy between the data driver and the storage cell array is also relatively simple, failing to fully consider the relationship between the path delay target and the number of buffer levels, making it difficult to achieve optimal timing performance.
[0006] The traditional way of dealing with hot spots in memory design is relatively passive, usually relying on global power management or simple power balancing technology, and lacking a refined dynamic management mechanism for local hot spots. This method cannot effectively deal with local hot spots caused by uneven access patterns in the memory cell array, which can easily cause local overheating and affect chip reliability and life. In addition, the existing technology lacks a method to organically combine power distribution information with layout and routing strategies, making it difficult to achieve global optimization that takes into account both power consumption and performance. Summary of the invention
[0007] The embodiment of the present invention provides an EDA software collaborative optimization method based on a memory compiler entity IP, which can solve the problems in the prior art.
[0008] According to a first aspect of the embodiments of the present invention,
[0009] Provides EDA software co-optimization methods based on memory compiler entity IP, including:
[0010] Acquire layout information of the memory compiler IP, wherein the layout information includes location information of a memory cell array region, location information of a data driver, location information of a decoder, and location information of a clock tree;
[0011] Based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence of temperature on the timing path delay value, a timing netlist model of the memory compiler IP is established, wherein the timing netlist model includes path timing information and power consumption distribution information of the memory cell array;
[0012] According to the path timing information in the timing netlist model, the driving strength of the data driver is adjusted and a progressive driving strength allocation strategy is adopted so that the signal transmission delay increases linearly with the distance to achieve timing convergence; the hierarchical structure of the decoder is adjusted, the fan-out number of each layer is determined by optimizing the decoder tree structure, and a multi-level buffer is set between the data driver and the storage cell array area in combination with the target delay of each layer;
[0013] Based on the power consumption distribution information of the memory cell array, determining the hot spot distribution of the memory cell array area, inserting a low power consumption unit in the hot spot distribution area, wherein the low power consumption unit realizes a dynamic balance of local power consumption by dynamically adjusting the operating voltage and operating frequency of the memory cell;
[0014] The location information of the multi-level buffer and the location information of the low power consumption unit are updated into the layout information, and the updated layout information is optimized for wiring to generate an optimized memory compiler IP layout.
[0015] Based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence of temperature on the timing path delay value, establishing the timing netlist model of the memory compiler IP includes:
[0016] Extracting parasitic parameters of interconnects in the memory compiler IP based on the layout information, the parasitic parameters including resistance parameters and capacitance parameters of the interconnects, wherein the resistance parameters are calculated based on the resistivity, length, width and thickness of the wires, and the capacitance parameters include parallel plate capacitance, edge capacitance and coupling capacitance;
[0017] Determine a delay value of a timing path based on the parasitic parameters, wherein the delay value is obtained by calculating the sum of the products of the delay coefficients of each segment on the timing path and the corresponding equivalent resistance and equivalent capacitance;
[0018] Calculate the timing margin of the timing path according to the delay value, wherein the timing margin is obtained by subtracting the actual arrival time and the timing uncertainty from the required time of the timing path;
[0019] Calculating the dynamic power consumption and static power consumption of the memory compiler IP, wherein the dynamic power consumption is calculated according to the relationship between load capacitance, operating voltage and operating frequency, and the static power consumption is calculated according to the relationship between operating voltage and leakage current;
[0020] Determine the influence of the dynamic power consumption and the static power consumption on the timing, and obtain the influence of the temperature on the delay value of the timing path by calculating the product of the temperature coefficient and the difference between the actual temperature and the nominal temperature;
[0021] Based on the influence of the temperature on the delay value, the delay value of the timing path is corrected by multiplying the temperature-delay conversion coefficient with the power consumption density and the thermal resistance to generate a complete timing netlist model containing timing information and power consumption distribution characteristics.
[0022] According to the path timing information in the timing netlist model, the driving strength of the data driver and the hierarchical structure of the decoder are adjusted, by setting a multi-level buffer between the data driver and the storage cell array area, and adopting a progressive driving strength allocation strategy, including:
[0023] Calculating the load capacitance, logic delay and interconnection delay of the memory compiler, and determining the path timing margin from the clock cycle according to the logic delay and the interconnection delay; setting the voltage swing and the conversion time based on the load capacitance, and determining the driving requirement reference value according to the load capacitance, the voltage swing and the conversion time;
[0024] Determine a minimum driver width of the data driver according to the driving requirement reference value, and multiply the minimum driver width by the square root of the ratio of the load capacitance to the reference capacitance to obtain an actual driver width of the data driver; determine a driver reference strength based on the actual driver width, establish a driving strength increasing function by setting an increasing coefficient, and calculate a theoretical distribution value of multi-level driving strength according to the driving strength increasing function;
[0025] Acquire the distance information from the data driver to the storage unit, perform an exponential function operation on the distance information and a preset characteristic distance to obtain a distance weight coefficient; modify the theoretical distribution value of the multi-level driving strength according to the distance weight coefficient, and determine the actual driving strength of each level by calculating the ratio of the modified driving strength to the total driving strength;
[0026] The path timing margin is updated based on the actual driving strength, and when the path timing margin meets the timing requirement, a progressive driving strength configuration scheme of the data driver is generated, and the progressive driving strength configuration scheme is applied to the memory compiler.
[0027] Adjusting the hierarchical structure of the decoder, determining the fan-out number of each layer by optimizing the decoder tree structure, and setting a multi-level buffer between the data driver and the storage cell array area in combination with the target delay of each layer includes:
[0028] Obtain buffer load capacitance parameters, the load capacitance parameters including single-stage input capacitance, wiring capacitance and current-stage load capacitance, and add the load capacitance parameters to obtain a total load capacitance; calculate a theoretical optimal number of stages based on a logarithmic ratio of the total load capacitance to the single-stage input capacitance, and correct the theoretical optimal number of stages according to a preset margin coefficient to obtain an actual number of buffer stages;
[0029] Obtaining total layout length information, distributing the total layout length to buffers at each level according to the distance attenuation coefficient based on the actual number of buffer levels, and obtaining initial spacings of buffers at each level; calculating the tree structure hierarchy depth according to the total number of decoders and a preset fan-out coefficient, and evenly distributing the total number of decodes based on the hierarchy depth to obtain the fan-out number of each layer;
[0030] Obtain a total delay constraint of the decoder, determine a layer weight coefficient according to the fan-out number of each layer, distribute the total delay constraint to each layer according to the layer weight coefficient, and obtain a target delay of each layer;
[0031] Calculating the actual delay corresponding to the initial spacing and the fan-out number of each layer, and taking the difference between the actual delay and the target delay as the delay compensation amount; adjusting the initial spacing and the fan-out number of each layer according to the delay compensation amount, and calculating the adjusted total delay of the signal transmission path;
[0032] The power consumption density distribution information of each physical position is collected, and the actual spacing is locally fine-tuned based on the power consumption density distribution information until the total delay of the signal transmission path meets the timing requirements, thereby generating a final decoder hierarchical structure design scheme.
[0033] The tree structure hierarchy depth is calculated according to the total number of decoders and the preset fan-out coefficient, and the total number of decodes is evenly distributed based on the hierarchy depth to obtain the fan-out number of each layer, including:
[0034] Obtaining a total number of decodes and a fan-out coefficient of the decoder, dividing the total number of decodes by the fan-out coefficient to perform a logarithmic operation to obtain an initial level value, and rounding up the initial level value to obtain an actual level depth of the tree structure;
[0035] Performing a square root operation of the actual layer depth on the total number of decodes to obtain a uniform inter-layer fan-out reference value, and calculating an ideal fan-out number of each layer according to the uniform inter-layer fan-out reference value;
[0036] Comparing the ideal fan-out number with the fan-out coefficient, and when the ideal fan-out number is greater than the fan-out coefficient, increasing the actual layer depth and recalculating the ideal fan-out number until the ideal fan-out number is less than or equal to the fan-out coefficient;
[0037] A tree structure of a decoder is constructed based on the actual layer depth and the ideal fan-out number that are finally determined, and it is verified whether the number of decoding outputs of the tree structure is equal to the total number of decodings, so as to generate a balanced hierarchical structure design scheme that meets the decoding requirements.
[0038] Based on the power consumption distribution information of the memory cell array, determining the hotspot distribution of the memory cell array region includes:
[0039] Obtaining the operating voltage, leakage current and cell density of each memory cell in the memory cell array, calculating the static power consumption density distribution, and collecting the load capacitance, flip activity factor and operating frequency of the memory cell; calculating the dynamic power consumption density distribution according to the operating voltage, load capacitance, flip activity factor and operating frequency, and superimposing the static power consumption density distribution with the dynamic power consumption density distribution to obtain the total power consumption density distribution;
[0040] Based on the ambient temperature and thermal resistance coefficient of the storage cell array, the actual temperature distribution of the storage cell array is calculated; the actual temperature distribution is analyzed according to a preset temperature threshold, the position coordinates of the hot spot area exceeding the preset temperature threshold are determined, and the temperature gradient at the position coordinates of the hot spot area is calculated.
[0041] Inserting a low-power unit in the hot spot distribution area, wherein the low-power unit dynamically adjusts the operating voltage and operating frequency of the storage unit to achieve dynamic balance of local power consumption includes:
[0042] Calculating a voltage regulation coefficient based on the temperature gradient, multiplying the nominal operating voltage by the voltage regulation coefficient to obtain an optimized operating voltage, and determining an optimized operating frequency according to a ratio of the optimized operating voltage to the nominal operating voltage;
[0043] Calculating a power consumption balance factor of a storage cell array, the power consumption balance factor being a ratio of an average power consumption density of the storage cell array to the total power consumption density distribution, and inserting a low power consumption unit at the position coordinates of the hot spot area according to the power consumption balance factor;
[0044] The low power consumption unit receives the power consumption balance factor, and adjusts the optimized operating voltage and the optimized operating frequency of the storage unit at the corresponding position in real time according to the power consumption balance factor.
[0045] A second aspect of an embodiment of the present invention provides an EDA software collaborative optimization system based on a memory compiler entity IP, comprising:
[0046] A first unit is used to obtain layout information of a memory compiler IP, wherein the layout information includes location information of a memory cell array area, location information of a data driver, location information of a decoder, and location information of a clock tree;
[0047] A second unit is used to establish a timing netlist model of the memory compiler IP based on the layout information, in combination with the dynamic power consumption and static power consumption of the memory compiler IP, and taking into account the influence of temperature on the timing path delay value, wherein the timing netlist model includes path timing information and power consumption distribution information of the memory cell array;
[0048] The third unit is used to adjust the driving strength of the data driver and adopt a progressive driving strength allocation strategy according to the path timing information in the timing netlist model, so that the signal transmission delay increases linearly with the distance to achieve timing convergence; adjust the hierarchical structure of the decoder, determine the fan-out number of each layer through the decoder tree structure optimization, and set a multi-level buffer between the data driver and the storage cell array area in combination with the target delay of each layer;
[0049] A fourth unit is used to determine the hot spot distribution of the storage cell array area based on the power consumption distribution information of the storage cell array, insert a low power consumption unit in the hot spot distribution area, and the low power consumption unit realizes the dynamic balance of local power consumption by dynamically adjusting the operating voltage and operating frequency of the storage cell;
[0050] The fifth unit is used to update the location information of the multi-level buffer and the location information of the low power consumption unit into the layout information, and perform wiring optimization on the updated layout information to generate an optimized memory compiler IP layout.
[0051] A third aspect of the embodiments of the present invention
[0052] An electronic device is provided, comprising:
[0053] processor;
[0054] a memory for storing processor-executable instructions;
[0055] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0056] A fourth aspect of the embodiments of the present invention is:
[0057] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0058] The beneficial effects of this application are as follows: The present invention obtains the layout information of the memory compiler IP and establishes an accurate timing netlist model to achieve the coordinated optimization of timing and power consumption. The method can accurately reflect the impact of temperature on the delay value of the timing path, providing a reliable basis for subsequent optimization, thereby improving the overall performance and stability of the memory compiler IP.
[0061] The present invention adopts a progressive drive strength allocation strategy and decoder hierarchical structure optimization to make the signal transmission delay increase linearly with distance, effectively solving the problem of difficult timing convergence in traditional methods. By setting up a multi-level buffer between the data driver and the memory cell array area, the signal transmission path is further optimized, significantly improving the timing performance and working reliability of the memory compiler IP.
[0062] The present invention identifies hotspot distribution areas based on power consumption distribution information and inserts low-power units, and achieves dynamic balance of local power consumption by dynamically adjusting the operating voltage and frequency of the storage unit. This method effectively reduces the overall power consumption of the memory compiler IP, improves energy utilization efficiency, and reduces the negative impact of hotspot areas on surrounding circuits, prolongs the service life of the chip, and greatly improves the adaptability and reliability of the memory compiler IP in various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1It is a flowchart of an EDA software collaborative optimization method based on a memory compiler entity IP according to an embodiment of the present invention;
[0064] Figure 2 A data table comparing the parasitic parameter extraction accuracy of different methods in the embodiments of the present invention;
[0065] Figure 3 A schematic diagram of the relationship between driving strength and timing margin according to an embodiment of the present invention;
[0066] Figure 4 A schematic diagram of a tree structure balanced hierarchical design system for a decoder according to an embodiment of the present invention;
[0067] Figure 5 This is a comparison data table of power consumption balance factors and temperature hotspot distribution according to an embodiment of the present invention. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0069] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0070] Figure 1 FIG. 1 is a flow chart of an EDA software collaborative optimization method based on a memory compiler entity IP according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0071] Acquire layout information of the memory compiler IP, wherein the layout information includes location information of a memory cell array region, location information of a data driver, location information of a decoder, and location information of a clock tree;
[0072] Based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence of temperature on the timing path delay value, a timing netlist model of the memory compiler IP is established, wherein the timing netlist model includes path timing information and power consumption distribution information of the memory cell array;
[0073] According to the path timing information in the timing netlist model, the driving strength of the data driver is adjusted and a progressive driving strength allocation strategy is adopted so that the signal transmission delay increases linearly with the distance to achieve timing convergence; the hierarchical structure of the decoder is adjusted, the fan-out number of each layer is determined by optimizing the decoder tree structure, and a multi-level buffer is set between the data driver and the storage cell array area in combination with the target delay of each layer;
[0074] Based on the power consumption distribution information of the memory cell array, determining the hot spot distribution of the memory cell array area, inserting a low power consumption unit in the hot spot distribution area, wherein the low power consumption unit realizes a dynamic balance of local power consumption by dynamically adjusting the operating voltage and operating frequency of the memory cell;
[0075] The location information of the multi-level buffer and the location information of the low power consumption unit are updated into the layout information, and the updated layout information is optimized for wiring to generate an optimized memory compiler IP layout.
[0076] In an optional implementation, based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence of temperature on the timing path delay value, establishing the timing netlist model of the memory compiler IP includes:
[0077] Extracting parasitic parameters of interconnects in the memory compiler IP based on the layout information, the parasitic parameters including resistance parameters and capacitance parameters of the interconnects, wherein the resistance parameters are calculated based on the resistivity, length, width and thickness of the wires, and the capacitance parameters include parallel plate capacitance, edge capacitance and coupling capacitance;
[0078] Determine a delay value of a timing path based on the parasitic parameters, wherein the delay value is obtained by calculating the sum of the products of the delay coefficients of each segment on the timing path and the corresponding equivalent resistance and equivalent capacitance;
[0079] Calculate the timing margin of the timing path according to the delay value, wherein the timing margin is obtained by subtracting the actual arrival time and the timing uncertainty from the required time of the timing path;
[0080] Calculating the dynamic power consumption and static power consumption of the memory compiler IP, wherein the dynamic power consumption is calculated according to the relationship between load capacitance, operating voltage and operating frequency, and the static power consumption is calculated according to the relationship between operating voltage and leakage current;
[0081] Determine the influence of the dynamic power consumption and the static power consumption on the timing, and obtain the influence of the temperature on the delay value of the timing path by calculating the product of the temperature coefficient and the difference between the actual temperature and the nominal temperature;
[0082] Based on the influence of the temperature on the delay value, the delay value of the timing path is corrected by multiplying the temperature-delay conversion coefficient with the power consumption density and the thermal resistance to generate a complete timing netlist model containing timing information and power consumption distribution characteristics.
[0083] Extract the parasitic parameters of the interconnect based on the layout information of the memory compiler IP. The parasitic parameters include the resistance parameters and capacitance parameters of the interconnect. The calculation of the resistance parameters is based on the resistivity, length, width and thickness of the wire. For example, for a copper wire with a length of 100 microns, a width of 0.5 microns and a thickness of 0.2 microns, its resistivity is 1.68 micro-ohms cm, and its resistance value can be calculated to be 6.72 ohms. The capacitance parameters include the parallel plate capacitance value, the edge capacitance value and the coupling capacitance value. For two adjacent parallel wires, the parallel plate capacitance can be calculated by dividing the overlapping area of the wire by the thickness of the insulating layer and multiplying by the dielectric constant. For example, if the overlapping area of the two wires is 10 square microns, the thickness of the insulating layer is 0.1 microns, and the dielectric constant is 3.9, the parallel plate capacitance is about 3.45 femtofarads. The edge capacitance depends on the edge shape of the wire and the surrounding environment, and is usually obtained by looking up a table or a dedicated extraction tool. The coupling capacitance reflects the electromagnetic coupling effect between adjacent wires, and its value is related to the wire spacing and parallel length.
[0084] Determine the delay value of the timing path based on the extracted parasitic parameters. The delay value of the timing path is obtained by calculating the sum of the products of the delay coefficient of each segment on the path and the corresponding equivalent resistance and equivalent capacitance. Specifically, for each segment on the timing path, first determine the delay coefficient of its driving unit, then calculate the equivalent resistance and equivalent capacitance, and finally take the product of the three as the delay value of the segment. For example, for an inverter with a drive strength of 4, its delay coefficient is 0.15 nanoseconds / picofarad ohm, the load equivalent resistance is 100 ohms, and the equivalent capacitance is 20 picofarads, then the delay value of the segment is 0.3 nanoseconds. Add the delay values of all segments on the timing path to obtain the total delay value of the entire path.
[0085] The timing margin of the timing path is calculated based on the calculated delay value. The timing margin is obtained by subtracting the actual arrival time and timing uncertainty from the required time of the timing path. For example, if the required time of a timing path is 2 nanoseconds, the actual arrival time is 1.7 nanoseconds, and the timing uncertainty is 0.2 nanoseconds, then the timing margin of the path is 0.1 nanoseconds. Timing uncertainty includes factors such as clock skew, clock jitter, and delay changes caused by process changes.
[0086] Calculate the dynamic and static power consumption of the memory compiler IP. Dynamic power consumption is calculated based on the relationship between load capacitance, operating voltage, and operating frequency. For example, for a load capacitance of 10 pF, an operating voltage of 1.2 volts, and an operating frequency of 500 MHz, the dynamic power consumption is approximately 3.6 milliwatts. Static power consumption is calculated based on the relationship between operating voltage and leakage current, for example, if the leakage current is 100 microamperes and the operating voltage is 1.2 volts, the static power consumption is 0.12 milliwatts. The total power consumption is the sum of dynamic power consumption and static power consumption, which is 3.72 milliwatts in this case.
[0087] Determine the impact of dynamic and static power consumption on timing. This is done by calculating the temperature coefficient multiplied by the difference between the actual temperature and the nominal temperature to obtain the degree of temperature impact on the delay value of the timing path. For example, if the temperature coefficient is 0.2% / degree Celsius, the actual temperature is 85 degrees Celsius, and the nominal temperature is 25 degrees Celsius, the delay caused by temperature increases by 12%. For a path with an original delay of 1 nanosecond, the delay will increase to 1.12 nanoseconds at high temperatures.
[0088] Based on the degree of influence of temperature on the delay value, the delay value of the timing path is corrected by multiplying the temperature-delay conversion coefficient with the power density and thermal resistance to generate a complete timing netlist model containing timing information and power distribution characteristics. The temperature-delay conversion coefficient is usually provided by the process library. For example, the temperature-delay conversion coefficient of a 65-nanometer process is 0.15 nanoseconds / watt·square millimeter·degree Celsius. The power density can be obtained by dividing the total power consumption by the chip area. For example, if the total power consumption is 10 milliwatts and the chip area is 0.5 square millimeters, the power density is 20 watts / square millimeter. Thermal resistance represents the temperature rise caused by unit power and is usually related to the package type. For example, the thermal resistance of a BGA package is 20 degrees Celsius / watt. Multiply these parameters to obtain the delay increase value, which is used to correct the original delay value. For example, if the original delay value is 1 nanosecond and the delay increase value is 0.06 nanoseconds, the corrected delay value is 1.06 nanoseconds.
[0089] Through the above steps, a complete timing netlist model containing timing information and power distribution characteristics is finally generated. This model not only takes into account the layout information and interconnect parasitic parameters, but also covers the impact of power consumption on temperature and the impact of temperature on timing, thereby providing more accurate timing analysis results. Using this model, the performance of the memory compiler IP can be predicted in the early stages of the design, assisting designers in optimizing the memory architecture and layout, shortening the design cycle, and improving the reliability and performance of the memory IP.
[0090] Figure 2 The following is a comparison table of parasitic parameter extraction accuracy of different methods in the embodiments of the present invention:
[0091] This picture is a comparison table of interconnect parameters, showing the electrical parameter measurement results of interconnects (M1-M5) of different lengths (1μm to 200μm). The table includes two key parameters: resistance (Ω) and capacitance (fF), and lists the actual measured value, technical solution value, and comparison data and error rate with NE and NF processes for each parameter.
[0092] From the data, as the length of the interconnect increases, the resistance and capacitance values show an upward trend. Taking the short line M1 (L=1μm) as an example, its measured resistance is 8.75Ω and its capacitance is 0.95fF; when it comes to the longest line M5 (L=200μm), the resistance increases to 67.80Ω and the capacitance reaches 36.40fF. The error rate of each parameter generally increases with the increase of line length. The error rate of the short line M1 is relatively small (resistance 0.8%, capacitance 2.1%), while the error rate of the long line M5 is larger (resistance 2.2%, capacitance 2.6%).
[0093] Compared with the existing technologies NE and NF, the parameters of this solution are generally better. Taking the M3 long line (L=50μm) as an example, its measured resistance is 105.30Ω, while NE and NF are 111.62Ω and 114.78Ω respectively; the measured capacitance is 12.75fF, while NE and NF are 13.71fF and 14.23fF respectively. This shows that this technical solution has obvious advantages in controlling interconnection line parameters, especially in reducing parasitic resistance and capacitance.
[0094] The present application has made innovative optimizations by improving the preparation process of interconnects, especially in metal layer deposition, patterning and etching processes. In the prior art, the traditional preparation process of interconnects mainly relies on conventional physical vapor deposition (PVD) and chemical vapor deposition (CVD) processes. This method is prone to problems such as uneven metal grain size and insufficient line width control accuracy during the preparation process, resulting in large fluctuations in the resistance and capacitance parameters of the interconnects, especially in long line structures.
[0095] The method optimizes the deposition conditions of the metal layer and adopts a multi-step deposition process to effectively control the growth process of the metal grains; at the same time, by improving the photolithography and etching process parameters, the patterning accuracy and the control ability of the sidewall profile are improved. In addition, the application also introduces a new surface treatment process to effectively reduce the interface effect between the metal layer and the dielectric layer.
[0096] The preparation accuracy of interconnection lines is significantly improved, so that the deviation between the actual parameters of each layer of interconnection lines and the design values is significantly reduced; the conductivity of the interconnection lines is improved, and the line resistance is reduced overall; by optimizing the interface characteristics, the parasitic capacitance is effectively reduced; the technical solution of this application shows more obvious advantages in long-line structures, and provides a better process foundation for the design and manufacture of high-performance integrated circuits. Compared with the prior art, the solution of this application has made significant progress in the stability and consistency of interconnection line parameter control, especially in reducing parasitic effects, which is of great significance to improving the overall performance and reliability of integrated circuits.
[0097] In an optional implementation, according to the path timing information in the timing netlist model, the driving strength of the data driver and the hierarchical structure of the decoder are adjusted, by setting a multi-level buffer between the data driver and the storage cell array area, and adopting a progressive driving strength allocation strategy including:
[0098] Calculating the load capacitance, logic delay and interconnection delay of the memory compiler, and determining the path timing margin from the clock cycle according to the logic delay and the interconnection delay; setting the voltage swing and the conversion time based on the load capacitance, and determining the driving requirement reference value according to the load capacitance, the voltage swing and the conversion time;
[0099] Determine a minimum driver width of the data driver according to the driving requirement reference value, and multiply the minimum driver width by the square root of the ratio of the load capacitance to the reference capacitance to obtain an actual driver width of the data driver; determine a driver reference strength based on the actual driver width, establish a driving strength increasing function by setting an increasing coefficient, and calculate a theoretical distribution value of multi-level driving strength according to the driving strength increasing function;
[0100] Acquire the distance information from the data driver to the storage unit, perform an exponential function operation on the distance information and a preset characteristic distance to obtain a distance weight coefficient; modify the theoretical distribution value of the multi-level driving strength according to the distance weight coefficient, and determine the actual driving strength of each level by calculating the ratio of the modified driving strength to the total driving strength;
[0101] The path timing margin is updated based on the actual driving strength, and when the path timing margin meets the timing requirement, a progressive driving strength configuration scheme of the data driver is generated, and the progressive driving strength configuration scheme is applied to the memory compiler.
[0102] According to the path timing information in the memory timing netlist model, the load capacitance, logic delay and interconnection delay are calculated. For a typical 256KB SRAM memory, its load capacitance is usually in the range of 10-20pF. Path timing analysis includes the entire signal transmission path from the driver to the memory cell. For example, in a system running at 1GHz, the clock cycle is 1ns, if the logic delay is 0.3ns and the interconnection delay is 0.4ns, the available path timing margin is 0.3ns.
[0103] Based on the calculated load capacitance value, set the appropriate voltage swing and transition time. In the technology with a process node of 28nm, the typical voltage swing value can be set to 0.9V and the transition time target can be set to 150ps. Based on these parameters, the reference drive requirement value can be determined. For example, for a load capacitance of 15pF, a voltage swing of 0.9V and a transition time of 150ps, the reference drive requirement value may be 90μA / V.
[0104] , the minimum driver width of the data driver is determined based on the driving requirement reference value. Assuming that the minimum driver unit width is 120nm under the process used and provides a driving capability of 4μA / V, the minimum number of drivers required is the driving requirement reference value divided by the driving capability of a single driver, that is, 90μA / V÷4μA / V=22.5, rounded to 23 units.
[0105] The actual driver width of the data driver is obtained by multiplying the minimum driver width by the square root of the ratio of the load capacitance to the reference capacitance. Assuming the reference capacitance is 1pF and the load capacitance is 15pF, the square root of the ratio is √(15 / 1)=3.87. The actual driver width is calculated as 2760nm×3.87=10681.2nm, rounded to 10680nm.
[0106] The driver base strength is determined based on the actual driver width. The driver base strength is set to the driving capability corresponding to the actual width, that is, 10680nm / 120nm×4μA / V=356μA / V. Then the increasing coefficient is set to establish the driving strength increasing function. Assuming that an exponential increasing strategy is adopted and the increasing coefficient is set to 1.5, for a 4-level driver configuration, the theoretical allocation value can be calculated as follows: the first level is 1 times the driver base strength, the second level is 1.5 times, the third level is 1.5² times, and the fourth level is 1.5³ times.
[0107] The distance information from the data driver to the storage unit is obtained, and the distance information is subjected to an exponential function operation with the preset characteristic distance to obtain a distance weight coefficient. For example, assuming that the distance from the driver to the storage unit is 800 μm and the preset characteristic distance is 200 μm, the distance weight coefficient can be expressed as e^(800 / 200)=54.6.
[0108] The theoretical allocation values of multi-level drive strength are corrected according to the distance weight coefficient. The correction method is to multiply the theoretical allocation value by an adjustment factor related to the distance weight coefficient. For example, for long distances, it may be necessary to increase the strength of the subsequent driver. Assume that the adjusted drive strength allocation ratio is: 10% for the first level, 15% for the second level, 25% for the third level, and 50% for the fourth level. The corresponding actual drive strength is: 35.6μA / V for the first level, 53.4μA / V for the second level, 89μA / V for the third level, and 178μA / V for the fourth level.
[0109] Update the path timing margin based on the actual drive strength. Perform timing analysis on the modified driver configuration to calculate the new signal propagation delay. For example, if the original path delay is 0.7ns and the optimized configuration reduces the delay to 0.65ns, the path timing margin increases by 0.05ns, from the original 0.3ns to 0.35ns.
[0110] When the path timing margin meets the timing requirements, a progressive drive strength configuration scheme for the data driver is generated. The configuration scheme includes the specific specifications of each level of driver, such as transistor width, drive strength, etc. For example, for the four-level driver calculated above, its configuration can be expressed as: the first level driver width is 1068nm, the second level is 1602nm, the third level is 2670nm, and the fourth level is 5340nm.
[0111] The progressive drive strength configuration scheme is applied to the memory compiler. The memory compiler automatically generates the corresponding layout and circuit design according to the configuration scheme. In practical applications, for different specifications of memory, the optimal driver configuration can be automatically calculated according to its specific parameters. For example, for a 1MB SRAM, the load capacitance may increase to 30pF. At this time, the above method can be used to calculate a stronger drive strength requirement and a more detailed hierarchical structure.
[0112] The above method can significantly improve the quality of data drive signals, reduce signal reflection and noise, and improve the read and write speed and reliability of memory. In an actual test case, compared with the traditional fixed drive strength configuration, the use of a progressive drive strength allocation strategy reduced the memory access time by 15%, power consumption by 12%, and signal integrity was significantly improved, with the error rate reduced by more than 20%.
[0113] This method is also highly adaptable and can automatically adjust the driver configuration according to different process nodes, different memory specifications, and different timing requirements, improving the versatility and efficiency of the memory compiler. For example, when the process node shrinks from 28nm to 16nm, by re-evaluating the load capacitance and timing requirements, the drive strength allocation strategy can be adjusted accordingly to maintain optimal performance.
[0114] Figure 3 A schematic diagram of the relationship between driving strength and timing margin according to an embodiment of the present invention is shown below:
[0115] This picture shows the response time comparison curve of three configuration schemes under different drive strength coefficients. The horizontal axis of the figure represents the drive strength coefficient (range 1-10), and the vertical axis represents the response time (unit ps). The three schemes are represented by different marks: this technology scheme (black triangle), traditional linear allocation scheme (grey dots) and uniform allocation scheme (grey squares).
[0116] Judging from the curve trend, as the drive strength coefficient increases, the response time of the three solutions all show the characteristics of first rising rapidly and then gradually flattening. Specifically, when the drive strength coefficient is 1, the response time of the present technical solution is about 45ps, the traditional solution is about 32ps, and the uniform distribution solution is about 25ps. When the drive strength coefficient increases to 5, the present technical solution rises to about 120ps, the traditional solution reaches about 82ps, and the uniform distribution solution is about 72ps. When the drive strength coefficient reaches 10, the response time of the three solutions stabilizes at about 140ps, 95ps, and 82ps, respectively.
[0117] From the overall performance point of view, the response time of this technical solution at each drive strength coefficient is significantly higher than that of the other two solutions, and the growth curve is steeper, indicating that it is more sensitive to changes in drive strength. In contrast, the response time of the uniform distribution solution is always the lowest, and the curve is relatively flat, showing good stability. The performance of the traditional linear distribution solution is between the two.
[0118] In an optional implementation, adjusting the hierarchical structure of the decoder, determining the fan-out number of each layer by optimizing the decoder tree structure, and setting a multi-level buffer between the data driver and the storage cell array area in combination with the target delay of each layer includes:
[0119] Obtain buffer load capacitance parameters, the load capacitance parameters including single-stage input capacitance, wiring capacitance and current-stage load capacitance, and add the load capacitance parameters to obtain a total load capacitance; calculate a theoretical optimal number of stages based on a logarithmic ratio of the total load capacitance to the single-stage input capacitance, and correct the theoretical optimal number of stages according to a preset margin coefficient to obtain an actual number of buffer stages;
[0120] Obtaining total layout length information, distributing the total layout length to buffers at each level according to the distance attenuation coefficient based on the actual number of buffer levels, and obtaining initial spacings of buffers at each level; calculating the tree structure hierarchy depth according to the total number of decoders and a preset fan-out coefficient, and evenly distributing the total number of decodes based on the hierarchy depth to obtain the fan-out number of each layer;
[0121] Obtain a total delay constraint of the decoder, determine a layer weight coefficient according to the fan-out number of each layer, distribute the total delay constraint to each layer according to the layer weight coefficient, and obtain a target delay of each layer;
[0122] Calculating the actual delay corresponding to the initial spacing and the fan-out number of each layer, and taking the difference between the actual delay and the target delay as the delay compensation amount; adjusting the initial spacing and the fan-out number of each layer according to the delay compensation amount, and calculating the adjusted total delay of the signal transmission path;
[0123] The power consumption density distribution information of each physical position is collected, and the actual spacing is locally fine-tuned based on the power consumption density distribution information until the total delay of the signal transmission path meets the timing requirements, thereby generating a final decoder hierarchical structure design scheme.
[0124] Get the buffer load capacitance parameters, including single-stage input capacitance, wiring capacitance, and current-stage load capacitance. For example, in a 64KB SRAM memory cell design, the single-stage input capacitance is 5fF, the wiring capacitance is 0.2fF / μm, and the current-stage load capacitance is 50fF. Add these parameters to get the total load capacitance, which is 55.2fF in this example.
[0125] The theoretical optimal number of stages is calculated based on the logarithmic ratio of the total load capacitance to the single-stage input capacitance. Specifically, by calculating the logarithmic ratio of 55.2fF to 5fF, the theoretical optimal number of stages is about 2.7. Considering the stability of the actual circuit design, a preset margin factor of 1.2 is introduced to correct the theoretical optimal number of stages, and the actual number of buffer stages is 3.
[0126] Get the total layout length information. In this example, the total layout length between the data driver and the memory cell array area is 500μm. According to the actual number of buffer levels (3 levels), the total layout length is distributed to each level of buffer according to the distance attenuation coefficient. The distance attenuation coefficient can be set to {0.25, 0.35, 0.4}, corresponding to the first to third level buffers, respectively. Therefore, the initial spacing of each level of buffers is calculated to be 125μm, 175μm, and 200μm, respectively.
[0127] In the decoder tree structure design, the tree structure level depth is calculated according to the total number of decoders and the preset fan-out coefficient. Assuming the total number of decoders is 256 and the preset fan-out coefficient is 4, the tree structure level depth is log 4 (256) = 4 layers. Based on this layer depth, the total number of decodes is evenly distributed to obtain the fan-out number of each layer. In this example, the fan-out number of each layer is 4 after even distribution.
[0128] Get the total delay constraint of the decoder, for example, 500ps in this design. Determine the layer weight coefficient according to the fan-out number of each layer. Generally, the larger the fan-out number, the higher the weight coefficient. In this example, since the fan-out number of each layer is 4, the same weight coefficient of 0.25 can be set. Distribute the total delay constraint to each layer according to the layer weight coefficient, and obtain a target delay of 125ps for each layer.
[0129] Calculate the actual delay corresponding to the initial spacing and the fan-out number of each layer. Assume that through circuit simulation, the actual delay of the first layer is 140ps, the second layer is 130ps, the third layer is 120ps, and the fourth layer is 110ps, totaling 500ps. The difference between the actual delay and the target delay is used as the delay compensation amount, which is +15ps for the first layer, +5ps for the second layer, -5ps for the third layer, and -15ps for the fourth layer.
[0130] Adjust the initial spacing and fan-out number of each layer according to the delay compensation amount. For example, reduce the fan-out number of the first layer to 3 and increase it to 5 for the fourth layer. At the same time, adjust the spacing of each level of buffers, reduce the first level to 115μm, and increase the third level to 210μm. Recalculate the total delay of the adjusted signal transmission path, and perform iterative optimization until the total delay meets the constraints.
[0131] Collect power density distribution information at each physical location. In this case, the data shows that the power density is higher in the area close to the memory cell array, about 10mW / μm², while it is 5mW / μm² in the distant area. Based on the power density distribution information, the actual spacing is locally fine-tuned. For example, the buffer spacing is appropriately increased by 5% in areas with high power density to reduce local hot spots. After fine-tuning, the final buffer spacing at each level is 115μm, 175μm, and 220μm, respectively. The fan-out numbers of each layer are 3, 4, 4, and 5, respectively. The total delay of the signal transmission path is 485ps, which meets the timing constraint of 500ps.
[0132] Each parameter can be flexibly adjusted according to actual design requirements. For example, for high-performance memory design, the number of buffer levels can be appropriately increased and the number of fan-outs at each level can be reduced to further optimize latency performance. For low-power design, the number of buffer levels can be appropriately reduced and the number of fan-outs at each level can be increased to reduce overall power consumption.
[0133] In the physical implementation stage, the decoder hierarchical structure can be further optimized by combining the timing-driven layout technology. By accurately laying out the critical path, the wiring delay can be reduced and the decoding speed can be improved. At the same time, dynamic power management technology can be used to adaptively adjust the drive strength of each level of the decoder buffer according to different working modes to achieve a balance between power consumption and performance.
[0134] Through the above method, the decoder hierarchical structure design scheme generated can effectively balance multi-dimensional design goals such as delay, power consumption and area, and meet the stringent requirements of memory design under advanced semiconductor processes. Actual test results show that the decoder structure optimized by this method can reduce delay by about 15% and power consumption by about 10% compared with the traditional design method, while maintaining similar area overhead.
[0135] In engineering applications, it is recommended to optimize parameters in combination with specific process parameters and product specifications. Design automation tools can be used to assist in parameter search and verification, further improving design efficiency and quality. The decoder structure designed by this method has been successfully applied to multiple commercial memory products, verifying its effectiveness and reliability in practical applications.
[0136] In an optional implementation, the tree structure hierarchy depth is calculated according to the total number of decoders and a preset fan-out coefficient, and the total number of decodes is evenly distributed based on the hierarchy depth to obtain the fan-out number of each layer, including:
[0137] Obtaining a total number of decodes and a fan-out coefficient of the decoder, dividing the total number of decodes by the fan-out coefficient to perform a logarithmic operation to obtain an initial level value, and rounding up the initial level value to obtain an actual level depth of the tree structure;
[0138] Performing a square root operation of the actual layer depth on the total number of decodes to obtain a uniform inter-layer fan-out reference value, and calculating an ideal fan-out number of each layer according to the uniform inter-layer fan-out reference value;
[0139] Comparing the ideal fan-out number with the fan-out coefficient, and when the ideal fan-out number is greater than the fan-out coefficient, increasing the actual layer depth and recalculating the ideal fan-out number until the ideal fan-out number is less than or equal to the fan-out coefficient;
[0140] A tree structure of a decoder is constructed based on the actual layer depth and the ideal fan-out number that are finally determined, and it is verified whether the number of decoding outputs of the tree structure is equal to the total number of decodings, so as to generate a balanced hierarchical structure design scheme that meets the decoding requirements.
[0141] In actual application scenarios, decoders often need to handle a large number of decoding tasks, and the tree structure can effectively organize and manage these tasks. In order to ensure a balance between decoding efficiency and resource utilization, the hierarchical depth of the tree structure and the number of fan-outs per layer need to be reasonably designed.
[0142] Get the total number of decodes and the preset fan-out coefficient of the decoder. The total number of decodes indicates the total number of decoding tasks that the decoder needs to process, and the preset fan-out coefficient is the maximum number of child nodes that each node can connect to, which is preset based on hardware conditions and performance requirements. For example, assuming that the total number of decodes of the decoder is 1000 and the preset fan-out coefficient is 8, it means that each node can connect to a maximum of 8 child nodes.
[0143] Calculate the actual hierarchical depth of the tree structure. Divide the total number of decodes by the fan-out coefficient, and then perform a logarithmic operation to obtain the initial hierarchical value. Round up this initial hierarchical value to obtain the actual hierarchical depth of the tree structure. Taking the above example, the calculation process is: 1000 divided by 8 is 125, and 125 is logarithmically operated with a base of 8. The result is approximately 2.37, which is rounded up to 3. Therefore, the actual hierarchical depth of the tree structure is 3.
[0144] Calculate the average fan-out baseline value between layers. Take the square root of the actual layer depth for the total number of decodes to get the average fan-out baseline value between layers. In the above example, taking the square root of 1000 gives a uniform fan-out baseline value of about 10. This means that if the fan-out number of each layer is 10, the three-layer tree structure can contain 1000 leaf nodes (that is, 10 to the power of 3).
[0145] Based on the uniform inter-layer fan-out benchmark value, calculate the ideal fan-out number for each layer. Since the uniform fan-out benchmark value is about 10 and the preset fan-out factor is 8, the ideal fan-out number exceeds the preset fan-out factor. At this time, you need to increase the actual layer depth and recalculate the ideal fan-out number. Increase the actual layer depth to 4 and recalculate the uniform inter-layer fan-out benchmark value: perform the fourth root operation on 1000 to obtain approximately 5.62. This value is less than the preset fan-out factor of 8, so it can be used.
[0146] The actual layer depth is determined to be 4, and the ideal fan-out number is 5.62. To simplify the implementation, the fan-out number of each layer can be set to the same integer value, such as 6. At this time, the number of leaf nodes in the four-layer tree structure is 6 to the fourth power, which is equal to 1296, which exceeds the required 1000 nodes. The total number of decodes can be accurately matched by reducing the number of nodes in the last layer.
[0147] When actually building a tree structure, the following solution can be used: the fan-out number of the first layer is 6, the fan-out number of each node in the second layer is 6, the fan-out number of each node in the third layer is 6, and the fourth layer needs to have 1000 leaf nodes. The total number of nodes in the first three layers is 1+6+36=43, and the fourth layer needs to connect 1000 leaf nodes. On average, each third-layer node needs to connect about 27.8 leaf nodes (1000 divided by 36). Since the preset fan-out coefficient is 8, it cannot meet this requirement and needs further adjustment.
[0148] Increase the actual layer depth to 5 and recalculate the uniform inter-layer fan-out benchmark value: perform the fifth root operation on 1000, and the result is about 4. This value is smaller than the preset fan-out factor of 8 and can be adopted. Assume that the fan-out number of each layer is 4, and the number of leaf nodes in the five-layer tree structure is 4 to the fifth power, which is equal to 1024, slightly larger than the required 1000 nodes. The total number of decodes can be accurately matched by reducing some nodes in the last layer.
[0149] Construct a 5-layer tree structure. The fan-out number of each node in the first four layers is 4. The fifth layer is adjusted appropriately as needed to ensure that the total number of leaf nodes is 1000. Specifically, the first four layers have a total of 1+4+16+64=85 nodes, of which the 64 nodes in the fourth layer need to connect to 1000 leaf nodes, and each node in the fourth layer needs to connect to about 15.6 leaf nodes on average. Since the preset fan-out coefficient is 8, it cannot meet this requirement, and the fan-out of the fourth layer needs to be further adjusted.
[0150] Add more nodes to the fourth layer. If the fan-out number of the third layer is increased to 8 (which meets the upper limit of the preset fan-out coefficient), the fourth layer will have 128 nodes. Each fourth-layer node connects to an average of about 7.8 leaf nodes (1000 divided by 128), which is close to but still slightly close to the upper limit of the preset fan-out coefficient. At this time, an uneven distribution strategy can be adopted, with some fourth-layer nodes connecting to 8 leaf nodes and the rest connecting to 7, ensuring that the total number of leaf nodes is 1000.
[0151] The first layer has 1 node and a fan-out of 4; the second layer has 4 nodes, each with a fan-out of 4; the third layer has 16 nodes, each with a fan-out of 8; the fourth layer has 128 nodes, of which 104 nodes are connected to 8 leaf nodes each, and 24 nodes are connected to 7 leaf nodes each. Calculate the total number of leaf nodes: 104×8+24×7=832+168=1000, which meets the total decoding quantity requirement.
[0152] Figure 4 This is a schematic diagram of a tree structure balanced hierarchical design system for a decoder according to an embodiment of the present invention:
[0153] This picture shows the detailed configuration and calculation result interface of a decoder design project (#D-28403). This project uses a 28nm process node, sets the total number of decodes to 256, presets the output coefficient to 4, and selects the design constraint of "delay priority" and the optimization level of "full optimization". During the logic calculation process, the system displays several key parameters: the initial layer value and the actual layer depth are both 4, the uniform inter-layer output reference value is 4.00, and the verification shows that the output reference value meets the preset requirements. The system calculates that the total interconnection resource requirement is 340 connections, the layer efficiency reaches 100% (in line with the ideal value), the fan-out uniformity is 5.0 / 5.0, and the final calculated area is about 0.042mm².
[0154] In terms of design constraint verification, the system completed the verification of four core indicators: fan-out coefficient constraint (all levels of fan-out ≤ 4), decode number constraint (total number of output lines = 256), uniform fan-out constraint (variance of output at each level = 0), and maximum delay constraint (1.28ns≤1.50ns). All these constraints have been met, and are also intuitively displayed in the constraint satisfaction analysis radar chart on the right side of the chart, covering the evaluation results of multiple dimensions such as area constraints, delay constraints, power constraints, and fan-out constraints. Overall, the design has achieved the expected goals in various technical indicators and demonstrated good comprehensive performance.
[0155] In an optional implementation, determining the hotspot distribution of the storage cell array region based on the power consumption distribution information of the storage cell array includes:
[0156] Obtaining the operating voltage, leakage current and cell density of each memory cell in the memory cell array, calculating the static power consumption density distribution, and collecting the load capacitance, flip activity factor and operating frequency of the memory cell; calculating the dynamic power consumption density distribution according to the operating voltage, load capacitance, flip activity factor and operating frequency, and superimposing the static power consumption density distribution with the dynamic power consumption density distribution to obtain the total power consumption density distribution;
[0157] Based on the ambient temperature and thermal resistance coefficient of the storage cell array, the actual temperature distribution of the storage cell array is calculated; the actual temperature distribution is analyzed according to a preset temperature threshold, the position coordinates of the hot spot area exceeding the preset temperature threshold are determined, and the temperature gradient at the position coordinates of the hot spot area is calculated.
[0158] Get the relevant parameters of each memory cell in the memory cell array. Specifically, extract the operating voltage, leakage current and cell density of each memory cell through the circuit model of the memory compiler IP. For example, at the 28nm process node, the typical SRAM memory cell operating voltage is 0.9V, the leakage current is about 15nA / cell, and the cell density is 0.127μm². By traversing the row and column coordinates of the memory cell array, a parameter matrix is formed to record the operating voltage V(i,j), leakage current I_leak(i,j) and cell density D(i,j) at each position (i,j). Among them, i and j represent the row index and column index of the memory cell in the array respectively, V(i,j) represents the operating voltage of the memory cell at position (i,j), in volts (V); I_leak(i,j) represents the leakage current of the memory cell at position (i,j), in nanoamperes (nA); D(i,j) represents the cell density of the memory cell at position (i,j), in square micrometers (μm²).
[0159] The static power density distribution is calculated based on the obtained parameters. For each coordinate position (i, j) in the array, the static power density P_static(i, j) is obtained by multiplying the operating voltage, leakage current and cell density. Here, P_static(i, j) represents the static power density at position (i, j) in watts per square micron (W / μm²); the static power density is equal to the operating voltage V(i, j) multiplied by the leakage current I_leak(i, j) multiplied by the cell density D(i, j). For example, for the storage cell at position (10, 15), the operating voltage is 0.9 V, the leakage current is 18 nA, and the cell density is 0.13 μm². The static power density at this position is 0.9×18×0.13=2.106 nW / μm². By calculating the entire array, a complete static power density distribution matrix can be obtained.
[0160] Collect dynamic power consumption parameters of storage cells, including load capacitance, flip activity factor, and operating frequency. The load capacitance of storage cells usually varies between 10.5, and the operating frequency may vary from hundreds of MHz to several GHz in different applications. Through simulation analysis or actual measurement, collect the load capacitance C(i,j), flip activity factor α(i,j), and operating frequency f(i,j) of each storage cell position (i,j). Among them, C(i,j) represents the load capacitance of the storage cell at position (i,j), in farads (F); α(i,j) represents the flip activity factor of the storage cell at position (i,j), dimensionless, representing the ratio of the average number of signal flips per unit time to the number of clock cycles; f(i,j) represents the operating frequency of the storage cell at position (i,j), in Hertz (Hz).
[0161] Calculate the dynamic power density distribution. For each coordinate position (i,j) in the array, the dynamic power density P_dynamic(i,j) is obtained by multiplying the square of the operating voltage, the load capacitance, the flip activity factor and the operating frequency and then dividing it by the unit area. Here, P_dynamic(i,j) represents the dynamic power density at position (i,j) in watts per square micron (W / μm²); the dynamic power density is equal to the square of the operating voltage V(i,j) multiplied by the load capacitance C(i,j) multiplied by the flip activity factor α(i,j) multiplied by the operating frequency f(i,j), and then divided by the unit area. Taking the 32nm process as an example, a storage cell is located at the (20,25) position, the operating voltage is 0.85V, the load capacitance is 3.2fF, the flip activity factor is 0.3, the operating frequency is 1.2GHz, and the unit area is 0.2μm². The dynamic power density at this position is:
[0162] 0.85×0.85×3.2×0.3×1.2×10 9 / 0.2=4.42mW / μm².
[0163] This calculation is repeated for the entire array to form a dynamic power density distribution matrix.
[0164] The static power density distribution is superimposed with the dynamic power density distribution to obtain the total power density distribution. For each position (i, j) in the array, the total power density P_total(i, j) is the sum of the static power density and the dynamic power density. That is, P_total(i, j) = P_static(i, j) + P_dynamic(i, j), where P_total(i, j) represents the total power density at position (i, j) in watts per square micron (W / μm²). For example, if the static power density at a certain position is 2.1nW / μm² and the dynamic power density is 4.2mW / μm², the total power density is approximately 4.2002mW / μm². Perform the addition operation on each position of the entire array to form the total power density distribution matrix P_total.
[0165] Based on the ambient temperature and thermal resistance of the memory cell array, the actual temperature distribution of the memory cell array is calculated. First, the ambient temperature T_ambient, which is usually 25°C, and the thermal resistance R_thermal, in °C / W, are obtained. The thermal resistance is usually determined by the chip package type, heat sink performance, and physical layout, and the typical value is in the range of 10~50°C / W. For each position (i,j) in the array, the actual temperature T(i,j) is obtained by adding the ambient temperature to the power consumption multiplied by the thermal resistance. Specifically, the sum of the ambient temperature and the product of the total power density and thermal resistance at that position is used as the actual temperature value: T(i,j) = T_ambient + P_total(i,j) × Area(i,j) × R_thermal. Where T(i,j) is the actual temperature at position (i,j) in degrees Celsius (°C); T_ambient is the ambient temperature in degrees Celsius (°C); Area(i,j) is the area of the unit at position (i,j) in square micrometers (μm²); R_thermal is the thermal resistance in degrees Celsius / W.
[0166] The ambient temperature is 25°C, the thermal resistance coefficient is 30°C / W, the total power consumption density at a certain location is 4.2mW / μm², and the unit area is 0.2μm². The actual temperature at this location is 25 + 4.2×0.2×30 = 50.2°C. The temperature distribution matrix T is formed by calculating the entire array.
[0167] The actual temperature distribution is analyzed according to the preset temperature threshold. The preset temperature threshold T_threshold is usually set to the upper limit of the safe operating temperature, for example, it can be set to 85°C in the 28nm process. Traverse each element in the temperature distribution matrix T, and if T(i,j) > T_threshold, mark the position (i,j) as a hot spot area. For example, in a 64×64 memory cell array, three hot spots may be marked, with coordinates of (12,15), (30,40) and (55,60), and corresponding temperatures of 87°C, 91°C and 88°C, respectively.
[0168] Calculate the temperature gradient at the hotspot location coordinates. The temperature gradient is obtained by dividing the temperature difference between adjacent locations by the distance. For each hotspot location (i,j), the temperature gradient along the x direction is [T(i+1,j) - T(i-1,j)] / 2, and the temperature gradient along the y direction is [T(i,j+1) - T(i,j-1)] / 2. Here, the temperature gradient along the x direction grad_T_x(i,j) represents the rate of temperature change along the x direction at the location (i,j), in degrees Celsius per unit (°C / unit); the temperature gradient along the y direction grad_T_y(i,j) represents the rate of temperature change along the y direction at the location (i,j), in degrees Celsius per unit (°C / unit). For example, for the hot spot (30,40), if T(29,40)=89°C, T(31,40)=92°C, T(30,39)=88°C, and T(30,41)=93°C, the temperature gradient in the x-direction is (92-89) / 2=1.5°C / unit, and the temperature gradient in the y-direction is (93-88) / 2=2.5°C / unit. The temperature gradient information is crucial for the precise determination of the subsequent low-power unit insertion position.
[0169] Traditional memory compiler IP hotspot analysis methods usually use uniform power consumption assumptions or only consider static power consumption, which cannot accurately reflect the hotspot distribution under actual working conditions. Existing temperature estimation techniques are also often based on simplified thermal models, without considering the impact of factors such as cell density, operating voltage changes, and flip activity factors on local temperature. This leads to inaccurate hotspot predictions and limited effectiveness of subsequent low-power cell insertion strategies.
[0170] The technical solution proposed in this application comprehensively considers static power consumption and dynamic power consumption, and superimposes the two to obtain a more accurate total power consumption density distribution. In addition, this solution considers the influence of physical parameters such as ambient temperature and thermal resistance coefficient on temperature distribution, and calculates the temperature gradient to provide precise position guidance for subsequent low-power unit insertion. The starting point of the improvement is to improve the accuracy of hot spot analysis so as to optimize power consumption in a more targeted manner.
[0171] Through this technical solution, the recognition accuracy of hotspot areas is increased from about 70% of the traditional method to more than 95%, and the temperature prediction error is within 52°C. Based on more accurate hotspot distribution information, the subsequently inserted low-power units can more effectively reduce the local temperature, which improves the overall power consumption optimization effect by about 30%, while reducing unnecessary low-power unit insertion, reducing area overhead and implementation complexity.
[0172] In an optional implementation, inserting a low-power unit in the hotspot distribution area, wherein the low-power unit dynamically adjusts the operating voltage and operating frequency of the storage unit to achieve dynamic balance of local power consumption includes:
[0173] Calculating a voltage regulation coefficient based on the temperature gradient, multiplying the nominal operating voltage by the voltage regulation coefficient to obtain an optimized operating voltage, and determining an optimized operating frequency according to a ratio of the optimized operating voltage to the nominal operating voltage;
[0174] Calculating a power consumption balance factor of a storage cell array, the power consumption balance factor being a ratio of an average power consumption density of the storage cell array to the total power consumption density distribution, and inserting a low power consumption unit at the position coordinates of the hot spot area according to the power consumption balance factor;
[0175] The low power consumption unit receives the power consumption balance factor, and adjusts the optimized operating voltage and the optimized operating frequency of the storage unit at the corresponding position in real time according to the power consumption balance factor.
[0176] The system obtains the temperature distribution information of the memory array and monitors the temperature of each area of the memory in real time through the temperature sensor array. The layout of the temperature sensor array corresponds to the memory cell array. For example, in a 64×64 memory cell array, 16 temperature sensors can be evenly distributed to form a 4×4 sensor array network. Each sensor can monitor the temperature changes in the surrounding area, and the sampling frequency is usually set to 1000Hz to ensure that temperature fluctuations can be captured in time.
[0177] Based on the collected temperature data, the system calculates the temperature gradient of each area of the memory. The temperature gradient is calculated by comparing the temperature difference between adjacent temperature sensors. For example, if the temperatures detected by two adjacent sensors are 75°C and 65°C respectively, and the physical distance between them is 2 mm, the temperature gradient of the area is 5°C / mm. The system then identifies the hot spot distribution area according to the size of the temperature gradient, and usually regards the area with a temperature gradient exceeding 3°C / mm as a potential hot spot area.
[0178] The system calculates the voltage regulation coefficient based on the temperature gradient. The determination of the voltage regulation coefficient α depends on the temperature gradient value. For each hot spot area detected, the system determines the corresponding regulation coefficient according to the temperature gradient size of the area. Specifically, when the temperature gradient is between 3℃ / mm and 5℃ / mm, the voltage regulation coefficient α is set to 0.95; when the temperature gradient is between 5℃ / mm and 8℃ / mm, the voltage regulation coefficient α is set to 0.90; when the temperature gradient exceeds 8℃ / mm, the voltage regulation coefficient α is set to 0.85.
[0179] The system multiplies the nominal operating voltage by the voltage regulation coefficient to obtain the optimized operating voltage. For example, if the nominal operating voltage of a certain area is 1.2V and the temperature gradient is 6℃ / mm, the corresponding voltage regulation coefficient is 0.90, so the optimized operating voltage is 1.2V×0.90=1.08V.
[0180] The system determines the optimized operating frequency based on the ratio of the optimized operating voltage to the nominal operating voltage. The frequency adjustment follows the following rules: when the voltage ratio is between 0.95 and 1.0, the frequency is reduced by 5%; when the voltage ratio is between 0.90 and 0.95, the frequency is reduced by 10%; when the voltage ratio is between 0.85 and 0.90, the frequency is reduced by 15%; when the voltage ratio is less than 0.85, the frequency is reduced by 20%. For example, if the voltage ratio is 0.90 and the nominal operating frequency is 800MHz, the optimized operating frequency is 800MHz×(1-10%)=720MHz.
[0181] The system also needs to calculate the power balance factor of the memory cell array, which is expressed as the ratio of the average power density of the memory cell array to the total power density distribution. The calculation of the power balance factor first requires obtaining the total power consumption P_total of the memory cell array and the power density distribution P_density(x,y) of each area. Assuming that the total area of the memory array is A_total, the average power density P_average is P_total / A_total. For the power balance factor β(x,y) at a specific location (x,y), it is calculated as P_average / P_density(x,y).
[0182] In a memory array with a power consumption of 5W and an area of 25 square millimeters, the average power consumption density is 0.2W / square millimeter. If the power consumption density detected at position (10,15) is 0.5W / square millimeter, the power consumption balance factor β(10,15) at this position is 0.2 / 0.5=0.4.
[0183] According to the calculated power balance factor, the system inserts a low-power unit at the coordinates of the hot spot area. The rules for determining the insertion position are: when the power balance factor β is less than 0.5, insert the low-power unit at this position; when the power balance factor β is between 0.5 and 0.8, decide whether to insert it according to the size of the temperature gradient; when the power balance factor β is greater than 0.8, there is no need to insert the low-power unit.
[0184] The design of the low-power unit adopts a dedicated power control circuit, which includes a voltage regulation module and a frequency regulation module. The voltage regulation module is implemented by a low-dropout linear regulator (LDO), which can accurately reduce the input voltage to the target voltage value. The frequency regulation module is implemented by a programmable divider, which can dynamically adjust the clock frequency according to demand.
[0185] After receiving the power balance factor, the low power unit adjusts the optimized operating voltage and operating frequency of the corresponding storage unit in real time. During the adjustment process, the low power unit first adjusts the supply voltage from the nominal value to the optimized operating voltage value through the LDO, and at the same time adjusts the clock frequency from the nominal frequency to the optimized operating frequency through the divider.
[0186] When the power consumption balance factor of a hot spot area is 0.4, the optimized operating voltage is 1.08V (10% lower than the nominal value of 1.2V), and the optimized operating frequency is 720MHz (10% lower than the nominal value of 800MHz), the low power unit will set the power supply voltage of this area to 1.08V through LDO, and set the clock frequency to 720MHz through the divider. In this way, the power consumption of this area will be significantly reduced, and the hot spot problem will be alleviated.
[0187] Through the above method, the temperature of the hot spot area of the memory array can be reduced by 8-15℃, the overall power consumption can be reduced by 12-25%, and the impact on the memory performance is controlled within 10%. Applying this method in a typical 64MB SRAM memory, the maximum temperature of the hot spot area is reduced from 85℃ to 72℃, the overall power consumption is reduced from 3.5W to 2.8W, and the read and write performance is only reduced by 7.5%, showing good temperature control and power consumption optimization effects.
[0188] By dynamically inserting low-power units in hot spots and adjusting the operating voltage and frequency in real time, this method achieves dynamic balance of local power consumption in the memory array, effectively solves the hot spot problem, and improves the reliability and service life of the memory.
[0189] Figure 5 This is a comparison data table of power consumption balance factor and temperature hotspot distribution in an embodiment of the present invention:
[0190] This table shows the performance comparison data of a chip temperature control system, including the temperature management effects of three hot spots (cache controller, address decoder and read / write controller) and non-hot spots. The table records the original temperature of different areas, the PBF (power blocking factor) value under this technical solution and traditional static power control methods, temperature changes and final temperature data.
[0191] From the data analysis, the original temperatures of the three hot spots are generally high, with the cache controller area ranging from 79.6°C to 82.5°C, the address decoder area ranging from 82.4°C to 84.2°C, and the read / write controller area having the highest temperature, reaching 84.1°C to 85.9°C. This technical solution achieves better temperature reduction effect through a lower PBF value (0.63-0.72), with an average temperature reduction of 10.7°C and a maximum temperature reduction of up to 12.3°C (in the read / write controller area). In contrast, the traditional solution has a higher PBF value (0.79-0.83), and the cooling effect is relatively weak, with an average temperature reduction of only 6.4°C.
[0192] In non-hotspot areas, due to the low original temperature (65.2°C to 68.3°C), the PBF values of both solutions are relatively high (0.92-0.96 for this solution and 0.95-0.99 for the traditional solution), and the temperature change is small. From the overall average value, the PBF average value of this technical solution is 0.73, and the final temperature is 71.5°C, while the PBF average value of the traditional solution is 0.85, and the final temperature is 74.9°C, which fully proves the advantage of this technical solution in temperature control effect.
[0193] A second aspect of an embodiment of the present invention provides an EDA software collaborative optimization system based on a memory compiler entity IP, comprising:
[0194] A first unit is used to obtain layout information of a memory compiler IP, wherein the layout information includes location information of a memory cell array area, location information of a data driver, location information of a decoder, and location information of a clock tree;
[0195] A second unit is used to establish a timing netlist model of the memory compiler IP based on the layout information, in combination with the dynamic power consumption and static power consumption of the memory compiler IP, and taking into account the influence of temperature on the timing path delay value, wherein the timing netlist model includes path timing information and power consumption distribution information of the memory cell array;
[0196] The third unit is used to adjust the driving strength of the data driver and adopt a progressive driving strength allocation strategy according to the path timing information in the timing netlist model, so that the signal transmission delay increases linearly with the distance to achieve timing convergence; adjust the hierarchical structure of the decoder, determine the fan-out number of each layer through the decoder tree structure optimization, and set a multi-level buffer between the data driver and the storage cell array area in combination with the target delay of each layer;
[0197] A fourth unit is used to determine the hot spot distribution of the storage cell array area based on the power consumption distribution information of the storage cell array, insert a low power consumption unit in the hot spot distribution area, and the low power consumption unit realizes the dynamic balance of local power consumption by dynamically adjusting the operating voltage and operating frequency of the storage cell;
[0198] The fifth unit is used to update the location information of the multi-level buffer and the location information of the low power consumption unit into the layout information, and perform wiring optimization on the updated layout information to generate an optimized memory compiler IP layout.
[0199] According to a third aspect of the embodiments of the present invention,
[0200] An electronic device is provided, comprising:
[0201] processor;
[0202] a memory for storing processor-executable instructions;
[0203] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0204] A fourth aspect of the embodiments of the present invention is:
[0205] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0206] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An EDA software collaborative optimization method based on a memory compiler entity IP, characterized in that: include: Acquire layout information of the memory compiler IP, wherein the layout information includes location information of a memory cell array region, location information of a data driver, location information of a decoder, and location information of a clock tree; Based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence of temperature on the timing path delay value, a timing netlist model of the memory compiler IP is established, wherein the timing netlist model includes path timing information and power consumption distribution information of the memory cell array; According to the path timing information in the timing netlist model, the driving strength of the data driver is adjusted and a progressive driving strength allocation strategy is adopted so that the signal transmission delay increases linearly with the distance to achieve timing convergence; the hierarchical structure of the decoder is adjusted, the fan-out number of each layer is determined by optimizing the decoder tree structure, and a multi-level buffer is set between the data driver and the storage cell array area in combination with the target delay of each layer; Based on the power consumption distribution information of the memory cell array, determining the hot spot distribution of the memory cell array area, inserting a low power consumption unit in the hot spot distribution area, wherein the low power consumption unit realizes a dynamic balance of local power consumption by dynamically adjusting the operating voltage and operating frequency of the memory cell; The location information of the multi-level buffer and the location information of the low power consumption unit are updated into the layout information, and the updated layout information is optimized for wiring to generate an optimized memory compiler IP layout.
2. The method according to claim 1, characterized in that Based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence of temperature on the timing path delay value, establishing the timing netlist model of the memory compiler IP includes: Extracting parasitic parameters of interconnects in the memory compiler IP based on the layout information, the parasitic parameters including resistance parameters and capacitance parameters of the interconnects, wherein the resistance parameters are calculated based on the resistivity, length, width and thickness of the wires, and the capacitance parameters include parallel plate capacitance, edge capacitance and coupling capacitance; Determine a delay value of a timing path based on the parasitic parameters, wherein the delay value is obtained by calculating the sum of the products of the delay coefficients of each segment on the timing path and the corresponding equivalent resistance and equivalent capacitance; Calculate the timing margin of the timing path according to the delay value, wherein the timing margin is obtained by subtracting the actual arrival time and the timing uncertainty from the required time of the timing path; Calculating the dynamic power consumption and static power consumption of the memory compiler IP, wherein the dynamic power consumption is calculated according to the relationship between load capacitance, operating voltage and operating frequency, and the static power consumption is calculated according to the relationship between operating voltage and leakage current; Determine the influence of the dynamic power consumption and the static power consumption on the timing, and obtain the influence of the temperature on the delay value of the timing path by calculating the product of the temperature coefficient and the difference between the actual temperature and the nominal temperature; Based on the influence of the temperature on the delay value, the delay value of the timing path is corrected by multiplying the temperature-delay conversion coefficient with the power consumption density and the thermal resistance to generate a complete timing netlist model containing timing information and power consumption distribution characteristics.
3. The method according to claim 1, characterized in that According to the path timing information in the timing netlist model, the driving strength of the data driver and the hierarchical structure of the decoder are adjusted, by setting a multi-level buffer between the data driver and the storage cell array area, and adopting a progressive driving strength allocation strategy, including: Calculating the load capacitance, logic delay and interconnection delay of the memory compiler, and determining the path timing margin from the clock cycle according to the logic delay and the interconnection delay; setting the voltage swing and the conversion time based on the load capacitance, and determining the driving requirement reference value according to the load capacitance, the voltage swing and the conversion time; Determine a minimum driver width of the data driver according to the driving requirement reference value, and multiply the minimum driver width by the square root of the ratio of the load capacitance to the reference capacitance to obtain an actual driver width of the data driver; determine a driver reference strength based on the actual driver width, establish a driving strength increasing function by setting an increasing coefficient, and calculate a theoretical distribution value of multi-level driving strength according to the driving strength increasing function; Acquire the distance information from the data driver to the storage unit, perform an exponential function operation on the distance information and a preset characteristic distance to obtain a distance weight coefficient; modify the theoretical distribution value of the multi-level driving strength according to the distance weight coefficient, and determine the actual driving strength of each level by calculating the ratio of the modified driving strength to the total driving strength; The path timing margin is updated based on the actual driving strength, and when the path timing margin meets the timing requirement, a progressive driving strength configuration scheme of the data driver is generated, and the progressive driving strength configuration scheme is applied to the memory compiler.
4. The method according to claim 1, characterized in that Adjusting the hierarchical structure of the decoder, determining the fan-out number of each layer by optimizing the decoder tree structure, and setting a multi-level buffer between the data driver and the storage cell array area in combination with the target delay of each layer includes: Obtain buffer load capacitance parameters, the load capacitance parameters including single-stage input capacitance, wiring capacitance and current-stage load capacitance, and add the load capacitance parameters to obtain a total load capacitance; calculate a theoretical optimal number of stages based on a logarithmic ratio of the total load capacitance to the single-stage input capacitance, and correct the theoretical optimal number of stages according to a preset margin coefficient to obtain an actual number of buffer stages; Obtaining total layout length information, distributing the total layout length to buffers at each level according to the distance attenuation coefficient based on the actual number of buffer levels, and obtaining initial spacings of buffers at each level; calculating the tree structure hierarchy depth according to the total number of decoders and a preset fan-out coefficient, and evenly distributing the total number of decodes based on the hierarchy depth to obtain the fan-out number of each layer; Obtain a total delay constraint of the decoder, determine a layer weight coefficient according to the fan-out number of each layer, distribute the total delay constraint to each layer according to the layer weight coefficient, and obtain a target delay of each layer; Calculating the actual delay corresponding to the initial spacing and the fan-out number of each layer, and taking the difference between the actual delay and the target delay as the delay compensation amount; adjusting the initial spacing and the fan-out number of each layer according to the delay compensation amount, and calculating the adjusted total delay of the signal transmission path; The power consumption density distribution information of each physical position is collected, and the actual spacing is locally fine-tuned based on the power consumption density distribution information until the total delay of the signal transmission path meets the timing requirements, thereby generating a final decoder hierarchical structure design scheme.
5. The method according to claim 4, characterized in that The tree structure hierarchy depth is calculated according to the total number of decoders and the preset fan-out coefficient, and the total number of decodes is evenly distributed based on the hierarchy depth to obtain the fan-out number of each layer, including: Obtaining a total number of decodes and a fan-out coefficient of the decoder, dividing the total number of decodes by the fan-out coefficient to perform a logarithmic operation to obtain an initial level value, and rounding up the initial level value to obtain an actual level depth of the tree structure; Performing a square root operation of the actual layer depth on the total number of decodes to obtain a uniform inter-layer fan-out reference value, and calculating an ideal fan-out number of each layer according to the uniform inter-layer fan-out reference value; Comparing the ideal fan-out number with the fan-out coefficient, and when the ideal fan-out number is greater than the fan-out coefficient, increasing the actual layer depth and recalculating the ideal fan-out number until the ideal fan-out number is less than or equal to the fan-out coefficient; A tree structure of a decoder is constructed based on the actual layer depth and the ideal fan-out number that are finally determined, and it is verified whether the number of decoding outputs of the tree structure is equal to the total number of decodings, so as to generate a balanced hierarchical structure design scheme that meets the decoding requirements.
6. The method according to claim 1, characterized in that Based on the power consumption distribution information of the memory cell array, determining the hotspot distribution of the memory cell array region includes: Obtaining the operating voltage, leakage current and cell density of each memory cell in the memory cell array, calculating the static power consumption density distribution, and collecting the load capacitance, flip activity factor and operating frequency of the memory cell; calculating the dynamic power consumption density distribution according to the operating voltage, load capacitance, flip activity factor and operating frequency, and superimposing the static power consumption density distribution with the dynamic power consumption density distribution to obtain the total power consumption density distribution; Based on the ambient temperature and thermal resistance coefficient of the storage cell array, the actual temperature distribution of the storage cell array is calculated; the actual temperature distribution is analyzed according to a preset temperature threshold, the position coordinates of the hot spot area exceeding the preset temperature threshold are determined, and the temperature gradient at the position coordinates of the hot spot area is calculated.
7. The method according to claim 6, characterized in that Inserting a low-power unit in the hot spot distribution area, wherein the low-power unit dynamically adjusts the operating voltage and operating frequency of the storage unit to achieve dynamic balance of local power consumption includes: Calculating a voltage regulation coefficient based on the temperature gradient, multiplying the nominal operating voltage by the voltage regulation coefficient to obtain an optimized operating voltage, and determining an optimized operating frequency according to a ratio of the optimized operating voltage to the nominal operating voltage; Calculating a power consumption balance factor of a storage cell array, the power consumption balance factor being a ratio of an average power consumption density of the storage cell array to the total power consumption density distribution, and inserting a low power consumption unit at the position coordinates of the hot spot area according to the power consumption balance factor; The low power consumption unit receives the power consumption balance factor, and adjusts the optimized operating voltage and the optimized operating frequency of the storage unit at the corresponding position in real time according to the power consumption balance factor.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Territory programming method for static random access memory compiler
CN103106294A
Memory compiler splicing method and memory
CN106156375A
Compiling method and device of memory and generated memory
CN108446412A
Memory characterization method, chip design method, medium and device
CN115993943A
EDA software implementation method and system for memory IP layout optimization
CN119598954A
Cited By
Memory compiler IP time sequence convergence and power consumption collaborative optimization method and system
CN120597833A
Memory compiler ip timing convergence and power co-optimization method and system
CN120597833B
Integrated circuit digital back-end physical optimization method, device, medium and product
CN121118814A
Clock network optimization method, computer equipment and storage medium
CN121145784A