EDA Software Co-Optimization Method Based on Memory Compiler Entity IP
By adjusting the layout layout and driver strength of the memory compiler IP, optimizing the decoder hierarchical structure, and inserting low-power units, the timing convergence difficulties and hot issues in traditional memory compiler IP are solved, and linear delay and local power consumption balance of signal transmission are achieved, improving performance and reliability.
Patent Information
- Application Number
- CN202510479248.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the design of traditional memory compiler IP, fixed driving strength leads to timing violations and waste of power consumption, decoder design lacks systematic optimization, passive handling of hot issues, and difficult to achieve global power consumption and performance optimization.
By obtaining the layout information of the memory compiler IP, establishing a timing netlist model, adjusting the driver strength and decoder hierarchical structure, inserting low-power units, the signal transmission delay increases linearly with distance, and dynamic balance of local power consumption.
It improves the overall performance and stability of the memory compiler IP, reduces overall power consumption, extends the chip service life, and improves adaptability and reliability.
Smart Images

Figure CN119990049B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to memory technologies, and in particular, to an EDA software co - optimization method based on a memory compiler entity IP. Background Art
[0002] The memory compiler IP is a key component in integrated circuit design, which can automatically generate memory cells that meet specific specifications according to design requirements. With the continuous development of integrated circuit processes and the increase in design complexity, the memory compiler IP has occupied an increasingly important position in chip design. The memory compiler IP usually includes key components such as a memory cell array region, data drivers, decoders, and clock trees. The performance and power consumption of these components directly affect the performance and reliability of the entire chip.
[0003] At the current semiconductor process nodes, chip design faces severe power consumption and timing challenges. Electronic design automation (EDA) tools play an indispensable role in the design and optimization process of the memory compiler IP. Traditional EDA tools usually optimize the layout, power consumption distribution, and timing paths independently. However, with the continuous progress of the process, this isolated optimization method is difficult to meet the design requirements of high performance and low power consumption. Especially at advanced process nodes, the influence of temperature on the timing path delay is more significant, and the hot - spot problem caused by the uneven power consumption distribution is more prominent, which poses higher requirements for the design of the memory compiler IP.
[0004] In the design process of traditional memory compiler IP, data drivers with fixed drive strengths are usually adopted, and the drive ability cannot be dynamically adjusted according to the change of signal transmission distance. This may lead to timing violations during long - distance signal transmission, while power consumption waste may occur during short - distance signal transmission, making it difficult to achieve global optimization of timing and power consumption. In addition, the fixed drive strength strategy makes the relationship between signal transmission delay and distance non - linear, increasing the difficulty of timing convergence.
[0005] In the existing decoder design, a simple multi - stage structure is often adopted, lacking systematic optimization of the fan - out numbers at each stage. This design method is difficult to balance the delay distribution among the decoder stages, resulting in uneven distribution of timing margins on the critical path and ultimately affecting the overall performance. At the same time, the buffering strategy between the data driver and the memory cell array is also relatively simple, failing to fully consider the relationship between the path delay target and the number of buffer stages, making it difficult to achieve optimal timing performance.
[0006] In traditional memory design, the handling of hot spot problems is relatively passive. It usually relies on global power management or simple power consumption balancing techniques, lacking a refined dynamic management mechanism for local hot spots. This method cannot effectively address the local hot spot problems caused by uneven access patterns in the memory cell array, easily leading to local overheating and affecting the reliability and lifespan of the chip. In addition, existing technologies lack a method to organically combine power consumption distribution information with layout and routing strategies, making it difficult to achieve global optimization that takes both power consumption and performance into account. Summary of the Invention
[0007] The embodiments of the present invention provide an EDA software co-optimization method based on a memory compiler entity IP, which can solve the problems in the prior art.
[0008] In the first aspect of the embodiments of the present invention,
[0009] Provide an EDA software co-optimization method based on a memory compiler entity IP, including:
[0010] Obtain the layout information of the memory compiler IP, where the layout information includes the position information of the memory cell array area, the position information of the data driver, the position information of the decoder, and the position information of the clock tree;
[0011] Based on the layout information, combine the dynamic power consumption and static power consumption of the memory compiler IP, and consider the influence degree of temperature on the timing path delay value to establish a timing netlist model of the memory compiler IP, where the timing netlist model includes path timing information and the power consumption distribution information of the memory cell array;
[0012] According to the path timing information in the timing netlist model, adjust the driving strength of the data driver and adopt a progressive driving strength allocation strategy to make the signal transmission delay increase linearly with the distance to achieve timing convergence; adjust the hierarchical structure of the decoder, determine the fan-out number of each layer through the optimization of the decoder tree structure, and set multiple-level buffers between the data driver and the memory cell array area in combination with the target delay of each layer;
[0013] Based on the power consumption distribution information of the memory cell array, determine the hot spot distribution of the memory cell array area, and insert low-power units in the hot spot distribution area. The low-power units achieve local power consumption dynamic balance by dynamically adjusting the operating voltage and operating frequency of the memory cells;
[0014] Update the position information of the multiple-level buffers and the position information of the low-power units to the layout information, and perform routing optimization on the updated layout information to generate an optimized memory compiler IP layout.
[0015] Based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence degree of temperature on the timing path delay value, establishing a timing netlist model of the memory compiler IP includes:
[0016] Extracting the parasitic parameters of the interconnects in the memory compiler IP based on the layout information, where the parasitic parameters include the resistance parameter and capacitance parameter of the interconnects. The resistance parameter is calculated according to the resistivity, length, width, and thickness of the wire, and the capacitance parameter includes the parallel plate capacitance value, fringe capacitance value, and coupling capacitance value;
[0017] Determining the delay value of the timing path based on the parasitic parameters, where the delay value is obtained by calculating the sum of the products of the delay coefficients of each segment on the timing path and the corresponding equivalent resistance and equivalent capacitance;
[0018] Calculating the timing margin of the timing path according to the delay value, where the timing margin is obtained by subtracting the actual arrival time and timing uncertainty from the required time of the timing path;
[0019] Calculating the dynamic power consumption and static power consumption of the memory compiler IP, where the dynamic power consumption is calculated according to the relationship between the load capacitance, operating voltage, and operating frequency, and the static power consumption is calculated according to the relationship between the operating voltage and leakage current;
[0020] Determining the influence of the dynamic power consumption and static power consumption on timing, and obtaining the influence degree of temperature on the timing path delay value by calculating the product of the temperature coefficient and the difference between the actual temperature and the nominal temperature;
[0021] Based on the influence degree of temperature on the delay value, using the product of the temperature-delay conversion coefficient, power consumption density, and thermal resistance to correct the delay value of the timing path, and generating a complete timing netlist model including timing information and power consumption distribution characteristics.
[0022] According to the path timing information in the timing netlist model, adjusting the driving strength of the data driver and the hierarchical structure of the decoder, by setting multiple-level buffers between the data driver and the storage cell array area, and adopting a progressive driving strength allocation strategy, including:
[0023] Calculating the load capacitance, logic delay, and interconnect delay of the memory compiler, determining the path timing margin from the clock cycle according to the logic delay and the interconnect delay; setting the voltage swing and transition time based on the load capacitance, and determining the driving requirement reference value according to the load capacitance, the voltage swing, and the transition time;
[0024] Determine the minimum driver width of the data driver according to the driving requirement reference value, multiply the minimum driver width by the square root of the ratio of the load capacitance to the reference capacitance to obtain the actual driver width of the data driver; determine the driver reference strength based on the actual driver width, establish a driving strength increment function by setting an increment coefficient, and calculate the theoretical allocation values of multiple-level driving strengths according to the driving strength increment function;
[0025] Obtain the distance information from the data driver to the storage unit, perform an exponential function operation on the distance information and a preset characteristic distance to obtain a distance weight coefficient; correct the theoretical allocation values of the multiple-level driving strengths according to the distance weight coefficient, and determine the actual driving strength of each level by the ratio of the corrected driving strength to the total driving strength;
[0026] Update the path timing margin based on the actual driving strength. When the path timing margin meets the timing requirements, generate a progressive driving strength configuration scheme for the data driver, and apply the progressive driving strength configuration scheme to the memory compiler.
[0027] Adjust the hierarchical structure of the decoder, determine the fan-out numbers of each layer through the optimization of the decoder tree structure, and set multiple-level buffers between the data driver and the storage unit array area in combination with the target delay of each layer, including:
[0028] Obtain the buffer load capacitance parameters, where the load capacitance parameters include the single-stage input capacitance, the wiring capacitance, and the current-stage load capacitance, add the load capacitance parameters to obtain the total load capacitance; calculate the theoretical optimal number of stages based on the logarithmic ratio of the total load capacitance to the single-stage input capacitance, and correct the theoretical optimal number of stages according to a preset margin coefficient to obtain the actual buffer number of stages;
[0029] Obtain the total layout length information, allocate the total layout length to each stage of the buffer according to the distance attenuation coefficient according to the actual buffer number of stages to obtain the initial spacing of each stage of the buffer; calculate the hierarchical depth of the tree structure according to the total decoding number of the decoder and the preset fan-out coefficient, and evenly distribute the total decoding number based on the hierarchical depth to obtain the fan-out numbers of each layer;
[0030] Obtain the total delay constraint of the decoder, determine the layer weight coefficients according to the fan-out numbers of each layer, and allocate the total delay constraint to each layer according to the layer weight coefficients to obtain the target delay of each layer;
[0031] Calculate the actual delay corresponding to the initial spacing and the fan-out numbers of each layer, and use the difference between the actual delay and the target delay as the delay compensation amount; adjust the initial spacing and the fan-out numbers of each layer according to the delay compensation amount, and calculate the total delay of the adjusted signal transmission path;
[0032] Collect the power consumption density distribution information of each physical location, and locally fine-tune the actual pitch based on the power consumption density distribution information until the total delay of the signal transmission path meets the timing requirements, and generate a final decoder hierarchical structure design scheme.
[0033] Calculate the hierarchical depth of the tree structure according to the total number of decodings of the decoder and the preset fan-out coefficient, and evenly distribute the total number of decodings based on the hierarchical depth to obtain the fan-out numbers of each layer, including:
[0034] Obtain the total number of decodings and the fan-out coefficient of the decoder, perform a logarithmic operation on dividing the total number of decodings by the fan-out coefficient to obtain an initial hierarchical value, and round up the initial hierarchical value to obtain the actual hierarchical depth of the tree structure;
[0035] Perform a root operation of the total number of decodings to the power of the actual hierarchical depth to obtain a uniform inter-layer fan-out reference value, and calculate the ideal fan-out number of each layer according to the uniform inter-layer fan-out reference value;
[0036] Compare the ideal fan-out number with the fan-out coefficient. When the ideal fan-out number is greater than the fan-out coefficient, increase the actual hierarchical depth and recalculate the ideal fan-out number until the ideal fan-out number is less than or equal to the fan-out coefficient;
[0037] Construct a tree structure of the decoder based on the finally determined actual hierarchical depth and the ideal fan-out number, verify whether the decoding output number of the tree structure is equal to the total number of decodings, and generate a balanced hierarchical structure design scheme that meets the decoding requirements.
[0038] Based on the power consumption distribution information of the memory cell array, determine the hot spot distribution in the memory cell array area, including:
[0039] Obtain the operating voltage, leakage current and cell density of each memory cell in the memory cell array, calculate the static power consumption density distribution, and collect the load capacitance, switching activity factor and operating frequency of the memory cell; calculate the dynamic power consumption density distribution according to the operating voltage, load capacitance, switching activity factor and operating frequency, and superimpose the static power consumption density distribution and the dynamic power consumption density distribution to obtain the total power consumption density distribution;
[0040] Based on the ambient temperature and thermal resistance coefficient of the memory cell array, calculate the actual temperature distribution of the memory cell array; analyze the actual temperature distribution according to the preset temperature threshold, determine the position coordinates of the hot spot area exceeding the preset temperature threshold, and calculate the temperature gradient at the position coordinates of the hot spot area.
[0041] Insert low-power units in the hot spot distribution area. The low-power units achieve local power consumption dynamic balance by dynamically adjusting the operating voltage and operating frequency of the memory cells, including:
[0042] Calculate the voltage adjustment coefficient based on the temperature gradient, multiply the nominal operating voltage by the voltage adjustment coefficient to obtain the optimized operating voltage, and determine the optimized operating frequency according to the ratio of the optimized operating voltage to the nominal operating voltage;
[0043] Calculate the power consumption balance factor of the memory cell array. The power consumption balance factor is the ratio of the average power consumption density of the memory cell array to the total power consumption density distribution. Insert low-power units at the position coordinates of the hot spot area according to the power consumption balance factor;
[0044] The low-power unit receives the power consumption balance factor and adjusts the optimized operating voltage and the optimized operating frequency of the memory cells at the corresponding positions in real time according to the power consumption balance factor.
[0045] In the second aspect of the embodiments of the present invention, an EDA software co-optimization system based on a memory compiler entity IP is provided, including:
[0046] The first unit is used to obtain the layout information of the memory compiler IP. The layout information includes the position information of the memory cell array area, the position information of the data driver, the position information of the decoder, and the position information of the clock tree;
[0047] The second unit is used to establish a timing netlist model of the memory compiler IP based on the layout information, combine the dynamic power consumption and static power consumption of the memory compiler IP, and consider the influence degree of temperature on the timing path delay value. The timing netlist model includes path timing information and power consumption distribution information of the memory cell array;
[0048] The third unit is used to adjust the driving strength of the data driver according to the path timing information in the timing netlist model and adopt a progressive driving strength allocation strategy to make the signal transmission delay increase linearly with the distance to achieve timing convergence; adjust the hierarchical structure of the decoder, determine the fan-out number of each layer through the optimization of the decoder tree structure, and set multi-level buffers between the data driver and the memory cell array area in combination with the target delay of each layer;
[0049] The fourth unit is used to determine the hot spot distribution of the memory cell array area based on the power consumption distribution information of the memory cell array, insert low-power units in the hot spot distribution area, and the low-power units achieve local power consumption dynamic balance by dynamically adjusting the operating voltage and operating frequency of the memory cells;
[0050] A fifth unit is configured to update the position information of the multi-level buffer and the position information of the low-power unit to the layout information, and perform routing optimization on the updated layout information to generate an optimized memory compiler IP layout.
[0051] In a third aspect of the embodiments of the present invention
[0052] There is provided an electronic device, including:
[0053] A processor;
[0054] A memory for storing instructions executable by the processor;
[0055] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0056] In a fourth aspect of the embodiments of the present invention,
[0057] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0058] The beneficial effects of this application are as follows:
[0059] By obtaining the layout information of the memory compiler IP and establishing an accurate timing netlist model, the present invention realizes the co-optimization of timing and power consumption. This method can accurately reflect the influence of temperature on the timing path delay value, providing a reliable basis for subsequent optimization, thereby improving the overall performance and stability of the memory compiler IP.
[0060] The present invention adopts a progressive drive strength allocation strategy and decoder hierarchical structure optimization, making the signal transmission delay increase linearly with distance, effectively solving the problem of difficult timing convergence in traditional methods. By setting a multi-level buffer between the data driver and the memory cell array area, the signal transmission path is further optimized, significantly improving the timing performance and working reliability of the memory compiler IP.
[0061] Based on the power consumption distribution information, the present invention identifies the hot spot distribution area and inserts low-power units, and realizes the dynamic balance of local power consumption by dynamically adjusting the working voltage and frequency of the memory cells. This method effectively reduces the overall power consumption of the memory compiler IP, improves the energy utilization efficiency, reduces the negative impact of the hot spot area on the surrounding circuits, extends the chip service life, and greatly improves the adaptability and reliability of the memory compiler IP in various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1Schematic flowchart of the EDA software co - optimization method based on the memory compiler entity IP according to an embodiment of the present invention;
[0063] Figure 2 Data comparison table of parasitic parameter extraction accuracy of different methods according to an embodiment of the present invention;
[0064] Figure 3 Schematic diagram of the relationship between drive strength and timing margin according to an embodiment of the present invention;
[0065] Figure 4 Schematic diagram of the decoder tree - shaped structure balanced hierarchical design system according to an embodiment of the present invention;
[0066] Figure 5 Data comparison table of the power consumption balance factor and temperature hot - spot distribution according to an embodiment of the present invention. Detailed implementation manners
[0067] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0069] Figure 1 Schematic flowchart of the EDA software co - optimization method based on the memory compiler entity IP according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0070] Obtain the layout information of the memory compiler IP, where the layout information includes the position information of the storage cell array area, the position information of the data driver, the position information of the decoder, and the position information of the clock tree;
[0071] Based on the layout information, combine the dynamic power consumption and static power consumption of the memory compiler IP, and consider the influence degree of temperature on the timing path delay value to establish a timing netlist model of the memory compiler IP, where the timing netlist model includes path timing information and power consumption distribution information of the storage cell array;
[0072] According to the path timing information in the timing netlist model, adjust the driving strength of the data driver and adopt a progressive driving strength allocation strategy to make the signal transmission delay increase linearly with distance, achieving timing convergence; adjust the hierarchical structure of the decoder, determine the fan-out number of each layer through the optimization of the decoder tree structure, and set multiple-level buffers between the data driver and the storage cell array area in combination with the target delay of each layer;
[0073] Based on the power consumption distribution information of the storage cell array, determine the hot spot distribution in the storage cell array area, and insert low-power cells in the hot spot distribution area. The low-power cells achieve local power consumption dynamic balance by dynamically adjusting the working voltage and working frequency of the storage cells;
[0074] Update the position information of the multiple-level buffers and the position information of the low-power cells to the layout information, and perform routing optimization on the updated layout information to generate an optimized memory compiler IP layout.
[0075] In an alternative embodiment, based on the layout information, combined with the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence degree of temperature on the timing path delay value, establishing a timing netlist model of the memory compiler IP includes:
[0076] Extract the parasitic parameters of the interconnect lines in the memory compiler IP based on the layout information. The parasitic parameters include the resistance parameter and capacitance parameter of the interconnect lines, where the resistance parameter is calculated according to the resistivity, length, width, and thickness of the wire, and the capacitance parameter includes the parallel plate capacitance value, fringe capacitance value, and coupling capacitance value;
[0077] Determine the delay value of the timing path based on the parasitic parameters. The delay value is obtained by calculating the sum of the products of the delay coefficients of each segment on the timing path and the corresponding equivalent resistance and equivalent capacitance;
[0078] Calculate the timing margin of the timing path. The timing margin is obtained by subtracting the actual arrival time and the timing uncertainty from the required time of the timing path;
[0079] Calculate the dynamic power consumption and static power consumption of the memory compiler IP, where the dynamic power consumption is calculated according to the relationship between the load capacitance, working voltage, and working frequency, and the static power consumption is calculated according to the relationship between the working voltage and the leakage current;
[0080] Determine the influence of the dynamic power consumption and static power consumption on the timing. By calculating the product of the temperature coefficient and the difference between the actual temperature and the nominal temperature, obtain the influence degree of temperature on the timing path delay value;
[0081] Based on the influence degree of the temperature on the delay value, the delay value of the timing path is corrected by using the product of the temperature-delay conversion coefficient and the power consumption density and the thermal resistance, and a complete timing netlist model including timing information and power consumption distribution characteristics is generated.
[0082] Extract the parasitic parameters of the interconnecting wires based on the layout information of the memory compiler IP. The parasitic parameters include the resistance parameter and the capacitance parameter of the interconnecting wires. The calculation of the resistance parameter is based on the resistivity, length, width and thickness of the wire. For example, for a copper wire with a length of 100 microns, a width of 0.5 microns and a thickness of 0.2 microns, its resistivity is 1.68 micro-ohm·centimeter, and its resistance value can be calculated to be 6.72 ohms. The capacitance parameters include the parallel plate capacitance value, the fringe capacitance value and the coupling capacitance value. For two adjacent parallel wires, the parallel plate capacitance can be calculated by dividing the overlapping area of the wires by the thickness of the insulating layer and multiplying by the dielectric constant. For example, if the overlapping area of two wires is 10 square microns, the thickness of the insulating layer is 0.1 microns, and the dielectric constant is 3.9, the parallel plate capacitance is about 3.45 femtofarads. The fringe capacitance depends on the edge shape of the wire and the surrounding environment, and is usually obtained by looking up a table or using a dedicated extraction tool. The coupling capacitance reflects the electromagnetic coupling effect between adjacent wires, and its value is related to the wire spacing and the parallel length.
[0083] Determine the delay value of the timing path based on the extracted parasitic parameters. The delay value of the timing path is obtained by calculating the sum of the products of the delay coefficients of each segment on the path and the corresponding equivalent resistance and equivalent capacitance. Specifically, for each segment on the timing path, first determine the delay coefficient of its driving unit, then calculate the equivalent resistance and equivalent capacitance, and finally take the product of the three as the delay value of this segment. For example, for an inverter with a driving strength of 4, its delay coefficient is 0.15 nanoseconds / picofarad·ohm, the load equivalent resistance is 100 ohms, and the equivalent capacitance is 20 picofarads, then the delay value of this segment is 0.3 nanoseconds. Add up the delay values of all segments on the timing path to obtain the total delay value of the entire path.
[0084] Calculate the timing margin of the timing path according to the calculated delay value. The timing margin is obtained by subtracting the actual arrival time and the timing uncertainty from the required time of the timing path. For example, if the required time of a certain timing path is 2 nanoseconds, the actual arrival time is 1.7 nanoseconds, and the timing uncertainty is 0.2 nanoseconds, then the timing margin of this path is 0.1 nanosecond. The timing uncertainty includes factors such as clock skew, clock jitter and delay variations caused by process variations.
[0085] Calculate the dynamic power consumption and static power consumption of the memory compiler IP. The dynamic power consumption is calculated based on the relationship among the load capacitance, operating voltage, and operating frequency. For example, for a load capacitance of 10 picofarads, an operating voltage of 1.2 volts, and an operating frequency of 500 megahertz, the dynamic power consumption is approximately 3.6 milliwatts. The static power consumption is calculated based on the relationship between the operating voltage and the leakage current. For example, if the leakage current is 100 microamps and the operating voltage is 1.2 volts, the static power consumption is 0.12 milliwatts. The total power consumption is the sum of the dynamic power consumption and the static power consumption, which is 3.72 milliwatts in this example.
[0086] Determine the impact of dynamic power consumption and static power consumption on timing. This is mainly achieved by calculating the product of the temperature coefficient and the difference between the actual temperature and the nominal temperature to obtain the degree of influence of temperature on the timing path delay value. For example, if the temperature coefficient is 0.2% / °C, the actual temperature is 85°C, and the nominal temperature is 25°C, the increase in delay caused by temperature is 12%. For a path with an original delay of 1 nanosecond, the delay will increase to 1.12 nanoseconds at high temperature.
[0087] Based on the degree of influence of temperature on the delay value, use the product of the temperature-delay conversion coefficient, power consumption density, and thermal resistance to correct the delay value of the timing path, and generate a complete timing netlist model that includes timing information and power consumption distribution characteristics. The temperature-delay conversion coefficient is usually provided by the process library. For example, the temperature-delay conversion coefficient of a certain 65-nanometer process is 0.15 nanoseconds / watt·square millimeter·°C. The power consumption density can be obtained by dividing the total power consumption by the chip area. For example, if the total power consumption is 10 milliwatts and the chip area is 0.5 square millimeters, the power consumption density is 20 watts / square millimeter. The thermal resistance represents the temperature increase caused by unit power and is usually related to the package type. For example, the thermal resistance of a certain BGA package is 20°C / watt. Multiply these parameters to obtain the increase in delay, which is used to correct the original delay value. For example, if the original delay value is 1 nanosecond and the increase in delay is 0.06 nanoseconds, the corrected delay value is 1.06 nanoseconds.
[0088] Through the above steps, a complete timing netlist model that includes timing information and power consumption distribution characteristics is finally generated. This model not only considers the layout information and interconnect parasitic parameters, but also covers the impact of power consumption on temperature and the impact of temperature on timing, thus providing more accurate timing analysis results. Using this model, the performance of the memory compiler IP can be predicted in the early stage of design, assisting designers to optimize the memory architecture and layout, shorten the design cycle, and improve the reliability and performance of the memory IP.
[0089] Figure 2 The following is a comparison data table of the parasitic parameter extraction accuracy of different methods in the embodiments of the present invention:
[0090] This picture is a comparison table of interconnect parameters, showing the measurement results of the electrical parameters of interconnects (M1 - M5) with different lengths (from 1μm to 200μm). The table includes two key parameters: resistance (Ω) and capacitance (fF), and lists the actual measured values, the values of the technical solution, as well as the comparison data and error rates with the NE and NF processes for each parameter.
[0091] From the data, as the length of the interconnect increases, both the resistance and capacitance values show an upward trend. Taking the short M1 line (L = 1μm) as an example, the measured resistance value is 8.75Ω and the capacitance is 0.95fF; while for the longest M5 line (L = 200μm), the resistance increases to 67.80Ω and the capacitance reaches 36.40fF. The error rates of each parameter generally increase with the increase of the line length. The error rate of the short M1 line is relatively small (resistance 0.8%, capacitance 2.1%), while the error rate of the long M5 line is relatively large (resistance 2.2%, capacitance 2.6%).
[0092] Compared with the existing technologies NE and NF, the parameters of this solution are generally better. Taking the long M3 line (L = 50μm) as an example, the measured resistance value is 105.30Ω, while for NE and NF they are 111.62Ω and 114.78Ω respectively; the measured capacitance value is 12.75fF, while for NE and NF they are 13.71fF and 14.23fF respectively. This shows that this technical solution has obvious advantages in the control of interconnect parameters, especially in achieving good results in reducing parasitic resistance and capacitance.
[0093] This application improves the manufacturing process of the interconnect, especially innovatively optimizes aspects such as metal layer deposition, patterning, and etching processes. In the existing technology, the traditional manufacturing process of interconnects mainly relies on conventional physical vapor deposition (PVD) and chemical vapor deposition (CVD) processes. This method is prone to problems such as uneven metal grain size and insufficient line width control accuracy during the manufacturing process, resulting in large fluctuations in the resistance and capacitance parameters of the interconnect, especially being more obvious in long - line structures.
[0094] This method effectively controls the growth process of metal grains by optimizing the deposition conditions of the metal layer and adopting a multi - step deposition process; at the same time, by improving the lithography and etching process parameters, the patterning accuracy and the control ability of the sidewall profile are improved. In addition, this application also introduces a new surface treatment process, effectively reducing the interface effect between the metal layer and the dielectric layer.
[0095] The preparation accuracy of the interconnecting lines is significantly improved, resulting in a remarkable reduction in the deviation between the actual parameters and the design values of the interconnecting lines in each layer; the electrical conductivity of the interconnecting lines is improved, and the line resistance is generally reduced; by optimizing the interface characteristics, the parasitic capacitance is effectively reduced; the technical solution of the present application shows more obvious advantages in the long-line structure, providing a better process basis for the design and manufacture of high-performance integrated circuits. Compared with the prior art, the solution of the present application has made remarkable progress in terms of the stability and consistency of the control of interconnecting line parameters, especially having obvious advantages in reducing parasitic effects, which is of great significance for improving the overall performance and reliability of integrated circuits.
[0096] In an alternative embodiment, according to the path timing information in the timing netlist model, the driving strength of the data driver and the hierarchical structure of the decoder are adjusted. By setting multiple-level buffers between the data driver and the memory cell array region and adopting a progressive driving strength allocation strategy, it includes:
[0097] Calculate the load capacitance, logical delay, and interconnect delay of the memory compiler, and determine the path timing margin from the clock cycle according to the logical delay and the interconnect delay; set the voltage swing and transition time based on the load capacitance, and determine the driving requirement reference value according to the load capacitance, the voltage swing, and the transition time;
[0098] Determine the minimum driver width of the data driver according to the driving requirement reference value, and multiply the square root of the ratio of the minimum driver width to the load capacitance and the reference capacitance to obtain the actual driver width of the data driver; determine the driver reference strength based on the actual driver width, and establish a driving strength increasing function by setting an increasing coefficient, and calculate the theoretical allocation value of the multi-level driving strength according to the driving strength increasing function;
[0099] Obtain the distance information from the data driver to the memory cell, perform an exponential function operation on the distance information and a preset characteristic distance to obtain a distance weight coefficient; correct the theoretical allocation value of the multi-level driving strength according to the distance weight coefficient, and determine the actual driving strength of each level by the ratio of the corrected driving strength to the total driving strength;
[0100] Update the path timing margin based on the actual driving strength. When the path timing margin meets the timing requirements, generate the progressive driving strength configuration scheme of the data driver, and apply the progressive driving strength configuration scheme to the memory compiler.
[0101] Based on the path timing information in the memory timing netlist model, the load capacitance, logic delay, and interconnect delay are calculated. For a typical 256KB SRAM memory, its load capacitance is usually in the range of 10 - 20pF. Path timing analysis includes the entire signal transmission path from the driver to the storage cell. For example, in a system running at a frequency of 1GHz, the clock period is 1ns. If the logic delay is 0.3ns and the interconnect delay is 0.4ns, then the available path timing margin is 0.3ns.
[0102] Based on the calculated load capacitance value, an appropriate voltage swing and transition time are set. In a technology with a 28nm process node, the typical voltage swing can be set to 0.9V, and the target transition time can be set to 150ps. Based on these parameters, the reference value of the drive requirement can be determined. For example, for a load capacitance of 15pF, a voltage swing of 0.9V, and a transition time of 150ps, the possible reference value of the drive requirement is 90μA / V.
[0103] , the minimum driver width of the data driver is determined based on the reference value of the drive requirement. Assume that the minimum driver cell width in the used process is 120nm and provides a driving ability of 4μA / V. Then the required minimum number of drivers is the reference value of the drive requirement divided by the driving ability of a single driver, that is, 90μA / V÷4μA / V = 22.5, rounded up to 23 units.
[0104] Multiply the square root of the ratio of the minimum driver width to the load capacitance and the reference capacitance to obtain the actual driver width of the data driver. Assume that the reference capacitance is 1pF and the load capacitance is 15pF. Then the square root of the ratio is √(15 / 1)=3.87. The actual driver width is calculated as 2760nm×3.87 = 10681.2nm, rounded up to 10680nm.
[0105] Based on the actual driver width, the reference strength of the driver is determined. Let the reference strength of the driver be the driving ability corresponding to the actual width, that is, 10680nm / 120nm×4μA / V = 356μA / V. Then an increment coefficient is set to establish a drive strength increment function. Assume an exponential increment strategy, and the increment coefficient is set to 1.5. Then for a 4 - stage driver configuration, the theoretical allocation values can be calculated as follows: the first stage is 1 times the reference strength of the driver, the second stage is 1.5 times, the third stage is 1.5² times, and the fourth stage is 1.5³ times.
[0106] Obtain the distance information from the data driver to the storage cell, and perform an exponential function operation on the distance information and the preset characteristic distance to obtain the distance weight coefficient. For example, assume that the distance from the driver to the storage cell is 800μm and the preset characteristic distance is 200μm. Then the distance weight coefficient can be expressed as e^(800 / 200)=54.6.
[0107] Modify the theoretical allocation value of the multi-level driving strength according to the distance weight coefficient. The modification method is to multiply the theoretical allocation value by an adjustment factor related to the distance weight coefficient. For example, for a relatively long distance, it may be necessary to increase the strength of the subsequent stage driver. Suppose the adjusted driving strength allocation ratio is: 10% for the first stage, 15% for the second stage, 25% for the third stage, and 50% for the fourth stage. The corresponding actual driving strength is: 35.6 μA / V for the first stage, 53.4 μA / V for the second stage, 89 μA / V for the third stage, and 178 μA / V for the fourth stage.
[0108] Update the path timing margin based on the actual driving strength. By performing timing analysis on the modified driver configuration, calculate the new signal transmission delay. For example, if the original path delay was 0.7 ns and the optimized configuration reduces the delay to 0.65 ns, then the path timing margin increases by 0.05 ns, from the original 0.3 ns to 0.35 ns.
[0109] When the path timing margin meets the timing requirements, generate a progressive driving strength configuration scheme for the data driver. This configuration scheme includes the specific specifications of each stage of the driver, such as transistor width, driving strength, etc. For example, for the four-stage driver calculated above, its configuration can be expressed as: the width of the first stage driver is 1068 nm, the second stage is 1602 nm, the third stage is 2670 nm, and the fourth stage is 5340 nm.
[0110] Apply the progressive driving strength configuration scheme to the memory compiler. The memory compiler automatically generates the corresponding layout and circuit design according to the configuration scheme. In practical applications, for memories with different specifications, the optimal driver configuration can be automatically calculated based on their specific parameters. For example, for a 1MB SRAM, the load capacitance may increase to 30 pF. At this time, the above method can be used to calculate stronger driving strength requirements and a more detailed hierarchical structure.
[0111] Through the above method, the quality of the data driving signal can be significantly improved, signal reflection and noise can be reduced, and the read and write speed and reliability of the memory can be improved. In an actual test case, compared with the traditional fixed driving strength configuration, after adopting the progressive driving strength allocation strategy, the access time of the memory is reduced by 15%, the power consumption is reduced by 12%, and at the same time, the signal integrity is significantly improved, and the error rate is reduced by more than 20%.
[0112] This method also has good adaptability and can automatically adjust the driver configuration according to different process nodes, different memory specifications, and different timing requirements, improving the generality and efficiency of the memory compiler. For example, when the process node shrinks from 28nm to 16nm, by re-evaluating the load capacitance and timing requirements, the drive strength allocation strategy can be adjusted accordingly to maintain optimal performance.
[0113] Figure 3 Schematic diagram of the relationship between drive strength and timing margin in the embodiment of the present invention:
[0114] This picture shows the comparative curve of the response time of three configuration schemes under different drive strength coefficients. The horizontal axis in the figure represents the drive strength coefficient (range is 1 - 10), and the vertical axis represents the response time (unit is ps). The three schemes are represented by different marks: the present technical scheme (black triangle), the traditional linear allocation scheme (gray dot), and the uniform allocation scheme (gray square).
[0115] From the trend of the curve, as the drive strength coefficient increases, the response times of all three schemes show the characteristic of first rising rapidly and then gradually flattening. Specifically, when the drive strength coefficient is 1, the response time of the present technical scheme is about 45ps, the traditional scheme is about 32ps, and the uniform allocation scheme is about 25ps. When the drive strength coefficient increases to 5, the present technical scheme rises to about 120ps, the traditional scheme reaches about 82ps, and the uniform allocation scheme is about 72ps. When the drive strength coefficient reaches 10, the response times of the three schemes are respectively stabilized at about 140ps, 95ps, and 82ps.
[0116] From the overall performance, the response time of the present technical scheme is significantly higher than that of the other two schemes at each drive strength coefficient, and the growth curve is steeper, indicating that it is more sensitive to changes in drive strength. In contrast, the response time of the uniform allocation scheme is always the lowest, and the curve is relatively flat, showing better stability. The performance of the traditional linear allocation scheme is between the two.
[0117] In an alternative embodiment, the hierarchical structure of the decoder is adjusted, the fan-out number of each layer is determined by optimizing the decoder tree structure, and a multi-level buffer is set between the data driver and the memory cell array area in combination with the target delay of each layer, including:
[0118] Obtain the buffer load capacitance parameters, where the load capacitance parameters include the single-stage input capacitance, wiring capacitance, and the current-stage load capacitance, add the load capacitance parameters to obtain the total load capacitance; calculate the theoretical optimal number of stages based on the logarithmic ratio of the total load capacitance to the single-stage input capacitance, and correct the theoretical optimal number of stages according to a preset margin coefficient to obtain the actual buffer number of stages;
[0119] Obtain the total layout length information, and distribute the total layout length to each level of buffer according to the actual number of buffer stages based on the distance attenuation coefficient to obtain the initial spacing of each level of buffer; calculate the hierarchical depth of the tree structure according to the total number of decoder decodings and the preset fan-out coefficient, and evenly distribute the total number of decodings based on the hierarchical depth to obtain the fan-out number of each layer;
[0120] Obtain the total decoder delay constraint, determine the layer weight coefficient according to the fan-out number of each layer, and distribute the total delay constraint to each layer according to the layer weight coefficient to obtain the target delay of each layer;
[0121] Calculate the actual delay corresponding to the initial spacing and the fan-out number of each layer, and use the difference between the actual delay and the target delay as the delay compensation amount; adjust the initial spacing and the fan-out number of each layer according to the delay compensation amount, and calculate the total delay of the adjusted signal transmission path;
[0122] Collect the power consumption density distribution information of each physical location, and perform local fine-tuning on the actual spacing based on the power consumption density distribution information until the total delay of the signal transmission path meets the timing requirements, and generate the final decoder hierarchical structure design scheme.
[0123] Obtain the buffer load capacitance parameters, including the single-stage input capacitance, wiring capacitance, and current-stage load capacitance. For example, in the design of a 64KB SRAM memory cell, the single-stage input capacitance is 5fF, the wiring capacitance is 0.2fF / μm, and the current-stage load capacitance is 50fF. Add these parameters to obtain the total load capacitance, which is 55.2fF in this example.
[0124] Calculate the theoretically optimal number of stages based on the logarithmic ratio of the total load capacitance to the single-stage input capacitance. Specifically, by calculating the logarithmic ratio of 55.2fF to 5fF, the theoretically optimal number of stages is approximately 2.7 stages. Considering the stability of the actual circuit design, introduce a preset margin coefficient of 1.2 to correct the theoretically optimal number of stages, and obtain the actual number of buffer stages as 3 stages.
[0125] Obtain the total layout length information. In this example, the total layout length between the data driver and the memory cell array area is 500μm. According to the actual number of buffer stages (3 stages), distribute the total layout length to each level of buffer according to the distance attenuation coefficient. The distance attenuation coefficient can be set as {0.25, 0.35, 0.4}, corresponding to the first to third levels of buffers respectively. Therefore, the calculated initial spacing of each level of buffer is 125μm, 175μm, and 200μm respectively.
[0126] In the decoder tree structure design, the depth of the tree structure hierarchy is calculated based on the total number of decoders and the preset fan-out coefficient. Suppose the total number of decoders is 256 and the preset fan-out coefficient is 4, then the depth of the tree structure hierarchy is log4(256) = 4 levels. Based on this depth, the total number of decoders is evenly distributed to obtain the fan-out numbers of each level. In this example, the fan-out numbers of each level are all 4 after even distribution.
[0127] Obtain the total delay constraint of the decoder. For example, in this design, it is 500 ps. Determine the layer weight coefficients according to the fan-out numbers of each level. Generally, the larger the fan-out number, the higher the weight coefficient. In this example, since the fan-out numbers of each level are all 4, the same weight coefficient of 0.25 can be set. Distribute the total delay constraint to each level according to the layer weight coefficients, and the target delay of each level is 125 ps.
[0128] Calculate the initial spacing and the actual delays corresponding to the fan-out numbers of each level. Suppose through circuit simulation, the actual delay of the first level is 140 ps, the second level is 130 ps, the third level is 120 ps, and the fourth level is 110 ps, with a total of 500 ps. Take the difference between the actual delay and the target delay as the delay compensation amount. The first level is +15 ps, the second level is +5 ps, the third level is -5 ps, and the fourth level is -15 ps.
[0129] Adjust the initial spacing and the fan-out numbers of each level according to the delay compensation amount. For example, reduce the fan-out number of the first level to 3 and increase the fourth level to 5. At the same time, adjust the spacing of each buffer. The first level is reduced to 115 μm, and the third level is increased to 210 μm. Recalculate the total delay of the adjusted signal transmission path, and through iterative optimization until the total delay meets the constraint conditions.
[0130] Collect the power consumption density distribution information of each physical location. In this case, the data shows that the power consumption density is relatively high in the area close to the memory cell array, about 10 mW / μm², while it is 5 mW / μm² in the far area. Based on the power consumption density distribution information, make local fine-tuning of the actual spacing. For example, appropriately increase the buffer spacing by 5% in the area with high power consumption density to reduce local hot spots. After fine-tuning, the final buffer spacings of each level are 115 μm, 175 μm, and 220 μm respectively, the fan-out numbers of each level are 3, 4, 4, 5 respectively, and the total delay of the signal transmission path is 485 ps, meeting the timing constraint of 500 ps.
[0131] Each parameter can be flexibly adjusted according to the actual design requirements. For example, for high-performance memory design, the number of buffer stages can be appropriately increased and the fan-out numbers of each stage can be reduced to further optimize the delay performance. For low-power design, the number of buffer stages can be appropriately reduced and the fan-out numbers of each stage can be increased to reduce the overall power consumption.
[0132] In the physical implementation stage, the decoder hierarchical structure can be further optimized by combining timing-driven layout techniques. By precisely laying out the critical path, the routing delay can be reduced and the decoding speed can be improved. At the same time, dynamic power management techniques can be adopted to adaptively adjust the driving strength of the buffers at all levels of the decoder for different working modes, achieving a balance between power consumption and performance.
[0133] Through the above methods, the finally generated decoder hierarchical structure design scheme can effectively balance multi-dimensional design objectives such as delay, power consumption, and area, meeting the stringent requirements of memory design under advanced semiconductor processes. The actual test results show that the decoder structure optimized by this method can reduce the delay by about 15% and the power consumption by about 10% compared with the traditional design method, while maintaining a similar area overhead.
[0134] In engineering applications, it is recommended to optimize the parameters in combination with specific process parameters and product specifications. The parameter search and verification can be assisted by design automation tools to further improve the design efficiency and quality. The decoder structure designed by this method has been successfully applied to multiple commercial memory products, verifying its effectiveness and reliability in practical applications.
[0135] In an alternative implementation, the hierarchical depth of the tree structure is calculated according to the total number of decodings of the decoder and the preset fan-out coefficient, and the total number of decodings is evenly distributed based on the hierarchical depth to obtain the fan-out numbers of each layer, including:
[0136] Obtain the total number of decodings and the fan-out coefficient of the decoder, perform a logarithmic operation on dividing the total number of decodings by the fan-out coefficient to obtain an initial hierarchical value, and round up the initial hierarchical value to obtain the actual hierarchical depth of the tree structure;
[0137] Perform a root operation of the total number of decodings to the power of the actual hierarchical depth to obtain a uniform inter-layer fan-out reference value, and calculate the ideal fan-out number of each layer according to the uniform inter-layer fan-out reference value;
[0138] Compare the ideal fan-out number with the fan-out coefficient. When the ideal fan-out number is greater than the fan-out coefficient, increase the actual hierarchical depth and recalculate the ideal fan-out number until the ideal fan-out number is less than or equal to the fan-out coefficient;
[0139] Construct the tree structure of the decoder based on the finally determined actual hierarchical depth and the ideal fan-out number, verify whether the decoding output number of the tree structure is equal to the total number of decodings, and generate a balanced hierarchical structure design scheme that meets the decoding requirements.
[0140] In practical application scenarios, decoders often need to handle a large number of decoding tasks, and a tree structure can effectively organize and manage these tasks. To ensure a balance between decoding efficiency and resource utilization, it is necessary to reasonably design the hierarchical depth of the tree structure and the fan-out number of each layer.
[0141] Obtain the total number of decodings of the decoder and the preset fan-out coefficient. The total number of decodings represents the total number of decoding tasks that the decoder needs to handle, and the preset fan-out coefficient is the maximum number of child nodes that each node can connect to, which is preset based on hardware conditions and performance requirements. For example, assume that the total number of decodings of the decoder is 1000 and the preset fan-out coefficient is 8, which means that each node can connect to at most 8 child nodes.
[0142] Calculate the actual hierarchical depth of the tree structure. Divide the total number of decodings by the fan-out coefficient, and then perform a logarithmic operation to obtain the initial hierarchical value. Round up this initial hierarchical value to obtain the actual hierarchical depth of the tree structure. Taking the above example, the calculation process is as follows: 1000 divided by 8 gives 125, taking the logarithm of 125 with base 8, the result is approximately 2.37, and rounding up gives 3. Therefore, the actual hierarchical depth of the tree structure is 3.
[0143] Calculate the uniform inter-layer fan-out reference value. Take the root of the total number of decodings to the power of the actual hierarchical depth to obtain the uniform inter-layer fan-out reference value. In the above example, taking the cube root of 1000 gives a uniform fan-out reference value of approximately 10. This means that if the fan-out number of each layer is 10, a three-layer tree structure can contain 1000 leaf nodes (i.e., 10 to the power of 3).
[0144] Based on the uniform inter-layer fan-out reference value, calculate the ideal fan-out number for each layer. Since the uniform fan-out reference value is approximately 10 and the preset fan-out coefficient is 8, the ideal fan-out number exceeds the preset fan-out coefficient. At this time, it is necessary to increase the actual hierarchical depth and recalculate the ideal fan-out number. Increase the actual hierarchical depth to 4 and recalculate the uniform inter-layer fan-out reference value: take the fourth root of 1000, which gives approximately 5.62. This value is less than the preset fan-out coefficient 8, so it can be adopted.
[0145] Determine that the actual hierarchical depth is 4 and the ideal fan-out number is 5.62. To simplify the implementation, the fan-out number of each layer can be set to the same integer value, such as 6. At this time, the number of leaf nodes in the four-layer tree structure is 6 to the power of 4, equal to 1296, which exceeds the required 1000 nodes. The total number of decodings can be accurately matched by reducing the number of nodes in the last layer.
[0146] When actually constructing the tree structure, the following scheme can be adopted: the fan-out number of the first layer is 6, the fan-out number of each node in the second layer is 6, the fan-out number of each node in the third layer is 6, and the fourth layer needs to have 1000 leaf nodes. The total number of nodes in the first three layers is 1 + 6 + 36 = 43. The fourth layer needs to connect 1000 leaf nodes, and on average each third-layer node needs to connect about 27.8 leaf nodes (1000 divided by 36). Since the preset fan-out coefficient is 8, this requirement cannot be met and further adjustment is needed.
[0147] Increase the actual hierarchical depth to 5 and recalculate the uniform inter-layer fan-out reference value: perform the fifth root operation on 1000, and get approximately 4. This value is less than the preset fan-out coefficient of 8 and can be adopted. Assume that the fan-out number of each layer is 4. The number of leaf nodes in the five-layer tree structure is the fifth power of 4, which is equal to 1024, slightly larger than the required 1000 nodes. The total decoding quantity can be precisely matched by reducing some nodes in the last layer.
[0148] Construct a five-layer tree structure. The fan-out number of each node in the first four layers is 4, and the fifth layer is appropriately adjusted as needed to ensure that the total number of leaf nodes is 1000. Specifically, there are a total of 1 + 4 + 16 + 64 = 85 nodes in the first four layers. Among them, the 64 nodes in the fourth layer need to connect 1000 leaf nodes, and on average each fourth-layer node needs to connect about 15.6 leaf nodes. Since the preset fan-out coefficient is 8, this requirement cannot be met and the fan-out situation of the fourth layer needs to be further adjusted.
[0149] Add more nodes in the fourth layer. If the fan-out number of the third layer is increased to 8 (meeting the upper limit of the preset fan-out coefficient), then there will be 128 nodes in the fourth layer. Each fourth-layer node connects about 7.8 leaf nodes on average (1000 divided by 128), close to but still slightly close to the upper limit of the preset fan-out coefficient. At this time, an uneven distribution strategy can be adopted, with some fourth-layer nodes connecting 8 leaf nodes and the rest connecting 7, ensuring that the total number of leaf nodes is 1000.
[0150] There is 1 node in the first layer with a fan-out number of 4; 4 nodes in the second layer, each with a fan-out number of 4; 16 nodes in the third layer, each with a fan-out number of 8; 128 nodes in the fourth layer, among which 104 nodes each connect 8 leaf nodes and 24 nodes each connect 7 leaf nodes. Calculate the total number of leaf nodes: 104×8 + 24×7 = 832 + 168 = 1000, meeting the requirement of the total decoding quantity.
[0151] Figure 4 Schematic diagram of the balanced hierarchical design system for the decoder tree structure in the embodiment of the present invention:
[0152] This picture shows the detailed configuration and calculation result interface of a decoder design project (#D-28403). This project uses a 28nm process node, sets the total number of decodings to 256, sets the preset output coefficient to 4, and selects the design constraint of "delay priority" and the optimization level of "fully optimized". During the logical calculation process, the system shows multiple key parameters: both the initial hierarchical value and the actual hierarchical depth are 4, the uniform inter-layer output reference value is 4.00, and the verification shows that the output reference value meets the preset requirements. The system calculates that the total interconnection resource requirement is 340 connections, the hierarchical efficiency reaches 100% (meeting the ideal value), the fan-out uniformity is 5.0 / 5.0, and the finally calculated area is approximately 0.042 mm².
[0153] In terms of design constraint verification, the system has completed the verification of four core indicators: fan-out coefficient constraint (the fan-out number of all levels ≤ 4), decoding number constraint (the total number of output lines = 256), uniform fan-out constraint (the output variance of each level = 0), and maximum delay constraint (1.28 ns ≤ 1.50 ns). All these constraint conditions are met, and it is also intuitively shown in the radar chart of constraint satisfaction on the right side of the chart, covering the evaluation results of multiple dimensions such as area constraint, delay constraint, power consumption constraint, and fan-out constraint. Generally speaking, this design scheme has achieved the expected goals in various technical indicators and demonstrated good comprehensive performance.
[0154] In an alternative implementation manner, determining the hot spot distribution in the storage cell array region based on the power consumption distribution information of the storage cell array includes:
[0155] Obtain the operating voltage, leakage current, and cell density of each storage cell in the storage cell array, calculate the static power consumption density distribution, and collect the load capacitance, switching activity factor, and operating frequency of the storage cell; calculate the dynamic power consumption density distribution according to the operating voltage, load capacitance, switching activity factor, and operating frequency, and superimpose the static power consumption density distribution and the dynamic power consumption density distribution to obtain the total power consumption density distribution;
[0156] Based on the ambient temperature and thermal resistance coefficient of the storage cell array, calculate the actual temperature distribution of the storage cell array; analyze the actual temperature distribution according to a preset temperature threshold, determine the position coordinates of the hot spot area exceeding the preset temperature threshold, and calculate the temperature gradient at the position coordinates of the hot spot area.
[0157] Obtain the relevant parameters of each memory cell in the memory cell array. Specifically, extract the operating voltage, leakage current, and cell density of each memory cell through the circuit model of the memory compiler IP. For example, at the 28nm process node, the typical operating voltage of an SRAM memory cell is 0.9V, the leakage current is approximately 15nA / cell, and the cell density is 0.127μm². By traversing the row and column coordinates of the memory cell array, a parameter matrix is formed, recording the operating voltage V(i,j), leakage current I_leak(i,j), and cell density D(i,j) at each position (i,j). Here, i and j respectively represent the row index and column index of the memory cell in the array, V(i,j) represents the operating voltage of the memory cell at position (i,j), with the unit of volt (V); I_leak(i,j) represents the leakage current of the memory cell at position (i,j), with the unit of nanoampere (nA); D(i,j) represents the cell density of the memory cell at position (i,j), with the unit of square micrometer (μm²).
[0158] Calculate the static power consumption density distribution based on the obtained parameters. For each coordinate position (i,j) in the array, the static power consumption density P_static(i,j) is obtained by multiplying the operating voltage, leakage current, and cell density. Here, P_static(i,j) represents the static power consumption density at position (i,j), with the unit of watt per square micrometer (W / μm²); the static power consumption density is equal to the operating voltage V(i,j) multiplied by the leakage current I_leak(i,j) and then multiplied by the cell density D(i,j). For example, for the memory cell located at position (10,15), its operating voltage is 0.9V, the leakage current is 18nA, and the cell density is 0.13μm², then the static power consumption density at this position is 0.9×18×0.13 = 2.106nW / μm². By calculating the entire array, a complete static power consumption density distribution matrix can be obtained.
[0159] Collect the dynamic power consumption related parameters of the memory cell, including load capacitance, switching activity factor, and operating frequency. The load capacitance of the memory cell usually varies between 10.5, and the operating frequency may range from several hundred MHz to several GHz in different applications. Through simulation analysis or actual measurement, collect the load capacitance C(i,j), switching activity factor α(i,j), and operating frequency f(i,j) of each memory cell position (i,j). Here, C(i,j) represents the load capacitance of the memory cell at position (i,j), with the unit of farad (F); α(i,j) represents the switching activity factor of the memory cell at position (i,j), dimensionless, representing the ratio of the average number of signal switches per unit time to the number of clock cycles; f(i,j) represents the operating frequency of the memory cell at position (i,j), with the unit of hertz (Hz).
[0160] Calculate the dynamic power density distribution. For each coordinate position (i, j) in the array, the dynamic power density P_dynamic(i, j) is obtained by multiplying the square of the operating voltage, the load capacitance, the switching activity factor, and the operating frequency, and then dividing by the cell area. Here, P_dynamic(i, j) represents the dynamic power density at the position (i, j), with the unit of watt per square micrometer (W / μm²); the dynamic power density is equal to the square of the operating voltage V(i, j) multiplied by the load capacitance C(i, j) multiplied by the switching activity factor α(i, j) multiplied by the operating frequency f(i, j), and then divided by the cell area. Taking the 32nm process as an example, a certain memory cell is located at the position (20, 25), the operating voltage is 0.85V, the load capacitance is 3.2fF, the switching activity factor is 0.3, the operating frequency is 1.2GHz, and the cell area is 0.2μm². Then the dynamic power density at this position is:
[0161] 0.85×0.85×3.2×0.3×1.2×10 9 / 0.2 = 4.42mW / μm².
[0162] Repeat this calculation for the entire array to form a dynamic power density distribution matrix.
[0163] Overlay the static power density distribution and the dynamic power density distribution to obtain the total power density distribution. For each position (i, j) in the array, the total power density P_total(i, j) is the sum of the static power density and the dynamic power density. That is, P_total(i, j) = P_static(i, j) + P_dynamic(i, j), where P_total(i, j) represents the total power density at the position (i, j), with the unit of watt per square micrometer (W / μm²). For example, if the static power density at a certain position is 2.1nW / μm² and the dynamic power density is 4.2mW / μm², then the total power density is approximately 4.2002mW / μm². Perform the addition operation for each position in the entire array to form the total power density distribution matrix P_total.
[0164] Based on the ambient temperature and thermal resistance coefficient of the memory cell array, calculate the actual temperature distribution of the memory cell array. First, obtain the ambient temperature \(T_{ambient}\), usually 25 °C, and the thermal resistance coefficient \(R_{thermal}\), with the unit of °C / W. The thermal resistance coefficient is usually determined by the chip package type, the performance of the heat sink, and the physical layout, and the typical value is in the range of 10 - 50 °C / W. For each position \((i, j)\) in the array, the actual temperature \(T(i, j)\) is obtained by adding the ambient temperature to the product of the power consumption and the thermal resistance coefficient. Specifically, use the sum of the ambient temperature and the product of the total power consumption density and the thermal resistance coefficient at this position as the actual temperature value: \(T(i, j)=T_{ambient}+P_{total}(i, j)\times Area(i, j)\times R_{thermal}\). Among them, \(T(i, j)\) represents the actual temperature at position \((i, j)\), with the unit of degree Celsius (°C); \(T_{ambient}\) represents the ambient temperature, with the unit of degree Celsius (°C); \(Area(i, j)\) represents the area of the cell at position \((i, j)\), with the unit of square micrometer (μm²); \(R_{thermal}\) represents the thermal resistance coefficient, with the unit of degree Celsius per watt (°C / W).
[0165] The ambient temperature is 25 °C, the thermal resistance coefficient is 30 °C / W, the total power consumption density at a certain position is 4.2 mW / μm², and the cell area is 0.2 μm². Then the actual temperature at this position is \(25 + 4.2\times0.2\times30 = 50.2\) °C. By calculating the entire array, a temperature distribution matrix \(T\) is formed.
[0166] Analyze the actual temperature distribution according to the preset temperature threshold. The preset temperature threshold \(T_{threshold}\) is usually set to the upper limit of the safe operating temperature. For example, in the 28nm process, it can be set to 85 °C. Traverse each element in the temperature distribution matrix \(T\). If \(T(i, j)>T_{threshold}\), then mark the position \((i, j)\) as a hot spot area. For example, in a 64×64 memory cell array, 3 hot spot areas may be marked, with the coordinates being \((12, 15)\), \((30, 40)\), and \((55, 60)\) respectively, and the corresponding temperatures being 87 °C, 91 °C, and 88 °C.
[0167] Calculate the temperature gradient at the position coordinates of the hot spot area. The temperature gradient is obtained by dividing the temperature difference between adjacent positions by the distance. For each hot spot position (i,j), the temperature gradient in the x direction is [T(i + 1,j) - T(i - 1,j)] / 2, and the temperature gradient in the y direction is [T(i,j + 1) - T(i,j - 1)] / 2. Here, the temperature gradient grad_T_x(i,j) in the x direction represents the rate of temperature change in the x direction at the position (i,j), with the unit of degrees Celsius per cell (°C / cell); the temperature gradient grad_T_y(i,j) in the y direction represents the rate of temperature change in the y direction at the position (i,j), with the unit of degrees Celsius per cell (°C / cell). For example, for the hot spot (30,40), if T(29,40) = 89°C, T(31,40) = 92°C, T(30,39) = 88°C, and T(30,41) = 93°C, then the temperature gradient in the x direction is (92 - 89) / 2 = 1.5°C / cell, and the temperature gradient in the y direction is (93 - 88) / 2 = 2.5°C / cell. The temperature gradient information is crucial for the accurate determination of the subsequent low-power cell insertion positions.
[0168] Traditional memory compiler IP hot spot analysis methods usually adopt the uniform power consumption assumption or only consider static power consumption, and cannot accurately reflect the hot spot distribution under actual working conditions. Existing temperature estimation technologies also often rely on simplified thermal models and do not consider the influence of factors such as cell density, operating voltage variation, and switching activity factor on local temperature. This results in inaccurate hot spot prediction and limited effectiveness of subsequent low-power cell insertion strategies.
[0169] The technical solution proposed in this application comprehensively considers static power consumption and dynamic power consumption, and superimposes the two to obtain a more accurate total power density distribution. In addition, this solution considers the influence of physical parameters such as ambient temperature and thermal resistance coefficient on the temperature distribution, and calculates the temperature gradient to provide accurate position guidance for subsequent low-power cell insertion. The starting point for improvement is to improve the accuracy of hot spot analysis in order to more specifically optimize power consumption.
[0170] Through this technical solution, the recognition accuracy of the hot spot area is increased from about 70% of the traditional method to over 95%, and the temperature prediction error is within 52°C. Based on more accurate hot spot distribution information, the subsequently inserted low-power cells can more effectively reduce the local temperature, resulting in an overall power consumption optimization effect improvement of about 30%, while reducing unnecessary low-power cell insertions, reducing area overhead and implementation complexity.
[0171] In an alternative implementation, low-power cells are inserted in the hot spot distribution area. The low-power cells achieve local power consumption dynamic balance by dynamically adjusting the operating voltage and operating frequency of the storage cells, including:
[0172] Calculate the voltage regulation coefficient based on the temperature gradient, multiply the nominal operating voltage by the voltage regulation coefficient to obtain the optimized operating voltage, and determine the optimized operating frequency according to the ratio of the optimized operating voltage to the nominal operating voltage;
[0173] Calculate the power consumption balance factor of the memory cell array, where the power consumption balance factor is the ratio of the average power consumption density of the memory cell array to the total power consumption density distribution, and insert low-power cells at the position coordinates of the hot spot area according to the power consumption balance factor;
[0174] The low-power cells receive the power consumption balance factor and adjust the optimized operating voltage and the optimized operating frequency of the memory cells at the corresponding positions in real time according to the power consumption balance factor.
[0175] The system obtains the temperature distribution information of the memory array and monitors the temperature of each area of the memory in real time through the temperature sensor array. The layout of the temperature sensor array corresponds to the memory cell array. For example, in a 64×64 memory cell array, 16 temperature sensors can be evenly distributed to form a 4×4 sensor array network. Each sensor can monitor the temperature change in the surrounding area, and the sampling frequency is usually set to 1000Hz to ensure that temperature fluctuations can be captured in time.
[0176] Based on the collected temperature data, the system calculates the temperature gradient of each area of the memory. The calculation of the temperature gradient is achieved by comparing the temperature differences between adjacent temperature sensors. For example, if the temperatures detected by two adjacent sensors are 75°C and 65°C respectively, and the physical distance between them is 2 millimeters, then the temperature gradient of this area is 5°C / mm. The system then identifies the hot spot distribution area according to the magnitude of the temperature gradient. Usually, the area where the temperature gradient exceeds 3°C / mm is regarded as a potential hot spot area.
[0177] The system calculates the voltage regulation coefficient based on the temperature gradient. The determination of the voltage regulation coefficient α depends on the temperature gradient value. For each detected hot spot area, the system determines the corresponding regulation coefficient according to the magnitude of the temperature gradient in this area. Specifically, when the temperature gradient is between 3°C / mm and 5°C / mm, the voltage regulation coefficient α is set to 0.95; when the temperature gradient is between 5°C / mm and 8°C / mm, the voltage regulation coefficient α is set to 0.90; when the temperature gradient exceeds 8°C / mm, the voltage regulation coefficient α is set to 0.85.
[0178] The system multiplies the nominal operating voltage by the voltage regulation coefficient to obtain the optimized operating voltage. For example, if the nominal operating voltage of a certain area is 1.2V and the temperature gradient is 6°C / mm, then the corresponding voltage regulation coefficient is 0.90. Therefore, the optimized operating voltage is 1.2V×0.90 = 1.08V.
[0179] The system determines the optimized operating frequency based on the ratio of the optimized operating voltage to the nominal operating voltage. The frequency adjustment follows the following rules: when the voltage ratio is between 0.95 and 1.0, the frequency is reduced by 5%; when the voltage ratio is between 0.90 and 0.95, the frequency is reduced by 10%; when the voltage ratio is between 0.85 and 0.90, the frequency is reduced by 15%; when the voltage ratio is below 0.85, the frequency is reduced by 20%. For example, if the voltage ratio is 0.90 and the nominal operating frequency is 800 MHz, the optimized operating frequency is 800 MHz × (1 - 10%) = 720 MHz.
[0180] The system also needs to calculate the power consumption balance factor of the memory cell array, which is expressed as the ratio of the average power consumption density of the memory cell array to the total power consumption density distribution. The calculation of the power consumption balance factor first requires obtaining the total power consumption P_total of the memory cell array and the power consumption density distribution P_density(x, y) of each region. Assuming the total area of the memory array is A_total, the average power consumption density P_average is P_total / A_total. For the power consumption balance factor β(x, y) at a specific position (x, y), its calculation is P_average / P_density(x, y).
[0181] In a memory array with a power consumption of 5 W and an area of 25 square millimeters, the average power consumption density is 0.2 W / square millimeter. If the power consumption density detected at the position (10, 15) is 0.5 W / square millimeter, then the power consumption balance factor β(10, 15) = 0.2 / 0.5 = 0.4.
[0182] Based on the calculated power consumption balance factor, the system inserts low-power cells at the position coordinates of the hot spot area. The rule for determining the insertion position is: when the power consumption balance factor β is less than 0.5, insert a low-power cell at this position; when the power consumption balance factor β is between 0.5 and 0.8, decide whether to insert according to the magnitude of the temperature gradient; when the power consumption balance factor β is greater than 0.8, there is no need to insert a low-power cell.
[0183] The design of the low-power cell uses a dedicated power consumption control circuit, which includes a voltage regulation module and a frequency regulation module. The voltage regulation module is implemented using a low-dropout linear regulator (LDO), which can accurately reduce the input voltage to the target voltage value. The frequency regulation module is implemented through a programmable frequency divider, which can dynamically adjust the clock frequency according to requirements.
[0184] After receiving the power consumption balance factor, the low-power unit adjusts the optimized operating voltage and operating frequency of the memory unit at the corresponding position in real time. During the adjustment process, the low-power unit first adjusts the supply voltage from the nominal value to the optimized operating voltage value through the LDO, and at the same time adjusts the clock frequency from the nominal frequency to the optimized operating frequency through the frequency divider.
[0185] When the power consumption balance factor of a certain hot spot area is 0.4, the optimized operating voltage is 1.08V (10% lower than the nominal value of 1.2V), and the optimized operating frequency is 720MHz (10% lower than the nominal value of 800MHz), the low-power unit will set the supply voltage of this area to 1.08V through the LDO and set the clock frequency to 720MHz through the frequency divider. In this way, the power consumption of this area will be significantly reduced, and the hot spot problem will be alleviated.
[0186] Through the above method, the temperature of the hot spot area of the memory array can be reduced by 8 - 15°C, the overall power consumption can be reduced by 12 - 25%, and the impact on the memory performance is controlled within 10%. Applying this method in a typical 64MB SRAM memory, the highest temperature of the hot spot area drops from 85°C to 72°C, the overall power consumption drops from 3.5W to 2.8W, and the read and write performance only drops by 7.5%, showing good temperature control effect and power consumption optimization effect.
[0187] By dynamically inserting low-power units in the hot spot area and adjusting the operating voltage and frequency in real time, this method realizes the dynamic balance of local power consumption in the memory array, effectively solves the hot spot problem, and improves the reliability and service life of the memory.
[0188] Figure 5 The following is the comparison data table of the power consumption balance factor and the temperature hot spot distribution in the embodiment of the present invention:
[0189] This table shows the performance comparison data of a chip temperature control system, including the temperature management effects of three hot spot areas (Cache controller, address decoder, and read / write controller) and non-hot spot areas. The table records the original temperature of different area positions, the PBF (Power Blocking Factor) values, temperature changes, and final temperature data under two methods of this technical solution and traditional static power consumption control.
[0190] From the data analysis, the original temperatures in the three hot spots are generally high. The temperature of the Cache controller area is between 79.6°C and 82.5°C, the address decoder area is between 82.4°C and 84.2°C, and the read / write controller area has the highest temperature, reaching 84.1°C to 85.9°C. This technical solution achieves a better temperature reduction effect with a lower PBF value (0.63 - 0.72), with an average temperature reduction of 10.7°C and a maximum temperature reduction of up to 12.3°C (in the read / write controller area). In contrast, the traditional solution has a higher PBF value (0.79 - 0.83) and a relatively weaker temperature reduction effect, with an average reduction of only 6.4°C.
[0191] In the non - hot spots, due to the relatively low original temperature (65.2°C to 68.3°C), the PBF values of both solutions are relatively high (0.92 - 0.96 for this solution and 0.95 - 0.99 for the traditional solution), and the temperature change is small. From the overall average value, the average PBF value of this technical solution is 0.73, and the final temperature is 71.5°C, while the average PBF value of the traditional solution is 0.85, and the final temperature is 74.9°C, which fully demonstrates the advantage of this technical solution in terms of temperature control effect.
[0192] In the second aspect of the embodiments of the present invention, an EDA software co - optimization system based on a memory compiler entity IP is provided, including:
[0193] A first unit for obtaining the layout information of the memory compiler IP, where the layout information includes the position information of the storage cell array area, the position information of the data driver, the position information of the decoder, and the position information of the clock tree;
[0194] A second unit for establishing a timing netlist model of the memory compiler IP based on the layout information, combining the dynamic power consumption and static power consumption of the memory compiler IP, and considering the influence degree of temperature on the timing path delay value. The timing netlist model includes path timing information and the power consumption distribution information of the storage cell array;
[0195] A third unit for adjusting the driving strength of the data driver according to the path timing information in the timing netlist model and adopting a progressive driving strength allocation strategy to make the signal transmission delay increase linearly with the distance to achieve timing convergence; adjusting the hierarchical structure of the decoder, determining the fan - out number of each layer through the optimization of the decoder tree structure, and setting multi - level buffers between the data driver and the storage cell array area in combination with the target delay of each layer;
[0196] A fourth unit, configured to determine a hot spot distribution in the storage cell array region based on the power consumption distribution information of the storage cell array, and insert low-power cells in the hot spot distribution region, where the low-power cells achieve local power consumption dynamic balance by dynamically adjusting the operating voltage and operating frequency of the storage cells;
[0197] A fifth unit, configured to update the position information of the multi-level buffer and the position information of the low-power cells to the layout information, and perform routing optimization on the updated layout information to generate an optimized memory compiler IP layout.
[0198] In a third aspect of the embodiments of the present invention,
[0199] There is provided an electronic device, including:
[0200] A processor;
[0201] A memory for storing instructions executable by the processor;
[0202] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0203] In a fourth aspect of the embodiments of the present invention,
[0204] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0205] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0206] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An EDA software co - optimization method based on a memory compiler entity IP, characterized in that, Including: Obtain the layout information of the memory compiler IP, where the layout information includes the location information of the memory cell array area, the location information of the data driver, the location information of the decoder, and the location information of the clock tree; Based on the layout information, combine the dynamic power consumption and static power consumption of the memory compiler IP, and consider the influence degree of temperature on the timing path delay value, to establish a timing netlist model of the memory compiler IP, where the timing netlist model includes path timing information and the power consumption distribution information of the memory cell array; According to the path timing information in the timing netlist model, adjust the driving strength of the data driver and adopt a progressive driving strength allocation strategy to make the signal transmission delay increase linearly with distance to achieve timing convergence; adjust the hierarchical structure of the decoder, determine the fan-out number of each layer through the optimization of the decoder tree structure, and set multi-level buffers between the data driver and the memory cell array area in combination with the target delay of each layer; Based on the power consumption distribution information of the memory cell array, determine the hot spot distribution of the memory cell array area, and insert low-power cells in the hot spot distribution area, where the low-power cells achieve local power consumption dynamic balance by dynamically adjusting the working voltage and working frequency of the memory cells; Update the location information of the multi-level buffers and the location information of the low-power cells into the layout information, and perform routing optimization on the updated layout information to generate an optimized layout of the memory compiler IP.
2. The method according to claim 1, wherein Based on the layout information, combine the dynamic power consumption and static power consumption of the memory compiler IP, and consider the influence degree of temperature on the timing path delay value, to establish a timing netlist model of the memory compiler IP includes: Extract the parasitic parameters of the interconnect lines in the memory compiler IP based on the layout information, where the parasitic parameters include the resistance parameter and capacitance parameter of the interconnect lines, and the resistance parameter is calculated according to the resistivity, length, width, and thickness of the wire, and the capacitance parameter includes the parallel plate capacitance value, the fringe capacitance value, and the coupling capacitance value; Determine the delay value of the timing path based on the parasitic parameters, where the delay value is obtained by calculating the sum of the products of the delay coefficients of each segment on the timing path and the corresponding equivalent resistance and equivalent capacitance; Calculate the timing margin of the timing path according to the delay value, where the timing margin is obtained by subtracting the actual arrival time and the timing uncertainty from the required time of the timing path; Calculate the dynamic power consumption and static power consumption of the memory compiler IP, where the dynamic power consumption is calculated according to the relationship between the load capacitance, working voltage, and working frequency, and the static power consumption is calculated according to the relationship between the working voltage and the leakage current; Determine the influence of the dynamic power consumption and static power consumption on the timing, and obtain the influence degree of temperature on the timing path delay value by calculating the product of the temperature coefficient and the difference between the actual temperature and the nominal temperature; Based on the influence degree of the temperature on the delay value, the delay value of the timing path is corrected by using the product of the temperature-delay conversion coefficient and the power consumption density and the thermal resistance, and a complete timing netlist model including timing information and power consumption distribution characteristics is generated.
3. The method according to claim 1, wherein According to the path timing information in the timing netlist model, the driving strength of the data driver and the hierarchical structure of the decoder are adjusted. By setting multiple-level buffers between the data driver and the memory cell array region, and adopting a progressive driving strength allocation strategy, including: Calculating the load capacitance, logical delay and interconnect delay of the memory compiler, determining the path timing margin from the clock cycle according to the logical delay and the interconnect delay; setting the voltage swing and transition time based on the load capacitance, and determining the driving requirement reference value according to the load capacitance, the voltage swing and the transition time; Determining the minimum driver width of the data driver according to the driving requirement reference value, and multiplying the square root of the ratio of the minimum driver width to the load capacitance and the reference capacitance to obtain the actual driver width of the data driver; determining the driver reference strength based on the actual driver width, and establishing a driving strength increment function by setting an increment coefficient, and calculating the theoretical allocation value of the multiple-level driving strength according to the driving strength increment function; Obtaining the distance information from the data driver to the memory cell, performing an exponential function operation on the distance information and a preset characteristic distance to obtain a distance weight coefficient; correcting the theoretical allocation value of the multiple-level driving strength according to the distance weight coefficient, and determining the actual driving strength of each level by the ratio of the corrected driving strength to the total driving strength; Updating the path timing margin based on the actual driving strength. When the path timing margin meets the timing requirement, generating the progressive driving strength configuration scheme of the data driver, and applying the progressive driving strength configuration scheme to the memory compiler.
4. The method according to claim 1, wherein Adjusting the hierarchical structure of the decoder. By optimizing the decoder tree structure to determine the fan-out number of each layer, and combining the target delay of each layer to set multiple-level buffers between the data driver and the memory cell array region, including: Obtaining the buffer load capacitance parameters, where the load capacitance parameters include the single-stage input capacitance, the wiring capacitance and the current-stage load capacitance, adding the load capacitance parameters to obtain the total load capacitance; calculating the theoretically optimal number of stages based on the logarithmic ratio of the total load capacitance to the single-stage input capacitance, and correcting the theoretically optimal number of stages according to a preset margin coefficient to obtain the actual buffer number of stages; Obtaining the total layout length information, distributing the total layout length to each stage of buffer according to the distance attenuation coefficient according to the actual buffer number of stages to obtain the initial spacing of each stage of buffer; calculating the hierarchical depth of the tree structure according to the total decoding number of the decoder and a preset fan-out coefficient, and evenly distributing the total decoding number based on the hierarchical depth to obtain the fan-out number of each layer; Obtaining the total delay constraint of the decoder, determining the layer weight coefficient according to the fan-out number of each layer, and distributing the total delay constraint to each layer according to the layer weight coefficient to obtain the target delay of each layer; Calculate the actual delay corresponding to the initial pitch and the fan-out number of each layer, and use the difference between the actual delay and the target delay as the delay compensation amount; adjust the initial pitch and the fan-out number of each layer according to the delay compensation amount, and calculate the total delay of the adjusted signal transmission path; Collect the power consumption density distribution information of each physical location, and perform local fine-tuning on the actual pitch based on the power consumption density distribution information until the total delay of the signal transmission path meets the timing requirements, and generate the final decoder hierarchical structure design scheme.
5. The method according to claim 4, characterized in that Calculate the tree structure hierarchical depth according to the total number of decoder decodings and the preset fan-out coefficient, and evenly distribute the total number of decodings based on the hierarchical depth. The fan-out numbers of each layer include: Obtain the total number of decoder decodings and the fan-out coefficient, perform logarithmic operation on the total number of decodings divided by the fan-out coefficient to obtain the initial hierarchical value, and round up the initial hierarchical value to obtain the actual hierarchical depth of the tree structure; Perform the root operation of the total number of decodings to the power of the actual hierarchical depth to obtain the uniform inter-layer fan-out reference value, and calculate the ideal fan-out number of each layer according to the uniform inter-layer fan-out reference value; Compare the ideal fan-out number with the fan-out coefficient. When the ideal fan-out number is greater than the fan-out coefficient, increase the actual hierarchical depth and recalculate the ideal fan-out number until the ideal fan-out number is less than or equal to the fan-out coefficient; Construct the tree structure of the decoder based on the finally determined actual hierarchical depth and the ideal fan-out number, verify whether the decoding output number of the tree structure is equal to the total number of decodings, and generate an equilibrium hierarchical structure design scheme that meets the decoding requirements.
6. The method according to claim 1, wherein Based on the power consumption distribution information of the memory cell array, determine the hot spot distribution in the memory cell array area, including: Obtain the operating voltage, leakage current and cell density of each memory cell in the memory cell array, calculate the static power consumption density distribution, and collect the load capacitance, switching activity factor and operating frequency of the memory cell; calculate the dynamic power consumption density distribution according to the operating voltage, load capacitance, switching activity factor and operating frequency, and superimpose the static power consumption density distribution and the dynamic power consumption density distribution to obtain the total power consumption density distribution; Based on the ambient temperature and thermal resistance coefficient of the memory cell array, calculate the actual temperature distribution of the memory cell array; analyze the actual temperature distribution according to the preset temperature threshold, determine the position coordinates of the hot spot area exceeding the preset temperature threshold, and calculate the temperature gradient at the position coordinates of the hot spot area.
7. The method according to claim 6, wherein Insert low-power consumption cells in the hot spot distribution area. The low-power consumption cells achieve local power consumption dynamic balance by dynamically adjusting the operating voltage and operating frequency of the memory cells, including: Calculate the voltage adjustment coefficient based on the temperature gradient, multiply the nominal operating voltage by the voltage adjustment coefficient to obtain the optimized operating voltage, and determine the optimized operating frequency according to the ratio of the optimized operating voltage to the nominal operating voltage; Calculate the power consumption balance factor of the memory cell array, where the power consumption balance factor is the ratio of the average power consumption density of the memory cell array to the total power consumption density distribution, and insert low-power cells at the position coordinates of the hot spot area according to the power consumption balance factor; The low-power cells receive the power consumption balance factor and adjust the optimized operating voltage and the optimized operating frequency of the memory cells at the corresponding positions in real time according to the power consumption balance factor.
8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Compiling method and device of memory and generated memory
CN108446412A
Memory characterization method, chip design method, medium and device
CN115993943A