Integrated circuit chip particle time sequence optimization method and device, equipment and medium
Through logical hierarchical analysis and physical core particle division combined with the use of DREAMPlace tool, the problem of lack of hierarchical optimization and physical and timing optimization separation of core particle division in the existing technology has been solved, and efficient and accurate core particle timing optimization and performance balance have been achieved.
Patent Information
- Application Number
- CN202510379565.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing technology lacks a hierarchical optimization strategy in core particle division, which leads to the inability to effectively optimize core particles, and the physical and timing optimization separation may lead to poor timing performance when the layout scheme meets physical constraints.
Logical hierarchical analysis is used to obtain the performance parameters, power consumption parameters and area parameters of the integrated circuit module, perform preliminary division, and combine physical core particle division, use DREAMPlace tool to optimize core particle layout and timing, and iteratively adjust the division plan until the timing meets the requirements.
It improves the efficiency and accuracy of core particle timing optimization, and can comprehensively consider physical and timing constraints, balance the physical layout and timing performance of core particle, reduce performance trade-offs, and shorten the design cycle.
Smart Images

Figure CN120145993A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of integrated circuits, and particularly to a method for optimizing the timing of integrated circuit chiplets, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the continuous development of integrated circuit technology, Moore's Law is gradually approaching its physical limit, which poses unprecedented challenges and opportunities to the global semiconductor industry, especially the chip industry in China. In the era dominated by US enterprises, Moore's Law promoted the construction of the software and hardware ecosystems of its IC industry. However, with the continuous shrinking of technology nodes, the manufacturing cost has risen sharply, and problems such as chip yield, thermal management, poor scalability, verification challenges, and integration difficulties have become increasingly prominent.
[0003] In the post-Moore era, in order to continue Moore's Law and solve problems such as short-channel effects, high leakage currents, and subthreshold swing limitations faced by nanoscale devices, new device technologies marked by FinFET technology have emerged. The transition of CMOS devices from planar to three-dimensional FinFET is an inevitable choice to solve the power consumption problem, while traditional CMOS scaling may face an end. Devices with new principles, new structures, or new materials will break through the constraints of Moore's Law and bring new solutions for performance improvement.
[0004] Against this background, the research focus in the academic community in recent years has shifted to Chiplet technology, which has proposed a new chip building solution, structurally changing the future chip design direction from the bottom up. Chiplet technology optimally combines different functional chiplets to achieve the optimization of performance and cost. This modular design not only improves the flexibility and scalability of the design but also promotes innovation and diversification in chip design.
[0005] However, in the key technical link of chiplet partitioning and placement, the existing technologies mainly focus on how to make effective partitions according to the physical and timing constraints of chiplets to optimize the performance, power consumption, and area of the entire chip. However, the current chiplet partitioning is manually done based on design experience. This method lacks a hierarchical optimization strategy, resulting in the inability to effectively optimize chiplets at different levels. At the same time, physical and timing optimizations are separated, and there is a lack of a method that comprehensively considers both, which may lead to unsatisfactory timing performance when the placement scheme meets the physical constraints. In addition, existing technologies mostly adopt a one-time partitioning method, lacking an iterative optimization process and unable to make dynamic adjustments according to the timing results, making it difficult to achieve the optimal timing performance.
[0006] In summary, the current die partitioning in the prior art is all manually performed based on design experience. This method lacks a hierarchical optimization strategy, resulting in the inability to effectively optimize dies at different levels. At the same time, physical and timing optimizations are separated, and there is a lack of a method for comprehensively considering both, which may lead to problems such as unsatisfactory timing performance when the layout scheme meets physical constraints. In consideration of solving this problem, the applicant has made corresponding explorations. Summary of the Invention
[0007] An object of the present application is to solve the above problems and provide an integrated circuit die timing optimization method, a corresponding device, an electronic device, and a computer-readable storage medium.
[0008] To meet the various objectives of the present application, the following technical solutions are adopted:
[0009] An integrated circuit die timing optimization method proposed for one of the objectives of the present application includes:
[0010] In response to an instruction for integrated circuit die timing optimization, obtain the performance parameters, power consumption parameters, and area parameters of each integrated circuit module, and perform a preliminary partition of each integrated circuit module according to the performance parameters, power consumption parameters, and area parameters using logical hierarchical analysis to determine a preliminary logical partition scheme;
[0011] Based on the preliminary logical partition scheme, combined with physical die partitioning, generate a preliminary die partition scheme according to area constraints, call the DREAMPlace tool for die layout, and optimize the total negative timing slack and worst negative timing slack on all paths inside the die in the timing optimization module to determine the die backend layout result;
[0012] Based on the die backend layout result, analyze the timing of the die according to the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst case of setup time among all paths, and the negative timing slack in the worst case of hold time among all paths, and readjust the preliminary die partition scheme when violations occur to determine an iterative partition scheme until the timing of the die meets the requirements;
[0013] When the iterative partition scheme meets the timing optimization requirements, record and save the current iterative partition scheme, and continue to perform regular partitioning on the integrated circuit modules at the same level until all preliminary die partition schemes at this level are completely traversed.
[0014] Optionally, after the steps of obtaining the performance parameters, power consumption parameters, and area parameters of each integrated circuit module, and performing a preliminary division of the integrated circuit modules according to the performance parameters, power consumption parameters, and area parameters by using logical hierarchical analysis to determine a preliminary logical division scheme, the following steps are included:
[0015] Obtain a preliminary logical division scheme in the initial stage and use it as an initial classification set;
[0016] Integrate the preliminary logical division scheme into the initial classification set by using hierarchical information and the input module division restrictions, and combine physical die division to generate the preliminary die division scheme according to the area limit.
[0017] Optionally, the steps of analyzing the timing of the die according to the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, and readjusting the preliminary die division scheme when violations occur to determine an iterative division scheme until the timing of the die meets the requirements include:
[0018] Write and run the corresponding tcl script in the OpenSTA tool to retrieve the underlying files of the required environment and call the corresponding timing analysis module inside OpenSTA to perform the final timing analysis, and then obtain the corresponding quantitative results;
[0019] By analyzing the four data of the total negative timing slack of all paths generated by OpenSTA in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, determine whether the timing of the die reaches the ideal state.
[0020] Optionally, the steps of analyzing the four data of the total negative timing slack of all paths generated by OpenSTA in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths to determine whether the timing of the die reaches the ideal state include:
[0021] Judge whether the four data of the total negative timing slack of all paths generated by OpenSTA in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths are all positive;
[0022] If all are positive, it means that the timing of all paths meets the requirements, and the timing of the die reaches the ideal state.
[0023] If a negative number appears, it means that at least one path of the die is violated, and it is determined that the timing of the die does not reach the ideal state.
[0024] Optionally, when the iterative partitioning scheme meets the timing optimization requirements, record and save the current iterative partitioning scheme, and continue to perform regular partitioning on the integrated circuit modules at the same level until all preliminary die partitioning schemes at this level are completely traversed. The steps include:
[0025] If at least one path of the die is violated, return to the original preliminary logic partitioning scheme and readjust the partitioning strategy.
[0026] Re - execute the step of performing preliminary partitioning on each integrated circuit module according to the performance parameters, power consumption parameters, and area parameters using logical hierarchical analysis to determine the preliminary logic partitioning scheme.
[0027] Then, detect four data: the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst - case setup time among all paths, and the negative timing slack in the worst - case hold time among all paths, and so on until the timing of the die meets the requirements.
[0028] When the result reaches the optimum, the current partitioning scheme is recorded and saved, and continue to perform regular partitioning on the modules at the same level until all possible partitions at this level are completely traversed.
[0029] Optionally, after the step of recording and saving the current iterative partitioning scheme and continuing to perform regular partitioning on the integrated circuit modules at the same level until all preliminary die partitioning schemes at this level are completely traversed when the iterative partitioning scheme meets the timing optimization requirements, it includes:
[0030] Obtain the preset full - chip gate - level netlist file, process file, layout file, timing constraint file, physical constraint file, and partitioning constraint file.
[0031] According to the full - chip gate - level netlist file, process file, layout file, timing constraint file, physical constraint file, and partitioning constraint file, gradually optimize the timing of the total negative timing slack and the worst - case negative timing slack on all paths inside the die in hierarchical iteration to determine the corresponding layout file and die partitioning file of the die.
[0032] Optionally, when the iterative partitioning scheme meets the timing optimization requirements, record and save the current iterative partitioning scheme, and continue to perform regular partitioning on the integrated circuit modules at the same level until all the preliminary die partitioning schemes at this level are completely traversed. The steps include:
[0033] Calculate the connection ratio in the module set based on the floorplan data before input layout and the positions of the given chip I / Os as the preference candidate set index.
[0034] Combine the size of the die core area, the minimum area of any die, the utilization rate of each die, and the aspect ratio constraint, perform proportional allocation according to the total area of the partitioned modules, and add a slack coefficient during the allocation to achieve the generality of the layout scheme. Then, preferentially partition the modules with high preference candidate index values in adjacent parts and output the corresponding physical area scheme.
[0035] An integrated circuit die timing optimization device provided to meet another object of the present application includes:
[0036] A preliminary logic partitioning module, configured to respond to an instruction for performing die timing optimization on an integrated circuit, obtain the performance parameters, power consumption parameters, and area parameters of each integrated circuit module, and perform preliminary partitioning on each integrated circuit module according to the performance parameters, the power consumption parameters, and the area parameters by using logical hierarchical analysis to determine a preliminary logic partitioning scheme.
[0037] A die timing optimization module, configured to generate a preliminary die partitioning scheme based on the preliminary logic partitioning scheme in combination with physical die partitioning, call the DREAMPlace tool for die layout, and optimize the total negative timing slack and the worst negative timing slack on all paths inside the die in the die timing optimization module to determine the die back-end layout result.
[0038] A partitioning scheme adjustment module, configured to analyze the timing of the die based on the die back-end layout result according to the total negative timing slack of all paths in the setup time generated by the OpenSTA tool, the total negative timing slack of all paths in the hold time, the negative timing slack in the worst case of the setup time among all paths, and the negative timing slack in the worst case of the hold time among all paths, and readjust the preliminary die partitioning scheme when a violation occurs to determine an iterative partitioning scheme until the timing of the die meets the requirements.
[0039] A partitioning scheme traversal module, configured to record and save the current iterative partitioning scheme when the iterative partitioning scheme meets the timing optimization requirements, and continue to perform regular partitioning on the integrated circuit modules at the same level until all the preliminary die partitioning schemes at this level are completely traversed.
[0040] An electronic device provided to meet another object of the present application includes a central processing unit and a memory. The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the integrated circuit die timing optimization method described in the present application.
[0041] A computer-readable storage medium provided to meet another object of the present application stores a computer program implemented based on the integrated circuit die timing optimization method in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.
[0042] Compared with the prior art, in the prior art, the current die partitioning is all manually partitioned based on design experience. This method lacks a hierarchical optimization strategy, resulting in the inability to effectively optimize dies at different levels. At the same time, physical and timing optimizations are separated, and there is a lack of a method to comprehensively consider both, which may cause the timing performance to be unsatisfactory when the layout scheme meets physical constraints. The present application includes, but is not limited to, the following beneficial effects:
[0043] First, the present application can greatly improve the efficiency of integrated circuit die timing optimization. Existing die partitioning and layout optimization methods often rely on manual experience and simple optimization algorithms. This manual partitioning method has limitations, especially when facing complex designs and variable constraint conditions, with low efficiency and prone to missing the optimal solution. Through the hierarchical iteration strategy, the present application first makes a rough partition at a high level and then gradually delves into the low level for refinement, which can efficiently solve complex problems. This phased optimization method greatly improves the speed of die partitioning and also improves the accuracy during the die partitioning process;
[0044] Second, the integrated circuit die timing optimization method of the present application can comprehensively optimize physical and timing constraints. In the traditional die partitioning process, physical constraints and timing constraints are often considered separately. In this way, timing performance may be ignored while meeting the physical layout requirements, and vice versa, resulting in the final design being unable to meet both key requirements simultaneously. However, the present application can consider physical and timing constraints simultaneously. By comprehensively considering the optimizations of both during the design process, it can better balance the physical layout and timing performance of the die, thereby reducing the phenomenon of performance trade-off. This comprehensive consideration method ensures the feasibility of the design and achieves a more optimized overall performance, meeting the requirements for high performance and low power consumption in modern chip design;
[0045] Thirdly, the integrated circuit die timing optimization method of the present application can significantly improve automation and flexibility. Existing die partitioning methods usually require a large amount of manual intervention and manual adjustment. Although some optimizations can be made based on design experience, there are still limitations highly dependent on manual work and it is not easy to adapt to different design requirements and constraints. In contrast, the present application can quickly respond to changes in different designs and automatically complete partitioning and optimization. This not only improves the efficiency of the design process but also greatly enhances the flexibility of the system, enabling it to adapt to more diverse requirements and design environments. At the same time, it reduces manual intervention and the risk of human errors.
[0046] Fourthly, the integrated circuit die timing optimization method of the present application can greatly shorten the design cycle. In traditional design processes, especially in the stages of die partitioning and layout, it may take multiple iterations to find a relatively satisfactory solution, which is time-consuming and prone to design rework. By introducing hierarchical optimization and a strategy that comprehensively considers physical and timing constraints, the present application reduces the number of design iterations and can identify and solve potential problems at the initial stage of design. Through early optimization, the need for late-stage modifications is reduced, thus significantly shortening the design cycle. This has important commercial value for the rapidly developing semiconductor industry and can help enterprises gain an advantage in the highly competitive market.
[0047] Furthermore, compared with the prior art, the present application provides a significant improvement in the process of die partitioning and layout optimization by introducing key technologies such as hierarchical iterative optimization, comprehensive physical and timing constraints, and an automated design process. These improvements not only increase the design efficiency and quality but also shorten the design cycle, reduce human errors, and enhance the flexibility and adaptability of the design. These effects enable chip design to better cope with the challenges of the post-Moore era and meet the requirements of higher performance, lower power consumption, and more complex design needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0049] Figure 1 is a schematic flowchart of the integrated circuit die timing optimization method in the embodiment of the present application;
[0050] Figure 2 is an exemplary network architecture adopted by the integrated circuit die timing optimization system in the embodiment of the present application;
[0051] Figure 3 is a schematic diagram showing modules with similar functions placed in the same die in the embodiment of the present application;
[0052] Figure 4It is a principle block diagram of an integrated circuit die timing optimization device in an embodiment of the present application;
[0053] Figure 5 It is a structural schematic diagram of a computer device in an embodiment of the present application. Specific embodiments
[0054] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.
[0055] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0056] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0057] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with wireless signal receivers that only have the ability to receive and no ability to transmit, and devices with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or cellular or other communication devices without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed form at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be devices such as a smart TV, a set-top box, etc.
[0058] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.
[0059] It should be noted that the concept of "server" in this application can similarly be extended to the case applicable to a server cluster. According to the network deployment principle understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through interfaces, or integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this variation and should not be restricted by this in the implementation manner of the network deployment method of this application.
[0060] One or several technical features of this application, unless expressly specified, can either be deployed on the server for implementation and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for implementation and access.
[0061] The neural network models cited or possibly cited in this application, unless expressly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client capable of handling the device for direct invocation. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.
[0062] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.
[0063] Those skilled in the art should be aware of this: Although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for each embodiment disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.
[0064] For each embodiment to be disclosed in this application, unless expressly pointed out that there is a mutually exclusive relationship between them, otherwise, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this variation.
[0065] Please refer to Figure 1 , in one embodiment of the integrated circuit die timing optimization method of this application, it includes:
[0066] Step S10: In response to an instruction for optimizing the core chip timing of an integrated circuit, obtain the performance parameters, power consumption parameters, and area parameters of each integrated circuit module, and perform a preliminary division of each integrated circuit module according to the performance parameters, the power consumption parameters, and the area parameters by using logical hierarchical analysis to determine a preliminary logical division scheme;
[0067] Please refer to Figure 2 , the core chip timing optimization system in the terminal device can respond to an instruction for optimizing the core chip timing of an integrated circuit, obtain the performance parameters, power consumption parameters, and area parameters of each integrated circuit module, and perform a preliminary division of each integrated circuit module according to the performance parameters, the power consumption parameters, and the area parameters by using logical hierarchical analysis to determine a preliminary logical division scheme;
[0068] In some embodiments, the hierarchical iterative timing optimization architecture of the core chip mainly includes: The hierarchical iterative timing optimization architecture of the core chip is a precise integrated circuit design process aimed at optimizing performance, power consumption, and area (PPA). This architecture improves the efficiency and quality of core chip design through modules that work together. Specifically, the information acquisition module is responsible for processing input data and constraints, including data format conversion and information collation, to ensure the accuracy of subsequent processing. The division module allocates the modules in the def file to the corresponding Chiplets according to the predefined hierarchical structure and constraints, constructing the basic framework for modular design and optimization. The layout module performs timing optimization on the preliminary division scheme to ensure that the physical placement scheme meets the timing requirements of signal transmission, so as to reduce latency and improve the overall performance. The iterative module further refines on this basis, optimizing the division scheme through hierarchical iteration to obtain the optimal backend division and module placement scheme. In addition, an interruption mechanism for optimizing the number of iterations is integrated into the architecture, which can automatically terminate the iteration process when preset conditions are met, thereby effectively controlling the design cycle and cost while ensuring the design quality.
[0069] Step S20: Based on the preliminary logical division scheme, combined with physical core chip division, generate a preliminary core chip division scheme according to area constraints, call the DREAMPlace tool for core chip layout, and optimize the total negative timing slack and the worst negative timing slack on all paths inside the core chip in the timing optimization module to determine the core chip backend layout result;
[0070] Obtain the performance parameters, power consumption parameters, and area parameters of each integrated circuit module. After performing a preliminary division of each integrated circuit module according to the performance parameters, power consumption parameters, and area parameters using logical hierarchical analysis to determine a preliminary logical division scheme, based on the preliminary logical division scheme, combined with physical die partitioning, generate a preliminary die partitioning scheme according to the area limit. Invoke the DREAMPlace tool for die placement, and optimize the total negative timing slack and worst negative timing slack on all paths inside the die in the timing optimization module to determine the die backend placement result;
[0071] In some embodiments, after the step of obtaining the performance parameters, power consumption parameters, and area parameters of each integrated circuit module and performing a preliminary division of each integrated circuit module according to the performance parameters, power consumption parameters, and area parameters using logical hierarchical analysis to determine a preliminary logical division scheme, it includes:
[0072] Step S201: Obtain a preliminary logical division scheme in the initial stage and use it as an initial classification set;
[0073] Step S202: Integrate the preliminary logical division scheme into the initial classification set using hierarchical information and the input module division limit, and combine physical die partitioning to generate the preliminary die partitioning scheme according to the area limit.
[0074] Specifically, in the front-end design stage, this architecture performs a preliminary division of the integrated circuit module through logical hierarchical analysis to meet the performance, power consumption, and area (PPA) standards, obtains the preliminary logical division scheme in the initial stage, and uses it as the initial classification set; by using hierarchical information, the algorithm can effectively compress the search space and reduce the computational complexity; then the algorithm integrates the logical division scheme into the initial classification set according to the input module division limit, and combines physical die partitioning to generate an initial die partitioning scheme according to the area limit.
[0075] After obtaining the preliminary division scheme, use the DREAMPlace tool for die placement in combination with the process files (lef, lib) to perform physical placement on the divided die scheme. At the same time, the built-in timing optimization module optimizes the TNS (Total Negative Slack) and WNS (Worst Negative Slack) inside the die to obtain the die backend placement result, where TNS (Total Negative Slack) represents the total negative timing slack on all paths inside the die, and WNS (Worst Negative Slack) represents the worst negative timing slack on all paths inside the die.
[0076] Step S30: Based on the die backend layout result, analyze the timing of the die according to the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths. When a violation occurs, readjust the preliminary die partitioning scheme to determine an iterative partitioning scheme until the timing of the die meets the requirements;
[0077] Based on the preliminary logic partitioning scheme, combine physical die partitioning, generate a preliminary die partitioning scheme according to the area limit, call the DREAMPlace tool for die placement, and optimize the total negative timing slack and the worst negative timing slack on all paths inside the die in the timing optimization module to determine the die backend layout result. Then, based on the die backend layout result, analyze the timing of the die according to the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths. When a violation occurs, readjust the preliminary die partitioning scheme to determine an iterative partitioning scheme until the timing of the die meets the requirements;
[0078] In some embodiments, the step of analyzing the timing of the die according to the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, and when a violation occurs, readjusting the preliminary die partitioning scheme to determine an iterative partitioning scheme until the timing of the die meets the requirements includes:
[0079] Step S301: Write and run the corresponding tcl script in the OpenSTA tool to retrieve the underlying files of the required environment and call the corresponding timing analysis module inside OpenSTA to perform the final timing analysis, and then obtain the corresponding quantitative results;
[0080] Step S302: By analyzing the four data of the total negative timing slack of all paths generated by OpenSTA in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, determine whether the timing of the die reaches the ideal state.
[0081] In a further embodiment, the step of determining whether the timing of the die reaches an ideal state by analyzing four data: the total negative timing slack of all paths generated by OpenSTA in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, includes:
[0082] Step S3021: Determine whether all four data, namely, the total negative timing slack of all paths generated by OpenSTA in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, are positive.
[0083] Step S3022: If all are positive, it means that the timing of all paths meets the requirements and the timing of the die reaches an ideal state.
[0084] Step S3023: If a negative number appears, it means that at least one path of the die is violated, and thus it is determined that the timing of the die does not reach an ideal state.
[0085] Specifically, write and apply the corresponding tcl script in the OpenSTA tool to retrieve the underlying files of the required environment and call the corresponding timing analysis module inside OpenSTA for the final timing analysis, then obtain the corresponding quantitative results. By analyzing the four data of TNS_early, TNS_late, WNS_early, and WNS_late generated by OpenSTA, determine whether the timing of the die reaches an ideal state, providing an accurate timing basis for iterative optimization. Among them, WNS_late represents the negative timing slack in the worst-case setup time among all paths, TNS_late represents the total negative timing slack of all paths in setup time, WNS_early represents the negative timing slack in the worst-case hold time among all paths, and TNS_early represents the total negative timing slack of all paths in hold time.
[0086] Determine whether all four data of TNS_early, TNS_late, WNS_early, and WNS_late generated by the OpenSTA tool are positive. If all are positive, it means that the timing of all paths of the die meets the requirements. If a negative number appears, it means that at least one path of the die is violated, and thus determine whether the timing of the die reaches an ideal state, providing an accurate timing basis for iterative optimization.
[0087] In the die partitioning stage, the hierarchical iteration is completed, allowing for layout optimization across dies. The iterative algorithm utilizes the module hierarchical information, taking the hierarchy of the module information as set information, and partitioning each partition into sets according to the hierarchical information. As Figure 3 shown, the hierarchical modules form sets layer by layer from low to high, gradually merging multiple low-level module sets, and placing modules with similar functions together within the same die to optimize the timing performance.
[0088] Step S40: When the iterative partitioning scheme meets the timing optimization requirements, record and save the current iterative partitioning scheme, and continue to perform regular partitioning on the integrated circuit modules at the same level until all the preliminary die partitioning schemes at this level are completely traversed.
[0089] Based on the die back-end layout result, analyze the timing of the die according to the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths. When a violation occurs, readjust the preliminary die partitioning scheme to determine the iterative partitioning scheme. After the timing of the die meets the requirements, when the iterative partitioning scheme meets the timing optimization requirements, record and save the current iterative partitioning scheme, and continue to perform regular partitioning on the integrated circuit modules at the same level until all the preliminary die partitioning schemes at this level are completely traversed.
[0090] In some embodiments, the step of, when the iterative partitioning scheme meets the timing optimization requirements, recording and saving the current iterative partitioning scheme, and continuing to perform regular partitioning on the integrated circuit modules at the same level until all the preliminary die partitioning schemes at this level are completely traversed includes:
[0091] Step S401: If there is at least one path violation in the die, return to the initial preliminary logic partitioning scheme and readjust the partitioning strategy.
[0092] Step S402: Re-execute the step of performing preliminary partitioning on each integrated circuit module according to the performance parameters, the power consumption parameters, and the area parameters using logical hierarchical analysis to determine the preliminary logic partitioning scheme.
[0093] Step S403: Then detect the four data of the total negative timing slack of all paths generated by the OpenSTA tool in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, and so on until the timing of the die meets the requirements.
[0094] Step S404: When the result reaches the optimum, the current partitioning scheme is recorded and saved, and the regular partitioning of modules at the same level continues until all possible partitions at this level are fully traversed.
[0095] In a further embodiment, after the step of recording and saving the current iterative partitioning scheme when the iterative partitioning scheme meets the timing optimization requirements and continuing the regular partitioning of the integrated circuit modules at the same level until all preliminary die partitioning schemes at this level are fully traversed, it includes:
[0096] Step S4001: Obtain a preset full-chip gate circuit netlist file, process file, layout file, timing constraint file, physical constraint file, and partitioning constraint file;
[0097] Step S4002: According to the full-chip gate circuit netlist file, process file, layout file, timing constraint file, physical constraint file, and partitioning constraint file, gradually optimize the total negative timing margin and the worst negative timing margin on all paths inside the die in hierarchical iteration to determine the corresponding layout file and die partitioning file of the die.
[0098] In a further embodiment, the step of recording and saving the current iterative partitioning scheme when the iterative partitioning scheme meets the timing optimization requirements and continuing the regular partitioning of the integrated circuit modules at the same level until all preliminary die partitioning schemes at this level are fully traversed includes:
[0099] Step S100: Calculate the connection ratio in the module set based on the floorplan data before inputting the layout and the positions of the given chip I / Os as the preference candidate set index;
[0100] Step S200: Combine the size of the die core area, the minimum area of any die, the utilization rate of each die, and the aspect ratio constraint, allocate proportionally according to the total area of the partitioning modules, and add a relaxation coefficient during the allocation to achieve the generality of the layout scheme. Then, preferentially partition the modules with high preference candidate index values in adjacent parts and output the corresponding physical area scheme.
[0101] Specifically, after forming the preliminary combination schemes, these schemes are incorporated into the candidate set and iterated in the compressed hierarchical design space to explore the partitioning schemes of each die. This iterative process aims to achieve cross-die timing optimization. In each iteration, the partitioning scheme is laid out by the DREAMPlace tool and the corresponding judgment indexes are generated.
[0102] When the result reaches the optimum, the current partitioning scheme is recorded and saved. Subsequently, the algorithm continues to perform regular partitioning on modules at the same level until all possible partitions at this level have been fully traversed. This process achieves rapid convergence of the partitioning boundary based on the hierarchical structure, providing a basis for the iterative advancement of the partitioning scheme to a higher-level stage. At the same time, this optimization process also involves optimizing the layout of the within-die and cross-die collaborations within the die system, considering the clustering characteristics of the die system. During the entire iterative process, this architecture comprehensively considers the full-chip gate-level netlist file, process file, layout file, timing constraint file, physical constraint file, and partitioning constraint file, and realizes the gradual timing optimization of TNS and WNS in hierarchical iteration.
[0103] By analyzing four data: the total negative timing slack of all paths in setup time generated by the OpenSTA tool, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, if a violation occurs, it returns to the initial partitioning scheme module and readjusts the partitioning strategy.
[0104] Through the above steps S3021 to S3023, then detect the total negative timing slack of all paths in setup time, the total negative timing slack of all paths in hold time, the negative timing slack in the worst-case setup time among all paths, and the negative timing slack in the worst-case hold time among all paths, and so on, until the timing of the die meets the requirements.
[0105] When the result reaches the optimum, the current partitioning scheme is recorded and saved. Subsequently, the algorithm continues to perform regular partitioning on modules at the same level until all possible partitions at this level have been fully traversed. This process achieves rapid convergence of the partitioning boundary based on the hierarchical structure, providing a basis for the iterative advancement of the partitioning scheme to a higher-level stage. At the same time, this optimization process also involves optimizing the layout of the within-die and cross-die collaborations within the die system, considering the clustering characteristics of the die system. During the entire iterative process, this architecture comprehensively considers the layout file, gate-level netlist information, process file, and physical and timing constraint files, and realizes the gradual timing optimization of TNS and WNS in hierarchical iteration.
[0106] In the area partitioning scheme, the connection ratio in the module set is calculated based on the floorplan data before the input layout and the positions of the given chip I / Os, serving as the preference candidate set index. Then, combined with the constraints of area information (the size of the die core area, the minimum area of any die, the utilization rate and aspect ratio constraints of each die), proportional allocation is performed according to the total area of the partitioned modules, and a slack coefficient is added during the allocation to achieve the generality of the layout scheme. Modules with higher preference candidate index values are preferentially partitioned into adjacent parts, and the corresponding physical area scheme is output.
[0107] To balance the optimization iteration time and the backend design time, the algorithm designs a stopping condition, that is, it automatically stops after iterating to the preset number of times, obtains the current optimal partitioning scheme as the timing optimization scheme for the modules inside and outside the die. This method not only improves the design efficiency but also ensures the design quality and performance.
[0108] In some embodiments, the full-chip gate-level netlist file contains the gate-level netlist of the entire chip, and the netlist is organized hierarchically by module; the process file contains the technical parameters and timing models involved in the design; the layout file contains the floorplan data before the layout and the positions of the given chip I / Os; the timing constraint file contains all relevant timing constraints; the physical constraint file contains the size of the die core area, the minimum area of any die, the utilization rate and aspect ratio constraints of each die; the partitioning constraint file contains the partitioning constraints of the die; the layout file is used to specify the physical cell placement data after the layout is completed; the die partitioning file is used to specify the shape and position information of each die, as well as the integrated circuit modules included in each die.
[0109] In some embodiments, the DREAMPlace tool is a deep learning-based integrated circuit (IC) physical layout tool, mainly used in the die placement stage. It is developed by researchers at Stanford University in the United States, aiming to improve the layout optimization efficiency and quality in IC design. The tool uses deep learning (especially deep reinforcement learning) to optimize the chip layout process, aiming to improve the optimization effects in aspects such as chip performance, power consumption, and area.
[0110] OpenSTA (Open Source Static Timing Analysis) is an open-source static timing analysis tool widely used in integrated circuit (IC) design, especially in the timing verification stage of digital circuit design. Its main function is to help designers analyze and verify the timing characteristics of the circuit to ensure that the circuit can work stably under the given clock constraints.
[0111] Floorplan data refers to the layout planning data in integrated circuit (IC) design, which describes the physical locations and arrangements of various functional modules within the chip (such as logic units, memories, input / output ports, etc.). In IC design, floorplan is the first step of the entire physical design, providing a basic framework for subsequent steps of the chip (such as placement, routing, timing analysis, etc.).
[0112] As can be seen from the above embodiments, compared with the prior art, in the prior art, the current chiplet partitioning is all manually performed based on design experience. This method lacks a hierarchical optimization strategy, resulting in the inability to effectively optimize chiplets at different levels. At the same time, physical and timing optimizations are separated, and there is a lack of a method to comprehensively consider both, which may lead to problems such as unsatisfactory timing performance when the layout scheme meets physical constraints. The present application includes, but is not limited to, the following beneficial effects:
[0113] First, the present application can greatly improve the efficiency of timing optimization of integrated circuit chiplets. Existing chiplet partitioning and layout optimization methods often rely on manual experience and simple optimization algorithms. This manual partitioning method has limitations, especially when facing complex designs and variable constraint conditions, with low efficiency and prone to missing the optimal solution. Through the hierarchical iterative strategy, the present application first makes a rough partition at a high level and then gradually delves into the low level for refinement, capable of efficiently solving complex problems. This phased optimization method greatly improves the speed of chiplet partitioning and also improves the accuracy during the chiplet partitioning process;
[0114] Second, the integrated circuit chiplet timing optimization method of the present application can comprehensively optimize physical and timing constraints. In the traditional chiplet partitioning process, physical constraints and timing constraints are often considered separately. In this way, when meeting the physical layout requirements, timing performance may be ignored, and vice versa, resulting in the final design being unable to meet both key requirements simultaneously. However, the present application can consider physical and timing constraints simultaneously. By comprehensively considering the optimizations of both during the design process, it can better balance the physical layout and timing performance of chiplets, thereby reducing the phenomenon of performance compromise. This comprehensive consideration method ensures the feasibility of the design and achieves a more optimized overall performance, meeting the requirements for high performance and low power consumption in modern chip design;
[0115] Thirdly, the integrated circuit die timing optimization method of the present application can significantly improve automation and flexibility. Existing die partitioning methods usually require a large amount of manual intervention and manual adjustment. Although some optimizations can be made based on design experience, there are still limitations that highly rely on manual work and it is not easy to adapt to different design requirements and constraints. In contrast, the present application can quickly respond to changes in different designs and automatically complete partitioning and optimization. This not only improves the efficiency of the design process but also greatly enhances the flexibility of the system, enabling it to adapt to more diverse requirements and design environments. At the same time, it reduces manual intervention and the risk of human errors.
[0116] Fourthly, the integrated circuit die timing optimization method of the present application can greatly shorten the design cycle. In the traditional design process, especially in the stages of die partitioning and layout, it may take multiple iterations to find a relatively satisfactory solution, which is time-consuming and prone to design rework. By introducing hierarchical optimization and a strategy that comprehensively considers physical and timing constraints, the present application reduces the number of design iterations and can identify and solve potential problems at the early stage of design. Through early optimization, the need for late-stage modification is reduced, thus significantly shortening the design cycle. This has important commercial value for the rapidly developing semiconductor industry and can help enterprises gain an advantage in the highly competitive market.
[0117] Furthermore, compared with the prior art, the present application provides a significant improvement in the die partitioning and layout optimization process by introducing key technologies such as hierarchical iterative optimization, comprehensive physical and timing constraints, and an automated design process. These improvements not only increase the design efficiency and quality but also shorten the design cycle, reduce human errors, and enhance the flexibility and adaptability of the design. These effects enable chip design to better cope with the challenges of the post-Moore era and meet the requirements of higher performance, lower power consumption, and more complex design needs.
[0118] Please refer to Figure 4, An integrated circuit die timing optimization device provided to meet one of the purposes of the present application, including a preliminary logic division module 1100, a die timing optimization module 1200, a division scheme adjustment module 1300, and a division scheme traversal module 1400. Among them, the preliminary logic division module 1100 is configured to respond to an instruction for optimizing the timing of integrated circuit dies, obtain the performance parameters, power consumption parameters, and area parameters of each integrated circuit module, and perform a preliminary division of each integrated circuit module according to the performance parameters, the power consumption parameters, and the area parameters by using logical hierarchical analysis to determine a preliminary logic division scheme; the die timing optimization module 1200 is configured to generate a preliminary die division scheme based on the preliminary logic division scheme, in combination with physical die division, generate a preliminary die division scheme according to area constraints, call the DREAMPlace tool for die placement, and optimize the total negative timing slack and the worst negative timing slack on all paths inside the die in the timing optimization module to determine the die backend placement result; the division scheme adjustment module 1300 is configured to analyze the timing of the die based on the die backend placement result according to the total negative timing slack of all paths generated by the OpenSTA tool in the setup time, the total negative timing slack of all paths in the hold time, the negative timing slack in the worst case of the setup time among all paths, and the negative timing slack in the worst case of the hold time among all paths, and readjust the preliminary die division scheme when a violation occurs to determine an iterative division scheme until the timing of the die meets the requirements; the division scheme traversal module 1400 is configured to record and save the current iterative division scheme when the iterative division scheme meets the timing optimization requirements, and continue to perform regular division on the integrated circuit modules at the same level until all the preliminary die division schemes at this level are completely traversed.
[0119] Based on any embodiment of the present application, please refer to Figure 5 , Another embodiment of the present application further provides an electronic device, which can be implemented by a computer device, as Figure 5 shown, the internal structure schematic diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement an integrated circuit die timing optimization method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the integrated circuit die timing optimization method of the present application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand,Figure 5 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0120] In this embodiment, the processor is used to execute Figure 4 the specific functions of each module in. The memory stores the program codes and various types of data required to execute the above-mentioned modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules in the integrated circuit die timing optimization device of this application, and the server can call the program codes and data of the server to execute the functions of all modules.
[0121] This application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the integrated circuit die timing optimization method according to any embodiment of this application.
[0122] This application also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by one or more processors, the steps of the integrated circuit die timing optimization method according to any embodiment of this application are implemented.
[0123] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of this application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium may be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0124] The above are only some embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for optimizing timing of integrated circuit chips, characterized in that: include: In response to an instruction to optimize chip timing of an integrated circuit, performance parameters, power consumption parameters and area parameters of each integrated circuit module are obtained, and each integrated circuit module is preliminarily divided according to the performance parameters, the power consumption parameters and the area parameters by using logic level analysis to determine a preliminary logic division scheme; Based on the preliminary logic partitioning scheme, combined with the physical core particle partitioning, a preliminary core particle partitioning scheme is generated according to the area limit, the DREAMPlace tool is called to perform core particle layout, and the total negative timing margin and the worst negative timing margin on all paths inside the core particle are optimized in the timing optimization module to determine the core particle back-end layout result; Based on the back-end layout result of the coregrain, the timing of the coregrain is analyzed according to the total negative timing margin of all paths in setup time, the total negative timing margin of all paths in hold time, the negative timing margin of the worst-case setup time in all paths, and the negative timing margin of the worst-case hold time in all paths generated by the OpenSTA tool, and when a violation occurs, the preliminary coregrain partitioning scheme is readjusted to determine an iterative partitioning scheme until the timing of the coregrain meets the requirements; When the iterative partitioning scheme meets the timing optimization requirement, the current iterative partitioning scheme is recorded and saved, and the integrated circuit modules at the same level are continuously partitioned regularly until all preliminary core grain partitioning schemes at the level are completely traversed.
2. The integrated circuit chip timing optimization method according to claim 1, characterized in that: After obtaining the performance parameters, power consumption parameters and area parameters of each integrated circuit module, and performing preliminary division of each integrated circuit module according to the performance parameters, power consumption parameters and area parameters by using logic level analysis to determine a preliminary logic division scheme, the method includes: In the initial stage, a preliminary logical partitioning scheme is obtained and used as an initial classification set; The preliminary logic partitioning scheme is integrated into the initial classification set by using the hierarchical information and the input module partitioning restriction, and combined with the physical core grain partitioning, the preliminary core grain partitioning scheme is generated according to the area restriction.
3. The integrated circuit chip timing optimization method according to claim 1, characterized in that: The steps of analyzing the timing of the core grain according to the total negative timing margin of all paths in setup time, the total negative timing margin of all paths in hold time, the negative timing margin of the worst case setup time in all paths, and the negative timing margin of the worst case hold time in all paths generated by the OpenSTA tool, and readjusting the preliminary core grain partitioning scheme to determine an iterative partitioning scheme when a violation occurs until the timing of the core grain meets the requirements include: Write and use the corresponding tcl script in the OpenSTA tool, call the underlying files of the required environment and call the corresponding timing analysis module in OpenSTA to perform the final timing analysis, and then obtain the corresponding quantitative results; By analyzing the four data of the total negative timing margin of all paths generated by OpenSTA in the setup time, the total negative timing margin of all paths in the hold time, the negative timing margin of the worst-case setup time in all paths, and the negative timing margin of the worst-case hold time in all paths, it is determined whether the timing of the chip reaches the ideal state.
4. The integrated circuit chip timing optimization method according to claim 3, characterized in that: By analyzing the four data of the total negative timing margin of all paths generated by OpenSTA in the setup time, the total negative timing margin of all paths in the hold time, the negative timing margin of the worst-case setup time in all paths, and the negative timing margin of the worst-case hold time in all paths, the steps of judging whether the timing of the chiplet reaches the ideal state include: Determine whether the total negative timing margin of all paths generated by OpenSTA in setup time, the total negative timing margin of all paths in hold time, the negative timing margin of the worst-case setup time in all paths, and the negative timing margin of the worst-case hold time in all paths are all positive numbers; If they are all positive, it means that the timing of all paths meets the requirements and the timing of the core particle reaches the ideal state. If a negative number appears, it means that at least one path of the chip has violated, which means that the timing of the chip has not reached the ideal state.
5. The integrated circuit chip timing optimization method according to claim 4, characterized in that: When the iterative partitioning scheme meets the timing optimization requirement, the current iterative partitioning scheme is recorded and saved, and the integrated circuit modules at the same level are continued to be regularly partitioned until all preliminary core grain partitioning schemes at the level are completely traversed, including: If at least one path of the core particle violates the rule, the initial preliminary logic partitioning scheme is returned to and the partitioning strategy is readjusted; Re-execute the step of performing preliminary partitioning of the integrated circuit modules according to the performance parameter, the power consumption parameter and the area parameter by using logic level analysis to determine a preliminary logic partitioning scheme; Then detect four data generated by the OpenSTA tool: the total negative timing margin of all paths in the setup time, the total negative timing margin of all paths in the hold time, the negative timing margin of the worst-case setup time in all paths, and the negative timing margin of the worst-case hold time in all paths, and so on, until the timing of the core particle meets the requirements; When the result reaches the optimal level, the current partitioning scheme is recorded and saved, and the modules at the same level continue to be divided regularly until all possible partitions at that level are completely traversed.
6. The integrated circuit chip timing optimization method according to claim 1, characterized in that: When the iterative partitioning scheme meets the timing optimization requirements, the current iterative partitioning scheme is recorded and saved, and the integrated circuit modules at the same level are continuously partitioned regularly until all the preliminary core grain partitioning schemes at the level are completely traversed, including: Obtaining a preset full-chip gate circuit netlist file, process file, layout file, timing constraint file, physical constraint file, and partition constraint file; According to the full-chip gate circuit netlist file, process file, layout file, timing constraint file, physical constraint file and partition constraint file, step-by-step timing optimization of the total negative timing margin and the worst negative timing margin on all paths inside the core grain is implemented in hierarchical iteration to determine the layout file and core grain partition file corresponding to the core grain.
7. The integrated circuit chip timing optimization method according to any one of claims 1 to 6, characterized in that: When the iterative partitioning scheme meets the timing optimization requirements, the current iterative partitioning scheme is recorded and saved, and the integrated circuit modules at the same level are continued to be regularly partitioned until all preliminary core grain partitioning schemes at the level are completely traversed, including: The connection ratio in the module set is calculated based on the floorplan data before input layout and the location of the given chip IO as the preference candidate set index; Combined with the size of the core area of the core particle, the minimum area of any core particle, the utilization rate of each core particle and the aspect ratio constraint, the modules are allocated proportionally according to their total area, and a relaxation coefficient is added during the allocation to achieve the versatility of the layout plan. Modules with high preference candidate index values are preferentially divided into adjacent parts, and the corresponding physical area plan is output.
8. An integrated circuit chip timing optimization device, characterized in that: include: a preliminary logic partitioning module, configured to respond to an instruction to optimize the timing of a chiplet of an integrated circuit, obtain performance parameters, power consumption parameters, and area parameters of each integrated circuit module, and perform preliminary partitioning of each integrated circuit module according to the performance parameters, the power consumption parameters, and the area parameters using logic level analysis to determine a preliminary logic partitioning scheme; A chip timing optimization module is configured to generate a preliminary chip partitioning scheme based on the preliminary logic partitioning scheme and in combination with the physical chip partitioning according to the area restriction, call the DREAMPlace tool to perform chip layout, and optimize the total negative timing margin and the worst negative timing margin on all paths inside the chip in the timing optimization module to determine the chip backend layout result; a partitioning scheme adjustment module, configured to analyze the timing of the core particles based on the back-end layout result of the core particles according to the total negative timing margin of all paths on the setup time, the total negative timing margin of all paths on the hold time, the negative timing margin of the worst-case setup time in all paths, and the negative timing margin of the worst-case hold time in all paths generated by the OpenSTA tool, and readjust the preliminary core particle partitioning scheme to determine an iterative partitioning scheme when a violation occurs, until the timing of the core particles meets the requirements; The partitioning scheme traversal module is configured to record and save the current iterative partitioning scheme when the iterative partitioning scheme meets the timing optimization requirements, and continue to perform regular partitioning on the integrated circuit modules at the same level until all preliminary core grain partitioning schemes at the level are completely traversed.
9. An electronic device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
Citation Information
Patent Citations
Core particle algorithm scheduling method and system, electronic equipment and storage medium
CN115860081A
Parallel simulation method and device for integrated chip system and computer equipment
CN118013690A
Communication method and device for interconnection packaging of memristor core particles and silicon-based core particles
CN118467449A
Core particle dividing method
CN118504507A
Optical phased array structure and fabrication techniques
US20210103199A1
Cited By
Integrated circuit time sequence drive overall layout method based on path optimization
CN120975021A
ASIC layout optimization method based on hybrid shaping programming
CN121457426A
Core particle division and mixed bonding pad distribution method based on 3D layout device
CN121525623A
A 3D placer based die partition and hybrid bonding pad assignment method
CN121525623B