Layout and wiring method of PCIe interface chip
By precisely placing and constraining register positions in the back-end design of the PCIe 3.0 interface chip and optimizing the clock tree length, the timing violation problem caused by the randomness of register positions on the receiving side was solved, achieving lower power consumption, smaller area and shorter R&D cycle.
Patent Information
- Application Number
- CN202510811297.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-19
AI Technical Summary
In the back-end design of PCIe 3.0 interface chips, there is a lack of artificial constraints on the position of the receiving-side registers and no unified standard for the length of the clock tree, which leads to a large number of violations during timing checks and makes it difficult to optimize the chip power consumption and area.
By importing the chip backend design files, planning the layout, and accurately placing the registers associated with the IP core receiving side, the clock tree is used to comprehensively constrain the register clock length, and the ladder structure and area constraint method are used to fix the register position and optimize the wiring process.
It effectively reduces the number of timing violations, improves chip performance, reduces power consumption and area, shortens the R&D cycle, and reduces R&D costs.
Smart Images

Figure CN120671625A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chip design automation, and in particular to a layout and wiring method for a PCIe interface chip. Background Art
[0002] In modern computer systems, the rapid development of high-performance computing, data centers, gaming, industrial control, medical equipment, and autonomous driving has led to an increasing demand for high-speed data transmission. To meet this demand, the PCI-SIG has introduced the Peripheral Component Interconnect Express (PCIe). PCIe 3.0, by inheriting the advantages of the previous two generations and significantly improving data transfer speeds and efficiency, has become a key technology for high-speed data transmission today and in the future.
[0003] Despite the ongoing evolution of PCIe technology, PCIe 3.0 boasts significant market demand and existing capacity, as its performance is sufficient to meet daily user needs and its hardware compatibility and cost-effectiveness make it a significant market force. In scientific computing, artificial intelligence, and big data analytics, PCIe 3.0's high-speed data transmission capabilities significantly improve computing efficiency. In data centers, its high throughput and low latency make it the preferred interface for servers and storage devices. In gaming, it supports faster graphics processing and data transmission, optimizing gaming performance. PCIe 3.0 also plays a vital role in demanding real-time and reliability scenarios such as industrial control, medical equipment, and autonomous driving.
[0004] In the field of digital chip design, performance, power consumption, and area (PPA) constitute the three key indicators guiding design. With the continuous advancement of process technology, chip power consumption has become an urgent issue that needs to be addressed. High power consumption not only reduces device endurance but can also cause a series of problems such as heat dissipation and reliability. Therefore, reducing power consumption is one of the important goals of chip design.
[0005] To provide higher-performance, lower-cost system solutions, improving chip performance to meet the needs of various application scenarios is also a key factor in the chip design process. Furthermore, chip area is a significant cost driver in integrated circuit manufacturing. Reducing chip area not only reduces costs but also increases chip integration to enable more functionality. Therefore, optimizing area is also a key goal in chip design.
[0006] Currently, during the back-end design of PCIe 3.0-based interface chips, the industry typically lacks constraints on the location of receive (RX) registers, and there are no unified standards for clock tree length. This results in random distribution of logic circuits, especially registers, on the receive side. Back-end optimization tools may scatter these logic circuits at the intellectual property (IP) core interface or throughout the chip. Because PCIe 3.0 IP cores are larger than PCIe 4.0 / 5.0, shortening the receive-side clock tree length during subsequent clock tree synthesis increases the difficulty.
[0007] At the same time, this type of chip generally has the problem that the internal clock length of the IP core is longer than the data path and the frequency is higher than PCIe 1.0 / 2.0. This makes it easy for a large number of setup time violations to occur when timing checks are performed on registers that are far away from the IP core interface and the IP core; when timing checks are performed on registers that are close to the IP core interface and the IP core, a large number of hold time violations are easy to occur.
[0008] Given the clock frequency limitations of such chips, chip designers often have no choice but to resort to brute force when faced with these violations, incurring a significant cost in power consumption, area, and R&D time.
[0009] This invention aims to provide an efficient method to reduce power consumption, area, and R&D costs by optimizing the back-end design of the interface receiving side chip, while also improving performance and shortening the R&D cycle. In addition to being applicable to PCIe 3.0, this invention can also be applied to other PCIe protocols to achieve even better technical results. Summary of the Invention
[0010] In order to alleviate or partially alleviate the above technical problems, the solutions of the present invention are as follows:
[0011] A layout and routing method for a PCIe interface chip comprises the following steps:
[0012] Step S1: Import the files required for chip backend design;
[0013] Step S2: layout planning;
[0014] Step S3: placing, which includes:
[0015] Capturing the IP core receiving side interface signals and register information associated with the receiving side inside the IP core, wherein the interface signals include x-axis and y-axis coordinates and the IP core direction;
[0016] Arrange the registers associated with the receiving side inside the IP core according to the direction of the IP core;
[0017] Step S4: clock tree synthesis, including:
[0018] Constrain the clock length of the registers associated with the receiving side based on the delay length of the data path of the IP core's internal receiving side interface, the delay length of the clock path of the IP core's internal receiving side interface, and the delay length of the data path between registers associated with the IP core's internal receiving side.
[0019] Step S5: wiring;
[0020] Step S6: ECO.
[0021] Furthermore, the capturing of the IP core receiving side interface signal specifically involves capturing the x-axis and y-axis coordinates and the IP core direction in the interface signal respectively through the IP core receiving side interface signal keyword.
[0022] Furthermore, the register information associated with the receiving side inside the IP core is specifically the name of the register of the corresponding endpoint device captured through the fan-out information of the interface signal.
[0023] Furthermore, the registers associated with the receiving side inside the IP core are placed according to the direction of the IP core, specifically:
[0024] When the IP core is placed on the PCIe 3.0 interface chip, the register corresponding to the interface pin of the first IP core is placed at the coordinates of (x, yd); when the register corresponding to the interface pin of the nth IP core is placed, the register coordinates are (x, yd-(n-1)×h);
[0025] When the IP core is placed under the PCIe 3.0 interface chip, when placing the register corresponding to the interface pin of the first IP core, the coordinates of the register are (x, y+d); when placing the register corresponding to the interface pin of the nth IP core, the coordinates of the register are (x, y+d+(n-1)×h), where d is the first distance, the coordinate values x and y are real values, h is the height of a register, and n is the number of a group of signal channels.
[0026] Furthermore, the registers associated with the receiving side inside the IP core are placed according to the direction of the IP core, specifically:
[0027] When the IP core is placed on the left side of the PCIe 3.0 interface chip, the register corresponding to the interface pin of the first IP core is placed at the coordinates of (x+d,y); when the register corresponding to the interface pin of the nth IP core is placed, the register coordinates are (x,y+d+(n-1)×w);
[0028] When the IP core is placed on the right side of the PCIe 3.0 interface chip, when placing the register corresponding to the interface pin of the first IP core, the coordinates of the register are (xd, y); when placing the register corresponding to the interface pin of the nth IP core, the coordinates of the register are (xd-(n-1)×w, y), where d is the first distance, the coordinate values x and y are real values, w is the width of a register, and n is the number of a group of signal channels.
[0029] Furthermore, the first distance d is 50-200 microns.
[0030] Furthermore, when the delay length of the clock path of the receiving side interface inside the IP core is greater than the delay length of the data path of the receiving side interface inside the IP core, the clock length of the register associated with the receiving side is shortened.
[0031] Furthermore, when the delay length of the clock path of the IP core internal receiving side interface is less than the delay length of the data path of the IP core internal receiving side interface, and the delay length of the data path of the IP core internal receiving side interface plus the delay length of the data path between the registers associated with the IP core internal receiving side minus the delay length of the clock path of the IP core internal receiving side interface is less than the clock length of the registers associated with the receiving side, the clock length of the registers associated with the receiving side is shortened;
[0032] When the delay length of the IP core internal receiving side interface clock path is less than the delay length of the IP core internal receiving side interface data path, and the delay length of the IP core internal receiving side interface data path plus the delay length of the data path between registers associated with the IP core internal receiving side minus the delay length of the IP core internal receiving side interface clock path is greater than the clock length of the registers associated with the receiving side, the clock length of the registers associated with the receiving side is shortened, but the clock length of the registers associated with the receiving side is not less than: the delay length of the IP core internal receiving side interface data path plus the delay length of the data path between registers associated with the IP core internal receiving side minus the delay length of the IP core internal receiving side interface clock path minus the clock length of the registers associated with the receiving side.
[0033] Furthermore, after executing step S2 and before executing step S3, the delay length of the data path of the receiving side interface inside the IP core is obtained through the timing library of the IP core;
[0034] After executing step S3 and before executing step S4, the delay length of the clock path of the receiving side interface inside the IP core is obtained through the timing library of the IP core.
[0035] A PCIe 3.0 interface chip is obtained according to any one of the aforementioned layout and routing methods for a PCIe interface chip.
[0036] The technical solution of the present invention has one or more of the following beneficial technical effects:
[0037] (1) During the placement phase, the corresponding registers are accurately captured using the direction of the IP core and the IP interface signal keywords. A ladder-shaped placement strategy is then used to effectively control the position of the registers associated with the receiving side of the IP interface. This approach not only fixes the position of the registers but also reduces the interference of noise on timing and the adverse effects of layout congestion on routing. At the same time, the IP clock is used to capture the remaining registers on the receiving side, and the area constraint method is used to clarify the physical location of the remaining registers.
[0038] (2) In the clock tree synthesis stage, the required clock length of the register associated with the receiving side is constrained based on the difference in the delay length of the data path of the internal interface of the IP core, the delay length of the clock path of the internal interface of the IP core, and the delay length of the data path between the IP core and the register associated with the receiving side.
[0039] The method proposed in the present invention can effectively reduce the number of overall timing violations, so that the chip can achieve a better PPA index.
[0040] In addition, other beneficial effects of the present invention will be mentioned in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flow chart of the integrated circuit physical layout design method of the present invention;
[0042] Figure 2 This is a schematic diagram of the physical placement of registers in one embodiment of the present invention;
[0043] Figure 3 This is a schematic diagram after completing the physical location constraints of the register. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0045] To facilitate a clear description of the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish between identical or similar items having substantially the same functions and effects. Those skilled in the art will appreciate that terms such as "first" and "second" do not limit the quantity or order of execution.
[0046] Explanation of terms:
[0047] Floorplanning: A key step in the digital backend design process, it primarily encompasses component planning and placement, power network planning, and physical unit insertion. Proper floorplanning can avoid issues like poor layout and difficult routing during subsequent design stages, thereby reducing design iterations and improving design efficiency.
[0048] Clock Tree Synthesis (CTS): In digital integrated circuits, clock signals are used to synchronize the operation of sequential components (such as flip-flops and latches), ensuring that data is sampled and processed at the correct time. The goal of clock tree synthesis is to build an efficient clock distribution network, such as common topologies such as X-tree and H-tree. This network evenly distributes clock signals from a clock source (such as a crystal oscillator) to each sequential component on the chip, ensuring that each sequential component is accurately triggered on the same clock edge, thereby ensuring normal chip operation.
[0049] Engineering Change Orders (ECOs): These are important tools for making partial design modifications in the later stages of chip design when design issues arise or optimization is needed. These include logical ECOs and physical ECOs. Logical ECOs typically make logic modifications to the netlist during the physical design phase. Physical ECOs include timing ECOs and design rule check ECOs. The former can be achieved by inserting buffers and adjusting wire lengths, while the latter can be achieved by modifying layout and routing to meet process rules.
[0050] Front-end design: It is a part of chip design, focusing on the functional implementation and logic design of the chip. It is the process from chip specification definition to logic circuit description. It usually includes: according to the chip specifications and other requirements, the chip logic circuit is described using hardware description language (HDL) to form register transfer level (RTL) code, and then the RTL code is converted into a gate-level netlist during logic synthesis. Finally, the logic function of the design is simulated and verified.
[0051] Back-end design is a step in chip design that focuses on the physical implementation of the chip. It is the process of converting the gate-level netlist generated by the front-end design into the actual chip physical layout. The verified physical layout is submitted to the wafer foundry in GDSII format for tape-out, resulting in the final chip.
[0052] Receiving side: The receiving side mentioned in the present invention refers to the receiving side of the internal interface of the IP core.
[0053] Chip design can generally be divided into front-end design and back-end design, and the present invention focuses on the back-end design part.
[0054] The present invention discloses a layout and routing method for a PCIe interface chip, which mainly includes the following steps: importing required files, layout planning, placement, clock tree synthesis, routing, ECO and other steps commonly found in the art. Parts of these steps that are not described in detail can be implemented using conventional means in the art, and the present invention will not elaborate on them.
[0055] The present invention can reduce the overall number of violations by constraining the register positions associated with the receiving side inside the IP core and rationally planning the length of the clock tree, thereby achieving lower power consumption, better performance, smaller area, shorter R&D cycle and lower R&D cost.
[0056] Figure 1 The flowchart of the integrated circuit physical layout design method of the present invention is as follows. To achieve the above technical objectives, the integrated circuit physical layout design method of the present invention specifically includes the following steps:
[0057] Step S1: Import the files required for chip back-end design.
[0058] This step is the starting point and foundation of the entire digital backend design process. It requires importing all the files required for the backend design process. For example, it may include:
[0059] (a) Netlist after comprehensive design for testability (DFT): This is a specific representation of the chip logic circuit, including all logic units and their connections.
[0060] (b) Synopsys Design Constraints (SDC) file: This file specifies the timing requirements for each signal in the chip, such as clock cycle, setup time, hold time, etc., providing a basis for subsequent timing analysis and optimization.
[0061] (c) Library file: Contains the timing and physical characteristics information of the logical units and physical units, which is used by the tool for accurate analysis and layout and routing.
[0062] (d) Verification (signoff) design document: specifies the standards and conditions for final verification and delivery of the chip.
[0063] (e) Various process files required by Electronic Design Automation (EDA) tools: These files contain information related to the chip manufacturing process, such as metal layer thickness, line width, spacing, etc., and are the basis for the tool to perform physical design.
[0064] Step S2: Layout planning.
[0065] In this step, at least the following components should be planned:
[0066] (a) I / O pad planning: Determine the location of input and output pins, considering the connection with external devices and the signal transmission path.
[0067] (b) Macro unit planning: Macro units usually refer to larger functional modules, such as processor cores and memory controllers, which need to be reasonably laid out based on their functions and interactions with other modules.
[0068] (c) Power and ground planning (powerplan): Design a reasonable power and ground network to ensure stable power supply to all parts of the chip and reduce the impact of power supply noise on circuit performance.
[0069] (d) Inserting physical cells: These physical cells include buffers, inverters, etc., which are used to improve the driving capability and timing characteristics of the signal.
[0070] Figure 2 This is a schematic diagram of the physical placement of registers in one embodiment of the present invention. For example, Figure 2 The IP cores shown in the figure include a first IP core IP-1, a second IP core IP-2, a third IP core IP-3, and a fourth IP core IP-4, which are located at different positions of the chip.
[0071] Furthermore, after executing step S2 and before executing step S3, the delay length (Ldata) of the receiving side interface data path inside the IP core is obtained through the timing library of the IP core, laying the foundation for the subsequent adjustment of the register position and register clock length associated with the receiving side.
[0072] Step S3: Place.
[0073] This step primarily involves automatically placing standard cells using existing tooling. Incorporating the rules defined in this invention can guide the tool to achieve better placement results. For example, the positions of key cells can be pre-specified based on the chip's functional modules and signal flow, or layout constraints can be set, such as minimum spacing between cells and alignment requirements.
[0074] As one of the innovations of the present invention, after completing the above step S2, the physical location of the associated registers is realized by capturing the internal receiving side interface position information of the IP core and the associated register information, which specifically includes the following implementation steps:
[0075] Step S31: Capture the IP core receiving side interface signal and register information associated with the internal receiving side of the IP core.
[0076] Use the IP core receiving side interface signal keyword (such as rx*_data*) to capture the x-axis and y-axis coordinates of the interface signal and record the IP core direction.
[0077] The name of the register of the corresponding endpoint device (EP) is captured through the fanout information of the interface signal.
[0078] Step S32: Arrange the registers according to the direction of the IP core. Specifically, the registers can be arranged in the following four ways.
[0079] (i) On the one hand, when the IP core is placed on the chip, the direction attributes are MX and R180:
[0080] When placing the register corresponding to the interface pin of the first IP core, the x-axis coordinate remains unchanged, and the y-axis coordinate is reduced by the first distance d (for example, 100 μm), that is, the coordinate is (x, yd), where x, y, and d are all real values.
[0081] When placing the registers corresponding to the interface pins of the second IP core, the x-axis coordinate remains unchanged, and the y-axis coordinate is subtracted from the first distance d (for example, 100 μm) and then from the height (h) of a register, that is, the coordinate is (x, ydh).
[0082] Similarly, when placing the registers corresponding to the interface pins of the nth IP core, the x-axis coordinate remains unchanged, and the y-axis coordinate is reduced by the first distance d (for example, 100 μm) and then by the height of n-1 registers, that is, the coordinate is (x, yd - (n-1) × h). Taking n = 10 as an example, the registers associated with the IP core interface can be arranged in a trapezoidal structure, where n is a positive integer such as 3, 4, 8, 10, 16, 32, etc.
[0083] (ii) On the other hand, when the IP core is placed under the chip, the direction attributes are MY, R0:
[0084] When placing the registers corresponding to the interface pins of the first IP core, the x-axis coordinate remains unchanged, and the y-axis coordinate is increased by the first distance d (for example, 100 μm), that is, the coordinate is (x, y+d).
[0085] When placing the registers corresponding to the interface pins of the second IP core, the x-axis coordinate remains unchanged, and the y-axis coordinate is added with the first distance d (for example, 100 μm) plus the height of a register, that is, the coordinate is (x, y+d+h).
[0086] Similarly, when placing the registers corresponding to the interface pins of the nth IP core, the x-axis coordinate remains unchanged, and the y-axis coordinate is added with the first distance d (for example, 100 μm) plus the height of n-1 registers, that is, the coordinate is (x, y + d + (n-1) × h).
[0087] (iii) On the other hand, when the IP core is placed on the left side of the chip, the orientation attributes are MX90 and R270:
[0088] When placing the registers corresponding to the interface pins of the first IP core, the y-axis coordinate remains unchanged, and the x-axis coordinate is added with the first distance d (for example, 100 μm), that is, the coordinate is (x+d, y).
[0089] When placing the registers corresponding to the interface pins of the second IP core, the y-axis coordinate remains unchanged, and the x-axis coordinate is added with the first distance d (for example, 100 μm) plus the width (w) of a register, that is, the coordinate is (x+d+w,y).
[0090] Similarly, when placing the registers corresponding to the interface pins of the nth IP core, the y-axis coordinate remains unchanged, and the x-axis coordinate is added with the first distance d (for example, 100 μm) plus the width of n-1 registers, that is, the coordinate is (x+d+(n-1)×w,y).
[0091] (iv) On the other hand, when the IP core is placed on the right side of the chip, the orientation attributes are MY90 and R90:
[0092] When placing the registers corresponding to the interface pins of the first IP core, the y-axis coordinate remains unchanged, and the x-axis coordinate is reduced by the first distance d (for example, 100 μm), that is, the coordinate is (xd, y).
[0093] When placing the registers corresponding to the interface pins of the second IP core, the y-axis coordinate remains unchanged, and the x-axis coordinate is subtracted from the first distance d (for example, 100 μm) and then from the width of a register, that is, the coordinate is (xdw, y).
[0094] Similarly, when placing the registers corresponding to the interface pins of the nth IP core, the y-axis coordinate remains unchanged, and the x-axis coordinate is reduced by the first distance d (for example, 100 μm) and then by the width of n-1 registers, that is, the coordinate is (xd-(n-1)×w,y).
[0095] The positive integer n is the number of a group of signal channels, which can be optimized and adjusted according to the result of place congestion and the result of design rule check (DRC) of routing.
[0096] It is worth mentioning that the terms "up", "down", "left" and "right" in the above descriptions regarding the chip are all relative position concepts, and the present invention is not limited to a specific position.
[0097] In addition, the above-mentioned first distance d=100 μm is only an example and not an absolute limitation. The length can be any reasonable length value, for example, the first distance d is 50-200 microns, and the present invention is not limited to this example.
[0098] Step S33: Capture the remaining registers and constrain the positions.
[0099] According to the receiving side clock of different lanes of the IP core, the register names of the corresponding endpoint devices (EP) are captured through fan-out, and the registers that have been placed in the above step S32 are filtered out.
[0100] Create instance groups of the captured registers and place them around 100 μm from the IP core interface according to different channels of the IP core (for congestion considerations).
[0101] According to the above steps, the physical location constraints of all associated registers on the receiving side of the IP core are completed. Figure 3 The diagram shows the physical location constraints of the registers, which include 16 lanes: Lane-0, Lane-1, Lane-2, Lane-3, Lane-4, ..., Lane-13, Lane-14, and Lane-15.
[0102] Furthermore, based on the placement result of step S3, check whether the local density meets the requirements, whether there is local congestion problem, and whether the delay length of the data path of the register associated with the IP core and the receiving side is within the delay range corresponding to the physical distance.
[0103] For example, in the layout, the physical distance is about 100μm, and the theoretical data path delay is about 100ps (with slight deviations in different processes).
[0104] If the above points meet the requirements, then proceed to the next step S4. Otherwise, it is necessary to adjust the physical positions of the receiving-side registers and the associated combinational logic circuits, optimize the local density, congestion, and the delay length (Lrxdata) of the data path between the registers associated with the internal receiving side of the IP core. During this process, Lrxdata can be obtained.
[0105] It should be noted that for the optimized Lrxdata here, it is necessary to ensure whether the positions of the receiving-side registers and the associated combinational logic circuits are near the interface of the IP core. Otherwise, it will directly cause high delay.
[0106] Furthermore, after completing step S3 and before executing step S4, obtain the delay length (Lclock) of the clock path of the internal receiving-side interface of the IP core through the timing library of the IP core, and combine the previously obtained delay length (Ldata) of the data path of the internal receiving-side interface of the IP core to provide a basis for adjusting the clock length of the registers associated with the receiving side subsequently.
[0107] Step S4: Clock Tree Synthesis (CTS).
[0108] Step S41: Constrain the clock length according to the delay relationship.
[0109] In the previous steps, three values are obtained respectively: the delay length (Ldata) of the data path of the internal receiving-side interface of the IP core, the delay length (Lclock) of the clock path of the internal receiving-side interface of the IP core, and the delay length (Lrxdata) of the data path between the registers associated with the internal receiving side of the IP core.
[0110] When Lclock > Ldata: The clock length of the registers associated with the receiving side should be shortened as much as possible according to the physical distance constraint. In the layout planning, the physical distance between the clock source and the associated registers is about 100 μm, and theoretically the clock length can be controlled at about 100 picoseconds (ps) (with slight deviations for different processes).
[0111] When Lclock < Ldata and Ldata + Lrxdata - Lclock < Tclock: The clock length of the registers associated with the receiving side also needs to be shortened as much as possible. Theoretically, the clock length Tclock of the registers associated with the receiving side can be controlled at about 100 ps (with slight deviations for different processes).
[0112] When Lclock < Ldata and Ldata + Lrxdata - Lclock > Tclock: The clock length Tclock of the register associated with the receiving side is shortened as much as possible, but it should not be lower than Ldata + Lrxdata - Lclock - Tclock.
[0113] Step S42: Clock tree length analysis and adjustment.
[0114] After CTS, analyze whether the clock tree length in various cases meets the requirements. If it meets, proceed to the next step S5; otherwise, adjust the clock length of the register associated with the receiving side, and then re-execute step S4.
[0115] In summary, as another innovative point of the present invention, in the CTS stage, according to the delay length (Ldata) of the data path of the receiving side interface inside the IP core, the delay length (Lclock) of the clock path of the receiving side interface inside the IP core, and the delay length (Lrxdata) of the data path between the registers associated with the receiving side inside the IP core, based on the difference between these delay lengths, the clock length of the register associated with the receiving side required is constrained. This helps to reduce the number of violations, reduce the engineering development volume related to late repair, especially, and improve the PPA metrics of the chip.
[0116] Step S5: Routing.
[0117] This step includes optimization after routing. Mainly call the Place and Route (PR) tool's solution to automatically route the nets in the circuit design, and continue to optimize timing, area, power consumption, etc. after routing. For routing, the most important thing is whether it can be routed through, that is, whether it can minimize or even reduce to 0 the cases of short circuits and other violations of design rules after routing.
[0118] Optionally, the objects of routing in this step do not include special nets such as power supply components and analog components, because these nets usually have special constraints and need to be designed by the designer according to the process, layout planning, and other constraints.
[0119] Furthermore, in step S5, it is necessary to confirm whether it can be routed through during the routing process, and whether there are local short circuits and other non-compliances with DRC rules introduced by the positions of the registers associated with the receiving side and the combinational logic. If so, return to the placement stage in step S3, and obtain a new placement result by adjusting the positions of the registers associated with the receiving side and the associated combinational logic circuit, such as adjusting the positions of local cells, etc.; otherwise, proceed to the next step S6.
[0120] Step S6: ECO.
[0121] This step is mainly to fix the problem that the tool cannot completely fix. There are two main types of ECO:
[0122] (a) Logic ECO: Modifying the logic functions of the netlist. In the later stages of chip design, front-end engineers may discover design flaws (bugs) that require circuit modifications. However, the schedule at this time does not allow for resynthesis. Therefore, they will choose to modify the logic of the netlist in the PR tool. Generally, this involves adding some logic or rewiring certain logic nets.
[0123] (b) Physical ECO: Manually correct issues that cannot be completely fixed by PR tools, generally including timing ECO and DRC correction.
[0124] When all the above steps are in compliance with the requirements, it can be found that the number of related violations on the receiving side is much less than that of the tool-driven version, and the size of the violations is very small, which can be repaired with lower power consumption, smaller area, less time and labor cost. In other words, the PCI based on the present invention is e 3.0 interface chips can achieve lower power consumption, better performance, smaller area, shorter development cycles, and lower R&D costs. Table 1 shows the progress of an example of the present invention compared to the existing technology. In addition to effectively reducing the number and severity of violations, the present invention also achieves certain improvements in power consumption and area.
[0125] Table 1: Comparison of the benefits of the present invention and the prior art on some indicators
[0126] index Existing technology The present invention income Hold maximum violation size -1.12ns .0.086ns 92.30% Hold violation count 2437 152 93.76% Setup maximum violation size -0.187ns .0.084ns 55.08% Number of Setup violations 852 105 87.68% Power consumption 334mW 312mW 6.59% Occupied receiving side chip area <![CDATA[0.88μm 2 ]]> <![CDATA[0.81μm 2 ]]> 7.95% Number of Eco rounds required to repair violations 5 2 / Eco time per round 8.4 hours 6.5 hours 69.05%
[0127] To better illustrate the present invention, numerous specific details are provided in the detailed description above. Those skilled in the art will appreciate that the present invention can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main purpose of the present invention.
[0128] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A layout and routing method for a PCIe interface chip, characterized in that: The steps include: Step S1: Import the files required for chip backend design; Step S2: layout planning; Step S3: placing, which includes: Capturing the IP core receiving side interface signals and register information associated with the receiving side inside the IP core, wherein the interface signals include x-axis and y-axis coordinates and the IP core direction; Arrange the registers associated with the receiving side inside the IP core according to the direction of the IP core; Step S4: clock tree synthesis, including: Constrain the clock length of the registers associated with the receiving side based on the delay length of the data path of the IP core's internal receiving side interface, the delay length of the clock path of the IP core's internal receiving side interface, and the delay length of the data path between registers associated with the IP core's internal receiving side. Step S5: wiring; Step S6: ECO.
2. The layout and routing method for a PCIe interface chip according to claim 1, wherein: The capturing of the IP core receiving side interface signal specifically involves capturing the x-axis and y-axis coordinates and the IP core direction in the interface signal through the IP core receiving side interface signal keywords.
3. The layout and routing method for a PCIe interface chip according to claim 2, wherein: The register information associated with the internal receiving side of the IP core is specifically the name of the register of the corresponding endpoint device captured through the fan-out information of the interface signal.
4. The layout and routing method for a PCIe interface chip according to claim 3, wherein: The registers associated with the receiving side inside the IP core are arranged according to the direction of the IP core, specifically: When the IP core is placed on the PCIe 3.0 interface chip, the register corresponding to the interface pin of the first IP core is placed at the coordinates of (x, yd); when the register corresponding to the interface pin of the nth IP core is placed, the coordinates of the register are (x, yd-(n-1)×h); When the IP core is placed under the PCIe 3.0 interface chip, when placing the register corresponding to the interface pin of the first IP core, the coordinates of the register are (x, y+d); when placing the register corresponding to the interface pin of the nth IP core, the coordinates of the register are (x, y+d+(n-1)×h), where d is the first distance, the coordinate values x and y are real values, h is the height of a register, and n is the number of a group of signal channels.
5. The layout and routing method for a PCIe interface chip according to claim 4, wherein: The registers associated with the receiving side inside the IP core are arranged according to the direction of the IP core, specifically: When the IP core is placed on the left side of the PCIe 3.0 interface chip, the register corresponding to the interface pin of the first IP core is placed at the coordinates of (x+d,y); when the register corresponding to the interface pin of the nth IP core is placed, the register coordinates are (x,y+d+(n-1)×w); When the IP core is placed on the right side of the PCIe 3.0 interface chip, when placing the register corresponding to the interface pin of the first IP core, the coordinates of the register are (xd, y); when placing the register corresponding to the interface pin of the nth IP core, the coordinates of the register are (xd-(n-1)×w, y), where d is the first distance, the coordinate values x and y are real values, w is the width of a register, and n is the number of a group of signal channels.
6. The layout and routing method for a PCIe interface chip according to claim 5, wherein: The first distance d is 50-200 microns.
7. The layout and routing method for a PCIe interface chip according to claim 5, wherein: When the delay length of the clock path of the IP core's internal receiving-side interface is greater than the delay length of the data path of the IP core's internal receiving-side interface, the clock length of the register associated with the receiving side is shortened.
8. The layout and routing method for a PCIe interface chip according to claim 7, wherein: If the delay length of the IP core internal receiving side interface clock path is less than the delay length of the IP core internal receiving side interface data path, and the delay length of the IP core internal receiving side interface data path plus the delay length of the data path between the registers associated with the IP core internal receiving side minus the delay length of the IP core internal receiving side interface clock path is less than the clock length of the registers associated with the receiving side, the clock length of the registers associated with the receiving side is shortened; When the delay length of the IP core internal receiving side interface clock path is less than the delay length of the IP core internal receiving side interface data path, and the delay length of the IP core internal receiving side interface data path plus the delay length of the data path between registers associated with the IP core internal receiving side minus the delay length of the IP core internal receiving side interface clock path is greater than the clock length of the registers associated with the receiving side, the clock length of the registers associated with the receiving side is shortened, but the clock length of the registers associated with the receiving side is not less than: the delay length of the IP core internal receiving side interface data path plus the delay length of the data path between registers associated with the IP core internal receiving side minus the delay length of the IP core internal receiving side interface clock path minus the clock length of the registers associated with the receiving side.
9. The layout and routing method for a PCIe interface chip according to claim 8, wherein: After executing step S2 and before executing step S3, the delay length of the receiving side interface data path inside the IP core is obtained through the timing library of the IP core; After executing step S3 and before executing step S4, the delay length of the clock path of the receiving side interface inside the IP core is obtained through the timing library of the IP core.
10. The layout and routing method for a PCIe interface chip according to claim 9, wherein: The PCIe interface chip is a PCIe 3.0 interface chip.
Citation Information
Cited By
Layout method and system of digital circuit pipeline register and storage medium
CN121303048A
A method, system, and storage medium for placement of digital circuit pipeline registers
CN121303048B