Timing sequence optimization system for eliminating influence of OCV (Open Circuit Voltage) of chip
By adopting a serial cascaded modular architecture and asynchronous read/write mechanism, the negative impact of OCV on chip timing under deep nanometer process was resolved, achieving rapid convergence of chip timing and improved yield.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 奕行智能科技(广州)有限公司
- Filing Date
- 2026-03-24
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies cannot effectively eliminate the negative impact of OCV on chip timing under deep nanometer processes, resulting in difficulties in chip timing convergence, high design costs, and low production yield.
It adopts a modular architecture with serial cascading, including a global clock source module, front-end and back-end clock tree modules, front-end and back-end data transmission modules, and a FIFO core processing module. It achieves timing isolation and data transfer through asynchronous read and write mechanisms, completely eliminating cross-clock tree delay deviations caused by OCV.
It achieves rapid convergence of chip timing, reduces design costs, improves production yield, and avoids frequency non-compliance and functional failure caused by OCV deviation.
Smart Images

Figure CN121900583A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated circuit chip design technology, and in particular to a timing optimization system for eliminating the effects of chip OCV. Background Technology
[0002] As integrated circuit technology continues to evolve towards deep nanometer nodes such as 7nm and 14nm, the requirements for chip integration density and operating frequency are constantly increasing. Timing design has become a core critical link in chip physical design, directly determining chip performance, production yield, and reliability in practical applications. In the actual production and operation of chips, on-chip process variations are an unavoidable phenomenon. Affected by the precision deviations in wafer manufacturing, etching, doping, and other process stages, the actual delays of devices (cells) and interconnects (nets) within the chip will deviate from the theoretical values calculated during the design phase using process libraries and RC parasitic parameters. This deviation is directly reflected in the delays of the clock path and data path, significantly negatively impacting chip timing convergence. Simultaneously, the proportion of interconnect delay in the overall delay increases significantly under deep nanometer processes, further amplifying the interference of OCV (Optical Characteristic Value) on timing. How to eliminate the timing impact of OCV and achieve rapid timing convergence has become a core technical problem that urgently needs to be solved in deep nanometer process chip design.
[0003] To mitigate the impact of OCV on chip timing, various timing optimization schemes have been proposed in existing technologies, becoming the mainstream implementation method in the industry. On the one hand, the mainstream design approach is to cover the process deviations caused by OCV by setting pessimistic delay correction coefficients (derate values) during the chip timing constraint stage. Specifically, a derate value greater than 1 is applied to the delay path of the launch clock, and a derate value less than 1 is applied to the delay path of the capture clock. By artificially amplifying the delay deviation of the clock path, the timing constraints are made more redundant, thereby ensuring the timing correctness of the chip in actual operation after production. On the other hand, the industry also optimizes from the physical design level of the clock tree. For example, in the clock tree synthesis (CTS) stage, clock synchronization optimization (CCOPT) technology, unbalanced clock tree design, and flexible configurable H-tree (FCHT) clock structure are used. By shortening the clock transmission path, increasing the proportion of the common path between the launch clock and the capture clock, and optimizing the layout and driving capability of the clock buffer, the cumulative effect of OCV deviation on the clock path is reduced. At the same time, in the chip synthesis and placement and routing stages, the combinational logic of the data channel is adjusted by iterative optimization to reduce data latency, and clock tree optimization is used to mitigate the timing impact of OCV. The above-mentioned technical solutions have been widely used in conventional chip design scenarios, which have reduced the interference of OCV on timing to a certain extent and improved the success rate of timing convergence.
[0004] However, existing technical solutions are all passive adaptation methods to the impact of OCV (Optical Characteristic Variation), failing to eliminate the negative timing effects of OCV on long data channel chip designs at their root. In practical applications, they still suffer from numerous unresolved technical defects, resulting in high design costs, long cycles, and limited optimization effects for chip timing convergence. Firstly, while setting a pessimistic derate value can cover OCV deviations, excessively pessimistic timing constraints significantly compress the chip's timing margin. To meet convergence requirements, designers need to perform numerous complex iterative optimizations on the clock and data paths, increasing not only the manpower and time costs of chip design but also potentially introducing more buffers and driver units due to over-optimization, leading to a significant increase in chip area and power consumption, sacrificing actual chip performance. Secondly, existing clock tree physical optimization methods are mostly concentrated in a single stage of chip design (such as the clock tree synthesis stage) and only make local adjustments to the clock tree itself. For long data channels... In scenarios with long transmission distances, the inherent delays of both the clock and data paths are already high. The non-common path delay deviations caused by OCV (Optical Characteristic Transmission) continue to accumulate, making it impossible to fundamentally cut off the cross-path propagation of these deviations, thus keeping timing convergence extremely difficult. Finally, existing technologies have not addressed the strong dependence of clock phase in synchronization circuits. Since the transmit and capture clocks are always within the same large clock tree, the delay deviations caused by OCV continue to propagate throughout the entire clock tree. For highly integrated chips manufactured using deep nanometer processes, this problem can lead to the risk of substandard operating frequencies and functional failures after chip production, significantly reducing chip yield. In summary, for chip design scenarios with long data channels, current technologies lack an effective solution to fundamentally eliminate the timing impact of OCV and achieve rapid timing convergence. A novel technical approach is urgently needed to address these technical problems. Summary of the Invention
[0005] This invention provides a timing optimization system to eliminate the impact of OCV (On-Chip Process Variation) in chip design. This system is suitable for timing-critical scenarios in chip design where the data path is long and the transmission distance is far. It can solve the problem of timing convergence difficulties caused by the delay deviation of the non-common part of the clock introduced by OCV, achieve rapid convergence of chip timing, reduce the timing optimization cost of chip design, and improve chip production yield.
[0006] This invention provides a timing optimization system for eliminating the effects of chip OCV, comprising: The front-end clock tree module is configured to provide a synchronous write clock for the front-end data transmission module and the FIFO core processing module. The front-end data transmission module is configured to receive, process, and transmit front-end data under write clock synchronization, and send the data to the FIFO core processing module buffer. The FIFO core processing module is configured to perform timing isolation and data transfer between the front-end and back-end clock tree modules through an asynchronous read / write mechanism, in order to separate the original clock tree and data channel; The subsequent clock tree module is configured to provide a synchronous read clock for the subsequent data transmission module and the FIFO core processing module; and The subsequent data transmission module is configured to receive, process, and transmit cached data under read clock synchronization, and then send the data to the target unit at the back end of the chip.
[0007] In one embodiment of the present invention, it further includes: The global clock source module is configured to generate and distribute raw reference clock signals to the front-end and rear-end clock tree modules.
[0008] In one embodiment of the present invention, the front-end clock tree module includes: The front-end clock source node is configured as the root node of the front-end small clock tree, receiving the original reference clock signal from the global clock source module and completing the first-level distribution of the clock signal. The clock buffer unit is configured to perform multi-level buffering and drive enhancement on the clock signal output by the front-end clock source node, ensuring that the clock signal can be stably transmitted to the write-side interface of the front-end data transmission module and the FIFO core processing module. The front-end clock routing network is configured to enable the physical transmission of clock signals from the front-end clock source node to each target module.
[0009] In one embodiment of the present invention, the front-end data transmission module includes: The data transmission unit is configured to receive raw business data from the chip front-end data source, perform data latching and synchronization processing, and convert asynchronous input data into a data stream synchronized with the write clock. The combinational logic processing unit is configured to preprocess the synchronized data to remove redundant information; The front-end data cabling network is configured to enable the physical transfer of processed data from combinational logic units to the write data interface of the FIFO core processing module.
[0010] In one embodiment of the present invention, the input terminal of the front-end clock tree module is electrically connected to the first output terminal of the global clock source module to receive the original reference clock signal; The first output terminal of the front-end clock tree module is electrically connected to the clock input terminal of the front-end data transmission module, providing a synchronous working clock for the front-end data transmission module; The second output terminal of the front-end clock tree module is electrically connected to the write clock input terminal of the FIFO core processing module, providing a clock reference for the write operation of the FIFO core processing module. The first input terminal of the front-end data transmission module is electrically connected to the data source at the front end of the chip to receive raw business data; The second input terminal of the front-end data transmission module is electrically connected to the first output terminal of the front-end clock tree module to receive the synchronous working clock; The output of the front-end data transmission module is electrically connected to the write data input of the FIFO core processing module, and the processed data to be transmitted is sent to the FIFO core processing module for buffering.
[0011] In one embodiment of the present invention, the FIFO core processing module includes: The write control unit is configured to receive input data from the preceding data transmission module under the control of the write clock, generate write address and write enable signal, control data to be written sequentially to the on-chip memory array, and adjust the write operation according to the empty / full status flag to prevent data overflow. The read control unit is configured to generate read address and read enable signal under read clock control, read cached data sequentially from the on-chip memory array and send it to the subsequent data transmission module, and adjust the read operation according to empty / full status flags to prevent empty reads. The on-chip storage array is configured to temporarily store data written to the previous stage, achieving physical isolation between read and write operations; The empty / full flag generation unit is configured to monitor the storage capacity of the on-chip memory array in real time, generate empty, full, and half-full status flags, and feed them back to the write control unit and the read control unit.
[0012] In one embodiment of the present invention, the subsequent clock tree module and the preceding clock tree module have a symmetrical structure, including: The subsequent clock source node is configured to receive the original reference clock signal from the global clock source module and complete the first-level distribution of the clock signal. The clock buffer unit is configured to buffer and enhance the clock signal output by the subsequent clock source node, ensuring that the clock signal is stably transmitted to the read-side interface of the subsequent data transmission module and the FIFO core processing module. The subsequent clock routing network is configured to enable the physical transmission of clock signals from the subsequent clock source node to each target module.
[0013] In one embodiment of the present invention, the subsequent data transmission module includes: The data receiving unit is configured to receive buffered data output from the FIFO core processing module, perform data latching and synchronization processing, and convert asynchronous input data into a data stream synchronized with the read clock. Combinational logic processing unit, which is configured to post-process synchronized data to restore the original business data format; The back-end data routing network is configured to enable the physical transfer of processed data from combinational logic units to target units at the back end of the chip.
[0014] In one embodiment of the present invention, the input terminal of the subsequent clock tree module is electrically connected to the second output terminal of the global clock source module to receive the original clock signal; The first output of the subsequent clock tree module is electrically connected to the read clock interface of the FIFO core processing module, providing a clock reference for the read operation of the FIFO core processing module. The second output terminal of the subsequent clock tree module is electrically connected to the second input terminal of the subsequent data transmission module to provide a synchronous working clock for the subsequent data transmission module; The first input terminal of the subsequent data transmission module is electrically connected to the read data interface of the FIFO core processing module to receive valid data from the FIFO buffer. The second input terminal of the subsequent data transmission module is electrically connected to the second output terminal of the subsequent clock tree module to receive the synchronous working clock; The output of the subsequent data transmission module is electrically connected to the target data receiving unit at the back end of the chip to complete the final data transmission.
[0015] The present invention also provides a method for operating the system according to the above, comprising: The front-end clock tree module provides a synchronous write clock for the front-end data transmission module and the FIFO core processing module. The front-end data transmission module performs the reception, processing and transmission of front-end data under write clock synchronization, and sends the data to the FIFO core processing module buffer; The FIFO core processing module performs timing isolation and data transfer between the front-end and back-end clock tree modules through an asynchronous read / write mechanism to separate the original clock tree and data channel; The subsequent clock tree module provides a synchronous read clock for the subsequent data transmission module and the FIFO core processing module; and The subsequent data transmission module receives, processes, and transmits the cached data under read clock synchronization, and then sends the data to the target unit at the back end of the chip.
[0016] The present invention has the following beneficial effects: (1) It solves the problem of clock non-common part delay deviation caused by OCV, and fundamentally overcomes the chip timing convergence difficulty caused by OCV.
[0017] (2) It significantly improves timing convergence efficiency, eliminates the need for complex traditional timing optimization operations, simplifies the design difficulty of timing convergence, and achieves rapid convergence of chip timing.
[0018] (3) It avoids the high optimization costs of traditional methods for convergence timing, and reduces the time and manpower costs of chip design.
[0019] (4) It gets rid of the extreme pessimistic timing constraints, avoids the problem of the chip's actual operating frequency not meeting the standard due to OCV deviation, or even chip failure, and improves the yield after chip production. Attached Figure Description
[0020] Figure 1 A system block diagram of a timing optimization system according to an embodiment of the present invention is shown. Detailed Implementation
[0021] In the following description, the invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments may be practiced without one or more specific details or with other alternatives and / or additional methods, materials, or components. In other instances, well-known structures, materials, or operations are not shown or described in detail so as not to obscure the inventive points of the invention. Similarly, for illustrative purposes, specific quantities, materials, and configurations are set forth to provide a comprehensive understanding of embodiments of the invention. However, the invention is not limited to these specific details.
[0022] In this invention, the various embodiments are merely intended to illustrate the solutions of the invention and should not be construed as limiting.
[0023] In this specification, references to "an embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment in all instances.
[0024] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0025] This system adopts a serially cascaded modular architecture, consisting of six functional units: a global clock source module 100, a front-end clock tree module 200, a front-end data transmission module 300, a FIFO core processing module 400, a rear-end clock tree module 500, and a rear-end data transmission module 600. Among them, the FIFO core processing module 400 is the core functional unit of the system, undertaking the core role of timing isolation and data relay between the front-end and rear-end subsystems. The front-end subsystem (front-end clock tree module 200 + front-end data transmission module 300) and the rear-end subsystem (rear-end clock tree module 500 + rear-end data transmission module 600) have no direct clock and data interaction. The asynchronous buffering and transmission of data are only completed through the FIFO core processing module 400, which completely cuts off the clock phase dependence of the original synchronous circuit at the physical and timing layers, and eliminates the impact of cross-clock tree delay deviation caused by OCV.
[0026] The following section, in conjunction with the accompanying drawings, provides a detailed explanation of the composition, core functions, and connections of each functional module of this system, while also outlining the overall workflow and connection architecture of the system.
[0027] Figure 1 A system block diagram of a timing optimization system according to an embodiment of the present invention is shown.
[0028] like Figure 1 As shown, in this embodiment, the timing optimization system includes: The global clock source module 100 serves as the source of the chip's clock system, responsible for generating and distributing the original reference clock signal. It provides the same clock input to the preceding and following clock tree modules, and simultaneously achieves physical segmentation of the clock tree through independent routing, avoiding the superposition and amplification of OCV deviations along long paths. This includes: The on-chip clock generator (PLL / OSC) generates the original high-frequency reference clock signal required by the chip, ensuring the stability and accuracy of the clock frequency and providing a unified clock frequency reference for the entire system. The global clock buffer unit buffers and enhances the original clock signal, improving the clock driving capability and ensuring that the clock signal can be stably transmitted to the front-end clock tree module 200 and the back-end clock tree module 500, while isolating signal interference between different sub-clock trees.
[0029] The connection relationships are as follows: The input terminal is electrically connected to the chip's on-chip clock generator (PLL / crystal oscillator, etc.) to receive the original oscillation clock signal; The first output terminal is electrically connected to the input terminal of the front-end clock tree module 200 to provide the original clock to the front-end subsystem; The second output terminal is electrically connected to the input terminal of the subsequent clock tree module 500, providing the original clock to the subsequent subsystem; The clock tree is wired independently from the pre-stage clock tree module 200 and the post-stage clock tree module 500, with no common clock path, ensuring the physical independence of the pre-stage and post-stage clock trees.
[0030] The front-end clock tree module 200, as the core of the front-end subsystem's clock supply, constructs an independent small clock tree to provide synchronous write clocks for the write side of the front-end data transmission module 300 and the FIFO core processing module 400. By shortening clock paths, increasing the proportion of common paths, it reduces the cumulative effect of OCV deviations on clock paths and simplifies clock timing constraints, including: The front-end clock source node, as the root node of the front-end small clock tree, receives the original clock signal from the global clock source module 100, completes the first-level distribution of the clock signal, and is the starting point of all clock signals in the front-end subsystem. The clock buffer unit performs multi-level buffering and drive enhancement on the clock signal output by the front-end clock source node to ensure that the clock signal can be stably transmitted to the write-side interface of the front-end data transmission module 300 and the FIFO core processing module 400, while compensating for wiring delay and maintaining the integrity of the clock signal. The front-end clock routing network uses high-layer metal traces on the chip to realize the physical transmission of clock signals from the front-end clock source node to each target module. Dedicated routing reduces signal crosstalk, shortens the transmission path, and reduces the impact of base delay and OCV deviation.
[0031] The connection relationships are as follows: The input terminal is electrically connected to the first output terminal of the global clock source module 100 to receive the original clock signal; The first output terminal is electrically connected to the clock input terminal of the left front-end data transmission module 300 to provide a synchronous working clock for the front-end data transmission module 300; The second output terminal is electrically connected to the write clock input terminal (W_CLK) of the FIFO core processing module 400, providing a clock reference for the write operation of the FIFO core processing module 400.
[0032] The front-end data transmission module 300, as the front-end segment of the long data channel, completes the reception, processing, and transmission of front-end data under write clock synchronization, sends the data into the FIFO buffer, reduces transmission latency by shortening the data path, and works with the front-end clock tree module 200 to alleviate the timing pressure of "clock + data" dual delay, ensuring the timing correctness of the front-end subsystem, including: The data transmission unit receives raw business data from the chip front-end data source 10, completes data latching and synchronization processing, and converts asynchronous input data into a data stream synchronized with the write clock, providing stable input for subsequent logic processing. The combinational logic processing unit performs combinational logic operations and data format conversion on the synchronized data, filters valid data, removes redundant information, optimizes data bit width and transmission efficiency, and prepares for writing to the FIFO. The front-end data cabling network enables the physical transmission of processed data from combinational logic units to the FIFO write data interface. Short-distance cabling is used to reduce data latency, signal attenuation and crosstalk, and ensure the integrity and timing stability of data transmission.
[0033] The connection relationships are as follows: The first input terminal is electrically connected to the data source (such as the arithmetic unit, register array, signal acquisition unit, etc.) of the chip front-end 10 to receive raw service data; The second input terminal is electrically connected to the first output terminal of the front-end clock tree module 200 to receive the synchronous working clock; The output terminal is electrically connected to the write data input terminal (W_DIN) of the FIFO core processing module 400, and sends the processed data to be transmitted into the FIFO core processing module 400 for buffering.
[0034] The FIFO core processing module 400 is the core unit of the system. It achieves timing isolation and data transfer between the front-end and rear-end subsystems through an asynchronous read / write mechanism. It separates the original large clock tree and long data channel, eliminating the cumulative effect of cross-clock tree OCV deviations at the source, ensuring conflict-free data transmission, and isolating timing fluctuations between the two subsystems, including: The write control unit, under the control of the write clock (W_CLK), receives input data from the front-end data transmission module 300, generates write address and write enable signal, controls data to be written sequentially to the on-chip memory array, and adjusts the write operation according to the empty / full flag to prevent data overflow; The read control unit, under the control of the read clock (R_CLK), generates the read address and read enable signal, reads the cached data sequentially from the on-chip memory array, and sends it to the subsequent data transmission module 600. At the same time, it adjusts the read operation according to the empty / full flag to prevent empty reads. The on-chip memory array, as a data cache carrier, temporarily stores the data written by the previous stage, realizes physical isolation between read and write operations, ensures stable storage of data in the asynchronous clock domain, and provides buffer space for data transmission across clock domains; The flag generation unit monitors the storage capacity of the on-chip memory array in real time, generates status flags such as empty, full, and half-full, and feeds them back to the write control unit and read control unit to control the read and write enable signals, so as to avoid data loss or errors caused by FIFO overflow or empty read.
[0035] This module is a unidirectional transmission relay unit with no reverse data / clock interaction. Its external interface consists of a write-side input and a read-side output. Internally, it uses a modular serial connection, as detailed below: Write-side input terminal: The write clock interface (W_CLK) is electrically connected to the second output of the front-end clock tree module 200 to receive the front-end write clock signal; The write data interface (W_DIN) is electrically connected to the output of the left front-end data transmission module 300 to receive the buffered data transmitted from the front-end; Read-side output: The read clock interface (R_CLK) is electrically connected to the first output terminal of the subsequent clock tree module 500 to receive the subsequent read clock signal; The read data interface (R_DOUT) is electrically connected to the first input terminal of the right-side subsequent data transmission module 600 to transmit the buffered valid data to the subsequent stage.
[0036] The downstream clock tree module 500 serves as the core of the downstream subsystem's clock supply. It constructs an independent small clock tree to provide a synchronous read clock for the downstream data transmission module 600 and the FIFO read side. Through a short-path, high common-path ratio design, it reduces the cumulative effect of OCV deviation, supports asynchronous operation with the upstream write clock, and completely eliminates cross-clock tree OCV impacts, including: The subsequent clock source node, as the root node of the subsequent small clock tree, receives the original clock signal from the global clock source module 100, completes the first-level distribution of the clock signal, and is the starting point of all clock signals in the subsequent subsystem. It is physically independent of the preceding clock source node and has no phase synchronization requirement. The clock buffer unit buffers and enhances the clock signal output by the clock source node of the subsequent stage, ensuring that the clock signal is stably transmitted to all subsequent data transmission modules 600 and FIFO read-side interface. The drive strength can be independently configured to adapt to the timing requirements of the subsequent stage. The subsequent clock routing network uses high-layer metal traces on the chip to realize the physical transmission of clock signals from the subsequent clock source node to each target module. Dedicated routing reduces signal crosstalk, shortens the transmission path, and reduces the impact of basic delay and OCV deviation. It has no physical intersection with the preceding clock routing network.
[0037] The connection relationships are as follows: The input terminal is electrically connected to the second output terminal of the global clock source module 100 to receive the original clock signal; The first output terminal is electrically connected to the read clock interface (R_CLK) of the FIFO core processing module 400, providing a clock reference for the read operation of the FIFO core processing module 400; The second output terminal is electrically connected to the second input terminal of the right-side subsequent data transmission module 600, providing a synchronous operating clock for the subsequent data transmission module 600.
[0038] The subsequent data transmission module 600, acting as a segment of the long data channel, receives, processes, and transmits FIFO buffered data under read clock synchronization, sending the data to the target unit 20 at the chip's back end. By shortening the data path, it reduces transmission latency and, in conjunction with the subsequent clock tree module, achieves rapid timing convergence of the subsequent subsystems, unaffected by the OCV deviation of the preceding stage. This includes: The data receiving unit receives the buffered data output from the FIFO read side, performs data latching and synchronization processing, and converts asynchronous input data into a data stream synchronized with the read clock, providing stable input for subsequent logic processing. The combinational logic processing unit performs combinational logic operations, data verification, format conversion, and other post-processing on the synchronized data to restore the original business data format, verify data integrity, and prepare for transmission to the target unit 20 at the back end of the chip.
[0039] The back-end data routing network enables the physical transmission of processed data from combinational logic units to the target units at the back end of the chip. It uses short-distance routing to reduce data latency, ensure the integrity and timing stability of data transmission, and has no physical intersection with the front-end data routing network.
[0040] The connection relationships are as follows: The first input terminal is electrically connected to the read data interface (R_DOUT) of the FIFO core processing module 400 to receive valid data from the FIFO buffer; The second input terminal is electrically connected to the second output terminal of the subsequent clock tree module 500 to receive the synchronous working clock; The output terminal is electrically connected to the target data receiving unit (such as register array, chip output interface, back-end arithmetic unit, etc.) of the chip back end 20 to complete the final data transmission.
[0041] In this embodiment, the system's connection architecture strictly follows Figure 1 Layout shown: The global clock source module 100 provides raw clock signals to the preceding clock tree node and the following clock tree node respectively. The preceding clock tree node and the following clock tree node are completely independent small clock trees with no common wiring or drive unit, thus completely cutting off the OCV deviation transmission path across the clock tree. The left front-end data transmission module 300 → FIFO core processing module 400 (write side) → FIFO core processing module 400 (read side) → right rear-end data transmission module 600 form a unidirectional serial data path. The FIFO serves as the only relay node to achieve timing isolation between the front-end and rear-end subsystems. The front-end subsystem completes data processing and writing synchronously based on the write clock, while the back-end subsystem completes data reading and output synchronously based on the read clock. The FIFO decouples the two-level subsystems through an asynchronous read-write mechanism. The OCV deviation only affects the internal operation of each subsystem and does not overlap across levels.
[0042] In this embodiment, all functional modules of the system are electrically connected, and the specific workflow is as follows: 1) The global clock source module outputs the original clock signal to the front-stage clock tree module and the back-stage clock tree module respectively; the front-stage and back-stage clock tree modules buffer and drive the original clock signal respectively, and then complete the clock allocation of their own small clock tree. The front-stage clock tree module outputs the write clock (W_CLK) and the front-stage synchronization clock, and the back-stage clock tree module outputs the read clock (R_CLK) and the back-stage synchronization clock. There is no phase synchronization requirement between the front-stage and back-stage clocks. 2) The chip front-end data source outputs the original business data to the left front-end data transmission module. Under the reference of the front-end synchronous clock, the front-end data transmission module completes the combinational logic processing of the data and writes the valid data to be transmitted into the FIFO core processing module through the write data interface (W_DIN). 3) Under the control of the write clock (W_CLK), the FIFO core processing module temporarily stores the data from the previous stage to the on-chip memory array through the write control unit. The empty / full flag generation unit detects the storage status in real time to prevent data overflow. 4) Under the control of the read clock (R_CLK), the rear FIFO core processing module reads the cached data from the on-chip memory array through the read control unit and transmits it to the right-side back-end data transmission module through the read data interface (R_DOUT); 5) Under the reference of the synchronous clock of the back-end, the subsequent data transmission module completes the subsequent combinational logic processing of the cached data, and transmits the processed valid data to the target receiving unit at the back end of the chip, thus completing the entire data transmission process.
[0043] In the above workflow, all operations of the front-end subsystem are completed based on the clock of the front-end clock tree module, and all operations of the back-end subsystem are completed based on the clock of the back-end clock tree module. The two are asynchronously decoupled through the FIFO core processing module. The delay deviation caused by OCV only affects the internal subsystems of their respective subsystems and has no cross-level superposition, thus completely eliminating the negative impact of OCV on the overall timing.
[0044] Although various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitations. It will be apparent to those skilled in the art that various combinations, modifications, and alterations can be made without departing from the spirit and scope of the invention. Therefore, the breadth and scope of the invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely by the appended claims and their equivalents.
Claims
1. A timing optimization system for eliminating the influence of chip OCV, characterized in that, include: The front-end clock tree module is configured to provide a synchronous write clock for the front-end data transmission module and the FIFO core processing module; The front-end data transmission module is configured to receive, process, and transmit front-end data under write clock synchronization, and send the data to the FIFO core processing module buffer. The FIFO core processing module is configured to perform timing isolation and data transfer between the front-end and back-end clock tree modules through an asynchronous read / write mechanism, in order to separate the original clock tree and data channel; The subsequent clock tree module is configured to provide a synchronous read clock for the subsequent data transmission module and the FIFO core processing module. as well as The subsequent data transmission module is configured to receive, process, and transmit cached data under read clock synchronization, and then send the data to the target unit at the back end of the chip.
2. The system according to claim 1, characterized in that, Also includes: The global clock source module is configured to generate and distribute raw reference clock signals to the front-end and rear-end clock tree modules.
3. The system according to claim 1, characterized in that, The front-end clock tree module includes: The front-end clock source node is configured as the root node of the front-end small clock tree, receiving the original reference clock signal from the global clock source module and completing the first-level distribution of the clock signal. The clock buffer unit is configured to perform multi-level buffering and drive enhancement on the clock signal output by the front-end clock source node, ensuring that the clock signal can be stably transmitted to the write-side interface of the front-end data transmission module and the FIFO core processing module. The front-end clock routing network is configured to enable the physical transmission of clock signals from the front-end clock source node to each target module.
4. The system according to claim 1, characterized in that, The front-end data transmission module includes: The data transmission unit is configured to receive raw business data from the chip front-end data source, perform data latching and synchronization processing, and convert asynchronous input data into a data stream synchronized with the write clock. The combinational logic processing unit is configured to preprocess the synchronized data to remove redundant information; The front-end data cabling network is configured to enable the physical transfer of processed data from the combinational logic unit to the write data interface of the FIFO core processing module.
5. The system according to claim 1, characterized in that, The input terminal of the front-end clock tree module is electrically connected to the first output terminal of the global clock source module to receive the original reference clock signal; The first output terminal of the front-end clock tree module is electrically connected to the clock input terminal of the front-end data transmission module, providing a synchronous working clock for the front-end data transmission module; The second output terminal of the front-end clock tree module is electrically connected to the write clock input terminal of the FIFO core processing module, providing a clock reference for the write operation of the FIFO core processing module. The first input terminal of the front-end data transmission module is electrically connected to the data source at the front end of the chip to receive raw business data; The second input terminal of the front-end data transmission module is electrically connected to the first output terminal of the front-end clock tree module to receive the synchronous working clock; The output of the front-end data transmission module is electrically connected to the write data input of the FIFO core processing module, and the processed data to be transmitted is sent to the FIFO core processing module for buffering.
6. The system according to claim 1, characterized in that, The FIFO core processing module includes: The write control unit is configured to receive input data from the preceding data transmission module under the control of the write clock, generate write address and write enable signal, control data to be written sequentially to the on-chip memory array, and adjust the write operation according to the empty / full status flag to prevent data overflow. The read control unit is configured to generate read address and read enable signal under read clock control, read cached data sequentially from the on-chip memory array and send it to the subsequent data transmission module, and adjust the read operation according to empty / full status flags to prevent empty reads. The on-chip storage array is configured to temporarily store data written to the previous stage, achieving physical isolation between read and write operations; The empty / full flag generation unit is configured to monitor the storage capacity of the on-chip memory array in real time, generate empty, full, and half-full status flags, and feed them back to the write control unit and the read control unit.
7. The system according to claim 1, characterized in that, The subsequent clock tree module and the preceding clock tree module have a symmetrical structure, including: The subsequent clock source node is configured to receive the original reference clock signal from the global clock source module and complete the first-level distribution of the clock signal. The clock buffer unit is configured to buffer and enhance the clock signal output by the subsequent clock source node, ensuring that the clock signal is stably transmitted to the read-side interface of the subsequent data transmission module and the FIFO core processing module. The subsequent clock routing network is configured to enable the physical transmission of clock signals from the subsequent clock source node to each target module.
8. The system according to claim 1, characterized in that, The subsequent data transmission module includes: The data receiving unit is configured to receive buffered data output from the FIFO core processing module, perform data latching and synchronization processing, and convert asynchronous input data into a data stream synchronized with the read clock. Combinational logic processing unit, which is configured to post-process synchronized data to restore the original business data format; The back-end data routing network is configured to enable the physical transfer of processed data from combinational logic units to target units at the back end of the chip.
9. The system according to claim 1, characterized in that, The input terminal of the subsequent clock tree module is electrically connected to the second output terminal of the global clock source module to receive the original clock signal; The first output of the subsequent clock tree module is electrically connected to the read clock interface of the FIFO core processing module, providing a clock reference for the read operation of the FIFO core processing module. The second output terminal of the subsequent clock tree module is electrically connected to the second input terminal of the subsequent data transmission module to provide a synchronous working clock for the subsequent data transmission module; The first input terminal of the subsequent data transmission module is electrically connected to the read data interface of the FIFO core processing module to receive valid data from the FIFO buffer. The second input terminal of the subsequent data transmission module is electrically connected to the second output terminal of the subsequent clock tree module to receive the synchronous working clock; The output of the subsequent data transmission module is electrically connected to the target data receiving unit at the back end of the chip to complete the final data transmission.
10. A method for operating the system according to any one of claims 1-9, characterized in that, include: The front-end clock tree module provides a synchronous write clock for the front-end data transmission module and the FIFO core processing module. The front-end data transmission module performs the reception, processing and transmission of front-end data under write clock synchronization, and sends the data to the FIFO core processing module buffer; The FIFO core processing module performs timing isolation and data transfer between the front-end and back-end clock tree modules through an asynchronous read / write mechanism to separate the original clock tree and data channel; The subsequent clock tree module provides a synchronous read clock for the subsequent data transmission module and the FIFO core processing module. as well as The subsequent data transmission module receives, processes, and transmits the cached data under read clock synchronization, and then sends the data to the target unit at the back end of the chip.
Citation Information
Patent Citations
Multi-channel AFE signal acquisition method and system based on adaptive adjustment
CN120389752A
Power Supply Current Spike Reduction Techniques for an Integrated Circuit
US20090183019A1
CPU Current Ripple and OCV Effect Mitigation
US20140258765A1