Programmable device configuration memory system

The pipelined configuration memory system addresses inefficiencies in programmable devices by segmenting data lines, improving write/read bandwidth and reducing latency for applications like AI and data centers.

JP7689117B2Active Publication Date: 2025-06-05XILINX INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022525486
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-31
Filing Date
2020-08-11
Publication Date
2025-06-05
Estimated Expiration
2040-08-11

AI Technical Summary

Technical Problem

The performance of configuration memory systems in programmable devices is limited by distributed memory systems, where data lines stretch across the entire device width, leading to inefficient write/read operations.

Method used

A configuration memory system with a pipelined bidirectional data line structure and source clocking, which segments data lines between configuration memory read/write pipeline units to improve performance and minimize area and cost.

Benefits of technology

The pipelined configuration memory system enhances write/read bandwidth and reduces latency, benefiting applications requiring fast device state reads in AI, data centers, and automotive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689117000001
    Figure 0007689117000001
  • Figure 0007689117000002
    Figure 0007689117000002
  • Figure 0007689117000003
    Figure 0007689117000003
Patent Text Reader

Abstract

An exemplary configuration system for a programmable device includes: a configuration memory read / write unit configured to receive configuration data for storage in a configuration memory of the programmable device, the configuration memory comprising a plurality of frames; a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; and a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Technical Field TECHNICAL FIELD Examples of the present disclosure relate generally to programmable devices, and more particularly to configuration memory systems of programmable devices. [Background technology]

[0002] background Programmable devices such as field programmable gate arrays (FPGAs) and systems-on-chips (SoCs) with FPGA programmable fabric are gaining momentum in artificial intelligence (AI), datacenter, and automotive applications. One technique useful for these applications is partial reconfiguration of programmable devices. Partial reconfiguration is the ability to dynamically modify logic blocks of a programmable device by downloading partial configuration bit files while the remaining logic continues to operate uninterrupted. Traditionally, partial reconfiguration performance is limited by distributed memory systems within programmable devices, where data lines stretch across the entire device width and the memory controller must perform a series of operations before initiating the next write / read. Previously Therefore, it is desirable to improve the performance of configuration memory systems in programmable devices. Summary of the Invention [Means for solving the problem]

[0003] overview Techniques are described for providing a configuration memory system in a programmable device. In one example, the configuration system of the programmable device includes a configuration memory read / write unit configured to receive configuration data for storage in a configuration memory of the programmable device, the configuration memory comprising a plurality of frames, a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit, and a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs.

[0004] In another example, a programmable device includes a programmable fabric, a configuration memory for storing data for configuring the programmable fabric, the configuration memory having a plurality of frames, a configuration memory read / write unit configured to receive configuration data for storage in the configuration memory, a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit, and a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs.

[0005] In another example, a method for configuring a programmable device includes the steps of: a configuration memory read / write unit receiving configuration data for storage in a configuration memory of the programmable device, the configuration memory comprising a plurality of frames; providing the configuration data to a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; and providing the configuration data from the plurality of configuration memory read / write controllers to a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs.

[0006] Other non-limiting examples of the disclosed technology are provided as follows. Example 1. A configuration system for a programmable device, comprising: 1. A system comprising: a configuration memory read / write unit configured to receive configuration data for storage in a configuration memory of a programmable device, the configuration memory comprising a plurality of frames; a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; and a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs.

[0007] Example 2. The configuration system of Example 1, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes a flip-flop, a multiplexer having an output coupled to an input of the flip-flop, a first input coupled to a respective one of the plurality of configuration memory read / write controllers, and a second input coupled to one of the buffers in the respective FSR, and a first buffer coupled to the output of the flip-flop, the output of the first buffer being coupled to a respective one of the plurality of configuration memory read / write controllers.

[0008] Example 3. The configuration system of Example 2, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes: a first inverter having an output coupled to the second input of the multiplexer; a second inverter having an input coupled to the input of the first inverter and an output coupled to one of the buffers; and a third inverter having an output coupled to the input of the first inverter and an input coupled to one of the buffers.

[0009] Example 4. The configuration system of Example 3, wherein the configuration memory read / write pipeline integrating each of the plurality of FSRs includes a second buffer having an input coupled to the output of the flip-flop and an input coupled to one of the buffers.

[0010] Example 5. The configuration system of example 4, wherein each of the first and second buffers comprises a three-state buffer, and the third inverter comprises a three-state inverter.

[0011] Example 6. The configuration system of Example 1, wherein the pipeline in each of the plurality of FSRs includes a data line pipe, a frame address register (FAR) pipe, and a control pipe, the data line pipe propagating configuration data, the FAR pipe propagating address information, and the control pipe propagating control signals for latching the data line pipe and the FAR pipe.

[0012] Example 7. The configuration system of example 1, wherein the pipeline in each of the plurality of FSRs includes a tag pipe configured to latch read data on the data line pipe.

[0013] Example 8. A programmable device comprising: a programmable fabric; a configuration memory for storing data for configuring the programmable fabric, the configuration memory comprising a plurality of frames; a configuration memory read / write unit configured to receive configuration data for storage in the configuration memory of the programmable device, the configuration memory comprising a plurality of frames; a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; and a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs.

[0014] Example 9. The programmable device of Example 8, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes a flip-flop, a multiplexer having an output coupled to an input of the flip-flop, a first input coupled to a respective one of the plurality of configuration memory read / write controllers, and a second input coupled to one of the buffers in the respective FSR, and a first buffer coupled to the output of the flip-flop, an output of the first buffer being coupled to a respective one of the plurality of configuration memory read / write controllers.

[0015] Example 10. The programmable device of Example 9, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes: a first inverter having an output coupled to the second input of the multiplexer; a second inverter having an input coupled to the input of the first inverter and an output coupled to one of the buffers; and a third inverter having an output coupled to the input of the first inverter and an input coupled to one of the buffers.

[0016] Example 11. The programmable device of example 10, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes a second buffer having an input coupled to the output of the flip-flop and an input coupled to one of the buffers.

[0017] Example 12. The programmable device of example 11, wherein each of the first and second buffers comprises a three-state buffer and the third inverter comprises a three-state inverter.

[0018] Example 13. The programmable device of example 8, wherein the pipeline in each of the plurality of FSRs includes a data line pipe, a frame address register (FAR) pipe, and a control pipe, the data line pipe carrying configuration data, the FAR pipe carrying address information, and the control pipe carrying control signals for latching the data line pipe and the FAR pipe.

[0019] Example 14. The programmable device of example 8, wherein the pipeline in each of the plurality of FSRs includes a tag pipe configured to latch the read data on the data line pipe.

[0020] Example 15. A method for configuring a programmable device comprising the steps of: receiving, at a configuration memory read / write unit, configuration data for storing in a configuration memory of the programmable device, the configuration memory comprising a plurality of frames; providing the configuration data to a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; and providing the configuration data from the plurality of configuration memory read / write controllers to a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs.

[0021] Example 16. The method of Example 15, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes a flip-flop, a multiplexer having an output coupled to an input of the flip-flop, a first input coupled to a respective one of the plurality of configuration memory read / write controllers, and a second input coupled to one of the buffers in the respective FSR, and a first buffer coupled to the output of the flip-flop, the output of the first buffer being coupled to a respective one of the plurality of configuration memory read / write controllers.

[0022] Example 17. The method of Example 16, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes: a first inverter having an output coupled to the second input of the multiplexer; a second inverter having an input coupled to the input of the first inverter and an output coupled to one of the buffers; and a third inverter having an output coupled to the input of the first inverter and an input coupled to one of the buffers.

[0023] Example 18. The method of Example 17, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes a second buffer having an input coupled to the output of the flip-flop and an input coupled to one of the buffers.

[0024] Example 19. The method of example 18, wherein each of the first and second buffers comprises a three-state buffer and the third inverter comprises a three-state inverter.

[0025] Example 20. The method of example 15, wherein the pipeline in each of the plurality of FSRs includes a data line pipe, a frame address register (FAR) pipe, and a control pipe, the data line pipe propagating configuration data, the FAR pipe propagating address information, and the control pipe propagating control signals for latching the data line pipe and the FAR pipe.

[0026] These and other aspects can be understood with reference to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS So that the above features may be understood in detail, a more particular description than that briefly summarized above can be had by reference to exemplary implementations, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings depict only typical exemplary implementations and therefore should not be considered as limiting the scope thereof. [Brief description of the drawings]

[0027] [Figure 1A] FIG. 1 is a block diagram illustrating a programmable device according to an example. [Figure 1B] FIG. 1 is a block diagram illustrating a programmable IC according to an example. [Figure 1C] FIG. 1 is a block diagram illustrating a SOC implementation of a programmable IC according to an example. [Figure 1D] 1 illustrates a field programmable gate array (FPGA) implementation of a programmable IC including a PL, according to an example. [Diagram 2]FIG. 2 is a block diagram illustrating a configuration subsystem according to an example. [Diagram 3] FIG. 2 is a block diagram illustrating a configuration pipeline according to an example. [Figure 4] FIG. 2 is a block diagram illustrating a configuration memory read / write pipeline unit according to an example. [Diagram 5] FIG. 13 is a schematic diagram illustrating a write operation according to an example. [Figure 6] FIG. 13 is a schematic diagram illustrating a read operation according to an example. [Figure 7] FIG. 1 is a flow diagram illustrating a method for configuring a programmable device according to an example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] For ease of understanding, wherever possible, identical reference numbers have been used to designate identical elements common to the figures, and it is believed that elements of one example may be beneficially incorporated in other examples.

[0029] Detailed Description Various features are described below with reference to the drawings. It should be noted that the drawings may or may not be drawn to scale, and that elements of similar structure or function are represented by similar reference numbers throughout the drawings. It should be noted that the drawings are intended only to facilitate the description of features. They are not intended as an exhaustive description of the claimed invention, or as limitations on the scope of the claimed invention. Furthermore, the illustrated example need not have all aspects or advantages shown. An aspect or advantage described in connection with a particular example is not necessarily limited to that example, and may be implemented in any other example, even if not so shown or explicitly described.

[0030] Techniques are described for providing a configuration memory system within a programmable device. In an example, the configuration memory system uses a unique structure to pipeline bidirectional data lines with source clocking to achieve improved performance over previous configuration memory systems while minimizing area and cost. The configuration memory system described herein can benefit a variety of applications that utilize partial reconfiguration, including applications that require fast device state reads via the configuration memory system, such as artificial intelligence (AI), data centers, automotive applications, and emulation applications. These and other aspects are described below with reference to the drawings.

[0031] 1A is a block diagram illustrating a programmable device 54 according to an example. The programmable device 54 includes a plurality of programmable integrated circuits (ICs) 1, e.g., programmable ICs 1A, 1B, 1C, and 1D. In the example, each programmable IC 1 is an IC die disposed on an interposer 90. Each programmable IC 1 comprises a super logic region (SLR) 53, e.g., SLRs 53A, 53B, 53C, and 53D, of the programmable device 54. The programmable ICs 1 are interconnected via conductors (referred to as super long lines (SLLs) 52) on the interposer 90.

[0032] FIG. 1B is a block diagram illustrating a programmable IC 1 according to an example. The programmable IC 1 can be used to implement one of the programmable ICs 1A-1D in the programmable device 128 or the programmable device 54. The programmable IC 1 includes a programmable logic 3 (also called a programmable fabric), a configuration logic 25, and a configuration memory 26. The programmable IC 1 can be connected to external circuits such as a non-volatile memory 27, a DRAM 28, and other circuits 29. The programmable logic 3 includes logic cells 30, support circuits 31, and programmable interconnects 32. The logic cells 30 include circuits that can be configured to perform a general logic function of multiple inputs. The support circuits 31 include dedicated circuits such as transceivers, input / output blocks, digital signal processors, memories, and the like. The logic cells and the support circuits 31 can be interconnected using the programmable interconnects 32. Information for programming the logic cells 30, setting parameters of the support circuits 31, and programming the programmable interconnects 32 is stored in the configuration memory 26 by the configuration logic 25. The configuration memory 26 is organized into a number of frames 95. The configuration logic 25 can obtain the configuration data from a non-volatile memory 27 or any other source (e.g., from a DRAM 28, or other circuitry 29). In some examples, the programmable IC 1 includes a processing system 2. The processing system 2 can include a microprocessor, memory, support circuitry, IO circuitry, etc. In some examples, the programmable IC 1 includes a network on chip (NOC) 55 and a data processing engine (DPE) array 56. The NOC 55 is configured to provide communication between subsystems of the programmable IC 1, such as between the PS 2, the PL 3, and the DPE array 56. The DPE array 56 can include an array of DPEs configured to perform data processing, such as an array of vector processors.

[0033] FIG. 1C is a block diagram showing a SOC implementation of a programmable IC 1 according to an example. In this example, the programmable IC 1 comprises a processing system 2 and a programmable logic 3. The processing system 2 includes various processing units such as a real-time processing unit (RPU) 4, an application processing unit (APU) 5, a graphics processing unit (GPU) 6, a configuration and security unit (CSU) 12, and a platform management unit (PMU) 11. The processing system 2 also includes various support circuits such as an on-chip memory (OCM) 14, a transceiver 7, peripherals 8, an interconnect 16, a DMA circuit 9, a memory controller 10, peripherals 15, and a multiplexed IO (MIO) circuit 13. The processing units and the support circuits are interconnected by the interconnect 16. The PL 3 is also connected to the interconnect 16. The transceiver 7 is coupled to an external pin 24. The PL 3 is coupled to an external pin 23. The memory controller 10 is coupled to an external pin 22. The MIO 13 is coupled to an external pin 20. The PS2 is generally coupled to external pins 21. The APU 5 may include a CPU 17, a memory 18, and support circuits 19.

[0034] With reference to the PS2, each processing unit includes one or more central processing units (CPUs) and associated circuitry such as memory, interrupt controllers, direct memory access (DMA) controllers, memory management units (MMUs), floating point units (FPUs), etc. Interconnect 16 includes various switches, buses, communication links, etc. configured to interconnect each processing unit as well as to interconnect other components within the PS2 to each processing unit.

[0035] The OCM 14 includes one or more RAM modules that may be distributed throughout the PS2. For example, the OCM 14 may include a battery-backed RAM (BBRAM), a tightly coupled memory (TCM), and the like. The memory controller 10 may include a DRAM interface for accessing an external DRAM. The peripherals 8, 15 may include one or more components that provide an interface to the PS2. For example, the peripherals 15 may include a graphics processing unit (GPU), a display interface (e.g., a DisplayPort, a high-definition multimedia interface (HDMI) port, and the like), a universal serial bus (USB) port, an Ethernet port, a universal asynchronous transceiver (UART) port, a serial peripheral interface (SPI) port, a general-purpose IO (GPIO) port, a serial advanced technology attachment (SATA) port, a PCIe port, and the like. The peripherals 15 may be coupled to the MIO 13. The peripherals 8 may be coupled to the transceiver 7. The transceiver 7 may include a serializer / deserializer (SERDES) circuit, a multi-gigabit transceiver (MGT), and the like.

[0036] FIG. 1D illustrates a field programmable gate array (FPGA) implementation of a programmable IC 1 including PL3. The PL3 illustrated in FIG. 1D can be used in any of the examples of programmable devices described herein. For example, PL3 can include a number of different programmable tiles including transceivers 37, configurable logic blocks ("CLBs") 33, random access memory blocks ("BRAMs") 34, input / output blocks ("IOBs") 36, configuration and clocking logic ("CONFIG / CLOCKS") 42, digital signal processing blocks ("DSPs") 35, specialized input / output blocks ("I / Os") 41 (e.g., configuration and clock ports), and other programmable logic 39 such as digital clock managers, analog-to-digital converters, system monitoring logic, etc. PL3 can also include a PCIe interface 40, an analog-to-digital converter (ADC) 38, etc.

[0037] In some PLs, each programmable tile may include at least one programmable interconnect element ("INT") 43 having connections to input and output terminals 48 of programmable logic elements within the same tile, as shown by the example included at the top of FIG. 1D. Each programmable interconnect element 43 may also include connections for interconnecting segments 49 of adjacent programmable interconnect elements within the same tile or other tiles. Each programmable interconnect element 43 may also include connections for interconnecting segments 50 of general-purpose routing resources between logic blocks (not shown). The general-purpose routing resources may include routing channels between logic blocks (not shown) that comprise tracks of interconnect segments (e.g., interconnect segments 50) and switch blocks (not shown) for connecting the interconnect segments. The interconnect segments of the general-purpose routing resources (e.g., interconnect segments 50) may span one or more logic blocks. The programmable interconnect elements 43 together with the general-purpose routing resources implement a programmable interconnect structure ("programmable interconnect") for the illustrated PL.

[0038] In an exemplary implementation, the CLB 33 may include a configurable logic element ("CLE") 44 that can be programmed to implement a single programmable interconnect element ("INT") 43 in addition to user logic. The BRAM 34 may include a BRAM logic element ("BRL") 45 in addition to one or more programmable interconnect elements. Typically, the number of interconnect elements included in a tile depends on the height of the tile. In the illustrated example, the BRAM tile has the same height as five CLBs, although other numbers (e.g., four) may be used. The DSP tile 35 may include a DSP logic element ("DSPL") 46 in addition to any suitable number of programmable interconnect elements. The IOB 36 may include, for example, two instances of an input / output logic element ("IOL") 47 in addition to one instance of a programmable interconnect element 43. As will be apparent to one skilled in the art, the actual I / O pads connected to, for example, the I / O logic element 47 are typically not limited to the area of ​​the input / output logic element 47.

[0039] In the illustrated example, a horizontal region near the center of the die (shown in FIG. 3D) is used for configuration, clocks, and other control logic. Vertical columns 51 extending from this horizontal region or column are used to distribute clock and configuration signals across the width of the PL.

[0040] Some PLs utilizing the architecture shown in FIG. 1D include additional logic blocks that break up the regular columnar structure that constitutes most of the PL. The additional logic blocks can be programmable blocks and / or dedicated logic. Note that FIG. 1D is intended to illustrate merely an exemplary PL architecture. For example, the number of logic blocks in the rows included in the upper portion of FIG. 1D, the relative widths of the rows, the number and order of the rows, the types of logic blocks included in the rows, the relative sizes of the logic blocks, and the interconnect / logic implementation are purely exemplary. For example, in an actual PL, multiple adjacent CLB rows will typically be included where CLBs appear to facilitate efficient implementation of user logic, although the number of adjacent CLB rows will vary depending on the overall size of the PL.

[0041] 2 is a block diagram illustrating a configuration subsystem 200 according to an example. The configuration subsystem 200 is disposed, for example, in a programmable IC1 to configure programmable logic. The configuration subsystem 200 includes a configuration memory read / write unit (herein referred to as a Cframe unit (CFU) 202), a number of configuration memory read / write controllers (herein referred to as a Cframe engine 204), a number of configuration memory read / write pipeline units (herein referred to as a Cpipe 206), and configuration memory cells in a fabric sub-region (FSR) 208. The CFU 202 may be disposed in a platform management controller or similar component (e.g., PMU 11) of the programmable IC1. The CFU 202 is configured to receive input configuration data of the programmable IC. The CFU 202 serves as a master configuration controller for the configuration subsystem 200.

[0042] The CFU 202 is coupled to each of the Cframe engines 204. Each Cframe engine 204 comprises a configuration frame write / read controller. A "frame" is a unit of configuration data that is stored in or read from a set of configuration memory cells. A frame has a "height" based on the number of configuration memory cells it contains data in. Each Cframe engine 204 provides data to one or more FSRs 208 via a pipeline that comprises a Cpipe 206 and an FSR 208. The Cpipe 206 is further described below. Each FSR 208 is a region of associated configuration memory having programmable logic and the height of the frame.

[0043] FIG. 3 is a block diagram illustrating a configuration pipeline 300 according to an example. The configuration pipeline 300 includes a Cframe engine 204 and one or more FSRs 208 (e.g., two are shown). Each FSR 208 includes a buffer (Cbrk 302), configuration memory cells (mem cells 304), and a Cpipe 206. Each Cbrk 302 includes a bidirectional buffer. The mem cells 304 are disposed between the Cbrks 302. In general, each FSR 208 includes one or more Cbrks 302 with mem cells 304 disposed between them. In operation, configuration data is provided from the Cframe engine 204 to the mem cells 304 via the Cbrk 302. In an example, each FSR 208 includes a Cpipe 206 (e.g., disposed at the end of the Cbrk memory cell chain). Without the Cpipe 206, data lines would stretch across the entire width of the FSR 208, which limits the write / read bandwidth of the configuration memory. By adding a Cpipe 206 for each FSR 208, the data lines are segmented between consecutive Cpipes 206. The data line segments are smaller than the width of the FSR 208, which improves the write / read bandwidth of the configuration memory.

[0044] 4 is a block diagram illustrating a Cpipe 206 according to an example. The Cpipe 206 includes a buffer 402, a flip-flop 404, a multiplexer 406, an inverter 408, an inverter 410, an inverter 412, and a buffer 414. An input ("1") of the multiplexer 406 is coupled to the Cframe engine 204 (e.g., directly or through other components). Another input ("0") of the multiplexer 406 is coupled to the output of the inverter 408. A control input of the multiplexer 406 is coupled to a control signal C2. An output of the multiplexer 406 is coupled to an input of the flip-flop 404. An output of the flip-flop 404 is coupled to an input of the buffer 402. An output of the buffer 402 is coupled to the Cframe engine 204 (e.g., directly or through other components). In the example, the buffer 402 is a three-state buffer and includes a control input coupled to a control signal C1.

[0045] The output of flip-flop 404 is coupled to the input of buffer 414. The output of buffer 414 is coupled to Cbrk 302. In the example, buffer 414 is a three-state buffer and includes a control input coupled to control signal C3. The input of inverter 412 is coupled to Cbrk 302. The output of inverter 412 is coupled to the input of inverter 408. The input of inverter 410 is coupled to the output of inverter 412. The output of inverter 410 is coupled to Cbrk 302. In the example, inverter 410 is a three-state inverter and includes a control input coupled to signal C4. The circuit for one data line is shown in FIG. 4. The circuit is repeated for each data line passing through Cpipe 206.

[0046] In operation, during a write, configuration data is coupled to the "1" input of multiplexer 406. Control signal C2 is set to select the "1" input of multiplexer 406. After the configuration data is stored in flip-flop 404, it is received by Cbrk 302 via buffer 414. Control signal C3 is set to enable buffer 414. Control signal C4 is set to disable inverter 410. In this manner, configuration data is passed from the Cframe engine 204 through the Cpipe 206 to Cbrk 302 for writing to the configuration memory.

[0047] During a read, the read data is coupled to the "0" input of the multiplexer 406. In particular, inverter 412 and inverter 410 form a latch for latching the read data from Cbrk 302. The latched read data is coupled to the "0" input of the multiplexer 406 via inverter 408. Control signal C4 is set to enable inverter 410 and thus the latch. Control signal C2 is set to select the "0" input of the multiplexer 406. This read data is stored in flip-flop 404 and read by the Cframe engine 204 via buffer 402. Control signal C1 is set to enable buffer 402. Control signal C3 is set to disable buffer 414. In this manner, the read data is passed from Cbrk 302 to the Cframe engine 204 for reading from the configuration memory.

[0048] FIG. 5 is a schematic diagram showing a write operation according to an example. The central circle (labeled CTRL pipe) of the diagram is a symbol pipeline stage for matching data line propagation delay. In one example, there are three types of pipes: a data line pipe 502 (labeled Data pipe), a frame address register (FAR) pipe 504 (labeled FAR pipe), and a CTRL pipe 506. The data pipe 502 carries configuration data, the FAR pipe 504 carries address information for configuring the configuration memory, and the CTRL pipe 506 carries control signals for latching the data pipe and the FAR pipe 504.

[0049] During a write operation, the Cframe engine 204 generates a write waveform that includes the frame data and the necessary control signals / sequences. Now the entire waveform must propagate through the Cpipe 206 in lock step. The control signal path includes additional pipeline stages that match the data line propagation time. The data line pipeline 502 is a multi-cycle path, so the tag / token is used to latch the data line value into the Cpipe 206 after it has stabilized. Each FSR 208 decodes the frame address locally to determine if the waveform is for it or not. The waveform flows from the Cframe engine 204 to the edge of the device regardless of the target frame position.

[0050] FIG. 6 is a schematic diagram showing a read operation according to an example. The read waveform (no data), FAR 504, control 506, and rdata_tag 602 generated by the Cframe engine propagate to the edge of the device. The data lines 502 propagate back to the Cframe engine 204. Each frame decodes the read FAR and only active frames are taken from the read frame. The read data is captured by rdata_tag into the cpipe and then propagates back to the Cframe engine 204. In FIG. 6, the FAR and control do not show an extra circle in the middle of the figure because the data line propagation does not need to be aligned during a read operation as it is during a write operation. In the example, additional pipeline stages can be added for the FAR and control as well as the write operation for timing or noise reduction purposes. A multiplexer (Mux) is provided to multiplex the rdata_tag and read waveforms per FSR.

[0051] Since rdata_tag runs against a clock (sourced from the Cframe engine 204 to the edge of the device), the tag must be at least two clocks wide to ensure it is not missed by the synchronizer of the next Cframe. After synchronization, rdata_tag is pulled back to at least two clocks wide. Rdata_tag is used to latch read data on the slowly propagating data lines. Additional circles of rdata_tag are present to match the propagation delay.

[0052] The configuration system described herein uses source clocking. The configuration system does not use clock trees due to its distributed nature. Once a transaction leaves the Cframe engine 204, it is difficult to stall the pipeline. Therefore, the Cframe engine 204 must analyze the incoming transactions it receives and police the traffic to the pipeline to ensure that the pipeline does not overrun. A distributed pipeline can also generate read hazard conditions. If a new read is closer to the Cframe engine 204 than a previous read, the read data may collide. Therefore, the Cframe engine 204 can detect such hazards and delay the new transaction if necessary.

[0053] For frame addressing, in previous systems, the frame address is column / primary address based; that is, each block has its unique column / primary address and has N frames within each column. When N frames are reached, the column / primary address is incremented based on a feedback signal. The configuration system described herein uses a linear addressing scheme, which eliminates the performance limitations associated with the previous schemes mentioned above.

[0054] 7 is a flow diagram illustrating a method 700 for configuring a programmable device according to an example. The method 700 begins at step 702, where a CFU 202 receives configuration data for storage in a configuration memory 26 of the programmable device 1. The configuration memory 26 comprises a plurality of frames 95. At step 704, the CFU 202 provides the configuration data to a plurality of Cframe engines 204 coupled to the CFU 202. At step 706, the Cframe engines 204 provide the configuration data to a plurality of FSRs 208. Each FSR 208 includes a pipeline of memory cells 304 of the configuration memory arranged between a buffer (Cbrk 302) and a Cpipe circuit 206 coupled between the pipeline and the next of the FSRs 208.

[0055] While the above is directed to particular examples, other and further examples may be devised without departing from the basic scope thereof, which is determined by the claims that follow.

Claims

1. 1. A configuration system for a programmable device, comprising: a configuration memory read / write unit configured to receive configuration data for storage in a configuration memory of the programmable device, the configuration memory comprising a plurality of frames; a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs; The configuration memory read / write pipeline unit in each of the plurality of FSRs comprises: Flip-flops and a multiplexer having an output coupled to an input of the flip-flop, a first input coupled to a respective one of the plurality of configuration memory read / write controllers, and a second input coupled to one of the buffers in the respective FSR; a first buffer coupled to an output of the flip-flop, an output of the first buffer being coupled to the respective one of the plurality of configuration memory read / write controllers.

2. The configuration memory read / write pipeline unit in each of the plurality of FSRs comprises: a first inverter having an output coupled to the second input of the multiplexer; an input coupled to an input of the first inverter and to one of the buffers; a second inverter having a combined output; a third inverter having an output coupled to the input of the first inverter and an input coupled to the one of the buffers.

3. The configuration memory read / write pipeline unit in each of the plurality of FSRs comprises:

3. The configuration system of claim 2 including a second buffer having an input coupled to the output of the flip-flop and an input coupled to the one of the buffers.

4. 4. The configuration system of claim 3, wherein each of the first and second buffers comprises a three-state buffer, and the third inverter comprises a three-state inverter.

5. 2. The configuration system of claim 1, wherein the pipeline in each of the plurality of FSRs includes a data line pipe, a frame address register (FAR) pipe, and a control pipe, the data line pipe carrying configuration data, the FAR pipe carrying address information, and the control pipe carrying control signals for latching the data line pipe and the FAR pipe.

6. 2. The configuration system of claim 1, wherein the pipeline in each of the plurality of FSRs includes a tag pipe configured to latch read data on a data line pipe.

7. 1. A programmable device, comprising: A programmable fabric; a configuration memory for storing data for configuring the programmable fabric, the configuration memory comprising a plurality of frames; a configuration memory read / write unit configured to receive configuration data for storage in a configuration memory of the programmable device, the configuration memory comprising a plurality of frames; a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs; The configuration memory read / write pipeline unit in each of the plurality of FSRs comprises: Flip-flops and a multiplexer having an output coupled to an input of the flip-flop, a first input coupled to a respective one of the plurality of configuration memory read / write controllers, and a second input coupled to one of the buffers in the respective FSR; a first buffer coupled to an output of the flip-flop, an output of the first buffer being coupled to the respective one of the plurality of configuration memory read / write controllers.

8. The configuration memory read / write pipeline unit in each of the plurality of FSRs comprises: a first inverter having an output coupled to the second input of the multiplexer; a second inverter having an input coupled to the input of the first inverter and an output coupled to one of the buffers; a third inverter having an output coupled to the input of the first inverter and an input coupled to the one of the buffers.

9. 9. The programmable device of claim 8, wherein the configuration memory read / write pipeline unit in each of the plurality of FSRs includes a second buffer having an input coupled to the output of the flip-flop and an input coupled to the one of the buffers.

10. 10. The programmable device of claim 9, wherein each of the first and second buffers comprises a three-state buffer and the third inverter comprises a three-state inverter.

11. 8. The programmable device of claim 7, wherein the pipeline in each of the plurality of FSRs includes a data line pipe, a frame address register (FAR) pipe, and a control pipe, the data line pipe carrying configuration data, the FAR pipe carrying address information, and the control pipe carrying control signals for latching the data line pipe and the FAR pipe.

12. 8. The programmable device of claim 7, wherein the pipeline in each of the plurality of FSRs includes a tag pipe configured to latch read data on a data line pipe.

13. 1. A method of configuring a programmable device, comprising the steps of: receiving, in a configuration memory read / write unit, configuration data for storage in a configuration memory of the programmable device, the configuration memory comprising a plurality of frames; providing the configuration data to a plurality of configuration memory read / write controllers coupled to the configuration memory read / write unit; providing the configuration data from the plurality of configuration memory read / write controllers to a plurality of fabric sub-regions (FSRs) respectively coupled to the plurality of configuration memory read / write controllers, each FSR including a pipeline of memory cells of the configuration memory arranged between buffers and a configuration memory read / write pipeline unit coupled between the pipeline and a next one of the plurality of FSRs; The configuration memory read / write pipeline unit in each of the plurality of FSRs comprises: Flip-flops and a multiplexer having an output coupled to an input of the flip-flop, a first input coupled to a respective one of the plurality of configuration memory read / write controllers, and a second input coupled to one of the buffers in the respective FSR; a first buffer coupled to an output of the flip-flop, an output of the first buffer being coupled to the respective one of the plurality of configuration memory read / write controllers.

Citation Information

Patent Citations

  • Configurable Mixed-Memory System

    JP2016504650A

  • High speed FPGA boot-up through concurrent multi-frame configuration scheme

    US20160307612A1

  • Sector-Aligned Memory Accessible to Programmable Logic Fabric of Programmable Logic Device

    US20190043536A1

  • Interface for parallel configuration of programmable devices

    US20190103872A1

  • Dedicated input / output first in / first out module for a field programmable gate array

    US6867615B1