A memory sequencer system and a memory sequencing method applying the same
By combining command and address sequencers with data sequencers to form complex sequences, the limitations of existing memory sequencer systems and the problem of synchronization expansion are solved, enabling support for different memory protocols and efficient memory interface training, calibration and debugging.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-16
- Publication Date
- 2026-04-07
AI Technical Summary
Existing memory sequencer systems only support pattern-based address and data sequencers, which limits pattern complexity, cannot support different memory protocols, and makes it difficult to synchronize and expand different sequencers in a central on-chip network.
It employs a combination of command and address sequencers and data sequencers to form complex sequences, supports partitioning between protocol-independent blocks and protocol-specific blocks, achieves scalability between different memory protocols through the on-chip network of the control center, and utilizes a timestamp synchronization system for self-enumeration and triggering.
It enables support for different memory protocols, improves the efficiency of memory interface training, calibration and debugging, supports concurrent command sequencing for protocols such as high-bandwidth memory and double data rate, and enhances the programmability and scalability of the control center.
Smart Images

Figure CN114518902B_ABST
Abstract
Description
Technical Field
[0001] This invention relates generally to the field of memory management technology, and specifically to a memory sequencer system and a memory sequencing method using the system. Background Technology
[0002] Modern computer architecture combines three main storage technologies that dominate supercomputing: Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), and magnetic storage media. Magnetic storage media include hard disk drives and magnetic tape. Main memory is the primary component used to store data within the computer and is almost entirely composed of DRAM technology.
[0003] In modern computer systems, one or more DRAM memory controllers (DMCs) may be contained within a processor package or integrated into a system controller residing outside the processor package. Regardless of the location of the DRAM memory controller, its function is to accept read and write requests for a given address in memory, translate the request into one or more commands and place them in the storage system, transmit these commands to the DRAM device in the appropriate order and at the appropriate time, and retrieve or store data in the name of the processor or I / O device in the system.
[0004] US2005257109A1 discloses a Built-in Self-Test (BIST) architecture with distributed algorithm interpretation. This architecture comprises three levels: a centralized BIST controller, a set of sequencers, and a set of memory interfaces. The BIST controller stores a set of commands that typically define algorithms for testing storage modules, regardless of the physical characteristics or timing requirements of the storage modules. The sequencers interpret the commands according to a command protocol and generate a sequence of storage operations. The memory interfaces apply the storage operations to the storage modules based on their physical characteristics, such as by translating address and data signals based on the row-column arrangement of the storage modules to implement the bit patterns described by the commands. This command protocol allows for the description of powerful algorithms in an extremely concise manner, applicable to storage modules with diverse characteristics.
[0005] US2008219112A1 discloses an apparatus for generating digital signal patterns, comprising a memory, a program sequencer, first and second circuits, and an event execution unit. The memory may store multiple instructions that, when executed, cause the generation of digital signal patterns on multiple nodes. The program sequencer may be used to control the order in which multiple instructions are retrieved from the memory and executed. The first circuit may sequentially pass through multiple different output states in response to a clock signal. The second circuit may identify an output event when the output state of the first circuit corresponds to an output state identified by a retrieved specific type of instruction. The event execution unit may, in response to the output event identified by the second circuit, control the signal states on the multiple nodes in a manner specified by the retrieved specific type of instruction.
[0006] The aforementioned references aim to support pattern-based address and data sequencers. However, they have many limitations and drawbacks. For example, the memory sequencer architectures in the aforementioned references only support pattern-based address and data sequencers, which limit the pattern complexity that can be driven due to shallow pattern queues. Existing memory sequencers are designed to support only one or two memory protocols and are not reprogrammable to support different memory protocols. Furthermore, the aforementioned references typically support non-concurrent single addressing streams for row or column addressing. In addition, the biggest challenge in expanding memory sequencer systems is solving the ability to synchronize different sequencers connected to a central control network-on-a-chip (CCNOC) in order to precisely orchestrate them to form signals for the widest memory interface protocol.
[0007] Therefore, a memory sequencer system that overcomes the above problems and shortcomings is still needed. Summary of the Invention
[0008] The following brief summary of the invention provides a basic understanding of certain aspects of the invention. This summary is not a broad overview of the invention, and its sole purpose is to present some concepts of the invention in a simplified form as a prelude to a more detailed description that follows.
[0009] One object of the present invention is to provide a memory sequencer system that uses command and address sequencers and data sequencers as means to arrange signal ordering.
[0010] Another object of the present invention is to provide a memory sequencer system that allows sequencers to be linked together with triggers to form complex sequences.
[0011] Another object of the present invention is to provide a memory sequencer system that supports partitioning between protocol-independent blocks and protocol-specific blocks to support scalability across different memory protocols, such as concurrent row and column command sequencing for high bandwidth memory (HBMx), or single address stream command sequencing for double data rate (DDRx) and low power double data rate (LPDDRx).
[0012] Another object of the present invention is to provide a memory sequencer system that supports a highly programmable control center and a highly scalable control center on-chip network (CCNOC) that can self-enumerate and trigger between different sequencers through its timestamp synchronization system.
[0013] Another object of the present invention is to provide a memory ordering method for external memory protocols.
[0014] Therefore, these objectives can be achieved through the teachings of this invention. This invention relates to a memory sequencer system for an external memory protocol, comprising: a control center including a microcontroller; an on-chip network of the control center including nodes with point-to-point connections for synchronizing and coordinating communication; characterized in that:
[0015] Command and address sequencers are used to generate command, control, and address commands for specific memory protocols; and
[0016] At least one data sequencer is used to generate pseudo-random or deterministic data patterns for each byte channel of the memory interface;
[0017] The command and address sequencer and the data sequencer are linked together to form complex address and data sequences for training, calibration, and debugging of the memory interface; the control center on-chip network interconnects the control center with the command and address sequencer and the data sequencer to provide firmware controllability.
[0018] This invention also relates to a memory sequencing method for an external memory protocol, characterized by comprising the following steps: generating a command, address, and data sequence for each entry; selecting one or more address sequencers to generate addresses; comparing trigger values to trigger the increment of an adder or comparator at the address sequencer; linking the adder or comparator to create an address sequence; encoding and decoding the command and address to trigger a data path; implementing data latency using a shift register based on the number of clock cycles; transmitting data including write data and read data to a data sequencer; converting AXI-lite read or write commands into control center on-chip network read or write commands; transmitting the control center on-chip network read or write commands to the target slave node of the control center on-chip network based on the address and identifier of the slave node; enumerating the slave nodes to assign an identifier to each slave node; and synchronizing the control center on-chip network timestamp and alarm register.
[0019] The foregoing and other objects, features, aspects and advantages of the present invention will become more readily understood in conjunction with the detailed description provided below and with appropriate reference to the accompanying drawings. Attached Figure Description
[0020] To provide a detailed understanding of the above-described features of the invention, a more specific description of the invention, which has been briefly summarized above, can be derived through embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only show typical embodiments of the invention and should not be considered as limiting the scope of the invention, as other equivalent embodiments are permissible.
[0021] These and other features, benefits, and advantages of the invention will become apparent from the following accompanying drawings, wherein the same reference numerals refer to the same structures throughout the views, wherein:
[0022] Figure 1 A block diagram of a memory sequencer system within a memory physical layer subsystem according to an embodiment of the present invention is shown;
[0023] Figure 2 A block diagram of the control center according to an embodiment of the present invention is shown;
[0024] Figure 3 A block diagram of a command and address sequencer and a data sequencer according to an embodiment of the present invention is shown;
[0025] Figure 4 A block diagram of an address sequencer according to an embodiment of the present invention is shown;
[0026] Figure 5 A tree topology of a command center on-chip network (CCNOC) according to an embodiment of the present invention is shown;
[0027] Figure 6 The diagram illustrates the connections between nodes in a command center on-chip network (CCNOC) according to an embodiment of the present invention;
[0028] Figure 7 A memory sequencing method for an external memory protocol according to an embodiment of the present invention is illustrated;
[0029] Figure 8 The packet types of a command center on-chip network (CCNOC) according to an embodiment of the present invention are shown;
[0030] Figure 9 An example of packet transmission in a command center on-chip network (CCNOC) according to an embodiment of the present invention is shown;
[0031] Figure 10 A block diagram of a programmable linear feedback shift register (LFSR) or a multiple-input signature register (MISR) according to an embodiment of the present invention is shown;
[0032] Figure 11 An enumeration of a command center on-chip network (CCNOC) according to an embodiment of the present invention is shown;
[0033] Figure 12 illustrates alternative methods for timestamp synchronization among all slave nodes, including (a) a connection between a CCNOC network with one master node and multiple slave nodes; (b) when the CCNOC master node is programmed to send write packets to slave nodes 1, 2, and 4; (c) when slave nodes 1, 2, and 4 receive write packets and forward them to slave nodes 3 and 5; (d) when slave nodes 3 and 5 receive write packets and forward them to slave node 6; and (e) when slave node 6 receives a write packet. Detailed Implementation
[0034] Detailed embodiments of the invention are disclosed herein as needed. However, it should be understood that the disclosed embodiments are merely examples of the invention, which can be implemented in various forms. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but only as the basis for the claims. It should be understood that the drawings and their detailed description are not intended to limit the invention to the specific forms disclosed; rather, the invention will cover all modifications, equivalents, and alternatives falling within the scope of the invention as defined in the claims. Throughout this application, the term “may” signifies permissible (i.e., possible) rather than mandatory (i.e., required). Similarly, the terms “include, including, include” signify including but not limited to. Furthermore, unless otherwise stated, the term “a, an” means “at least one,” and the term “plurality” means one or more. In the case of the use of abbreviations or technical terms, they represent the generally accepted meanings in the art.
[0035] In the following description, the invention is illustrated with reference to the accompanying drawings, in which reference numerals used correspond to similar elements throughout the specification. However, the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, the embodiments provided make this disclosure sufficient and complete, and will fully convey the scope of the invention to those skilled in the art. In the following detailed description, numerical values and ranges are provided for various aspects of the described embodiments. These numerical values and ranges should be considered as examples only and are not intended to limit the scope of the claims. Furthermore, many materials are identified as suitable for various aspects of implementation. These materials will be considered exemplary and are not intended to limit the scope of the invention.
[0036] Now refer to Figure 1 to Figure 1 The invention is described in more detail with reference to the accompanying drawings shown in Figure 2.
[0037] Reference Figure 1The present invention relates to a memory sequencer system for an external memory protocol, comprising: a control center (CC) (6) including a microcontroller (62); a control center on-chip network (CCNOC) (5) including nodes (51) connected in a point-to-point manner for synchronizing and coordinating communication; characterized in that: a command and address sequencer (4) for generating commands, control and address commands for a specific memory protocol; and at least one data sequencer (3) for generating pseudo-random or deterministic data patterns for each byte channel of the memory interface; wherein the command and address sequencer (4) and the data sequencer (3) are linked to form composite address and data sequences for training, calibration and debugging of the memory interface; the control center on-chip network (5) interconnects the control center (6) with the command and address sequencer (4) and the data sequencer (3) to provide firmware controllability.
[0038] According to an embodiment of the present invention, the memory sequencer system (100) is inserted between the DDRPHY interface (DFI) (2) of the memory controller (1) and the physical layer transmission or reception path (7) for memory interface commands and data.
[0039] According to an embodiment of the invention, firmware running on the microcontroller of the control center (6) is loaded into the static random access memory (SRAM) of the control center (6). This firmware is responsible for configuring and setting the command and address sequencer (4) and the data sequencer (3) according to the expected interface memory protocol via the control center on-chip network (CCNOC) (5). The main function of the control center (6) is to execute the external memory interface initialization sequence to initialize the memory host and the attached memory device. The control center (6) also performs calibration and training to enable the physical layer to reliably transmit and receive data from the external memory via its command and data buses. Specific sequences representing the memory devices of the memory controller are also executed by the control center (6). The core of the control center is a microcontroller (62) that includes a two- or three-stage pipelined single-issue instruction processor. In addition to the microcontroller (62), the other two hosts on the AXI-lite interconnect are the Joint Test Action Group (JTAG) controller (61) and the Direct Memory Access (DMA) (63).
[0040] Reference Figure 2The JTAG controller provides JTAG access to an external host (e.g., a personal computer or test instrument) to access the AXI-lite interconnect. The JTAG controller (61) can issue AXI reads or writes to any memory-mapped location. The JTAG controller (61) can be used to probe or overwrite register values for debugging. The JTAG controller (61) can also be used to download firmware to the instruction SRAM (64), which the microcontroller (62) can execute. DMA (63) is designed to free the microcontroller (62) from moving large blocks of data. Its primary purpose is to move data into or out of all channels and data paths of the host memory physical layer, specifically to move data into or out of the data SRAM (65). The data SRAM (65) is used for general storage, while the instruction SRAM (65) is used to store the firmware of the microcontroller (62). Other peripheral devices, such as programmable input or output (I / O), interrupt controller (67), timer (69), and monitor (66), provide auxiliary services to the microcontroller (62) to make the independent subsystem complete.
[0041] Reference Figure 3 The command and address sequencer (4) includes a command sequencer (41) and an address sequencer (43). The command and address sequencer (4) arranges signaling sequences with the help of a distributed data sequencer.
[0042] According to an embodiment of the invention, the command and address sequencer (4) includes a command sequence table (42) for interpreting each entry by iterating through the table to coordinate the generation of command, address, or data sequences. The command sequence table (42) may be of depth, comprising multiple entries, each entry having multiple fields. For example, the command sequence table (42) may be eight entries deep, each entry having the fields shown in Table 1 of the example, and Table 2 of the example shows an example of the use of the command sequence table (42).
[0043] Reference Figure 4The entries in the command sequencer (41) will select one or more address sequencers responsible for address generation. The number of address sequencers (43) follows the number of entries available for row or column address generation in the command sequence table. For example, if the command sequence table includes eight entries, the address sequencer should include eight address sequencers. Each address sequencer (43) includes a set of adders or comparators to trigger increments of different adder or comparator groups within the same group based on a trigger value. For example, each address sequencer (43) includes, but is not limited to, three sets of adders or comparators. This allows the address sequencer (43) to increment by a fixed address offset in each cycle it is indexed, or only when it is triggered by another set of adders or comparators. Linking adders or comparators in different orders will construct different addressing sequences.
[0044] According to an embodiment of the present invention, the command and address sequencer (4) includes a command encoder and a command decoder.
[0045] According to an embodiment of the present invention, the data sequencer (3) further includes a read data memory having a read data buffer.
[0046] According to an embodiment of the present invention, the control center on-chip network (CCNOC) (5) is connected to AXI-lite via the CCNOC master device (68) to receive and convert read or write commands and transmit the read or write commands to the network within the memory sequencer system (100).
[0047] Reference Figure 5 The Control Center Network-on-Chip (CCNOC) (5) comprises a master node (51) and multiple slave nodes (52) organized in a tree topology. Each node includes a set of registers as memory mapped to the AXI memory space. Configuration registers are used to control the functions of the memory sequencer and other modules in the storage physical layer (7) and memory controller (1). For example, the register can be used to set the delays for transmit first-in-first-out (TX FIFO) and receive first-in-first-out (RX FIFO). The Command and Address Sequencer (CAS) and the Data Sequencer (DS) are also mapped to the configuration register of one of the slave nodes. Each register is 32 bits wide. The maximum number of registers supported by each slave node is 1024 x 32-bit registers or 4096 bytes.
[0048] According to an embodiment of the present invention, the master node (51) has multiple downstream ports, and each slave node (52) has one upstream port and multiple downstream ports. Each upstream or downstream port has a pin list, as shown in Table 3 in the example. Furthermore, the direct connections between the ports of the CCNOC nodes are as follows: Figure 6As shown. However, the actual implementation of the connection may insert pipe stages between nodes for timing convergence.
[0049] Reference Figure 7 The present invention also relates to a method (200) for memory sequencing for an external storage protocol, characterized by comprising the following steps: generating a command, address, and data sequence for each entry; selecting one or more address sequencers (43) to generate addresses; comparing trigger values to trigger increments of adders or comparators at the address sequencer (43); linking adders or comparators to create address sequences; encoding and decoding commands and addresses to trigger data paths; implementing data delays using shift registers based on the number of clock cycles; transmitting data including write data and read data to a data sequencer; converting AXI-lite read or write commands into control center on-chip network read or write commands; transmitting the control center on-chip network read or write commands to the target slave node (52) of the control center on-chip network (5) based on the address and identifier of the slave node (52); enumerating the slave node (52) to assign identifiers to each slave node (52); and synchronizing the control center on-chip network timestamp and alarm register.
[0050] According to an embodiment of the present invention, a sequence of commands, addresses and data is generated based on the programming values of the command sequence table (42).
[0051] According to an embodiment of the present invention, the programmed value determines the delay amount of the command, address, and data sequence.
[0052] According to embodiments of the present invention, an encoded write command triggers the write transmission data path, and an encoded read command triggers the read reception data path to capture and upload data from the memory device. Depending on the external memory protocol, the command encoder may employ any one of phases P0 or P1; two phases P0 and P1; or three phases P0, P1, and P0 of the next cycle to encode row or column commands.
[0053] According to an embodiment of the present invention, decoding a read command triggers a write to the transmit data path, and decoding a write command triggers a read to the receive data path to capture and unload data from the memory host. If the memory sequencer is used as a memory device instead of a memory host, the command decoder is used to decode the incoming command. The command decoder is used to receive commands from the RX FIFO of the receive data path of the command and address pin rows or columns.
[0054] According to an embodiment of the invention, both the command encoder and decoder use the CMD2DATA interface (45) to control the associated transmit or receive data paths. The CMD2DATA interface (45) essentially implements read or write latency via shift registers. When the command encoder or decoder triggers a read, the read trigger is delayed by N memory cycles before being passed to the corresponding data sequencer (3). When the command encoder or decoder triggers a write, the write trigger is delayed by M memory cycles before being passed to the corresponding data sequencer (3). Since the read and write latency is typically several clock cycles, this is sufficient to pipeline the CMD2DATA bus multiple times as needed to reach the furthest data sequencer (3) from the command and address sequencer (4). Each data sequencer (3) closer to the pipeline is configured to have less internal latency to ensure that all data sequencers (3) can send or receive data to or from the data bus pins in the same number of memory cycles.
[0055] According to an embodiment of the present invention, the step of transmitting write data includes iterating through a write data sequence table to generate a write data transmission pattern. The write data sequence table may be of depth including multiple entries, each entry having multiple fields. Table 4 shows an example of a write data sequence table of depth including eight entries, each entry having fields.
[0056] According to an embodiment of the present invention, the step of transmitting read data includes iterating through a read data sequence list to generate a read data transmission pattern. The read data sequence list is a depth comprising multiple entries, each entry having multiple fields. Table 5 shows an example of a read data sequence list comprising 8 entries, each entry having fields.
[0057] According to an embodiment of the invention, a target slave node (52) is assigned to the packet destination ID field based on the address, and the packet is then sent downstream. When the packet arrives at the slave node (52), the slave node (52) checks the identifier of the packet destination. If the identifier belongs to the slave node (52), it performs a read or write operation on the addressing configuration register attached to it. If the identifier does not belong to the slave node (52), the slave node (52) sends the packet to all its downstream ports. Ultimately, the packet will be acquired by a more downstream slave node (52). If the identifier is for a broadcast write or selective broadcast write, and the slave node (52) is part of a selective broadcast set, the slave node (52) will acquire the packet. Regardless of whether it has acquired the packet for itself, the slave node (52) will additionally broadcast the broadcast write or selective broadcast write to all its downstream ports.
[0058] According to an embodiment of the invention, commands are transmitted in a group comprising a write group (53), a read group (54), a completion group (55), and a message group (56).
[0059] Reference Figure 8 The master node (51) always publishes write packets (53) under CCNOC (5). More specifically, the target slave node (52) will not respond with a write response or a completion packet (55). After the master node (51) sends a write packet (53) to CCNOC (5), it will immediately send an AXI write response. If the destination ID on the packet is all 1s, the write packet (53) can be broadcast to all slave nodes (52). Each slave node (52) will also maintain an ID mask to allow selective broadcasting. Read packets (54) under CCNOC are always split transactions; in other words, the target slave node (52) does not need to respond to the read completion packet (55) immediately. If the master node (51) sends multiple read packets (54) to different slave nodes (52), the read completion packets (55) will return to the master node (51) out of order. The master node (51) is responsible for reordering the read completions and sending the AXI read responses in order. The destination IDs of read packets (54) will not be equal because read packets (54) do not support broadcasting. Read complete packets (55) can always reach the master node (51) because there is only one upstream port. If the address sent by the AXI master node for reading or writing is within the address range of the CCNOC slave node but is not mapped to any slave node (52), the CCNOC master node (51) will still respond to the read with 0xDEAD_BEEF or some other fixed data mode. For writes, the CCNOC master device will not perform any operation and will return an OKAY write response. The status register in the CCNOC master node (51) will be set to indicate that the AXI access has been mapped to a slave ID that does not exist.
[0060] According to an embodiment of the present invention, the steps for synchronizing the timestamps of the control center on-chip network include: sending a read packet (54) from the control center on-chip network master node (51) to a slave node (52); reading the timestamp register of the slave node (52); recording the sending time of the read packet (54); when the slave node (52) receives the read packet (54), sending a read response to the read packet (54); recording the time when the control center on-chip network master node (51) receives the read response; sending a write packet (53) with the current timestamp and pipeline delay data to the timestamp register of the slave node (52), and repeating the previous steps from sending the read packet (54) to the slave node (52) until all slave nodes have the same timestamp as the control center on-chip network master node (51). The pipeline delay is half of the round-trip delay, which is the difference between the receiving time of the read response and the sending time of the read packet.
[0061] Referring to Figure 12, the steps for synchronizing the timestamps of the control center on-chip network further include programming the timestamp offset registers of the control center on-chip network master node (51) and slave nodes (52) with the number of clock cycles. A write packet (53) is broadcast from the control center on-chip network master node (51) to all slave nodes (52) with packet data loads; the timestamp register value of the slave node (52) is set with the received data; the written timestamp value is sent to the downstream slave node (52), and all steps starting from setting the timestamp register value are repeated until all slave nodes have the same timestamp as the control center on-chip network master node (51). The write packet (53) travels from the control center on-chip network master node (51) to the control center on-chip network slave node (52) according to the number of clock cycles and the data load containing the timestamp value of the control center on-chip network master node (51).
[0062] According to an embodiment of the invention, a timestamp is used for alarm triggering to perform synchronization actions among all slave nodes (52). The slave nodes (52) in timestamp synchronization and alarm triggering implement the registers listed in Table 6, wherein the timestamp register is a free-running rollover counter that increments on each rising edge of the clock.
[0063] Examples of the invention will be provided below for more detailed explanation. It should be understood that the examples described below are not intended to limit the scope of the invention.
[0064] Example
[0065] Generating command, address, and data sequences
[0066] Table 1:
[0067] Name Symbol #bytes Description Delay count delay 8 Number of DFI delay cycles to apply before / after a command cycle Pre / post delay post_delay 1 If 1, the delay will be applied after the command cycle Row command phase 0 row_sel_p0 3 Select address sequencer corresponding to phase P0 Row command phase 1 row_sel_p1 3 Select address sequencer corresponding to phase P1 Column command phase 0 col_sel_p0 3 Select address sequencer corresponding to phase P0 Column command phase 1 col_sel_p1 3 Select address sequencer corresponding to phase P1 Loop counter select loop_select 2 Select 1 of 4 loop counters Branch equal to zero pointer bez 3 Next table entry index if selected loop counter is 0 Branch not equal to zero pointer bnz 3 Next table entry index if selected loop counter is not 0 End of sequence end 1 Indicate last entry of sequence
[0068] Table 1 shows the command sequence table fields in the command sequence table entries. The command sequencer uses values programmed in the command sequence table (42) to interpret the nature of the operation to be performed. The "Delay Count" or "Pre / Post Delay" setting can jointly determine the amount of delay to be inserted before or after the operation. If "Pre / Post Delay" is set to 1, the "Delay Count" is applied after the operation is completed; otherwise, it is applied before the operation is completed. The delay count calculates the number of DFI cycles to wait before or after the operation. No-Operation (NOP) cycles are inserted within the delay count period.
[0069] The Phase 0 and Phase 1 fields for row and column commands allow different commands to be initiated during different phases of the clock. For example, if a command is to be inserted during the high phase of the clock (Phase 0), row or column command Phase 0 will be used. If a command is to be inserted during the low phase of the clock (Phase 1), row or column command Phase 1 will be used. Not all memory interface protocols allow different commands to be used during different phases of the clock. For memory interface protocols that only allow the insertion of clock-level granular commands, the row and column commands for Phase 1 can be programmed as 7 (3'b111) to ignore or remove them. For memory interface protocols that only support non-concurrent row and column access, such as Double Data Rate (DDRx) and Low Power Double Data Rate (LPDDRx), row or column command phases will be used.
[0070] After a command is completed, the loop counter selection, the pointer field for branches equal to zero, and the pointer field for branches not equal to zero work together to determine the next entry in the command sequence list to be executed. The loop counter selection is used to select one of the four loop counters to be used. Loop counter 0 is always 0, while the other three loop counters can be preset to different values by the control center (6). After the command is completed, the loop counter pointed to by the loop counter selection is decremented by 1. If the decremented loop counter is zero, the next entry in the command sequence list to be executed will be the entry pointed to by the pointer for branches equal to zero. If the decremented loop counter is not zero, the next entry in the command sequence list to be executed will be the entry pointed to by the pointer for branches not equal to zero. Therefore, even limited to eight entries in the table, the firmware can configure the command sequence list to implement loops and complex sequences.
[0071] Table 2:
[0072]
[0073] Table 2 shows an example of the use of the command sequence list (42). In cycle 0, the pointer (current_ptr) of the currently executing entry is at entry 0, and the line command output for stage 0 is selected from address sequencer 1. In cycles 1 through 4, all command outputs are 7 (3'b111), and a NOP loop is inserted due to the late delay programmed for 5. In cycle 5, the loop counter is decremented by 1 because it is not zero, and the execution of entry 0 is repeated 4 more times. In cycle 30, the loop counter 1 expires, so current_ptr is now at entry 1, and the command sequencer (41) will wait for two more cycles before executing entry 1 because the pre / post delay is 0. In cycle 32, the column command output for stage 0 is selected from address sequencer 2. Since the loop counter 2 is 1 and the non-zero pointer (BNZ) is 0, current_ptr will jump back to entry 0. Since address sequencer 1 is still in progress, the address will continue to increment, so the sequence from cycle 0 to cycle 32 will be reissued. At period 65, loop counter 2 is now 0 and ends at 1, so the sequence will end. Therefore, the entire command flow is: <Start> → <line cmd 0@p0> → <5-cycle delay> → <line cmd 0@p0> → <5-cycle delay> → <line cmd 0@p0> → <5-cycle delay> → <line cmd0@p0> → <5-cycle delay> → <line cmd 0@p0> → <5-cycle delay> → <2-cycle delay> → <column cmd1@p0> → <line cmd 0@p0> → <5-cycle delay> → <line cmd0@p0> → <5-cycle delay> → <line cmd 0@p0> → <5-cycle delay> → <line cmd 0@p0> → <5-cycle delay> → <2-cycle delay> → <column cmd1@p0> → <End>.
[0074] Build address sequence
[0075] Reference Figure 4The address sequencer (43) can be used to cycle between the library and SID at row address 0 and row address 1. The SID is called the stack ID address of the High Bandwidth Memory (HBM) memory device. Assume adder 0 is mapped to the library address, adder 1 is mapped to the row address, and adder 2 is mapped to the SID of the HBM channel. The address register of adder 0 is loaded with the starting value of the library address, library 0; the address register of adder 1 is loaded with the starting value of the row address, address 0; and the address register of adder 2 is loaded with the starting value of the SID, SID 0. Adder 0 is configured to be triggered when adder 1 (row address) is triggered, adder 1 is configured to be triggered in each cycle triggered by the command sequencer (41), and adder 2 is configured to be triggered when adder 0 (library address) is triggered. When triggered, the increment value of all three adders is set to 1 to increment by one count. Next, the end address of adder 0 is configured to the value of library 1, the end address of adder 1 is configured to the value of row address 1, and the end address of adder 2 is configured to the value of SID 3. When the command sequencer (41) triggers the address sequencer (43) for the first time, it will output library address 0, row address 0, and SID 0. The next value of the row address is 1. When the command sequencer (41) triggers the address sequencer (43) for the second time, it will output library address 0, row address 1, and SID 0. Since row address 1 matches the end address of the row, it will reset its value back to the starting value, i.e., row address 0, and enable the trigger to trigger adder 0 (library address) to increment to library address 1. The next time the command sequencer (41) triggers the address sequencer (43), it will output library address 1, row address 0, and SID 0. When the command sequencer (41) triggers the address sequencer (43) for the third time, it will output library address 1, row address 1, and SID 0. Since library address 1 matches the end address of the library and line address 1 matches the end address of the line, the library address will reset its value back to its starting value, library address 0, and trigger the SID. The next time the command sequencer (41) triggers the address sequencer (43), it will output library address 0, line address 0, and SID1. This process is then repeated.
[0076] CCNOC port pin
[0077] Table 3:
[0078]
[0079]
[0080] Table 3 lists the port pinouts for the upstream and downstream ports of CCNOC, such as... Figure 6 As shown.
[0081] Transfer data to data sequencer
[0082] Table 4:
[0083]
[0084] Table 4 shows the write data sequence table fields for each write data sequence table entry. Looping within the write data sequence table is similar to looping within the command sequence table. The only exception is that the source of the data pattern to be transferred depends on either the output of the Linear Feedback Shift Register (LFSR) or the contents of the write data buffer pointed to by the data_sel field. If lfsr_sel is 1, the LFSR output is used as the write data. Otherwise, the contents of the write data buffer pointed to by data_sel are used. Furthermore, the write sequence table has no pre / post delay selection, so a delay is always applied after the data transfer. The write data buffer is a local first-in-first-out (FIFO) memory, eight entries deep, allowing each pin to drive any custom pattern. data_sel points to one of the entries containing the data to be transferred.
[0085] Reference Figure 9 The LFSR structure used by the write data sequencer is fully programmable, allowing it to be transmitted with any LFSR polynomial. The LFSR feedback path is fully configurable by cfg_lfsr_poly[N-1:0]. When cfg_lfsr_poly[N] is 1, the feedback path from dataout[0] is XORed to FF[N]. When cfg_lfsr_init occurs, the LFSR seed is initialized in the LFSR shift register. This LFSR structure can also be used with a Multiple-Input Signature Register (MISR), but cfg_misr_en is always invalidated in the write sequencer and is therefore only used in the read sequencer. The write data sequencer is also used to generate data patterns when the memory sequencer is used as a memory device.
[0086] Table 5:
[0087]
[0088] Table 5 shows the read data sequence table fields for read data sequence table entries. The read data storage area contains a read data buffer, which is used to temporarily buffer received data. To operate on the received data, the read data sequence table shown in Table 5 is required. The read data sequence table functions exactly the same as the write data sequence table and can be shared physically and logically if the read and write data sequencers do not need to run simultaneously. Data captured in the read data buffer can be stored in the read data buffer for firmware to read for future operation, debugging, and inspection. This data can be fed into a programmable MISR chain to generate a unique signature for comparison. The data can also be compared with a programmable LFSR chain to determine whether it is a match or a precise bit mismatch. The same programmable LFSR-MISR chain is used to generate a unique MISR signature or to generate LFSR data to compare with the data in the read data buffer for reading.
[0089] CCNOC packet transfer
[0090] Reference Figure 10 CCNOC(5) uses a credit scheme to regulate packet transmission from one port to another. Credits are sent from the receiver to the transmitter on the rx_credit bus to indicate the number of free slots it must accept CCNOC packets from. This value is displayed and remains constant. One credit unit is equivalent to 1 byte transmission on tx_data, which is 8 bits wide from the transmitter to the receiver. When a reset is active, the transmitter initializes its downstream port with zero credits and therefore cannot transmit any data. After the reset is inactive, the receiver can begin to advertise the number of FIFO free slots it has in the form of credits to the transmitter. The receiver can send up to three credits per cycle and may take up to twenty cycles to advertise its total credits to the transmitter. For example, if the FIFO implemented in the receiver has eight entries equal to 8 bytes, the receiver displays 2'b11 in the first two cycles and then 2'b10 in the third cycle. It should then drive 2'b00 until more credits are available to be released. Upon receiving a credit, the transmitter increments the internal counter of its corresponding downstream port by the amount of credits released by the receiver. Each time the transmitter transmits a payload byte, it needs to increment its internal counter by 1. As long as the credit counter on that port of the transmitter is not zero, the transmitter is allowed to send a transaction to the receiver.
[0091] In another credit scheme, in addition to cyclically sending released credits, the receiver can optionally use an N-bit credit bus to announce the total number of credits. For example, if the receiver has eight entries equal to 8 bytes, it can display a pseudo-static value of 8 on a 4-bit bus. Whenever the receiver releases more credits, it only needs to change the value of the credit bus once, for example, from 8 to 12. The advantage of doing this is that the transmitter does not need to increment its internal counter, as it can directly obtain the total credit value from this credit bus. Another advantage is that the credit bus is not time-critical because its value is pseudo-static.
[0092] CCNOC enumeration
[0093] When the CCNOC master enumerates all connected nodes, it checks if a port is connected to a slave. If a downstream port is not connected to a slave, its tx_stub association is set to 1. If the tx_stub association is 0, the master sends a write packet of N=1 to the slave port with ID=0, where N = the write payload of the write packet received from the upstream node. On the slave node, the port that receives the first packet is assigned as the upstream port, and all other ports of the slave node are assigned as downstream ports, while the slave node assigns itself ID=N. If one or more downstream ports have tx_stub associations of 0, the slave node selects the first downstream port and forwards the write packet with ID=0 and data=N+1. If there are no downstream ports, causing all ports to have tx_stub associations set to 1, the slave node generates a message packet with an enumeration response, whose data is its own slave node ID, and sends the packet upstream. This notifies the upstream node that the enumeration has reached the leaf slave node. The slave node cycles through all its downstream ports according to the prescribed steps. If the slave node receives a "write packet", it will look for downstream slave nodes and forward the "write packet" with ID=0 and data=N+1. If the slave node receives a "message packet" with an "enumeration response": if there is another downstream port with tx_stub association 0, then send a "write packet" to that downstream port with ID=0 and data=(data of the "message packet with enumeration response")+1; otherwise, if there are no other downstream ports with tx_stub association 0, or all downstream ports with tx_stub association 1, then generate a "message packet" with an "enumeration response", where data=(data of the "message packet with enumeration response"). If the CCNOC master node receives a "message packet" with an "enumeration response" but there are no other downstream ports that have not yet sent data with the "write packet" with ID=0, then the enumeration process is considered complete. The data in the "Message Group" containing "Enumerated Response" equals the total number of enumerated slave devices.
[0094] Reference Figure 11 The CCNOC master device first sends a "write packet" with ID=0 and data=1 to the leftmost downstream port. The leftmost slave device receives this message, sets its own ID to 1, and then forwards the "write packet" with ID=0 and data=2 to its leftmost slave device. The next slave device sets its own ID to 2. Since there are no more downstream ports (all tx_stubs are associated with 1), slave device 2 returns a "message packet" with data=2 and an "enumeration response" upstream. When slave device 1 receives the "message packet" from slave device 2, it sends a "message packet" with data=3 to the next leftmost slave device. This slave device sets its own ID to 3. Since there are no more downstream ports, slave device 3 returns a "message packet" with data=3 and an "enumeration response" upstream. When slave device 1 receives the "message packet" from slave device 3, it sends a "write packet" with data=4 to the rightmost slave device. The slave device sets its ID to 4. Since there are no more downstream ports, slave device 4 returns a message packet with an "enumeration response" to the upstream with data=4. When slave device 1 receives the message packet from slave device 4, it sends a message packet with an "enumeration response" to the upstream with data=4. When the CCNOC master device receives the message packet with an "enumeration response" from slave device 4, it sends a "write packet" to the rightmost slave device with data=5. This slave device sets its ID to 5 and sends a message packet with data=6 to its single downstream port. The last slave device sets its ID to 6. Because there are no more downstream ports, slave device 6 returns a message packet with an "enumeration response" to the upstream with data=6. When slave device 5 receives the message packet from slave device 6, it sends a message packet with an "enumeration response" with data=6. The CCNOC master device receives this message packet from slave device 5. Since the CCNOC master device has enumerated all its downstream ports, the enumeration process is complete, and it is known that there are a total of 6 slave devices in the NOC topology.
[0095] CCNOC timestamp synchronization
[0096] Using CCNOC timestamp synchronization, different sequencers connected to different CCNOC nodes can learn and understand the global time. Once the global time is known, alarms can be set on different sequencers to perform time synchronization operations. Time synchronization is completed when the CCNOC bus is idle (i.e., there are no pending reads or writes) and CCNOC enumeration is complete. Each slave device also executes the registers listed in Table 6, where the timestamp register is a free-running rolling counter that increments on each rising edge of the clock.
[0097] Table 6:
[0098]
[0099]
[0100] For example, the CCNOC master device has a timestamp value of 1000, and there are 5 pipelines between the CCNOC master device and slave device 1. The CCNOC master device begins by sending a "read packet" to slave device 1. The CCNOC master device should receive a "read response" after 10 cycles. At timestamp 1020, the CCNOC master device sends a "write packet" to the timestamp register of slave device 1, whose timestamp register data is incremented by 5 relative to the current timestamp, i.e., 1025. After 5 cycles, slave device 1 receives this packet and sets its timestamp register to 1025. Therefore, the timestamp registers of both the CCNOC master device and slave device 1 are at 1025, and can be said to be perfectly synchronized.
[0101] Timestamps are used in conjunction with alarm registers to perform synchronization operations on all slave nodes. To do this, if a write operation is intended to be used in conjunction with the expected operation of the target slave node, the CCNOC master device first writes to the alarm operation register or alarm data register. Then, an alarm timestamp is written with a future timestamp, which will be used to synchronize events such as initiating a sequencer on different or the same channel. When the timestamp registers of each slave device reach the count in the "alarm timestamp," an "alarm operation" will be triggered simultaneously on the target slave node.
[0102] Referring to Figure 12(a), a CCNOC network comprises a master device and multiple slave devices connected in a hierarchical manner. This hierarchical arrangement is an alternative method for timestamp synchronization among all slave devices relative to the CCNOC master device. Slave devices have different initial timestamps because they exit the reset state at different times. However, the physical connections between slave devices are static, and the number of clock cycles required to transfer write packets between slave devices is known. Each slave device includes a timestamp offset register, representing an additional offset to be added to the data payload of a write packet when it is transmitted from that slave device. The timestamp offset register is programmed based on the number of clock cycles required to transfer a write packet from one slave device to another. If identical slave devices are implemented, the same values can be used to program the timestamp offset registers for all slave devices. For simplicity, the timestamp offset register can also be implemented as a constant value.
[0103] Referring to Figure 12(b), the timestamp offset register value for each slave device is programmed to, but is not limited to, 9. The CCNOC master device also includes a timestamp offset register, which can be programmed to a value equal to the number of clock cycles required to transfer a write packet from the CCNOC master device to a CCNOC slave device directly connected to it. Timestamp synchronization can be accomplished via a single broadcast write packet, which the CCNOC master device broadcasts to all slave devices, where the packet's data payload includes the master device's timestamp value. In this example, the CCNOC master device's timestamp value is programmed to, but is not limited to, 1000, and since the CCNOC master device's timestamp offset value is programmed to, but is not limited to, 9, the write packet's data payload is modified by adding 9 to 1000, thus creating a modified write packet with a data payload of 1009 when sent to slave devices 1, 2, and 4 in cycle 0 (C0).
[0104] Referring to Figure 12(c), the write packet is received nine clock cycles after period 9 (C9), with a write value of 1009 in its timestamp register. Therefore, slave device 1's timestamp can be considered synchronized with the CCNOC master, as both nodes contain the same timestamp value of 1009. Meanwhile, slave devices 2 and 4 both contain updated timestamp registers with a timestamp value of 1009 at the ninth cycle (C9). Both slave devices 2 and 4 add their values to their respective timestamp offset registers before forwarding the write packet to slave devices 3 and 5; this write packet will contain a data payload value of 1018 due to the increment of 9.
[0105] Referring to Figure 12(d), slave devices 3 and 5 receive write packets at cycle 18 (C18), and both slave devices set their timestamp registers with a write packet data payload value of 1018. Slave devices 3 and 5 are synchronized with the CCNOC master because both slave devices have the same timestamp as the master. Next, slave device 5 adds this value to its timestamp offset and forwards the write packet with a data payload of 1027 to slave device 6.
[0106] Referring to Figure 12(e), slave device 6 receives the write packet from slave device 5 and sets its timestamp register to 1027. Finally, since all slave devices contain the same timestamp as the CCNOC master device, timestamp synchronization is completed.
[0107] In this invention, command and address sequencers and data sequencers are used as means to coordinate signal ordering. Furthermore, this invention forms composite memory sequences triggered by the linking of sequencer clusters. In addition, this memory sequencer system supports concurrent row and column command ordering or single-address stream command ordering. Moreover, this memory sequencer system can synchronize different sequencers attached to a CCNOC with self-enumeration capabilities, thereby deterministically arranging them to form signals to a wide memory interface protocol.
[0108] The exemplary embodiments described above are illustrated with specific features, but the scope of the invention may also include various other features.
[0109] Various modifications to these embodiments will be apparent to those skilled in the art from the specification and accompanying drawings. The principles associated with the various embodiments described herein can be applied to other embodiments. Therefore, the description of the invention is not limited to those described in the accompanying drawings. Figure 1 The embodiments shown are provided to represent the broadest scope consistent with the principles, novelty, and inventiveness disclosed or taught herein. Therefore, all alternatives, modifications, and variations made in accordance with this invention fall within the scope of this invention and the appended claims.
[0110] It should be understood that any prior art publications referred to herein are not an admission that such publications constitute part of common general knowledge in the field.
Claims
1. A method (200) for memory ordering for an external memory protocol, characterized in that, Includes the following steps: Generate a command, address, and data sequence for each entry; Choose one or more address sequencers to generate addresses; Compare the trigger value to trigger the increment of the adder or comparator at the address sequencer; Link the adders or comparators to create an address sequence; The commands and addresses are encoded and decoded to trigger the data path; Data delay is achieved using shift registers based on the number of clock cycles; The data, including both written and read data, is transmitted to the data sequencer. Convert AXI-lite read or write commands into on-chip network read or write commands for the control center; Based on the address and identifier of the slave node (52), the read or write command of the control center on-chip network is transmitted to the target slave node (52) of the control center on-chip network (5); Enumerate the subordinate nodes (52) to assign an identifier to each subordinate node (52); as well as Synchronize the on-chip network timestamp and alarm register in the control center; The steps for synchronizing the on-chip network timestamps in the control center include: A read packet (54) is sent from the master node (51) of the on-chip network in the control center to a slave node (52); Read the timestamp register (52) of the slave node; Record the transmission time of the read packet (54); When the slave node (52) receives the read packet (54), it sends a read response to the read packet (54); Record the time when the on-chip network master node (51) of the control center receives and reads the response; Send the write packet (53) containing the current timestamp and pipeline delay data to the timestamp register of the slave node (52); Repeat the above steps, starting from sending the read packet (54) to the slave node (52), until all slave nodes have the same timestamp as the control center on-chip network master node (51); The pipeline delay mentioned above is half of the round-trip delay; The round-trip delay is the difference between the time it takes to receive the read response and the time it takes to transmit the read packet.
2. The method according to claim 1, characterized in that, Generate command, address, and data sequences based on the programming values of the command sequence table (42).
3. The method according to claim 2, characterized in that, The programmed value determines the amount of delay for the command, address, and data sequence.
4. The method according to claim 1, characterized in that, The encoding write command triggers the writing of the data transmission path, and the encoding read command triggers the reading of the data reception path to capture and upload data from the memory device.
5. The method according to claim 1, characterized in that, The decode read command triggers the write to the transmit data path, and the decode write command triggers the read to the receive data path to capture and unload data from the memory host.
6. The method according to claim 1, characterized in that, The step of transmitting and writing data further includes iterating through the write data sequence list to generate a write data transmission pattern.
7. The method according to claim 1, characterized in that, The step of transmitting and reading data further includes iterating through the read data sequence list to generate a read data transmission pattern.
8. The method according to claim 1, characterized in that, The step of transmitting data, including written data and read data, to the data sequencer further includes: capturing read data in the read data buffer for debugging and checking operations, signature generation, or bit comparison.
9. The method according to claim 1, characterized in that, The command is transmitted in a group that includes a write group (53), a read group (54), a completion group (55), and a message group (56).
10. The method according to claim 1, characterized in that, The step of synchronizing the on-chip network timestamps in the control center also includes: The timestamp offset registers of the master node (51) and slave node (52) of the on-chip network are programmed using the number of clock cycles; The write packet (53) is broadcast from the master node (51) of the on-chip network in the control center to all slave nodes (52) with packet data load; Use the received data to set the timestamp register value of the slave node (52); Send the written timestamp value to the downstream slave node (52); Repeat the above steps, starting from setting the timestamp register value, until all slave nodes have the same timestamp as the control center on-chip network master node (51); The write packet (53) is transmitted from the master node (51) of the control center on-chip network to the slave node (52) of the control center on-chip network according to the number of clock cycles. The data payload includes the timestamp value of the on-chip network master node (51) of the control center.
11. The method according to claim 9 or 10, characterized in that, The timestamp is used to trigger an alarm to perform synchronization actions among all slave nodes (52).
12. A memory sequencer system (100) for an external memory protocol, used for the memory sequencing method according to any one of claims 1-11, the system comprising: The control center (6) includes a microcontroller (62); The control center on-chip network (5) includes nodes (51) that are connected in a point-to-point manner for synchronizing and coordinating communication; Its features are: Command and address sequencer (4), used to generate command, control, and address commands for a specific memory protocol; and At least one data sequencer (3) is used to generate pseudo-random or deterministic data patterns for each byte channel of the memory interface; The command and address sequencer (4) and the data sequencer (3) are linked to form a composite address and data sequence for training, calibration and debugging of the memory interface; The control center on-chip network (5) interconnects the control center (6) with the command and address sequencer (4) and the data sequencer (3) to provide firmware controllability; The memory sequencer system (100) is inserted between the DDR PHY interface (2) of the memory controller (1) and the physical layer transmission or reception path (7) for the memory interface commands and data.
13. The memory sequencer system (100) according to claim 12, characterized in that, The command and address sequencer (4) includes a command sequencer (41) and an address sequencer (43).
14. The memory sequencer system (100) according to claim 13, characterized in that, The command and address sequencer (4) includes a command sequence table (42) for interpreting each entry by iterating through the table to coordinate the generation of command, address, or data sequences.
15. The memory sequencer system (100) according to claim 13, characterized in that, The address sequencer (43) includes a set of adders or comparators for triggering increments of different adder or comparator groups within the same group based on a trigger value.
16. The memory sequencer system (100) according to claim 13, characterized in that, The command and address sequencer (4) includes a command encoder and a command decoder.
17. The memory sequencer system (100) according to claim 12, characterized in that, The data sequencer (3) also includes a read data memory with a read data buffer.
18. The memory sequencer system (100) according to claim 12, characterized in that, The control center on-chip network (5) is connected to AXI-lite and is used to receive, convert and transmit read or write commands in the network within the memory sequencer system.
19. The memory sequencer system (100) according to claim 12, characterized in that, The control center on-chip network (5) includes a master node (51) and multiple slave nodes (52) organized in a tree topology, each node including a set of registers as memory mapped to the AXI memory space.
20. The memory sequencer system (100) according to claim 19, characterized in that, The master node (51) has multiple downstream ports, and each slave node (52) has one upstream port and multiple downstream ports.
Citation Information
Patent Citations
Built-in self-test (BIST) architecture having distributed interpretation and generalized command protocol
US20050257109A1
Software programmable timing architecture
US20080219112A1
Integrated controller for training memory physical layer interface
CN106133710A