A configurable topology architecture based on DRAM
The configurable DRAM topology with adaptive routing addresses flexibility and bandwidth issues, enhancing system adaptability and efficiency by optimizing data transmission and reducing latency.
Patent Information
- Application Number
- CN202410723723.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-06-05
AI Technical Summary
The immobility and unconfigurability of the existing memory topology limit the scalability and adaptability of the system. Insufficient bandwidth optimization in multiple DDR environments, it is impossible to effectively deal with congestion and bandwidth bottlenecks, resulting in limited memory performance.
Using a configurable topology based on DRAM, the configuration signals are generated through FPGA units, and the transmission algorithms and paths between nodes are dynamically configured. Combined with X, Y routing and Circle routing algorithms, data transmission paths are optimized to improve flexibility and bandwidth utilization.
It improves the adaptability of memory in different application scenarios, solves the congestion and conflict problems in data transmission, improves data transmission efficiency, reduces memory access latency, and enhances the bandwidth utilization of multi-DDR systems.
Smart Images

Figure CN118689837B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of memories, and particularly relates to a configurable topology architecture based on DRAM. Background Art
[0002] Currently, most solutions in the technical field of memories rely on traditional bus architectures and fixed topologies, such as ring, star, or fully interconnected topologies. These solutions are usually designed for specific application scenarios or performance requirements at the beginning, so their adaptability and scalability are limited when facing variable workloads or system expansion requirements. For example, in the ring topology, when the number of nodes increases, the data transmission delay will increase significantly; while the star topology is limited by the bandwidth and processing capacity of the central node.
[0003] In addition, the bandwidth optimization in the existing technology in a multi-DDR (Double Data Rate) environment is usually limited to static scheduling strategies and fixed crossbar switch networks. These methods cannot effectively adapt to the bandwidth requirements under dynamically changing load conditions, resulting in insufficient memory bandwidth utilization. In terms of DRAM (Dynamic Random Access Memory) memories, traditional architectures usually do not fully consider the impact of instruction sequences and access patterns on performance, resulting in poor instruction execution efficiency and access latency.
[0004] The defects of the existing technology are mainly reflected in the following aspects:
[0005] The fixity and non-configurability of the memory topology structure limit the scalability and adaptability of the system, and cannot adapt to the changes of various application scenarios and workloads.
[0006] In the context of multi-DDR, the existing routing algorithms and traffic control strategies cannot effectively cope with congestion and bandwidth bottlenecks, resulting in limited memory performance.
[0007] The existing DRAM bandwidth utilization optimization strategies are insufficient and cannot fully exploit the bandwidth potential of multi-channel DRAM, especially in the case of high concurrency and large data volumes.
[0008] The existing SOC systems under multi-DDR fail to effectively optimize instruction sequences and bus encoding and decoding, resulting in low memory access efficiency.
[0009] In summary, the existing technology has obvious deficiencies in the flexibility of the memory topology structure, the bandwidth optimization in the multi-DDR environment, and the DRAM bandwidth utilization. These defects limit the further improvement of memory performance. Summary of the Invention
[0010] The object of the present invention is to provide a configurable topology architecture based on DRAM to solve the problems of obvious deficiencies in the flexibility of the memory topology structure, bandwidth optimization in a multi-DDR environment, and DRAM bandwidth utilization.
[0011] To achieve the above object, the present invention adopts the following technical solutions: The present invention provides a configurable topology architecture based on DRAM, and the architecture includes: a multi-node data transmission network, and the multi-node data transmission network includes: several nodes, each node is assigned at least three adjacent nodes, and each node is respectively communicatively connected to at least three adjacent nodes, and the at least three adjacent nodes are any at least three nodes among the remaining nodes other than the node to which adjacent nodes need to be assigned among all nodes;
[0012] Each node is configured with a memory unit and an FPGA unit, the memory unit and the FPGA unit on each node are both communicatively connected to the node, and the FPGA unit on each node is used to generate a configuration signal, and the configuration signal is used to assign a transmission algorithm to the node.
[0013] Preferably, the number of nodes in the multi-node data transmission network is an even number.
[0014] Preferably, there are 16 nodes in total, and all nodes are distributed in a 4*4 array.
[0015] Preferably, each node includes:
[0016] A data queue unit, which is used to cache a request for target data currently input to the node, and takes the node as an initial node, and the target data is data to be written into a preset memory unit;
[0017] An algorithm multiplexing unit, which is used to assign a transmission algorithm to the target data of the node based on the configuration signal generated by the FPGA unit corresponding to the node, and the transmission algorithm is used to transmit the target data from the initial node to a target node, and the target node is a node connected to a preset memory unit;
[0018] A transmission path gating unit, which is used to select a transmission path for the target data based on the transmission algorithm, and the transmission path is the optimal transmission route from the initial node to the target node;
[0019] A local recording unit, which is used to record the transmission path of the target data.
[0020] Preferably, an assignment instruction for assigning a transmission algorithm to the target data of the node is generated by the FPGA unit of the node.
[0021] Preferably, the transmission algorithm includes: an X,Y routing algorithm and / or a Circle routing algorithm.
[0022] Preferably, when the transmission algorithm is the Circle routing algorithm, the FPGA unit is further configured to configure at least two nodes in the multi-node data transmission network into a ROUND topology; when the transmission algorithm is the X,Y routing algorithm, the FPGA unit is further configured to configure at least two nodes in the multi-node data transmission network into a MESH topology.
[0023] Preferably, each memory unit is a DDR memory.
[0024] Preferably, each node further includes:
[0025] An address encoder, configured to select a write target address of a preset memory unit for the target data of the target node;
[0026] A DDR controller, configured to write the target data of the target node into a preset memory unit according to the write target address.
[0027] Advantageous effects:
[0028] The multi-node data transmission network of the present invention can be flexibly configured through the FPGA unit, improving the adaptability of the memory in different application scenarios, effectively solving the congestion and conflict problems in data transmission, enhancing the data transmission efficiency, and reducing the memory access latency. Description of the Drawings
[0029] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0030] Figure 1 is a schematic diagram of the overall structure of a configurable topology architecture based on DRAM provided by an embodiment of the present invention. Detailed Embodiments
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structure of the drawings is only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. It should be noted here that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation to the present invention.
[0032] Such as Figure 1As shown in the figure, this embodiment provides a configurable topology architecture based on DRAM. The architecture includes: a multi-node data transmission network, and the multi-node data transmission network includes: several nodes, each node is assigned at least three adjacent nodes, and each node is respectively communicatively connected to at least three adjacent nodes. The at least three adjacent nodes are any at least three nodes among the remaining nodes other than the node to which adjacent nodes need to be assigned among all nodes;
[0033] Each node is configured with a memory unit and an FPGA unit. The memory unit and the FPGA unit on each node are both communicatively connected to this node. The FPGA unit on each node is used to generate a configuration signal to allocate a transmission algorithm for this node with the configuration signal.
[0034] In this example, the number of nodes in the multi-node data transmission network is an even number. Therefore, the number of nodes in the multi-node data transmission network can be 4, 6, 8, 10, 12, 14, or 16. When the number of nodes in the multi-node data transmission network is 4, the 4 nodes are distributed in a 2*2 array; when the number of nodes in the multi-node data transmission network is 6, the 6 nodes are distributed in a 3*2 array; when the number of nodes in the multi-node data transmission network is 8, the 8 nodes are distributed in a 4*2 array; when the number of nodes in the multi-node data transmission network is 10, the 10 nodes are distributed in a 5*2 array; when the number of nodes in the multi-node data transmission network is 12, the 12 nodes are distributed in a 6*2 array or a 4*3 array; when the number of nodes in the multi-node data transmission network is 14, the 14 nodes are distributed in a 7*2 array; when the number of nodes in the multi-node data transmission network is 16, the 16 nodes are distributed in an 8*2 array or a 4*4 array.
[0035] Taking 16 nodes as an example, as Figure 1 shown, the 16 nodes are respectively: R0, R1, R2, R3, R4, R5, R6, R7, R8, R9, R10, R11, R12, R13, R14, and R15. The connection relationship of the 16 nodes is:
[0036] Node R0 is respectively communicatively connected to node R1, node R4, and node R12. Node R1 is also communicatively connected to node R2 and node R5. Node R4 is also communicatively connected to node R5 and node R8.
[0037] Construct a coordinate system for the multi-node data transmission network. Taking the first node R0 as the origin, the coordinates of the first node R0 are (0, 0), the coordinates of the second node R1 are (1, 0), the coordinates of the third node R2 are (2, 0), the coordinates of the fourth node R3 are (3, 0), the coordinates of the fifth node R4 are (0, 1), the coordinates of the sixth node R5 are (1, 1), the coordinates of the seventh node R6 are (2, 1), the coordinates of the eighth node R7 are (3, 1), the coordinates of the ninth node R8 are (0, 2), the coordinates of the tenth node R9 are (1, 2), the coordinates of the eleventh node R10 are (2, 2), the coordinates of the twelfth node R11 are (3, 2), the coordinates of the thirteenth node R12 are (0, 3), the coordinates of the fourteenth node R13 are (1, 3), the coordinates of the fifteenth node R14 are (2, 3), and the coordinates of the sixteenth node R15 are (3, 3).
[0038] Each node is connected to a memory unit, namely DIMM (Dual In-line Memory Module), and the storage capacity of each memory unit is 64GB * 4. Each node is also connected to an FPGA unit.
[0039] In the prior art, multi-DDR systems often adopt ROUND topology (circular topology) and MESH topology (mesh topology). Among them, the ROUND topology is a circular structure formed by connecting all nodes end to end in sequence, and the MESH topology is a rectangular array structure with m * n nodes, and the nodes at each position are connected to the nodes at all adjacent coordinate positions of the node; however, the data transmission paths of the ROUND topology and the MESH topology are fixed and rather inflexible, lacking flexibility.
[0040] In this embodiment, the multi-node data transmission network of the present invention can, according to the configuration signal, configure some nodes to form a ROUND topology, or form a MESH topology, or simultaneously form a ROUND topology and a MESH topology, greatly improving the flexibility of the multi-node data transmission network, enhancing the adaptability of the memory in different application scenarios, effectively solving the congestion and conflict problems in data transmission, improving the data transmission efficiency, and reducing the memory access latency.
[0041] Taking 16 nodes as an example, nodes R0, R4, R8, and R12 can form a ROUND - type topology structure, and nodes R3, R7, R11, and R15, as well as nodes R4, R5, R6, R7, R11, R10, R9, and R8 can all form ROUND - type topology structures. Moreover, the number of nodes forming the ROUND - type topology structure can also be different and can be flexibly configured according to actual needs, featuring high flexibility and practicality.
[0042] For another example, nodes R0, R1, R2, R3, R4, R5, R6, and R7 can form a 2 * 4 MESH - type topology structure. Therefore, the multi - node data transmission network of the present invention has high flexibility. The more nodes there are, the higher the flexibility. And this structure takes into account both performance and scalability, adapting to changing application requirements.
[0043] As a further optimization of this embodiment, each node includes: a data queue unit, an algorithm multiplexing unit, a transmission path gating unit, and a local recording unit.
[0044] Among them, the data queue unit is used to cache the request for the target data currently input to this node. Taking this node as the initial node, the target data is the data that needs to be written into a preset memory unit. The target data in this embodiment can be data input by an external device, can be data read from other memory units, and the data in this memory unit needs to be transferred and stored to a preset memory unit.
[0045] The algorithm multiplexing unit is used to allocate a transmission algorithm for the target data of this node based on the configuration signal generated by the FPGA unit corresponding to this node. The transmission algorithm is used to transfer the target data from the initial node to the target node, and the target node is the node connected to the preset memory unit.
[0046] The transmission path gating unit is used to select a transmission path for the target data based on the transmission algorithm, and the transmission path is the optimal transmission route from the initial node to the target node;
[0047] The local recording unit is used to record the transmission path of the target data.
[0048] In this embodiment, after the data queue unit of a certain node caches a request for target data, the memory unit to which the target data needs to be written is used as the preset memory unit. Then, the optimal transmission route in the multi-node data transmission network is determined based on the initial node and the target node. Next, the FPGA unit assigns transmission functions to the nodes on the optimal transmission route. The transmission function can select the direction of the transmission path of each node so that the target data is transmitted along the optimal transmission route. After the direction of the transmission path of the target data is selected at each node, the local recording unit records the transmission path of the target data.
[0049] In this embodiment, the allocation instruction for allocating the transmission algorithm for the target data of the node is generated by the FPGA unit of the node. The transmission algorithm includes: X,Y routing algorithm and / or Circle routing algorithm. When the transmission algorithm is the Circle routing algorithm, the FPGA unit is further configured to configure at least two nodes in the multi-node data transmission network into a ROUND topology structure; when the transmission algorithm is the X,Y routing algorithm, the FPGA unit is further configured to configure at least two nodes in the multi-node data transmission network into a MESH topology structure.
[0050] Take Figure 1 a multi-node data transmission network with 16 nodes as an example. For example: when node R0 inputs target data and the target data needs to be transmitted to node R12, at this time, the FPGA units of node R0 and node R12 generate allocation instructions to configure node R0 and node R12 into a ROUND topology structure. At this time, node R0 directly sends the target data to node R12, and node R12 receives the target data from node R0.
[0051] Another example is that when node R0 inputs target data and the target data needs to be transmitted to node R6, at this time, the FPGA units of node R0, node R1, node R2, and node R6 generate allocation instructions to configure node R0, node R1, node R2, and node R6 into a MESH topology structure. At this time, node R0 transmits the target data to node R1, node R1 transmits the target data to node R2, and node R2 then transmits the target data to node R6 to complete the transmission of the target data.
[0052] The data transmission strategy of the present invention effectively solves the problems of congestion and conflict in data transmission and improves the data transmission efficiency.
[0053] As a further optimization of this embodiment, each memory unit is a DDR memory, and each node further includes: an address encoder and a DDR controller;
[0054] The address encoder is used to select a preset write target address for the target data of the target node, that is, to write it into the DDR memory;
[0055] The DDR controller is used to write the target data of the target node into a preset memory cell according to the write target address.
[0056] In this embodiment, the present invention can significantly improve the bandwidth utilization rate in a multi-DDR system. By optimizing the architecture of the multi-node data transmission network, the memory access latency is reduced, and the system performance is improved.
[0057] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0058] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0059] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A configurable topology architecture based on DRAM, characterized in that, The architecture includes: a multi-node data transmission network, which includes: several nodes, each node is assigned three adjacent nodes, and each node is communicatively connected to the three adjacent nodes, and the three adjacent nodes are any three nodes among the remaining nodes except the node that needs to be assigned adjacent nodes; there are 16 nodes in total, and all nodes are distributed in a 4*4 array; Each node is configured with a memory unit and an FPGA unit. The memory unit and the FPGA unit on each node are both communicatively connected to the node. The FPGA unit on each node is used to generate a configuration signal to allocate a transmission algorithm for the node with the configuration signal; The transmission algorithm includes: X,Y routing algorithm and / or Circle routing algorithm. When the transmission algorithm is the Circle routing algorithm, the FPGA unit is further used to configure at least two nodes in the multi-node data transmission network into a ROUND topology; when the transmission algorithm is the X,Y routing algorithm, the FPGA unit is further used to configure at least two nodes in the multi-node data transmission network into a MESH topology; The multi-node data transmission network configures some nodes to form a ROUND topology, or a MESH topology, or both a ROUND topology and a MESH topology according to the configuration signal.
2. The configurable topology architecture based on DRAM according to claim 1, wherein The number of nodes in the multi-node data transmission network is an even number.
3. The configurable topology architecture based on DRAM according to claim 1, characterized in that, Each node includes: A data queue unit, which is used to cache the request of the target data currently input to the node. Taking the node as the initial node, the target data is the data that needs to be written into a preset memory unit; An algorithm multiplexing unit, which is used to allocate a transmission algorithm for the target data of the node based on the configuration signal generated by the FPGA unit corresponding to the node. The transmission algorithm is used to transmit the target data from the initial node to the target node, and the target node is the node connected to the preset memory unit; A transmission path gating unit, which is used to select a transmission path for the target data based on the transmission algorithm, and the transmission path is the optimal transmission route from the initial node to the target node; A local recording unit, which is used to record the transmission path of the target data.
4. The configurable topology based on DRAM according to claim 3, characterized in that, The allocation instruction for allocating the transmission algorithm for the target data of the node is generated by the FPGA unit of the node.
5. The configurable topology architecture based on DRAM according to claim 3, wherein, Each memory unit is a DDR memory.
6. The configurable topology architecture based on DRAM according to claim 5, characterized in that, Each node further includes: An address encoder, which is used to select a write target address of a preset memory unit for the target data of the target node; A DDR controller, which is used to write the target data of the target node into a preset memory unit according to the write target address.
Citation Information
Patent Citations
Mixed interconnection Mesh topological structure for on-chip network and routing algorithm thereof
CN103986664A
Data transmission circuit, chip, system and electronic equipment
CN117194310A