Controller for accelerated computing, and accelerated computing system
By grouping accelerated computing units and optimizing data transmission paths using internal buses and multi-level switching modules, the problem of high latency in on-chip interconnect technology is solved, improving the communication efficiency and overall performance of the accelerated computing system.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2025-10-21
- Publication Date
- 2026-07-30
AI Technical Summary
In existing technologies, on-chip interconnect technology is constrained by network topology, resulting in significant delays and reduced computing efficiency during some transmission processes for accelerated computing tasks.
The accelerated computing units are grouped based on a group bus-based on-chip interconnect structure. Each group shares memory and is connected within the group via an internal bus. Different groups are connected diagonally via on-chip interconnect technology, which reduces the number of jumps in the switching process and optimizes the data transmission path through multi-level switching modules and routing modules.
It significantly reduces data transmission latency, improves communication efficiency between accelerated computing units, and enhances the overall performance of the accelerated computing system.
Smart Images

Figure CN2025129057_30072026_PF_FP_ABST
Abstract
Description
A controller and an accelerated computing system for accelerating computing
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202510101194.2, filed on January 22, 2025, entitled “A controller and a computing system for accelerating computing”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of computers, and more specifically, to a controller and a computing acceleration system for accelerating computing. Background Technology
[0004] Heterogeneous acceleration architectures typically combine controllers with other processors to leverage the strengths of different processors for high-performance computing. In heterogeneous acceleration architectures, the interconnection and communication between the acceleration computing units within the controller has a critical impact on the overall system performance.
[0005] In related technologies, bus-based interconnect and on-chip interconnect technologies are generally used to interconnect accelerated computing units. However, bus-based interconnect suffers from insufficient bandwidth allocation and increased latency due to bus contention when multiple computing units simultaneously request local memory. Therefore, this drawback is unavoidable in large-scale systems. On the other hand, on-chip interconnect technology relies on network topology and link length limitations. In some transmission processes, there are often many hops, resulting in significant latency. Furthermore, on-chip interconnect technology requires a large number of routers and links to build the network, leading to increased chip area and power consumption. In addition, when transmitting across devices, the link width has a significant impact on the transmission rate, affecting the overall performance of the accelerated computing system.
[0006] In summary, the accelerated computing systems in related technologies suffer from a problem where the on-chip interconnect technology is constrained by the network topology, resulting in significant latency in some transmission processes and a decrease in the computational efficiency of accelerated computing tasks. Summary of the Invention
[0007] This application provides a controller and an accelerated computing system for accelerating computing, in order to at least solve the problem in the related art where on-chip interconnect technology is constrained by network topology, resulting in large delays in some transmission processes, which leads to a decrease in the computing efficiency of accelerated computing tasks.
[0008] According to one aspect of the embodiments of this application, a controller for accelerated computing is provided, including: a plurality of accelerated computing unit groups, adjacent accelerated computing unit groups being connected in pairs for communication and arranged in a completely symmetrical ring, accelerated computing units within the same accelerated computing unit group being connected sequentially through an internal bus, and accelerated computing units in different accelerated computing unit groups being connected in communication with each other at symmetrical positions on the same diameter of the ring.
[0009] Each accelerated computing unit group also includes shared memory, which is connected to each accelerated computing unit via an internal bus.
[0010] In one exemplary embodiment, the controller further includes: a device-side controller configured to receive data packets in CXL format, parse the data packets to obtain accelerated computing tasks, and send them to the accelerated computing unit; and a switching module connected to the device-side controller via a CXL bus and communicating with the accelerated computing unit via an internal bus, configured to receive accelerated computing tasks via direct memory access and send them to the accelerated computing unit.
[0011] In one exemplary embodiment, the switching module includes a multi-level switching module, including a switching module that is connected to the device-side controller via a CXL bus and connected to other switching modules via multiple output signals in master mode, and a switching module that is connected to other switching modules, connected to shared memory via an output signal in slave mode, and connected to an accelerated computing unit via an output signal in master mode.
[0012] In one exemplary embodiment, the controller further includes: an on-chip routing module, which is communicatively connected to each accelerated computing unit group and configured to implement routing jumps within the same controller; and an inter-chip routing module, which is communicatively connected to the on-chip routing module and configured to implement routing jumps between different controllers.
[0013] In one exemplary embodiment, the on-chip routing module includes a first cross switch and a first cross switch control unit, wherein the first cross switch is configured to connect each accelerated computing unit group and the inter-chip routing module, and the first cross switch control unit is configured to control the on / off state of the path connected to the first cross switch.
[0014] In one exemplary embodiment, the inter-chip routing module includes a second cross switch and a second cross switch control unit, wherein the second cross switch is configured to connect the intra-chip routing module and a multiplexer, and the second cross switch control unit is configured to control the on / off state of the path connected to the second cross switch.
[0015] In one exemplary embodiment, the intra-chip routing module or inter-chip routing module includes a virtual channel control unit, configured to divide a physical channel into multiple virtual channels to prevent deadlock caused by channel closure.
[0016] In an exemplary embodiment, the virtual channel control unit is configured to construct multiple independent micro-slice buffers in the output queue buffer corresponding to the physical channel to obtain multiple virtual channels; the virtual channels are divided into a first channel group and a second channel group in a clockwise direction; the first channel group is used when the routing coordinates of the source node are greater than the routing coordinates of the target node; and the second channel group is used when the routing coordinates of the source node are less than or equal to the routing coordinates of the target node.
[0017] In one exemplary embodiment, when the accelerated computing unit group is in a first operating mode, the accelerated computing units within the group can be configured to perform different computing functions.
[0018] In one exemplary embodiment, when the accelerated computing unit group is in a second operating mode, the accelerated computing units within the group are configured to perform the same computing function so that if any accelerated computing unit within the group fails, it can be replaced by another accelerated computing unit within the group.
[0019] In one exemplary embodiment of this application, the controller is one of a Field Programmable Gate Array (FPGA), a Reduced Instruction Set Computer (ARM), and an Enhanced Reduced Instruction Set Controller (AVR).
[0020] According to another aspect of the embodiments of this application, an accelerated computing system is provided, comprising: a host; multiple controller clusters configured to execute accelerated computing tasks issued by the host, each controller cluster including multiple controllers; and a switching chip connected to the host and the controller clusters via a bus, configured to realize data interaction between the host and the controller clusters.
[0021] In one exemplary embodiment, the controllers in the controller cluster are configured as a fully connected architecture.
[0022] In one exemplary embodiment, the switching chip supports at least the CXL communication protocol.
[0023] In one exemplary embodiment, the controller cluster further includes a multiplexer configured to receive signals generated within the controller and transmit them to other controllers or a host.
[0024] According to another aspect of the embodiments of this application, a routing method within a controller is provided, comprising: configuring the routing coordinates of an accelerated computing unit within the controller, the routing coordinates including a controller ID, an accelerated computing unit group ID, and an accelerated computing unit ID; determining the routing coordinates of a source node and a target node based on a routing request, and calculating an intra-chip routing step count based on the routing coordinates of the source node and the target node; determining the routing routing direction based on the intra-chip routing step count, and performing a routing redirection based on the routing routing direction.
[0025] In one exemplary embodiment, routing based on the routing direction includes: obtaining the output buffer length of the controller port; and performing reverse routing if the output buffer length is greater than a fifth threshold.
[0026] In an exemplary embodiment, determining the routing direction based on the number of intra-chip jump steps includes: determining the routing direction based on a preset range in which the number of intra-chip jump steps is located, wherein the routing direction includes clockwise direction, counterclockwise direction, opposite direction, or a combination of clockwise direction, counterclockwise direction, and opposite direction.
[0027] In one exemplary embodiment, the controller includes multiple accelerated computing unit groups, wherein the accelerated computing units within the same accelerated computing unit group are sequentially connected via an internal bus and connected to the same shared memory via the internal bus. Accelerated computing units in adjacent accelerated computing unit groups and accelerated computing units in opposite accelerated computing unit groups are also connected via communication connections. The routing jump direction is determined based on a preset range of intra-chip jump steps. The method includes: when the intra-chip jump step number equals a first threshold, determining a clockwise direction as the routing jump direction; when the intra-chip jump step number equals a second threshold, determining the routing jump direction... The direction of routing is determined as follows: If the number of intra-chip hops equals the third threshold, the counter-clockwise direction is determined as the routing direction; if the number of intra-chip hops is less than the first threshold, the counter-clockwise direction is determined as the routing direction; if the number of intra-chip hops is greater than the first threshold and less than the second threshold, the routing direction is determined first as the face direction and then as the clockwise direction; if the number of intra-chip hops is greater than the second threshold and less than the third threshold, the routing direction is determined first as the face direction and then as the counter-clockwise direction; if the number of intra-chip hops is greater than the third threshold and less than the fourth threshold, the clockwise direction is determined as the routing direction.
[0028] According to another aspect of the embodiments of this application, a routing jump method between controllers is provided. The controller includes an intra-chip routing module and an inter-chip routing module. The intra-chip routing module is communicatively connected to each accelerated computing unit group, and the inter-chip routing module is communicatively connected to the intra-chip routing module. The method includes: configuring the routing coordinates of the accelerated computing units inside the controller, the routing coordinates including controller ID, accelerated computing unit group ID, and accelerated computing unit ID; determining the routing coordinates of the source node and the routing coordinates of the target node based on a routing jump request; jumping from the intra-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the source node based on the routing coordinates of the source node, jumping from the inter-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the target node based on the routing coordinates of the target node, and jumping from the inter-chip routing module corresponding to the target node to the intra-chip routing module corresponding to the target node based on the routing coordinates of the target node, and transmitting the data to the target node.
[0029] According to another aspect of the embodiments of this application, a method for scheduling accelerated computing units for accelerating computing tasks is provided. The accelerated computing unit scheduling method is applied to an accelerated computing system. The method includes: obtaining an accelerated computing unit status table and a memory access latency table. The computing unit status table includes at least the routing coordinates, unit function, unit status, and storage occupancy of the accelerated computing unit. The memory access latency table includes the latency of each accelerated computing unit accessing each shared cache. The method also includes obtaining computing tasks to be assigned and determining a scheduling strategy based on the computing unit status table and the memory access latency table. During the determination of the scheduling strategy, accelerated computing units with different storage occupancy can run in parallel. The scheduling strategy corresponds to the computing task to be assigned with the lowest latency for each accelerated computing unit.
[0030] This application addresses the issue that the controller employs an improved on-chip interconnect structure based on a grouped bus. This structure groups the internal acceleration computing units, with each group sharing memory. Routing jumps within a group can be directly accomplished through this shared memory. Acceleration computing units in different groups are diagonally connected via on-chip interconnect technology, and the groups are sequentially connected, eliminating the need for sequential jumps and reducing jump steps. Furthermore, the shared memory reduces global bus contention. Therefore, this application solves the problem in related technologies where on-chip interconnect technology is constrained by network topology, resulting in significant latency in some transmission processes and decreased computational efficiency of acceleration tasks. This improves communication efficiency within and between controllers, reduces communication latency, and enhances the overall performance of the acceleration computing system. Attached Figure Description
[0031] Figure 1 is a schematic diagram of the topology of a controller for accelerating computing according to an embodiment of this application;
[0032] Figure 2 is a schematic diagram of the connection structure of a controller for accelerating computing according to an embodiment of this application;
[0033] Figure 3 is a schematic diagram of an on-chip routing module according to an embodiment of this application;
[0034] Figure 4 is a schematic diagram of an inter-chip routing module according to an embodiment of this application;
[0035] Figure 5 is a logic diagram of a cross switch according to an embodiment of this application;
[0036] Figure 6 is a schematic diagram of a controller-based accelerated computing system according to an embodiment of this application;
[0037] Figure 7 is a schematic diagram of a fully connected structure of a controller cluster according to an embodiment of this application;
[0038] Figure 8 is a flowchart illustrating a routing method within a controller according to an embodiment of this application;
[0039] Figure 9 is a flowchart illustrating a routing method between controllers according to an embodiment of this application;
[0040] Figure 10 is a flowchart illustrating a method for scheduling accelerated computing units for an accelerated computing task according to an embodiment of this application. Detailed Implementation
[0041] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0043] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:
[0044] Field Programmable Gate Array (FPGA) is a programmable semiconductor device that allows users to configure its hardware as needed. An FPGA contains programmable logic blocks and configurable interconnects, and users can implement different hardware logic by loading different configuration files.
[0045] Direct Memory Access (DMA) is a technology that allows hardware subsystems to exchange data directly with system memory without the direct control of the CPU. A DMA controller can directly transfer data from input devices to memory or from memory to output devices without CPU intervention, thereby reducing the CPU's workload and improving data transfer efficiency.
[0046] Peripheral Component Interconnect Express (PCIe) is a general-purpose serial expansion bus standard used for connecting internal computer hardware components. It is widely used for high-speed data communication and peripheral device connections due to its high data transfer rate and flexibility.
[0047] Compute Express Link (CXL) is an open, high-speed interconnect technology standard designed to provide higher data throughput and lower communication latency for modern data centers, high-performance computing systems, and other compute-intensive applications.
[0048] With the continuous development of technologies such as artificial intelligence, big data analytics, and cloud computing, the amount of system data is growing exponentially, leading to a surge in demand for data processing and analysis. Among these technologies, traditional single-architecture database systems face significant challenges, proving insufficient in their ability to handle massive amounts of data, ensure data real-time performance, and cope with diverse workloads.
[0049] Heterogeneous acceleration architectures typically combine controllers with other processors, leveraging the strengths of different processors to achieve high-performance computing. In heterogeneous acceleration architectures, the interconnection and communication between the accelerated computing units within the controller has a critical impact on the overall system performance. Improving the interconnection structure within the controller can shorten data transmission paths and increase bandwidth, effectively reducing system latency. This allows data to be scheduled and transmitted efficiently and systematically among the accelerated computing units, fully utilizing the strengths of each unit and improving overall system performance.
[0050] Some solutions address the issue that controllers employing on-chip interconnect technology are constrained by network topology, resulting in significant latency during some transmission processes and consequently reducing the computational efficiency of accelerated computing tasks. To resolve this technical problem, this embodiment provides a controller for accelerated computing, an accelerated computing system, a routing method within the controller, a routing method between controllers, and a method for scheduling accelerated computing units for accelerated computing tasks.
[0051] In this embodiment of the application, a controller for accelerated computing is provided, the topology of which is shown in Figure 1. The controller includes: multiple accelerated computing unit groups, adjacent accelerated computing unit groups are connected in pairs for communication, arranged in a completely symmetrical ring, accelerated computing units within the same accelerated computing unit group are connected sequentially through an internal bus, and accelerated computing units in different accelerated computing unit groups located at symmetrical positions on the same diameter of the ring are connected for communication; each accelerated computing unit group also includes shared memory, and the shared memory is connected to each accelerated computing unit through an internal bus.
[0052] In accelerated computing systems related to this technology, the accelerated computing units in the controller used for accelerated computing and the memory typically employ a single bus or a complex on-chip interconnect network. These interconnect methods suffer from increased latency and bandwidth bottlenecks due to bus contention and excessive skipping during multi-unit parallel computing and data access. Taking FPGA as an example, if multiple accelerated computing units simultaneously access Double Data Rate Synchronous Dynamic Random Access Memory (DDR) / High Bandwidth Memory (HBM), traditional bus interconnects can lead to uneven bandwidth distribution and data transmission delays caused by skipping between accelerated computing units.
[0053] In the above embodiments of this application, as shown in FIG1, the overall topology of the controller is ring-shaped and completely symmetrical. Within each accelerated computing unit group, the accelerated computing units are connected through an internal bus and share the same memory resources to ensure that data is efficiently transmitted and accessed within the group. The accelerated computing units between different groups are connected in a specific way. Specifically, adjacent accelerated computing unit groups are connected sequentially, and each accelerated computing unit communicates with the accelerated computing unit located opposite it to minimize the jumps and delays in cross-group communication and optimize the overall interconnection performance.
[0054] In an optional embodiment, as shown in Figure 1, taking an FPGA chip as an example, in the topology, every four nodes are divided into groups (group 0, group 1, group 2, and group 3) in a clockwise direction. Taking group 0 as an example, the accelerated computing units 0, 1, 2, and 3 within group 0 are all connected to the same shared memory DDR / HSM_0 via an internal bus. This structural design allows the computing units within the group to efficiently access local memory. It is understood that, using this structure, accelerated computing units 0 and 3 can jump between each other via the shared memory DDR / HSM_0. Compared to related technologies where acceleration units 0 jump to acceleration units 3 one by one, this reduces the access latency caused by intra-group jumps.
[0055] In an optional embodiment, as shown in Figure 1, taking an FPGA chip as an example, in the topology, for example, in groups 0 and 1, adjacent accelerated computing units 3 and 4 communicate directly; for example, in groups 0 and 2, accelerated computing units 2 and 10 located opposite each other communicate directly. Through this structure, accelerated computing units 0 and 7 can jump via a path 0-8-7, while accelerated computing units 1 and 5 can jump via a path 1-3-4-5. Compared to clockwise or counterclockwise routing in related technologies, this significantly reduces access latency caused by intra-group jumps.
[0056] In summary, this application significantly reduces data transmission latency and improves the efficiency of data exchange between accelerated computing units by setting specific intra-group and cross-group communication paths. Furthermore, the use of an internal bus within the group to implement shared memory ensures that accelerated computing units can efficiently utilize local resources, avoiding resource utilization degradation caused by insufficient bandwidth allocation. Grouping accelerated computing units also better supports large-scale parallel computing. In addition, the circular and completely symmetrical topology ensures that the routing rules and effects are completely consistent regardless of the location of the accelerated computing units within the architecture, eliminating the need to consider performance deviations caused by the distribution of accelerated computing units with different functions during the design process. Moreover, the topology has a low node density and network diameter, which is beneficial for optimizing transmission latency. Using the above controller for accelerated computing can provide more powerful heterogeneous computing capabilities for fields such as database acceleration and deep learning.
[0057] In this embodiment, the interconnection logic block diagram of the controller is shown in Figure 2. The controller includes a device-side controller configured to receive CXL format data packets, parse the data packets to obtain accelerated computing tasks, and send them to the accelerated computing unit.
[0058] In related technologies, heterogeneous computing systems connect components via the CXL bus. It is understood that CXL, as an open high-speed interconnect technology standard, is widely used to connect devices such as Central Processing Units (CPUs), Graphics Processing Units (GPUs), accelerators, and memory to achieve low-latency, high-bandwidth data transmission. Therefore, to ensure data transmission in CXL format, this application configures a device-side controller in the controller, configured to receive and parse CXL format data packets. As shown in Figure 2, the device-side controller (CXL device-side controller) within the chip, in the controller architecture of this application, connects to the host via the CXL bus and is responsible for receiving data packets and parsing tasks via the CXL bus. Specifically, when the host sends accelerated computing tasks via the CXL bus in CXL format data packets, the device-side controller receives the data packets, parses them, extracts the optional accelerated computing tasks and related parameters required for accelerated computing, and then sends the accelerated computing tasks to the corresponding specific accelerated computing units.
[0059] In one optional embodiment, assume the host needs to perform an accelerated computing task involving database queries, which includes steps such as data packet decompression, full table scan, and expression filtering. During the distribution of the accelerated computing task, the host encapsulates these tasks into CXL format data packets and sends them to the device-side controller via the CXL bus. Upon receiving the data packets, the device-side controller parses them according to the CXL protocol, extracting information such as the task type, data element location, and target memory address. Then, the device-side controller selects an accelerated computing unit with the corresponding function based on the running status of the accelerated computing unit and sends the relevant information to that accelerated computing unit via the internal bus.
[0060] The aforementioned device-side controller, via a direct connection between the CXL bus and the internal bus, enables the accelerated computing unit to receive and process relevant data and control information for the accelerated computing task with minimal latency. Furthermore, the device-side controller can dynamically schedule tasks based on the functionality and status of the computing unit, thereby achieving parallel execution of tasks and optimal allocation of resources to the accelerated computing unit, ensuring minimal latency for the accelerated computing task.
[0061] The switching module is connected to the device-side controller via the CXL bus and communicates with the accelerated computing unit via an internal bus. It is configured to receive accelerated computing tasks and send them to the accelerated computing unit via direct memory access.
[0062] The switching module acts as a bridge between the device-side controller and the accelerated computing unit. This application connects the device-side controller and the switching module via a CXL bus, enabling the device-side controller to transfer data to the switching module via direct memory access (DMA) without CPU intervention. This significantly improves data transmission efficiency, reduces computational latency, and ensures high-speed data transmission and protocol compatibility. Furthermore, the accelerated computing unit communicates with the switching module via an internal bus, enabling efficient scheduling of accelerated computing tasks among the accelerated computing units.
[0063] The introduction of DMA technology between the switching module and the device-side controller ensures high-speed, low-power data transfer between the device-side controller and the accelerated computing unit, improving overall computing efficiency.
[0064] In summary, the combined use of the device-side controller and the switching module enables efficient reception, parsing, and execution of accelerated computing tasks in the controller-accelerated computing system, as well as fast, low-latency data transmission between the accelerated computing unit and memory, significantly improving the controller's performance and response speed in database accelerated computing tasks.
[0065] In this embodiment of the application, as shown in FIG2, the switching module includes a multi-level switching module, including a switching module that is connected to the device-side controller via a CXL bus and connected to other switching modules via multiple interfaces with output signals in master mode, and a switching module that is connected to other switching modules, connected to shared memory via an interface with output signals in slave mode, and connected to an accelerated computing unit via an interface with output signals in master mode.
[0066] To optimize data transmission paths and reduce the number of ports, this application proposes an interconnection mechanism based on multi-level switching modules, employing master-mode and slave-mode interfaces to distinguish links for different functions. The multi-level switching module includes at least two levels: a top-level switching module and a bottom-level switching module. The top-level switching module is directly connected to the device-side controller and receives CXL format data packets from the controller via the CXL bus. The output signal interfaces of the top-level switching module are set to master mode, connecting to the next-level intra-group switching modules through these interfaces to distribute data packets to different groups. The bottom-level switching modules are connected to shared memory (slave-mode interface) and accelerated computing units within the accelerated computing unit group (master-mode interface), respectively, responsible for sending data packets to the designated accelerated computing units according to task requirements.
[0067] In an optional embodiment, as shown in Figure 2, in the 16-node controller, the accelerated computing units are divided into four groups, each with four accelerated computing units, and each group includes a shared memory block. The switching module consists of only two layers, including the aforementioned top-level switching module (Crossebar) and bottom-level switching modules (Crossebar_0 to Crossebar_3). Each bottom-level switching module is connected to a group of accelerated computing units, and all bottom-level switching modules are directly connected to the top-level switching module. In an optional workflow, the host sends accelerated computing tasks to the device-side controller via the CXL bus. These tasks may include data reading, page parsing, expression filtering, multi-table joins, etc. After parsing the received CXL data packets, the device-side controller decomposes the tasks using DMA technology via the CXL bus and forwards them to the top-level switching module. Upon receiving the instructions, the top-level switching module distributes the data packets to the corresponding group-level switching modules through the master mode interface, based on the task requirements and the distribution of computing units. After receiving the task data assigned to its group, the group-level switching module uses the slave mode interface to exchange data with the shared memory, loading the required data into memory. Simultaneously, the main mode interface of the intra-group switching module sends data packets to the accelerated computing unit within the group to execute optional tasks. After the task is completed, the accelerated computing unit stores the results in shared memory, and then the intra-group switching module sends the result data to the top-level switching module through the slave mode interface, or transmits it directly within the chip. Ultimately, the data can be accessed by other computing units or the host.
[0068] By setting up multi-level switching and DMA data transfer, the controller of this application avoids the bottleneck problem in bus interconnection in related technologies and improves the output transmission efficiency. In the multi-level switching mechanism, the underlying switching module and the accelerated computing unit are grouped and connected, which facilitates the transmission of data packets along the optimal path. At the same time, the design of the master mode and slave mode interface simplifies the data scheduling between the computing unit and the shared memory, and realizes the efficient allocation of computing resources.
[0069] In this embodiment of the application, as shown in FIG3, the controller further includes: an on-chip routing module, which is communicatively connected to each accelerated computing unit group and is configured to implement routing jumps within the same controller.
[0070] In related technologies, in controller-based heterogeneous acceleration systems, the routing module plays an indispensable role in achieving high-performance and low-latency data transmission. It is responsible for correctly transmitting data packets from the source node to the destination node, a crucial process in complex system architectures. To ensure correct data packet transmission within the chip, this application establishes an on-chip routing module responsible for routing data packets between computing unit groups within a single controller. The on-chip routing module is tightly connected to the accelerated computing units and switching modules of each group, achieving fast data transmission through a low-latency interconnect structure. When a computing unit needs to access the memory of another group or exchange data with computing units in another group, the on-chip routing module efficiently completes the data packet routing according to a deterministic routing algorithm, ensuring the shortest data transmission path and lowest latency.
[0071] In an optional embodiment, as shown in Figure 3, in a 16-node controller architecture, the on-chip routing module is designed as a router with five ports, optionally including a clockwise port, a counter-clockwise port, a cross-directional port, a local port (connected to the local accelerated computing unit), and an inter-chip port for connecting inter-chip routing modules. When data packets need to be transmitted between different groups within the chip, the on-chip routing module first parses the routing ID of the data packet to determine the location of the target node, and then transmits the data through the most suitable port (such as a clockwise or counter-clockwise port) to achieve the shortest path hopping. For example, if the destination of the data packet is another group located in the clockwise direction, the on-chip routing module will perform routing hopping through the clockwise port, which can avoid unnecessary detours and improve data transmission efficiency.
[0072] The on-chip routing module is designed to ensure that data packets are transmitted within the controller architecture with the shortest path and lowest latency, improving the overall system response speed. The on-chip routing module connects to the accelerated computing unit in groups, and during transitions, it can optimize data transmission paths through deterministic and adaptive routing algorithms, avoiding resource waste and path bottlenecks, and ensuring that the computing unit can efficiently utilize local and remote resources.
[0073] In this embodiment of the application, as shown in FIG4, the controller further includes an inter-chip routing module, which is communicatively connected to the intra-chip routing module and is configured to implement routing jumps between different controllers.
[0074] The inter-chip routing module is primarily configured to handle output transmission between different controllers. In complex accelerated computing tasks, multiple controllers configured to work collaboratively may be required. In this case, the inter-chip routing module acts as a communication bridge between different controllers, responsible for cross-chip data packet routing and hopping, achieving efficient data transmission and routing selection.
[0075] In an optional embodiment, as shown in Figure 4, in a 16-node controller architecture, the inter-chip routing module is designed as a router with five ports. Four ports are connected to different intra-chip routing modules within the controller, and one port is connected to an inter-chip routing module within another controller. The inter-chip routing module is configured to receive data packets sent by the intra-chip routing modules, decode them according to the controller ID and accelerated computing unit group ID in the routing ID, and select the corresponding controller node for data transmission. For example, if the source node of the data packet is located in group 1 of controller 0, and the destination node is located in group 3 of controller 2, then inter-chip routing module 0 first sends the data packet to inter-chip routing module 2 of controller 2 through the inter-chip fully connected port. Then, the intra-chip routing module of controller 2 performs the final routing hop of the data packet according to the group ID and intra-group ID of the destination node, based on the shortest path principle, to ensure that the data directly reaches its destination.
[0076] The on-chip routing module ensures that data packets are transmitted between controller architectures with the shortest path and lowest latency, improving the overall system response speed. The inter-chip routing module connects to the on-chip routing module in a group. During transitions, it optimizes data transmission paths through deterministic and adaptive routing algorithms, avoiding resource waste and path bottlenecks, and ensuring that computing units can efficiently utilize local and remote resources. Furthermore, the inter-chip routing module supports efficient communication between multiple controllers, enabling the expansion of heterogeneous acceleration systems and allowing the system to support a larger number of accelerated computing units to meet more complex and larger-scale acceleration computing needs.
[0077] In the embodiments of this application, the on-chip routing module includes a first cross switch and a first cross switch control unit, wherein the first cross switch is configured to connect each accelerated computing unit group and the inter-chip routing module, and the first cross switch control unit is configured to control the on / off state of the path connected to the first cross switch.
[0078] In accelerated computing systems based on related technologies, on-chip communication efficiency has a crucial impact on the overall performance of the accelerated computing system. A crossbar switch is a key component for data transmission routing, effectively connecting different nodes and enabling rapid packet forwarding. This application sets a first crossbar switch responsible for establishing connections between different accelerated computing unit groups within the chip and between the crossbar switch and the chip's routing module, while the first crossbar switch control unit ensures the correctness and efficiency of these connections.
[0079] In an optional embodiment, as shown in Figure 5, the first crossbar switch in the on-chip routing module is designed as a 5×5 switching matrix, which can be connected to clockwise, counterclockwise, opposite, local accelerated computing unit groups, and inter-chip routing modules. This design allows data packets to flow freely in different directions. For example, in a 16-node controller, four accelerated computing unit groups are connected through the first crossbar switch, and the inter-chip routing module also establishes connections with these groups through the first crossbar switch. The first crossbar switch control unit is responsible for parsing the routing information of data packets and controlling the opening and closing of the connection paths of the first crossbar switch based on this information to achieve accurate data routing. In particular, when a data packet needs to be transmitted from one accelerated computing unit group to another group or the inter-chip routing module, the first crossbar switch control unit will determine whether the data packet should be transmitted through a clockwise port, counterclockwise port, opposite port, or local port based on the routing ID of the data packet and the current link load. If there is link congestion, the control unit will automatically select the less loaded path to achieve adaptive load balancing. For example, in the logic diagram of Figure 5, the first crossbar switch control unit needs to monitor the load of all inputs [0-4] and outputs [0-4]. When a data packet enters from input [0] (clockwise direction), if the buffer of output [1] (counterclockwise direction) is not full, the first cross switch is controlled so that the data packet can flow directly from input [0] to output [1], thereby completing the data transmission in the clockwise direction.
[0080] By incorporating an on-chip routing module that includes a first cross switch and a first cross switch control unit, this application achieves efficient data transmission and adaptive routing selection. This not only optimizes communication between computing units but also improves the interaction efficiency with the inter-chip routing module, providing powerful hardware support for data-intensive tasks such as database acceleration computing and deep learning.
[0081] In the embodiments of this application, the inter-chip routing module includes a second cross switch and a second cross switch control unit, wherein the second cross switch is configured to connect the intra-chip routing module with other controllers, and the second cross switch control unit is configured to control the on / off state of the path connected to the second cross switch.
[0082] The second crossbar switch is the same high-performance switching matrix as the first crossbar switch. In the inter-chip routing module, the second crossbar switch is configured to connect the intra-chip routing module and the multiplexer, thereby enabling cross-chip packet routing. As shown in Figure 4, it contains five ports, corresponding to the inter-chip ports of the FPGA intra-chip routing module and the inputs of the multiplexer. This design allows packets to flow freely between different FPGA chips. The second crossbar switch control unit is responsible for parsing the packet routing information and controlling the path selection of the second crossbar switch based on this information to achieve fast transmission of packets from the source chip to the target chip. That is, under the deterministic routing strategy, the control unit needs to determine the transmission path of the packet based on the controller ID, the accelerated computing unit group ID, and the accelerated computing unit ID, and ensure that it can intelligently select other paths when encountering link congestion. In particular, when a packet enters the second crossbar switch from the inter-chip port of the intra-chip routing module, the control unit checks the ID information of the controller corresponding to the target node and the current network status. If the link load in the direction corresponding to the target controller is low, the control unit will guide the packet to be transmitted through the output port in that direction.
[0083] By incorporating an inter-chip routing module that includes a second crossbar switch and a control unit for the second crossbar switch, this application ensures fast and efficient path selection for cross-chip data transmission, significantly reducing latency during data packet transmission between chips. In the event of link failure or congestion, the second crossbar switch control unit can intelligently select an alternative path, improving the overall robustness and reliability of the network. This ensures that data transmission can continue even under adverse network conditions, achieving efficient cross-chip data transmission and adaptive routing. This not only optimizes inter-chip communication and improves data transmission speed but also enhances the network robustness and reliability of the accelerated computing system when handling tasks such as database acceleration computing and large-scale data processing.
[0084] In the embodiments of this application, as shown in FIG5, the intra-chip routing module or inter-chip routing module includes a virtual channel control unit, which is configured to divide a physical channel into multiple virtual channels to prevent deadlock caused by channel closure.
[0085] In accelerated computing systems based on related technologies, the complexity and high concurrency of data communication can lead to network congestion and deadlock, especially in bus- and ring-based interconnection structures. This application proposes to avoid deadlock and improve network efficiency by using virtual channel technology to divide physical channels into multiple logically independent virtual channels.
[0086] In an optional embodiment, as shown in Figure 5, both the intra-chip routing module and the inter-chip routing module integrate a virtual channel control unit, which divides each physical channel into four independent virtual channels (VCs): VC0, VC1, VC2, and VC3. In this way, different virtual channels can be used when data packets are transmitted in different directions, thus avoiding the channel loop problem common in ring networks and effectively preventing deadlock. Specifically, the virtual channel control unit is located at the front end of each input and output port and is configured to manage the channel allocation for inbound and outbound data packets. By dividing the physical channel into four independent virtual channels, when a data packet enters from a certain direction, the control unit selects a virtual channel for transmission according to predefined rules or the current network state, ensuring that data packets do not create loops in the channel under any circumstances, thus avoiding deadlock.
[0087] The virtual channel control unit improves the transmission efficiency of the on-chip network by preventing deadlock and achieving adaptive load balancing, ensuring that data packets arrive at their destination in the shortest possible time. The deadlock avoidance mechanism guarantees the continuity of data transmission and the stability of system operation, especially in high-concurrency or network congestion environments, effectively preventing system performance degradation. In summary, this application introduces a virtual channel control unit into the intra-chip routing module and the inter-chip routing module. This application achieves high efficiency, stability, and resource optimization of data transmission in accelerated computing systems, providing strong hardware support for demanding tasks such as accelerated database computing and large-scale data processing, ensuring high performance and reliability when handling complex interconnected communications.
[0088] In the embodiments of this application, the virtual channel control unit is configured to: construct multiple independent micro-piece buffers in the output queue buffer corresponding to the physical channel to obtain multiple virtual channels; divide the virtual channels into a first channel group and a second channel group in a clockwise direction; use the first channel group when the routing coordinates of the source node are greater than the routing coordinates of the target node; and use the second channel group when the routing coordinates of the source node are less than or equal to the routing coordinates of the target node.
[0089] The virtual channel control unit transforms a single physical channel into multiple logically independent virtual channels by creating multiple independent micro-chip buffers in the output queue buffer of the physical channel, thereby improving the communication efficiency and stability of the network.
[0090] In one optional embodiment, the virtual channel control unit creates four independent micro-buffers, namely VC0, VC1, VC2, and VC3, in the output queue buffer of each physical channel, collectively forming four virtual channels. Each micro-buffer is responsible for processing data packets in a specific direction, ensuring that data packet transmission does not form a closed loop and avoiding deadlock. Specifically, when a data packet needs to be transmitted from one computing unit group to another, it will be assigned by the virtual channel control unit to a specific virtual channel for transmission, rather than through a single physical channel. The virtual channels are equally divided into a first channel group (containing VC0 and VC1) and a second channel group (containing VC2 and VC3) in a clockwise direction. This division method is associated with the transmission direction of the data packets, making data transmission in a specific direction more orderly and efficient. For example, in the on-chip routing module, data packets transmitted in a clockwise direction will use the first channel group (VC0 and VC1), while data packets transmitted in a counterclockwise direction will use the second channel group (VC2 and VC3). In actual transmission, the choice between using the first or second channel accelerated computing unit group is based on the routing coordinates (controller ID, accelerated computing unit group ID, accelerated computing unit ID) of the source and destination nodes. When the routing coordinates of the source node are greater than those of the destination node, the first channel accelerated computing unit group is used for transmission; when the routing coordinates of the source node are less than or equal to those of the destination node, the second channel accelerated computing unit group is used. Specifically, in a controller cluster architecture, if a data packet is transmitted from a node in accelerated computing unit group 0 (e.g., controller ID = 0, accelerated computing unit group ID = 0, accelerated computing unit ID = 0) to a node in accelerated computing unit group 3 (controller ID = 0, accelerated computing unit group ID = 3, accelerated computing unit ID = 0), since the accelerated computing unit group ID of the source node is less than that of the destination node, the second channel accelerated computing unit group (VC2 and VC3) should be used for transmission. Conversely, if a data packet is transmitted from a node in Accelerated Computing Unit Group 3 to a node in Accelerated Computing Unit Group 0, the first channel accelerated computing unit group (VC0 and VC1) will be used for transmission because the accelerated computing unit group ID of the source node is greater than that of the target node.
[0091] By dividing the physical channel into multiple virtual channels and selecting appropriate channel groups for data transmission based on direction and routing coordinates, this application effectively avoids the deadlock problem common in ring networks, ensuring the continuity of data packet transmission and system stability. The channel grouping and routing coordinate-based channel selection strategy simplify the routing algorithm of the on-chip communication network, making the system design simpler, while reducing data transmission latency and improving data transmission speed and overall system performance.
[0092] In the embodiments of this application, when the accelerated computing unit group is in the first working mode, the accelerated computing units within the group can be configured to different computing functions.
[0093] The first working mode allows accelerated computing units within the same group to be configured to perform different computational functions. This means that different types of accelerated computing units can be flexibly allocated and combined according to application requirements to achieve optimal performance optimization. This is crucial for system flexibility and task execution efficiency. In particular, each accelerated computing unit can be dynamically configured to perform different computational tasks, such as data decompression, full table scan, expression filtering, sorting, hashing, etc. This configuration method provides great flexibility, enabling the system to select the most suitable accelerated computing unit for acceleration based on the characteristics of different tasks.
[0094] The aforementioned first operating mode allows accelerated computing units to work collaboratively to complete a series of complex database acceleration computing tasks without requiring each accelerated computing unit to possess all the functions, significantly improving the system's resource utilization efficiency. Furthermore, because the accelerated computing units within the group share memory, in the first operating mode, although the computing units within the group can be configured to perform different functions, these units directly access the shared memory during task execution, reducing data transmission latency and energy consumption. For example, after the expression filtering computing unit completes its task, it can directly write the result to the shared memory within the group without needing to transmit data through the on-chip network. In this way, the sorting computing units within the same group can immediately access these results for the next calculation, reducing data transmission overhead and improving the overall system efficiency. Moreover, the shared group memory design reduces unnecessary data transmission, lowers energy consumption during data transmission, and also reduces communication latency between computing units, providing strong technical support for building low-power, high-performance computing systems.
[0095] In the embodiments of this application, when the accelerated computing unit group is in the second working mode, the accelerated computing units in the group are configured to have the same computing function so that if any accelerated computing unit in the group fails, other accelerated computing units in the group can be used to replace it.
[0096] In large-scale parallel computing scenarios, hardware failures are inevitable. In order to improve the stability and fault tolerance of the system, this application proposes a second working mode in addition to the first working mode. The second working mode is designed to run when all accelerated computing units in the group are configured to perform the same computing function. In this way, even if one computing unit fails, the system can still replace the function through other computing units in the group, ensuring the continuity of tasks and the stability of system performance.
[0097] In one optional embodiment, in the second operating mode, all computing units within the same accelerated computing unit group are configured to perform the same computing function. For example, all computing units within a group can be set to perform tasks such as data decompression or full table scans to ensure that if any computing unit fails, other computing units can immediately take over its work without reconfiguration or scheduling. Specifically, assuming that (0,0,0), (0,0,1), (0,0,2), and (0,0,3) within group 0 are configured in the second operating mode and set to perform the same function, such as data decompression, then if any computing unit (e.g., (0,0,1)) experiences a hardware failure, other computing units (e.g., (0,0,0), (0,0,2), and (0,0,3)) can immediately take over the task and continue executing the data decompression task using the shared memory resources within the group to ensure uninterrupted task execution.
[0098] In the second operating mode, since all computing units within the group are configured to perform the same function, they can share data and resources more effectively, reducing the steps required to retransmit data during fault replacement and further improving computational efficiency and resource utilization. The instant task reallocation and data sharing mechanism provided by shared memory ensures that the real-time performance of tasks and the overall stability of the system are not affected by hardware failures, which is particularly important for tasks with high real-time requirements such as accelerated database computing and large-scale data processing.
[0099] In embodiments of this application, the controller is one of an FPGA, an Advanced RISC Machine (ARM), and an Enhanced RISC Controller (Atmel Virtual Research, AVR).
[0100] FPGAs, ARM microprocessors, and AVR microprocessors each have their own characteristics and are suitable for different scenarios and needs. When an FPGA acts as a controller, it can provide highly customized control logic, suitable for system designs that require high flexibility and programmability. As a controller, an ARM microprocessor can provide powerful processing capabilities and rich peripheral interfaces, suitable for tasks requiring high-performance computing and complex system management. AVR microprocessors are ideal for some cost-sensitive and power-constrained applications, and their low cost and low power consumption make it possible to deploy sensor networks on a large scale while ensuring system stability and efficiency.
[0101] Depending on the design requirements and application scenarios of the accelerated computing system, selecting FPGA, ARM microprocessor or AVR microprocessor as the controller can achieve functional optimization and performance improvement at different levels, providing a variety of hardware platform options for building efficient, stable and economical electronic devices and accelerated computing systems.
[0102] In this embodiment of the application, an accelerated computing system is provided, as shown in Figure 6. The accelerated computing system includes a host, multiple controller clusters, and a switching chip (which can be a CXL Switch). The host is responsible for task scheduling and data management, the controller clusters are responsible for executing accelerated computing tasks, and the switching chip is responsible for efficient data exchange between the host and the controller clusters.
[0103] In one alternative embodiment, assuming the host has a large-scale data processing task, such as database acceleration computation, it decomposes the task into multiple subtasks and distributes them to different controller clusters according to the computational requirements of the subtasks. Controllers within each controller cluster efficiently execute the corresponding subtasks based on their configured computational functions; for example, a controller in one cluster might be configured specifically for data decompression, while a controller in another cluster might handle full table scans. The switching chip, acting as a data bridge between the host and the controller clusters, is responsible for achieving efficient bus connectivity and data exchange. It can handle large data streams, ensuring fast data transmission speeds and low latency between the host and controller clusters, thus improving the overall system performance. For example, the host connects to the switching chip via a high-speed bus such as PCIe or CXL, and the switching chip then connects to each controller cluster via a CXL bus. When the host needs to send data to a controller cluster for computation, the data is first sent to the switching chip via the host-side bus. The switching chip repackages the data based on the destination controller cluster and group information and sends it to the corresponding controller cluster. Within the controller cluster, the data is transmitted to the target acceleration computing unit via an on-chip network or packet bus to execute the corresponding computational task. The aforementioned switching chip is a heterogeneous cache coherency switching chip, which solves the problem of insufficient cache coherency support in on-chip interconnects in related technologies.
[0104] By employing an accelerated computing system architecture, this application can allocate computing resources according to task requirements, thereby accelerating data processing and improving overall system performance and data processing speed. The controller cluster design allows the system to dynamically adjust and expand according to application needs, increasing system flexibility and scalability, enabling the system to adapt to various computing tasks and data volumes. Through the configuration of multiple controller clusters and the improved internal interconnection structure, even in the event of a controller or accelerated computing unit failure, the system can still continue to complete tasks using other available resources, improving system reliability and fault tolerance.
[0105] In an embodiment of this application, as shown in FIG7, the controllers in the controller cluster are configured as a fully connected structure.
[0106] A fully connected architecture refers to a controller cluster where all controllers are directly interconnected. This means each controller can communicate directly with other controllers without complex routing strategies. As shown in Figure 7, each controller comprises 16 nodes, and the full connectivity between controllers is achieved through routing. This architecture improves data transmission efficiency, reduces communication latency, and enhances the coordination capabilities between system resources.
[0107] In one alternative embodiment, within the controller cluster, the fully connected structure can be implemented through the CXL interface. Through the CXL interface, each controller can directly access the main memory and the memory resources of other controller clusters, reducing data transmission latency and increasing bandwidth. In addition, the system-wide unified addressing feature of the CXL protocol enables the accelerated computing units of different controller clusters to access remote memory as if they were accessing local memory, thereby realizing data flow within the fully connected structure.
[0108] A fully connected controller cluster enables dynamic task allocation and data sharing, which is particularly important for tasks such as database acceleration. The host can allocate data and tasks to any controller as needed, and these controllers can access other controllers with only a single hop, without additional routing control. In summary, a fully connected controller cluster significantly improves data transmission efficiency, reduces communication latency, and simplifies task scheduling and data management. This allows the host to allocate tasks more flexibly, while acceleration computing units can directly access the required data, improving the utilization of computing resources and the parallelism of task execution. Efficient and convenient mutual access also enhances the collaborative capabilities of the system containing the controller cluster, providing key technical support for building low-latency, high-bandwidth, high-performance computing systems.
[0109] In this embodiment, the switching chip supports at least the CXL communication protocol.
[0110] The CXL communication protocol is a high-speed, low-latency interconnect technology standard designed to meet the needs of modern data centers and high-performance computing systems for high-speed data transmission and cache coherency support. The CXL protocol supports three main modes: CXL.io, CXL.cache, and CXL.memory, used for inter-device input / output communication, cache coherency communication, and shared memory access, respectively. In this application, the switch chip is designed to support the CXL communication protocol, thereby enabling efficient data transmission and cache coherency management within the controller cluster architecture. This means that data can flow efficiently between various components in the system without increasing latency or reducing bandwidth. As shown in Figure 6, the host communicates with the switch chip via the CXL bus, while the switch chip connects to the CXL device-side controllers in each FPGA cluster via the CXL bus. This design allows the host to directly access the DDR / HBM memory in the controller cluster without complex routing and conversion processes, thus greatly improving data access speed and overall system performance. Furthermore, CXL communication between controller clusters can omit additional conversions or intermediate nodes. For example, in the interconnection diagram of routing modules between FPGAs shown in Figure 6, it can be seen that different FPGAs are fully connected through switching chips supported by the CXL protocol. This means that data transmission between any two FPGA clusters can be completed directly through the CXL bus without complex routing decisions and packet redirection. Switching chips supporting the CXL protocol can achieve cache consistency between different FPGA clusters, ensuring data consistency and integrity, and avoiding data conflicts and inconsistencies.
[0111] In summary, switching chips supporting the CXL communication protocol provide high-speed, low-latency data transmission capabilities, which are crucial for high-performance computing and data acceleration tasks in controller cluster architectures. Through CXL protocol support, switching chips can achieve inter-chip cache coherency, ensuring data integrity and consistency in multi-controller cluster architectures and improving system reliability. Switching chips under the CXL protocol simplify the communication architecture of the system, reduce intermediate links in data transmission, make data flow more direct and efficient, reduce system latency, and improve overall processing speed.
[0112] In an embodiment of this application, as shown in FIG4, the controller cluster further includes a multiplexer configured to receive signals generated internally by the controller and transmit them to other controllers or a host.
[0113] A multiplexer is a circuit component that can select one signal from multiple input signals for transmission. In a multi-controller architecture, the multiplexer is configured to select the correct output path in inter-chip communication to send data packets or signals to a specific target chip.
[0114] In an optional embodiment, as shown in Figure 4, a multiplexer is connected to multiple controllers after being connected to the inter-chip routing module. Each controller has a corresponding input terminal. By decoding the controller ID, the multiplexer can select the correct output path and send data packets or signals to the specified target controller. As shown in Figure 6, assuming there is a cluster containing four FPGAs, each FPGA has a port connected to the multiplexer. When a data packet needs to be transmitted from the current controller to another controller, the on-chip routing module sends the data packet to the inter-chip fully connected port, which then passes the data packet to the multiplexer. After receiving the signal from the inter-chip fully connected port, the multiplexer decodes the target controller ID in the data packet and transmits the signal to the output port that matches the target controller ID.
[0115] The introduction of the multi-to-one selector provides direct and efficient path selection and data transmission capabilities, greatly simplifies the data transmission process between different controllers, reduces the path selection delay of data packets during routing hops, and avoids unnecessary detours of data packets in the interconnection network of multiple controllers, thereby significantly improving the overall data transmission efficiency and speed of the accelerated computing system.
[0116] In the embodiments of this application, a routing method within a controller is proposed, which can run on the accelerated computing system shown in Figure 6. Figure 8 is a flowchart of the aforementioned routing method within the controller. As shown in Figure 8, the above process includes:
[0117] Step S201: Configure the routing coordinates of the accelerated computing units inside the controller. The routing coordinates include the controller ID, the accelerated computing unit group ID, and the accelerated computing unit ID.
[0118] In some embodiments of this application, each accelerated computing unit is assigned corresponding routing coordinates, which include three dimensions: the controller ID, the accelerated computing unit group ID, and the accelerated computing unit ID. For example, accelerated computing unit (0, 0, 0) represents the first accelerated computing unit in the first group located on controller 0. This coordinate configuration makes the location and connection relationship of each computing unit in the system clearly identifiable, providing a basis for subsequent routing hops.
[0119] Step S202: Determine the routing coordinates of the source node and the target node based on the routing jump request, and calculate the number of intra-chip jump steps based on the routing coordinates of the source node and the target node.
[0120] Optionally, when one accelerated computing unit needs to transmit data to another accelerated computing unit, the accelerated computing system first determines the routing coordinates of the source node and the target node, and then calculates the intra-chip jump number based on the routing coordinates of the two, using the following formula:
[0121] Intra-chip jump steps = ((Accelerated computing unit group ID of the destination node * number of accelerated computing units in the group + accelerated computing unit ID of the destination node) - (Accelerated computing unit group ID of the source node * number of accelerated computing units in the group + accelerated computing unit ID of the source node) + number of accelerated computing nodes) mod (i.e., modulo operation) number of accelerated computing nodes.
[0122] Step S203: Determine the routing direction based on the number of intra-chip jump steps, and perform routing jump based on the routing direction.
[0123] Based on the calculated number of intra-chip jump steps, the accelerated computing system determines the routing jump direction according to the shortest path principle, and then controls the transmission of data packets according to the routing jump direction to complete the routing jump process.
[0124] In one optional embodiment, the source node coordinates are set to (0, 0, 0), and the target node coordinates are set to (0, 3, 0). In a 16-node FPGA cluster, the on-chip jump step count is calculated to be 12. According to the routing strategy, the system determines that this jump step count falls within the interval (3 / 4, 1), and therefore chooses to jump in a clockwise direction. This means that the data packet will travel through the path (0, 0, 0) → (0, 1, 0) → (0, 2, 0) → (0, 3, 0) to directly reach the target node without passing through unnecessary intermediate nodes, thus reducing latency.
[0125] Through the above steps, based on the calculation of the number of jump steps and the determination of the route jump direction according to the routing coordinates, this application can optimize the transmission path of data packets within the FPGA, ensuring that data packets arrive at the target node with the shortest path and lowest latency, thereby significantly improving data transmission efficiency. The deterministic routing strategy and shortest path selection reduce latency during data transmission and minimize unnecessary data rerouting, thus reducing power consumption and improving the overall system performance and energy efficiency. The coordinate-based routing jump method allows the system to flexibly adapt to FPGA cluster architectures of different sizes, and when expanding the system, it is easy to maintain and adjust the routing strategy, ensuring that newly added computing units can be smoothly integrated into the existing network, improving system flexibility and scalability.
[0126] The entity performing the above steps can be a host, controller, etc., but is not limited to these.
[0127] To avoid insufficient bandwidth and data congestion, and to reduce the computational efficiency of the accelerated computing system, in the embodiments of this application, routing is performed based on the routing direction, including:
[0128] Step S2031: Obtain the output buffer length of the controller port;
[0129] Step S2032: If the output buffer length is greater than the fifth threshold, reverse routing is performed.
[0130] Accelerating communication between computing units requires complex routing networks. However, the congestion of these networks is closely related to communication latency and packet loss. To avoid performance degradation caused by network congestion, this application proposes a routing strategy with a buffer length of the aforementioned technology. When the buffer length is long, reverse routing is performed.
[0131] In one alternative embodiment, as shown in Figure 5, the first crossbar switch control unit needs to monitor the load of all inputs [0-4] and outputs [0-4]. When a data packet enters from input [0] (clockwise), the control unit checks the load in the output direction. If the buffer of output [1] is full (i.e., greater than the fifth threshold mentioned above), the control unit will guide the data packet to be routed in the reverse direction through output [2] (opposite direction) or output [3] (clockwise direction) to avoid the data packet encountering congestion during transmission.
[0132] Through the above steps, this application monitors the output buffer length of the port and employs a reverse routing strategy when the port buffer is full. This effectively balances the communication load in the network, avoids local network congestion, and ensures smooth and efficient data transmission in the accelerated computing system. Therefore, the reverse routing strategy avoids data transmission through high-load ports, thereby reducing communication latency and maintaining the system's high performance and response speed. Furthermore, dynamic routing adjustment reduces the queuing time of data packets on congested links, lowers the risk of data packet loss, and improves communication reliability.
[0133] In embodiments of this application, determining the routing jump direction based on the number of intra-chip jump steps includes:
[0134] Step S2033: Based on the preset range of the number of jump steps within the chip, determine the route jump direction. The route jump direction includes clockwise direction, counterclockwise direction, opposite direction, or a combination of clockwise direction, counterclockwise direction, and opposite direction.
[0135] To satisfy the shortest path principle, this application provides a method for determining the routing direction based on the number of intra-chip jump steps. Specifically, by analyzing the preset range of intra-chip jump steps, the most suitable routing direction is intelligently selected, thereby reducing communication latency and improving data transmission efficiency.
[0136] Based on the pre-configured routing strategy of this application, the accelerated computing system can intelligently select the shortest path for data packets within the controller's internal network, significantly reducing communication latency and improving data transmission efficiency. The preset jump step range and routing jump direction strategy allows the system to flexibly adjust under different communication needs and network conditions, improving communication stability and reliability.
[0137] The execution order of step S2033, step S2031, and step S2032 can be interchanged. That is, step S2033 can be executed first, followed by steps S2031 and S2032, or steps S2033, S2031, and S2032 can be executed simultaneously.
[0138] In the embodiments of this application, determining the routing jump direction based on a preset range of intra-chip jump steps includes:
[0139] Step S20331: If the number of intra-chip jump steps is equal to the first threshold, the clockwise direction is determined as the routing jump direction;
[0140] Step S20332: If the number of intra-chip jump steps is equal to the second threshold, the opposite direction is determined as the routing jump direction;
[0141] Step S20333: If the number of intra-chip jump steps is equal to the third threshold, the counterclockwise direction is determined as the routing jump direction.
[0142] Step S20334: If the number of intra-chip jump steps is less than the first threshold, the counterclockwise direction is determined as the routing jump direction;
[0143] Step S20335: If the number of intra-chip jump steps is greater than the first threshold and less than the second threshold, the opposite direction followed by the clockwise direction will be determined as the route jump direction.
[0144] Step S20336: If the number of intra-chip jump steps is greater than the second threshold and less than the third threshold, the route jump direction will be determined as first the opposite direction and then the counterclockwise direction.
[0145] Step S20337: If the number of intra-chip jump steps is greater than the third threshold and less than the fourth threshold, the clockwise direction is determined as the routing jump direction.
[0146] Based on the controller architecture shown in Figure 1, taking a controller with a scale of 16 nodes as an example, the first threshold of this application is set to 4 (i.e., 1 / 4 of the number of computing nodes), the second threshold is set to 8 (i.e., 1 / 2 of the number of computing nodes), the third threshold is set to 12 (i.e., 3 / 4 of the number of computing nodes), and the fourth threshold is set to 16 (i.e., the number of computing nodes). When the number of intra-chip jump steps equals the first threshold (4), the system determines the clockwise direction as the routing jump direction; when the number of intra-chip jump steps equals the second threshold (8), the system determines the opposite direction as the routing jump direction; when the number of intra-chip jump steps equals the third threshold (12), the system determines the counterclockwise direction as the routing jump direction; when the number of intra-chip jump steps is less than the first threshold (i.e., 0 to 3), the system determines the counterclockwise direction as the routing jump direction; when the number of intra-chip jump steps is greater than the first threshold (i.e., 4 to 7) and less than the second threshold (i.e., 8), the system determines the opposite direction first and then the clockwise direction as the routing jump direction; when the number of intra-chip jump steps is greater than the second threshold (i.e., 8 to 11) and less than the third threshold (i.e., 12), the system determines the opposite direction first and then the counterclockwise direction as the routing jump direction; when the number of intra-chip jump steps is greater than the third threshold (i.e., 12 to 15) and less than the fourth threshold (i.e., 16), the system determines the clockwise direction as the routing jump direction. Specifically, assuming a data packet needs to be transmitted from (0, 0, 0) in group 0 to (0, 3, 0) in group 3, the intra-chip jump step count is 12 in a 16-node controller. According to the strategy of this application, the transmission direction is determined to be counter-clockwise, meaning the data packet will be transmitted counter-clockwise from (0, 0, 0) through the path (0, 3, 3) → (0, 2, 3) → (0, 1, 3) → (0, 3, 0), ensuring efficient data scheduling.
[0147] By employing a routing strategy based on the number of intra-chip jump steps and a preset threshold range, this application ensures that data packets are transmitted along the shortest path within the controller's internal network, significantly reducing communication latency and improving data transmission efficiency. The preset threshold range and routing strategy dynamically adapt to different communication needs, achieving intelligent and efficient data scheduling and improving the overall system performance and reliability. In summary, the above steps design accelerate the bus connection within the computing unit group and the communication connection between groups to achieve low-latency, high-bandwidth data transmission. Simultaneously, by setting first to fourth thresholds, and based on the range of intra-chip jump steps, clockwise, counterclockwise, or opposite-side routing jumps, as well as combinations of these directions, are intelligently determined, thereby optimizing the data packet transmission path within the controller's internal network. Through this flexible routing strategy, the system can effectively cope with network congestion, ensuring fast and reliable data transmission, thereby improving the overall computing performance and efficiency of the system.
[0148] In an embodiment of this application, a routing method between controllers is proposed, which can run on the accelerated computing system shown in Figure 6. Figure 9 is a flowchart of the aforementioned routing method between controllers. As shown in Figure 9, the above process includes:
[0149] Step S301: Configure the routing coordinates of the accelerated computing units inside the controller. The routing coordinates include the controller ID, the accelerated computing unit group ID, and the accelerated computing unit ID.
[0150] Step S302: Determine the route coordinates of the source node and the route coordinates of the target node based on the route redirection request;
[0151] Step S303: Based on the routing coordinates of the source node, jump from the intra-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the source node; based on the routing coordinates of the target node, jump from the inter-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the target node; and based on the routing coordinates of the target node, jump from the inter-chip routing module corresponding to the target node to the intra-chip routing module corresponding to the target node, and transmit to the target node.
[0152] When data needs to be transmitted between accelerated computing units of different controllers, intra-chip routing is insufficient to complete the data transmission. In this case, the system will utilize the inter-chip routing module to perform cross-controller routing. For example, if the source node coordinates are (0, 0, 0) and the target node coordinates are (1, 0, 0), the data packet will jump from the source node's intra-chip routing module to the inter-chip routing module of controller 0, then to the inter-chip routing module of controller 1, and finally reach the target node's intra-chip routing module before being transmitted to the target node.
[0153] In the above steps, by configuring the routing coordinates of the accelerated computing units, the system can accurately calculate the number of intra-chip jump steps, providing clear path guidance for data transmission and reducing routing errors and communication latency. The combined use of the intra-chip routing module and the inter-chip routing module enables this application to support communication within the same controller and communication across controllers, meeting the requirements for efficient data exchange between heterogeneous accelerated computing units in a controller cluster architecture.
[0154] In the embodiments of this application, a method for scheduling accelerated computing units for accelerated computing tasks is proposed. This method can run on the accelerated computing system shown in Figure 6. Figure 10 is a flowchart of the routing and hopping method between the controllers described above. As shown in Figure 10, the above process includes:
[0155] Step S401: Obtain the accelerated computing unit status table and the memory access latency table. The computing unit status table is used to include at least the routing coordinates, unit functions, unit status and storage usage of the accelerated computing unit. The memory access latency table includes the latency of each accelerated computing unit accessing each shared cache.
[0156] The accelerated computing unit status table records detailed information about existing accelerated computing units (as shown in Table 1), including but not limited to route coordinates, unit function, unit status (such as idle or occupied), and storage usage (i.e., whether the current unit is using its shared cache space). The memory access latency table records latency information when each accelerated computing unit accesses different shared caches, used to analyze data transfer efficiency and optimize scheduling strategies (as shown in Table 2).
[0157] Table 1
[0158] Table 2
[0159] Step S402: Obtain the computing tasks to be assigned, and determine the scheduling strategy based on the computing unit status table and the memory access latency table. During the process of determining the scheduling strategy, accelerated computing units with different storage occupancy can run in parallel. In the scheduling strategy, each accelerated computing unit corresponds to the computing task to be assigned with the lowest latency.
[0160] The system acquires accelerated computing tasks to be executed. These tasks may include database acceleration operations such as data decompression, full table scan, expression filtering, and sorting. For example, in an accelerated computing system, the order of tasks to be assigned might be data decompression, full table scan, expression filtering, and sorting. The accelerated computing system will intelligently allocate tasks to the most suitable accelerated computing unit based on the status of the accelerated computing unit, the requirements of the tasks to be assigned, and information from the memory access latency table. This process considers the following factors:
[0161] Unit Function Matching: The system assigns tasks to accelerated computing units whose functions match the task's. Unit Status: The system prioritizes idle accelerated computing units to avoid overlapping task execution. Storage Usage: The system considers the storage usage of different computing units to achieve efficient execution of parallel tasks. Memory Access Latency: Based on a memory access latency table, the system selects the computing unit and memory pair with the lowest latency to reduce data transfer time.
[0162] In one optional embodiment, assuming the system needs to perform a data decompression task, firstly, according to the computing unit status table, the system finds that the computing unit at coordinates (0, 1, 2) is idle and has decompression capabilities. Next, the system checks the memory access latency table and determines that this computing unit has the shortest latency for accessing memory in group 1. Therefore, the system selects the accelerated computing unit at coordinates (0, 1, 2) to perform the data decompression task. After this unit completes the task and stores the calculation results in memory, the system continues to check the status table and latency table to select the accelerated computing unit with the lowest latency and matching functionality for subsequent full table scan tasks. This process is repeated until all tasks are completed.
[0163] By using an accelerated computing unit status table, the accelerated computing system can accurately identify the function and status of each unit, thereby allocating tasks to the most suitable computing unit and improving task execution efficiency. A memory access latency table provides latency information, enabling the system to intelligently select the path with the least latency for data transmission, reducing communication latency and increasing data transmission speed. Furthermore, during task scheduling, the accelerated computing system considers the possibility of parallel operation of accelerated computing units with different storage occupancy rates. By rationally allocating tasks, it achieves efficient parallel execution of multiple tasks, further improving the overall computing performance of the system.
[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0165] The optional examples in this embodiment can refer to the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0166] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0167] The above are merely optional embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A controller for accelerating computation, characterized by, include: Multiple accelerated computing unit groups are arranged in a completely symmetrical ring, with adjacent accelerated computing unit groups communicating with each other. Accelerated computing units within the same accelerated computing unit group are connected sequentially via an internal bus. Accelerated computing units in different accelerated computing unit groups that are located at symmetrical positions on the same diameter of the ring are also connected with each other. Each accelerated computing unit group also includes shared memory, which is connected to each of the accelerated computing units via the internal bus.
2. The controller according to claim 1, characterized in that, The controller further includes: The device-side controller is configured to receive data packets in CXL format for computing fast connection, parse the data packets to obtain accelerated computing tasks, and send them to the accelerated computing unit. The switching module is connected to the device-side controller via a CXL bus and communicates with the accelerated computing unit via an internal bus. It is configured to receive the accelerated computing tasks via direct memory access and send them to the accelerated computing unit.
3. The controller according to claim 2, characterized in that, The switching module includes a multi-level switching module, including a switching module that is connected to the device-side controller via the CXL bus and connected to other switching modules via multiple interfaces with master output signals, and a switching module that is connected to other switching modules, connected to the shared memory via an interface with slave output signals, and connected to the accelerated computing unit via an interface with master output signals.
4. The controller according to claim 1, characterized in that, The controller also includes: The on-chip routing module is communicatively connected to each of the aforementioned accelerated computing unit groups and is configured to enable routing jumps within the same controller; The inter-chip routing module, which is communicatively connected to the intra-chip routing module, is configured to enable routing jumps between different controllers.
5. The controller according to claim 4, characterized in that, The intra-chip routing module includes a first cross switch and a first cross switch control unit. The first cross switch is configured to connect each of the accelerated computing unit groups and the inter-chip routing module. The first cross switch control unit is configured to control the on / off state of the path connected to the first cross switch.
6. The controller according to claim 4, characterized in that, The inter-chip routing module includes a second cross switch and a second cross switch control unit, wherein the second cross switch is configured to connect the intra-chip routing module to other controllers, and the second cross switch control unit is configured to control the on / off state of the path connected to the second cross switch.
7. The controller according to claim 4, characterized in that, The intra-chip routing module or the inter-chip routing module includes a virtual channel control unit, which is configured to divide a physical channel into multiple virtual channels to prevent deadlock caused by channel closure.
8. The controller according to claim 7, characterized in that, The virtual channel control unit is set to: Multiple independent micro-piece buffers are constructed in the output queue buffer corresponding to the physical channel to obtain multiple virtual channels; The virtual channel is divided into a first channel group and a second channel group in a clockwise direction; If the route coordinates of the source node are greater than the route coordinates of the target node, the first channel group is used; If the routing coordinates of the source node are less than or equal to the routing coordinates of the target node, the second channel group is used.
9. The controller according to claim 1, characterized in that, When the accelerated computing unit group is in the first operating mode, the accelerated computing units within the group can be configured to different computing functions.
10. The controller according to claim 1, characterized in that, When the accelerated computing unit group is in the second operating mode, the accelerated computing units within the group are configured to perform the same computing function so that if any accelerated computing unit within the group fails, it can be replaced by another accelerated computing unit within the group.
11. The controller according to claim 1, characterized in that, The controller is one of a Field Programmable Gate Array (FPGA), a Reduced Instruction Set Computer (ARM), and an Enhanced Reduced Instruction Set Controller (AVR).
12. An accelerated computing system, characterized in that, The accelerated computing system includes: Host; Multiple controller clusters are configured to execute accelerated computing tasks issued by the host, and each controller cluster includes multiple controllers as described in any one of claims 1 to 11; The switching chip, connected to the host and the controller cluster via a bus, is configured to enable data interaction between the host and the controller cluster.
13. The accelerated computing system according to claim 12, characterized in that, The controllers in the controller cluster are configured as a fully connected architecture.
14. The accelerated computing system according to claim 12, characterized in that, The switching chip supports at least the CXL communication protocol.
15. The accelerated computing system according to claim 12, characterized in that, The controller cluster also includes a multiplexer configured to receive signals generated within the controller and transmit them to other controllers or the host.
16. A routing method within a controller, characterized in that, include: Configure the routing coordinates of the accelerated computing units inside the controller, wherein the routing coordinates include the controller identifier ID, the accelerated computing unit group ID, and the accelerated computing unit ID; The routing coordinates of the source node and the routing coordinates of the target node are determined based on the routing jump request, and the number of intra-chip jump steps is calculated based on the routing coordinates of the source node and the routing coordinates of the target node. The routing direction is determined based on the number of intra-chip jump steps, and the routing is performed based on the routing direction.
17. The method according to claim 16, characterized in that, Based on the stated routing direction, routing is performed, including: Obtain the output buffer length of the controller port; If the output buffer length is greater than the fifth threshold, reverse routing is performed.
18. The method according to claim 16, characterized in that, Determining the routing direction based on the number of intra-chip jump steps includes: Based on the preset range of the number of jump steps within the chip, the routing jump direction is determined. The routing jump direction includes clockwise direction, counterclockwise direction, opposite direction, or a combination of clockwise direction, counterclockwise direction, and opposite direction.
19. The method according to claim 18, characterized in that, The controller includes multiple accelerated computing unit groups, wherein the accelerated computing units within the same accelerated computing unit group are sequentially connected via an internal bus and connected to the same shared memory via the internal bus. Accelerated computing units in adjacent groups of different accelerated computing unit groups communicate with each other, as do those in opposite groups. The routing jump direction is determined based on a preset range of the in-chip jump steps. The method includes: If the number of intra-chip jump steps is equal to the first threshold, the clockwise direction will be determined as the routing jump direction; If the number of jump steps within the slice is equal to the second threshold, the opposite direction is determined as the route jump direction; If the number of intra-chip jump steps is equal to the third threshold, the counterclockwise direction will be determined as the routing jump direction; If the number of intra-chip jump steps is less than the first threshold, the counterclockwise direction will be determined as the routing jump direction; If the number of jump steps within the slice is greater than the first threshold and less than the second threshold, the route jump direction will be determined as first facing the opposite direction and then clockwise. If the number of jump steps within the slice is greater than the second threshold and less than the third threshold, the route jump direction will be determined as first facing the opposite direction and then counterclockwise. If the number of intra-chip jump steps is greater than the third threshold and less than the fourth threshold, the clockwise direction will be determined as the routing jump direction.
20. A routing hopping method between controllers, characterized in that, The controller includes an on-chip routing module and an inter-chip routing module. The on-chip routing module is communicatively connected to each accelerated computing unit group, and the inter-chip routing module is communicatively connected to the on-chip routing module. The method includes: Configure the routing coordinates of the accelerated computing units inside the controller, wherein the routing coordinates include the controller ID, the accelerated computing unit group ID, and the accelerated computing unit ID; The route coordinates of the source node and the route coordinates of the target node are determined based on the route redirection request; Based on the routing coordinates of the source node, the system jumps from the intra-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the source node; based on the routing coordinates of the target node, the system jumps from the inter-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the target node; and based on the routing coordinates of the target node, the system jumps from the inter-chip routing module corresponding to the target node to the intra-chip routing module corresponding to the target node, and then transmits the data to the target node.
21. A method for scheduling accelerated computing units for accelerated computing tasks, characterized in that, The accelerated computing unit scheduling method is applied to the accelerated computing system of claim 12, and the method includes: Obtain the accelerated computing unit status table and the memory access latency table. The computing unit status table is used to include at least the routing coordinates, unit functions, unit status and storage usage of the accelerated computing unit. The memory access latency table includes the latency of each accelerated computing unit accessing each shared cache. Obtain the computing tasks to be assigned, and determine the scheduling strategy based on the computing unit status table and the memory access latency table. In the process of determining the scheduling strategy, the accelerated computing units with different storage occupancy can run in parallel. In the scheduling strategy, each accelerated computing unit corresponds to the computing task to be assigned with the minimum latency.