A controller for accelerating calculations and an accelerated computing system
By adopting an on-chip interconnect structure based on packet bus and a virtual channel control unit in the accelerated computing system, the delay problem caused by on-chip interconnection technology is solved, and the computing efficiency and system performance of the accelerated computing task are improved.
Patent Information
- Application Number
- CN202510101194.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In the existing accelerated computing systems, due to the constraints of the network topology, the on-chip interconnection technology has a large delay in part of the transmission process, which in turn reduces the computing efficiency of the accelerated computing task.
Adopting an on-chip interconnection improved structure based on packet bus, the acceleration computing units inside the controller are grouped, and each group shares memory. Through the internal bus connection, the acceleration computing units are connected by diagonally to reduce the number of jump steps, and the virtual channel control unit prevents the channel closed loop from causing deadlock.
It significantly reduces the delay of data transmission, improves the efficiency of data exchange between accelerated computing units, reduces communication delay, and improves the overall performance of the accelerated computing system.
Smart Images

Figure CN119537294B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computers, and more specifically, to a controller for accelerating calculations, an accelerated computing system, a routing jump method inside the controller, a routing jump method between controllers, and an accelerated computing unit scheduling method for accelerated computing tasks. Background Art
[0002] With the continuous development of technologies such as artificial intelligence, big data analysis, and cloud computing, the amount of system data has shown an exponential growth trend, which has led to an increasing demand for data processing and analysis. In related technologies, traditional single-architecture database systems face huge challenges, and their capabilities in processing massive data, ensuring data real-time performance, and coping with diverse workloads are relatively insufficient.
[0003] Heterogeneous acceleration architectures usually combine controllers with other processors to achieve high-performance computing by leveraging the advantages of different processors. In a heterogeneous acceleration architecture, the interconnection and communication of the accelerated computing units inside the controller have a key impact on the overall performance of the system. By improving the internal interconnection structure of the controller, the data transmission path can be made shorter and the bandwidth can be higher, which can effectively reduce the system transmission delay. Enabling data to be scheduled and transmitted orderly and efficiently among the accelerated computing units, giving full play to the advantages of each accelerated computing unit, and enhancing the overall performance of the system.
[0004] In related technologies, bus-based interconnection and on-chip interconnection technologies are generally used to achieve the interconnection of accelerated computing units. However, bus-based interconnection has problems such as insufficient bandwidth allocation and increased latency due to bus contention when multiple computing units simultaneously request access to memory. Therefore, this shortcoming is inevitable in large-scale systems. On the other hand, on-chip interconnection technology depends on network topology and link length limitations. In some transmission processes, there are often many skipped steps, resulting in obvious latency. Moreover, on-chip interconnection technology requires a large number of routers and links to build a network in terms of hardware, leading to an increase in chip area and power consumption. Also, when transmitting across devices, the link width has a great impact on the transmission rate, which will affect the overall performance of the accelerated computing system.
[0005] In summary, in related technologies, the accelerated computing system has the problem that due to the restriction of the on-chip interconnection technology by the network topology structure, there is a large delay in some transmission processes, resulting in a decrease in the computing efficiency of the accelerated computing tasks. Summary of the Invention
[0006] Embodiments of the present invention provide a controller for accelerated computing, an accelerated computing system based on the controller, a routing jump method inside the controller, a routing jump method between controllers, and an accelerated computing unit scheduling method for accelerated computing tasks, so as to at least solve the problem in the related art that due to the restriction of the on-chip interconnection technology by the network topology, there is a large delay in some transmission processes, resulting in a decrease in the computing efficiency of the accelerated computing tasks.
[0007] According to an embodiment of the present invention, a controller for accelerated computing is provided, including: a plurality of accelerated computing unit groups, which are pairwise communicatively connected between adjacent accelerated computing unit groups and are arranged in a completely symmetric ring shape. The accelerated computing units within the same accelerated computing unit group are sequentially connected through an internal bus, and the accelerated computing units located at symmetric positions on the same diameter of the ring between different accelerated computing unit groups are communicatively connected;
[0008] Each accelerated computing unit group further includes a shared memory, and the shared memory is connected to each accelerated computing unit through an internal bus.
[0009] In an exemplary embodiment, the controller further includes: a device-side controller, configured to receive a CXL-format data packet, parse the data packet to obtain an accelerated computing task, and send it to the accelerated computing unit; a switching module, connected to the device-side controller through a CXL bus and communicatively connected to the accelerated computing unit through an internal bus, configured to receive the accelerated computing task through direct memory access and send it to the accelerated computing unit.
[0010] In an exemplary embodiment, the switching module includes a multi-level switching module, where the switching module includes a switching module connected to the device-side controller through a CXL bus and connected to other switching modules through interfaces with multiple output signals in the master mode, and a switching module connected to other switching modules, connected to the shared memory through an interface with an output signal in the slave mode, and connected to the accelerated computing unit through an interface with an output signal in the master mode.
[0011] In an exemplary embodiment, the controller further includes: an on-chip routing module, communicatively connected to each accelerated computing unit group respectively, configured to implement routing jumps inside the same controller; an inter-chip routing module, communicatively connected to the on-chip routing module, configured to implement routing jumps between different controllers.
[0012] In an exemplary embodiment, the on-chip routing module includes a first crossbar switch and a first crossbar switch control unit, where the first crossbar switch is used to connect each accelerated computing unit group and the inter-chip routing module, and the first crossbar switch control unit is used to control the on-off of the paths connected by the first crossbar switch.
[0013] In an exemplary embodiment, the inter-chip routing module includes a second crossbar switch and a second crossbar switch control unit. Among them, the second crossbar switch is used to connect the intra-chip routing module and the multiplexer selector, and the second crossbar switch control unit is used to control the on / off of the connected path of the second crossbar switch.
[0014] In an exemplary embodiment, the intra-chip routing module or the inter-chip routing module includes a virtual channel control unit, which is used to divide a physical channel into multiple virtual channels to prevent deadlock caused by channel closed-loop.
[0015] In an exemplary embodiment, the virtual channel control unit is specifically used to construct multiple independent micro-slice buffers in the output queue buffer corresponding to the physical channel to obtain multiple virtual channels; equally divide the virtual channels into a first channel group and a second channel group in the clockwise direction; use the first channel group when the routing coordinate of the source node is greater than the routing coordinate of the target node; use the second channel group when the routing coordinate of the source node is less than or equal to the routing coordinate of the target node.
[0016] In an exemplary embodiment, when the acceleration computing unit group is in the first working mode, the acceleration computing units within the group are allowed to be configured with different computing functions.
[0017] In an exemplary embodiment, when the acceleration computing unit group is in the second working mode, the acceleration computing units within the group are configured with the same computing function, so that in the case of any acceleration computing unit failure within the group, other acceleration computing units within the group can be used for replacement.
[0018] In an exemplary embodiment, in one implementation manner of the present invention, the controller is one of FPGA, ARM, and AVR.
[0019] According to another embodiment of the present invention, an acceleration computing system is provided. The acceleration computing system includes: a host; a plurality of controller clusters for executing the acceleration computing tasks issued by the host, and each controller cluster includes a plurality of controllers; a switching chip connected to the host and the controller clusters through a bus for realizing data interaction between the host and the controller clusters.
[0020] In an exemplary embodiment, the controllers in the controller cluster are configured in a fully connected structure.
[0021] In an exemplary embodiment, the switching chip at least supports the CXL communication protocol.
[0022] In an exemplary embodiment, the controller cluster further includes a multiplexer selector for receiving the signals generated inside the controller and transmitting them to other controllers or the host.
[0023] According to another embodiment of the present invention, a routing jump method inside a controller is provided, including: configuring routing coordinates of an acceleration computing unit inside the controller, where the routing coordinates include a controller ID, an acceleration computing unit group ID, and an acceleration computing unit ID; determining the routing coordinates of a source node and the routing coordinates of a target node based on a routing jump request, and calculating the number of on-chip jump steps according to the routing coordinates of the source node and the routing coordinates of the target node; determining a routing jump direction based on the number of on-chip jump steps, and performing a routing jump based on the routing jump direction.
[0024] In an exemplary embodiment, performing a routing jump based on the routing jump direction includes: obtaining the output buffer length of a controller port; and performing a reverse route when the output buffer length is greater than a fifth threshold.
[0025] In an exemplary embodiment, determining the routing jump direction based on the number of on-chip jump steps includes: determining the routing jump direction based on a preset range where the number of on-chip jump steps is located, and the routing jump direction includes a clockwise direction, a counterclockwise direction, a face-to-face direction, or a combination of the clockwise direction, the counterclockwise direction, and the face-to-face direction.
[0026] In an exemplary embodiment, the controller includes multiple acceleration computing unit groups. Among them, the acceleration computing units within the same acceleration computing unit group are sequentially connected through an internal bus and are connected to the same shared memory through the internal bus. The acceleration computing units between different acceleration computing unit groups are communicatively connected to the acceleration computing units in another adjacent group and the acceleration computing units in another group located opposite. Determining the routing jump direction based on the preset range where the number of on-chip jump steps is located, the method includes: when the number of on-chip jump steps is equal to a first threshold, determining the clockwise direction as the routing jump direction; when the number of on-chip jump steps is equal to a second threshold, determining the face-to-face direction as the routing jump direction; when the number of on-chip jump steps is equal to a third threshold, determining the counterclockwise direction as the routing jump direction; when the number of on-chip jump steps is less than the first threshold, determining the counterclockwise direction as the routing jump direction; when the number of on-chip jump steps is greater than the first threshold and less than the second threshold, determining the face-to-face direction first and then the clockwise direction as the routing jump direction; when the number of on-chip jump steps is greater than the second threshold and less than the third threshold, determining the face-to-face direction first and then the counterclockwise direction as the routing jump direction; when the number of on-chip jump steps is greater than the third threshold and less than a fourth threshold, determining the clockwise direction as the routing jump direction.
[0027] According to another embodiment of the present invention, a routing jump method between controllers is provided.
[0028] The controller includes an on-chip routing module and an inter-chip routing module. Among them, the on-chip routing module is communicatively connected to each acceleration computing unit group respectively, and the inter-chip routing module is communicatively connected to the on-chip routing module. The method includes: configuring the routing coordinates of the acceleration computing units inside the controller, where the routing coordinates include the controller ID, the acceleration computing unit group ID, and the acceleration computing unit ID; determining the routing coordinates of the source node and the routing coordinates of the target node based on a routing jump request; jumping from the on-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the source node based on the routing coordinates of the source node, jumping from the inter-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the target node according to the routing coordinates of the target node, and jumping from the inter-chip routing module corresponding to the target node to the on-chip routing module corresponding to the target node according to the routing coordinates of the target node, and transmitting to the target node.
[0029] According to another embodiment of the present invention, there is provided a scheduling method for acceleration computing units of an acceleration computing task. The acceleration computing unit scheduling method is applied to an acceleration computing system, and the method includes: obtaining an acceleration computing unit status table and a memory access latency table. The computing unit status table is used to at least include the routing coordinates, unit function, unit status, and storage occupancy of the acceleration computing units, and the memory access latency table includes the latency of each acceleration computing unit accessing each shared cache; obtaining a computing task to be allocated, and determining a scheduling policy according to the computing unit status table and the memory access latency table. Among them, during the process of determining the scheduling policy, acceleration computing units with different storage occupancies can run in parallel, and each acceleration computing unit in the scheduling policy corresponds to the computing task to be allocated with the minimum latency.
[0030] With the present invention, since the controller of the present invention adopts an improved on-chip interconnection structure based on a grouped bus, the acceleration computing units inside the controller are grouped, and each group shares memory. When performing routing jumps within the group, it can be directly completed through the shared memory. The acceleration computing units of different groups are diagonally connected through on-chip interconnection technology, and the groups are connected in sequence. There is no need to jump one by one in order, reducing the number of hops in the jump process, and the shared memory reduces the global bus contention. Therefore, it can solve the problem in the related art that due to the on-chip interconnection technology being restricted by the network topology structure, there is a large delay in some transmission processes, resulting in a decrease in the computing efficiency of the acceleration computing task, and achieve the effects of improving the communication efficiency between the inside of the controller and between controllers, reducing the communication latency, and enhancing the overall performance of the acceleration computing system. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic diagram of the topology structure of a controller for acceleration computing according to an embodiment of the present invention;
[0032] Figure 2 is a schematic diagram of the connection structure of a controller for acceleration computing according to an embodiment of the present invention;
[0033] Figure 3 It is a schematic structural diagram of an on-chip routing module according to an embodiment of the present invention;
[0034] Figure 4 It is a schematic structural diagram of an inter-chip routing module according to an embodiment of the present invention;
[0035] Figure 5 It is a schematic logic diagram of a crossbar switch according to an embodiment of the present invention;
[0036] Figure 6 It is a schematic structural diagram of an acceleration computing system based on a controller according to an embodiment of the present invention;
[0037] Figure 7 It is a schematic diagram of a fully connected structure of a controller cluster according to an embodiment of the present invention;
[0038] Figure 8 It is a schematic flow diagram of a routing jump method inside a controller according to an embodiment of the present invention;
[0039] Figure 9 It is a schematic flow diagram of a routing jump method between controllers according to an embodiment of the present invention;
[0040] Figure 10 It is a schematic flow diagram of a scheduling method for an acceleration computing unit of an acceleration computing task according to an embodiment of the present invention. Detailed implementation manners
[0041] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments.
[0042] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.
[0043] For the convenience of description, some nouns or terms related to the embodiments of the present invention will be described below:
[0044] Field Programmable Gate Array: abbreviated as FPGA, is a programmable semiconductor device that allows users to configure its hardware according to needs. The FPGA contains programmable logic blocks and configurable interconnections, and users can implement different hardware logics by loading different configuration files.
[0045] Direct Memory Access technology: Direct Memory Access, abbreviated as DMA, is a technology that allows a hardware subsystem to directly exchange data with system memory without the direct control of the CPU. The DMA controller can directly transfer data from an input device to memory or from memory to an output device without the intervention of the CPU, thereby reducing the burden on the CPU and improving data transfer efficiency.
[0046] Peripheral Component Interconnect Express: Peripheral Component Interconnect Express, abbreviated as PCIe, is a general-purpose serial expansion bus standard used for the connection between internal hardware components of a computer. It is widely used in high-speed data communication and peripheral device connection due to its high data transfer rate and flexibility.
[0047] Compute Express Link: Compute Express Link, abbreviated as CXL, is an open high-speed interconnect technology standard designed to provide higher data throughput and lower communication latency for modern data centers, high-performance computing systems, and other compute-intensive applications.
[0048] In some solutions, the controller using on-chip interconnect technology is restricted by the network topology, resulting in a large delay in some transmission processes, which leads to a decrease in the computing efficiency of the accelerated computing task. To solve the above technical problems, in this embodiment, a controller for accelerated computing, an accelerated computing system, a routing jump method inside the controller, a routing jump method between controllers, and an accelerated computing unit scheduling method for accelerated computing tasks are provided.
[0049] In the embodiment of the present invention, a controller for accelerated computing is provided, and its topology diagram is as Figure 1 shown. The controller includes: a plurality of accelerated computing unit groups, which are pairwise communicatively connected between adjacent accelerated computing unit groups and are arranged in a completely symmetric ring. The accelerated computing units within the same accelerated computing unit group are sequentially connected through an internal bus, and the accelerated computing units located at symmetric positions on the same diameter of the ring between different accelerated computing unit groups are communicatively connected; each accelerated computing unit group further includes a shared memory, and the shared memory is connected to each accelerated computing unit through an internal bus.
[0050] In an acceleration computing system in the related art, in an acceleration computing unit in a controller for acceleration computing, the acceleration computing unit and the memory usually adopt a single bus or a complex on-chip interconnection network. The above interconnection methods have problems of increased latency and bandwidth bottlenecks due to bus contention and excessive skipping during multi-unit parallel computing and data access. Taking an FPGA as an example, if multiple acceleration computing units access a DDR / HBM memory simultaneously, the traditional bus interconnection will cause uneven bandwidth allocation and data transmission latency due to skipping between acceleration computing units.
[0051] In the above embodiments of the present invention, as Figure 1 shown, the overall topology of the controller is circular and completely symmetric. Inside each acceleration computing unit group, the acceleration computing units are connected through an internal bus and share the same memory resource to ensure efficient data transfer and access within the group. The acceleration computing units between different groups are connected in a specific manner. Specifically, adjacent acceleration computing unit groups are connected in sequence, and each acceleration computing unit is communicatively connected to the acceleration computing unit located opposite to minimize the skipping and latency of cross-group communication and optimize the overall interconnection performance.
[0052] In a specific embodiment, as Figure 1 shown, taking an FPGA chip as an example, in the topology, in the clockwise direction, every four nodes are divided into a group (group 0, group 1, group 2, and group 3). Taking group 0 as an example, the acceleration computing units 0, 1, 2, and 3 inside group 0 are all connected to the same shared memory DDR / HSM_0 through an internal bus. The above structural design enables the computing units within the group to efficiently access the local memory. It can be understood that with the above structure, the acceleration computing unit 0 and the acceleration computing unit 3 can jump through the shared memory DDR / HSM_0, which reduces the access latency caused by intra-group skipping compared with the prior art where the acceleration computing unit 0 jumps to the acceleration computing unit 3 one by one.
[0053] In a specific embodiment, as Figure 1 shown, taking an FPGA chip as an example, in the topology, taking group 0 and group 1 as an example, the adjacent acceleration computing units 3 and 4 are directly connected for communication. Taking group 0 and group 2 as an example, the acceleration computing units 2 and 10 located opposite are directly connected for communication. Through the above structure, the acceleration computing unit 0 and the acceleration computing unit 7 can jump through the path 0-8-7, and the acceleration computing unit 1 and the acceleration computing unit 5 can jump through the path 1-3-4-5, which greatly reduces the access latency caused by intra-group skipping compared with routing and jumping in the clockwise or counterclockwise direction in the prior art.
[0054] In summary, by setting specific intra-group and cross-group communication paths, the present invention can significantly reduce the latency of data transmission in the data center and improve the efficiency of data exchange between acceleration computing units. Moreover, by using an internal bus to implement shared memory within the group, it ensures that the acceleration computing units can efficiently utilize local resources and avoids the decline in resource utilization efficiency caused by insufficient bandwidth allocation. Additionally, grouping the acceleration computing units can better support large-scale parallel computing. Furthermore, setting a ring-shaped and completely symmetric topology structure makes the routing rules and effects of the acceleration computing units exactly the same regardless of their positions in the above architecture, ensuring that performance deviations caused by the distribution positions of acceleration computing units with different functions do not need to be considered during the involved process. And the above topology structure has a low node degree and network diameter, which is beneficial to optimizing the transmission delay. Using the above controller for acceleration computing can provide more powerful heterogeneous computing capabilities for fields such as database acceleration and deep learning.
[0055] In an embodiment of the present invention, the interconnection logic block diagram of the above controller is as Figure 2 shown. The above controller includes: a device-side controller, which is used to receive CXL-format data packets, parse the data packets to obtain acceleration computing tasks, and send them to the acceleration computing units.
[0056] In the related art, heterogeneous computing systems are connected between components through the CXL bus. It can be understood that as an open high-speed interconnection technology standard, CXL is widely used to connect devices such as CPUs, GPUs, accelerators, and memories to achieve low-latency and high-bandwidth data transmission. Therefore, to ensure data transmission in the CXL format, the present invention configures a device-side controller in the controller to receive and parse CXL-format data packets. As Figure 2 shown, the device-side controller within the chip (CXL device-side controller) is connected to the host through the CXL bus in the controller architecture of the present invention and is responsible for receiving data packets and task parsing through the CXL bus. Specifically, when the host sends an acceleration computing task in the form of a CXL-format data packet through the CXL bus, the device-side controller receives the above data packet, parses it, extracts the specific acceleration computing tasks and related parameters required for acceleration computing, and then sends the acceleration computing tasks to the corresponding specific acceleration computing units.
[0057] In a specific embodiment, it is assumed that the host needs to execute an accelerated computing task for database query, and this task includes steps such as data packet decompression, full table scan, and expression filtering. During the distribution process of the accelerated computing task, the host encapsulates these tasks into data packets in CXL format and distributes them to the device-side controller through the CXL bus. After receiving the data packets, the device-side controller parses them according to the CXL protocol and extracts information such as task type, data element location, and target memory address. Then, the device-side controller selects an accelerated computing unit with corresponding functions according to the operating state of the accelerated computing unit and distributes the relevant information to this accelerated computing unit through the internal bus.
[0058] Through the direct connection of the above-mentioned device-side controller via the CXL bus and the internal bus, the relevant data and control information of the accelerated computing task can be received and processed by the accelerated computing unit with the minimum delay. And the device-side controller can dynamically schedule tasks according to the functions and states of the computing units, thereby realizing the parallel execution of tasks and the optimal allocation of the resources of the accelerated computing units, ensuring the minimization of the delay of the accelerated computing task.
[0059] The switching module is connected to the device-side controller via the CXL bus and communicates with the accelerated computing unit via the internal bus, and is used to receive the accelerated computing task in the way of direct memory access and send it to the accelerated computing unit.
[0060] The function of the switching module is between the device-side controller and the accelerated computing unit, playing the role of a bridge. In the present invention, the device-side controller and the switching module are connected via the CXL bus to enable the device-side controller to transmit data to the switching module in the form of direct memory access (DMA) without the access of the CPU, thereby greatly improving the data transmission efficiency, reducing the computing delay, and ensuring high-speed data transmission and protocol compatibility; the accelerated computing unit and the switching module communicate via the internal bus, realizing the efficient scheduling of the accelerated computing task among the accelerated computing units.
[0061] The introduction of the DMA technology between the switching module and the device-side controller ensures the high-speed and low-power consumption transmission of data between the device-side controller and the accelerated computing unit, improving the overall computing efficiency.
[0062] In summary, the combined use of the device-side controller and the switching module realizes the efficient reception, parsing, and execution of the accelerated computing task in the controller accelerated computing system, as well as the fast and low-latency transmission of data between the accelerated computing unit and the memory, significantly improving the performance and response speed of the controller in the database accelerated computing task.
[0063] In the embodiment of the present invention, as Figure 2As shown, the switching module includes a multi-level switching module. Among them, there is a switching module connected to the device-side controller via the CXL bus and connected to other switching modules through interfaces with multiple output signals in the master mode, and a switching module connected to other switching modules, connected to the shared memory through an interface with an output signal in the slave mode, and connected to the acceleration computing unit through an interface with an output signal in the master mode.
[0064] To optimize the data transmission path and reduce the number of ports, the present invention proposes an interconnection mechanism based on a multi-level switching module, in which master mode and slave mode interfaces are used to distinguish links with different functions. The multi-level switching module includes at least two levels of switching modules, namely, a top-level switching module and a bottom-level switching module. Among them, the top-level switching module is directly connected to the device-side controller and receives CXL format data packets from the device-side controller via the CXL bus. The output signal interfaces of the top-level switching module are set in the master mode, and it is connected to the intra-group switching modules at the next level through these interfaces to distribute the data packets to different groups. The bottom-level switching module is respectively connected to the shared memory (slave mode interface) and the acceleration computing units within the acceleration computing unit group (master mode interface), and is responsible for sending them to the specified acceleration computing unit according to the task requirements.
[0065] In a specific embodiment, as Figure 2 shown, in a 16-node controller, the acceleration computing units are divided into four groups, with 4 acceleration computing units in each group, and each group includes a shared memory. The switching module only includes two layers, including the above-mentioned top-level switching module (Crossebar) and bottom-level switching modules (Crossebar_0 to Crossebar_3). Each bottom-level switching module is connected to a group of acceleration computing units, and all bottom-level switching modules are directly connected to the top-level switching module. In the specific working process, the host sends acceleration computing tasks to the device-side controller via the CXL bus. These tasks may include data reading, page parsing, expression filtering, multi-table joining, etc. After parsing the received CXL data packets, the device-side controller decomposes and forwards the tasks to the top-level switching module via the CXL bus using DMA technology. After receiving the instructions, the top-level switching module distributes the data packets to the corresponding intra-group switching modules through the master mode interfaces according to the task requirements and the distribution of the computing units. After receiving the task data assigned to its group, the intra-group switching module exchanges data with the shared memory using the slave mode interface and loads the required data into the memory. At the same time, the master mode interface of the intra-group switching module sends the data packets to the acceleration computing units within the group to execute specific tasks. After completing the tasks, the acceleration computing units store the results in the shared memory, and then the intra-group switching module sends the result data to the top-level switching module through the slave mode interface, or directly transmits it inside the chip. Finally, the data can be accessed by other computing units or the host.
[0066] By setting up multi-level switching and DMA data transmission, the controller of the present invention avoids the bottleneck problem in bus interconnection in the related art, improves the output transmission efficiency. In the multi-level switching mechanism, the underlying switching module is grouped and connected to the acceleration computing units, facilitating the transmission of data packets along the optimal path. Meanwhile, the design of the master mode and slave mode interfaces simplifies the data scheduling between the computing units and the shared memory, realizing the efficient allocation of computing resources.
[0067] In the embodiment of the present invention, as Figure 3 shown, the controller further includes: an on-chip routing module, which is communicatively connected to each acceleration computing unit group respectively, and is used to realize the routing jump within the same controller.
[0068] In the related art, for a heterogeneous acceleration system based on a controller to achieve high-performance and low-latency data transmission, the routing module plays an indispensable role. It is responsible for correctly transmitting data packets from the source node to the target node, which is very important in a complex system architecture during this process. To ensure the correct transmission of data packets within the chip, the present invention sets that the on-chip routing module is responsible for the routing jump of data packets between the computing unit groups within a single controller. The on-chip routing module is tightly connected to the acceleration computing units and switching modules of each group, and realizes the fast transmission of data through a low-latency interconnection structure. When a computing unit needs to access the memory of another group or exchange data with a computing unit of another group, the on-chip routing module efficiently completes the jump routing of data packets according to the deterministic routing algorithm, ensuring the shortest data transmission path and the lowest latency.
[0069] In a specific embodiment, as Figure 3 shown, in a 16-node controller architecture, the on-chip routing module is designed as a router with 5 ports, specifically including a clockwise port, a counterclockwise port, an opposite port, a local port (connected to the local acceleration computing unit), and an inter-chip port for connecting to the inter-chip routing module. When a data packet needs to be transmitted between different groups inside the chip, the on-chip routing module first parses the routing ID of the data packet to determine the position of the target node, and then transmits the data through the most suitable port (such as the clockwise or counterclockwise port) to achieve the jump of the shortest path. For example, if the destination of the data packet is another group located in the clockwise direction, the on-chip routing module will perform routing jump through the clockwise port, which can avoid unnecessary detours and improve the data transmission efficiency.
[0070] The involvement of the on-chip routing module ensures that data packets are transmitted within the controller architecture with the shortest path and the lowest latency, improving the overall response speed of the system. The on-chip routing module is connected to the acceleration computing units in groups. During the jump process, it can optimize the data transmission path through deterministic and adaptive routing algorithms, avoid resource waste and path bottlenecks, and ensure that the computing units can efficiently utilize local and remote resources.
[0071] In an embodiment of the present invention, as Figure 4 shown, the controller further includes: an inter-chip routing module, which is communicatively connected to the on-chip routing module and is used to implement routing jumps between different controllers.
[0072] The inter-chip routing module is mainly used to handle the output transmission between different controllers. In complex acceleration computing tasks, multiple controllers may be configured to work together. At this time, the inter-chip routing module serves as a communication bridge between different controllers, responsible for cross-chip data packet routing jumps, and realizing efficient data transmission and routing selection.
[0073] In a specific embodiment, as Figure 4 shown, in a 16-node controller architecture, the inter-chip routing module is designed as a router with 5 ports. Four of the ports are respectively connected to different on-chip routing modules inside the controller, and one port is connected to the inter-chip routing module inside other controllers. The inter-chip routing module is used to receive the data packets sent by the on-chip routing module, decode them according to the controller ID and the acceleration computing unit group ID in the routing ID, and select the corresponding controller node for data transmission. For example, if the source node of the data packet is in group 1 of controller 0 and the destination node is in group 3 of controller 2, then the inter-chip routing module 0 first sends the data packet to the inter-chip routing module 2 of controller 2 through the inter-chip full-connection port, and then the on-chip routing module of controller 2 performs the final routing jump of the data packet according to the group ID and the intra-group ID of the target node, ensuring that the data directly reaches the destination.
[0074] The involvement of the on-chip routing module ensures that data packets are transmitted between controller architectures with the shortest path and the lowest latency, improving the overall response speed of the system. The inter-chip routing module is connected to the on-chip routing module in groups. During the jump process, it can optimize the data transmission path through deterministic and adaptive routing algorithms, avoid resource waste and path bottlenecks, and ensure that the computing units can efficiently utilize local and remote resources. Moreover, the existence of the inter-chip routing module supports efficient communication between multiple controllers, provides the possibility for the expansion of heterogeneous acceleration systems, enables the system to support a larger number of acceleration computing units, and meets more complex and larger-scale acceleration computing requirements.
[0075] In an embodiment of the present invention, the on-chip routing module includes a first crossbar switch and a first crossbar switch control unit. Among them, the first crossbar switch is used to connect each acceleration computing unit group and the inter-chip routing module, and the first crossbar switch control unit is used to control the on-off of the connection paths connected by the first crossbar switch.
[0076] In the acceleration computing system of the related art, the on-chip communication efficiency has a key impact on the overall performance of the acceleration computing system. The crossbar switch is a key component for data transmission routing selection, which can effectively connect different nodes and achieve fast forwarding of data packets. In the present invention, the first crossbar switch is set to be responsible for establishing connections between different acceleration computing unit groups inside the chip and the routing module between chips, while the first crossbar switch control unit ensures the correctness and efficiency of these connections.
[0077] In a specific embodiment, as Figure 5 shown, the first crossbar switch in the on-chip routing module is designed as a 5×5 switching matrix, which can be connected to the clockwise direction, counterclockwise direction, opposite direction, local acceleration computing unit group, and the inter-chip routing module. This design allows data packets to flow freely in different directions. For example, in a 16-node controller, four acceleration computing unit groups are connected through the first crossbar switch, and at the same time, the inter-chip routing module is also connected to these groups through the first crossbar switch. The first crossbar switch control unit is responsible for parsing the routing information of the data packet, and according to this information, controls the on-off of the connection paths of the first crossbar switch to achieve precise routing of data. Specifically, when a data packet needs to be transmitted from a certain acceleration computing unit group to another group or the inter-chip routing module, the first crossbar switch control unit will decide whether the data packet should be transmitted through the clockwise port, counterclockwise port, opposite direction port, or local port according to the routing ID of the data packet and the current link load situation. If there is link congestion, the control unit will automatically select a path with lighter load to achieve adaptive load balancing. For example, in Figure 4 the logic diagram, the first crossbar switch control unit needs to monitor the load conditions of all inputs [0-4] and outputs [0-4]. When the data packet enters from input [0] (clockwise direction), if the buffer of output [1] (counterclockwise direction) is not full, the first crossbar switch is controlled to enable the data packet to flow directly from input [0] to output [1], thereby completing the data transmission in the clockwise direction.
[0078] Through the on-chip routing module including the first crossbar switch and the first crossbar switch control unit, the present invention realizes efficient data transmission and adaptive routing selection, not only optimizing the communication between computing units, but also improving the interaction efficiency with the inter-chip routing module, providing strong hardware support for high data-intensive tasks such as database acceleration computing and deep learning.
[0079] In an embodiment of the present invention, the inter-chip routing module includes a second crossbar switch and a second crossbar switch control unit. Among them, the second crossbar switch is used to connect the intra-chip routing module to other controllers, and the second crossbar switch control unit is used to control the on / off of the paths connected by the second crossbar switch.
[0080] The second crossbar switch is the same high-performance switching matrix as the first crossbar switch. In the inter-chip routing module, the second crossbar switch is configured to connect the intra-chip routing module and the multiplexer selector, so as to achieve cross-chip data packet routing. As Figure 4 shown, it contains 5 ports, which respectively correspond to: the inter-chip port of the FPGA intra-chip routing module and the input end of the multiplexer selector. This design allows data packets to flow freely between different FPGA chips. The second crossbar switch control unit is responsible for parsing the routing information of the data packets, and controlling the path selection of the second crossbar switch according to this information to achieve the fast transmission of data packets from the source chip to the target chip. That is, under the deterministic routing strategy, the control unit needs to determine the transmission path of the data packets according to the controller ID, the acceleration computing unit group ID, and the acceleration computing unit ID, and ensure that other paths can be intelligently selected when link congestion occurs. Specifically, when a data packet enters the second crossbar switch from the inter-chip port of the intra-chip routing module, the control unit will check the ID information of the controller corresponding to the target node and the current network state. If the load of the direction link corresponding to the target controller is low, the control unit will direct the data packet to be transmitted through the output port in this direction.
[0081] Through the inter-chip routing module including the second crossbar switch and the second crossbar switch control unit, the present invention ensures that the path selection for cross-chip data transmission is both fast and efficient, significantly reducing the delay of data packets during inter-chip transmission. When link failures or congestion occur, the second crossbar switch control unit can intelligently select alternative paths, improving the overall robustness and reliability of the network, ensuring that data transmission can still proceed even under adverse network conditions, realizing efficient cross-chip data transmission and adaptive routing selection. It not only optimizes inter-chip communication, improves data transmission speed, but also enhances the network robustness and reliability of the acceleration computing system when processing tasks such as database acceleration computing and large-scale data processing.
[0082] In an embodiment of the present invention, as Figure 5 shown, the intra-chip routing module or the inter-chip routing module includes a virtual channel control unit, which is used to divide a physical channel into multiple virtual channels to prevent the occurrence of deadlock due to channel closed-loop.
[0083] In the acceleration computing system of the related art, the complexity and high concurrency of data communication can lead to network congestion and deadlocks. Especially in the interconnection structures based on buses and ring networks, the present invention sets up a virtual channel technology to divide physical channels into multiple logically independent virtual channels to avoid deadlocks and improve network efficiency.
[0084] In a specific embodiment, as Figure 5 shown, both the on-chip routing module and the inter-chip routing module integrate a virtual channel control unit, which divides each physical channel into four independent virtual channels (VCs), namely VC0, VC1, VC2, and VC3. In this way, when data packets are transmitted in different directions, different virtual channels can be used, thus avoiding the common channel closed-loop problem in the ring network and effectively preventing the occurrence of deadlocks. Specifically, the virtual channel control unit is located at the front end of each input and output port and is used to manage the channel allocation of incoming and outgoing data packets. By dividing the physical channel into four independent virtual channels, when a data packet enters from a certain direction, the control unit will select a virtual channel for transmission according to predefined rules or the current network state, ensuring that the data packet will not generate a closed loop in the channel under any circumstances and avoiding deadlocks.
[0085] The virtual channel control unit can improve the transmission efficiency of the on-chip network by preventing deadlocks and implementing adaptive load balancing, ensuring that data packets can reach the destination in the shortest time. The mechanism of avoiding deadlocks guarantees the continuity of data transmission and the stability of system operation. Especially in a high-concurrency or network-congested environment, it can effectively prevent the degradation of system performance. In summary, the present invention introduces a virtual channel control unit in the on-chip routing module and the inter-chip routing module. The present invention realizes the high efficiency, stability, and resource optimization of data transmission in the acceleration computing system, provides strong hardware support for high-demand tasks such as database acceleration computing and large-scale data processing, and ensures the high performance and reliability of the system when dealing with complex interconnection communications.
[0086] In the embodiment of the present invention, the virtual channel control unit is specifically used for: constructing multiple independent micro-slice buffers in the output queue buffer corresponding to the physical channel to obtain multiple virtual channels; equally dividing the virtual channels into a first channel group and a second channel group in the clockwise direction; using the first channel group when the routing coordinate of the source node is greater than the routing coordinate of the target node; and using the second channel group when the routing coordinate of the source node is less than or equal to the routing coordinate of the target node.
[0087] The virtual channel control unit converts a single physical channel into multiple logically independent virtual channels by creating multiple independent micro-slice buffers in the output queue buffer of the physical channel to improve the communication efficiency and stability of the network.
[0088] In a specific embodiment, the virtual channel control unit creates four independent micro - slice buffers, namely VC0, VC1, VC2, and VC3, in the output queue buffer of each physical channel, which together constitute four virtual channels. Each micro - slice buffer is responsible for processing data packets in a specific direction, ensuring that the transmission of data packets does not form a closed loop and avoiding the occurrence of deadlock. Specifically, when a data packet needs to be transmitted from one group of computing units to another group, it will be assigned by the virtual channel control unit to a specific virtual channel for transmission, rather than through a single physical channel. The virtual channels are equally divided into a first channel group (including VC0 and VC1) and a second channel group (including VC2 and VC3) in the clockwise direction. This division method is associated with the transmission direction of data packets, making the data transmission in a specific direction more orderly and efficient. For example, in the on - chip routing module, data packets transmitted in the clockwise direction will use the first channel group (VC0 and VC1), while data packets transmitted in the counter - clockwise direction will use the second channel group (VC2 and VC3). During the actual transmission process, that is, according to the routing coordinates (controller ID, accelerated computing unit group ID, accelerated computing unit ID) of the source node and the destination node, to select whether to use the first channel accelerated computing unit group or the second channel accelerated computing unit group. When the routing coordinates of the source node are greater than those of the destination node, the first channel accelerated computing unit group is used for transmission; when the routing coordinates of the source node are less than or equal to those of the destination node, the second channel accelerated computing unit group is used. Specifically, in the controller cluster architecture, if a data packet is transmitted from a node in accelerated computing unit group 0 (for example, controller ID = 0, accelerated computing unit group ID = 0, accelerated computing unit ID = 0) to a node in accelerated computing unit group 3 (controller ID = 0, accelerated computing unit group ID = 3, accelerated computing unit ID = 0), since the accelerated computing unit group ID of the source node is less than that of the destination node, the second channel accelerated computing unit group (VC2 and VC3) should be used for transmission at this time. On the contrary, if the data packet is transmitted from a node in accelerated computing unit group 3 to a node in accelerated computing unit group 0, since the accelerated computing unit group ID of the source node is greater than that of the destination node, the first channel accelerated computing unit group (VC0 and VC1) will be used for transmission at this time.
[0089] By dividing the physical channel into multiple virtual channels and selecting an appropriate channel group for data transmission based on the direction and routing coordinates, the present invention effectively avoids the common deadlock problem in the ring network, ensuring the transmission continuity of data packets and the system stability. The division of the channel group and the channel selection strategy based on routing coordinates simplify the routing algorithm of the on - chip communication network, making the system design simpler. At the same time, it reduces the delay during data transmission, improves the data transmission speed and the overall performance of the system.
[0090] In an embodiment of the present invention, when the acceleration computing unit group is in the first working mode, the acceleration computing units within the group are allowed to be configured with different computing functions.
[0091] The first working mode allows the acceleration computing units within the same group to be configured to perform different computing functions, which means that different types of acceleration computing units can be flexibly allocated and combined according to specific application requirements to achieve optimal performance optimization. This is crucial for system flexibility and task execution efficiency. Specifically, each acceleration computing unit can be dynamically configured to perform different computing tasks, such as data decompression, full table scan, expression filtering, sorting, hashing, etc. This configuration method provides great flexibility, enabling the system to select the most suitable acceleration computing unit for acceleration according to the characteristics of different tasks.
[0092] The above-mentioned first working mode allows the acceleration computing units to work together to complete a series of complex database acceleration computing tasks, without each acceleration computing unit having all functions, greatly improving the resource utilization efficiency of the system. Moreover, since the acceleration computing units within the acceleration computing unit group share memory, in the first working mode, although the computing units within the group can be configured to perform different functions, the computing units with different functions directly access the shared memory within the group when performing tasks, reducing the latency and energy consumption of data transmission. For example, when the expression filtering computing unit completes its task, it can directly write the result into the shared memory within the group, without transmitting data through the on-chip network. In this way, the sorting computing unit in the same group can immediately access these results for the next step of calculation, reducing the overhead of data transmission and improving the working efficiency of the entire system. Additionally, the design of sharing group memory reduces unnecessary data transmission, can also reduce the energy consumption during data transmission, and at the same time reduces the communication latency between computing units, providing strong technical support for building a low-power and high-performance computing system.
[0093] In an embodiment of the present invention, when the acceleration computing unit group is in the second working mode, the acceleration computing units within the group are configured with the same computing function to replace with other acceleration computing units within the group in case any acceleration computing unit fails.
[0094] In large-scale parallel computing scenarios, hardware failures are inevitable. To improve the stability and fault tolerance of the system, the present invention proposes a second working mode in addition to the first working mode. The second working mode is designed to run when all the acceleration computing units within the group are configured to perform the same computing function. In this way, even if one computing unit fails, the system can still replace the function through other computing units within the group to ensure the continuity of tasks and the stability of system performance.
[0095] In a specific embodiment, in the second working mode, all the computing units within the same acceleration computing unit group will be configured to execute the same computing function. For example, all the computing units within a group can be set to specifically perform tasks such as data decompression or full table scan, etc., to ensure that when any computing unit fails, other computing units can immediately take over its work without the need for reconfiguration or scheduling. Specifically, assume that (0, 0, 0), (0, 0, 1), (0, 0, 2), and (0, 0, 3) within group 0 are configured in the second working mode and set to execute the same function, such as data decompression. Then, when any computing unit (such as (0, 0, 1)) has a hardware failure, other computing units (such as (0, 0, 0), (0, 0, 2), and (0, 0, 3)) can immediately take over the task and continue to execute the data decompression task using the shared memory resources within the group to ensure the uninterrupted execution of the task.
[0096] In the second working mode, since all the computing units within the group are configured to execute the same function, they can more effectively share the data and resources within the group, reducing the steps of retransmitting data during the fault replacement process, and further improving the computing efficiency and resource utilization efficiency. With the instant task reallocation and data sharing mechanism provided by the shared memory, the present invention can ensure that the real-time performance of task execution and the overall stability of the system are not affected when there is a hardware failure, which is particularly important for tasks with high real-time requirements such as database acceleration computing and large-scale data processing.
[0097] In an embodiment of the present invention, the controller is one of FPGA, ARM, and AVR.
[0098] FPGA, ARM microprocessor, and AVR microprocessor each have their own characteristics and are suitable for different scenarios and requirements. When FPGA plays the role of the controller, it can provide highly customized control logic and is suitable for system designs that require high flexibility and programmability. As the controller, the ARM microprocessor can provide powerful processing capabilities and rich peripheral interfaces and is suitable for tasks that require high-performance computing and complex system management. The AVR microprocessor is an ideal choice in some cost-sensitive and power consumption-limited application scenarios, and the low-cost and low-power characteristics of the AVR microprocessor make it possible to deploy large-scale sensor networks while ensuring the stability and efficiency of the system.
[0099] According to the design requirements and application scenarios of the acceleration computing system, selecting FPGA, ARM microprocessor, or AVR microprocessor as the controller can achieve functional optimization and performance improvement at different levels, providing a diverse selection of hardware platforms for building efficient, stable, and economical electronic devices and acceleration computing systems.
[0100] In an embodiment of the present invention, an acceleration computing system is provided, such as Figure 6 shown, the acceleration computing system includes: a host, a plurality of controller clusters, and a switching chip (which can be a CXL Switch). Among them, the host is responsible for task scheduling and data management, the controller clusters are responsible for executing acceleration computing tasks, and the switching chip is responsible for efficient data exchange between the host and the controller clusters.
[0101] In a specific embodiment, assuming that the host has a large-scale data processing task, such as database acceleration computing, it will decompose the task into multiple subtasks and distribute them to different controller clusters according to the computing requirements of the subtasks. The controllers within each controller cluster will efficiently execute the corresponding subtasks according to their configured computing functions. For example, the controllers in one cluster may be configured to specifically perform data decompression, while the controllers in another cluster may be responsible for full table scanning. The switching chip, as the data bridge between the host and the controller clusters, is responsible for implementing efficient bus connections and data exchanges. It can handle a large amount of data streams, ensuring fast data transmission speed and low latency between the host and the controller clusters, and improving the overall performance of the system. For example, the host is connected to the switching chip through a high-speed bus such as PCIe or CXL, and then the switching chip is connected to each controller cluster through the CXL bus. When the host needs to send data to a certain controller cluster for computing, the data will first be sent to the switching chip through the bus on the host side. The switching chip repackages the data according to the destination controller cluster and group information and sends it to the corresponding controller cluster. Inside the controller cluster, the data will be transmitted to the target acceleration computing unit through the on-chip network or packet bus to execute the corresponding computing tasks. The above switching chip is a heterogeneous cache coherence switching chip to solve the problem of insufficient cache coherence support in on-chip interconnection in related technologies.
[0102] Through the acceleration computing system architecture, the present invention can allocate computing resources according to task requirements, achieve acceleration of data processing, improve the overall performance of the system and the data processing speed. The design of the controller clusters allows the system to be dynamically adjusted and expanded according to application requirements, increasing the flexibility and scalability of the system, so that the system can adapt to various computing tasks and data volumes. Through the configuration of multiple controller clusters and the improved internal interconnection structure, even in the case of failures of controllers or acceleration computing units, the system can still continue to complete tasks through other available resources, improving the reliability and fault tolerance of the system.
[0103] In an embodiment of the present invention, as Figure 7 shown, the controllers in the controller cluster are configured in a fully connected structure.
[0104] The fully connected structure means that all the controllers within the controller cluster are directly connected to each other, that is, each controller can directly communicate with other controllers without the need for complex routing strategies. For example Figure 7 As shown, each controller includes 16 nodes, and the fully connected structure is realized through routing between controllers. This structure improves the efficiency of data transmission, reduces communication latency, and enhances the collaborative ability between system resources.
[0105] In a specific embodiment, within the controller cluster, the fully connected structure can be implemented through the CXL interface. Through the CXL interface, each controller can directly access the main memory and the memory resources of other controller clusters, reducing data transmission latency while increasing bandwidth. In addition, the unified addressing feature of the CXL protocol across the entire system enables the acceleration computing units of different controller clusters to access remote memory as if they were accessing local memory, so as to achieve data circulation within the fully connected structure.
[0106] The controller cluster under the fully connected structure can achieve data sharing for dynamic task allocation, which is particularly important in tasks such as database acceleration computing. The host side can allocate data and tasks to any one of the controllers as needed, and these controllers only need a single hop when accessing other controllers without additional routing control. In summary, the controller cluster under the fully connected structure can significantly improve the efficiency of data transmission, reduce communication latency, and the fully connected structure simplifies the complexity of task scheduling and data management, enabling the host to allocate tasks more flexibly. At the same time, the acceleration computing units can directly access the required data, improving the utilization rate of computing resources and the parallelism of task execution. The efficient and convenient mutual access also improves the collaborative ability of the system where the controller cluster is located, providing key technical support for building a high-performance computing system with low latency and high bandwidth.
[0107] In the embodiment of the present invention, the switching chip at least supports the CXL communication protocol.
[0108] The CXL communication protocol is a high-speed, low-latency interconnection technology standard designed to meet the requirements of modern data centers and high-performance computing systems for high-speed data transmission and cache coherence support. The CXL protocol supports three main modes: CXL.io, CXL.cache, and CXL.memory, which are used for input / output communication between devices, cache coherence communication, and shared memory access respectively. In the present invention, the switching chip is designed to support the CXL communication protocol, so as to achieve efficient data transmission and cache coherence management within the controller cluster architecture. This means that data can flow efficiently between various components in the system without increasing latency and reducing bandwidth. For example Figure 6As shown in the figure, the host communicates with the switching chip via the CXL bus, and the switching chip is connected to the CXL device-side controllers in each FPGA cluster via the CXL bus. This design allows the host to directly access the DDR / HBM memory in the controller cluster without going through complex routing and conversion processes, thus greatly improving the data access speed and the overall system performance. The CXL communication between controller clusters can omit additional conversions or intermediate nodes. For example, in Figure 6 In the schematic diagram of the inter-FPGA routing module interconnection shown, it can be seen that different FPGAs are fully connected through a switching chip supported by the CXL protocol, which means that data transmission between any two FPGA clusters can be directly completed via the CXL bus without complex routing judgment and packet redirection. The switching chip supporting the CXL protocol can achieve cache coherence between different FPGA clusters, ensuring data consistency and integrity and avoiding data conflict and inconsistency problems.
[0109] In summary, the switching chip supporting the CXL communication protocol can provide high-speed and low-latency data transmission capabilities, which are crucial for high-performance computing and data acceleration tasks in the controller cluster architecture. With the support of the CXL protocol, the switching chip can achieve cache coherence between chips, ensuring data integrity and consistency in the multi-controller cluster architecture and improving the system reliability. The switching chip under the CXL protocol simplifies the communication architecture in the system, reduces intermediate links in data transmission, makes the data flow more direct and efficient, reduces system latency, and improves the overall processing speed.
[0110] In an embodiment of the present invention, as shown in Figure 4 the figure, the controller cluster further includes a multiplexer for receiving signals generated inside the controller and transmitting them to other controllers or the host.
[0111] The multiplexer is a circuit component that can select one signal from multiple input signals for transmission. In the controller of a multi-controller architecture, the multiplexer is used to select the correct output path in inter-chip communication to send packets or signals to a specific target chip.
[0112] In a specific embodiment, as shown in Figure 4 the figure, after being connected to the inter-chip routing module, the multiplexer is connected to multiple controllers. Each controller has a corresponding input terminal. By decoding the controller ID, the multiplexer can select the correct output path to send packets or signals to the specified target controller. As shown in Figure 6As shown in the figure, assume there is a cluster containing 4 FPGAs, and each FPGA has a port connected to a multiplexer. When a data packet needs to be transmitted from the current controller to other controllers, the on-chip routing module will send the data packet to the inter-chip full-connection port, which transmits the data packet to the multiplexer. After receiving the signal from the inter-chip full-connection port, the selector decodes it according to the destination controller ID in the data packet and transmits the signal to the output port that matches the destination controller ID.
[0113] The introduction of the multiplexer provides direct and efficient path selection and data transmission capabilities, greatly simplifies the data transmission process between different controllers, reduces the path selection delay of data packets during routing jumps, and avoids unnecessary detours of data packets in the multi-controller interconnected network, thereby significantly improving the overall data transmission efficiency and speed of the acceleration computing system.
[0114] In an embodiment of the present invention, a routing jump method inside the controller is proposed, and this method can run on an acceleration computing system as Figure 6 shown in the figure. Figure 8 It is a flowchart of the routing jump method inside the above-mentioned controller. As Figure 8 shown in the figure, the above process includes:
[0115] Step S201, configure the routing coordinates of the acceleration computing units inside the controller, and the routing coordinates include the controller ID, the acceleration computing unit group ID, and the acceleration computing unit ID;
[0116] In some embodiments of the present invention, each acceleration computing unit is assigned corresponding routing coordinates, and the routing coordinates include three dimensions, namely the above-mentioned controller ID, the acceleration computing unit group ID, and the acceleration computing unit ID. For example, the acceleration computing unit (0, 0, 0) represents the first acceleration computing unit in the first group on controller 0. This coordinate configuration makes the position and connection relationship of each computing unit in the system clearly distinguishable and provides a basis for subsequent routing jumps.
[0117] Step S202, determine the routing coordinates of the source node and the routing coordinates of the target node based on the routing jump request, and calculate the number of on-chip jump steps according to the routing coordinates of the source node and the routing coordinates of the target node;
[0118] Specifically, when an acceleration computing unit needs to transmit data to another acceleration computing unit, the acceleration computing system first determines the routing coordinates of the source node and the target node, and then calculates the number of on-chip jump steps based on the routing coordinates of both. The formula is as follows:
[0119] The number of on-chip jump steps = ((Acceleration computing unit group ID of the destination node * Number of acceleration computing units in the group + Acceleration computing unit ID of the destination node) - (Acceleration computing unit group ID of the source node * Number of acceleration computing units in the group + Acceleration computing unit ID of the source node) + Number of acceleration computing nodes) mod Number of acceleration computing nodes.
[0120] Step S203: Determine the routing jump direction based on the number of on-chip jump steps, and perform routing jumps based on the routing jump direction.
[0121] Based on the calculated number of on-chip jump steps, the acceleration computing system determines the routing jump direction according to the shortest path principle, and then controls the transmission of data packets according to the routing jump direction to complete the routing jump process.
[0122] In a specific embodiment, the coordinates of the source node are set to (0, 0, 0), and the coordinates of the target node are (0, 3, 0). In a 16-node scale FPGA cluster, the calculated number of on-chip jump steps is 12. According to the routing policy, the system will determine that this number of steps falls within the interval (3 / 4, 1), so it chooses to jump in the clockwise direction. This means that the data packet will pass through the path of (0, 0, 0) → (0, 1, 0) → (0, 2, 0) → (0, 3, 0) and directly reach the target node without passing through unnecessary intermediate nodes, reducing latency.
[0123] Through the above steps, based on the calculation of the jump steps based on the routing coordinates and the determination of the routing jump direction, the present invention can optimize the transmission path of data packets inside the FPGA, ensure that the data packets reach the target node with the shortest path and the lowest latency, thereby significantly improving the data transmission efficiency. The deterministic routing policy and the shortest path selection reduce the latency during the data transmission process, and at the same time reduce unnecessary data rerouting, thereby reducing power consumption and improving the overall performance and energy efficiency of the system. The routing jump method based on the coordinate system enables the system to flexibly adapt to FPGA cluster architectures of different scales, and when expanding the system, it is easy to maintain and adjust the routing policy to ensure that newly added computing units can be smoothly integrated into the existing network, improving the flexibility and scalability of the system.
[0124] Among them, the execution subject of the above steps can be a host, a controller, etc., but is not limited thereto.
[0125] To avoid insufficient bandwidth and data congestion and reduce the operation efficiency of the acceleration computing system, in the embodiment of the present invention, performing routing jumps based on the routing jump direction includes:
[0126] Step S2031: Obtain the output buffer length of the controller port;
[0127] Step S2032, when the output buffer length is greater than the fifth threshold, perform reverse routing.
[0128] The communication between the acceleration computing units needs to be carried out through a complex routing network. However, the congestion degree of the routing network is closely related to the communication delay and packet loss. To avoid the performance degradation caused by network congestion, the present invention proposes the above routing strategy based on the buffer length. When the buffer length is relatively long, perform reverse routing.
[0129] In a specific embodiment, as Figure 4 shown, the first crossbar switch control unit needs to monitor the load conditions of all inputs [0-4] and outputs [0-4]. When a data packet enters from input [0] (clockwise), the control unit checks the load in the output direction. If the buffer of output [1] is full (i.e., greater than the fifth threshold as mentioned above), the control unit will direct the data packet to perform reverse routing through output [2] (opposite direction) or output [3] (clockwise) to avoid the data packet being blocked during transmission.
[0130] Through the above steps, the present invention monitors the output buffer length of the ports and adopts the reverse routing strategy when the port buffer is full, effectively balancing the communication load in the network, avoiding local network congestion, and ensuring the smooth and efficient data transmission in the acceleration computing system. Therefore, the reverse routing strategy avoids data transmission through high-load ports, thereby reducing the communication delay and maintaining the high performance and response speed of the system. Moreover, the dynamic routing adjustment reduces the queuing time of data packets on the congested link, reduces the risk of packet loss, and improves the reliability of communication.
[0131] In the embodiment of the present invention, determining the routing jump direction based on the number of on-chip jump steps includes:
[0132] Step S2033, determine the routing jump direction based on the preset range where the number of on-chip jump steps is located. The routing jump direction includes the clockwise direction, the counterclockwise direction, the opposite direction, or a combination of the clockwise direction, the counterclockwise direction, and the opposite direction.
[0133] To meet the shortest path principle, the present invention sets a method for determining the routing jump direction based on the number of on-chip jump steps. Specifically, by analyzing the preset range where the number of on-chip jump steps is located, intelligently select the most appropriate routing jump direction, thereby reducing the communication delay and improving the data transmission efficiency.
[0134] Based on the pre-configured routing policy of the present invention, the acceleration computing system can intelligently select the shortest path of data packets in the internal network of the controller, significantly reducing the communication delay and improving the data transmission efficiency. The policy of preset jump step range and routing jump direction enables the system to flexibly adjust under different communication requirements and network states, improving the stability and reliability of communication.
[0135] The execution order of step S2033, step S2031, and step S2032 can be interchanged, that is, step S2033 can be executed first, and then step S2031 and step S2032, or step S2033, step S2031, and step S2032 can be executed synchronously.
[0136] In the embodiment of the present invention, determining the routing jump direction based on the preset range of the on-chip jump step includes:
[0137] Step S20331, when the on-chip jump step is equal to the first threshold, determining the clockwise direction as the routing jump direction;
[0138] Step S20332, when the on-chip jump step is equal to the second threshold, determining the opposite direction as the routing jump direction;
[0139] Step S20333, when the on-chip jump step is equal to the third threshold, determining the counterclockwise direction as the routing jump direction;
[0140] Step S20334, when the on-chip jump step is less than the first threshold, determining the counterclockwise direction as the routing jump direction;
[0141] Step S20335, when the on-chip jump step is greater than the first threshold and less than the second threshold, determining the routing jump direction as the opposite direction first and then the clockwise direction;
[0142] Step S20336, when the on-chip jump step is greater than the second threshold and less than the third threshold, determining the routing jump direction as the opposite direction first and then the counterclockwise direction;
[0143] Step S20337, when the on-chip jump step is greater than the third threshold and less than the fourth threshold, determining the clockwise direction as the routing jump direction.
[0144] Based on such as Figure 1For the controller architecture shown, taking a 16-node controller as an example, the first threshold of the present invention is set to 4 (i.e., 1 / 4 of the number of computing nodes), the second threshold is set to 8 (i.e., 1 / 2 of the number of computing nodes), the third threshold is set to 12 (i.e., 3 / 4 of the number of computing nodes), and the fourth threshold is set to 16 (i.e., the number of computing nodes). When the number of in-chip jump steps is equal to the first threshold (4), the system determines that the clockwise direction is the routing jump direction; when the number of in-chip jump steps is equal to the second threshold (8), the system determines that the opposite direction is the routing jump direction; when the number of in-chip jump steps is equal to the third threshold (12), the system determines that the counterclockwise direction is the routing jump direction; when the number of in-chip jump steps is less than the first threshold (i.e., 0 to 3), the system determines the counterclockwise direction as the routing jump direction; when the number of in-chip jump steps is greater than the first threshold (i.e., 4 to 7) and less than the second threshold (i.e., 8), the system determines the routing jump direction as the opposite direction first and then the clockwise direction; when the number of in-chip jump steps is greater than the second threshold (i.e., 8 to 11) and less than the third threshold (i.e., 12), the system determines the routing jump direction as the opposite direction first and then the counterclockwise direction; when the number of in-chip jump steps is greater than the third threshold (i.e., 12 to 15) and less than the fourth threshold (i.e., 16), the system determines that the clockwise direction is the routing jump direction. Specifically, assuming that a data packet needs to be transmitted from (0, 0, 0) in group 0 to (0, 3, 0) in group 3, in a 16-node controller, the number of in-chip jump steps is calculated as 12. According to the strategy of the present invention, it will be determined as the counterclockwise direction, that is, the data packet will be transmitted along the counterclockwise direction from (0, 0, 0) through the path of (0, 3, 3) → (0, 2, 3) → (0, 1, 3) → (0, 3, 0), ensuring efficient data scheduling.
[0145] Through the routing jump direction strategy based on the number of in-chip jump steps and the preset threshold range, the present invention can ensure that data packets are transmitted along the shortest path in the internal network of the controller, significantly reducing communication latency and improving data transmission efficiency. The strategy of the preset threshold range and the routing jump direction can dynamically adapt to different communication requirements, realizing intelligent and efficient data scheduling, and improving the overall performance and reliability of the system. In summary, the above steps design the bus connection inside the acceleration computing unit group and the communication connection between groups to achieve low-latency and high-bandwidth data transmission. At the same time, by setting the first to fourth thresholds, according to the range where the number of in-chip jump steps is located, the routing jump in the clockwise, counterclockwise or opposite direction, as well as the combination of the above directions, is intelligently determined, so as to optimize the transmission path of data packets in the internal network of the controller. Through this flexible routing strategy, the system can effectively cope with network congestion, ensure the fast and reliable transmission of data, and thus improve the overall computing performance and efficiency of the system.
[0146] In an embodiment of the present invention, a routing jump method between controllers is proposed, and this method can run as Figure 6On the acceleration computing system shown above. Figure 9 It is a flowchart of the routing jump method between the above-mentioned controllers. As Figure 9 shown, the above process includes:
[0147] Step S301, configure the routing coordinates of the internal acceleration computing unit of the controller, where the routing coordinates include the controller ID, the acceleration computing unit group ID, and the acceleration computing unit ID;
[0148] Step S302, determine the routing coordinates of the source node and the routing coordinates of the target node based on the routing jump request;
[0149] Step S303, based on the routing coordinates of the source node, jump from the on-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the source node, jump from the inter-chip routing module corresponding to the source node to the inter-chip routing module corresponding to the target node according to the routing coordinates of the target node, and jump from the inter-chip routing module corresponding to the target node to the on-chip routing module corresponding to the target node according to the routing coordinates of the target node, and transmit to the target node.
[0150] When data needs to be transmitted between the acceleration computing units of different controllers, the on-chip routing jump is not sufficient to complete the data transmission. At this time, the system will use the inter-chip routing module for routing jumps across controllers. For example, if the source node coordinates are (0, 0, 0) and the target node coordinates are (1, 0, 0), the data packet will jump from the on-chip routing module of the source node to the inter-chip routing module of controller 0, then jump to the inter-chip routing module of controller 1, and finally reach the on-chip routing module of the target node and then be transmitted to the above-mentioned target node.
[0151] In the above steps, by configuring the routing coordinates of the acceleration computing unit, the system can accurately calculate the on-chip jump steps, provide a clear path guidance for data transmission, and reduce routing errors and communication delays. The combined use of the on-chip routing module and the inter-chip routing module enables the present invention to support communication within the same controller and communication across controllers, meeting the high-efficiency data exchange requirements between heterogeneous acceleration computing units under the controller cluster architecture.
[0152] In an embodiment of the present invention, an acceleration computing unit scheduling method for acceleration computing tasks is proposed, and this method can run on an acceleration computing system as Figure 6 shown above. Figure 10 It is a flowchart of the routing jump method between the above-mentioned controllers. As Figure 10 shown, the above process includes:
[0153] Step S401: Obtain the accelerated computing unit status table and the memory access latency table. The computing unit status table is used to include at least the routing coordinates, unit functions, unit status, and storage occupancy of the accelerated computing units. The memory access latency table includes the latency of each accelerated computing unit accessing each shared cache.
[0154] The accelerated computing unit status table records the detailed information of existing accelerated computing units (as shown in Table 1), including but not limited to routing coordinates, unit functions, unit status (such as idle or occupied), and storage occupancy (i.e., whether the current unit is using its shared cache space). The memory access latency table records the latency information of each accelerated computing unit accessing different shared caches, which is used to analyze data transfer efficiency and optimize scheduling strategies (as shown in Table 2).
[0155] Table 1
[0156]
[0157] Table 2
[0158]
[0159] Step S402: Obtain the computing tasks to be allocated, and determine the scheduling strategy according to the computing unit status table and the memory access latency table. Among them, during the process of determining the scheduling strategy, accelerated computing units with different storage occupancies can run in parallel, and the computing tasks to be allocated with the minimum latency corresponding to each accelerated computing unit in the scheduling strategy.
[0160] The system obtains the accelerated computing tasks to be executed, which may include database acceleration operations such as data decompression, full table scan, expression filtering, sorting, etc. For example, in an accelerated computing system, the order of the computing tasks to be allocated is data decompression, full table scan, expression filtering, and sorting. The accelerated computing system will intelligently allocate tasks to the most suitable accelerated computing units according to the status of the accelerated computing units and the requirements of the computing tasks to be allocated, combined with the information in the memory access latency table. This process takes into account the following factors:
[0161] Unit function matching: The system will allocate tasks to the accelerated computing units that match their functions. Unit status: The system preferentially selects accelerated computing units in the idle state to avoid overlapping execution of tasks. Storage occupancy: The system considers the storage occupancy of different computing units to achieve efficient execution of parallel tasks. Memory access latency: Based on the memory access latency table, the system will select the computing unit - memory pair with the minimum latency to reduce data transfer time.
[0162] In a specific embodiment, assume that the system needs to perform a data decompression task. First, according to the computing unit status table, the system finds that the computing unit with coordinates (0, 1, 2) is in an idle state and has a decompression function. Then, the system checks the memory access latency table and determines that the latency for this computing unit to access Group 1 memory is the shortest. Therefore, the system selects the acceleration computing unit with coordinates (0, 1, 2) to perform the data decompression task. After this unit completes the task and stores the computing result in memory, the system continues to check the status table and latency table to select the acceleration computing unit with the minimum latency and matching function for the subsequent full-table scanning task, and so on in a loop until all tasks are completed.
[0163] Through the acceleration computing unit status table, the acceleration computing system can accurately identify the functions and statuses of each unit, thereby allocating tasks to the most suitable computing unit, improving the efficiency of task execution. The memory access latency table provides latency information for the system, enabling the system to intelligently select the path with the minimum latency for data transmission, reducing communication latency and increasing data transmission speed. And during the task scheduling process, the acceleration computing system considers the possibility of parallel operation of acceleration computing units with different storage occupancies. By reasonably allocating tasks, it achieves parallel and efficient execution of multiple tasks, further improving the overall computing performance of the system.
[0164] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0165] The specific examples in this embodiment can refer to the examples described in the above embodiments and exemplary embodiments, and will not be elaborated here.
[0166] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0167] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A controller for accelerating computing, characterized in that: include: A plurality of accelerated computing unit groups, wherein adjacent accelerated computing unit groups are connected to each other in communication with each other, and are arranged in a completely symmetrical ring. Accelerated computing units in the same accelerated computing unit group are connected in sequence through an internal bus, and accelerated computing units in different accelerated computing unit groups that are located at symmetrical positions of the same ring diameter are connected to each other in communication with each other; Each accelerated computing unit group also includes a shared memory, and the shared memory is connected to each of the accelerated computing units through the internal bus; Among them, the controller also includes: an on-chip routing module, which is communicated with each of the accelerated computing unit groups and is used to realize routing jumps within the same controller; an inter-chip routing module, which is communicated with the on-chip routing module and is used to realize routing jumps between different controllers.
2. The controller according to claim 1, characterized in that: The controller further includes: A device-side controller, configured to receive a data packet in a CXL format, parse the data packet to obtain an accelerated computing task, and send the task to the accelerated computing unit; The switching module is connected to the device-side controller via a CXL bus, and communicates with the accelerated computing unit via an internal bus, and is used to receive the accelerated computing task by direct memory access and send it to the accelerated computing unit.
3. The controller according to claim 2, characterized in that: The switching module includes a multi-stage switching module, including the switching module connected to the device-side controller through the CXL bus and connected to other switching modules through multiple interfaces with output signals in the main mode, and the switching module connected to other switching modules, connected to the shared memory through an interface with output signals in the slave mode, and connected to the accelerated computing unit through an interface with output signals in the main mode.
4. The controller according to claim 1, characterized in that: The on-chip routing module includes a first cross switch and a first cross switch control unit, wherein the first cross switch is used to connect each of the accelerated computing unit groups and the inter-chip routing module, and the first cross switch control unit is used to control the on-off of the path connected by the first cross switch.
5. The controller according to claim 1, characterized in that: The inter-chip routing module includes a second cross switch and a second cross switch control unit, wherein the second cross switch is used to connect the intra-chip routing module with other controllers, and the second cross switch control unit is used to control the on-off of the path connected by the second cross switch.
6. The controller according to claim 1, characterized in that: The intra-chip routing module or the inter-chip routing module includes a virtual channel control unit, which is used to divide a physical channel into multiple virtual channels to prevent deadlock caused by channel closed loop.
7. The controller according to claim 6, characterized in that: The virtual channel control unit is specifically used for: Constructing a plurality of mutually independent micro-slice buffers in the output queue buffer corresponding to the physical channel to obtain a plurality of the virtual channels; Divide the virtual channels into a first channel group and a second channel group in a clockwise direction; In the case where the routing coordinates of the source node are greater than the routing coordinates of the target node, using the first channel group; In a case where the routing coordinates of the source node are less than or equal to the routing coordinates of the target node, the second channel group is used.
8. The controller according to claim 1, characterized in that: When the accelerated computing unit group is in the first operating mode, the accelerated computing units in the group are allowed to be configured to perform different computing functions.
9. The controller according to claim 1, characterized in that: When the accelerated computing unit group is in the second working mode, the accelerated computing units in the group are configured to perform the same computing function so that if any accelerated computing unit in the group fails, other accelerated computing units in the group can be used to replace it.
10. The controller according to claim 1, characterized in that: The controller is one of FPGA, ARM and AVR.
11. An accelerated computing system, characterized in that: The accelerated computing system comprises: Host; A plurality of controller clusters, used to execute the accelerated computing tasks issued by the host, each of the controller clusters comprising a plurality of controllers described in any one of claims 1 to 10; The switching chip is connected to the host and the controller cluster via a bus, and is used to implement data interaction between the host and the controller cluster.
12. The accelerated computing system according to claim 11, characterized in that: The controllers in the controller cluster are configured as a fully connected structure.
13. The accelerated computing system according to claim 11, characterized in that: The switching chip at least supports the CXL communication protocol.
14. The accelerated computing system according to claim 11, characterized in that: The controller cluster further includes a multiple-to-one selector for receiving a signal generated inside the controller and transmitting the signal to other controllers or the host.
15. A routing jump method inside a controller, characterized in that: The controller includes: A plurality of accelerated computing unit groups, wherein adjacent accelerated computing unit groups are connected to each other in communication with each other, and are arranged in a completely symmetrical ring. Accelerated computing units in the same accelerated computing unit group are connected in sequence through an internal bus, and accelerated computing units in different accelerated computing unit groups that are located at symmetrical positions of the same ring diameter are connected to each other in communication with each other; The method comprises: Configure the routing coordinates of the accelerated computing unit in the controller, wherein the routing coordinates include the controller ID, the accelerated computing unit group ID, and the accelerated computing unit ID; Determine the routing coordinates of the source node and the routing coordinates of the target node based on the routing jump request, and calculate the number of intra-chip jump steps according to the routing coordinates of the source node and the routing coordinates of the target node; A route jump direction is determined based on the number of intra-chip jump steps, and a route jump is performed based on the route jump direction.
16. The method according to claim 15, characterized in that Based on the route jump direction, route jump is performed, including: Obtaining the output buffer length of the controller port; When the output buffer length is greater than a fifth threshold, reverse routing is performed.
17. The method according to claim 15, characterized in that Determining a routing jump direction based on the number of jump steps within the chip includes: The routing jump direction is determined based on a preset range of the number of jump steps within the chip, and the routing jump direction includes a clockwise direction, a counterclockwise direction, an opposite direction, or a combination of the clockwise direction, the counterclockwise direction, and the opposite direction.
18. The method according to claim 17, characterized in that The controller includes a plurality of accelerated computing unit groups, wherein the accelerated computing units in the same accelerated computing unit group are sequentially connected through an internal bus and connected to the same shared memory through the internal bus, the accelerated computing units in another adjacent group of the accelerated computing units between different accelerated computing unit groups are communicated and connected to each other, and the accelerated computing units in another opposite group are communicated, and the routing jump direction is determined based on a preset range of the number of jump steps within the chip, and the method includes: When the number of jump steps within the chip is equal to the first threshold, determining the clockwise direction as the routing jump direction; When the number of jump steps within the slice is equal to the second threshold, determining the opposite direction as the routing jump direction; When the number of jump steps within the chip is equal to a third threshold, determining a counterclockwise direction as the routing jump direction; When the number of jump steps within the chip is less than the first threshold, determining the counterclockwise direction as the routing jump direction; When the number of jump steps within the chip is greater than the first threshold and less than the second threshold, firstly determining the opposite direction and then the clockwise direction as the routing jump direction; When the number of jump steps within the chip is greater than the second threshold and less than the third threshold, firstly facing the direction and then the counterclockwise direction are determined as the routing jump direction; When the number of intra-chip jump steps is greater than the third threshold and less than the fourth threshold, the clockwise direction is determined as the routing jump direction.
19. A method for routing jumps between controllers, characterized in that: The controller comprises: a plurality of accelerated computing unit groups, wherein adjacent accelerated computing unit groups are connected to each other in communication with each other, and are arranged in a completely symmetrical ring, wherein the accelerated computing units in the same accelerated computing unit group are connected in sequence through an internal bus, and the accelerated computing units in different accelerated computing unit groups that are located at symmetrical positions of the same ring diameter are connected to each other in communication; an intra-chip routing module and an inter-chip routing module, wherein the intra-chip routing module is connected to each accelerated computing unit group in communication with each other, and the inter-chip routing module is connected to the intra-chip routing module in communication with each other, and the method comprises: Configure the routing coordinates of the accelerated computing unit in the controller, wherein the routing coordinates include the controller ID, the accelerated computing unit group ID, and the accelerated computing unit ID; Determining the routing coordinates of the source node and the routing coordinates of the target node based on the routing jump request; Based on the routing coordinates of the source node, the intra-chip routing module corresponding to the source node is jumped to the inter-chip routing module corresponding to the source node; according to the routing coordinates corresponding to the target node, the inter-chip routing module corresponding to the source node is jumped to the inter-chip routing module corresponding to the target node; and according to the routing coordinates corresponding to the target node, the inter-chip routing module corresponding to the target node is jumped to the intra-chip routing module corresponding to the target node, and transmitted to the target node.
20. A method for scheduling accelerated computing units for accelerated computing tasks, characterized in that: The accelerated computing unit scheduling method is applied to the accelerated computing system according to claim 11, and the method comprises: Obtain an accelerated computing unit state table and a memory access delay table, wherein the computing unit state table is used to include at least the routing coordinates, unit functions, unit states, and storage occupancy of the accelerated computing unit, and the memory access delay table includes the delay of each of the accelerated computing units accessing each shared cache; Obtain the computing task to be assigned, and determine the scheduling strategy according to the computing unit status table and the memory access delay table, wherein, in the process of determining the scheduling strategy, the accelerated computing units with different storage occupancy can run in parallel, and each of the accelerated computing units in the scheduling strategy corresponds to the computing task to be assigned with the smallest delay.
Citation Information
Patent Citations
Network-on-chip routing centralized control system and device and adaptive routing control method
CN102546406A
Network-on-chip structure construction method and system, network-on-chip structure use method and system, equipment and storage medium
CN113986813A