An acceleration unit and electronic device
By designing an accelerator card matrix and a fully connected square network topology in the acceleration unit, the data transmission bandwidth challenge in multi-card networks was solved, achieving efficient massive data processing and improved computing power.
Patent Information
- Application Number
- CN202010969307.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-15
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2040-09-15
AI Technical Summary
In multi-card networks composed of multiple ASIC acceleration chips, high data throughput poses a significant challenge to the data transmission bandwidth of the ASICs. Designing interconnection schemes between chips to improve the computing power of the entire system becomes crucial.
An acceleration unit design is adopted, in which each acceleration card connects to other acceleration cards through an internal port to form an L*N-sized acceleration card matrix. By utilizing a fully connected square network topology and a dual-ring structure, interconnection between acceleration cards is achieved, thereby improving computing power and data processing speed.
It effectively improves the computing power of the acceleration unit, meets the system's real-time requirements, achieves the goal of high-speed processing of massive amounts of data, and enhances the stability and reliability of the system.
Smart Images

Figure CN114185831B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of processor technology. More specifically, this disclosure relates to an acceleration unit, acceleration component, acceleration device, circuit board, and electronic device. Background Technology
[0002] Currently, with the rapid development of Artificial Intelligence (AI) and Machine Learning, the demand for ultra-high-performance processors will continue to grow, while the era of big data places even higher demands on data processing. High-performance processors and clusters need to handle massive amounts of data in real time and complete the training and inference of complex models within a specified timeframe. ASICs (Application Specific Integrated Circuits) are dedicated acceleration chips that can be used to train deep neural networks. ASICs can complete tasks in a much shorter time, requiring far less data center infrastructure than non-parallel processing supercomputers.
[0003] However, when faced with massive amounts of data, even the most powerful single ASIC is inevitably insufficient. To achieve greater computing power, a common approach is to use multiple ASIC acceleration chips. However, for multi-GPU networks composed of interconnected ASICs, the extremely high data throughput presents a significant challenge to the data transmission bandwidth of the ASICs. Therefore, designing interconnection schemes between chips to improve the overall system's computing power and achieve efficient processing of massive amounts of data has become a key technical issue in building high-performance processor clusters. Summary of the Invention
[0004] To address the aforementioned technical issues, this disclosure provides an acceleration unit, acceleration component, acceleration device, circuit board, and electronic device capable of improving computing power.
[0005] In one aspect, this disclosure provides an acceleration unit comprising M accelerator cards, each accelerator card including an internal port, each accelerator card being connected to other accelerator cards via the internal port, wherein the M accelerator cards are logically formed into an L*N accelerator card matrix, where L and N are integers not less than 2.
[0006] In another aspect, this disclosure provides an electronic device including the acceleration unit as described above.
[0007] In this disclosed scheme, the acceleration unit consists of multiple acceleration cards. Each acceleration card connects to other acceleration cards through its internal ports, achieving interconnection between the acceleration cards. This configuration effectively improves the computing power of the acceleration unit, which is beneficial for increasing the speed of processing massive amounts of data. Furthermore, the interconnection between acceleration units minimizes the latency of the entire system, maximizing the real-time performance requirements while processing massive amounts of data. This enhances the overall computing power of the system and enables it to achieve the goal of high-speed processing of massive amounts of data. Attached Figure Description
[0008] The above and other objects, features, and advantages of exemplary embodiments of this disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:
[0009] Figure 1a To disclose a schematic diagram of the acceleration unit structure in one embodiment.
[0010] Figure 1b , Figure 2 , Figure 3 , Figure 4 as well as Figures 5a-5c These are schematic diagrams of multiple structures of the acceleration unit in embodiments of this disclosure;
[0011] Figures 6-11 These are schematic diagrams of multiple structures of the acceleration components in embodiments of this disclosure;
[0012] Figures 12a-12c A schematic diagram representing the network topology to accelerate components;
[0013] Figure 13 This is a schematic diagram of an acceleration device including multiple acceleration units, according to an embodiment of this disclosure.
[0014] Figure 14 This is a schematic diagram of the network topology corresponding to the acceleration device in one embodiment;
[0015] Figure 15 This is a schematic diagram of the network topology corresponding to the acceleration device in another embodiment;
[0016] Figures 16-20 These are several schematic diagrams of an acceleration device including multiple acceleration components, as described in embodiments of this disclosure;
[0017] Figure 21 This is a schematic diagram of the network topology for yet another type of acceleration device;
[0018] Figure 22A schematic diagram of a matrix network topology based on wireless extension of an acceleration device;
[0019] Figure 23 This is a schematic diagram of the acceleration device in yet another embodiment of this disclosure;
[0020] Figure 24 This is a schematic diagram of the network topology for yet another type of acceleration device;
[0021] Figure 25 This is a schematic diagram of the network topology for yet another type of acceleration device;
[0022] Figure 26 This is a schematic diagram of the combined device structure in one embodiment of this disclosure;
[0023] Figure 27 This is a schematic diagram of the circuit board structure in one embodiment of this disclosure. Detailed Implementation
[0024] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0025] Several embodiments disclosed herein will now be described in detail with reference to the accompanying drawings.
[0026] Figure 1a This document discloses a schematic diagram of an acceleration unit structure in one embodiment. According to one embodiment of this disclosure, an acceleration unit is provided, comprising M local acceleration cards. Each local acceleration card includes an internal port, and each local acceleration card is connected to other local acceleration cards through the internal port. The M local acceleration cards logically form an L*N scale acceleration card matrix, where L and N are integers not less than 2.
[0027] like Figure 1a As shown, multiple accelerator cards can be used to form an accelerator card matrix, which is interconnected to enable data or instruction transmission and communication. For example, the accelerator card MC... 00 To MC 0N The 0th row of the acceleration card matrix is formed, and the acceleration card MC is formed. 10 To MC 1N This forms the first row of the acceleration card matrix, and so on, forming the acceleration card MC. L0 To MC LN This forms the Lth row of the accelerator card matrix.
[0028] It should be understood that, for ease of understanding, accelerator cards within the same acceleration unit are referred to as "accelerator cards of this unit," while accelerator cards in other acceleration units are referred to as "accelerator cards of other units." This terminology is merely for ease of description and does not constitute a limitation on the technical solution disclosed herein.
[0029] Each accelerator card can have multiple ports, which can connect to accelerator cards within the same unit or to accelerator cards in other units. In this disclosure, the connection ports between accelerator cards within the same unit can be referred to as internal ports, while the connection ports between an accelerator card within the same unit and an accelerator card in other units can be referred to as external ports. It should be understood that the terms "external port" and "internal port" are used merely for convenience, and the same port can be used for both. This will be described below.
[0030] It is important to understand that M can be any integer. M accelerator cards can be arranged into a 1*M or M*1 matrix, or the M matrices can be arranged into other types of matrices. The accelerator unit disclosed herein does not limit the specific matrix size and format.
[0031] Furthermore, accelerator cards can be connected via one or more communication paths, such as between accelerator cards within the same unit or between an accelerator card in the same unit and an accelerator card in an external unit. This will be described in detail later.
[0032] It is also important to understand that, in the context of this disclosure, although the positions of multiple accelerator cards are described using a rectangular network, the resulting matrix is not necessarily a matrix in physical space. It can be located anywhere; for example, multiple accelerator cards can form a straight line or be arranged irregularly. The aforementioned matrix is merely logical; as long as the connections between the accelerator cards form a matrix relationship, it is acceptable.
[0033] According to one embodiment of this disclosure, M can be 4, whereby 4 accelerator cards of this unit can logically form a 2*2 accelerator card matrix; M can be 9, whereby 9 accelerator cards of this unit can logically form a 3*3 accelerator card matrix; M can be 16, whereby 16 accelerator cards of this unit can logically form a 4*4 accelerator card matrix. M can also be 6, whereby 6 accelerator cards of this unit can logically form a 2*3 or 3*2 accelerator card matrix; M can also be 8, whereby 8 accelerator cards of this unit can logically form a 2*4 or 4*2 accelerator card matrix.
[0034] According to one embodiment of this disclosure, each local accelerator card is connected to at least one other local accelerator card via two paths.
[0035] In the topology described in this disclosure, two accelerator cards in this unit can be connected via a single communication path or via multiple paths (e.g., two), provided there are enough ports. Connecting via multiple communication paths helps ensure the reliability of communication between the accelerator cards, which will be explained and described in more detail in the examples below.
[0036] According to one embodiment of this disclosure, the diagonal accelerator cards in the four corners of the accelerator card matrix are connected by two paths. For a matrix, it is preferable to connect two pairs of accelerator cards located at opposite corners. For certain topologies, connecting accelerator cards at diagonal positions helps to form two complete communication loops. This will be explained and described in more detail in the examples below.
[0037] More specifically, according to one embodiment of this disclosure, at least one of the unit accelerator cards may include an external port. For example, each accelerator unit may include four unit accelerator cards, each unit accelerator card may include six ports, and four ports of each unit accelerator card are internal ports for connecting to the other three unit accelerator cards; the remaining two ports of at least one unit accelerator card are external ports for connecting to external unit accelerator cards.
[0038] It's important to understand that of the six ports on each accelerator card in this unit, four ports can be used to connect to the accelerator card within that unit, while the remaining two ports can be used to connect to accelerator cards in other accelerator units. These remaining ports can also remain unused, without connecting to any external devices, or they can be directly or indirectly connected to other devices or ports.
[0039] For illustrative and simplification purposes, the acceleration units, acceleration components, acceleration devices, and electronic devices described below are all based on the example of each acceleration unit comprising four acceleration cards. It should be understood that each acceleration unit may include more or fewer acceleration cards.
[0040] For ease of description, the acceleration unit may include four acceleration cards, namely the first acceleration card, the second acceleration card, the third acceleration card and the fourth acceleration card. Each acceleration card is provided with an internal port and an external port, and each acceleration card is connected to the other three acceleration cards through the internal port.
[0041] Figure 1bThis is a schematic diagram of the acceleration unit structure in one embodiment of this disclosure. The acceleration unit 100 includes four acceleration cards: MC0, MC1, MC2, and MC3. Each acceleration card may include an external port and an internal port. The internal port of acceleration card MC0 is connected to the internal ports of acceleration cards MC1, MC2, and MC3; the internal port of acceleration card MC1 is connected to the internal ports of acceleration cards MC2 and MC3; and the internal port of acceleration card MC2 is connected to the internal port of acceleration card MC3. In other words, the internal port of each acceleration card is connected to the internal ports of the other three acceleration cards. Information exchange between the four acceleration cards can be achieved through the interconnection of their internal ports. This embodiment of the disclosure utilizes the interconnection between the four acceleration cards in the acceleration unit to improve the computing power of the acceleration unit and achieve high-speed processing of massive amounts of data, while minimizing the path between each acceleration card and the other acceleration cards, resulting in the lowest communication latency.
[0042] As described above, the number of accelerator cards in this disclosure is not limited to four, but can be any other number. For example, in one embodiment, the number of accelerator cards N equals 3, each accelerator card is provided with an internal port and an external port, and each accelerator card is connected to the other two accelerator cards through its internal port, realizing interconnection between the three accelerator cards. In another embodiment, the number of accelerator cards N equals 5, each accelerator card is provided with an internal port and an external port, and each accelerator card is connected to the other four accelerator cards through its internal port, realizing interconnection between the five accelerator cards, thereby improving the computing power of the acceleration unit and achieving high-speed processing of massive amounts of data. In yet another embodiment, the number of accelerator cards N is greater than 5, each accelerator card is provided with an internal port and an external port, and each accelerator card is connected to all other accelerator cards through its internal port, realizing interconnection between N accelerator cards and achieving high-speed processing of massive amounts of data.
[0043] based on Figure 1b The provided acceleration unit 100 further allows each acceleration card to be connected to at least one other acceleration card via two paths. Specifically, there are three possible connection methods: the first method is that each acceleration card can be connected to one of the other three acceleration cards via two paths; the second method is that each acceleration card can be connected to two of the other three acceleration cards via two paths; and the third method is that each acceleration card can be connected to all three acceleration cards via two paths, in which case it is possible that each acceleration card has more ports. To facilitate understanding of the above-mentioned two-path connection methods, the first connection method will be used as an example below, combined with... Figure 2 An exemplary description is provided.
[0044] Figure 2This is a schematic diagram of the acceleration unit structure in another embodiment of this disclosure. Figure 2 In the acceleration unit 200 shown, each acceleration card can be connected to at least one other acceleration card via two paths. For example, acceleration card MC0 and acceleration card MC2 can be connected via two paths, as can acceleration card MC1 and acceleration card MC3. With this configuration, there can be two links (or paths) for information exchange between two acceleration cards. Thus, if one link fails, the two acceleration cards can still be connected via the other link, effectively improving the security of the acceleration unit.
[0045] The above is combined with Figure 1 and Figure 2 The acceleration unit and the connection methods between its plurality of acceleration cards according to this disclosure have been described exemplaryly. Those skilled in the art should understand that the above description is exemplary and not restrictive; for example, the arrangement of the acceleration cards in the acceleration unit may not be limited to that shown in Figure 1 and... Figure 2 As shown in the diagram, in one embodiment, the four accelerator cards of the acceleration unit can be logically arranged in a quadrilateral pattern, which will be discussed below. Figure 3 Describe it.
[0046] Figure 3 This is a schematic diagram of the acceleration unit structure in yet another embodiment of this disclosure. Figure 3 In the acceleration unit 300 shown, the four acceleration cards MC0, MC1, MC2, and MC3 can logically be arranged in a quadrilateral pattern, with each card occupying one of the four vertices of the quadrilateral. The wiring between the acceleration cards MC0, MC1, MC2, and MC3 also forms a quadrilateral, making the wiring arrangement clearer and easier to set up. It should be noted that... Figure 3 The four accelerator cards shown are arranged in a rectangle or a 2x2 matrix. However, this is a logic interconnection diagram, and the rectangular form is used for ease of description. The specific quadrilateral shape can be freely set, such as a parallelogram, trapezoid, or square. In actual layout and routing, the four accelerator cards can also be arranged arbitrarily. For example, in an actual system, the four accelerator cards are arranged side by side in a straight line, and the order could be MC0, MC1, MC2, and MC3. It should also be understood that the logical quadrilateral shown in this embodiment is exemplary. In reality, the arrangement shape of multiple accelerator cards can vary greatly, and the quadrilateral is just one of them. For example, when there are five accelerator cards, they can be logically arranged in a pentagon.
[0047] based on Figure 2 For further details regarding the connection relationships of the provided acceleration unit 200, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram of the acceleration unit structure in yet another embodiment of this disclosure. Figure 4In the acceleration unit 400 shown, the four acceleration cards MC0, MC1, MC2, and MC3 can logically be arranged in a quadrilateral, with each card occupying one of the four vertices of the quadrilateral. As further illustrated, the internal ports of acceleration cards MC1 and MC3 can be connected via two paths, as can the internal ports of MC0 and MC2. This not only simplifies the wiring configuration of the acceleration unit 400 but also enhances its security.
[0048] Figure 5a This is a schematic diagram of the acceleration unit structure in one embodiment of this disclosure. Figure 5a In the acceleration unit 500 shown, the numerical markings on each acceleration card represent ports. Each acceleration card may include six ports: port 0, port 1, port 2, port 3, port 4, and port 5. Ports 1, 2, 4, and 5 are internal ports, while ports 0 and 3 are external ports. For the four acceleration cards MC0, MC1, MC2, and MC3, each card's two external ports can be connected to other acceleration units for interconnection between multiple acceleration units. Each acceleration card's four internal ports can be used to interconnect with the other three acceleration cards in the same acceleration unit.
[0049] like Figure 5a As further illustrated, the four accelerator cards can logically be arranged, for example, in a quadrilateral. Accelerator cards MC0 and MC2 can be diagonally connected, with port 2 of MC0 and port 2 of MC2 connected, and port 5 of MC0 and port 5 of MC2 connected. This means there can be two communication links between accelerator cards MC0 and MC2. Similarly, accelerator cards MC1 and MC3 can be diagonally connected, with port 2 of MC1 and port 2 of MC3 connected, and port 5 of MC1 and port 5 of MC3 connected. This also means there can be two communication links between accelerator cards MC1 and MC3.
[0050] With this configuration, each accelerator card has two external ports and four internal ports. Furthermore, in two diagonally arranged pairs of accelerator cards, each pair can be connected via two internal ports to form two links, effectively improving the security and stability of the acceleration unit. The logical quadrilateral arrangement of the four accelerator cards also makes the wiring layout of the entire acceleration unit clear and logical, facilitating wiring operations within each unit. It should be further noted that... Figure 5bIn the interconnection lines between the four accelerator cards shown, the connection lines between port 1 of accelerator card MC1 and port 1 of MC0, port 2 of accelerator card MC0 and port 2 of MC2, port 1 of accelerator card MC2 and port 1 of MC3, and port 2 of accelerator card MC3 and port 2 of MC1, these four lines form a vertical figure-eight network, as shown below. Figure 5b As shown. The connection lines between port 4 of accelerator card MC1 and port 4 of MC2, the connection line between port 5 of accelerator card MC2 and port 5 of MC0, the connection line between port 4 of accelerator card MC0 and port 4 of MC3, and the connection line between port 5 of accelerator card MC3 and port 5 of MC1, these four lines form a horizontal figure-eight network, as shown. Figure 5c As shown, these two fully connected square networks can form a double-ring structure, providing redundancy and enhancing system reliability.
[0051] According to one embodiment of this disclosure, the accelerator card described herein may be a Mezzanine Card (MC Card), which can be a separate circuit board. The MC Card may house an ASIC chip and some necessary peripheral control circuitry. The MC Card can be connected to the substrate via a snap-on connector. Power and control signals from the substrate can be transmitted to the MC Card through the snap-on connector. According to another embodiment of this disclosure, the internal and / or external ports described herein may be SerDes ports. For example, in one embodiment, each MC Card can provide six bidirectional SerDes ports, each with eight channels and a data transmission rate of 56Gbps. Therefore, the total bandwidth of each port can reach up to 400Gbps, which can support massive data exchange between accelerator cards, facilitating high-speed processing of massive amounts of data by the acceleration unit.
[0052] The SerDes mentioned above is a portmanteau of the English words serializer and de-serializer, and is referred to as a serial deserializer. SerDes interfaces can be used to build high-performance processor clusters. The main function of SerDes is to convert multiple low-speed parallel signals into serial signals at the transmitting end, transmit them through the transmission medium, and finally convert the high-speed serial signals back into low-speed parallel signals at the receiving end. Therefore, it is very suitable for end-to-end long-distance high-speed transmission requirements. In another embodiment, the external port in the accelerator card can be connected to the QSFP-DD interface of other accelerator units. The QSFP-DD interface is a commonly used optical module interface in SerDes technology, which, when used with cables, can be used to interconnect with other external devices.
[0053] Furthermore, according to yet another embodiment of this disclosure, an acceleration unit can house four acceleration cards, and the interconnection of the four acceleration cards can be achieved using printed circuit board (PCB) traces. On a high-speed board with a low dielectric constant, signal integrity can be maximized through reasonable layout and wiring, thereby ensuring that the communication bandwidth between the four acceleration cards approaches the theoretical value.
[0054] The acceleration unit disclosed herein comprises four acceleration cards. Each acceleration card connects to the other three acceleration cards via its internal port. Each acceleration card can directly communicate with the other three acceleration cards. This communication architecture is a fully connected quad network topology. The advantage of this fully connected network architecture is that the path between each acceleration card and the other acceleration cards is the shortest, the total number of hops is the smallest, and the latency is the lowest. This disclosure uses hops to describe the system latency. Hop in communication represents the number of hops, i.e., the number of communications. Specifically, Hop represents the shortest path from a node, traversing all nodes in the network, and returning to the initial node. The interconnection of the four acceleration cards forms a fully connected quad network topology with the shortest latency. Furthermore, the double-ring structure formed by the interconnection of two diagonally opposite acceleration cards improves the robustness of the system, ensuring that services can still operate normally even if a single acceleration card fails. During various arithmetic and logical operations, each ring in the double-ring structure can complete a portion of the operation, thereby improving overall computational efficiency and maximizing the utilization of topology bandwidth.
[0055] The above combination Figures 1a-5c Several embodiments of the acceleration unit according to this disclosure have been described. Based on the above-described acceleration unit, this disclosure also discloses an acceleration component that may include multiple of the above-described acceleration units. Several embodiments of the acceleration component will be described exemplarily below.
[0056] Figure 6 This is a schematic diagram of the acceleration component structure in one embodiment of this disclosure. Figure 6 As shown, the acceleration component 600 may include n acceleration units, namely acceleration unit A1, acceleration unit A2, acceleration unit A3, ..., acceleration unit An. Acceleration unit A1 and acceleration unit A2 are connected through an external port, and acceleration unit A2 and acceleration unit A3 are connected through an external port. That is, each acceleration unit is connected to the others through its external port. In one embodiment, the external port of acceleration card MC0 in acceleration unit A1 can be connected to the external port of acceleration card MC0 in acceleration unit A2, and the external port of acceleration card MC0 in acceleration unit A2 can be connected to the external port of acceleration card MC0 in acceleration unit A3. That is, each acceleration unit is connected through the external port of acceleration card MC0.
[0057] Those skilled in the art will understand that the connection between acceleration units in this disclosure is not limited to the connection of the external port of acceleration card MC0, but may also include one or more of the following: the connection of the external ports of acceleration card MC1, acceleration card MC2, and acceleration card MC3. That is, in this disclosure, the connection method between acceleration unit A1 and acceleration unit A2 may include one or more of the following: connecting the external port of MC0 in A1 to the external port of MC0 in A2; connecting the external port of MC1 in A1 to the external port of MC1 in A2; connecting the external port of MC2 in A1 to the external port of MC2 in A2; and connecting the external port of MC3 in A1 to the external port of MC3 in A2. Similarly, the connection methods between acceleration unit A2 and acceleration unit A3 can include one or more of the following: connecting the external port of MC0 in A2 to the external port of MC0 in A3; connecting the external port of MC1 in A2 to the external port of MC1 in A3; connecting the external port of MC2 in A2 to the external port of MC2 in A3; and connecting the external port of MC3 in A2 to the external port of MC3 in A3. This can be extrapolated to the connection between acceleration unit An-1 and acceleration unit An. It should be noted that the above description is exemplary; for example, the connection between different acceleration units is not limited to the connection of acceleration cards corresponding to the labels, but can be set to the connection of acceleration cards with the same label as needed.
[0058] It should be noted that, Figure 6 The diagram shows n acceleration units, where n is greater than 3. However, the number of acceleration units is not limited to the number greater than 3 shown in the diagram. It can also be set to, for example, 2 or 3. The connection relationship between two acceleration units is the same as or similar to the connection relationship between acceleration units A1 and A2 mentioned above. The connection relationship between three acceleration units is the same as or similar to the connection relationship between acceleration units A1, A2, and A3 mentioned above. These details will not be repeated here.
[0059] Furthermore, the structures of multiple acceleration units within an acceleration component can be identical or different. Figure 6 For ease of demonstration, the structures of the multiple acceleration units shown are identical. However, in reality, the structures of multiple acceleration units can be different. For example, in some acceleration units, the multiple accelerator cards are arranged in a polygonal layout; in others, they are arranged in a line; in still others, they are connected by a single line; and in some acceleration units, they are connected by two links. Some acceleration units include four accelerator cards, while others include three or five. In other words, the structure of each acceleration unit can be set independently, and different acceleration units can have the same or different structures.
[0060] The acceleration component disclosed herein allows for interconnection not only between accelerator cards within the acceleration unit but also between accelerator cards in different acceleration units, thereby enabling the construction of a hybrid three-dimensional network. With this configuration, each accelerator card can process data while simultaneously sharing data through interconnection between acceleration units. Since data sharing allows for direct data acquisition, reducing data propagation paths and time, it significantly improves data processing efficiency.
[0061] Figure 7 This is a schematic diagram of the acceleration component structure in another embodiment of this disclosure. Figure 7 As shown, the acceleration component 700 may include n aforementioned acceleration units, namely acceleration unit A1, acceleration unit A2, acceleration unit A3, ..., acceleration unit An. The multiple acceleration units in the acceleration component 700 can logically form a multi-layered structure (shown as dashed lines in the diagram). Each layer may include one acceleration unit, and the acceleration card of each acceleration unit is connected to the acceleration card in another acceleration unit via an external port. This progressively layered configuration allows each acceleration card to share data via a high-speed serial link while processing data at high speed, enabling unlimited interconnection of acceleration cards to meet customizable computing power requirements and achieve flexible configuration of the processor cluster's hardware computing power. As further shown in the diagram, each layer of acceleration units may include four acceleration cards, which can logically be arranged in a quadrilateral configuration, with the four acceleration cards positioned at the four vertices of the quadrilateral.
[0062] Those skilled in the art should understand that the above combination Figure 7 The described acceleration components are exemplary and not limiting. For example, the structures of multiple acceleration units can be the same or different. The number of layers in the acceleration component can be 2, 3, 4, or more, and the number of layers can be freely set as needed. For each pair of connected acceleration units, the number of connection paths between them can be 1, 2, 3, or 4. For ease of understanding, the following will combine... Figure 8 - Figure 12 provides an exemplary description.
[0063] Figure 8 This is a schematic diagram of the acceleration component structure in yet another embodiment of this disclosure. Figure 8 As shown, the acceleration component 701 can have two acceleration units, which are connected by a path. Specifically, they can be connected through, for example, the external port of the acceleration card MC0 in acceleration unit A1 and the external port of the acceleration card MC0 in acceleration unit A2, so that information exchange between acceleration unit A1 and acceleration unit A2 can be realized.
[0064] like Figure 9As shown, the acceleration component 702 can contain two acceleration units, connected via two paths. The external port of acceleration card MC0 in acceleration unit A1 is connected to the external port of acceleration card MC0 in acceleration unit A2, and the external port of acceleration card MC1 in acceleration unit A1 is connected to the external port of acceleration card MC1 in acceleration unit A2. This ensures that if one path fails, the other path still supports communication between the acceleration units, further enhancing the security of the acceleration component.
[0065] Please refer to the following. Figure 10 , Figure 10 This is a schematic diagram of the acceleration component structure in yet another embodiment of this disclosure. Figure 10 In the acceleration component 703 shown, there can be two acceleration units. The two acceleration units are connected via three paths: the external port of acceleration card MC0 in acceleration unit A1 is connected to the external port of acceleration card MC0 in acceleration unit A2; the external port of acceleration card MC1 in acceleration unit A1 is connected to the external port of acceleration card MC1 in acceleration unit A2; and the external port of acceleration card MC2 in acceleration unit A1 is connected to the external port of acceleration card MC2 in acceleration unit A2. Thus, even if two of the paths fail, there is still a third path to support communication between the acceleration units, further improving the security of the acceleration component.
[0066] Please refer to the following. Figure 11 , Figure 11 This is a schematic diagram of the acceleration component structure in yet another embodiment of this disclosure. Figure 11 In the acceleration component 704 shown, there can be two acceleration units. The two acceleration units can be connected via four paths. For example, the external port of acceleration card MC0 in acceleration unit A1 can be connected to the external port of acceleration card MC0 in acceleration unit A2; the external port of acceleration card MC1 in acceleration unit A1 can be connected to the external port of acceleration card MC1 in acceleration unit A2; the external port of acceleration card MC2 in acceleration unit A1 can be connected to the external port of acceleration card MC2 in acceleration unit A2; and the external port of acceleration card MC3 in acceleration unit A1 can be connected to the external port of acceleration card MC3 in acceleration unit A2. Thus, even if three of these paths fail, there is still one path to support communication between the acceleration units, further improving the security of the acceleration component.
[0067] Figure 12a A schematic diagram illustrating the network topology to accelerate component representation. For example... Figure 12aAs shown, the acceleration component 705 may include two acceleration units, each acceleration unit may include four acceleration cards, and there may be two links between acceleration cards MC1 and MC3 in each acceleration unit, and two links between acceleration cards MC0 and MC2. Figure 12a The acceleration device 705 in the left figure can form the three-dimensional representation shown in the right figure. Figure 12a In the right-hand diagram, circles represent accelerator cards, and lines represent link connections. Within each circle, number 0 represents accelerator card MC0, number 1 represents accelerator card MC1, number 2 represents accelerator card MC2, and number 3 represents accelerator card MC3. The right-hand diagram still represents the accelerator component 705, but as a different representation, showing a network topology. The numbers embedded in the vertical lines in the right-hand diagram indicate the port numbers for connections. For example, MC0 units are connected via port 0, MC1 units are connected via port 0, MC2 units are connected via port 3, and MC3 units are connected via port 3.
[0068] for Figure 12a The right figure shows an acceleration unit as a node. Two nodes have eight acceleration cards, forming an eight-card interconnect. The one-machine-four-card interconnection relationship within each node is fixed. When two nodes are interconnected, MC0 and MC1 in the upper node (acceleration unit A1) are connected to MC0 and MC1 in the lower node (acceleration unit A2) through port 0, respectively; MC2 and MC3 in the upper node are connected to MC2 and MC3 in the lower node through port 3, respectively. This node topology is called a hybrid cube mesh network topology, meaning that the acceleration component 705 is a hybrid cube mesh network topology.
[0069] exist Figure 12a In the 8-card topology shown, two independent rings can also be formed. For example... Figure 12b and Figure 12c As shown, this allows for maximum utilization of the topology bandwidth for reduction operations.
[0070] exist Figure 12b In the acceleration unit A1, acceleration cards MC1 and MC3 are connected via their respective internal ports 5, acceleration cards MC0 and MC2 are connected via their respective internal ports 5, and acceleration cards MC2 and MC3 are connected via their respective internal ports 1. Acceleration cards MC1 in acceleration unit A1 and MC1 in acceleration unit A2 are connected via their respective external ports 0, and acceleration cards MC0 in acceleration unit A1 and MC0 in acceleration unit A2 are connected via their respective external ports 0. Thus, an independent loop is formed among the eight cards in Figure 12.
[0071] exist Figure 12cIn the acceleration unit A1, acceleration cards MC1 and MC3 are connected via their respective internal ports 2, acceleration cards MC0 and MC2 are connected via their respective internal ports 2, and acceleration cards MC0 and MC1 are connected via their respective internal ports 1. Acceleration cards MC2 in acceleration unit A1 and MC2 in acceleration unit A2 are connected via their respective external ports 3, and acceleration cards MC3 in acceleration unit A1 and MC3 in acceleration unit A2 are connected via their respective external ports 3. Thus, another independent loop is formed among the eight cards in Figure 12.
[0072] The above only shows two exemplary connection methods. In reality, the four connection paths between the two acceleration units are actually equivalent. Therefore, any one to three of these four paths can be used to connect the two acceleration units and form a ring connection with the acceleration card within each acceleration unit. This will not be elaborated further here.
[0073] Figure 13 This is a schematic diagram of the acceleration device in yet another embodiment of this disclosure. Figure 13 As shown, the acceleration device 800 may include n acceleration units, namely acceleration unit A1, acceleration unit A2, acceleration unit A3, ..., acceleration unit An. The multiple acceleration units in the acceleration device 800 are logically arranged in a multi-layered structure (shown as dashed lines in the figure). This multi-layered structure may include an odd number of layers or an even number of layers. Each layer may include one acceleration unit. The acceleration card of each acceleration unit is connected to the acceleration card of another acceleration unit through an external port. Specifically, acceleration unit A1 and acceleration unit A2 are connected through an external port, acceleration unit A2 and acceleration unit A3 are connected through an external port, and so on, with acceleration unit An-1 and acceleration unit An connected through an external port. Furthermore, the last acceleration unit can be connected to the first acceleration unit, thus forming a ring structure with the multiple acceleration units connected end-to-end. For example, in the figure, the external port of acceleration card MC0 of acceleration unit An is connected to the external port of acceleration card MC0 of acceleration unit A1. This layered configuration allows each accelerator card to share data via a high-speed serial link while processing data at high speed, enabling unlimited interconnection of accelerator cards to meet customizable computing power requirements and achieve flexible configuration of the hardware computing power of the processor cluster.
[0074] It should be noted that there are various connection relationships among the acceleration units in the acceleration device disclosed herein, which have been described in detail above. For specific details, please refer to examples above. Figure 6The details of the connection relationships between the acceleration units are not repeated here. Furthermore, there are several ways to connect the last acceleration unit to the first acceleration unit, specifically including one or more of the following: connecting the external port of MC0 in acceleration unit A1 to the external port of MC0 in An; connecting the external port of MC1 in acceleration unit A1 to the external port of MC1 in An; connecting the external port of MC2 in acceleration unit A1 to the external port of MC2 in An; and connecting the external port of MC3 in acceleration unit A1 to the external port of MC3 in An. For ease of understanding, the following will combine... Figure 14 and Figure 15 An exemplary description will be provided. In the following description, those skilled in the art will understand that... Figure 14 and Figure 15 The acceleration device shown is Figure 13 The acceleration device 800 shown has various specific manifestations, therefore regarding Figure 13 The description of the accelerator 800 can also be applied to Figure 14 and Figure 15 The acceleration device in the middle.
[0075] refer to Figure 14 , Figure 14 This is a schematic diagram of the network topology corresponding to the acceleration device in one embodiment. For example... Figure 14 The acceleration device 801 shown can consist of four acceleration units. Circles represent acceleration cards, and lines represent link connections. Within each circle, number 0 represents acceleration card MC0, number 1 represents MC1, number 2 represents MC2, and number 3 represents MC3. The numbers embedded in the vertical lines indicate the port numbers of the connections. The last acceleration unit is connected to the first, with a total of 5 hops. Each acceleration unit is a node, and interconnection between nodes allows for the interconnection of 4 nodes (16 cards). The four acceleration units form a small cluster, internally interconnected, called a super pod. This topology is the preferred form for ultra-large-scale clusters, using high-speed SerDes ports, with a total of 5 hops and minimal latency. The cluster exhibits good manageability and robustness.
[0076] refer to Figure 15 , Figure 15 This is a schematic diagram of the network topology corresponding to the acceleration device in another embodiment. Figure 15 and Figure 14 The difference is that, Figure 15 The acceleration device 802 shown has a greater number of acceleration units. As can be seen from the diagram, the last acceleration unit of acceleration device 802 is connected to the first acceleration unit. With this configuration, the total number of hops is the number of nodes plus one, meaning the total number of hops is the number of acceleration units plus one.
[0077] The above combination Figures 13-15 An exemplary acceleration device including multiple acceleration units has been described. According to the technical solution disclosed herein, an acceleration device including multiple of the aforementioned acceleration components is also provided. The following will describe the invention in detail with reference to several embodiments.
[0078] Figure 16 This is a schematic diagram of an acceleration device in another embodiment of this disclosure. The acceleration device 900 may include m of the aforementioned acceleration components. Each acceleration component, in addition to the external ports that need to connect the acceleration units within the component, also has idle external ports. The acceleration components are interconnected through these idle external ports. Specifically, the external port of the acceleration card MC1 of acceleration unit A1 in acceleration component B1 can be connected to the external port of the acceleration card MC1 of acceleration unit A1 in acceleration component B2, the external port of the acceleration card MC1 of acceleration unit A1 in acceleration component B2 can be connected to the external port of the acceleration card MC1 of acceleration unit A1 in acceleration component B3, and so on, with multiple acceleration components interconnected. It is understood that... Figure 16 The acceleration device shown is exemplary and not limiting; for example, the structures of multiple acceleration components may be the same or different. Furthermore, the method by which different acceleration components are connected via unused external ports is not limited to... Figure 16 The methods shown may also include other methods. For ease of understanding, the following will combine... Figures 17-25 An exemplary description is provided.
[0079] based on Figure 16 The provided acceleration device, further, refers to Figure 17 , Figure 17 This is a schematic diagram of the network topology corresponding to the acceleration device in another embodiment. Acceleration device 901 may include two acceleration components. Acceleration component B1 may include four acceleration units, and acceleration component B2 may include four acceleration units. The first acceleration unit in acceleration component B1 is connected to the first acceleration unit in acceleration component B2, and the last acceleration unit in acceleration component B1 is connected to the last acceleration unit in acceleration component B2. The total number of hops in this network topology is 9. Those skilled in the art will understand that... Figure 17 The network structure consisting of multiple acceleration units in each acceleration component is logical; in practical applications, the arrangement of these acceleration units can be adjusted as needed. The number of acceleration units in each acceleration component is not limited to the four shown in the diagram; it can be set to more or fewer as required, such as six or eight.
[0080] based on Figure 16 The provided acceleration device, further, refers to Figure 18 , Figure 18This is a schematic diagram of an acceleration device in another embodiment of this disclosure. Acceleration device 902 may include four acceleration components, namely acceleration components B1, B2, B3, and B4. Each of the four acceleration components may include two acceleration units A1 and A2. Each acceleration component can be interconnected with one of the acceleration units A1 and A2 of the other acceleration components. For example, acceleration unit A1 in acceleration component B1 is connected to acceleration unit A1 in acceleration component B2, acceleration unit A1 in acceleration component B2 is connected to acceleration unit A1 in acceleration component B3, and acceleration unit A1 in acceleration component B3 is connected to acceleration unit A1 in acceleration component B4. These connections are all made through external ports of the acceleration units.
[0081] It should be noted that, in addition to the connection methods between acceleration components, Figure 18 Besides the connection method shown, there can be many other types. For example, the connection methods between acceleration components can specifically include: acceleration unit A1 or A2 in acceleration component B1 is connected to acceleration unit A1 or A2 in acceleration component B2; acceleration unit A1 or A2 in acceleration component B2 is connected to acceleration unit A1 or A2 in acceleration component B3; and acceleration unit A1 or A2 in acceleration component B3 is connected to acceleration unit A1 or A2 in acceleration component B4.
[0082] based on Figure 18 For further information on the provided acceleration device, please refer to [link / reference]. Figure 19 , Figure 19 This is a schematic diagram of the acceleration device in yet another embodiment of this disclosure. Figure 19 In the acceleration device 903 shown, each acceleration component can be interconnected with one of the first acceleration unit and the second acceleration unit of other acceleration components via two paths. For example, the first acceleration unit (e.g., acceleration unit A1) in acceleration component B1 and the first acceleration unit (e.g., acceleration unit A1) in acceleration component B2 can be connected via two paths, acceleration unit A1 in acceleration component B2 and acceleration unit A1 in acceleration component B3 can be connected via two paths, and acceleration unit A1 in acceleration component B3 and acceleration unit A1 in acceleration component B4 can be connected via two paths.
[0083] It should be noted that, Figure 19 The symbol indicates a connection between two paths, but it can also include connections between more than two paths. Besides the methods for accelerating connections between components... Figure 19The connection method shown may also include other methods, such as acceleration unit A1 or A2 in acceleration component B1 being connected to acceleration unit A1 or A2 in acceleration component B2 via two paths, acceleration unit A1 or A2 in acceleration component B2 being connected to acceleration unit A1 or A2 in acceleration component B3 via two paths, and acceleration unit A1 or A2 in acceleration component B3 being connected to acceleration unit A1 or A2 in acceleration component B4 via two paths.
[0084] based on Figure 16 For further information on the provided acceleration device, please refer to [link / reference]. Figure 20 , Figure 20 This is a schematic diagram of an acceleration device in another embodiment of this disclosure. The acceleration device 904 includes four acceleration components: acceleration component B1, acceleration component B2, acceleration component B3, and acceleration component B4. Each acceleration component includes two acceleration units, and each acceleration unit includes two pairs of acceleration cards. In each acceleration unit, MC0 and MC1 form a first pair of acceleration cards, and MC2 and MC3 form a second pair of acceleration cards. Specifically, the second pair of acceleration cards in acceleration unit A1 of acceleration component B1 is connected to the second pair of acceleration cards in acceleration unit A2 of acceleration component B2; the first pair of acceleration cards in acceleration unit A2 of acceleration component B2 is connected to the first pair of acceleration cards in acceleration unit A1 of acceleration component B3; the second pair of acceleration cards in acceleration unit A2 of acceleration component B3 is connected to the second pair of acceleration cards in acceleration unit A1 of acceleration component B4; and the first pair of acceleration cards in acceleration unit A1 of acceleration component B4 is connected to the first pair of acceleration cards in acceleration unit A2 of acceleration component B1.
[0085] refer to Figure 21 , Figure 21 This is a network topology diagram of yet another type of acceleration device. Figure 21 The acceleration device 905 shown is Figure 20 This is a specific embodiment of the acceleration device 904 shown, therefore the above description of the acceleration device 904 can also be applied to... Figure 21 The acceleration device 905 in the middle. For example... Figure 21 As shown, each acceleration component of the acceleration device 905 can form a hybrid three-dimensional network unit. The interconnection relationship within each hybrid three-dimensional network unit can be as shown in the figure, realizing the interconnection of 8 nodes and 32 cards in the acceleration device 905. The four acceleration components can be interconnected through, for example, QSFP-DD interfaces and cables to form a matrix network topology.
[0086] Specifically, in this embodiment, ports 0 of the accelerator cards MC2 and MC3 of the upper-level nodes of acceleration component B1 can be connected to the accelerator cards MC2 and MC3 of the lower-level nodes of acceleration component B2, respectively. Ports 3 of MC0 and MC1 of the lower-level nodes of acceleration component B2 can be connected to MC0 and MC1 of the upper-level nodes of acceleration component B3, respectively. Ports 0 of MC2 and MC3 of the lower-level nodes of acceleration component B3 can be connected to MC2 and MC3 of the upper-level nodes of acceleration component B4, respectively. Ports 3 of MC0 and MC1 of the upper-level nodes of acceleration component B4 can be connected to MC0 and MC1 of the lower-level nodes of acceleration component B1, respectively. This interconnection of the hybrid three-dimensional networks can form two bidirectional ring structures (as described above). Figure 5b , Figure 5c , Figure 12b and Figure 12c As described, it has advantages such as good reliability and security, and is suitable for deep learning training with high computational efficiency. In the accelerator 905, the matrix network topology consisting of 8 nodes has a total of 11 hops.
[0087] Furthermore, such as Figure 21 As shown, the first pair of accelerator cards and the second pair of accelerator cards in different accelerator units within the same accelerator assembly can be indirectly connected. For example, accelerator cards MC0 and MC1 of the upper-level accelerator unit in accelerator assembly B1 are indirectly connected to accelerator cards MC2 and MC3 of the lower-level accelerator unit.
[0088] exist Figure 21 Based on the existing network topology, matrix network topology can be further extended into larger network topologies. Figure 22 This is a schematic diagram of a matrix network topology based on wireless extension of an acceleration device. (Example:) Figure 22 As shown, the acceleration device 906 may include multiple acceleration components, and each acceleration component (shown in a box in the figure) may include multiple acceleration units (not shown in perspective, but can be referenced). Figure 21 The acceleration component structure can include, for example, four interconnected acceleration cards as shown in the figure, so the matrix network topology can theoretically be expanded indefinitely.
[0089] based on Figure 16 For further information on the provided acceleration device, please refer to [link / reference]. Figure 23 , Figure 23This is a schematic diagram of an acceleration device in another embodiment of this disclosure. The acceleration device 908 may include m (m≥2) acceleration components, each acceleration component may include n (n≥2) acceleration units, and the m acceleration components may be connected in a ring. Specifically, the acceleration unit An of acceleration component B1 may be connected to the acceleration unit A1 of acceleration component B2, the acceleration unit An of acceleration component B2 may be connected to the acceleration unit A1 of acceleration component B3, and so on down to acceleration component Bm. The acceleration unit An of acceleration component Bm may be connected to the acceleration unit A1 of acceleration component B1, thus these m acceleration components are connected end-to-end in a ring.
[0090] based on Figure 23 Please refer to Figure 24 , Figure 24 This is a network topology diagram of another acceleration device. The acceleration device 909 may include 6 acceleration components, each acceleration component may include two acceleration units, and the second acceleration unit of each acceleration component may be connected to the first acceleration unit of the next acceleration component, forming an interconnection of 12 nodes and 48 cards, forming a larger matrix network topology. The total Hop under this network topology is 13 times.
[0091] based on Figure 24 Please refer to Figure 25 , Figure 25 This is a network topology diagram of another acceleration device. The acceleration device 910 includes 8 acceleration components, each of which includes two acceleration units. The second acceleration unit of each acceleration component can be connected to the first acceleration unit of the next acceleration component, forming an interconnection of 16 nodes and 64 cards, forming a larger matrix network topology. The total Hop under this network topology is 17.
[0092] exist Figure 25 Based on this, it can be continuously scaled vertically to form ultra-large-scale matrix networks, such as 20 nodes with 80 cards, 24 nodes with 96 cards, etc. Theoretically, it can be expanded indefinitely, with the total number of hops being the number of nodes plus one. By optimizing the interconnection method between nodes, the latency of the entire system can be minimized, maximizing the ability to meet the real-time requirements of the system while processing massive amounts of data.
[0093] The above combination Figures 16-25 An exemplary description has been provided of an acceleration device comprising multiple acceleration components. Those skilled in the art will understand that the above description is exemplary and not restrictive; for example, the number, structure, and connection relationships between the acceleration components can be adjusted as needed. Those skilled in the art can also combine the above embodiments to form an acceleration device as needed, which is also within the scope of this disclosure.
[0094] Additionally, it should be noted that the accelerator card matrix, fully connected square network (topology), hybrid three-dimensional network (topology), matrix network (topology) and other similar terms mentioned in this disclosure are all logical, and the specific deployment can be adjusted as needed.
[0095] The topology disclosed herein can also perform data reduction operations. These reduction operations can be performed on each accelerator card, each accelerator unit, and within the accelerator apparatus. Specific operational steps are as follows.
[0096] Taking reduction summation as an example, the reduction operation process in an acceleration unit may include: transferring the data stored in the first acceleration card to the second acceleration card, and performing addition on the data originally stored in the second acceleration card and the data received from the first acceleration card; then, transferring the addition result in the second acceleration card to the third acceleration card, and performing addition again, and so on, until all the data stored in the acceleration cards has been added and each acceleration card has received the final result.
[0097] by Figure 4 Taking the acceleration unit shown as an example, acceleration card MC0 stores data (0,0), acceleration card MC1 stores data (1,2), acceleration card MC2 stores data (3,1), and acceleration card MC3 stores data (2,4). The data (0,0) in acceleration card MC0 can be passed to acceleration card MC1, and after addition, the result (1,2) is obtained. Next, the result (1,2) is passed to acceleration card MC2 to obtain the next result (4,3). Then, the next result (4,3) is passed to acceleration card MC3 to obtain the final result (6,7).
[0098] Subsequently, in the reduction operation disclosed herein, the final result (6,7) is passed to each accelerator card MC0, MC1, MC2 and MC3, so that data (6,7) is stored in all accelerator cards, thus completing the reduction operation in one accelerator unit.
[0099] Figure 4 The acceleration unit can form two independent rings, each of which can complete half of the data reduction operations, thereby speeding up the operation and improving the efficiency.
[0100] In addition, the aforementioned acceleration unit can also achieve concurrent computation of multiple acceleration cards during reduction operations, thereby accelerating the computation speed. For example, acceleration card MC0 stores data (0,0), acceleration card MC1 stores data (1,2), acceleration card MC2 stores data (3,1), and acceleration card MC3 stores data (2,4). Part of the data (0) in acceleration card MC0 can be transferred to acceleration card MC1, and after addition, the result (1) is obtained. Simultaneously, part of the data (2) in acceleration card MC1 can be transferred to acceleration card MC2, and after addition, the result (3) is obtained. This achieves concurrent computation between acceleration cards MC1 and MC2. And so on, to complete the entire reduction operation.
[0101] The aforementioned concurrent computation can also include grouping acceleration units to perform addition operations first, and then performing a reduction operation between the results of the current group of acceleration units and the results of the operations of another group of acceleration units. For example, if acceleration card MC0 stores data (0,0), acceleration card MC1 stores data (1,2), acceleration card MC2 stores data (3,1), and acceleration card MC3 stores data (2,4), the data in acceleration card MC0 can be transferred to acceleration card MC1 for computation to obtain the first set of results (1,2); synchronously or asynchronously, the data in acceleration card MC2 can be transferred to acceleration card MC3 for computation to obtain the second set of results (5,5). Next, the first set of results and the second set of results are then processed to obtain the final reduction result (6,7).
[0102] Similarly, in addition to performing reduction operations within an acceleration unit, reduction operations can also be performed within an acceleration component or acceleration device. It should be understood that an acceleration device can also be considered as an acceleration component connected end-to-end.
[0103] When performing reduction operations in an acceleration component or acceleration device, it may include: performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result in each acceleration unit; and performing a second reduction operation on the first reduction results in multiple acceleration units to obtain a second reduction result.
[0104] Taking reduction summation as an example, the first step mentioned above has already been described. For an acceleration device that includes multiple acceleration units, local reduction operations can be performed in each acceleration unit first. After the reduction operation in each acceleration unit is completed, the acceleration card in the same acceleration unit will obtain the result of the local reduction operation, which is referred to as the first reduction result.
[0105] Next, the first reduction results from all acceleration units can be passed between adjacent acceleration units and added together. Thus, similar to performing a reduction operation within a single acceleration unit, the first acceleration unit passes the first reduction result to the second acceleration unit, where addition operations are performed on the acceleration cards. Then, the result is passed and added again. After the final addition operation, the final result is passed to each acceleration unit.
[0106] It should be noted that, since the acceleration components mentioned above are not necessarily connected end-to-end, the final result can be transmitted in reverse order to each acceleration unit, rather than in a loop as when the acceleration units are connected end-to-end. The technical solution disclosed herein does not impose specific limitations on how the final result is transmitted.
[0107] Furthermore, according to one embodiment of this disclosure, the acceleration device may also be configured to perform reduction operations, including: performing a first reduction operation on data in the acceleration card of the same acceleration unit to obtain a first reduction result; performing an intermediate reduction operation on the first reduction result in multiple acceleration units of the same acceleration component to obtain an intermediate reduction result; and performing a second reduction operation on the intermediate reduction result in multiple acceleration components to obtain a second reduction result.
[0108] In this implementation, reduction operations can be performed in the same acceleration unit first, as described above, and will not be repeated here.
[0109] Next, reduction operations can be performed in each acceleration component, so that each accelerator card in each acceleration component obtains the local reduction result of its own acceleration component. Then, reduction operations are performed in multiple acceleration components, so that each accelerator card obtains the global reduction result of the acceleration device.
[0110] It should be understood that the above transmission order is only for the convenience of description and is not necessarily the correct transmission order. Figure 26 This is a schematic diagram of a combined processing device structure in one embodiment of this disclosure. As shown in the figure, the combined processing device 2600 may include an acceleration unit 2601, specifically the acceleration unit shown in Figures 1 to 5. Additionally, the combined processing device may also include an interconnection interface 2602 and other processing devices 2603. According to this disclosure, the acceleration unit 2601 can interact with other processing devices 2603 through the interconnection interface 2602 to jointly complete user-specified operations.
[0111] According to the scheme disclosed herein, the other processing device may include one or more types of processors such as microprocessor units (MCUs), board controllers (BMCs), and central processing units, and the number of such processors is not limited but determined according to actual needs. In one or more embodiments, the other processing device may serve as an interface between the acceleration unit disclosed herein and external data and control, performing tasks including but not limited to data transfer, and completing basic control such as starting and stopping the acceleration unit; the other processing device may also cooperate with the acceleration unit to jointly complete computational tasks.
[0112] Optionally, the combined processing apparatus 2600 may further include a storage device 2604, which may be connected to the acceleration unit 2601, the interconnect interface 2602, and other processing devices 2603, respectively. In one or more embodiments, the storage device 2604 may be used to store data of the acceleration unit 2601 and other processing devices 2603, especially data that cannot be fully stored in the internal or on-chip storage of the acceleration unit 2601 and other processing devices 2603.
[0113] In some application scenarios, the combined processing device 2600 disclosed herein can be used in, for example, large-scale data centers, supercomputing centers, cloud computing centers, etc., to build high-performance processor clusters, thereby realizing real-time processing of massive amounts of data.
[0114] In some embodiments, this disclosure also discloses a circuit board that may include the aforementioned acceleration unit. See also... Figure 27 The present invention provides an exemplary circuit board 2700, which, in addition to including one or more acceleration units 2706 (two are shown in the figure), may also include other supporting components, including but not limited to: storage device 2701, interface device 2707 and controller 2705.
[0115] The storage device 2701 can be connected to the acceleration unit 2706 via a bus for storing data. The storage device 2701 may include multiple sets of storage units 2702. Each set of storage units 2702 can be connected to the acceleration unit 2706 via a bus. It is understood that each set of storage units 2702 can be at least one of DDR SDRAM (Double Data Rate SDRAM), HBM (High Bandwidth Memory), etc.
[0116] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read on both the rising and falling edges of the clock pulse. DDR is twice as fast as standard SDRAM. In one embodiment, the memory device 2701 may include four groups of memory cells 2702. Each group of memory cells 2702 may include multiple DDR4 chips. In one embodiment, the chip may internally include four 72-bit DDR4 controllers, of which 64 bits are used for data transmission and 8 bits are used for ECC verification.
[0117] In one embodiment, each group of storage units 2702 may include multiple parallel-connected Double Data Rate (DDR) synchronous dynamic random access memories (DDRs). DDRs can transmit data twice within one clock cycle. A controller for controlling the DDRs is provided in the acceleration unit 2706 for controlling data transmission and data storage in each storage unit. The interface device 2707 can be connected to the acceleration unit 2706. The interface device 2707 is used to realize data transmission between the acceleration unit 2706 and an external device 2708 (e.g., a server or computer). For example, in one embodiment, the interface device 2707 can be a standard PCIe interface. For instance, data to be processed is transferred from the server to the acceleration unit 2706 via a standard PCIe interface, realizing data transfer. In another embodiment, the interface device 2707 can also be other interfaces; this disclosure does not limit the specific form of the other interfaces mentioned above, as long as the interface device can realize the switching function. Furthermore, the calculation results of the acceleration unit 2706 can still be transmitted back to the external device (e.g., the server) by the interface device 2707. The controller 2705 can be connected to the acceleration unit 2706. The controller 2705 can be used to monitor the status of the acceleration unit 2706. Specifically, the acceleration unit 2706 and the controller 2705 can be electrically connected via an SPI interface. The controller 2705 may include a microcontroller (MCU).
[0118] In some embodiments, this disclosure also discloses an electronic device or apparatus that includes the aforementioned acceleration unit. In some embodiments, this disclosure also discloses yet another electronic device or apparatus that includes the aforementioned acceleration component. In some embodiments, this disclosure also discloses yet another electronic device or apparatus that includes the aforementioned acceleration device. In some embodiments, this disclosure also discloses yet another electronic device or apparatus that includes the aforementioned circuit board.
[0119] Depending on the application scenario, electronic devices or apparatuses may include, for example, data processing devices, data centers, supercomputing centers, cloud computing centers, servers, and cloud servers.
[0120] In the above embodiments disclosed herein, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant descriptions in other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as these combinations of technical features do not contradict each other, they should be considered within the scope of this specification.
[0121] The foregoing can be better understood in accordance with the following terms:
[0122] Clause 1. An acceleration unit comprising M local acceleration cards, each local acceleration card including an internal port, each local acceleration card being connected to other local acceleration cards via the internal port, wherein,
[0123] The M accelerator cards in this unit are logically formed into an L*N accelerator card matrix, where L and N are integers not less than 2.
[0124] Clause 2. The acceleration unit as described in Clause 1, wherein the M acceleration cards of this unit are logically formed as a 2*2, 3*3 or 4*4 acceleration card matrix.
[0125] Clause 3. The acceleration unit as described in Clause 1 or 2, wherein each of the unit's acceleration cards is connected to at least one other unit's acceleration card via two paths.
[0126] Clause 4. An acceleration unit according to any one of Clauses 1-3, wherein the diagonal acceleration cards of the unit located at the four corners of the acceleration card matrix are connected by two paths.
[0127] Clause 5. The acceleration unit according to any one of Clauses 1-4, wherein at least one of the M acceleration cards of this unit includes an external port.
[0128] Clause 6. An acceleration unit according to any one of Clauses 1-5, wherein when the acceleration unit includes four local acceleration cards, each local acceleration card includes six ports, and wherein four ports of each local acceleration card are internal ports for connecting to the other three local acceleration cards; and the remaining two ports of at least one local acceleration card are external ports for connecting to external acceleration cards.
[0129] Clause 7. The acceleration unit according to any one of Clauses 1-6, wherein the internal port and the external port are SerDes ports.
[0130] Clause 8. An acceleration unit according to any one of Clauses 1-7, wherein the acceleration unit is configured to perform a reduction operation on data in the acceleration card of the acceleration unit to obtain a reduction result.
[0131] Clause 9. The acceleration unit as described in Clause 4, wherein the acceleration card matrix comprises two independent rings, each ring performing a portion of the computation.
[0132] Clause 10. An electronic device comprising an acceleration unit as described in any one of Clauses 1-9.
[0133] It should be understood that the terms "first," "second," "third," and "fourth," etc., in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0134] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0135] The embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this disclosure. Furthermore, any changes or modifications made by those skilled in the art based on the ideas of this disclosure, and on the specific implementation methods and application scope of this disclosure, are all within the scope of protection of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.
Claims
1. An acceleration unit, comprising M accelerator cards, each accelerator card including an internal port, each accelerator card being connected to other accelerator cards via the internal port, wherein, The M accelerator cards in this unit are logically formed into an accelerator card matrix of size L*N, where L and N are integers not less than 2; In the accelerator card matrix, the accelerator cards located at the four corners of the unit are connected by two paths, and the accelerator card matrix includes two independent rings.
2. The acceleration unit according to claim 1, wherein, The M accelerator cards in this unit are logically arranged into an accelerator card matrix of 2*2, 3*3, or 4*4.
3. The acceleration unit according to claim 1 or 2, wherein, Each local accelerator card is connected to at least one other local accelerator card via two paths.
4. The acceleration unit according to claim 1, wherein, At least one of the M accelerator cards in this unit includes an external port.
5. The acceleration unit according to claim 1, wherein, When the acceleration unit includes four local acceleration cards, each local acceleration card includes six ports, and four of the ports of each local acceleration card are internal ports for connecting to the other three local acceleration cards; the remaining two ports of at least one local acceleration card are external ports for connecting to external acceleration cards.
6. The acceleration unit according to claim 1, wherein, The internal and external ports are SerDes ports.
7. The acceleration unit according to claim 1, wherein, The acceleration unit is configured to perform reduction operations on the data in the acceleration card of the acceleration unit to obtain the reduction result.
8. The acceleration unit according to claim 1, wherein, Each ring performs a portion of the computation, and the data reduction operation is performed in the acceleration unit, with each ring completing half of the data reduction operation.
9. An electronic device comprising the acceleration unit as described in any one of claims 1-8.
Citation Information
Patent Citations
Data acceleration processing system
CN110413561A
Acceleration unit and electronic device
CN212846786U