Acceleration component, acceleration device and electronic equipment

By building acceleration units and acceleration components in multiple ASIC acceleration chips, and using inline and external port connections to form an acceleration card matrix and ring structure, the challenge of data transmission bandwidth in multiple ASIC chip interconnection networks is solved, and efficient massive data processing is achieved.

CN114185832BActive Publication Date: 2025-08-29ANHUI CAMBRICON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010970850.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-15
Publication Date
2025-08-29
Estimated Expiration
2040-09-15

AI Technical Summary

Technical Problem

In the interconnection network of multiple ASIC acceleration chips, high data throughput poses a major challenge to data transmission bandwidth, and how to design interconnection solutions between chips to improve the computing power of the entire system has become a key issue.

Method used

Multiple acceleration units are adopted, each acceleration unit includes multiple acceleration cards, which are connected through internal and external ports to form a logically L*N-scale acceleration card matrix, and multiple acceleration units are connected through external ports to build a multi-layer structure and a ring structure to realize interconnection between acceleration cards.

Benefits of technology

The computing power of the acceleration unit is improved, communication delay is reduced, real-time requirements for processing massive data are met, and the computing power and data processing speed of the system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114185832B_ABST
    Figure CN114185832B_ABST
Patent Text Reader

Abstract

This disclosure relates to an acceleration unit, an acceleration component, an acceleration device, and an electronic device. The acceleration unit may be included in a combined processing device, which may also include an interconnection interface and other processing devices. The acceleration unit interacts with the other processing devices to jointly complete user-specified computing operations. The combined processing device may also include a storage device, which is connected to the acceleration unit and the other processing devices, respectively, to serve data from the acceleration unit and the other processing devices. By utilizing the contents of this disclosure, high-speed processing of massive amounts of data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of processor technology. More specifically, the present disclosure relates to an acceleration unit, an acceleration component, an acceleration device, a board, and an electronic device. Background Art

[0002] With the rapid development of artificial intelligence (AI) and machine learning, the demand for ultra-high-performance processors will continue to grow. The big data era also places even higher demands on data processing. High-performance processors and clusters must handle massive amounts of data in real time, training and inferencing complex models within a specified timeframe. ASICs (Application Specific Integrated Circuits) are specialized acceleration chips that can be used to train deep neural networks. ASICs can complete tasks in a fraction of the time, using significantly less data center infrastructure than non-parallel processing supercomputers.

[0003] However, even the most powerful individual ASICs are often insufficient when faced with massive amounts of data. To achieve even greater computing power, a common approach is to use multiple ASIC accelerator chips. However, the ultra-high data throughput of interconnected ASICs in a multi-card network poses significant challenges to the ASICs' data transmission bandwidth. Therefore, designing interconnection schemes between chips to improve the overall system's computing power and efficiently process massive amounts of data has become a key technical challenge in building high-performance processor clusters. Summary of the Invention

[0004] In order to solve the above technical problems, the present disclosure provides an acceleration unit, an acceleration component, an acceleration device, a board and an electronic device that can improve computing power.

[0005] In one aspect, the present disclosure provides an acceleration component, comprising multiple acceleration units, each acceleration unit comprising M local unit acceleration cards, each local unit acceleration card comprising an internal port, each local unit acceleration card being connected to other local unit acceleration cards through the internal port, wherein the M local unit acceleration cards logically form an acceleration card matrix of an L*N scale, L and N being integers not less than 2, at least one of the M local unit acceleration cards comprising an external port, and the acceleration units being connected through the external ports.

[0006] In another aspect, the present disclosure provides an acceleration device, comprising multiple acceleration units, each acceleration unit comprising M local unit acceleration cards, each local unit acceleration card comprising an internal port, each local unit acceleration card being connected to other local unit acceleration cards through the internal port, wherein the M local unit acceleration cards logically form an acceleration card matrix of L*N scale, L and N being integers not less than 2, at least one of the M local unit acceleration cards comprising an external port, the acceleration units being connected to each other through the external ports, the multiple acceleration units logically forming a multi-layer structure, each layer comprising an acceleration unit, the local unit acceleration card of each acceleration unit being connected to the external unit acceleration card through the external port, and the last acceleration unit being connected to the first acceleration unit, so that the multiple acceleration units are connected end to end to form a ring structure.

[0007] In yet another aspect, the present disclosure provides an acceleration device, comprising a plurality of acceleration components as described above, wherein the acceleration components are interconnected via idle external ports.

[0008] In another aspect, the present disclosure provides an electronic device comprising the acceleration component as described above, or the acceleration device as described above.

[0009] In the disclosed solution, the acceleration unit is composed of multiple accelerator cards. Each accelerator card is connected to other accelerator cards via its internal port, achieving interconnection between the accelerator cards. This configuration effectively improves the computing power of the acceleration unit and facilitates increased processing speed for massive amounts of data. Furthermore, for the acceleration components and acceleration devices, the interconnection between the acceleration units minimizes latency across the entire system, maximizing the system's real-time requirements for processing massive amounts of data. This improves the computing power of the entire system and enables the system to process massive amounts of data at high speeds. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0011] Figure 1a To disclose a schematic diagram of the acceleration unit structure in one embodiment

[0012] Figure 1b 、 Figure 2 、 Figure 3 、 Figure 4 as well as Figure 5a-5c are multiple structural schematic diagrams of the acceleration unit according to the embodiment of the present disclosure;

[0013] Figures 6-11 are multiple structural schematic diagrams of acceleration components according to an embodiment of the present disclosure;

[0014] Figure 12a-12c A schematic diagram showing the acceleration components as a network topology;

[0015] Figure 13 A schematic diagram of an acceleration device including a plurality of acceleration units according to an embodiment of the present disclosure;

[0016] Figure 14 A schematic diagram of a network topology corresponding to an acceleration device in one embodiment;

[0017] Figure 15 A schematic diagram of a network topology corresponding to an acceleration device in another embodiment;

[0018] Figures 16-20 are multiple schematic diagrams of an acceleration device including multiple acceleration components according to an embodiment of the present disclosure;

[0019] Figure 21 A schematic diagram of the network topology of another acceleration device;

[0020] Figure 22 This is a schematic diagram of the matrix network topology based on wireless expansion of the acceleration device;

[0021] Figure 23 This is a schematic diagram of an acceleration device in another embodiment of the present disclosure;

[0022] Figure 24 A schematic diagram of the network topology of another acceleration device;

[0023] Figure 25 A schematic diagram of the network topology of another acceleration device;

[0024] Figure 26 This is a schematic structural diagram of a combined device in one embodiment of the present disclosure;

[0025] Figure 27 This is a schematic diagram of the structure of a board in one embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.

[0027] Several embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0028] Figure 1a A schematic diagram of an acceleration unit structure in one embodiment is disclosed. According to one embodiment of the present disclosure, an acceleration unit is provided, comprising M local unit acceleration cards, each local unit acceleration card including an internal port, each local unit acceleration card connected to other local unit acceleration cards via the internal port. The M local unit acceleration cards logically form an L*N acceleration card matrix, where L and N are integers not less than 2.

[0029] like Figure 1a As shown, multiple accelerator cards can be used to form an accelerator card matrix, and the accelerator cards are connected to each other, so that data or instructions can be transmitted and communicated. 00 To MC 0N Forming the 0th row of the accelerator card matrix, accelerator card MC 10 To MC 1N Forming the first row of the accelerator card matrix, and so on, the accelerator card MC L0 To MC LN This forms the Lth row of the accelerator card matrix.

[0030] It should be understood that, for ease of understanding in the context, accelerator cards in the same acceleration unit are referred to as "local unit accelerator cards," while accelerator cards in other acceleration units are referred to as "external unit accelerator cards." Such designations are merely for ease of description and do not limit the technical solution disclosed herein.

[0031] Each accelerator card can have multiple ports, which can connect to the accelerator card of the same unit or to an external accelerator card. In this disclosure, the connection ports between the accelerator cards of the same unit can be referred to as internal ports, while the connection ports between the accelerator card of the same unit and an external accelerator card can be referred to as external ports. It should be understood that the terms external ports and internal ports are merely for convenience of description, and both can use the same port. This will be described below.

[0032] It should be understood that M can be any integer, and the M accelerator cards can be formed into a 1*M or M*1 matrix, or the M matrices can be formed into other types of matrices. The acceleration unit disclosed herein is not limited to a specific matrix size and form.

[0033] Furthermore, accelerator cards, for example, accelerator cards in one unit, or accelerator cards in one unit and accelerator cards in another unit, can be connected via a single or multiple communication paths, which will be described in detail later.

[0034] It's also important to understand that, although the positions of multiple accelerator cards are described as a rectangular network in the context of this disclosure, the resulting matrix is ​​not necessarily physically arranged in a matrix. It can be arranged in any position, for example, multiple accelerator cards can form a straight line or be arranged irregularly. The aforementioned matrix is ​​merely a logical concept; as long as the connections between the accelerator cards form a matrix relationship, it will suffice.

[0035] According to one embodiment of the present disclosure, M can be 4, whereby four accelerator cards of this unit can logically form a 2*2 accelerator card matrix; M can be 9, whereby nine accelerator cards of this unit can logically form a 3*3 accelerator card matrix; M can be 16, whereby 16 accelerator cards of this unit can logically form a 4*4 accelerator card matrix. M can also be 6, whereby six accelerator cards of this unit can logically form a 2*3 or 3*2 accelerator card matrix; M can also be 8, whereby eight accelerator cards of this unit can logically form a 2*4 or 4*2 accelerator card matrix.

[0036] According to one embodiment of the present disclosure, each local unit acceleration card is connected to at least one other local unit acceleration card through two paths.

[0037] In the topology described in this disclosure, two local accelerator cards can be connected via a single communication path or via multiple (e.g., two) paths, as long as the number of ports is sufficient. Using multiple communication paths helps ensure reliable communication between accelerator cards, which will be explained and described in more detail in the examples below.

[0038] According to one embodiment of the present disclosure, diagonal accelerator cards in the four corners of the accelerator card matrix are connected via two paths. For a matrix, it is preferable to connect two pairs of accelerator cards at opposite corners. For certain topologies, connecting diagonally positioned accelerator cards helps form two complete communication loops. This will be explained and described in more detail in the examples below.

[0039] More specifically, according to one embodiment of the present disclosure, at least one of the local accelerator cards may include an external port. For example, each acceleration unit may include four local accelerator cards, each of which may include six ports. Four of the ports on each local accelerator card are internal ports for connecting to the other three local accelerator cards, while the remaining two ports on at least one local accelerator card are external ports for connecting to an external accelerator card.

[0040] It should be understood that of the six ports on each accelerator card, four can be used to connect to the accelerator card in that unit, while the remaining two ports can be used to connect to accelerator cards in other accelerator units. These remaining ports can also be idle, not connected to any external devices, or can be directly or indirectly connected to other devices or ports.

[0041] For the purpose of illustration and simplicity, the following description of the acceleration unit, acceleration assembly acceleration device, and electronic device is based on the example of each acceleration unit including four acceleration cards. It should be understood that each acceleration unit may include more or fewer acceleration cards.

[0042] For ease of description, the acceleration unit may include four acceleration cards, namely a first acceleration card, a second acceleration card, a third acceleration card and a fourth acceleration card. Each acceleration card is provided with an internal port and an external port, and each acceleration card is connected to the other three acceleration cards through the internal port.

[0043] Figure 1b This is a schematic diagram of the acceleration unit structure in one embodiment of the present disclosure. The acceleration unit 100 includes four acceleration cards, which are acceleration card MC0, acceleration card MC1, acceleration card MC2 and acceleration card MC3. For the four acceleration cards, each acceleration card may include an external port and an internal port. The internal port of acceleration card MC0 is connected to the internal ports of acceleration cards MC1, MC2 and MC3, the internal port of acceleration card MC1 is connected to the internal ports of acceleration cards MC2 and MC3, and the internal port of acceleration card MC2 is connected to the internal port of acceleration card MC3, that is, the internal port of each acceleration card is connected to the internal ports of the other three acceleration cards. Information interaction between the four acceleration cards can be achieved through the interconnection of the internal ports of the four acceleration cards. The embodiment of the present disclosure utilizes the interconnection between the four acceleration cards in the acceleration unit to improve the computing power of the acceleration unit and achieve the purpose of high-speed processing of massive data, and makes the path between each acceleration card and other acceleration cards as short as possible and the communication delay as low as possible.

[0044] As described above, the number of accelerator cards in the present disclosure may not be limited to four, but may be other numbers. For example, in one embodiment, the number N of accelerator cards is equal to 3, and each accelerator card is provided with an internal port and an external port, and each accelerator card is connected to the other two accelerator cards through the internal port, so as to realize the interconnection between the three accelerator cards. In another embodiment, the number N of accelerator cards is equal to 5, and each accelerator card is provided with an internal port and an external port, and each accelerator card is connected to the other four accelerator cards through the internal port, so as to realize the interconnection between the five accelerator cards, thereby improving the computing power of the acceleration unit and realizing high-speed processing of massive data. In another embodiment, the number N of accelerator cards is greater than 5, and each accelerator card is provided with an internal port and an external port, and each accelerator card is connected to all other accelerator cards through the internal port, so as to realize the interconnection between the N accelerator cards and realize high-speed processing of massive data.

[0045] based on Figure 1b The provided acceleration unit 100, further, each acceleration card can be connected to at least one other acceleration card through two paths. Specifically, there can be, for example, three connection methods: the first connection method is that each acceleration card can be connected to one of the other three acceleration cards through two paths; the second method is that each acceleration card can be connected to two of the other three acceleration cards through two paths; the third method is that each acceleration card can be connected to the other three acceleration cards through two paths. In this case, it is not ruled out that each acceleration card has more ports. In order to facilitate understanding of the above-mentioned connection method of the two paths, the following will take the first connection method as an example and combine it with Figure 2 An exemplary description is given.

[0046] Figure 2 This is a schematic diagram of the acceleration unit structure in another embodiment of the present disclosure. Figure 2 In the illustrated acceleration unit 200, each accelerator card can connect to at least one other accelerator card via two paths. For example, accelerator cards MC0 and MC2 can be connected via two paths, and accelerator cards MC1 and MC3 can be connected via two paths. This configuration allows two accelerator cards to exchange information over two links (or paths). This ensures that if one link fails, another link remains between the two accelerator cards, effectively improving the safety of the acceleration unit.

[0047] The above is combined with Figure 1 and Figure 2 The connection method between the acceleration unit and its multiple acceleration cards according to the present disclosure is described in an exemplary manner. It should be understood by those skilled in the art that the above description is exemplary and not restrictive. For example, the arrangement of the acceleration cards in the acceleration unit may not be limited to that shown in Figures 1 and 2. Figure 2 In the form shown in FIG, in one embodiment, the four acceleration cards of the acceleration unit can be logically arranged in a quadrilateral arrangement. Figure 3 Provide a description.

[0048] Figure 3 FIG. 1 is a schematic diagram of the acceleration unit structure in another embodiment of the present disclosure. Figure 3 In the acceleration unit 300 shown, the four acceleration cards MC0, MC1, MC2, and MC3 can be logically arranged in a quadrilateral arrangement, with the four acceleration cards occupying the four vertices of the quadrilateral. The lines between the acceleration cards MC0, MC1, MC2, and MC3 are arranged in a quadrilateral, making the line arrangement clearer and facilitating line setting. It should be noted that Figure 3 The four accelerator cards shown in the figure are arranged in a rectangular or 2*2 matrix, but this is a logical interconnection diagram. It is drawn in the form of a rectangle for the convenience of description. The specific quadrilateral can be freely set, such as a parallelogram, a trapezoid, a square, etc. In the actual layout and wiring, the four accelerator cards can also be arranged in any arrangement. For example, in an actual complete machine, the four accelerator cards are arranged in a straight line in parallel, and the order can be MC0, MC1, MC2, and MC3. It should also be understood that the logical quadrilateral described in this embodiment is exemplary. In fact, the arrangement shape of multiple accelerator cards can be varied, and the quadrilateral is only one of them. For example, when the number of accelerator cards is five, they can be logically arranged in a pentagon.

[0049] based on Figure 2 The connection relationship of the acceleration unit 200 is provided, further see Figure 4 , Figure 4 FIG. 1 is a schematic diagram of the acceleration unit structure in another embodiment of the present disclosure. Figure 4 In the illustrated acceleration unit 400, the four accelerator cards MC0, MC1, MC2, and MC3 can be logically arranged in a quadrilateral, with each accelerator card occupying one of the four vertices of the quadrilateral. As further illustrated, the internal port of accelerator card MC1 can be connected to the internal port of accelerator card MC3 via two paths, and the internal port of accelerator card MC0 can be connected to the internal port of accelerator card MC2 via two paths. This facilitates circuit configuration and improves security for acceleration unit 400.

[0050] Figure 5a This is a schematic diagram of the acceleration unit structure in one embodiment of the present disclosure. Figure 5aIn the illustrated acceleration unit 500, the numbers on each accelerator card represent ports. Each accelerator card can include six ports: Port 0, Port 1, Port 2, Port 3, Port 4, and Port 5. Port 1, Port 2, Port 4, and Port 5 are internal ports, while Port 0 and Port 3 are external ports. For the four accelerator cards MC0, MC1, MC2, and MC3, each accelerator card's two external ports can be connected to other acceleration units, enabling interconnection between multiple acceleration units. Each accelerator card's four internal ports can be used to interconnect with the other three accelerator cards in the same acceleration unit.

[0051] like Figure 5a As further shown in FIG, the four accelerator cards can be logically arranged in a quadrilateral, for example. Accelerator card MC0 and accelerator card MC2 can be in a diagonal relationship, with port 2 of MC0 connected to port 2 of MC2, and port 5 of MC0 connected to port 5 of MC2. That is, two links can be used for communication between accelerator card MC0 and accelerator card MC2. Accelerator card MC1 and accelerator card MC3 can be in a diagonal relationship, with port 2 of MC1 connected to port 2 of MC3, and port 5 of MC1 connected to port 5 of MC3. That is, two links can be used for communication between accelerator card MC1 and accelerator card MC3.

[0052] According to this configuration, since each accelerator card has two external ports and four internal ports, and in two pairs of accelerator cards in a diagonal relationship, the two accelerator cards in each pair can be connected using two internal ports to form two links, the safety and stability of the acceleration unit can be effectively improved. In addition, the four accelerator cards are logically arranged in a quadrilateral, which makes the wiring layout of the entire acceleration unit reasonable and clear, and facilitates the wiring operations within each acceleration unit. It should be further noted that, if Figure 5b Among the interconnection lines between the four accelerator cards shown in FIG, the connection line between port 1 of accelerator card MC1 and port 1 of MC0, the connection line between port 2 of accelerator card MC0 and port 2 of MC2, the connection line between port 1 of accelerator card MC2 and port 1 of MC3, and the connection line between port 2 of accelerator card MC3 and port 2 of MC1 form a vertical figure 8 network, as shown in FIG. Figure 5b As shown. The connection line between port 4 of accelerator card MC1 and port 4 of MC2, the connection line between port 5 of accelerator card MC2 and port 5 of MC0, the connection line between port 4 of accelerator card MC0 and port 4 of MC3, and the connection line between port 5 of accelerator card MC3 and port 5 of MC1 form a horizontal figure 8 network, as shown. Figure 5c As shown in Figure 2, two fully connected square networks can form a dual-ring structure, which provides redundancy and enhances system reliability.

[0053] According to one embodiment of the present disclosure, the acceleration card described in the present disclosure may be a Mezzanine Card (MC card for short), which may be a separate circuit board. The MC card may be equipped with an ASIC chip and some necessary peripheral control circuits. The MC card may be connected to the baseboard through a gusset connector. The power supply and control signals on the baseboard may be transmitted to the MC card through the gusset connector. According to another embodiment of the present disclosure, the internal port and / or external port described in the present disclosure may be a SerDes port. For example, in one embodiment, each MC card may provide 6 bidirectional SerDes ports, each SerDes port having 8 channels and a data transmission rate of 56Gbps, and the total bandwidth of each port may be as high as 400Gbps, which may support massive data exchange between acceleration cards and help the acceleration unit process massive data at high speed.

[0054] The SerDes mentioned above is a compound word of the English words serializer and deserializer, and is called a serial deserializer. The SerDes interface can be used to build a high-performance processor cluster. The main function of Serdes is to convert multiple low-speed parallel signals into serial signals at the transmitting end, transmit them through the transmission medium, and finally convert the high-speed serial signals back into low-speed parallel signals at the receiving end. Therefore, it is very suitable for end-to-end long-distance high-speed transmission requirements. In another embodiment, the external port in the acceleration card can be connected to the QSFP-DD interface of other acceleration units, wherein the QSFP-DD interface is an optical module interface commonly used in SerDes technology, which can be used in conjunction with a cable to interconnect with other external devices.

[0055] Furthermore, according to another embodiment of the present disclosure, a single acceleration unit can accommodate four accelerator cards, interconnected by printed circuit board (PCB) wiring. On a high-speed, low-dielectric-constant substrate, proper layout and routing maximize signal integrity, ensuring the communication bandwidth between the four accelerator cards approaches the theoretical value.

[0056] The acceleration unit disclosed in the present disclosure has four acceleration cards inside the acceleration unit. Each acceleration card is connected to the other three acceleration cards through the internal port of the acceleration card. Each acceleration card can communicate directly with the other three acceleration cards. Such a communication architecture is a fully connected square network topology (fully connected quad). The advantage of this fully connected network architecture is that the path between each acceleration card and other acceleration cards is the shortest, the total number of hops is the smallest, and the delay is the lowest. The present disclosure uses Hop to describe the system's delay. Hop represents the number of hops in communication, that is, the number of communications. Hop specifically represents the shortest path from a node to the initial node after traversing all the nodes in the network. The four acceleration cards are interconnected, and the fully connected square network topology formed has the shortest delay, and the dual-ring structure formed by the interconnection of the two diagonal acceleration cards can improve the robustness of the system, and the business can still operate normally when a single acceleration card fails. When performing various arithmetic and logical operations, each ring in the dual-ring structure can complete a part of the operation separately, thereby improving the overall operation efficiency and maximizing the use of the topology bandwidth.

[0057] Combination of the above Figure 1a-Figure 5c Multiple embodiments of the acceleration unit according to the present disclosure are described. Based on the above-mentioned acceleration unit, the present disclosure also discloses an acceleration component that can include multiple of the above-mentioned acceleration units. The following will provide an exemplary description in combination with multiple embodiments of the acceleration component.

[0058] Figure 6 This is a schematic diagram of the acceleration component structure in one embodiment of the present disclosure. Figure 6 As shown in , the acceleration assembly 600 may include n acceleration units as described above, namely, acceleration unit A1, acceleration unit A2, acceleration unit A3, ..., acceleration unit An, wherein acceleration unit A1 and acceleration unit A2 are connected via an external port, and acceleration unit A2 and acceleration unit A3 are connected via an external port, i.e., each acceleration unit is connected via an external port of each acceleration unit. In one embodiment, the external port of the acceleration card MC0 in acceleration unit A1 can be connected to the external port of the acceleration card MC0 in acceleration unit A2, and the external port of the acceleration card MC0 in acceleration unit A2 can be connected to the external port of the acceleration card MC0 in acceleration unit A3, i.e., each acceleration unit is connected via the external port of the acceleration card MC0.

[0059] Those skilled in the art will appreciate that the connection between acceleration units in the present disclosure may not be limited to the connection of the external port of the acceleration card MC0, but may also include, for example, one or more of the connection of the external port of the acceleration card MC1, the connection of the external port of the acceleration card MC2, and the connection of the external port of the acceleration card MC3. That is, in the present disclosure, the connection method between the acceleration unit A1 and the acceleration unit A2 may include: the external port of MC0 in A1 is connected to the external port of MC0 in A2, the external port of MC1 in A1 is connected to the external port of MC1 in A2, the external port of MC2 in A1 is connected to the external port of MC2 in A2, and the external port of MC3 in A1 is connected to the external port of MC3 in A2. Similarly, the connection method between acceleration unit A2 and acceleration unit A3 may include: the external port of MC0 in A2 is connected to the external port of MC0 in A3, the external port of MC1 in A2 is connected to the external port of MC1 in A3, the external port of MC2 in A2 is connected to the external port of MC2 in A3, and the external port of MC3 in A2 is connected to the external port of MC3 in A3. The same can be said for the connection between acceleration unit An-1 and acceleration unit An. It should be noted that the above description is exemplary. For example, the connection between different acceleration units is not limited to the connection of acceleration cards with corresponding numbers, and can be set to the connection of acceleration cards with the same numbers as needed.

[0060] It should be noted that Figure 6 The figure shows n acceleration units, where n is greater than 3. However, the number of acceleration units is not limited to greater than 3 as shown in the figure, and can also be set to, for example, 2 or 3. The connection relationship between two acceleration units is the same as or similar to the connection relationship between the above-mentioned acceleration units A1 and A2, and the connection relationship between three acceleration units is the same as or similar to the connection relationship between the above-mentioned acceleration units A1, A2, and A3, which will not be repeated here.

[0061] In addition, the structures of the multiple acceleration units in the acceleration assembly can be the same or different. Figure 6 For ease of illustration, the structures of the multiple acceleration units shown are identical. However, in practice, the structures of the multiple acceleration units can be different. For example, in some acceleration units, the multiple accelerator cards are arranged in a polygonal pattern, in some acceleration units, the multiple accelerator cards are arranged in a line, in some acceleration units, the multiple accelerator cards are connected by a single line, in some acceleration assemblies, the multiple accelerator cards are connected by two links, and some acceleration units include four accelerator cards, while others include three or five accelerator cards, etc. In other words, the structure of each acceleration unit can be individually configured, and the structures of different acceleration units can be the same or different.

[0062] The acceleration assembly disclosed herein not only interconnects the accelerator cards within the acceleration unit within the assembly, but also the accelerator cards of different acceleration units, thereby building a hybrid three-dimensional network. With this setup, while each accelerator card processes data, it can also share data through the interconnections between acceleration units. Since data sharing allows for direct access to data, it reduces the data transmission path and time, significantly improving data processing efficiency.

[0063] Figure 7 This is a schematic diagram of the acceleration component structure in another embodiment of the present disclosure. Figure 7 As shown in , the acceleration component 700 may include n aforementioned acceleration units, namely acceleration unit A1, acceleration unit A2, acceleration unit A3, ..., acceleration unit An. The multiple acceleration units in the acceleration component 700 can logically be in a multi-layer structure (shown by dotted lines in the figure), and each layer may include an acceleration unit. The acceleration card of each acceleration unit is connected to the acceleration card in another acceleration unit through an external port. This layer-by-layer configuration combination allows each acceleration card to share data through a high-speed serial link while processing data at high speed, thereby realizing unlimited interconnection of the acceleration cards to meet customizable computing power requirements and realize flexible configuration of the computing power of the processor cluster hardware. As further shown in the figure, the acceleration units of each layer may include four acceleration cards, and the acceleration units can be logically arranged in a quadrilateral arrangement, with the four acceleration cards respectively arranged at the four vertex positions of the quadrilateral.

[0064] It should be understood by those skilled in the art that the above Figure 7 The acceleration assembly described is exemplary and non-restrictive. For example, the structures of the multiple acceleration units can be the same or different. The number of layers of the acceleration assembly can be 2, 3, 4, or more, and the number of layers can be freely set as needed. For each two connected acceleration units, the number of connection paths between the two can be 1, 2, 3, or 4. For ease of understanding, the following will be combined with Figure 8 - Figure 12 provides an exemplary description.

[0065] Figure 8 FIG. 1 is a schematic diagram of the acceleration component structure in another embodiment of the present disclosure. Figure 8 As shown in the figure, the number of acceleration units in the acceleration component 701 can be 2, and the two acceleration units are connected through a path. Specifically, they can be connected through, for example, the external port of the acceleration card MC0 in the acceleration unit A1 and the external port of the acceleration card MC0 in the acceleration unit A2, so that information interaction can be realized between the acceleration unit A1 and the acceleration unit A2.

[0066] like Figure 9As shown in FIG, the number of acceleration units in acceleration assembly 702 can be two, and the two acceleration units are connected via two paths: the external port of acceleration card MC0 in acceleration unit A1 is connected to the external port of acceleration card MC0 in acceleration unit A2, and the external port of acceleration card MC1 in acceleration unit A1 is connected to the external port of acceleration card MC1 in acceleration unit A2. In this way, if one path fails, the other path still supports communication between the acceleration units, further improving the security of the acceleration assembly.

[0067] Please refer to the following Figure 10 , Figure 10 FIG. 1 is a schematic diagram of the acceleration component structure in another embodiment of the present disclosure. Figure 10 In the illustrated acceleration assembly 703, there can be two acceleration units, each connected via three paths: the external port of the acceleration card MC0 in acceleration unit A1 is connected to the external port of the acceleration card MC0 in acceleration unit A2; the external port of the acceleration card MC1 in acceleration unit A1 is connected to the external port of the acceleration card MC1 in acceleration unit A2; and the external port of the acceleration card MC2 in acceleration unit A1 is connected to the external port of the acceleration card MC2 in acceleration unit A2. This ensures that even if two of the paths fail, another path still supports communication between the acceleration units, further improving the security of the acceleration assembly.

[0068] Please refer to the following Figure 11 , Figure 11 FIG. 1 is a schematic diagram of the acceleration component structure in another embodiment of the present disclosure. Figure 11 In the illustrated acceleration assembly 704, there can be two acceleration units, and two acceleration units can be connected via four paths. For example, the external port of the acceleration card MC0 in acceleration unit A1 is connected to the external port of the acceleration card MC0 in acceleration unit A2, the external port of the acceleration card MC1 in acceleration unit A1 is connected to the external port of the acceleration card MC1 in acceleration unit A2, the external port of the acceleration card MC2 in acceleration unit A1 is connected to the external port of the acceleration card MC2 in acceleration unit A2, and the external port of the acceleration card MC3 in acceleration unit A1 is connected to the external port of the acceleration card MC3 in acceleration unit A2. In this way, even if three of the paths fail, another path still supports communication between the acceleration units, further improving the security of the acceleration assembly.

[0069] Figure 12a The diagram below shows the acceleration components as a network topology. Figure 12aAs shown in , the acceleration component 705 may include two acceleration units, each acceleration unit may include four acceleration cards, and there may be two links between the acceleration card MC1 and the acceleration card MC3 in each acceleration unit, and there may be two links between the acceleration card MC0 and the acceleration card MC2. Figure 12a The acceleration device 705 in the left figure can form a three-dimensional expression as shown in the right figure. Figure 12a In the right figure, the circles represent accelerator cards, and the lines represent link connections. The number 0 in the circle represents accelerator card MC0, the number 1 represents accelerator card MC1, the number 2 represents accelerator card MC2, and the number 3 represents accelerator card MC3. The right figure still shows the acceleration component 705, but in another form of expression, that is, it shows the form of network topology. The numbers embedded in the vertical lines in the right figure represent the connected port numbers. For example, the MC0s in the two acceleration units are connected using port 0, the MC1s are connected using port 0, the MC2s are connected using port 3, and the MC3s are connected using port 3.

[0070] for Figure 12a In the right figure, an acceleration unit is considered a node. Two nodes have eight acceleration cards, meaning the two nodes form a so-called eight-card interconnect. The one-machine, four-card interconnection relationship within each node is fixed. When two nodes are interconnected, MC0 and MC1 in the upper node (acceleration unit A1) are connected to MC0 and MC1 in the lower node (acceleration unit A2) via port 0, respectively. MC2 and MC3 in the upper node are connected to MC2 and MC3 in the lower node via port 3, respectively. This node topology is called a hybrid cube mesh, meaning that acceleration component 705 is a hybrid cube mesh.

[0071] exist Figure 12a In the topology with 8 cards shown, two independent rings can also be formed. Figure 12b and Figure 12c As shown, this can maximize the use of topological bandwidth for reduction operations.

[0072] exist Figure 12b In Figure 12 , accelerator cards MC1 and MC3 in acceleration unit A1 are connected via their respective internal port 5, accelerator cards MC0 and MC2 are connected via their respective internal port 5, and accelerator cards MC2 and MC3 are connected via their respective internal port 1. Accelerator card MC1 in acceleration unit A1 and MC1 in acceleration unit A2 are connected via their respective external single port 0, and accelerator card MC0 in acceleration unit A1 and MC0 in acceleration unit A2 are connected via their respective external port 0. Thus, an independent ring is formed among the eight cards in Figure 12 .

[0073] exist Figure 12cIn Figure 12 , accelerator cards MC1 and MC3 in acceleration unit A1 are connected via their respective internal port 2, accelerator cards MC0 and MC2 are connected via their respective internal port 2, and accelerator cards MC0 and MC1 are connected via their respective internal port 1. Accelerator cards MC2 in acceleration unit A1 and MC2 in acceleration unit A2 are connected via their respective external single port 3, and accelerator cards MC3 in acceleration unit A1 and MC3 in acceleration unit A2 are connected via their respective external port 3. Thus, another independent loop is formed among the eight cards in Figure 12 .

[0074] The above only shows two exemplary connection methods. However, in reality, the four connection paths between two acceleration units are actually equivalent. Therefore, any one to three of these four paths can be used to connect the two acceleration units, forming a ring connection with the accelerator card in each acceleration unit. This will not be further described here.

[0075] Figure 13 FIG. 1 is a schematic diagram of an acceleration device in another embodiment of the present disclosure. Figure 13 As shown in FIG, the acceleration device 800 may include n acceleration units described above, namely, acceleration unit A1, acceleration unit A2, acceleration unit A3, ..., acceleration unit An. The multiple acceleration units in the acceleration device 800 logically form a multi-layer structure (shown by dotted lines in the figure). The multiple layers here may include odd or even layers, and each layer may include an acceleration unit. The acceleration card of each acceleration unit is connected to the acceleration card of another acceleration unit via an external port. Acceleration unit A1 and acceleration unit A2 are connected via an external port, acceleration unit A2 and acceleration unit A3 are connected via an external port, and so on, acceleration unit An-1 and acceleration unit An are connected via an external port. The last acceleration unit can be connected to the first acceleration unit, so that the multiple acceleration units are connected end to end to form a ring structure. For example, the external port of acceleration card MC0 of acceleration unit An is connected to the external port of acceleration card MC0 of acceleration unit A1. This progressive configuration combination allows each accelerator card to share data through high-speed serial links while processing data at high speed, achieving unlimited interconnection of accelerator cards to meet customizable computing power requirements and realize flexible configuration of the processor cluster hardware computing power.

[0076] It should be noted that there are many situations in which the connection relationship of the acceleration units in the acceleration device disclosed in this disclosure has been described in detail above. For details, please refer to the above Figure 6The description of the connection relationship of the acceleration units in the acceleration unit will not be repeated here. In addition, there are many ways to connect the last acceleration unit to the first acceleration unit, which may include: the external port of MC0 in the acceleration unit A1 is connected to the external port of MC0 in An, the external port of MC1 in the acceleration unit A1 is connected to the external port of MC1 in An, the external port of MC2 in the acceleration unit A1 is connected to the external port of MC2 in An, and the external port of MC3 in the acceleration unit A1 is connected to the external port of MC3 in An. For ease of understanding, the following will be combined with Figure 14 and Figure 15 In the following description, it can be understood by those skilled in the art that Figure 14 and Figure 15 The accelerator shown is Figure 13 The various embodiments of the acceleration device 800 are shown, so Figure 13 The relevant description of the acceleration device 800 can also be applied to Figure 14 and Figure 15 The accelerator in.

[0077] refer to Figure 14 , Figure 14 FIG. 1 is a schematic diagram of the network topology corresponding to the acceleration device in one embodiment. Figure 14 The acceleration device 801 shown can be composed of four acceleration units, the circles represent acceleration cards, the lines represent link connections, the number 0 in the circle represents acceleration card MC0, the number 1 represents acceleration card MC1, the number 2 represents acceleration card MC2, and the number 3 represents acceleration card MC3; the numbers embedded in the vertical lines in the figure represent the connected port numbers. The last acceleration unit is connected to the first acceleration unit, and the total number of hops is 5 times. Each acceleration unit is a node, and the interconnection between nodes can achieve the interconnection of 4 nodes and 16 cards. The four acceleration units form a small cluster, which is internally interconnected and is called a supercomputing cluster super pod. This topology is the main form of ultra-large-scale clusters, using high-speed SerDes ports, with a total number of hops of 5 times and the lowest latency. The cluster has good manageability and robustness.

[0078] refer to Figure 15 , Figure 15 FIG. 1 is a schematic diagram of a network topology corresponding to an acceleration device in another embodiment. Figure 15 and Figure 14 The difference is that Figure 15 The acceleration device 802 shown in FIG has a larger number of acceleration units. As can be seen from the figure, the last acceleration unit of the acceleration device 802 is connected to the first acceleration unit. With this configuration of the acceleration device, the total hop count is the number of nodes plus one, that is, the total hop count is the number of acceleration units plus one.

[0079] Combined with the above Figure 13-15 An acceleration device including a plurality of acceleration units is exemplarily described. According to the technical solution disclosed herein, an acceleration device including a plurality of the aforementioned acceleration components is also provided, which will be described in detail below in conjunction with a plurality of embodiments.

[0080] Figure 16 Here is a schematic diagram of an acceleration device in another embodiment of the present disclosure. The acceleration device 900 may include m aforementioned acceleration components. In each acceleration component, in addition to the external connection ports required for connection between acceleration units within the acceleration component, there are also idle external connection ports. The acceleration components are connected to each other through the idle external connection ports. The external connection port of the acceleration card MC1 of the acceleration unit A1 in the acceleration component B1 can be connected to the external connection port of the acceleration card MC1 of the acceleration unit A1 in the acceleration component B2. The external connection port of the acceleration card MC1 of the acceleration unit A1 in the acceleration component B2 can be connected to the external connection port of the acceleration card MC1 of the acceleration unit A1 in the acceleration component B3. And so on. Multiple acceleration components are connected to each other. It can be understood that Figure 16 The acceleration device shown is exemplary and non-restrictive. For example, the structures of the multiple acceleration components can be the same or different. For another example, the connection between different acceleration components through the idle external ports can be limited to Figure 16 The method shown in , may also include other methods. For ease of understanding, the following will be combined with Figures 17-25 An exemplary description is given.

[0081] based on Figure 16 The accelerator provided, further reference Figure 17 , Figure 17 This is a schematic diagram of the network topology corresponding to an acceleration device in another embodiment. The acceleration device 901 may include two acceleration components. The acceleration component B1 may include four acceleration units. The acceleration component B2 may include four acceleration units. The first acceleration unit in the acceleration component B1 is connected to the first acceleration unit in the acceleration component B2, and the last acceleration unit in the acceleration component B1 is connected to the last acceleration unit in the acceleration component B2. The total hop count under this network topology is 9. It will be understood by those skilled in the art that Figure 17 The network structure consisting of multiple acceleration units in each acceleration assembly is logical; in actual applications, the arrangement of multiple acceleration units can be adjusted as needed. The number of acceleration units in each acceleration assembly is not limited to the four shown in the figure, and can be increased or decreased as needed, for example, six or eight.

[0082] based on Figure 16 The accelerator provided, further reference Figure 18 , Figure 18This is a schematic diagram of an acceleration device in another embodiment of the present disclosure. The acceleration device 902 may include four acceleration components, namely acceleration components B1, B2, B3, and B4. Among the four acceleration components, each acceleration component may include two acceleration units A1 and A2, and each acceleration component may be interconnected with one of the acceleration units A1 and A2 of other acceleration components. For example, the acceleration unit A1 in the acceleration component B1 is connected to the acceleration unit A1 in the acceleration component B2, the acceleration unit A1 in the acceleration component B2 is connected to the acceleration unit A1 in the acceleration component B3, and the acceleration unit A1 in the acceleration component B3 is connected to the acceleration unit A1 in the acceleration component B4. The connections here are all made through the external ports of the acceleration units.

[0083] It should be noted that the connection between acceleration components is Figure 18 There are many other connection methods besides those shown in the figure. For example, the connection methods between acceleration components may specifically include: acceleration unit A1 or A2 in acceleration component B1 is connected to acceleration unit A1 or A2 in acceleration component B2, acceleration unit A1 or A2 in acceleration component B2 is connected to acceleration unit A1 or A2 in acceleration component B3, and acceleration unit A1 or A2 in acceleration component B3 is connected to acceleration unit A1 or A2 in acceleration component B4.

[0084] based on Figure 18 For further information on the accelerator provided, please refer to Figure 19 , Figure 19 FIG. 1 is a schematic diagram of an acceleration device in another embodiment of the present disclosure. Figure 19 In the illustrated acceleration device 903, each acceleration assembly can be interconnected with one of the first and second acceleration units of another acceleration assembly via two paths. For example, the first acceleration unit (e.g., acceleration unit A1) in acceleration assembly B1 can be connected to the first acceleration unit (e.g., acceleration unit A1) in acceleration assembly B2 via two paths, the acceleration unit A1 in acceleration assembly B2 can be connected to the acceleration unit A1 in acceleration assembly B3 via two paths, and the acceleration unit A1 in acceleration assembly B3 can be connected to the acceleration unit A1 in acceleration assembly B4 via two paths.

[0085] It should be noted that Figure 19 The connection between two paths is marked in , but it can also include more than two paths. Figure 19The connection method shown in the figure may also include other methods. For example, the acceleration unit A1 or A2 in the acceleration component B1 can be connected to the acceleration unit A1 or A2 in the acceleration component B2 using two paths, the acceleration unit A1 or A2 in the acceleration component B2 can be connected to the acceleration unit A1 or A2 in the acceleration component B3 using two paths, and the acceleration unit A1 or A2 in the acceleration component B3 can be connected to the acceleration unit A1 or A2 in the acceleration component B4 using two paths.

[0086] based on Figure 16 For further information on the accelerator provided, please refer to Figure 20 , Figure 20 This is a schematic diagram of an acceleration device in another embodiment of the present disclosure. The acceleration device 904 includes four acceleration components, namely, acceleration component B1, acceleration component B2, acceleration component B3, and acceleration component B4. Each acceleration component includes two acceleration units, and each acceleration unit includes two pairs of acceleration cards. In each acceleration unit, MC0 and MC1 are the first pair of acceleration cards, and MC2 and MC3 are the second pair of acceleration cards. Among them, the second pair of acceleration cards of acceleration unit A1 of acceleration component B1 is connected to the second pair of acceleration cards of acceleration unit A2 of acceleration component B2; the first pair of acceleration cards of acceleration unit A2 of acceleration component B2 is connected to the first pair of acceleration cards of acceleration unit A1 of acceleration component B3; the second pair of acceleration cards of acceleration unit A2 of acceleration component B3 is connected to the second pair of acceleration cards of acceleration unit A1 of acceleration component B4; the first pair of acceleration cards of acceleration unit A1 of acceleration component B4 is connected to the first pair of acceleration cards of acceleration unit A2 of acceleration component B1.

[0087] refer to Figure 21 , Figure 21 This is a network topology diagram of another acceleration device. Figure 21 The acceleration device 905 shown is Figure 20 The acceleration device 904 is a specific form of the acceleration device 904, so the above description of the acceleration device 904 can also be applied to Figure 21 The acceleration device 905. Figure 21 As shown in the figure, each acceleration component of the acceleration device 905 can form a hybrid three-dimensional network unit. The internal interconnection relationship of each hybrid three-dimensional network unit can be shown as shown in the figure, realizing the interconnection of 8 nodes and 32 cards of the acceleration device 905. The four acceleration components can realize multi-card and multi-node interconnection through, for example, QSFP-DD interfaces and cables, forming a matrix network topology.

[0088] Specifically, in this embodiment, the ports 0 of the acceleration cards MC2 and MC3 of the upper node of the acceleration component B1 can be connected to the acceleration cards MC2 and MC3 of the lower node of the acceleration component B2, respectively; the ports 3 of the MC0 and MC1 of the lower node of the acceleration component B2 can be connected to the MC0 and MC1 of the upper node of the acceleration component B3, respectively; the ports 0 of the MC2 and MC3 of the lower node of the acceleration component B3 can be connected to the MC2 and MC3 of the upper node of the acceleration component B4, respectively; the ports 3 of the MC0 and MC1 of the upper node of the acceleration component B4 can be connected to the MC0 and MC1 of the lower node of the acceleration component B1, respectively. The interconnection between the hybrid three-dimensional networks thus set up can form two bidirectional ring structures (as shown in the above combined Figure 5b 、 Figure 5c , Figure 12b and Figure 12c As described above, it has advantages such as good reliability and security, is suitable for deep learning training, and has high computational efficiency. For the matrix network topology consisting of 8 nodes in the acceleration device 905, the total number of hops is 11.

[0089] Furthermore, if Figure 21 As shown in , the first pair of accelerator cards and the second pair of accelerator cards in different acceleration units of the same acceleration assembly can be indirectly connected. For example, the accelerator cards MC0 and MC1 of the upper acceleration unit in the acceleration assembly B1 are indirectly connected to the accelerator cards MC2 and MC3 of the lower acceleration unit.

[0090] exist Figure 21 Based on the network topology of the matrix network topology as the basic unit, it can be further expanded into a larger network topology. Figure 22 This is a schematic diagram of the matrix network topology based on wireless expansion of the acceleration device. Figure 22 As shown, the acceleration device 906 may include multiple acceleration components, each of which (shown as a block in the figure) may include multiple acceleration units (not shown in the stereogram, refer to Figure 21 Each acceleration unit may include, for example, four acceleration cards interconnected as shown in the figure, so the matrix network topology can theoretically be expanded infinitely.

[0091] based on Figure 16 For further information on the accelerator provided, please refer to Figure 23 , Figure 23This is a schematic diagram of an acceleration device in another embodiment of the present disclosure. The acceleration device 908 may include m (m ≥ 2) acceleration assemblies, each of which may include n (n ≥ 2) acceleration units, and the m acceleration assemblies may be connected in a ring. Specifically, the acceleration unit An of the acceleration assembly B1 may be connected to the acceleration unit A1 of the acceleration assembly B2, and the acceleration unit An of the acceleration assembly B2 may be connected to the acceleration unit A1 of the acceleration assembly B3. This is analogous to the acceleration assembly Bm, where the acceleration unit An of the acceleration assembly Bm may be connected to the acceleration unit A1 of the acceleration assembly B1. Thus, the m acceleration assemblies are connected end to end in a ring-like connection.

[0092] based on Figure 23 , please refer to Figure 24 , Figure 24 This is a network topology diagram of another acceleration device. The acceleration device 909 can include 6 acceleration components, each of which can include two acceleration units. The second acceleration unit of each acceleration component can be connected to the first acceleration unit of the next acceleration component, forming an interconnection of 12 nodes and 48 cards, forming a larger matrix network topology. The total hop under this network topology is 13 times.

[0093] based on Figure 24 , please refer to Figure 25 , Figure 25 This is a network topology diagram of another acceleration device. The acceleration device 910 includes 8 acceleration components, each of which includes two acceleration units. The second acceleration unit of each acceleration component can be connected to the first acceleration unit of the next acceleration component, forming an interconnection of 16 nodes and 64 cards, forming a larger matrix network topology. The total hop under this network topology is 17 times.

[0094] exist Figure 25 Based on this, vertical expansion can be achieved, forming ultra-large-scale matrix networks, such as 20 nodes with 80 cards or 24 nodes with 96 cards. Theoretically, this expansion is unlimited, with the total number of hops equal to the number of nodes plus one. By optimizing the interconnection between nodes, the overall system latency can be minimized, maximizing the real-time performance required by the system while processing massive amounts of data.

[0095] Combined with the above Figure 16-Figure 25 While an accelerator device including multiple accelerator assemblies has been described as an example, those skilled in the art will appreciate that the above description is illustrative and non-limiting. For example, the number and structure of the accelerator assemblies, as well as the connection relationships between the accelerator assemblies, can be adjusted as needed. Those skilled in the art may also combine the above embodiments to form an accelerator device as needed, which is also within the scope of protection of this disclosure.

[0096] In addition, it should be noted that the accelerator card matrix, fully connected square network (topology), hybrid three-dimensional network (topology), matrix network (topology), etc. described in this disclosure are all logical, and the specific layout form can be adjusted as needed.

[0097] The topology disclosed herein can also perform data reduction operations. The reduction operations can be performed on each accelerator card, each accelerator unit, and the accelerator device. The specific operation steps can be as follows.

[0098] Taking the reduction sum operation as an example, the reduction operation process performed in an acceleration unit may include: transferring the data stored in the first accelerator card to the second accelerator card, and performing an addition operation on the data originally stored in the second accelerator card and the data received from the first accelerator card in the second accelerator card; then, transferring the addition operation result in the second accelerator card to the third accelerator card, and performing an addition operation again, and so on, until all the data stored in the accelerator cards have been added and each accelerator card has received the final operation result.

[0099] by Figure 4 Taking the acceleration unit shown in the figure as an example, the data (0,0) is stored in the acceleration card MC0, the data (1,2) is stored in the acceleration card MC1, the data (3,1) is stored in the acceleration card MC2, and the data (2,4) is stored in the acceleration card MC3. The data (0,0) in the acceleration card MC0 can be transferred to the acceleration card MC1, and after the addition operation, the result (1,2) is obtained. Next, the result (1,2) is transferred to the acceleration card MC2, and the next result (4,3) is obtained. Then, the next result (4,3) is transferred to the acceleration card MC3, and the final result (6,7) is obtained.

[0100] Thereafter, in the reduction operation disclosed herein, the final result (6, 7) is passed to each acceleration card MC0, MC1, MC2 and MC3, so that the data (6, 7) is stored in all acceleration cards, thereby completing the reduction operation in one acceleration unit.

[0101] Figure 4 The acceleration unit can form two independent rings, each of which can complete the reduction operation of half of the data, thereby accelerating the operation speed and improving the operation efficiency.

[0102] In addition, when performing reduction operations, the acceleration unit can also implement concurrent calculations on multiple acceleration cards, thereby speeding up the calculation speed. For example, the acceleration card MC0 stores data (0,0), the acceleration card MC1 stores data (1,2), the acceleration card MC2 stores data (3,1), and the acceleration card MC3 stores data (2,4). Part of the data (0) in the acceleration card MC0 can be transferred to the acceleration card MC1, and after the addition operation, the result (1) is obtained. At the same time, part of the data (2) in the acceleration card MC1 can be transferred to the acceleration card MC2, and after the addition operation, the result (3) is obtained. In this way, the concurrent calculation of the acceleration cards MC1 and MC2 is realized; and so on, the entire reduction operation is completed.

[0103] The aforementioned concurrent computing can also include performing addition operations on grouped acceleration units, followed by a reduction operation on the results of the acceleration units in the group and the results of the acceleration units in another group. For example, if the data (0, 0) is stored in accelerator card MC0, the data (1, 2) is stored in accelerator card MC1, the data (3, 1) is stored in accelerator card MC2, and the data (2, 4) is stored in accelerator card MC3, the data in accelerator card MC0 can be transferred to accelerator card MC1 for calculation to obtain the first set of results (1, 2). Synchronously or asynchronously, the data in accelerator card MC2 can be transferred to accelerator card MC3 for calculation to obtain the second set of results (5, 5). Next, the first and second sets of results are combined to obtain the final reduced result (6, 7).

[0104] Similarly, in addition to performing reduction operations in an acceleration unit, reduction operations can also be performed in an acceleration component or an acceleration device. It should be understood that an acceleration device can also be considered as an acceleration component connected end to end.

[0105] When performing a reduction operation in an acceleration component or an acceleration device, it may include: performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result in each acceleration unit; performing a second reduction operation on the first reduction results in multiple acceleration units to obtain a second reduction result.

[0106] Taking the reduction sum operation as an example, the first step has been described above. For an acceleration device including multiple acceleration units, a local reduction operation can be first performed in each acceleration unit. After the reduction operation in each acceleration unit is completed, the accelerator card in the same acceleration unit will obtain the result of the local reduction operation, which is called the first reduction result.

[0107] Next, the first reduction results from all acceleration units can be transferred to adjacent acceleration units and added. Similar to performing a reduction operation within one acceleration unit, the first acceleration unit transfers the first reduction result to the second acceleration unit. After the accelerator card in the second acceleration unit performs the addition operation, the results are transferred and added. After the final addition operation, the final result is transmitted to each acceleration unit.

[0108] It should be noted that since the acceleration components described above are not necessarily connected end to end, the final result can be transmitted to each acceleration unit in a reverse direction, rather than in a circular manner as when the acceleration units are connected end to end. The technical solution disclosed herein does not specifically limit how the final result is transmitted.

[0109] Furthermore, according to one embodiment of the present disclosure, the acceleration device can also be configured to perform reduction operations, including: performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result; performing an intermediate reduction operation on the first reduction result in multiple acceleration units of the same acceleration component to obtain an intermediate reduction result; performing a second reduction operation on the intermediate reduction results in multiple acceleration components to obtain a second reduction result.

[0110] In this embodiment, the reduction operation may be first performed in the same acceleration unit, which has been described above and will not be repeated here.

[0111] Next, reduction operations can be performed in each acceleration component so that each acceleration card in each acceleration component obtains the local reduction results in the acceleration component; next, reduction operations can be performed in multiple acceleration components based on the acceleration component, so that each acceleration card obtains the global reduction results in the acceleration device.

[0112] It should be understood that the above-mentioned transmission order is only for the convenience of description and is not necessarily the correct transmission order. Figure 26 This is a schematic diagram of the structure of a combined processing device according to an embodiment of the present disclosure. As shown in the figure, the combined processing device 2600 may include an acceleration unit 2601, which may specifically be the acceleration unit shown in Figures 1 to 5. In addition, the combined processing device may also include an interconnection interface 2602 and other processing devices 2603. According to the present disclosure, the acceleration unit 2601 can interact with the other processing devices 2603 through the interconnection interface 2602 to jointly complete user-specified operations.

[0113] According to the disclosed solution, the other processing devices may include one or more types of processors, such as a microprocessor unit (MCU), a baseboard controller (BMC), and a central processing unit. The number of such processors is not limited and is determined based on actual needs. In one or more embodiments, the other processing devices may serve as an interface between the disclosed acceleration unit and external data and control, performing basic control functions including but not limited to data transfer and starting and stopping the acceleration unit. The other processing devices may also collaborate with the acceleration unit to jointly complete computing tasks.

[0114] Optionally, the combined processing device 2600 may further include a storage device 2604, which may be connected to the acceleration unit 2601, the interconnection interface 2602, and the other processing device 2603. In one or more embodiments, the storage device 2604 may be used to store data of the acceleration unit 2601 and the other processing device 2603, particularly data that cannot be fully stored in the internal or on-chip storage devices of the acceleration unit 2601 and the other processing device 2603.

[0115] In some application scenarios, the combined processing device 2600 disclosed herein can be used in, for example, large-scale data centers, supercomputing centers, cloud computing centers, etc., and can build high-performance processor clusters to achieve real-time processing of massive data.

[0116] In some embodiments, the present disclosure further discloses a circuit board, which may include the above-mentioned acceleration unit. Figure 27 , which provides an exemplary circuit board 2700. In addition to the above-mentioned one or more acceleration units 2706 (two are shown as an example), the above-mentioned circuit board 2700 may also include other supporting components, which include but are not limited to: a storage device 2701, an interface device 2707 and a control device 2705.

[0117] The storage device 2701 can be connected to the acceleration unit 2706 via a bus for storing data. The storage device 2701 can include multiple groups of storage units 2702. Each group of storage units 2702 can be connected to the acceleration unit 2706 via a bus. It is understood that each group of storage units 2702 can be at least one of DDR SDRAM (Double Data Rate SDRAM) and HBM (High Bandwidth Memory).

[0118] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read on both the rising and falling edges of the clock pulse. DDR is twice as fast as standard SDRAM. In one embodiment, the memory device 2701 may include four groups of memory cells 2702. Each group of memory cells 2702 may include multiple DDR4 particles (chips). In one embodiment, the chip may include four 72-bit DDR4 controllers, of which 64 bits are used for data transmission and 8 bits are used for ECC verification.

[0119] In one embodiment, each group of storage units 2702 may include multiple double-bit rate synchronous dynamic random access memories (DDRs) connected in parallel. DDR can transmit data twice within one clock cycle. A controller for controlling DDR is provided in the acceleration unit 2706 to control data transmission and data storage in each storage unit. The interface device 2707 may be connected to the acceleration unit 2706. The interface device 2707 is used to facilitate data transmission between the acceleration unit 2706 and an external device 2708 (e.g., a server or computer). For example, in one embodiment, the interface device 2707 may be a standard PCIE interface. For example, data to be processed is transmitted from the server to the acceleration unit 2706 via the standard PCIE interface to achieve data transfer. In another embodiment, the interface device 2707 may also be another interface. This disclosure does not limit the specific form of such another interface, as long as the interface device can implement a switching function. In addition, the calculation results of the acceleration unit 2706 can still be transmitted back to the external device (e.g., a server) by the interface device 2707. The control device 2705 may be connected to the acceleration unit 2706. The control device 2705 can be used to monitor the status of the acceleration unit 2706. Specifically, the acceleration unit 2706 and the control device 2705 can be electrically connected via an SPI interface. The control device 2705 can include a microcontroller (MCU).

[0120] In some embodiments, the present disclosure further discloses an electronic device or apparatus comprising the aforementioned acceleration unit. In some embodiments, the present disclosure further discloses yet another electronic device or apparatus comprising the aforementioned acceleration component. In some embodiments, the present disclosure further discloses yet another electronic device or apparatus comprising the aforementioned acceleration device. In some embodiments, the present disclosure further discloses yet another electronic device or apparatus comprising the aforementioned circuit board.

[0121] Depending on different application scenarios, electronic devices or devices may include, for example, data processing devices, data centers, supercomputing centers, cloud computing centers, servers, and cloud servers.

[0122] In the above embodiments of the present disclosure, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] The foregoing content can be better understood in accordance with the following terms:

[0124] Clause 1. An acceleration component, comprising a plurality of acceleration units, each acceleration unit comprising M local unit acceleration cards, each local unit acceleration card comprising an internal port, each local unit acceleration card being connected to other local unit acceleration cards via the internal port, wherein:

[0125] The M local unit acceleration cards logically form an acceleration card matrix of L*N scale, where L and N are integers not less than 2. At least one of the M local unit acceleration cards includes an external port, and the acceleration units are connected through the external ports.

[0126] Clause 2. The acceleration component according to clause 1, wherein the M local unit acceleration cards are logically formed into a 2*2, 3*3 or 4*4 acceleration card matrix.

[0127] Clause 3. The acceleration component according to clause 1 or 2, wherein each local unit acceleration card is connected to at least one other local unit acceleration card via two paths.

[0128] Clause 4. The acceleration assembly according to clause 2 or 3, wherein the diagonal accelerator cards at the four corners of the accelerator card matrix are connected by two paths.

[0129] Clause 5. An acceleration component according to any one of Clauses 1-4, wherein, when the acceleration unit includes four local-unit acceleration cards, each local-unit acceleration card includes six ports, and wherein four ports of each local-unit acceleration card are internal ports for connecting to the other three local-unit acceleration cards; and the remaining two ports of at least one local-unit acceleration card are external ports for connecting to an external-unit acceleration card.

[0130] Clause 6. The acceleration component according to any one of clauses 1-5, wherein the internal port and the external port are SerDes ports.

[0131] Clause 7. An acceleration component according to any one of clauses 1-6, wherein the multiple acceleration units are logically arranged in a multi-layer structure, each layer including an acceleration unit, and the acceleration card of each acceleration unit is connected to an external unit acceleration card through an external port.

[0132] Clause 8. The acceleration assembly according to any one of clauses 1 to 7, wherein the plurality of acceleration units have the same structure.

[0133] Clause 9. The acceleration component according to any one of clauses 1-8, wherein the acceleration unit is configured to perform a reduction operation on the data in the acceleration card of the acceleration unit to obtain a reduction result.

[0134] Clause 10. The acceleration component of any one of clauses 1-9, wherein the acceleration component is configured to perform a reduction operation, comprising:

[0135] Performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result in each acceleration unit;

[0136] A second reduction operation is performed on the first reduction results in the plurality of acceleration units to obtain a second reduction result.

[0137] Clause 11. An acceleration device, comprising a plurality of acceleration units, each acceleration unit comprising M local acceleration cards, each local acceleration card comprising an internal port, each local acceleration card being connected to other local acceleration cards via the internal port, wherein:

[0138] The M local unit acceleration cards logically form an acceleration card matrix of L*N scale, where L and N are integers not less than 2. At least one of the M local unit acceleration cards includes an external port, and the acceleration units are connected through the external ports. The multiple acceleration units logically present a multi-layer structure, each layer including an acceleration unit. The local unit acceleration card of each acceleration unit is connected to the external unit acceleration card through the external port, and the last acceleration unit is connected to the first acceleration unit, so that the multiple acceleration units are connected end to end to form a ring structure.

[0139] Clause 12. The acceleration device according to Clause 11, wherein the M local unit acceleration cards logically form a 2*2, 3*3 or 4*4 acceleration card matrix.

[0140] Clause 13. The acceleration device according to clause 11 or 12, wherein each local acceleration card is connected to at least one other local acceleration card via two paths.

[0141] Clause 14. The acceleration device according to clause 12 or 13, wherein the diagonal accelerator cards at the four corners of the accelerator card matrix are connected by two paths.

[0142] Clause 15. An acceleration device according to any one of Clauses 1 to 14, wherein, when the acceleration unit includes four local-unit acceleration cards, each local-unit acceleration card includes six ports, and wherein four ports of each local-unit acceleration card are internal ports for connecting to the other three local-unit acceleration cards; and the remaining two ports of at least one local-unit acceleration card are external ports for connecting to an external-unit acceleration card.

[0143] Clause 16. The acceleration device according to any one of clauses 11-15, wherein the internal port and the external port are SerDes ports.

[0144] Clause 17. The acceleration device according to any one of clauses 1-16, wherein the acceleration component is configured to perform a reduction operation, comprising:

[0145] Performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result in each acceleration unit;

[0146] A second reduction operation is performed on the first reduction results in the plurality of acceleration units to obtain a second reduction result.

[0147] Clause 18. An acceleration device comprising a plurality of acceleration components as described in any one of clauses 1-10, wherein the acceleration components are interconnected via idle external ports.

[0148] Clause 19. An acceleration device according to Clause 18, wherein the acceleration device includes four acceleration assemblies, each acceleration assembly includes a first acceleration unit and a second acceleration unit, and each acceleration assembly is interconnected with one of the first acceleration unit and the second acceleration unit of other acceleration assemblies through one of the first acceleration unit and the second acceleration unit.

[0149] Clause 20. The acceleration device according to Clause 19, wherein each acceleration assembly is interconnected with one of the first acceleration unit and the second acceleration unit of other acceleration assemblies through one of the first acceleration unit and the second acceleration unit using at least two paths.

[0150] Clause 21. The acceleration device according to clause 19, wherein the four acceleration assemblies include a first acceleration assembly, a second acceleration assembly, a third acceleration assembly, and a fourth acceleration assembly, and

[0151] The second pair of acceleration cards of the first acceleration unit of the first acceleration assembly is connected to the second pair of acceleration cards of the second acceleration unit of the second acceleration assembly;

[0152] The first pair of acceleration cards of the second acceleration unit of the second acceleration assembly is connected to the first pair of acceleration cards of the first acceleration unit of the third acceleration assembly;

[0153] The second pair of acceleration cards of the second acceleration unit of the third acceleration assembly is connected to the second pair of acceleration cards of the first acceleration unit of the fourth acceleration assembly;

[0154] The first pair of acceleration cards of the first acceleration unit of the fourth acceleration assembly is connected to the first pair of acceleration cards of the second acceleration unit of the first acceleration assembly.

[0155] Clause 22. The acceleration device according to Clause 21, wherein the first pair of acceleration cards and the second pair of acceleration cards in different acceleration units in the same acceleration assembly are indirectly connected.

[0156] Clause 23. The acceleration device according to clause 18, wherein the acceleration device comprises a plurality of acceleration assemblies connected in a ring shape, wherein the second acceleration unit of each acceleration assembly is connected to the first acceleration unit of the next acceleration assembly, so that the plurality of acceleration assemblies are connected end to end.

[0157] Clause 24. The acceleration device according to any one of clauses 18-23, wherein the acceleration device is configured to perform a reduction operation, comprising:

[0158] Performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result;

[0159] Performing an intermediate reduction operation on the first reduction results of the multiple acceleration units of the same acceleration component to obtain an intermediate reduction result;

[0160] The intermediate reduction results in the plurality of acceleration components are subjected to a second reduction operation to obtain a second reduction result.

[0161] Clause 25. An electronic device comprising the acceleration component as described in any one of Clauses 1-10, or the acceleration device as described in any one of Clauses 11-24.

[0162] It should be understood that the terms "first," "second," "third," and "fourth," etc. in the claims, specification, and drawings of the present disclosure are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0163] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0164] The above is a detailed introduction to the embodiments of the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, changes or modifications made by those skilled in the art based on the ideas of the present disclosure, on the specific implementation methods and application scope of the present disclosure, all fall within the scope of protection of the present disclosure. In summary, the contents of this specification should not be understood as limiting the present disclosure.

Claims

1. An acceleration device, comprising a plurality of acceleration units, each of which comprises M acceleration cards of the same unit, each of which comprises an internal port, each of which is connected to other acceleration cards of the same unit via the internal port, wherein: The M local unit acceleration cards logically form an acceleration card matrix of L*N scale, where L and N are integers not less than 2. At least one of the M local unit acceleration cards includes an external port, and the acceleration units are connected through the external ports. The multiple acceleration units logically present a multi-layer structure, each layer including an acceleration unit. The local unit acceleration card of each acceleration unit is connected to the external unit acceleration card through the external port, and the last acceleration unit is connected to the first acceleration unit, so that the multiple acceleration units are connected end to end to form a ring structure.

2. The acceleration device according to claim 1, wherein: The M local unit acceleration cards logically form a 2*2, 3*3 or 4*4 acceleration card matrix.

3. The acceleration device according to claim 1, wherein: Each local unit accelerator card is connected to at least one other local unit accelerator card through two paths.

4. The acceleration device according to claim 1, wherein: The diagonal unit accelerator cards at the four corners of the accelerator card matrix are connected via two paths.

5. The acceleration device according to claim 1, wherein: When the acceleration unit includes four local unit acceleration cards, each local unit acceleration card includes six ports, and the four ports of each local unit acceleration card are internal ports for connecting to the other three local unit acceleration cards; the remaining two ports of at least one local unit acceleration card are external ports for connecting to an external unit acceleration card.

6. The acceleration device according to claim 1, wherein: The internal port and the external port are SerDes ports.

7. The acceleration device according to any one of claims 1 to 6, wherein: The acceleration device is configured to perform a reduction operation, including: Performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result in each acceleration unit; A second reduction operation is performed on the first reduction results in the plurality of acceleration units to obtain a second reduction result.

8. An acceleration device comprising a plurality of acceleration components, wherein the acceleration components are interconnected via idle external ports; in, The acceleration component includes multiple acceleration units, each acceleration unit includes M local unit acceleration cards, each local unit acceleration card includes an internal port, and each local unit acceleration card is connected to other local unit acceleration cards through the internal port, wherein, The M local unit accelerator cards logically form an L*N accelerator card matrix, where L and N are integers not less than 2. At least one of the M local unit accelerator cards includes an external port, and the accelerator units are connected to each other via the external ports. Wherein, the acceleration device includes a plurality of acceleration components connected in a ring shape, so that the plurality of acceleration components are connected end to end.

9. The acceleration device according to claim 8, wherein: The M local unit acceleration cards logically form a 2*2, 3*3 or 4*4 acceleration card matrix.

10. The acceleration device according to claim 8, wherein: Each local unit accelerator card is connected to at least one other local unit accelerator card through two paths.

11. The acceleration device according to claim 8, wherein: The diagonal unit accelerator cards at the four corners of the accelerator card matrix are connected via two paths.

12. The acceleration device according to claim 8, wherein: When the acceleration unit includes four local unit acceleration cards, each local unit acceleration card includes six ports, and the four ports of each local unit acceleration card are internal ports for connecting to the other three local unit acceleration cards; the remaining two ports of at least one local unit acceleration card are external ports for connecting to an external unit acceleration card.

13. The acceleration device according to claim 8, wherein: The internal port and the external port are SerDes ports.

14. The acceleration device according to claim 8, wherein The multiple acceleration units are logically in a multi-layer structure, each layer includes an acceleration unit, and the acceleration card of each acceleration unit is connected to the acceleration card of an external unit through an external port.

15. The acceleration device according to claim 8, wherein The plurality of acceleration units have the same structure.

16. The acceleration device according to any one of claims 8 to 15, wherein: The acceleration component is configured to perform a reduction operation on the data in the acceleration card of the acceleration unit to obtain a reduction result.

17. The acceleration device according to any one of claims 8 to 15, wherein: The acceleration component is configured to perform a reduction operation, including: Performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result in each acceleration unit; A second reduction operation is performed on the first reduction results in the plurality of acceleration units to obtain a second reduction result.

18. The acceleration device according to claim 8, wherein: The acceleration device includes four acceleration assemblies, each of which includes a first acceleration unit and a second acceleration unit. Each acceleration assembly is connected to one of the first acceleration unit and the second acceleration unit of other acceleration assemblies through one of the first acceleration unit and the second acceleration unit.

19. The acceleration device according to claim 18, wherein: Each acceleration assembly is connected to one of the first acceleration unit and the second acceleration unit of other acceleration assemblies through at least two paths through one of the first acceleration unit and the second acceleration unit.

20. The acceleration device according to claim 18, wherein The four accelerating assemblies include a first accelerating assembly, a second accelerating assembly, a third accelerating assembly and a fourth accelerating assembly, and The second pair of acceleration cards of the first acceleration unit of the first acceleration assembly is connected to the second pair of acceleration cards of the second acceleration unit of the second acceleration assembly; The first pair of acceleration cards of the second acceleration unit of the second acceleration assembly is connected to the first pair of acceleration cards of the first acceleration unit of the third acceleration assembly; The second pair of acceleration cards of the second acceleration unit of the third acceleration assembly is connected to the second pair of acceleration cards of the first acceleration unit of the fourth acceleration assembly; The first pair of acceleration cards of the first acceleration unit of the fourth acceleration assembly is connected to the first pair of acceleration cards of the second acceleration unit of the first acceleration assembly.

21. The acceleration device according to claim 20, wherein: The first pair of acceleration cards and the second pair of acceleration cards in different acceleration units in the same acceleration component are indirectly connected.

22. The acceleration device according to claim 8, wherein The second accelerating unit of each accelerating assembly is connected to the first accelerating unit of the next accelerating assembly.

23. The acceleration device according to claim 8, wherein: The acceleration device is configured to perform a reduction operation, including: Performing a first reduction operation on the data in the acceleration card of the same acceleration unit to obtain a first reduction result; Performing an intermediate reduction operation on the first reduction results of the multiple acceleration units of the same acceleration component to obtain an intermediate reduction result; The intermediate reduction results in the plurality of acceleration components are subjected to a second reduction operation to obtain a second reduction result.

24. An electronic device comprising the acceleration device according to any one of claims 1 to 23.

Citation Information

Patent Citations

  • Data acceleration processing system

    CN110413561A

  • Acceleration assembly, acceleration device and electronic equipment

    CN212846785U