Processor unit set communication method, processor system and accelerator card

By dividing the processor unit array into multiple subarrays and using a bidirectional ring algorithm, using the bidirectional connection of the Mesh network, the problem of inefficient communication of large-scale PE arrays in the prior art is solved, and more efficient collective communication is achieved.

CN120234285AActive Publication Date: 2025-07-01BEIJING TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510683279.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-01
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Existing ring-based ensemble communication algorithms lead to reduced communication efficiency in large-scale PE arrays and fail to fully utilize the bidirectional connection bandwidth of Mesh networks.

Method used

A processor unit collection communication method is proposed. By dividing the target processor unit array into multiple subarrays, and using a bidirectional ring algorithm for collective communication in each subarray, the communication efficiency is improved by using the bidirectional connection of the Mesh network.

Benefits of technology

Through hierarchical communication division strategy and bidirectional ring algorithm, the ensemble communication efficiency of large PE arrays is significantly improved, and the problem of inefficient communication in large-scale arrays is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234285A_ABST
    Figure CN120234285A_ABST
Patent Text Reader

Abstract

The invention discloses a processor unit set communication method, a processor system and an accelerator card, and the method comprises the steps: dividing a target processor unit array for executing a calculation task into m first processor unit sub-arrays and n second processor unit sub-arrays, and enabling the size of the target processor unit array to be 2m * 2n, the size of the first processor unit subarray is 2 * 2n, and the size of the second processor unit subarray is 2m * 2; controlling the first processor unit sub-array to perform set communication based on a bidirectional ring algorithm so as to execute a calculation task; and controlling the second processor unit sub-array to perform set communication based on the bidirectional ring algorithm so as to execute the computing task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of processor data communication, and more specifically, to a communication method for a set of processor units, a processor system, and an acceleration card. Background Art

[0002] With the rapid development of deep learning and artificial intelligence (AI) technologies, the parameter scale of AI models has become increasingly large. The computing and memory access bandwidth resources of traditional AI processors are difficult to meet the training and inference requirements of current AI models with such parameter scales. To address the bottleneck problems in large model training, such as limited computing power and memory access bandwidth, various new processor architectures have been proposed by the academic and industrial communities. Among them, the wafer-level processor chip stands out among many new processor architectures due to its unique architecture design, which enables it to have much higher computing power and memory access bandwidth than traditional AI chips, and thus has been widely studied and applied in the industry.

[0003] A wafer-level processor chip consists of an on-chip interconnection network and a large number of processor cores (PEs). These PE cores are arranged in an array, and each PE core forms a two-way interconnection with adjacent PE cores in a 2D Mesh structure, thus constituting a wafer-level processor chip. The scale of the PE core array in a wafer-level processor chip is usually relatively large. A wafer-level processor chip often contains thousands or even more PE cores. Therefore, for a processor with a large array of PE cores, optimizing the communication efficiency between internal processor cores is particularly important. Currently in an AI cluster, the commonly used communication algorithms are mainly divided into Tree-based algorithms and Ring-based algorithms. Since a PE core is only directly connected to its adjacent PE cores, if it wants to access a non-adjacent PE core, it must rely on other PE cores for relay routing. Therefore, if a Tree-type algorithm is used for collective communication, link contention is likely to occur, resulting in serious network congestion on some paths and affecting communication efficiency. Therefore, for a large-scale PE core array, a ring algorithm is usually used to complete the collective communication between PE cores.

[0004] Collective communication refers to the efficient data exchange and collaborative work among each PE core according to specific rules. Therefore, the communication efficiency of the collective communication operator will directly affect the computing efficiency of the AI cluster.

[0005] Large-scale PE arrays will cause the time consumption of collective communication operations between cores to increase dramatically. For 2DMesh arrays, the existing collective communication algorithm is basically implemented using the ring algorithm, which has obvious advantages in small and medium-sized node scales, with a small amount of transmitted data and no bottleneck; however, once the node scale becomes larger, a large number of nodes will form an extremely long ring, and the collective communication time consumption will increase linearly, resulting in reduced communication efficiency. Therefore, the existing technology has the following technical problems: 1. In the existing ring-based collective communication algorithm, all computing nodes in the ring perform collective communication operations in a single direction; the bidirectional connection between computing nodes in the Mesh network topology is not utilized, resulting in insufficient bandwidth utilization of the Mesh network. 2. There are many processor cores in the wafer-level processor, and the array scale is huge, while the collective communication time consumption of the traditional Ring algorithm will increase linearly with the scale of the PE array, and cannot meet the demand for high collective communication efficiency. Summary of the invention

[0006] In view of the deficiencies in the prior art, the present invention provides a processor unit aggregate communication method, a processor system and an acceleration card.

[0007] According to one aspect of the present invention, there is provided a processor unit set communication method, comprising: Divide the target processor unit array for performing the computing task into m first processor unit sub-arrays and n second processor unit sub-arrays, the size of the target processor unit array is 2m×2n, the size of the first processor unit sub-array is 2×2n, and the size of the second processor unit sub-array is 2m×2; controlling the first processor unit subarray to perform collective communication based on a bidirectional ring algorithm to perform a computing task; and The second processor unit subarray is controlled to perform collective communication based on a bidirectional ring algorithm to execute a computing task.

[0008] Optionally, the method further comprises: When at least one of the number of rows and the number of columns of the processor unit array assigned to the computing task is an odd number, the processor unit array assigned to the computing task is reduced according to a preset reduction rule to form a target processor unit array; When the first processor unit sub-array and the second processor unit sub-array complete collective communication, the calculation result of the calculation task is sent to the processor unit other than the target processor unit array in the processor unit array assigned to the calculation task.

[0009] Optionally, when the number of columns of the processor unit array allocated for the computing task is 2n + 1, control each processor unit in a specific column to send the data of the computing task to the processor units in the adjacent column in the same row, so that the processor units in the adjacent column in the same row perform reduction processing on the data of the computing task of this processor unit; When the number of rows of the processor unit array allocated for the computing task is 2m + 1, control each processor unit in a specific row to send the data of the computing task to the processor units in the adjacent row in the same column, so that the processor units in the adjacent row in the same column perform reduction processing on the data of the computing task of this processor unit.

[0010] Optionally, when the number of columns of the processor unit array allocated for the computing task is 2n + 1, the specific column is the first column or the (2n + 1)-th column; When the number of rows of the processor unit array allocated for the computing task is 2m + 1, the specific row is the first row or the (2m + 1)-th row.

[0011] Optionally, controlling the first processor unit sub-array to perform collective communication based on the bidirectional ring algorithm includes: Controlling the first processor unit group in the first processor unit sub-array to perform collective communication based on the unidirectional ring algorithm, and controlling the second processor unit group in the first processor unit sub-array to perform collective communication based on the unidirectional ring algorithm. The first processor unit sub-array is divided into the first processor unit group and the second processor unit group in the form of a processor unit interval, and the directions of the unidirectional ring algorithms used by the first processor unit group and the second processor unit group are opposite; Controlling the second processor unit sub-array to perform collective communication based on the bidirectional ring algorithm includes: Controlling the third processor unit group in the second processor unit sub-array to perform collective communication based on the unidirectional ring algorithm, and controlling the fourth processor unit group in the second processor unit sub-array to perform collective communication based on the unidirectional ring algorithm. The second processor unit sub-array is divided into the third processor unit group and the fourth processor unit group in the form of a processor unit row interval, and the directions of the unidirectional ring algorithms used by the third processor unit group and the fourth processor unit group are opposite.

[0012] Optionally, controlling the first processor unit sub-array to perform collective communication based on the bidirectional ring algorithm includes: Controlling the fifth processor unit group in the first processor unit sub-array to perform collective communication based on the unidirectional ring algorithm, and controlling the sixth processor unit group in the first processor unit sub-array to perform collective communication based on the unidirectional ring algorithm. The first processor unit sub-array is divided into the fifth processor unit group and the sixth processor unit group in the form of processor unit column intervals, and the directions of the unidirectional ring algorithms used by the fifth processor unit group and the sixth processor unit group are opposite; Controlling the second processor unit sub-array to perform collective communication based on the bidirectional ring algorithm, including: Controlling the seventh processor unit group in the second processor unit sub-array to perform collective communication based on the unidirectional ring algorithm, and controlling the eighth processor unit group in the second processor unit sub-array to perform collective communication based on the unidirectional ring algorithm. The second processor unit sub-array is divided into the seventh processor unit group and the eighth processor unit group in the form of processor unit intervals, and the directions of the unidirectional ring algorithms used by the seventh processor unit group and the eighth processor unit group are opposite.

[0013] Optionally, m first processor unit sub-arrays perform collective communication in parallel; n second processor unit sub-arrays perform collective communication in parallel.

[0014] Optionally, controlling the second processor unit sub-array to perform collective communication based on the bidirectional ring algorithm, including: When the m first processor unit sub-arrays complete collective communication, controlling the second processor unit sub-array to perform collective communication based on the bidirectional ring algorithm.

[0015] According to another aspect of the present invention, a processor system is provided, including a compiler and a processor unit array, and the compiler is configured to execute the method described in any of the above aspects of the present invention.

[0016] According to another aspect of the present invention, an acceleration card is provided, including a compiler and a processor unit array, and the compiler is configured to execute the processor system described above in the present invention.

[0017] According to another aspect of the present invention, a server is provided, including a compiler and a processor unit array, and the compiler is configured to execute the acceleration card described above in the present invention.

[0018] Therefore, the present invention proposes a communication method for a set of processor units. Based on the characteristic of bidirectional connection of computing nodes in a 2D Mesh network topology, a bidirectional ring collective communication method is proposed, which can improve the utilization rate of the Mesh network bandwidth. And a hierarchical communication division strategy is proposed; it can effectively solve the disadvantage that the conventional ring algorithm is not applicable to large PE arrays; by hierarchically dividing the nodes in the PE array, the PE array can be sliced into multiple groups of small arrays, and then according to the rules, the PE cores in the small arrays form rings, and the scale of the number of PE cores in the rings is significantly reduced. At the same time, these rings can perform collective communication operations in parallel; thus, the collective communication efficiency of large PE arrays can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] By referring to the following drawings, the exemplary embodiments of the present invention can be more fully understood: Figure 1 FIG. is a schematic flowchart of a hierarchical optimization method for collective communication of processor units based on a bidirectional ring provided by an exemplary embodiment of the present invention; Figure 2 FIG. is a schematic diagram of the implementation of the bidirectional ring-A method provided by an exemplary embodiment of the present invention; Figure 3 FIG. is a schematic diagram of the implementation of the bidirectional ring-B method provided by an exemplary embodiment of the present invention; Figure 4 FIG. is a schematic diagram of reducing an array of odd rows or columns by reduction processing provided by an exemplary embodiment of the present invention; Figure 5 FIG. is a schematic diagram of a hierarchical communication division strategy provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.

[0021] It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps described in these embodiments do not limit the scope of the present invention.

[0022] Those skilled in the art can understand that the terms "first", "second", etc. in the embodiments of the present invention are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.

[0023] It should also be understood that in the embodiments of the present invention, "a plurality" may refer to two or more, and "at least one" may refer to one, two, or more.

[0024] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present invention, in the absence of a clear limitation or a contrary indication in the context, it is generally understood as one or more.

[0025] In addition, the term "and / or" in the present invention is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.

[0026] It should also be understood that the description of each embodiment of the present invention emphasizes the differences between the embodiments, and their similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0027] At the same time, it should be understood that for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0028] The following description of at least one exemplary embodiment is actually merely illustrative and in no way limits the present invention or its application or use.

[0029] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0030] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0031] The solution of the present invention runs on AI chips and computing power chips, and the product form can be these chips and the products equipped with these chips.

[0032] In an embodiment of the present invention, Figure 1 is a schematic flowchart of a communication method for a set of processor units. This embodiment can be applied to AI chips and computing power chips, such as Figure 1 shown, the communication method 100 for a set of processor units includes the following steps: Step 101, divide the target processor unit array for executing a computing task into m first processor unit sub-arrays and n second processor unit sub-arrays. The size of the target processor unit array is 2m×2n, the size of the first processor unit sub-array is 2×2n, and the size of the second processor unit sub-array is 2m×2; Step 102, control the first processor unit sub-array to perform collective communication based on the bidirectional ring algorithm to execute a computing task; and Step 103, control the second processor unit sub-array to perform collective communication based on the bidirectional ring algorithm to execute a computing task.

[0033] Specifically, in view of the technical problems existing in the prior art, the present application proposes a method for collective communication of processor units, which is as follows: Among them, the bidirectional ring algorithm includes: the bidirectional ring has two forms, namely the bidirectional ring-A method and the bidirectional ring-B method; 1) The bidirectional ring-A method: For a 2*N or N*2 PE array (i.e., the processor unit array, hereinafter simply referred to as the PE array), the two types of arrays are equivalent; taking 2*N as an example, as shown in Figure 2 , first number the PEs in the array clockwise, numbered 0~2N-1; group all the PEs with even numbers into a group of forward ring A. In the forward ring A, the data flow of collective communication all flows in the clockwise direction; group all the PEs with odd numbers into a group of reverse ring A'. In the reverse ring A', the data flow of collective communication all flows in the counterclockwise direction. Since the adjacent PEs in the array are bidirectionally connected, the data flows of the forward ring and the reverse ring will not conflict with each other, and the collective communication operations can be executed in parallel.

[0034] All the PEs in the forward ring A use the conventional ring algorithm to perform the all-reduce operation clockwise; all the PEs in the reverse ring A' use the conventional ring algorithm to perform the all-reduce operation counterclockwise.

[0035] 2) The bidirectional ring-B method: For a 2*2M or 2M*2 PE array, the two arrays are equivalent; taking the 2M*2 PE array as an example, as shown in Figure 3 , group all the PEs in the even rows into a group of forward ring B. In the forward ring B, the data flow of collective communication all flows in the clockwise direction; group all the PEs in the odd rows into a group of reverse ring B'. In the reverse ring B', the data flow of collective communication all flows in the counterclockwise direction. Since the adjacent PEs in the array are bidirectionally connected, the data flows of the forward ring B and the reverse ring B' will not conflict with each other, and the collective communication operations can be executed in parallel.

[0036] All the PEs in the forward ring B use the conventional ring algorithm to perform the all-reduce operation clockwise; all the PEs in the reverse ring B' use the conventional ring algorithm to perform the all-reduce operation counterclockwise.

[0037] In an embodiment of the present application, the hierarchical communication partitioning strategy for collective communication based on the bidirectional ring algorithm includes: 1). The hierarchical communication partitioning strategy is applicable to a PE array with both even rows and columns (i.e., 2m * 2n); if the number of rows or columns of the original array is odd, the array needs to be trimmed to the form of 2m * 2n first; specifically, as Figure 4 shown, when the number of columns is odd, the data of each PE core in the (2n + 1)-th column is first sent to the PE core in the adjacent 2n-th column, and then reduced with the data of the PE core itself in the 2n-th column, and the processing result is saved in the corresponding PE core in the 2n-th column; when the number of rows is odd, the data of each PE core in the (2m + 1)-th row is first sent to the PE core in the adjacent 2m-th row, and then reduced with the data of the PE core itself in the 2m-th row, and the processing result is saved in the corresponding PE core in the 2m-th row; when both the number of rows and columns are odd, the rows can be reduced first, and then the columns can be reduced, or the columns can be reduced first, and then the rows can be reduced, so as to obtain an array of even processor units.

[0038] 2). Hierarchically partition the PE array with a size of 2m * 2n; it can be divided into two communication levels, namely the first communication level and the second communication level. Refer to Figure 5 shown, where: The first communication level: The 2m * 2n PE array is sliced by rows, and the slicing unit is an array of 2 * 2n, denoted as slicing unit A, and a total of m slicing units A can be obtained; The second communication level: The 2m * 2n PE array is sliced by columns, and the slicing unit is an array of 2m * 2, denoted as slicing unit B, and a total of n slicing units B can be obtained; When performing an all-reduce communication on the 2m * 2n array, the operation steps can be decomposed into two steps: First, complete the all-reduce operation in the first communication level, where m slicing units A execute the all-reduce operation in parallel, and each slicing unit A uses the bidirectional ring-A method for collective communication; then complete the all-reduce operation in the second communication level, where n slicing units B execute the all-reduce operation in parallel, and each slicing unit B uses the bidirectional ring-B method for collective communication.

[0039] 3) When both the number of rows and columns of the original array are even, the all-reduce operation of the original array is completed, and each processing unit obtains the reduced data. When the number of rows or columns of the original array is odd, the result after all-reduce (i.e., the intermediate result of the calculation) also needs to be sent to the (2m + 1)-th row or the (2n + 1)-th column. Specifically, when the number of columns is odd, first send the reduced data of each PE core in the 2n-th column to the adjacent PE core in the (2n + 1)-th column. When the number of rows is odd, first send the reduced data of each PE core in the 2m-th row to the adjacent PE core in the (2m + 1)-th row. When both the number of rows and columns are odd, restore according to the rules of reduction processing so that each processing unit in the original processing unit array obtains the reduced data.

[0040] Therefore, for a large-scale PE array, applying this technology in the present invention can effectively improve the communication efficiency of the all-reduce collective communication of the PE array. Compared with the conventional ring all-reduce, the collective communication efficiency is significantly improved. The main application scenarios include wafer-level processor chips, wafer-level processor boards, and wafer-level processor servers and clusters. In an embodiment of the present invention, the task scheduler assigns the target computing task to a specific processing unit array, and the assignment result is given to the compiler (the compiler can sense the even-scale and odd-scale configurations (0 / 1)). The target computing task will be split into each PE core in the array. According to the size of the actually used processor array, the selected processor array is divided into a first communication layer and a second communication layer. During the operation of the computing task, the collective communication between PE cores can be performed according to the hierarchical collective communication method provided by the present invention.

[0041] Therefore, based on the characteristic of bidirectional connection of computing nodes in the 2D Mesh network topology, the present invention proposes a bidirectional ring collective communication method, which can improve the utilization rate of the Mesh network bandwidth. And a hierarchical communication division strategy is proposed, which can effectively solve the disadvantage that the conventional ring algorithm is not applicable to large PE arrays. By hierarchically dividing the nodes in the PE array, the PE array can be split into multiple groups of small arrays, and then according to the rules, the PE cores in the small arrays form rings, and the scale of the number of PE cores in the rings is significantly reduced. At the same time, these rings can perform collective communication operations in parallel, thereby significantly improving the collective communication efficiency of large PE arrays.

[0042] In another embodiment of the present invention, a processor system is provided, including a compiler and a processing unit array, and the compiler is configured to execute the method described in any one of the above aspects of the present invention.

[0043] In another embodiment of the present invention, an acceleration card is provided, which includes a compiler and an array of processor units, and the compiler is configured to execute the processor system described above in the present invention.

[0044] In another embodiment of the present invention, a server is provided, which includes a compiler and an array of processor units, and the compiler is configured to execute the acceleration card described above in the present invention.

[0045] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present invention to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some variations, modifications, alterations, additions, and subcombinations thereof.

Claims

1. A communication method for a set of processor units, characterized in that Including: Dividing a target processor unit array for executing a computing task into m first processor unit sub-arrays and n second processor unit sub-arrays, where the size of the target processor unit array is 2m×2n, the size of the first processor unit sub-array is 2×2n, and the size of the second processor unit sub-array is 2m×2; Controlling the first processor unit sub-array to perform collective communication based on the bidirectional ring algorithm to execute the computing task; And Controlling the second processor unit sub-array to perform collective communication based on the bidirectional ring algorithm to execute the computing task.

2. The method according to claim 1, wherein The method further includes: When at least one of the number of rows and the number of columns of the processor unit array allocated for the computing task is odd, performing a reduction process on the processor unit array allocated for the computing task according to a preset reduction rule to form the target processor unit array; When the first processor unit sub-array and the second processor unit sub-array complete collective communication, sending the calculation result of the computing task to the processor units outside the target processor unit array in the processor unit array allocated for the computing task.

3. The method according to claim 2, wherein: When the number of columns of the processor unit array allocated for the computing task is 2n + 1, controlling each processor unit in a specific column to send the data of the computing task to the processor unit in the adjacent column of the same row, so that the processor unit in the adjacent column of the same row performs a reduction process on the data of the computing task of this processor unit; When the number of rows of the processor unit array allocated for the computing task is 2m + 1, controlling each processor unit in a specific row to send the data of the computing task to the processor unit in the adjacent row of the same column, so that the processor unit in the adjacent row of the same column performs a reduction process on the data of the computing task of this processor unit.

4. The method according to claim 3, wherein: When the number of columns of the processor unit array allocated for the computing task is 2n + 1, the specific column is the first column or the (2n + 1)-th column; When the number of rows of the processor unit array allocated for the computing task is 2m + 1, the specific row is the first row or the (2m + 1)-th row.

5. The method according to claim 1, wherein Controlling the first processor unit sub-array to perform collective communication based on the bidirectional ring algorithm includes: Controlling the first processor unit group in the first processor unit sub-array to perform collective communication based on the unidirectional ring algorithm, and controlling the second processor unit group in the first processor unit sub-array to perform collective communication based on the unidirectional ring algorithm. The first processor unit sub-array is divided into the first processor unit group and the second processor unit group in a processor unit interval manner, and the directions of the unidirectional ring algorithms used by the first processor unit group and the second processor unit group are opposite; Controlling the second processor unit sub-array to perform collective communication based on the bidirectional ring algorithm includes: Controlling the third processor unit group in the second processor unit sub-array to perform collective communication based on the one-way ring algorithm, and controlling the fourth processor unit group in the second processor unit sub-array to perform collective communication based on the one-way ring algorithm. The second processor unit sub-array is divided into the third processor unit group and the fourth processor unit group in the form of a processor unit row interval, and the directions of the one-way ring algorithms used by the third processor unit group and the fourth processor unit group are opposite.

6. The method according to claim 1, characterized in that, Controlling the first processor unit sub-array to perform collective communication based on the two-way ring algorithm, including: Controlling the fifth processor unit group in the first processor unit sub-array to perform collective communication based on the one-way ring algorithm, and controlling the sixth processor unit group in the first processor unit sub-array to perform collective communication based on the one-way ring algorithm. The first processor unit sub-array is divided into the fifth processor unit group and the sixth processor unit group in the form of a processor unit column interval, and the directions of the one-way ring algorithms used by the fifth processor unit group and the sixth processor unit group are opposite; Controlling the second processor unit sub-array to perform collective communication based on the two-way ring algorithm, including: Controlling the seventh processor unit group in the second processor unit sub-array to perform collective communication based on the one-way ring algorithm, and controlling the eighth processor unit group in the second processor unit sub-array to perform collective communication based on the one-way ring algorithm. The second processor unit sub-array is divided into the seventh processor unit group and the eighth processor unit group in the form of a processor unit interval, and the directions of the one-way ring algorithms used by the seventh processor unit group and the eighth processor unit group are opposite.

7. The method according to claim 5 or 6, characterized in that, The m first processor unit sub-arrays perform collective communication in parallel; the n second processor unit sub-arrays perform collective communication in parallel.

8. The method according to claim 5 or 6, characterized in that The controlling the second processor unit sub-array to perform collective communication based on the two-way ring algorithm includes: When the m first processor unit sub-arrays complete collective communication, controlling the second processor unit sub-array to perform collective communication based on the two-way ring algorithm.

9. A processor system, characterized in that, Including a compiler and a processor unit array, the compiler being configured to execute the method according to any one of claims 1 to 8.

10. An acceleration card, characterized in that, Including the processor system according to claim 9.

11. A server, characterized in that, Including the acceleration card according to claim 10.

Citation Information

Patent Citations

  • Hardware accelerator architecture and template for web-scale k-means clustering

    CN108268320A

  • Method and device for executing communication task in accelerator card system

    CN114764374A

  • Data processor core, data processor, electronic device, and storage medium

    CN119003001A

  • Load balancing method of low-power AI processor, chip and storage medium

    CN119271418A

  • Techniques for collective operations in distributed systems

    US20190042527A1