Processor chip, collective communication method and electronic device
By dividing the processor core array into hierarchical parallel loop communications, the problem of low communication efficiency in the existing technology is solved, and efficient communication of large-scale processor chips is achieved.
Patent Information
- Application Number
- CN202510652717.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-05-20
AI Technical Summary
In existing technologies, tree-based algorithms are prone to link contention and network congestion in AI clusters, resulting in low communication efficiency of processor chips in 2D Mesh arrays, especially in the case of large-scale nodes, where communication time increases linearly.
A hierarchical communication strategy is adopted to divide the processor core array into multiple division units in row and column directions, and a parallel loop communication is formed in each division unit to communicate through the loops of the first and second communication levels.
The collective communication efficiency of processor chips is improved, especially the communication performance is significantly improved in large-scale processor core arrays.
Smart Images

Figure CN120179420B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a processor chip, a collective communication method, and an electronic device. Background Art
[0002] Collective communication refers to the data exchange and collaborative work between various processor cores (Processing Elements, PEs) based on specific rules. The communication efficiency of collective communication operators will directly affect the computing efficiency of the artificial intelligence (AI) cluster.
[0003] In existing technologies, the communication algorithms commonly used in AI clusters are mainly divided into tree-based algorithms and ring-based algorithms. Since tree-based algorithms are used for collective communication, link contention is prone to occur, leading to severe network congestion on some paths and affecting collective communication efficiency. Therefore, for large-scale PE arrays, ring algorithms are often used to achieve collective communication between PEs. For PEs in a two-dimensional mesh (2D mesh) array, existing collective communication algorithms are generally implemented using ring algorithms. However, for large-scale nodes, the large number of nodes will form extremely long loops, and the collective communication time increases linearly, resulting in reduced communication efficiency. Therefore, how to improve the communication efficiency of processor chips with PEs in a 2D mesh array has become an important issue that needs to be addressed in this field. Summary of the Invention
[0004] In response to the problems in the prior art, embodiments of the present invention provide a processor chip, a collective communication method, and an electronic device, which can at least partially solve the problems in the prior art.
[0005] In a first aspect, the present invention provides a processor chip, comprising a plurality of processor cores arranged in an array, wherein:
[0006] The multiple processor cores are hierarchically divided according to a first communication layer division rule to obtain multiple first slicing units of the first communication layer in a first array direction; each first slicing unit includes multiple processor cores, and the processor cores included in each first slicing unit form a corresponding first loop according to a first loop formation rule;
[0007] The multiple processor cores are hierarchically divided according to a second communication hierarchical division rule to obtain a plurality of second slicing units of the second communication hierarchical level in a second array direction; each second slicing unit includes a plurality of processor cores, and the processor cores included in each second slicing unit form a corresponding second loop according to a second loop formation rule;
[0008] Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates based on a first loop; each second slicing unit of the second communication level communicates in parallel, and each processor core included in each second slicing unit communicates based on a second loop.
[0009] Furthermore, the first array direction is a row direction and the second array direction is a column direction; accordingly, the first communication level division rule includes dividing the multiple processor cores arranged in the array into rows, and the processor cores in each row serve as a first dividing unit; the second communication level division rule includes dividing the multiple processor cores arranged in the array into columns, and the processor cores in each column serve as a second dividing unit.
[0010] Furthermore, the first ring-forming rule includes:
[0011] The multiple processor cores included in the first splitting unit are numbered in sequence along the row direction, with the numbers starting from 1; if the maximum value n of the numbering is an even number, the processor cores with odd numbers are connected in series in the row direction, the n-1th processor core is connected in series with the nth processor core, the processor cores with even numbers are connected in series in the opposite direction of the row direction, and the second processor core is connected in series with the first processor core to form a first loop, and the direction of the first loop is clockwise or counterclockwise; if the maximum value n of the numbering is an odd number, the processor cores with odd numbers are connected in series in the row direction, the nth processor core is connected in series with the n-1th processor core in the opposite direction of the row direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a first loop; wherein the direction of the first loop is clockwise or counterclockwise.
[0012] Furthermore, the second ring-forming rule includes:
[0013] The multiple processor cores included in the second splitting unit are numbered in sequence along the column direction, with the numbers starting from 1; if the maximum value m of the numbering is an even number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series with the mth processor core, the processor cores with even numbers are connected in series in the opposite direction of the column direction, and the second processor core is connected in series with the first processor core to form a second loop; if the maximum value m of the numbering is an odd number, the processor cores with odd numbers are connected in series in the column direction, the mth processor core is connected in series with the m-1th processor core in the opposite direction of the column direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a second loop; wherein the direction of the second loop is clockwise or counterclockwise.
[0014] Furthermore, the first array direction is a column direction and the second array direction is a row direction; accordingly, the first communication level division rule includes dividing the multiple processor cores arranged in the array into columns, and the processor cores in each column serve as a first dividing unit; the second communication level division rule includes dividing the multiple processor cores arranged in the array into rows, and the processor cores in each row serve as a second dividing unit.
[0015] Furthermore, the first ring-forming rule includes:
[0016] The multiple processor cores included in the first splitting unit are numbered in sequence along the column direction, with the numbers starting from 1; if the maximum value m of the numbering is an even number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series with the mth processor core, the processor cores with even numbers are connected in series in the opposite direction of the column direction, and the second processor core is connected in series with the first processor core to form a first loop; if the maximum value m of the numbering is an odd number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series in the opposite direction of the column direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a first loop; wherein the direction of the first loop is clockwise or counterclockwise.
[0017] Furthermore, the second ring-forming rule includes:
[0018] The multiple processor cores included in the second splitting unit are numbered in sequence along the row direction, with the numbers starting from 1; if the maximum value n of the numbering is an even number, the processor cores numbered odd are connected in series in the row direction, the n-1th processor core is connected in series with the nth processor core, the processor cores numbered even are connected in series in the opposite direction of the row direction, and the second processor core is connected in series with the first processor core to form a second loop, and the direction of the second loop is clockwise or counterclockwise; if the maximum value n of the numbering is an odd number, the processor cores numbered odd are connected in series in the row direction, the nth processor core is connected in series with the n-1th processor core in the opposite direction of the row direction, the processor cores numbered even are connected in series, and the second processor core is connected in series with the first processor core to form a second loop; wherein the direction of the second loop is clockwise or counterclockwise; wherein the direction of the second loop is clockwise or counterclockwise.
[0019] Furthermore, the multiple processor cores are arranged in a 2m×2n array; the first array direction is a row direction and the second array direction is a column direction; accordingly, the first communication level division rule includes dividing the 2m rows of processor cores into units of 2 rows, and every 2 rows of processor cores serve as a first splitting unit; the second communication level division rule includes dividing the 2n columns of processor cores by columns, and the processor cores in the i-th column and the i+n-th column serve as a second splitting unit, where i is greater than or equal to 1 and less than or equal to n.
[0020] Furthermore, the first ring-forming rule includes:
[0021] The 2×2n processor cores included in the first split unit are evenly divided into two sub-units, the first sub-unit includes the processor cores in the 1st to nth columns in the first split unit, and the second sub-unit includes the processor cores in the n+1th to 2nth columns in the first split unit; the processor cores included in the first sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop, and the processor cores included in the second sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop.
[0022] Furthermore, the second ring-forming rule includes:
[0023] The 2m×2 processor cores included in the second split unit are divided into two sub-units according to odd and even rows. The third sub-unit includes the processor cores in the odd rows of the second split unit, and the fourth sub-unit includes the processor cores in the even rows of the second split unit. The processor cores included in the third sub-unit are serially connected in a first direction to form a loop, and the processor cores included in the fourth sub-unit are serially connected in a second direction to form a loop. The first direction and the second direction are opposite.
[0024] Furthermore, the multiple processor cores are arranged in a 2m×2n array; the first array direction is the column direction and the second array direction is the row direction; accordingly, the first communication level division rule includes dividing the 2n columns of processor cores into 2 columns, with the processor cores in every 2 columns serving as a first splitting unit; the second communication level division rule includes dividing the 2m rows of processor cores into rows, with the processor cores in the jth row and the j+mth row serving as a second splitting unit, where j is greater than or equal to 1 and less than or equal to m.
[0025] Furthermore, the first ring-forming rule includes:
[0026] The 2m×2 processor cores included in the first split unit are evenly divided into two sub-units, the first sub-unit includes the processor cores in the 1st to mth rows in the first split unit, and the second sub-unit includes the processor cores in the m+1th to 2mth rows in the first split unit; the processor cores included in the first sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop, and the processor cores included in the second sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop.
[0027] Furthermore, the second ring-forming rule includes:
[0028] The 2×2n processor cores included in the second split unit are divided into two sub-units according to odd and even columns. The third sub-unit includes the processor cores in the odd columns of the second split unit, and the fourth sub-unit includes the processor cores in the even columns of the second split unit. The processor cores included in the third sub-unit are connected in series in a first direction to form a loop, and the processor cores included in the fourth sub-unit are connected in series in a second direction to form a loop. The first direction and the second direction are opposite.
[0029] In a second aspect, the present invention provides a collective communication method, applied to the processor chip described in any one of the above embodiments, comprising:
[0030] Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates according to a ring full reduce operation based on the first loop;
[0031] Each second slicing unit of the second communication level communicates in parallel, and each processor core included in each second slicing unit communicates according to a ring full reduce operation based on the second loop.
[0032] Furthermore, the multiple processor cores are arranged in a 2m×2n array, and each first split unit includes a first sub-unit and a second sub-unit; accordingly, the processor cores included in each first split unit communicate based on the first loop according to the ring full reduce operation, including:
[0033] Each processor core in a first subunit included in each first splitting unit communicates according to a ring full reduce operation based on a clockwise loop to process the 1st to 2xth shares of communication data of each processor core, and communicates according to a ring full reduce operation based on a counterclockwise loop to process the 2x+1st to 4xth shares of data of each processor core; wherein the communication data of each processor core in the first subunit is divided into 4x shares, where x is the total number of processor cores included in the first subunit;
[0034] Each processor core in the second sub-unit included in each first slicing unit communicates according to a ring full reduce operation based on a clockwise loop to process the 1st to 2xth shares of communication data of each processor core, and communicates according to a ring full reduce operation based on a counterclockwise loop to process the 2x+1st to 4xth shares of data of each processor core; wherein the communication data of each processor core in the second sub-unit is divided into 4x shares, where x is the total number of processor cores included in the second sub-unit;
[0035] Each first splitting unit includes a first subunit and a second subunit that perform a parallel ring full reduction operation.
[0036] In a third aspect, the present invention provides a collective communication method, applied to the processor chip described in any one of the above embodiments, comprising:
[0037] Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates according to a reduce-scatter operation based on the first loop;
[0038] Each second slicing unit of the second communication level communicates in parallel, and each processor core included in each second slicing unit communicates according to a ring full reduce operation based on the second loop;
[0039] Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates according to an aggregation operation based on the first loop.
[0040] Furthermore, the multiple processor cores are arranged in a 2m×2n array, and each first slicing unit includes a first sub-unit and a second sub-unit; accordingly, the processor cores included in each first slicing unit communicate based on the first loop according to the reduce-scatter operation, including:
[0041] Each processor core in a first subunit included in each first splitting unit communicates according to a reduce-scatter operation based on a clockwise loop to process the 1st to 2yth shares of communication data of the processor core, and communicates according to a reduce-scatter operation based on a counterclockwise loop to process the 2y+1st to 4yth shares of data of the processor core; wherein the communication data of each processor core in the first subunit is divided into 4y shares, where y is the total number of processor cores included in the first subunit;
[0042] Each processor core in the second sub-unit included in each first slicing unit communicates according to a reduce-scatter operation based on a clockwise loop to process the 1st to 2yth shares of communication data of the processor core, and communicates according to a reduce-scatter operation based on a counterclockwise loop to process the 2y+1st to 4yth shares of data of the processor core; wherein the communication data of each processor core in the second sub-unit is divided into 4y shares, where y is the total number of processor cores included in the second sub-unit;
[0043] Each first segmentation unit includes a first subunit and a second subunit that perform a reduction and scatter operation in parallel.
[0044] Furthermore, the processor cores included in each first slicing unit communicate according to the aggregation operation based on the first loop, including:
[0045] The processor cores in the first sub-units included in each first slicing unit communicate according to the aggregation operation based on a clockwise loop, and communicate according to the aggregation operation based on a counterclockwise loop;
[0046] The processor cores in the second sub-unit included in each first slicing unit communicate according to the aggregation operation based on the clockwise loop, and communicate according to the aggregation operation based on the counterclockwise loop;
[0047] The first subunit and the second subunit included in each first splitting unit are aggregated in parallel.
[0048] In a fourth aspect, the present invention provides an electronic device comprising the processor chip described in any one of the above embodiments.
[0049] Embodiments of the present invention provide a processor chip, a collective communication method, and an electronic device. The processor chip includes multiple processor cores, which are arranged in an array. The multiple processor cores are hierarchically divided according to a first communication layer division rule, and multiple first slicing units of the first communication layer are obtained in a first array direction; each first slicing unit includes multiple processor cores, and the processor cores included in each first slicing unit form a corresponding first loop according to a first ring formation rule; the multiple processor cores are hierarchically divided according to a second communication layer division rule, and multiple second slicing units of the second communication layer are obtained in a second array direction; each second slicing unit includes multiple processor cores, and the processor cores included in each second slicing unit form a corresponding second loop according to a second ring formation rule; each first slicing unit of the first communication layer communicates in parallel, and the processor cores included in each first slicing unit communicate based on the first loop; each second slicing unit of the second communication layer communicates in parallel, and the processor cores included in each second slicing unit communicate based on the second loop. Since the hierarchical division into multiple loops communicates in parallel, the collective communication efficiency of the processor chip is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 This is a schematic diagram of the all-reduce communication process provided by one embodiment of the present invention.
[0052] Figure 2 It is a structural diagram of a processor chip provided by one embodiment of the present invention.
[0053] Figure 3 This is a schematic diagram of communication layer division provided by an embodiment of the present invention.
[0054] Figure 4 is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0055] Figure 5 is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0056] Figure 6 is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0057] Figure 7is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0058] Figure 8 This is a schematic diagram of communication layer division provided by an embodiment of the present invention.
[0059] Figure 9 is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0060] Figure 10 is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0061] Figure 11 is a schematic diagram of a second loop provided by an embodiment of the present invention.
[0062] Figure 12 is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0063] Figure 13 is a schematic diagram of a first loop provided by an embodiment of the present invention.
[0064] Figure 14 is a schematic diagram of a second loop provided by an embodiment of the present invention.
[0065] Figure 15 It is a flowchart of a collective communication method provided by one embodiment of the present invention.
[0066] Figure 16 It is a flowchart of a collective communication method provided by one embodiment of the present invention.
[0067] Figure 17 This is a schematic diagram of communication layer division provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the embodiments of the present invention are further described in detail with reference to the accompanying drawings. Here, the schematic embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention. It should be noted that, in the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other in any way. The acquisition, storage, use, processing, etc. of data in the technical solutions in this application comply with the relevant provisions of laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant channels, and the acquisition, storage, use, processing, etc. of user information are agreed to by the customer.
[0069] In order to facilitate understanding of the technical solution provided by this application, the relevant contents of the technical solution of this application are first explained below.
[0070] With the rapid development of deep learning and artificial intelligence technologies, the parameter scale of AI models is becoming increasingly large. The computational and memory bandwidth resources of traditional AI processors are no longer sufficient to meet the training and inference requirements of these large-scale AI models. To address the bottlenecks of large-model training, such as limited computing power and memory bandwidth, a variety of new processor architectures have been proposed in academia and industry. Among them, wafer-level processor chips, due to their unique architectural design, have far superior computing power and memory bandwidth than traditional AI chips. This makes them stand out among many new processor architectures and has been widely researched and applied in the industry.
[0071] Wafer-scale processor chips consist of an on-chip interconnect network and a large number of processor cores. These PEs are arranged in an array, with each PE forming a bidirectional connection with adjacent PEs using a 2D mesh structure, forming a wafer-level processor chip. The PE arrays in wafer-level processor chips are typically large, often including tens of thousands or even more PEs. Therefore, for processors with large PE arrays, optimizing communication efficiency between the internal processor cores is particularly important.
[0072] Large-scale PE arrays lead to a dramatic increase in the time spent on collective communication operations between PEs. For 2DMesh arrays, existing collective communication algorithms typically use ring algorithms. Large numbers of PEs form extremely long loops, which linearly increase the time spent on collective communication and reduce communication efficiency.
[0073] Wafer-scale processor chips contain numerous processor cores, resulting in massive arrays. However, the collective communication time of the traditional Ring algorithm scales linearly with the size of the PE array, making it inefficient. This paper proposes a hierarchical communication strategy that effectively addresses the low communication efficiency of the conventional Ring algorithm in large-scale PE arrays.
[0074] Common collective communication operations include broadcast, reduce, all-reduce, gather, scatter, and all-gather. Ring algorithms are used in collective communication. Taking all-reduce based on the ring algorithm as an example, the communication process can be divided into two phases: reduce-scatter and gather. Both reduce-scatter and all-gather are implemented using the ring algorithm. Reduce-scatter is a collective communication operation that combines reduce and scatter. In a reduce-scatter operation, each node sends 1 / N of its data to its neighboring nodes in a specific direction (clockwise or counterclockwise). Upon receiving this data, the neighboring nodes perform reduction processing, which includes but is not limited to summing, finding the maximum, minimum, and average values. Ultimately, each node receives 1 / N of the reduced data. All-gather is also a collective communication operation. During each communication, each node sends 1 / N of its data to adjacent nodes in a specific direction. Upon receiving the data, the adjacent nodes directly replace their own corresponding data. Ultimately, each node receives the complete reduced data. N is the total number of nodes participating in the collective communication.
[0075] like Figure 1 The figure shows the all-reduce communication process of four nodes, a, b, c, and d. The four nodes form a loop, splitting each node's data into N parts (N represents the number of nodes in the loop, here 4). Node a splits the data to be communicated into four parts: a0, a1, a2, and a3; node b splits the data to be communicated into four parts: b0, b1, b2, and b3; node c splits the data to be communicated into four parts: c0, c1, c2, and c3; and node d splits the data to be communicated into four parts: d0, d1, d2, and d3.
[0076] Reduce-Scatter process: Each time 4 nodes communicate, they communicate in a certain direction (clockwise or counterclockwise, Figure 1 The next node will perform reduction processing after receiving the data (reduction processing can be summation, maximum value, minimum value, average value, etc., here it is summation processing), and finally each node obtains 1 / 4 of the reduced data;
[0077] All-Gather process: Each time the four nodes communicate, they send 1 / 4 of the data to the adjacent node in the direction determined above. After receiving the data, the next node directly replaces the corresponding part of its own data. Finally, each node obtains the complete reduced data.
[0078] The time consumption of collective communication based on the Ring algorithm increases linearly with the increase in the number of nodes in the loop, which cannot meet the requirements of collective communication efficiency. Therefore, this application divides the processor cores in the processor core array into multiple groups of small arrays, and then forms a loop with the processor cores in each group of small arrays. The number of processor cores in the loop is significantly reduced. At the same time, the loop composed of each small array can also perform collective communication operations in parallel, thereby significantly improving the collective communication efficiency of the processor core array, especially for large-scale processor core arrays. The improvement in collective communication efficiency is very obvious.
[0079] Figure 2 FIG. 1 is a schematic diagram of the structure of a processor chip provided by an embodiment of the present invention. Figure 2 As shown, the processor chip provided by the embodiment of the present invention includes multiple processor cores 1, and the multiple processor cores 1 are arranged in an array, wherein:
[0080] The multiple processor cores 1 are hierarchically divided according to a first communication hierarchical division rule, to obtain a plurality of first slicing units 10 of a first communication hierarchical level in a first array direction; each first slicing unit 10 includes a plurality of processor cores 1, and the processor cores 1 included in each first slicing unit 10 form a corresponding first loop according to a first loop formation rule;
[0081] The multiple processor cores 1 are hierarchically divided according to the second communication hierarchical division rule, and multiple second slicing units 20 of the second communication hierarchical are obtained in the second array direction; each second slicing unit 20 includes multiple processor cores 1, and the processor cores 1 included in each second slicing unit 20 form a corresponding second loop according to the second loop formation rule;
[0082] Each first slicing unit 10 of the first communication level communicates in parallel, and each processor core 1 included in each first slicing unit communicates based on the first loop; each second slicing unit 20 of the second communication level communicates in parallel, and each processor core 1 included in each second slicing unit 20 communicates based on the second loop.
[0083] Specifically, the processor cores 1 included in the processor chip are arranged in an array. In order to divide the processor cores 1 arranged in the overall array into multiple small arrays, the multiple processor cores 1 are hierarchically divided in the first array direction according to the first communication level division rule to obtain multiple first segmentation units 10 of the first communication level. The processor core 1 in each first segmentation unit 10 forms a corresponding first loop according to the first ring-forming rule. The number of processor cores 1 in the first loop is significantly reduced, which is conducive to improving the collective communication efficiency of the processor cores 1 based on the first loop. In addition, each first segmentation unit 10 is independent of each other and can communicate in parallel, further improving the collective communication efficiency. Among them, the first communication level division rule and the first ring-forming rule are set according to actual needs, and are not limited in the embodiment of the present invention.
[0084] In the second array direction, the multiple processor cores 1 are hierarchically divided according to the second communication layer division rule to obtain multiple second segmentation units 20 of the second communication layer. The processor core 1 in each second segmentation unit 20 forms a corresponding second loop according to the second ring formation rule. The number of processor cores 1 in the second loop is significantly reduced, which is conducive to improving the collective communication efficiency of the processor cores 1 based on the second loop. In addition, each second segmentation unit 20 is independent of each other and can communicate in parallel, further improving the collective communication efficiency. Among them, the second communication layer division rule and the second ring formation rule are set according to actual needs and are not limited in the embodiment of the present invention.
[0085] The first array direction may be the row direction or the column direction of the array arrangement, and the second array direction may be the row direction or the column direction of the array arrangement. When the first array direction is the row direction, the second array direction is the column direction; when the first array direction is the column direction, the second array direction is the row direction.
[0086] A processor chip provided by an embodiment of the present invention includes multiple processor cores, which are arranged in an array. The multiple processor cores are hierarchically divided according to a first communication layer division rule, and multiple first slicing units of the first communication layer are obtained in the first array direction; each first slicing unit includes multiple processor cores, and the processor cores included in each first slicing unit form a corresponding first loop according to a first ring rule; the multiple processor cores are hierarchically divided according to a second communication layer division rule, and multiple second slicing units of the second communication layer are obtained in the second array direction; each second slicing unit includes multiple processor cores, and the processor cores included in each second slicing unit form a corresponding second loop according to the second ring rule; each first slicing unit of the first communication layer communicates in parallel, and the processor cores included in each first slicing unit communicate based on the first loop; each second slicing unit of the second communication layer communicates in parallel, and the processor cores included in each second slicing unit communicate based on the second loop. Since the hierarchical division into multiple loops communicates in parallel, the collective communication efficiency of the processor chip is improved.
[0087] On the basis of the above embodiments, further, the first array direction is a row direction and the second array direction is a column direction; accordingly, the first communication level division rule includes dividing the multiple processor cores arranged in the array into rows, and the processor cores in each row serve as a first dividing unit; the second communication level division rule includes dividing the multiple processor cores arranged in the array into columns, and the processor cores in each column serve as a second dividing unit.
[0088] Specifically, when the first array direction is a row direction and the second array direction is a column direction, the first communication level division rule includes dividing the multiple processor cores arranged in the array by rows, with the processor cores in each row serving as a first division unit, that is, the processor cores in each row constitute a first division unit. The second communication level division rule includes dividing the multiple processor cores arranged in the array by columns, with the processor cores in each column serving as a second division unit, that is, the processor cores in each column constitute a second division unit.
[0089] For example, Figure 3As shown, the processor cores arranged in an m×n array are divided by rows, with the PEs in each row forming a first split unit A, for a total of m first split units: first split unit A-1, first split unit A-2, first split unit A-3, ..., first split unit Am-1, first split unit Am. The processor cores arranged in an m×n array are divided by columns, with the PEs in each column forming a second split unit B, for a total of n second split units: second split unit B-1, second split unit B-2, second split unit B-3, ..., second split unit Bn-1, second split unit Bn.
[0090] Based on the above embodiments, further, the first ring-forming rule includes:
[0091] The multiple processor cores included in the first splitting unit are numbered in sequence along the row direction, with the numbers starting from 1; if the maximum value n of the numbering is an even number, the processor cores with odd numbers are connected in series in the row direction, the n-1th processor core is connected in series with the nth processor core, the processor cores with even numbers are connected in series in the opposite direction of the row direction, and the second processor core is connected in series with the first processor core to form a first loop, and the direction of the first loop is clockwise or counterclockwise; if the maximum value n of the numbering is an odd number, the processor cores with odd numbers are connected in series in the row direction, the nth processor core is connected in series with the n-1th processor core in the opposite direction of the row direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a first loop; wherein the direction of the first loop is clockwise or counterclockwise.
[0092] Specifically, when the first communication layer division rule performs layer division in the row direction, the first ring formation rule is used to determine the ring formation mode of each processor core in each first segmentation unit.
[0093] For example, Figure 4 As shown in the figure, the formation process of the first loop is described using a first split unit as an example. The n PEs included in the first split unit are numbered from left to right along the row direction, with n being an even number. The odd-numbered PEs are serially connected from left to right. When connected to the n-1th PE, the n-1th PE is serially connected to the nth PE. Then, the even-numbered PEs are serially connected from right to left. When connected to the second PE, the second PE is serially connected to the first PE, forming the first loop. Figure 4 The direction of the first loop is clockwise.
[0094] For example, Figure 5As shown in the figure, the formation process of the first loop is described using a first split unit as an example. The n PEs included in the first split unit are numbered from left to right along the row direction, with n being an odd number. The odd-numbered PEs are serially connected from left to right until the nth PE is connected. The PEs are then serially connected from right to left, connecting the nth PE to the n-1th PE. The even-numbered PEs are serially connected from right to left until the second PE is connected. The second PE is then serially connected to the first PE, forming the first loop. Figure 5 The direction of the first loop is clockwise.
[0095] Based on the above embodiments, further, the second ring-forming rule includes:
[0096] The multiple processor cores included in the first splitting unit are numbered in sequence along the row direction, with the numbers starting from 1; if the maximum value n of the numbering is an even number, the processor cores with odd numbers are connected in series in the row direction, the n-1th processor core is connected in series with the nth processor core, the processor cores with even numbers are connected in series in the opposite direction of the row direction, and the second processor core is connected in series with the first processor core to form a second loop, and the direction of the second loop is clockwise or counterclockwise; if the maximum value n of the numbering is an odd number, the processor cores with odd numbers are connected in series in the row direction, the nth processor core is connected in series with the n-1th processor core in the opposite direction of the row direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a second loop; wherein the direction of the second loop is clockwise or counterclockwise.
[0097] Specifically, when the second communication level division rule performs level division in the row direction, the second ring formation rule is used to determine the ring formation mode of each processor core in each second segmentation unit.
[0098] On the basis of the above embodiments, further, the first array direction is a column direction and the second array direction is a row direction; accordingly, the first communication level division rule includes dividing the multiple processor cores arranged in the array into columns, and the processor cores in each column serve as a first dividing unit; the second communication level division rule includes dividing the multiple processor cores arranged in the array into rows, and the processor cores in each row serve as a second dividing unit.
[0099] Specifically, the division of the first communication level and the second communication level is symmetrical, that is, the division of the first communication level and the division of the second communication level are interchangeable. When the first array direction is the column direction and the second array direction is the row direction, then the first communication level division rule includes dividing the multiple processor cores arranged in the array by columns, with the processor cores in each column serving as a first division unit, that is, the processor cores in each column constitute a first division unit. The second communication level division rule includes dividing the multiple processor cores arranged in the array by rows, with the processor cores in each row serving as a second division unit, that is, the processor cores in each row constitute a second division unit.
[0100] Based on the above embodiments, further, the first ring-forming rule includes:
[0101] The multiple processor cores included in the first splitting unit are numbered in sequence along the column direction, with the numbers starting from 1; if the maximum value m of the numbering is an even number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series with the mth processor core, the processor cores with even numbers are connected in series in the opposite direction of the column direction, and the second processor core is connected in series with the first processor core to form a first loop; if the maximum value m of the numbering is an odd number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series in the opposite direction of the column direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a first loop; wherein the direction of the first loop is clockwise or counterclockwise.
[0102] Specifically, when the first communication level division rule performs level division in the column direction, the first ring formation rule is used to determine the ring formation mode of each processor core in each first segmentation unit.
[0103] For example, Figure 6 As shown in the figure, the formation process of the first loop is described using a first split unit as an example. The m PEs included in the first split unit are numbered from top to bottom in the row direction, with m being an even number. The odd-numbered PEs are serially connected from top to bottom. When connected to the m-1th PE, the n-1th PE is serially connected to the nth PE. Then, the even-numbered PEs are serially connected from bottom to top. When connected to the second PE, the second PE is serially connected to the first PE, forming the first loop. Figure 6 The direction of the first loop is counterclockwise.
[0104] For example, Figure 7As shown in the figure, the formation process of the first loop is explained using a first split unit as an example. The m PEs included in the first split unit are numbered from top to bottom in the row direction, with m being an odd number. The odd-numbered PEs are serially connected from top to bottom, until they are connected to the mth PE. Then, the PEs are serially connected from bottom to top, connecting the mth PE to the m-1th PE. The even-numbered PEs are serially connected from bottom to top, until they are connected to the second PE. The second PE is then serially connected to the first PE, forming the first loop. Figure 7 The direction of the first loop is counterclockwise.
[0105] Based on the above embodiments, further, the second ring-forming rule includes:
[0106] The multiple processor cores included in the second splitting unit are numbered in sequence along the column direction, with the numbers starting from 1; if the maximum value m of the numbering is an even number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series with the mth processor core, the processor cores with even numbers are connected in series in the opposite direction of the column direction, and the second processor core is connected in series with the first processor core to form a second loop; if the maximum value m of the numbering is an odd number, the processor cores with odd numbers are connected in series in the column direction, the mth processor core is connected in series with the m-1th processor core in the opposite direction of the column direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a second loop; wherein the direction of the second loop is clockwise or counterclockwise.
[0107] Specifically, when the second communication level division rule performs level division in the column direction, the second ring formation rule is used to determine the ring formation mode of each processor core in each second segmentation unit.
[0108] Based on the above embodiments, further, the multiple processor cores are arranged in a 2m×2n array;
[0109] The first array direction is a row direction and the second array direction is a column direction; accordingly, the first communication level division rule includes dividing the 2m rows of processor cores into units of 2 rows, with the processor cores in every 2 rows serving as a first division unit; the second communication level division rule includes dividing the 2n columns of processor cores into columns, with the processor cores in the i-th column and the i+n-th column serving as a second division unit, where i is greater than or equal to 1 and less than or equal to n;
[0110] Specifically, the multiple processor cores are arranged in a 2m×2n array, where m and n are positive integers. When the first array direction is the row direction and the second array direction is the column direction, the first communication level division rule includes dividing the 2m rows of processor cores into units of 2 rows, with every 2 rows of processor cores serving as a first splitting unit, i.e., every 2 rows of processor cores constitute a first splitting unit, resulting in a total of m first splitting units. The second communication level division rule includes dividing the 2n columns of processor cores by column, with the processor cores in the i-th column and the i+n-th column serving as a second splitting unit, i.e., each second splitting unit includes 2 columns of processor cores, resulting in a total of n second splitting units. Wherein, i is greater than or equal to 1 and less than or equal to n.
[0111] For example, Figure 8 As shown, the processor cores arranged in a 2m×2n array are divided by rows, with every two rows of PEs forming a first split unit C, for a total of m first split units: first split unit C-1, first split unit C-2, ..., first split unit Cm. The processor cores arranged in a 2m×2n array are divided by columns, with the processor cores in columns i and i+n forming a second split unit. Each second split unit D includes two columns of PEs, for a total of n second split units: second split unit D-1, second split unit D-2, ..., second split unit Dn.
[0112] Based on the above embodiments, further, the first ring-forming rule includes:
[0113] The 2×2n processor cores included in the first split unit are evenly divided into two sub-units, the first sub-unit includes the processor cores in the 1st to nth columns in the first split unit, and the second sub-unit includes the processor cores in the n+1th to 2nth columns in the first split unit; the processor cores included in the first sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop, and the processor cores included in the second sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop.
[0114] Specifically, when the first communication layer division rule performs layer division in the row direction, the first ring formation rule is used to determine the ring formation mode of each processor core in each first segmentation unit.
[0115] For example, Figure 9The following figure illustrates the formation of the first loop, using a first split unit as an example. For the 2n columns of PEs included in the first split unit, the PEs in columns 1 through n constitute the first subunit, and the PEs in columns (n+1) through 2n constitute the second subunit. The PEs in the first subunit are serially connected in a clockwise direction to form loop CR-1, and the PEs in the second subunit are serially connected in a clockwise direction to form loop CR-2.
[0116] For example, Figure 10 As shown, the process of forming the first loop is described using a first splitting unit as an example. For the 2n columns of PEs included in the first splitting unit, the PEs in the 1st to nth columns constitute the first subunit, and the PEs in the n+1st to 2nth columns constitute the second subunit. The PEs included in the first subunit are serially connected in a clockwise direction to form a loop CR-1, and serially connected in a counterclockwise direction to form a loop CR-1'; the PEs included in the second subunit are serially connected in a clockwise direction to form a loop CR-2, and serially connected in a counterclockwise direction to form a loop CR-2'. Since the first subunit corresponds to two loops, the data can be divided into 2n parts. The 1st to nth parts of data are transmitted along loop CR-1, and the n+1st to 2nth parts of data are transmitted along loop CR-1'. Data transmission along loop CR-1 and along loop CR-1' can be carried out in parallel, which can improve data transmission efficiency compared to the case where the first subunit corresponds to one loop for data transmission.
[0117] Based on the above embodiments, further, the second ring-forming rule includes:
[0118] The 2m×2 processor cores included in the second split unit are divided into two sub-units according to odd and even rows. The third sub-unit includes the processor cores in the odd rows of the second split unit, and the fourth sub-unit includes the processor cores in the even rows of the second split unit. The processor cores included in the third sub-unit are serially connected in a first direction to form a loop, and the processor cores included in the fourth sub-unit are serially connected in a second direction to form a loop. The first direction and the second direction are opposite.
[0119] Specifically, when the second communication hierarchy division rule performs hierarchy division in the row direction, the second ring formation rule is used to determine the ring formation of each processor core within each second segmentation unit. When the first direction is clockwise, the second direction is counterclockwise; when the first direction is counterclockwise, the second direction is clockwise.
[0120] For example, Figure 11The following figure illustrates the formation of the second loop, using a second split unit D-1 as an example. For the 2m×2 PEs included in second split unit D-1, the odd-numbered rows of PEs constitute the first subunit, and the even-numbered rows of PEs constitute the second subunit. The PEs in the first subunit are serially connected in a clockwise direction to form loop DR-1, while the PEs in the second subunit are serially connected in a counterclockwise direction to form loop DR-2. The second loop includes loops DR-1 and DR-2.
[0121] On the basis of the above embodiments, further, the multiple processor cores are arranged in a 2m×2n array; the first array direction is the column direction and the second array direction is the row direction; accordingly, the first communication level division rule includes dividing the 2n columns of processor cores into 2 columns, and the processor cores in every 2 columns serve as a first splitting unit; the second communication level division rule includes dividing the 2m rows of processor cores by rows, and the processor cores in the jth row and the j+mth row serve as a second splitting unit, where j is greater than or equal to 1 and less than or equal to m, and m and n are natural numbers.
[0122] Specifically, when the first array direction is the column direction and the second array direction is the row direction, the first communication level division rule includes dividing the 2n columns of processor cores into units of 2 columns, with each 2 columns of processor cores serving as a first division unit, i.e., each 2 columns of processor cores constitute a first division unit, resulting in a total of n first division units. The second communication level division rule includes dividing the 2m rows of processor cores into rows, with the processor cores in the jth row and the j+mth row serving as a second division unit, i.e., each second division unit includes 2 rows of processor cores, resulting in a total of m second division units. Wherein, j is greater than or equal to 1 and less than or equal to m.
[0123] Based on the above embodiments, further, the first ring-forming rule includes:
[0124] The 2m×2 processor cores included in the first split unit are evenly divided into two sub-units, the first sub-unit includes the processor cores in the 1st to mth rows in the first split unit, and the second sub-unit includes the processor cores in the m+1th to 2mth rows in the first split unit; the processor cores included in the first sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop, and the processor cores included in the second sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop.
[0125] Specifically, when the first communication layer division rule performs layer division in the row direction, the first ring formation rule is used to determine the ring formation mode of each processor core in each first segmentation unit.
[0126] For example, Figure 12 The following figure illustrates the formation of the first loop, using a first split unit as an example. For the 2m rows of PEs included in the first split unit, the PEs in rows 1 through m constitute the first subunit, and the PEs in rows m+1 through 2m constitute the second subunit. The PEs in the first subunit are serially connected in a clockwise direction to form loop CR-1, and the PEs in the second subunit are serially connected in a clockwise direction to form loop CR-2.
[0127] For example, Figure 13 As shown, the process of forming the first loop is described using a first splitting unit as an example. For the 2m rows of PEs included in the first splitting unit, the PEs in rows 1 to n constitute the first subunit, and the PEs in rows m+1 to 2m constitute the second subunit. The PEs included in the first subunit are serially connected in a clockwise direction to form loop CR-1, and serially connected in a counterclockwise direction to form loop CR-1'; the PEs included in the second subunit are serially connected in a clockwise direction to form loop CR-2, and serially connected in a counterclockwise direction to form loop CR-2'. Since the first subunit corresponds to two loops, the data can be divided into 2m parts. The 1st to mth parts of data are transmitted along loop CR-1, and the m+1st to 2mth parts of data are transmitted along loop CR-1'. Data transmission along loop CR-1 and along loop CR-1' can be carried out in parallel, which can improve data transmission efficiency compared to the case where the first subunit corresponds to one loop for data transmission.
[0128] Based on the above embodiments, further, the second ring-forming rule includes:
[0129] The 2×2n processor cores included in the second split unit are divided into two sub-units according to odd and even columns. The third sub-unit includes the processor cores in the odd columns of the second split unit, and the fourth sub-unit includes the processor cores in the even columns of the second split unit. The processor cores included in the third sub-unit are connected in series in a first direction to form a loop, and the processor cores included in the fourth sub-unit are connected in series in a second direction to form a loop. The first direction and the second direction are opposite.
[0130] Specifically, when the second communication hierarchy division rule performs hierarchy division in the column direction, the second ring formation rule is used to determine the ring formation of each processor core within each second segmentation unit. When the first direction is clockwise, the second direction is counterclockwise; when the first direction is counterclockwise, the second direction is clockwise.
[0131] For example, Figure 14The following figure illustrates the formation of the second loop, using a second split unit D-1 as an example. For the 2×2n PEs included in second split unit D-1, the odd-numbered columns of PEs constitute the first subunit, and the even-numbered columns of PEs constitute the second subunit. The PEs in the first subunit are serially connected in a clockwise direction to form loop DR-1, while the PEs in the second subunit are serially connected in a counterclockwise direction to form loop DR-2. The second loop includes loops DR-1 and DR-2.
[0132] An embodiment of the present invention provides an electronic device, comprising the processor chip described in any one of the above embodiments.
[0133] The processor chip may be a wafer-level processor chip, the electronic device may be a server, and multiple servers may form a server cluster.
[0134] Figure 15 FIG. 1 is a flow chart of a collective communication method provided by an embodiment of the present invention. Figure 15 As shown, the collective communication method provided by an embodiment of the present invention is applied to the processor chip described in any of the above embodiments, including:
[0135] S1501: Each first slicing unit of the first communication layer communicates in parallel, and each processor core included in each first slicing unit communicates based on the first loop according to a ring-reduce operation;
[0136] Specifically, the processor cores within each first sharding unit communicate using a ring algorithm based on the first loop, using an All-Reduce operation. The All-Reduce operation is divided into two phases: a Reduce-Scatter operation and an All-Gather operation. The ring all-reduce operations within each first sharding unit can be performed in parallel to improve communication efficiency. This ring all-reduce operation is implemented using the ring algorithm.
[0137] S1502: Each second split unit of the second communication level communicates in parallel, and each processor core included in each second split unit communicates according to a ring full reduce operation based on the second loop.
[0138] Specifically, after the first slicing units at the first communication level complete communication, the processor cores within each second slicing unit communicate using a ring algorithm based on the second loop, using an All-Reduce operation. The All-Reduce operation is divided into two phases: a Reduce-Scatter operation and an All-Gather operation. The ring all-reduce operations of each second slicing unit can be performed in parallel to improve communication efficiency.
[0139] In the collective communication method provided by an embodiment of the present invention, each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates based on the first loop according to a ring-wide reduction operation; each second slicing unit of the second communication level communicates in parallel, and each processor core included in each second slicing unit communicates based on the second loop according to a ring-wide reduction operation. Since the array-arranged processor cores are divided into the first slicing unit and the second slicing unit, the number of processor cores in the loop is reduced, and each first slicing unit can communicate in parallel, and each second slicing unit can communicate in parallel, thereby improving the collective communication efficiency of the processor chip.
[0140] For processor chip X, the structure is as follows Figure 2 As shown, it includes m×n array-arranged processor cores and follows Figure 3 As shown in the figure, each first segmentation unit of the first communication level is divided according to Figure 4 The corresponding first ring-forming rule forms a first loop. Each second splitting unit of the second communication layer, in accordance with the second ring-forming rule, sequentially numbers the multiple processor cores included in the second splitting unit in the column direction, with the numbering starting from 1; if the maximum value m of the numbering is an even number, then sequentially connect the odd-numbered processor cores in the column direction, connect the m-1th processor core in series with the mth processor core, sequentially connect the even-numbered processor cores in series in the opposite direction of the column direction, and connect the second processor core in series with the first processor core to form a second loop; if the maximum value m of the numbering is an odd number, then sequentially connect the odd-numbered processor cores in the column direction, connect the m-1th processor core in series in the opposite direction of the column direction, connect the even-numbered processor cores in series, and connect the second processor core in series with the first processor core to form a second loop.
[0141] When processor chip X performs collective communication, first, within each first-slice unit at the first communication level, each processor core communicates based on the first loop according to the ring-all-reduce operation, and all first-slice units complete the all-reduce operation in parallel. Then, within each second-slice unit at the second communication level, each processor core communicates based on the second loop according to the ring-all-reduce operation, and all second-slice units complete the all-reduce operation in parallel.
[0142] Based on the above embodiments, further, the multiple processor cores are arranged in a 2m×2n array, and each first slicing unit includes a first sub-unit and a second sub-unit; accordingly, the processor cores included in each first slicing unit communicate based on the first loop according to the ring full reduce operation, including:
[0143] Each processor core in a first subunit included in each first slicing unit communicates according to a ring full reduce operation based on a clockwise loop to process the 1st to 2xth shares of communication data of the processor core, and communicates according to a ring full reduce operation based on a counterclockwise loop to process the 2x+1st to 4xth shares of data of the processor core; wherein the communication data of each processor core in the first subunit is divided into 4x shares, where x is the total number of processor cores included in the first subunit;
[0144] Each processor core in the second sub-unit included in each first slicing unit communicates according to a ring full reduce operation based on a clockwise loop to process the 1st to 2xth shares of communication data of the processor core, and communicates according to a ring full reduce operation based on a counterclockwise loop to process the 2x+1st to 4xth shares of data of the processor core; wherein the communication data of each processor core in the second sub-unit is divided into 4x shares, where x is the total number of processor cores included in the second sub-unit;
[0145] Each first splitting unit includes a first subunit and a second subunit that perform a parallel ring full reduction operation.
[0146] Specifically, the multiple processor cores are arranged in a 2m×2n array. If the first array direction is the row direction and the second array direction is the column direction, then the first communication level division rule is used: the 2m rows of processor cores are divided into units of 2 rows, and every 2 rows of processor cores are used as a first splitting unit, thereby obtaining m first splitting units. And the second communication level division rule is used: the 2n columns of processor cores are divided by columns, and the processor cores in the i-th column and the i+n-th column are used as a second splitting unit, where i is greater than or equal to 1 and less than or equal to n, thereby obtaining n second splitting units.
[0147] Based on the first ring formation rule, the 2×2n processor cores included in the first split unit are evenly divided into two subunits. The first subunit includes the processor cores in columns 1 to n of the first split unit, and the second subunit includes the processor cores in columns n+1 to 2n of the first split unit. The processor cores in the first subunit are serially connected in a clockwise and counterclockwise direction to form a loop, and the processor cores in the second subunit are serially connected in a clockwise and counterclockwise direction to form a loop. This results in a first subunit and a second subunit for each first split unit, and a clockwise and counterclockwise loop formed by each processor core in the first subunit, as well as a clockwise and counterclockwise loop formed by each processor core in the second subunit. The first subunit and the second subunit of each first split unit can communicate in parallel, and each processor core in the first subunit can communicate simultaneously in a clockwise and counterclockwise loop.
[0148] Since the structures of the m first splitting units are similar, they all include a first subunit and a second subunit, and the first subunit and the second subunit each have a clockwise loop and a counterclockwise loop. The communication process of the first splitting unit is taken as an example to illustrate the communication process of the first splitting unit.
[0149] Each processor core in the first subunit included in the first segmentation unit divides the communication data into 4x shares, where x equals n because the first subunit includes 2n processor cores. The processor cores in the first subunit communicate using a ring-reduce operation in a clockwise loop to process shares 1 through 2x of the communication data for each processor core. Simultaneously, the processor cores in the first subunit communicate using a ring-reduce operation in a counterclockwise loop to process shares 2x+1 through 4x of the communication data for each processor core.
[0150] Each processor core in the second subunit included in the first segmentation unit divides the communication data into 4x shares. The processor cores in the second subunit communicate based on a clockwise loop according to a ring-full-reduce operation to process the 1st to 2xth shares of the communication data of each processor core. Simultaneously, the processor cores in the second subunit communicate based on a counterclockwise loop according to a ring-full-reduce operation to process the 2x+1st to 4xth shares of the communication data of each processor core.
[0151] Since the first subunit and the second subunit are independent of each other, the first subunit and the second subunit can communicate in parallel.
[0152] If the first array direction is the column direction and the second array direction is the row direction, then according to the first communication level division rule: the 2n columns of processor cores are divided into 2 columns, and the processor cores of every 2 columns serve as a first division unit, and n first division units are obtained. And according to the second communication level division rule: the 2m rows of processor cores are divided into rows, and the processor cores of the jth row and the j+mth row serve as a second division unit, where j is greater than or equal to 1 and less than or equal to m, and m second division units are obtained.
[0153] Based on the first ring formation rule: the 2m×2 processor cores included in the first split unit are evenly divided into two subunits. The first subunit includes the processor cores in rows 1 to m of the first split unit, and the second subunit includes the processor cores in rows m+1 to 2m of the first split unit. The processor cores in the first subunit are serially connected in a clockwise and counterclockwise direction to form a loop, and the processor cores in the second subunit are serially connected in a clockwise and counterclockwise direction to form a loop. This results in a first subunit and a second subunit for each first split unit, and a clockwise and counterclockwise loop formed by each processor core in the first subunit, as well as a clockwise and counterclockwise loop formed by each processor core in the second subunit. The first subunit and the second subunit of each first split unit can communicate in parallel, and each processor core in the first subunit can communicate simultaneously in a clockwise and counterclockwise loop.
[0154] Since the structures of the n first splitting units are similar, they all include a first subunit and a second subunit, and the first subunit and the second subunit each have a clockwise loop and a counterclockwise loop. The communication process of the first splitting unit is taken as an example to illustrate the communication process of the first splitting unit.
[0155] Each processor core in the first subunit included in the first segmentation unit divides the communication data into 4x shares. Since the first subunit includes 2m processor cores, x is equal to m. The processor cores in the first subunit communicate based on a clockwise loop according to a ring-full-reduce operation to process the 1st to 2xth shares of the communication data of each processor core. Simultaneously, the processor cores in the first subunit communicate based on a counterclockwise loop according to a ring-full-reduce operation to process the 2x+1st to 4xth shares of the data of each processor core.
[0156] Each processor core in the second subunit included in the first segmentation unit divides the communication data into 4x shares. The processor cores in the second subunit communicate based on a clockwise loop according to a ring-full-reduce operation to process the 1st to 2xth shares of the communication data of each processor core. Simultaneously, the processor cores in the second subunit communicate based on a counterclockwise loop according to a ring-full-reduce operation to process the 2x+1st to 4xth shares of the communication data of each processor core.
[0157] Since the first subunit and the second subunit are independent of each other, the first subunit and the second subunit can communicate in parallel.
[0158] Since the first segmentation unit is further divided into a first sub-unit and a second sub-unit to form a loop, the number of processor cores in the loop is further reduced. The first sub-unit and the second sub-unit can each form two loops to communicate in parallel, and the amount of single-step communication in the loop is reduced, which is conducive to further improving the efficiency of collective communication.
[0159] For example, for processor chip Z, the structure is as follows Figure 2 As shown, it includes 2m×2n array-arranged processor cores and follows Figure 8 As shown in the figure, each first segmentation unit of the first communication level is divided according to Figure 10 The corresponding first ring rule obtains the first subunit and the clockwise loop 1-1 and the counterclockwise loop 1-2 corresponding to the first subunit, and obtains the second subunit and the clockwise loop 2-1 and the counterclockwise loop 2-2 corresponding to the second subunit. Each second split unit of the second communication layer, according to Figure 11 The corresponding second ring-forming rule forms a second loop.
[0160] When processor chip Z performs collective communication, the 2n processor cores in each first sub-unit within each first-slice unit of the first communication layer first divide the data to be communicated into 4n shares. These 2n processor cores communicate using an all-reduce operation along the corresponding clockwise loop 1-1, transmitting shares 1 through 2n. Simultaneously, they communicate using an all-reduce operation along the corresponding counterclockwise loop 1-2, transmitting shares 2n+1 through 4n. The 2n processor cores in each second sub-unit divide the data to be communicated into 4n shares. These 2n processor cores communicate using an all-reduce operation along the corresponding clockwise loop 2-1, transmitting shares 1 through 2n. Simultaneously, they communicate using an all-reduce operation along the corresponding counterclockwise loop 2-2, transmitting shares 2n+1 through 4n. The first and second sub-units within each first sub-unit perform the all-reduce operation in parallel, and each first sub-unit communicates in parallel. Then, in each second split unit at the second communication level, each processor core communicates according to the ring all-reduce operation based on the second loop, and all second split units complete the all-reduce operation in parallel.
[0161] Figure 16 FIG. 1 is a flow chart of a collective communication method provided by an embodiment of the present invention. Figure 16 As shown, the collective communication method provided by an embodiment of the present invention is applied to the processor chip described in any of the above embodiments, including:
[0162] S1601: Each first slicing unit of the first communication layer communicates in parallel, and each processor core included in each first slicing unit communicates according to a reduce-scatter operation based on a first loop.
[0163] Specifically, each processor core included in each first slicing unit communicates using a ring algorithm based on the first loop and in accordance with a Reduce-Scatter operation, and each first slicing unit communicates in parallel to improve communication efficiency.
[0164] S1602: Each second split unit of the second communication level communicates in parallel, and each processor core included in each second split unit communicates according to an all-reduce operation based on the second loop;
[0165] Specifically, after each first slicing unit communicates using the Reduce-Scatter operation, each second slicing unit communicates using the ring algorithm based on the second loop using the All-Reduce operation. The ring all-reduce operations of each second slicing unit can be performed in parallel to improve communication efficiency.
[0166] S1603: Each first slicing unit of the first communication layer communicates in parallel, and each processor core included in each first slicing unit communicates according to an aggregation operation based on the first loop.
[0167] Specifically, after the second slicing units at the second communication level complete communication, the processor cores included in each first slicing unit communicate using a ring algorithm based on the first loop in an all-gather operation. Furthermore, the first slicing units communicate in parallel to improve communication efficiency.
[0168] In the collective communication method provided by an embodiment of the present invention, each first split unit of the first communication level communicates in parallel, and each processor core included in each first split unit communicates according to a reduce-scatter operation based on the first loop; each second split unit of the second communication level communicates in parallel, and each processor core included in each second split unit communicates according to a ring-full reduce operation based on the second loop; each first split unit of the first communication level communicates in parallel, and each processor core included in each first split unit communicates according to an aggregation operation based on the first loop. Since the array-arranged processor cores are divided into the first split unit and the second split unit, the number of processor cores in the loop is reduced, and each first split unit can communicate in parallel, and each second split unit can communicate in parallel, thereby improving the collective communication efficiency of the processor chip.
[0169] In addition, each first split unit in the first communication layer first communicates according to the reduce-scatter operation, and then each second split unit in the second communication layer communicates according to the full-reduce operation. This can reduce the amount of single-step communication processed by the second split unit and reduce the risk of link contention when the second split units in the second communication layer collectively communicate, which is conducive to improving communication efficiency.
[0170] For processor chip Y, the structure is as follows Figure 2 As shown, it includes 2m×2n array-arranged processor cores and follows Figure 8 As shown in the figure, each first segmentation unit of the first communication level is divided according to Figure 9 The corresponding first ring rule forms the first ring. Each second segmentation unit of the second communication layer, according to Figure 13 The corresponding second ring-forming rule forms a second loop.
[0171] When processor chip Y performs collective communication, first, within each first split unit of the first communication layer, each processor core within the first split unit communicates according to the Reduce-Scatter operation based on the first loop, and all first split units complete the Reduce-Scatter operation in parallel; since the 2×2n processor cores included in the first split unit are evenly divided into two sub-units, the first sub-unit includes the processor cores from the 1st column to the nth column in the first split unit, which communicate according to the Reduce-Scatter operation based on serial connection to form a loop; the second sub-unit includes the processor cores from the n+1th column to the 2nth column in the first split unit, which communicate according to the Reduce-Scatter operation based on serial connection to form a loop, and the first sub-unit and the second sub-unit communicate in parallel.
[0172] Then, in each second split unit at the second communication level, each processor core communicates according to the ring all-reduce operation based on the second loop, and all second split units complete the all-reduce operation in parallel.
[0173] Finally, within each first-slice unit at the first communication level, each processor core within the first-slice unit performs an all-gather operation based on the first loop, and each first-slice unit completes the all-gather operation in parallel. Because the 2×2n processor cores included in the first-slice unit are evenly divided into two sub-units, the first sub-unit includes the processor cores in columns 1 to n of the first-slice unit, which communicate based on a serial connection to form a loop and perform an all-gather operation. The second sub-unit includes the processor cores in columns n+1 to 2n of the first-slice unit, which communicate based on a serial connection to form a loop and perform an all-gather operation. The first and second sub-units communicate in parallel.
[0174] Based on the above embodiments, further, the multiple processor cores are arranged in a 2m×2n array, and each first slicing unit includes a first sub-unit and a second sub-unit; accordingly, the processor cores included in each first slicing unit communicate based on the first loop according to the reduce-scatter operation, including:
[0175] Each processor core in a first subunit included in each first splitting unit communicates according to a reduce-scatter operation based on a clockwise loop to process the 1st to 2yth shares of communication data of each processor core, and communicates according to a reduce-scatter operation based on a counterclockwise loop to process the 2y+1st to 4yth shares of data of each processor core; wherein the communication data of each processor core in the first subunit is divided into 4y shares, where y is the total number of processor cores included in the first subunit;
[0176] Each processor core in the second sub-unit included in each first splitting unit communicates according to a reduce-scatter operation based on a clockwise loop to process the 1st to 2yth shares of communication data of each processor core, and communicates according to a reduce-scatter operation based on a counterclockwise loop to process the 2y+1st to 4yth shares of data of the processor core; wherein the communication data of each processor core in the second sub-unit is divided into 4y shares, where y is the total number of processor cores included in the second sub-unit;
[0177] Each first segmentation unit includes a first subunit and a second subunit that perform a reduction and scatter operation in parallel.
[0178] Specifically, the multiple processor cores are arranged in a 2m×2n array. If the first array direction is the row direction and the second array direction is the column direction, then the first communication level division rule is used: the 2m rows of processor cores are divided into units of 2 rows, and every 2 rows of processor cores are used as a first splitting unit, thereby obtaining m first splitting units. And the second communication level division rule is used: the 2n columns of processor cores are divided by columns, and the processor cores in the i-th column and the i+n-th column are used as a second splitting unit, where i is greater than or equal to 1 and less than or equal to n, thereby obtaining n second splitting units.
[0179] Based on the first ring formation rule, the 2×2n processor cores included in the first split unit are evenly divided into two subunits. The first subunit includes the processor cores in columns 1 to n of the first split unit, and the second subunit includes the processor cores in columns n+1 to 2n of the first split unit. The processor cores in the first subunit are serially connected in a clockwise and counterclockwise direction to form a loop, and the processor cores in the second subunit are serially connected in a clockwise and counterclockwise direction to form a loop. This results in a first subunit and a second subunit for each first split unit, and a clockwise and counterclockwise loop formed by each processor core in the first subunit, as well as a clockwise and counterclockwise loop formed by each processor core in the second subunit. The first subunit and the second subunit of each first split unit can communicate in parallel, and each processor core in the first subunit can communicate simultaneously in a clockwise and counterclockwise loop.
[0180] Since the structures of the m first splitting units are similar, they all include a first subunit and a second subunit, and the first subunit and the second subunit each have a clockwise loop and a counterclockwise loop. The communication process of the first splitting unit is taken as an example to illustrate the communication process of the first splitting unit.
[0181] Each processor core in the first subunit included in the first segmentation unit divides the communication data into 4y shares. Since the first subunit includes 2n processor cores, y equals n. The processor cores in the first subunit communicate using a Reduce-Scatter operation based on a clockwise loop to process shares 1 to 2x of the communication data for each processor core. Simultaneously, the processor cores in the first subunit communicate using a Reduce-Scatter operation based on a counterclockwise loop to process shares 2y+1 to 4y of the communication data for each processor core.
[0182] Each processor core in the second subunit included in the first segmentation unit divides the communication data into 4y parts. The processor cores in the second subunit communicate based on a clockwise loop according to a ring-full-reduce operation to process the 1st to 2yth parts of the communication data of each processor core. At the same time, the processor cores in the second subunit communicate based on a counterclockwise loop according to a ring-full-reduce operation to process the 2y+1st to 4yth parts of the communication data of each processor core.
[0183] Since the first subunit and the second subunit are independent of each other, the first subunit and the second subunit can communicate in parallel.
[0184] If the first array direction is the column direction and the second array direction is the row direction, then according to the first communication level division rule: the 2n columns of processor cores are divided into 2 columns, and the processor cores of every 2 columns serve as a first division unit, and n first division units are obtained. And according to the second communication level division rule: the 2m rows of processor cores are divided into rows, and the processor cores of the jth row and the j+mth row serve as a second division unit, where j is greater than or equal to 1 and less than or equal to m, and m second division units are obtained.
[0185] Based on the first ring formation rule: the 2m×2 processor cores included in the first split unit are evenly divided into two subunits. The first subunit includes the processor cores in rows 1 to m of the first split unit, and the second subunit includes the processor cores in rows m+1 to 2m of the first split unit. The processor cores in the first subunit are serially connected in a clockwise and counterclockwise direction to form a loop, and the processor cores in the second subunit are serially connected in a clockwise and counterclockwise direction to form a loop. This results in a first subunit and a second subunit for each first split unit, and a clockwise and counterclockwise loop formed by each processor core in the first subunit, as well as a clockwise and counterclockwise loop formed by each processor core in the second subunit. The first subunit and the second subunit of each first split unit can communicate in parallel, and each processor core in the first subunit can communicate simultaneously in a clockwise and counterclockwise loop.
[0186] Since the structures of the n first splitting units are similar, they all include a first subunit and a second subunit, and the first subunit and the second subunit each have a clockwise loop and a counterclockwise loop. The communication process of the first splitting unit is taken as an example to illustrate the communication process of the first splitting unit.
[0187] Each processor core in the first subunit included in the first segmentation unit divides the communication data into 4x shares. Since the first subunit includes 2m processor cores, y equals m. The processor cores in the first subunit communicate using a ring-reduce operation in a clockwise loop to process shares 1 to 2y of their communication data. Simultaneously, the processor cores in the first subunit communicate using a ring-reduce operation in a counterclockwise loop to process shares 2y+1 to 4y of their communication data.
[0188] Each processor core in the second subunit included in the first segmentation unit divides the communication data into 4y shares. The processor cores in the second subunit communicate using a Reduce-Scatter operation based on a clockwise loop to process shares 1 to 2y of the communication data for each processor core. Simultaneously, the processor cores in the second subunit communicate using a Reduce-Scatter operation based on a counterclockwise loop to process shares 2y+1 to 4y of the communication data for each processor core.
[0189] Since the first subunit and the second subunit are independent of each other, the first subunit and the second subunit can communicate in parallel.
[0190] Since the first segmentation unit is further divided into a first sub-unit and a second sub-unit to form a loop, the number of processor cores in the loop is further reduced. The first sub-unit and the second sub-unit can each form two loops to communicate in parallel, and the amount of single-step communication in the loop is reduced, which is conducive to further improving the efficiency of collective communication.
[0191] Based on the above embodiments, further, each processor core included in each first slicing unit communicates according to the aggregation operation based on the first loop, including:
[0192] The processor cores in the first sub-units included in each first slicing unit communicate according to the aggregation operation based on a clockwise loop, and communicate according to the aggregation operation based on a counterclockwise loop;
[0193] The processor cores in the second sub-unit included in each first slicing unit communicate according to the aggregation operation based on the clockwise loop, and communicate according to the aggregation operation based on the counterclockwise loop;
[0194] The first subunit and the second subunit included in each first splitting unit are aggregated in parallel.
[0195] Specifically, for each first split unit of the first communication level, the communication data of the aggregation operation of each processor core in the first sub-unit of the first split unit is the data obtained after communication at the second communication level, that is, the communication data of all processor cores in the second split unit to which the processor core belongs. The processor cores in the first sub-unit communicate according to the aggregation operation based on a clockwise loop to process the 1st to 2yth portions of the communication data of all processor cores in the second split unit to which each processor core in the first sub-unit belongs. At the same time, the processor cores in the first sub-unit communicate according to the aggregation operation based on a counterclockwise loop to process the 2y+1st to 4yth portions of the communication data of all processor cores in the second split unit to which each processor core in the first sub-unit belongs.
[0196] For each first split unit at the first communication level, the communication data aggregated by each processor core in the second sub-unit of the first split unit is the data obtained after communication at the second communication level, that is, the communication data of all processor cores in the second split unit to which the processor core belongs. The processor cores in the second sub-unit communicate according to the aggregate operation based on a clockwise loop to process the 1st to 2yth portions of the communication data of all processor cores in the second split unit to which each processor core in the first split unit belongs. At the same time, the processor cores in the second sub-unit communicate according to the aggregate operation based on a counterclockwise loop to process the 2y+1st to 4yth portions of the communication data of all processor cores in the second split unit to which each processor core in the second split unit belongs.
[0197] The aggregation operation of the first subunit and the aggregation operation of the second subunit included in each first segmentation unit are performed in parallel to improve the efficiency of aggregate communication.
[0198] For example, Figure 17 As shown, the processor cores of the processor chip are arranged in a 6×4 array. The first communication level includes the first split units C-1 and C-2, and the second communication level includes the second split units D-1, D-2 and D-3. The first split units C-1 and C-2 include the first sub-unit and the second sub-unit respectively, and are arranged in accordance with Figure 10 The second segmentation units D-1, D-2 and D-3 are formed in the counterclockwise direction and the clockwise direction. Figure 11 A second loop is formed.
[0199] Each PE in the first subunit divides the communication data into 12 shares. The six PEs in the first subunit communicate using a reduce-scatter operation in a clockwise loop to process the first through sixth shares of the communication data for the six PEs in the first subunit. They also communicate using a reduce-scatter operation in a counterclockwise loop to process the seventh through twelfth shares of the communication data for the six PEs in the first subunit. Similarly, each PE in the second subunit divides the communication data into 12 shares. The six PEs in the second subunit communicate using a reduce-scatter operation in a clockwise loop to process the first through sixth shares of the communication data for the six PEs in the second subunit. They also communicate using a reduce-scatter operation in a counterclockwise loop to process the seventh through twelfth shares of the communication data for the six PEs in the second subunit. The first split units C-1 and C-2 communicate in parallel. The first and second subunits in the first split unit C-1 communicate in parallel, and the first and second subunits in the first split unit C-2 communicate in parallel.
[0200] After the reduction and scatter operation at the first communication level is completed, each PE obtains 1 / 6 of the communication data (labeled as data G), of which 1 / 12 of the data (labeled as data G-1) is obtained through the counterclockwise loop, and 1 / 12 of the data (labeled as data G-2) is obtained through the clockwise loop.
[0201] Each PE in second split unit D-1 divides the communication data into eight parts. The eight PEs in second split unit D-1 communicate using a ring full-reduce operation based on the second loop to process the data G obtained by the eight PEs in second split unit D-1 during communication at the first communication level. After the full-reduce operation is completed, each PE in second split unit D-1 obtains the data G of each PE in second split unit D-1. The specific communication process at the second communication level for second split units D-2 and D-3 is similar to that for second split unit D-1 and is not further described here. Second split units D1, D-2, and D-3 communicate in parallel.
[0202] The six PEs in the first subunit communicate using an aggregation operation along a clockwise loop to process data G-2 from each PE in the second split unit, obtained by each PE in the first subunit during communication at the second communication level. They also communicate using an aggregation operation along a counterclockwise loop to process data G-1 from each PE in the second split unit, obtained by each PE in the first subunit during communication at the second communication level. Similarly, the six PEs in the second subunit communicate using an aggregation operation along a clockwise loop to process data G-2 from each PE in the second subunit during communication at the second communication level. They also communicate using an aggregation operation along a counterclockwise loop to process data G-1 from each PE in the second subunit during communication at the second communication level. The first split units C-1 and C-2 communicate in parallel. The first and second subunits included in the first split unit C-1 communicate in parallel, and the first and second subunits included in the first split unit C-2 communicate in parallel.
[0203] After the above aggregation operation communication, each PE obtains the communication data of all PEs.
[0204] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0205] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0206] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0207] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0208] Throughout this specification, reference to terms such as "one embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0209] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A processor chip, characterized in that: The processor chip includes multiple processor cores, which are arranged in an array, wherein: The multiple processor cores are hierarchically divided according to a first communication layer division rule to obtain multiple first slicing units of the first communication layer in a first array direction; each first slicing unit includes multiple processor cores, and the processor cores included in each first slicing unit form a corresponding first loop according to a first loop formation rule; The multiple processor cores are hierarchically divided according to a second communication hierarchical division rule to obtain a plurality of second slicing units of the second communication hierarchical level in a second array direction; each second slicing unit includes a plurality of processor cores, and the processor cores included in each second slicing unit form a corresponding second loop according to a second loop formation rule; Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates based on a first loop; each second slicing unit of the second communication level communicates in parallel, and each processor core included in each second slicing unit communicates based on a second loop.
2. The processor chip according to claim 1, wherein: The first array direction is a row direction and the second array direction is a column direction; accordingly, the first communication level division rule includes dividing the multiple processor cores arranged in the array into rows, and the processor cores in each row serve as a first dividing unit; the second communication level division rule includes dividing the multiple processor cores arranged in the array into columns, and the processor cores in each column serve as a second dividing unit.
3. The processor chip according to claim 2, wherein: The first ring-forming rule includes: The multiple processor cores included in the first splitting unit are numbered in sequence along the row direction, with the numbers starting from 1; if the maximum value n of the numbering is an even number, the processor cores with odd numbers are connected in series in the row direction, the n-1th processor core is connected in series with the nth processor core, the processor cores with even numbers are connected in series in the opposite direction of the row direction, and the second processor core is connected in series with the first processor core to form a first loop, and the direction of the first loop is clockwise or counterclockwise; if the maximum value n of the numbering is an odd number, the processor cores with odd numbers are connected in series in the row direction, the nth processor core is connected in series with the n-1th processor core in the opposite direction of the row direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a first loop; wherein the direction of the first loop is clockwise or counterclockwise.
4. The processor chip according to claim 2, wherein: The second ring-forming rule includes: The multiple processor cores included in the second splitting unit are numbered in sequence along the column direction, with the numbers starting from 1; if the maximum value m of the numbering is an even number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series with the mth processor core, the processor cores with even numbers are connected in series in the opposite direction of the column direction, and the second processor core is connected in series with the first processor core to form a second loop; if the maximum value m of the numbering is an odd number, the processor cores with odd numbers are connected in series in the column direction, the mth processor core is connected in series with the m-1th processor core in the opposite direction of the column direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a second loop; wherein the direction of the second loop is clockwise or counterclockwise.
5. The processor chip according to claim 1, wherein: The first array direction is the column direction and the second array direction is the row direction; accordingly, the first communication level division rule includes dividing the multiple processor cores arranged in the array into columns, and the processor cores in each column serve as a first dividing unit; the second communication level division rule includes dividing the multiple processor cores arranged in the array into rows, and the processor cores in each row serve as a second dividing unit.
6. The processor chip according to claim 5, characterized in that: The first ring-forming rule includes: The multiple processor cores included in the first splitting unit are numbered in sequence along the column direction, with the numbers starting from 1; if the maximum value m of the numbering is an even number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series with the mth processor core, the processor cores with even numbers are connected in series in the opposite direction of the column direction, and the second processor core is connected in series with the first processor core to form a first loop; if the maximum value m of the numbering is an odd number, the processor cores with odd numbers are connected in series in the column direction, the m-1th processor core is connected in series in the opposite direction of the column direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a first loop; wherein the direction of the first loop is clockwise or counterclockwise.
7. The processor chip according to claim 6, wherein: The second ring-forming rule includes: The multiple processor cores included in the second splitting unit are numbered in sequence along the row direction, with the numbers starting from 1; if the maximum value n of the numbering is an even number, the processor cores with odd numbers are connected in series in the row direction, the n-1th processor core is connected in series with the nth processor core, the processor cores with even numbers are connected in series in the opposite direction of the row direction, and the second processor core is connected in series with the first processor core to form a second loop, and the direction of the second loop is clockwise or counterclockwise; if the maximum value n of the numbering is an odd number, the processor cores with odd numbers are connected in series in the row direction, the nth processor core is connected in series with the n-1th processor core in the opposite direction of the row direction, the processor cores with even numbers are connected in series, and the second processor core is connected in series with the first processor core to form a second loop; wherein the direction of the second loop is clockwise or counterclockwise.
8. The processor chip according to claim 1, wherein: The multiple processor cores are arranged in a 2m×2n array; the first array direction is the row direction and the second array direction is the column direction; accordingly, the first communication level division rule includes dividing the 2m rows of processor cores into units of 2 rows, and every 2 rows of processor cores serve as a first splitting unit; the second communication level division rule includes dividing the 2n columns of processor cores by columns, and the processor cores in the i-th column and the i+n-th column serve as a second splitting unit, where i is greater than or equal to 1 and less than or equal to n.
9. The processor chip according to claim 8, wherein: The first ring-forming rule includes: The 2×2n processor cores included in the first split unit are evenly divided into two sub-units, the first sub-unit includes the processor cores in the 1st to nth columns in the first split unit, and the second sub-unit includes the processor cores in the n+1th to 2nth columns in the first split unit; the processor cores included in the first sub-unit are connected in series in a clockwise and / or counterclockwise direction to form a loop, and the processor cores included in the second sub-unit are connected in series in a clockwise and / or counterclockwise direction to form a loop.
10. The processor chip according to claim 8, wherein: The second ring-forming rule includes: The 2m×2 processor cores included in the second split unit are divided into two sub-units according to odd and even rows. The third sub-unit includes the processor cores in the odd rows of the second split unit, and the fourth sub-unit includes the processor cores in the even rows of the second split unit. The processor cores included in the third sub-unit are serially connected in a first direction to form a loop, and the processor cores included in the fourth sub-unit are serially connected in a second direction to form a loop. The first direction and the second direction are opposite.
11. The processor chip according to claim 1, wherein: The multiple processor cores are arranged in a 2m×2n array; the first array direction is the column direction and the second array direction is the row direction; accordingly, the first communication level division rule includes dividing the 2n columns of processor cores into 2 columns, and the processor cores in every 2 columns serve as a first splitting unit; the second communication level division rule includes dividing the 2m rows of processor cores into rows, and the processor cores in the jth row and the j+mth row serve as a second splitting unit, where j is greater than or equal to 1 and less than or equal to m.
12. The processor chip according to claim 5, wherein: The first ring-forming rule includes: The 2m×2 processor cores included in the first split unit are evenly divided into two sub-units, the first sub-unit includes the processor cores in the 1st to mth rows in the first split unit, and the second sub-unit includes the processor cores in the m+1th to 2mth rows in the first split unit; the processor cores included in the first sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop, and the processor cores included in the second sub-unit are serially connected in a clockwise and / or counterclockwise direction to form a loop.
13. The processor chip according to claim 5, wherein: The second ring-forming rule includes: The 2×2n processor cores included in the second split unit are divided into two sub-units according to odd and even columns. The third sub-unit includes the processor cores in the odd columns of the second split unit, and the fourth sub-unit includes the processor cores in the even columns of the second split unit. The processor cores included in the third sub-unit are connected in series in a first direction to form a loop, and the processor cores included in the fourth sub-unit are connected in series in a second direction to form a loop. The first direction and the second direction are opposite.
14. A collective communication method, characterized in that: The processor chip according to any one of claims 1 to 13, comprising: Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates according to a ring full reduce operation based on the first loop; Each second slicing unit of the second communication level communicates in parallel, and each processor core included in each second slicing unit communicates according to a ring full reduce operation based on the second loop.
15. The method according to claim 14, characterized in that The multiple processor cores are arranged in a 2m×2n array, and each first split unit includes a first sub-unit and a second sub-unit. Accordingly, the processor cores included in each first split unit communicate based on the first loop according to the ring full reduce operation, including: Each processor core in a first subunit included in each first splitting unit communicates according to a ring full reduce operation based on a clockwise loop to process the 1st to 2xth shares of communication data of each processor core, and communicates according to a ring full reduce operation based on a counterclockwise loop to process the 2x+1st to 4xth shares of data of each processor core; wherein the communication data of each processor core in the first subunit is divided into 4x shares, where x is the total number of processor cores included in the first subunit; Each processor core in the second sub-unit included in each first slicing unit communicates according to a ring full reduce operation based on a clockwise loop to process the 1st to 2xth shares of communication data of each processor core, and communicates according to a ring full reduce operation based on a counterclockwise loop to process the 2x+1st to 4xth shares of data of each processor core; wherein the communication data of each processor core in the second sub-unit is divided into 4x shares, where x is the total number of processor cores included in the second sub-unit; Each first splitting unit includes a first subunit and a second subunit that perform a parallel ring full reduction operation.
16. A collective communication method, characterized in that: The processor chip according to any one of claims 1 to 13, comprising: Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates according to a reduce-scatter operation based on the first loop; Each second slicing unit of the second communication level communicates in parallel, and each processor core included in each second slicing unit communicates according to a ring full reduce operation based on the second loop; Each first slicing unit of the first communication level communicates in parallel, and each processor core included in each first slicing unit communicates according to an aggregation operation based on the first loop.
17. The method according to claim 16, characterized in that The multiple processor cores are arranged in a 2m×2n array, and each first slicing unit includes a first sub-unit and a second sub-unit. Accordingly, the processor cores included in each first slicing unit communicate based on the first loop according to the reduce-scatter operation, including: Each processor core in a first subunit included in each first splitting unit communicates according to a reduce-scatter operation based on a clockwise loop to process the 1st to 2yth shares of communication data of the processor core, and communicates according to a reduce-scatter operation based on a counterclockwise loop to process the 2y+1st to 4yth shares of data of the processor core; wherein the communication data of each processor core in the first subunit is divided into 4y shares, where y is the total number of processor cores included in the first subunit; Each processor core in the second sub-unit included in each first slicing unit communicates according to a reduce-scatter operation based on a clockwise loop to process the 1st to 2yth shares of communication data of the processor core, and communicates according to a reduce-scatter operation based on a counterclockwise loop to process the 2y+1st to 4yth shares of data of the processor core; wherein the communication data of each processor core in the second sub-unit is divided into 4y shares, where y is the total number of processor cores included in the second sub-unit; Each first segmentation unit includes a first subunit and a second subunit that perform a reduction and scatter operation in parallel.
18. The method according to claim 17, characterized in that The processor cores included in each first slicing unit communicate according to the aggregation operation based on the first loop, including: The processor cores in the first sub-units included in each first slicing unit communicate according to the aggregation operation based on a clockwise loop, and communicate according to the aggregation operation based on a counterclockwise loop; The processor cores in the second sub-unit included in each first slicing unit communicate according to the aggregation operation based on the clockwise loop, and communicate according to the aggregation operation based on the counterclockwise loop; The first subunit and the second subunit included in each first splitting unit are aggregated in parallel.
19. An electronic device, characterized in that: A processor chip comprising any one of claims 1 to 13.
Citation Information
Patent Citations
RNN (Recurrent Neural Network) parallel model and implementation method and system thereof on multi-core CPU (Central Processing Unit)
CN114154616A
System and method for accelerated processing of image data, and storage medium
CN117437113A