Artificial intelligence chip, data multicasting method, device and storage medium
By establishing data multicast paths between computing units, cross-unit data sharing and reuse are achieved, solving the problem of mismatch between the growth of computing power and storage bandwidth, and improving computing efficiency and energy efficiency ratio.
Patent Information
- Application Number
- CN202511715568.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-11-21
AI Technical Summary
In existing technologies, the growth in computing power of computing cores is not matched with the local storage bandwidth, leading to performance bottlenecks. This is especially true in high-performance computing chips, where the increase in memory bandwidth is limited and cannot meet the data consumption needs of computing cores.
Establish data multicast paths between computing units, and realize data sharing and reuse across units through horizontal and vertical multicast paths to collaboratively supply the data required by the computing core.
It effectively reduces the bandwidth requirements of each local shared memory, improves computing efficiency and energy efficiency, solves the memory access bandwidth bottleneck problem caused by the growth of computing power of computing cores, and supports the operation of computing cores with higher computing power.
Smart Images

Figure CN121166611B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of chip design, and in particular to an artificial intelligence chip, a data multicasting method and device, equipment and a storage medium. BACKGROUND
[0002] In recent years, artificial intelligence, deep learning and high-performance computing technologies have developed rapidly, and the market demand for artificial intelligence chip computing power has increased sharply. To meet the demand, modern artificial intelligence chips mostly use a highly parallel architecture, integrating a large number of computing units to cooperatively process massive computing tasks to improve overall performance.
[0003] At present, a typical computing unit usually has a computing core for performing core operations and a local shared memory for supplying data to the computing core. To ensure that the computing core is not idle due to waiting for data, a common design guideline is to have the local shared memory provide slightly more data than the computing core consumes within one computing cycle. In this way, under normal workloads, the computing core can be fully supplied with data, thereby maximizing its performance.
[0004] However, the evolution of chip design technology is not balanced. The computing power of the computing core can be easily doubled through various means, while the memory bandwidth is physically limited and costly to improve. When the computing power of the computing core of a new generation of chips increases dramatically and the data consumption doubles, the originally small local shared memory is insufficient in supply, and its bandwidth becomes a bottleneck that limits the performance of the computing unit. SUMMARY
[0005] The present application provides an artificial intelligence chip, a data multicasting method and device, equipment and a storage medium to solve the performance bottleneck problem caused by the mismatch between the growth of the computing power of the computing core and the bandwidth of the local memory in the high-performance computing chip of the prior art.
[0006] The present application provides an artificial intelligence chip, comprising a plurality of computing unit groups, each computing unit group comprising a plurality of computing units and a data multicasting path;
[0007] Each computing unit comprises a computing core and a local shared memory, and the local shared memory is used to store data required for the corresponding computing core to perform a computing task;
[0008] The data multicasting path is used to connect the local shared memory of any computing unit in the corresponding computing unit group with the computing core of the adjacent other computing unit;
[0009] When the local shared memory of the any computing unit supplies data to the computing core of the any computing unit, the corresponding data is multicasted to the computing core of the adjacent other computing unit through the data multicasting path.
[0010] According to the artificial intelligence chip provided by the application, each computing unit group comprises at least four computing units arranged in a two-dimensional grid structure in the corresponding computing unit group.
[0011] The data multicast path comprises a horizontal multicast path and a vertical multicast path.
[0012] The horizontal multicast path is used for data multicast between computing units adjacent in the horizontal direction, and the vertical multicast path is used for data multicast between computing units adjacent in the vertical direction.
[0013] According to the artificial intelligence chip provided by the application, the local shared memory comprises a first storage area and a second storage area; the first storage area is used for storing data multicast through the horizontal multicast path; and the second storage area is used for storing data multicast through the vertical multicast path.
[0014] According to the artificial intelligence chip provided by the application, any computing unit and the other adjacent computing units are orthogonally adjacent in the corresponding two-dimensional grid structure; the orthogonally adjacent means directly adjacent to the any computing unit in the upper, lower, left or right direction.
[0015] The application further provides a data multicast method applied to the artificial intelligence chip as described in any one of the above, and the method comprises:
[0016] When the computing core of any computing unit performs a computing task, first data is read from the local shared memory of the any computing unit, and second data multicast by the local shared memory of the other computing units adjacent to the any computing unit is received through the data multicast path in the computing unit group where the any computing unit is located;
[0017] The computing task is performed based on the first data and the second data.
[0018] According to the data multicast method provided by the application, the data multicast path comprises a horizontal multicast path and a vertical multicast path; and the second data multicast by the local shared memory of the other computing units adjacent to the any computing unit is received through the data multicast path in the computing unit group where the any computing unit is located, comprising:
[0019] Second horizontal data multicast by the local shared memory of the other computing units adjacent to the any computing unit in the horizontal direction is received through the horizontal multicast path;
[0020] Second vertical data multicast by the local shared memory of the other computing units adjacent to the any computing unit in the vertical direction is received through the vertical multicast path.
[0021] determining the second data based on the second horizontal data and the second vertical data.
[0022] According to the data multicasting method provided by the application, the second data multicasted by the local shared memory of the other computing units adjacent to the any computing unit is received through the data multicasting path in the computing unit group where the any computing unit is located, and the method comprises the following steps of:
[0023] receiving a multicast enabling instruction used for controlling the opening or closing of the data multicasting path;
[0024] in the case that the multicast enabling instruction indicates the opening of the data multicasting path, the second data multicasted by the local shared memory of the other computing units adjacent to the any computing unit is received through the data multicasting path.
[0025] The application further provides a data multicasting device applied to the artificial intelligence chip as described in any one of the above, and the device comprises:
[0026] a data acquisition unit used for reading the first data from the local shared memory of the any computing unit when the computing core of the any computing unit executes a computing task, and receiving the second data multicasted by the local shared memory of the other computing units adjacent to the any computing unit through the data multicasting path in the computing unit group where the any computing unit is located;
[0027] a task execution unit used for executing the computing task based on the first data and the second data.
[0028] The application further provides an electronic device comprising a memory, a processor and a computer program stored in the memory and running on the processor, and the processor implements the data multicasting method as described in any one of the above when executing the computer program.
[0029] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the data multicasting method as described in any one of the above.
[0030] The artificial intelligence chip, the data multicasting method, the device, the equipment and the storage medium provided by the application realize the cross-unit sharing and reuse of the data in the local shared memory by establishing the data multicasting path between the computing units, so that the data required by a single computing core can be cooperatively supplied by multiple local shared memories, thereby effectively reducing the bandwidth demand that each local shared memory must bear, successfully resolving the memory bandwidth bottleneck problem caused by the rapid growth of the computing core power without significantly increasing the hardware cost and power consumption, and greatly improving the overall computing efficiency and energy efficiency ratio of the artificial intelligence chip. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0032] Figure 1 This is a schematic diagram of the structure of the artificial intelligence chip provided by the present invention;
[0033] Figure 2 This is an example diagram of the artificial intelligence chip provided by the present invention;
[0034] Figure 3 This is a structural example diagram of the SPC provided by the present invention;
[0035] Figure 4 This is a structural example diagram of the computing unit group provided by the present invention;
[0036] Figure 5 This is an example diagram of the data storage and multicast process provided by the present invention;
[0037] Figure 6 This is a flowchart illustrating the data multicast method provided by the present invention;
[0038] Figure 7 This is a schematic diagram of the data multicast device provided by the present invention;
[0039] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention.
[0040] Figure label:
[0041] 110: Computing unit group; 111: Computing unit; 112: Data multicast path;
[0042] 1111: Computational core; 1112: Local shared memory; 710: Data acquisition unit;
[0043] 720: Task execution unit; 810: Processor; 820: Communication interface;
[0044] 830: Memory; 840: Communication bus. Detailed Implementation
[0045] In order to make the objects, technical solutions and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application with reference to the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0046] With the rapid development of technologies such as artificial intelligence AI (Artificial Intelligence), deep learning, and high-performance computing, the demand for artificial intelligence chip computing power in the market is growing at an unprecedented speed. In order to meet this demand, modern artificial intelligence chips usually adopt a highly parallel architecture, which integrates hundreds or thousands of compute units cu (Compute Unit) inside, which work together to process massive computing tasks.
[0047] In the prior art, a typical compute unit usually contains a compute core responsible for performing core operations, such as a tensor compute core tcore (Tensor Core), and a local shared memory smem (Shared Memory) that provides high-speed data for it. In order to ensure that the compute core will not be idle due to waiting for data, the current general design guideline is that the data provided by the local shared memory in a computing cycle needs to be slightly larger than the data consumed by the compute core. This design with a certain margin ensures that the compute core can be fully supplied with data under normal workloads, thereby maximizing its performance.
[0048] However, the evolution of chip design technology presents an unbalanced trend, that is, the computing power of the compute core can be relatively easily doubled through architecture optimization, process improvement, etc., but the memory bandwidth improvement faces greater physical limitations, often accompanied by high chip area, power consumption and design complexity costs. Therefore, when the computing power of the compute core of a new generation of chips increases by one hundred percent compared to the previous generation of products, its data consumption rate also doubles. At this time, the data supply capability of the originally only slightly margin local shared memory is immediately insufficient to support the greatly increased computing power of the compute core, resulting in the bandwidth of the local shared memory becoming the bottleneck that limits the performance of the entire compute unit.
[0049] To this end, the application provides an artificial intelligence chip, aiming to establish a data multicast mechanism among different computing units, so that the data required by a single computing core can be provided by the local shared memory of itself and adjacent computing units, thereby sharing the bandwidth pressure of a single local shared memory, effectively solving the local shared memory bandwidth bottleneck problem caused by the growth of computing core computing power, and further improving the overall operation performance and energy efficiency ratio of the chip.
[0050] Figure 1 is a structural schematic diagram of the artificial intelligence chip provided by the application, as Figure 1 shown, the chip includes a plurality of computing unit groups 110, each computing unit group including a plurality of computing units 111 and a data multicast path 112;
[0051] Each computing unit includes a computing core 1111 and a local shared memory 1112, and the local shared memory is used to store data required for the corresponding computing core to perform a computing task;
[0052] The data multicast path is used to connect the local shared memory of any computing unit in the corresponding computing unit group with the computing cores of other adjacent computing units.
[0053] Among them, when the local shared memory of the computing unit supplies data to the computing core of the computing unit, the corresponding data is multicast to the computing cores of other adjacent computing units through the data multicast path.
[0054] Specifically, considering the current high-performance computing field of artificial intelligence, graphics processing, etc., the computing power of the computing core iterates very fast, often exceeding the bandwidth growth rate of the local memory that supplies data to it, thereby forming a performance bottleneck. In view of this, in the embodiment of the application, an artificial intelligence chip based on a data multicast mechanism is proposed to solve the problem of insufficient local storage bandwidth caused by the growth of computing core computing power in the current high-performance computing chip.
[0055] In practical applications, the artificial intelligence chip can be a standalone accelerator chip, or a processing module integrated in a more complex system-level chip. Figure 2 is an example diagram of the artificial intelligence chip provided by the application, as Figure 2As shown, the top layer architecture of the chip can include a plurality of streaming processor clusters SPC, a level 2 cache L2, a high bandwidth memory HBM, a high bandwidth memory controller HBM ctrl, a central controller cp, and various communication interfaces, such as a PCIe (Peripheral Component Interconnect Express) interface for communication with external systems, a UCIe (Universal Chiplet Interconnect Express) interface for connection with an input / output die IO-die, and the like. On the data flow path inside the chip, data is transferred from the HBM to the L2 on the chip via the HBM ctrl, and then distributed to each SPC through a memory node MNode and a network on chip for efficient execution of final computing tasks.
[0056] Each SPC includes a plurality of computing unit groups cu-group, and each computing unit group is a basic logical unit for computing task division and resource management inside the chip. Each computing unit group includes a plurality of computing units cu and data multicast paths. Figure 3 is a structural example diagram of the SPC provided by the present application, as Figure 3 As shown, one SPC can be divided into a plurality of cu-group, each cu-group includes a plurality of cu, and each cu is connected with a cluster bus interface cbi line to ensure the data path and control path between it and the L2 / HBM.
[0057] Specifically, inside the chip, computing tasks are usually performed in computing unit groups. The computing unit group is a basic logical unit for parallel computing. The computing unit group includes a plurality of computing units, and different computing units work together. The computing unit is the smallest hardware unit for performing specific computing tasks. In the embodiment of the present application, each computing unit includes a computing core tcore and a local shared memory smem. The computing core is a hardware engine responsible for performing core computing tasks, such as performing large-scale matrix multiplication, convolution, and other artificial intelligence related operations. The local shared memory is a high-speed on-chip memory tightly coupled with the computing core, which is used to temporarily store the data required by the corresponding computing core for performing computing tasks, and provides low-latency and high-bandwidth data access for the computing core to ensure that the computing core can work continuously and efficiently.
[0058] And notably, unlike current high-performance computing chips, in the embodiments of the present application, the interconnection mode within the computing unit group in the artificial intelligence chip is redesigned, and a dedicated hardware connection line, i.e., a data multicast path, is introduced to break the data barrier between computing units through the data multicast path, thereby realizing efficient data sharing.
[0059] Specifically, Figure 4 is a structural example diagram of the computing unit group provided by the present application, as Figure 4 shown, the data multicast path physically or logically connects the output end of the local shared memory (smem0) of any computing unit (such as cu0) in the computing unit group to the input end of the computing core (such as tcore1 and tcore2) of the other computing unit (such as cu1 and cu2) adjacent to it. It should be noted that the adjacent here can be the physical location of the two computing units adjacent to each other on the chip layout, or the logical adjacent in the communication topology of the computing unit group, and the embodiments of the present application do not make specific limitation on this.
[0060] And based on this data multicast path, a data multicast mechanism can be established between different computing units in each computing unit group. Specifically, when the local shared memory (smem0) of any computing unit (such as cu0) in the computing unit group provides data to its own computing core (tcore0), in addition to being read by its own computing core (tcore0) along the inherent local data path within the computing unit (such as cu0), at the same time, a complete copy of the same data will be captured by the data multicast path and multicast to the computing cores (such as tcore1 and tcore2) of the adjacent other computing units (such as cu1 and cu2).
[0061] The following will take the computing unit group shown in Figure 4 as an example to illustrate the data multicast process:
[0062] If tcore1 needs to perform a certain computing task, it needs data A and data B. In a traditional design, tcore1 has to read data A and data B completely from its own local shared memory smem1, which puts a high demand on the bandwidth of smem1. In the embodiment of the present application, however, tcore1 can read only a part of data A and a part of data B from smem1, and the other part of data can be obtained directly from the local shared memories smem0 and smem3 of the adjacent cu0 and cu3 through the data multicast path. At the same time, smem0 and smem3 provide the multicast data to tcore0 and tcore3. In this way, the data required by tcore1 to complete the task is supplied by the local shared memories of itself and the adjacent cu, and the bandwidth pressure of smem1 is greatly reduced.
[0063] The artificial intelligence chip provided by the present application establishes a data multicast path between the computing units, realizes cross-unit sharing and reuse of data in the local shared memory, so that the data required by a single computing core can be supplied by multiple local shared memories, thereby effectively reducing the bandwidth demand that each local shared memory must bear. Without significantly increasing the hardware cost and power consumption, the present application successfully resolves the memory bandwidth bottleneck problem caused by the rapid growth of computing core computing power, greatly improving the overall computing efficiency and energy efficiency ratio of the artificial intelligence chip.
[0064] Based on the above embodiment, each computing unit group includes at least four computing units, and the at least four computing units are arranged in a two-dimensional grid structure in the corresponding computing unit group.
[0065] The data multicast path includes a horizontal multicast path and a vertical multicast path. The horizontal multicast path is used for data multicast between adjacent computing units in the horizontal direction, and the vertical multicast path is used for data multicast between adjacent computing units in the vertical direction.
[0066] Specifically, as a preferred embodiment, the present application provides that each computing unit group includes at least four computing units. Referring to Figure 3 It can be seen that in each cu-group, the four or more computing units are not randomly organized, but arranged in a two-dimensional grid structure.
[0067] For example, in the embodiment of the present application, the computing units in each cu-group are arranged in a two-dimensional grid structure. Figure 4The cu-group shown is an example, and the four computing units cu0, cu1, cu2 and cu3 inside the cu-group form a 2x2 grid array in the cu-group, in which cu0 and cu1 are in the same row, and cu2 and cu3 are in the same row; at the same time, cu0 and cu2 are in the same column, and cu1 and cu3 are in the same column. The grid structure can clearly determine the adjacent relationship between the computing units. That is, each computing unit has horizontal and vertical neighbors. For example, cu0 is horizontally adjacent to cu1 and vertically adjacent to cu2.
[0068] Correspondingly, in the embodiment of the present application, the data multicast path also includes a horizontal data path and a vertical data path, that is, a horizontal multicast path and a vertical multicast path. Here, the horizontal multicast path is dedicated to data multicast between horizontally adjacent computing units. For example, Figure 4 In the cu-group shown, the data multicast path includes multiple horizontal multicast paths from smem0 to tcore1, from smem1 to tcore0, from smem2 to tcore3, and from smem3 to tcore2.
[0069] Similarly, the vertical multicast path is dedicated to data multicast between vertically adjacent computing units. For example, Figure 4 In the cu-group shown, the data multicast path includes multiple vertical multicast paths from smem0 to tcore2, from smem2 to tcore0, from smem1 to tcore3, and from smem3 to tcore1.
[0070] Based on the above embodiment, the local shared memory includes a first storage area and a second storage area; the first storage area is used to store data multicast through the horizontal multicast path; and the second storage area is used to store data multicast through the vertical multicast path.
[0071] Specifically, to enable the two directional multicast paths to operate efficiently, the storage strategy of the data can be configured in the embodiment of the present application to ensure that the data can be accurately matched with its predetermined multicast direction, thereby realizing efficient data flow.
[0072] In detail, in the embodiment of the present application, the local shared memory of each computing unit is logically divided into two areas, that is, a first storage area and a second storage area. This division can be managed by a hardware memory controller or can be agreed upon by an upper software (such as a compiler or a driver) when allocating memory, and the embodiment of the present application does not make specific limitations.
[0073] Specifically, the use of the two storage areas is closely bound to the direction of data multicasting. That is, the first storage area is used to store data that is planned to be shared through horizontal multicasting channels, and when the computing core reads data from this area, the multicasting hardware logic automatically directs the data to the horizontal multicasting channel and sends it to the computing unit adjacent in the horizontal direction. Correspondingly, the second storage area is used to store data that is planned to be shared through vertical multicasting channels, and when the computing core reads data from this area, the data is directed to the vertical multicasting channel and sent to the computing unit adjacent in the vertical direction.
[0074] Through the cooperative design of the above structure, channel and storage area, an extremely efficient data flow can be realized. Taking a typical matrix multiplication (for example, C = AB) as an example, Figure 5 is an example diagram of the data storage and multicasting process provided by the present application, as Figure 5 shown:
[0075] Before the task starts, the compiler or runtime scheduler will perform data scheduling, as shown in (a) of Figure 5 , it will load the data block of matrix A required for performing the computing task into the first storage area of the local shared memory of each computing unit. At the same time, the data block of matrix B is loaded into the second storage area of the local shared memory.
[0076] In the task execution phase, as shown in (b) of Figure 5 , when the computing core (tcore0) of cu0 needs data of matrix A, it will initiate a read request to the first storage area of the local shared memory. This data is sent to the computing core of cu0 itself, and at the same time, it is captured by the horizontal multicasting channel and multicasted to the computing core (tcore1) of cu1.
[0077] Correspondingly, when the computing core of cu0 needs data of matrix B, it will initiate a read request to the second storage area of the local shared memory, and this data is sent to the computing core of cu0 itself, and at the same time, it is multicasted to the computing core (tcore2) of cu2 by the vertical multicasting channel.
[0078] In the embodiment of the present application, the computing units are organized into a two-dimensional grid structure and are equipped with corresponding horizontal and vertical multicasting channels, and in combination with the region division strategy of the local shared memory, a highly structured and directional data sharing network can be constructed. In this way, not only can the bandwidth pressure of the local shared memory be reduced, but also hardware support can be provided for algorithms with two-dimensional data dependency characteristics, such as matrix multiplication, so that data of different dimensions can flow and be reused along the optimal path on the physical structure of the chip, thereby maximizing the efficiency of data sharing and greatly enhancing the ability of artificial intelligence chips to handle complex parallel computing tasks.
[0079] Based on the above embodiment, the computing unit and other computing units adjacent thereto are orthogonally adjacent in the corresponding two-dimensional grid structure; orthogonally adjacent means directly adjacent above, below, left or right of the computing unit.
[0080] Specifically, in the embodiment of the present application, the relationship between adjacent computing units in each computing unit group is orthogonally adjacent. That is, the effective adjacent relationship in each computing unit group is limited to the orthogonally adjacent computing units. Here, orthogonally adjacent means that in the two-dimensional grid structure, a computing unit only considers the computing units directly adjacent above, below, left or right thereof as adjacent units.
[0081] For example, as shown in the 2x2 grid structure: Figure 4 For the computing unit cu0, the orthogonally adjacent computing units thereof only include cu1 located right thereof and cu2 located below thereof.
[0082] Similarly, for the computing unit cu1, the orthogonally adjacent computing units thereof are cu0 located left thereof and cu3 located below thereof.
[0083] The computing units located in the diagonal direction, such as cu0 and cu3, are not orthogonally adjacent computing units, and there is no direct data multicast path therebetween.
[0084] Such a setting can make the physical wiring of the horizontal multicast path exactly used for connecting the orthogonally adjacent computing units in the horizontal direction (such as cu0 and cu1), and the physical wiring of the vertical multicast path used for connecting the orthogonally adjacent computing units in the vertical direction (such as cu0 and cu2). Thus, the data of the local shared memory of a computing unit can be multicast to the computing cores of the left, right, above and below adjacent computing units, but there is no hardware path for directly multicasting to the diagonal neighbor computing cores.
[0085] In the embodiment of the present application, the orthogonally adjacent adjacent relationship setting can greatly simplify the on-chip network wiring complexity and data arbitration logic inside the chip, avoid the problems of additional wiring resource consumption, timing convergence difficulty and power consumption increase caused by implementing long-distance and irregular diagonal connections. At the same time, a highly deterministic data sharing model is provided for the upper software, so that the compiler can more easily perform data arrangement and computing scheduling optimization, and provide optimal hardware support for executing structured data with row and column dependency characteristics, such as matrix operation, and further improve the computing efficiency.
[0086] The present application also provides a data multicast method,
[0087] is a flowchart of the data multicast method provided by the present application, as shown in Figure 6 Figure 6 As shown, the method runs on the artificial intelligence chip described in any of the preceding embodiments, aiming to provide a data collaborative supply strategy to efficiently perform a computing task. The method comprises:
[0088] Step 610, when the computing core of any computing unit performs a computing task, reading first data from the local shared memory of the computing unit, and receiving second data multicast from the local shared memory of other computing units adjacent to the computing unit through the data multicast path in the computing unit group to which the computing unit belongs;
[0089] Step 620, performing the computing task based on the first data and the second data.
[0090] Specifically, when the computing core (tcore0) of any computing unit (such as cu0) in a computing unit group starts to perform a computing task, the data flow required by the computing core is divided into two parts, i.e. first data and second data. The first data can be directly read by tcore0 through the internal data path between tcore0 and the local shared memory (smem0) of tcore0, and the second data is obtained from the local shared memory of other computing units adjacent to tcore0.
[0091] In detail, during the execution of the computing task by tcore0, the computing units adjacent to tcore0 (such as cu2) are also executing associated tasks in the same window period when the first data is read. At this time, the local shared memory (smem2) of cu2 is providing the second data required by the computing core (tcore2) of cu2. With the help of the data multicast path, the second data output by smem2 is broadcast in real time. At this time, the hardware logic of tcore0 directly receives the second data multicast through the data multicast path. This process is an efficient passive receiving process, and tcore0 does not need to initiate a new and independent read request to smem2, thereby avoiding additional arbitration and delay.
[0092] Finally, after the computing core of tcore0 collects the first data read from the local shared memory and the second data received from the adjacent computing units, the computing core has all the data required to perform the computing task, and can immediately perform the computing task based on the two data.
[0093] Taking a typical matrix multiplication (for example, C=AB) as an example:
[0094] If tcore0 is responsible for computing a sub-block of the result matrix C, the data it needs for the computation task includes a row data of matrix A (which can be regarded as first data) and a column data of matrix B (which can be regarded as second data). Before the computation task is performed, the row data of matrix A is pre-loaded into smem0. tcore0 reads the row data from its local shared memory. At the same time, the column data of matrix B is pre-loaded into smem2 of cu2 vertically adjacent to cu0 and is being used by tcore2. When the column data is output by smem2, it is also broadcast through the vertical multicast path, so that tcore0 can receive the column data. After obtaining the row data and the column data, tcore0 can immediately perform the matrix multiplication operation.
[0095] The data multicast method provided by the application disperses the data source required by a computation task to a plurality of local shared memories working cooperatively, greatly reduces the instantaneous bandwidth requirement of a single local shared memory, effectively solves the bandwidth bottleneck problem of a high-performance computing chip, supports a computing core with higher computing power without forcibly increasing the bit width and cost of a single local shared memory, and successfully decouples the strong binding relationship between the computing power growth of a computing unit and the bandwidth growth of a single storage unit.
[0096] Based on the above embodiment, the data multicast path includes a horizontal multicast path and a vertical multicast path;
[0097] In step 610, the second data multicast by the local shared memory of the other computing units adjacent to the computing unit is received through the data multicast path in the computing unit group where the computing unit is located, including:
[0098] The second horizontal data multicast by the local shared memory of the other computing units adjacent to the computing unit in the horizontal direction is received through the horizontal multicast path.
[0099] The second vertical data multicast by the local shared memory of the other computing units adjacent to the computing unit in the vertical direction is received through the vertical multicast path.
[0100] The second data is determined based on the second horizontal data and the second vertical data.
[0101] Specifically, in the embodiment of the application, the data multicast path in each computing unit group includes a data path in the horizontal direction and a data path in the vertical direction, i.e., a horizontal multicast path and a vertical multicast path. Here, the horizontal multicast path is used for data multicast between the computing units adjacent in the horizontal direction. For example, Figure 4The data multicast paths in the cu-group shown include multiple horizontal multicast paths from smemO to tcorel, from smel to tcoreO, from smem2 to tcore3, and from smem3 to tcore2.
[0102] Similarly, the vertical multicast paths are dedicated for data multicast between the computing units that are adjacent in the vertical direction. For example, Figure 4 The data multicast paths in the cu-group shown include multiple horizontal multicast paths from smemO to tcorel, from smel to tcoreO, from smem2 to tcore3, and from smem3 to tcore2.
[0103] When a computing core (e.g., tcoreO) performs a computing task, the data it needs is not from a single source, but is supplied by multiple parties in collaboration. That is, when the computing starts, tcoreO first reads the first data from smemO, at the same time, it also receives the second horizontal data, which is multicast by the local shared memory (smel) of the computing unit (cul) adjacent to it in the horizontal direction, through the horizontal multicast path inside the chip. Similarly, it also receives the second vertical data, which is multicast by the local shared memory (smem2) of the computing unit (cu2) adjacent to it in the vertical direction, through the vertical multicast path. Finally, tcoreO combines the second horizontal data received from cul and the second vertical data received from cu2, to determine the complete second data it needs to complete the computing task.
[0104] For example, referring to Figure 5 It can be seen that in a typical matrix multiplication scenario, the complete data tcoreO needs to perform the operation is the data of a part of matrix A (A_0_0) and a part of matrix B (B_0_0) (A_0_0 and B_0_0 together constitute the first data) read by tcoreO from smemO, the data of another part of matrix A (second horizontal data A_0_1) provided by smel received through the horizontal multicast path, and the data of another part of matrix B (second vertical data B_0_1) provided by smem2 received through the vertical multicast path.
[0105] In the embodiment of the present application, the data supply is further decomposed to the horizontally and vertically adjacent computing units, so that the pressure of data supply can be evenly distributed to the plurality of local shared memory units, thereby reducing the bandwidth demand for a single local shared memory to a lower level. More importantly, the horizontal and vertical multicast setting can enable the data of different dimensions, such as the rows and columns of a matrix, to be propagated and shared along the most efficient physical path, realizing the deep collaboration between software and hardware, and greatly optimizing the data flow efficiency.
[0106] Based on the above embodiment, in step 610, the second data of the local shared memory multicast of the other computing units adjacent to the computing unit is received through the data multicast path in the computing unit group where the computing unit is located, including:
[0107] The multicast enabling instruction is received, and the multicast enabling instruction is used to control the opening or closing of the data multicast path.
[0108] In the case where the multicast enabling instruction indicates to open the data multicast path, the second data of the local shared memory multicast of the other computing units adjacent to the computing unit is received through the data multicast path.
[0109] Specifically, in actual application, not all computing tasks can benefit from the data sharing of adjacent computing units. For some algorithms with weak data dependency or irregular computing mode, forcibly opening the data multicast function may not bring any gain, and even cause unnecessary data interference or power consumption.
[0110] Therefore, in the embodiment of the present application, when the second data of the local shared memory of the adjacent computing unit is received through the data multicast path, a judgment is first made. That is, the hardware logic of the computing unit receives the multicast enabling instruction. The multicast enabling instruction is a control command issued by the upper software, such as a compiler, a driver or an application itself, which is used to control the opening or closing of the data multicast path. Then, the hardware logic makes a conditional judgment according to the state of the instruction. That is, in the case where the multicast enabling instruction indicates to open the data multicast path, the data sharing function of the chip is activated. At this time, when the computing core of the computing unit needs to obtain data from the adjacent computing unit, the second data of the local shared memory multicast of the other adjacent computing units is received through the data multicast path.
[0111] On the contrary, if the multicast enabling instruction indicates to close the data multicast path, the data multicast path will be disconnected, and the data sharing between the computing units will be stopped. In this mode, the computing core of the computing unit will return to the traditional working mode, and all the data (including the first data and the second data) required for the computing task will be read from its own local shared memory.
[0112] For example, when the artificial intelligence chip needs to perform a large-scale matrix multiplication task, the compiler analyzes that the task has a high degree of data reuse when generating executable code, and therefore inserts a multicast enable instruction in an open state. After the hardware logic receives the instruction, it opens the data multicast channel and starts the data multicast function to efficiently perform the calculation. When the chip needs to perform a vector addition task in which each thread processes independent data, the compiler inserts a multicast enable instruction in a closed state, and the hardware logic disconnects the data multicast channel to allow each computing unit to work independently, avoiding unnecessary data transmission.
[0113] In the embodiments of the present application, the data multicast function is programmable and flexible, and software controllable, so that the developer can decide whether to enable this function according to the characteristics of the specific algorithm, thereby greatly widening the algorithm application range and application scenarios of the artificial intelligence chip. In this way, not only the performance and energy efficiency advantages brought by data multicast are retained, but also good compatibility with traditional algorithms that are not suitable for this feature is ensured, and unnecessary power consumption can be saved when the function is closed, making the overall design of the chip more efficient, universal and powerful.
[0114] The data multicast device provided by the present application is described below. The data multicast device described below can be referred to in conjunction with the data multicast method described above.
[0115] Figure 7 is a structural schematic diagram of the data multicast device provided by the present application, as Figure 7 shown, the device is applied to the artificial intelligence chip described in any of the above, and the device comprises:
[0116] The data acquisition unit 710 is configured to read first data from the local shared memory of any computing unit when the computing core of the computing unit performs a computing task, and receive second data multicast from the local shared memory of other computing units adjacent to the computing unit through the data multicast channel in the computing unit group to which the computing unit belongs.
[0117] The task execution unit 720 is configured to perform the computing task based on the first data and the second data.
[0118] The data multicast device provided by the present application disperses the data source required for a computing task to multiple local shared memories working in cooperation, greatly reduces the instantaneous bandwidth requirement of a single local shared memory, thereby effectively resolving the bandwidth bottleneck problem of the current high-performance computing chip, supporting computing cores with higher computing power without forcibly increasing the bit width and cost of a single local shared memory, and successfully decoupling the strong binding relationship between the computing unit computing power growth and the single storage unit bandwidth growth.
[0119] Based on the above embodiment, the data multicast path includes a horizontal multicast path and a vertical multicast path;
[0120] The data acquisition unit 710 is configured to:
[0121] Through the horizontal multicast path, receive second horizontal data which is multicast by local shared memory of other computing units adjacent to the computing unit in horizontal direction;
[0122] Through the vertical multicast path, receive second vertical data which is multicast by local shared memory of other computing units adjacent to the computing unit in vertical direction;
[0123] Based on the second horizontal data and the second vertical data, determine the second data.
[0124] Based on the above embodiment, the data acquisition unit 710 is configured to:
[0125] Receive a multicast enabling instruction, the multicast enabling instruction being used to control opening or closing of the data multicast path;
[0126] In a case where the multicast enabling instruction indicates opening of the data multicast path, receive second data which is multicast by local shared memory of other computing units adjacent to the computing unit through the data multicast path.
[0127] Figure 8 An example of a schematic diagram of a physical structure of an electronic device is shown in FIG. 1. Figure 8As shown, the electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 complete mutual communication through the communications bus 840. The processor 810 can invoke a logic instruction in the memory 830 to execute a data multicasting method applied to an artificial intelligence chip, the artificial intelligence chip including a plurality of calculation unit groups, each calculation unit group including a plurality of calculation units and a data multicasting path; each calculation unit including a calculation core and a local shared memory for storing data required for the corresponding calculation core to perform a calculation task; the data multicasting path being used to connect the local shared memory of any calculation unit in the corresponding calculation unit group with the calculation cores of other adjacent calculation units; wherein the local shared memory of the any calculation unit multicasts corresponding data to the calculation cores of the adjacent other calculation units through the data multicasting path when supplying the calculation core of the any calculation unit. The data multicasting method includes: reading first data from the local shared memory of the any calculation unit when the calculation core of the any calculation unit performs a calculation task, and receiving second data multicasted by the local shared memories of other calculation units adjacent to the any calculation unit through the data multicasting path in the calculation unit group where the any calculation unit is located; and performing the calculation task based on the first data and the second data.
[0128] In addition, the logic instruction in the memory 830 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0129] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions that, when executed by a computer, enable the computer to perform the data multicasting method provided by any of the above methods, which is applied to an artificial intelligence chip comprising a plurality of computing unit groups, each computing unit group comprising a plurality of computing units and a data multicasting path; each computing unit comprises a computing core and a local shared memory, the local shared memory being configured to store data required by the corresponding computing core to perform a computing task; the data multicasting path is configured to connect the local shared memory of any computing unit in the corresponding computing unit group with the computing core of an adjacent other computing unit; wherein the local shared memory of the any computing unit multicasts corresponding data to the computing core of the any computing unit via the data multicasting path when the local shared memory supplies the computing core with the data. The data multicasting method comprises: reading first data from the local shared memory of the any computing unit when the computing core of the any computing unit performs the computing task, and receiving second data multicasted by the local shared memory of the other computing unit adjacent to the any computing unit via the data multicasting path in the computing unit group where the any computing unit is located; and performing the computing task based on the first data and the second data.
[0130] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data multicasting method provided by any of the above methods, which is applied to an artificial intelligence chip comprising a plurality of computing unit groups, each computing unit group comprising a plurality of computing units and a data multicasting path; each computing unit comprises a computing core and a local shared memory, the local shared memory being configured to store data required by the corresponding computing core to perform a computing task; the data multicasting path is configured to connect the local shared memory of any computing unit in the corresponding computing unit group with the computing core of an adjacent other computing unit; wherein the local shared memory of the any computing unit multicasts corresponding data to the computing core of the any computing unit via the data multicasting path when the local shared memory supplies the computing core with the data. The data multicasting method comprises: reading first data from the local shared memory of the any computing unit when the computing core of the any computing unit performs the computing task, and receiving second data multicasted by the local shared memory of the other computing unit adjacent to the any computing unit via the data multicasting path in the computing unit group where the any computing unit is located; and performing the computing task based on the first data and the second data.
[0131] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0132] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0133] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An artificial intelligence chip, characterized in that, It includes multiple computing unit groups, and each computing unit group includes multiple computing units and data multicast paths; Each computing unit includes a computing core and local shared memory, wherein the local shared memory is used to store the data required by the corresponding computing core to perform computing tasks; The data multicast path is used to connect the local shared memory of any computing unit in the corresponding computing unit group to the computing core of other adjacent computing units; When the local shared memory of any computing unit supplies data to the computing core of any computing unit, it multicasts the corresponding data to the computing cores of the adjacent computing units through the data multicast path, so that the data required by a single computing core to perform a computing task is provided collaboratively by the local shared memory of the corresponding computing unit and the local shared memory of the adjacent computing units.
2. The artificial intelligence chip according to claim 1, characterized in that, Each computing unit group includes at least four computing units, and the at least four computing units are arranged in a two-dimensional grid structure in the corresponding computing unit group; The data multicast path includes a horizontal multicast path and a vertical multicast path; The horizontal multicast path is used to perform data multicast between adjacent computing units in the horizontal direction; the vertical multicast path is used to perform data multicast between adjacent computing units in the vertical direction.
3. The artificial intelligence chip according to claim 2, characterized in that, The local shared memory includes a first storage area and a second storage area; the first storage area is used to store data multicast through the horizontal multicast path; the second storage area is used to store data multicast through the vertical multicast path.
4. The artificial intelligence chip according to claim 2, characterized in that, The computing unit is orthogonally adjacent to the other adjacent computing units in the corresponding two-dimensional grid structure; the orthogonal adjacent means that it is directly adjacent to any computing unit above, below, to the left or to the right.
5. A data multicast method, characterized in that, The method, applied to an artificial intelligence chip as described in any one of claims 1 to 4, comprises: When a computing core of any computing unit executes a computing task, it reads first data from the local shared memory of the computing unit and receives second data multicast from the local shared memory of other computing units adjacent to the computing unit through the data multicast path in the computing unit group to which the computing unit is located. The computation task is performed based on the first data and the second data.
6. The data multicast method according to claim 5, characterized in that, The data multicast path includes a horizontal multicast path and a vertical multicast path; receiving second data multicast from other computing units adjacent to the computing unit via the data multicast path in the computing unit group where the computing unit is located includes: Through the horizontal multicast path, second horizontal data multicast from local shared memory of other computing units adjacent to any of the computing units in the horizontal direction is received; Through the vertical multicast path, second vertical data multicast from local shared memory of other computing units adjacent to any of the computing units in the vertical direction is received; The second data is determined based on the second horizontal data and the second vertical data.
7. The data multicast method according to claim 5, characterized in that, The step of receiving second data multicast from local shared memory of other computing units adjacent to any computing unit through the data multicast path in the computing unit group to which any computing unit is located includes: Receive a multicast enable command, which is used to control the opening or closing of the data multicast path; When the multicast enable instruction indicates that the data multicast path is enabled, second data multicast from local shared memory of other computing units adjacent to any of the computing units is received through the data multicast path.
8. A data multicast device, characterized in that, The device, applied to an artificial intelligence chip as described in any one of claims 1 to 4, comprises: The data acquisition unit is used to read first data from the local shared memory of any computing unit when the computing core of any computing unit is executing a computing task, and to receive second data multicast from the local shared memory of other computing units adjacent to the computing unit through the data multicast path in the computing unit group to which the computing unit is located. A task execution unit is used to execute the computation task based on the first data and the second data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the data multicast method as described in any one of claims 5 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data multicast method as described in any one of claims 5 to 7.
Citation Information
Patent Citations
Data multicast in compute core clusters
US20240220254A1
Many-core processor architecture and many-core operating system
WO2016159765A1