Computing device, data processing method, and related products
Patent Information
- Application Number
- CN202110480510.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2041-04-30
AI Technical Summary
因此,传统的适合于密集型数据的卷积神经网络在应用于这种稀疏型数据时,效率将变得非常低,尤其是涉及卷积运算时,会在零值数据点上浪费大量的算力等资源
[0010]通过如上所提供的计算装置、使用计算装置来处理数据的方法、芯片和板卡,本披露实施例针对可用于稀疏型数据的卷积运算处理中的数据融合处理步骤,提供了一种多核处理器架构上的实施方案,其可以有效拆分任务,从而缩短处理时间,提高整体效率。进一步地,在一些实施例中,通过在多核处理器上实现多轮流水处理,可以快速完成融合处理任务。在多轮流水处理中,可以采用桶排序方式来分配各轮流水处理的数据部分,进而实现各轮流水处理的输出数据的有序拼接。在一些实施例中,还可以在不同的输出通道维度上复用索引,从而减小数据吞吐量,进一步加速处理。
Smart Images

Figure CN115221103B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of data processing. More specifically, this disclosure relates to computing devices, data processing methods, chips, and circuit boards. Background Technology
[0002] In recent years, significant progress has been made in object detection, instance segmentation, and keypoint detection technologies based on convolutional neural networks. These detection methods are typically based on LiDAR or RGB-D data and can be applied to fields such as autonomous driving and robot vision.
[0003] Unlike dense image data, LiDAR point cloud data is typically sparse, and the point density varies dramatically due to factors such as non-uniform sampling in 3D space, the effective range of the sensor, occlusion, and relative pose. Therefore, traditional convolutional neural networks, which are well-suited for dense data, become very inefficient when applied to this type of sparse data, especially when convolution operations are involved, wasting significant computational resources on zero-value data points.
[0004] Therefore, it is desirable to provide an improved data processing scheme suitable for sparse data such as point cloud data, thereby improving processing efficiency. Summary of the Invention
[0005] In order to at least partially solve one or more of the technical problems mentioned in the background art, the present disclosure provides a computing device, a data processing method, a chip, and a board.
[0006] In a first aspect, this disclosure discloses a computing device including a master device and a slave device, the slave device including a plurality of processing cores, wherein: the master device is configured to launch a first task, the first task being used to perform a fusion process, the fusion process instructing data elements in multiple streams of data to be fused to be merged into one ordered fused data stream according to their corresponding indices, wherein data elements with the same index are merged into one fused data element, the data element including any of scalar, vector, or higher-dimensional data; and the slave device is configured to schedule a corresponding number of the processing cores to execute the first task according to a splitting strategy of the first task, wherein each processing core performs the fusion process on a portion of the multiple streams of data to be fused.
[0007] In a second aspect, this disclosure provides a chip that includes a computing device according to any of the embodiments of the first aspect.
[0008] In a third aspect, this disclosure provides a board including the chip of any of the embodiments of the second aspect above.
[0009] In a fourth aspect, this disclosure provides a method for processing data using a computing device according to any of the embodiments of the first aspect.
[0010] Using the computing device, method for processing data using the computing device, chip, and board provided above, this disclosure provides an implementation scheme on a multi-core processor architecture for the data fusion processing step in convolution operations for sparse data. This scheme can effectively split tasks, thereby shortening processing time and improving overall efficiency. Furthermore, in some embodiments, by implementing multi-round pipelined processing on a multi-core processor, the fusion processing task can be completed quickly. In multi-round pipelined processing, a bucket sort method can be used to allocate the data portions of each round of pipelined processing, thereby achieving ordered concatenation of the output data from each round of pipelined processing. In some embodiments, indexes can also be reused across different output channel dimensions, thereby reducing data throughput and further accelerating processing. Attached Figure Description
[0011] The above and other objects, features, and advantages of exemplary embodiments of this disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding portions wherein:
[0012] Figure 1 This diagram shows the structure of the board card according to an embodiment of this disclosure;
[0013] Figure 2 This diagram illustrates the structure of the combined processing apparatus according to an embodiment of the present disclosure.
[0014] Figure 3 This diagram illustrates the internal structure of a single-core computing device according to an embodiment of the present disclosure.
[0015] Figure 4 This diagram illustrates the internal structure of a multi-core computing device according to an embodiment of the present disclosure.
[0016] Figure 5 This illustrates the operational principle of a conventional convolution scheme;
[0017] Figure 6 An exemplary schematic diagram of the sparse convolution scheme of this disclosure embodiment is shown;
[0018] Figure 7 This diagram illustrates the preprocessing of high-dimensional sparse input data according to an embodiment of this disclosure.
[0019] Figure 8 This illustrates the meaning of the multiplication operation in the embodiments disclosed herein;
[0020] Figures 9A-9CThis shows the index mapping relationship between the product result and the convolution operation result;
[0021] Figure 10 This illustrates the impact of padding in convolution operations on the input data index.
[0022] Figure 11 This illustrates an exemplary principle of the MERGE instruction;
[0023] Figure 12 An exemplary structural diagram of a computing device that can implement the embodiments of this disclosure is shown;
[0024] Figure 13 An exemplary schematic diagram of bucket sort is shown;
[0025] Figure 14 This schematically illustrates the buffer space partitioning within the memory core; and
[0026] Figure 15 An exemplary pipeline process for fusion processing according to an embodiment of this disclosure is shown. Detailed Implementation
[0027] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0028] It should be understood that the terms "first," "second," "third," and "fourth," etc., that may be used in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0029] It should also be understood that the terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0030] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection."
[0031] The specific embodiments disclosed herein will now be described in detail with reference to the accompanying drawings.
[0032] Figure 1 A schematic diagram of the structure of a board 10 according to an embodiment of this disclosure is shown. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 10 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities.
[0033] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. The calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface forms, such as a PCIe interface.
[0034] The board 10 also includes a storage device 104 for storing data, which includes one or more memory cells 105. The storage device 104 is connected to and transmits data with the controller 106 and the chip 101 via a bus. The controller 106 in the board 10 is configured to regulate the state of the chip 101. Therefore, in one application scenario, the controller 106 may include a microcontroller (MCU).
[0035] Figure 2 This is a structural diagram illustrating the combined processing device in chip 101 of this embodiment. (As shown) Figure 2 As shown, the combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203, and a storage device 204.
[0036] The computing device 201 is configured to perform user-specified operations. It is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.
[0037] Interface device 202 is used to transmit data and control commands between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to on-chip storage device of computing device 201. Further, computing device 201 can obtain control commands from processing device 203 via interface device 202 and write them to on-chip control cache of computing device 201. Alternatively or optionally, interface device 202 can also read data from storage device of computing device 201 and transmit it to processing device 203.
[0038] Processing device 203, as a general-purpose processing device, performs basic control including but not limited to data transfer, and starting and / or stopping computing device 201. Depending on the implementation, processing device 203 may be one or more types of processors, including but not limited to digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, computing device 201 disclosed herein can be considered as having a single-core structure or a homogeneous multi-core structure. However, when computing device 201 and processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure.
[0039] Storage device 204 is used to store data to be processed. It may be DRAM or DDR memory, typically 16G or larger in size, and is used to store data of computing device 201 and / or processing device 203.
[0040] Figure 3The diagram shows the internal structure of the single-core computing device 201. The single-core computing device 301 is used to process input data such as computer vision, speech, natural language, and data mining. The single-core computing device 301 includes three main modules: a control module 31, a processing module 32, and a storage module 33.
[0041] The control module 31 coordinates and controls the operation of the computation module 32 and the storage module 33 to complete the deep learning task. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 fetches instructions from the processing device 203, and the instruction decode unit 312 decodes the fetched instructions and sends the decoding result as control information to the computation module 32 and the storage module 33.
[0042] The computation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformations; the matrix operation unit 322 is responsible for the core computations of deep learning algorithms, namely matrix multiplication and convolution.
[0043] Storage module 33 is used to store or move relevant data, including neuron RAM (NRAM) 331, weight RAM (WRAM) 332, and direct memory access (DMA) module 333. NRAM 331 is used to store input neurons, output neurons, and intermediate results after computation; WRAM 332 is used to store the convolution kernels of the deep learning network, i.e., the weights; DMA 333 is connected to DRAM 204 through bus 34 and is responsible for data transfer between single-core computing device 301 and DRAM 204.
[0044] Figure 4 The diagram shows the internal structure of the computing device 201 as a multi-core system. The multi-core computing device 400 adopts a hierarchical structure design. As a system-on-a-chip, the multi-core computing device 400 includes at least one cluster, and each cluster includes multiple processor cores. In other words, the multi-core computing device 400 is constructed in a hierarchical structure of system-on-a-chip, cluster, and processor cores.
[0045] From the perspective of system-on-a-chip hierarchy, such as Figure 4 As shown, the multi-core computing device 400 includes an external storage controller 41, a peripheral communication module 42, an on-chip interconnect module 43, a synchronization module 44, and multiple clusters 45.
[0046] There can be multiple external storage controllers 41; two are shown as an example in the figure. These controllers are used to access external storage devices, such as those issued by the processor core, in response to access requests from the processor core. Figure 2 The DRAM 204 in the chip allows data to be read from or written to external storage devices. The peripheral communication module 42 receives control signals from the processing device 203 via the interface device 202, initiating the computing device 201 to execute tasks. The on-chip interconnect module 43 connects the external storage controller 41, the peripheral communication module 42, and multiple clusters 45 to transmit data and control signals between modules. The synchronization module 44 is a global barrier controller (GBC) used to coordinate the working progress of each cluster and ensure information synchronization. The multiple clusters 45 are the computing cores of the multi-core computing device 400. Four are shown exemplary in the figure; however, with hardware development, the multi-core computing device 400 of this disclosure can also include 8, 16, 64, or even more clusters 45. The clusters 45 are used to efficiently execute deep learning algorithms.
[0047] From the perspective of cluster hierarchy, such as Figure 4 As shown in the upper right corner, each cluster 45 includes a processing unit 402 and a memory core (MEM core) 404. The processing unit 402 performs various computational tasks. In some implementations, the processing unit may be a multi-core architecture, for example, including multiple processing cores (IPU cores) 411-1 to 411-n, to perform tasks such as large-scale vector computation. This disclosure does not limit the number of processing cores 411.
[0048] The internal architecture of the processing core 411 is as follows Figure 4 As shown below. Each processing core 411 may have multiple computing modules 424-1 to 424-m for performing computing tasks, and a local storage module 423 required for performing computing tasks. It should be noted that the local storage module 423 may include various communication modules for exchanging data with external storage units. For example, the local storage module 423 may include a communication module 421 for communicating with the shared storage module 415 in the storage core 404. The communication module 421 may be, for example, a move direct memory access (MVDMA) module. The local storage module 423 may also include a communication module 422 for exchanging data with off-chip memory, such as DRAM 408. The communication module 422 may be, for example, an input / output direct memory access (IODMA) module. The IODMA 422 controls the NRAM / WRAM in the local storage module 423. Figure 4 Not shown, see Figure 3The MVDMA 421 controls the access of NRAM / WRAM in the local storage module 423 and the shared storage module 415.
[0049] continue Figure 4 In the upper right view, storage core 404 is mainly used for storage and communication, namely storing shared data or intermediate results between processing cores 411, and performing communication between cluster 45 and DRAM 408, communication between clusters 45, and communication between processing cores 411. In other embodiments, storage core 404 has scalar operation capabilities to perform scalar operations to implement computational tasks in data communication.
[0050] The storage core 404 includes a large shared memory module (SRAM) 415, a broadcast bus 414, a cluster direct memory access (CDMA) module 418, a global direct memory access (GDMA) module 416, and a communication-time computation module 417. The SRAM 415 acts as a high-performance data relay station. Data multiplexed between different processing cores 411 within the same cluster 45 does not need to be obtained from the DRAM 408 by each processing core 411. Instead, it is relayed between processing cores 411 via the SRAM 415. The storage core 404 only needs to quickly distribute the multiplexed data from the SRAM 415 to multiple processing cores 411 to improve inter-core communication efficiency and greatly reduce on-chip and off-chip input / output access.
[0051] Broadcast bus 414, CDMA 418, and GDMA 416 are used to perform communication between processing cores 411, communication between clusters 45, and data transfer between cluster 45 and DRAM 408, respectively. These will be explained below.
[0052] The broadcast bus 414 is used to complete high-speed communication between the processing cores 411 within the cluster 45. In this embodiment, the broadcast bus 414 supports inter-core communication methods including unicast, multicast, and broadcast. Unicast refers to point-to-point (e.g., data transmission from one processing core to another) data transmission. Multicast is a communication method that transmits a copy of data from SRAM 415 to several specific processing cores 411. Broadcast is a communication method that transmits a copy of data from SRAM 415 to all processing cores 411, and is a special case of multicast.
[0053] CDMA 418 is used to control SRAM 415 access between different clusters 45 within the same computing device 201.
[0054] GDMA 416 works in conjunction with external memory controller 41 to control memory access from SRAM 415 to DRAM 408 in cluster 45, or to read data from DRAM 408 into SRAM 415. As described above, communication between DRAM 408 and NRAM / WRAM in local storage module 423 can be achieved through two channels. The first channel is a direct connection between DRAM 408 and local storage module 423 via IODMA 422; the second channel involves first transmitting data between DRAM 408 and SRAM 415 via GDMA 416, and then transmitting data between SRAM 415 and local storage module 423 via MVDMA 421. Although the second channel appears to require more components and has a longer data flow, in some embodiments, the bandwidth of the second channel is significantly greater than that of the first channel. Therefore, communication between DRAM 408 and local storage module 423 may be more efficient via the second channel. Embodiments of this disclosure can select the data transmission channel based on their hardware capabilities.
[0055] In some embodiments, storage core 404 can serve as a cache layer within cluster 45, thereby expanding communication bandwidth. Furthermore, storage core 404 can also facilitate communication with other clusters 45. For example, storage core 404 can implement communication functions such as broadcast, scatter, gather, reduce, and all-reduce between clusters 45. Broadcast refers to distributing the same data to all clusters; scatter refers to distributing different data to different clusters; gather refers to aggregating data from multiple clusters; reduce refers to processing data from multiple clusters according to a specified mapping function to obtain the final result and sending it to a specific cluster; and the difference between all-reduce and scatter is that the latter only sends the final result to one cluster, while all-reduce sends it to all clusters.
[0056] The communication-time computing module 417 can be used to complete computational tasks in communication processes, such as those involving the aforementioned protocols and full protocols, without relying on the processing unit 402, thereby improving communication efficiency and achieving the effect of "in-memory computing". Depending on different hardware implementations, the communication-time computing module 417 and the shared storage module 415 can be integrated in the same or different components. This disclosure embodiment is not limited in this respect, as long as the functions implemented and the technical effects achieved are similar to those disclosed herein, they are all within the protection scope of this disclosure.
[0057] Based on the aforementioned hardware environment, in one aspect, this disclosure provides a computing device that utilizes a multi-core processor architecture to achieve multi-channel data fusion processing. In another aspect, this disclosure also provides a convolution operation scheme suitable for sparse data, which can employ the fusion processing of the first aspect of this disclosure. To better understand the role of multi-channel data fusion processing, a convolution operation scheme suitable for sparse data is first described below. This sparse convolution scheme is applicable to multi-dimensional convolution operations, including but not limited to two-dimensional and three-dimensional convolution. For simplicity and ease of understanding, two-dimensional convolution is used as an example in some embodiments.
[0058] In this disclosure, "N-dimensional convolution" refers to the number of convolutional dimensions in which sliding accumulation is performed. For example, when N=2, the convolutional kernel performs translational accumulation along two dimensions (e.g., width W and height H) according to the corresponding convolutional stride. When N=3, the convolutional kernel performs translational accumulation along three dimensions (e.g., width W, height H, and depth D) according to the corresponding convolutional stride. When N=4, the convolutional kernel performs translational accumulation along four dimensions (e.g., width W, height H, depth D, and batch) according to the corresponding convolutional stride. In this disclosure, "non-convolutional dimension" refers to a dimension in which the convolutional kernel does not perform sliding accumulation.
[0059] To better understand the convolution scheme of this disclosure embodiment, the operational principle of a conventional convolution scheme will be described first using two-dimensional convolution as an example.
[0060] Figure 5 This illustrates the operational principle of a conventional convolution scheme. In this example, the convolution kernel 510 is dense, a 3×3 matrix, with the numbers in the kernel representing the corresponding weights. The input data 520 is a 6×6 matrix that is sparse, containing only three non-zero values: 2, 3, and 5, as shown by the dark squares. For simplicity, in this exemplary convolution process, the stride for both dimensions is set to 1, with zero padding (no dilation). The 3×3 gray squares in the diagram represent the sliding accumulation process of the convolution kernel across the input data. 530 shows the computation at the start of the convolution, 540 shows the computation one step to the right, and 550 shows the computation one step down. In each step, the weights of the convolution kernel are multiplied bitwise with the input data and accumulated. 560 shows the final computation result as the output data. The output data is a 4×4 matrix. It can be seen that the calculation of 530 corresponds to the data at coordinate (1,1) in the output data, the calculation of 540 corresponds to the data at coordinate (1,2) in the output data, and the calculation of 550 corresponds to the data at coordinate (2,1) in the output data.
[0061] In the sparse convolution operation of the embodiments disclosed herein, the convolution kernel is dense, and its input format can be the same as that of conventional convolution; while the input data is sparse, and its input format can be different from that of conventional convolution input data, thereby saving storage space. In some embodiments, the input data is sparse data, and each input data element has index and numerical information, which can be represented as (index, value). Depending on the application, data type, or scale, the index here can be one-dimensional, two-dimensional, or more-dimensional, and this disclosure is not limited in this respect. Indices of different dimensions can be converted to each other, for example, two-dimensional or more-dimensional indices can be converted to one-dimensional indices according to a predetermined traversal rule, and vice versa. Similarly, depending on the application, data type, or scale, the numerical information here can include scalars, vectors, or higher-dimensional data.
[0062] by Figure 5 Taking the example in the example, the input data has three non-sparse points: ((1,4),2), ((3,3),3), and ((5,6),5). For the first non-sparse point, (1,4) represents the two-dimensional index coordinates, and 2 represents the specific numerical value of the position pointed to by the index, which is a scalar in this example; the meanings of the other non-sparse points are similar.
[0063] from Figure 5 As can be seen from the description, the final result of sparse convolution depends only on the results of operations on non-zero input data elements. Therefore, multiplication and addition operations with the convolution kernel can be performed only on these non-zero input data elements. Furthermore, from... Figure 5 As can be seen from the output data 560, these multiply-accumulate operations can be further broken down into multiplication operations and corresponding positional accumulation operations. Accordingly, in the sparse convolution scheme of this disclosure embodiment, the sparse convolution operation can include three steps: calculating the operation result of the convolution kernel and the sparse input data (e.g., the product result or the multiply-accumulate result, considering that some dimensions of the data also need to be accumulated, such as the input channel dimension Ci described later); determining the index of each operation result; and merging the multi-way operation results into one fused data according to the index order, as the result of the sparse convolution operation. During the merging process, operation results with the same index are accumulated. In the following description, depending on the dimension of the data, in scenarios where the input channel dimension Ci is not mentioned, the multi-way operation result is sometimes directly referred to as the multi-way multiplication result. Those skilled in the art can understand its corresponding meaning based on the context.
[0064] Figure 6 An exemplary principle of a sparse convolution scheme according to an embodiment of this disclosure is shown. Figure 6 Still with Figure 5 The sparse convolution operation scheme implemented in this disclosure will be described using data as an example.
[0065] like Figure 6 As shown, in the MAC step, the input data element 620 is multiplied by the convolution kernel 610 to obtain a multi-way product result 630. When calculating the product result of the convolution kernel and the sparse input data, the weight data of each convolution dimension of the convolution kernel can be merged into one dimension, with the dimension size being offset, where offset is the total number of weight data in the convolution dimension.
[0066] As mentioned above, the sparse convolution scheme provided in this disclosure embodiment can be applied to multidimensional convolution operations, including but not limited to two-dimensional convolution and three-dimensional convolution. Figure 6 The diagram illustrates merging the weight data from the two convolutional dimensions (width W and height H) of a two-dimensional convolutional kernel into a single dimension. For example, the weight data from a 3×3 matrix can be expanded into a column containing nine weight data points, which can also be referred to as nine weight scalars.
[0067] It's understandable that when the convolution kernel is a three-dimensional kernel, the weight data across its three convolutional dimensions (width W, height H, and depth D) can be merged into a single dimension. For example, the 27 weight data from a 3×3×3 cubic convolution kernel can be unfolded into a single column, resulting in 27 weight scalars. Other cases can be deduced similarly. If the convolution kernel also has non-convolutional dimensions, such as the input channel Ci dimension and / or the output channel Co dimension, these dimensions can be retained since convolution operations are not performed on them. In this case, for example, if the convolution kernel includes a 3×3 convolutional dimension and also a non-convolutional dimension Ci = 2, then after merging these dimensions, the convolution kernel becomes 9×2, which can be viewed as nine weight vectors of length 2. For example, if the convolution kernel includes a 3×3×3 convolution dimension, as well as non-convolution dimensions Ci=2 and Co=4, then after merging the above dimensions, the convolution kernel becomes 27×2×4, which can be regarded as 4 sets of weights on the Co dimension, each set including 27 weight vectors of length 2.
[0068] The input data to be processed by convolution can include multidimensional data, and it is sparse across multiple dimensions. For example, in target detection based on LiDAR data, the input data is detection data in three-dimensional space, which may represent the grayscale value, RGB values, signal strength, etc., of each three-dimensional spatial coordinate point. Therefore, depending on the information content to be represented, the input data element at each coordinate point can be one-dimensional, two-dimensional, three-dimensional, or higher-dimensional data. Due to the characteristics of point cloud data, coordinate points with non-zero value data elements are sparse, that is, they are sparse in three spatial dimensions (e.g., width W, height H, and depth D).
[0069] Depending on the initial state of the input data, preprocessing can be performed before the sparse input data is provided to the computation circuit for processing. In some embodiments, such preprocessing may include, for example, merging multiple sparse dimensions into one dimension; densifying sparse data points in the input data along the merged dimension to form input data elements; and representing each input data element using indexes and numerical information. The index can be a one-dimensional or multi-dimensional index, and the numerical information can include any of scalar, vector, or higher-dimensional data.
[0070] In one example, for instance, reference Figure 5 The input data is a two-dimensional 6×6 matrix, sparse in both width (W) and height (H). During preprocessing, W and H are merged into a single dimension, and the sparse data points (2, 3, and 5 in this example) are compacted along this merged dimension, resulting in three compacted input data elements. Each input data element is then represented using indexes and numerical information. The index of a data element indicates its position within the sparse input data before compaction. For example, Figure 5 The indices of the three input data elements in the example are (1,4), (3,3), and (5,6). These indices are two-dimensional, but can be converted to one-dimensional indices, such as 4, 15, and 29. The numerical values of these three data elements are 2, 3, and 5, representing three scalars.
[0071] In another example, for example, reference Figure 7This diagram illustrates the preprocessing of high-dimensional sparse input data according to an embodiment of this disclosure. As shown, the sparse input data 710 comprises five dimensions: a batch (B) dimension, a three-dimensional HWD space dimension, and an input channel Ci dimension. The input data is sparse in the B dimension and the HWD three-dimensional space. In the diagram, dark squares in the HWD matrix represent numerical values, while other parts are all zero values. Multiple such HWD matrices exist in the B dimension, and the sparsity pattern (i.e., the position of the dark squares) on each matrix can be different. The input data is dense in the Ci dimension, which is the lowest dimension. Due to the limited representation capabilities of the accompanying figures, only four dimensions are shown in Figure 610, but the Ci dimension can be understood as the thickness of each dark square. The size of the Ci dimension is uniform, meaning the thickness of each dark square is the same. During preprocessing, the four sparse dimensions of the input data (B dimension and HWD three-dimensional space dimension) can be merged into a single dimension Ni, and the sparse data points (dark squares in the diagram) can be densified in the merged dimension to form dense input data elements. That is, each HWD 3D matrix in the B dimension undergoes the same dimension merging and compaction process to obtain the preprocessed compact form of the input data 720, which is a two-dimensional matrix with Ci as the low dimension and Ni as the merged dimension of BHWD as the high dimension.
[0072] Next, index and numerical information are used to represent each densified input data element. Figure 7 The indices of the input data elements in the example can be represented using their coordinates in the uncompacted BHWD four-dimensional space, or they can be converted into one-dimensional indices. The numerical information of each input data element can be viewed as a Ci vector.
[0073] The input data, after the above preprocessing, can be provided to the arithmetic circuit for further processing.
[0074] The input data elements can be arranged into an input vector to perform multiplication with the convolution kernel. In some embodiments, the input data elements can be arranged in their index order (e.g., ascending order) to accommodate subsequent merging. Each vector element in the input vector comprises one input data element. As described above, each input data element can be a scalar, vector, or higher-dimensional data.
[0075] Back Figure 6 Next, a multiplication operation can be performed on the convolutional dimension of the convolutional kernel 610 after dimension merging and the input vector 620 composed of input data elements, resulting in offset data 630, where each data path includes several product results, and each product result includes any of the following: scalar, vector, or higher-dimensional data. Figure 6In the example, the convolutional kernel 610 (the nine scalars in the diagram) after dimension merging performs scalar vector multiplication with the input vector 620 (a vector of length 3 consisting of 2, 3, and 5), resulting in nine vectors, or nine-way product results 630. In this example, each product result is a scalar.
[0076] The MAC step is equivalent to performing a 1×1 point-level convolution on the input vector for each row of data in the dimension-merged convolution kernel, resulting in a convolution result. To better understand the meaning of the multiplication operation in the MAC step above, Figure 8 The meaning of several of these multiplication operations is illustrated by example.
[0077] As shown in the figure, for Figure 6 The operation on the first weight scalar in the diagram: 1*[2,3,5]=[2,3,5], can be understood as corresponding to operations 810, 820, and 830 in the diagram, respectively, that is, the product results generated when 2, 3, and 5 are located at the first position of the convolution kernel. Similarly, for... Figure 6 The operation of the second weight scalar in the diagram: 2*[2,3,5]=[4,6,10], can be understood as the operation corresponding to 840, 850 and 860 in the diagram, that is, the product results generated when 2, 3 and 5 are located at the second position of the convolution kernel.
[0078] As mentioned earlier, input data can also include non-convolutional dimensions, such as those described in the previous reference. Figure 7 The input channel Ci dimension is described. In some embodiments, the input data may include N convolutional dimensions and at least one non-convolutional dimension, and correspondingly, the convolutional kernel may also include N convolutional dimensions and at least one non-convolutional dimension. In this case, in the MAC step described above, the corresponding operation processing can be performed on the non-convolutional dimensions according to the specific operation requirements. For example, in some examples, the non-convolutional dimensions of the input data include the input channel Ci dimension, and the non-convolutional dimensions of the convolutional kernel include the input channel Ci dimension and the output channel Co dimension. The operation processing on these non-convolutional dimensions may include, but is not limited to: keeping the results of the Ci dimension from being accumulated (e.g., in depthwise convolution); performing positional accumulation of the multiplication results on the convolutional dimensions on the Ci dimension (e.g., accumulation on the Ci dimension); and / or stacking the accumulation results on the Ci dimension on the Co dimension (e.g., accumulation on the Ci dimension, but not on the Co dimension) to obtain the multi-way operation result, where each operation result is a vector on the Co dimension.
[0079] continue Figure 6The MAC step obtains the operation results related to non-zero values. To obtain the final convolution operation result in the subsequent MERGE step, it is necessary to determine the indices of these operation results for corresponding accumulation. Therefore, the corresponding indices of these operation results can be obtained in the INDEX step. Specifically, in some embodiments, the index of each operation result in the multiplication results is determined according to the index mapping relationship. Each operation result is obtained by multiplying or multiplying and adding the weight data in the convolution kernel with the input data elements. Therefore, the index mapping relationship indicates the relationship between the position of the weight data, the index of the input data element, and the corresponding result element in the convolution operation result. In other words, the index of the operation result obtained by multiplying or multiplying and adding the weight data can be determined based on the position of the weight data in the convolution kernel and the index of the input data element.
[0080] Figure 9A This example illustrates the index mapping relationship between partial product results and convolution operation results in the MAC step. The meanings of 910A, 920A, and 930A in the figure are... Figure 8 Similar to 810, 820, and 830, these represent the product operation of the first weight scalar and the input vector: 1*[2,3,5]=[2,3,5]. The arrows in the diagram indicate the corresponding positions of each product result (910A, 920A, and 930A) in the convolution result (940A). Specifically, the product result of 910A corresponds to position (1,4) in the 4×4 convolution result, the product result of 920A corresponds to position (3,4) in the convolution result, and the product result of 930A overflows the convolution result range and is therefore invalid.
[0081] from Figure 9A It can be seen that when the input vector is multiplied with the first weight data of the convolution kernel, the index has the following mapping relationship: Assuming that the index of the vector element (i.e., data points 2, 3 and 5) in the input vector is (x, y), then the index of the product result of the product operation with the first weight data is also (x, y).
[0082] Figure 9B This example illustrates the index mapping between partial product results and convolution operation results in the MAC step. In the figure, 910B, 920B, and 930B represent the product operation of the second weight scalar and the input vector: 2*[2,3,5]=[4,6,10]. The arrows in the figure indicate the corresponding positions of each product result (910B, 920B, and 930B) in the convolution operation result (940B). Specifically, the product result of 910B corresponds to position (1,3) in the 4×4 convolution result, the product result of 920B corresponds to position (3,2) in the convolution result, and the product result of 930B overflows the range of the convolution result and is an invalid result.
[0083] from Figure 9B It can be seen that when the input vector is multiplied with the second weight data of the convolution kernel, the index has the following mapping relationship: Assuming that the index of the vector element (i.e., data points 2, 3 and 5) in the input vector is (x, y), then the index of the product result of the product operation with the second weight data is (x, y-1).
[0084] Figure 9C This example illustrates the index mapping between partial product results and convolution operation results in the MAC step. In the figure, 910C, 920C, and 930C represent the product operation of the 9th weight scalar and the input vector: 1*[2,3,5]=[2,3,5]. The arrows in the figure indicate the corresponding positions of each product result (910C, 920C, and 930C) in the convolution operation result (940C). Specifically, the product result of 910C overflows the convolution result range and is invalid; the product result of 920C corresponds to position (1,1) in the 4×4 convolution result, and the product result of 930C corresponds to position (3,4) in the convolution result.
[0085] from Figure 9C It can be seen that when the input vector is multiplied by the 9th weight data of the convolution kernel, the index has the following mapping relationship: Assuming that the index of the vector element (i.e., data points 2, 3 and 5) in the input vector is (x, y), then the index of the product result of the product operation with the 9th weight data is (x-2, y-2).
[0086] comprehensive Figures 9A-9C As can be seen, each vector element in the input vector (i.e., data points 2, 3, and 5) sequentially traverses the 3×3 convolution kernel. Therefore, the offset of each data point relative to the center point of the convolution kernel (here, the center of the 3×3 convolution kernel is the 5th weight data) is fixed. Based on this characteristic, the index of the center point of the convolution kernel can be obtained sequentially based on the specific location of the data points. Then, the index of the center point can be mapped to the index of the output point. Thus, the index of the product result associated with each vector element in the input vector can be directly determined based on the index of that vector element. That is, knowing only the index of the input data element allows us to determine the index of the product result of that input data element multiplied by all the weight data.
[0087] For example, for a 3×3 two-dimensional convolution kernel, the coordinate offset relative to the center point of this two-dimensional convolution kernel is a constant when data points traverse the kernel. In this example, these 9 coordinate offsets can be constructed as follows:
[0088] (1,1),(0,1),(-1,1);
[0089] (1,0),(0,0),(-1,0);
[0090] (1,-1),(0,-1),(-1,-1).
[0091] For example, for a 3×3×3 convolution kernel, the coordinate offset relative to the center point of this 3D convolution kernel is a constant when data points traverse the kernel. In this example, these 27 coordinate offsets can be constructed as follows:
[0092] (1,1,1),(0,1,1),(-1,1,1),(1,0,1),(0,0,1),(-1,0,1),(1,-1,1),(0,-1,1),(-1,-1,1);
[0093] (1,1,0),(0,1,0),(-1,1,0),(1,0,0),(0,0,0),(-1,0,0),(1,-1,0),(0,-1,0),(-1,-1,0);
[0094] (1,1,-1),(0,1,-1),(-1,1,-1),(1,0,-1),(0,0,-1),(-1,0,-1),(1,-1,-1),(0,-1,-1),(-1,-1,-1).
[0095] Therefore, based on the indices of the input data points and the aforementioned fixed coordinate offsets, the indices of the convolution kernel center points corresponding to each convolution kernel traversal can be sequentially calculated. Then, mapping the center point indices to the output point indices determines the index of each product result generated by the input data points.
[0096] from Figures 9A-9C The diagram also shows that some product results have overflowed the range of the convolution results and are therefore invalid. For these cases, the indexes exceeding the range of the convolution results (i.e., the range of the output data dimensions) can be set to a predetermined value, such as -1, to identify these invalid results in subsequent processing and prevent them from being merged.
[0097] In some embodiments, convolution operations involve padding. For example, in object detection algorithms based on LiDAR data, same padding is required, meaning that padding ensures the shape of the input data is the same as the shape of the output data after the convolution operation. It is understandable that different padding rules may exist in other applications of convolution operations.
[0098] Figure 10 The effect of padding in convolution operations on the input data index is shown.
[0099] As shown in the figure, matrix 1010 represents the initial input data, and matrix 1020 represents the padded input data. The initial input data is, for example, a 2×3 matrix. The padded input data is processed according to the padded rules, by adding one column to the left, three columns to the right, four rows to the top, and one row to the bottom of the initial input data. The padded data can be, for example, zero.
[0100] For any data point (x, y) in the initial input data, its coordinates in the input data after padding become (x + pad_left, y + pad_top), where pad_left is the padding amount on the left and pad_top is the padding amount on the top. Therefore, the indices of the input data can be adjusted using simple addition operations and according to the padding rules.
[0101] In this embodiment, the indices of the input data elements can be adjusted based on the padding rules of the convolution operation before determining the index of the product result according to the index mapping relationship.
[0102] Those skilled in the art will understand that index adjustment processing can also be performed after or during index mapping, as long as the impact of the filling rules is taken into account, and the embodiments disclosed herein are not limited in this respect.
[0103] Back Figure 6 Figure 640 shows the indices corresponding to the 9-way product results determined by the INDEX step, with each product result having a corresponding index. Dark squares indicate invalid results, whose indices are set to -1.
[0104] After obtaining the multi-way product results through the MAC step and determining the index of each product result through the INDEX step, these multi-way product results can be further fused to obtain the convolution operation result.
[0105] Specifically, in the MERGE step, the multiplication results are merged and combined according to their index order to obtain the convolution operation result, where multiplication results with the same index are accumulated. Figure 650 shows the merged indexes, where duplicate indices, such as indices 2 and 3, have been removed. Figure 660 shows the merged data, where multiplication results with the same index are accumulated; for example, the data corresponding to two indices 2, 6 and 12, are accumulated, and the data corresponding to two indices 3, 4 and 3, are accumulated.
[0106] By comparison Figure 5 and Figure 6 The calculation results show that, based on Figure 6 The fused index 650 and fused data 660 can recover the sparse form of the convolution operation result, which completely corresponds to Figure 5The result of a regular 4×4 convolution operation is 560.
[0107] The above mainly describes a sparse two-dimensional convolution scheme, whose operational principle can be similarly extended to sparse three-dimensional or higher-dimensional convolution operations. For example, in three-dimensional sparse convolution, the convolution kernel can also have non-convolutional dimensions: an input channel Ci dimension and an output channel Co dimension. In the convolution operation of this example, the results of the operation in the Ci dimension need to be accumulated, while the results of the operation in the Co dimension are not accumulated.
[0108] After the aforementioned MAC processing, a multi-way product result is obtained. It can be understood that in this MAC processing, in addition to performing multiplication operations along the convolution dimension, accumulation is also performed along the Ci dimension. Since no operations are performed along the Co dimension, each product result in the obtained multi-way product result can be viewed as a vector along the Co direction.
[0109] As described above regarding the INDEX step, after determining the fixed offset based on the convolution kernel, the index of the product result is only related to the index of the input data. Since the input data does not have a Co dimension, the indices of each product result can be reused along the Co direction.
[0110] Therefore, the above describes a convolution operation scheme for sparse data, in which the effective product results can be sorted and accumulated through data fusion processing (MERGE step). This data fusion step can be implemented using a specially designed hardware instruction, the MERGE instruction. The basic function of the MERGE instruction is to merge multiple streams of data into a single stream according to their index order. The MERGE instruction can have multiple processing modes to adapt to different needs. The MERGE instruction can include a mode flag to indicate different processing modes.
[0111] Figure 11 The diagram illustrates the exemplary principle of the MERGE command. The figure exemplifies four streams of data to be fused, each stream comprising six data elements. Data elements can be scalars, vectors, or higher-dimensional tensors. In the figure, data elements are exemplarily shown as vectors, such as D11, D12, ..., D46. These vectors have a uniform length; for example, D11 is (d1, d2, d3, ..., dn) with a length of n. Each data element has an associated index indicating its position within the corresponding stream of data. For example, the original stream of data might contain 1000 data elements, but only some of these elements are valid. In this case, these valid elements can be extracted to form the data to be fused, and their corresponding indices can be extracted to indicate their positions in the original data; these indices form the fusion index.
[0112] The diagram schematically illustrates the four indices to be merged, with each index corresponding to one channel of data to be merged. The first index identifies the position of each data element in the first channel, the second index identifies the position of each data element in the second channel, and so on. Furthermore, the index elements in each channel are stored in an ordered manner and correspond one-to-one with the data elements in the corresponding channel. In the example shown, the index elements in each channel are arranged in a first order (e.g., ascending order), and the data elements in each channel are also arranged in the order of their corresponding indices. For example, the first index element in the first channel indicates that the index of the first data element in the first channel is 0, i.e., the first element; the second index element in the first channel indicates that the index of the second data element in the first channel is 2, i.e., the third element; and so on.
[0113] The figure shows exemplary results of the MERGE command in different processing modes.
[0114] In the first processing mode, Mode 1, also known as the "merge sort mode," only the indices of the aforementioned data are processed. Specifically, after the merging process, the indices of the data from each path are merged into a single merged index, and the merged index elements are arranged in a second order (e.g., ascending order). Duplicate index elements are retained in this merge sort process. As shown in the figure, the four indices to be merged are merged into a single merged index, comprising 24 data elements.
[0115] In the second processing mode, Mode 2, also known as the "sorting and accumulating mode," data elements from multiple streams of data to be merged are grouped into a single ordered merged data stream according to their corresponding indices. Data elements with the same index are accumulated and merged into a single merged data element. As shown in the figure, four streams of data to be merged are grouped into a single ordered merged data stream according to their corresponding indices, and data elements with the same index are accumulated and merged into a single merged data element. In this example, the merged index includes 16 index elements, arranged in a second order (e.g., ascending order), where duplicate index elements in the original merged indexes have been removed, as shown by the dark squares in the figure. Correspondingly, the merged data also includes 16 data elements, arranged in the order of their corresponding indices, and data elements with the same index are accumulated and merged into a single merged data element, as shown by the dark squares in the figure. The second processing mode, Mode 2, is commonly used in sparse matrix multiplication operations; therefore, it can also be called the "matrix multiplication mode."
[0116] In the third processing mode, Mode3, similar to the second processing mode, sorting and accumulation processing is also required. Figure 10The second and third processing modes are illustrated using the same processing result. The difference between these two modes lies in the output format. In the second processing mode, for cases where accumulation exists in the fused data elements, the accumulated result is directly output. In the third processing mode, for at least cases where accumulation exists in the fused data elements, the accumulated result is not output; instead, the relevant accumulation expression is output. In some implementations, all fused data elements can be output as expressions for easier, uniform processing. For example, fused data elements without accumulation can be represented as an accumulation expression with zero. This accumulation expression can be called an operation structure. In some implementations, each operation structure indicates an in-situ addition operation, including addresses pointing to the two addends. The third processing mode is particularly suitable for cases where the data elements to be fused are vectors or higher-dimensional tensors, such as in radar algorithms based on LiDAR data; therefore, the third processing mode can also be called the "radar algorithm mode."
[0117] Those skilled in the art will understand that the first order and the second order mentioned above may be the same or different, and both may be selected from either: an ascending order or a descending order. Those skilled in the art will also understand that although the diagram shows that each data path has an equal number of data elements, the number of data elements in each data path may be the same or different, and this disclosure is not limited in this respect. Furthermore, those skilled in the art will understand that since the MERGE instruction can have multiple processing modes, the required parameters may also vary accordingly in different processing modes. For example, in the first processing mode, it is not necessary to provide the data to be merged because only the index of the data is sorted. In the third processing mode, the output parameters also include the operation structure.
[0118] To accelerate the processing efficiency of MERGE instructions, this disclosed embodiment provides an implementation scheme on a multi-core processor architecture that supports parallel processing of MERGE instructions.
[0119] Figure 12 An exemplary structural diagram of a computing device that can implement embodiments of the present disclosure is shown. As shown, the computing device 1200 includes a master device 1210 and a slave device 1220, wherein the slave device 1220 may include a plurality of processing cores 1221.
[0120] The master device 1210 can be, for example, a general-purpose processor, which acts as a control device (referred to as the host device) and is responsible for complex control and scheduling tasks. The master device 1210 can be, for example, a general-purpose processor. Figure 2 The processing unit 203 is located within the processing unit. The slave device 1220 may be, for example, a domain-specific processor responsible for large-scale parallel computing or domain-specific computing tasks. The slave device 1220 may be, for example, a... Figure 2 The computing device 201 in the middle. The two work together to complete the computing task.
[0121] In some embodiments, the computing device 1200 can be used to implement the fusion process described above. Specifically, the master device 1210 can issue a first task. The first task is used to perform a fusion process that instructs the merging of data elements from multiple streams of data to be merged into a single ordered fused data stream according to their corresponding indices, wherein data elements with the same index are merged into a single fused data element. The data element can include any of scalar, vector, or higher-dimensional data. The slave device 1220 can schedule a corresponding number of processing cores 1221 to execute the first task according to the splitting strategy of the first task. Each processing core performs the fusion process described above on a portion of the multiple streams of data to be merged.
[0122] In some embodiments, a bucket sorting method can be used to divide the data to be fused into parts for parallel processing on multiple processing cores. Specifically, the master device 1210 can be further configured to determine the splitting strategy for the first task to support the bucket sorting implementation. The splitting strategy may include: the number of processing cores required to complete the first task, and the index range of the data to be fused to be processed corresponding to each core count, wherein the index ranges corresponding to each core count are ordered. Here, "core count" refers to how many times a task needs to be executed by processing cores to complete it, that is, the spatial or temporal expansion of the corresponding task.
[0123] The device 1220 may also include a storage core 1222, which may be shared by multiple processing cores 1221 to store data before, during and / or after processing by the processing cores.
[0124] Figure 13 This diagram illustrates an exemplary principle of bucket sort. The basic working principle of bucket sort is to divide the data to be sorted into a finite number of buckets, and then sort each bucket separately.
[0125] As shown in the diagram, assume the original array to be sorted contains 12 data items in a random order, and there are 4 buckets. Each bucket is responsible for sorting a specific range of data, and the data ranges within each bucket are ordered. For example, in the example shown, the data ranges of the four buckets are roughly evenly distributed, in ascending order: Bucket 1 has a data range of 0-25, Bucket 2 has a data range of 26-50, Bucket 3 has a data range of 51-75, and Bucket 4 has a data range of 76-100. Therefore, based on the data range of each bucket, the data in the original array can be assigned to the corresponding bucket. The diagram shows that Bucket 1 has 5 numbers, Bucket 2 has 2 numbers, Bucket 3 has 1 number, and Bucket 4 has 4 numbers. Next, sorting is performed within each bucket. Finally, the sorted arrays from each bucket are concatenated according to the bucket order to obtain the final sorted result. The diagram shows the concatenated sorted result.
[0126] The process of allocating data to buckets can also be represented as a mapping function f. Through the mapping function f, the key k to be sorted is mapped to the i-th bucket, and at this time the key k is the element in bucket B[i].
[0127] In some embodiments, each processing core in a multi-core processor can be considered as a bucket in the bucket sort described above, with a limited processing capacity at any given time. Distributing the data to be fused across multiple processing cores is equivalent to placing the data to be fused into multiple buckets for processing. The index ranges corresponding to the processing in each bucket are ordered. Then, the fusion processing results from each processing core are directly stored back to external storage circuitry (e.g., Figure 2 On the storage device 204 (e.g., DRAM), the results of sorting all buckets can be sequentially concatenated. The merging operation performed by the MERGE instruction can be viewed as a sorting operation within each bucket.
[0128] In some embodiments, the slave device can be configured to schedule L rounds, with N per round, according to a splitting strategy for the first task. core Each processing core is used to execute the first task. The number of rounds L for executing the first task, and the number of processing cores N in each round... core The number of processing cores required to complete the first task and the number of currently available processing cores can be determined, for example, by the master device or the slave device. For example, if completing the first task requires 16 processing cores and there are 16 currently available processing cores, then only one round of processing is needed to complete it; if there are 8 currently available processing cores, then 2 rounds of processing are needed to complete it; if there are 4 currently available processing cores, then 4 rounds of processing are needed to complete it.
[0129] In some embodiments, such as when the fusion process is part of the aforementioned sparse convolution operation, in the case of three-dimensional sparse convolution, since the indices of each product result can be reused in the Co direction (i.e., the product results in the Co direction correspond to the same index), the Co dimension can be split to achieve index reuse; for example, the same index can be reused Co times. Specifically, in the same round of processing, each processing kernel performs fusion processing on the data to be fused in different output channel dimensions, and the data to be fused in these different output channel dimensions reuse the same index.
[0130] In some embodiments, during each processing round, the slave device can load indices falling within the index range corresponding to the current processing round to a storage core shared by multiple processing cores. These indices are shared among the scheduled processing cores. Therefore, the indices can be broadcast to these processing cores. Each scheduled processing core can load the values of the data to be fused on the Co dimension allocated to this processing core, corresponding to the broadcasted indices. Then, the processing core performs fusion processing on the loaded values and indices of the data to be fused. Finally, the processing core can store the fused data back to external storage circuitry. In some implementations, the processing core can store the processed data back to external storage circuitry via a storage core, for example... Figure 2 In the storage device (e.g., DDR) 204. In other implementations, when the interface between the storage core and the processing core is busy, the processed data can be directly written back to the external storage circuit via another interface.
[0131] To make the most of the storage space in each round of processing, it can be divided according to the index distribution of the data to be merged, thereby ensuring that the storage space is filled as much as possible in each round of processing.
[0132] In some embodiments, the storage core in the slave device may be configured with at least two buffers to support data access between one buffer and external storage circuitry, while simultaneously accessing data between the other buffer and the processing core. These two buffers may be referred to as the ping-pong buffer space and the pong buffer space, i.e., employing a ping-pong pipelined approach.
[0133] Specifically, when the processing core performs operations on data in the ping-buffer space of the memory core, the memory core can access data from external storage circuits (e.g., Figure 2 The storage device 204, for example DRAM, loads the next computation data into its pong buffer space. The memory access interface between the memory core and the external memory circuitry is different from the memory access interface between the memory core and the processing core, thus supporting the above-mentioned parallel processing and forming a pipelined process.
[0134] The memory core has limited space, such as 512KB, and space needs to be reserved for the compiler, for example, 128KB. Therefore, the usable space during merging is only 512-128=384KB. Furthermore, to support pipelining, the available space for the memory circuitry is divided into two buffer spaces: a ping-pong buffer and a pong-pong buffer. In some implementations, these two buffer spaces are evenly distributed to maximize pipelining efficiency. In the aforementioned example, the space available for each merging operation is RAM_merge_size, which is 384 / 2 = 192KB.
[0135] Limited by the available space on the storage cores, space management is necessary to efficiently perform fusion processing. Two factors need to be considered in space management: first, space allocations must not be polluted or stacked; second, the buffer space must be large enough to hold the data processed each time.
[0136] Based on the above considerations, in some embodiments, the maximum number of indexes Nmax that can be processed in one fusion process can be determined according to the available space size of the storage core and the relevant parameters of the fusion process. Then, according to the determined Nmax, buffer space is allocated in each buffer for each relevant parameter of the fusion process.
[0137] The relevant parameters for fusion processing mainly include at least one of the following: the size of the multiple data streams to be fused (size_addr); the value of the multiple data streams to be fused (merge_input_mac_result); the index of the multiple data streams to be fused (merge_input_output_index); the value of the fused data (output_data); the index of the fused data (output_index); the operation structure representing the fused data elements (out_op_addr); and the data to be fused in each round of fusion processing (Compute_buffer).
[0138] The size of the multiple streams of data to be fused refers to the size of the K streams of input data that need to be fused, which can be indicated, for example, by the starting address of each stream. This address is a first-level pointer, which can be labeled size_addr, and includes K elements, where the i-th element represents the number of data elements in the i-th stream, and 0 < i ≤ K. It can be understood that when the fusion process is used for the sparse convolution operation in the aforementioned embodiment, K = offset. A buffer space needs to be reserved in the storage kernel for size_addr, with a size of K * index_data_type = offset * index_data_type, where index_data_type represents the data type of the element at that address.
[0139] The number of other parameters is related to the maximum number of indices Nmax that can be processed in a single fusion process. Therefore, the maximum number of indices Nmax that can be processed can be determined based on the available space of the storage kernel and the requirements of these parameters, thereby further determining the space occupied by each parameter. In the following description, the space occupied by each parameter is described using the sparse convolution operation scenario of the aforementioned embodiment as an example.
[0140] For the numerical values of multiple data streams to be fused, i.e., the input K data streams, their space can be calculated as: Nmax * Co * input_data_type, where Co represents the output channel dimension and input_data_type represents the data type of the input data. In the example of sparse convolution operation, the numerical values of these multiple data streams to be fused are... Figure 6 The product result calculated in the MAC step can therefore be represented as merge_input_mac_result.
[0141] For the index of multiple streams of data to be merged, since there is a one-to-one correspondence between data and index, the space occupied by the K indices corresponding to the input K streams of data can be calculated as: Nmax * index_data_type, where index_data_type represents the data type of the index. In the example of sparse convolution operation, the index of the multiple streams of data to be merged is... Figure 6 The index calculated in the INDEX step can therefore be represented as merge_input_output_index. In some embodiments, the index elements in the input K-way index are ordered in each way, for example, in ascending order.
[0142] Regarding the numerical values of the fused data, it's understandable that since the data to be fused undergoes fusion processing, the number of output data items after fusion will always be less than or equal to the number of data items to be fused. Its maximum space requirement is: Nmax * Co * output_data_type, where output_data_type represents the data type of the output data. The numerical values of the fused data can be represented using output_data.
[0143] Similarly, the index of the merged data, i.e. the output index (denoted as output_index), occupies a maximum space of: Nmax*index_data_type.
[0144] For an operation structure representing fused data elements (denoted as out_op_addr), its maximum space requirement is Nmax*2*8. In this case, the operation structure has at most Nmax elements, each element being a structure, and each operation structure element indicates an in-place addition operation, including addresses pointing to the two addends. Each address can, for example, use 8 bytes.
[0145] For the data to be merged in each round of fusion processing (denoted as Compute_buffer), it represents the input index of the actual execution of the MERGE instruction, also known as the computation buffer space, the meaning of which will be described in detail later. The space occupied by this part can be calculated to be at most: Nmax * index_data_type.
[0146] The above analysis of the space occupied by the relevant parameters of the fusion processing shows that the total space occupied, which is the sum of the space occupied by each item, is at most equal to the available space of the storage core. This relationship can be expressed as the following formula (1):
[0147] K*index_data_type+Nmax*Co*input_data_type+Nmax*index_data_type+Nmax*Co*output_data_type+Nmax*index_data_type+Nmax*2*8+
[0148] Nmax*index_data_type=RAM_merge_size (1)
[0149] Therefore, the maximum number of indices that can be processed in a single fusion process, Nmax, can be determined as follows: Nmax = (RAM_merge_size - K * index_data_type) / (Co * input_data_type + index_data_type + Co * output_data_type + index_data_type + 2 * 8 + ...
[0150] index_data_type) (2)
[0151] Once Nmax is determined, the space occupied by each of the above parameters can also be determined.
[0152] Figure 14 The diagram illustrates the buffer space partitioning within the memory core. As shown, the available space of the memory core is evenly divided into two buffer spaces: a ping buffer space 1410 and a pong buffer space 1420. Each buffer space is allocated corresponding buffer space for the aforementioned parameters, thus adhering to the principle of preventing space contamination and interference.
[0153] After allocating space resources for each parameter, the MERGE instruction can be executed in a pipelined manner. As described above regarding the principle of MERGE, this fusion process is strongly correlated with the K-way index (i.e., merge_input_output_index) of the data to be merged, and therefore its pipeline is also strongly correlated with the K-way index of the data to be merged.
[0154] In some embodiments, the master device may allocate the index range for each round of processing as follows, such that the amount of data processed in each round is approximately the same and that the available storage space is filled as much as possible. In some implementations, the master device may determine the number N of data that a single processing core can process at one time based on the storage capacity of the storage cores in the slave device and the parameter configuration of the fusion processing. spacing ; and based on the number of data N spacing Given the total number of data points to be merged from multiple sources, determine the number of cores required to complete the first task (fusion processing).
[0155] Specifically, based on the space management calculation method described above, the maximum buffer space allocated to the fused data values in the buffer can be determined as: Nmax * Co * output_data_type. Therefore, Nmax * Co can be used as the number of data items that can be processed in each round of processing (i.e., each bucket), which can also be called the bucket spacing N. spacing .
[0156] Once the distance between buckets is determined, the number of buckets N can be determined accordingly. bucket Divide the total number of data points by the bucket spacing N spacing When applied to sparse convolution operations, the total number of data points to be fused can be represented as Nin * offset, where Nin represents the number of non-sparse points in the input data, and offset represents the total number of weight data points in the convolution dimension of the convolution kernel. Therefore, the number of buckets can be represented as:
[0157] N bucket =Nin*offset / N spacing (3)
[0158] As can be seen from the above formula, the number of rounds (buckets) processed is directly proportional to the number of non-sparse points in the input data; in other words, it is directly proportional to the sparsity of the input data.
[0159] After determining the bucket spacing and number of buckets, the indexes of the data to be merged can be sorted, and then the sorted indexes can be ordered according to the bucket spacing N. spacing Sequential partitioning is performed to determine the index range corresponding to each round of processing, or the index range or index interval of each bucket. This partitioning method ensures that the data to be merged in each round of processing fills as much of the available space in the storage circuit as possible.
[0160] In some embodiments, sorting the indexes of the data to be merged can be performed using the first processing mode of the MERGE instruction described above. Specifically, the merge sort mode of the merging instruction is invoked to sort the indexes of the multiple data streams to be merged.
[0161] In some implementations, the master device can launch a second task before launching the first task (fusion processing). This second task is used to pre-sort the indices of data elements in the multiple streams of data to be fused. Then, the master device can, based on the pre-sorted indices, process the data according to the aforementioned number N. spacing Sequential splitting is performed to determine the index range of the data to be fused corresponding to each core.
[0162] It is understandable that other sorting methods can be used to pre-sort the index, thereby dividing the index range of each of the above buckets.
[0163] After determining the index range for each round of processing based on the aforementioned method, the indexes and associated data falling within the corresponding index range can be retrieved in each round of processing and then fusion processing can be performed.
[0164] As mentioned earlier, the indexes of the multiple streams of data to be merged are ordered within each stream, for example, arranged in ascending order. Therefore, during each round of data retrieval, each stream can be selected sequentially to identify all indexes falling within the corresponding index range.
[0165] Specifically, in some embodiments, indices falling within the index range corresponding to the current processing round and their corresponding data to be fused are selected sequentially from the indices of multiple streams of data to be fused. The number of indices extracted from each stream does not exceed the remaining processable number. Here, the remaining processable number is equal to the number of data points N. spacing The difference between the number of indexes selected and the number of indexes already selected.
[0166] For example, in each round of processing, the storage core reserves space for 20 (bucket spacing) numbers. Therefore, 20 numbers can be retrieved for each index, and then a size comparison operation (e.g., __bang_ge (greater than) / __bang_le (less than) function) can be used to select the index that falls within the range of 0 to 13, and the result is stored in the computation buffer space.
[0167] In some implementations, the data retrieval steps can be as follows: 20 numbers are retrieved from the first index, and 2 numbers meet the requirements; then 20-2=18 numbers are retrieved from the second index, and 2 numbers meet the requirements; then 20-2-2=16 numbers are retrieved from the third index, and so on. As can be seen from the above steps, since the number of indices retrieved gradually decreases each time, the amount of I / O can be reduced.
[0168] Furthermore, as described above regarding sparse convolution operations, the multiplication-addition result of the MAC step may be invalid, with its corresponding index set to a fixed value, such as -1. In this case, when executing the MERGE instruction, the hardware can avoid outputting any data when encountering an index of -1, thus preventing invalid processing.
[0169] When there is an invalid index (e.g., -1) in the index of the data to be merged, the bucket spacing can be adjusted appropriately when the bucket sort method is used for processing.
[0170] As mentioned earlier, Nmax is a fixed value calculated based on the available space of the storage cores and the relevant parameters of the fusion processing; that is, the maximum number of indexes that can be processed at one time. If there are many invalid indexes (-1), the bucket distance will decrease. This is because the number of invalid indexes also occupies the number of indexes in Nmax. Therefore, after pre-sorting the indexes of the data to be fused, the number of invalid indexes can be counted, which can be denoted as N. invaild Then, reserve the corresponding buffer space in the storage core. This way, each time the index to be merged is retrieved and stored in the storage core, N buffer space must always be reserved in the storage core. invaild The number of spaces is used to prevent data overflow.
[0171] From the previous reference Figure 8 As shown in Figure 9, only the data within the four outermost borders (top, bottom, left, and right) of the input data will have an invalid index "-1". Therefore, even with a large input data set, only the data within these four outermost borders will have an invalid index of "-1". This demonstrates that the number of invalid indices will be very small, and reserving a small amount of space is sufficient.
[0172] In practical applications, such as radar algorithms, convolution operations often require padding with zeros around the edges (as mentioned in the previous reference). Figure 10 (As described), at least one layer of zeros will be added to the top, bottom, left, and right edges (e.g., padding amount = 1 or 2). After padding, the four outermost borders of the input data are all 0, thus significantly reducing the number of indices containing "-1". Furthermore, the radar algorithm processes point cloud data. Observations of real data show that point cloud data becomes sparser towards the edges, meaning point cloud objects are generally located in the center of the input image. Therefore, the probability of invalid indices appearing at the borders is low, and reserving a small amount of space for invalid indices is sufficient.
[0173] Therefore, in some embodiments, bucket sorting is applied to achieve the fusion processing of all data to be fused on a multi-core processor architecture. In each round of processing, the slave device can select input indices that fall within the index range corresponding to the current round based on the indices of the multiple data to be fused, and load them into the compute buffer space of the storage core (i.e., the previously allocated compute_buffer). These indices can then be broadcast to each processing core participating in the computation via the broadcast bus. The processing core can execute the fusion processing indicated by the MERGE instruction on the indices in the compute buffer space and the data to be fused corresponding to these indices. Finally, the fused data is stored back to a designated location, such as off-chip memory circuit DDR. It can be understood that since the processing rounds are also ordered, that is, the buckets are ordered, the fused data obtained from each round of processing can be concatenated in the round order to obtain the final result.
[0174] The processing core in the device is also equipped with local storage circuitry. In order to further support pipelined processing, in some embodiments, the local storage circuitry may be configured with at least two storage areas to support data access between one storage area and the storage core or external storage circuitry, while data access between the other storage area and the processing core is performed simultaneously.
[0175] Figure 15 An exemplary pipeline of fusion processing according to an embodiment of this disclosure is illustrated. The figure shows the storage core divided into ping-pong and pong-pong buffer spaces, and the local storage circuitry of the processing core also divided into ping-pong and pong-pong buffer spaces, thereby supporting a five-stage pipeline. The five-stage pipeline in the figure may include data loading L, first data transfer MV1, computation C, second data transfer MV2, and write-back S. The time slices are listed on the left side of the figure, and the steps flow in chronological order.
[0176] At time slice 0, an (L1) index (e.g., Id1) is loaded into the ping buffer space of the memory core. This index may, for example, come from the last-level cache LLC of the computing device.
[0177] In time slice 1, an index (e.g., Id2) is loaded into the storage core's ping buffer space (L2). At the same time, the index Id1 in the storage core's ping buffer space can be transferred (MV11) to the processing core's ping buffer space via the broadcast bus.
[0178] In time slice 2, an index (e.g., Id3) is loaded (L3) into the ping-pong buffer space of the memory core. Simultaneously, index Id2 in the memory core's ping-pong buffer space can be transmitted (MV12) to the processing core's ping-pong buffer space via the broadcast bus. Meanwhile, the processing core performs fusion processing on index Id1 in its local ping-pong buffer space, and the processing result (e.g., Out1) remains stored in the local ping-pong buffer space. The processing may also include loading the value of the data to be fused corresponding to index Id1, for example, this can be done by loading a pre-stored value of the data to be fused from local memory circuitry. The steps of the computational processing are not broken down in the figure.
[0179] In time slice 3, an index (e.g., Id4) is loaded (L4) into the storage core's pong buffer. Simultaneously, the processing result Out1 in the processing core's pong buffer can be moved (MV21) to the storage core's pong buffer, and then the index Id3 in the storage core's pong buffer can be broadcast (MV13) to the processing core's pong buffer. Within the processing core, the index Id2 in the local pong buffer is simultaneously fused, and the processing result (e.g., Out2) is stored in the local pong buffer. The processing also includes operations such as loading the value of the data to be fused corresponding to index Id2.
[0180] In time slice 4, the processing result Out1 in the storage core ping-pong buffer is written back to the external storage circuit, and then an index (e.g., Id5) is loaded into the storage core ping-pong buffer (L5). Simultaneously, the processing result Out2 in the processing core ping-pong buffer can be moved (MV22) to the storage core ping-pong buffer, and then the index Id4 in the storage core ping-pong buffer can be broadcast (MV14) to the processing core ping-pong buffer. Meanwhile, in the processing core, the index Id3 in the local ping-pong buffer is simultaneously merged, and the processing result (e.g., Out3) is stored in the local ping-pong buffer.
[0181] In subsequent time slices, the above pipeline processing can be cycled sequentially to complete the processing of all data. As can be seen from time slice 4 onwards, within the same time slice, five operations in the five-stage pipeline can be performed simultaneously, as well as different operation steps in the five rounds of processing, thereby shortening processing time and improving processing efficiency. It is understood that the above illustrations are merely illustrative, and the data-level pipeline of this disclosed embodiment can be configured with pipelines of different scales according to different situations, thus flexibly adapting to different scenarios.
[0182] The implementation scheme of data fusion processing of the embodiments of this disclosure has been described above in conjunction with a computing device with a multi-core processor architecture. The computing device described above can also be used to implement the sparse convolution operation scheme provided in another aspect of this disclosure. As described above, the sparse convolution operation may include three steps: a MAC step, which calculates the product of the convolution kernel and the sparse input data; an INDEX step, which determines the index of each product result; and a MERGE step, which merges the multiple product results into one fused data stream in index order as the result of the sparse convolution operation.
[0183] In some implementations, the above three steps can be implemented by different tasks or operators.
[0184] Specifically, in some embodiments, the master device can be configured to launch a third task before launching the second task (index pre-sorting). This third task is used to compute indices for data elements in the multiple streams of data to be fused; that is, the third task is used to perform the INDEX step described above. In sparse convolution operations, the aforementioned index corresponds to the index of the multiple product results generated during the convolution operation processing of the sparse input data and the convolution kernel. At this time, the slave device can be configured to perform this third task and store the computed index in a specified location. Typically, the index occupies less space than the last-level cache (LLC) of the computing device; therefore, it is preferable to store the index on the LLC to allow for high-speed transmission using the high bandwidth between the LLC and the storage kernel.
[0185] As described above regarding the INDEX step, this third task involves element-level computation, and therefore can be processed using a typical Load-Compute-Store (LCS) pipeline. To reduce data transfer time, some implementations store the computed index on the last-level cache LLC, allowing the index to be directly loaded from the LLC to the storage core during subsequent fusion processing, so that it can be broadcast to different processing cores for fusion.
[0186] Furthermore, in some embodiments, the master device may be configured to launch a fourth task before launching the first task (fusion processing). This fourth task is used to calculate the values of data elements in the multiple streams of data to be fused; that is, the fourth task is used to perform the MAC steps described above. In sparse convolution operations, these values are the multiplicative results generated during the convolution operation processing of the sparsed input data and the convolution kernel. At this time, the slave device may be configured to perform this fourth task. Due to the limited space in the last-level cache, the calculated multiplicative results are stored in the slave device's external storage circuitry, such as DRAM.
[0187] As described above regarding the MAC step, this fourth task is to calculate the product result, which is equivalent to point-level convolution operations and can also be processed using the LCS pipeline.
[0188] In the above embodiments, since the MAC, INDEX, and MERGE steps are split into different tasks for execution in sparse convolution operations, each can be optimized using multi-stage pipelined processing without interruption in the pipeline process or interference with spatial resources. For example, the MAC and INDEX steps can be processed using a three-stage LCS pipeline, while the MERGE step can be processed using a five-stage LMCMS pipeline.
[0189] This disclosure also provides a chip that may include the computing device of any of the embodiments described above in conjunction with the accompanying drawings. Furthermore, this disclosure also provides a board that may include the aforementioned chip.
[0190] Depending on the application scenario, the electronic devices or apparatus disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablets, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, dashcams, navigators, sensors, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, autonomous driving terminals, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound machines, and / or electrocardiographs. The electronic devices or apparatus disclosed herein can also be applied in fields such as the Internet, IoT, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Furthermore, the electronic devices or apparatus disclosed herein can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as cloud computing, edge computing, and terminal applications. In one or more embodiments, the high-computing-power electronic devices or apparatuses according to the present disclosure can be applied to cloud devices (e.g., cloud servers), while the low-power electronic devices or apparatuses can be applied to terminal devices and / or edge devices (e.g., smartphones or cameras). In one or more embodiments, the hardware information of the cloud devices and the hardware information of the terminal devices and / or edge devices are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud devices to simulate the hardware resources of the terminal devices and / or edge devices based on the hardware information of the terminal devices and / or edge devices, so as to complete the unified management, scheduling and collaborative work of end-to-cloud or cloud-edge-end integration.
[0191] It should be noted that, for the sake of brevity, this disclosure describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art will understand that the solutions disclosed herein are not limited by the order of the described actions. Therefore, based on the disclosure or teachings of this document, those skilled in the art will understand that some steps can be performed in a different order or simultaneously. Furthermore, those skilled in the art will understand that the embodiments described in this disclosure can be considered optional embodiments, that is, the actions or modules involved are not necessarily essential for the implementation of one or more solutions disclosed herein. In addition, depending on the solution, the description of some embodiments in this disclosure may have different emphases. In view of this, those skilled in the art will understand that parts not described in detail in a certain embodiment of this disclosure can also be referred to the relevant descriptions of other embodiments.
[0192] In terms of specific implementation, based on the disclosure and teachings of this document, those skilled in the art will understand that several embodiments disclosed herein can also be implemented in other ways not disclosed herein. For example, regarding the various units in the electronic device or apparatus embodiments described above, this document has divided them based on logical functions, but in actual implementation, there may be other ways of division. As another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. Regarding the connection relationships between different units or components, the connections discussed above in conjunction with the accompanying drawings can be direct or indirect couplings between units or components. In some scenarios, the aforementioned direct or indirect couplings involve communication connections utilizing interfaces, where the communication interface can support electrical, optical, acoustic, magnetic, or other forms of signal transmission.
[0193] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network units. Furthermore, depending on actual needs, some or all of the units can be selected to achieve the purpose of the solution described in the embodiments of this disclosure. Additionally, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically independently.
[0194] In other implementation scenarios, the integrated units described above can also be implemented in hardware, i.e., as specific hardware circuits, which may include digital circuits and / or analog circuits. The physical implementation of the circuit's hardware structure may include, but is not limited to, physical devices, which may include, but are not limited to, transistors or memristors. Therefore, the various devices described herein (e.g., computing devices or other processing devices) can be implemented using appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage units or storage devices can be any suitable storage medium (including magnetic storage media or magneto-optical storage media), such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), ROM, and RAM.
[0195] The embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.
Claims
1. A computing device comprising a master device and a slave device, the slave device comprising a plurality of processing cores, wherein: The master device is configured to launch a first mission, which is used to perform fusion processing. The fusion processing instructs the grouping of data elements from multiple streams of data to be fused into a single ordered fused data stream according to their corresponding indices. Data elements with the same index are merged into a single fused data element. The data element includes any of scalar, vector, or higher-dimensional data. The slave device is configured to schedule a corresponding number of the processing cores to execute the first task according to the splitting strategy of the first task, wherein each processing core performs the fusion processing on a portion of the multi-channel data to be fused.
2. The computing device of claim 1, wherein the master device is further configured to determine a splitting strategy for the first task, the splitting strategy comprising: The number of cores required to complete the first task, and the index range of the data to be fused to be processed corresponding to each core, wherein the index ranges are ordered.
3. The computing device of claim 2, wherein the main device is further configured to: Based on the storage capacity of the slave device and the parameter configuration of the fusion processing, determine the number N of data that a single processing core can process at one time. spacing ;as well as Based on the number of data N spacing The number of core iterations is determined by the total number of data items to be fused from the multiple data streams.
4. The computing device of claim 3, wherein the main device is further configured to: Before launching the first mission, a second mission is launched, the second mission being used to pre-sort the indices of data elements in the multi-channel data to be fused; and Based on the pre-sorted index, and according to the number of data N spacing Sequential splitting is performed to determine the index range of the data to be fused corresponding to each core.
5. The computing device according to any one of claims 2-4, wherein the slave device is further configured to: schedule L rounds, each round containing N, according to the splitting strategy of the first task. core Each processing core performs the first task, wherein in the same round of processing, each processing core performs the fusion processing on the data to be fused on different output channel dimensions.
6. The computing device of claim 5, wherein the slave device further comprises a memory core shared by the plurality of processing cores, and the slave device is further configured to: Load the indexes falling within the index range corresponding to the current processing round from the last-level cache (LLC) of the computing device to the memory core; and The index is broadcast to the scheduled processing cores.
7. The computing device of claim 6, wherein each of the processing cores is configured to: Load the numerical values of the data to be fused, corresponding to the index transmitted in the broadcast mode, on the output channel dimension of this processing core; The fusion process is performed on the numerical values and indexes of the loaded data to be fused; and The fused data is then stored back to the external storage circuit.
8. The computing device according to any one of claims 6-7, wherein the storage core is configured with at least two buffers to support data access between one buffer and external storage circuitry while simultaneously accessing data between the other buffer and the processing core.
9. The computing device of claim 8, wherein each of the processing cores is configured with a local storage circuit, the local storage circuit being configured with at least two storage areas to support data access between one storage area and the storage core or an external storage circuit, while simultaneously performing data access between the other storage area and the processing core.
10. The computing device according to claim 4, wherein: The master device is further configured to: before launching the second task, launch a third task, the third task being used to calculate the index of data elements in the multi-path data to be fused, wherein the index is the index corresponding to the multi-path product result generated during the convolution operation processing of the sparse input data and the convolution kernel; and The slave device is further configured to: perform the third task and store the calculated index on the last-level cache (LLC) of the computing device.
11. The computing device according to any one of claims 1-4, 6-7 or 9, wherein: The master device is further configured to: before launching the first task, launch a fourth task, the fourth task being used to calculate the values of data elements in the multi-path data to be fused, wherein the values are multi-path product results generated during the convolution operation processing of the sparse input data and the convolution kernel; and The slave device is further configured to: perform the fourth task and store the calculated multi-way product result in the external storage circuit of the slave device.
12. A chip comprising a computing device according to any one of claims 1-11.
13. A circuit board comprising the chip according to claim 12.
14. A method for processing data using a computing device according to any one of claims 1-11.
Citation Information
Patent Citations
Calculation engine and electronic equipment
CN106126481A
Neural network processing method and device, computer device and storage medium
CN110674936A