Data processing circuit, data processing method and related product
Patent Information
- Application Number
- CN202111642096.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-12-29
AI Technical Summary
因此,传统的适合于密集型数据的卷积神经网络在应用于这种稀疏型数据时,效率将变得非常低,尤其是涉及卷积运算时,会在零值数据点上浪费大量的算力等资源
[0013]通过如上所提供的数据处理电路、使用数据处理电路来处理数据的方法、芯片和板卡,本披露实施例提供了一种适合于稀疏型数据的卷积方案,其通过仅将非零/非空数据与卷积核进行运算,可以极大节省运算量,提高处理效率。本披露实施例提供的稀疏卷积方案可以适用于多维卷积运算,包括但不限于二维卷积和三维卷积,由此可以适用于LiDAR点云数据的处理。
Smart Images

Figure CN114329324B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of data processing. More specifically, this disclosure relates to data processing circuits, data processing methods, chips, and circuit boards. Background Technology
[0002] In recent years, significant progress has been made in object detection, instance segmentation, and keypoint detection based on convolutional neural networks. These detection methods are typically based on LiDAR or RGB-D data and can be applied to fields such as autonomous driving and robot vision.
[0003] Unlike dense image data, LiDAR point cloud data is typically sparse, and the point density varies dramatically due to factors such as non-uniform sampling in 3D space, the effective range of the sensor, occlusion, and relative pose. Therefore, traditional convolutional neural networks, which are well-suited for dense data, become very inefficient when applied to this type of sparse data, especially when convolution operations are involved, wasting significant computational resources on zero-value data points.
[0004] Therefore, it is desirable to provide an improved convolution scheme that is suitable for sparse data such as point cloud data, thereby improving processing efficiency. Summary of the Invention
[0005] In order to at least partially solve one or more of the technical problems mentioned in the background art, the present disclosure provides a data processing circuit, a data processing method, a chip, and a board.
[0006] In a first aspect, this disclosure discloses a data processing circuit, comprising: a control circuit, a storage circuit, and an arithmetic circuit, wherein:
[0007] The control circuit is configured to control the storage circuit and the computing circuit to perform N-dimensional convolution operations on the input data and the convolution kernel, where N>1, and N represents the number of convolution dimensions for sliding accumulation in the convolution operation, wherein the input data is sparse data and is represented in a dense form;
[0008] The storage circuitry is configured to store information, which includes at least pre-processing, during-processing, and / or post-processing information; and
[0009] The computation circuit is configured to perform multiple one-dimensional convolution operations on the input data and the convolution kernel under the control of the control circuit, to obtain multiple operation results and the corresponding output point coordinates on the first convolution dimension; and to merge the multiple operation results into one fused data according to their corresponding output point coordinates as the result of the convolution operation, wherein operation results with the same output point coordinates are accumulated.
[0010] In a second aspect, this disclosure provides a chip that includes the data processing circuitry of any of the embodiments of the first aspect.
[0011] In a third aspect, this disclosure provides a board including the chip of any of the embodiments of the second aspect above.
[0012] In a fourth aspect, this disclosure provides a method for processing data using the aforementioned data processing circuitry.
[0013] Through the data processing circuit, the method for processing data using the data processing circuit, the chip, and the board provided above, this disclosure provides a convolution scheme suitable for sparse data. By only performing operations with non-zero / non-empty data and the convolution kernel, it can greatly save computational load and improve processing efficiency. The sparse convolution scheme provided in this disclosure is applicable to multi-dimensional convolution operations, including but not limited to two-dimensional and three-dimensional convolution, and therefore can be applied to the processing of LiDAR point cloud data. Attached Figure Description
[0014] The above and other objects, features, and advantages of exemplary embodiments of this disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding portions wherein:
[0015] Figure 1 This diagram shows the structure of the board card according to an embodiment of this disclosure;
[0016] Figure 2 This diagram illustrates the structure of the combined processing apparatus according to an embodiment of the present disclosure.
[0017] Figure 3a This diagram illustrates the internal structure of a single-core computing device according to an embodiment of the present disclosure.
[0018] Figure 3b This diagram illustrates the internal structure of a multi-core computing device according to an embodiment of the present disclosure.
[0019] Figure 4a This illustrates the operational principle of a conventional convolution scheme;
[0020] Figure 4b This illustrates an exemplary principle of a sparse convolution operation scheme according to embodiments of this disclosure;
[0021] Figure 5 This illustrates an exemplary representation of input data according to an embodiment of this disclosure;
[0022] Figure 6 An exemplary process for a sparse convolution scheme according to embodiments of this disclosure is shown;
[0023] Figure 7 An example of splitting an input data block according to an embodiment of this disclosure is shown;
[0024] Figure 8 A flowchart illustrating an exemplary method for filtering valid input data points according to embodiments of this disclosure is shown.
[0025] Figure 9 This diagram illustrates a scanning traversal of the third input parameter according to an embodiment of this disclosure;
[0026] Figure 10 A schematic diagram illustrating the construction of the Q matrix according to an embodiment of this disclosure is shown;
[0027] Figure 11 This illustrates exemplary logic for calculating the wo coordinates of the output data according to embodiments of this disclosure; and
[0028] Figure 12 A schematic diagram of the structure of a data processing circuit according to an embodiment of this disclosure is shown. Detailed Implementation
[0029] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0030] It should be understood that the terms "first," "second," "third," and "fourth," etc., in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0031] It should also be understood that the terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0032] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection."
[0033] The specific embodiments disclosed herein will now be described in detail with reference to the accompanying drawings.
[0034] Exemplary hardware environment
[0035] Figure 1 A schematic diagram of the structure of a board 10 according to an embodiment of this disclosure is shown. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 10 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities.
[0036] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. The calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface forms, such as a PCIe interface.
[0037] The board 10 also includes a storage device 104 for storing data, which includes one or more memory cells 105. The storage device 104 is connected to and transmits data with the controller 106 and the chip 101 via a bus. The controller 106 in the board 10 is configured to regulate the state of the chip 101. Therefore, in one application scenario, the controller 106 may include a microcontroller (MCU).
[0038] Figure 2 This is a structural diagram illustrating the combined processing device in chip 101 of this embodiment. (As shown) Figure 2 As shown, the combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203, and a storage device 204.
[0039] The computing device 201 is configured to perform user-specified operations. It is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.
[0040] Interface device 202 is used to transmit data and control commands between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to on-chip storage device of computing device 201. Further, computing device 201 can obtain control commands from processing device 203 via interface device 202 and write them to on-chip control cache of computing device 201. Alternatively or optionally, interface device 202 can also read data from storage device of computing device 201 and transmit it to processing device 203.
[0041] Processing device 203, as a general-purpose processing device, performs basic controls including but not limited to data transfer and starting / stopping computing device 201. Depending on the implementation, processing device 203 may be one or more types of processors, such as a central processing unit (CPU), graphics processing unit (GPU), or other general-purpose and / or special-purpose processors. These processors include, but are not limited to, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, computing device 201 disclosed herein can be considered as having a single-core structure or a homogeneous multi-core structure. However, when computing device 201 and processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure.
[0042] Storage device 204 is used to store data to be processed. It may be DRAM or DDR memory, typically 16G or larger in size, and is used to store data of computing device 201 and / or processing device 203.
[0043] Figure 3aThe diagram shows the internal structure of the processing core when the computing device 201 is a single-core device. The computing device 301 is used to process input data such as computer vision, speech, natural language, and data mining. The computing device 301 includes three main modules: a control module 31, an arithmetic module 32, and a storage module 33.
[0044] The control module 31 coordinates and controls the operation of the computation module 32 and the storage module 33 to complete the deep learning task. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 fetches instructions from the processing device 203, and the instruction decode unit 312 decodes the fetched instructions and sends the decoding result as control information to the computation module 32 and the storage module 33.
[0045] The computation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformations; the matrix operation unit 322 is responsible for the core computations of deep learning algorithms, namely matrix multiplication and convolution.
[0046] The storage module 33 is used to store or move relevant data, including a neuron RAM (NRAM) 331, a weight RAM (WRAM) 332, and a direct memory access (DMA) module 333. The NRAM 331 stores the input neurons, output neurons, and intermediate results after computation; the WRAM 332 stores the convolution kernels of the deep learning network, i.e., the weights; the DMA 333 is connected to the DRAM 204 via bus 34 and is responsible for data transfer between the computing device 301 and the DRAM 204.
[0047] Figure 3b A simplified schematic diagram of the internal structure of the computing device 201 as a multi-core processor is shown. The multi-core computing device can be abstracted using a hierarchical hardware model. As shown, the multi-core computing device can be abstracted into four levels: Card level 350, Chip level 360, Cluster level 370, and Core level 380. This disclosure primarily concerns the data transmission and computing unit portions of the storage unit; therefore, the accompanying drawings and description briefly illustrate and introduce the relevant computing structure, omitting other parts.
[0048] At the board level, each board contains local DDR memory, and each processor chip serves as a computing and control unit.
[0049] At the chip level, each processor chip contains multiple multiprocessors as computing units.
[0050] At the computing cluster level, each multiprocessor includes multiple accelerator cores as control and computing units, as well as shared SRAM as storage units.
[0051] At the processor core level, each accelerator core contains local memory and an array of local processing units. NFU stands for Neuron Function Unit, used for performing convolution calculations.
[0052] In this multi-core computing device, the storage model includes global memory on the board, SRAM (shared memory) on the cluster, NRAM, WRAM, and registers on the core. To achieve better performance, data movement between storage levels below the card and the balance between memory access and computation can be explicitly controlled. SRAM is contained within the Memory Processing Unit (MPU, or Mem Core). A Core refers to the Intelligent Processing Unit (IPU Core, or Core) in a multi-core computing device. One IPU Core contains NRAM, WRAM, NFU, etc. A Cluster refers to a cluster of processors or computing clusters; typically, a multi-core computing device contains several Clusters, and one Cluster contains one Mem Core + N IPU Cores.
[0053] Exemplary Convolution Operation Principle
[0054] Based on the aforementioned hardware environment, the embodiments disclosed herein provide a data processing circuit that supports convolution operations on sparse data. By providing an optimized convolution scheme, convolution processing related to sparse data such as LiDAR point cloud data can be simplified and accelerated. The sparse convolution scheme provided in the embodiments of this disclosure is applicable to multidimensional convolution operations, including but not limited to two-dimensional convolution and three-dimensional convolution. For simplicity and ease of understanding, two-dimensional convolution is used as an example in some embodiments.
[0055] In this disclosure, "N-dimensional convolution" refers to the number of convolutional dimensions in which sliding accumulation is performed. For example, when N=2, the convolutional kernel performs translational accumulation along two dimensions (e.g., width W and height H) according to the corresponding convolutional stride. When N=3, the convolutional kernel performs translational accumulation along three dimensions (e.g., width W, height H, and depth D) according to the corresponding convolutional stride. In this disclosure, "non-convolutional dimension" refers to a dimension in which the convolutional kernel does not perform sliding accumulation. Different non-convolutional dimensions may have different operational requirements. For example, for regular convolution, the input channel dimension Ci is required to be accumulated, while the output channel dimension Co is not accumulated; similarly, for depth-wise convolution, the input channel dimension Ci is also not accumulated.
[0056] To better understand the convolution scheme of this disclosure embodiment, the operational principle of a conventional convolution scheme will be described first using two-dimensional convolution as an example.
[0057] Figure 4a This illustrates the operational principle of a conventional convolution scheme. In this example, the convolution kernel 410 is dense, a 3×3 matrix, with the numbers in the kernel representing the corresponding weights. The input data 420 is a 7×7 matrix that is sparse, containing only four non-zero values: 2, 3, 5, and 6, as shown by the dark squares. In this exemplary convolution process, the stride for both dimensions is set to 2, with zero padding (no dilation). The 3×3 gray squares in the diagram represent the sliding accumulation process of the convolution kernel across the input data. 430 shows the initial calculation at the start of the convolution, 440 shows the calculation after sliding one step to the right (stride 2), and 450 shows the calculation after sliding one step down (stride 2). In each step, the weights of the convolution kernel are multiplied bitwise with the input data and accumulated. 460 shows the final calculation result as the output data. The output data is a 3×3 matrix. It can be seen that the calculation of 430 corresponds to the data at coordinate (0,0) in the output data, the calculation of 440 corresponds to the data at coordinate (0,1) in the output data, and the calculation of 450 corresponds to the data at coordinate (1,0) in the output data.
[0058] In the sparse convolution operation of this disclosed embodiment, the convolution kernel is dense, and its input format can be the same as that of conventional convolution; while the input data is sparse, and its input format can be different from that of conventional convolution input data, thereby saving storage space.
[0059] from Figure 4a As can be seen from the description, the final result of sparse convolution operation is only related to the operation result of non-zero input data elements. Therefore, the multiplication and addition operation with the convolution kernel can be performed only on these non-zero input data elements, thereby reducing invalid operations.
[0060] Furthermore, from Figure 4a As can be seen from the output data 460, these multiplication and addition operations can be further broken down into partial operations performed by row-wise one-dimensional convolution and the result of column-wise positional summation. In other words, N-dimensional convolution operations can be broken down into multiple one-dimensional convolution operations.
[0061] Figure 4b An exemplary principle of a sparse convolution scheme according to an embodiment of this disclosure is shown. Figure 4b Still with Figure 4a The sparse convolution operation scheme of this disclosure embodiment is described using data as an example.
[0062] like Figure 4b As shown, the entire convolution operation can be broken down into calculating the convolution result row by row. In this example, one row of the result 460 corresponds to three rows of input data 420. For example, the first three rows of input data can be used to calculate the first row of output data, the middle three rows (which overlap with the first three and last three rows) of input data can be used to calculate the second row of output data, and the last three rows of input data can be used to calculate the last row of output data.
[0063] As can be seen from the breakdown of the above calculation process, sparse data can be filtered at a granularity (referred to here as the "first filtering granularity") based on the input data required for a corresponding row of calculation results (three rows of input data in the example in the diagram). This reduces invalid calculations. For example, in Figure 4b In the example, the middle three rows of input data are all 0s, meaning there are no non-sparse points, so the convolution operation of the next three rows can be omitted.
[0064] Furthermore, during the calculation of each row of output results, the two-dimensional (H and W dimensions) convolution operation can be further divided into three one-dimensional (W dimension) convolution operations.
[0065] As shown in the figure, the original method of using a 3*3 convolution window (shown in the black box) to slide and calculate over three rows of input data to obtain one row of output data can be transformed into using three 1*3 convolution windows (shown in the dashed boxes) to slide and calculate over three rows of input data respectively (as shown in 470), obtaining three rows of partial sums (as shown in 480), and then accumulating them bit by bit to obtain the final row of output data (460).
[0066] As can be seen from the above decomposition of convolution operations, by breaking down the two-dimensional convolution operation into multiple one-dimensional convolution operations, sparse data can be filtered at a granularity of one dimension (referred to here as the "second filtering granularity") during the convolution operation, thereby reducing invalid computations. For example, in the example in Figure 4b, the third row of input data is all 0s, that is, there are no non-sparse points, so the convolution operation in this row can be omitted.
[0067] Furthermore, within each one-dimensional convolution operation, sparse data can be filtered at the granularity of a one-dimensional convolution window (referred to here as the "third filtering granularity"), thereby further reducing invalid computations. For example, in Figure 4b In the example, there are no non-sparse points in the last two one-dimensional convolution windows of the first row of input data, so the operation of these two convolution windows can be omitted; and there are no non-sparse points in the last one-dimensional convolution window of the second row of input data, so the operation of this convolution window can also be omitted.
[0068] Therefore, by using three different granularity levels, sparse data can be filtered quickly and effectively, minimizing invalid computations.
[0069] Accordingly, in the sparse convolution scheme of this disclosure embodiment, the sparse convolution operation may include the following steps: performing multiple one-dimensional convolution operations on the convolution kernel and sparse input data to obtain multiple operation results (e.g., product results or multiply-add results, considering that some dimensions of the data also need to be accumulated, such as the input channel dimension Ci) and the corresponding output point coordinates on the first convolution dimension; then merging the multiple operation results into one fused data according to their corresponding output point coordinates, as the result of the sparse convolution operation. During the merging process, operation results with the same output point coordinates are accumulated.
[0070] It is understood that although the above example of two-dimensional convolution illustrates the exemplary principle of the sparse convolution operation scheme of this disclosure embodiment, the above scheme can also be applied to three-dimensional convolution operations or even higher-dimensional convolution operations.
[0071] Based on the above principle, when applied to N-dimensional convolution operations (N>1), the N-dimensional convolution operation can be decomposed into M one-dimensional convolution operations, where M equals the product of the sizes of the convolution kernel in the N-1 convolution dimensions excluding the first convolution dimension containing the one-dimensional convolution operation. For example, in the two-dimensional convolution example above, the convolution kernel size is kx*ky, and the first convolution dimension is W(kx), then M = ky; while in the three-dimensional convolution example, the convolution kernel size is kx*ky*kz, and the first convolution dimension is W(kx), then M = ky*kz.
[0072] The computation process of implementing the sparse convolution scheme of this disclosure embodiment is described in detail below.
[0073] Example structure of input and output data
[0074] In sparse convolution operations, input data, convolution kernel, and output data are involved. The input data is also called input neurons in neural networks, and the output data is called output neurons.
[0075] In convolution operations involving LiDAR point cloud data, the convolution kernel is dense, and its input format can be the same as that of regular convolution. In 2D convolution, the kernel size is typically 3*3, and a single convolution requires summing 3*3*ci values. In 3D convolution, the kernel size is typically 3*3*3, and a single convolution requires summing 3*3*3*ci values. The stride of the convolution is typically 2 in each dimension, for example, S. H =S W =S D =2.
[0076] The input data to be processed by convolution can include multidimensional data, and it is sparse across multiple dimensions. For example, in target detection based on LiDAR data, the input data is detection data in three-dimensional space, which may represent the grayscale value, RGB values, signal strength, etc., of each three-dimensional spatial coordinate point. Therefore, depending on the information content to be represented, the input data element at each coordinate point can be one-dimensional, two-dimensional, three-dimensional, or higher-dimensional data. Due to the characteristics of point cloud data, coordinate points with non-zero value data elements are sparse, that is, they are sparse in three spatial dimensions (e.g., width W, height H, and depth D).
[0077] Depending on the initial state of the input data, preprocessing can be performed before the sparse input data is provided to the computing circuit for computation. In some embodiments, such preprocessing may include, for example, merging multiple sparse dimensions into one dimension; densifying the sparse data points in the input data along the merged dimension; and using several input parameters to represent the numerical values and coordinates of the densified input data.
[0078] Figure 5 An exemplary representation of input data (sparse neurons) is shown. The input data here can be data that has been padded according to the requirements of the convolution operation. In this example, a preprocessing operator can be used to transform the sparse input data into a dense input data representation, thereby saving storage space.
[0079] As shown in the figure, the sparse input data 510 comprises five dimensions: the batch (B) dimension, the HWD 3D spatial dimension, and the input channel Ci dimension. The input data is sparse in the B dimension and the HWD 3D space. In the figure, the dark squares in the HWD 3D matrix represent locations with values (called valid input points), while all other parts are zero values. Multiple such HWD 3D matrices exist in the B dimension, and the sparsity pattern (i.e., the positions of the dark squares) can be different in each matrix. The input data is dense in the input channel Ci dimension, which is the lowest dimension. Due to the limited representation capabilities of the attached figure, only four dimensions are shown in Figure 510, but the Ci dimension can be understood as the thickness of each dark square. The size of the Ci dimension is uniform, meaning that the thickness of each dark square is the same.
[0080] In some embodiments, when converting sparse input data 510 into dense input data, it can be represented by the CSR format in the sparse matrix.
[0081] In the storage of sparse matrices, to achieve compression, only the non-zero element values (sometimes called valid element values) are stored. However, the positions of the non-zero elements must also be preserved for easy recovery. Therefore, the storage of sparse matrices not only stores the non-zero element values but also their coordinate positions (row index and column index).
[0082] CSR (Compressed Sparse Row Format) storage uses three arrays to store the sparse matrix: row pointers, column indices, and values. The lengths of the column index array and the value array are the number of non-zero elements in the sparse matrix. The row pointer array stores the offset of the first non-zero element in each row from the first non-zero element in the sparse matrix, and its last element stores the total number of non-zero elements in the sparse matrix. Therefore, the length of the row pointer array is the number of rows in the sparse matrix plus one. In essence, according to the definition of a row pointer array, subtracting the pointer value of the previous row from the pointer value of the next row gives the number of non-zero elements in the previous row.
[0083] Similarly, in this disclosed embodiment, when converting sparse input data 510 into dense input data, three input parameters can be used to represent it.
[0084] The first input parameter is the effective dense data, that is, the compactly arranged sparse data, denoted by Min. The shape of Min is ain*ci, where ain is the number of non-sparse points in the input data and ci is the size of the input channel dimension.
[0085] During preprocessing, the four sparse dimensions of the input data (B dimension and HWD 3D spatial dimension) can be merged into a single dimension ain. The sparse data points (darker squares in the diagram) are then compacted along this merged dimension, resulting in denser input data elements. That is, each HWD 3D matrix along the B dimension undergoes the same dimension merging and compaction process, yielding the preprocessed, denser form of the input data 521, a two-dimensional matrix with the lower dimension Ci and the higher dimension being the merged BHWD dimension ain. In this example, ain = 4.
[0086] In some embodiments, for cases with multiple batches, a batch-by-batch processing method can be adopted, splitting the batches across different processor cores (e.g., ...). Figure 3b Processing is done on the core (of the process). For example, SWIFT has 12 batches, so based on... Figure 3b The example hardware environment can process up to 12 batches in parallel at a time. In these embodiments, for each batch, it is only necessary to merge the valid input points of the three sparse dimensions HWD, at which point ain corresponds to the number of valid input points of the three dimensions HWD.
[0087] The second input parameter is the coordinate or index of each valid input point in the W dimension, denoted by wi_coord, and its shape is 1*ain. As shown in Figure 522, for the four valid input points in the figure, wi_coord is [1,2,0,6].
[0088] The third input parameter is the input data in CSR (Compressed Sparse Row) format along the H dimension or H direction, denoted by hin. hin stores the offset of the first non-zero element in each row from the first non-zero element in the input data, and its last element stores the total number of non-zero elements in the input data. Depending on the number of dimensions of the input data, the third input parameter may have multiple dimensions. For example, when the input data is... Figure 5 In the example of five-dimensional data, the third input parameters from high dimension to low dimension are: batch(B), depth in_d, and height in_h, with a shape of B*Din*(Hin+1), where B, Din, and Hin are the size of the B, D, and H dimensions of the sparse input data, respectively.
[0089] Without loss of generality, the shape of the third input parameter hin can be represented as X*(Hin+1), where X represents the product of the sizes of the other dimensions of the input data besides H, W, and Ci. For example, when the input data is three-dimensional data including H, W, and Ci, X does not exist or X = 1; while when the input data is four-dimensional data including D, H, W, and Ci, X = D.
[0090] As shown in Figure 523, for B=0 and Din=0, hin in the example is [0,1,2,2,2,2,2,4]. Specifically, the first element is "0", which means that the first non-zero element in the row hi=0 is offset by 0 from the first non-zero element in the input data, since it is the first non-zero element itself. The second element "1" means that the first non-zero element in the row hi=1 is offset by 1 from the first non-zero element in the input data, which is equal to the number of non-zero elements in the row hi=0. The third element "2" means that the first non-zero element (if any) in the row hi=2 is offset by 2 from the first non-zero element in the input data, which is equal to the number of non-zero elements in the first two rows hi=0 and hi=1. The fourth element "2" means that the first non-zero element (if any) in the row hi=3 is offset by 2 from the first non-zero element in the input data, which is equal to the number of non-zero elements hi=0 to hi=2 in the first three rows. And so on, with the last element "4" representing the total number of all non-zero elements.
[0091] Similarly, in this disclosed embodiment, since the output data of the convolution operation is also sparse, three output parameters can also be used to represent the output data.
[0092] The first output parameter is the dense data of the effective output, that is, the compactly arranged sparse data, denoted by Mout. The shape of Mout is aout*co, where aout is the number of non-sparse points in the output data and co is the size of the output channel dimension.
[0093] The second output parameter is the coordinates or index of each valid output point in the W dimension, denoted by wo_coord, and its shape is 1*aout.
[0094] The third output parameter is the CSR (Compressed Sparse Line) format of the output data in the H direction, denoted by `hout`. `hout` stores the offset of the first non-zero element in each line of the output data from the first non-zero element in the output data, and its last element stores the total number of non-zero elements in the output data. Similar to the third input parameter, the third output parameter may also have multiple dimensions, for example, from high to low dimensions: batch(B), depth `out_d`, and height `out_h`, with a shape of B*Dout*(Hout+1), where B, Dout, and Hout are the B, D, and H dimensions of the sparse output data, respectively.
[0095] For example, continue Figure 5In the example, assuming the convolution kernel is 3*3 and the convolution stride is 2, the output data includes 4 valid output points, whose corresponding W-dimensional coordinates wo_coord 532 are [0,1,0,2], and the corresponding H-direction CSR format data hout 533 are [0,2,2,4]. The first output parameter is not shown in the figure.
[0096] Exemplary sparse convolution operation process
[0097] As can be seen from the data structures of the input and output data, given the rules of convolution (e.g., kernel size, stride), the coordinates of valid output points in the output data can be determined based on the coordinates of the input data. Therefore, in the sparse convolution operation scheme of this disclosed embodiment, the operation process can be broken down into several steps: coordinate calculation and numerical calculation.
[0098] Figure 6 An exemplary process for a sparse convolution scheme according to an embodiment of this disclosure is shown.
[0099] As shown in the figure, in step 610, the data that each processor core needs to process each time is first filtered based on the coordinates of the input data. The input data here is CSR format data that has already undergone zero-padding and sparse-to-dense transformation.
[0100] Based on Figure 3b In the case of multi-core computing devices shown to perform sparse convolution operations, given the memory (e.g. Figure 3b The SRAM in the database has limited storage space, necessitating data splitting. In some embodiments, splitting can be performed according to the dimensions of the output data, as described above. Figure 4b As described in the convolution operation principle, the result of the convolution operation can be calculated row by row. Therefore, the splitting method can be: each processor core calculates one row of output data points in dimension Wo at a time. Correspondingly, the shape of the input data block corresponding to one row of output data points in dimension Wo is (kz*ci)*ky*wi, where wi is the size of the W dimension of the input data, ci is the size of the Ci dimension of the input data, and the sizes of the convolution kernel in the W, H, and D dimensions are kx, ky, and kz, respectively.
[0101] Figure 7 An example of splitting the input data block is shown, where the gray input data block, after undergoing convolution and summation, corresponds exactly to a row of output data points in dimension Wo.
[0102] Considering the sparsity of the data, in order to extract the valid input data points corresponding to the output data points of a row of Wo dimension, some embodiments can use the third input parameter of the input data, namely hin, for filtering. As hin means, subtracting the ith value from the (i+1)th value yields the number of non-zero elements (valid input data points) in the ith row. Therefore, this characteristic of hin can be used to determine whether valid input data points exist, thereby performing filtering.
[0103] Figure 8 A flowchart illustrating an exemplary method for filtering valid input data points according to an embodiment of this disclosure is shown.
[0104] As shown in the figure, in step 811, the third input parameter is first loaded from the external storage circuit onto the on-chip memory (e.g., SRAM). The storage space required for the third input parameter is (Hin+1+2*ph)*(Din+2*pd)*dwidth, where Din and Hin are the D-dimensional and H-dimensional dimensions of the sparse input data, respectively, ph and pd are the padding amounts on one side of the H-dimensional and D-dimensional dimensions, respectively, and dwidth is the data bit width. The shape of the third input parameter can be a two-dimensional matrix (Din+2*pd)*(Hin+1+2*ph). Here, it is assumed that batch B=1 because processing is performed batch by batch.
[0105] Next, at step 812, the third input parameter is traversed with a specified scan window (first scan window) and a specified scan step size (first scan step size) to find valid input data points.
[0106] The size of the first scanning window corresponds to the input data required to calculate the output data points of a row in dimension Wo, which is also the "first filtering granularity" mentioned earlier. The size of the first scanning window is determined by the size of the convolution kernel. Specifically, the size of the first scanning window is kz*(ky+1), where kz corresponds to the size of the scanning window in dimension D, and ky+1 is the size of the scanning window in dimension H. This is because the third input parameter uses CSR format, where ky rows of data in dimension H require ky+1 data points to represent. The first scanning stride is equal to the convolution stride Sy in dimension H.
[0107] During the scanning traversal, the system checks if each row of data in the third input parameter has changed. If any row changes, it indicates the existence of a valid input data point. The scanning window that detects a valid input data point can be called a range block. After detecting a range block, the corresponding hi and di coordinates can be recorded. From these hi and di coordinates, the corresponding ho and do coordinates in the output data can be calculated.
[0108] When N is continuously detected IPUAfter finding the N blocks, in step 813, the found blocks can be sent to the processor core (IPU), for example, these N blocks... IPU Each block is sent to N. IPU There are several different IPUs. Each IPU can be informed of the H and D dimension coordinates (ho and do coordinates) of its processed output point in the output data. It can be understood that when each IPU calculates the value of the output point (wo_x,ho,do), wo_x varies, ranging from 0 to (wo-1), while ho and do are fixed.
[0109] Figure 9 The diagram illustrates the scanning process for the third input parameter. It's important to understand that the diagram only shows a portion of the third input parameter, such as the first three rows of hin data (910) from di = 0 to 2. In this example, the scanning window size is 3*4, the scanning step size is 2, and the scan proceeds sequentially along the hin direction. It can be seen that in the first scanning window (901) corresponding to the output point coordinate ho = 0, all three rows of hin data have changed, indicating the presence of valid input data points. Therefore, scanning window 901 is recorded as block 0, and its corresponding ho = 0 and do = 0 can be calculated. Next, moving two rows to the right, in this scanning window (902), all three rows of hin data remain unchanged, indicating the absence of valid input data points. No further processing is needed, and the next scan can continue. In this way, four consecutive blocks can be detected and sent to four different IPUs, i.e., N in this example. IPU =4. The do coordinates of these 4 blocks are all 0, and the ho coordinates are 0, 2, 3, and 4 respectively.
[0110] continue Figure 6 In step 620, each IPU assigned data to be processed (indicated by the aforementioned blocks) can retrieve the corresponding data according to the block's indication and construct the matrix to be convolved, hereinafter referred to as the Q matrix, while simultaneously calculating the output point coordinate information. For example, the Q matrix can be constructed by retrieving data from a shared memory SRAM that has been pre-loaded from an external storage circuit.
[0111] In some embodiments, the wi_coord vector of valid input data points in the input data block corresponding to the allocated block can be extracted first from the second input parameter wi_coord. Then, based on the wi_coord vector, the input data corresponding to the vector can be traversed with a second scanning window and a second scanning step size to extract the corresponding input data points from the valid input dense data Min (i.e., the first input parameter) to construct the Q matrix.
[0112] As mentioned earlier, the block is derived from the third input parameter hin, which records the distance from the first valid input data point in each row to the specified point (i.e., the first valid input data point in the entire input data). Therefore, based on this information, a specified number of data points can be extracted from the specified position of the second input parameter wi_coord to form the corresponding wi_coord vector. Here, the wi_coord vector refers to the vector composed of the wi coordinates of all valid input data points in a row of W dimension in the input data. For example, assuming there are 34 valid input data points in a row of W dimension, the vector length is 34, and the value of each vector element is the wi coordinate of the corresponding data point.
[0113] The number of wi_coord vectors taken is kz (D dimension) * ky (H dimension), which corresponds to the data range indicated by the block, and also to the number M of the transformed one-dimensional convolution operations described above. Thus, these M wi_coord vectors can cover the range of a convolution window in the H and D dimensions, while taking the entire wi dimension. For example, in the previous example, kz = ky = 3, so 9 wi_coord vectors are taken. The taken kz * ky wi_coord vectors can be concatenated to construct the matrix to be convolved.
[0114] Similarly, based on the meaning of hin, the wi_coord vector can be filtered during its construction. Since the two elements in hin represent the number of valid input data points in the corresponding row, the wi_coord vector is empty for rows with a difference of 0, thus eliminating empty wi_coord vectors. This filtering step corresponds to the "second filtering granularity" described earlier.
[0115] Next, in some embodiments, the input data corresponding to the extracted wi_coord vector can be traversed with a second scanning window and a second scanning step size to extract the corresponding input data points from the first input parameter Min to construct matrix Q.
[0116] Figure 10 A schematic diagram illustrating the construction of a Q matrix according to an embodiment of this disclosure is shown. For simplicity, only the construction of the Q matrix with 3 rows w is shown in the figure; the other rows w can be constructed similarly.
[0117] As shown in the figure, for the sake of clarity, 1010 shows the sparse form of the input data, and also shows the extracted wi_coord vector 1020 (corresponding to the three rows of di=0, hi=4~6) and the part hin 1030 of di=0 in the block allocated to the current IPU. 1040 shows the Q matrix constructed based on the three rows of w of di=0, hi=4~6.
[0118] As mentioned earlier, the wi_coord vector can be constructed based on the information in the third input parameter hin, which determines whether there are valid input data points in the current row to be scanned. Specifically, the number of valid input data points in the i-th row can be determined based on the difference between the (i+1)th value and the i-th value in hin.
[0119] For example, in the example shown in the diagram, based on the information in block (1030): 4-2=2, it can be determined that there are 2 valid input data points in the row hi=4. Based on the information in block 1030: 4-4=0, it can be determined that there are no valid input data points in the row hi=5. Based on 7-4=3, it can be determined that there are 3 valid input data points in the row hi=6. Therefore, only the rows hi=4 and hi=6 need to be scanned. Thus, in this example, since the data in the row hi=5 is empty, there is no corresponding wi_coord vector. According to the starting position and number indicated by hin 1030, the corresponding data can be extracted from wi_coord to construct the wi_coord vector. As shown in diagram 1020, there are only 2 wi_coord vectors.
[0120] Next, based on the extracted wi_coord vector, the input data corresponding to the wi_coord vector is traversed with the second scanning window and the second scanning step size to extract the corresponding valid input data points and construct matrix Q.
[0121] Specifically, the rows of the Q matrix are constructed by scanning row by row. In the example shown in the figure, the row hi=4 is scanned first, then hi=5 is skipped, and the row hi=6 is scanned. During the scan, the data covered by the second scan window that detects valid input data points is extracted and sequentially tiled to form the corresponding rows of the Q matrix, while the second scan window that does not detect valid input data points is skipped. It can be understood that because it is a row-by-row sliding scan, the size of the second scan window corresponds to the size of the convolution window in the first convolution dimension (e.g., the size of the W dimension, kx), and the second scan stride corresponds to the convolution stride Sx in the W dimension. This scanning and filtering step corresponds to the "third filtering granularity" described above.
[0122] During scanning, a second scanning window (1x3 in this example) and a convolution stride Sx along the W dimension (Sx = 2 in this example) are used to scan and traverse the input data row by row along the W dimension. When a valid input data point exists within the scanning window, the input data corresponding to that scanning window is extracted. In this way, the extracted window data is sequentially unfolded and tiled along the W dimension to construct the Q matrix.
[0123] like Figure 10As shown, we can first scan the row of data where hi=4, as shown in the first scan window 1001. If a valid input data point is found in window 1001, the data from this window is extracted to form the first three columns of the first row of the Q matrix 1040. Next, we shift two data points to the right and scan the next scan window 1002. This window also contains valid input data points, and similarly, the data from this window is extracted to form the next three columns of the first row of the Q matrix. We continue shifting two data points to the right and scan the next scan window 1003. This window also contains valid input data points, and similarly, the data from this window is extracted to form the last three columns of the first row of the Q matrix. The scanning of this row of data is now complete. During the scan, if no valid input data point is found in a scan window, we jump to the next window.
[0124] Next, skip the row where hi=5 and scan the row where hi=6. Similarly, use a 1*3 scanning window with a step size of 2 data points. The scanning results are: there are 2 valid input data points in scanning window 1004, 2 valid input data points in scanning window 1005, and no valid input data points in scanning window 1006. The result of extracting the data and constructing the Q matrix is shown in the second row of 1040, consisting of the data covered by scanning windows 1004 and 1005.
[0125] Since the input data is in a dense form, when determining whether there are valid input data points within a scanning window, the coordinate information of the input data can be used to determine whether it falls within a certain scanning window. Specifically, this can be determined based on the block (1030) allocated to the IPU and the constructed wi_coord vector (1020). It can be understood that a scanning window essentially corresponds to a partial sum of output points; therefore, the wo coordinates of the output data points contributed by the valid input data points can be inferred from the wi coordinates, thereby determining whether it falls within one or more scanning windows.
[0126] Figure 11 An exemplary logic for calculating the wo coordinates of output data points according to an embodiment of this disclosure is shown. In this example, it is assumed that the kernel size is 3*3*3 and the convolution stride is 2 in both the HWD direction.
[0127] As can be seen from the figure, according to the convolution operation rules and the corresponding convolution parameters (including the kernel size and stride), there is the following mapping relationship from wi_coord to wo_coord:
[0128] If wi_coord is odd, then the mapping is one-to-one, wo_coord = (wi_coord-1) / 2;
[0129] If wi_coord is even, the mapping can be one-to-one (boundary point wo_coord = 0 or wi - 1) or one-to-two (non-boundary point) depending on whether the wi coordinate is a boundary. For example, the mapping relationship can be wo_coord = wi_coord / 2 - 1 and wo_coord = wi_coord / 2.
[0130] Based on this mapping relationship, wo_coord can be calculated. See, for example... Figure 10 For example, based on the first wi_coord vector = [1,4], we can calculate wo_coord = [0,1,2], indicating 3 valid output data points. Based on the second wi_coord vector = [0,2,3], we can calculate wo_coord = [0,0,1,1], which, after deduplication, is [0,1], indicating 2 valid output data points. During the calculation process, we can count the number of wo coordinates in each row, i.e., the number of valid output data points in each row, for later use.
[0131] It is understandable that the mapping relationship may change depending on the convolution parameters. Those skilled in the art can deduce the mapping relationship between the wi coordinates of the input data and the wo coordinates of the output data based on the specific convolution parameters (specifically, the kernel size kx and the stride Sx).
[0132] Once the coordinates of the output data points to which the valid input data points contribute are inferred from the coordinates of the input data points (wi), it can be determined which one or more second scan windows the input data points fall into.
[0133] For example, in Figure 10 In the example, for the first wi_coord vector, it is inferred that the valid input data point with coordinate 1 falls into the second scan window with coordinate wo=0, and the valid input data point with coordinate 4 falls into the two second scan windows with coordinates wo=1 and wo=2.
[0134] In the example above, when extracting valid input data points to construct the Q matrix, the Q matrix can also be constructed based on whether wi_coord is odd or even. Specifically, the position of the input data point in the second scan window can be determined based on whether the wi coordinate is odd or even. For example, an input data point with an odd wi coordinate will necessarily fall in the middle of the second scan window, while an input data point with an even wi coordinate will fall in two adjacent positions between two adjacent second scan windows.
[0135] Therefore, based on this rule, the value of the valid input data point can be read from, for example, shared memory SRAM and stored in on-chip memory (e.g., NRAM).
[0136] Specifically, for valid input data points with even-numbered wi coordinates, two copies need to be stored because they will fall into two adjacent scan windows and be located at the end of the current scan window and the beginning of the next scan window; for valid input data points with odd-numbered wi coordinates, only one copy is needed, and its position is in the middle of the current scan window.
[0137] Therefore, by scanning each of the valid input data points row by row, a Q matrix can be constructed. For each of the M non-empty wi_coord vectors, it can be processed sequentially in the manner described above. The Q matrix constructed in this way has M rows, each row consisting of Li second scan windows, where Li depends on the number of output data points in that row, as counted earlier when calculating the wo coordinates.
[0138] The count of output data points can be used to calculate the third output parameter `hout` in the output data. According to the definition of the third output parameter `hout`, the i-th value represents the total number of valid output points in the preceding i-1 rows. Therefore, based on the count of coordinates for each row of `wo` calculated using `wo_coord`, the values of `hout` can be obtained through integration and summation / accumulation. For example, in... Figure 5 In the example, since the output data has two wo coordinates in the 0th row, no wo coordinates in the 1st row, and two wo coordinates in the 2nd row, the corresponding hout = [0, 2, 2, 4].
[0139] Back Figure 6 In step 630, after constructing matrix Q, M one-dimensional convolution operations can be performed on matrix Q. The M one-dimensional convolution kernels are obtained by splitting the original N-dimensional convolution kernel according to the first convolution dimension. The convolution stride of the one-dimensional convolution operation is equal to the size of the second scanning window, which is also equal to the size of the N-dimensional convolution kernel in the first convolution dimension. In this disclosed embodiment, the first convolution dimension is W, therefore the convolution stride of the one-dimensional convolution operation is equal to kx.
[0140] Therefore, M partial sums can be obtained through M one-dimensional convolution operations, each corresponding to a single output point (wo) in the same row. The coordinates of each part and its corresponding wo are also determined within each partial sum.
[0141] Next, in step 640, the M-path portions and results are merged into one fused data stream according to their corresponding output data point coordinates to obtain the final result of the corresponding row of wo output data points. During the merging process, portions and results with the same wo coordinates are accumulated.
[0142] The above merging process can be carried out in several ways.
[0143] In some embodiments, the merging and fusion process can be implemented in hardware. In these embodiments, this can be achieved using a hardware instruction called the MERGE instruction. The basic function of the MERGE instruction is to merge multiple streams of data to be merged into a single stream of merged data, according to their index order, and to accumulate data with the same index.
[0144] In other embodiments, the merge and fusion process can be implemented in software. In these embodiments, for example, a fully vectorized sorting algorithm based on a multi-core processor can be used to perform the sorting during the merge and fusion process. Then, the `bang_add` operator is called to traverse the sorted data; when coordinates are the same, they are directly added; if they are different, no addition is needed, and the traversal continues.
[0145] The convolution operation scheme for sparse data in this disclosure embodiment has been described above from multiple aspects. Compared with conventional convolution schemes, the scheme in this disclosure embodiment only performs operations on non-zero / non-empty data in the sparse data, which can avoid excessive invalid operations, greatly save computational load, and improve processing efficiency. Furthermore, by filtering the input data at different levels (e.g., three levels), the data that needs to be subjected to convolution operations can be extracted quickly and efficiently. The sparse convolution scheme provided in this disclosure embodiment is particularly suitable for processing LiDAR point cloud data.
[0146] This disclosure also provides a data processing circuit for performing convolution operations on the sparse data described above, and a data processing method implemented by the data processing circuit.
[0147] Figure 12 A schematic structural diagram of a data processing circuit that can implement embodiments of this disclosure is shown as an example. Figure 12 As shown, the data processing circuit 1200 includes a control circuit 1210, a storage circuit 1220, and an arithmetic circuit 1230.
[0148] The control circuit 1210 is responsible for handling various functions on the data processing circuit 1200, including but not limited to control, instruction fetching, decoding, and calculation. The control circuit 1210 may include, for example, the control module 31 shown in Figure 3.
[0149] In some embodiments, the control circuit 1210 may be configured to control the storage circuit 1220 and the arithmetic circuit 1230 to perform N-dimensional convolution operations on the input data and the convolution kernel, where N>1, and N represents the number of convolution dimensions in which sliding accumulation is performed. In this convolution operation, the input data is sparse data and represented in a dense form.
[0150] Furthermore, the control circuit 1210 can be configured to: filter the input data blocks allocated to the computation circuit 1230 for computation in the current round according to the input parameters of the input data, wherein the computation circuit 1230 calculates one row of output data in dimension W in each round.
[0151] The storage circuit 1220 can be used to store information, including at least pre-processing and / or post-processing information, and may also include intermediate information that needs to be cached during processing, such as various RAMs shown in FIG3, or on-chip caches. In some embodiments, the storage circuit 1220 can be configured to store input data, convolution kernels, convolution operation results, and / or cache intermediate results, such as cached portions and results, or provide cache space required during the execution of MERGE instructions.
[0152] The arithmetic circuit 1230 can be configured to perform various arithmetic operations according to relevant instructions. Specifically, the arithmetic circuit 1230 can be configured to perform multiple one-dimensional convolution operations on the input data and the convolution kernel under the control of the control circuit 1210, to obtain multiple operation results and the corresponding output point coordinates on the first convolution dimension; and to merge the multiple operation results into one fused data as the result of the convolution operation according to their corresponding output point coordinates, wherein the operation results with the same output point coordinates are accumulated.
[0153] In some embodiments, the above N-dimensional convolution operation is broken down into M one-dimensional convolution operations, where M is equal to the product of the sizes of the convolution kernels in the N-1 convolution dimensions other than the first convolution dimension.
[0154] In one embodiment, the arithmetic circuit 1230 may further include an arithmetic processing circuit (not shown), which can be configured to preprocess the data before the arithmetic circuit performs the operation or postprocess the data after the operation according to the arithmetic instructions. In some application scenarios, the aforementioned preprocessing and postprocessing may include, for example, data splitting and / or data concatenation operations.
[0155] In some embodiments, the arithmetic circuit 1230 may include multiple processor cores, each of which can process the input data block allocated by the control circuit 1210 at a time, for example, calculating one line of output points W at a time.
[0156] Specifically, each processor core can be further configured to perform the following operations: construct a matrix Q for performing a one-dimensional convolution operation based on the allocated blocks indicating the input data blocks to be processed; calculate the coordinates of each part and result of the output points of the one-dimensional convolution operation in the first convolution dimension; perform multiple one-dimensional convolution operations on matrix Q to obtain multiple partial sums and results; and merge and fuse the multiple partial sums and results to obtain the final convolution operation result.
[0157] Although the coordinate determination step is described above as being performed by computational circuitry, those skilled in the art will understand that this step can also be performed by software, such as by control circuitry. Furthermore, while the various processing steps are generally described above as being performed on computational circuitry, this computational circuitry can also be distributed, for example, including computational circuitry in heterogeneous systems, so that some operations are performed, for example, on a CPU, and others, for example, on a GPU. In one implementation, preprocessing of the input data can be performed, for example, on a CPU; this preprocessing may include, for example, densification of sparse input data, etc. One-dimensional convolution operations with the input data and convolution kernels, multi-path subprocessing, and merging of results can be performed on a GPU, thereby fully leveraging the advantages of heterogeneous systems.
[0158] Those skilled in the art will understand that the description of the convolution operation processing of sparse data in the embodiments of this disclosure described above in conjunction with the accompanying drawings can also be applied to... Figure 12 The data processing circuit is not described again.
[0159] This disclosure also provides a chip that may include the data processing apparatus of any of the embodiments described above in conjunction with the accompanying drawings. Furthermore, this disclosure also provides a circuit board that may include the aforementioned chip.
[0160] Depending on the application scenario, the electronic devices or apparatus disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablets, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, dashcams, navigators, sensors, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, autonomous driving terminals, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound machines, and / or electrocardiographs. The electronic devices or apparatus disclosed herein can also be applied in fields such as the Internet, IoT, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Furthermore, the electronic devices or apparatus disclosed herein can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as cloud computing, edge computing, and terminal applications. In one or more embodiments, the high-computing-power electronic devices or apparatuses according to the present disclosure can be applied to cloud devices (e.g., cloud servers), while the low-power electronic devices or apparatuses can be applied to terminal devices and / or edge devices (e.g., smartphones or cameras). In one or more embodiments, the hardware information of the cloud devices and the hardware information of the terminal devices and / or edge devices are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud devices to simulate the hardware resources of the terminal devices and / or edge devices based on the hardware information of the terminal devices and / or edge devices, so as to complete the unified management, scheduling and collaborative work of end-to-cloud or cloud-edge-end integration.
[0161] It should be noted that, for the sake of brevity, this disclosure describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art will understand that the solutions disclosed herein are not limited by the order of the described actions. Therefore, based on the disclosure or teachings of this document, those skilled in the art will understand that some steps can be performed in a different order or simultaneously. Furthermore, those skilled in the art will understand that the embodiments described in this disclosure can be considered optional embodiments, that is, the actions or modules involved are not necessarily essential for the implementation of one or more solutions disclosed herein. In addition, depending on the solution, the description of some embodiments in this disclosure may have different emphases. In view of this, those skilled in the art will understand that parts not described in detail in a certain embodiment of this disclosure can also be referred to the relevant descriptions of other embodiments.
[0162] In terms of specific implementation, based on the disclosure and teachings of this document, those skilled in the art will understand that several embodiments disclosed herein can also be implemented in other ways not disclosed herein. For example, regarding the various units in the electronic device or apparatus embodiments described above, this document has divided them based on logical functions, but in actual implementation, there may be other ways of division. As another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. Regarding the connection relationships between different units or components, the connections discussed above in conjunction with the accompanying drawings can be direct or indirect couplings between units or components. In some scenarios, the aforementioned direct or indirect couplings involve communication connections utilizing interfaces, where the communication interface can support electrical, optical, acoustic, magnetic, or other forms of signal transmission.
[0163] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network units. Furthermore, depending on actual needs, some or all of the units can be selected to achieve the purpose of the solution described in the embodiments of this disclosure. Additionally, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically independently.
[0164] In other implementation scenarios, the integrated units described above can also be implemented in hardware, i.e., as specific hardware circuits, which may include digital circuits and / or analog circuits. The physical implementation of the circuit's hardware structure may include, but is not limited to, physical devices, which may include, but are not limited to, transistors or memristors. Therefore, the various devices described herein (e.g., computing devices or other processing devices) can be implemented using appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage units or storage devices can be any suitable storage medium (including magnetic storage media or magneto-optical storage media), such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), ROM, and RAM.
[0165] The embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.
Claims
1. A data processing circuit, applied in the field of target detection based on lidar, comprising a control circuit, a storage circuit, and a processing circuit, wherein: The control circuit is configured to control the storage circuit and the computing circuit to perform N-dimensional convolution operations on the input data and convolution kernel to implement target detection, where N>1, and N represents the number of convolution dimensions in which sliding accumulation is performed. The input data is sparse point cloud data representing detection data in three-dimensional space obtained from LiDAR target detection and is characterized in a dense form to save storage space. The storage circuit is configured to store information, which includes at least pre-processing, during-processing, and / or post-processing information, wherein the information includes input data characterized by the dense form; as well as The computing circuit is configured to perform multiple one-dimensional convolution operations on the input data and the convolution kernel under the control of the control circuit, so as to obtain multiple operation results and the corresponding output point coordinates on the first convolution dimension. The multi-path operation results are then grouped into a single fused data stream according to their corresponding output point coordinates, with the operation results having the same output point coordinates being accumulated.
2. The data processing circuit according to claim 1, wherein the N-dimensional convolution operation is split into M one-dimensional convolution operations, M being equal to the product of the sizes of the convolution kernel in N-1 convolution dimensions other than the first convolution dimension, and the first convolution dimension being the width W dimension.
3. The data processing circuit according to claim 2, wherein the input data includes at least a width W, a height H, and an input channel Ci dimension, the input data is dense and uniform in size at least in the Ci dimension, and the dense form of the input data includes three input parameters: The first input parameter, Min, represents the valid dense input data, with shape ain. ci, where ain is the number of valid input data points in the input data, and ci is the size of the Ci dimension of the input data; The second input parameter wi_coord represents the coordinates of each valid input data point in dimension W, with a shape of 1. ain; and The third input parameter hin represents the input data in H-dimensional compressed sparse row (CSR) format, with shape X. (Hin+1), where Hin represents the size of the H dimension of the input data, and X represents the product of the sizes of the other dimensions that the input data may have besides H, W, and Ci.
4. The data processing circuit according to claim 3, wherein the control circuit is further configured to: Based on the input parameters of the input data, the input data blocks allocated to the computing circuit for the current round are filtered, wherein the computing circuit calculates one row of output data in dimension W in each round.
5. The data processing circuit of claim 4, wherein the control circuit is further configured to filter input data blocks as follows: The third input parameter is traversed using a first scanning window and a first scanning stride to detect whether there are valid input data points within the scanning window. The first scanning window corresponds to the size of the input data block required to compute one row of output data in dimension W, and the first scanning stride is equal to the convolution stride of the convolution operation in dimension H. The block in the third input parameter corresponding to the first scan window that detects a valid input data point is assigned to the arithmetic circuit.
6. The data processing circuit according to claim 5, wherein the arithmetic circuit includes N IPU One processor core, and the control circuitry is further used for: The continuously detected N IPU Each block is sent to the N IPU Each processor core processes the data.
7. The data processing circuit according to any one of claims 5-6, wherein the arithmetic circuit is further configured to: Based on the received block, construct the matrix Q for which the one-dimensional convolution operation is to be performed; and Calculate the coordinates of each part of the output point and the result of the one-dimensional convolution operation in the first convolution dimension.
8. The data processing circuit according to claim 7, wherein the arithmetic circuit is further configured to construct the matrix Q as follows: Extract the wi_coord vector of the valid input data points in the input data block corresponding to the block from the second input parameter wi_coord; Based on the wi_coord vector, the input data corresponding to the wi_coord vector is traversed with a second scanning window and a second scanning step size to extract the corresponding input data points from the first input parameter Min to construct the matrix Q.
9. The data processing circuit according to claim 8, wherein the arithmetic circuit is further configured to: According to the instructions of the block, a specified number of data are extracted from the specified position of the wi_coord to form the wi_coord vector, and the number of wi_coord vectors extracted is equal to M; and Filter out any empty wi_coord vectors.
10. The data processing circuit according to any one of claims 8-9, wherein the arithmetic circuit is further configured to construct the corresponding rows of the matrix Q by scanning row by row according to the wi_coord vector as follows: The data covered by the second scanning window that detected valid input data points is extracted and sequentially tiled to construct the corresponding rows of the matrix Q; and Skip the second scan window where no valid input data points were detected; The second scanning window corresponds to the size of the N-dimensional convolution window of the convolution operation in the W dimension, and the second scanning stride is equal to the convolution stride of the convolution operation in the W dimension.
11. The data processing circuit according to any one of claims 7-10, wherein the arithmetic circuit is further configured to: Based on the kernel size and stride of the convolution operation, determine the mapping relationship between the W-dimensional coordinates of the input data and the W-dimensional coordinates of the output data; and Based on the mapping relationship, the W-dimensional coordinates of one or more corresponding output points are determined according to the W-dimensional coordinates of each valid input point, and used as the W-dimensional coordinates of the part and the result.
12. The data processing circuit according to any one of claims 7-11, wherein the arithmetic circuit is further configured to: The one-dimensional convolution operation is performed on each row of the matrix Q to obtain a multi-path partial sum result, wherein the one-dimensional convolution kernel of the one-dimensional convolution operation corresponds to the corresponding W-dimensional row of the N-dimensional convolution kernel of the N-dimensional convolution operation, and the convolution stride of the one-dimensional convolution operation is equal to the size of the N-dimensional convolution kernel in the W-dimensional dimension.
13. The data processing circuit according to claim 12, wherein the arithmetic circuit is further configured to: The multiple parts and results are merged into one fused data stream according to their corresponding W-dimensional coordinates. In the merging process, the parts and results with the same W-dimensional coordinates are accumulated to obtain the final result of the output data points of the corresponding row.
14. A chip comprising a data processing circuit according to any one of claims 1-13.
15. A circuit board comprising the chip according to claim 14.
16. A method for processing data using the data processing circuit according to any one of claims 1-13.
Citation Information
Patent Citations
Load-balanced sparse convolutional neural network accelerator and acceleration method thereof
CN109993297A
Data processing method, data processing device and electronic equipment
CN110399972A