Data processing apparatus, data processing method and related product

By employing a submanifold sparse convolution scheme, effective data elements are extracted from sparse data and processed with convolution kernels, thus solving the problem of low efficiency in sparse data processing and achieving efficient LiDAR point cloud data processing.

CN115221105BActive Publication Date: 2026-01-02CAMBRICON SINGGO (NANJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110482877.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-30
Publication Date
2026-01-02
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

Existing convolutional neural networks are inefficient when processing sparse data, especially during convolution operations, which waste a lot of computing power and cannot effectively utilize sparse point cloud data.

Method used

The submanifold sparse convolution scheme is adopted. By extracting the effective data elements in the sparse data and performing operations with the convolution kernel, it is converted into a second convolution operation, which reduces the amount of computation and improves the processing efficiency.

Benefits of technology

It effectively improves the processing efficiency of sparse data, maintains the sparsity of the output data, and is suitable for multidimensional convolution operations, especially for processing LiDAR point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115221105B_ABST
    Figure CN115221105B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a data processing apparatus, a data processing method and related products. The data processing apparatus can be implemented as a computing apparatus included in a combined processing apparatus, which can further include an interface apparatus and other processing apparatuses. The computing apparatus interacts with the other processing apparatuses to jointly complete a user-specified computing operation. The combined processing apparatus can further include a storage apparatus connected with the computing apparatus and the other processing apparatuses respectively, for storing data of the computing apparatus and the other processing apparatuses. The scheme of the present disclosure provides a convolution processing scheme for sparse data, which can simplify the processing and improve the processing efficiency of the machine.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to the field of data processing. More specifically, the present disclosure relates to a data processing apparatus, a data processing method, a chip and a board card. BACKGROUND

[0002] In recent years, great progress has been made in target detection, instance segmentation and key point detection based on convolutional neural networks. These detections are usually based on LiDAR data or RGB-D data, which can be applied in the fields of autonomous driving, robot vision, etc.

[0003] Unlike image data, LiDAR point cloud data is usually sparse, and the point density varies dramatically due to factors such as non-uniform sampling of 3D space, effective range of sensors, occlusion and relative pose, etc. Therefore, the traditional convolutional neural network suitable for dense data will become very inefficient when applied to such sparse data, especially when involving convolution operations, which will waste a lot of resources such as computing power on zero-value data points.

[0004] In view of this, it is desirable to provide an improved convolution scheme suitable for sparse data such as point cloud data to improve processing efficiency. SUMMARY

[0005] To at least partially solve one or more technical problems mentioned in the background, the present disclosure provides a data processing apparatus, a data processing method, a chip and a board card.

[0006] In a first aspect, the present disclosure discloses a data processing apparatus, comprising a storage circuit and a processing circuit, wherein: the storage circuit is configured to store information, the information at least comprising pre-processing and / or post-processing information; the processing circuit is configured to access the storage circuit and perform a first convolution operation on input data and a convolution kernel, wherein the input data is sparse data, each input data element having an index and a value: extracting valid data elements corresponding to valid output points from the input data elements, wherein the valid output points represent convolution output points whose receptive field centers have input data elements, and the valid data elements are input data elements that contribute to the valid output points; and performing a second convolution operation on the valid data elements and the convolution kernel to obtain each valid output point as the operation result of the first convolution operation.

[0007] In a second aspect, the present disclosure provides a chip comprising the data processing apparatus of any one of the preceding first aspect.

[0008] In a third aspect, the present disclosure provides a board card comprising the chip of any one of the preceding second aspect.

[0009] In a fourth aspect, the present disclosure provides a method of processing data using the data processing apparatus of any one of the preceding first aspect.

[0010] By means of the data processing apparatus, the method of processing data using the data processing apparatus, the chip and the board card as provided above, the embodiments of the present disclosure provide a convolution scheme suitable for sparse data, which can greatly save the amount of calculation and improve the processing efficiency by only operating the effective data elements in the input data elements with the convolution kernel. Further, in the convolution operation described above, the first convolution operation for the original input data and the convolution kernel is converted into the second convolution operation for the effective data elements and the convolution kernel, which can speed up the data processing. In addition, the sparse convolution scheme provided by the embodiments of the present disclosure can be applied to multi-dimensional convolution operations, including but not limited to two-dimensional convolution and three-dimensional convolution, so as to be applicable to the processing of LiDAR point cloud data. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0012] Figure 1 A structural diagram of a board card of an embodiment of the present disclosure is shown;

[0013] Figure 2 A structural diagram of a combined processing apparatus of an embodiment of the present disclosure is shown;

[0014] Figure 3 An internal structural diagram of a processor core of a single-core or multi-core computing apparatus of an embodiment of the present disclosure is shown;

[0015] Figure 4 An operation principle of a conventional convolution scheme is shown;

[0016] Figure 5 An exemplary diagram showing that the use of a conventional convolution scheme results in a decrease in sparsity is shown;

[0017] Figure 6 An exemplary principle of a sub-manifold sparse convolution scheme according to an embodiment of the present disclosure is shown;

[0018] Figure 7 An exemplary operation process of a sub-manifold sparse convolution scheme according to an embodiment of the present disclosure is shown;

[0019] Figure 8 An exemplary operation using the same padding in a sub-manifold sparse convolution operation according to an embodiment of the present disclosure is shown;

[0020] Figure 9 a schematic diagram showing pre-processing of high-dimensional sparse input data according to embodiments of the present disclosure;

[0021] Figure 10 a structural schematic diagram of a data processing apparatus according to embodiments of the present disclosure; and

[0022] Figure 11 an exemplary flow chart of a data processing method according to embodiments of the present disclosure. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of, rather than all of, the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative effort fall within the protection scope of the present disclosure.

[0024] It should be understood that the terms “first”, “second”, “third”, and “fourth” and the like in the claims, specification, and drawings of the present disclosure are used to distinguish different objects, and are not used to describe a particular order. The terms “include” and “contain” used in the specification and claims of the present disclosure indicate the presence of described features, whole, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components, and / or sets thereof.

[0025] It should also be understood that the terms used in the specification of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. As used in the specification and claims of the present disclosure, the singular forms “a”, “an” and “the” are intended to include the plural forms, unless the context clearly indicates otherwise. It should be further understood that the term “and / or” used in the specification and claims of the present disclosure refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0026] As used in the specification and claims of the present disclosure, the term “if’ can be interpreted as “when” or “upon” or “in response to a determination” or “in response to detecting” depending on the context.

[0027] The specific embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings.

[0028] Figure 1 A structural schematic diagram of a board card 10 according to embodiments of the present disclosure is shown. As shown in FIG. 1, the board card 10 includes a plurality of input data 11, a plurality of output data 12, a plurality of data processing units 13, and a plurality of data processing units 14. Figure 1As shown, the board card 10 includes a chip 101, which is a system on chip (SoC) integrated with one or more combined processing devices, which is an artificial intelligence operation unit to support various deep learning and machine learning algorithms to meet the intelligent processing needs in complex scenarios in the fields of computer vision, speech, natural language processing, data mining, etc. In particular, deep learning technology is widely used in cloud intelligent fields. A significant feature of cloud intelligent applications is the large amount of input data, which has high requirements for the storage capacity and computing capacity of the platform. The board card 10 of this embodiment is suitable for cloud intelligent applications and has a large off-chip storage, on-chip storage and strong computing capacity.

[0029] The chip 101 is connected with an external device 103 through an external interface device 102. The external device 103 is, for example, a server, a computer, a camera, a display, a mouse, a keyboard, a network card or a wifi interface, etc. The data to be processed can be transmitted from the external device 103 to the chip 101 through the external interface device 102. The computing result of the chip 101 can be transmitted back to the external device 103 through the external interface device 102. According to different application scenarios, the external interface device 102 can have different interface forms, such as a PCIe interface, etc.

[0030] The board card 10 further includes a storage device 104 for storing data, which includes one or more storage units 105. The storage device 104 is connected and transmits data with the control device 106 and the chip 101 through a bus. The control device 106 in the board card 10 is configured to regulate the state of the chip 101. For this purpose, in one application scenario, the control device 106 can include a micro controller unit (MCU).

[0031] Figure 2 is a structural diagram of the combined processing device in the chip 101 of this embodiment. As shown in Figure 2 The combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203 and a storage device 204.

[0032] The computing device 201 is configured to perform user-specified operations, mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations, which can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.

[0033] The interface device 202 is used to transmit data and control instructions between the computing device 201 and the processing device 203. For example, the computing device 201 can obtain input data from the processing device 203 via the interface device 202 and write into the storage device on the computing device 201. Further, the computing device 201 can obtain control instructions from the processing device 203 via the interface device 202 and write into the control buffer on the computing device 201. Alternatively or additionally, the interface device 202 can also read data from the storage device of the computing device 201 and transmit to the processing device 203.

[0034] The processing device 203 is a general-purpose processing device, which performs basic control including but not limited to data transfer, start and / or stop of the computing device 201, etc. Depending on the implementation, the processing device 203 can be one or more types of processors, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), or other general-purpose and / or special-purpose processors, including but not limited to a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc., and the number thereof can be determined according to actual needs. As mentioned above, only in terms of the computing device 201 of the present disclosure, it can be considered as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 201 and the processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure.

[0035] The storage device 204 is used to store data to be processed, which can be a DRAM, a DDR memory, usually with a size of 16G or more, for saving data of the computing device 201 and / or the processing device 203.

[0036] Figure 3 The internal structure of the processor core of the computing device 201 is shown when the computing device 201 is a single-core or multi-core device. The computing device 301 is used to process input data of computer vision, voice, natural language, data mining, etc. The computing device 301 includes three modules: a control module 31, a computation module 32, and a storage module 33.

[0037] The control module 31 is configured to coordinate and control the operations of the arithmetic module 32 and the storage module 33 to complete the task of deep learning, which includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 is configured to fetch instructions from the processing device 203, and the instruction decode unit 312 is configured to decode the fetched instructions and send the decoded results as control information to the arithmetic module 32 and the storage module 33.

[0038] The arithmetic module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is configured to perform vector operations, which can support complex operations such as vector multiplication, addition, and nonlinear transformation; and the matrix operation unit 322 is responsible for the core calculation of the deep learning algorithm, i.e., matrix multiplication and convolution.

[0039] The storage module 33 is configured to store or transfer related data, including a neuron storage unit (NRAM) 331, a parameter storage unit (WRAM) 332, and a direct memory access module (DMA) 333. The NRAM 331 is configured to store input neurons, output neurons, and intermediate results after calculation; the WRAM 332 is configured to store the convolution kernel of the deep learning network, i.e., the weight; and the DMA 333 is connected to the DRAM 204 through the bus 34 and is responsible for data transfer between the computing device 301 and the DRAM 204.

[0040] Based on the foregoing hardware environment, embodiments of the present disclosure provide a data processing device supporting convolution operation of sparse data. By providing an optimized convolution scheme, the convolution processing related to sparse data such as LiDAR point cloud data can be simplified and accelerated. The sparse convolution scheme provided by the embodiments of the present disclosure can be applied to multi-dimensional convolution operations, including but not limited to two-dimensional convolution and three-dimensional convolution. For the sake of simplicity and easy understanding, two-dimensional convolution is used as an example for illustration in some embodiments.

[0041] The“N-dimensional convolution” mentioned in the embodiments of the present disclosure, where N represents the number of convolution dimensions in which sliding accumulation is performed in the convolution operation. For example, when N = 2, the convolution kernel is shifted and accumulated in two dimensions (for example, width W and height H) according to the corresponding convolution step. When N = 3, the convolution kernel is shifted and accumulated in three dimensions (for example, width W, height H, and depth D) according to the corresponding convolution step. When N = 4, the convolution kernel is shifted and accumulated in four dimensions (for example, width W, height H, depth D, and batch) according to the corresponding convolution step. The“non-convolution dimension” mentioned in the embodiments of the present disclosure refers to a dimension in which the convolution kernel does not slide and accumulate.

[0042] In order to more clearly understand the convolution scheme of the embodiments of the present disclosure, the operation principle of the conventional convolution scheme is first described by taking a two-dimensional convolution as an example.

[0043] Figure 4 The operation principle of the conventional convolution scheme is shown. In this example, the convolution kernel 410 is dense, which is a 3x3 matrix, and the numbers in the convolution kernel are the corresponding weight data. The input data 420 is a 6x6 matrix, which is sparse, and only has four non-zero data: 2, 3, 4, and 5, as shown by the dark squares. For the sake of simplicity, in this exemplary convolution process, the convolution step of the two dimensions is set to 1, the padding is 0, and there is no dilation. The 3x3 size gray square in the figure represents the sliding accumulation process of the convolution kernel on the input data. 430 shows the calculation at the beginning of the convolution, 440 shows the calculation after sliding one step to the right, and 450 shows the calculation after sliding one step downward. In each step of calculation, the weight data of the convolution kernel is multiplied with the input data in position and accumulated. 460 is the final calculation result as the output data. The output data is a 4x4 size matrix. It can be seen that the calculation of 430 corresponds to the data at coordinate (1, 1) in the output data, the calculation of 440 corresponds to the data at coordinate (1, 2) in the output data, and the calculation of 450 corresponds to the data at coordinate (2, 1) in the output data.

[0044] From the description of Figure 4 It can be seen that the final result of the sparse convolution operation is only related to the operation result of the non-zero input data elements, and therefore, the multiplication and accumulation operation with the convolution kernel can be performed only for these non-zero input data elements. Further, by comparing the input data 420 and the output data 460 of Figure 4 It can be seen that the sparsity of the operation result after such sparse convolution operation is reduced compared to the sparsity of the original input data, that is, the number of non-sparse points increases. Specifically, there are only 4 non-sparse points in the input data 420, while there are 13 non-sparse points in the output data 460. When in a convolution network, with the increase of the convolution layers, the result of such sparse convolution operation will diverge, and the sparsity will quickly disappear.

[0045] Figure 5 An exemplary diagram showing that the conventional convolution scheme results in a reduction of sparsity is shown. Figure 5 The left graph (a) in the middle shows the original sparse input, which is a hand-drawn circle, a one-dimensional curve embedded on a two-dimensional grid. Figure 5 The middle graph (b) in the middle shows the result after one 3x3 conventional convolution, Figure 5 The right graph (c) in the middle shows the result after two 3x3 conventional convolutions. As can be seen from the graph, the sparsity of the original data is very high, but after the conventional convolution scheme, the sparsity is quickly lost.

[0046] In view of this, there are documents that propose a submanifold sparse convolution scheme, in which only when there is a non-zero input data element at the position corresponding to the center of the convolution kernel, the corresponding output result is valid, otherwise it is processed as zero. Under this sparse convolution scheme, the sparsity of the input and output can be kept consistent. The present disclosure provides a data processing apparatus and related data processing method for implementing such a sparse convolution scheme.

[0047] Figure 6 An exemplary principle of the submanifold sparse convolution scheme according to the present disclosure is shown. Figure 6 Still taking the data of Figure 4 as an example to describe the submanifold sparse convolution operation scheme of the present disclosure.

[0048] As shown in Figure 6 , only the case where there is a non-zero input data element at the position corresponding to the center of the convolution kernel is calculated. The graph shows four non-zero data elements in the input data, which are respectively located at the center of the convolution kernel, and the corresponding operations 610, 620, 630 and 640. The operation results of 610 and 640 have overflowed the convolution result range and are invalid results. The operation result of 620 corresponds to the (2, 2) position in the output data 650, and the operation result of 630 corresponds to the (2, 3) position in the output data 650.

[0049] As can be seen from the above exemplary principle diagram, in the operation based on the principle of submanifold sparse convolution, only the position multiplication and addition operation when there is a non-zero input data element at the position corresponding to the center of the convolution kernel needs to be calculated, thereby greatly reducing the amount of calculation compared to the conventional convolution operation. Further, since the valid output points in the output result correspond to the non-sparse points (i.e. non-zero input data elements) in the input data elements, the number of valid output points does not exceed the number of non-sparse points in the input data elements, so that the sparsity of the input result does not decrease compared to the input result.

[0050] In the description herein, the "convolution output point" mentioned in the embodiments of the present disclosure refers to each output point in the result of a convolution operation, for example, Figure 6 The output result 650 includes 16 convolution output points. Each convolution output point has a corresponding receptive field, and the shape of the receptive field is equal to the shape of the convolution kernel, for example Figure 6 the 3x3 gray box in 630.

[0051] The "effective output point" mentioned in the embodiments of the present disclosure refers to a convolution output point whose receptive field center has an input data element. For example, Figure 6 The output result 650 includes 2 effective output points. The receptive field of the effective output point at the (2, 2) position corresponds to the gray box in 620, and the center of the receptive field has an input data element 3; the receptive field of the effective output point at the (2, 3) position corresponds to the gray box in 630, and the center of the receptive field has an input data element 4.

[0052] As can be seen from the figure, when the receptive field of the effective output point has other input data elements at non-central positions, these input data elements also contribute to the result of the effective output point. In the embodiments of the present disclosure, such input data elements that contribute to the effective output point are referred to as "effective data elements". For example, for the effective output point (2, 2), its effective data elements include 3 at the center of the receptive field and 4 at the non-central position of the receptive field. Similarly, for the effective output point (2, 3), its effective data elements include 4 at the center of the receptive field and 3 at the non-central position of the receptive field.

[0053] In addition, as can be seen from the above definition, the effective data element is relative to the effective output point, one effective output point can include one or more effective data elements, and one effective data element can contribute to one or more effective output points.

[0054] As can be seen from the operation principle of Figure 6 , the input data elements 3 and 4 at the non-central positions of the receptive field still contribute to the corresponding convolution output points. Therefore, in the convolution operation scheme of the embodiments of the present disclosure, the effective data elements corresponding to the effective output points can be extracted from the input data elements first, and then the effective data elements are subjected to convolution operation with the convolution kernel, so as to obtain the corresponding effective output points.

[0055] Figure 7 An exemplary operation process of the sub-manifold sparse convolution scheme according to the embodiments of the present disclosure is shown. Still taking the data of Figure 4 as an example. Similarly to Figure 6, 710, 720, 730 and 740 correspond to the operations when the four non-zero data elements in the input data are located at the center of the convolution kernel, respectively. The operation results of 710 and 740 have overflowed the range of the convolution result and are invalid. The operation results of 720 and 730 are valid and correspond to one valid output point, for example, the coordinate points (2, 2) and (2, 3) in the output result, respectively.

[0056] In the sparse convolution operation process of the embodiments of the present disclosure, the convolution kernel is dense, and its input format can be the same as that of the conventional convolution. The input data is sparse, and its input format can be different from that of the conventional convolution, thereby saving storage space. In some embodiments, the input data is sparse data, each input data element has index and value information, which can be represented as (index, value). Depending on different applications, data types or scales, the index here can be one-dimensional, two-dimensional or more-dimensional, and the present disclosure is not limited in this respect. Different dimensions of the index can be converted into each other, for example, two-dimensional or more-dimensional indexes are converted into one-dimensional indexes according to a predetermined traversal rule, and vice versa. Similarly, depending on different applications, data types or scales, the value information here can include scalar, vector or higher-dimensional data.

[0057] For example, as shown in Figure 7 , the input data has four non-sparse points, which are ((1, 4), 2), ((3, 3), 3), ((3, 4), 4) and ((5, 6), 5). For the first non-sparse point, (1, 4) here represents a two-dimensional index coordinate, and 2 represents the specific value of the position pointed to by the index, which is a scalar in this example. The meanings of the remaining non-sparse points are similar.

[0058] In some embodiments, for each valid output point, the corresponding valid data element can be extracted from the input data elements. As explained before, a valid output point represents a convolution output point whose center of the receptive field has an input data element, and a valid data element is an input data element that contributes to the valid output point.

[0059] Specifically, in some embodiments, the associated valid data element corresponding to the valid output point can be extracted by scanning each input data element.

[0060] As shown in Figure 7 , each input data element can be scanned in turn. For example, input data point 2 can be scanned first. According to the foregoing description, input data point 2 can be represented in the form of index and value: ((1, 4), 2).

[0061] Then, it is determined whether the input data element is located at the center point of the receptive field of any convolution output point. In other words, it can be determined whether the operation result with the input data element as the center of the convolution kernel falls within the range of the convolution operation result, or in other words, it can be determined whether there is a valid output point associated with the input data element. If not, it means that its operation result is out of the range of the convolution result, which is invalid convolution, and the next input data element can continue to be scanned. If yes, the corresponding convolution output point is a valid output point, and then the valid data element corresponding to the valid output point can be extracted.

[0062] In actual implementation, although the input data is sparse, its input format can be dense, such as the input format of the combination of index and numerical value mentioned above. Therefore, whether the input data element is located at the center point of the receptive field of the convolution output point can be determined according to the index of the input data element.

[0063] There are various ways to determine whether the input data element is located at the center point of the receptive field of any convolution output point.

[0064] In some embodiments, the center point positions of the receptive fields of all convolution output points of the convolution operation can be determined, and then the index of the input data element is compared with the center point positions of the receptive fields to determine whether the input data element is located at the center point of the receptive field of any convolution output point.

[0065] It can be understood that, given the input data shape, the convolution kernel shape, and the convolution stride, the positions of all convolution output points of the convolution operation can be determined. For example, Figure 7 The convolution operation result of Figure 6 is the same as that of , and the shape of the output result is 4x4, including 16 convolution output points. Each convolution output point has a corresponding receptive field, and the shape of the receptive field is equal to the shape of the convolution kernel. Accordingly, the center point positions of the receptive fields of all convolution output points can also be determined.

[0066] In some implementations, the center point positions of the receptive fields can be represented in the same coordinate system as the input data element, so that whether the input data element is located at the center point of the receptive field of the convolution output point can be determined by comparing the coordinates or indexes of the two.

[0067] For example, in the Figure 7 example, the center points of the receptive fields of the 16 convolution output points are shown by the dashed boxes of 4x4. By comparing the index of the input data element with the index of the center point of the receptive field, it can be seen that input data points 2 and 5 are not located at the center point of any receptive field, and input data points 3 and 4 fall into the center of the receptive field of a convolution output point, respectively.

[0068] In some embodiments, it can be determined whether the input data element is located at the center point of the receptive field of any of the convolution output points based on the index of the input data element.

[0069] For example, as can be seen from 710, the receptive field with the input data point 2 as the center has exceeded the shape range of the input data, i.e., its corresponding operation result has overflowed the convolution result range, thus it can be determined that the input data point 2 is not located at the center point of the receptive field of any of the convolution output points, i.e., the input data point 2 does not exist an associated valid output point.

[0070] Then, in response to determining that the input data element is located at the center point of the receptive field of a convolution output point, the input data elements within the receptive field of the convolution output point (which is a valid output point at this time) are extracted as the corresponding valid data elements of the convolution output point. Similarly, whether the input data element falls within the currently determined receptive field range is judged based on the index of the input data element, so as to perform corresponding extraction.

[0071] In some embodiments, the sparse form of the valid data elements within the receptive field of the corresponding valid output point can be further generated based on the extracted valid data elements.

[0072] For example, when the input data element 3 is scanned, it can be determined that it is located at the center of the receptive field of the convolution output point (2, 2), at this time all the input data elements within the receptive field of the convolution output point (2, 2) are extracted as the corresponding valid data elements of the convolution output point (2, 2). As can be seen from 720, the valid data elements corresponding to the convolution output point (2, 2) include the input data elements 3 and 4. The input data elements 3 and 4 within the receptive field of the convolution output point (2, 2) are expressed as a sparse form, which is a 3x3 sparse matrix, as shown in 751.

[0073] Similarly, when the input data element 4 is scanned, it can be determined that it is located at the center of the receptive field of the convolution output point (2, 3). As can be seen from 730, the valid data elements corresponding to the convolution output point (2, 3) include the input data elements 3 and 4. The input data elements 3 and 4 within the receptive field of the convolution output point (2, 3) are expressed as a sparse form, which is a 3x3 sparse matrix, as shown in 752.

[0074] Finally, the convolution operation can be performed on the extracted valid data elements and the convolution kernel to obtain corresponding respective valid output points as the final operation result. In order to distinguish from the original convolution operation, the original convolution operation is referred to as the "first convolution operation" and the convolution operation of the extracted valid data elements and the convolution kernel is referred to as the "second convolution operation". It can be understood that the first convolution operation refers to the convolution operation of the input data elements and the convolution kernel based on the sub-manifold sparse convolution principle, which has a first convolution stride, for example, stride1 = 1 in the example. Figure 7 The second convolution operation refers to the conventional convolution operation between the extracted valid data elements and the convolution kernel, which has a second convolution stride. From the analysis of Figure 7 , it can be seen that since the valid data elements are extracted according to the shape of the convolution kernel of the first convolution operation and tiled into a sparse form, the second convolution operation can be regarded as using the convolution kernel to perform a second convolution operation on the valid data elements in the sparse form with a second convolution stride. The second convolution stride is equal to the shape of the convolution kernel, and in the above example, stride2 = 3.

[0075] The output of each step of convolution corresponds to a valid output point. For example, in Figure 7 , the result of the second convolution operation on the sparse valid data 751 corresponds to the valid output point (2, 2); and the result of the second convolution operation on the sparse valid data 752 corresponds to the valid output point (2, 3).

[0076] From the above description, it can be seen that the embodiments of the present disclosure provide an implementation scheme of sub-manifold sparse convolution operation, which can simplify the operation and improve the processing efficiency by converting the first convolution operation based on the sub-manifold sparse convolution principle into the second convolution operation between the valid data elements and the convolution kernel.

[0077] In some embodiments, there is a padding operation in the convolution operation. For example, in a target detection algorithm based on LiDAR data, same padding needs to be performed, that is, by padding, the shape of the input data is made the same as the shape of the output data after the convolution operation. It can be understood that in other application scenarios of convolution operation, there can be different padding rules.

[0078] Figure 8 An exemplary operation after using same padding in the sub-manifold sparse convolution operation according to the embodiments of the present disclosure is shown.

[0079] As shown in the figure, the original input data is a 6×6 matrix. The convolution kernel is a 3×3 matrix, and the stride is 1. To ensure the output data has the same shape as the original input data, the input data needs to be padded. As shown in 820, the white squares around the 6×6 matrix represent the padding area, that is, adding one row / column of data to the top, bottom, left, and right sides to form an 8×8 matrix. The padding data can be, for example, zero. After performing submanifold sparse convolution with the convolution kernel 810, the padded input data 820 is output data 830, which is also a 6×6 matrix.

[0080] As can be seen from the figure, the valid output points (dark squares) of the output data 830 correspond one-to-one with the non-zero data element points (dark squares) of the input data 820, thus not reducing the sparsity of the input data.

[0081] Since the convolution output point changes with the padding operation, in this embodiment, the index of the receptive field center can be adjusted at least according to the padding rule of the first convolution operation before determining whether the input data element is located at the center of the receptive field of the convolution output point.

[0082] For example, in Figure 8 As shown, with the same padding, the receptive field center point generally corresponds to the shape of the original input data, and therefore its index range also corresponds to the shape range of the original input data. In this embodiment, it can be determined that all input data elements will fall into the receptive field center point of a certain convolution output point, so it is only necessary to extract the corresponding valid data elements for the convolution output point associated with each input data element.

[0083] Those skilled in the art will understand that the index adjustment process can be performed at any stage of determining whether an input data element falls within the receptive field center point, as long as the impact of the filling rule is taken into account, and the embodiments disclosed herein are not limited in this respect.

[0084] In some embodiments, the input data elements can be sorted based on their indices before scanning. For example, the input data elements can be arranged in ascending order of their indices. Therefore, the aforementioned scanning process can also scan the input data elements sequentially according to their order of arrangement.

[0085] Since the input data elements are ordered by their indices, and based on the principle of submanifold sparse convolution, each input data element corresponds to one output point. For example, in the case of the same padding, each input data element corresponds to one valid output point. Therefore, in some embodiments, each input data element can be scanned in sequence, and the valid data element of the associated valid output point can be extracted, and the above-mentioned second convolution operation can be performed accordingly, so that the valid output points can be obtained in sequence. It can be understood that the index of the valid output point is associated with the index of the input data element at the center of the receptive field, and can be determined according to the index mapping relationship. The above-mentioned scanning, extracting, operation and output process can meet the flow sequence of load-compute-store (LCS), thereby accelerating the processing process.

[0086] In addition, as mentioned above, the sparse convolution scheme provided by the embodiments of the present disclosure can be applied to multi-dimensional convolution operations, including but not limited to two-dimensional convolution and three-dimensional convolution.

[0087] The input data to be convolved can include multi-dimensional data, and it is sparse in multiple dimensions. For example, in target detection based on LiDAR data, the input data is detection data in a three-dimensional space, which represents, for example, the gray value, RGB, signal strength, etc. at each three-dimensional coordinate point, so according to the information content it represents, the input data element at each coordinate point can be one-dimensional, two-dimensional, three-dimensional or higher-dimensional data. Due to the characteristics of point cloud data, the coordinate points with non-zero value data elements are sparse, that is, they are sparse in three spatial dimensions (for example, width W, height H and depth D).

[0088] Depending on the initial state of the input data, preprocessing can be performed before the sparse input data is provided to the processing circuit for processing. In some embodiments, such preprocessing can include, for example, merging the sparse multiple dimensions into one dimension, densifying the sparse data points in the input data in the merged dimension to form input data elements, and using index and numerical information to represent each input data element. The index can be one-dimensional or multi-dimensional, and the numerical information can include any of scalar, vector or higher-dimensional data.

[0089] In one example, for example, referring to Figure 4 , the input data is a two-dimensional 6x6 matrix, which is sparse in two dimensions of width W and height H. When preprocessing, W and H are merged into one dimension, and the sparse data points (2, 3, 4 and 5 in this example) are densified in the merged dimension, thereby forming four dense input data elements. Then, index and numerical information are used to represent each input data element. The index of the data element can represent its position relationship in the sparse form of the input data before densification. For example, Figure 4The indices of the four input data elements in the example are (1, 4), (3, 3), (3, 4) and (5, 6), respectively. The indices in this example are two-dimensional indices, which can also be converted to one-dimensional indices, for example, 4, 15, 16 and 29, respectively. The numerical information of the four data elements are 2, 3, 4 and 5, respectively, i.e., four scalars.

[0090] In another example, for example, referring to Figure 9 which shows a schematic diagram of the preprocessing of high-dimensional sparse input data according to embodiments of the present disclosure. As shown in the figure, the input data 910 in sparse form includes five dimensions, a batch dimension B, a HWD three-dimensional space dimension and an input channel dimension Ci. The input data is sparse in the B dimension and the HWD three-dimensional space, and the dark squares in the HWD solid matrix in the figure represent places with numerical values, and the other parts are all zero values. There are multiple such HWD solid matrices in the B dimension, and the sparse patterns (i.e., the positions of the dark squares) on each solid matrix can be different. The input data is dense in the Ci dimension, and the Ci dimension is the lowest dimension. Due to the limited performance of the drawing, only four dimensions are shown in 610 in the figure, but the Ci dimension can be understood as the thickness of each dark square. The size of the Ci dimension is uniform, i.e., the thickness of each dark square is the same. In the preprocessing process, the four sparse dimensions (the B dimension and the HWD three-dimensional space dimension) of the input data can be merged into one dimension Ni, and the sparse data points (dark squares in the figure) are densified in the merged dimension, thereby forming dense input data elements. That is, each HWD solid matrix in the B dimension performs the same dimension merging and densification processing, thereby obtaining the preprocessed input data 920 in dense form, which is a two-dimensional matrix, with the low dimension being Ci and the high dimension being the merged dimension Ni of BHWD.

[0091] Next, the indices and numerical information are used to represent each densified input data element. Figure 9 The indices of the input data elements in the example can be represented using their coordinates in the BHWD four-dimensional space before densification, or can be converted to one-dimensional indices. The numerical information of each input data element can be regarded as a Ci vector.

[0092] The input data after the above preprocessing can be provided to the processing circuit for subsequent processing.

[0093] The convolution operation scheme of sparse data of the embodiments of the present disclosure is described above from multiple aspects. Compared with the conventional convolution scheme, the convolution operation scheme of the embodiments of the present disclosure is based on the operation principle of sub-manifold sparse convolution, avoids the problem of reduced sparsity, and also reduces the amount of computation. Further, by converting the first convolution operation based on sub-manifold sparse convolution into the second convolution operation of the effective data elements and the convolution kernel, an innovative scheme suitable for the data processing device to execute is provided to implement the sub-manifold sparse convolution operation. In some implementations, by decomposing the convolution operation into steps such as extracting effective data elements, multiplication and addition operation, and output, it can be implemented through LCS pipelining, which is especially suitable for the hardware environment of the embodiments of the present disclosure described in conjunction with the accompanying drawings, and makes full use of the high efficiency of parallel processing. In addition, the sparse convolution scheme provided by the embodiments of the present disclosure can be especially suitable for processing based on LiDAR point cloud data.

[0094] The embodiments of the present disclosure also provide a data processing device for performing the convolution operation of sparse data described above, and a data processing method implemented by the data processing device.

[0095] Figure 10 An exemplary schematic structural diagram of a data processing device that can implement the embodiments of the present disclosure is shown. As shown, the data processing device 1000 includes a processing circuit 1010 and a storage circuit 1020. Figure 10 The processing circuit 1010 is responsible for processing various functions on the data processing device 1000, including but not limited to control, instruction fetching, decoding, operation, etc. The processing circuit 1010 may, for example, include the control module 31 and / or the operation module 32 in the

[0096] The processing circuit 1010 is responsible for processing various functions on the data processing device 1000, including but not limited to control, instruction fetching, decoding, operation, etc. The processing circuit 1010 may, for example, include the control module 31 and / or the operation module 32 in the Figure 3

[0097] In some embodiments, the processing circuit 1010 can be configured to access the storage circuit 1020 and perform a first convolution operation on input data and a convolution kernel. The first convolution operation is a convolution operation based on the principle of sub-manifold sparse convolution, wherein the input data is sparse data, and each input data element has index and value information.

[0098] In some embodiments, the processing circuit 1010 can be configured to perform the first convolution operation as follows: extracting, from the input data elements, effective data elements corresponding to effective output points, wherein the effective output points represent convolution output points whose receptive field centers have input data elements, and the effective data elements are input data elements that contribute to the effective output points; and performing a second convolution operation on the effective data elements and the convolution kernel to obtain each effective output point as an operation result of the first convolution operation.

[0099] ​Further, in some embodiments, the processing circuit 1010 can be configured to extract valid data elements as follows: sequentially scan each input data element; during the scanning, determine whether the input data element is located at the center point of the receptive field of any convolution output point; and in response to determining that the input data element is located at the center point of the receptive field of one convolution output point, extract the input data elements within the receptive field of the convolution output point as the valid data elements of the convolution output point.

[0100] Further, in some embodiments, the processing circuit 1010 can be configured to sort the input data elements according to their index sizes before scanning the input data elements. Thus, the input data elements can be sequentially scanned in the order of, for example, the index from small to large, and the valid data elements of the associated valid output points can be extracted accordingly for the second convolution operation to obtain the corresponding valid output points. Such sequential processing manner can be implemented in the hardware environment of the embodiments of the present disclosure in the manner of LCS pipeline, thereby improving the processing efficiency.

[0101] In some embodiments, the processing circuit 1010 can be configured to determine whether the input data element is located at the center point of the receptive field of any convolution output point as follows: determine the index of the center point of the receptive field of all convolution output points based on the input data shape, the kernel shape and the first convolution stride of the first convolution operation; compare the index of the input data element with the index of the center point of the receptive field; and in response to the index of the input data element being the same as the index of one center point of the receptive field, determine that the input data element is located at the center point of the receptive field of the convolution output point.

[0102] Optionally or additionally, in some embodiments, when the first convolution operation has a corresponding padding rule, the index of the center point of the receptive field can be adjusted according to the padding rule.

[0103] In some embodiments, the processing circuit 1010 can be further configured to generate a sparse form of the valid data elements within the receptive field of the associated convolution output point according to the extracted valid data elements. In such embodiments, the subsequent second convolution operation processing can be performed on the sparse form of the valid data elements, thereby simplifying the processing process, which can be performed as a regular convolution processing only.

[0104] In particular, in these embodiments, the processing circuit 1010 can be further configured to perform a second convolution operation on the valid data elements and the convolution kernel as follows: determine a second convolution stride of the second convolution operation based on a shape of the convolution kernel; and perform the second convolution operation on the valid data elements in the sparsified form according to the second convolution stride using the convolution kernel to obtain an operation result. In some embodiments, the second convolution stride of the second convolution operation is equal to a size of the convolution kernel. The convolution output points of this second convolution operation correspond one-to-one to the valid output points mentioned above, and their indices are associated with the indices of the input data elements at the center of their receptive fields.

[0105] In some embodiments, the input data comprises multi-dimensional data that is sparse in multiple dimensions. In these embodiments, the input data is pre-processed before being provided to the processing circuit. Such pre-processing can include, for example, merging the sparse multiple dimensions into one dimension; densifying the sparse data points in the input data in the merged dimension to form input data elements; and representing each input data element using index and value information, where the index is one or more dimensional indices, and the value information comprises any of a scalar, a vector, or a higher dimensional data.

[0106] The storage circuit 1020 can be configured to store information or carry relevant data, which at least includes pre-processing and / or post-processing information, and can also include intermediate information that needs to be cached during processing, which can be, for example, Figure 3 The various RAMs shown, or on-chip caches. In some embodiments, the storage circuit 1020 can be configured to store input data, convolution kernels, convolution operation results, and / or cache possible intermediate results.

[0107] Figure 11 An exemplary flowchart of a data processing method implemented by a data processing apparatus according to embodiments of the present disclosure is shown. In this data processing method, the processing circuit accesses the storage circuit to perform a first convolution operation on input data and a convolution kernel, where the input data is sparsified data, each input data element having index and value information.

[0108] More specifically, in step 1110, the processing circuit extracts, from the input data elements, valid data elements corresponding to valid output points, where a valid output point represents a convolution output point whose center of receptive field exists an input data element, and a valid data element is an input data element that contributes to the valid output point.

[0109] Then, in step 1120, a second convolution operation is performed on the valid data elements and the convolution kernel to obtain each valid output point as an operation result of the first convolution operation.

[0110] Although in the above description, the steps of extracting the valid data elements and performing the second convolution operation are separated and described in a sequential order, one skilled in the art can understand that when performed in a pipelined manner, the two steps can also be performed simultaneously, and the embodiments of the present disclosure have no limitation in this respect. In addition, although in the above description, the steps of extracting the valid data elements and performing the second convolution operation are both described as performed by the processing circuit, one skilled in the art can understand that the step of extracting the valid data elements can be performed by a control module in the processing circuit, for example, by program software, while the step of performing the second convolution operation can be performed by an operation module in the processing circuit, for example, by a multiplication-addition circuit or the like hardware, and the embodiments of the present disclosure have no limitation in this respect. Further, although in the above description, each processing step is generally described as performed on the processing circuit, the processing circuit here can also be distributed, for example, containing processing circuits in a heterogeneous system, so that a part of the operations are performed on a CPU, for example, while another part of the operations are performed on a GPU. In one implementation, the preprocessing of the input data can be performed on a CPU, for example, which can include, for example, the densification of the input data in a sparse form, the index sorting of the data elements after densification, and the like. The extraction of the valid data elements, the second convolution operation with the convolution kernel, and the like processing can be performed on a GPU, thereby taking full advantage of the heterogeneous system.

[0111] One skilled in the art can understand that the description of the embodiments of the present disclosure in the foregoing description in connection with the drawings in relation to the convolution operation processing of sparse data can also be applied to the data processing apparatus of Figure 10 and the data processing method of Figure 11 , and thus repeated description is not performed.

[0112] The present disclosure also provides a chip, which can include the data processing apparatus of any of the embodiments described in the foregoing description in connection with the drawings. Further, the present disclosure also provides a board card, which can include the aforementioned chip.

[0113] According to different application scenarios, the electronic device or apparatus of the present disclosure can include a server, a cloud server, a server cluster, a data processing apparatus, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a PC device, an Internet of Things terminal, a mobile terminal, a mobile phone, a vehicle record device, a navigator, a sensor, a camera, a camera, a video camera, a projector, a watch, a headset, a mobile storage, a wearable device, a visual terminal, an autonomous driving terminal, a vehicle, a household appliance, and / or a medical device. The vehicle includes an airplane, a ship, and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, an electric rice cooker, a humidifier, a washing machine, an electric lamp, a gas stove, an oil smoke exhaust fan; the medical device includes a nuclear magnetic resonance instrument, a B-ultrasound instrument, and / or an electrocardiograph. The electronic device or apparatus of the present disclosure can also be applied to the fields of Internet, Internet of Things, data center, energy, transportation, public management, manufacturing, education, power grid, telecommunications, finance, retail, construction site, medical treatment, etc. Further, the electronic device or apparatus of the present disclosure can also be used in cloud, edge, terminal, etc. application scenarios related to artificial intelligence, big data, and / or cloud computing. In one or more embodiments, the electronic device or apparatus with high computing power according to the present disclosure can be applied to a cloud device (such as a cloud server), and the electronic device or apparatus with small power consumption can be applied to a terminal device and / or an edge device (such as a smart phone or a camera). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or the edge device are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud device according to the hardware information of the terminal device and / or the edge device to simulate the hardware resources of the terminal device and / or the edge device, so as to complete the unified management, scheduling and collaborative work of end-cloud integration or cloud-edge integration.

[0114] It should be noted that for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art can understand that the schemes of the present disclosure are not limited by the order of the described actions. Therefore, those skilled in the art can understand that some steps can be executed in other orders or simultaneously according to the disclosure or teaching of the present disclosure. Further, those skilled in the art can understand that the described embodiments of the present disclosure can be regarded as optional embodiments, i.e. the actions or modules involved therein are not necessarily essential for the implementation of one or more schemes of the present disclosure. In addition, the description of some embodiments of the present disclosure also focuses on different schemes. Therefore, those skilled in the art can understand that the parts not described in detail in one embodiment of the present disclosure can also refer to the relevant description of other embodiments.

[0115] In terms of specific implementation, based on the disclosure and teaching of the present disclosure, those skilled in the art can understand that the several embodiments disclosed by the present disclosure can also be implemented in other manners not disclosed herein. For example, in terms of the units described in the foregoing electronic device or apparatus embodiments, the units can be split into other forms in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions of the units or components can be selectively disabled. In terms of the connection relationship between the units or components, the connection discussed above can be direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling refers to the communication connection using an interface, and the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.

[0116] In the present disclosure, the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units. The foregoing components or units can be located in the same place or distributed on multiple network units. In addition, according to actual needs, part or all of the units can be selected to achieve the purpose of the scheme described in the embodiments of the present disclosure. In addition, in some scenarios, multiple units in the embodiments of the present disclosure can be integrated into one unit or physically exist separately.

[0117] In some other implementation scenarios, the above-mentioned integrated units can also be implemented in the form of hardware, i.e., specific hardware circuits, which can include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit can include but is not limited to physical devices, and the physical devices can include but are not limited to transistors or memristors, etc. In view of this, various devices described herein (e.g., computing devices or other processing devices) can be implemented by appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, ASICs, etc. Further, the aforementioned storage units or storage devices can be any appropriate storage medium (including magnetic storage medium or magneto-optical storage medium, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM, RAM, etc.

[0118] The above detailed description of the embodiments of the present disclosure is made with specific examples applied to the principles and implementation modes of the present disclosure. The above description of the embodiments is only used to help understand the method of the present disclosure and its core idea; at the same time, for those skilled in the art, according to the idea of the present disclosure, the specific implementation mode and application range will be changed, and the above description should not be understood as a limitation of the present disclosure.

Claims

1. A data processing apparatus, comprising a storage circuit and a processing circuit, wherein: The storage circuit is configured to store information, which includes at least pre-processing and / or post-processing information. The processing circuit is configured to access the storage circuit and perform a first convolution operation on the input data and the convolution kernel as follows, wherein the input data is sparse data, and each input data element has an index and a value: Extract the valid data elements corresponding to the valid output points from the input data elements, wherein the valid output point represents the convolution output point where there are input data elements at the center of its receptive field, and the valid data element is the input data element that contributes to the valid output point; as well as A second convolution operation is performed on the valid data elements and the convolution kernel to obtain each valid output point, which is used as the result of the first convolution operation. The processing circuitry is further configured to extract valid data elements as follows: Scan each of the input data elements sequentially; During the scan, it is determined whether the input data element is located at the center of the receptive field of any convolution output point; as well as In response to determining that the input data element is located at the center point of the receptive field of a convolution output point, the input data element within the receptive field of the convolution output point is extracted as the valid data element of the convolution output point.

2. The data processing apparatus according to claim 1, wherein the processing circuit is further configured to: Before scanning the input data elements, the input data elements are sorted according to their index size.

3. The data processing apparatus according to any one of claims 1-2, wherein the processing circuit is further configured to determine whether the input data element is located at the center point of the receptive field of any convolution output point as follows: Based on the input data shape, kernel shape, and stride of the first convolution operation, determine the index of the receptive field center point of all convolution output points; Compare the index of the input data element with the index of the receptive field center point; and In response to the fact that the index of the input data element is the same as the index of a receptive field center point, it is determined that the input data element is located at the center point of the receptive field of the convolution output point.

4. The data processing apparatus according to claim 3, wherein the processing circuit is further configured to: Further, the index of the receptive field center point is adjusted according to the filling rule of the first convolution operation.

5. The data processing apparatus according to any one of claims 1-2, wherein the processing circuitry is further configured to: Based on the extracted valid data elements, a sparse form of the valid data elements is generated within the receptive field of the associated convolution output point.

6. The data processing apparatus of claim 5, wherein the processing circuitry is further configured to perform a second convolution operation on the valid data elements and the convolution kernel as follows: Based on the shape of the convolution kernel, the second convolution stride of the second convolution operation is determined; Using the convolution kernel and according to the second convolution stride, the second convolution operation is performed on the effective data elements in the sparse form to obtain the operation result.

7. The data processing apparatus according to any one of claims 1-2, wherein the second convolution stride of the second convolution operation is equal to the size of the convolution kernel.

8. The data processing apparatus according to any one of claims 1-2, wherein the convolution output point of the second convolution operation corresponds one-to-one with the effective output point, and its index is associated with the index of the input data element at the center of its receptive field.

9. The data processing apparatus according to any one of claims 1-2, wherein the input data comprises multidimensional data that is sparse in multiple dimensions, and the input data is preprocessed before being provided to the processing circuit, the preprocessing comprising: Combine multiple sparse dimensions into one dimension; The sparse data points in the input data are densified in the merging dimension to form input data elements; as well as Each input data element is represented using an index and numerical information, wherein the index is a one-dimensional or multi-dimensional index, and the numerical information includes any of scalar, vector, or higher-dimensional data.

10. A chip comprising a data processing apparatus according to any one of claims 1-9.

11. A circuit board comprising the chip according to claim 10.

12. A method for processing data using the data processing apparatus according to any one of claims 1-9.

Citation Information

Patent Citations

  • Sparse tensor calculation method and device, equipment and storage medium

    CN109857744A

  • Analyzing spatially-sparse data based on sub-manifold sparse convolutional neural networks

    CN111615706A