Data processing apparatus, data processing method and related product

By introducing a sparse flag in the convolution instruction, the computation circuit is configured to perform structured sparse convolution operations, which solves the problem of supporting sparsity processing on hardware resource-constrained devices, improves processing efficiency, and reduces computation and storage requirements.

CN114692846BActive Publication Date: 2025-11-28CAMBRICON TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011566148.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-25
Publication Date
2025-11-28
Estimated Expiration
2040-12-25

AI Technical Summary

Technical Problem

Existing hardware and/or instruction sets cannot effectively support sparsification and related operations, making it difficult to apply deep learning models on computationally intensive and storage-intensive hardware resources-constrained devices.

Method used

A data processing apparatus and method are provided, which configure the arithmetic circuit to perform structured sparse convolution operations by introducing a sparse flag bit into the convolution instruction, using the sparse flag bit to indicate whether sparse processing is performed, and reusing the instruction field of the convolution instruction to simplify the processing flow.

Benefits of technology

It improves processing efficiency on hardware-restricted devices, simplifies the sparsity processing flow, and reduces computing and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692846B_ABST
    Figure CN114692846B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a data processing apparatus, a data processing method and related products. The data processing apparatus can be implemented as a computing apparatus included in a combined processing apparatus, which can further include an interface apparatus and other processing apparatuses. The computing apparatus interacts with the other processing apparatuses to jointly complete a user-specified computing operation. The combined processing apparatus can further include a storage apparatus connected with the computing apparatus and the other processing apparatuses respectively, for storing data of the computing apparatus and the other processing apparatuses. The scheme of the present disclosure provides a special instruction for structured sparse convolution operation, which can simplify processing and improve the processing efficiency of the machine.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to the field of processors. More specifically, the present disclosure relates to a data processing apparatus, a data processing method, a chip and a board card. BACKGROUND

[0002] In recent years, with the rapid development of deep learning, the algorithm performance in a series of fields such as computer vision and natural language processing has made a leap. However, deep learning algorithm is a kind of computing-intensive and storage-intensive tool. With the increasing complexity of information processing tasks, the real-time performance and accuracy of the algorithm are constantly increasing. The neural network is often designed to be deeper and deeper, so that the computing amount and storage space requirement is larger and larger, which makes it difficult for existing artificial intelligence technology based on deep learning to be directly applied to mobile phones, satellites or embedded devices with limited hardware resources.

[0003] Therefore, the compression, acceleration and optimization of deep neural network model become particularly important. A large number of studies try to reduce the computing and storage requirements of neural network without affecting the model accuracy, which has great significance for the engineering application of deep learning technology in embedded and mobile terminals. Sparsification is one of the methods for model lightening.

[0004] Network parameter sparsification is to reduce the redundant components in a large network by appropriate methods to reduce the demand of the network on computing amount and storage space. The existing hardware and / or instruction set cannot effectively support sparsification processing and operations related to sparsification processing. SUMMARY

[0005] In order to at least partially solve one or more technical problems mentioned in the background, the scheme of the present disclosure provides a data processing apparatus, a data processing method, a chip and a board card.

[0006] In a first aspect, the present disclosure discloses a data processing apparatus, comprising: a control circuit configured to parse a convolution instruction, the convolution instruction comprising a sparse flag bit for indicating whether to perform a structured sparse convolution operation; a storage circuit configured to store information before and / or after convolution; and an operation circuit configured to perform a corresponding convolution operation according to the convolution instruction.

[0007] In a second aspect, the present disclosure provides a chip comprising the data processing apparatus of any one of the preceding first aspect.

[0008] In a third aspect, the present disclosure provides a board card comprising the chip of any one of the preceding second aspect.

[0009] In a fourth aspect, the present disclosure provides a data processing method, comprising: parsing a convolution instruction, the convolution instruction comprising a sparse flag bit for indicating whether to perform a structured sparse convolution operation; reading corresponding operands according to the convolution instruction; and performing a corresponding convolution operation on the operands according to the convolution instruction.

[0010] By means of the data processing apparatus, the data processing method, the integrated circuit chip and the board card provided as above, the present disclosure provides a convolution instruction comprising a sparse flag bit for indicating whether to perform a structured sparse convolution operation. By setting the sparse flag bit, the corresponding operation circuit can be configured to perform a corresponding convolution operation according to the value of the flag bit. In some embodiments, when the sparse flag bit indicates to perform a structured sparse convolution operation, the operation circuit can be configured to perform structured sparse processing and then perform convolution on the data after sparse processing. By adding an enabling flag bit of structured sparse to the instruction domain of the convolution instruction, the processing can be simplified, thereby improving the processing efficiency of the machine. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0012] Figure 1 is a structural diagram of a board card of an embodiment of the present disclosure;

[0013] Figure 2 is a structural diagram of a combined processing apparatus of an embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram of the internal structure of a single-core computing apparatus of an embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of the internal structure of a multi-core computing apparatus of an embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram of the internal structure of a processor core of an embodiment of the present disclosure;

[0017] Figure 6 is a structural diagram of a data processing apparatus of an embodiment of the present disclosure;

[0018] Figures 7A-7C is a partial structural diagram of an operation circuit of an embodiment of the present disclosure;

[0019] Figure 8An exemplary flowchart illustrating a data processing method according to an embodiment of the present disclosure.

[0020] Figure 9 An exemplary flowchart illustrating a data processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of, rather than all of, the embodiments of the present disclosure. Based on the embodiments of the present disclosure, any other embodiments obtained by a person of ordinary skill in the art without creative effort fall within the protection scope of the present disclosure.

[0022] It should be understood that the terms "first", "second", "third", and "fourth" and the like in the claims, the specification and the drawings of the present disclosure are used to distinguish different objects, and are not used to describe a particular order. The terms "comprise", "comprising", "include", "including" and "contains" used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0023] It should also be understood that the terms used in the specification of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used in the specification and claims of the present disclosure, the singular forms "a", "an" and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should be further understood that the term "and / or" used in the specification and claims of the present disclosure means any combination of one or more of the associated listed items and all possible combinations thereof, and includes these combinations.

[0024] As used in the specification and claims of this document, the term "if' can be interpreted as meaning "when" or "once" or "in response to a determination" or "in response to detecting" depending on the context.

[0025] The specific embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0026] Figure 1 A structural schematic diagram of a board card 10 according to an embodiment of the present disclosure is shown. As shown in FIG. 1, the board card 10 includes a plurality of processing units 100, a plurality of memory units 200, a plurality of input / output units 300, and a plurality of bus units 400. Figure 1As shown, the board card 10 includes a chip 101, which is a system on chip (SoC) integrated with one or more combined processing devices, which is an artificial intelligence operation unit to support various deep learning and machine learning algorithms to meet the intelligent processing needs in complex scenarios in the fields of computer vision, speech, natural language processing, data mining, etc. In particular, deep learning technology is widely used in cloud intelligent fields. A significant feature of cloud intelligent applications is the large amount of input data, which has high requirements for the storage capacity and computing capacity of the platform. The board card 10 of this embodiment is suitable for cloud intelligent applications and has a large off-chip storage, on-chip storage and strong computing capacity.

[0027] The chip 101 is connected with an external device 103 through an external interface device 102. The external device 103 is, for example, a server, a computer, a camera, a display, a mouse, a keyboard, a network card or a wifi interface, etc. The data to be processed can be transmitted from the external device 103 to the chip 101 through the external interface device 102. The computing result of the chip 101 can be transmitted back to the external device 103 through the external interface device 102. According to different application scenarios, the external interface device 102 can have different interface forms, such as a PCIe interface, etc.

[0028] The board card 10 further includes a storage device 104 for storing data, which includes one or more storage units 105. The storage device 104 is connected and transmits data with the control device 106 and the chip 101 through a bus. The control device 106 in the board card 10 is configured to regulate the state of the chip 101. For this purpose, in one application scenario, the control device 106 can include a micro controller unit (MCU).

[0029] Figure 2 is a structural diagram of the combined processing device in the chip 101 of this embodiment. As Figure 2 As shown in the figure, the combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203 and a storage device 204.

[0030] The computing device 201 is configured to perform user-specified operations, mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations, which can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.

[0031] The interface device 202 is used to transmit data and control instructions between the computing device 201 and the processing device 203. For example, the computing device 201 can obtain input data from the processing device 203 via the interface device 202 and write into the storage device on the computing device 201. Further, the computing device 201 can obtain control instructions from the processing device 203 via the interface device 202 and write into the control buffer on the computing device 201. Alternatively or additionally, the interface device 202 can also read data from the storage device of the computing device 201 and transmit to the processing device 203.

[0032] The processing device 203 is a general-purpose processing device, which performs basic controls including but not limited to data transfer, start and / or stop of the computing device 201, etc. Depending on the implementation, the processing device 203 can be one or more types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), or other general-purpose and / or special-purpose processors, including but not limited to a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc., and the number thereof can be determined according to actual needs. As mentioned above, only in terms of the computing device 201 of the present disclosure, it can be considered as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 201 and the processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure.

[0033] The storage device 204 is used to store data to be processed, which can be a DRAM, a DDR memory, usually with a size of 16G or more, for saving data of the computing device 201 and / or the processing device 203.

[0034] Figure 3 An internal structure diagram of the computing device 201 as a single-core is shown. The single-core computing device 301 is used to process input data such as computer vision, voice, natural language, data mining, etc., and the single-core computing device 301 includes three major modules: a control module 31, a computation module 32, and a storage module 33.

[0035] The control module 31 is configured to coordinate and control the operation of the operation module 32 and the storage module 33 to complete the task of deep learning, which includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 is configured to fetch instructions from the processing device 203, and the instruction decode unit 312 decodes the fetched instructions and sends the decoding results as control information to the operation module 32 and the storage module 33.

[0036] The operation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is configured to perform vector operations, which can support complex operations such as vector multiplication, addition, and nonlinear transformation; and the matrix operation unit 322 is responsible for the core calculation of the deep learning algorithm, i.e., matrix multiplication and convolution.

[0037] The storage module 33 is configured to store or transfer related data, which includes a neuron RAM (NRAM) 331, a weight RAM (WRAM) 332, and a direct memory access (DMA) 333. The NRAM 331 is configured to store input neurons, output neurons, and intermediate results after calculation; the WRAM 332 is configured to store the convolution kernel of the deep learning network, i.e., the weight; and the DMA 333 is connected to the DRAM 204 through the bus 34 and is responsible for data transfer between the single-core computing device 301 and the DRAM 204.

[0038] Figure 4 An internal structure schematic diagram of the multi-core computing device 41 is shown. The multi-core computing device 41 adopts a hierarchical structure design. The multi-core computing device 41 is a system on chip, which includes at least one cluster, and each cluster includes a plurality of processor cores. In other words, the multi-core computing device 41 is composed of a system on chip-cluster-processor core hierarchy.

[0039] From the perspective of the system on chip hierarchy, as shown in Figure 4 The multi-core computing device 41 includes an external storage controller 401, a peripheral communication module 402, an on-chip interconnection module 403, a synchronization module 404, and a plurality of clusters 405.

[0040] The external storage controller 401 can be multiple, and two are exemplarily shown in the figure. The external storage controller 401 is configured to respond to an access request issued by the processor core to access an external storage device, such as a DRAM 204. Figure 2The DRAM 204 in the chip allows data to be read from or written to external devices. The peripheral communication module 402 receives control signals from the processing device 203 via the interface device 202, initiating the computing device 201 to execute tasks. The on-chip interconnect module 403 connects the external storage controller 401, the peripheral communication module 402, and multiple clusters 405 to transmit data and control signals between modules. The synchronization module 404 is a global barrier controller (GBC) used to coordinate the working progress of each cluster and ensure information synchronization. The multiple clusters 405 are the computing cores of the multi-core computing device 41. Four are shown exemplary in the figure; however, with hardware development, the multi-core computing device 41 disclosed herein may also include 8, 16, 64, or even more clusters 405. The clusters 405 are used to efficiently execute deep learning algorithms.

[0041] From the perspective of cluster hierarchy, such as Figure 4 As shown, each cluster 405 includes multiple processor cores (IPU cores) 406 and one memory core (MEM core) 407.

[0042] Four processor cores 406 are shown in the figure as an example; this disclosure does not limit the number of processor cores 406. Its internal architecture is as follows: Figure 5 As shown. Each processor core 406 is similar to Figure 3 The single-core computing device 301 also includes three main modules: a control module 51, an arithmetic module 52, and a storage module 53. The functions and structures of the control module 51, arithmetic module 52, and storage module 53 are largely the same as those of the control module 31, arithmetic module 32, and storage module 33, and will not be described again. It should be noted that the storage module 53 includes an input / output direct memory access (IODMA) module 533 and a move direct memory access (MVDMA) module 534. The IODMA 533 controls the memory access of NRAM 531 / WRAM 532 and DRAM 204 via the broadcast bus 409; the MVDMA 534 controls the memory access of NRAM 531 / WRAM 532 and SRAM 408.

[0043] Back Figure 4The storage core 407 is mainly used to store and communicate, i.e., store shared data or intermediate results among the processor cores 406, perform communication between the execution cluster 405 and the DRAM 204, perform communication among the clusters 405, perform communication among the processor cores 406, etc. In other embodiments, the storage core 407 has the capability of scalar operations to perform scalar operations.

[0044] The storage core 407 includes an SRAM 408, a broadcast bus 409, a cluster direct memory access (CDMA) 410, and a global direct memory access (GDMA) 411. The SRAM 408 serves as a high-performance data relay station. Data reused among different processor cores 406 within the same cluster 405 does not need to be obtained by the processor cores 406 from the DRAM 204 individually, but is relayed among the processor cores 406 through the SRAM 408. The storage core 407 only needs to quickly distribute the reused data from the SRAM 408 to the multiple processor cores 406, so as to improve the efficiency of inter-core communication and greatly reduce on-chip and off-chip input / output access.

[0045] The broadcast bus 409, the CDMA 410, and the GDMA 411 are respectively used to perform communication among the processor cores 406, communication among the clusters 405, and data transmission between the clusters 405 and the DRAM 204. The following will be described respectively.

[0046] The broadcast bus 409 is used to complete high-speed communication among the processor cores 406 within the cluster 405. The broadcast bus 409 of the embodiment supports inter-core communication modes including unicast, multicast, and broadcast. Unicast refers to point-to-point (e.g., single processor core to single processor core) data transmission, multicast refers to a communication mode of transmitting a piece of data from the SRAM 408 to specific processor cores 406, and broadcast refers to a communication mode of transmitting a piece of data from the SRAM 408 to all processor cores 406, which is a special case of multicast.

[0047] The CDMA 410 is used to control access to the SRAM 408 among different clusters 405 within the same computing device 201.

[0048] The GDMA 411 cooperates with the external memory controller 401 to control the access of the SRAM 408 to the DRAM 204 or the reading of data from the DRAM 204 to the SRAM 408. As mentioned above, the communication between the DRAM 204 and the NRAM 431 or the WRAM 432 can be achieved through two channels. The first channel is to directly contact the DRAM 204 and the NRAM 431 or the WRAM 432 through the IODMA 433; the second channel is to first transfer data between the DRAM 204 and the SRAM 408 through the GDMA 411, and then transfer data between the SRAM 408 and the NRAM 431 or the WRAM 432 through the MVDMA 534. Although the second channel seems to involve more components and the data flow is longer, in some embodiments, the bandwidth of the second channel is much greater than that of the first channel, and thus the communication between the DRAM 204 and the NRAM 431 or the WRAM 432 through the second channel can be more efficient. The embodiments of the present disclosure can select the data transmission channel according to the hardware conditions.

[0049] In other embodiments, the functions of the GDMA 411 and the functions of the IODMA 533 can be integrated in the same component. For the convenience of description, the GDMA 411 and the IODMA 533 are regarded as different components in the present disclosure, and for those skilled in the art, as long as the functions achieved and the technical effects achieved are similar to those of the present disclosure, they belong to the protection scope of the present disclosure. Further, the functions of the GDMA 411, the functions of the IODMA 533, the functions of the CDMA 410, and the functions of the MVDMA 534 can also be achieved by the same component.

[0050] Based on the hardware environment described above, one embodiment of the present disclosure provides a data processing scheme for performing structured sparse convolution operation according to the sparse flag included in the convolution instruction.

[0051] Figure 6 A structural block diagram of a data processing apparatus 600 according to an embodiment of the present disclosure is shown. The data processing apparatus 600 can be implemented in, for example, the computing device 201 of Figure 2 As shown in the figure, the data processing apparatus 600 can include a control circuit 610, a storage circuit 620, and an operation circuit 630.

[0052] The functions of the control circuit 610 can be similar to those of the control module 31 of Figure 3 or the control module 51 of Figure 5 which can include, for example, an instruction fetching unit to fetch instructions from, for example, the memory 32 of Figure 2instructions of the processing device 203, and an instruction decoding unit configured to decode the obtained instructions and send the decoding results as control information to the operation circuit 630 and the storage circuit 620.

[0053] In one embodiment, the control circuit 610 can be configured to parse a convolution instruction, wherein the convolution instruction includes a sparse flag bit indicating whether to perform a structured sparse convolution operation. In one implementation, the sparse flag bit can take a value of "1" to indicate that the current convolution instruction performs a structured sparse convolution operation; correspondingly, the sparse flag bit can take a value of "0" to indicate that the current convolution instruction performs a regular convolution operation; vice versa.

[0054] The storage circuit 620 can be configured to store pre- and / or post-convolution information. In one embodiment, the operands of the convolution instruction include weight values of a convolution layer in a neural network and neuron data of the neural network. In this embodiment, the storage circuit can include, for example, Figure 3 a WRAM 332 of the processing device 203, or Figure 5 a WRAM 532 of the processing device 502, configured to store the weight values; and Figure 3 a NRAM 331 of the processing device 203, or Figure 5 a NRAM 531 of the processing device 502, configured to store the neuron data.

[0055] The operation circuit 630 can be configured to perform a corresponding convolution operation according to the convolution instruction.

[0056] In some embodiments, the operation circuit 630 can include a structured sparse circuit 632 and a convolution circuit 633.

[0057] When the sparse flag bit in the convolution instruction indicates that the current convolution instruction needs to perform a structured sparse convolution operation, the operation circuit 630 can be configured accordingly. For example, the structured sparse circuit 632 in the operation circuit 630 can be configured to perform structured sparse processing on at least one input data and output the sparsified input data to the convolution circuit 633. The convolution circuit 633 can be configured to receive data to be convolved and perform a convolution operation thereon. The data to be convolved includes at least the sparsified input data received from the structured sparse circuit 632. Thus, by the structured sparse circuit 632 and the convolution circuit 633, a structured sparse convolution processing can be implemented when the sparse flag bit is set.

[0058] The structured sparse circuit 632 is configured to perform structured sparse processing, which includes selecting n data elements as valid data elements from every m data elements, where m > n. In one implementation, m = 4 and n = 2. In other implementations, m = 4 and n can also take other values, such as 1 or 3.

[0059] The convolution circuit 633 is configured to perform convolution operation on input data. When the sparse flag is "1", i.e., the structured sparse convolution operation is performed, the data received by the convolution circuit 633 includes at least the sparse input data from the structured sparse circuit 632. When the sparse flag is "0", i.e., the regular convolution operation is performed, the data received by the convolution circuit 633 is the non-sparse input data.

[0060] Depending on different application scenarios, the input data to be convolved can have various forms, and thus the structured sparse circuit and the convolution circuit can need to perform structured sparse convolution processing according to different requirements.

[0061] Figures 7A-7C A partial structural diagram of the operation circuit of the embodiments of the present disclosure is shown. As shown in the figure, the structured sparse circuit 710 can include a first structured sparse sub-circuit 712 and / or a second structured sparse sub-circuit 714. The first structured sparse sub-circuit 712 can be configured to perform structured sparse processing on input data according to a specified sparse mask. The second structured sparse sub-circuit 714 can be configured to perform structured sparse processing on input data according to a predetermined sparse rule.

[0062] In the first scenario, one of the data to be convolved (not assumed to be the first data) can have been previously subjected to structured sparse processing, and the other of the data to be convolved (assumed to be the second data) needs to be subjected to structured sparse according to the sparse manner of the first data. At this time, the first structured sparse sub-circuit can be used to perform sparse processing on the second data.

[0063] Figure 7AA partial structure diagram of the operation circuit in the first scenario is shown. As shown in the figure, in this first scenario, the structured sparse circuit 710 includes a first structured sparse sub-circuit 712 which receives the second data and the index part of the first data which has been previously structured sparse. The first structured sparse sub-circuit 712 uses the index part of the first data as a sparse mask to perform structured sparse processing on the second data. Specifically, the first structured sparse sub-circuit 712 extracts the data at the corresponding positions from the first data as valid data according to the valid data positions indicated by the index part of the first data. In these embodiments, the first structured sparse sub-circuit 712 can be implemented by, for example, a vector multiplication or matrix multiplication circuit or the like. The convolution circuit 720 receives the first data which has been structured sparse and the second data which has been processed by the first structured sparse sub-circuit 712 and performs convolution on the two data.

[0064] In the second scenario, neither of the two data to be convolved has been structured sparse processed, and it is necessary to perform structured sparse processing on the two data respectively before convolution. At this time, the second structured sparse sub-circuit can be used to perform sparse processing on the first data and the second data.

[0065] Figure 7B A partial structure diagram of the operation circuit in the second scenario is shown. As shown in the figure, in this second scenario, the structured sparse circuit 710 can include two second structured sparse sub-circuits 714 which respectively receive the first data and the second data to be convolved, so as to simultaneously and independently perform structured sparse processing on the first data and the second data respectively, and output the sparse processed data to the convolution circuit 720. The second structured sparse sub-circuit 714 can be configured to perform structured sparse processing according to a predetermined screening rule, for example, according to a rule of screening n data elements with larger absolute values from every m data elements as valid data elements. In these embodiments, the second structured sparse sub-circuit 714 can be implemented by, for example, configuring a multi-stage operation pipeline composed of comparators and the like to implement the above processing. Those skilled in the art can understand that the structured sparse circuit 710 can also include only one second structured sparse sub-circuit 714 which sequentially performs structured sparse processing on the first data and the second data.

[0066] In the third scenario, neither of the two data to be convolved has been structured sparse processed, and it is necessary to perform structured sparse processing on the two data respectively before convolution, and one of the data (for example, the first data) needs to use the index part of the sparse processed other data (for example, the second data) as a sparse mask.

[0067] Figure 7CA partial structure diagram of an operation circuit in a third scenario is shown. As shown, in this third scenario, the structured sparse circuit 710 can include a first structured sparse sub-circuit 712 and a second structured sparse sub-circuit 714. The second structured sparse sub-circuit 714 can be configured to perform structured sparse processing on, for example, the second data according to a predetermined screening rule, and provide an index part of the sparse second data to the first structured sparse sub-circuit 712. The first structured sparse sub-circuit 712 uses the index part of the second data as a sparse mask to perform structured sparse processing on the first data. The convolution circuit 720 receives the structured sparse first data and the second data from the first structured sparse sub-circuit 712 and the second structured sparse sub-circuit 714, respectively, and performs convolution on both.

[0068] Other application scenarios can also be considered by those skilled in the art, and the structured sparse circuit can be designed accordingly. For example, it can be required to apply the same sparse mask to both data to be convolved, in which case two first structured sparse sub-circuits can be included in the structured sparse circuit to process.

[0069] Figure 8 An exemplary operation pipeline of structured sparse processing according to an embodiment of the present disclosure is shown. The pipeline can be used to implement, for example, the second structured sparse sub-circuit described above. In the pipeline, the first data and the second data are input to the pipeline, and the pipeline performs structured sparse processing on the first data and the second data according to a predetermined screening rule. Figure 8 In an embodiment, when m = 4 and n = 2, structured sparse processing of screening 2 data elements with larger absolute values from 4 data elements A, B, C and D is shown. As shown, the structured sparse processing can be performed by a multi-stage pipelined operation circuit composed of absolute value operators, comparators, etc.

[0070] The first pipeline stage can include m (4) absolute value operators 810 for synchronously performing absolute value operations on the 4 input data elements A, B, C and D, respectively. In order to facilitate the output of valid data elements at the end, in some embodiments, the first pipeline stage outputs both the original data elements (i.e., A, B, C and D) and the data after the absolute value operation (i.e., |A|, |B|, |C| and |D|) at the same time.

[0071] The second pipeline stage can include a permutation and combination circuit 820 for performing permutation and combination on the m absolute values to generate m groups of data, each of which includes the m absolute values and the positions of the m absolute values in each group are different.

[0072] In some embodiments, the permutation and combination circuit can be a cyclic shifter, which performs m-1 cyclic shifts on the permutation of the m absolute values (e.g., |A|, |B|, |C| and |D|), thereby generating m groups of data. For example, in the example shown in the figure, 4 groups of data are generated, which are: {|A|, |B|, |C|, |D|}, {|B|, |C|, |D|, |A|}, {|C|, |D|, |A|, |B|} and {|D|, |A|, |B|, |C|}, respectively. Similarly, the original data elements corresponding to each group of data are also outputted at the same time, one original data element for each group of data.

[0073] The third pipeline stage includes a comparison circuit 830 for comparing the absolute values in the m groups of data and generating comparison results.

[0074] In some embodiments, the third pipeline stage can include m comparison circuits, each of which includes m-1 comparators (831, 832, 833), and the m-1 comparators in the i-th comparison circuit are used to compare one absolute value in the i-th group of data with the other three absolute values in turn and generate comparison results, where 1≤i≤m.

[0075] As can be seen from the figure, the third pipeline stage can also be considered as m-1 (3) sub-pipeline stages. Each sub-pipeline stage includes m comparators for comparing the corresponding one absolute value with the other absolute values. The m-1 sub-pipeline stages are used to compare the corresponding one absolute value with the other m-1 absolute values in turn.

[0076] For example, in the example shown in the figure, the 4 comparators 831 in the first sub-pipeline stage are used to compare the first absolute value with the second absolute value in the 4 groups of data respectively, and output comparison results w0, x0, y0 and z0, respectively. The 4 comparators 832 in the second sub-pipeline stage are used to compare the first absolute value with the third absolute value in the 4 groups of data respectively, and output comparison results w1, x1, y1 and z1, respectively. The 4 comparators 833 in the third sub-pipeline stage are used to compare the first absolute value with the fourth absolute value in the 4 groups of data respectively, and output comparison results w2, x2, y2 and z2, respectively. In this way, the comparison results of each absolute value with the other m-1 absolute values can be obtained.

[0077] In some embodiments, the comparison results can be represented using bitmaps. For example, at the first comparator of the first path comparison circuit, w0 = 1 when |A| ≥ |B|; at the second comparator of the first path, w1 = 0 when |A| < |C|; at the third comparator of the first path, w2 = 1 when |A| ≥ |D|. Thus, the output of the first path comparison circuit is {A, w0, w1, w2}, which is {A, 1, 0, 1} in this case. Similarly, the output of the second path comparison circuit is {B, x0, x1, x2}, the output of the third path comparison circuit is {C, y0, y1, y2}, and the output of the fourth path comparison circuit is {D, z0, z1, z2}.

[0078] The fourth pipeline stage includes a screening circuit 840 for selecting n data elements with larger absolute values from the m data elements as valid data elements according to the comparison results of the third stage, and outputting the valid data elements and corresponding indexes. The indexes are used to indicate the positions of the valid data elements in the input m data elements. For example, when A and C are selected from the four data elements A, B, C, and D, the corresponding indexes can be 0 and 2.

[0079] According to the comparison results, appropriate logic can be designed to select the n data elements with larger absolute values. Considering that there can be multiple data elements with the same absolute value, in further embodiments, when there are data elements with the same absolute value, selection is made according to a specified priority order. For example, the priority of A can be set to be the highest and the priority of D can be set to be the lowest in a fixed priority order from low to high index. In one example, when the absolute values of A, C, and D are all the same and larger than the absolute value of B, the selected data are A and C.

[0080] From the foregoing comparison results, it can be seen that |A| is larger than the absolute values of |B|, |C|, and |D| according to w0, w1, and w2. If w0, w1, and w2 are all 1, it indicates that |A| is larger than |B|, |C|, and |D|, and is the largest among the four numbers, so A is selected. If two of w0, w1, and w2 are 1, it indicates that |A| is the second largest among the four absolute values, so A is also selected. Otherwise, A is not selected. Thus, in some embodiments, the numbers can be analyzed and determined according to the number of occurrences.

[0081] In one implementation, the valid data elements can be selected based on the following logic. First, the number of times each data is larger than the other data can be counted. For example, define N A = sum_w = w0 + w1 + w2, N B = sum_x = x0 + x1 + x2, N C = sum_y = y0 + y1 + y2, N D= sum_z = z0 + z1 + z2. Then, the selection is determined by the following conditions.

[0082] The condition for selection A is: N A = 3, or N A = 2 and N B / N C / N D has only one 3;

[0083] The condition for selection B is: N B = 3, or N B = 2 and N A / N C / N D has only one 3, and N A ≠ 2;

[0084] The condition for selection C is: N C = 3, and N A / N B has at most one 3, or N C = 2 and N A / N B / N D has only one 3, and N A / N B has no 2;

[0085] The condition for selection D is: N D = 3, and N A / N B / N C has at most one 3, or N D = 2 and N A / N B / N C has only one 3, and N A / N B / N C has no 2.

[0086] Those skilled in the art can understand that there is a certain redundancy in the above logic in order to ensure the selection according to the predetermined priority. Based on the size and order information provided by the comparison result, those skilled in the art can design other logic to realize the screening of effective data elements, and the present disclosure is not limited in this respect. Thus, the multi-stage pipelined operation circuit Figure 8 can realize a four-to-two structured sparse processing.

[0087] Those skilled in the art can understand that other forms of pipelined operation circuits can also be designed to realize structured sparse processing, and the present disclosure is not limited in this respect.

[0088] The result of the sparsification processing includes two parts: a data part and an index part. The data part includes the data after the sparsification processing, i.e., the valid data elements extracted according to the filtering rule of the structured sparsification processing. The index part is used to indicate the data after the sparsification, i.e., the original position of the valid data elements in the original data (i.e., the data to be sparsified) before the sparsification.

[0089] The structured sparsification processed data can be represented and / or stored in various forms. In one implementation, the structured sparsification processed data can be in the form of a structure. In the structure, the data part and the index part are bound to each other. In some embodiments, each bit in the index part can correspond to one data element. For example, when the data type is fix8, one data element is 8 bits, then each bit in the index part can correspond to 8 bits of data. In other embodiments, considering the hardware level implementation when using the structure later, each bit in the index part in the structure can be set to correspond to the position of N bits of data, N being determined at least partially based on the hardware configuration. For example, it can be set that each bit in the index part in the structure corresponds to the position of 4 bits of data. For example, when the data type is fix8, each 2 bits in the index part corresponds to one data element of fix8 type. In some embodiments, the data part in the structure can be aligned according to a first alignment requirement, the index part in the structure can be aligned according to a second alignment requirement, so that the entire structure also satisfies the alignment requirement. For example, the data part can be aligned according to 64B, the index part can be aligned according to 32B, and the entire structure is aligned according to 96B (64B+32B). Through such alignment requirements, the number of memory accesses can be reduced when used later, improving processing efficiency.

[0090] By using such a structure, the data part and the index part can be used uniformly. Since in the structured sparsification processing, the proportion of valid data elements occupying the original data elements is fixed, e.g., n / m, the size of the data after the sparsification processing is also fixed or predictable. Therefore, the structure can be stored densely in the storage circuit without performance loss.

[0091] In other implementations, the data part and the index part obtained after the sparsification processing can also be represented and / or stored separately for separate use. For example, the index part of the second input data of the structured sparsification processing can be provided to the first structured sparsification circuit 712 to be used as a mask to perform structured sparsification processing on the first input data. At this time, in order to use different data types, each bit in the separately provided index part can indicate whether a data element is valid.

[0092] The convolution circuit can be implemented using a variety of circuit configurations to perform convolution operations. For example, the convolution circuit can share the same processing circuit for both regular convolution and structured sparse convolution, or the convolution circuit can be configured with separate processing circuits for regular convolution and structured sparse convolution operations. The present embodiments are not limited in this respect.

[0093] Returning to Figure 6 In some embodiments, the arithmetic circuit 630 can further include a pre-processing circuit 631 and a post-processing circuit 634. The pre-processing circuit 631 can be configured to pre-process data before performing operations on the structured sparse circuit 632 and / or the convolution circuit 633 according to the instructions, and the post-processing circuit 634 can be configured to post-process data after the convolution circuit 633 performs operations.

[0094] In some implementations, when the sparse flag in the convolution instruction indicates performing a structured sparse convolution operation, the pre-processing circuit 631 can read input data from the storage circuit 620 and output the input data to the structured sparse circuit 632 at a first rate. When the sparse flag indicates regular convolution, the pre-processing circuit 631 can read input data from the storage circuit 620 and output the input data to the convolution circuit 633 at a second rate. The first rate is greater than the second rate, for example, by a ratio equal to the sparsity ratio in the structured sparse processing, such as m / n. For example, in a 4-of-2 structured sparse processing, the first rate is twice the second rate. Thus, the first rate is determined based at least in part on the processing capability of the convolution circuit 633 and the sparsity ratio of the structured sparse processing.

[0095] In some application scenarios, the aforementioned pre-processing and post-processing can further include, for example, data splitting and / or data stitching operations. For example, the post-processing circuit 634 can perform fusion processing on the output results of the convolution circuit, such as addition, subtraction, multiplication, etc.

[0096] As mentioned earlier, the operands of convolution instructions can be data within a neural network, such as weights and neurons. In other words, convolution instructions are used for structured sparse convolution operations within neural networks. Data in neural networks typically contains multiple dimensions. For example, in a convolutional neural network, the data may have four dimensions: input channels, output channels, length, and width. In some embodiments, the structured sparsity in the aforementioned convolution instructions can be performed on at least one dimension of the multidimensional data in the neural network. Specifically, in one implementation, convolution instructions can be used for structured sparse convolution operations in the forward process of a neural network (e.g., inference or forward training), where the structured sparsity processing is performed on the input channel dimension of the multidimensional data in the neural network. In another implementation, convolution instructions can be used for structured sparse convolution operations in the reverse process of a neural network (e.g., reverse training), where the structured sparsity processing is performed simultaneously on both the input and output channel dimensions of the multidimensional data in the neural network.

[0097] In the context of this disclosure, the aforementioned convolution instruction may be a microinstruction or control signal that runs within one or more multi-stage computational pipelines, and may include (or instruct) one or more computational operations that need to be performed by the multi-stage computational pipelines.

[0098] Figure 9 An exemplary flowchart of a data processing method 900 according to an embodiment of this disclosure is shown.

[0099] like Figure 9 As shown, in step 910, the convolution instruction is parsed. This convolution instruction includes a sparse flag to indicate whether a structured sparse convolution operation is performed. This step can be, for example, performed by... Figure 6 The control circuit 610 is used to execute this.

[0100] Next, in step 920, the corresponding operands are read according to the convolution instructions. This step can be, for example, performed by... Figure 6 The control circuit 610 controls the storage circuit 620 to execute.

[0101] Finally, in step 930, the corresponding convolution operation is performed on the read operands according to the convolution instruction. This step can be, for example, by... Figure 6 The operation is performed by the 630 arithmetic circuit.

[0102] Different convolution operations can be performed depending on the value of the sparse flag in the convolution instruction. For example, when the sparse flag is set to "0", a regular convolution operation can be performed. When the sparse flag is set to "1", a structured sparse convolution operation is performed.

[0103] At this time, the structured sparse convolution operation can include performing a structured sparse processing on the at least one input data, and performing a convolution operation on the input data after the sparse processing.

[0104] In particular, in some implementations, performing the structured sparse processing on the at least one input data includes any one of the following: performing the structured sparse processing on the first input data and the second input data to be convolved respectively, and outputting the first input data and the second input data after the sparse processing to the convolution circuit to perform the convolution operation; or performing the structured sparse processing on the first or the second input data, and outputting the first input data after the sparse processing to the convolution circuit to perform the convolution operation with the second input data after the sparse processing, wherein the index part corresponding to the first or the second input data after the structured sparse processing is used as a sparse mask to perform the structured sparse processing on the second or the first input data, and the index part indicates the positions of the valid data elements to be performed in the structured sparse processing. Accordingly, the structured sparse processing includes selecting n data elements from every m data elements as valid data elements according to the indication of the index part, where m>n.

[0105] The first or the second input data after the structured sparse processing can be pre-processed and stored in a storage circuit, or the first or the second input data after the structured sparse processing can be directly provided to the convolution operation after being processed online.

[0106] The data after the structured sparse processing can be provided in various forms. In one implementation, the data after the structured sparse processing is in the form of a structure, and the structure includes a data part and an index part bound to each other, the data part includes the valid data elements after the structured sparse processing, and the index part is used to indicate the original positions of the data after the sparse processing in the data before the sparse processing.

[0107] Further, before performing the structured sparse processing, the method can further include: delivering the input data at a first rate to perform the structured sparse processing, wherein the first rate is determined based at least in part on the processing capability of the hardware performing the convolution operation and the sparse ratio of the structured sparse processing.

[0108] In some embodiments, the first input data described above can be neuron data of a neural network, and the second input data can be weight values of a convolution layer in the neural network; or vice versa.

[0109] Those skilled in the art can understand that the steps described in the method flowchart are the same as the steps described in the foregoing Figure 6 Figure 6 The various circuits of the data processing apparatus described correspondingly, and therefore the features described above are also applicable to the method steps, which are not repeated here.

[0110] From the above description, the embodiments of the present disclosure provide a convolution instruction, which includes a sparse flag bit for indicating whether to perform a structured sparse convolution operation. By setting the sparse flag bit, the corresponding operation circuit can be configured to perform the corresponding convolution operation according to the value of the flag bit. In some embodiments, when the sparse flag bit indicates to perform a structured sparse convolution operation, the operation circuit can be configured to perform structured sparse processing and then perform convolution on the data after the sparse processing. By adding an enabling flag bit of structured sparse to the instruction domain of the convolution instruction, the processing can be simplified, thereby improving the processing efficiency of the machine.

[0111] According to different application scenarios, the electronic device or apparatus of the present disclosure can include a server, a cloud server, a server cluster, a data processing apparatus, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a PC device, an Internet of Things terminal, a mobile terminal, a mobile phone, a vehicle record instrument, a navigation instrument, a sensor, a camera, a camera, a video camera, a projector, a watch, a headset, a mobile storage, a wearable device, a visual terminal, an autonomous driving terminal, a vehicle, a household appliance, and / or a medical device. The vehicle includes an airplane, a ship, and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, an electric rice cooker, a humidifier, a washing machine, an electric lamp, a gas stove, an extractor hood; the medical device includes a nuclear magnetic resonance instrument, a B-ultrasound instrument, and / or an electrocardiograph. The electronic device or apparatus of the present disclosure can also be applied to the fields of Internet, Internet of Things, data center, energy, transportation, public management, manufacturing, education, power grid, telecommunications, finance, retail, construction site, medical treatment, etc. Further, the electronic device or apparatus of the present disclosure can also be used in cloud, edge, terminal, etc. application scenarios related to artificial intelligence, big data, and / or cloud computing. In one or more embodiments, the electronic device or apparatus with high computing power according to the present disclosure scheme can be applied to a cloud device (such as a cloud server), and the electronic device or apparatus with small power consumption can be applied to a terminal device and / or an edge device (such as a smart phone or a camera). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or the edge device are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud device according to the hardware information of the terminal device and / or the edge device to simulate the hardware resources of the terminal device and / or the edge device, so as to complete the unified management, scheduling and collaborative work of end-cloud integration or cloud-edge integration.

[0112] It should be noted that, for the purpose of clarity, the disclosure describes some methods and embodiments thereof as a series of acts and / or combinations thereof, but those skilled in the art will understand that the present disclosure is not limited to the order of the acts described. Those skilled in the art will understand and appreciate that some steps of the methods can be decided to be executed in other orders or at the same time with other steps. Further, those skilled in the art will understand and appreciate that some of the embodiments described in the disclosure can be considered optional, i.e., the acts or modules involved therein are not necessarily essential for the implementation of one or more of the aspects of the present disclosure. In addition, the disclosure describes some embodiments with different focuses according to different aspects. In view of this, those skilled in the art will understand that the parts not described in detail in some embodiments of the disclosure can also be seen from the relevant description of other embodiments.

[0113] In specific implementation aspects, based on the disclosure and teachings of the present disclosure, those skilled in the art will understand that some of the embodiments disclosed in the present disclosure can also be implemented in other ways not disclosed herein. For example, as for each unit in the electronic device or apparatus embodiments described above, the units are split based on the logical functions considered herein, and there can be other splitting manners in actual implementation. For another example, a plurality of units or components can be combined or integrated into another system, or some features or functions of the units or components can be selectively disabled. As for the connection relationship between different units or components, the connections discussed above in conjunction with the drawings can be direct or indirect coupling between the units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, where the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.

[0114] In the present disclosure, the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units. The aforementioned components or units can be located in the same place or distributed on a plurality of network units. In addition, according to actual needs, some or all of the units can be selected to achieve the purpose of the aspects described in the embodiments of the present disclosure. In addition, in some scenarios, a plurality of units in the embodiments of the present disclosure can be integrated into one unit or each unit physically exists separately.

[0115] In some other implementation scenarios, the above-mentioned integrated units can also be implemented in the form of hardware, i.e., specific hardware circuits, which can include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit can include but is not limited to physical devices, and the physical devices can include but are not limited to transistors or memristors, etc. In view of this, various apparatuses (such as computing apparatuses or other processing apparatuses) described herein can be implemented by appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs, etc. Further, the aforementioned storage units or storage devices can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM, and RAM, etc.

[0116] The foregoing can be better understood in light of the following clauses:

[0117] Clause 1, a data processing apparatus, comprising:

[0118] a control circuit configured to parse a convolution instruction, the convolution instruction comprising a sparse flag indicating whether to perform a structured sparse convolution operation;

[0119] a storage circuit configured to store pre-convolution and / or post-convolution information; and

[0120] an operation circuit configured to perform a corresponding convolution operation according to the convolution instruction.

[0121] Clause 2, the data processing apparatus according to clause 1, wherein the operation circuit comprises a structured sparse circuit and a convolution circuit, when the sparse flag indicates to perform a structured sparse convolution operation,

[0122] the structured sparse circuit is configured to perform structured sparse processing on at least one input data and output the sparse input data to the convolution circuit; and

[0123] The convolution circuit is configured to receive data to be convolved and perform convolution operation on the data, wherein the data to be convolved comprises at least the sparse input data.

[0124] Clause 3, the data processing apparatus according to clause 2, wherein the structured sparse circuit comprises:

[0125] a first structured sparse sub-circuit configured to perform structured sparse processing on input data according to a specified sparse mask; and / or

[0126] a second structured sparse sub-circuit configured to perform structured sparse processing on input data according to a predetermined sparse rule.

[0127] Clause 4, the data processing apparatus according to clause 3, wherein the structured sparse circuit is further configured to perform any one of:

[0128] performing structured sparse processing on the first and second input data to be convolved respectively by the second structured sparse sub-circuit, and outputting the sparse first and second input data to the convolution circuit to perform convolution operation; or

[0129] performing structured sparse processing on the first or second input data to be convolved by the first structured sparse sub-circuit using the index part of the structured sparse processed first or second input data as a sparse mask, wherein the index part indicates the position of valid data elements in the structured sparse processing, and outputting the sparse first input data to the convolution circuit to perform convolution operation with the structured sparse processed second input data.

[0130] Clause 5, the data processing apparatus according to any one of clauses 4, wherein:

[0131] the structured sparse processed first or second input data is pre-processed and stored in the storage circuit, or

[0132] the structured sparse processed first or second input data is generated on-line by the second structured sparse sub-circuit.

[0133] Clause 6, the data processing apparatus according to any one of clauses 4-5, wherein the structured sparse processed first or second input data is in the form of a structure, the structure comprising a data part and an index part, the data part comprising valid data elements after structured sparse processing, and the index part indicating the position of the sparse data in the data before sparse processing.

[0134] Clause 7, the data processing apparatus according to any one of clauses 3-6, wherein the structured sparsity processing comprises selecting n data elements from every m data elements as valid data elements, wherein m>n.

[0135] Clause 8, the data processing apparatus according to any one of clauses 3-7, wherein the second structured sparsity sub-circuit further comprises at least one multi-stage pipelined operation circuit comprising a plurality of operators arranged stage by stage and configured to perform the structured sparsity processing of selecting n data elements with larger absolute values from m data elements as valid data elements.

[0136] Clause 9, the data processing apparatus according to clause 8, wherein the multi-stage pipelined operation circuit comprises four pipelined stages, wherein:

[0137] the first pipelined stage comprises m absolute value operators for respectively taking absolute values of m data elements to be sparsified to generate m absolute values;

[0138] the second pipelined stage comprises a permutation circuit for permuting the m absolute values to generate m groups of data, wherein each group of data comprises the m absolute values and the m absolute values are in different positions in each group of data;

[0139] the third pipelined stage comprises m comparison circuits for comparing the absolute values in the m groups of data and generating comparison results; and

[0140] the fourth pipelined stage comprises a screening circuit for selecting n data elements with larger absolute values as valid data elements according to the comparison results, and outputting the valid data elements and corresponding indexes indicating positions of the valid data elements in the m data elements.

[0141] Clause 10, the data processing apparatus according to clause 9, wherein each comparison circuit in the third pipelined stage comprises m-1 comparators, and the m-1 comparators in the i-th comparison circuit are configured to sequentially compare one absolute value in the i-th group of data with other three absolute values and generate comparison results, 1≤i≤m.

[0142] Clause 11, the data processing apparatus according to any one of clauses 9-10, wherein the screening circuit is further configured to select according to a specified priority order when there are data elements with the same absolute value.

[0143] Clause 12, the data processing apparatus according to any one of clauses 2-11, wherein the operation circuit further comprises a pre-processing circuit configured to, when the sparsity flag indicates to perform a structured sparse convolution operation,

[0144] The pre-processing circuit reads input data from the storage circuit and outputs the input data to the structured sparse circuit at a first rate, wherein the first rate is based at least in part on a processing capability of the convolution circuit and a sparsity ratio of the structured sparse processing.

[0145] Clause 13, the data processing apparatus of any of clauses 2-12, wherein the input data comprises neuron data and weights of a neural network.

[0146] Clause 14, the data processing apparatus of any of clauses 1-13, wherein the convolution instruction is for a structured sparse convolution operation in a neural network, and the structured sparse is performed for at least one dimension of multi-dimensional data in the neural network.

[0147] Clause 15, the data processing apparatus of clause 14, wherein:

[0148] the at least one dimension is selected from an input channel dimension and an output channel dimension.

[0149] Clause 16, a chip comprising the data processing apparatus of any of clauses 1-15.

[0150] Clause 17, a board card comprising the chip of clause 15.

[0151] Clause 18, a data processing method comprising:

[0152] parsing a convolution instruction, the convolution instruction comprising a sparse flag indicating whether to perform a structured sparse convolution operation;

[0153] reading corresponding operands according to the convolution instruction; and

[0154] performing a corresponding convolution operation on the operands according to the convolution instruction.

[0155] Clause 19, the data processing method of clause 18, when the sparse flag indicates to perform a structured sparse convolution operation, the method further comprising:

[0156] performing structured sparse processing on at least one input data using a structured sparse circuit to obtain sparse input data; and

[0157] performing a convolution operation on data to be convolved using a convolution circuit, the data to be convolved comprising at least the sparse input data.

[0158] Clause 20, the data processing method of clause 19, wherein the structured sparse processing comprises:

[0159] performing structured sparse processing on input data according to a specified sparse mask using a first structured sparse sub-circuit; and / or

[0160] performing structured sparse processing on input data according to a predetermined sparse rule using a second structured sparse sub-circuit.

[0161] Clause 21, the data processing method according to clause 20, wherein the structured sparse processing further comprises any one of the following:

[0162] performing structured sparse processing on the first and second input data to be convolved respectively using a second structured sparse sub-circuit, and outputting the sparsified first and second input data to the convolution circuit to perform convolution operation; or

[0163] performing structured sparse processing on the first or second input data to be convolved using a first structured sparse sub-circuit, and outputting the sparsified first or second input data to the convolution circuit to perform convolution operation with the second or first input data to be convolved, wherein the index part indicates the position of the valid data element to be executed in the structured sparse processing.

[0164] Clause 22, the data processing method according to clause 21, wherein:

[0165] the first or second input data to be processed is pre-processed and stored in a storage circuit, or

[0166] the first or second input data to be processed is generated by processing online using the second structured sparse sub-circuit.

[0167] Clause 23, the data processing method according to any one of clauses 20-21, wherein the first or second input data to be processed is in the form of a structure, the structure comprising a data part and an index part bound to each other, the data part comprising valid data elements after structured sparse processing, and the index part being used to indicate the position of the data after sparsification in the data before sparsification.

[0168] Clause 24, the data processing method according to any one of clauses 20-23, wherein the structured sparse processing comprises selecting n data elements from every m data elements as valid data elements, wherein m>n.

[0169] Clause 25. The data processing method of any of clauses 20-24, wherein the second structured sparse sub-circuit further comprises: at least one multi-stage pipelined operation circuit comprising a plurality of operators arranged stage by stage and configured to perform the structured sparse processing of selecting n data elements with larger absolute values from m data elements as valid data elements.

[0170] Clause 26. The data processing method of clause 25, wherein the multi-stage pipelined operation circuit comprises four pipelined stages, wherein:

[0171] the first pipelined stage comprises m absolute value operators for respectively taking absolute values of m data elements to be sparsified to generate m absolute values;

[0172] the second pipelined stage comprises a permutation circuit for permuting the m absolute values to generate m groups of data, wherein each group of data comprises the m absolute values and the m absolute values are in different positions in each group of data;

[0173] the third pipelined stage comprises m-way comparison circuits for comparing the absolute values in the m groups of data and generating comparison results; and

[0174] the fourth pipelined stage comprises a screening circuit for selecting n data elements with larger absolute values as valid data elements according to the comparison results, and outputting the valid data elements and corresponding indices indicating positions of the valid data elements in the m data elements.

[0175] Clause 27. The data processing method of clause 26, wherein each way comparison circuit in the third pipelined stage comprises m-1 comparators, and the m-1 comparators in the i-th way comparison circuit are configured to sequentially compare one absolute value in the i-th group of data with other three absolute values and generate comparison results, 1≤i≤m.

[0176] Clause 28. The data processing method of any of clauses 26-27, wherein the screening circuit is further configured to select according to a specified priority order when there are data elements with the same absolute value.

[0177] Clause 29. The data processing method of any of clauses 19-28, further comprising:

[0178] delivering the input data at a first rate for performing the structured sparse processing, wherein the first rate is based at least in part on processing capability of hardware performing convolution operations and sparsity ratio of the structured sparse processing.

[0179] Clause 30. The data processing method of any of clauses 19-29, wherein the input data comprises neuron data and weights of a neural network.

[0180] Clause 31. The data processing method of any of clauses 18-30, wherein the convolution instruction is for a structured sparse convolution operation in a neural network, and the structured sparse is performed for at least one dimension of multi-dimensional data in the neural network.

[0181] Clause 32. The data processing method of clause 31, wherein:

[0182] the at least one dimension is selected from an input channel dimension and an output channel dimension.

[0183] The above detailed description of the embodiments of the present disclosure is made with specific examples applied to the principles and implementation modes of the present disclosure, and the above example is only used to help understand the method of the present disclosure and its core idea; at the same time, for those skilled in the art, according to the idea of the present disclosure, the specific implementation mode and application range will be changed, and the above description should not be understood as a limitation of the present disclosure.

Claims

1. A data processing apparatus comprising: a control circuit configured to parse a convolution instruction, the convolution instruction comprising a sparsity flag indicating whether to perform a structured sparse convolution operation; a storage circuit configured to store pre- and / or post-convolution information; and an arithmetic circuit configured to perform a corresponding convolution operation according to the convolution instruction; the arithmetic circuit comprising a structured sparse circuit configured to perform a structured sparse process on at least one input data; the structured sparse circuit comprising a second structured sparse sub-circuit configured to perform a structured sparse process on input data according to a predetermined sparsity rule; wherein the second structured sparse sub-circuit further comprises at least one multi-stage pipelined arithmetic circuit comprising a plurality of arithmetic units arranged in stages and configured to perform a structured sparse process of selecting n data elements with larger absolute values from m data elements as valid data elements. 2.The data processing apparatus of claim 1, wherein the arithmetic circuit comprises a convolution circuit, when the sparsity flag indicates to perform a structured sparse convolution operation, the structured sparse circuit is configured to output the sparsified input data to the convolution circuit; and the convolution circuit is configured to receive data to be convolved and perform a convolution operation thereon, wherein the data to be convolved comprises at least the sparsified input data. 3.The data processing apparatus of claim 2, wherein the structured sparse circuit comprises: a first structured sparse sub-circuit configured to perform a structured sparse process on input data according to a specified sparsity mask. 4.The data processing apparatus of claim 3, wherein the structured sparse circuit is further configured to perform either of: performing a structured sparse process on first and second input data to be convolved respectively by the second structured sparse sub-circuit and outputting the sparsified first and second input data to the convolution circuit to perform a convolution operation; or performing a structured sparse process on the first or second input data to be convolved by the first structured sparse sub-circuit using an index part of the structured sparse processed first or second input data as a sparsity mask, wherein the index part indicates positions of valid data elements in the structured sparse process to be performed, and outputting the sparsified first input data to the convolution circuit to perform a convolution operation with the structured sparse processed second input data. 5.The data processing apparatus of claim 4, wherein: the structured sparse processed first or second input data is pre-structured sparse processed and stored in the storage circuit, or the structured sparse processed first or second input data is generated on-line by performing a structured sparse process using the second structured sparse sub-circuit. ​ 6.The data processing apparatus of any of claims 4-5, wherein the first or second input data of the structured sparsification is in a structure including a data portion and an index portion, the data portion including valid data elements after the structured sparsification, and the index portion indicating positions of the data after the sparsification in the data before the sparsification. 7.The data processing apparatus of any of claims 3-5, wherein the structured sparsification includes selecting n data elements from every m data elements as valid data elements, where m>n. 8.The data processing apparatus of claim 3, wherein the multi-stage pipelined operation circuit includes four pipelined stages, wherein: the first pipelined stage includes m absolute value operators for taking absolute values of m data elements to be sparsified to generate m absolute values; the second pipelined stage includes permutation and combination circuitry for permuting and combining the m absolute values to generate m groups of data, each group of data including the m absolute values and the m absolute values being in different positions in each group of data; the third pipelined stage includes m comparison circuits for comparing the absolute values in the m groups of data and generating comparison results; and the fourth pipelined stage includes a screening circuit for selecting n data elements with larger absolute values as valid data elements according to the comparison results, and outputting the valid data elements and corresponding indexes indicating positions of the valid data elements in the m data elements. 9.The data processing apparatus of claim 8, wherein each comparison circuit in the third pipelined stage includes m-1 comparators, and the m-1 comparators in the i th comparison circuit are configured to compare one absolute value in the i th group of data with other three absolute values in turn and generate comparison results, where 1≤i≤m. 10.The data processing apparatus of any of claims 8-9, wherein the screening circuit is further configured to select according to a specified priority order when there are data elements with the same absolute value. 11.The data processing apparatus of any of claims 2-5, wherein the operation circuit further includes a pre-processing circuit configured to, when the sparsity flag indicates that a structured sparsified convolution operation is to be performed, read input data from the storage circuit and output the input data to the structured sparsification circuit at a first rate, wherein the first rate is based at least in part on a processing capability of the convolution circuit and a sparsity ratio of the structured sparsification. 12.The data processing apparatus of any of claims 2-5, wherein the input data includes neuron data and weights of a neural network. 13.The data processing apparatus of any of claims 1-5, wherein the convolution instruction is for a structured sparsified convolution operation in a neural network, and the structured sparsification is performed for at least one dimension of multi-dimensional data in the neural network. 14.The data processing apparatus of claim 13, wherein: ​ ​ ​ ​ ​ ​ The at least one dimension is selected from an input channel dimension and an output channel dimension.

15. A chip comprising the data processing apparatus according to any one of claims 1-14.

16. A board card comprising the chip according to claim 15.

17. A data processing method comprising: parsing a convolution instruction, the convolution instruction comprising a sparse flag for indicating whether to perform a structured sparse convolution operation; reading corresponding operands according to the convolution instruction; and performing a corresponding convolution operation on the operands according to the convolution instruction; when the sparse flag indicates to perform a structured sparse convolution operation, the method further comprises performing a structured sparse processing on at least one input data using a structured sparse circuit to obtain a sparse input data; wherein the structured sparse processing comprises performing a structured sparse processing on the input data according to a predetermined sparse rule using a second structured sparse sub-circuit; wherein the second structured sparse sub-circuit further comprises at least one multi-stage pipelined operation circuit comprising a plurality of operators arranged in stages and configured to perform a structured sparse processing of selecting n data elements with larger absolute values from m data elements as valid data elements.

18. The data processing method according to claim 17, when the sparse flag indicates to perform a structured sparse convolution operation, the method further comprises: performing a convolution operation on data to be convolved using a convolution circuit, the data to be convolved comprising at least the sparse input data.

19. The data processing method according to claim 18, wherein the structured sparse processing comprises: performing a structured sparse processing on the input data according to a specified sparse mask using a first structured sparse sub-circuit.

20. The data processing method according to claim 19, wherein the structured sparse processing further comprises any one of: performing a structured sparse processing on first and second input data to be convolved respectively using a second structured sparse sub-circuit, and outputting the sparse first and second input data to the convolution circuit to perform a convolution operation; or performing a structured sparse processing on the second or first input data using a first structured sparse sub-circuit with an index part of the first or second input data as a sparse mask, wherein the index part indicates positions of valid data elements in the structured sparse to be performed, and outputting the sparse first input data to the convolution circuit to perform a convolution operation with the second input data.

21. The data processing method according to claim 20, wherein: the first or second input data that has been structured sparse processed is pre-processed and stored in a storage circuit, or the first or second input data that has been structured sparse processed is generated on-line using the second structured sparse sub-circuit. ​ 22.The data processing method of claim 20, wherein the first or second input data of the structured sparse processing is in a form of a struct, the struct including a data portion and an index portion that are bound together, the data portion including valid data elements after the structured sparse processing, and the index portion indicating positions of the data after the sparsification in the data before the sparsification. 23.The data processing method of any one of claims 19-21, wherein the structured sparse processing includes selecting n data elements from every m data elements as valid data elements, where m>n. 24.The data processing method of claim 19, wherein the multi-stage pipelined operation circuit includes four pipelined stages, wherein: a first pipelined stage includes m absolute value operators for taking absolute values of m data elements to be sparsified to generate m absolute values, respectively; a second pipelined stage includes a permutation circuit for permuting the m absolute values to generate m groups of data, each group of data including the m absolute values and the m absolute values being in different positions in each group of data; a third pipelined stage includes m comparison circuits for comparing the absolute values in the m groups of data and generating comparison results; and a fourth pipelined stage includes a screening circuit for selecting n data elements with larger absolute values as valid data elements according to the comparison results, and outputting the valid data elements and corresponding indexes indicating positions of the valid data elements in the m data elements. 25.The data processing method of claim 24, wherein each comparison circuit in the third pipelined stage includes m-1 comparators, and m-1 comparators in an i-th comparison circuit are configured to compare one absolute value in an i-th group of data with other three absolute values in turn and generate comparison results, 1≤i≤m. 26.The data processing method of any one of claims 24-25, wherein the screening circuit is further configured to select according to a specified priority order when there are data elements with the same absolute value. 27.The data processing method of any one of claims 18-21, further comprising: delivering the input data at a first rate to perform the structured sparse processing, wherein the first rate is based at least in part on processing capability of hardware performing the convolution operation and a sparsity ratio of the structured sparse processing. 28.The data processing method of any one of claims 18-21, wherein the input data includes neuron data and weights of a neural network. 29.The data processing method of any one of claims 17-21, wherein the convolution instruction is for a structured sparse convolution operation in a neural network, and the structured sparse is performed for at least one dimension of multi-dimensional data in the neural network. 30.The data processing method of claim 29, wherein: the at least one dimension is selected from an input channel dimension and an output channel dimension. ​ ​ ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • Neural network computing device, neural network computing method and related products

    CN109740739A