Operation unit architecture, operation unit cluster, and method for performing convolution operation

By introducing a delayed queue connection computing unit architecture into the deep neural network accelerator, the data transmission path is optimized, and the storage space and power consumption problems are solved, and circuit area and power consumption are saved without reducing computing efficiency.

CN114692853BActive Publication Date: 2025-07-11IND TECH RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111173336.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-29
Filing Date
2021-10-08
Publication Date
2025-07-11
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

When designing existing deep neural network accelerators, how to maintain computing performance while reducing storage space and power consumption, especially how to design data transmission paths suitable for a large number of computing units to reduce the usage of storage components has become an important issue.

Method used

A computing unit architecture is adopted, which includes a first computing unit and a second computing unit. Through the delay queue connection, the storage device of the second computing unit is reduced. The delay queue is used to transmit shared data after the delay period for convolutional operations, and the computing unit cluster and busbar are designed to optimize data transmission.

Benefits of technology

Without reducing the computing performance, the storage space and power consumption are significantly reduced, especially when the number of computing units in the second computing group increases, further saving circuit area and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692853B_ABST
    Figure CN114692853B_ABST
Patent Text Reader

Abstract

An operation unit architecture applicable to convolution operations includes: a plurality of operation units and a delay queue. Among the operation units, there are a first operation unit and a second operation unit that perform convolution operations based on at least shared data. The delay queue is connected to the first operation unit and the second operation unit. The delay queue receives the shared data transmitted by the first operation unit, and after receiving the shared data and passing through a delay period, transmits the shared data to the second operation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence and an artificial intelligence accelerator for running a deep neural network. Background Art

[0002] In recent years, deep neural networks (DNNs) have developed rapidly. The accuracy of using DNNs for image recognition has gradually improved and is even more accurate than human recognition. To meet the computational requirements of DNNs, artificial intelligence accelerators (i.e., processors for running DNN models) must improve hardware performance. For artificial intelligence systems used in wearable devices, mobile communication devices, self-driving cars, and cloud servers, the required amount of computation grows exponentially with the device scale.

[0003] Generally, processors dedicated to DNNs must meet the requirements of both computing power and input / output bandwidth. Increasing the number of processing elements (PEs) can theoretically improve computing power, but a data network architecture suitable for a large number of processing elements is also required to send input data to each processing element in real time. For a processing element, the storage element occupies the largest proportion of its circuit area, followed by control logic and arithmetic logic. Considering the power consumption and circuit area associated with a large number of processing elements, how to design a good data transmission path to reduce the usage of storage elements has become an important issue in designing artificial intelligence accelerators. Summary of the Invention

[0004] In view of this, the present invention proposes an arithmetic unit architecture, an arithmetic unit cluster, and an execution method for convolution operations, which reduce the required storage space while maintaining the original computing performance of the artificial intelligence accelerator and have scalability.

[0005] An arithmetic unit architecture according to an embodiment of the present invention is applicable to a convolution operation. The architecture includes: a plurality of arithmetic units, among which there is a first arithmetic unit and a second arithmetic unit, and the first arithmetic unit and the second arithmetic unit perform the convolution operation at least based on a shared data; and a delay queue connected to the first arithmetic unit and the second arithmetic unit, the delay queue receiving the shared data transmitted by the first arithmetic unit and transmitting the shared data to the second arithmetic unit after receiving the shared data and after a delay period.

[0006] An arithmetic unit cluster according to an embodiment of the present invention is applicable to a convolution operation. The cluster includes: a first arithmetic group having a plurality of first arithmetic units; a second arithmetic group having a plurality of second arithmetic units; a bus connecting the first arithmetic group and the second arithmetic group, the bus providing a plurality of shared data to each of the first arithmetic units; and a plurality of delay queues, one of the delay queues connecting one of the first arithmetic units and one of the second arithmetic units, another of the delay queues connecting two of the second arithmetic units, and each of the delay queues transmitting one of the shared data; wherein each of the first arithmetic units in the first arithmetic group includes a storage device for storing the corresponding one of the shared data; and each of the second arithmetic units in the second arithmetic group does not include the storage device.

[0007] A method for performing a convolution operation according to an embodiment of the present invention is applicable to the arithmetic unit architecture of an embodiment of the present invention. The method includes: receiving an input data and the shared data by the first arithmetic unit and performing the convolution operation according to the input data and the shared data; transmitting the shared data from the first arithmetic unit to the delay queue; waiting for the delay period by the delay queue; after the delay queue waits for the delay period, transmitting the shared data from the delay queue to the second arithmetic unit; and receiving another input data by the second arithmetic unit and performing the convolution operation according to the another input data and the shared data.

[0008] The above description of the content of the present invention and the following description of the embodiments are used to demonstrate and explain the spirit and principle of the present invention, and provide a further explanation of the patent protection scope of the present invention. Description of the Drawings

[0009] Figure 1 is a block diagram of an arithmetic unit architecture according to an embodiment of the present invention;

[0010] Figure 2 is a block diagram of an arithmetic unit architecture according to another embodiment of the present invention;

[0011] Figure 3 is a block diagram of an arithmetic unit cluster according to an embodiment of the present invention; and

[0012] Figure 4 is a flowchart of a method for performing a convolution operation according to an embodiment of the present invention.

[0013]

Description of the Reference Numerals in the Drawings

[0014] Operation unit architectures 10, 10'; First operation unit PE1; Second operation units PE2, PE2a, PE2b; Operation circuit MAC; First storage device M1; Second storage device M2; Delay queues Q, Q1, Q2; Operation unit cluster 20; First operation group 21; Second operation group 22; Bus 23; Steps S1 to S8. Detailed implementation

[0015] The detailed features and characteristics of the present invention are described in detail in the following embodiments. The content is sufficient for those skilled in the art to understand the technical content of the present invention and implement it accordingly. According to the content, claims and drawings disclosed in this specification, those skilled in the art can easily understand the related concepts and characteristics of the present invention. The following embodiments further illustrate the viewpoints of the present invention, but do not limit the scope of the present invention in any way.

[0016] The present invention relates to a processing element array (PEArray) in an artificial intelligence accelerator. The processing element array is used to process one or more convolution operations. The processing element array receives the input data required for convolution operations, such as input feature maps (ifmap), kernel maps, and partial sums, from a global buffer (GLB). The processing element array contains multiple processing elements. Generally, each processing element includes a scratch pad memory (spad) for temporarily storing the aforementioned input data, a multiply accumulate (MAC) unit, and control logic.

[0017] The operation unit architecture proposed by the present invention includes two types of operation units: a first operation unit and a second operation unit. Among them, the number of the first operation units PE1 is 1, and the number of the second operation units PE2 is at least 1 or more. Figure 1 and Figure 2 Two embodiments showing one second operation unit and two second operation units are respectively illustrated. Embodiments with more than two second operation units can be deduced according to Figure 1 and Figure 2 oneself.

[0018] Figure 1 is a block diagram of an operation unit architecture according to an embodiment of the present invention. The described operation unit architecture is applicable to convolution operations and includes multiple operation units and a delay queue. Figure 1 The shown operation unit architecture 10 includes a first operation unit PE1, a second operation unit PE2, and a delay queue Q.

[0019] The first arithmetic unit PE1 and the second arithmetic unit PE2 perform convolution operations based on at least one shared data. In an embodiment, the shared data is a convolution kernel or a filter. The first arithmetic unit PE1 includes a first storage device M1, a second storage device M2, and an arithmetic circuit MAC. The hardware structure of the second arithmetic unit PE2 is similar to that of the first arithmetic unit PE1, except that the second arithmetic unit PE2 does not have the first storage device M1. In actual application, the first storage device M1 is used to temporarily store the shared data, such as a convolution kernel or a filter. The second storage device M2 is used to temporarily store non-shared data, such as an input feature map or a partial sum. The arithmetic circuit MAC is, for example, a multiply-accumulate operator. The arithmetic circuit MAC performs convolution operations based on data such as a convolution kernel taken from the first storage device M1, an input feature map taken from the second storage device M2, and a partial sum taken from the second storage device M2. The convolution kernel belongs to the shared data, and the input feature map and the partial sum belong to the non-shared data. In actual application, the input feature map and the partial sum can be stored in two different storage devices respectively, or in different storage spaces under one storage device, and the present invention does not limit this.

[0020] A delayed-control queue Q connects the first arithmetic unit PE1 and the second arithmetic unit PE2. The delayed queue Q is used to receive the shared data transmitted by the first arithmetic unit PE1, and after receiving the shared data and passing through a delay period P, it transmits the shared data to the second arithmetic unit PE2. In actual application, the delayed queue Q has a First In-First Out (FIFO) data structure. An example is as follows, where T k represents the kth unit time;

[0021] At T k the first arithmetic unit PE1 transmits the shared data F1 to the delayed queue Q;

[0022] At T k+1 the first arithmetic unit PE1 transmits the shared data F2 to the delayed queue Q; Therefore,

[0023] At the T k+P the second arithmetic unit PE2 receives the shared data F1 from the delayed queue Q: and

[0024] At the T k+1+P the second arithmetic unit PE2 receives the shared data F2 from the delayed queue Q.

[0025] In an embodiment of the present invention, the order of magnitude of the delay period P is the same as the value of the stride of the convolution operation. For example, if the stride is 2, the delay period is also 2 unit times.

[0026] In an embodiment of the present invention, the size of the storage space of the delay queue Q is not less than the stride of the convolution operation. The following is an example. If the stride of the convolution operation is 3, and the first arithmetic unit PE1 obtains the shared data F1 at time T k and performs the first convolution operation, the first arithmetic unit PE1 will obtain the shared data F4 at time T k+3 and perform the second convolution operation. However, during the period from T k+1 to T k+2 , the delay queue Q still needs to temporarily store the shared data F2 and F3 from the first arithmetic unit PE1, and at time T k+3 , the delay queue Q transmits the shared data F1 to the second arithmetic device PE2. Therefore, the delay queue Q requires at least 3 unit spaces to store the shared data F1~F3.

[0027] Figure 2 FIG. is a block diagram of an arithmetic unit architecture 10' according to another embodiment of the present invention. Compared with the previous embodiment, the arithmetic unit architecture 10' of this embodiment includes a first arithmetic unit PE1, a second arithmetic unit PE2a, another second arithmetic unit PE2b, a delay queue Q1, and another delay queue Q2. The second arithmetic unit PE2a and another second arithmetic unit PE2b perform convolution operations at least based on the shared data. Another delay queue Q2 is connected to the second arithmetic unit PE2a and another second arithmetic unit PE2b. This another delay queue Q2 receives the shared data transmitted by the second arithmetic unit PE2a, and transmits the shared data to another second arithmetic unit PE2b after receiving the shared data and passing through the delay period. In the actual application process, multiple second arithmetic units PE2 connected in series after the first arithmetic unit PE1 and the delay queues Q corresponding to these second arithmetic units PE2 can be added according to requirements. As can be seen from the above, the number of delay queues Q in the arithmetic unit architecture 10 is the same as the number of second arithmetic units PE2.

[0028] Figure 3 FIG. is a block diagram of an arithmetic unit cluster 20 according to an embodiment of the present invention. The described arithmetic unit cluster 20 is applicable to convolution operations, and includes a first arithmetic group 21, a second arithmetic group 22, a bus 23, and multiple delay queues Q. The first arithmetic group 21 and the second arithmetic group 22 are arranged in a two-dimensional array of M columns and N rows. Each of the M columns has one of the multiple first arithmetic units and ones of the multiple second arithmetic units. In Figure 3In the illustrated example, M = 3 and N = 7. However, the present invention does not limit the values of M and N. The delay queue Q has M groups, and each of these M groups has delay queues Q.

[0029] The first operation group 21 has M first operation units PE1. Each first operation unit PE1 in the first operation group 21 is the same as the first operation unit PE1 described in the previous embodiment. The first operation unit PE1 has a first storage device M1 for storing shared data.

[0030] The second operation group 22 has second operation units PE2. Each second operation unit PE2 in the second operation group 22 does not include the first storage device M1.

[0031] The bus 23 connects the first operation group 21 and the second operation group 22. In an embodiment of the present invention, the bus 23 is connected to each first operation unit PE1, and the bus 23 is connected to each second operation unit PE2. The bus 23 provides a plurality of shared data to each first operation unit PE1. The bus 23 provides a plurality of non-shared data to each first operation unit PE1 and each second operation unit PE2. The sources of the shared data and the non-shared data are, for example, the GLB.

[0032] Please refer to Figure 3 , the number of delay queues Q of the operation unit cluster 20 is ones, and each delay queue Q is used to transfer shared data.

[0033] One of these delay queues Q connects one of these first operation units PE1 and one of these second operation units PE2. Another one of these delay queues Q connects two of these second operation units PE2, and each of these delay queues Q transfers one of these shared data. In other words, each first operation unit PE1 in the first operation group 21 is connected to a second operation unit PE2 in the second operation group 22 through a delay queue Q. Two second operation units PE2 in the same column and adjacent rows in the second operation group 22 are connected to each other through one of these delay queues.

[0034] Figure 4 is a flowchart of an execution method of a convolution operation according to an embodiment of the present invention. Figure 4 The illustrated execution method of the convolution operation is applicable to Figure 1 the operation unit architecture 10 shown in Figure 2 the operation unit architecture 10' shown in or Figure 3 the operation unit cluster 20 shown in.

[0035] Step S1 is that the first arithmetic unit PE1 receives input data and shared data, and performs a convolution operation based on the input data and the shared data. The input data and the shared data are transmitted to the first arithmetic unit PE1 by the bus 23, for example.

[0036] Step S2 is that the first arithmetic unit PE1 transmits the shared data to the k-th delay queue Q, where k = 1. k represents both the number of the delay queue and the number of the second processing unit. Steps S1 and S2 do not limit the order of execution. Steps S1 and S2 can be executed simultaneously.

[0037] Step S3 is that the k-th delay queue Q waits for a delay time. The length of the delay time depends on the stride of the convolution operation.

[0038] Step S4 is that the k-th delay queue Q transmits the shared data to the k-th second arithmetic unit PE2.

[0039] Step S5 is that the k-th second arithmetic unit PE2 receives another input data, and performs a convolution operation based on the another input data and the shared data.

[0040] Step S6 is to determine whether the k-th second arithmetic unit PE2 is the last second arithmetic unit PE2. If the determination result of step S6 is yes, the execution method of the convolution operation in an embodiment of the present invention ends. If the determination result of step S6 is no, step S7 is executed.

[0041] Step S7 is that the k-th second arithmetic unit PE2 transmits the shared data to the (k + 1)-th delay queue Q. Step S7 is similar to step S2. Both step S7 and step S2 are that the first arithmetic unit PE1 or the second arithmetic unit PE2 transmits the shared data to the next-level delay queue Q.

[0042] Step S8 is k = k + 1, that is, the value of k is incremented. According to the number of the second arithmetic units PE2 in the arithmetic unit architecture 10 or 10´, steps S3 to S8 may be repeatedly executed a plurality of times.

[0043] In summary, the arithmetic unit architecture, the arithmetic unit cluster, and the execution method of the convolution operation proposed by the present invention can save a large amount of storage devices originally used to store shared data through the design of the second arithmetic unit and the delay queue. The more the number of the second arithmetic units belonging to the second arithmetic group in the artificial intelligence accelerator, the larger the circuit area that can be saved by applying the present invention, and thus a large amount of power consumption is also saved.

Claims

1. An operation unit architecture applicable to a convolution operation, the architecture comprising: A plurality of operation units, among which there is a first operation unit and a second operation unit, and the first operation unit and the second operation unit perform the convolution operation at least based on a shared data; and A delay queue connecting the first operation unit and the second operation unit, the delay queue receiving the shared data transmitted by the first operation unit and transmitting the shared data to the second operation unit after receiving the shared data and after a delay period; wherein the order of magnitude of the delay period is the same as the stride value of the convolution operation.

2. The arithmetic unit architecture according to claim 1, wherein, Among the operation units there is another second operation unit, and the second operation unit and the other second operation unit perform the convolution operation at least based on the shared data; and The operation unit architecture further includes another delay queue, the other delay queue connecting the second operation unit and the other second operation unit, the other delay queue receiving the shared data transmitted by the second operation unit and transmitting the shared data to the other second operation unit after receiving the shared data and after the delay period.

3. The arithmetic unit architecture according to claim 1, wherein, The storage space of the delay queue is not less than the stride of the convolution operation.

4. An operation unit cluster applicable to a convolution operation, the cluster comprising: A first operation group having a plurality of first operation units; A second operation group having a plurality of second operation units; A bus connecting the first operation group and the second operation group, the bus providing a plurality of shared data to each of the first operation units; and A plurality of delay queues, one of the delay queues connecting one of the first operation units and one of the second operation units, another of the delay queues connecting two of the second operation units, and each of the delay queues transmitting one of the shared data; wherein Each of the first operation units in the first operation group includes a storage device for storing the corresponding one of the shared data; and Each of the second operation units in the second operation group does not include the storage device.

5. The arithmetic unit cluster as claimed in claim 4, wherein, The storage device is a first storage device, and each of the first operation units and each of the second operation units further includes: A second storage device for storing a non-shared data; and An operation circuit electrically connected to the first storage device and the second storage device, and the operation circuit performs the convolution operation based on the corresponding one of the shared data and the non-shared data.

6. The operation unit cluster according to claim 4, wherein, The first operation group and the second operation group form a two-dimensional array of M columns and N rows, and each of the M columns has one of the first operation units and ones of the second operation units; There are M groups of these delay queues, and each of the M groups has delay queues.

7. A method for executing an operation unit architecture applicable to the operation unit architecture according to claim 1, the method comprising: Receiving an input data and the shared data by the first operation unit and performing the convolution operation based on the input data and the shared data; Transmitting the shared data from the first operation unit to the delay queue; Waiting for the delay period by the delay queue; After the delay queue waits for the delay period, transmitting the shared data from the delay queue to the second operation unit; and The second arithmetic unit receives another input data, and performs the convolution operation according to the another input data and the shared data.

8. The execution method of the arithmetic unit architecture according to claim 7, wherein, The shared data is a convolution kernel, and the input data includes an input feature map and a partial sum.

Citation Information

Patent Citations

  • Computing convolutions using a neural network processor

    TWI645301B