Neural network arithmetic unit

The neural network arithmetic unit controls clock supply to processor elements, addressing power fluctuations in AI accelerators by gradual activation and deactivation, enhancing power efficiency and reducing costs.

JP7707977B2Active Publication Date: 2025-07-15DENSO CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022047092
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-07-15
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

Existing AI accelerators with large PE arrays experience significant power consumption fluctuations due to startup and shutdown, leading to voltage drops and inefficiencies, with conventional solutions either increasing costs or causing unnecessary power consumption.

Method used

A neural network arithmetic unit with a main controller that controls the timing of clock supply to processor elements, allowing for gradual activation and deactivation of processor elements to manage power consumption.

Benefits of technology

The solution effectively suppresses steep power fluctuations, optimizing power usage and maintaining system efficiency without the need for external capacitors or prolonged startup/shutdown times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007707977000001
    Figure 0007707977000001
  • Figure 0007707977000002
    Figure 0007707977000002
  • Figure 0007707977000003
    Figure 0007707977000003
Patent Text Reader

Abstract

To provide a neural network arithmetic device that prevents steep power fluctuations.SOLUTION: A SOC 1 comprises: a plurality of processor elements 104 arranged in array; an activation memory 103 storing input activation data supplied to the processor elements; and a main controller 101 controlling the operation of the processor elements. The processor elements arranged in array have a function to transfer a command from the main controller between the adjacent processor elements. A clock line of a clock generator is connected to every group obtained by logically arranging the processor elements. The main controller synchronizes with supply of data for neural network operation to the processor elements, supplies clocks to the processor elements for every group, and increases the number of groups to be supplied with clocks with the lapse of time.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an arithmetic unit, and more particularly to a neural network arithmetic unit that performs arithmetic operations with a plurality of processor elements.

Background Art

[0002] In general, in an AI accelerator, a PE array in which a large number of PEs (Processing Elements) are arranged in a two-dimensional array is processed in parallel. When an AI accelerator with a large number of these PE arrays is configured in an implementation form that occupies a dominant area in the entire SOC, a large instantaneous power consumption fluctuation occurs due to the startup and stop of the AI accelerator, resulting in problems such as voltage drop.

[0003] As a conventional technique for solving this problem, (1) a method of taking measures such as arranging a large-capacity capacitor that can withstand instantaneous power consumption fluctuations outside the chip, or (2) a method of gradually turning on and off the power in the entire SOC to maintain an operating state in which no fluctuations occur during operation, etc. are considered. Also, regarding the control of power fluctuations during processing, the technique described in Patent Document 1 is also known.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the method of (1), a capacitor is required outside the chip, which not only increases the additional cost, but also poses new problems in terms of its reliability and durability, especially in harsh system requirements such as in-vehicle applications. On the other hand, in the method of (2), since a certain amount of time is required for startup and shutdown, unnecessary power consumption occurs during that period, and the power efficiency of the entire system may be impaired.

[0006] Although the method described in Patent Document 1 is effective for controlling power fluctuations during processing, it does not serve as a solution for gradually increasing or decreasing power consumption during startup.

[0007] Therefore, in view of the above background, an object of the present invention is to provide a neural network arithmetic unit capable of suppressing steep power fluctuations.

Means for Solving the Problems

[0008] The present invention employs the following technical means to solve the above problems. The claims and the reference numerals in parentheses described in this section are an example showing the correspondence relationship with the specific means described in the embodiments to be described later as one aspect, and do not limit the technical scope of the present invention.

[0009] The neural network arithmetic unit of the present invention is a neural network arithmetic unit (10) that performs neural network processing, and includes a plurality of processor elements (104) arranged in an array, an activation memory (103) that stores input activation data supplied to the processor elements, and a main controller (101) that controls the operations of the respective processor elements. The main controller (101) has a configuration for controlling the timing of supplying the clock supplied from the clock generator (11) to each processor element.

[0010] With this configuration, by controlling the clock supply timing to the processor element by the main controller, it is possible to effectively utilize power and suppress steep power fluctuations.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Embodiments for Carrying Out the Invention

[0012] Hereinafter, the neural network arithmetic unit according to the embodiment will be described with reference to the drawings. (First Embodiment) FIG. 1 is a diagram showing an LSI including the neural network arithmetic unit 10 according to the first embodiment.

[0013] 1 indicates the entire SOC. The SOC 1 often includes a CPU (not shown) and a bus system that controls the entire system, and functions as a system. 10 is the neural network arithmetic unit according to the present embodiment. 11 is a clock generator that supplies a clock to the entire SOC 1. The clock generator also supplies a clock to the neural network arithmetic unit 10. This clock may supply one type of frequency depending on the design, or may supply clocks of multiple types of frequencies. In the present embodiment, for simplicity, the case of supplying one type of clock will be described, but the present invention can also be realized with multiple types of clocks.

[0014] Depending on its configuration, the neural network arithmetic unit 10 may occupy more than 50% of the scale of the entire SOC 1. Therefore, the power consumption of the neural network arithmetic unit 10 can also become a main part of the entire SOC 1 accordingly.

[0015] Next, the internal configuration of the neural network arithmetic unit 10 will be described. 101 is the main controller. The main controller 101 is a module that controls the processing sequence of the entire neural network arithmetic unit 10, manages input / output data, and issues operation requests to the processor elements described later. In addition, the main controller 101 also performs power control based on the information on the operating state of the neural network arithmetic unit 10.

[0016] 102 is the clock controller. The clock controller 102 supplies clocks to the processor elements 104 described later and all other internal modules of the neural network arithmetic unit 10. The clock controller 102 is equipped with a mechanism for individually supplying and stopping clocks to a plurality of processor elements 104. In addition, a mechanism for controlling the clock frequency for each individual processor element PEij may be provided.

[0017] 103 is the activation memory. The activation memory 103 is a module that stores the activation data to be supplied to the processor elements 104 described later. The main controller 101 transfers the necessary activation data to this activation memory 103 in advance using a DMA controller (not shown), etc., and issues an instruction to supply the data to the processor elements 104 at the start of processing. The activation memory 103 supplies the necessary activation data to the processor elements 104 based on the instruction.

[0018] Regarding the implementation method of this part, it is also acceptable that the main controller 101 issues an instruction to the activation memory 103 and the activation memory 103 actively supplies data to the processor elements, or that the main controller 101 issues an instruction to the processor elements 104, causing the processor elements 104 to read the necessary data from the activation memory 103.

[0019] 104 is a processor element. The processor element 104 includes one or more convolutional arithmetic units each consisting of a multiplier and an adder that perform a convolutional operation with the input activation and the Weight value from the Weight input memory (not shown) as inputs. In many cases, a plurality of convolutional multipliers are implemented in the processor element 104, and it is configured to be able to execute a plurality of convolutional operations simultaneously in parallel.

[0020] The term "processor element 104" is used when collectively referring to processor elements, and when referring to an individual processor element, it is called a processor element "PEij". Here, ij is a number that identifies the processor element PE by its position, i (i = 0, ···, m) identifies the row, and j (j = 0, ···, n) identifies the column.

[0021] The output activation, which is the arithmetic output of the processor element 104, is either held again in the activation memory 103 or output outside the module.

[0022] These processor elements 104 are logically arranged in a two-dimensional array inside the neural network arithmetic unit 10. The processor elements adjacent to each other vertically, horizontally, transfer activation data to each other during the process of performing a convolutional operation, thereby realizing a convolutional operation of one or more kernel sizes.

[0023] For example, when PE11 in FIG. 1 performs a convolutional operation with a kernel size of 2x2, PE11 can obtain the activation data of the left adjacent element to the processing target by acquiring the activation data input from the activation memory 103 to PE10. Similarly, it can receive the data of the upper adjacent element from PE01 and the activation data once transmitted from PE00 to PE01 and use them for the operation. Also, as another configuration, it may acquire the activation data of the adjacent 2x2 region from the activation memory 103.

[0024] Next, the clock processing flow of the LSI including the neural network arithmetic unit 10 will be described. FIG. 2 is a diagram showing a signal for instructing the timing at which the main controller 101 supplies a clock to the clock controller 102. 201 to 204 in FIG. 2 are Enable signals output from the main controller 101 to the clock controller 102. 205 to 208 are clock signals output from the clock controller 102 to the processor element 104.

[0025] Among the processor elements 104 arranged two-dimensionally, one clock is supplied to the processor elements 104 in the same column. For example, 205 is the clock signal of the clock line connected to the processor element in the first column (PEx0 in FIG. 1). Similarly, 206 is the signal of the clock line connected to the processor element in the second column (PEx1 in FIG. 1). In the present embodiment, groups are configured by the processor elements 104 in the same column.

[0026] The main controller 101 sequentially adjusts the timing of applying the clock for each column. More specifically, in the convolution operation process, the clock of the next column is activated by shifting by the number of cycles required for one element's multiply-accumulate operation to be completed. FIG. 2 is shown on the premise that one element's multiply-accumulate operation is completed in one cycle, so the clock is applied by shifting one cycle for each column. As a result, the next column can be activated at the timing immediately after the activation is used in the processing of adjacent cycles.

[0027] By controlling the clock in this way, the processor elements PEij are activated one column at a time without delaying the data between adjacent processor elements PEij. As a result, as shown in FIG. 3, the power consumption increases more gently than starting to supply the clock to all the processor elements PEij at once.

[0028] (Second Embodiment) The basic configuration of the neural network arithmetic unit according to the second embodiment is the same as that of the neural network arithmetic unit according to the first embodiment. In the second embodiment, the increase in power consumption is made even gentler compared to the first embodiment.

[0029] FIG. 4 is a diagram showing a signal for instructing the clock input timing in the LSI including the neural network arithmetic unit according to the second embodiment. That is, it is a diagram showing a signal for instructing the timing at which the main controller 101 inputs a clock to the clock controller 102.

[0030] 401 to 404 in FIG. 4 are Enable signals output from the main controller 101 to the clock controller 102. 405 to 408 are clock signals output from the clock controller 102 to the processor element 104.

[0031] At this time, among the processor elements 104 arranged two-dimensionally, one clock is supplied to the processor elements 104 in the same column. For example, 405 is the clock signal of the clock line connected to the processor element in the first column (PEx0 in FIG. 1). Similarly, 406 is the clock signal of the clock line connected to the processor element in the second column (PEx1 in FIG. 1).

[0032] At this time, the main controller 101 sequentially adjusts the timing of applying the clock for each column. More specifically, in the convolution operation process, the clock of the next column is started with a shift corresponding to the number of cycles required for the product-sum operation of one element to be completed. FIG. 4 is shown on the premise that the product-sum operation of one element is completed in one cycle, so the clock is applied with a shift of one cycle for each column. As a result, the next column can be started at the timing immediately after the activation is used in the processing of adjacent cycles. Further, in the present embodiment, the activation interval of each column is thinned out by the number of cycles of the number of columns of the array, and the thinning interval is gradually reduced.

[0033] Specifically, when the number of columns is 4, after applying the first clock, the next clock is applied after 4 cycles. Then, after applying the clock with a 2-cycle interval, normal clocks are supplied. By controlling the clocks in this way, while one arithmetic operation is being executed, the number of operating processor elements is 1 / 4 of the total. In the subsequent 2 cycles, the number of operating processors is 1 / 4 of the total, and in the next cycle, 1 / 2 of the total operates. In the following 2 cycles, 3 / 4 of the total operates, and thereafter all processor elements operate.

[0034] By performing control in this manner, the processor elements are activated one column at a time without causing data delay between adjacent elements, and the operating rate of all processor elements can be sequentially increased. As a result, as shown in FIG. 5, compared to the clock supply method shown in the first embodiment, the power consumption increases more gently.

[0035] (Third Embodiment) FIG. 6 is a diagram showing an LSI including the neural network arithmetic unit of the third embodiment.

[0036] 6 indicates the entire SOC. The SOC 6 often includes a CPU and a bus system (not shown) that control the entire system and functions as a system. 60 is the neural network arithmetic unit of this embodiment. 61 is a clock generator that supplies clocks to the entire neural network arithmetic unit 60. The clock generator 61 also supplies clocks to the neural network arithmetic unit 60. Depending on the design, this clock may supply one type of frequency or multiple types of frequencies. In this embodiment, for simplicity, the case of supplying one type of clock will be described, but the present invention can also be realized with multiple types of clocks.

[0037] Depending on its configuration, the neural network arithmetic unit 60 may occupy more than 50% of the entire scale of the SOC6. Therefore, the power consumption of the neural network arithmetic unit 60 can also become a major part of the entire SOC accordingly.

[0038] Next, the internal configuration of the neural network arithmetic unit 60 will be described. 601 is the main controller. The main controller 601 is a module that controls the processing sequence of the entire neural network arithmetic unit 60, manages input / output data, and issues operation requests to the processor elements described later. In addition, the main controller 601 also performs power control based on the information on the operating state of the neural network arithmetic unit 60.

[0039] 602 is the clock controller. The clock controller 602 supplies clocks to the processor elements 604 described later and all other internal modules of the neural network arithmetic unit 60. The clock controller 602 is equipped with a mechanism for individually supplying and stopping clocks to a plurality of processor elements 604. In addition, a mechanism for controlling the clock frequency for each individual processor element PEij may be provided.

[0040] 603 is the activation memory. The activation memory 603 is a module that stores the activation data to be supplied to the processor elements 604 described later. The main controller 601 uses a DMA controller (not shown) or the like to transfer the necessary activation data to this activation memory 603 in advance and issues an instruction to supply the data to the processor elements at the start of processing. The activation memory 603 supplies the necessary activation data to the processor elements 604 based on this instruction.

[0041] As for the implementation method of this part, the main controller 601 may issue an instruction to the activation memory 603, and the activation memory 603 may actively supply data to the processor element. Alternatively, the main controller 601 may issue an instruction to the processor element 604, and the processor element 604 may read the required data from the activation memory 603.

[0042] 604 is a processor element. The processor element 604 includes one or more convolutional arithmetic units composed of a multiplier and an adder that perform a convolutional operation with the input activation and the Weight value from the Weight input memory (not shown) as inputs. In many cases, a plurality of convolutional multipliers are implemented in the processor element, and it is configured to be able to execute a plurality of convolutional operations simultaneously in parallel.

[0043] The term "processor element 604" is used when collectively referring to the processor elements. When referring to an individual processor element, it is called a processor element "PEij". Here, ij is a number that identifies the processor element PE by its position, i (i = 0, ···, m) identifies the row, and j (j = 0, ···, n) identifies the column.

[0044] The output activation, which is the arithmetic output of the processor element 604, is either retained again in the activation memory 603 or output outside the module.

[0045] These processor elements 604 are logically arranged in a two-dimensional array inside the neural network arithmetic unit 60. The processor elements adjacent to each other vertically, horizontally transfer activation data to each other during the process of performing the convolutional operation, thereby realizing a convolutional operation with a kernel size of 1 or more.

[0046] For example, when PE11 in FIG. 6 performs a convolution operation with a kernel size of 2x2, PE11 can obtain the activation data of the left adjacent activation data to be processed by acquiring the activation data input from the activation memory 603 to PE10. Similarly, it also receives the data of the upper adjacent data from PE01 and the activation data once transmitted from PE00 to PE01, and has a structure for use in the operation. As another configuration, activation data in an adjacent 2x2 region may be acquired from the activation memory 603.

[0047] Next, the clock processing flow of the LSI including the neural network arithmetic unit 60 will be described. Since the timing of the clock processing in the third embodiment is the same as that in the first embodiment, it will be described with reference to FIG. 2. In the third embodiment, the group of processor elements 604 to which the same clock is supplied is different. FIG. 2 is a diagram showing a signal for the main controller 601 to instruct the timing of supplying a clock to the clock controller 602.

[0048] 201 to 204 in FIG. 2 are Enable signals output from the main controller 601 to the clock controller 602. 205 to 208 are clock signals output from the clock controller 602 to the processor element 604.

[0049] Among the two-dimensionally arranged processor elements 604, as shown in FIG. 6, clock groups are configured radially from PE00 according to the separation distances in the vertical and horizontal directions, and one clock is supplied to each clock group. For example, 605 is a clock group connected to the processor element at the upper left end (PE00 in FIG. 6). 606 is a clock group connected to the clock group formed by the group of processor elements adjacent to PE00 in the vertical and horizontal directions in FIG. 6. FIG. 6 shows groups from 605 to 609, but the number of groups is not limited to this.

[0050] The main controller 601 sequentially adjusts the timing of applying the clock for each clock group. More specifically, in the convolution operation processing, the clock of the next group is started by shifting by the number of cycles required for the sum-of-products operation of one element to be completed. FIG. 2 is shown on the premise that the sum-of-products operation of one element is completed in one cycle. Therefore, the clock is applied by shifting by one cycle for each column.

[0051] As a result, the next row and column can be activated at the timing immediately after the activation is used in the processing of adjacent cycles. By controlling the clock in this way, the data between adjacent processor elements PEij can be sequentially activated without being delayed, so that the power consumption increases more gently than starting the clock supply to all processor elements at once.

[0052] (Fourth Embodiment) FIG. 7 is a diagram showing an example of the physical configuration of an LSI including the neural network arithmetic unit according to the fourth embodiment. On the other hand, FIG. 8 is a diagram showing the corresponding logical structure of the LSI. Each internal module will be described with reference to FIG. 8.

[0053] 8 in FIG. 8 indicates the entire SOC. The SOC 8 often includes a CPU and a bus system (not shown) that control the entire system and functions as a system. 80 is the neural network arithmetic unit according to the fourth embodiment. 81 is a clock generator that supplies a clock to the entire SOC 8. The clock generator also supplies a clock to the neural network arithmetic unit 80. This clock may supply one type of frequency depending on the design, or may supply clocks of multiple types of frequencies. In this embodiment, for simplicity, the case of supplying one type of clock will be described, but the present invention can also be realized with multiple types of clocks.

[0054] Next, the internal configuration of the neural network arithmetic unit 80 will be described. 801 is the main controller. The main controller 801 is a module that controls the processing sequence of the entire neural network arithmetic unit 80, manages input / output data, and issues operation requests to the processor elements described later. In addition, the main controller 801 also performs power control based on the operation state information of the neural network arithmetic unit 80.

[0055] 803 is the activation memory. The activation memory 803 is a module that stores the activation data to be supplied to the processor element 804 described later. The main controller 801 uses a DMA controller (not shown) or the like to transfer the necessary activation data to this activation memory 803 in advance, and issues an instruction to supply the data to the processor element 804 at the start of processing. Based on the instruction, the activation memory 803 supplies the necessary activation data to the processor element 804. FIG. 7 shows an example in which the activation memory 803 is arranged side by side on the left side in terms of layout.

[0056] 804 is the processor element. The processor element 804 includes one or more convolutional arithmetic units composed of a multiplier and an adder that execute a convolutional operation with the input activation and the Weight value from the Weight input memory (not shown) as inputs. In many cases, a plurality of convolutional multipliers are implemented in the processor element 804, and a configuration is provided in which a plurality of convolutional operations can be executed simultaneously in parallel.

[0057] The term "processor element 804" is used when collectively referring to the processor elements, and when referring to an individual processor element, it is called a processor element "PEij". Here, ij is a number that specifies the processor element PE according to the position, i (i = 0, ···, m) specifies the row, and j (j = 0, ···, n) specifies the column.

[0058] The output activation, which is the calculation output, is either retained again in the activation memory 803 or output outside the module.

[0059] These processor elements 804 are logically arranged in a two-dimensional array inside the neural network arithmetic unit 80. The processor elements adjacent to each other vertically and horizontally transfer activation data to each other during the convolution operation process to realize a convolution operation with a kernel size of 1 or more.

[0060] For example, when PE11 in FIG. 8 performs a convolution operation with a kernel size of 2x2, PE11 can obtain the activation data of the left adjacent activation data to be processed by obtaining the activation data input from the activation memory 803 to PE10. Similarly, it also receives the activation data of the upper adjacent data from PE01 and the activation data once transmitted from PE00 to PE01, and has a structure for use in the operation. Also, as another configuration, activation data in an adjacent 2x2 region may be obtained from the activation memory 803.

[0061] At this time, among the processor elements 804 arranged two-dimensionally, when the arrangement wiring is performed as shown in FIG. 7, a clock group is configured for the adjacent PE groups, and a clock is supplied to each clock group.

[0062] 805 to 807 in FIG. 8 logically represent clock groups. For example, 805 is a clock group connected to the processor element at the upper left end. Since these are in a proximity relationship in the physical arrangement wiring of FIG. 7, they form one clock group. 806 is a clock group formed by the processor elements 804 from the second column in the first row onwards. In the example of FIG. 7, PE1x is arranged in a strip shape in terms of arrangement wiring, and in this case, the clock group is separated.

[0063] As described above, the clock group is set according to the layout wiring state as shown in FIG. 7. At this time, one clock line may be connected to the clock group to form a physically fixed clock group. However, since the layout wiring is often determined in the latter stage of the design process and may be changed depending on the process, a clock line is connected independently to each processor element, and the clock supply timing is logically adjusted as the same group by the clock controller 802, and it may be implemented to obtain the same effect.

[0064] Next, the clock processing flow of the LSI including the neural network arithmetic unit 80 will be described. Since the timing of the clock processing in the fourth embodiment is the same as that in the first embodiment, it will be described with reference to FIG. 2. FIG. 2 is a diagram showing a signal for instructing the timing at which the main controller 601 supplies a clock to the clock controller 602.

[0065] The main controller 801 sequentially adjusts the timing of applying the clock for each group. More specifically, in the convolution operation process, the clock of the next column is started by shifting by the number of cycles required for the completion of the product-sum operation of one element. FIG. 2 is shown on the premise that the product-sum operation of one element is completed in one cycle. Therefore, the clock is applied by shifting by one cycle for each column.

[0066] As a result, the next row and column can be started at the timing immediately after the activation is used in the processing of adjacent cycles. By controlling the clock in this way, the processor elements are sequentially activated without delaying the data between adjacent elements. As a result, the power consumption increases more gently than supplying the clock to all the processor elements.

[0067] In the present embodiment, an example where the clock timing is the same as that of the first embodiment has been described. However, the clock timing may adopt a method of thinning out the activation intervals of each column by the number of cycles of the number of columns of the array and gradually reducing the thinning interval as in the second embodiment (see Fig. 4).

[0068] (Fifth Embodiment) Fig. 9 is a diagram showing an LSI including the neural network arithmetic unit of the fifth embodiment. In the above-described embodiments, the configuration for delaying the supply start timing of the clock at the start of the operation has been described. In the present embodiment, an example of suppressing fluctuations in the power consumption of the neural network arithmetic unit by using the power consumption in the diagnostic circuit will be described.

[0069] That is, the neural network arithmetic unit of the present embodiment is a neural network arithmetic unit that performs neural network processing, and includes a plurality of processor elements arranged in an array, an activation memory that stores input activation data supplied to the processor elements, and a main controller that controls the operations of the respective processor elements. Either one or both of the respective processor elements or the main controller have an operation inspection processing function, and the main controller has a configuration for selectively instructing the processor elements to execute arithmetic processing and operation inspection processing. Hereinafter, it will be described in detail with reference to the drawings.

[0070] 9 shows the entire SOC. The SOC 9 often includes a CPU and a bus system (not shown) that control the entire system and functions as a system. 90 is the neural network arithmetic unit of the present embodiment. 91 is a clock generator that supplies a clock to the entire SOC9. This clock generator 91 also supplies a clock to the neural network arithmetic unit 90. Depending on the design, this clock may supply one type of frequency or multiple types of frequencies. In this embodiment, for simplicity, the case of supplying one type of clock will be described, but the present invention can also be realized with multiple types of clocks.

[0071] Depending on its configuration, the neural network arithmetic unit 90 may occupy more than 50% of the scale of the entire SOC. Therefore, the power consumption of the neural network arithmetic unit 90 can also become a main part of the entire SOC9 accordingly.

[0072] Next, the internal configuration of the neural network arithmetic unit 90 will be described. 901 is a main controller. The main controller 901 is a module that controls the processing sequence of the entire neural network arithmetic unit 90, manages input / output data, and issues operation requests to the processor elements described later. In addition, the main controller 901 also performs power control based on the information on the operating state of the neural network arithmetic unit 90.

[0073] 902 is a clock controller. The clock controller 902 supplies a clock to the processor element 904 described later and all other internal modules of the neural network arithmetic unit 90. The clock controller 902 is equipped with a mechanism for individually supplying and stopping clocks to a plurality of processor elements 904. In addition, a mechanism for controlling the clock frequency for each individual processor element PEij may be provided.

[0074] 903 is an activation memory. The activation memory 903 is a module that stores activation data to be supplied to the processor element 904 described later. The main controller 901 transfers the required activation data to this activation memory 903 in advance using a DMA controller or the like (not shown), and issues an instruction to supply data to the processor element 904 at the start of processing. Based on this instruction, the activation memory 903 supplies the required activation data to the processor element 904.

[0075] As for the implementation method of this part, the main controller 901 may issue an instruction to the activation memory 903 and the activation memory 903 may actively supply data to the processor element, or the main controller 901 may issue an instruction to the processor element 904, and the processor element 904 may read the required data from the activation memory 903.

[0076] 904 is a processor element. The processor element 904 includes one or more convolutional calculators composed of a multiplier and an adder that execute a convolutional operation with the input activation and the Weight value from the Weight input memory (not shown) as inputs. In many cases, a plurality of convolutional multipliers are implemented in the processor element, and it is configured to be able to execute a plurality of convolutional operations simultaneously in parallel.

[0077] The term "processor element 904" is used when collectively referring to the processor elements. When referring to an individual processor element, it is called a processor element "PEij". Here, ij is a number that identifies the processor element PE by its position, i (i = 0, ···, m) identifies the row, and j (j = 0, ···, n) identifies the column.

[0078] The output activation, which is the calculation output, is either retained again in the activation memory 903 or output outside the module.

[0079] These processor elements 904 are logically arranged in a two-dimensional array inside the neural network arithmetic unit 90. The processor elements adjacent to each other vertically and horizontally transfer activation data to each other during the execution process of the convolution operation, thereby realizing a convolution operation with a kernel size of 1 or more.

[0080] For example, when PE11 in FIG. 9 performs a convolution operation with a kernel size of 2x2, PE11 can obtain the activation data input from the activation memory 903 to PE10, thereby obtaining the activation data of the left neighbor of the processing target. Similarly, it also receives the activation data of the upper neighbor from PE01 and the activation data once transmitted from PE00 to PE01, and has a structure for using in the operation. Alternatively, the activation data of the adjacent 2x2 region may be obtained from the activation memory 903.

[0081] 905 is a diagnostic circuit. In FIG. 9, the diagnostic circuit is described as "BIST". This is an abbreviation of "Build in soft test". The diagnostic circuit 905 is a circuit that generates or holds the operation expected value corresponding to the input pattern when performing the diagnostic process, and compares the operation result of the processor element 904 with respect to the input pattern with the operation expected value. One such circuit is arranged for each processor element 904. By periodically executing this function, it becomes possible to detect hardware failures of the processor element 904.

[0082] The diagnostic circuit 905 is respectively connected to the main controller 901. The main controller 901 can select whether to supply normal activation data to each processor element 904 to execute the operation or to supply pattern data from the diagnostic circuit 905 to execute the operation.

[0083] The diagnostic circuit 905 is arranged adjacent to each processor element in the figure, and when a pattern from the diagnostic circuit 905 is selected by a selection signal from the main controller 901, it can supply a diagnostic data pattern to the corresponding processor element 904.

[0084] On the other hand, the diagnostic circuit 905 may be implemented inside the main controller 901 so that a diagnostic pattern can be supplied to each processor element 904. In the case of this configuration, when at least one processor element 904 executes a diagnostic process, the diagnostic circuit 905 generates a diagnostic pattern and sends the diagnostic pattern together with a selection signal generated by the main controller 901 to the target processor element 904.

[0085] The diagnostic process executed by the diagnostic circuit 905 does not require data transfer between the processor elements 904 compared to the case of computing normal activation data, and the processing is completed for each processor element 904. In addition, the pattern used for diagnosis only needs to achieve a predetermined toggle rate, and the power consumed by each processor element 904 can be controlled to a certain extent according to the toggle rate.

[0086] Next, the processing flow of the clock and the diagnostic process using these configurations will be described. In this process, the main controller 901 instructs the clock controller 902 to supply a clock. Among the processor elements 904 arranged two-dimensionally, one clock is supplied to the processor elements 904 in the same column.

[0087] With the above configuration, the main controller 901 enables the diagnostic circuit 905 and sequentially adjusts the timing of applying the clock for each column. In the present embodiment, the timing of applying the clock may be simultaneous, may be shifted by one cycle for each PE column, or may be shifted by a larger number of cycles. At this time, in the case of one cycle, the toggle pattern is set such that the power consumed by the diagnostic circuit is likely to be smaller than the power consumed by the actual arithmetic processing.

[0088] Also, in the case of multiple cycles, the diagnostic pattern generated by the diagnostic circuit 905 is preferably a pattern in which the toggle rate gradually increases. Further, when the arithmetic unit of the processor element 904 is composed of a plurality of arithmetic units, it is not necessary to start all of them in a way that can realize a series of convolution arithmetic processes. It is sufficient to confirm the validity of each arithmetic unit, and each arithmetic unit may be activated.

[0089] Thereafter, the main controller 901 sequentially switches to the actual multiplication and accumulation arithmetic processing from the processor element 904 in which the processing of the diagnostic circuit 905 has been completed. As a result, before the normal convolution arithmetic processing, the power consumption of the module can be gradually increased without depending on the data flow of the convolution arithmetic and without wasting power consumption.

[0090] The above will be described taking FIG. 10 as an example. FIG. 10 is a diagram showing the operation in the column direction of the processor element 904 with time on the horizontal axis and the vertical axis. As an example, in FIG. 10, the number of columns (i.e., the number of groups) of the processor element 904 is 8. In this example, for simplicity, the main controller 901 supplies all the clocks to the clock controller 902 simultaneously.

[0091] On the above premise, the main controller 901 first activates the diagnostic circuit in the first column. The toggle rate of the pattern of this diagnostic circuit is adjusted to be around 25% for example. At the timing when the processing of the diagnostic circuit in the first column is completed, the main controller 902 activates the diagnostic circuit in the second column. The toggle rate of the pattern of this diagnostic circuit is adjusted to be around 50% for example. Similarly, the diagnostic circuit in the third column with a toggle rate of 75% and the diagnostic circuit in the fourth column with a toggle rate of 100% are sequentially activated.

[0092] After that, in the example of FIG. 10, for those where the diagnosis is completed, the actual convolution operation is sequentially started. In the process of this convolution operation, since data transfer occurs between the processor elements 904 as described above, a startup deviation that matches the latency of the data transfer between the processor elements will occur. On the other hand, since the diagnostic process does not depend on the data transfer, the activation timing can be set more freely.

[0093] In the example of FIG. 10, the diagnostic processes and the convolution operation processes from the fourth column to the eighth column overlap. In this case, the toggle rate of the diagnostic patterns from the fourth column to the eighth column may be implemented with a lower value. Also, if the column patterns are always fixed, there may be a difference in the defect detection rate between the diagnostic circuits with a low toggle rate and those with a high toggle rate. Therefore, by appropriately swapping the order in which the diagnostic circuits are activated, the activation rate can be averaged.

[0094] In this way, the neural network arithmetic unit 80 of the present embodiment can suppress power fluctuations by controlling the combination of the clock supply to the processor element 904 and the diagnostic circuit 905.

[0095] (Sixth Embodiment) The basic configuration of the neural network arithmetic unit of the sixth embodiment is the same as that of the neural network arithmetic unit 50 of the fifth embodiment. The neural network arithmetic unit of the sixth embodiment sequentially increases the power by changing the number of activated diagnostic circuits in the neural network arithmetic unit 90 of the fifth embodiment.

[0096] FIG. 11 is a diagram for explaining the processing by the neural network arithmetic unit according to the sixth embodiment. In FIG. 11, the main controller 901 performs a diagnosis process on the processor elements 904 in the 0th column and the 4th column using the diagnosis circuit 905. Therefore, the number of processor element columns operating at this time is 2. At the same timing when this process is completed, the main controller 901 performs the same process on the 2nd column and the 6th column. Therefore, the number of processor element columns operating at this timing is also 2.

[0097] Thereafter, the main controller 901 activates a diagnosis process for the remaining four processor element columns. Thereafter, the main controller 901 executes a convolution operation process on all the processor elements. As a result, the operating rate of the processor elements gradually increases, and it is possible to gently increase the power consumption. In the present embodiment, the diagnosis circuit is activated in the above pattern, but the number of operating processor elements 904 may gradually increase, and the pattern is not limited to the above.

[0098] In addition, in the fifth and sixth embodiments described above, an example in which groups are formed in units of columns and clock signals are supplied to the processor elements 904 has been given. Needless to say, it is also possible to adopt a configuration in which rows and columns are interchanged and groups are formed in units of rows to supply clock signals. Furthermore, for these embodiments, the delay control of clock supply described in the first to fourth embodiments can also be combined and implemented.

[0099] (Seventh Embodiment) The basic configuration of the neural network arithmetic unit according to the seventh embodiment is the same as that of the neural network arithmetic unit 90 according to the fifth embodiment (see FIG. 9). In the seventh embodiment, the power is stabilized by changing the number of activated diagnosis circuits according to the state of the neural network process.

[0100] That is, the neural network arithmetic unit according to the present embodiment is a neural network arithmetic unit that performs neural network processing, and includes a plurality of processor elements arranged in an array, an activation memory (103) that stores input activation data supplied to the processor elements, and a main controller (101) that controls the operations of the respective processor elements. The main controller detects the type of neural network processing, and has a configuration that instructs a part of the processor elements to perform an arithmetic inspection process based on the power consumption defined in advance according to the type of neural network processing. Hereinafter, a detailed description will be given with reference to the drawings.

[0101] The main controller instructs the processor elements to perform arithmetic processing. At this time, depending on the neural network processing, there is processing that does not use all of the processor elements. For example, there is a process of simply connecting two networks in the channel direction without performing a convolution operation. In such a process, compared with the case of executing a convolution operation, the power consumption may be drastically reduced by the amount that the arithmetic unit does not operate.

[0102] Also, as another example, in the case of the edge of an image, etc., not all of the processor elements perform processing, and only a part may be activated. In this case as well, the power consumption may decrease.

[0103] In the present embodiment, in view of the above background, according to the neural network processing, the main controller instructs the processor elements or the arithmetic processing units of the processor elements to execute diagnostic processing.

[0104] FIG. 12 is a diagram for explaining the processing of the neural network arithmetic processing unit according to the seventh embodiment. The neural network arithmetic unit according to the seventh embodiment, similar to the fifth embodiment, supplies clocks to the processor elements in groups on a column-by-column basis.

[0105] In the example shown in FIG. 12, in the initial process, only the processor elements in the even rows perform processing, and the processor elements in the odd rows only execute transfer. At the next processing timing, only the odd rows perform processing, and the even rows perform transfer processing. At this timing, the main controller can initiate diagnostic processing for the processor elements in the even rows. At this time, if there is a time difference between the actual arithmetic processing and the diagnostic processing, the diagnostic processing may be executed at any time.

[0106] By the processing described above, it is possible to suppress a sudden change in power consumption caused by the processing content of the processor elements.

[0107] (Eighth Embodiment) The basic configuration of the neural network arithmetic unit according to the eighth embodiment is the same as that of the neural network arithmetic unit 50 according to the fifth embodiment. In the above-described embodiments, the suppression of power fluctuations associated with starting clock supply all at once when starting the operation of the neural network arithmetic unit has been described. However, power fluctuations can also occur when stopping clock supply.

[0108] In the neural network arithmetic unit according to the eighth embodiment, at the end of processing, the main controller gradually stops the clock for the processor elements in the reverse order of FIG. 5. Also, at the end of processing, the main controller gradually stops the clock for the processor elements and the diagnostic circuit in the reverse order of FIG. 10 or FIG. 11.

[0109] Further, when the currently executing process is completed, the main controller determines whether the current process is the final layer of the neural network process, and only when the neural network process is the final layer, proceeds to gradual clock stop. On the other hand, when it is not the final layer, the procedure for gradual clock stop may be canceled and the clock supply may continue.

[0110] With this function, when the layer processing continues, the processing can continue without degrading the performance, while avoiding a sudden change in power consumption when the layer processing ends.

[0111] (The Ninth Embodiment) FIG. 13 is a diagram showing the configuration of the neural network arithmetic unit according to the ninth embodiment. In the above-described embodiments, the unit of clock supply has been described assuming a processor element. In this embodiment, a more detailed control for suppressing fluctuations in power consumption will be described.

[0112] The neural network arithmetic unit according to this embodiment is a neural network arithmetic unit (10) that performs neural network processing, and includes a plurality of processor elements (104) arranged in an array, an activation memory (103) that stores input activation data supplied to the processor elements, and a main controller (101) that controls the operations of the respective processor elements. Each of the processor elements has a plurality of processing units that execute different functions, and has a configuration in which a clock can be supplied independently to each processing unit. The main controller individually controls the activation timings for the plurality of processing units of each processor element at the start of arithmetic processing, and has a configuration in which the processor elements are activated step by step. Hereinafter, a detailed description will be given with reference to the drawings.

[0113] The basic configuration of the neural network arithmetic unit according to the ninth embodiment is the same as that of the first embodiment. However, in the neural network arithmetic unit 130 according to this embodiment, the processor element 1304 is divided into an arithmetic unit 13041 and a data transfer unit 13042 as an internal structure. In this embodiment, the processor element is divided and controlled into the arithmetic unit 13041 and the data transfer unit 13042, and the power is sequentially increased by a combination of data calculation and arithmetic processing.

[0114] 13041 is an arithmetic unit, which is an arithmetic unit that performs a multiply-accumulate operation based on an input activation and a weight value. The arithmetic unit 13041 is composed of at least one or more multiply-accumulate units, and usually, a plurality of units are implemented according to the required performance. Also, depending on the supported layer processing, operations other than the multiply-accumulate operation may be executable.

[0115] 13042 is a data transfer unit. When performing a convolution operation using a kernel of 2x2 or more, the data transfer unit transfers activation data in order to share the activation data between adjacent processor elements.

[0116] The arithmetic unit 13041 and the data transfer unit 13042 are structured such that clock supply can be executed independently from the outside of the processor element 1304. Also, the reset control may be configured to be independent. Alternatively, in the present embodiment, although it is premised that the clock supply is controlled from outside the processor element, a configuration in which the clock is controlled inside the processor element 1304 may be used.

[0117] During normal convolution operation processing, both the arithmetic unit 13041 and the data transfer unit 13042 are executed simultaneously. If only one of the arithmetic unit 13041 or the data transfer unit 13042 is executed, the power consumption becomes smaller compared to normal processing. Also, when comparing the circuits of the arithmetic unit 13041 and the data transfer unit 13042, generally, the arithmetic unit 13041 has a larger logic scale, and thus the power consumption amount is also larger. Utilizing the above characteristics, in the present embodiment, processing is executed in the following sequence.

[0118] FIG. 14 is a diagram showing the processing sequence of the present embodiment. In FIG. 14, for simplicity, it is assumed that eight columns of processor elements are arranged, but the present invention is not limited to eight columns. FIG. 14 shows the operating state of each processor element when a plurality of convolutional layers are continuously processed. The horizontal axis represents time, and the vertical axis represents the column of processor elements. FIG. 14 shows the sequence when three layers are continuously executed, but this sequence is valid as long as at least two or more layers of layer processing can be executed.

[0119] When the main controller 1301 executes the Conv0-th layer processing, first, for the even-numbered columns (columns 0, 2, 4, 6), it instructs both the arithmetic unit 13041 and the data transfer unit 13042 to operate, and at the same time, for the odd-numbered columns (columns 1, 3, 5, 7), it instructs only the data transfer unit 13042 to operate and does not operate the arithmetic unit 13041. At this time, since the data transfer unit 13042 is executed by all the processor elements 1304, the data required by all the processor elements 1304 is supplied as usual, while the arithmetic processing is performed only on the even-numbered columns, and no arithmetic is performed on the odd-numbered columns.

[0120] Next, the main controller 1301 instructs both the arithmetic unit 13041 and the data transfer unit 13042 to operate for the odd-numbered columns (columns 1, 3, 5, 7), and at the same time, for the even-numbered columns (columns 0, 2, 4, 6), it instructs only the data transfer unit 13042 to operate and does not operate the arithmetic unit 13041. By this combination, the arithmetic processing of the odd-numbered columns is executed, and the Conv0-th layer processing is completed in combination with the previous processing.

[0121] The processor element 1304 continues to execute the Conv1 layer, while the main controller 1301 issues an instruction to activate both the arithmetic unit 13041 and the data transfer unit 13042 for all the processor elements 1304 after the first Conv1 layer. As a result, all of the processor elements operate. Since the operating rate of the arithmetic unit during the processing of the Conv0 layer is limited to 50% of the normal rate, the power consumption can be controlled to be lower compared to the power during the processing of the layers after Conv1.

Explanation of Signs

[0122] 1 SOC, 10 neural network arithmetic units, 11 clock generators, 101 main controller, 102 clock controller, 103 activation memory, 104 processing element, 201~204 Enable signals, 205~208 clock signals, 401~404 Enable signals, 405~408 clock signals, 6 SOC, 60 neural network arithmetic units, 61 clock generators, 601 main controller, 602 clock controller, 603 activation memory, 604 processing element, 605~609 groups, 8 SOC, 80 neural network arithmetic units, 81 clock generators, 801 main controller, 802 clock controller, 803 activation memory, 804 processing element, 805~807 groups, 9 SOC, 90 neural network arithmetic units, 91 clock generators, 901 main controller, 902 clock controller, 903 activation memory, 904 processing element, 905 diagnostic circuit, 13 SOC, 130 neural network arithmetic unit, 131 clock generator, 1301 main controller, 1302 clock controller, 1303 activation memory, 1304 processing element, 13041 arithmetic unit, 13042 data transfer unit.

Claims

1. A neural network arithmetic unit (10) that performs neural network processing, comprising: a plurality of processor elements (104) arranged in an array; an activation memory (103) that stores input activation data supplied to the processor elements; a main controller (101) that controls the operations of each of the processor elements; and clock lines of a clock generator radially form groups (605 to 609) that logically group processor elements (604) from the endpoints of the processor elements arranged in an array, and clock lines are independently connected to each of the groups, wherein the main controller (601) activates the arithmetic processing for each of the radially connected groups and instructs the clock generator (61) to supply a clock to the group of the processor elements, and sequentially supplies a clock from the endpoints of the processor elements in synchronization with the supply of data for neural network arithmetic, a neural network arithmetic unit.

2. A neural network arithmetic unit (10) that performs neural network processing, comprising: a plurality of processor elements (104) arranged in an array; an activation memory (103) that stores input activation data supplied to the processor elements; a main controller (101) that controls the operations of each of the processor elements; and either one or both of each of the processor elements or the main controller has an arithmetic inspection processing function, wherein the main controller controls the timing of supplying a clock supplied from a clock generator (11) to each of the processor elements, and selectively instructs the processor elements to execute arithmetic processing and arithmetic inspection processing, a neural network arithmetic unit.

3. A neural network arithmetic unit (10) that performs neural network processing, comprising: a plurality of processor elements (104) arranged in an array; an activation memory (103) that stores input activation data supplied to the processor elements; a main controller (101) that controls the operations of each of the processor elements; comprising; The main controller (101) controls the timing of supplying the clock supplied from the clock generator (11) to each processor element, and before the start of neural network processing, gives an execution instruction so that the number of execution times of the arithmetic inspection process gradually increases. A neural network arithmetic unit. **Claim 4** A neural network arithmetic unit (10) that performs neural network processing, a plurality of processor elements (104) arranged in an array; an activation memory (103) that stores input activation data supplied to the processor element; a main controller (101) that controls the operations of the respective processor elements; comprising; The main controller (101) controls the timing of supplying the clock supplied from the clock generator (11) to each processor element, detects the type of neural network processing, and based on the power consumption defined in advance according to the type of the neural network processing, A neural network arithmetic unit that instructs an arithmetic inspection process for a part of the processor elements. **Claim 5** The processor elements arranged in an array have a function of transferring instructions from the main controller between adjacent processor elements, The clock lines of the clock generator are connected for each group obtained by logically grouping the processor elements, The main controller supplies a clock to the processor elements for each group in synchronization with the supply of data for neural network operations to the processor elements, and increases the group that supplies the clock as time elapses from the start of the arithmetic process. The neural network arithmetic unit according to any one of claims 2 to 4. **Claim 6** The processor elements arranged in an array have a function of transferring instructions from the main controller between adjacent processor elements, The clock lines of the clock generator are connected for each column of the processor elements, The main controller starts arithmetic processing for each column, instructs the clock generator to supply the clock for the column, and sequentially supplies the clock from one end side of the column of the processor elements in synchronization with the supply of data for neural network arithmetic. The neural network arithmetic device according to any one of claims 2 to 4.

7. The clock lines of the clock generator form groups of processor elements according to the physical arrangement of the processor elements (804) arranged and wired on the semiconductor silicon, and the clock lines are independently connected to each of the groups. The main controller starts arithmetic processing for each of the groups and instructs the clock generator to supply the clock for the processor elements, and sequentially supplies the clock so that the power consumption on the semiconductor silicon becomes uniform. The neural network arithmetic device according to any one of claims 2 to 4.

8. The main controller has a function of detecting the progress state of neural network processing, and determines whether to continue clock supply or stop the clock according to the progress state of neural network processing. The neural network arithmetic device according to any one of claims 1 to 7.

9. Each of the processor elements has a plurality of processing units that execute different functions, and has a configuration in which clocks can be independently supplied to each processing unit. The main controller individually controls the activation timings for the plurality of processing units of each processor element at the time of starting arithmetic processing, and activates the processor elements step by step. The neural network arithmetic device according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Semiconductor integrated circuit

    JP2006048467A

  • Semiconductor integrated circuit, and its power saving control method and power saving control program

    JP2006065471A

  • Clock control device for semiconductor device

    JP2006079505A

  • Control device and semiconductor integrated circuit

    JP2007233718A

  • Power management device and power management method

    JP2009037335A