Semiconductor device

The semiconductor device stabilizes current consumption and power supply fluctuations through dummy operations, addressing the challenges of miniaturization-induced power supply instability and cost increases in neural network processing.

JP7701296B2Active Publication Date: 2025-07-01RENESAS ELECTRONICS CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022043264
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-07-01
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

The miniaturization of semiconductor manufacturing processes and increased circuit maturity lead to higher arithmetic operation efficiency in neural networks, resulting in sudden changes in consumption current and fluctuations in power supply voltage, which complicates power supply design and increases design and manufacturing costs.

Method used

A semiconductor device with n multiply-accumulate units, one or more memories, and DMA controllers, including a dummy circuit, performs dummy operations to stabilize current consumption, reducing sudden changes and fluctuations in power supply voltage.

Benefits of technology

The solution effectively suppresses sudden decreases and fluctuations in power consumption, simplifies power supply design, and reduces manufacturing costs by controlling current consumption rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701296000001
    Figure 0007701296000001
  • Figure 0007701296000002
    Figure 0007701296000002
  • Figure 0007701296000003
    Figure 0007701296000003
Patent Text Reader

Abstract

To provide a semiconductor device capable of suppressing sudden decrease fluctuations in current consumption in processing of a neural network.SOLUTION: A dummy circuit 22 executes a dummy operation by outputting dummy data DTd to at least a portion of n MAC circuits 25(1) to 25(n), and outputs dummy output data DToD. An output side DMA controller DMAC2oB transfers normal output data DTo from the n MAC circuits to a memory by using n channels CH(1) to CH(n), and does not transfer the dummy output data DToD to the memory. Here, at least a portion of the n MAC circuits executes a dummy operation within a period from the end of data transfer by an output side DMA controller DMAC2oB to the memory till a start of data transfer from the memory by an input side DMA controller DMAC2i.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a semiconductor device, for example, a semiconductor device that executes neural network processing.

Background Art

[0002] Patent Document 1 discloses a technique that enables reduction of an operating current flowing on a signal bus and accurate capture of a large amount of data during data transfer in a semiconductor device including a logic device and a memory device. In the semiconductor device, a data signal having an amplitude smaller than the amplitude of a power supply voltage, a first clock signal, and a second clock signal phase-shifted by a predetermined amount from the first clock signal are used. Each of the logic device and the memory device captures data in synchronization with the rising edges of the first and second clock signals.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For example, in the processing of neural networks such as CNN (Convolutional Neural Network), a huge amount of arithmetic processing is executed using a plurality of DMA (Direct Memory Access) controllers and a plurality of multiply-accumulate units (referred to as MAC (Multiply ACcumulate) circuits) mounted on a semiconductor device. Specifically, the plurality of DMA controllers transfer image data and coefficient data of a certain layer stored in the memory to the plurality of MAC circuits, causing the plurality of MAC circuits to perform multiply-accumulate operations. Also, the plurality of DMA controllers transfer the multiply-accumulate operation results by the plurality of MAC circuits to the memory as image data of the next layer. The semiconductor device repeatedly executes such processing.

[0005] On the other hand, in semiconductor devices, the miniaturization of the manufacturing process and the maturation of circuits are progressing. As a result, the processing efficiency of neural networks has increased, and the number of arithmetic operations that can be executed within a unit time has increased. Along with this, the consumption current has a tendency to increase. Here, when the period during which arithmetic operations are being performed is defined as the active period, and the waiting period for transitioning to the active period is defined as the idle period, usually, in a plurality of MAC circuits, the idle period and the active period are switched simultaneously. Thereby, the time required for neural network processing can be shortened to the maximum extent.

[0006] However, when such simultaneous switching is performed, a sudden change in the consumption current occurs, and fluctuations in the power supply voltage can occur due to parasitic inductance components of the power supply wiring. The fluctuations in the power supply voltage can become larger as the consumption current increases, and further, as the change rate of the consumption current becomes larger. In order to suppress the fluctuations in the power supply voltage, for example, it is necessary to strengthen the power supply design of the semiconductor device. However, in this case, the difficulty of the design increases, and there is a risk that the design cost and the manufacturing cost will increase.

[0007] The embodiments described below are made in view of such circumstances, and other problems and novel features will become apparent from the description of this specification and the accompanying drawings.

Means for Solving the Problems

[0008] A semiconductor device according to an embodiment executes neural network processing, and includes n multiply-accumulate units, one or more memories, a first DMA controller, a second input-side DMA controller, a dummy circuit, and a second output-side DMA controller. The n multiply-accumulate units perform a multiply-accumulate operation on input data and parameters. The one or more memories store the input data and parameters. The first DMA controller transfers the parameters stored in the memory to the n multiply-accumulate units. The second input-side DMA controller transfers the input data stored in the memory to the n multiply-accumulate units respectively using n channels, causing the n multiply-accumulate units to execute an operation and output normal output data as an operation result. The dummy circuit outputs predetermined dummy data to at least a part of the n multiply-accumulate units, causing at least a part of the n multiply-accumulate units to execute a dummy operation and output dummy output data as an operation result. The second output-side DMA controller transfers the normal output data from the n multiply-accumulate units to the memory respectively using n channels, and does not transfer the dummy output data from at least a part of the n multiply-accumulate units to the memory. Here, at least a part of the n multiply-accumulate units executes a dummy operation within a period from when the second output-side DMA controller finishes data transfer to the memory until the second input-side DMA controller starts data transfer from the memory.

Advantages of the Invention

[0009] By using the semiconductor device according to an embodiment, it is possible to suppress a sudden decrease and fluctuation in power consumption.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0011] In the following embodiments, when necessary for convenience, the description will be divided into a plurality of sections or embodiments. However, unless otherwise specified, they are not independent of each other, and one is related to a partial or complete modification, detail, supplementary explanation, etc. of the other. Also, in the following embodiments, when referring to the number of elements, etc. (including the number, numerical value, quantity, range, etc.), unless otherwise specified or limited to a specific number in principle, it is not limited to that specific number, and it may be more than or less than the specific number. Furthermore, in the following embodiments, it goes without saying that the constituent elements (including element steps, etc.) are not necessarily essential, unless otherwise specified or considered to be essential in principle. Similarly, in the following embodiments, when referring to the shape, positional relationship, etc. of the constituent elements, unless otherwise specified or considered not to be so in principle, it includes those substantially approximating or similar to the shape, etc. This also applies to the above numerical values and ranges.

[0012] Hereinafter, the embodiments will be described in detail with reference to the drawings. In all the drawings for explaining the embodiments, members having the same function are denoted by the same reference numerals, and repeated descriptions thereof are omitted. Also, in the following embodiments, the description of the same or similar parts will not be repeated in principle unless particularly necessary.

[0013] (Embodiment 1) <Schematic of the semiconductor device> FIG. 1 is a schematic diagram showing a configuration example of a main part in a semiconductor device according to Embodiment 1. The semiconductor device 10 shown in FIG. 1 is, for example, an SoC (System on Chip) composed of one semiconductor chip. Typically, the semiconductor device 10 is mounted on an ECU (Electronic Control Unit) of a vehicle or the like and provides functions of an ADAS (Advanced Driver Assistance System).

[0014] The semiconductor device 10 shown in FIG. 1 includes a neural network engine (NNE) 15a, a processor 17 such as a CPU (Central Processing Unit), one or more memories MEM1 and MEM2, and a system bus 16. The system bus 16 connects the neural network engine 15a, the memories MEM1 and MEM2, and the processor 17 to each other. The neural network engine 15a executes processing of a neural network represented by CNN. The processor 17 executes a predetermined program stored in the memory MEM1 to cause the semiconductor device 10 to perform a predetermined function including control of the neural network engine 15a.

[0015] The memory MEM1 is a DRAM (Dynamic Random Access Memory) or the like, and the memory MEM2 is a SRAM (Static Random Access Memory) for cache or the like. The memory MEM1 stores, for example, data DT consisting of pixel values, a parameter PR, and a command CMD. The parameter PR includes a weight parameter WP and a bias parameter BP. The command CMD is for controlling the sequence operation of the neural network engine 15a. The memory MEM2 is used as a high-speed cache memory of the neural network engine 15a. For example, a plurality of data DT in the memory MEM1 are copied to the memory MEM2 in advance and then used by the neural network engine 15a.

[0016] The neural network engine 15a includes a plurality of DMA (Direct Memory Access) controllers DMAC1 and DMAC2, a MAC unit 20, and a sequence controller 21a. The MAC unit 20 includes a plurality of MAC circuits 25, that is, a plurality of multiplier-accumulators. The DMA controller DMAC1 controls data transfer via the system bus 16, for example, between the memory MEM1 and the plurality of MAC circuits 25 in the MAC unit 20. The DMA controller DMAC2 controls data transfer between the memory MEM2 and the plurality of MAC circuits 25 in the MAC unit 20.

[0017] Specifically, the DMA controller DMAC1 transfers the parameter PR stored in the memory MEM1 to the plurality of MAC circuits 25 in the MAC unit 20. Also, the DMA controller DMAC1 transfers the command CMD stored in the memory MEM1 to the sequence controller 21a.

[0018] On the other hand, the DMA controller DMAC2 transfers the data stored in the memory MEM2 as input data DTi to the plurality of MAC circuits 25 in the MAC unit 20, causing the plurality of MAC circuits 25 to execute operations. Specifically, the plurality of MAC circuits 25 execute operations such as multiplying and accumulating the input data DTi from the DMA controller DMAC2 and the weight parameter WP from the DMA controller DMAC1, and adding the bias parameter BP from the DMA controller DMAC1.

[0019] As a result, the multiple MAC circuits 25 output output data DTo which is the result of the operation. The output data DTo represents, for example, the pixel values of the feature maps obtained from each layer of the neural network. The DMA controller DMAC2 transfers the output data DTo to the memory MEM2. The output data DTo transferred to the memory MEM2 is used as input data DTi to the next layer of the neural network. That is, for example, the input data DTi to the first layer of the neural network is determined by the data DT stored in the memory MEM1, and the input data DTi to the second and subsequent layers is determined by the output data DTo from the multiple MAC circuits 25.

[0020] The sequence controller 21a controls the operation sequence etc. of the neural network engine 15a based on the command CMD from the DMA controller DMAC1. As one of them, the sequence controller 21a outputs a read start signal to the DMA controller DMAC2 to start data transfer from the memory MEM2. Also, the sequence controller 21a performs transfer settings on the DMA controller DMAC2, for example, setting the address range of the memory MEM2 where the input data DTi is stored, setting the address range of the memory MEM2 where the output data DTo is stored, etc. <Configuration of Neural Network Engine>

[0021] FIG. 2 is a diagram showing a detailed configuration example of the neural network engine in FIG. 1. In FIG. 2, the MAC unit 20 has n (n is an integer of 2 or more) MAC circuits 25[1] to 25[n]. The value of n is, for example, 16 or the like. The DMA controller DMAC1 reads information from the memory MEM1 every control cycle based on a preset address range. The read information appropriately includes the parameter PR and the command CMD. The DMA controller DMAC1 transfers the read parameter PR to the n MAC circuits 25[1] to 25[n], and stores the read command CMD in the register REG.

[0022] The DMA controller DMAC2 shown in FIG. 1 specifically includes an input-side DMA controller DMAC2i and an output-side DMA controller DMAC2oA as shown in FIG. 2. Each of the input-side DMA controller DMAC2i and the output-side DMA controller DMAC2oA has n channels CH[1] to CH[n].

[0023] The input-side DMA controller DMAC2i transfers the input data DTi stored in the memory MEM2 to each of the n MAC circuits 25[1] to 25[n] using the n channels CH[1] to CH[n], causing the n MAC circuits 25[1] to 25[n] to execute operations. Address ranges for reading from the memory MEM2 are set for each of the n channels CH[1] to CH[n].

[0024] Specifically, for example, the MAC circuit 25[1] performs a multiply-accumulate operation on a plurality of input data DTi from the channel CH[1] of the input-side DMA controller DMAC2i and a plurality of weight parameters WP from the DMA controller DMAC1. The MAC circuit 25[1] also adds the bias parameter BP from the DMA controller DMAC1 to the multiply-accumulate operation result to output the output data DTo as the operation result.

[0025] As a more detailed configuration example, the channel CH[1] of the input-side DMA controller DMAC2i reads, for example, "M×K" input data DTi, where "M" is the number of input channels of the neural network and "K" is the kernel size, and transfers them to the MAC circuit 25[1]. On the other hand, the DMA controller DMAC1 also reads "M×K" weight parameters WP and transfers them to the MAC circuit 25[1].

[0026] The MAC circuit 25[1] includes, for example, "M×K" multipliers and an adder that adds the multiplication results of these multipliers. Thereby, the MAC circuit 25[1] performs "M×K" multiply-accumulate operations, and by separately adding the bias parameter BP to the multiply-accumulate operation result, it outputs output data DTo representing the value of one coordinate in the feature map. The same applies to the other MAC circuits 25[2] to 25[n].

[0027] At this time, the other MAC circuits 25[2] to 25[n] may perform operations on different input data DTi, that is, input data DTi with different coordinate ranges due to the convolution operation, or may perform operations on the same input data DTi. In the former case, a common parameter PR is used in the n MAC circuits 25[1] to 25[n]. On the other hand, in the latter case, different parameters PR are used in the n MAC circuits 25[1] to 25[n]. That is, in the latter case, the n MAC circuits 25[1] to 25[n] are respectively assigned to different output channels in the neural network.

[0028] The output-side DMA controller DMAC2oA transfers the output data DTo from the n MAC circuits 25[1] to 25[n] to the memory MEM2 using the n channels CH[1] to CH[n], respectively. Address ranges for writing to the memory MEM2 are set for each of the n channels CH[1] to CH[n].

[0029] The sequence controller 21a controls the operation sequences of the input-side DMA controller DMAC2i and the output-side DMA controller DMAC2oA based on the command CMD stored in the register REG. Specifically, the sequence controller 21a uses the control signal CS2i to perform transfer settings in the input-side DMA controller DMAC2i, such as setting the address range to be read from the memory MEM2. Similarly, the sequence controller 21a uses the control signal CS2o to perform transfer settings in the output-side DMA controller DMAC2oA, such as setting the address range to be written to the memory MEM2.

[0030] Furthermore, the sequence controller 21a can divide the n channels CH[1] to CH[n] in the input-side DMA controller DMAC2i and the output-side DMA controller DMAC2oA into m (where m is an integer smaller than n) groups GR[1] to GR[m] using the control signals CS2i and CS2o. By dividing the n channels CH[1] to CH[n] into m groups GR[1] to GR[m], as a result, the n MAC circuits 25[1] to 25[n] are also divided into m groups GR[1] to GR[m]. For example, when the value of n is 16 and the value of m is 4, each of the four groups GR[1] to GR[4] will have four channels and four MAC circuits belonging to it.

[0031] The sequence controller 21a can output read start signals RDS[1] to RDS[m] for each of the m groups GR[1] to GR[m] to the input-side DMA controller DMAC2i at mutually different timings. The read start signals RDS[1] to RDS[m] are signals for starting data transfer from the memory MEM2 to the m groups GR[1] to GR[m], respectively. Thereby, the sequence controller 21a can control the timings of a series of operations including the read operation by the input-side DMA controller DMAC2i, the arithmetic operation by the MAC unit 20, and the write operation by the output-side DMA controller DMAC2oA to be different from each other among the m groups GR[1] to GR[m].

[0032] To group such n channels CH[1] to CH[n], the input-side DMA controller DMAC2i includes a grouping circuit 26. The grouping circuit 26 groups the n channels CH[1] to CH[n] into m groups based on the control signal CS2i from the sequence controller 21a. That is, the m groups GR[1] to GR[m] can be changed by setting via the control signal CS2i. The grouping circuit 26 determines the correspondence between the n channels CH[1] to CH[n] and the read start signals RDS[1] to RDS[m] based on this setting.

[0033] <Operation of Neural Network Engine (Comparative Example)> FIG. 12 is a timing chart showing an operation example of a neural network engine as a comparative example. The neural network engine as a comparative example does not have the grouping function as described in FIG. 2. In this case, as shown in FIG. 12, a series of operations including the read operation in period T1, the arithmetic operation in period T2, and the write operation in period T3 are executed at the same timing for the n (n = 16 in this example) channels CH[1] to CH

[16] .

[0034] Specifically, during period T1, the 16 channels CH[1] to CH

[16] in the input-side DMA controller simultaneously transfer the input data DTi from the memory MEM2 to the 16 MAC circuits 25[1] to 25

[16] . During period T2, the 16 MAC circuits 25[1] to 25

[16] simultaneously execute operations. During period T3, the 16 channels CH[1] to CH

[16] in the output-side DMA controller simultaneously transfer the output data DTo from the 16 MAC circuits 25[1] to 25

[16] to the memory MEM2. After that, through period T4 which is an idle period, a series of operations are performed again in active periods T1 to T3. During period T4, for example, in the input-side / output-side DMA controller, changes in transfer settings, that is, changes in the address range of the memory MEM2, etc. are made.

[0035] However, when using such an operation, when switching between the idle period and the active period, that is, when transitioning from period T3 to period T4, or from period T4 to period T1, the consumption current changes rapidly. When the consumption current changes rapidly, fluctuations in the power supply voltage can occur due to the parasitic inductor components of the power supply wiring, etc. In order to suppress fluctuations in the power supply voltage, it is necessary to strengthen the power supply design of the semiconductor device, typically represented by methods such as providing a MIM (Metal Insulator Metal) capacitor, strengthening the power bumps and power trunks. However, in this case, the design difficulty increases, and the design cost and manufacturing cost may increase.

[0036] <Operation of Neural Network Engine (Embodiment 1)> Figure 3 is a timing chart showing an operation example of the neural network engine shown in Figure 2. Using the configuration example of Figure 2, as shown in Figure 3, the timings of a series of operations consisting of the read operation in period T1, the operation operation in period T2, and the write operation in period T3 can be controlled to be different from each other in m groups (in this example, m = 4), namely groups GR[1] to GR[4].

[0037] Specifically, the start timing of period T1 in groups GR[1] to GR[4] is determined based on read start signals RDS[1] to RDS[4], respectively. The sequence controller 21a outputs the read start signals RDS[1] to RDS[4] in order while shifting the timing by a fixed period each time. As a result, the start timing and end timing of a series of active periods consisting of periods T1 to T3 are controlled to be different from each other among the four groups GR[1] to GR[4].

[0038] Taking group GR[1] as an example, in period T1, four out of 16 channels CH[1] to CH[4] in the input-side DMA controller DMAC2i simultaneously transfer input data DTi from the memory MEM2 to four out of 16 MAC circuits 25[1] to 25[4]. In period T2, the four MAC circuits 25[1] to 25[4] simultaneously execute operations. In period T3, four out of 16 channels CH[1] to CH[4] in the output-side DMA controller DMAC2oA simultaneously transfer output data DTo from the four MAC circuits 25[1] to 25[4] to the memory MEM2. After that, through period T4, which is an idle period, a series of operations are performed again in the active periods (periods T1 to T3).

[0039] In this way, by controlling the start timing and end timing of the active periods (periods T1 to T3) to be different from each other among the four groups GR[1] to GR[4], as shown in FIG. 3, it becomes possible to suppress a sudden change in the consumption current. In other words, it becomes possible to reduce the rate of change of the consumption current. Here, four groups are used, but the number of such groups can be determined, for example, in units of powers of two. The setting of the groups is performed, for example, by a command CMD before starting the processing of a predetermined layer in the neural network and is maintained while the processing of the predetermined layer is being executed.

[0040] <Main effects of Embodiment 1> In the method of Embodiment 1 described above, n channels and n MAC circuits in the DMA controller are divided into m groups, and the m groups are operated at different timings, so that it becomes possible to suppress a sharp decrease in the consumption current. As a result, it is possible to suppress fluctuations in the power supply voltage, facilitate the power supply design of the semiconductor device 10, and suppress an increase in the design cost and the manufacturing cost. Such an effect is obtained more remarkably especially as the number of arithmetic operations that can be executed per unit time increases due to miniaturization or the like of the semiconductor device 10.

[0041] (Embodiment 2) <Outline of the semiconductor device> FIG. 4 is a schematic diagram showing a configuration example of a main part in the semiconductor device according to Embodiment 2. The semiconductor device 10 shown in FIG. 4 has a different configuration of the neural network engine (NNE) 15b as compared with the configuration example of FIG. 1. In the neural network engine 15b shown in FIG. 4, a dummy circuit 22 is added as compared with the neural network engine 15a shown in FIG. 1. Along with this, the sequence controller 21b controls the dummy circuit 22 in addition to the same operation as in the case of FIG. 1.

[0042] The dummy circuit 22 outputs predetermined dummy data DTd to at least a part of the plurality of MAC circuits 25, so that dummy operations are executed on at least a part of the plurality of MAC circuits 25, and dummy output data as an operation result is output. However, the DMA controller DMAC2 does not transfer the dummy output data from at least a part of the plurality of MAC circuits 25 to the memory MEM2. That is, the DMA controller DMAC2 transfers the normal output data DTo from the plurality of MAC circuits 25 corresponding to the input data DTi to the memory MEM2, but does not transfer the dummy output data corresponding to the dummy data DTd to the memory MEM2.

[0043] <Configuration of the neural network engine> FIG. 5 is a diagram showing a detailed configuration example of the neural network engine in FIG. 4. Here, the description will focus on the differences between the neural network engine (NNE) 15b shown in FIG. 5 and the neural network engine (NNE) 15a shown in FIG. 2, and the description of matters overlapping with FIG. 2 will be omitted.

[0044] In FIG. 5, the output-side DMA controller DMAC2oB includes a grouping circuit 27. The grouping circuit 27 groups n channels CH[1] to CH[n] into m groups based on the control signal CS2o from the sequence controller 21b. That is, the m groups GR[1] to GR[m] can be changed by setting via the control signal CS2o.

[0045] The output-side DMA controller DMAC2oB outputs a write end signal when the data transfer to the memory MEM2 is completed. Specifically, the output-side DMA controller DMAC2oB outputs write end signals WTE[1] to WTE[m] at the end of data transfer for each of the m groups GR[1] to GR[m]. The grouping circuit 27 determines the correspondence between the n channels CH[1] to CH[n] and the write end signals WTE[1] to WTE[m] based on the setting via the control signal CS2o.

[0046] The dummy circuit 22 outputs dummy data DTd to at least a part of the n MAC circuits 25[1] to 25[n] in response to the write end signals WTE[1] to WTE[m] from the output-side DMA controller DMAC2oB. Also, the dummy circuit 22 stops the output of the dummy data DTd in response to the read start signals RDS[1] to RDS[m] from the sequence controller 21b, and outputs the input data DTi from the input-side DMA controller DMAC2i to the n MAC circuits 25[1] to 25[n].

[0047] As a result, at least some of the n MAC circuits 25[1] to 25[n] execute dummy operations within the period from when the output-side DMA controller DMAC2oB finishes data transfer to the memory MEM2 until the input-side DMA controller DMAC2i starts data transfer from the memory MEM2. However, as described in FIG. 4, the output-side DMA controller DMAC2oB does not transfer to the memory MEM2 the dummy output data DToD obtained by the dummy operations.

[0048] Note that, although details will be described later, the dummy circuit 22 performs grouping similar to that in the case of the input-side DMA controller DMAC2i based on the control signal CS2i from the sequence controller 21b. Also, the dummy circuit 22 can determine the number of MAC circuits 25 for which dummy operations are to be performed, etc., based on the control signal CS2d from the sequence controller 21b.

[0049] FIG. 6 is a diagram showing a schematic configuration example of the dummy circuit in FIG. 5. The dummy circuit 22 shown in FIG. 6 includes m sub-circuits 30[1] to 30[m] respectively corresponding to m groups GR[1] to GR[m], a dummy data generation circuit 31, a grouping circuit 32, and a switch controller 33. The dummy data generation circuit 31 generates dummy data DTd. The switch controller 33 includes, for example, an RS flip-flop or the like, inputs read start signals RDS[1] to RDS[m] and write end signals WTE[1] to WTE[m], and outputs normal data selection signals ISL[1] to ISL[m] and dummy data selection signals DSL[1] to DSL[m].

[0050] For example, the normal data selection signal ISL[1] is a signal that is set at the falling edge of the read start signal RDS[1] and reset at the rising edge of the write end signal WTE[1]. The dummy data selection signal DSL[1] is a signal that is set at the falling edge of the write end signal WTE[1] and reset at the rising edge of the read start signal RDS[1]. Similarly, the normal data selection signal ISL[m] is a signal that is set at the falling edge of the read start signal RDS[m] and reset at the rising edge of the write end signal WTE[m]. The dummy data selection signal DSL[m] is a signal that is set at the falling edge of the write end signal WTE[m] and reset at the rising edge of the read start signal RDS[m].

[0051] The partial circuit 30[1] receives the input data DTi from the channels CH[1], CH[2],... belonging to the group GR[1] in the input-side DMA controller DMAC2i and the dummy data DTd. The partial circuit 30[1] selects the input data DTi during the set period of the normal data selection signal ISL[1] of the group GR[1] and selects the dummy data DTd during the set period of the dummy data selection signal DSL[1] of the group GR[1] as the data to be sent to the MAC circuits 25[1], 25[2],... belonging to the group GR[1]. When the dummy data DTd is selected, the MAC circuits 25[1], 25[2],... belonging to the group GR[1] execute dummy operations.

[0052] Similarly, the partial circuit 30[m] receives the input data DTi from the channels CH[n], CH[n - 1],... belonging to the group GR[m] in the input-side DMA controller DMAC2i and the dummy data DTd. The partial circuit 30[m] selects the input data DTi during the set period of the normal data selection signal ISL[m] and selects the dummy data DTd during the set period of the dummy data selection signal DSL[m] of the group GR[m] as the data to be sent to the MAC circuits 25[n], 25[n - 1],... belonging to the group GR[m]. When the dummy data DTd is selected, the MAC circuits 25[n], 25[n - 1],... belonging to the group GR[m] execute dummy operations.

[0053] In this way, the dummy circuit 22 causes the MAC circuits 25 for each of the m groups GR[1] to GR[m] to execute dummy operations based on the write end signals WTE[1] to WTE[m] and the read start signals RDS[1] to RDS[m] for each of the m groups GR[1] to GR[m]. The grouping circuit 32 determines the correspondence between the n channels CH[1] to CH[n] and the read start signals RDS[1] to RDS[m] and the write end signals WTE[1] to WTE[m] based on the setting via the control signal CS2i from the sequence controller 21b.

[0054] <Operation of Neural Network Engine (Embodiment 2)> FIG. 7 is a timing chart showing an operation example of the neural network engine shown in FIG. 5. In the operation example shown in FIG. 7, the start timings of the active periods (periods T1 to T3) in the m groups GR[1] to GR[4] (m = 4 in this example) are the same, and the end timings of the active periods are also the same. In this case, the write end signals WTE[1] to WTE[4] in the groups GR[1] to GR[4] are output simultaneously at the end timing of the active period, that is, the end timing of period T3.

[0055] The dummy circuit 22 simultaneously starts outputting the dummy data DTd to the n MAC circuits 25[1] to 25[n] in response to the write end signals WTE[1] to WTE[4]. Accordingly, in period T4, the n MAC circuits 25[1] to 25[n] simultaneously start dummy operations. Thereafter, the read start signals RDS[1] to RDS[4] in the groups GR[1] to GR[4] are simultaneously input to the dummy circuit 22.

[0056] The dummy circuit 22 terminates the dummy operations in the n MAC circuits 25[1] to 25[n] by stopping the output of the dummy data DTd to the n MAC circuits 25[1] to 25[n] according to the lead start signals RDS[1] to RDS[4]. Then, instead of outputting the dummy data DTd, the dummy circuit 22 simultaneously starts outputting the normal input data DTi from the input-side DMA controller DMAC2i to the n MAC circuits 25[1] to 25[n] during the period T1.

[0057] In this way, by causing the n MAC circuits 25[1] to 25[n] to perform dummy operations, as shown in FIG. 7, it becomes possible to suppress a sharp change in the consumption current. In other words, it becomes possible to reduce the rate of change of the consumption current, specifically the transient current. Here, for the sake of convenience of explanation, the grouping described in the first embodiment is performed. However, when using the operation example as shown in FIG. 7, it is not always necessary to perform the grouping.

[0058] FIG. 8 is a timing chart showing an operation example different from that of FIG. 7. When using the operation example of FIG. 7, the rate of change of the consumption current can be reduced. On the other hand, due to the dummy operations, the consumption current may increase unnecessarily. Therefore, an operation example as shown in FIG. 8 may be used. In the operation example of FIG. 8, different from the operation example of FIG. 7, during the period T4, not all of the groups GR[1] to GR[4], but some groups, in this example, the MAC circuits 25 belonging to two of the GR[1] and GR[2] are performing dummy operations.

[0059] The some groups can be changed by setting via the control signal CS2d. That is, it is possible to set which group of the MAC circuits 25 belonging to will perform the dummy operation. The setting of the group to perform the dummy operation is performed by a command CMD, for example, before starting the processing of a predetermined layer in the neural network, and is maintained while the processing of the predetermined layer is being executed, similar to the setting of the grouping.

[0060] By causing only some of the MAC circuits 25, rather than all of them, to execute dummy operations, it becomes possible to suppress a rapid change in the current consumption rate, that is, to reduce the rate of change of the current consumption, while suppressing an increase in unnecessary current consumption. Note that suppressing an increase in unnecessary current consumption and reducing the rate of change of the current consumption are in a trade-off relationship. That is, the more the number of MAC circuits 25 that perform dummy operations is increased, the smaller the rate of change of the current consumption can be made. On the other hand, unnecessary current consumption increases.

[0061] <Principal effects of Embodiment 2> As described above, in the method of Embodiment 2, by providing the dummy circuit 22 and causing at least some of the n MAC circuits 25[1] to 25[n] to execute dummy operations, it becomes possible to suppress a rapid decrease in the current consumption. As a result, similar to the case of Embodiment 1, it is possible to suppress fluctuations in the power supply voltage, facilitate the power supply design of the semiconductor device 10, and suppress an increase in the design cost and the manufacturing cost. In addition, by causing only some of the MAC circuits 25, rather than all of them, to execute dummy operations, it becomes possible to suppress an increase in unnecessary current consumption.

[0062] (Embodiment 3) <Operation of Neural Network Engine (Embodiment 3)> FIG. 9 is a timing chart showing an operation example of the neural network engine shown in FIG. 5 in the semiconductor device according to Embodiment 3. The operation example shown in FIG. 9 is an operation that combines the operation shown in FIG. 3 and the operation shown in FIG. 7. That is, in FIG. 9, similar to the case of FIG. 3, the start timing and the end timing of a series of active periods consisting of periods T1 to T3 are controlled to be different from each other in four groups GR[1] to GR[4]. In addition, in FIG. 9, in period T4, dummy operations are being executed, similar to the case of FIG. 7.

[0063] FIG. 10 is a timing chart showing an operation example different from that of FIG. 9. The operation example shown in FIG. 10 is obtained by applying the same method as in the case of FIG. 8 to the operation example of FIG. 9. That is, in FIG. 10, in period T4, not all but some of the MAC circuits, in this example, the MAC circuits 25[1] and 25[3] belonging to groups GR[1] and GR[3], are executing dummy operations.

[0064] When using an operation example such as that in FIG. 9, for example, compared with the cases of FIG. 3 and FIG. 7, the rate of change of the consumption current associated with the switching between the active periods (periods T1 to T3) and the idle period (period T4) in each of the groups GR[1] to GR[4] can be made smaller. Also, when using an operation example such as that in FIG. 10, while reducing the rate of change of the consumption current in the same manner as in the case of FIG. 9, it becomes possible to suppress an increase in unnecessary consumption current in the same manner as in the case of FIG. 8.

[0065] (Embodiment 4) <Setting of Groups and Dummy Circuits> FIG. 11 is a flowchart showing an example of a method for determining the setting contents of groups and the setting contents of dummy circuits in the semiconductor device according to Embodiment 4. For example, when using the operation example shown in FIG. 10, the degrees of the effect A of reducing the rate of change of the consumption current and the effect B of suppressing an increase in unnecessary consumption current vary depending on the setting contents of the groups, that is, the number of groups, and the setting contents of the dummy circuit 22, that is, the number and combination of the groups for which dummy operations are to be executed.

[0066] As described in FIG. 8, Effect A and Effect B have a trade-off relationship. Therefore, it is desirable to determine the optimal setting content by some method. As a method for determining the optimal setting content, for example, a method using simulation can be considered. However, the optimal setting content may vary depending on the configuration of the neural network to be processed, that is, how to operate the neural network engine (NNE). Also, an error may occur between the simulation result and the actual measurement. Therefore, here, the optimal setting content is determined using the flow as shown in FIG. 11.

[0067] The flow shown in FIG. 11 is executed by the processor 17 based on, for example, a calibration program stored in the memory MEM1 in FIG. 4. In FIG. 11, the processor 17 starts the operation of the neural network engine (NNE) 15b and the measurement of the consumption current (step S101).

[0068] Specifically, the processor 17 causes the neural network engine (NNE) 15b to perform the processing of a certain target layer in the neural network, for example. More specifically, the processor 17 sequentially causes the sequence controller 21b of the neural network engine 15b to read a series of commands CMD, etc. representing the operation sequence of the target layer stored in the memory MEM1. Also, the processor 17 measures the consumption current using, for example, a current sensor installed on the power supply wiring of the semiconductor device 10.

[0069] Subsequently, the processor 17 ends the operation of the neural network engine 15b and the measurement of the consumption current (step S102). Here, the processing of the target layer executed during the operation period of the neural network engine 15b, that is, the period from step S101 to S102, may be processing for a very small coordinate area within the target layer. Specifically, during the operation period, for example, processing of about several cycles may be performed with the periods T1 to T4 shown in FIG. 10 as one cycle.

[0070] After step S102, the processor 17 calculates the maximum change rate of the current consumption (Max(di / dt)) and the average current (Iave) based on the current consumption measured during the operation period of the neural network engine 15b (steps S103, S104). Next, the processor 17 determines whether or not all the setting contents of the dummy circuit 22, that is, all the numbers and combinations of groups for which dummy operations are to be executed, have been covered (step S105).

[0071] If all the setting contents of the dummy circuit 22 have not been covered (step S105: No), the processor 17 changes the setting contents of the dummy circuit 22 and returns to step S101 (step S108). On the other hand, if all the setting contents of the dummy circuit 22 have been covered (step S105: Yes), the processor 17 determines whether or not all the setting contents of the groups, that is, all the numbers of the groups that can be set, have been covered (step S106). If all the setting contents of the groups have not been covered (step S106: No), the processor 17 changes the setting contents of the groups and returns to step S101 (step S109).

[0072] In steps S108 and S109, the processor 17 changes the setting contents of the dummy circuit 22 and the setting contents of the groups by outputting, for example, a command CMD representing each of the changed setting contents to the sequence controller 21b of the neural network engine 15b. The setting contents of the groups, that is, the numbers of the groups that can be set, are determined in advance with a plurality of options, and any one of the options is selected based on the command CMD. Also, the options for the setting contents of the dummy circuit 22 are determined according to the setting contents of the groups, that is, the number of the selected groups.

[0073] When all the setting contents of the group are covered (step S106: Yes), the processor 17 determines the optimal setting content based on the maximum change rate (Max(di / dt)) and the average current (Iave) of the consumed current calculated in steps S103 and S104 for each different setting content (step S107). Here, the optimal setting content is the setting content in which both the maximum change rate of the consumed current and the average current, which are in a trade-off relationship, become small. Therefore, for example, the processor 17 may set the setting content in which the value obtained by weighting and adding the maximum change rate and the average current is the minimum value as the optimal setting content.

[0074] The optimal setting content is determined, for example, for each layer of the neural network. For example, in the calibration process before actually starting the processing of the neural network, the optimal setting content for each layer is determined using the flow as shown in FIG. 11. In the actual processing of the neural network executed thereafter, the optimal setting content determined in the calibration process is applied. Specifically, for example, the processor 17 may store the optimal setting content determined for each layer in the memory MEM1 or the like as a command CMD associated with each layer, and cause the sequence controller 21b to read it at the start of the actual processing of each layer.

[0075] <Main effects of Embodiment 4> As described above, by using the method of Embodiment 4, in addition to the various effects described in Embodiments 1 to 3, it becomes possible to optimize the setting contents of the group and the dummy circuit 22. That is, it becomes possible to balance and suppress the sudden decrease in the consumed current and the increase in unnecessary power consumption.

[0076] As described above, the invention made by the present inventor has been specifically described based on the embodiments. However, it goes without saying that the present invention is not limited to the above embodiments, and various modifications can be made without departing from the gist thereof.

Explanation of reference numerals

[0077] 10 Semiconductor device 15a, 15b Neural Network Engine (NNE) 16 System bus 17 Processor 20 MAC unit 21a, 21b Sequence controller 22 Dummy circuit 25 MAC circuit CH Channel DMAC1, DMAC2 DMA controller DTd Dummy data DTi Input data DTo Output data (normal output data) DToD Output data (dummy output data) GR Group MEM1, MEM2 Memory PR Parameter RDS Read start signal WTE Write end signal

Claims

1. A semiconductor device that executes neural network processing, comprising: n (where n is an integer of 2 or more) multiplier-accumulators that perform a multiply-accumulate operation on input data and parameters; one or more memories that store the input data and the parameters; a first DMA (Direct Memory Access) controller that transfers the parameters stored in the memory to the n multiplier-accumulators; a second input-side DMA controller that transfers the input data stored in the memory to the n multiplier-accumulators respectively using n channels, causing the n multiplier-accumulators to execute an operation and outputting normal output data as an operation result; a dummy circuit that outputs predetermined dummy data to at least a part of the n multiplier-accumulators, causing at least a part of the n multiplier-accumulators to execute a dummy operation and outputting dummy output data as an operation result; a second output-side DMA controller that transfers the normal output data from the n multiplier-accumulators to the memory respectively using n channels and does not transfer the dummy output data from at least a part of the n multiplier-accumulators to the memory; and at least a part of the n multiplier-accumulators executes the dummy operation within a period from when the second output-side DMA controller finishes data transfer to the memory until the second input-side DMA controller starts data transfer from the memory. A semiconductor device.

2. The semiconductor device according to claim 1, further comprising: a sequence controller that outputs a read start signal for causing the second input-side DMA controller to start data transfer from the memory, wherein the second output-side DMA controller outputs a write end signal when it finishes data transfer to the memory, and the dummy circuit outputs the dummy data to at least a part of the n multiplier-accumulators in response to the write end signal from the second output-side DMA controller, and outputs the input data from the second input-side DMA controller to the n multiplier-accumulators in response to the read start signal from the sequence controller. A semiconductor device.

3. The semiconductor device according to claim 2, wherein The sequence controller further divides the n channels in the second input-side DMA controller and the second output-side DMA controller and the n multiplier-accumulators into m (m is an integer smaller than n) groups, and outputs the read start signals for each of the m groups to the second input-side DMA controller at mutually different timings, thereby controlling the timings of a series of operations including the read operation by the second input-side DMA controller, the operation by the multiplier-accumulator, and the write operation by the second output-side DMA controller to be different from each other in the m groups. The second output-side DMA controller outputs the write end signal for each of the m groups. Based on the write end signals for each of the m groups and the read start signals for each of the m groups, the dummy circuit causes the multiplier-accumulators for each of the m groups to execute the dummy operation. Semiconductor device.

4. In the semiconductor device according to claim 3, the dummy circuit executes the dummy operation for some of the m groups. Semiconductor device.

5. A semiconductor device that executes neural network processing, n (n is an integer of 2 or more) multiplier-accumulators that perform a multiply-accumulate operation on input data and parameters, one or more memories that store the input data and the parameters, a first DMA (Direct Memory Access) controller that transfers the parameters stored in the memory to the n multiplier-accumulators, a second input-side DMA controller that transfers the input data stored in the memory to the n multiplier-accumulators respectively using n channels, causes the n multiplier-accumulators to execute an operation, and outputs output data as an operation result, a second output-side DMA controller that transfers the output data from the n multiplier-accumulators to the memory respectively using n channels, a sequence controller that outputs a read start signal for starting data transfer from the memory to the second input-side DMA controller, and includes The sequence controller divides the n channels in the second input-side DMA controller and the second output-side DMA controller and the n multiplier-accumulators into a plurality of groups, and outputs the read start signals for each of the plurality of groups to the second input-side DMA controller at different timings, thereby controlling the timings of a series of operations including the read operation by the second input-side DMA controller, the operation by the multiplier-accumulator, and the write operation by the second output-side DMA controller to be different from each other in the plurality of groups. Semiconductor device.

6. In the semiconductor device according to claim 5, the number of the plurality of groups can be changed by setting. Semiconductor device.

7. A semiconductor device composed of one semiconductor chip, a neural network engine that executes neural network processing, one or more memories that store input data and parameters, a processor, a bus that connects the neural network engine, the memory, and the processor to each other, and includes the neural network engine has n (n is an integer of 2 or more) multiplier-accumulators that perform a multiply-accumulate operation on the input data and the parameters, a first DMA (Direct Memory Access) controller that transfers the parameters stored in the memory to the n multiplier-accumulators, a second input-side DMA controller that transfers the input data stored in the memory to the n multiplier-accumulators respectively using n channels, causes the n multiplier-accumulators to execute an operation, and outputs normal output data as an operation result, a dummy circuit that outputs predetermined dummy data to at least a part of the n multiplier-accumulators, causes at least a part of the n multiplier-accumulators to execute a dummy operation, and outputs dummy output data as an operation result, a second output-side DMA controller that transfers the normal output data from the n multiplier-accumulators to the memory respectively using n channels and does not transfer the dummy output data from at least a part of the n multiplier-accumulators to the memory, and has At least a part of the n multiplier-accumulator units executes the dummy operation within a period from when the second output-side DMA controller finishes data transfer to the memory until the second input-side DMA controller starts data transfer from the memory. Semiconductor device.

8. In the semiconductor device according to Claim 7, the neural network engine further includes a sequence controller that outputs a read start signal for causing the second input-side DMA controller to start data transfer from the memory, the second output-side DMA controller outputs a write end signal when it finishes data transfer to the memory, the dummy circuit inputs the dummy data to at least a part of the n multiplier-accumulator units in response to the write end signal from the second output-side DMA controller, and transfers the input data from the second input-side DMA controller to the n multiplier-accumulator units in response to the read start signal from the sequence controller. Semiconductor device.

9. In the semiconductor device according to Claim 8, the sequence controller further divides the n channels in the second input-side DMA controller and the second output-side DMA controller and the n multiplier-accumulator units into m (m is an integer smaller than n) groups, and by outputting the read start signals for each of the m groups to the second input-side DMA controller at mutually different timings, controls the timings of a series of operations including a read operation by the second input-side DMA controller, an operation by the multiplier-accumulator units, and a write operation by the second output-side DMA controller to be mutually different for the m groups, the second output-side DMA controller outputs the write end signal for each of the m groups, the dummy circuit causes the multiplier-accumulator units for each of the m groups to execute the dummy operation based on the write end signal for each of the m groups and the read start signal for each of the m groups. Semiconductor device.

10. In the semiconductor device according to Claim 9, the dummy circuit causes the dummy operation to be executed for some of the m groups. Semiconductor device.

Citation Information

Patent Citations

  • Vector processor, information processor, and vector processing method

    JP2006039840A

  • Neural Network Instruction Set Architecture

    JP2019533868A

  • Arithmetic processing device, program, and method for controlling arithmetic processing device

    JP2020119213A

  • Semiconductor device

    JP2021064193A

  • Image processing device

    WO2020017026A1