Semiconductor device and method for controlling semiconductor device

By designing semiconductor devices including data processing units, parallel arithmetic units, holding circuits and data transmission units, the problem of insufficient processing performance in the prior art is solved, and efficient arithmetic processing capabilities are achieved, especially suitable for large-scale arithmetic processing such as deep learning processing.

CN110609804BActive Publication Date: 2025-06-24RENESAS ELECTRONICS CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201910453550.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-06-15
Filing Date
2019-05-28
Publication Date
2025-06-24
Estimated Expiration
2039-05-28

AI Technical Summary

Technical Problem

The processing performance of existing dynamic reconfiguration processors is not sufficient to perform large-scale arithmetic processing, such as deep learning processing.

Method used

A semiconductor device is designed, including a data processing unit, a parallel arithmetic unit, a holding circuit and a data transmission unit. The data processing unit sequentially performs data processing, the parallel arithmetic unit performs arithmetic processing between the data output by the data processing unit and a plurality of predetermined data in parallel, the holding circuit saves the arithmetic processing results, and the data transmission unit sequentially selects the arithmetic processing results held by the accelerator and outputs them.

Benefits of technology

It realizes efficient arithmetic processing capabilities and can provide efficient performance in large-scale arithmetic processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110609804B_ABST
    Figure CN110609804B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to semiconductor devices and methods for controlling semiconductor devices. A semiconductor device includes: a dynamically reconfigurable processor that performs data processing on sequentially input input data and sequentially outputs the results of the data processing as output data; an accelerator including a parallel arithmetic section that performs arithmetic operations in parallel between the output data from the dynamically reconfigurable processor and each of a plurality of predetermined data; and a data transfer unit that sequentially selects a plurality of arithmetic operation results of the accelerator and outputs them to the dynamically reconfigurable processor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] The disclosure (including the specification, drawings, and abstract) of Japanese Patent Application No. 2018-114861, filed on June 15, 2018, is incorporated herein by reference in its entirety. Technical Field

[0003] The present invention relates to a semiconductor device and a control method thereof, and more particularly to a semiconductor device and a control method thereof suitable for implementing high-speed arithmetic processing, for example. Background Art

[0004] In addition to a central processing unit (CPU), there are dynamic reconfiguration processors that execute high processing performance. The dynamic reconfiguration processor is referred to as a dynamically reconfigurable processor (DRP) or an array-type processor. The dynamic reconfiguration processor is a processor capable of dynamically reconfiguring a circuit by dynamically switching the operation content of each of a plurality of processor elements and the connection relationship between the plurality of processors in accordance with operation instructions sequentially given. Techniques related to the dynamic reconfiguration processor are disclosed, for example, in Japanese Patent No. 3674515 (Patent Document 1) as an array processor.

[0005] In addition, "SIMD" [Online] (searched on January 26, 2018), Internet <URL:https: / / ja.wikipedia.org / wiki / SIMD> (Non-Patent Document 1) and "Mechanisms for 30times fastermechanical learning with Google Tensor Processing Unit" [Online] (searched on January 26, 2030), Internet <URL:https: / / cloudplatform-jp.googleblog.com / 2017 / 05 / an-in-depth-look-at-googles-first-tensor-processing-unit-tpu.html> (Non-Patent Document 2) disclose techniques related to parallel arithmetic processing. Summary of the Invention

[0006] However, the processing performance of the dynamic reconfiguration processor disclosed in Patent Document 1 is insufficient to execute large-scale arithmetic processing such as deep learning processing. Other objects and novel features will become apparent from the description of the present specification and the drawings.

[0007] According to one embodiment, a semiconductor device includes: a data processing unit that performs data processing on first input data input sequentially and outputs the result of the data processing sequentially as first output data; a parallel arithmetic unit that performs arithmetic processing in parallel between the first output data output sequentially from the data processing unit and each of a plurality of predetermined data; a holding circuit that holds the result of the arithmetic processing; and a first data transmission unit that sequentially selects a plurality of arithmetic processing results held in order by an accelerator and outputs the arithmetic processing results sequentially as first input data.

[0008] According to another embodiment, a control method of a semiconductor device uses a data processing unit to perform arithmetic processing on first input data input sequentially, outputs the result of the arithmetic processing sequentially as first output data, uses an accelerator to perform arithmetic processing in parallel between the first output data output sequentially from the data processing unit and each of a plurality of predetermined data, sequentially selects a plurality of arithmetic processing results output from the accelerator, and outputs them sequentially as first input data.

[0009] According to the above embodiments, a semiconductor device and a control method thereof can be provided, and the semiconductor device can achieve efficient arithmetic processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a block diagram showing a configuration example of a semiconductor system in which the semiconductor device according to the first embodiment is mounted.

[0011] Figure 2 is a diagram illustrating Figure 1 a specific configuration example of the semiconductor device.

[0012] Figure 3 is a diagram illustrating Figure 2 a configuration example of the parallel arithmetic unit.

[0013] Figure 4 is a schematic diagram showing an example of a neural network structure.

[0014] Figure 5 is a schematic diagram showing the flow of the arithmetic process of the neural network;

[0015] Figure 6 is a timing diagram showing the process flow of the semiconductor system according to the first embodiment.

[0016] Figure 7 is a schematic diagram of matrix arithmetic.

[0017] Figure 8 is a schematic diagram showing default information stored in a local memory.

[0018] Figure 9 It is a schematic diagram showing a multiplication equation of the data matrix In and the first row of the matrix data W.

[0019] Figure 10 It is a schematic diagram of a configuration example of an accelerator according to the first embodiment;

[0020] Figure 11 It is a timing diagram for explaining the relationship between the data output and the data input of a dynamically reconfigurable processor.

[0021] Figure 12 It is a timing diagram for explaining the relationship between the arithmetic processing of matrix data for each layer by the accelerator.

[0022] Figure 13 It is a flowchart illustrating the operation of a semiconductor system according to the first embodiment.

[0023] Figure 14 It is a comparative example of the configuration of the accelerator.

[0024] Figure 15 It is a schematic diagram showing a configuration example of a parallel arithmetic unit.

[0025] Figure 16 It is a schematic diagram showing a first modification of the parallel arithmetic unit.

[0026] Figure 17 It is a schematic diagram showing a second modification of the parallel arithmetic unit;

[0027] Figure 18 It is a schematic diagram showing a third modification of the parallel arithmetic unit.

[0028] Figure 19 It is a schematic diagram showing a fourth modification of the parallel arithmetic unit.

[0029] Figure 20 It is a schematic diagram showing a fifth modification of the parallel arithmetic unit.

[0030] Figure 21 It is a schematic diagram showing a sixth modification of the parallel arithmetic unit.

[0031] Figure 22 It is a schematic diagram of the data transfer unit and the parallel arithmetic unit in the accelerator when the input mode is the first input mode.

[0032] Figure 23 It is a schematic diagram of the data transfer unit and the parallel arithmetic unit in the accelerator when the input mode is the second input mode.

[0033] Figure 24It is a schematic diagram showing the data transfer unit and the parallel arithmetic unit in the accelerator when the input mode is the third input mode.

[0034] Figure 25 It is a schematic diagram showing the data transfer unit and the parallel arithmetic unit in the accelerator when the input mode is the fourth input mode.

[0035] Figure 26 It is a schematic diagram showing the data transfer unit and the parallel arithmetic unit in the accelerator when the input mode is the fifth input mode.

[0036] Figure 27 It is a schematic diagram showing the data transfer unit and the parallel arithmetic unit in the accelerator when the input mode is the sixth input mode.

[0037] Figure 28 It is a schematic diagram showing the data transfer unit and the parallel arithmetic unit in the accelerator when the input mode is the seventh input mode.

[0038] Figure 29 It is a schematic diagram showing the parallel arithmetic unit and the data transfer unit in the accelerator when the output mode is the first output mode;

[0039] Figure 30 It is a schematic diagram showing the parallel arithmetic unit and the data transfer unit in the accelerator when the output mode is the second output mode;

[0040] Figure 31 It is a schematic diagram showing the parallel arithmetic unit and the data transfer unit in the accelerator when the output mode is the third output mode;

[0041] Figure 32 It is a schematic diagram showing the parallel arithmetic unit and the data transfer unit in the accelerator when the output mode is the fourth output mode.

[0042] Figure 33 It is a schematic diagram showing the parallel arithmetic unit and the data transfer unit 14 in the accelerator 12 when the output mode is the fifth output mode.

[0043] Figure 34 It is a schematic diagram showing the parallel arithmetic unit and the data transfer unit in the accelerator when the output mode is the sixth output mode.

[0044] Figure 35 It is a schematic diagram showing the parallel arithmetic unit and the data transfer unit in the accelerator when the output mode is the seventh output mode.

[0045] Figure 36 It is a schematic diagram showing the flow of the operation process of the parallel arithmetic part when the operation is performed with the input data set to the maximum degree of parallelization.

[0046] Figure 37 FIG. is a schematic diagram of the flow of the operation process of the parallel arithmetic section when the operation process is performed by minimizing the parallelization of the input data.

[0047] Figure 38 FIG. is a schematic diagram of the flow of the operation process of the parallel arithmetic section 121 when the operation process is performed by setting the input data in parallel to a medium extent;

[0048] Figure 39 FIG. is a schematic diagram of the flow of the operation process of the parallel arithmetic section when parallel arithmetic operations are performed for each of two input data.

[0049] Figure 40 FIG. is a block diagram showing an example of the configuration of a semiconductor system in which a semiconductor device according to the second embodiment is mounted. DETAILED DESCRIPTION

[0050] For clarity of explanation, the following description and drawings are appropriately omitted and simplified. Each element described as a functional block in the drawings can be configured by hardware through a CPU (Central Processing Unit), a memory, and other circuits, and can be implemented by software through a program loaded in the memory. Therefore, those skilled in the art should understand that these functional blocks can be implemented in various ways by only hardware, only software, or a combination thereof, and the present invention is not limited to any of them. In the drawings, the same elements are denoted by the same reference numerals, and repeated descriptions thereof are omitted as needed.

[0051] The programs described above can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., floppy disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROM (Compact Disc Read-Only Memory), CD-R, CD-R / W, solid-state memories (e.g., masked ROM, PROM (Programmable ROM), EPROM (Erasable PROM, flash ROM, RAM (Random Access Memory)). Programs can also be provided to a computer through various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can provide programs to a computer via wired or wireless communication paths such as wires and optical fibers.

[0052] First Embodiment

[0053] Figure 1FIG. is a block diagram showing a configuration example of a semiconductor system SYS1 on which a semiconductor device 1 according to a first embodiment of the present invention is mounted. The semiconductor device 1 according to the present embodiment includes: an accelerator having a parallel arithmetic section that executes parallel arithmetic operations; a data processing unit, such as a dynamic reconfiguration processor, that sequentially executes data exchange; and a data transfer unit that sequentially selects a plurality of arithmetic processing results of the accelerator and sequentially outputs them to the data processing unit. Therefore, the semiconductor device 1 and the semiconductor system SYS1 including the semiconductor device 1 according to the present embodiment can use the accelerator to execute a large amount of conventional data processing and use the data processing unit to execute other data processing, thereby achieving efficient arithmetic processing. A brief description will be given below.

[0054] As Figure 1 shown, the semiconductor system SYS1 includes a semiconductor device 1, a CPU 2, and an external memory 3. The semiconductor device 1 includes a dynamic reconfiguration processor (hereinafter referred to as DRP) 11, an accelerator 12, a data transfer unit 13, a data transfer unit 14, and a direct memory access (DMA) 15.

[0055] For example, the DRP 11 performs arithmetic processing on data sequentially input from the external memory 3 and sequentially outputs the results of the arithmetic processing as data DQout. In this way, the DRP 11 can send and receive data in each cycle. Here, the DRP 11 is a data processor capable of dynamically reconfiguring a circuit by dynamically switching the operation content of each of a plurality of processor elements and the connection between the plurality of processors according to operation instructions read from a configuration data memory provided in the DRP 11.

[0056] For example, the DRP 11 includes a plurality of processor elements provided in an array, a plurality of switching elements provided corresponding to the plurality of processor elements, and a state management unit. The state management unit issues an instruction pointer determined in advance by a program to each of the processor elements. Each of the processor elements includes, for example, at least an instruction memory and an arithmetic unit. The arithmetic unit executes arithmetic processing according to an operation instruction specified by the instruction pointer among a plurality of operation instructions stored in the instruction memory. The arithmetic unit can be, for example, a 16-bit arithmetic unit that performs arithmetic processing on 16-bit wide data, or an arithmetic unit that performs arithmetic processing on other bit-width numbers. Alternatively, the arithmetic unit can be configured by a plurality of arithmetic units. Each of the switching elements sets the connection relationship between the corresponding processor element and another processor element according to an operation instruction read from the instruction memory of the corresponding processor element. Thus, the DRP 11 can dynamically switch the circuit according to sequentially applied operation instructions.

[0057] In this embodiment, the DRP 11 is provided in the semiconductor device 1, but it is not limited thereto. For example, as long as the CPU performs arithmetic processing on the sequentially input data, a central processing unit (CPU) may be provided instead of the DRP 11.

[0058] The data transfer unit 13 distributes or serializes the data DQout according to, for example, the parallelization program of the arithmetic processing required by the parallel arithmetic section 121, and outputs the data as the data DPin.

[0059] The accelerator 12 performs an arithmetic operation in parallel between the data DPin sequentially output from the data transfer unit 13 and n (n is an integer equal to or greater than 2) pieces of predetermined data D_0 to D_(n - 1). In the following description, the predetermined data D_0 to D_(n - 1) are not distinguished and may be simply referred to as the predetermined data D.

[0060] Specifically, the accelerator 12 includes a parallel arithmetic section 121 and a local memory 122. The local memory 122 stores, for example, multiple pieces of predetermined data D_0 to D_(n - 1) read from the external memory 3 and initial setting information such as a bias value b.

[0061] For example, when k×m elements constituting matrix data having k rows and m columns are continuously input to the accelerator 12 as the data DPin, k rows each having m data are sequentially input to the accelerator 12, that is, k×m data. However, regardless of the value of k, the accelerator 12 uses the predetermined data D_0 to D_(n - 1) for each of the m data (which are the input data for one row) for arithmetic processing. Therefore, n pieces of predetermined data D_0 to D_(n - 1), that is, m×n pieces of data corresponding to the m data corresponding to one row of input data are stored in the local memory 122. The parallel arithmetic section 121 is configured by a plurality of arithmetic units that perform arithmetic processing in parallel. The parallel arithmetic section 121 performs an arithmetic operation in parallel between the data DPin and each of the multiple predetermined data D_0 to D_(n - 1), and outputs n arithmetic processing results as the data DPout.

[0062] The data transfer unit 14 sequentially selects n pieces of data DPout output in parallel from the accelerator 12, and sequentially outputs the selected data pieces as the data DQin.

[0063] The DRP 11 performs arithmetic processing on the data DQin sequentially output from the data transfer unit 14, and sequentially outputs the result of the arithmetic processing to, for example, the external memory 3.

[0064] For example, the CPU 2 controls the operation of the semiconductor device 1 according to control instructions read from the external memory 3. More specifically, the CPU 2 prepares a data string (descriptor) for instructing in detail the operations of the accelerator 12 and the data transfer units 13 and 14, and stores the data string (descriptor) in the external memory 3.

[0065] The DMA 15 reads the descriptor from the external memory 3, interprets the content, and issues operation instructions to the accelerator 12 and the data transfer units 13 and 14. For example, the DMA 15 transfers the initial setting information stored in the external memory 3 to the local memory 122 according to the instructions described in the descriptor. The DMA 15 instructs the data transfer unit 13 to distribute or serialize the data DPin according to the parallelization program of the arithmetic processing of the parallel arithmetic section 121. The DMA 15 instructs the data transfer unit 14 to combine or serialize the n pieces of data DPout output in parallel according to the parallelization program of the arithmetic processing of the parallel arithmetic section 121.

[0066] When the operation specified by one descriptor is completed, the DMA 15 reads the next descriptor from the external memory 3 and issues operation instructions to the accelerator 12 and the data transfer units 13 and 14. Preferably, the descriptor is read before the completion of the operation by the immediately preceding descriptor read. Thus, the processing delay can be hidden.

[0067] The descriptor can be applied from a program running on the DRP 11 instead of the CPU 2, or can be generated in advance.

[0068] Figure 2 is a block diagram showing a specific configuration example of the semiconductor device 1. In Figure 2 , the DRP 11 outputs data DQout with a width of 64 bits in 4 channels as data DQout_0 to DQout_3. The DRP 11 is not limited to outputting the data DQout_0 to DQout_3 in four channels, and can be appropriately changed to a configuration for outputting data with any number of channels and any number of bit widths.

[0069] In Figure 2 , the data transfer unit 13 transfers the 64-bit wide data DQout_0 to DQout_3 sequentially output from the DRP 11 as data DPin_0 to DPin_3. In Figure 2 , each of the data DPin_0 to DPin_3 forms 64-bit wide data by bundling four 16-bit wide operation results represented by the floating-point method, but the present invention is not limited thereto. For example, 16-bit wide data, 32-bit wide data, and 48-bit wide data can be configured by bundling 1 to 3 16-bit wide operation results.

[0070] The parallel arithmetic section 121 includes, for example, parallel arithmetic units MAC256_0 to MAC256_3. Each of the parallel arithmetic units MAC256_0 to MAC256_3 includes 256 arithmetic units that perform arithmetic processing in parallel. Data DPin_0 to DPin_3 are respectively input to the parallel arithmetic units MAC256_0 to MAC256_3.

[0071] The parallel arithmetic unit MAC256_0 performs arithmetic processing in parallel by using a maximum of 256 arithmetic units (four groups of 64 units) with respect to 64-bit wide (16-bit wide × 4 groups) data DPin_0 and outputs a maximum of 256 arithmetic processing results.

[0072] Similarly, the parallel arithmetic unit MAC256_1 performs arithmetic processing in parallel by using a maximum of 256 arithmetic units (four groups of 64 units) with respect to 64-bit wide (16-bit wide × 4 groups) data DPin_1 and outputs a maximum of 256 arithmetic processing results. The parallel arithmetic unit MAC256_2 performs arithmetic processing in parallel by using a maximum of 256 arithmetic units (four groups of 64 units) with respect to 64-bit wide (16-bit wide × 4 groups) data DPin_2 and outputs a maximum of 256 arithmetic processing results. The parallel arithmetic unit MAC256_3 performs arithmetic processing in parallel by using a maximum of 256 arithmetic units (four groups of 64 units) with respect to data DPin_3 having a 64-bit width (16-bit wide × 4 groups) and outputs a maximum of 256 arithmetic processing results.

[0073] Figure 3 is a block diagram showing a configuration example of the parallel arithmetic unit MAC256_0. Figure 3 Also shown are data transfer units 13 and 14 provided before and after the parallel arithmetic unit MAC256_0.

[0074] As Figure 3 shown, the parallel arithmetic unit MAC256_0 includes parallel arithmetic units MAC64_0 to MAC64_3. Each of the parallel arithmetic units MAC64_0 to MAC64_3 consists of 64 arithmetic units that perform arithmetic processing in parallel.

[0075] Bits 0 to 15 of the 64-bit wide data DPin_0 (hereinafter referred to as data DPin_00) are input to the parallel arithmetic unit MAC64_0. Bits 16 to 31 of the 64-bit wide data DPin_0 (hereinafter referred to as data DPin_01) are input to the parallel arithmetic unit MAC64_1. Bits 32 to 47 of the 64-bit wide data DPin_0 (hereinafter referred to as data DPin_02) are input to the parallel arithmetic unit MAC64_2. Bits 48 to 63 of the 64-bit wide data DPin_0 (hereinafter referred to as data DPin_03) are input to the parallel arithmetic unit MAC64_3.

[0076] The parallel arithmetic unit MAC64_0 uses up to 64 arithmetic units to perform arithmetic processing on the 16-bit wide data DPin_00 in parallel and outputs the arithmetic processing results of up to 64 arithmetic processing results, each arithmetic processing result having a width of 16 bits. The parallel arithmetic unit MAC64_1 uses up to 64 arithmetic units to perform arithmetic processing on the 16-bit wide data DPin_01 in parallel and outputs the arithmetic processing results of up to 64 arithmetic processing results, each arithmetic processing result having a width of 16 bits. The parallel arithmetic unit MAC64_2 can use up to 64 arithmetic units to perform arithmetic processing on the 16-bit wide data DPin_02 in parallel and outputs the arithmetic processing results of up to 64 arithmetic processing results, each arithmetic processing result having a width of 16 bits. The parallel arithmetic unit MAC64_3 can use up to 64 arithmetic units to perform arithmetic processing on the 16-bit wide data DPin_03 in parallel and outputs the arithmetic processing results of up to 64 arithmetic processing results, each arithmetic processing result having a width of 16 bits.

[0077] The parallel arithmetic units MAC256_1 to MAC256_3 have the same configuration as that of the parallel arithmetic unit MAC256_0, and thus their descriptions are omitted.

[0078] Return to Figure 2 , and the description will continue. The parallel arithmetic unit MAC256_0 performs arithmetic processing on the data DPin_0 having a width of 64 bits (16-bit wide × 4 groups), and outputs four groups of up to 64 arithmetic processing results as data DOut_0, each arithmetic processing result having a width of 16 bits.

[0079] Similarly, the parallel arithmetic unit MAC256_1 performs arithmetic processing on the data DPin_1 and outputs four sets of up to 64 arithmetic processing results as the data DPout_1, with each arithmetic processing result having a width of 16 bits. The parallel arithmetic unit MAC256_2 performs arithmetic processing on the data DPin_2 and outputs four sets of up to 64 arithmetic processing results as the data DPout_2, with each arithmetic processing result having a width of 16 bits. The parallel arithmetic unit MAC256_3 performs arithmetic processing on the data DPin_3 and outputs four sets of up to 64 arithmetic processing results as the data DPout_3, with each arithmetic processing result having a width of 16 bits.

[0080] The data transfer unit 14 sequentially selects, for example, each of the data from each of the four sets (each set having 64 16-bit wide data) included in the data DPout_0 output in parallel from the parallel processor MAC256_0 and sequentially outputs the data DQin_0 including the four sets, with each set having 16-bit wide data (i.e., the 64-bit wide data DQin_0). As described above, the data transfer unit 14 may sequentially select 16-bit wide data from each set and output it, or may sequentially output all the data for each set so as to output 64 16-bit wide data in one set and then output 64 16-bit wide data in the next set, but the present invention is not limited thereto. The data output method of the data transfer unit 14 may be switched depending on the mode.

[0081] Similarly, the data transfer unit 14 sequentially selects, for example, each of the data from each of the four sets (each set having 64 16-bit wide data) included in the data DPout_1 output in parallel from the parallel arithmetic unit MAC256_1 and sequentially outputs the data DQin_1 including four sets of 16-bit wide data (i.e., the 64-bit wide data DQin_1). In addition, the data transfer unit 14 sequentially selects each of the data from each of the four sets (each set having 64 16-bit wide data) included in the data DPout_2 output in parallel from the parallel processor MAC256_2 and sequentially outputs the data DQin_2 including four sets of 16-bit wide data (i.e., the 64-bit wide data DQin_2). The data transfer unit 14 sequentially selects, for example, each of the data from each of the four sets (each set having 64 16-bit wide data) in the DPout_3 output in parallel from the parallel processor MAC256_3 and outputs the data DQin_3 including four sets of 16-bit wide data (i.e., the 64-bit wide data DQin_3).

[0082] These 64-bit wide data of DQin_0 to DQin_3 are input to the DRP 11. The DRP 11 performs arithmetic processing on the data DQin_0 to DQin_3 and sequentially outputs the arithmetic processing results to the external memory 3. The data DQin_0 to DQin_3 can be used to calculate the data DQout_0 to DQout_3.

[0083] As described above, the semiconductor device 1 according to the present embodiment includes: an accelerator having a parallel arithmetic section that performs arithmetic processing in parallel; a data processing unit such as a DRP that sequentially transfers data; and a data transfer unit that sequentially selects a plurality of arithmetic processing results of the accelerator and outputs them to the data processing unit. Therefore, the semiconductor device according to the present embodiment and the semiconductor system including the same can use the accelerator to perform a large amount of conventional data processing and use the data processing unit to perform other data processing, enabling efficient arithmetic processing even in large-scale arithmetic processing such as deep learning processing.

[0084] Hereinafter, reference will be made to Figure 4 and Figure 5 describe a calculation method of a neural network using the semiconductor device 1 according to the present embodiment. Figure 4 is a schematic diagram showing an example of a neural network structure. Figure 5 is a schematic diagram schematically showing the flow of the operation processing of the neural network.

[0085] As Figure 4 shown, the operation of the neural network adopts the following process: performing a multiply-accumulate calculation operation of multiplying the input data by the weight w (w′), performing an operation such as activation on the result, and outputting the operation result.

[0086] As Figure 5As shown, DRP 11 reads out the data required for the arithmetic processing of accelerator 12 from external memory 3 (step S1), and rearranges the calculator and data as needed (step S2). Thereafter, the data read from external memory 3 is sequentially output from DRP 11 to accelerator 12 as the data input to accelerator 12. Accelerator 12 performs a parallel multiply-accumulate calculation operation (step S4) by multiplying the data sequentially output from DRP 11 in the order of the received data by the data stored in the local memory (corresponding to weights). Then, the arithmetic result of accelerator 12 is sequentially output to DRP 11 (step S5). DRP 11 performs operations such as addition and activation on the data received from accelerator 12 as needed (step S6). The processing result of DRP 11 is stored in external memory 3 (step S7). By implementing the processing of the neural network through such processing and repeating this processing, the arithmetic processing required for deep learning can be performed.

[0087] In this way, in the neural network, high-speed operation can be achieved by using accelerator 12 to perform the conventional parallel multiply-accumulate calculation operation among the required operations. In addition, as a data processor DRP 11 that can dynamically reconfigure the circuit, it performs arithmetic processing other than the conventional parallel multiply-accumulate calculation operation, making it possible to flexibly set processing such as activation in different layers (in Figure 5 the example, the first layer and the second layer). In addition, DRP 11 can reconfigure the circuit configuration so that the input data required for the multiply-accumulate calculation operation is divided according to the parallel operation size that can be simultaneously processed by accelerator 12 and read out from external memory 3 to be output to accelerator 12. Thereby, the degree of freedom of the operation format of parallel arithmetic section 121 can be provided.

[0088] Next, the operation of semiconductor system SYS1 will be described with reference to Figure 6 is a timing diagram showing the processing flow of semiconductor system SYS1. Figure 6 is a timing diagram showing the processing flow of semiconductor system SYS1.

[0089] Hereinafter, the case where the matrix operation is performed by accelerator 12 will be described as an example. Figure 7 is a schematic diagram schematically showing a matrix arithmetic expression. In Figure 7 a multiplication operation is performed on matrix data composed of elements of k rows × m columns and matrix data W composed of elements of m rows × n columns, and the result of adding the bias value b to each element of the multiplication result is output as matrix data Out composed of elements of k rows × n columns.

[0090] When the accelerator 12 performs a calculation operation on the matrix data In of the first layer, the initial setting information including the matrix data W and the bias value b corresponding to the matrix data In of the first layer is stored in the local memory 122 of the accelerator 12 ( Figure 6 at time t1 to t2 in Figure 8 ). More specifically, the DMA 15 transfers the initial setting information read from the external memory 3 to the local memory 122 according to the instruction of the descriptor generated by the CPU 2. Note that a DMA dedicated to the accelerator 12 (not shown) may be provided separately from the DMA 15, and the initial setting information read from the external memory 3 may be transferred to the local memory 122 using the DMA dedicated to the accelerator 12.

[0091] Thereafter, the first row data of the matrix data In (hereinafter also referred to as row data In1) is read from the external memory 3 ( Figure 6 at time t2 in Figure 6 ). The DRP 11 outputs the row data In read from the external memory 3 to the accelerator 12 after performing a predetermined process as needed (

[0092] at time t3 in Figure 6 ).

[0093] Figure 9 is a schematic diagram showing a specific example of the multiplication expression of the row data In1 (the first row data of the matrix data In) and the matrix data W. In Figure 9 , it is assumed that the row data In1 consists of 29 columns of elements b0 to b19. In the matrix data W, it is assumed that the first row data consists of 20 columns of elements a0,0 a0,1... a0,19, the second row data consists of 20 columns of elements a1,0 a1,1... a1,19, and the 20th row data (which is the last row) consists of 20 columns of elements a19,0 a19,1... a19,19.

[0094] Here, the accelerator 12 performs a multiplication operation in parallel on the elements of each column of the row data In1 (e.g., b0) and the 20 columns of elements of each row of the matrix data W (e.g., a0,0 a0,1... a0,19), and then adds the results of the 20 multiplications in each column to calculate the elements of each column of the matrix data Out.

[0095] Figure 10is a schematic diagram showing a specific configuration example of the accelerator 12. In Figure 10 the example, 20 arithmetic units 121_0 to 121_19 among the multiple arithmetic units provided in the parallel arithmetic section 121 are used. Each of the arithmetic units 121_0 to 121_19 includes a multiplier MX1, an adder AD1, a register RG1, and a register RG2.

[0096] In the arithmetic unit 121-j (j is any one of 0 to 19), the bias value b read from the local memory 122 is set as the initial value in the register RG1 (the bias value b is not shown in Figure 10 ).

[0097] Thereafter, the multiplier MX1 multiplies the element b0 of the first column data in the row data In1 (corresponding to 16-bit wide data DPin) by the element a0,j of the first row in the matrix data W read from the local memory 122 (corresponding to 16-bit wide predetermined data D_j). The adder AD1 adds the multiplication result (a0,j × b0) of the multiplier MX1 to the value (bias value b) stored in the register RG1 and transmits the addition result to the register RG1.

[0098] After that, the multiplier MX1 multiplies the element b1 of the second column of the subsequently input row data In1 by the element a1,j of the second row in the matrix data W read from the local memory 122. The adder AD1 adds the multiplication result (a1,j × b1) of the multiplier MX1 to the value (a0,j × b0) stored in the register RG1 and transmits the addition result to the register RG1.

[0099] Since the operations of multiplication, addition, and storage as described above are repeated 20 cycles, the register RG1 stores the element ((a0,j × b0) + (a1,j × b1) + · + · + (a19,j × b19)) of the first row in the matrix data Out. Thereafter, the value stored in the register RG1 is transmitted to the register RG2, and the value stored in the register RG2 is output as the element of the first row of the matrix data Out after Figure 6 the time t5 in

[0100] When the data transfer from the register RG1 to the register RG2 is completed ( Figure 6 the time t5 in Figure 6 ), the arithmetic operation on the data of the second row (also called row data In2), which is the next row in the matrix data In ( Figure 6during times t6 to t9 in) and simultaneously transfer the arithmetic operation result stored in register RG2 to data transfer unit 14 (corresponding to Figure 6 during times t7 to t10 in). Thus, the efficiency of parallel arithmetic operations can be improved.

[0101] Therefore, preferably, DRP 11 receives the arithmetic operation result of data In1 during the output period of row accelerator 12 for the second row data In2 (which is the period from the completion of the output of the first row data In1 in matrix data In to the start of the output of the third row data In3).

[0102] Data transfer unit 14 sequentially selects 20 arithmetic operation results (each with a width of 16 bits) output from arithmetic units 121_0 to 121_19 (corresponding to data DPout), and sequentially outputs them as 16-bit wide data DQin. In other words, data transfer unit 14 sequentially outputs the elements of the first row of twenty columns of matrix data Out as data DQin. The sequentially output data DQin is received by DRP 11 at Figure 6 times t7 to t10 in.

[0103] In DRP 11, for example, adder AD2 performs an addition process on the data DQin sequentially output from data transfer unit 14, arithmetic unit TN1 performs a predetermined arithmetic operation based on the hyperbolic tangent function, and multiplier MX2 performs a multiplication operation. The operation results are written to external memory 3, for example, at Figure 6 times t8 to t11 in.

[0104] When accelerator 12 finishes performing arithmetic operations on all row data from the first row to the k-th row of matrix data In of the first layer, the same arithmetic operations are then performed on the matrix data In of the second layer. Before performing arithmetic operations on the matrix data In of the second layer, the initial setting information corresponding to the matrix data In of the second layer (i.e., matrix data W and bias value b) is stored in local memory 122. Accelerator 12 repeats such parallel arithmetic operations.

[0105] Preferably, local memory 122 has a storage area for storing the initial setting information (i.e., matrix data W and bias value b) corresponding to at least two layers of matrix data In. Thus, during the matrix operation on the matrix data In of the first layer, the initial setting information for the operation on the matrix data In of the second layer can be transferred to the free area of local memory 122. Thus, after completing the arithmetic operation on the matrix data of the first layer, the matrix calculation for the matrix data of the second layer can be quickly performed without waiting for the transmission of the initial setting information, as Figure 12As shown. In this case, preferably, the local memory 122 is configured to be able to read and write data simultaneously.

[0106] On the other hand, even if the local memory 122 does not have sufficient storage space to store the initial setting information corresponding to the matrix data In of one layer or has sufficient space to store the initial setting information corresponding to the matrix data In of one layer, the initial setting information can be divided and stored. Hereinafter, reference will be made to Figure 13 for a brief description.

[0107] Figure 13 is a flowchart showing the operation of the semiconductor system SYS1. In Figure 13 example, it is assumed that the local memory 122 does not have a storage area sufficient to store the initial setting information corresponding to the matrix data In of the third layer.

[0108] As Figure 13 shown, in step S101, the initial setting information corresponding to the matrix data In of the first layer is stored in the local memory 122. Thereafter, in step S102, the parallel arithmetic section 121 performs an arithmetic operation on the matrix data In of the first layer. Thereafter, in step S103, the initial setting information corresponding to the matrix data In of the second layer is stored in the local memory 122. Thereafter, in step S104, the parallel arithmetic section 121 performs an arithmetic operation on the matrix data In of the second layer. Thereafter, in step S105, the initial setting information corresponding to a part of the matrix data In of the third layer is stored in the local memory 122. In step S106, the parallel arithmetic section 121 performs an arithmetic operation on the part of the matrix data In of the third layer. In step S107, the initial setting information corresponding to the remaining matrix data In of the third layer is stored in the local memory 122. In step S108, the parallel arithmetic section 121 performs an arithmetic operation on the remaining matrix data In of the third layer. Thereafter, in step S109, the result of the arithmetic operation performed in step S106 and the result of the arithmetic process performed in step S108 are added together. Thus, an arithmetic operation on the matrix data In of the third layer can be achieved.

[0109] As described above, the semiconductor device 1 according to the present embodiment includes: an accelerator having a parallel arithmetic section that performs arithmetic operations in parallel; a data processing unit such as a DRP that sequentially transfers data; and a data transfer unit that sequentially selects a plurality of arithmetic operation results of the accelerator and outputs them to the data processing unit. Therefore, the semiconductor device according to the present embodiment and the semiconductor system including the semiconductor device can use the accelerator to perform a large amount of conventional data processing, and use the data processing unit to perform other data processing, enabling efficient arithmetic processing even in large-scale arithmetic processing such as deep learning processing.

[0110] In the present embodiment, the case where each of the arithmetic units 121_0 to 121_19 includes the register RG2 in addition to the multiplier MX1, the adder AD1, and the register RG1 has been described as an example, but the present invention is not limited thereto. Each of the arithmetic units 121_0 to 121_19 may include the multiplier MX1, the adder AD1, and the register RG1, and may not include the register RG2. This further suppresses the circuit scale.

[0111] In the present embodiment, the case where the bias value b is stored in the local memory 122 has been described as an example, but the present invention is not limited thereto. For example, the bias value b may be stored in a register or the like provided separately from the local memory 122, or the bias value b may be a fixed value such as 0 and may not be stored in the local memory 122.

[0112] Figure 14 is a schematic diagram of a configuration example of the accelerator 52 according to the comparative example. As Figure 14 shown, in the accelerator 52, each of the arithmetic units 121_0 to 121_19 includes a multiplier MX1, an adder AD1, a register RG1, an adder AD2, an arithmetic unit TN1, and a multiplier MX2. That is, in the accelerator 52, the adder AD2, the arithmetic unit TN1, and the multiplier MX2 provided in the DPR 11 in the accelerator 12 are provided in the arithmetic units 121_0 to 121_19.

[0113] However, in the accelerator 52, after the arithmetic operation process performed by the multiplier MX1, the adder AD1, and the register RG1 is repeated 20 cycles in each arithmetic unit, the arithmetic operation process performed by the adder AD2, the arithmetic unit TN1, and the multiplier MX2 is executed only one cycle. That is, in the accelerator 52, since the less frequently used adder AD2, arithmetic unit TN1, and multiplier MX2 are provided in all of the plurality of arithmetic units, there is a problem of an increase in circuit scale.

[0114] On the other hand, in the accelerator 12, the arithmetic units 121_0 to 121_19 do not include the adder AD2, the arithmetic unit TN1, and the multiplier MX2 that are not frequently used, and these arithmetic units are configured and commonly used in the previous stage of the DRP 11. Thereby, an increase in the circuit scale can be suppressed.

[0115] Configuration example of parallel arithmetic unit

[0116] Next, a specific configuration example of a plurality of arithmetic units provided in the parallel arithmetic section 121 will be described. Figure 15 is a schematic diagram showing a specific configuration example of the parallel arithmetic unit MAC64_0. As Figure 15 shown, the parallel arithmetic unit MAC64_0 includes 64 arithmetic units 121_0 to 121_63 that perform arithmetic operation processing in parallel. Each of the arithmetic units 121_0 to 121_63 includes a multiplier MX1, an accelerator AD1, a register RG1, and a register RG2. Here, the paths of the multiplier MX1, the accelerator AD1, the register RG1, and the register RG2 in the arithmetic units 121_0 to 121_63 perform a predetermined arithmetic operation processing on 16-bit wide data and output 16-bit wide data.

[0117] Since the parallel arithmetic units MAC64_1 to MAC64_3 have the same configuration as that of the parallel arithmetic unit MAC64_0, their descriptions are omitted.

[0118] First modification of parallel arithmetic unit

[0119] Figure 16 is a schematic diagram showing a first modification of the parallel arithmetic unit MAC64_0 as the parallel operator MAC64a_0. As Figure 16 shown, the parallel arithmetic unit MAC64a_0 includes 64 parallel arithmetic units 121a_0 to 121a_63. Each of the arithmetic units 121a_0 to 121a_63 includes a selector SL1, a multiplier MX1, an adder AD1, a register RG1, and a register RG2.

[0120] The selector SL1 sequentially selects and outputs 16-bit wide data read from the local memory 122 bit by bit. The paths of the multiplier MX1, the adder AD1, the register RG1, and the register RG2 perform arithmetic operation processing using 1-bit wide data output from the selector SL1 and 16-bit wide data from the data transfer unit 13 and output 16-bit wide data.

[0121] In this way, even when the parallel arithmetic unit MAC64a_0 performs an arithmetic operation on data with a 1-bit width read from the local memory 122, it is possible to suppress an increase in the number of reads from the local memory 122 by reading data with a 16-bit width from the local memory 122 and then sequentially selecting one bit from the data with a 16-bit width and performing a parallel arithmetic operation process. Therefore, power consumption can be reduced.

[0122] The parallel arithmetic units MAC64a_1 to MAC64a_3 have the same configuration as that of the parallel arithmetic unit MAC64a_0, and thus the description thereof is omitted.

[0123] It should be noted that when performing an arithmetic operation process on 1-bit width data read from the local memory 122, multiple processing units multiply the data from the data transfer unit 13 by +1 or -1. Therefore, the multiply-accumulate calculation operation adds the data from the data transfer unit 13 to the data stored in the register RG1 or subtracts the data from the data transfer unit 13 from the data stored in the register RG1. This can also be achieved by the configuration of the parallel arithmetic unit as Figure 17 shown.

[0124] Second modification of the parallel arithmetic unit

[0125] Figure 17 is a schematic diagram showing a second modification of the parallel operator MAC64_0 as the parallel operator MAC64b_0. As Figure 17 shown, the parallel arithmetic unit MAC64b_0 includes 64 parallel arithmetic units 121b_0 to 121b_63. Each of the arithmetic units 121b_0 to 121b_63 includes a selector SL1, an adder AD1, a subtractor SB1, a selector SL2, a register RG1, and a register RG2.

[0126] Here, the selector SL1 sequentially selects and outputs 16-bit width data read from the local memory 122 bit by bit. The adder AD1 adds the 16-bit width data from the data transfer unit 13 to the data stored in the register RG1. The subtractor AD1 subtracts the data stored in the register RG1 from the 16-bit width data from the data transfer unit 13. The selector SL2 selects and outputs the addition result of the adder AD1 or the subtraction result of the subtractor SB1 based on the value of the 1-bit width data output from the selector SL1. The data output from the selector SL2 is stored in the register RG1. Thereafter, the data stored in the register RG1 is stored in the register RG2 and then output to the data transfer unit 14.

[0127] The parallel arithmetic unit MAC64b_0 can perform the same operations as the parallel arithmetic unit MAC64a_0.

[0128] The parallel arithmetic units MAC64b_1 to MAC64b_3 have the same configuration as that of the parallel arithmetic unit MAC64b_0, and thus the description thereof is omitted.

[0129] The third modification of the parallel arithmetic unit

[0130] Figure 18 (The third modification of a plurality of arithmetic units including the parallel arithmetic unit) shows the third modification of the parallel arithmetic unit MAC64_0 as the parallel operator MAC64c_0. As Figure 18 shown, the parallel arithmetic unit MAC64c_0 includes 64 parallel arithmetic units 121c_0 to 121c_63. Each of the arithmetic units 121c_0 to 121c_63 performs arithmetic operation processing in units of 1 bit between 16 pieces of 1-bit data from the data transfer unit 13 and 16 pieces of 1-bit data read from the local memory 122.

[0131] Each of the arithmetic units 121c_0 to 121c_63 includes 16 paths including a multiplier MX1, an adder AD1, a register RG1, and a register RG2. Here, each path performs arithmetic operation processing by using one of the 16 pieces of 1-bit data from the data transfer unit 13 and one of the 16 pieces of 1-bit data read from the local memory 122 and outputs 1-bit data. The 1-bit data is represented by binary values of 1 and 0 in hardware, and these values of 1 and 0 are used to calculate +1 and -1 respectively in meaning.

[0132] As described above, even when the calculation process is performed using the 1-bit data from the data transfer unit 131 and the 1-bit data read from the local memory 122, the parallel calculator MAC64c_0 can perform 16 arithmetic operation processes for 1-bit data by using 16-bit data paths to transfer and read data.

[0133] Figure 18 The operation of the configuration shown in Figure 19 can also be achieved by the configuration of the parallel arithmetic unit as shown.

[0134] The fourth modification of the parallel arithmetic unit

[0135] Figure 19 is a schematic diagram showing the fourth modification of the parallel operator MAC64_0 as the parallel operator MAC64d_0. As Figure 19As shown, the parallel arithmetic unit MAC64d_0 includes 64 parallel arithmetic units 121d_0 to 121d_63. The arithmetic units 121d_0 to 121d_63 include an XNOR circuit XNR1, a pop counter CNT1, an adder AD1, a register RG1, and a register RG2.

[0136] The XNOR circuit XNR1 performs a negative exclusive-OR operation on 16 pieces of 1-bit data from the data transfer unit 13 and 16 pieces of 1-bit data read from the local memory 122 in units of 1 bit. When the output value of the XNOR circuit XNR1 is observed in units of two, the pop counter CNT1 counts the number of "1" output values. Here, when the 16-bit data from the data transfer unit 13 and the 16-bit data read from the local memory 122 are regarded as binary numbers, when the output value of the pop counter CNT1 represents the number of bits with the same value, the output value of the pop counter CNT1 represents the number of bits with the same output value. The output data of the pop counter CNT1 is added to the data stored in the register RG1 by the adder AD1. However, since the values for +1 and -1 were originally calculated as 1 and 0, it is necessary to correct the output value. This problem can also be handled by preprocessing the bias value required for correction.

[0137] As described above, the parallel arithmetic unit MAC64d_0 performs arithmetic operation processing in units of 1 bit in parallel for 16 pieces between the 16 pieces of 1-bit data from the data transfer unit 13 and the 16 pieces of 1-bit data read from the local memory 122, adds up the arithmetic operation processing of these pieces, and outputs the result as 16-bit data. Thus, the parallel arithmetic unit MAC64d_0 can achieve the same operation as the operation of the parallel arithmetic unit MAC64d_0.

[0138] The parallel arithmetic units MAC64d_1 to MAC64d_3 have the same configuration as that of the parallel arithmetic unit MAC64d_0, and thus the description thereof is omitted.

[0139] Fifth modification of the parallel arithmetic unit

[0140] Figure 20 (Fifth modification of multiple operators including a parallel calculator) shows a fifth modification of the parallel calculator MAC64_0 as the parallel calculator MAC64e_0. The parallel calculator MAC64e_0 includes operators 121e_0 to 121e_63.

[0141] Compared with the arithmetic units 121d_0 to 121d_63, the arithmetic units 121e_0 to 121e_63 further include a 1-bit conversion circuit CNV1 for converting 16-bit wide data stored in the register RG1 into 1-bit wide data. The 1-bit conversion circuit CNV1 can output an active value as a 1-bit value by outputting 0 when the operation result is negative and otherwise outputting 1 by using a bias value, for example. In this case, 64 pieces of 1-bit data from the arithmetic units 121e_0 to 121e_63 are input to the data transfer unit 14. It should be noted that the data transfer unit 14 can also output 64 pieces of 1-bit data as 16-bit wide data by bundling them together. Therefore, the data transfer unit 14 can output 64 pieces of 1-bit data in four cycles.

[0142] Sixth modification of the parallel arithmetic unit

[0143] Figure 21 is a schematic diagram showing a sixth modification of the parallel operator MAC64_0 as the parallel arithmetic unit MAC64f_0. The parallel arithmetic unit MAC64f_0 includes 64 parallel arithmetic units 121e_0 to 121e_63.

[0144] The arithmetic unit 121e_0 includes the arithmetic units 121_0, 121a_0, 121c_0, and 121e_0 and a selector SL3. The selector SL3 selects one of the arithmetic units 121_0, 121a_0, 121c_0, and 121e_0 according to a mode and outputs the selected one. The arithmetic units MAC64e_1 to MAC64e_3 have the same configuration as that of the arithmetic unit MAC64e_0, and thus the description thereof is omitted. Note that a part of the arithmetic unit 121e_0 and a part of the arithmetic unit 121c_01 can have a common circuit, and it can be selected whether to output 16 bits as it is or via the 1-bit conversion circuit. The mode can be fixedly specified, for example, by setting a register by the CPU, or can be specified for each descriptor by describing information on the mode to be specified in the descriptor.

[0145] In this way, the parallel arithmetic unit MAC64f_0 can switch the content of the arithmetic operation according to the required arithmetic accuracy, memory usage, and throughput. The arithmetic units MAC64e_1 to MAC64e_3 have the same configuration as that of the parallel arithmetic unit MAC64e_0, and thus the description thereof is omitted.

[0146] Example of data transfer through the data transfer unit 13

[0147] Next, an example of data transfer from the DRP 11 to the accelerator 12 via the data transfer unit 13 will be described. Hereinafter, an example of data transfer via the data transfer unit 13 according to an operation mode (hereinafter referred to as the input mode) in which data is input from the DRP 11 to the accelerator 12 via the data transfer unit 13 will be described.

[0148] Figure 22 is a schematic diagram showing the parallel arithmetic unit MAC256_0, the data transfer unit 13, and the accelerator 12 when the input mode is the first input mode. In this case, the data transfer unit 13 uses the selection circuit 131 to output 64-bit (16 bits × 4) data DQout_0 as DPin_0. The 16-bit data DPin_00 to DPin_03 that make up the 64-bit data DPin_0 are respectively input to the parallel arithmetic units MACMAC64_0 to MAC64_3.

[0149] The relationship between the data transfer unit 13 and the parallel arithmetic units MAC256_1 to MAC256_3 is the same as the relationship between the data transfer unit 13 and the parallel arithmetic unit MAC256_0, and the description thereof is omitted.

[0150] Figure 23 is a schematic diagram showing the parallel arithmetic unit MAC256_0, the data transfer unit 13, and the accelerator 12 when the input mode is the second input mode. In this case, the data transfer unit 13 uses the selection circuit 131 to divide the data DQout_00 into two 16-bit data slices DQout_00 and DQout_02 that make up 32-bit (16 bits × 2) data DQout_0, and outputs the divided 16-bit data slices DPin_00 and DPin_01, and also divides the data DQout_02 into two slices and outputs the divided 16-bit data slices DPin_02 and DPin_03. These 16-bit data DPin_00 to DPin_03 are respectively input to the parallel arithmetic units MACMAC64_0 to MAC64_3.

[0151] The relationship between the data transfer unit 13 and the parallel arithmetic units MAC256_1 to MAC256_3 is the same as the relationship between the data transfer unit 13 and the parallel arithmetic unit MAC256_0, and the description thereof is omitted.

[0152] Figure 24FIG. is a schematic diagram showing the parallel arithmetic unit MAC256_0, the data transfer unit 13, and the accelerator 12 when the input mode is the third input mode. In this case, the data transfer unit 13 distributes 16-bit data DQout_0 into four data slices using the selection circuit 131, and outputs the divided data slices as 16-bit data DPin_00 to DPin_03. These 16-bit data DPin_00 to DPin_03 are respectively input to the parallel arithmetic units MAC64_0 to MAC64_3.

[0153] The relationship between the data transfer unit 13 and the parallel arithmetic units MAC256_1 to MAC256_3 is the same as the relationship between the data transfer unit 13 and the parallel arithmetic unit MAC256_0, and the description thereof is omitted.

[0154] Figure 25 FIG. is a schematic diagram showing the parallel arithmetic unit MAC256_0, the data transfer unit 13, and the accelerator 12 when the input mode is the fourth input mode. In this case, the data transfer unit 13 alternatively selects 16-bit data DQout_00 and DQout_01 (in the example shown in Figure 25 B1, B2, B3, and B4 are selected in this order) from the 16-bit data DQout_00 to DQout_03 that make up the 64-bit data DQout_0 (16 bits × 4) using the selection circuit 131, distributes the selection result into two, and outputs 16-bit data DPin_00 and DPin_01. The remaining 16-bit data DQout_02 and DQout_03 are alternatively selected (in the example shown in Figure 25 A1, A2, A3, and A4 are selected in this order), and the selection result is divided into two and output as 16-bit data DPin_02 and DPin_03. These 16-bit data DPin_00 to DPin_03 are respectively input to the parallel arithmetic units MAC64_0 to MAC64_3.

[0155] The relationship between the data transfer unit 13 and the parallel arithmetic units MAC256_1 to MAC256_3 is the same as the relationship between the data transfer unit 13 and the parallel arithmetic unit MAC256_0, and the description thereof is omitted.

[0156] At this time, two data slices to be output during one output process of the DRP 11 are input to each input terminal of the accelerator 12. Therefore, the processing speed of the accelerator 12 is balanced by doubling the processing speed of the DRP 11. In order to maximize the processing performance of the accelerator 12, it is preferable to adjust the processing speed of the accelerator 12 to be slightly slower than twice the processing speed of the DRP 11. When data is intermittently output from the DRP 11, it is preferable to increase the processing rate of the DRP 11 according to the degree of intermittency of the data, because the processing performance of the accelerator 12 can be maximized.

[0157] Figure 26 is a schematic diagram showing the parallel arithmetic unit MAC256_0, the data transfer unit 13, and the accelerator 12 when the input mode is the fifth input mode. In this case, the data transfer unit 13 alternately selects 16-bit data DQout_00 and DQout_01 (in the example shown in Figure 26 A1, A2, A3, and A4 are selected in this order), distributes the selection result into four, and outputs 16-bit data DPin_00 to DPin_03. These 16-bit data DPin_00 to DPin_03 are respectively input to the parallel arithmetic units MAC64_0 to MAC64_3.

[0158] The relationship between the data transfer unit 13 and the parallel arithmetic units MAC256_1 to MAC256_3 is the same as the relationship between the data transfer unit 13 and the parallel arithmetic unit MAC256_0, and the description thereof is omitted.

[0159] At this time, two data slices to be output during one output process of the DRP 11 are input to each input terminal of the accelerator 12. Therefore, the processing speed of the accelerator 12 is balanced by doubling the processing speed of the DRP 11. In order to maximize the processing performance of the accelerator 12, it is preferable to adjust the processing speed of the accelerator 12 to be slightly slower than twice the processing speed of the DRP 11. When data is intermittently output from the DRP 11, it is preferable to increase the processing rate of the DRP 11 according to the degree of intermittency of the data, because the processing performance of the accelerator 12 can be maximized.

[0160] Figure 27 is a schematic diagram showing the parallel arithmetic unit MAC256_0, the data transfer unit 13, and the accelerator 12 when the input mode is the sixth input mode. In this case, the data transfer unit 13 sequentially selects 16-bit data DQout_00 to DQout_02 (inFigure 27 In the example shown, A1, A2, A3, A4, A5, and A6 are selected in sequence), the selection results are distributed into four, and 16-bit data DPin_00 to DPin_03 are output. These 16-bit data DPin_00 to DPin_03 are respectively input to the parallel arithmetic units MACMAC64_0 to MAC64_3.

[0161] The relationship between the data transfer unit 13 and the parallel arithmetic units MAC256_1 to MAC256_3 is the same as the relationship between the data transfer unit 13 and the parallel arithmetic unit MAC256_0, and the description thereof is omitted.

[0162] At this time, three data slices to be output in one output process of the DRP 11 are input to each input terminal of the accelerator 12. Therefore, if the processing speed of the accelerator 12 is three times the processing speed of the DRP 11, it will be well balanced. To maximize the processing performance of the accelerator 12, it is preferable to adjust the processing speed of the accelerator 12 to be slightly slower than three times the processing speed of the DRP 11. When data is intermittently output from the DRP 11, it is preferable to increase the processing rate of the DRP 11 according to the degree of intermittency of the data, because the processing performance of the accelerator 12 can be maximized.

[0163] Figure 28 is a schematic diagram showing the parallel arithmetic unit MAC256_0, the data transfer unit 13, and the accelerator 12 when the input mode is the seventh input mode. In this case, the data transfer unit 13 uses the selection circuit 131 to sequentially select 16-bit data DQout_00 to DQout_03 (in Figure 28 In the example shown, A1, A2, A3, A4, A5, A6, A7, and A8 are selected in this order), the selection results are distributed into four, and 16-bit data DPin_00 to DPin_03 are output. These 16-bit data DPin_00 to DPin_03 are respectively input to the parallel arithmetic units MACMAC64_0 to MAC64_3.

[0164] The relationship between the data transfer unit 13 and the parallel arithmetic units MAC256_1 to MAC256_3 is the same as the relationship between the data transfer unit 13 and the parallel arithmetic unit MAC256_0, and the description thereof is omitted.

[0165] At this time, four data slices to be output in one DRP 11 output process are input to each input terminal of the accelerator 12. Therefore, if the processing speed of the accelerator 12 is four times that of the DRP 11, it will be well balanced. In order to maximize the processing performance of the accelerator 12, it is preferable to adjust the processing speed of the accelerator 12 to be slightly slower than four times the processing speed of the DRP 11. When data is output intermittently from the DRP 11, it is preferable to increase the processing rate of the DRP 11 according to the degree of intermittency of the data, because the processing performance of the accelerator 12 can be maximized.

[0166] As described above, the semiconductor device 1 according to the present embodiment can arbitrarily change the degree of parallelization of the parallel arithmetic processing of the data input from the DRP 11 to the accelerator 12 via the data transfer unit 13. It should be noted that data processing is efficient when the data output rate from the DRP 11 is adjusted to match the processing throughput of the accelerator 12. Specifically, if the data output rate from the DRP 11 is set to be slightly higher than the processing throughput of the accelerator 12, the processing performance of the accelerator 12 can be maximized.

[0167] Example of data transfer through the data transfer unit 14

[0168] Next, an example of data transfer from the accelerator 12 to the DRP 11 through the data transfer unit 14 will be described. Hereinafter, an example of data transfer through the data transfer unit 14 will be described according to an operation mode (hereinafter referred to as an output mode) in which data is output from the accelerator 12 to the DRP 11 via the data transfer unit 14. The data DPout_0 is composed of data DPout_00 to DPout_03, which will be described later.

[0169] Figure 29FIG. 0 is a schematic diagram showing the parallel arithmetic unit MAC256_0 and the data transfer unit 14 in the accelerator 12 when the output mode is the first output mode. In this case, the data transfer unit 14 sequentially selects one data from the maximum 64 16-bit data DPout_00 output in parallel from the parallel arithmetic unit MAC64_0 using the selection circuit 141, and sequentially outputs the selected data as 16-bit data DQin_00. In addition, 16-bit data DQin_01 is sequentially output by selecting one data from DPout_01 having the maximum 64 16-bit data output in parallel from the parallel processor MAC64_1. In addition, 16-bit data DQin_02 is sequentially output by selecting one data from DPout_02 having the maximum 64 16-bit data output in parallel from the parallel processor MAC64_2. In addition, the maximum 64 16-bit data DPout_03 output in parallel from the parallel processor MAC64_3 are sequentially selected, and the selected data are sequentially output as 16-bit data DQin_03. That is, the data transfer unit 14 sequentially outputs 64-bit wide data DQin_0 composed of 16-bit data DQin_00 to DQin_03.

[0170] The relationship between the parallel arithmetic units MAC256_1 to MAC256_3 and the data transfer unit 14 is the same as the relationship between the parallel arithmetic unit MAC256_0 and the data transfer unit 14, and the description thereof will be omitted.

[0171] Figure 30 FIG. 7 is a schematic diagram showing the parallel arithmetic unit MAC256_0 and the data transfer unit 14 in the accelerator 12 when the output mode is the second output mode. In this case, the data transfer unit 14 includes a selection circuit 141 composed of a first selection circuit 141_1 and a second selection circuit 141_2.

[0172] First, selection circuit 141_1 sequentially selects one data from the maximum 64 16-bit data DPout_00 output in parallel from parallel arithmetic unit MAC64_0, and sequentially outputs the selected data as 16-bit data DQin_00. Additionally, 16-bit data DQin_01 is sequentially output by selecting one by one from DPout_01 which has the maximum 64 16-bit data output in parallel from parallel processor MAC64_1. Additionally, 16-bit data DQin_02 is sequentially output by selecting one by one from DPout_02 which has the maximum 64 16-bit data output in parallel from parallel processor MAC64_2. Additionally, 16-bit data DQin_03 is sequentially output by selecting one by one from DPout_03 which has the maximum 64 16-bit data output in parallel from parallel processor MAC64_3.

[0173] After that, selection circuit 141_2 outputs 16-bit data DQin_00, and then outputs 16-bit data DQin_01. In parallel, 16-bit data DQin_2 is output, followed by 16-bit data DQin_3. That is, data transfer unit 14 sequentially outputs 32-bit wide data DQin_0 composed of one of data DQin_00 and DQin_01 output from selection circuit 141_2 and one of DQin_02 and DQin_03.

[0174] Data transfer unit 14 can alternatively use selection circuit 141_2 to output 16-bit data DQin_00 and 16-bit data DQin_01. 16-bit data DQin_02 and 16-bit data DQin_03 can be output alternatively.

[0175] The relationship between parallel arithmetic units MAC256_1 to MAC256_3 and data transfer unit 14 is the same as the relationship between parallel arithmetic unit MAC256_0 and data transfer unit 14, and its description will be omitted.

[0176] Figure 31 is a schematic diagram showing parallel arithmetic unit MAC256_0 and data transfer unit 14 of accelerator 12 when the output mode is the third output mode. In this case, data transfer unit 14 includes selection circuit 141 composed of first selection circuit 141_1 and second selection circuit 141_2.

[0177] First, selection circuit 141_1 sequentially selects one data from the maximum 64 16-bit data DPout_00 output in parallel from parallel arithmetic unit MAC64_0, and sequentially outputs the selected data as 16-bit data DQin_00. Additionally, 16-bit data DQin_01 is sequentially output by sequentially selecting one by one from DPout_01 which has the maximum 64 16-bit data output in parallel from parallel arithmetic unit MAC64_1. Additionally, 16-bit data DQin_02 is sequentially output by sequentially selecting one by one from DPout_02 which has the maximum 64 16-bit data output in parallel from parallel arithmetic unit MAC64_2. Additionally, 16-bit data DQin_03 is sequentially output by sequentially selecting one by one from DPout_03 which has the maximum 64 16-bit data output in parallel from parallel arithmetic unit MAC64_3.

[0178] Thereafter, selection circuit 141_2 sequentially selects one data from 16-bit data DQin_00 to DQin_03, and sequentially outputs the selected data as 16-bit wide data DQin_0.

[0179] The relationship between parallel arithmetic units MAC256_1 to MAC256_3 and data transfer unit 14 is the same as the relationship between parallel arithmetic unit MAC256_0 and data transfer unit 14, and its description will be omitted.

[0180] Figure 32 FIG. is a schematic diagram of parallel arithmetic unit MAC256_0 and data transfer unit 14 of accelerator 12 when the output mode is the fourth output mode. In this case, data transfer unit 14 includes selection circuit 141 composed of first selection circuit 141_1 and second selection circuit 141_2.

[0181] First, selection circuit 141_1 sequentially selects one data from the maximum 64 16-bit data DPout_00 output in parallel from parallel arithmetic unit MAC64_0, and sequentially outputs the selected data as 16-bit data DQin_00 (in the example of Figure 32 are C1, C2, C3, C4,...). Additionally, 16-bit data DPout_01 is sequentially selected one by one from the maximum 64 16-bit data DPout_01 output in parallel from parallel arithmetic unit MAC64_01, and is sequentially output as 16-bit data DQin_01 (in Figure 32In the example, they are D1, D2, D3, D4, ...). Additionally, the maximum 64 16-bit data DPout_02 output in parallel from the parallel arithmetic unit MAC64_2 are sequentially selected one by one, and the selected data is sequentially output as 16-bit data DQin_02 (E1, E2, E3, E4, ...). Additionally, the maximum 64 16-bit data DPout_03 output in parallel from the parallel arithmetic unit MAC64_3 are sequentially selected one by one, and the selected data is sequentially output as 16-bit data DQin_03 (in Figure 32 the example, they are F1, F2, F3, F4, ...).

[0182] Thereafter, the selection circuit 141_2 alternatively outputs 16-bit data DQin_00 and 16-bit data DQin_01 as 32-bit data. In parallel with this, 16-bit data DQin_02 and 16-bit data DQin_03 are output in order (in this example, four elements in order) and two data slices are collectively output as 32-bit data. That is, the data transfer unit 14 sequentially outputs 64-bit wide data DQin_0.

[0183] The relationship between the parallel arithmetic units MAC256_1 to MAC256_3 and the data transfer unit 14 is the same as the relationship between the parallel arithmetic unit MAC256_0 and the data transfer unit 14, and its description will be omitted.

[0184] At this time, data is input to the DRP 11 at a rate of 1 / 2 of the data output from the accelerator 12. Therefore, when the processing speed of the accelerator 12 is approximately twice the processing speed of the DRP 11, the data transfer speed of the data output from the accelerator 12 after efficiently performing the parallel arithmetic operation processing without being limited by the processing rate of the DRP 11 can be reduced to the processing speed of the DRP 11.

[0185] Figure 33 is a schematic diagram showing the parallel arithmetic unit MAC256_0 and the data transfer unit 14 of the accelerator 12 when the output mode is the fifth output mode. In this case, the data transfer unit 14 includes a selection circuit 141 composed of a first selection circuit 141_1 and a second selection circuit 141_2.

[0186] First, the selection circuit 141_1 sequentially selects one by one from the maximum 64 16-bit data DPout_00 output in parallel from the parallel arithmetic unit MAC64_0, and sequentially outputs the selected data as 16-bit data DQin_00 (in Figure 33In the example, they are C1, C2, C3, C4,...). Additionally, the maximum 64 16-bit data DPout_01 output in parallel from the parallel arithmetic unit MAC64_1 are sequentially selected one by one, and are sequentially output as 16-bit data DQin_01 (in Figure 33 In the example, they are C1, C2, C3, C4,...). Additionally, the maximum 64 16-bit data DPout_02 output in parallel from the parallel arithmetic unit MAC64_2 are sequentially selected one by one, and are sequentially output as 16-bit data DQin_02 (in Figure 33 In the example, they are E1, E2, E3, E4,...). Additionally, the maximum 64 16-bit data DPout_03 output in parallel from the parallel arithmetic unit MAC64_3 are sequentially selected one by one, and are sequentially output as 16-bit data DQin_03 (in Figure 33 In the example, they are F1, F2, F3, F4,...).

[0187] Thereafter, the selection circuit 141_2 sequentially outputs the 16-bit data DQin_00 to DQin_03 in order (in this example, in the order of four elements) and collects two data slices as 32-bit wide data DQin_0.

[0188] The relationship between the parallel arithmetic units MAC256_1 to MAC256_3 and the data transfer unit 14 is the same as the relationship between the parallel arithmetic unit MAC256_0 and the data transfer unit 14, and its description will be omitted.

[0189] At this time, the data is input to the DRP 11 at a rate of 1 / 2 of the data output from the accelerator 12. Therefore, in particular, when the processing speed of the accelerator 12 is approximately twice the processing speed of the DRP 11, the data transfer speed of the data output from the accelerator 12 after efficiently performing parallel arithmetic processing without being limited by the processing rate of the DRP 11 can be reduced to the DRP 11 processing speed.

[0190] Figure 34 is a schematic diagram showing the parallel arithmetic unit MAC256_0 and the data transfer unit 14 of the accelerator 12 in the case where the output mode is the sixth output mode. In this case, the data transfer unit 14 includes a selection circuit 141 composed of a first selection circuit 141_1 and a second selection circuit 141_2.

[0191] First, the selection circuit 141_1 sequentially selects one by one from the maximum 64 16-bit data DPout_00 output in parallel from the parallel arithmetic unit MAC64_0, and sequentially outputs the selected data as 16-bit data DQin_00 (inFigure 34 In the example of, they are C1, C2, C3, C4,...). In addition, the maximum 64 16-bit data DPout_01 output in parallel from the parallel arithmetic unit MAC64_1 are sequentially selected one by one, and are sequentially output as 16-bit data DQin_01 (in Figure 34 In the example of, they are C1, C2, C3, C4,...). In addition, the maximum 64 16-bit data DPout_02 output in parallel from the parallel arithmetic unit MAC64_2 are sequentially selected one by one, and are sequentially output as 16-bit data DQin_02 (in Figure 34 In the example of, they are E1, E2, E3, E4,...). In addition, the maximum 64 16-bit data DPout_03 output in parallel from the parallel arithmetic unit MAC64_3 are sequentially selected one by one, and are sequentially output as 16-bit data DQin_03 (in Figure 34 In the example of, they are F1, F2, F3, F4,...).

[0192] Thereafter, the selection circuit 141_2 sequentially outputs the 16-bit data DQin_00 to DQin_03 in order (in this example, in the order of four elements) and collects three data slices as 48-bit wide data DQin_0.

[0193] The relationship between the parallel arithmetic units MAC256_1 to MAC256_3 and the data transfer unit 14 is the same as the relationship between the parallel arithmetic unit MAC256_0 and the data transfer unit 14, and its description will be omitted.

[0194] At this time, the data is input to the DRP 11 at a rate one-third of the data output from the accelerator 12. Therefore, when the processing speed of the accelerator 12 is approximately three times the processing speed of the DRP 11, the data transfer speed of the data output from the accelerator 12 after efficiently performing parallel arithmetic processing without being limited by the processing rate of the DRP 11 can be reduced to the processing speed of the DRP 11.

[0195] Figure 35 is a schematic diagram showing the parallel arithmetic unit MAC256_0 and the data transfer unit 14 of the accelerator 12 when the output mode is the seventh output mode. In this case, the data transfer unit 14 includes a selection circuit 141 composed of a first selection circuit 141_1 and a second selection circuit 141_2.

[0196] First, selection circuit 141_1 sequentially selects one by one from the maximum 64 16-bit data DPout_00 output in parallel from parallel arithmetic unit MAC64_0, and sequentially outputs the selected data as 16-bit data DQin_00 (in the example of Figure 35 it is C1, C2, C3, C4,...). Additionally, the maximum 64 16-bit data DPout_01 output in parallel from parallel arithmetic unit MAC64_1 is sequentially selected one by one, and is sequentially output as 16-bit data DQin_01 (in the example of Figure 35 it is C1, C2, C3, C4,...). Additionally, 16-bit data DPout_02 is sequentially selected one by one from the maximum 64 16-bit data DPout_02 output in parallel from parallel arithmetic unit MAC64_02, and is sequentially output as 16-bit data DQin_02 (in the example of Figure 35 it is E1, E2, E3, E4,...). Additionally, 16-bit data DPout_03 is sequentially selected one by one from the maximum 64 16-bit data DPout_03 output in parallel from parallel arithmetic unit MAC64_03, and is sequentially output as 16-bit data DQin_03 (in the example of Figure 35 it is F1, F2, F3, F4,...).

[0197] Thereafter, selection circuit 141_2 sequentially outputs 16-bit data DQin_00 to DQin_03 in order (in this example, in the order of four elements) and collects four data slices as 64-bit wide data DQin_0.

[0198] The relationship between parallel arithmetic units MAC256_1 to MAC256_3 and data transfer unit 14 is the same as the relationship between parallel arithmetic unit MAC256_0 and data transfer unit 14, and its description will be omitted.

[0199] At this time, data is input to DRP 11 at a rate of 1 / 4 of the data output from accelerator 12. Therefore, when the processing speed of accelerator 12 is approximately four times the processing speed of DRP 11, the data transfer speed of the data output from accelerator 12 after efficiently performing parallel arithmetic processing without being limited by the processing rate of DRP 11 can be reduced to the processing speed of DRP 11.

[0200] As described above, in the semiconductor device 1 according to the present embodiment, the data output from the accelerator 12 to the DRP 11 via the data transfer unit 14 can be changed to data of any bit width. In order to maximize the performance of the accelerator 12, it is preferable that the data rate received by the DRP 11 is slightly higher than the data rate output from the accelerator 12.

[0201] Figure 36 is a schematic diagram showing the flow of arithmetic operation processing of the parallel arithmetic section 121 when performing arithmetic operations on input data with maximum parallelization. As Figure 36 shown, the data DQout_0 output from the DRP 11 is distributed by the data transfer unit 13 to the parallel arithmetic units MAC64_0 to MAC64_3 provided in the parallel arithmetic units MAC256_0 to MAC256_3 respectively and provided as data DPin_0 to DPin_3. At this time, the parallel arithmetic section 121 can perform arithmetic operations on the data DQout_0 (data DPin_0 to DPin_3) in parallel by using up to 1024 arithmetic units. Note that the data transfer unit 14 is configured to selectively output the arithmetic operation results output in parallel from each of the 1024 arithmetic units, so that these arithmetic operation results can be converted into data of a desired bit width and output to the DRP 11.

[0202] Figure 37 is a schematic diagram showing the flow of arithmetic operations of the parallel arithmetic section 121 when performing arithmetic operations on input data in the case where the degree of parallelization is the minimum unit. As Figure 37 shown, the data DQout_0 output from the DRP 11 is provided by the data transfer unit 13 to the parallel arithmetic unit MAC64_0 provided in the parallel arithmetic unit MAC256_0 as data DPin_0. At this time, the parallel operation section 121 can perform arithmetic operations on the data DQout_0 (data DPin_0) in parallel by using one to 64 arithmetic units among the 64 arithmetic units provided in the parallel arithmetic unit MAC64_0.

[0203] Figure 38 is a schematic diagram showing the flow of arithmetic operation processing of the parallel arithmetic section 121 when performing arithmetic operations on input data with a medium-level degree of parallelization. In Figure 38In the embodiment, data DQout_0 output from DRP 11 is distributed by data transmission unit 13 to parallel arithmetic units MAC64_0 to MAC64_3 provided in parallel arithmetic unit MAC256_0 and parallel arithmetic units MAC64_0 to MAC64_2 provided in parallel arithmetic unit MAC256_0, respectively, and provided as data DPin_0 and DPin_1. Here, parallel arithmetic section 121 can perform arithmetic operation processing on data DQout_0 (data DPin_0 and DPin_1) in parallel by using, for example, 400 arithmetic units.

[0204] Figure 39 is a schematic diagram showing the flow of arithmetic operations of parallel arithmetic section 121 when performing parallel arithmetic operations on each of two input data. In Figure 39 In the embodiment, data DQout_0 output from DRP 11 is distributed by data transmission unit 13 to parallel arithmetic units MAC64_0 to MAC64_3 provided in parallel arithmetic unit MAC256_0 and parallel arithmetic units MAC64_0 to MAC64_2 provided in parallel arithmetic unit MAC256_1, respectively, and provided as data DPin_0 and DPin_1. In addition, data DQout_2 output from DRP 11 is distributed by data transmission unit 13 to parallel arithmetic units MAC64_0 and MAC64_1 provided in parallel arithmetic unit MAC256_2 and provided as data DPin_2. At this time, parallel arithmetic section 121 can perform arithmetic operations on data DQout_0 (data DPin_0 and DPin_1) in parallel by using, for example, 400 arithmetic units, and perform arithmetic operations on data DQout_2 (data DPin_2) in parallel by using, for example, 120 arithmetic units.

[0205] For example, in the case of performing arithmetic operations on two or more input data using a plurality of arithmetic units different from each other, the plurality of arithmetic units for arithmetic operation processing for one input data and the plurality of arithmetic units for arithmetic operation processing for another input data may be provided with individual predetermined data read out from local memory 122, or may be provided with common predetermined data.

[0206] Second Embodiment

[0207] Figure 40 is a block diagram showing an exemplary configuration of semiconductor system SYS1a on which semiconductor device 1a according to the second embodiment is mounted. Compared with Figure 1 semiconductor device 1 shown in Figure 40 semiconductor device 1a shown in has DRP 11a instead of DRP 11.

[0208] The DRP 11a has, for example, two state management units (STC; state transition controllers) 111 and 112. One state management unit 111 performs arithmetic operations on the data read from the external memory 3 and outputs the arithmetic operation results to the accelerator 12. The other state management unit 112 performs arithmetic operations on the data output from the accelerator 12 and writes the arithmetic operation results to the external memory 3. That is, the DRP 11a independently operates the processing of the data to be sent to the accelerator 12 and the processing of the data received from the accelerator 12. As a result, in the DRP 11a, the operation instructions (applications) given when performing dynamic reconfiguration can be made simpler than the dynamic reconfiguration instructions (applications) when performing the dynamic reconfiguration operation (DRP 11). It also allows the DRP 11a to reconfigure the circuit more easily than the DRP 11.

[0209] In addition, the DRP 11a is provided with two state management units for independently operating the processing of the data to be sent to the accelerator 12 and the processing of the data received from the accelerator 12. Thus, for example, the degree of flexibility of the arrangement of the external input terminals to which the data read from the external memory 3 is input, the external output terminals to which the data directed to the accelerator 12 is output, the external input terminals to which the data from the accelerator 12 is input, and the external output terminals to which the write data directed to the external memory 3 is output can be increased.

[0210] As described above, the semiconductor devices according to the first and second embodiments include: an accelerator having a parallel arithmetic section that performs arithmetic operations in parallel; a data processing unit such as a DRP that sequentially transfers data; and a data transfer unit that sequentially selects multiple arithmetic operation processing results of the accelerator and outputs them to the data processing unit. Therefore, the semiconductor devices according to the first and second embodiments and the semiconductor systems including them can perform a large amount of conventional data processing by using the accelerator and perform other data processing by using the data processing unit, enabling efficient arithmetic operations even in large-scale arithmetic processing such as deep learning processing.

[0211] Although the present invention made by the present inventor has been specifically described based on the embodiments, the present invention is not limited to the described embodiments, and undoubtedly, various modifications can be made without departing from the object of the present invention.

[0212] In the first and second embodiments described above, it is described that individual predetermined data read from the local memory 122 is provided to a plurality of arithmetic units constituting the parallel arithmetic section 121, but the present invention is not limited thereto. The common predetermined data read from the local memory 122 may be provided to all or in groups of a plurality of arithmetic units constituting the parallel arithmetic section 121. In this case, the circuit scale and power consumption of the local memory 122 can be reduced.

[0213] Some or all of the above embodiments may be described as the following appendices, but the present invention is not limited to the following.

[0214] (Appendix 1)

[0215] A semiconductor device, comprising: a data processing unit that performs data processing on first input data sequentially input and sequentially outputs the result of the data processing as first output data; a parallel arithmetic unit that performs arithmetic processing in parallel between the first output data sequentially output from the data processing unit and each of a plurality of predetermined data; a holding circuit that holds the result of the arithmetic processing; and a first data transmission unit that sequentially selects a plurality of arithmetic processing results held by the accelerator and sequentially outputs the result of the arithmetic processing as the first input data.

[0216] (Appendix 2)

[0217] The semiconductor device according to Appendix 1, wherein the data processing unit is a processor capable of being dynamically reconfigured based on operation commands sequentially given.

[0218] (Appendix 3)

[0219] A semiconductor system, comprising: a semiconductor device as described in Appendix 3; an external memory; and a control unit that controls the operation of the semiconductor device based on control instructions read from the external memory.

Claims

1. A semiconductor device, comprising: A data processing unit (11); An accelerator (2); A first data transmission unit (13); A second data transmission unit (14), wherein the data processing unit (11) is configured to 1) sequentially receive first input data (DQin), 2) perform data processing on the first input data, and 3) sequentially output first output data (DQout) as a result of the data processing to the data transmission unit (13), wherein the first data transmission unit (13) is configured to 1) receive the first output data (DQout) from the data processing unit (11), and 2) output the first output data (DQout) to a plurality of arithmetic units grouped as a first arithmetic unit group among a plurality of arithmetic units in the parallel arithmetic part of the accelerator; wherein the first arithmetic unit group in the parallel arithmetic part (121) of the accelerator is configured to 1) sequentially receive the first output data (DQout) from the first data transmission unit (13), 2) perform arithmetic operations in parallel between the first output data (DQout) and each of a plurality of predetermined data, and 3) output a plurality of arithmetic operation results (DPout) to the second data transmission unit (14), and wherein the second data transmission unit (14) is configured to 1) receive the plurality of arithmetic operation results (DPout) from the accelerator, and 2) sequentially output the plurality of arithmetic operation results to the data processing unit as the first input data (DQin), wherein the data processing unit (11) is configured to sequentially output second output data in parallel with the first output data (DQout), wherein the first data transmission unit (13) is further configured to selectively output the second output data to a plurality of arithmetic units grouped as a second arithmetic unit group different from the first arithmetic unit group among the plurality of arithmetic units in the parallel arithmetic part (121) of the accelerator, wherein the second arithmetic unit group is configured to: 1) sequentially receive the second output data from the first data transmission unit (13); 2) perform arithmetic operations in parallel between the second output data and each of the plurality of predetermined data, wherein the first data transmission unit (13) is set to a first mode or a second mode based on the processing speed of the accelerator (2) relative to the processing speed of the data processing unit (11), wherein, when the first data transmission unit (13) is set to the first mode, the first data transmission unit (13) outputs the first output data (DQout) and the second output data to the first arithmetic unit group and the second arithmetic unit group respectively, Among them, when the first data transmission unit (13) is set to the second mode, the first data transmission unit (13) sequentially selects the first output data (DQout) and the second output data and outputs the selected results to the first arithmetic unit group and the second arithmetic unit group.

2. The semiconductor device according to claim 1, wherein the first output data (DQout) is commonly input to each of the plurality of arithmetic units in the first arithmetic unit group, wherein each of the plurality of arithmetic units in the first arithmetic unit group performs the arithmetic operation between the corresponding predetermined data among the plurality of predetermined data and the first output data (DQout).

3. The semiconductor device according to claim 1, wherein each of the plurality of arithmetic units in the parallel arithmetic units of the accelerator includes an adder and a multiplier to perform a multiply-accumulate operation.

4. The semiconductor device according to claim 1, wherein the second data transmission unit (14) sequentially outputs a plurality of second arithmetic operation results of the second arithmetic unit group to the data processing unit (11) as second input data, and wherein the data processing unit (11) performs the data processing on the second input data in parallel with the data processing on the first input data.

5. The semiconductor device according to claim 4, wherein the second data transmission unit (14) sequentially outputs the plurality of first arithmetic operation results and the plurality of second arithmetic operation results as the first input data (DQin).

6. The semiconductor device according to claim 1, wherein the second data transmission unit (14) is set to a third mode or a fourth mode based on the processing speed of the accelerator (2) relative to the processing speed of the data processing unit (11), Among them, when the second data transmission unit (14) is set to the third mode, the second data transmission unit (14) sequentially outputs the first arithmetic operation results of the first arithmetic unit as the first input data (DQin), and wherein, when the second data transmission unit (14) is set to the fourth mode, the second data transmission unit (14) collectively selects at least two arithmetic operation results among the first arithmetic operation results to output as the first input data (DQin).

7. The semiconductor device according to claim 1, wherein the first data transmission unit (13) is configured to selectively output the first output data (DQout) or the second output data to the second arithmetic unit group.

8. The semiconductor device according to claim 1, wherein the first data transmission unit (13) selects the first output data (DQout) and the second output data and sequentially outputs the first output data (DQout) and the second output data to the first arithmetic unit group and the second arithmetic unit group.

9. The semiconductor device according to claim 1, wherein each arithmetic unit in the arithmetic unit includes a plurality of arithmetic circuits and a selector, and the selector selectively outputs the results of the plurality of arithmetic circuits.

10. The semiconductor device according to claim 1, wherein the data processing unit sequentially outputs the first output data obtained by performing data processing on the data read from the external memory, and provides the result of performing data processing on the sequentially input first input data to the external memory.

11. The semiconductor device according to claim 10, wherein the data processing unit includes: A first state management unit that controls the arithmetic operation for generating the first output data, and a second state management unit that controls the arithmetic operation for the first input data, wherein the second state management unit is different from the first state management unit.

12. The semiconductor device according to claim 1, wherein the accelerator further includes a local memory that stores the plurality of predetermined data.

13. The semiconductor device according to claim 12, wherein the arithmetic units grouped into a first set of the first arithmetic unit group and the arithmetic units grouped into a second set of the second arithmetic unit group are commonly provided with the plurality of predetermined data read from the local memory.

14. The semiconductor device according to claim 12, wherein the arithmetic units grouped into a first set of the first arithmetic unit group and the arithmetic units grouped into a second set of the second arithmetic unit group are respectively provided with different predetermined data read from the local memory.

15. A method of controlling a semiconductor device, the method comprising: sequentially receiving first input data (DQin) at a data processing unit (11); performing data processing on the first input data (DQin) using the data processing unit; sequentially outputting the result of the data processing from the data processing unit as first output data (DQout); sequentially receiving the first output data (DQout) at the accelerator (2) by the data processing unit; using the accelerator (2) to perform arithmetic operations in parallel between the first output data (DQout) sequentially output from the data processing unit and each of the plurality of predetermined data, and outputting a plurality of arithmetic operation results; receiving the plurality of arithmetic operation results at a first data transmission unit from the accelerator (2); and sequentially outputting the plurality of arithmetic operation results from the first data transmission unit to the data processing unit (11) as the first input data (DQin), wherein the accelerator includes a parallel arithmetic section, and the parallel arithmetic section includes a first arithmetic unit group having a plurality of arithmetic units, wherein the first data transmission unit sequentially outputs a plurality of first arithmetic results of the first arithmetic unit group as the first input data, wherein the first data transmission unit (13) is set to a first mode or a second mode based on the processing speed of the accelerator (2) relative to the processing speed of the data processing unit (11). Wherein, when the first data transmission unit (13) is set to the first mode, the first data transmission unit (13) outputs the first arithmetic result of the first arithmetic unit in sequence as the first input data (DQin), and Wherein, when the first data transmission unit (13) is set to the second mode, the first data transmission unit (13) collectively selects at least two arithmetic results from the first arithmetic results to output as the first input data (DQin).

Citation Information

Patent Citations

  • Air guide

    JP2018114861A

  • System and Method for Parallelizing and Accelerating Learning Machine Training and Classification Using a Massively Parallel Accelerator

    US20090304268A1

  • Arithmetic processing device

    WO2017006512A1