Integrated circuit, and method
By allowing multiple processors to record and read from multiple memory banks concurrently, the neural network hardware accelerators overcome single-memory bank access limitations, speeding up inference processes.
Patent Information
- Application Number
- JP2024525511
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-26
- Filing Date
- 2022-06-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-22
AI Technical Summary
In neural network hardware accelerators, multiple processors are limited by a single memory bank access, causing delays as other processors have to wait to read data, thereby slowing down the inference process.
Configuring neural network hardware accelerators with multiple memory banks and processors to allow simultaneous data parallel recording and reading of results, enabling multiple processors to access and transmit values to multiple memory banks concurrently.
Significantly reduces the time required for neural network inference by minimizing waiting times for processors to read values from memory, enhancing processing efficiency.
Smart Images

Figure 0007700402000001 
Figure 0007700402000002 
Figure 0007700402000003
Abstract
Description
Background Art
[0001] Cross - reference to Related Applications This application claims priority to U.S. Non - Provisional Application No. 17 / 510,397, filed on October 26, 2021, the entire contents of which are incorporated herein by reference. Real - time neural network (NN) inference is becoming increasingly popular for computer vision or speech tasks on edge devices for applications such as autonomous vehicles, robotics, smartphones, portable healthcare devices, surveillance, etc. Dedicated NN inference hardware has become the mainstream method for providing power - efficient inference. [Problems to be Solved by the Invention] In a neural network hardware accelerator or the like, when multiple processors execute inference using a memory bank, a single memory bank can only be accessed by one processor at a time, and other processors have to wait to read data, resulting in a delay in the inference process.
Brief Description of the Drawings
[0002] Aspects of the present disclosure are best understood from the following detailed description of the invention when read in conjunction with the accompanying drawings. Note that, in accordance with standard practice in the industry, various features are not drawn to scale. In fact, the dimensions of various features may be arbitrarily increased or decreased for clarity of discussion.
[0003]
Figure 1
[0004]
Figure 2
[0005]
Figure 3
[0006]
Figure 4
[0007]
Figure 5
[0008]
Figure 6
[0009]
Figure 7
[0010]
Figure 8
[0011]
Figure 9
[0012]
Figure 10A
[0013]
Figure 10B
[0014]
Figure 11
[0015] The following disclosure provides many different embodiments or examples for implementing different features of the provided subject matter. Specific examples of components, values, operations, materials, or arrangements, etc. are described below to simplify the present disclosure. Of course, these are merely examples and are not intended to be limiting. Other components, values, operations, materials, or arrangements, etc. are contemplated. In addition, the present disclosure may repeat reference numerals and / or letters in various examples. This repetition is for the purpose of simplicity and clarity and does not in itself define the relationship between the various embodiments and / or configurations being discussed.
[0016] Some neural network hardware accelerators execute such operations by distributing inference operations among multiple processors. Such neural network hardware accelerators also include multiple memory banks for holding various values between operations. Each processor reads a value from one of the memory banks for calculation, but a single memory bank is read by only one processor at a time. The processor reading the memory bank places a lock on the memory bank to prevent other processors from interacting with and, in some cases, destroying the data on the memory bank. Similarly, each processor records the value obtained as a result of the calculation in one of the memory banks, but a single memory bank receives values from only one processor at a time to prevent data destruction.
[0017] The inference of some types of neural networks involves performing more than one subsequent calculation on a value obtained as a result of a single calculation. When performing an inference on a neural network hardware accelerator having multiple processors, different processors can execute subsequent operations simultaneously. However, since only one processor at a time can read the value obtained as a result of a single calculation, one of the processors has to wait for the other processor to read the value, so one subsequent process is delayed compared to the other subsequent process.
[0018] In at least some embodiments of the present specification, a neural network hardware accelerator comprising a plurality of processors and a plurality of memory banks is configured such that each processor can record the result value in more than one memory bank. In at least some embodiments, the neural network hardware accelerator comprises an external memory interface configured to record values in more than one memory bank. In at least some embodiments, the neural network hardware accelerator is configured to perform data parallel recording of values obtained as a result of at least some instructions. In at least some embodiments, by recording each of at least some result values in two memory banks while performing a neural network inference, the amount of time required for the neural network hardware accelerator to complete the inference task is significantly reduced due to the reduction in the waiting time for the processor to read the value from the memory.
[0019] In at least some embodiments of the present specification, a neural network hardware accelerator comprising a plurality of processors and a plurality of memory banks is configured such that each processor can read the value from the memory bank, whereby the other processor is configured to receive the transmission of the value from the memory bank when the value is read. In at least some embodiments, by configuring a plurality of processors to receive the same transmission from a read of a value from a memory bank, the time for the plurality of processors to individually read values from the same memory bank is reduced.
[0020] FIG. 1 is a schematic diagram of an integrated circuit 100 for neural network hardware acceleration data parallel processing according to at least one embodiment of the present invention. In at least some embodiments, the integrated circuit 100 is a field programmable gate array (FPGA) programmed as shown in FIG. 1. In at least some embodiments, the integrated circuit 100 is an application specific integrated circuit (ASIC) having a dedicated circuit as shown in FIG. 1. The integrated circuit 100 includes a data bus 102, a general purpose controller 104, an external memory interface 106, a plurality of computing units such as a computing unit 110, and a plurality of memory banks such as a memory bank 120.
[0021] The data bus 102 has a plurality of interconnects connecting the computing units, the memory banks, and the external memory interface 106. In at least some embodiments, the data bus 102 is configured to facilitate the transmission of values from any computing unit or external memory interface 106 to any memory bank. In at least some embodiments, the data bus 102 is configured to facilitate the transmission of values from any memory bank to any computing unit or external memory interface 106. In at least some embodiments, the data bus 102 has passive interconnects.
[0022] In at least some embodiments, integrated circuit 100 includes a controller, such as general-purpose controller 104, configured to receive instructions for performing neural network inference. General-purpose controller 104 has circuitry configured to execute instructions to cause integrated circuit 100 to perform neural network inference. In at least some embodiments, general-purpose controller 104 is configured to receive compiled instructions from a host processor. In at least some embodiments, the compiled instructions include a schedule of operations, specified compute units for performing each operation, specified memory banks and addresses for storing intermediate data, and any other details required for the integrated circuit to perform neural network inference.
[0023] External memory interface 106 includes circuitry configured to enable memory banks and compute units to exchange data with external memory. In at least some embodiments, the external memory is DRAM memory that communicates with a host processor. In at least some embodiments, the integrated circuit 100 stores a small working portion of the data for neural network inference while the DRAM memory stores the remainder of the data.
[0024] A computing unit, such as computing unit 110, has a circuit configured to perform mathematical operations on input values and stored weight values stored in a memory bank. In at least some embodiments, the computing unit outputs partial sums to the memory bank. In at least some embodiments, the computing unit performs accumulation using existing partial sums stored in the memory bank. In at least some embodiments, the computing unit has at least one processor configured to perform depthwise convolution or pointwise convolution. In at least some embodiments, the computing unit has a general-purpose convolution processor that supports combinations of depthwise convolution and pointwise convolution layers, such as Inverted Residual Blocks in the MobileNet architecture. In at least some embodiments, the computing unit has a processor configured to perform other operations for inference of a deep network or any other type of neural network. In at least some embodiments, integrated circuit 100 includes a plurality of computing units, and each computing unit of the plurality of computing units includes a processor having a circuit configured to perform mathematical operations on input data values and weight values to generate result data values, and a computing controller. In at least some embodiments, the computing unit is configured as shown in FIG. 2, which will be described below.
[0025] A memory bank, such as memory bank 120, has a circuit configured to store data. In at least some embodiments, the memory bank has volatile data storage. In at least some embodiments, integrated circuit 100 includes a plurality of memory banks, and each memory bank of the plurality of memory banks is configured to store values and transmit the stored values. In at least some embodiments, the memory bank is configured as shown in FIG. 2, which will be described below.
[0026] FIG. 2 is a schematic diagram of a calculation unit and a memory bank interconnection in an integrated circuit for neural network hardware acceleration data parallel processing according to at least one embodiment of the present invention. FIG. 2 shows a calculation unit 210 and a memory unit 222 connected through a data bus 202.
[0027] The calculation unit 210 includes a processor 212, a multiplexer 214, and a calculation controller 216. In at least some embodiments, the processor 212 includes a circuit configured to execute mathematical operations. In at least some embodiments, the processor 212 is configured to execute convolution operations such as point-wise convolution or depth-wise convolution. In at least some embodiments, the processor 212 is configured to execute point-wise convolution or depth-wise convolution. In at least some embodiments, the processor 212 is configured to directly support different parameters of mathematical operations such as a kernel size of height (KH) × width (KW), vertical and horizontal strides, dilation, padding, etc. The processor 212 includes a data input connected to the data bus 202 through the multiplexer 214 and includes a data output connected to the data bus 202.
[0028] The multiplexer 214 includes a plurality of inputs from the data bus 202 and a single output from the processor 212. In at least some embodiments, the multiplexer 214 is configured to select a data input connection from the data bus 202 to the processor 212. In at least some embodiments, the multiplexer 214 is configured to respond to a selection command, such as a signal, from the calculation controller 216. In at least some embodiments, the multiplexer 214 includes inputs from the data bus 202 for each memory bank within the integrated circuit. In at least some embodiments, each calculation unit includes a calculation multiplexer that can be configured to connect to one of a plurality of memory banks. In at least some embodiments, each memory bank of the plurality of memory banks is configured to store a value received through a corresponding bank multiplexer.
[0029] The calculation controller 216 includes a circuit configured to operate the calculation unit 210. In at least some embodiments, the calculation controller 216 is configured to receive a signal from a general-purpose controller of the integrated circuit and operate the calculation unit 210 according to the received signal. In at least some embodiments, the calculation controller 216 causes the multiplexer 214 to connect the processor 212 to a specific memory bank, causes the processor 212 to read one or more values from the connected memory bank, causes the processor 212 to perform a mathematical operation on the one or more values, and then causes the processor 212 to record the result value in one or more memory banks. In at least some embodiments, the calculation controller 216 receives an input data value from any one of a plurality of memory banks, receives a weight value from any one of a plurality of memory banks, causes the processor to perform a mathematical operation, and transmits the result data value to at least two of the plurality of memory banks such that a single transmission of the result data value is received at substantially the same time by the at least two memory banks. In at least some embodiments, the computing controller 216 is further configured to set a lock for each of at least two memory banks. In at least some embodiments, the computing controller 216 is further configured to apply a bank offset for one or more of at least two memory banks. In at least some embodiments, the computing controller 216 is further configured to apply a bank offset for one or more of at least two memory banks. In at least some embodiments, the computing controller 216 is further configured to release the lock for each of at least two memory banks. In at least some embodiments, the computing controller 216 synchronizes a second computing unit among a plurality of computing units to receive one of an input data value or a weight value, and to perform a single transmission of one of the input data value or the weight value such that a memory bank storing one of the input data value or the weight value is read by the computing controller and the second computing unit at substantially the same time, from among a plurality of memory banks, the memory bank storing one of the input data value or the weight value, to read out one of the input data value or the weight value.
[0030] Memory unit 222 includes memory bank 220, multiplexer 224, and memory controller 226. In at least some embodiments, memory bank 220 has substantially the same structure as memory bank 120 of FIG. 1 and performs substantially the same functions, unless the following description is different. Multiplexer 224 includes a plurality of inputs from data bus 202 and a single output to memory bank 220. In at least some embodiments, multiplexer 224 is configured to select a data input connection from data bus 202 to memory bank 220. In at least some embodiments, multiplexer 224 is configured to respond to a selection command such as a signal from memory controller 226. In at least some embodiments, multiplexer 224 includes an external memory interface and inputs from data bus 202 to each computing unit within the integrated circuit. In at least some embodiments, each bank multiplexer 224 can be configured to connect to one of the computing units among the plurality of computing units or an external memory.
[0031] Memory controller 226 includes a circuit configured to operate memory unit 222. In at least some embodiments, memory controller 226 is configured to receive a signal from a general-purpose controller of the integrated circuit and operate memory unit 222 according to the received signal. In at least some embodiments, memory controller 226 locks memory bank 220 in response to a signal received from a computing unit and causes memory bank 220 to transmit a recorded value to one or more computing units. In at least some embodiments, memory controller 226 causes multiplexer 224 to connect memory bank 220 to a specific computing unit and causes memory bank 220 to record one or more values transmitted from the connected computing unit.
[0032] Data bus 202 has substantially the same structure as data bus 102 in FIG. 1 and performs substantially the same functions, except where the following description is different. Data bus 202 has a plurality of interconnects such as interconnect 203 that connect a computing unit, a memory bank, and an external memory interface. Interconnect 203 connects memory bank 220 to multiplexer 214. In at least some embodiments, interconnect 203 is part of a line that connects memory bank 220 to the multiplexers of all other computing units within the integrated circuit.
[0033] FIG. 3 is an operation flow for neural network hardware accelerated inference according to at least one embodiment of the present invention. This operation flow may provide a method for neural network hardware acceleration data parallel processing. In at least some embodiments, this method is executed by a controller that includes a section for performing a specific operation, such as the calculation controller 216 shown in FIG. 2 or the general-purpose controller 104 shown in FIG. 1.
[0034] In S330, the receiving section receives an instruction for performing a mathematical operation on a value. In at least some embodiments, the instruction includes the position of each value, and each position is identified by a memory bank identifier and an address within the memory bank. In at least some embodiments, the instruction indicates one or more mathematical operations to be performed. In at least some embodiments, the instruction includes one or more positions for recording the result value, and each position is identified by a memory bank identifier and an address within the memory bank. In at least some embodiments, the receiving section receives, by an integrated circuit, instructions for performing an inference of a neural network. In at least some embodiments, the instructions for performing the inference include a plurality of instructions for performing mathematical operations. In at least some embodiments, the receiving section receives, by a computing unit, instructions from a general-purpose controller based on the instructions for performing the inference.
[0035] In S332, the extraction section or a sub-section thereof extracts a data value. In at least some embodiments, the extraction section extracts a data value from one or more memory banks through a data bus. In at least some embodiments, the extraction section selects an input of a multiplexer to connect to a memory bank storing the data value. In at least some embodiments, the extraction section extracts an input data value from a memory bank among a plurality of memory banks provided in the integrated circuit, by a first computing unit among a plurality of computing units provided in the integrated circuit configured to perform a neural network inference. In at least some embodiments, the extraction section configures a first multiplexer corresponding to the first computing unit to connect to a memory bank storing the input data value. In at least some embodiments, the process of extracting the data value is performed as described below with respect to FIG. 8.
[0036] In S333, the extraction section or its sub-section extracts the weight value. In at least some embodiments, the extraction section extracts the weight value from one or more memory banks through a data bus. In at least some embodiments, the extraction section selects the input of a multiplexer to connect to the memory bank storing the weight value. In at least some embodiments, the extraction section extracts the weight value from the memory bank storing the weight value among the plurality of memory banks by a first calculation unit. In at least some embodiments, the extraction section configures a first multiplexer to connect to one memory bank storing the weight value. In at least some embodiments, the weight value extraction process is executed as described below with respect to FIG. 8.
[0037] In S335, the operation section or its sub-section performs a mathematical operation. In at least some embodiments, the operation section performs a mathematical operation on the data value retrieved in S332 and the weight value retrieved in S333. In at least some embodiments, the operation section causes a processor, such as the processor 212 in FIG. 2, to perform a mathematical operation. In at least some embodiments, the operation section performs a mathematical operation on the input data value and the weight value by a first calculation unit to generate a result data value. In at least some embodiments, the operation section performs a mathematical operation on the result data value of a previous iteration of the operation in S335 and the weight value retrieved in S333.
[0038] In S336, the controller or its sub-section determines whether all the operations within the instruction received in S330 have been executed. In at least some embodiments, the result value of the operation in S335 is subject to further operations before being recorded in the memory bank. If the controller determines that the instruction includes further operations, the operation flow returns to the extraction of the weight value in S333. If the controller determines that all the operations within the instruction have been executed, the operation flow proceeds to the transmission of the result value in S338.
[0039] In S338, the transmission section or its sub-section transmits the result value to one or more memory banks. In at least some embodiments, the transmission section transmits the result value to one or more memory banks through a data bus. In at least some embodiments, the transmission section causes the result value to be transmitted to the calculation unit. In at least some embodiments, the transmission section instructs the multiplexer of each corresponding memory unit of one or more memory banks to connect to the calculation unit for recording. In at least some embodiments, the result value transmission process is executed as described below with respect to FIG. 4.
[0040] FIG. 4 is an operation flow for the transmission of the result value according to at least one embodiment of the present invention. The operation flow can provide a method for transmitting the result value, such as the operation executed in S338 in FIG. 3. In at least some embodiments, the method is executed by the transmission section of a controller such as the calculation controller 216 shown in FIG. 2 or the general-purpose controller 104 shown in FIG. 1.
[0041] At S440, the transmission section or its sub-section determines whether the transmission involves recording one or more duplicates of the result value. In at least some embodiments, the instructions received by the controller indicate more than one memory location for a transmission involving duplicate recording. If the controller determines that the transmission does not involve duplicate recording, the operation flow proceeds to a single bank lock at S442. If the controller determines that the transmission involves duplicate recording, the operation flow proceeds to multiple bank locks at S445.
[0042] At S442, the transmission section or its sub-section locks a single memory bank in which the result value is to be recorded. In at least some embodiments, the transmission section sets a lock for a single memory bank. In at least some embodiments, the transmission section instructs the memory controller of the corresponding memory unit to lock the memory bank for recording from the computing unit.
[0043] At S443, the transmission section or its sub-section sets a bank offset for the single memory bank in which the result value is to be recorded. In at least some embodiments, the transmission section applies a bank offset for a single memory bank. In at least some embodiments, the instructions include one or more bank offsets, each bank offset corresponding to a memory bank within the integrated circuit. In response to a single memory bank being associated with a bank offset within the instructions, in at least some embodiments, the transmission section adjusts the address within the memory bank based on the bank offset such that the result value is recorded at the specified address.
[0044] In S445, the transmission section or a sub-section thereof locks a plurality of memory banks in which result values are to be recorded. In at least some embodiments, the transmission section sets a lock for each of at least two memory banks. In at least some embodiments, the transmission section instructs the memory controller of each corresponding memory unit to lock the memory bank for recording from the calculation unit.
[0045] In S446, the transmission section or a sub-section thereof sets bank offsets for a plurality of memory banks in which result values are to be recorded. In at least some embodiments, the transmission section applies a bank offset to one or more of at least two memory banks. In at least some embodiments, the instruction includes one or more bank offsets, and each bank offset corresponds to a memory bank within the integrated circuit. In response to one or more of the memory banks being associated with the bank offset in the instruction, in at least some embodiments, the transmission section adjusts the addresses within those memory banks based on the associated bank offset so that the result value is recorded at the specified address.
[0046] In S448, the transmission section or a sub-section thereof transmits the result value to one or more banks locked to receive the result value in S442 or S445. In at least some embodiments, the transmission section causes the calculation unit to transmit the result value. In at least some embodiments, the transmission section causes the first calculation unit to transmit the result data value to at least two of the plurality of memory banks such that a single transmission of the result data value by the first calculation unit is received by the at least two memory banks at substantially the same time. In at least some embodiments, the transmission section causes the first calculation unit to transmit the result data value to a single memory bank.
[0047] In S449, the transmission section or its sub-section releases the locks of one or more banks locked in S442 or S445. In at least some embodiments, the transmission section instructs the memory controller of each corresponding memory unit to release the lock of the memory bank. In at least some embodiments, the transmission section releases the locks for each of at least two memory banks.
[0048] FIG. 5 is a schematic diagram of a computing unit and memory bank interconnection during data parallel recording according to at least one embodiment of the present invention. This figure shows an integrated circuit 500 including a data bus 502, a computing unit 510, a memory bank 520A, and a memory bank 520B. The data bus 502, the computing unit 510, and the memory banks 520A and 520B have substantially the same structure as the data bus 102, the computing unit 110, and the memory bank 120 of FIG. 1 respectively, and perform substantially the same functions, except where the description is different below.
[0049] During data parallel recording, the result value is transmitted to multiple memory banks, as in the operation at S448 in FIG. 4 where multiple memory banks are specified to record the result value thereon. In at least some embodiments where the calculation unit 510 is instructed to record the result value in memory bank 520A and memory bank 520B during inference, the transmission section instructs a multiplexer of each corresponding memory unit of memory bank 520A and memory bank 520B to connect to the calculation unit 510 for recording. In at least some embodiments, the interconnection through data bus 502 shown in FIG. 5 routes the output from the calculation unit 510 to memory bank 520A and memory bank 520B. Thus, a single transmission from the calculation unit 510 is received by both memory bank 520A and memory bank 520B. In at least some embodiments, the transmission is received by memory bank 520A at substantially the same time as the same transmission is received by memory bank 520B. In at least some embodiments, the time difference in the reception of the same transmission between memory bank 520A and memory bank 520B is due to the difference in the physical distance between the calculation unit 510 and each respective memory bank. In at least some embodiments, the time difference in reception is negligible. In at least some embodiments, errors in the reception of the transmission due to the time difference in reception are resolved by shortening the clock time of the integrated circuit.
[0050] FIG. 6 is a graph of the operation and time of the calculation unit during reading from a common memory bank according to at least one embodiment of the present invention. The horizontal axis represents the calculation unit, while the vertical axis represents time.
[0051] At T0, the first computing unit 610A starts to execute a series of operations including 660A, 662A, 664A, 666A, and 668A. At 660A, the first computing unit 610A receives a first input data value from a memory bank storing the first input data value. At 662A, the first computing unit 610A locks a first memory bank storing weight values. At 664A, the first computing unit 610A receives weight values from the first memory bank. At 666A, the first computing unit 610A executes a calculation on the first input data value and the weight values. At 668A, the first computing unit 610A releases the first memory bank. At T2, a series of operations executed by the first computing unit 610A is completed.
[0052] At T1, the second computing unit 610B starts to execute a series of operations including 660B, 662B, 664B, 666B, and 668B. At 660B, the second computing unit 610B receives a second input data value from a memory bank storing the second input data value. In at least some embodiments, the operations at 660A and 660B are executed simultaneously, i.e., T0 = T1. In at least some embodiments, there is a delay between the execution of the operation at 660A and the execution of the operation at 660B. In at least some embodiments, when the second computing unit 610B receives the second input data value from the memory bank storing the second input data value, the second computing unit 610B then proceeds to retrieve the weight values. However, while the first memory bank storing the weight values is locked by the first computing unit 610A, the second computing unit 610B cannot retrieve the weight values from the first memory bank. In at least some embodiments, the second computing unit 610B waits until the first computing unit 610A releases the lock on the first memory bank and then retrieves the weight values from the first memory bank.
[0053] At T2, the second computing unit 610B continues to execute a series of operations at 662B. At 662B, the second computing unit 610B locks the first memory bank. At 664B, the second computing unit 610B receives weight values from the first memory bank. At 666B, the second computing unit 610B executes calculations on the second input data values and the weight values. At 668B, the second computing unit 610B releases the first memory bank. In at least some embodiments where two computing units execute operations on the same values stored in the same memory bank, the total amount of time used is approximately twice the amount of time used by one computing unit to execute the operation.
[0054] Figure 7 is a graph of the operation and time of a computing unit during a read from a separate memory bank, according to at least one embodiment of the present invention. The horizontal axis represents the computing unit, while the vertical axis represents time.
[0055] At T0, the first computing unit 710A begins to execute a series of operations including 760A, 762A, 764A, 766A, and 768A. At 760A, the first computing unit 710A receives a first input data value from the memory bank storing the first input data value. At 762A, the first computing unit 710A locks the first memory bank storing the weight values. At 764A, the first computing unit 710A receives weight values from the first memory bank. At 766A, the first computing unit 710A executes calculations on the first input data value and the weight values. At 768A, the first computing unit 710A releases the first memory bank. At T2, the series of operations executed by the first computing unit 710A is completed.
[0056] At T1, the second computing unit 710B starts to execute a series of operations including 760B, 762B, 764B, 766B, and 768B. At 760B, the second computing unit 710B receives a second input data value from a memory bank that stores the second input data value. In at least some embodiments, the operations at 760A and 760B are executed simultaneously, i.e., T0 = T1. In at least some embodiments, there is a delay between the execution of the operation at 760A and the execution of the operation at 760B. In at least some embodiments, when the second computing unit 710B receives the second input data value from the memory bank that stores the second input data value, the second computing unit 710B then proceeds to retrieve the weight value. In at least some embodiments, the first memory bank that stores the weight value is locked to the first computing unit 710A, but the second computing unit 710B retrieves the weight value from the second memory bank without waiting for the first computing unit 710A to release the lock on the first memory bank.
[0057] At 762B, the second computing unit 710B locks the second memory bank. At 764B, the second computing unit 710B receives weight values from the second memory bank. At 766B, the second computing unit 710B performs calculations on the second input data values and the weight values. At 768B, the second computing unit 710B releases the second memory bank. In at least some embodiments where two computing units perform operations on the same values stored in separate memory banks, the total amount of time used is minimal or less than the amount of time used by one computing unit to perform the operations. In at least some embodiments, data parallel recording in which the result values are transmitted to multiple memory banks, such as the operation at S448 in FIG. 4 where multiple memory banks are designated to record the result values thereon, enables two computing units to perform operations on the same values stored in separate memory banks, thereby reducing the overall time used by the integrated circuit to complete neural network inference. In at least some embodiments, data parallel processing reduces dark silicon.
[0058] FIG. 8 is an operation flow for value extraction according to at least one embodiment of the present invention. The operation flow may provide a method for extracting values, such as the operations performed at S332 or S333 in FIG. 3. In at least some embodiments, the method is executed by an extraction section of a controller such as the calculation controller 216 shown in FIG. 2 or the general-purpose controller 104 shown in FIG. 1.
[0059] At S870, the fetch section or its sub-section determines whether the fetch involves broadcasting so that the secondary calculation unit can read the transmission from the memory bank caused by the primary calculation unit fetching a value from the memory bank. In at least some embodiments, the instructions received by the controller indicate more than one calculation unit for a fetch involving duplicate reads. If the controller determines that the transmission does not involve duplicate reads, the operation flow proceeds to the bank lock at S874. If the controller determines that the transmission involves duplicate reads, the operation flow proceeds to the calculation unit synchronization at S872.
[0060] At S872, the fetch section or its sub-section synchronizes the calculation units. In at least some embodiments, the fetch section synchronizes the secondary calculation unit with the primary calculation unit. In at least some embodiments, the fetch section synchronizes a second calculation unit to read one of the input data value and the weight value. In at least some embodiments, the fetch section instructs the secondary calculation unit to prepare to read a value. In at least some embodiments, the fetch section applies a timing offset for the secondary calculation unit when the physical distance of the secondary calculation unit from the memory bank is significantly different from the physical distance of the primary calculation unit from the memory bank.
[0061] At S874, the fetch section or its sub-section locks the memory bank from which the value is to be read. In at least some embodiments, the fetch section sets a lock on the memory bank for both the primary calculation unit and the secondary calculation unit. In at least some embodiments, the transmission section instructs the memory controller of the corresponding memory unit to lock the memory bank for the fetch of the value by both the primary calculation unit and the secondary calculation unit.
[0062] In S875, the fetch section or a sub-section thereof sets a bank offset to a memory bank from which a value is to be read. In at least some embodiments, the fetch section applies a bank offset to a memory bank. In at least some embodiments, an instruction includes one or more bank offsets, and each bank offset corresponds to a memory bank within an integrated circuit. In response to a memory bank being associated with a bank offset within an instruction, in at least some embodiments, the fetch section adjusts an address within the memory bank based on the bank offset so that a value is read from a specified address.
[0063] In S877, the fetch section or a sub-section thereof reads a value from a memory bank. In at least some embodiments, the fetch section causes a primary calculation unit to read a value from a memory bank. In at least some embodiments, the fetch section causes a single transmission of one of an input data value or a weight value to be performed by a first calculation unit from a memory bank storing one of the input data value or the weight value such that the memory bank storing one of the input data value or the weight value is read by a first calculation unit and a second calculation unit at substantially the same time.
[0064] In S879, the fetch section or a sub-section thereof releases the lock on the memory bank locked in S874. In at least some embodiments, the fetch section instructs a memory controller of a corresponding memory unit to release the lock on the memory bank. In at least some embodiments, the fetch section releases the lock on the memory bank.
[0065] FIG. 9 is a schematic diagram of a computing unit and a memory bank interconnect during data parallel readout, according to at least one embodiment of the present invention. This figure shows an integrated circuit 900 including a data bus 902, a computing unit 910A, a computing unit 910B, and a memory bank 920. The data bus 902, the computing units 910A and 910B, and the memory bank 920 have substantially the same structure as the data bus 102, the computing unit 110, and the memory bank 120 of FIG. 1, respectively, and perform substantially the same functions, except where the description differs below.
[0066] During a data parallel readout process, values are read from a single memory bank by multiple computing units, as in the operation at S877 of FIG. 8 where multiple computing units are specified to read values. In at least some embodiments where the computing unit 910A is instructed to synchronize with the computing unit 910B to receive a transmission of a value from the memory bank 920 when the value is read by the computing unit 910A during inference, the transmission section instructs the multiplexers of each of the computing unit 910A and the computing unit 910B to connect to the memory bank 920 for reading. In at least some embodiments, the interconnection through the data bus 902 shown in FIG. 9 routes the output from the memory bank 920 to the computing unit 910A and the computing unit 910B. Thus, a single transmission from the memory bank 920 is received by both the computing unit 910A and the computing unit 910B. In at least some embodiments, the transmission is received by the computing unit 910A at substantially the same time as the same transmission is received by the computing unit 910B. In at least some embodiments, the time difference in the reception of the same transmission between the computing unit 910A and the computing unit 910B is due to the difference in the physical distance between the memory bank 920 and each computing unit. In at least some embodiments, the time difference in reception is negligible. In at least some embodiments, errors in the reception of the transmission due to the time difference in reception are resolved by shortening the clock time of the integrated circuit.
[0067] Figure 10A shows an exemplary configuration of a depth unit convolution processor 1012 according to an embodiment of the present invention. The depth unit convolution processor 1012 includes a queue 1012Q, a main sequencer 1012MS, a window sequencer 1012WS, an activation feeder 1012AF, a weighting feeder 1012WF, a pipeline controller 1012PC, a convolution pipeline 1012CP, an external accumulation logic 1012A, and an accumulation memory interface 1012AI.
[0068] The queue 1012Q receives and transmits instructions. The queue 1012Q may receive instructions from a calculation controller such as the calculation controller 216 of FIG. 2 and transmit the instructions to the main sequencer 1012MS. The queue 1012Q can be a FIFO memory or any other memory suitable for queuing instructions.
[0069] The main sequencer 1012MS sequences control parameters for convolution. The main sequencer 1012MS receives instructions from the queue 1012Q and may output the instructions to the window sequencer 1012WS. The main sequencer 1012MS divides a KHxKW convolution into smaller convolutions of size 1x<window> and prepares instructions for activation data and weight values according to the order of the input regions within the kernel. Here, <window> refers to an architectural parameter that determines the line buffer length.
[0070] The window sequencer 1012WS sequences control parameters for one 1x<window> convolution. The window sequencer 1012WS receives instructions from the main sequencer 1012MS and may output a data sequence of activation data according to the order of the input regions within the kernel to the activation feeder 1012AF, and a data sequence of weight values according to the order of the input regions within the kernel to the weighting feeder 1012WF.
[0071] The activation feeder 1012AF supplies the activation data accessed from the memory bank through the data memory interface 1012DI to the convolution pipeline 1012CP according to the activation data indicated in the data sequence from the window sequencer 1012S. The activation feeder 1012AF can read out sufficient activation data for 1x<window> calculation from the memory bank to the line buffer of the convolution pipeline 1012CP.
[0072] The weighting feeder 1012WF preloads the weight values accessed from the memory bank through the weighting memory interface 1012WI to the convolution pipeline 1012CP according to the weight values indicated in the data sequence from the window sequencer 1012S. The weighting feeder 1012WF can read out sufficient weight values for 1x<window> calculation from the weighting memory into the weighting buffer of the convolution pipeline 1012CP.
[0073] The pipeline controller 1012PC controls the data transfer operation of the convolution pipeline 1012CP. When the current activation buffer content is processed, the pipeline controller 1012PC can start copying the data from the line buffer to the activation buffer of the convolution pipeline 1012CP. The pipeline controller 1012PC may control the convolution calculation performed by each channel pipeline 1012CH of the convolution pipeline 1012CP, where each channel pipeline 1012CH operates on one channel of the input to the depthwise convolution layer.
[0074] The convolutional pipeline 1012CP performs mathematical operations on the activation data supplied from the activation feeder 1012AF and the weight values preloaded from the weighting feeder 1012WF. The convolutional pipeline 1012CP is divided into channel pipelines 1012CH, and each channel pipeline 1012CH performs mathematical operations for one channel. In combination with the activation feeder 1012AF, the weighting feeder 1012WF, and the pipeline controller 1012PC, the convolutional pipeline logically performs convolutional calculations.
[0075] The external accumulation logic 1012A receives data from the convolutional pipeline 1012CP and stores the data in the memory bank through the accumulation memory interface 1012AI. The accumulation logic 1012A includes an adder 1012P for each channel pipeline 1012CH. The accumulation logic 1012A can be used for the per-point sum of the results of 1x<window> convolution by the contents of the memory bank.
[0076] In this embodiment, there are three channels exemplified by three window pipelines. However, other embodiments may have a different number of channels. Although possible, this embodiment shows three channels mainly for simplicity. Many embodiments include at least 16 channels to accommodate realistic applications.
[0077] FIG. 10B shows an exemplary configuration of a per-channel pipeline for a depthwise convolutional processor according to an embodiment of the present invention. The channel pipeline 1012CH includes a line buffer 1012LB, an activation buffer 1012AB, a weighting buffer 1012WB, a plurality of multipliers 1012X, a plurality of adders 1012P, a delay register 1012DR, and an internal accumulation register 1012NB.
[0078] The line buffer 1012LB stores the activation data received from the activation feeder 1012AF. The line buffer 1012LB may include a shift register that stores the activation data read out by the activation feeder 1012AF at one pixel per cycle.
[0079] The activation buffer 1012AB stores the activation data received from the line buffer 1012LB. The activation buffer 1012AB may include a set of registers that store the activation data to which the current convolution calculation is applied.
[0080] The weighting buffer 1012WB stores the weight values received from the weighting feeder 1012WF. The weighting buffer 1012WB may include a shift register that stores the weight values to which the current convolution calculation is applied.
[0081] The multiplier 1012X multiplies the activation data from the activation buffer 1012AB by the weight values from the weighting buffer 1012WB. In this embodiment, there are three multipliers 1012X, which means that the parallelism of the width or height dimension of the convolution kernel is 3. The adder 1012P that collectively forms an adder tree then adds together the product of the activation data and the weight values. During this process, the delay register 1012DR, which is also considered part of the adder tree, balances the adder tree. The internal accumulation register 1012IA assists in the addition by storing partial sums. For example, the internal accumulation register 1012IA may be used for the accumulation of partial sums when the number of buffer windows, which is six in this embodiment, and the width or height of the convolution filter are greater than the parallelism of 3.
[0082] When all the products are added together as a sum, the sum is output to the accumulation logic 1012A, which then stores the data in the memory bank through the accumulation memory interface 1012AI.
[0083] FIG. 11 shows an exemplary configuration of a dot-unit convolution processor according to an embodiment of the present invention. The dot-unit convolution processor 1112 includes a queue 1112Q, a main sequencer 1112S, a weighting memory interface 1112WI, a weighting feeder 1112WF, a weighting memory interface 1112WI, an activation feeder 1112AF, a data memory interface 1112DI, a systolic array 1112S, an accumulation logic 1112A, and an accumulation memory interface 1112AI.
[0084] The queue 1112Q receives and transmits instructions. The queue 1112Q may receive instructions from a calculation controller such as the calculation controller 216 of FIG. 2 and transmit the instructions to the main sequencer 1112S. The queue 1112Q may be a FIFO memory or any other memory suitable for queuing instructions.
[0085] The main sequencer 1112S sequences control parameters for convolution. The main sequencer 1112S may receive instructions from the queue 1112Q and output the control sequences to the weighting feeder 1112WF and the activation feeder 1112AF through the queues, respectively. In this embodiment, the main sequencer 1112S divides the KHxKW convolution into a sequence of 1x1 convolutions, which are supplied to the weighting feeder 1112WF and the activation feeder 1112AF as control parameters.
[0086] The weighting feeder 1112WF preloads the weight values accessed from the memory bank through the weighting memory interface 1112WI into the systolic array 1112SA according to the activation data indicated in the control parameters from the main sequencer 1112S.
[0087] The activation feeder 1112AF supplies the activation data accessed from the memory bank through the data memory interface 1112DI to the systolic array 1112SA according to the activation data indicated in the data sequence from the main sequencer 1112S.
[0088] The systolic array 1112SA includes a plurality of MAC elements 1112M. Each MAC element 1112M is preloaded with a weight value from a weighting feeder 1112WF before the calculation starts, and then receives an activation value from an activation feeder 1112F. A plurality of weighting buffers may be used to enable overlap of the calculation and the preloading of the weight values. The MAC elements 1112M are arranged in an array such that the product of the activation value and the weight output from a preceding MAC element 1112M is input to a subsequent MAC element 1112M. In this embodiment, for every cycle, each MAC element 1112M outputs an accumulated value equal to the value obtained by multiplying the weight value 1112W preloaded to the value output from the MAC element 1112M adjacent to its left, and the product is added to the value output from the MAC element 1112M adjacent to its upper. The MAC elements 1112M in the lowermost row output those products to the accumulation logic 1112A.
[0089] The accumulation logic 1112A receives the product from the systolic array 1112SA and stores the product in a memory bank. The accumulation logic 1112A includes an adder 1112P for each MAC element 1112M in the lowermost row. In this embodiment, when the accumulation required by the main sequencer 1112S reads the old value in the memory location where the writing will occur, the accumulation logic 1112A overwrites it with the sum with the new value. Otherwise, the accumulation logic 1112A writes the new value as it is.
[0090] The point-wise convolution module 1112 may be useful in performing point-wise convolution by dividing a single KHxKW convolution into a plurality of KHxKW 1x1 convolution partitions. For example, in the area of the memory bank corresponding to four different 1x1 convolutions, a 2x2 convolution may be replaced. The point-wise convolution module 1112 may calculate each 1x1 convolution as the dot product of the matrix of activation values in the MAC element and the matrix of weight values in the MAC element, and then sum the results of the 1x1 convolutions.
[0091] The convolutional processors of FIGS. 10A, 10B, and 11 are implemented in at least some embodiments configured to perform inference of a convolutional network. Other processors are used in at least some other embodiments configured to perform inference of other types of neural networks, including other types of deep networks.
[0092] At least some embodiments are described with reference to flowcharts and block diagrams that represent (1) steps of a process in which operations are performed, or (2) sections of a controller responsible for performing the operations. In at least some embodiments, certain steps and sections are implemented by dedicated circuits, programmable circuits supplied with computer-readable instructions stored on a computer-readable medium, and / or processors supplied with computer-readable instructions stored on a computer-readable medium. In at least some embodiments, the dedicated circuits include digital and / or analog hardware circuits, including integrated circuits (ICs) and / or discrete circuits. In at least some embodiments, the programmable circuits include reconfigurable hardware circuits that include logical AND, OR, XOR, NAND, NOR, and other logical operations, flip-flops, registers, memory elements, etc., such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.
[0093] In at least some embodiments, a computer-readable storage medium includes a tangible device that can maintain and store instructions for use by an instruction execution device. In some embodiments, a computer-readable storage medium includes, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other electromagnetic wave propagating freely, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0094] In at least some embodiments, the computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. In at least some embodiments, the network includes copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. In at least some embodiments, a network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.
[0095] In at least some embodiments, the computer-readable program instructions for performing the operations described above are any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk® or C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages, written in either source code or object code. In at least some embodiments, the computer-readable program instructions are executed entirely on a user's computer, executed partially on a user's computer as a stand-alone software package, executed partially on a user's computer and partially on a remote computer, or executed entirely on a remote computer or server. In at least some embodiments, in the latter scenario, the remote computer is connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection is made to an external computer (e.g., through the Internet using an Internet service provider). In at least some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), executes the computer-readable program instructions by utilizing state information of the computer-readable program instructions that configure the electronic circuit to perform aspects of the present invention.
[0096] While embodiments of the present invention have been described, the technical scope of any subject matter recited in the claims is not limited to the embodiments described above. Those skilled in the art will understand that various modifications and improvements to the embodiments described above are possible. Those skilled in the art will also understand that embodiments with such modifications or improvements added from the scope recited in the claims are included in the technical scope of the present invention.
[0097] The operations, procedures, steps, and stages of each process executed by the devices, systems, programs, and methods shown in the claims, embodiments, or figures may be executed in any order as long as the order is not indicated by "preceding" or "before" etc., and as long as the output from the previous process is not used in the subsequent process. Even when a process flow is described using phrases such as "first" or "next" in the claims, embodiments, or figures, such a description does not necessarily mean that the process must be executed in the described order.
[0098] Neural network hardware acceleration data parallel processing is executed by an integrated circuit having a plurality of memory banks, each memory bank of the plurality of memory banks being configured to store values and transmit the stored values, a plurality of calculation units, each calculation unit of the plurality of calculation units being configured to perform a mathematical operation on input data values and weight values to generate result data values, and a calculation controller configured to cause the transmission of values to be received by more than one calculation unit or bank.
[0099] The foregoing has outlined features of several embodiments so that those skilled in the art may better understand the aspects of the present disclosure. Those skilled in the art should appreciate that the present disclosure is readily available as a basis for designing or modifying other processes and structures for carrying out the same purposes and / or achieving the same advantages as the embodiments introduced herein. Those skilled in the art should also recognize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that various changes, substitutions, and alterations may be made therein without departing from the spirit and scope of the present disclosure.
Claims
1. A plurality of memory banks, each of the plurality of memory banks being configured to store a value and transmit the stored value; A plurality of computing units; and A plurality of interconnects connecting each of the plurality of computing units to each of the plurality of memory banks; Comprising: Each of the plurality of computing units: A processor including a circuit configured to perform a mathematical operation on an input data value and a weight value to generate a result data value; and A calculation controller: Receiving the input data value from any one of the plurality of memory banks Receiving the weight value from any one of the plurality of memory banks Causing the processor to perform the mathematical operation, and Transmitting the result data value to at least two of the plurality of memory banks such that a single transmission of the result data value is received by the at least two memory banks at substantially the same time A calculation controller configured as such Having The plurality of interconnects route an output from a single computing unit to the at least two memory banks to facilitate the single transmission of the result data value from the single computing unit to the at least two memory banks, An integrated circuit.
2. The integrated circuit according to claim 1, further comprising a controller configured to receive instructions for performing neural network inference.
3. The integrated circuit according to claim 1 or 2, wherein the calculation controller is further configured to set a lock for each of the at least two memory banks.
4. The integrated circuit according to claim 1 or 2, wherein the calculation controller is further configured to apply a bank offset to one or more of the at least two memory banks.
5. The integrated circuit according to claim 1 or 2, wherein the calculation controller is further configured to release the lock for each of the at least two memory banks.
6. The calculation controller of the first computing unit among the plurality of computing units Synchronizes a second computing unit among the plurality of computing units to receive one of the input data value or the weight value, and The memory bank that stores the one of the input data value or the weight value is configured to read out a single transmission of the one of the input data value or the weight value that will be read out by the calculation controller and the second calculation unit at substantially the same time, and reads out the one of the input data value or the weight value from the memory bank that stores the one of the input data value or the weight value among the plurality of memory banks. and is further configured as follows: The plurality of interconnects route the output from the memory bank that stores the one of the input data value or the weight value among the plurality of memory banks to the first calculation unit and the second calculation unit, and promote the single transmission of the one of the input data value or the weight value from the memory bank that stores the one of the input data value or the weight value among the plurality of memory banks to the first calculation unit and the second calculation unit. The integrated circuit according to claim 1 or 2.
7. The integrated circuit according to claim 1 or 2, wherein the processor is configured to perform point-wise convolution or depth-wise convolution.
8. The integrated circuit according to claim 1 or 2, wherein each memory bank among the plurality of memory banks is configured to store a value received through a corresponding bank multiplexer.
9. The integrated circuit according to claim 8, wherein the bank multiplexer is configurable to connect to a calculation unit or an external memory among the plurality of calculation units.
10. The integrated circuit according to claim 1 or 2, wherein each calculation unit among the plurality of calculation units further has a calculation multiplexer configurable to connect to one of the plurality of memory banks.
11. A first calculation unit among a plurality of calculation units provided in an integrated circuit configured to perform neural network inference extracts the input data value from a memory bank that stores the input data value among the plurality of memory banks provided in the integrated circuit; The first calculation unit extracts the weight value from a memory bank that stores the weight value among the plurality of memory banks. The first computing unit executes a mathematical operation on the input data value and the weight value to generate a result data value; and The first computing unit performs a single transmission of the result data value such that the result data value is received by at least two of the plurality of memory banks at substantially the same time, and the first computing unit transmits the result data value to the at least two memory banks comprising: A plurality of interconnects route the output from the first computing unit to the at least two memory banks to facilitate the single transmission of the result data value from the first computing unit to the at least two memory banks, a method. **Claim 12** The method according to claim 11, further comprising the integrated circuit receiving instructions for performing neural network inference. **Claim 13** The step of retrieving the input data value includes configuring a first multiplexer corresponding to the first computing unit to connect to the memory bank storing the input data value, and the step of retrieving the weight value includes configuring the first multiplexer to connect to the memory bank storing the weight value. The method according to claim 11 or 12. **Claim 14** The step of transmitting the result data value includes setting a lock for each of the at least two memory banks. The method according to claim 11 or 12. **Claim 15** The step of transmitting the result data value includes applying a bank offset to one or more of the at least two memory banks. The method according to claim 11 or 12. **Claim 16** The step of transmitting the result data value includes releasing the lock for each of the at least two memory banks. The method according to claim 11 or 12. **Claim 17** The step of retrieving one of the input data value or the weight value includes synchronizing a second computing unit to read out the one of the input data value or the weight value, and The first computing unit reads out one of the input data values or the weight values from the memory bank that stores one of the input data values or the weight values, such that the memory bank that stores one of the input data values or the weight values is read out by the first computing unit and the second computing unit at substantially the same time to perform a single transmission of one of the input data values or the weight values. It has The plurality of interconnects route the output from the memory bank that stores one of the input data values or the weight values among the plurality of memory banks to the first computing unit and the second computing unit, and facilitate the single transmission of one of the input data values or the weight values from the memory bank that stores one of the input data values or the weight values among the plurality of memory banks to the first computing unit and the second computing unit. The method according to claim 11 or 12.
18. A plurality of memory banks, each memory bank among the plurality of memory banks is configured to store a value and transmit the stored value; A plurality of computing units, each computing unit among the plurality of computing units Includes a processor having a circuit configured to perform a mathematical operation on an input data value and a weight value to generate a result data value, and A calculation controller Having; and A plurality of interconnects connecting each computing unit among the plurality of computing units to each memory bank among the plurality of memory banks; Comprising The calculation controller of the first computing unit among the plurality of computing units Synchronizes the second computing unit among the plurality of computing units to receive one of the input data values or the weight values, and The memory bank that stores the one of the input data value or the weight value is configured to read out the single transmission of the one of the input data value or the weight value that will be read out by the calculation controller and the second calculation unit at substantially the same time. Read out the one of the input data value or the weight value from the memory bank among the plurality of memory banks that stores the one of the input data value or the weight value. It is configured as follows. The plurality of interconnects route the output from the memory bank among the plurality of memory banks that stores one of the input data value or the weight value to the first calculation unit and the second calculation unit, and the plurality of memory banks The single transmission of one of the input data value or the weight value from the memory bank that stores one of the input data value or the weight value to the first calculation unit and the second calculation unit is promoted. Integrated circuit.
19. The calculation controller is Receives the input data value from any one of the plurality of memory banks Receives the weight value from any one of the plurality of memory banks Causes the processor to execute the mathematical operation, and Transmits the result data value to the at least two memory banks so that a single transmission of the result data value is received by the at least two memory banks among the plurality of memory banks at substantially the same time. The integrated circuit according to claim 18, further configured as follows.
20. The integrated circuit according to claim 18 or 19, wherein the calculation controller is further configured to set a lock for each of at least two memory banks among the plurality of memory banks.
Citation Information
Patent Citations
Neural network training under memory restraint
CN113469354A
Gradient compression for distributed training
DE102021107050A1
On-chip computational networks
JP2021506032A
Neural network accelerator run-time reconfigurability
JP2022105467A
Neural network accelerator run-time reconfigurability
US11144822B1