Data calculation system
The data calculation system addresses inefficiencies in conventional computing by using an accelerator to perform calculations independently, thereby enhancing processor efficiency and reducing overhead.
Patent Information
- Application Number
- JP2023188892
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-03-21
- Filing Date
- 2023-11-02
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2039-03-06
AI Technical Summary
Conventional computing methods are inefficient due to the processor's high demand for bus resources during calculations, which adversely affects execution efficiency.
A data calculation system that includes a memory, a processor, and an accelerator, where the processor controls the accelerator, and the accelerator performs data calculations independently, reducing the processor's reliance on bus resources.
This solution improves processor execution efficiency, reduces calculation overhead, and shortens data calculation time by allowing the processor to handle other tasks while the accelerator performs calculations.
Smart Images

Figure 0007696408000001 
Figure 0007696408000002 
Figure 0007696408000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications [1] This disclosure claims the benefit of priority of Chinese Patent Application No. 201810235312.9, filed on March 21, 2018, which is incorporated herein by reference in its entirety.
Background Art
[0002] Background [2] With the development of artificial intelligence (AI) technology, computing power and computing speed play an essential role in the field of AI. The conventional implementation method of computing is as follows: The processor accesses the memory via the bus to read data, performs calculations to obtain results, and then writes the calculation results back to the memory via the bus. One problem with the conventional implementation method of computing is that the processor occupies a large amount of bus resources because the processor needs to continuously access the memory during calculations. The execution efficiency of the processor is adversely affected.
Summary of the Invention
Means for Solving the Problems
[0003] Summary of the Disclosure [3] This disclosure provides a data calculation system including a memory, a processor, and an accelerator. The memory is communicatively coupled to the processor and configured to store calculation data, and the data is written by the processor. The processor is communicatively coupled to the accelerator and configured to control the accelerator. The accelerator is communicatively coupled to the memory and configured to access the memory according to pre - configured control information, perform data calculations, and write the calculation results back into the memory. This disclosure also provides an accelerator of the data calculation system and a method executed by the accelerator.
Brief Description of the Drawings
[0004] Brief Description of the Drawings
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0005] Detailed Description
[10] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described below with reference to the accompanying drawings in the embodiments of the present disclosure. Of course, the described embodiments are only a part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art without creative efforts based on the embodiments of the present disclosure shall be included in the protection scope of the present disclosure.
[0006]
[11] The data calculation system shown in this disclosure can improve the execution efficiency of the processor and reduce the calculation overhead of the processor. The data calculation system shown in some embodiments of this disclosure adds an accelerator based on the original memory and processor. The processor uses a bus to control the accelerator, and the accelerator accesses the memory, executes the data calculation, and writes the calculation result back to the memory after the calculation is completed. Compared with the conventional structure, the processor controls the accelerator, and the specific data calculation is completed by the accelerator. The calculation process is independently executed by the accelerator and does not occupy the calculation unit or bus resources of the processor. While the accelerator is executing the calculation process, the processor can process other events without being affected by the calculation performance of the accelerator. Therefore, the execution efficiency of the processor is improved, the calculation overhead of the processor is reduced, and the time spent on data calculation is also shortened.
[0007]
[12] Embodiments of this disclosure provide a data calculation system. FIG. 1 is a schematic diagram of an exemplary data calculation system according to some embodiments of this disclosure. As shown in FIG. 1, the data calculation system includes a memory 11, a processor 12, and an accelerator 13.
[0008]
[13] The memory 11 is communicatively coupled to the processor 12 and is configured to store calculation data. All calculation data is written into the memory 11 by the processor 12.
[0009]
[14] The processor 12 is communicatively coupled to the accelerator 13 and is configured to control the accelerator 13.
[0010]
[15] The accelerator 13 is communicatively coupled to the memory 11 and is configured to access the memory 11 according to preconfigured control information, perform a data calculation process, and write the calculation result back to the memory 11.
[0011]
[16] In some embodiments, when executing a data calculation process, the processor 12 of the data calculation system shown in the embodiments of the present disclosure controls the accelerator 13 but does not execute the data calculation process. The data calculation process is completed by the accelerator 13. Therefore, during the calculation process, the processor 12 does not need to access the memory 11, and thus does not occupy the bus, thereby improving the bus utilization rate. At the same time, when the accelerator 13 executes the calculation of data, the processor 12 can process other events, thus also improving the utilization rate of the processor. In addition, the accelerator 13 can be coupled to any type of memory for calculation.
[0012]
[17] In addition, experimental verification of the wake-on-voice algorithm can be performed using some embodiments of the present disclosure. Experimental data results show that in a conventional system, it is necessary to maintain a processing speed of 196 MCPS (Millions of Cycles Per Second) with the same wake-on-voice algorithm. In the present disclosure, using the accelerator 13, the processing speed can reach 90 MCPS. The performance can be improved by about 55%.
[0013]
[18] FIG. 2 is a schematic diagram of an exemplary accelerator according to some embodiments of the present disclosure. As shown in FIG. 2, the accelerator 13 includes a control register module 131 and a calculation module 132.
[0014]
[19] The control register module 131 is communicatively coupled to the processor 12 and is configured to store control information. The control information is pre-configured by the processor 12 using a bus for delivering instructions.
[0015]
[20] The calculation module 132 is communicatively coupled to the memory 11 and is configured to access the memory 11 according to the control information, perform data calculation, and write the calculation result back to the memory 11.
[0016]
[21] In some embodiments, the control information includes a start address for calculation data, the number of operands, a calculation method, a write-back address for the calculation result, and a calculation enable flag. The calculation methods include multiplication-accumulation operations, exponential functions, sigmoid functions, rectified linear unit (ReLU) functions, and softmax functions. In other words, the calculation module 132 can implement multiplication-accumulation operations, exponential functions, sigmoid functions, rectified linear unit (ReLU) functions, and softmax functions. However, the calculation methods of the present disclosure are not limited to the above several types, and the calculation methods can be customized according to the actual requirements of software applications. During execution, the calculation method can be controlled by the processor so that its use is flexible and convenient. Furthermore, in order to achieve the effect of optimizing the chip area, the hardware implementation of the calculation method can be added or deleted according to actual requirements.
[0017]
[22] After detecting that the calculation enable flag is 1, the calculation module 132 sequentially reads the calculation data and the number of operands from the memory 11 according to the start address for the calculation data, executes the calculation of the data according to the calculation method, and writes the calculation result back to the memory 11 according to the write-back address for the calculation result. At the same time, the calculation module 132 resets the calculation enable flag. After reading that the calculation enable flag is 0, the processor 12 can start the next calculation or read the calculation result from the memory 11.
[0018]
[23] FIG. 3 is a schematic diagram of an exemplary calculation module according to some embodiments of the present disclosure. As shown in FIG. 3, when implementing a multiplication-accumulation operation or a rectified linear unit (ReLU) function, the calculation module 132 includes a multiplication-accumulation unit 1321, a rectified linear unit (ReLU) calculation unit 1322, a first multiplexer 1323, and a second multiplexer 1324.
[0019]
[24] The multiply-accumulate unit 1321 includes a 2-channel 16-bit multiplier 13211, a 2-channel 8-bit multiplier 13212, an accumulator 13214, and a register 13213. The multiply-accumulate unit 1321 is configured to perform parallel calculations using the 2-channel 16-bit multiplier 13211 and the 2-channel 8-bit multiplier 13212, and store the multiply-accumulated calculation result in the register 13213.
[0020]
[25] The rectified linear unit (ReLU) calculation unit 1322 is configured to perform the calculation of the rectified linear unit (ReLU) function on the input data 1320 or the multiply-accumulated calculation result from the multiply-accumulate unit 1321.
[0021]
[26] The first multiplexer 1323 is configured to select, according to the ReLU_bypass signal, the multiply-accumulated calculation result from the multiply-accumulate unit 1321 or the input data 1320 as the data input to the rectified linear unit (ReLU) calculation unit 1322.
[0022]
[27] The second multiplexer 1324 is configured to select, according to the ReLU_bypass signal, whether to perform the calculation of the rectified linear unit (ReLU) function on the multiply-accumulated calculation result from the multiply-accumulate unit 1321.
[0023]
[28] FIG. 4 is a diagram of an exemplary 32-channel 8x8 data storage format and calculation process according to some embodiments of the present disclosure. To perform the 32-channel 8x8 multiply-accumulate calculation, the calculation process is shown below according to FIG. 4.
[0024]
[29] The processor 12 writes the data A and the data B into the memory 11 via the bus, and the data is not written until the subsequent calculation process is completed. If it is necessary to replace the calculation data after the completion of the calculation process, the processor 12 rewrites the calculation data.
[0025]
[30] After the calculation data is written into the memory 11, the processor 12 configures the control register module 131 of the accelerator 13, the start address for data A (DATA0_Start_addr), the start address for data B (DATA1_Start_addr), and the write-back address for the calculation result (Result_wb_addr).
[0026]
[31] Next, the processor 12 configures the calculation method to be a 32-channel 8x8 multiply-accumulate calculation (for example, the calculation in FIG. 4), sets the number of operands to 32, and sets the calculation enable flag to 1.
[0027]
[32] After detecting that the calculation enable flag is 1, the calculation module 132 of the accelerator 13 starts the calculation process, reads the calculation data from the memory 11 according to the start address for data A (DATA0_Start_addr), the start address for data B (DATA1_Start_addr), and the number of operands, and performs the multiply-accumulate calculation.
[0028]
[33] After the calculation is completed, the calculation result is written back into the memory 11 according to the write-back address (Result_wb_addr), and the calculation enable flag is reset.
[0029]
[34] After reading that the calculation enable flag is 0, the processor 12 can start the next calculation process or read the calculation result from the memory 11.
[0030]
[35] FIG. 5 is a diagram of an exemplary 4-channel 16x16 multiply-accumulate data storage format and calculation process according to some embodiments of the present disclosure. To perform the 4-channel 16x16 multiply-accumulate calculation, the calculation process is shown below according to FIG. 5.
[0031]
[36] The processor 12 writes data A and data B into the memory 11 via the bus, and no data is written until the subsequent calculation process is completed. If it is necessary to replace the calculation data after the completion of the calculation process, the processor 12 rewrites the calculation data.
[0032]
[37] After the data is written into the memory 11, the processor 12 configures the control register module 131 of the accelerator 13, the start address for data A (DATA0_Start_addr), the start address for data B (DATA1_Start_addr), and the write-back address for the calculation result (Result_wb_addr).
[0033]
[38] Next, the processor 12 configures the calculation method to be a 4-channel 16x16 multiplication and accumulation calculation (for example, the calculation in FIG. 5), sets the number of operands to 4, and sets the calculation enable flag to 1.
[0034]
[39] After detecting that the calculation enable flag is 1, the calculation module 132 of the accelerator 13 starts the calculation process, reads the calculation data from the memory 11 according to the start address for data A (DATA0_Start_addr), the start address for data B (DATA1_Start_addr), and the number of operands, and performs the multiplication and accumulation calculation.
[0035]
[40] After the calculation is completed, the calculation result is written back into the memory 11 according to the write-back address (Result_wb_addr), and the calculation enable flag is reset.
[0036]
[41] After reading that the calculation enable flag is 0, the processor 12 can start the next calculation process or read the calculation result from the memory 11.
[0037]
[42] FIG. 6 is a diagram of an exemplary data storage format and calculation process for an exponential function, a softmax function, and a sigmoid function according to some embodiments of the present disclosure. The calculation process of FIG. 6 is the same as the multiply-accumulate calculation process of FIGS. 4 and 5. Instead of the multiply-accumulate calculation, the processor 12 configures the calculation method to be an exponential function, a softmax function, or a sigmoid function.
[0038]
[43] Although some specific embodiments of the present disclosure have been described above, the protection scope of the present disclosure is not limited to those embodiments. Any changes or substitutions that can be easily devised by those skilled in the art within the technical scope disclosed by the present disclosure shall be included in the protection scope of the present disclosure. Therefore, the protection scope for protecting the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A memory configured to store calculation data, A processor communicably coupled to the memory and configured to write the calculation data to the memory, An accelerator communicably coupled to the memory and the processor, receiving control information from the processor, accessing the memory according to the control information, performing a calculation process that yields a calculation result, and configured to write the calculation result back to the memory comprising the calculation process being executed independently by the accelerator from the processor, the control information including a start address for the calculation data, the number of operands, a calculation method, a write-back address for the calculation result, and a calculation enable flag, after detecting that the calculation enable flag is enabled, the accelerator reads the calculation data from the memory according to the start address and the number of operands, performs the calculation process according to the calculation method, and is configured to write the calculation result back to the memory according to the write-back address, the calculation data stored in the memory not being updated during the calculation process, a data calculation system.
2. The accelerator is a control register module communicably coupled to the processor and configured to store the control information including instructions, a calculation module communicably coupled to the memory, accessing the memory according to the control information, performing the calculation process, and configured to write the calculation result back to the memory The data calculation system according to claim 1, comprising.
3. The data calculation system according to claim 2, wherein the control information is stored in the control register module.
4. The data calculation system according to claim 1, wherein the calculation method includes one of a multiplication and accumulation operation, an exponential function, a sigmoid function, a normalized linear function, or a softmax function.
5. The data calculation system according to claim 2, wherein the calculation module is configured to reset the calculation enable flag after the calculation process is completed.
6. The calculation module a multiplication and accumulation unit configured to perform a multiplication and accumulation operation to generate a result The data calculation system according to claim 2, comprising:
7. The calculation module a normalized linear calculation unit configured to perform a normalized linear function on the input data or the result from the multiplication and accumulation unit, and a first multiplexer configured to select the result from the multiplication and accumulation unit or the input data as data input to the normalized linear calculation unit The data calculation system according to claim 6, comprising:
8. The calculation module a second multiplexer configured to select the result from the multiplication and accumulation unit or the normalized linear calculation unit as the calculation result The data calculation system according to claim 7, comprising:
9. a control register module communicably coupled to an external processor and configured to receive control information from the external processor, and a calculation module communicably coupled to an external memory associated with the external processor, configured to access the external memory according to the control information, perform a calculation process to yield a calculation result, and write back the calculation result to the external memory comprising The control information includes a start address for calculation data, the number of operands, a calculation method, a write-back address for the calculation result, and a calculation enable flag. After detecting that the calculation enable flag is enabled, the calculation module is further configured to read the calculation data from the external memory according to the start address and the number of operands, perform the calculation process according to the calculation method, and write back the calculation result to the external memory according to the write-back address. An accelerator in which the calculation data stored in the memory is not updated during the calculation process.
10. The accelerator according to claim 9, wherein the calculation method includes one of a multiply-accumulate operation, an exponential function, a sigmoid function, a normalized linear function, or a softmax function.
11. The accelerator according to claim 9, wherein the calculation module is configured to reset the calculation enable flag after the calculation process is completed.
12. The calculation module A multiply-accumulate unit configured to execute a multiply-accumulate operation to generate a result The accelerator according to claim 9, including.
13. The calculation module A normalized linear calculation unit configured to execute a normalized linear function on input data or the result from the multiply-accumulate unit, and A first multiplexer configured to select the result from the multiply-accumulate unit or the input data as data input to the normalized linear calculation unit The accelerator according to claim 12, including.
14. The calculation module A second multiplexer configured to select the result from the multiplication accumulation unit or the normalization linear calculation unit as the calculation result The accelerator according to claim 13, comprising:
15. A data calculation method executed by an accelerator of a data calculation system, comprising: Receiving control information including a start address, the number of operands, a calculation method, and a write-back address for calculation data from a processor of the data calculation system by the accelerator of the data calculation system, wherein the accelerator is separated from the processor; Accessing, by the accelerator, a memory coupled to the processor according to the start address and the number of operands, wherein the accelerator is separated from the memory; Executing a calculation process on the calculation data according to the calculation method to obtain a calculation result, wherein the calculation process is independently executed by the accelerator from the processor, the calculation data stored in the memory is not updated during the calculation process, and Writing the calculation result to the memory according to the write-back address A data calculation method comprising:
16. The accelerator is A control register module communicatively coupled to the processor and configured to store the control information including instructions; A calculation module communicatively coupled to a memory, configured to access the memory according to the control information, perform the calculation process, and write back the calculation result to the memory The data calculation method according to claim 15, comprising:
17. The calculation module is A multiplication accumulation unit configured to execute a multiplication accumulation operation to generate a result The data calculation method according to claim 16, comprising
18. wherein the calculation module a normalization linear calculation unit configured to execute a normalization linear function on input data or the result from the multiplication and accumulation unit; a first multiplexer configured to select the result from the multiplication and accumulation unit or the input data as data input to the normalization linear calculation unit The data calculation method according to claim 17, comprising
Citation Information
Patent Citations
Digital signal processor
JP1998187599A
Composite arithmetic processor
JP2004046896A
Method for driving control device and control device with model calculation unit
JP2015015024A
Time division duplex (TDD) uplink downlink (UL-DL) reconfiguration
US20150003301A1
Deep processing unit (DPU) for implementing an artificial neural network (ANN)
US20180046903A1