Accelerator Based on Data Flow Architecture, Data Access Method and Device of Accelerator
By flexibly adjusting the read and write parallelism and calculation parallelism of the data flow architecture, the bandwidth and power consumption problems of traditional accelerators when data parallelism is not aligned, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202111642091.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-29
AI Technical Summary
Traditional accelerators based on data flow architecture need to fill zeros when data parallelism is not aligned, increasing data transmission bandwidth, storage time and power consumption, and low computing resource utilization.
Using flexible and variable preset read and write parallelism and calculation parallelism, the parallelism of the repository and data paths is dynamically adjusted through the read and write address generation unit and calculation unit, reducing data transmission and storage requirements, and reducing power consumption.
It effectively reduces the bandwidth requirements, data storage requirements and running time of the accelerator, while reducing power consumption and improving the utilization rate of computing resources.
Smart Images

Figure CN114327639B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of data processing, and in particular, to an accelerator based on a data flow architecture, a data access method and device for the accelerator. Background Art
[0002] Recent research has shown that compared with traditional feature extraction algorithms, neural network algorithms have great advantages in the field of computer vision. Neural networks have been widely used in fields such as image, speech, and video recognition. However, the computational and storage complexity of neural network algorithms has brought great difficulties to their applications. The CPU platform is difficult to provide sufficient computing power. The GPU platform is the preferred platform for neural network processing, with strong computing power and a simple and easy-to-use development framework. However, when the GPU processes neural networks, the utilization rate of computing resources is low, and the computing units are idle for a large part of the time. To improve the utilization rate of computing resources, an accelerator based on a data flow architecture has been proposed. In this architecture, data transmission and computing can be parallel, and different computing units can also execute in parallel.
[0003] To obtain greater computing power, the accelerator based on the data flow architecture needs to increase the parallelism of data fetching and computing. In traditional accelerators based on the data flow architecture, the parallelism of data fetching and computing is fixed. When the data is not aligned to the fixed parallelism, some zeros need to be filled to align the data to the fixed parallelism. After zero filling, the amount of data to be transmitted will increase, thereby increasing the data transmission bandwidth, data transmission time, and data storage time. In addition, storing and computing the filled zeros also increases the power consumption of the entire accelerator. Summary of the Invention
[0004] Embodiments of the present invention provide an accelerator based on a data flow architecture, a data access method and device for the accelerator, so as to reduce the bandwidth requirements, data storage requirements, and running time of the accelerator, and reduce power consumption.
[0005] In a first aspect, embodiments of the present invention provide an accelerator based on a data flow architecture, the accelerator including: a storage unit, a read / write address generation unit, and a computing unit; wherein,
[0006] The storage unit includes a plurality of storage repositories;
[0007] The read / write address generation unit is configured to generate a storage unit read / write address according to a preset read / write parallelism, so as to determine a target storage repository in the storage unit according to the storage unit read / write address, and read data to be processed from the target storage repository to the computing unit for operation;
[0008] The computing unit includes multiple data paths, which are used to determine a target data path according to a preset computing parallelism, so as to use the target data path to perform operations on the data to be processed to obtain processed data, and store the processed data into the target repository according to the read / write address of the storage unit.
[0009] Optionally, the number of repositories is an integer multiple of the preset read / write parallelism.
[0010] Optionally, the read / write address generation unit is further configured to generate an enable signal for the target repository, so as to enable the read / write of the target repository.
[0011] Optionally, the computing unit is further configured to generate an enable signal for the target data path, so as to enable the target data path.
[0012] Optionally, the preset read / write parallelism includes a preset read parallelism and a preset write parallelism;
[0013] Correspondingly, the read / write address of the storage unit includes a read address of the storage unit and a write address of the storage unit, and the target repository includes a target read repository and a target write repository;
[0014] The read / write address generation unit is specifically configured to generate the read address of the storage unit according to the preset read parallelism, and generate the write address of the storage unit according to the preset write parallelism, so as to determine the target read repository according to the read address of the storage unit, and read the data to be processed from the target read repository to the computing unit for operations;
[0015] The computing unit is specifically configured to use the target data path to perform operations on the data to be processed to obtain the processed data, and store the processed data into the target write repository according to the write address of the storage unit.
[0016] Optionally, the preset read parallelism, the preset write parallelism and the preset computing parallelism are the same, partially the same or different from each other.
[0017] In a second aspect, an embodiment of the present invention further provides a data access method for an accelerator, which is applied to the accelerator based on a data flow architecture provided in any embodiment of the present invention, and includes:
[0018] Determine an optimal read / write parallelism and an optimal computing parallelism according to the size of the data to be processed;
[0019] Configure the control register of the accelerator according to the optimal read / write parallelism and the optimal computing parallelism, where the configuration parameters of the control register include a preset read / write parallelism and a preset computing parallelism;
[0020] Read the data to be processed from the storage unit according to the configured preset read-write parallelism, and output it to the computing unit, so that the computing unit performs operations according to the configured preset computing parallelism;
[0021] Transfer the processed data obtained after the operation back to the storage unit according to the configured preset read-write parallelism.
[0022] Optionally, determining the optimal read-write parallelism and the optimal computing parallelism according to the size of the data to be processed includes:
[0023] Traverse all the configurable read-write parallelisms and computing parallelisms, and respectively determine the corresponding processing times;
[0024] Determine the read-write parallelism and the computing parallelism corresponding to the shortest processing time as the optimal read-write parallelism and the optimal computing parallelism respectively.
[0025] In a third aspect, an embodiment of the present invention further provides a computer device, which includes:
[0026] One or more processors;
[0027] A memory for storing one or more programs;
[0028] When the one or more programs are executed by the one or more processors, the one or more processors implement the data access method of the accelerator provided in any embodiment of the present invention.
[0029] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data access method of the accelerator provided in any embodiment of the present invention.
[0030] An embodiment of the present invention provides an accelerator based on a data flow architecture, including a storage unit, a read-write address generation unit, and a computing unit. The storage unit includes multiple storage repositories, and the computing unit includes multiple data paths. By using the read-write address generation unit to generate storage unit read-write addresses according to the preset read-write parallelism, and determining the target storage repository for reading the data to be processed according to the read-write addresses, then reading the data to be processed from the target storage repository to the computing unit, the computing unit determines the target data path for the operation according to the preset computing parallelism, and uses the target data path to perform operations on the data to be processed to obtain the processed data, and finally stores the processed data in the target storage repository according to the read-write addresses, thereby reducing the bandwidth requirements, data storage requirements, and running time of the accelerator, and also reducing the power consumption. Description of the Drawings
[0031] Figure 1 Schematic diagram of the accelerator based on the data flow architecture provided in the first embodiment of the present invention;
[0032] Figure 2 Flowchart of the data access method of the accelerator provided in the second embodiment of the present invention;
[0033] Figure 3 Schematic diagram of the computer device provided in the third embodiment of the present invention. Detailed implementation manners
[0034] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the convenience of description, only the parts related to the present invention are shown in the drawings, rather than all the structures.
[0035] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0036] Embodiment 1
[0037] Figure 1 Schematic diagram of the accelerator based on the data flow architecture provided in the first embodiment of the present invention. This embodiment is applicable to the situation of improving the utilization rate of computing resources when using a GPU to process neural networks. As Figure 1 shown, the accelerator includes: a storage unit 11, a read / write address generation unit 12, and a computing unit 13; wherein, the storage unit 11 includes a plurality of memory banks 111 ( Figure 1 taking 4 memory banks bank_1, bank_2, bank_3, and bank_4 as an example for illustration, but the number is not limited to 4 in this embodiment); the read / write address generation unit 12 is used to generate the read / write addresses of the storage unit according to a preset read / write parallelism, so as to determine the target memory bank in the storage unit 11 according to the read / write addresses of the storage unit, and read the data to be processed from the target memory bank to the computing unit 13 for operation; the computing unit 13 includes a plurality of data paths 131 ( Figure 1Taking four data paths, namely datapath_1, datapath_2, datapath_3, and datapath_4, as an example (but not limited to four in this embodiment), it is used to determine the target data path according to the pre-designed computing parallelism, perform operations on the data to be processed using the target data path to obtain the processed data, and store the processed data into the target repository according to the read / write address of the storage unit.
[0038] Specifically, the storage unit 11 is used to store the data to be processed and the processed data. To achieve configurable parallelism for reading data, the storage unit 11 can be divided into n storage banks, where n can be a positive integer greater than 1. The read / write address generation unit 12 can generate different read / write addresses of the storage unit according to different pre-set read / write parallelisms, so as to determine the target storage bank in the storage unit 11 that is actually used to read data, and read the data to be processed from the target storage bank to the computing unit 13 for operation. Exemplarily, when reading data, assuming that the storage unit 11 includes 16 storage banks (bank_1 to bank_16) and the pre-set read / write parallelism is 4, when the read / write address of the storage unit generated by the read / write address generation unit 12 is 0, bank_1 to bank_4 can be determined as the target storage banks, and the data is taken out from bank_1 to bank_4 and given to the computing unit 13. The computing unit 13 is used to perform operations on the data to be processed to obtain the processed data. To achieve configurable parallelism for computing, the computing unit 13 can be divided into m data paths (datapath), where m can be a positive integer greater than 1. The computing unit 13 can determine the target data path in the computing unit 13 that is actually used for operation according to the pre-designed computing parallelism. After obtaining the data to be processed taken out from the target storage bank, the target data path can be used to perform operations on the data to be processed. Exemplarily, when performing parallel computing on data, assuming that the computing unit 13 includes 8 data paths (datapath_1 to datapath_8) and the pre-designed computing parallelism is 4, datapath_1 to datapath_4 can be determined as the target data paths, and datapath_1 to datapath_4 can be used to perform operations on the data to be processed. Furthermore, the processed data can be stored into the target storage bank according to the determined read / write address of the storage unit above. Among them, both the pre-set read / write parallelism and the pre-designed computing parallelism are flexibly variable and do not need to be fixed, that is, the provided accelerator can adapt to various parallelisms. In addition, to prevent bank conflicts, if multiple data to be read are stored in one bank, multiple data cannot be taken out within one cycle. Optionally, the number of the storage banks is an integer multiple of the pre-set read / write parallelism, that is, it is ensured that the number of the storage banks can be divided evenly by the pre-set read / write parallelism.
[0039] Optionally, the preset read-write parallelism includes a preset read parallelism and a preset write parallelism; correspondingly, the read-write address of the storage unit includes a read address of the storage unit and a write address of the storage unit, and the target storage repository includes a target read storage repository and a target write storage repository; the read-write address generation unit 12 is specifically configured to generate the read address of the storage unit according to the preset read parallelism, and generate the write address of the storage unit according to the preset write parallelism, so as to determine the target read storage repository according to the read address of the storage unit, and read the data to be processed from the target read storage repository to the calculation unit 13 for operation; the calculation unit 13 is specifically configured to use the target data path to perform an operation on the data to be processed to obtain the processed data, and store the processed data into the target write storage repository according to the write address of the storage unit. Specifically, the read and write processes of the data can be performed separately, and the read address and write address of the storage unit are generated according to the preset read parallelism and the preset write parallelism respectively, so as to determine the target read storage repository and the target write storage repository respectively. Then, the calculation unit 13 can obtain the data to be processed from the target read storage repository for operation, and store the processed data obtained by the operation into the target write storage repository.
[0040] Further optionally, the preset read parallelism, the preset write parallelism, and the preset calculation parallelism are the same, partially the same, or different from each other. Specifically, the data to be processed read from one target read storage repository can be used by one or more target data paths, and the processed data obtained by the operation of multiple target data paths can also be stored in one or more target write storage repositories. That is, the number of target data paths used can be the same as or different from the number of target read storage repositories and the number of target write storage repositories. Therefore, the preset read parallelism, the preset write parallelism, and the preset calculation parallelism can be the same, partially the same, or different from each other.
[0041] Based on the above technical solution, optionally, the read-write address generation unit 12 is further configured to generate an enable signal for the target repository to enable the read-write enable of the target repository. And optionally, the calculation unit 13 is further configured to generate an enable signal for the target data path to enable the target data path. Specifically, the read-write address generation unit 12 can also generate an enable signal for the target repository, so that the read-write enable of the target repository is turned on, while the read-write enables of other repositories in the storage unit 11 can be turned off to further save power. Similarly, the calculation unit 13 can also generate an enable signal for the target data path, so that the target data path is enabled, while the enables of other data paths in the calculation unit 13 can be turned off to further save power. As in the above example, when reading data, the read enables of bank_1 to bank_4 can be turned on, while the read enables of bank_5 to bank_16 are not turned on. When performing parallel calculations, the enables of datapath_1 to datapath_4 can be turned on, while the enables of datapath_5 to datapath_8 are turned off.
[0042] The accelerator based on the data flow architecture provided by the embodiments of the present invention includes a storage unit, a read-write address generation unit, and a calculation unit. The storage unit includes multiple repositories, and the calculation unit includes multiple data paths. By using the read-write address generation unit to generate the read-write addresses of the storage unit according to the preset read-write parallelism, determining the target repository for reading the data to be processed based on the read-write addresses, and then reading the data to be processed from the target repository into the calculation unit, the calculation unit determines the target data path for operation according to the preset calculation parallelism, and uses the target data path to perform operations on the data to be processed to obtain the processed data, and finally stores the processed data into the target repository according to the read-write addresses, thereby reducing the bandwidth requirements, data storage requirements, and running time of the accelerator, and also reducing the power consumption.
[0043] Embodiment 2
[0044] Figure 2 It is a flowchart of the data access method of the accelerator provided by Embodiment 2 of the present invention. This embodiment is applicable to the situation of improving the utilization rate of computing resources when using a GPU to process a neural network. This method can be applied to the accelerator based on the data flow architecture provided by any embodiment of the present invention, and has the corresponding method flow and beneficial effects of the accelerator. As Figure 2 shown, it specifically includes the following steps:
[0045] S21. Determine the optimal read-write parallelism and the optimal calculation parallelism according to the size of the data to be processed.
[0046] S22. Configure the control register of the accelerator according to the optimal read-write parallelism and the optimal computing parallelism, where the configuration parameters of the control register include a preset read-write parallelism and a preset computing parallelism.
[0047] S23. Read the data to be processed from the storage unit according to the configured preset read-write parallelism and output it to the computing unit, so that the computing unit performs operations according to the configured preset computing parallelism.
[0048] S24. Transmit the processed data obtained after the operation back to the storage unit according to the configured preset read-write parallelism.
[0049] Optionally, determining the optimal read-write parallelism and the optimal computing parallelism according to the size of the data to be processed includes: traversing all configurable read-write parallelisms and computing parallelisms, and respectively determining the corresponding processing times; determining the read-write parallelism and the computing parallelism corresponding to the shortest processing time as the optimal read-write parallelism and the optimal computing parallelism, respectively.
[0050] Specifically, first, the optimal read-write parallelism and the optimal computing parallelism can be calculated according to the size of the data to be processed. Specifically, a traversal method can be adopted, that is, traversing all configurable read-write parallelisms and computing parallelisms, and respectively evaluating the corresponding required processing times under each parallelism setting, so that the read-write parallelism and the computing parallelism corresponding to the shortest processing time can be determined as the optimal read-write parallelism and the optimal computing parallelism, respectively. After obtaining the optimal read-write parallelism and the optimal computing parallelism, the control register of the accelerator can be configured. Specifically, the optimal read-write parallelism can be configured as the preset read-write parallelism, and the optimal computing parallelism can be configured as the preset computing parallelism for subsequent use by the accelerator during the data read-write process. The configuration parameters of the control register can also include the data fetching mode, etc., which can be configured together. After the configuration is completed, the data to be processed can be read from the storage unit according to the preset read-write parallelism and output to the computing unit, so that the computing unit performs operations according to the preset computing parallelism, and then the processed data obtained after the operation is transmitted back to the storage unit according to the preset read-write parallelism. Among them, the specific processes such as data reading, writing, and operation can refer to the description in the above embodiments and will not be repeated here.
[0051] The technical solution provided by the embodiments of the present invention, by using a self-designed accelerator and automatically determining the parallelism parameters required by the accelerator according to the size of the data to be processed, thereby performing data access and operation with more appropriate parallelism parameters, further improving the efficiency of data processing.
[0052] Embodiment III
[0053] Figure 3The structural schematic diagram of the computer device provided in Embodiment 3 of the present invention shows a block diagram of an exemplary computer device suitable for implementing the embodiments of the present invention. Figure 3 The displayed computer device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention. As Figure 3 shown, the computer device includes a processor 31, a memory 32, an input device 33, and an output device 34; the number of processors 31 in the computer device can be one or more, Figure 3 taking one processor 31 as an example, the processor 31, the memory 32, the input device 33, and the output device 34 in the computer device can be connected by a bus or other means, Figure 3 taking the connection by bus as an example.
[0054] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data access method of the accelerator in the embodiments of the present invention (for example, the storage unit 11, the read / write address generation unit 12, and the calculation unit 13 in the accelerator based on the data flow architecture). The processor 31 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 32, that is, implements the above-mentioned data access method of the accelerator.
[0055] The memory 32 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device. In addition, the memory 32 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 32 can further include a memory remotely set relative to the processor 31, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and their combinations.
[0056] The input device 33 can be used to obtain data to be processed, and generate key signal inputs related to the user settings and function controls of the computer device. The output device 34 can include devices such as a display, and can be used to display processing results to the user, etc.
[0057] Embodiment 4
[0058] Embodiment 4 of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a data access method of an accelerator when executed by a computer processor. The method includes:
[0059] Determine the optimal read-write parallelism and the optimal computing parallelism according to the size of the data to be processed;
[0060] Configure the control register of the accelerator according to the optimal read-write parallelism and the optimal computing parallelism, wherein the configuration parameters of the control register include a preset read-write parallelism and a preset computing parallelism;
[0061] Read the data to be processed from the storage unit according to the configured preset read-write parallelism and output it to the computing unit, so that the computing unit performs operations according to the configured preset computing parallelism;
[0062] Transmit the processed data obtained after the operation back to the storage unit according to the configured preset read-write parallelism.
[0063] The storage medium can be any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROM, floppy disk or magnetic tape devices; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium can also include other types of memory or combinations thereof. Additionally, the storage medium can be located in the computer system in which the program is executed, or can be located in a different second computer system that is connected to the computer system through a network (such as the Internet). The second computer system can provide program instructions to the computer for execution. The term "storage medium" can include two or more storage media that can reside in different locations (such as in different computer systems connected through a network). The storage medium can store program instructions (such as specifically implemented as a computer program) executable by one or more processors.
[0064] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the method operations described above, and can also execute related operations in the data access method of the accelerator provided by any embodiment of the present invention.
[0065] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0066] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0067] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and the necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disc of a computer, etc., including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0068] Note that the above are only the preferred embodiments of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it may include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. An accelerator based on a data flow architecture, characterized in that, Comprising: A storage unit, a read-write address generation unit, and a calculation unit; wherein, The storage unit includes multiple memory banks; The read-write address generation unit is used to generate a storage unit read-write address according to a preset read-write parallelism, so as to determine a target memory bank in the storage unit according to the storage unit read-write address, and read data to be processed from the target memory bank to the calculation unit for operation; The calculation unit includes multiple data paths, and is used to determine a target data path according to a preset calculation parallelism, so as to use the target data path to perform an operation on the data to be processed to obtain processed data, and store the processed data into the target memory bank according to the storage unit read-write address; the preset read-write parallelism and the preset calculation parallelism are both flexibly variable; The determination method of the preset read-write parallelism and the preset calculation parallelism is: according to the size of the data to be processed, traverse all settable read-write parallelisms and calculation parallelisms, and respectively determine the corresponding processing times; the read-write parallelism and the calculation parallelism corresponding to the shortest processing time are respectively determined as the preset read-write parallelism and the preset calculation parallelism.
2. The accelerator based on the data flow architecture according to claim 1, wherein The number of the memory banks is an integer multiple of the preset read-write parallelism.
3. The accelerator based on the data flow architecture according to claim 1, wherein The read-write address generation unit is further used to generate an enable signal for the target memory bank to enable the target memory bank to start reading and writing.
4. The accelerator based on a data flow architecture according to claim 1, characterized in that, The calculation unit is further used to generate an enable signal for the target data path to enable the target data path to start.
5. The accelerator based on the data flow architecture according to claim 1, wherein The preset read-write parallelism includes a preset read parallelism and a preset write parallelism; Correspondingly, the storage unit read-write address includes a storage unit read address and a storage unit write address, and the target memory bank includes a target read memory bank and a target write memory bank; The read-write address generation unit is specifically used to generate the storage unit read address according to the preset read parallelism, and generate the storage unit write address according to the preset write parallelism, so as to determine the target read memory bank according to the storage unit read address, and read the data to be processed from the target read memory bank to the calculation unit for operation; The calculation unit is specifically used to perform an operation on the data to be processed using the target data path to obtain the processed data, and store the processed data into the target write memory bank according to the storage unit write address.
6. The accelerator based on the data flow architecture according to claim 5, wherein, The preset read parallelism, the preset write parallelism, and the preset calculation parallelism are the same, partially the same, or different from each other.
7. A data access method for an accelerator, applied to the accelerator based on a data flow architecture as described in any one of claims 1-6, characterized in that, Comprising: Determine the optimal read-write parallelism and the optimal calculation parallelism according to the size of the data to be processed; Configure the control register of the accelerator according to the optimal read-write parallelism and the optimal calculation parallelism, wherein the configuration parameters of the control register include a preset read-write parallelism and a preset calculation parallelism; Read the data to be processed from the storage unit according to the configured preset read-write parallelism, and output it to the calculation unit, so that the calculation unit performs an operation according to the configured preset calculation parallelism; Transfer the processed data obtained after the operation back to the storage unit according to the configured preset read-write parallelism.
8. The data access method of the accelerator according to claim 7, characterized in that, Determining the optimal read-write parallelism and the optimal computing parallelism according to the size of the data to be processed includes: Traverse all the configurable read-write parallelisms and computing parallelisms, and respectively determine the corresponding processing times; Determine the read-write parallelism and the computing parallelism corresponding to the shortest processing time as the optimal read-write parallelism and the optimal computing parallelism respectively.
9. A computer device, characterized in that, Includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data access method of the accelerator as described in any one of claims 7-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the data access method of the accelerator as described in any one of claims 7-8.
Citation Information
Patent Citations
Parallelism degree adjustment algorithm for reducing power consumption of instruction-level parallel processor
CN106445678A
Control method and system for data transmission of data flow architecture neural network chip
CN111860821A
Parallelism degree determination method and system based on fine-grained convolution calculation structure
CN113592088A