Cache architecture and method for AI parameter access
By setting up a prefetch module and a cache module in the cache memory, the prefetch module prefetches and saves data not read from the external memory. The cache module stores data that has been read by the computing unit, solving the problem of extended data reading time of the computing unit in the neural network model, and achieving efficient data access and improvement of computing efficiency.
Patent Information
- Application Number
- CN202311425879.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-02
AI Technical Summary
With the continuous increase of neural network models, the time required for computing units to read external storage data gradually increases, and it is difficult for the prior art to design an efficient storage structure to reduce the time for computing units to read data.
Design a cache architecture for AI parameter access, including prefetch modules and cache modules. The prefetch module prefetches and stores data not read from the external memory, and the cache module stores data that has been read by the computing unit, thereby reducing the interaction between the computing unit and the external memory through this architecture.
It effectively reduces the number of times the calculation unit reads external memory, shortens the data reading time, and improves the computing efficiency.
Smart Images

Figure CN119917428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage technology, and in particular to a cache architecture and method for AI parameter access. Background Art
[0002] When using hardware computing units to run neural networks, the computing units need to read external storage data. As the size of neural network models continues to grow, the data required for calculation increases massively, and the time required for computing units to read external storage data gradually increases. Therefore, it is particularly important to design an efficient storage structure to reduce the time it takes for computing units to read data. Summary of the invention
[0003] Embodiments of the present invention provide a cache architecture and method for AI parameter access to reduce the time taken by a computing unit to read external storage data and improve computing efficiency.
[0004] In one aspect, an embodiment of the present invention provides a cache architecture for AI parameter access, the cache architecture comprising:
[0005] The pre-fetch module is used to pre-fetch and save data not read by the computing unit from the external memory; receive a read address request, obtain data from the local according to the read address request and return it to the computing unit, or send the read address request to the external memory.
[0006] Optionally, the cache architecture also includes: a cache module, which is used to store part of the data and its address in the external memory, receive a read address request sent by the computing unit, obtain data locally and return it to the computing unit, or send the read address request to the pre-fetch module, and send the data returned by the pre-fetch module to the computing unit.
[0007] Optionally, the cache module includes:
[0008] A first routing unit, used for address and data routing;
[0009] A cache unit, used to store part of the data in the external memory and the address corresponding to the data;
[0010] A cache control unit, used to receive a read address request sent by a computing unit, compare the read address with the address in the cache unit; when the two addresses are the same, read data from the cache unit; when the two addresses are different, control the first routing unit to transmit the read address request to the prefetch module; receive data returned by the prefetch module, and store the data and its address in the cache unit; and also to send the data read from the cache unit or the data returned by the prefetch module to the computing unit through the first routing unit.
[0011] Optionally, the cache module further includes: a first flag register;
[0012] The cache control unit is further configured to send a cache hit signal to the first routing unit when the two addresses are the same;
[0013] The first routing unit is further configured to generate a data source flag according to the cache hit signal, and write the data source flag into the first flag register;
[0014] The cache control unit sends the data read from the cache unit or the data returned by the pre-fetch module to the computing unit through the first routing unit according to the data source flag in the first flag register.
[0015] Optionally, the cache control unit is further configured to, after the data in the cache unit reaches a maximum capacity, overwrite old data and addresses that have not been hit for the longest time with new data and addresses that have been written.
[0016] Optionally, the pre-fetch module includes:
[0017] A second routing unit, used for address and data routing;
[0018] A first-in first-out memory, used for storing part of the data in the external memory in a first-in first-out manner;
[0019] A prefetch control unit, used for prefetching data from the external memory and saving it to the first-in-first-out memory; receiving a read address request sent by the cache module, reading data from the first-in-first-out memory according to the read address, or transmitting the read address to the external memory through the second routing unit, and receiving data returned by the external memory; and sending the data read from the first-in-first-out memory or the data returned by the external memory to the cache module through the second routing unit.
[0020] Optionally, the pre-fetch control unit suspends the pre-fetch data operation when the computing unit sends a read request.
[0021] Optionally, the pre-fetch control unit pre-fetches data of continuous addresses of a certain length from the external memory in order from low to high according to the read address request sent last by the calculation unit.
[0022] Optionally, the pre-fetch module further includes: an address memory;
[0023] The pre-fetch control unit is further configured to save the first address of the data of the continuous addresses to the address memory when saving the data of the continuous addresses to the first-in first-out memory;
[0024] After receiving the read address request sent by the cache module, the pre-fetch control unit determines whether it hits according to the address in the address memory; if it hits, data is read from the first-in first-out memory; otherwise, the read address request is transmitted to the external memory through the second routing unit.
[0025] Optionally, the pre-fetch module further includes: a second flag register;
[0026] The pre-fetch control unit is further configured to send a pre-fetch hit signal to the second routing unit after determining a hit;
[0027] The second routing unit is further configured to generate a data source flag according to the prefetch hit signal, and write the data source flag into the second flag register;
[0028] The prefetch control unit sends the data read from the FIFO memory or the data returned from the external memory to the cache module through the second routing unit according to the data source flag in the second flag register.
[0029] Optionally, the priority of the second routing unit pre-fetching data from the external memory is lower than the priority of obtaining data from the external memory according to the read address request sent by the cache module.
[0030] On the other hand, an embodiment of the present invention further provides a cache method for AI parameter access, the method comprising:
[0031] Prefetch and save data not read by the computing unit from the external memory using the prefetch module;
[0032] When data in the external memory needs to be read, determine whether the pre-fetch module stores the data to be read; if so, read the data from the pre-fetch module and return the data to the computing unit; otherwise, obtain the data from the external memory and return the data to the computing unit.
[0033] Optionally, the method also includes: when it is necessary to read data in the external memory, if a read address is stored in the cache module, reading data from the cache module; otherwise, reading data from the pre-fetch module or the external memory, returning the read data to the computing unit, and saving the read data and its address to the cache module.
[0034] Optionally, using the prefetch module to prefetch and save the data not read by the computing unit from the external memory includes: using the prefetch module to prefetch data of continuous addresses of a certain length from the external memory in order from low to high.
[0035] Optionally, the method further comprises: saving the first address of the data of the continuous addresses to the pre-fetch module, so that the pre-fetch module determines whether the data to be read is stored locally according to the first address.
[0036] The cache architecture and method for AI parameter access provided by the embodiment of the present invention, by setting a pre-fetch module in the cache memory, the pre-fetch module pre-fetches and saves the data not read by the computing unit from the external memory, so that when the computing unit reads data, the corresponding data can be read from the pre-fetch module, and only when there is no data to be read in the pre-fetch module, will it be read from the external memory. Since the data in the pre-fetch module is moved from the external memory to the pre-fetch module in advance when the computing unit does not read, the interaction with the external memory is effectively reduced. Compared with reading data from the external memory, the time for the computing unit to read data can be greatly reduced, thereby improving computing efficiency.
[0037] Furthermore, a cache module is set in the cache memory, and the cache module stores the data that the computing unit has read, so that when the computing unit reads data, the data can be obtained in the order of the cache module, the pre-fetch module, and the external memory, that is, when the cache module stores the data to be read, the corresponding data can be directly read from the cache module; when the cache module does not store the data to be read, the corresponding data can be obtained from the pre-fetch module; when the pre-fetch module does not store the data to be read, the corresponding data can be obtained from the external memory. The number of times the computing unit reads the external memory can be further reduced, and the computing efficiency can be improved.
[0038] Furthermore, in view of the characteristics of AI parameter access, the prefetch module can prefetch data of a certain length of continuous addresses from the external memory in order from low to high each time, so that there is no need to save the address of each data, only the first address of the data needs to be saved. After receiving the read address request from the cache module, it can be determined whether to save the data to be read based on the first address, thereby effectively saving storage space. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1It is a schematic diagram of the storage system structure of an existing neural network;
[0040] Figure 2 It is a schematic diagram of the structural relationship between the cache architecture for AI parameter access, neural network calculation and external memory provided by an embodiment of the present invention;
[0041] Figure 3 It is a structural diagram of a high-speed cache architecture for AI parameter access provided by an embodiment of the present invention;
[0042] Figure 4 yes Figure 2 The data flow processing timing diagram in the system shown;
[0043] Figure 5 It is a schematic diagram of the structure of a cache module in a high-speed cache architecture for AI parameter access provided by an embodiment of the present invention;
[0044] Figure 6 is a schematic diagram of the structure of a pre-fetch module in a cache architecture for AI parameter access provided by an embodiment of the present invention;
[0045] Figure 7 is a flow chart of a method for accessing AI parameters provided by an embodiment of the present invention;
[0046] Figure 8 This is another flow chart of a method for accessing AI parameters provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The principle and spirit of the present invention will be described below with reference to the exemplary embodiments shown in the accompanying drawings. It should be understood that these embodiments are described only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way.
[0048] like Figure 1 As shown, the storage system of the neural network can be divided into an on-chip memory 11 and an external memory 12. Usually, the interface reading frequency of the on-chip memory 11 is higher than that of the external memory 12, and the delay of the neural network computing unit 10 reading data from the on-chip memory 11 is much lower than the delay of reading data from the external memory 12. Therefore, in order to reduce the reading delay of the computing unit, an efficient control logic is usually used to increase the number of times the computing unit 10 reads the on-chip memory 11 and reduce the number of times the external memory 12 is read. The control logic can be controlled by hardware or software. Among them:
[0049] Hardware control: The on-chip memory is controlled by hardware logic designed inside.
[0050] Software control: The top-level software is controlled using instructions.
[0051] Cache memory (also called cache) is a typical on-chip storage system controlled by hardware. It is usually located between the computing unit and the external memory and is used to store part of the data. The computing unit directly accesses the cache when reading data, and the cache controls the reading of data in the cache or the reading of data in the external memory.
[0052] In order to shorten the time taken by a neural network computing unit to read external storage data, improve computing efficiency, and meet the computing requirements of large-scale data of neural networks, an embodiment of the present invention provides a cache architecture and method for AI parameter access, which can greatly reduce the number of times the computing unit reads external storage and improve computing efficiency.
[0053] like Figure 2 , which is a schematic diagram of the structural relationship between the cache architecture for AI parameter access, neural network calculation, and external memory provided by an embodiment of the present invention.
[0054] The cache architecture 21 is used to reduce the delay required for the neural network computing unit 20 to read parameters. The cache architecture 21 stores a portion of the data in the external memory 22. When the neural network computing unit 20 (hereinafter referred to as the computing unit) reads parameter data, it sends a read request to the cache architecture 21. The cache architecture 21 detects whether the required data is stored inside the cache architecture 21. If it is stored inside the cache architecture 21, the internal data is directly returned to the computing unit 20; if it is not stored inside the cache architecture 21, the external memory 22 is read and the data is returned.
[0055] The solution of the present invention only optimizes the delay of reading parameters of the computing unit 20. When applied, the neural network parameters in the external memory 22 remain unchanged, and the calculation results will not be wrong due to the difference between the data in the cache architecture 21 and the data in the external memory 22. Based on this feature, the computing unit 20 does not write data into the cache architecture 21, so the control logic can be simplified to achieve a better optimization effect.
[0056] The embodiment of the present invention provides a cache architecture for AI parameter access, in which a pre-fetch module is set in the cache memory to predict the data that the computing unit may read, and store the data in the pre-fetch module in advance. Accordingly, after receiving a read address request, data is obtained from the local according to the read address request and returned to the computing unit, or the read address request is sent to the external memory, thereby effectively reducing the number of times the computing unit reads the external memory.
[0057] Furthermore, a cache module can be set in the cache memory to store part of the data and its address in the external memory, receive a read address request sent by the computing unit, obtain data locally and return it to the computing unit, or send the read address request to the pre-fetch module.
[0058] It should be noted that the data stored in the prefetch module is the data in the external memory that has not been read by the computing unit, that is, the prefetch module moves the data from the external memory to the local storage when the computing unit has not read the data; and the data stored in the cache module is the data in the external memory that has been read by the computing unit, which can be read directly from the external storage or from the prefetch module. After reading, the corresponding data is saved locally, with the purpose of reducing the time required for the computing unit to read the data next time.
[0059] The following takes the simultaneous provision of a pre-fetch module and a cache module as an example to describe in detail the cache architecture for AI parameter access provided by an embodiment of the present invention.
[0060] like Figure 3 , which is a schematic diagram of a structure of a cache architecture for AI parameter access provided by an embodiment of the present invention.
[0061] The cache architecture includes: a cache module 31 and a pre-fetch module 32. Among them:
[0062] Combined with Figure 2 The cache module 211 is used to store part of the data and its address in the external memory 22, receive the read address request sent by the computing unit 20, obtain the data in the local memory and return it to the computing unit 20, or send the read address request to the pre-fetch module 212, and send the data returned by the pre-fetch module 212 to the computing unit 20;
[0063] The prefetch module 212 is used to prefetch and save data not read by the computing unit 20 from the external memory 22; receive the read address request sent by the cache module 211, obtain data from the local memory according to the read address request and return it to the cache module 211, so that the cache module 211 returns the data to the computing unit 20, or sends the read address request to the external memory 22.
[0064] Also refer to Figure 2 and Figure 3The read address request sent by the calculation unit 20 is first transmitted to the cache module 211. The cache module 211 determines whether the read data is in the memory of the module. If so, the stored data is returned. If not, the read address request is sent to the pre-fetch module 212. The pre-fetch module 212 also determines whether the data is in the memory of the module. If so, the stored data is returned. If not, the read address request is sent to the external memory 22.
[0065] It should be noted that, in the embodiment of the present invention, the read data may come from the cache module 211, the pre-fetch module 212, or the external memory 22, and the read data is returned to the computing unit 20 in sequence through the internal control of the cache architecture 21.
[0066] Figure 4 for Figure 2 The data flow processing timing diagram in the system is shown.
[0067] A data stream is a set of ordered data sequences with a start and an end. When reading data, the computing unit 20 sends out the read address in continuous clock cycles and sends an address end flag at the end. After several clock cycles, the read data is returned in sequence and a data end flag is sent. The cache architecture 21 designed in the embodiment of the present invention can process addresses and data in a data stream manner. The hardware design of the cache architecture 21 controls data and addresses separately. It does not need to wait for the current data to return before receiving the next read address, but processes batches of data and addresses. This processing method is more suitable for neural network accelerators and can reduce the delay in reading parameters.
[0068] like Figure 5 , which is a schematic diagram of the structure of a cache module in a high-speed cache architecture for AI parameter access provided by an embodiment of the present invention.
[0069] The cache module 211 includes: a first routing unit 51, a cache unit 52, and a cache control unit 53. Among them:
[0070] The first routing unit 51 is used for address and data routing;
[0071] The cache unit 52 is used to store Figure 2 Part of the data in the external memory 22, and the addresses corresponding to the data;
[0072] The cache control unit 53 is used to receive Figure 2 The calculation unit 20 sends a read address request, compares the read address with the address in the cache unit 52; when the two addresses are the same, reads data from the cache unit 52; when the two addresses are different, controls the first routing unit 51 to transmit the read address request to Figure 3receiving the data returned by the pre-fetch module 212, and storing the data and its address in the cache unit 52; and also used to send the data read from the cache unit 52 or the data returned by the pre-fetch module 212 to the computing unit 20 through the first routing unit 51.
[0073] Continue to refer to Figure 5 The cache module 211 may further include: a first flag register 510. The first flag register 510 may be a part of the first routing unit 51 (such as Figure 5 ), and may also be independent of the first routing unit 51, which is not limited in this embodiment of the present invention.
[0074] Correspondingly, the cache control unit 53 is also used to send a cache hit signal to the first routing unit 51 when the two addresses are the same. The first routing unit 51 can generate a data source flag according to the cache hit signal and write the data source flag into the first flag register 510. The data source flag can also be generated by the cache control unit 53 and written into the first flag register 510, which is not limited in this embodiment of the present invention. In addition, the cache control unit 53 can read the data from the cache unit 52 or Figure 3 The data returned by the pre-fetch module 212 in the first routing unit 51 is sent to Figure 2 The computing unit 20 in FIG.
[0075] Since the capacity of the cache unit 52 is limited, the cache control unit 53 is also used to overwrite the old data and addresses that have not been hit for the longest time with the new data and addresses that are written after the data in the cache unit 52 reaches the maximum capacity.
[0076] In the cache module 211, the first routing unit 51 generates a data source flag according to the cache hit signal and stores it in the first flag register 510. The data source flag is used to determine whether the read data comes from the next level or the current level cache unit 52. Figure 3 When the computing unit 20 in the buffer reads data, the data source flag is read to control the read data to be returned in order. Therefore, after the data source flag is stored in the first flag register 510, the cache module 212 can ensure that the returned data is correct, that is, the order of the returned data is the same as the read address order. At this time, the next address request of the computing unit 20 can be processed without waiting for the read data of the current request to be returned. Figure 6 , which is a schematic diagram of the structure of a pre-fetch module in a cache architecture for AI parameter access provided by an embodiment of the present invention.
[0077] The pre-fetch module 212 includes: a second routing unit 61, a first-in first-out memory 62, and a pre-fetch control unit 63. Among them:
[0078] The second routing unit 61 is used for address and data routing;
[0079] The first-in first-out memory 62 is used to store the Figure 2 Part of the data in the external memory 22, and the addresses corresponding to the data;
[0080] The pre-fetch control unit 63 is used to pre-fetch data from the external memory 22 and save it to the first-in first-out memory 62; receive Figure 3 The cache module 211 in the embodiment receives a read address request sent by the cache module 211, reads data from the first-in-first-out memory 62 according to the read address, or transmits the read address to the external memory 22 through the second routing unit 61, and receives data returned by the external memory 22; sends the data read from the first-in-first-out memory 62 or the data returned by the external memory 22 to the cache module 211 through the second routing unit 61.
[0081] It should be noted that the pre-fetch control unit 63 suspends the pre-fetch data operation when the computing unit 20 sends a read request. In addition, each time the pre-fetch control unit 63 pre-fetches data from the external memory 22, it can Figure 2 The last read address request sent by the calculation unit 20 in the pre-fetch module 212 is used to pre-fetch data of a certain length of continuous addresses from the external memory 22 in a low to high order. The length of the read data can be determined according to parameters such as the capacity of the first-in first-out memory 62 and the data transmission bandwidth, which is not limited in the embodiment of the present invention. Among them, the last read address request sent by the calculation unit 20 refers to the read address request sent to the pre-fetch module 212 when the last read address request sent by the calculation unit 20 but the cache module 211 does not hit.
[0082] In a non-limiting embodiment, the pre-fetch control unit 63 may save the data pre-fetched from the external memory 22 and the address thereof to the FIFO memory 62 .
[0083] Considering that the length and order of data pre-fetched by the pre-fetch control unit 63 from the external memory 22 each time are fixed, in order to further save storage space, in another non-limiting embodiment, the pre-fetched data can also be stored only in the first-in first-out memory 62, and the address is stored separately, and there is no need to store the addresses of all data.
[0084] For example, an address memory ( Figure 6(not shown). Accordingly, when the pre-fetch control unit 63 saves the data of the continuous addresses to the FIFO memory 60, it saves the first address of the data of the continuous addresses to the address memory. After receiving the read address request sent by the cache module 211, the pre-fetch control unit 63 determines whether it hits according to the address in the address memory; if it hits, it reads the data from the FIFO memory 62; otherwise, it transmits the read address request to the external memory 22 through the second routing unit 61.
[0085] like Figure 6 As shown, the pre-fetch module 212 may further include a second flag register 610. Accordingly, when the pre-fetch control unit 63 determines that the read address hits, it will also send a pre-fetch hit signal to the second routing unit 61. Accordingly, the second routing unit 61 generates a data source flag according to the pre-fetch hit signal, and writes the data source flag into the second flag register 610. The data source flag may also be generated by the pre-fetch control unit 63 and written into the second flag register 610, which is not limited in this embodiment of the present invention.
[0086] Accordingly, the pre-fetch control unit 63 may read the data from the first-in first-out memory 62 or Figure 2 The data returned by the external memory 22 in the second routing unit 61 is sent to Figure 3 The cache module 211 in.
[0087] It should be noted that the priority of the second routing unit 61 to pre-fetch data from the external memory 22 is lower than that according to Figure 3 The cache module 211 in the cache module 211 sends a read address request to obtain data from the external memory 22 with priority.
[0088] In the solution of the present invention, the pre-fetch module 212 predicts the data required by the computing unit and stores the data not read by the computing unit in advance into the first-in first-out memory of the pre-fetch module 212. When the computing unit reads the parameters, if the data has been stored in the pre-fetch module 212, the read delay can be reduced.
[0089] The cache architecture for AI parameter access provided by the present invention can process addresses and data in the form of data streams by setting the above-mentioned first register and second register. Processing addresses in the form of data streams requires processing one address and data per cycle. After receiving the address in each cycle, generating a data source flag and storing it in the first flag register 510, the address processing is completed. The data processing method is similar to the address processing method, except that a data source flag needs to be read each time. Therefore, the solution of the embodiment of the present invention can enable the hardware to process in the form of data streams.
[0090] The cache architecture for AI parameter access provided by the embodiment of the present invention can design a pre-fetching method based on the neural network parameter reading scenario. For example, the pre-fetching is performed according to the following two rules:
[0091] 1) Pre-fetch consecutive addresses based on the read address sent by the calculation unit last time.
[0092] 2) Data can be pre-fetched only when the computing unit does not send a read request. When the computing unit sends a read request, pre-fetching data is suspended. This ensures that pre-fetching does not block the computing unit from reading data, preventing the computing unit from increasing the delay in reading data.
[0093] The cache architecture for AI parameter access provided by the embodiment of the present invention is a hardware-controlled on-chip storage. Compared with software control, the hardware control logic can detect the state of the computing unit more accurately and predict the data that the computing unit needs to read more timely and reasonably. At the same time, no additional software configuration instructions are required, which is more friendly to software design.
[0094] In addition, compared with other cache architectures, the solution of the present invention can optimize the data access of the neural network computing unit to the external memory in the following aspects:
[0095] 1) The neural network computing unit usually reads data in the form of a data stream. The cache architecture in the solution of the present invention processes the read address and data in the form of a data stream, thereby reducing the return time of the computing unit reading data.
[0096] 2) The solution of the present invention is aimed at the characteristics of the neural network application scenario: when the neural network computing unit is running, the computing unit will not modify the parameters, and the parameters remain unchanged. Therefore, the cache architecture provided by the present invention can only use hardware to implement the read channel, not the write channel, reducing resource consumption; in addition, by combining hardware design with software control, a pre-fetching method of pre-fetching continuous addresses is adopted, and the software configuration can store the neural network parameters in continuous addresses. Continuous address pre-fetching can accurately predict the parameters required by the computing unit, and can also greatly simplify the hardware control logic.
[0097] 3) The solution of the present invention is targeted at the characteristics of the neural network application scenario: in the neural network, the parameter data is static and unchanged, and the instruction execution sequence of the common neural network processor is also pre-compiled. Therefore, when using the cache architecture provided by the embodiment of the present invention, the storage of the parameters in the storage can be rearranged in advance according to the order of accessing the parameters in the instruction sequence to ensure that the access to the parameters is always continuous during the operation.
[0098] Accordingly, an embodiment of the present invention also provides a cache method for AI parameter access, such as Figure 7As shown, it is a flow chart of the method, comprising the following steps:
[0099] Step 701: Use a pre-fetch module to pre-fetch and save data not read by a computing unit from an external memory.
[0100] Step 702, when it is necessary to read data in the external memory, determine whether the pre-fetch module stores the data to be read; if so, read the data from the pre-fetch module and return the data to the computing unit; otherwise, obtain the data from the external memory and return the data to the computing unit.
[0101] Specifically, the computing unit sends a read address request to the pre-fetch module, and the pre-fetch module determines whether data corresponding to the corresponding read address is stored locally.
[0102] It should be noted that, in a specific application, the pre-fetch module may store the pre-fetched data and its address locally, and after receiving the read address request, compare the locally stored address with the read address to determine whether the corresponding data is stored locally.
[0103] Furthermore, considering that the neural network computing unit has the characteristics of sequential access to AI parameters and access to multiple data at a time, in a specific application, the prefetch module can prefetch data of a certain length of continuous addresses from the external memory in order from low to high.
[0104] Accordingly, based on the above characteristics, for the data pre-fetched by the pre-fetch module from the external memory, it is not necessary to save the address of each data, and only the first address of the data of the read continuous address can be saved to the pre-fetch module. The pre-fetch module can determine whether the data to be read is stored locally based on the first address. In this way, the storage space of the pre-fetch module can be greatly saved.
[0105] like Figure 8 FIG. 2 is another flow chart of a cache method for AI parameter access provided by an embodiment of the present invention, comprising the following steps:
[0106] Step 801, using a pre-fetch module to pre-fetch and save data not read by a computing unit from an external memory.
[0107] Step 802, when it is necessary to read data in the external memory, if the read address is stored in the cache module, the data is read from the cache module; otherwise, the data is read from the pre-fetch module or the external memory, the read data is returned to the computing unit, and the read data and its address are saved to the cache module.
[0108] That is to say, the cache module is checked first, and then the pre-fetch module is checked. If there is no data to be read, the data is finally read from the external memory.
[0109] and Figure 7 Compared with the embodiment shown, Figure 8 In the embodiment, the pre-fetch module not only pre-fetches and saves the data not read by the computing unit from the external memory, but also saves the data and its address to the cache module after reading the data from the pre-fetch module or the external memory each time. Therefore, the next time the computing unit reads data, the cache module can first check whether the corresponding data is stored locally. If so, the data can be directly read from the cache module. This further reduces the number of times the computing unit reads data from the external memory, saves the data reading time, and improves the computing efficiency.
[0110] In a specific implementation, each module / unit included in each device or product described in the above embodiments may be a software module / unit or a hardware module / unit, or may be partly a software module / unit and partly a hardware module / unit.
[0111] For example, for each device or product applied to or integrated in a chip, each module / unit contained therein may be implemented in the form of hardware such as circuits, or at least some of the modules / units may be implemented in the form of software programs, which run on a processor integrated inside the chip, and the remaining (if any) modules / units may be implemented in the form of hardware such as circuits; for each device or product applied to or integrated in a chip module, each module / unit contained therein may be implemented in the form of hardware such as circuits, and different modules / units may be located in the same component (such as a chip, circuit module, etc.) or different components of the chip module, or at least some of the modules / units may be implemented in the form of software programs. The element can be implemented in the form of a software program, which runs on a processor integrated inside the chip module, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits; for various devices and products applied to or integrated in the terminal, the various modules / units contained therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (for example, chip, circuit module, etc.) or in different components in the terminal, or, at least some modules / units can be implemented in the form of a software program, which runs on a processor integrated inside the terminal, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits.
[0112] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope defined by the claims.
Claims
1. A cache architecture for AI parameter access, characterized in that: The cache architecture includes: The pre-fetch module is used to pre-fetch and save data not read by the computing unit from the external memory; receive a read address request, obtain data from the local according to the read address request and return it to the computing unit, or send the read address request to the external memory.
2. The cache architecture for AI parameter access according to claim 1, characterized in that: The cache architecture further comprises: The cache module is used to store part of the data and its address in the external memory, receive the read address request sent by the computing unit, obtain the data locally and return it to the computing unit, or send the read address request to the pre-fetch module and send the data returned by the pre-fetch module to the computing unit.
3. The cache architecture for AI parameter access according to claim 2, characterized in that: The cache module comprises: A first routing unit, used for address and data routing; A cache unit, used to store part of the data in the external memory and the address corresponding to the data; A cache control unit, used to receive a read address request sent by a computing unit, compare the read address with the address in the cache unit; when the two addresses are the same, read data from the cache unit; when the two addresses are different, control the first routing unit to transmit the read address request to the prefetch module; receive data returned by the prefetch module, and store the data and its address in the cache unit; and also to send the data read from the cache unit or the data returned by the prefetch module to the computing unit through the first routing unit.
4. The cache architecture for AI parameter access according to claim 3, characterized in that: The cache module also includes: a first flag register; The cache control unit is further configured to send a cache hit signal to the first routing unit when the two addresses are the same; The first routing unit is further configured to generate a data source flag according to the cache hit signal, and write the data source flag into the first flag register; The cache control unit sends the data read from the cache unit or the data returned by the pre-fetch module to the computing unit through the first routing unit according to the data source flag in the first flag register.
5. The cache architecture for AI parameter access according to claim 3, characterized in that: The cache control unit is further used to overwrite old data and addresses that have not been hit for the longest time with new data and addresses that have been written after the data in the cache unit reaches a maximum capacity.
6. The cache architecture for AI parameter access according to claim 2, characterized in that: The pre-fetch module comprises: A second routing unit, used for address and data routing; A first-in first-out memory, used for storing part of the data in the external memory in a first-in first-out manner; A prefetch control unit, used for prefetching data from the external memory and saving it to the first-in-first-out memory; receiving a read address request sent by the cache module, reading data from the first-in-first-out memory according to the read address, or transmitting the read address to the external memory through the second routing unit, and receiving data returned by the external memory; and sending the data read from the first-in-first-out memory or the data returned by the external memory to the cache module through the second routing unit.
7. The cache architecture for AI parameter access according to claim 6, characterized in that: The pre-fetch control unit suspends the pre-fetch data operation when the computing unit sends a read request.
8. The cache architecture for AI parameter access according to claim 6, characterized in that: The pre-fetch control unit pre-fetches data of continuous addresses of a certain length from the external memory in an ascending order according to the read address request sent last by the calculation unit.
9. The cache architecture for AI parameter access according to claim 8, characterized in that: The pre-fetch module also includes: an address memory; The pre-fetch control unit is further configured to save the first address of the data of the continuous addresses to the address memory when saving the data of the continuous addresses to the first-in first-out memory; After receiving the read address request sent by the cache module, the pre-fetch control unit determines whether it hits according to the address in the address memory; if it hits, data is read from the first-in first-out memory; otherwise, the read address request is transmitted to the external memory through the second routing unit.
10. The cache architecture for AI parameter access according to claim 9, characterized in that: The pre-fetch module further includes: a second flag register; The pre-fetch control unit is further configured to send a pre-fetch hit signal to the second routing unit after determining a hit; The second routing unit is further configured to generate a data source flag according to the prefetch hit signal, and write the data source flag into the second flag register; The prefetch control unit sends the data read from the FIFO memory or the data returned from the external memory to the cache module through the second routing unit according to the data source flag in the second flag register.
11. The cache architecture for AI parameter access according to claim 6, characterized in that: The priority of the second routing unit pre-fetching data from the external memory is lower than the priority of obtaining data from the external memory according to the read address request sent by the cache module.
12. A cache method for AI parameter access, characterized in that: The method comprises: Prefetch and save data not read by the computing unit from the external memory using the prefetch module; When data in the external memory needs to be read, determine whether the pre-fetch module stores the data to be read; if so, read the data from the pre-fetch module and return the data to the computing unit; otherwise, obtain the data from the external memory and return the data to the computing unit.
13. The cache method for AI parameter access according to claim 12, characterized in that: The method further comprises: When it is necessary to read data in the external memory, if the read address is stored in the cache module, the data is read from the cache module; otherwise, the data is read from the pre-fetch module or the external memory, the read data is returned to the computing unit, and the read data and its address are saved to the cache module.
14. The cache method for AI parameter access according to claim 12, characterized in that: The method of using the pre-fetch module to pre-fetch and save the data not read by the computing unit from the external memory comprises: The pre-fetch module is used to pre-fetch data of a certain length of continuous addresses from the external memory in order from low to high.
15. The cache method for AI parameter access according to claim 14, characterized in that: The method further comprises: The first address of the data of the continuous addresses is saved to the pre-fetch module, so that the pre-fetch module determines whether there is data to be read stored locally according to the first address.