Memory access system, method, storage medium, electronic device and processor
By employing an out-of-order access caching architecture and arbitration unit, the bank conflict problem in multi-bank caching architecture is resolved, achieving high bandwidth utilization and improved processor performance, and adapting to flexible tile width access requirements.
Patent Information
- Application Number
- CN202580001284.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies in devices such as multi-core processors, GPUs, and neural network accelerators suffer from bank conflicts in multi-bank caching architectures, leading to reduced bandwidth utilization and failing to meet the flexible tile width access requirements.
The out-of-order access caching architecture generates out-of-order responses and sequentially submitted read and write-back requests through write-back address generation and read address generation units. Combined with arbitration and caching units, it avoids Bank conflicts and achieves high bandwidth utilization.
It effectively reduces bank conflicts, improves memory bandwidth utilization, ensures data access correctness, enhances processor performance, and adapts to access requirements with different tile widths.
Smart Images

Figure CN121002473A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data storage, and particularly relates to a memory access system and method, a storage medium, an electronic device and a processor. BACKGROUND
[0002] In modern high-performance computing systems, memory access efficiency is one of the key factors determining the overall system performance. In particular, in multi-core processors, graphic processing units (GPUs) and neural network accelerators, optimizing memory access patterns to reduce latency and improve bandwidth utilization is crucial.
[0003] In cache design, to improve parallel access bandwidth, the cache is often divided into multiple independent banks, also known as Bank. Each Bank has independent data ports and address ports, and only one address can be read or written at the same time in a Bank. To improve memory bandwidth and parallel processing capability, modern processors usually use a multi-Bank cache architecture, i.e., dividing the SRAM cache into multiple independent Banks to achieve parallel memory access. However, in actual applications, multiple concurrent memory requests may be mapped to the same cache Bank, causing Bank conflict. Bank conflict requires serializing parallel access, thus reducing bandwidth utilization.
[0004] In existing solutions, Bank partitioning is used to place different arrays in different Bank partitions to solve conflicts between arrays; address remapping is used to solve conflicts within arrays due to stride access.
[0005] However, the existing solution has poor universality, and it strictly limits the access mode of the processor, i.e., the processor needs to access the image in a cross-row manner, accessing one data per row at the same time. Under this premise, the solution can eliminate Bank conflict. However, the common convolution calculation mode needs to access in Tile units, and the width and height of the Tile are flexible and variable. The solution is only applicable to the case where the width of the Tile is 1, i.e., only one data is accessed per row at the same time, which cannot meet the needs of actual applications.
[0006] SUMMARY
[0007] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a memory access system and method, a storage medium, an electronic device and a processor, which effectively reduce Bank conflict and improve the bandwidth utilization of the memory based on a cache architecture with out-of-order access.
[0008] In a first aspect, the application provides a memory access system, comprising a write-back address generation unit, a calculation unit, a read address generation unit and a memory; the read address generation unit is configured to generate a read address; the memory is configured to provide corresponding read data to the calculation unit in a manner of out-of-order response and in-order submission to avoid Bank conflict according to the read address; the write-back address generation unit is configured to generate a write-back address; and the calculation unit is configured to generate write-back data according to the read data and write the write-back data into the memory in a manner of out-of-order response according to the write-back address to avoid Bank conflict.
[0009] In an implementation form of the first aspect, the system further comprises a first preset number of cache units and a first arbitration unit; the first preset number is the number of Banks contained in the memory.
[0010] Each write-back address generated by the write-back address generation unit corresponds to a cache unit.
[0011] Each write-back data generated by the calculation unit corresponds to a cache unit.
[0012] The write-back address and the write-back data are written into the corresponding cache unit to form a write-back request.
[0013] The first arbitration unit detects whether there is a Bank conflict between the write-back requests, arbitrates when the Bank conflict occurs, and writes the corresponding write-back data into the memory according to the corresponding write-back address based on the first preset number of write-back requests that pass the arbitration.
[0014] In an implementation form of the first aspect, each cache unit can store a second preset number of write-back requests.
[0015] In an implementation form of the first aspect, the system further comprises a first multiple selector for selecting b from a, wherein a represents the product of the first preset number and the second preset number, and b represents the first preset number, the first multiple selector selects the first preset number of write-back requests that pass the arbitration for data write-back.
[0016] In an implementation form of the first aspect, when the write-back request does not pass the arbitration, the write-back request is temporarily stored in the corresponding cache unit and participates in arbitration in the next cycle.
[0017] In an implementation form of the first aspect, the system further comprises a first preset number of address cache units and a second arbitration unit; the first preset number is the number of Banks contained in the memory.
[0018] Each read address generated by the read address generation unit corresponds to an address cache unit and a read order.
[0019] The read address and the corresponding read order constitute a read request, and the read request is written into a corresponding address cache unit;
[0020] The second arbitration unit detects whether there is a Bank conflict between the read requests, arbitrates when the Bank conflict occurs, and obtains read data from the corresponding read address of the memory based on the first preset number of read requests that pass the arbitration;
[0021] Based on the read order corresponding to the read data, the read data is output to the calculation unit.
[0022] In an implementation form of the first aspect, the address cache unit can store a second preset number of read requests.
[0023] In an implementation form of the first aspect, the system further comprises a second multi-selector for selecting b from a, wherein a represents the product of the first preset number and the second preset number, and b represents the first preset number, and the second multi-selector selects the first preset number of read requests that pass the arbitration for data reading.
[0024] In an implementation form of the first aspect, the system further comprises a distributor and a first preset number of data cache units for selecting a from b, wherein a represents the product of the first preset number and the second preset number, and b represents the first preset number, and each data cache unit corresponds to an address cache unit.
[0025] The distributor writes the corresponding read data into the data cache unit according to the read order;
[0026] The data cache unit stores the read data provided by the corresponding Bank, and outputs the corresponding read data to the calculation unit in the read order.
[0027] In an implementation form of the first aspect, a counter is arranged in the data cache unit, and the counter is used to record the current output read order; only when all the data cache units obtain the read data corresponding to the current output read order, the read data is output to the calculation unit at the same time.
[0028] In an implementation form of the first aspect, when the read request does not pass the arbitration, the read request is temporarily stored in the corresponding address cache unit, and participates in arbitration in the next cycle.
[0029] In a second aspect, the application provides a processor comprising the memory access system described above.
[0030] In a third aspect, the present application provides a memory access method, the method comprising the following steps:
[0031] generating a read address based on a read address generating unit;
[0032] providing corresponding read data to the computing unit in a manner of out-of-order response and in-order submission to avoid Bank conflict based on the memory according to the read address;
[0033] generating a write-back address based on a write-back address generating unit;
[0034] generating write-back data according to the read data based on the computing unit, and writing the write-back data into the memory in a manner of out-of-order response according to the write-back address to avoid Bank conflict.
[0035] In a fourth aspect, the present application provides an electronic device, comprising a processor and a memory;
[0036] the memory is configured to store a computer program;
[0037] the processor is configured to execute the computer program stored in the memory, so that the electronic device executes the memory access method described above.
[0038] In a fifth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the program is executed by an electronic device to implement the memory access method described above.
[0039] As described above, the memory access system, method, storage medium, electronic device and processor of the present application have the following beneficial effects:
[0040] (1) Based on the cache architecture of out-of-order access, Bank conflict is effectively reduced;
[0041] (2) All access requests in the cache participate in arbitration at the same time, and when the previous group of access requests do not use all the Banks, the subsequent access can use the remaining Banks at the same time, achieving high bandwidth utilization;
[0042] (3) The correctness of data access is ensured, the performance of the processor is effectively improved, and the versatility is good. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 Fig. 1 shows a structural schematic diagram of the memory access system of the present application in an embodiment;
[0044] Figure 2 Fig. 2 shows a structural schematic diagram of the read address generating unit of the present application in an embodiment;
[0045] Figure 3Figure 6(a) shows a data writeback schematic diagram of a 2D array from cycle 0 to cycle 1 in an embodiment of the memory access system of the present application;
[0046] Figure 4 Figure 6(b) shows a data writeback schematic diagram of a 2D array from cycle 2 to cycle 5 in an embodiment of the memory access system of the present application;
[0047] Figure 6(c) shows a data writeback schematic diagram of a 2D array from cycle 6 to cycle 7 in an embodiment of the memory access system of the present application; Figure 4
[0048] Figure 7(a) shows a data read schematic diagram of a 2D array from cycle 0 to cycle 1 in an embodiment of the memory access system of the present application; Figure 4
[0049] Figure 7(b) shows a data read schematic diagram of a 2D array from cycle 2 to cycle 5 in an embodiment of the memory access system of the present application; Figure 4
[0050] Figure 6 Figure 7(c) shows a data read schematic diagram of a 2D array from cycle 6 to cycle 9 in an embodiment of the memory access system of the present application;
[0051] Figure 4 Figure 8 shows a schematic diagram of a processor in an embodiment of the present application;
[0052] Figure 4 Figure 9 shows a schematic diagram of a memory access method in an embodiment of the present application;
[0053] Figure 4 Figure 10 shows a schematic diagram of an electronic device in an embodiment of the present application.
[0054] Figure 8 DETAILED DESCRIPTION
[0055] Figure 9 Figure 9 shows a schematic diagram of a memory access method in an embodiment of the present application;
[0056] Figure 10 Figure 10 shows a schematic diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The present application is herein described, by way of example only, with reference to certain embodiments thereof. It is to be understood that variations and modifications of the embodiments can be made based on different viewpoints and applications without departing from the spirit of the present application. It is to be further understood that the embodiments and features of the embodiments can be combined with each other without departing from the spirit of the present application.
[0058] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concepts of the present application in a schematic manner, and only show the components related to the present application in the diagrams, not drawn according to the number, shape and size of the components in actual implementation. The shape, number and proportion of each component in actual implementation can be a random change, and the component layout pattern can also be more complex.
[0059] The memory access system, method, storage medium, electronic device and processor of the present application aim at the problem of memory bandwidth utilization. By using the out-of-order reading mode of "sequential application, out-of-order response, sequential submission" and the out-of-order write-back mode of "sequential application, out-of-order response", the Bank conflict in the memory is reduced and the bandwidth utilization is improved under the premise of ensuring correctness. It should be noted that the memory can be an SRAM memory or a multi-register stack.
[0060] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.
[0061] As shown in Figure 1 In an embodiment, the memory access system of the present application includes a write-back address generation unit 1 (WAG), a calculation unit 2, a read address generation unit 3 (RAG) and a memory 4. Preferably, the calculation unit 2 uses a process element array (PEA), and the memory 4 uses an on-chip SRAM, which contains a plurality of Banks.
[0062] When accessing data to the memory 4, the read address generation unit generates a read address; the memory provides corresponding read data to the calculation unit by the mode of out-of-order response and sequential submission to avoid Bank conflict. The read data is the data read from the memory by the out-of-order reading mode. The write-back address generation unit generates a write-back address; the calculation unit generates write-back data according to the read data and writes the write-back data into the memory by the out-of-order write-back mode according to the write-back address to avoid Bank conflict.
[0063] In an embodiment, as shown in Figure 2As shown, the read address generation unit includes three parts: a vector head coordinate generator, a vector head address calculator, and a vector address extender. The vector head coordinate generator includes cascaded multi-stage counters (e.g., counter 0, counter 1, counter 2) responsible for generating multi-dimensional coordinates. For example, for a three-dimensional array Data[C][H][W], coordinates (i, j, k) are generated, corresponding to data Data[i][j][k]. The vector head address calculator includes a number of adders and multipliers responsible for calculating data addresses according to the coordinates and taking the addresses as vector head addresses. For example, for a three-dimensional array Data[C][H][W] and coordinates (i, j, k), the address Addr = base address + i*H*W + j*W + k is calculated. The vector address extender is used to generate a plurality of addresses in parallel through a plurality of adders according to the vector head address. For example, for a vector head address Addr, the address combination (Addr0, Addr1, Addr2, Addr3) can be generated, corresponding to data (Data[i][j][k], Data[i][j][k+1], Data[i][j][k+2], Data[i][j][k+3]). Wherein, Addr0 = Addr, Addr1 = Addr + step, Addr2 = Addr + 2*step, Addr3 = Addr + 3*step, and step is the address stride. By controlling the address stride, the vector address generator can also be extended in any dimension. For example, if the address stride is W, the address combination (Addr, Addr+W, Addr+2*W, Addr+3*W) can be generated, corresponding to data (Data[i][j][k], Data[i][j+1][k], Data[i][j+2][k], Data[i][j+3][k]). For another example, if the address stride is H*W, the address combination (Addr, Addr+H*W, Addr+2*H*W, Addr+3*H*W) can be generated, corresponding to data (Data[i][j][k], Data[i+1][j][k], Data[i+2][j][k], Data[i+3][j][k]).
[0064] As Figure 3 As shown, in an embodiment, in order to perform data write back, the memory access system of the present application further includes a first preset number of cache units and a first arbitration unit (Arbiter1). The first preset number is the number of banks included in the memory. Preferably, 16 cache units are provided in the present application, i.e., Buffer0-Buffer15. The depth of each cache unit Buffer is a second preset number. Preferably, the second preset number is 4, i.e., each cache unit can store 4 write back addresses at the same time.
[0065] The following is a description of the data write-back in the above configuration. In the data write-back process of the present application, there are two stages, namely, sequential application and out-of-order response.
[0066] In the sequential application stage, each write-back address generated by the write-back address generation unit WAG corresponds to a cache unit. For example, write-back address Addr0 is written into cache unit Buffer0, write-back address Addr1 is written into cache unit Buffer2, and so on. Each write-back data generated by the calculation unit PEA corresponds to a cache unit. For example, write-back data Data0 is written into cache unit Buffer0, write-back data Data1 is written into cache unit Buffer2, and so on. Therefore, the write-back address Addr n and the write-back data Datan are written into the corresponding cache unit Buffern, thereby forming a write-back request. Among the 16 cache units Buffer, at most 64 write-back requests are stored.
[0067] In the out-of-order response stage, the first arbitration unit Arbiter detects whether there is a Bank conflict among the write-back requests in the 16 cache units, and arbitrates when a Bank conflict occurs. The first arbitration unit adopts a priority arbitration algorithm. The write-back request that enters the cache unit Buffer earlier has a higher priority, and in each cycle, a write-back request with the highest priority is selected from multiple write-back requests competing for the same Bank to be responded, and other requests are temporarily buffered and will be requested again in the next cycle. For the write-back request that passes the arbitration, 16 write-back requests are selected to write the corresponding write-back data into the memory according to the corresponding write-back address. Due to the existence of Bank conflict, it may occur that the write-back request that enters the cache unit Buffer later is written back first, so it is called out-of-order response. For the write-back request that does not pass the arbitration, the write-back request is temporarily stored in the corresponding cache unit, and in the next cycle Cycle, it continues to participate in arbitration. At the same time, since the depth of each cache unit Buffer is 4, at most 4 write-back requests can be stored, so even if the previous write-back request is stalled, it will not block the front pipeline until the cache unit Buffer is filled. In an embodiment, the memory access system of the present application further comprises a first multi-selector for a in b. The first multi-selector selects a in b write-back requests that pass the arbitration for data write-back, where a represents the product of the first preset number and the second preset number, and b represents the first preset number. For example, the first multi-selector is a 64-to-16 multi-selector Crossbar. The write-back requests in the cache unit Buffer pass through the 64-to-16 multi-selector Crossbar, and 16 write-back requests without Bank conflict are selected for data write-back on the memory.
[0068] The following example of out-of-order write-back of a 4-bank SRAM will be used to further illustrate data write-back. Assume the data to be written back is a 4x4 two-dimensional array, and its address layout in the SRAM is as follows... Figure 4 As shown. Assuming the processor's data access method is cross-row access, meaning the same cycle accesses different rows of the array, then the access process in each cycle is as follows: Figures 5(a)-5(c) As shown.
[0069] like Figures 5(a)-5(c) As shown in Figure 5(a), access to the 4x4 two-dimensional array is divided into 8 cycles. In Cycle 0, the first group of write-back requests (A0, B0, C0, D0) is loaded into the buffer. In Cycle 1, due to a bank conflict, only A0 passes arbitration. Since the buffer is not full, subsequent requests (A1, B1, C1, D1) can still enter the buffer without being blocked. In Cycle 1, the remaining (B0, C0, D0) of the first group and the second group (A1, B1, C1, D1) participate in arbitration together, and B0 and A1 pass. In Cycle 2, as shown in Figure 5(b), the remaining (C0, D0) of the first group, the second group (B1, C1, D1), and the third group (A2, B2, C2, D2) participate in arbitration together, and C0 and B1 / A2 pass. Cycle 6 and Cycle 7 follow the same pattern, as shown in Figure 5(c).
[0070] Regarding write-back efficiency, without the method described in this application, each request would require 4 cycles due to bank conflicts, totaling 16 cycles. Therefore, the write-back efficiency of this application is twice that of the traditional approach. In fact, the first 4 cycles of this application represent the initial buffer filling process, which can be considered a latency. If there are more memory access requests after Cycle 4, bandwidth utilization can be increased to 100%, completely eliminating bank conflicts, and the write-back efficiency is four times that of the traditional approach.
[0071] like Figure 6 As shown, in one embodiment, for data reading, the memory access system of this application further includes a first preset number of address cache units (AddrBuffer) and a second arbitration unit (Arbiter2). The first preset number is the number of banks contained in the memory. Preferably, this application provides 16 address cache units, namely AddrBuffer0-AddrBuffer15. The depth of each address cache unit (AddrBuffer) is a second preset number. Preferably, the second preset number is 4, that is, each address cache unit can store 4 read addresses simultaneously.
[0072] The following is a description of data reading in the above configuration. In the data reading process of the present application, there are three stages: sequential application, out-of-order response, and sequential submission.
[0073] In the sequential application stage, each read address generated by the read address generation unit corresponds to an address buffer unit and a read order. For example, read address Add0 corresponds to address buffer unit AddrBuffer0 and read order Order0; read address Add1 corresponds to address buffer unit AddrBuffer1 and read order Order1; and so on. The read address and the corresponding read order constitute a read request, which is written into the corresponding address buffer unit. Among the 16 address buffer units AddrBuffer, there are at most 64 read requests stored.
[0074] In the out-of-order response stage, the second arbitration unit detects whether there is a Bank conflict among the read requests in the 16 address buffer units AddrBuffer, and arbitrates when a Bank conflict occurs. The second arbitration unit uses a priority arbitration algorithm. The read request that enters the address buffer unit AddBuffer earliest has the highest priority, and in each cycle, a read request with the highest priority is selected from multiple read requests competing for the same Bank to be responded to, while other read requests are temporarily buffered in the address buffer unit and are requested again in the next cycle. For the read request that passes the arbitration, 16 read requests are selected to obtain read data from the corresponding read addresses in the memory. For the read request that does not pass the arbitration, the read request is temporarily stored in the corresponding address buffer unit and continues to participate in arbitration in the next cycle. At the same time, since the depth of each address buffer unit AddrBuffer is 4, at most 4 read requests can be stored, so even if the previous read request is stalled, it will not block the front pipeline until the address buffer unit AddrBuffer is filled. In an embodiment, the memory access system of the present application further includes a second multiplexer for a in b. The second multiplexer selects a in b read requests that pass the arbitration for data reading, where a represents the product of the first preset number and the second preset number, and b represents the first preset number. For example, the second multiplexer is a 64-to-16 multiplexer Crossbar. The read requests in the address buffer unit AddrBuffer pass through the 64-to-16 multiplexer Crossbar to select 16 read requests without Bank conflict for data reading on the corresponding Bank of the memory.
[0075] In the sequential submission stage, the read data is output to the computing unit based on the read order corresponding to the read data. In an embodiment, the memory access system of the present application further comprises a distributor and a first preset number of data buffer units DataBuffer for realizing the ordered rearrangement of the read data. Each data buffer unit corresponds to an address buffer unit. Preferably, 16 data buffer units, i.e. DataBuffer0-DataBuffer15, are provided in the present application. The distributor is a 16-to-64 distributor Crossbar. The distributor writes the corresponding read data into the data buffer units according to the read order. The data buffer units store the read data provided by the corresponding Bank and output the corresponding read data to the computing unit in turn according to the read order.
[0076] In an embodiment, a counter is provided in the data buffer unit, which is used to record the current output read order. The read data is output to the computing unit simultaneously only when all the data buffer units have obtained the read data corresponding to the current output read order. For example, for the data buffer unit DataBuffer0, assuming that the read request of read order Order 1 is responded by the memory before the read request of read order Order 0 in the address buffer unit AddrBuffer0, the read data of read order Order 1 is written into the address 1 of the data buffer unit DataBuffer0 first. At this time, the data buffer unit DataBuffer0 detects that the read data of read order Order 0 has not arrived, and therefore does not output the read data of read order Order 1, but continues to wait for the read data of read order Order 0. When the read data of read order Order 0 arrives, it is written into the address 0 of the data buffer unit DataBuffer0. At this time, the data buffer unit DataBuffer0 detects that the read data of the most leading order has arrived, and then checks whether all the data buffer units DataBuffer have collected the read data of the most leading order. If yes, the read data is output simultaneously, and if not, the read data continues to be waited for.
[0077] The following further describes the data read by taking the out-of-order read of the SRAM comprising 4 Banks as an example. Assuming that the data to be read is a 4x4 two-dimensional array, the address layout of which in the SRAM is as shown in Figure 4 . Assuming that the data read mode of the processor is cross-row read, i.e. accessing different rows of the array in the same cycle, the access process in each cycle is as shown in Figures 7(a)-7(c) .
[0078] As shown in Figures 7(a)-7(c)As shown, the reading of the 4x4 two-dimensional array takes 10 cycles Cycle. As shown in FIG. 7(a), Cycle 0 loads the first set of read requests (A0, B0, C0, D0) into the AddrBuffer. At Cycle 1, due to bank conflict, only A0 passes the arbitration. Since the AddrBuffer is not full, the following (A1, B1, C1, D1) requests can still enter the AddrBuffer and will not be blocked. At Cycle 1, the remaining (B0, C0, D0) of the first set and the (A1, B1, C1, D1) of the second set participate in arbitration together, and the results are that B0 and A1 pass the arbitration, and the SRAM returns the data A0, which is reordered by the Crossbar and enters the DataBuffer. As shown in FIG. 7(b), at Cycle 2, the remaining (C0, D0) of the first set and the (B1, C1, D1) of the second set and the (A2, B2, C2, D2) of the third set participate in arbitration together, and the results are that C0 and B1 A2 pass the arbitration, and the SRAM returns the data B0, A1, which is reordered by the Crossbar and enters the DataBuffer. In this way, as shown in FIG. 7(c), at Cycle 6, the data (A0, B0, C0, D0) corresponding to the first read request is collected in the DataBuffer and can be output together. At Cycle 7, the data (A1, B1, C1, D1) corresponding to the second read request is collected in the DataBuffer and can be output together, and so on. In terms of read efficiency, if the scheme of the present application is not used, each set of requests obviously needs 4 cycles due to bank conflict, and a total of 16 cycles are needed, so the write-back efficiency of the present application is 1.6 times that of the traditional scheme. In fact, the first 4 cycles of the present application are the initial filling process of the Buffer, which can be regarded as Latency. At Cycle 4 and thereafter, if there are more memory access requests, the bandwidth utilization rate can be improved to 100%, completely eliminating bank conflict, and the read efficiency is 4 times that of the traditional scheme.
[0079] Therefore, by using the memory access system of the present application, the bandwidth utilization rate can be effectively improved. By mapping the memory access process of convolution to the cache architecture of the present application, the corresponding number of access cycles can be obtained, and then the bandwidth utilization rate can be obtained. Among them, the bandwidth utilization rate = ideal cycle number / actual cycle number.
[0080] Several typical convolution layers in MobileNet are used as test cases, and the parameters of the convolution layers are shown in Table 1.
[0081] Table 1, test case convolution layer size parameters
[0082]
[0083] It is tested that the utilization rate of all convolution layers is greatly improved by 2-8 times. This shows that the cache architecture proposed in the application has a significant effect on eliminating Bank conflict. The more serious the memory conflict of the convolution layer, the greater the performance improvement. In particular, the bandwidth utilization rate of all convolution layers is close to 100%, which shows that the cache architecture proposed in the application has strong universality.
[0084] As shown in Figure 8 In an embodiment, the processor of the application includes the memory access system described above, thereby effectively improving the performance of the processor. The processor includes an NPU (Neural Network Processing Unit), a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a VPU (Video Processing Unit), a DPU (Display Unit), a DSP (Digital Signal Processor), etc. The processor effectively reduces Bank conflict through the memory access system described above, achieves high bandwidth utilization, ensures the correctness of data access, and effectively improves the performance of the processor.
[0085] As shown in Figure 9 The memory access method of the application includes steps S91-S94.
[0086] Step S91, generating a read address based on a read address generation unit.
[0087] Step S92, based on the memory, providing corresponding read data to the calculation unit in a way of out-of-order response and in-order submission to avoid Bank conflict.
[0088] Step S93, generating a write-back address based on a write-back address generation unit.
[0089] Step S94, based on the calculation unit, generating write-back data according to the read data, and writing the write-back data into the memory in an out-of-order response manner according to the write-back address to avoid Bank conflict.
[0090] The protection scope of the memory access method of the embodiments of the application is not limited to the execution order of the steps listed in the embodiments. Any scheme achieved by adding, replacing or replacing steps of the prior art according to the principles of the application is included in the protection scope of the application.
[0091] The embodiments of the present application further provide a computer readable storage medium. A person skilled in the art can understand that all or part of the steps of the memory access method described above can be completed by a program instructing a processor, and the program can be stored in a computer readable storage medium, which is a non-transitory medium, such as a random access memory, a read only memory, a flash memory, a hard disk, a solid state disk, a magnetic tape, a floppy disk, an optical disc and any combination thereof. The storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, a data center and the like, which includes one or more available medium sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disc (DVD)) or a semiconductor medium (for example, a solid state disk (SSD)) and the like.
[0092] The embodiments of the present application further provide an electronic device. The electronic device comprises a processor and a memory.
[0093] The memory is configured to store a computer program.
[0094] The memory comprises a ROM, a RAM, a disk, a U disk, a memory card or an optical disc and the like various medium capable of storing program codes.
[0095] The processor is connected with the memory and is configured to execute the computer program stored in the memory, so that the electronic device executes the memory access method described above.
[0096] Preferably, the processor can be a general processor, including a central processing unit (CPU), a network processor (NP) and the like; and can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0097] As Figure 10As shown, the electronic device of this application is embodied in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors or processing units 101, memory 102, and bus 103 connecting different system components (including memory 102 and processing unit 101).
[0098] Bus 103 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0099] Electronic devices typically include a variety of computer-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, and removable and non-removable media.
[0100] Memory 102 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 1021 and / or cache memory 1022. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 1023 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 10 Not shown; usually referred to as a "hard drive"). Although Figure 10 As not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 103 via one or more data media interfaces. Memory 102 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.
[0101] A program / utility 1024 having a set (at least one) of program modules 10241 may be stored, for example, in memory 102. Such program modules 10241 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 10241 typically perform the functions and / or methods described in the embodiments of this application.
[0102] The electronic device can also communicate with one or more external devices (e.g., keyboard, pointing device, display, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through input / output (I / O) interface 104. Furthermore, the electronic device can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) through network adapter 105. Figure 10 As shown, network adapter 105 communicates with other modules of the electronic device via bus 103. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0103] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A memory access system, characterized in that, It includes a write-back address generation unit, a calculation unit, a read address generation unit, and a memory; The read address generation unit is used to generate a read address; The memory is used to provide the corresponding read data to the computing unit according to the read address, through out-of-order response and sequential submission to avoid Bank conflicts; The write-back address generation unit is used to generate write-back addresses; The computing unit is used to generate write-back data based on the read data, and write the write-back data into the memory in an out-of-order response manner according to the write-back address to avoid Bank conflicts.
2. The memory access system according to claim 1, characterized in that: The system further includes a first preset number of cache units and a first arbitration unit; the first preset number is the number of banks contained in the memory; Each write-back address generated by the write-back address generation unit corresponds to a cache unit; Each write-back data generated by the computing unit corresponds to a cache unit; Write the write-back address and the write-back data into the corresponding cache unit to form a write-back request; The first arbitration unit detects whether there is a Bank conflict among the write-back requests. When a Bank conflict occurs, it arbitrates and writes the corresponding write-back data into the memory according to the corresponding write-back address based on the first preset number of write-back requests that pass the arbitration.
3. The memory access system according to claim 2, characterized in that: Each cache unit can store a second preset number of write-back requests.
4. The memory access system according to claim 3, characterized in that: The system further includes a first multi-selector for selecting b from a, wherein the first multi-selector selects the first preset number of write-back requests that have passed arbitration for data write-back, where a represents the product of the first preset number and the second preset number, and b represents the first preset number.
5. The memory access system according to claim 2, characterized in that: When the write-back request fails arbitration, the write-back request is temporarily stored in the corresponding cache unit and will participate in arbitration in the next cycle.
6. The memory access system according to claim 1, characterized in that: The system further includes a first preset number of address cache units and a second arbitration unit; the first preset number is the number of banks contained in the memory; Each read address generated by the read address generation unit corresponds to an address cache unit and a read order; The read address and the corresponding read order constitute a read request, and the read request is written to the corresponding address cache unit; The second arbitration unit detects whether there is a Bank conflict between the read requests, arbitrates when a Bank conflict occurs, and obtains read data from the corresponding read address in the memory based on the first preset number of read requests that have passed the arbitration. Based on the reading order corresponding to the read data, the read data is output to the computing unit.
7. The memory access system according to claim 6, characterized in that: The address cache unit can store a second preset number of read requests.
8. The memory access system according to claim 7, characterized in that: The system further includes a second multi-selector for selecting b from a, wherein the second multi-selector selects the first preset number of read requests that have passed arbitration for data reading, where a represents the product of the first preset number and the second preset number, and b represents the first preset number.
9. The memory access system according to claim 7, characterized in that: The system also includes an allocator for selecting a from b and a first preset number of data cache units; wherein each data cache unit corresponds to an address cache unit, a represents the product of the first preset number and the second preset number, and b represents the first preset number; The allocator writes the corresponding read data into the data cache unit according to the read order; The data caching unit stores the read data provided by the corresponding Bank, and outputs the corresponding read data to the computing unit in the order of read.
10. The memory access system according to claim 8, characterized in that: The data caching unit is equipped with a counter, which is used to record the current output reading order; the reading data is output to the computing unit only when all data caching units have obtained the reading data corresponding to the current output reading order.
11. The memory access system according to claim 7, characterized in that: When the read request fails the arbitration, the read request is temporarily stored in the corresponding address cache unit and will participate in the arbitration in the next cycle.
12. A processor, characterized in that: Includes the memory access system as described in any one of claims 1-11.
13. A memory access method, characterized in that, The method includes the following steps: The read address is generated based on the read address generation unit; Based on the memory, the corresponding read data is provided to the computing unit according to the read address through out-of-order response and sequential submission to avoid Bank conflicts; The write-back address is generated based on the write-back address generation unit; The computing unit generates write-back data based on the read data, and writes the write-back data into the memory using an out-of-order response method according to the write-back address to avoid Bank conflicts.
14. An electronic device, characterized in that, The electronic device includes: a processor and a memory; The memory is used to store computer programs; The processor is configured to execute a computer program stored in the memory to cause the electronic device to perform the memory access method as described in claim 13.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by an electronic device, it implements the memory access method of claim 13.