Data Remapping Strategy for Efficient BRAM Access Based on Distributed Storage

By adopting a distributed storage-based data remapping strategy in FPGAs, the BRAM storage space is remapping, the problem of BRAM multi-port access conflict is solved, efficient access and resource conservation are achieved, and system performance is improved.

CN114356801BActive Publication Date: 2025-06-24SOUTHEAST UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210014828.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-07
Publication Date
2025-06-24
Estimated Expiration
2042-01-07

AI Technical Summary

Technical Problem

The dual-port access feature of BRAM in FPGAs leads to conflict problems during multi-port access, affecting system performance. The existing technologies such as copying BRAM and joining XOR storage areas have problems such as wasting storage space and performance degradation.

Method used

The data remapping strategy based on distributed storage is adopted, and the BRAM storage space is remapped through the N-dimensional linear index and continuity principles, so that 2n to be accessed data are mapped into 2n-1 data areas without repeated repetition. Each data area contains two to be accessed data, and single-cycle efficient access is achieved using the BRAM dual-port feature.

Benefits of technology

It realizes efficient BRAM access, avoids data redundancy, saves BRAM resources, and in most cases, a single cycle of access to multiple address data simultaneously, thereby improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114356801B_ABST
    Figure CN114356801B_ABST
Patent Text Reader

Abstract

The present invention discloses a data remapping strategy for efficient access to BRAM based on distributed storage. Aiming at the problem of low read / write efficiency caused by limited read / write ports when accessing BRAM using traditional N-dimensional linear indexing, the present invention proposes an optimization strategy for realizing distributed storage of 2<supgt;N-1< / supgt; BRAM blocks based on the BRAM splitting strategy. This strategy first maps the data to be stored without repetition to 2<supgt;N-1< / supgt> BRAM data storage blocks according to the N-dimensional linear index of the BRAM access address. And maps the N-dimensional linear index of the original corresponding address for accessing BRAM to the addresses for accessing the corresponding data of 2<supgt;N-1< / supgt> BRAM blocks, thereby improving the access efficiency of BRAM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimizing memory access in FPGAs, and specifically relates to a data remapping strategy for realizing single-cycle efficient access of dual-port BRAMs based on distributed storage. Background Art

[0002] With the development of semiconductor technology, the number of logic gates in field-programmable gate arrays (FPGAs) has been continuously increasing. At the same time, the capacity of the block random-access memories (Block RAMs, BRAMs) in FPGAs has also been continuously rising, which means that a large amount of data to be accessed, such as trilinear interpolation, bilinear interpolation, and 3D-CNN, can be stored in BRAMs to improve the access speed and thus achieve the purpose of improving system performance. However, due to the characteristic that all commercially available FPGAs' BRAMs support at most dual-port access, there will be access conflict problems when accessing multiple data, and the access to multiple data needs to be divided into multiple access requests. Therefore, how to achieve efficient access to BRAMs has become one of the important means to improve system performance.

[0003] Currently, in the implementation of convolutional neural networks based on FPGAs, the problem of BRAM access conflicts has not been focused on, which means that longer access times will seriously slow down system performance. The strategies for efficient access to BRAMs in academia and industry mainly focus on two aspects: replicating BRAMs and adding an exclusive-OR storage area. For the strategy of replicating BRAMs, although this solution effectively realizes multi-port read and write, it only replicates the same data multiple times and stores them in BRAMs, which will undoubtedly cause a waste of a large amount of storage space; in addition, when the number of required access ports increases and the amount of stored data increases, the BRAM storage space will be quickly consumed. For the strategy of adding an exclusive-OR storage area, this solution divides the data into 4 data areas, each data area stores n data, and an exclusive-OR data area is added, and the value of this data area is obtained by exclusive-OR of the corresponding positions of the 4 data areas. Although this solution only adds an exclusive-OR data area to realize multi-port read and write, access conflict problems will inevitably occur when the amount of data to be accessed per cycle increases; at the same time, the added exclusive-OR data area will cause an overly long critical path and the write data cannot be completed within a single cycle, which will undoubtedly cause a decline in system performance. Summary of the Invention

[0004] The object of the present invention is to provide a data remapping strategy for realizing efficient access to BRAM based on distributed storage, so as to solve the problem of multi-port access BRAM conflict mentioned in the background technology. At the same time, this strategy will not introduce data redundancy, thus saving BRAM resources. The present invention is applicable to a general BRAM storage architecture with N-dimensional linear indexing and continuous access space for a single time. By remapping the BRAM storage space, the access efficiency of BRAM is improved on the premise of unchanged storage density.

[0005] To solve the above technical problems, the present invention proposes the following technical solutions:

[0006] A data remapping strategy for realizing efficient access to BRAM based on distributed storage, characterized in that the method includes the following steps:

[0007] Step S1, determine whether the current original data storage table is accessed through N-dimensional linear indexing, such as data being discretely arranged in a one-dimensional line, a two-dimensional matrix, or a three-dimensional cube, etc. Generally, the address access conforms to Addr = x n +a0*(x n-1 +a1*(x n-2 +a2*(x n-3 +a3*(…))));

[0008] Step S2, determine whether the data to be indexed follows the continuity principle, that is, the addresses of a certain dimension need to be continuously accessed at one time, and this address is denoted as m k and m k +1, m k ∈{x n , x n-1 , x n-2 ,......, x2, x1}, where the subscript k represents the k-dimensional address coordinate of the current data index. For N-dimensional linear indexing, it is necessary to continuously access {m n , m n-1 ,......, m2, m1}, {m n , m n-1 ,......, m2, 1+m1}, ……, {1+m n , 1+m n-1 ,......, 1+m2, m1}, {1+m n , 1+m n-1 ,......, 1+m2, 1+m1} corresponding to 2 n different addresses;

[0009] Step S3, remap any 2 n data to be accessed in Step S2 so that the 2 n data are mapped to 2 without repetitionn-1 In each data area, each data area contains two data to be accessed. The specific steps are as follows:

[0010] Step S301: Except for the nth dimension, any 2 n data points in Step S2 are grouped and arranged in a way of odd-even dispersion according to the corresponding dimensions, that is, within any group, x k are all even or all odd, k ∈ {n - 1, n - 2,..., 2, 1}. Assuming that '0' represents even and '1' represents odd, then {x n-1 , x n-2 ,......, x2, x1} can be divided into 2 n-1 groups in the form of {00…00, 00…01, 00…10,..., 11…10, 11…11} from the (n - 1)th dimension to the 1st dimension index value, that is, the parity of the address index from the (n - 1)th dimension to the 1st dimension within each group is determined and each group contains two of the 2 n data to be accessed.

[0011] Step S302: Traverse the entire N-dimensional data storage area in ascending order of Addr, and map the data corresponding to each Addr to 2 n-1 data storage areas according to the grouping method in Step S301.

[0012] Step S4: Construct a mapping from the index address Addr in Step S1 to the corresponding address in the corresponding data area in Step S3;

[0013] Step S5: Achieve the purpose of simultaneously accessing two data in each data area by the dual-port characteristic of BRAM, and achieve the purpose of efficiently accessing 2 n data in a single cycle in Step S2 according to the one-to-one mapping of the addresses in Step S4;

[0014] Further, the specific steps of Step S4 include:

[0015] Step S401: Assuming that '0' represents even and '1' represents odd, then 00…00, 00…01, ……, 11…10, 11…11, etc. correspond to the encodings of 2 n-1 data areas. Therefore, perform odd-even judgment on each bit of {x n-1 , x n-2 ,......, x2, x1} to locate which data area the currently accessed data exists in;

[0016] Step S402: Determine the current data {x n , x n-1 , x n-2,..., x2, x1} is the in-region offset address in the data area located in S401. Taking n = 3 as an example for illustration. Suppose the index address of step S1 is Addr = x3 + 65 * (x2 + 71 * x1), where x3 ∈ [0, 64], x2 ∈ [0, 70], x1 ∈ [0, 70], that is, the data is arranged in a 65 * 71 * 71 cube, and the data to be accessed {x2, x1} is {even, odd}. Thus, the address of the index address Addr in the corresponding data area is Generally, for the address of the N-dimensional index address Addr in the corresponding data area, it can be mapped to Addr1 = x n + a0 * (f(x n-1 ) + a1 * (f(x n-2 ) + a2 * (f(x n-3 ) + a3 * (...))). When x k is odd, When x k is even,

[0017] The beneficial effects of the present invention are:

[0018] 1. The present invention is applicable to all N-dimensional linear indexing methods, with high applicability and promotion potential.

[0019] 2. By remapping the data, compared with replicating BRAM and adding an XOR storage area, the present solution realizes the characteristic of zero data redundancy.

[0020] 3. Through the analysis of the characteristics of the data to be accessed, for the data continuously stored in the N-dimensional linear space that appears in most cases, the present solution can achieve simultaneous access to multiple address data in a single cycle, thus greatly improving the system performance. Description of the Drawings

[0021] Figure 1 is a schematic flowchart of a data remapping strategy for realizing efficient access to BRAM based on distributed storage proposed by the present invention.

[0022] Figure 2 is an original data layout diagram of a data remapping strategy for realizing efficient access to BRAM based on distributed storage proposed by the present invention.

[0023] Figure 3 is a distribution diagram of the data to be accessed of a data remapping strategy for realizing efficient access to BRAM based on distributed storage proposed by the present invention.

[0024] Figure 4 is a data remapping schematic diagram of a data remapping strategy for realizing efficient access to BRAM based on distributed storage proposed by the present invention.

[0025] Figure 5 This is the hardware design diagram corresponding to an embodiment of a data remapping strategy for efficient access to BRAM based on distributed storage proposed by the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Embodiment 1

[0028] Refer to Figures 1-5 , this embodiment provides a data remapping strategy for efficient access to BRAM based on distributed storage. For the sake of simplicity in description, this embodiment takes three-dimensional linear index access with N = 3 and data arranged in a three-dimensional space of 10 * 20 * 30 as an example for elaboration, where x ∈ [0, 9], y ∈ [0, 19], z ∈ [0, 29]:

[0029] Step A: According to Figure 1 the three-dimensional data arrangement rule and the linear index feature, rewrite the access address Addr = x + 10 * y + 200 * z as Addr = x + 10 * (y + 20 * z).

[0030] Step B: According to the data arrangement characteristics of the data to be indexed, determine whether the data to be indexed meets Figure 2 the area shown in red or Figure 2 the red area is a subset of the data to be indexed.

[0031] Specifically, data such as {(x, y, z), (x, y, z + 1), (x, y + 1, z), (x, y + 1, z + 1), (x + 1, y, z), (x + 1, y, z + 1), (x + 1, y + 1, z), (x + 1, y + 1, z + 1)} is a subset of the data to be indexed.

[0032] Step C: According to the parity characteristics of {y, z} at different data points, map one by one the data points arranged in a three-dimensional manner of 10 * 20 * 30 in Figure 2 to the corresponding data areas, which specifically includes the following steps:

[0033] Step C1: The data points are arranged in a way that they are scattered according to the parity of each dimension index except the x dimension. Assuming '0' represents even and '1' represents odd, then {y, z} can be divided into four groups in the form of {00, 01, 10, 11}, that is, the parity of each group's address index {y, z} is determined and each group contains two of the eight data points to be accessed.

[0034] Specifically, all data points are divided into four groups in the form of {y, z} being {odd, even}, {even, even}, {odd, odd}, {even, odd}. For a specific {y, z}, such as when {y, z} is {7, 8}, there are 10 x values corresponding to it.

[0035] Step C2: Traverse the entire three-dimensional data storage area in ascending order of Addr, and map the {x, y, z} corresponding to each Addr to four data storage areas according to the grouping method in step S301.

[0036] Specifically, traverse all Addrs in the order of x, y, z from low to high, and distribute all {y, z} to four data areas, namely four BRAMs, in the form of being scattered as {odd, even}, {even, even}, {odd, odd}, {even, odd}. The arrangement method in each data area is still in the form of x, y, z from low to high, that is, as shown in Figure 4 the three-dimensional arrangement form.

[0037] More specifically, for the data area where {y, z} is {odd, even}, first put the 10 points corresponding to {y, z} being {1, 0} into the target data area, and then put the 10 points corresponding to {3, 0} into the target data area, and so on until the 10 points corresponding to {19, 28} are put into the target data area. The other three data areas are filled in the same way until all the data points in the three-dimensional arrangement are traversed.

[0038] Up to this point, all data points have been mapped to four data storage areas without repetition. Corresponding to Figure 3 any data block to be accessed in the three-dimensional arrangement is evenly distributed to four data areas, and there are two data points to be accessed in each data area. Therefore, due to the characteristics of the BRAM dual port, store the four data areas in four BRAMs, thereby realizing Figure 3 the single-cycle fast access characteristic of 8 data points within the data block.

[0039] After the data remapping in step C, all data has been mapped to four data storage areas. Therefore, it is necessary to establish the address mapping from the original index address Addr = x + 10 * (y + 20 * z) to the address within the new index data area.

[0040] Step D: As Figure 5 shown, construct the address mapping from the index address Addr corresponding to different (x, y, z) in Step A to the corresponding data area in Step C, which specifically includes the following steps:

[0041] Step D1: Locate the data area according to the storage position (x, y, z) of the data to be accessed in the original data area, and determine the data area where the currently accessed data is located after mapping;

[0042] Specifically, perform odd-even judgment on {y, z} according to a method similar to that in Step C2 to achieve data area location, as Figure 5 shown. Assume that {y, z} is {odd, even}, then the target data area is stored in BRAM0; similarly, if {y, z} is {even, even}, the target data area is stored in BRAM1; if {y, z} is {odd, odd}, the target data area is stored in BRAM2; if {y, z} is {even, odd}, the target data is stored in BRAM3.

[0043] Step D2: Perform address indexing on the located data area in Step D1 according to the storage position (x, y, z) of the data to be accessed in the original data area, such as the address mapping logic Figure 5 shown;

[0044] Specifically, judge how many times the current data is written into the target BRAM according to the parity and specific values of {y, z}, and then determine the specific position of the currently accessed data in the target BRAM according to x.

[0045] More specifically, in this embodiment, assume that when {y, z} is {even, odd}, then the current data is the th time written into the target BRAM, and 10 data are written at a time. That is, the corresponding index address of the target BRAM is Generally, in this embodiment, the address mapping of the target BRAM can be performed according to the following formula.

[0046]

[0047] In summary, a data remapping strategy for realizing efficient access to BRAM based on distributed storage provided by this embodiment has the following benefits compared with the prior art:

[0048] 1. The present invention can be used for all 3D linear indexing methods and has high applicability.

[0049] 2. By remapping the data, compared with copying BRAM and adding an exclusive OR storage area, the present solution realizes the characteristic of zero data redundancy.

[0050] 3. Through the analysis of the characteristics of the data to be accessed, for the data continuously stored in the 3D linear space that appears in most cases, this solution can achieve simultaneous access to multiple address data in a single cycle, thus greatly improving the system performance.

[0051] Where the present invention is not described in detail, it is the well-known technology of those skilled in the art.

[0052] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in this technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A data remapping strategy for efficient access to BRAM based on distributed storage, characterized in that, including the following steps: Step S1, determine whether the current original data storage table is accessed through N-dimensional linear indexing; Step S2: Determine whether the data to be indexed follows the continuity principle, that is, it is necessary to continuously access the addresses of a certain dimension at one time, and this address is denoted as m k and m k +1, m k ∈{x n , x n-1 , x n-2 ,......, x2, x1}, where the subscript k represents the k-dimensional address coordinate of the current data index; for N-dimensional linear indexing, it is necessary to continuously access {m n , m n-1 ,......, m2, m1}, {m n , m n-1 ,......, m2, 1 + m1},......, {1 + m n , 1 + m n-1 ,......, 1 + m2, m1}, {1 + m n , 1 + m n-1 ,......, 1 + m2, 1 + m1} corresponding to 2 n different addresses; Step S3: Remap any two n of the to-be-accessed data in Step S2 so that the two n data are mapped to two n-1 data areas without repetition, and each data area contains two to-be-accessed data; Step S4, construct a mapping from the index address Addr in Step S1 to the corresponding address in the corresponding data area in Step S3; Step S5: Achieve the purpose of simultaneously accessing two data in each data area by the dual-port feature of BRAM, and achieve the purpose of efficiently accessing two data in a single cycle in Step S2 according to the one-to-one mapping of the addresses in Step S4 n data; In step S1, the address access conforms to Addr = x n + a0 * (x n-1 + a1 * (x n-2 + a2 * (x n-3 + a3 * (…)))); where {x n , x n-1 , x n-2 ,......, x2, x1} respectively correspond to the index addresses from low dimension to high dimension, and {a0, a1, a2,......, a n-3 , a n-2} respectively correspond to the linear mapping coefficients from the index addresses of each dimension to the storage address Addr.

2. The data remapping strategy for efficient BRAM access based on distributed storage according to claim 1, wherein Map a data storage space that conforms to N-dimensional linear indexing to 2 n-1 data regions through data remapping and store them in BRAM, so as to achieve single-cycle efficient access to the 2 n data that need to be continuously accessed in step S2.

3. A data remapping strategy for efficient access to BRAM based on distributed storage according to claim 1, characterized in that, The specific steps of Step S3 include the following steps: Step S301. Except for the n-th dimension, any two data points in Step S2 are grouped and arranged in the way of odd-even dispersion according to the corresponding dimension, that is, in any group, all x values are even or all are odd, where k ∈ {n - 1, n - 2,..., 2, 1}; assuming that '0' represents even and '1' represents odd, {x, x,..., x2, x1} can be divided into 2 groups in the form of {00…00, 00…01, 00…10,..., 11…10, 11…11} from the (n - 1)-th dimension to the 1st dimension index value, that is, the parity of the address index from the (n - 1)-th dimension to the 1st dimension within each group is determined and each group contains two of the two data to be accessed; n For any two data points in Step S2, except for the n-th dimension, they are grouped and arranged in the way of odd-even dispersion according to the corresponding dimension, that is, in any group, all x values are even or all are odd, where k ∈ {n - 1, n - 2,..., 2, 1}; assuming that '0' represents even and '1' represents odd, {x, x,..., x2, x1} can be divided into 2 groups in the form of {00…00, 00…01, 00…10,..., 11…10, 11…11} from the (n - 1)-th dimension to the 1st dimension index value, that is, the parity of the address index from the (n - 1)-th dimension to the 1st dimension within each group is determined and each group contains two of the two data to be accessed; k are all even or all odd, where k ∈ {n - 1, n - 2,..., 2, 1}; assuming that '0' represents even and '1' represents odd, {x, x,..., x2, x1} can be divided into 2 groups in the form of {00…00, 00…01, 00…10,..., 11…10, 11…11} from the (n - 1)-th dimension to the 1st dimension index value, that is, the parity of the address index from the (n - 1)-th dimension to the 1st dimension within each group is determined and each group contains two of the two data to be accessed; n-1 , x n-2 ,......, x2, x1} from the (n - 1)-th dimension to the 1st dimension index value in the form of {00…00, 00…01, 00…10,..., 11…10, 11…11} into 2 groups, that is, the parity of the address index from the (n - 1)-th dimension to the 1st dimension within each group is determined and each group contains two of the two data to be accessed; n-1 groups, that is, the parity of the address index from the (n - 1)-th dimension to the 1st dimension within each group is determined and each group contains two of the two data to be accessed; n of the two data to be accessed; Step S302: Traverse the entire N-dimensional data storage area in ascending order of Addr, and map the data corresponding to each Addr to 2 n-1 data storage areas according to the grouping method in Step S301.

4. A data remapping strategy for efficient access to BRAM based on distributed storage according to claim 1, characterized in that The specific steps of Step S4 include: Step S401: Represent even numbers with '0' and odd numbers with '1'. Then, 00…00, 00…01, ……, 11…10, 11…11, etc. correspond to the encodings of 2 n-1 data areas. Therefore, perform parity checks on each bit of {x n-1 , x n-2 ,......, x2, x1} to locate in which data area the currently accessed data exists; Step S402: Determine the in-zone offset address of the current data {x n , x n-1 , x n-2 ,......, x2, x1} in the data area located in S401.

5. A data remapping strategy for efficient access to BRAM based on distributed storage according to claim 4, characterized in that In step S402, the address of the N - dimensional index address Addr in the corresponding data area can be mapped to Addr1 = x n + a0 * (f(x n-1 ) + a1 * (f(x n-2 ) + a2 * (f(x n-3 ) + a3 * (…)))) ; When x k is odd, When x k is even, where {x n , f(x n-1 ), f(x n-2 ),......, f(x2), f(x1)} are respectively the index addresses from the low - dimension to the high - dimension in the corresponding data area, and {a0, a1, a2,......, a n-3 , a n-2} are respectively the linear mapping coefficients from each - dimension index address to the storage address Addr.

Citation Information

Patent Citations

  • Interlaced even and odd address mapping

    US20070050593A1