CNN data access method for FPGA
By optimizing the feature map fragmentation and BRAM of the FPGA platform and adopting a staggered caching strategy, the problem of low data transmission bandwidth utilization on the FPGA platform was solved, achieving efficient data access and improving system performance.
Patent Information
- Application Number
- CN202511032781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
Existing CNN high-parallel mapping methods fail to maximize the use of data transmission bandwidth on FPGA platforms, resulting in full utilization of computing resources but insufficient data supply, which limits system performance.
By segmenting feature maps and optimizing the number and layout of BRAMs, data is interleaved into different BRAMs. A staggered caching strategy is adopted to ensure that data is not stored in the same BRAM, thus designing a CNN data access method for FPGA.
Data is read out within one clock cycle, maximizing bandwidth utilization, resolving read address conflicts, and improving data transmission efficiency.
Smart Images

Figure CN120931475A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of CNN data access technology, and in particular to a CNN data access method for FPGA. Background Technology
[0002] In recent years, Convolutional Neural Networks (CNNs) have achieved remarkable results in fields such as image recognition and object detection. Their widespread application has driven the continuous pursuit of high performance and low power consumption in computing platforms. Field-Programmable Gate Arrays (FPGAs), due to their flexibility, low power consumption, and high parallelism, have become an important hardware platform for accelerating CNN computation. However, in practical applications, how to efficiently utilize FPGA resources to meet the high concurrency requirements of CNNs for computation and data transmission remains a pressing technical challenge.
[0003] Existing high-parallel mapping methods for CNNs primarily accelerate computation by dividing the network's computational tasks into numerous parallel processing units, fully leveraging the parallel processing capabilities of FPGAs. However, traditional data access strategies fail to simultaneously improve bandwidth utilization. Due to bottlenecks in data transfer between internal and external FPGA memory, high-parallel processing units often face the risk of insufficient data supply, thus limiting the overall system performance. Specifically, in high-parallel mapping, although computational resources are fully utilized, limitations in data access patterns and memory scheduling strategies prevent the maximization of data transfer bandwidth. This phenomenon is particularly pronounced when processing large-scale CNN models. Summary of the Invention
[0004] The purpose of this invention is to provide a CNN data access method for FPGA, optimize the CNN data access strategy on the FPGA platform, and break through the performance bottleneck of existing high-parallel mapping methods by maximizing bandwidth utilization.
[0005] To achieve the above objectives, this invention provides a CNN data access method for FPGA, comprising the following steps:
[0006] S1. Input the feature map and divide the feature map data into pieces, each piece including several feature maps;
[0007] S2. Determine the number of BRAMs in the CNN and optimize them;
[0008] S3. Select the first feature map group and stitch together all the feature pixel data in the first feature map group;
[0009] S4. Process the feature map groups using the same method as S3 until the feature pixel data of all feature map groups are stitched together.
[0010] S5. Interweave all the spliced feature pixel data together in the order of the segments.
[0011] Preferably, the process of S1 is as follows:
[0012] S11. Input all feature maps into DDR3;
[0013] S12, Parallelism N of the input parallel unit;
[0014] S13. Divide all feature maps into several pieces according to their parallelism, as follows:
[0015]
[0016] Where x represents the number of segments into which all feature maps are divided. If x is not an integer, the total number of feature maps is padded with zeros until x is an integer.
[0017] S14. Based on S13, divide all feature maps into x pieces, and the number of feature maps in each piece is the same as the parallelism value, which is N.
[0018] Preferably, the process of S2 is as follows:
[0019] S21. Set up N BRAMs in the CNN according to the parallelism N;
[0020] S22, Write the burst access data of DDR3 into different BRAMs;
[0021] S23. The data that needs to be interleaved is cached in different BRAMs in a staggered manner to ensure that the data that needs to be interleaved is not in the same BRAM, thus completing the optimization.
[0022] Preferably, the process of S3 is as follows:
[0023] S31. Select the first feature map group from DDR3;
[0024] S32. Cache the feature pixel data of the first row of each feature map into N BRAMs in sequence;
[0025] S33. Simultaneously extract the first feature pixel data of the first row through the input channels of N BRAMs;
[0026] S34. Concatenate the N feature pixel data extracted in S33 in sequence;
[0027] S35. Repeat S33-S34 until all feature pixel data in the first row of N BRAMs are concatenated separately, and the concatenated feature pixel data are interwoven in sequence to complete the interleaving of the first row.
[0028] S36. Repeat steps S32-S35 for the rows after N feature maps in the first feature map group until all rows are interleaved;
[0029] S37. Join all the interlacing lines in sequence to complete the first piece of interlacing.
[0030] Preferably, the process of S4 is as follows:
[0031] S41. Select the second feature map group;
[0032] S42. Use the same method as S32-S35 to complete the interleaving of the feature maps in the second feature map group;
[0033] S43. Repeat steps S41-S42 for the feature map groups of subsequent slices until all slices have been interleaved.
[0034] Preferably, the process of S5 is as follows:
[0035] S51. Join the interweaving of the first piece with the interweaving of the second piece, and attach the second piece behind the first piece;
[0036] S52. Join the interlacing of the third piece to the end of the second piece that is furthest from the first piece;
[0037] S53. All interleaving from the fourth piece to the xth piece is spliced using the method in S52 to complete the interleaving of all pieces and complete the storage and retrieval of feature map data.
[0038] Preferably, the optimized BRAM can simultaneously cache the data that needs to be interleaved into different BRAMs in a staggered manner, ensuring that the data that needs to be interleaved is not in the same BRAM.
[0039] Preferably, the number of BRAMs is designed according to the parallelism of the parallel units in the CNN.
[0040] Therefore, this invention adopts the above-mentioned structure to provide a CNN data access method for FPGA. It designs a CNN data access method for FPGA platform, which can maximize bandwidth utilization during high parallel processing. Furthermore, it uses an optimized BRAM to solve the read address conflict problem and can complete the reading of interleaved data within one clock cycle.
[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0042] Figure 1 This invention describes the overall process of a CNN data access method for FPGA.
[0043] Figure 2This invention describes the interleaving storage process of feature pixel data in a sliced feature map group for a CNN data access method for FPGA.
[0044] Figure 3 (a) is a diagram showing how to write N consecutive parallel data into one BRAM; (b) is a diagram showing how to split N consecutive parallel data into N BRAMs and write them into N BRAMs respectively; (c) is a diagram showing how to simultaneously cache the data that needs to be interleaved into different BRAMs in a staggered manner. Detailed Implementation
[0045] Example
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0047] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0048] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0049] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed when in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0050] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," and "connect" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0051] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0052] like Figures 1-3 As shown, the present invention provides a CNN data access method for FPGA, comprising the following steps:
[0053] S1. Input the feature map and divide the feature map data into pieces, each piece including several feature maps;
[0054] S11. Input all feature maps into DDR3;
[0055] S12, Parallelism N of the input parallel unit;
[0056] S13. Divide all feature maps into several pieces according to their parallelism, as follows:
[0057]
[0058] Where x represents the number of segments into which all feature maps are divided. If x is not an integer, the total number of feature maps is padded with zeros until x is an integer.
[0059] S14. Based on S13, divide all feature maps into x pieces, and the number of feature maps in each piece is the same as the parallelism value, which is N.
[0060] S2. Determine the number of BRAMs in the CNN and optimize them;
[0061] S21. Set up N BRAMs in the CNN according to the parallelism N;
[0062] S22, Write the burst access data of DDR3 into different BRAMs;
[0063] S23. The data that needs to be interleaved is cached in different BRAMs in a staggered manner to ensure that the data that needs to be interleaved is not in the same BRAM, thus completing the optimization.
[0064] S3. Select the first feature map group and concatenate the feature pixel data of N feature maps in the first feature map group together.
[0065] S31. Select the first feature map group from DDR3;
[0066] S32. Cache the feature pixel data of the first row of each feature map into N BRAMs in sequence;
[0067] S33. Simultaneously extract the first feature pixel data of the first row through the input channels of N BRAMs;
[0068] S34. Concatenate the N feature pixel data extracted in S33 in sequence;
[0069] S35. Repeat S33-S34 until all feature pixel data in the first row of N BRAMs are concatenated separately, and the concatenated feature pixel data are interwoven in sequence to complete the interleaving of the first row.
[0070] S36. Repeat steps S32-S35 for the rows after N feature maps in the first feature map group until all rows are interleaved;
[0071] S37. Join all the interlacing lines in sequence to complete the first piece of interlacing.
[0072] S4. Process the feature map groups using the same method as S3 until the feature pixel data of all feature map groups are stitched together.
[0073] S41. Select the second feature map group;
[0074] S42. Use the same method as S32-S35 to complete the interleaving of the feature maps in the second feature map group;
[0075] S43. Repeat steps S41-S42 for the feature map groups of subsequent slices until all slices have been interleaved.
[0076] S5. Interweave all the spliced feature pixel data together in the order of the segments.
[0077] S51. Join the interweaving of the first piece with the interweaving of the second piece, and attach the second piece behind the first piece;
[0078] S52. Join the interlacing of the third piece to the end of the second piece that is furthest from the first piece;
[0079] S53. All interleaving from the fourth piece to the xth piece is spliced using the method in S52 to complete the interleaving of all pieces and complete the storage and retrieval of feature map data.
[0080] Outputting parallel data from multiple adjacent consecutive memory cells within a single clock cycle leads to an imbalance in internal BRAM read / write bandwidth when the data is interleaved through N BRAMs within the FPGA. Traditionally, there are two ways to cache consecutive parallel data into N BRAMs within the FPGA: one is to write N consecutive parallel data into a single BRAM, and the other is to divide the N consecutive parallel data into N data segments and write them into the N BRAMs respectively. Regardless of the method used, the problem of uneven read / write access efficiency arises. A detailed analysis follows.
[0081] The method of writing N consecutive parallel data into a BRAM is as follows: Figure 3 As shown in (a), assuming data access occurs at a 200MHz clock, and each data point has a bit width of 1 byte. When parallel data is written to a BRAM sequentially, it takes multiple clock cycles to write to the BRAM, resulting in a write instantaneous bandwidth of 200MB / s. During reading, data can be read from multiple BRAMs simultaneously, so the read instantaneous bandwidth is N*200MB / s. The read bandwidth is N times the write bandwidth, making the BRAM's write bandwidth the bottleneck.
[0082] The method of splitting N consecutive parallel data and writing them into N BRAMs is as follows: Figure 3 As shown in (b), when parallel data is split and written to different BRAMs, it can be written in one clock cycle with an instantaneous bandwidth of N*200MB / s. However, when reading out the data to be interleaved, since the interleaved data is stored in the same BRAM, an address conflict occurs, and multiple interleaved data cannot be read and written at the same time. It takes multiple clock cycles to complete the data reading, and the instantaneous bandwidth of the BRAM is 200MB / s. At this time, the instantaneous read bandwidth of the BRAM becomes the bottleneck.
[0083] Therefore, both methods result in uneven read / write access efficiency. To address this, this invention proposes a method to optimize BRAM by writing burst access data from DDR3 into different BRAMs, while simultaneously staggering the data that needs to be interleaved into different BRAMs. This ensures that the data requiring interleaving is not in the same BRAM, such as... Figure 3 As shown in (c), the instantaneous write bandwidth is N*200MB / s, and the instantaneous read bandwidth is N*200MB / s. This resolves the read address conflict issue and allows interleaved data to be read within one clock cycle.
[0084] The number of BRAMs is designed based on the parallelism of the parallel units in the CNN.
[0085] In this embodiment, a hardware platform consisting of a Xilinx xc7vx690t FPGA and an MT41K256M16 DDR3 SDRAM is selected. The internal DDR access data width is 128 bits, and the clock speed is 200MHz. Using a traditional data read scheme, the bandwidth is only 1.2GB / s, and the bandwidth utilization rate is only 38.4%. However, using the data interleaving storage method proposed in this invention, the bandwidth reaches a rate of 2.8GB / s, and the bandwidth utilization rate reaches 89.6%.
[0086] Therefore, this invention adopts the above-mentioned structure to provide a CNN data access method for FPGA. It designs a CNN data access method for FPGA platform, which can maximize bandwidth utilization during high parallel processing. Furthermore, it uses an optimized BRAM to solve the read address conflict problem and can complete the reading of interleaved data within one clock cycle.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A CNN data access method for FPGA, characterized in that, Includes the following steps: S1. Input the feature map and divide the feature map data into pieces, each piece including several feature maps; S2. Determine the number of BRAMs in the CNN and optimize them; S3. Select the first feature map group and stitch together all the feature pixel data in the first feature map group; S4. Process the feature map groups using the same method as S3 until the feature pixel data of all feature map groups are stitched together. S5. Interweave all the spliced feature pixel data together in the order of the segments.
2. The CNN data access method for FPGA according to claim 1, characterized in that, The process of S1 is as follows: S11. Input all feature maps into DDR3; S12, Parallelism N of the input parallel unit; S13. Divide all feature maps into several pieces according to their parallelism, as follows: Where x represents the number of segments into which all feature maps are divided. If x is not an integer, the total number of feature maps is padded with zeros until x is an integer. S14. Based on S13, divide all feature maps into x pieces, and the number of feature maps in each piece is the same as the parallelism value, which is N.
3. The CNN data access method for FPGA according to claim 1, characterized in that, The process of S2 is as follows: S21. Set up N BRAMs in the CNN according to the parallelism N; S22, Write the burst access data of DDR3 into different BRAMs; S23. The data that needs to be interleaved is cached in different BRAMs in a staggered manner to ensure that the data that needs to be interleaved is not in the same BRAM, thus completing the optimization.
4. The CNN data access method for FPGA according to claim 3, characterized in that, The process of S3 is as follows: S31. Select the first feature map group from DDR3; S32. Cache the feature pixel data of the first row of each feature map into N BRAMs in sequence; S33. Simultaneously extract the first feature pixel data of the first row through the input channels of N BRAMs; S34. Concatenate the N feature pixel data extracted in S33 in sequence; S35. Repeat S33-S34 until all feature pixel data in the first row of N BRAMs are concatenated separately, and the concatenated feature pixel data are interwoven in sequence to complete the interleaving of the first row. S36. Repeat steps S32-S35 for the rows after N feature maps in the first feature map group until all rows are interleaved; S37. Join all the interlacing lines in sequence to complete the first piece of interlacing.
5. The CNN data access method for FPGA according to claim 4, characterized in that, The process of S4 is as follows: S41. Select the second feature map group; S42. Use the same method as S32-S35 to complete the interleaving of the feature maps in the second feature map group; S43. Repeat the steps S41-S42 for the feature map groups of subsequent slices until all slices have been interleaved.
6. The CNN data access method for FPGA according to claim 5, characterized in that, The process of S5 is as follows: S51. Join the interweaving of the first piece with the interweaving of the second piece, and attach the second piece behind the first piece; S52. Join the interlacing of the third piece to the end of the second piece that is furthest from the first piece; S53. All interleaving from the fourth piece to the xth piece is spliced using the method in S52 to complete the interleaving of all pieces and complete the storage and retrieval of feature map data.
7. A CNN data access method for FPGA according to claim 6, characterized in that: The optimized BRAM can simultaneously cache the data that needs to be interleaved into different BRAMs in a staggered manner, ensuring that the data that needs to be interleaved is not in the same BRAM.
8. The CNN data access method for FPGA according to claim 7, characterized in that: The number of BRAMs is designed based on the parallelism of the parallel units in the CNN.