Hardware acceleration device for high-concurrency hash addressing
In the hash addressing process of Instant-NGP, dynamic random access memory and multi-layer perceptron units are used to work together, and appropriate memory access mode is selected according to data bandwidth and memory access requirements, and the data in the hash table storage unit is rearranged, which solves the problems of frequent memory access and memory bank conflicts and improves system performance.
Patent Information
- Application Number
- CN202510376894.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
During the hash addressing process of Instant-NGP, frequent memory access and bank conflict problems lead to limited performance improvement, especially the bank conflict problems caused by irregular memory access of the hash table affect the overall operating efficiency.
It provides a hardware acceleration device with high concurrent hash addressing, including dynamic random access memory, host side, input/output interface, multi-layer perceptron unit and voxel processing unit. The access mode judgment unit selects a suitable access mode according to the data bandwidth and access requirements, and rearranges the data in the hash table storage unit according to the hash function mapping rules and the number of SRAM memory banks through the data rearrangement unit to reduce conflicts.
It optimizes memory access efficiency, improves system throughput, reduces hash conflicts, improves hash table search efficiency, reduces access delay, and solves the problems of frequent memory access and memory bank conflicts.
Smart Images

Figure CN120295938A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer hardware acceleration technology, and particularly to a hardware acceleration device for high-concurrency hash addressing. Background Art
[0002] With the continuous development of computer graphics technology, three-dimensional reconstruction technology has shown great application potential in fields such as virtual reality, augmented reality, and game development. Traditional three-dimensional reconstruction methods, such as rendering through the GPU, although able to synthesize realistic images, their complex calculation processes and high computational complexity limit their application in large-scale scenarios and real-time rendering. To overcome these limitations, the Instant-NGP method emerged. This method significantly improves the rendering speed and quality by using trainable multi-resolution hash encoding to replace traditional frequency encoding. However, during the hash addressing process of Instant-NGP, frequent memory access and reduced effective bandwidth have become the main bottlenecks restricting its performance improvement, especially the bank conflict problem caused by irregular memory access in the hash table, which has a significant impact on the overall operating efficiency.
[0003] Regarding the memory access conflict problem in hash addressing, in related technologies, the first solution can reduce single-point memory access conflicts by changing the data arrangement method and distributing different channels of the same feature vector of sampling points among different memory banks. The second solution can focus on the Nerf algorithm and optimize the storage strategy of three-dimensional scene features by storing adjacent features in different memories.
[0004] However, the first solution has limited improvement in the conflict problem of concurrent memory access to eight vertices. Especially when the feature channel size exceeds the number of memory banks, conflicts will still occur. Although the second solution alleviates the conflict to a certain extent, the read-write conflict problem in hash embedding still exists. Therefore, aiming at the frequent memory access and bank conflict problems in Instant-NGP hash addressing, a more effective data arrangement and memory access strategy needs to be proposed to improve memory access efficiency and bandwidth utilization, thereby enhancing the overall performance of Instant-NGP in edge hardware deployment. Summary of the Invention
[0005] This application provides a hardware acceleration device for high-concurrency hash addressing to solve the problems of frequent memory access and bank conflicts in high-concurrency hash addressing.
[0006] This application provides a hardware acceleration device for high-concurrency hash addressing, including: dynamic random access memory, host side, input / output interface, multi-layer perceptron unit, and voxel processing unit;
[0007] The host is respectively communicatively connected to the dynamic random access memory and the input / output interface, and the input / output interface is respectively communicatively connected to the multi-layer perceptron unit and the voxel processing unit;
[0008] The voxel processing unit includes a memory access mode judgment unit, a data rearrangement unit, and a hash table storage unit; the memory access mode judgment unit, the data rearrangement unit, and the hash table storage unit are communicatively connected in sequence;
[0009] The memory access mode judgment unit is used to select a memory access mode according to the data bandwidth and memory access requirements; the memory access modes include pipelined access combined layout and synchronous access combined layout;
[0010] The data rearrangement unit is used to rearrange the data in the hash table storage unit according to the mapping rule of the hash function, the data layout method, the number of memory banks of the SRAM, and the memory access mode to reduce conflicts; the data layout methods include sequential layout between memory banks, sequential layout within a memory bank, and interleaved storage layout.
[0011] The high-concurrency hash addressing hardware acceleration device dynamically selects a memory access mode through the memory access mode judgment unit according to the data bandwidth and memory access requirements, thereby optimizing the memory access efficiency and enhancing the system throughput. At the same time, the data rearrangement unit intelligently rearranges the data in the hash table storage unit according to the hash function mapping rule, the data layout method, the number of memory banks of the SRAM, and the memory access mode, effectively reducing hash conflicts, enhancing the hash table lookup efficiency, reducing access latency, and solving the problems of frequent memory access and memory bank conflicts in high-concurrency hash addressing.
[0012] Optionally, the calculation formula of the hash function is:
[0013] h i =H(x i , y i , z i )=(x i ×π1)⊕(y i ×π2)⊕(z i ×π2)modT;
[0014] where x i , y i , z i are input parameters; π1, π2, π3 are constants, and are respectively selected as 1, 2654435761, 805459861; ⊕ represents a bitwise exclusive OR operation; T is a prime number.
[0015] Optionally, the input parameter of the hash function is the coordinate of the sampling point obtained by radiometric sampling from 2D image sets from different perspectives; the output of the hash function is the index value of the 1D hash table.
[0016] The coordinates of the sampling points obtained by radiometric sampling from 2D image sets from different perspectives can quickly and accurately output the corresponding index values of the 1D hash table, improving the processing efficiency and accuracy.
[0017] Optionally, the data rearrangement unit is further configured to:
[0018] Obtain the coordinates of the vertices around the sampling points and provide the data storage method in the SRAM;
[0019] Group the coordinates of the surrounding vertices so that the y and z axis coordinates of each group are the same;
[0020] When the SRAM size is 8 banks, the data arrangement method selects the in-memory sequential arrangement;
[0021] When the SRAM size is 16 banks, the data arrangement method selects the interleaved storage arrangement;
[0022] When the SRAM size is 32 banks, the data arrangement method selects the inter-memory sequential arrangement.
[0023] The data rearrangement unit rearranges the data according to the SRAM size, so as to disperse the data access of the index address to different storage media at the physical level, thereby reducing access conflicts. When the SRAM size is 8 banks, the in-memory sequential arrangement is selected, which can reduce the potential conflicts of data access between groups; when the SRAM size is 16 banks, the interleaved storage arrangement is selected, which can reduce the access conflicts between adjacent features; when the SRAM size is 32 banks, the inter-memory sequential arrangement is selected, which can reduce the possible conflicts of data access within a group. Furthermore, the storage resources can be utilized more reasonably and the storage efficiency can be improved.
[0024] Optionally, the memory access mode judgment unit is further configured to:
[0025] If the data bandwidth is limited and the memory access demand exceeds the preset quota, the memory access mode selects the pipeline access combined arrangement;
[0026] If the data bandwidth is not limited and the memory access demand does not exceed the preset quota, the memory access mode selects the synchronous access combined arrangement;
[0027] Wherein, the memory access demand includes the amount of data accessed, the access frequency, and whether there are potential conflicts.
[0028] The memory access mode judgment unit flexibly selects an appropriate memory access mode according to the data bandwidth and memory access requirements, and the beneficial effects are remarkable. When the data bandwidth is limited, the pipelined access combined layout is selected, which can efficiently process the frequent access of a large amount of data, reduce access conflicts, and improve data throughput; when the data bandwidth is not limited, the synchronous access combined layout is selected to ensure data synchronization and accuracy.
[0029] Optionally, in the pipelined access combined layout, the addressing request is first searched in the first storage unit. If there is a possible conflict, the conflicting request is then transferred to the next storage unit for searching; the arrangement and combination methods of the storage units include the combination of sequential arrangement between memory banks and interleaved storage arrangement, the combination of sequential arrangement between memory banks and sequential arrangement within a memory bank, and the combination of sequential arrangement within a memory bank and interleaved storage arrangement.
[0030] In the synchronous access combined layout, the vertices are first divided into two parts, and the two storage units respectively access the two parts of vertices synchronously; the arrangement and combination methods of the storage units include the combination of sequential arrangement between memory banks and interleaved storage arrangement, the combination of sequential arrangement between memory banks and sequential arrangement within a memory bank, the combination of sequential arrangement within a memory bank and interleaved storage arrangement, and the combination of sequential arrangement within a memory bank and sequential arrangement within a memory bank.
[0031] The pipelined access combined layout searches the addressing request hierarchically. It first searches in the first storage unit and transfers to the next storage unit if there is a conflict, reducing access conflicts and improving the efficiency of data access. At the same time, the storage units adopt various arrangement and combination methods to adapt to the data access requirements in different scenarios, further optimizing the storage performance. The synchronous access combined layout divides the vertices into two parts, and the two storage units access them synchronously. Through various arrangement and combination methods of the storage units, efficient parallel processing can be achieved, improving the data processing speed and the overall system performance, making data access and processing more flexible and efficient.
[0032] Optionally, the host side includes: a central processing unit, a graphics processing unit, and a controller;
[0033] The controller is respectively communicatively connected with the central processing unit, the graphics processing unit, and the input / output interface; the central processing unit is used to execute instructions and detect and handle abnormal situations and interrupt requests; the controller is used to control the reading or writing of data from / to the dynamic random access memory; the graphics processing unit is used to process images according to the instructions.
[0034] The host side executes execution instructions through a central processing unit, and at the same time discovers and processes abnormal situations and interrupt requests, achieving efficient data processing and operations. The controller controls the reading or writing of data from the dynamic random access memory, ensuring the smoothness and stability of data transmission between components and improving the overall operation efficiency. The graphics processing unit processes images according to instructions, achieving high-quality image rendering and display effects. Each module works together, enabling the host side to exhibit powerful performance when processing complex tasks and enhancing the overall effectiveness of the system and the user experience.
[0035] Optionally, the multi-layer perceptron unit includes: a calculation module and an on-chip cache module;
[0036] The calculation module and the on-chip cache module are communicatively connected; the calculation module is used to calculate matrix multiplication operations in the neural network;
[0037] The on-chip cache module is used to store intermediate results during the operation of the multi-layer perceptron unit.
[0038] The multi-layer perceptron unit calculates matrix multiplication operations in the neural network through the calculation module, achieving efficient processing of data and being able to quickly calculate the forward and backward propagation processes of the MLP. The on-chip cache module is communicatively connected to the calculation module, stores calculation intermediate results during the operation, ensures fast data reading and writing, reduces data transmission latency, and improves the overall operation efficiency. The two work together, making the multi-layer perceptron unit more smooth and efficient when processing complex data, enhancing the speed and accuracy of data processing, and providing more reliable technical support for related applications.
[0039] Optionally, the voxel processing unit further includes a 3D coordinate cache unit, a hash function calculation unit, and an interpolation cache unit;
[0040] The 3D coordinate cache unit, the hash function calculation unit, and the interpolation cache unit are connected in sequence;
[0041] The 3D coordinate cache unit is used to temporarily store 3D coordinate data to be processed;
[0042] The hash function calculation unit is used to receive the 3D coordinate data and perform hash function calculation on it to generate a hash value;
[0043] The interpolation cache unit is used to store the addresses of the corresponding eight vertices in the embedded grid for the 3D coordinates to be accessed.
[0044] The voxel processing unit temporarily stores the 3D coordinate data to be processed through the 3D coordinate cache unit, which can effectively improve the fluency and stability of data processing. The hash function calculation unit receives the 3D coordinate data and performs hash function calculation to generate hash values, achieving fast data retrieval and matching, and improving the efficiency and accuracy of data processing. The interpolation cache unit stores the address of the interpolation data required for the 3D coordinates, facilitating quick call and reuse, and further enhancing the overall processing speed. The three are connected in sequence and work together to optimize the data processing flow, reduce the consumption of computing resources, and bring higher performance and accuracy to voxel processing.
[0045] Optionally, the voxel processing unit includes: an interpolation calculation unit and a gradient calculation unit;
[0046] The interpolation calculation unit and the gradient calculation unit are respectively connected to the hash table storage unit;
[0047] The interpolation calculation unit is used to perform interpolation calculation according to the data rearranged and stored in the hash table storage unit;
[0048] The gradient calculation unit is used to calculate gradient information according to the data rearranged and stored in the hash table storage unit.
[0049] Through the collaborative work of the interpolation calculation unit and the gradient calculation unit, the voxel processing unit realizes the efficient processing of the rearranged data. The interpolation calculation unit performs accurate interpolation calculation based on the data in the hash table storage unit, which can effectively improve the data integrity and accuracy; the gradient calculation unit calculates the gradient information according to the data also stored in the hash table storage unit, which helps to more comprehensively analyze the data change trend. The two are respectively connected to the hash table storage unit to ensure fast and accurate data transmission, and overall improve the efficiency and quality of voxel processing.
[0050] As can be seen from the above technical solutions, the present application provides a hardware acceleration device for high-concurrency hash addressing, including: a dynamic random access memory, a host, an input / output interface, a multi-layer perceptron unit, and a voxel processing unit; the host is respectively communicatively connected to the dynamic random access memory and the input / output interface, and the input / output interface is respectively communicatively connected to the multi-layer perceptron unit and the voxel processing unit; the voxel processing unit includes a memory access mode judgment unit, a data rearrangement unit, and a hash table storage unit; the memory access mode judgment unit, the data rearrangement unit, and the hash table storage unit are communicatively connected in sequence; the memory access mode judgment unit is used to select a memory access mode according to the data bandwidth and memory access requirements; the memory access mode includes a pipeline access combined layout and a synchronous access combined layout; the data rearrangement unit is used to rearrange the data in the hash table storage unit according to the mapping rule of the hash function, the data layout method, the number of memory banks of the SRAM, and the memory access mode to reduce conflicts; the data layout method includes sequential layout between memory banks, sequential layout within a memory bank, and interleaved storage layout to solve the problems of frequent memory access and memory bank conflicts in high-concurrency hash addressing. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 FIG. is a schematic structural connection diagram of the hardware acceleration device for high-concurrency hash addressing according to the embodiment of the present application;
[0053] Figure 2 FIG. is a schematic diagram of the hash function interpolation process in the hardware acceleration device for high-concurrency hash addressing according to the embodiment of the present application;
[0054] Figure 3 FIG. is a schematic diagram of the data layout method in the hardware acceleration device for high-concurrency hash addressing according to the embodiment of the present application;
[0055] Figure 4 FIG. is a schematic diagram of the memory access mode in the hardware acceleration device for high-concurrency hash addressing according to the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following will describe the embodiments in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application.
[0057] In the field of classical computer graphics, its core goal is to synthesize realistic and controllable images. This process usually relies on the graphics processing unit (GPU) for rendering. However, since the GPU mainly performs rendering tasks through complex geometric calculations, its computational complexity is relatively high. Traditional rendering algorithms simulate the interaction between light and objects by tracing the photon paths starting from light source objects and using the geometric and scattering distribution information of the objects, while inverse rendering algorithms follow the reverse of the rendering process. Neural rendering methods provide an alternative and efficient solution to traditional physical simulations.
[0058] In the relentless pursuit of improving 3D reconstruction quality, Instant-NGP uses neural network rendering methods to replace traditional rendering methods. To reconstruct a specific 3D scene, Instant-NGP starts from a set of multi-view 2D images. It performs radiative sampling on the rays passing through the pixel points, and then conducts trainable multi-resolution hash encoding on the sampled points. The purpose of the encoding is to map the neural network input to an encoding in a high-dimensional space, which is the key to extracting high-quality approximations from a compact model. Then, a customized deep neural network processes the features of the sampled points after hash encoding to achieve continuous volume encoding of density and backpropagation of loss gradients. That is, Instant-NGP learns the RGB color values of pixel points from the perspective and position information of different target scenes, and constructs an implicit model of the 3D scene to render images from a new perspective.
[0059] In Instant-NGP, for each sampled point along the pixel ray, the coordinates are mapped to trainable feature vectors at different resolution levels through a three-dimensional embedded hash grid, and the eight nearest vertices around this point in the 3D embedded grid at different resolution levels are found. The features of the sampled points are queried by storing a compact one-dimensional hash table. This way, the relatively costly multi-layer perceptron (MLP) inference operation is converted into a less costly embedding interpolation operation. Similar interpolation operations make parallel acceleration possible, but due to the large number of parameters in the hash table, the storage amount at different resolution levels exceeds 23MB, and the on-chip cache cannot store the entire hash table, resulting in a large number of hash lookups and frequent memory irregular accesses at the same time.
[0060] CUDA kernels are used for acceleration in Instant-NGP. In the CUDA architecture, to achieve high-throughput access, shared memory is divided into multiple independent storage regions. Consecutive words are assigned to consecutive memory banks. This is like arranging seats. A column of seats is equivalent to a memory bank. So, having multiple memory banks in each row is equivalent to having multiple seats. A 32-bit data can be placed in each seat, including combinations of multiple data smaller than 32 bits. Since multiple sampling points can share the same cube in the 3D hash grid, eigenvalue extraction may be required from the same or similar addresses in the same memory bank, which is likely to cause memory bank conflicts and result in a reduction in effective bandwidth.
[0061] To solve the problems of frequent memory access and memory bank conflicts in high-concurrency hash addressing, refer to Figure 1 , an embodiment of this application provides a hardware acceleration device for high-concurrency hash addressing, including: a dynamic random access memory, a host side, an input / output interface, a multi-layer perceptron unit, and a voxel processing unit; the host side is respectively communicatively connected to the dynamic random access memory and the input / output interface, and the input / output interface is respectively communicatively connected to the multi-layer perceptron unit and the voxel processing unit.
[0062] It should be understood that a dynamic random access memory (DRAM) is a semiconductor memory that uses the amount of charge stored in a capacitor to represent whether a binary bit (bit) is 1 or 0. A multi-layer perceptron unit (MLP) is a computational model inspired by the biological nervous system, consisting of an input layer, a hidden layer, and an output layer. It can automatically learn data features and perform classification or regression prediction by simulating neurons to perform weighted summation and non-linear transformation on input signals. An input / output interface (I / O interface) is a bridge connecting a computer system and external devices, responsible for enabling efficient data transmission between the internal processor and peripheral devices. It not only supports multiple communication protocols and standards but also ensures smooth interaction between devices with different speeds, solving the speed matching problem through technologies such as buffering and conversion.
[0063] The voxel processing unit includes a memory access mode judgment unit, a data rearrangement unit, and a hash table storage unit; the memory access mode judgment unit, the data rearrangement unit, and the hash table storage unit are communicatively connected in sequence;
[0064] The memory access mode judgment unit is used to select a memory access mode according to the data bandwidth and memory access requirements; the memory access modes include pipelined access combined layout and synchronous access combined layout;
[0065] The data rearrangement unit is used to rearrange the data in the hash table storage unit according to the mapping rule of the hash function, the data arrangement mode, the number of memory banks of the SRAM, and the memory access mode, so as to reduce conflicts; the data arrangement mode includes sequential arrangement between memory banks, sequential arrangement within a memory bank, and interleaved storage arrangement.
[0066] It should be understood that SRAM (Static Random Access Memory) is a semiconductor memory that stores data based on bistable flip-flops, and has the advantages of high speed, low power consumption, but high cost and low integration.
[0067] The high-concurrency hash addressing hardware acceleration device dynamically selects a memory access mode through the memory access mode judgment unit according to the data bandwidth and memory access requirements, thereby optimizing the memory access efficiency and improving the system throughput. At the same time, the data rearrangement unit intelligently rearranges the data in the hash table storage unit according to the hash function mapping rule, the data arrangement mode, the number of memory banks of the SRAM, and the memory access mode, effectively reducing hash conflicts, improving the hash table search efficiency, reducing access latency, and solving the problems of frequent memory access and memory bank conflicts in high-concurrency hash addressing.
[0068] In some embodiments, the calculation formula of the hash function is:
[0069] h i =H(x i , y i , z i )=(x i ×π1)⊕(y i ×π2)⊕(z i ×π3)modT;
[0070] Wherein, x i , y i , z i are input parameters; π1, π2, π3 are constants, and are respectively selected as 1, 2654435761, 805459861; ⊕ represents a bitwise exclusive OR operation; T is a prime number.
[0071] It should be understood that T is used for modulo operation to ensure that the hash value is within a specific range. The calculation formula of the hash function performs bitwise exclusive OR operations on the introduced constants π1, π2, π3 and the x, y, z coordinates of the vertex respectively. By analyzing the corresponding relationship between the values of π1, π2, π3 and the coordinates, the index values calculated by the hash function can be divided into two cases. When the coordinate difference between adjacent vertices is on the y-axis or z-axis, the difference will be amplified by π2 or π3; when the coordinate difference is on the x-axis, since π1 = 1 in the equation, the difference will not be amplified. See Figure 2, taking the vertices {111, 011, 001, 101, 000, 100, 010, 110} as an example, these 8 vertices can be divided into four groups. In each group, the y and z coordinates are the same, and only the x-axis coordinate is different. Therefore, according to the above rule, the hash index values within the same group should be close, while there are significant differences between the index values of different groups. During the process of hash function interpolation, by extracting and analyzing the real data of the eight vertices around the sampling point at different resolutions, four rules can be found:
[0072] (1) Approximately 90% of the intra-group distances are less than 5, among which approximately 70% are concentrated around 1, and the average distance of the index values is about 4. The main reason is that at the high-resolution level, the hash mapping is not a one-to-one mapping, and the index value is the result of taking the remainder of the hash value with respect to the size T of the hash table, which leads to a small number of cases with large intra-group distances, thus increasing the average value;
[0073] (2) Except that the average inter-group distance is 90 at the first-level resolution, the inter-group distances at other resolutions can be very large, with an average distance up to 91,000;
[0074] (3) Although the average distance of the intra-group index values is 4, the distances between the index values are all odd numbers;
[0075] (4) For groups with completely different y and z coordinates, the intra-group distances are the same. Through experiments, it is found that the data is almost axisymmetric, that is, the intra-group distances of the four groups are pairwise the same. For example, Figure 2 the intra-group distances of group 1 and group 3, and group 2 and group 4 are the same.
[0076] In some embodiments, the input parameter of the hash function is the coordinate of the sampling point obtained by radiometric sampling from 2D image sets from different perspectives; the output of the hash function is the index value of a 1D hash table.
[0077] The coordinate of the sampling point obtained by radiometric sampling from 2D image sets from different perspectives can quickly and accurately output the corresponding index value of the 1D hash table, improving the processing efficiency and accuracy.
[0078] In some embodiments, the data rearrangement unit is further configured to:
[0079] Obtain the coordinates of the vertices around the sampling point and provide the data storage method in SRAM;
[0080] Group the coordinates of the surrounding vertices so that the y and z axis coordinates of each group are the same;
[0081] When the SRAM size is 8 banks, the data arrangement method selects sequential arrangement within the memory bank; when the SRAM size is 16 banks, the data arrangement method selects interleaved storage arrangement; when the SRAM size is 32 banks, the data arrangement method selects sequential arrangement between memory banks.
[0082] It should be understood that when performing parallel queries on scene features stored on the same memory bank, an inappropriate storage strategy may cause memory bank conflicts, thereby leading to a reduction in the effective data bandwidth. In view of this, three types are divided based on the sequential arrangement within the memory bank (bank). See Figure 3 , first, Figure 3 (a) in shows the strategy of sequential arrangement between memory banks. This strategy, based on the law of hash addressing, ensures that in the intra-group access of eight points, related addresses are not in the same memory bank, thus improving the efficiency of intra-group search. Second, Figure 3 (b) in describes the strategy of sequential arrangement within the memory bank. This strategy fills the embedded values corresponding to the hash index into different memory banks in sequence, ensuring that in the inter-group access of eight vertices, related addresses are not in the same memory bank, thus optimizing the performance of inter-group search. Finally, Figure 3 (c) in shows the strategy of interleaved storage arrangement. This strategy, based on the law that the distance between two points within a group is odd, performs partial transposition on the data arranged sequentially between memory banks to ensure that adjacent features are stored in different memory banks.
[0083] The data rearrangement unit rearranges the data according to the SRAM size to disperse the data access of the index address to different storage media at the physical level, thereby reducing access conflicts. When the SRAM size is 8 banks, selecting sequential arrangement within the memory bank can reduce potential conflicts in inter-group data access; when the SRAM size is 16 banks, selecting interleaved storage arrangement can reduce access conflicts between adjacent features; when the SRAM size is 32 banks, selecting sequential arrangement between memory banks can reduce possible conflicts in intra-group data access. Furthermore, it can make more reasonable use of storage resources and improve storage efficiency.
[0084] In some embodiments, the memory access mode judgment unit is further configured to:
[0085] If the data bandwidth is limited and the memory access demand exceeds a preset limit, the memory access mode selects a pipelined access combination arrangement; if the data bandwidth is not limited and the memory access demand does not exceed the preset limit, the memory access mode selects a synchronous access combination arrangement;
[0086] wherein, the memory access demand includes the amount of data accessed, the access frequency, and whether there are potential conflicts.
[0087] It should be understood that in the memory access mode determination unit, to achieve the balance between the number of memory bank conflicts and the average access cycle, static random access memories (SRAMs) with different capacities and different access modes are considered respectively. Specifically, for the 8-bank configuration, the internal memory structure of the SRAM directly references the aforementioned three data arrangement methods. By capturing the coordinate index value data from the source code, the parameter scheme is implemented using the Python language, and the number of conflicts and the number of running cycles are counted. At the same time, comparisons are also made through SRAMs with different numbers of banks:
[0088] The number of conflicts and the number of cycles for the case where the SRAM has 8 banks are shown in the following table:
[0089] Sequential arrangement between banks Sequential arrangement within a bank Interleaved storage arrangement Number of conflicts 0.999488 0.622080 1.093404 Total number of cycles 262063 212598 274372
[0090] For the 16-bank configuration, two cases need to be considered: one is that the SRAM internally contains 16 banks, which is equivalent to expanding the 8-bank SRAM into a wider and shallower structure; the other case is that the SRAM internally contains two memory units A and B, each with 8 banks, where memory units A and B store the same data, but there may be differences in the access mode and layout. Specifically, for the access mode of the input eight vertex data, it can be divided into two processing strategies, named SRAM Logic 1 and SRAM Logic 2 respectively.
[0091] Figure 4 Figure (d) shows the first processing strategy, which sequentially accesses the two memory units through a pipeline method, that is, the addressing request is first searched in the SRAM A unit. If there is a potential conflict, the conflict request is transferred to the SRAM B unit for continued search. There are three layout combinations of the memory units A and B (SRAM Cell A and SRAM Cell B): (ab) the combination of inter-bank sequential layout and intra-bank sequential layout, (ac) the combination of inter-bank sequential layout and interleaved storage layout, and (bc) the combination of intra-bank sequential layout and interleaved storage layout. Figure 4Figure (e) shows the synchronous access combined layout. In this case, the eight vertex data are divided into two parts, and two storage units synchronously access four vertex data. Similarly, there are three different layout combination ways for storage units A and B, and a new combination way (bb) is added, that is, both storage units A and B adopt the in - memory sequential layout. The advantage of this combination way is that in the in - memory sequential layout storage method, memory bank conflicts may only be caused by the two vertex index values within a group being too close. The coordinate characteristics of the two vertices within a group are as follows: there are differences in the x - axis coordinates, while the y - axis and z - axis coordinates are the same. Therefore, vertices with different x - axis coordinate values are assigned to different storage units, that is, the eight vertices are recombined according to whether the x values are the same. For example, the four points {111, 101, 100, 110} will be searched in memory bank A, while {011, 001, 000, 010} will be searched in memory bank B. In this way, the index value distance of the four vertices within a group increases, and the two vertices that would originally conflict are assigned to different storage units for searching, thus reducing the possibility of conflicts.
[0092] For the case of 16 - bank SRAM, the number of conflicts and the number of cycles are shown in the following table:
[0093] Sequential arrangement between banks Sequential arrangement within a bank Interleaved storage arrangement Number of conflicts 0.995132 0.403387 0.339800 Total number of cycles 261492 183935 175601
[0094] Pipeline with 16 banks (ac) Pipeline with 16 banks (ab) Pipeline with 16 banks (bc) Number of conflicts 0.181200 0.559653 0.750253 Total number of cycles 154814 204416 229397
[0095]
[0096] For the 32 - memory - bank configuration, similar to the 16 - memory - bank configuration, the main difference is that the original identical storage units A and B are expanded from 8 memory banks to 16. At the same time, a new layout way (4a) is introduced, which divides the SRAM into four storage units each containing 8 memory banks and performs synchronous access, and they all adopt the inter - memory - bank sequential layout. In this configuration, each memory - bank block only needs to search the index values of two coordinates among the eight vertices. Since the inter - memory - bank sequential layout is beneficial for intra - group searching, in order to minimize conflicts, each memory - bank block searches for two adjacent values within the group, that is, only searches the coordinate index values with different x - coordinates.
[0097] For the case of 32 - bank SRAM, the number of conflicts and the number of cycles are shown in the following table:
[0098] Sequential arrangement between banks Sequential arrangement within a bank Interleaved storage arrangement Number of conflicts 0.993033 0.037126 0.165749 Total number of cycles 261217 135931 152789
[0099] Pipeline with 32 banks (ac) Pipeline with 32 banks (ab) Pipeline with 32 banks (bc) Number of conflicts 0.038629 0.395467 0.184549 Total number of cycles 154653 182897 155253
[0100]
[0101] From the above table, it can be clearly seen that by adopting the method of this solution, the number of conflicts can be significantly reduced.
[0102] From the perspective of the same bank size, comparing different data arrangement methods and memory access patterns, for 8 banks, the arrangement method with the smallest average number of conflicts is sequential arrangement in the memory; for 16 banks, first comprehensively compare, with the minimum conflict as the top priority. Among the synchronous access combined arrangements, the number of conflicts in the sequential arrangement in the memory is much lower than the other cases. Then compare the memory access methods of the pipeline. Among the three data arrangement methods, the ac combination has a relatively low average number of conflicts, that is, the A unit is arranged sequentially in the memory, and the B unit is arranged in an interleaved manner. For the synchronous access mode, except for bb, the average number of conflicts of the other three data arrangement combinations is not much different. For 32 banks, by adopting the synchronous 4a combination method, the number of conflicts can be almost eliminated. Based on extensive measurements, this method significantly reduces the number of conflicts and the total number of cycles by reasonably allocating the index values of eight sampling points.
[0103] Then make a horizontal comparison. From the perspective of the same memory access method, comparing different bank sizes, it can be found that the larger the size of the SRAM, the smaller the number of conflicts and the total number of cycles. But at the same time, it brings a greater area overhead. Therefore, this solution provides different optional methods for different requirements.
[0104] The memory access mode judgment unit flexibly selects an appropriate memory access mode according to the data bandwidth and memory access requirements, and the beneficial effects are significant. When the data bandwidth is limited, the pipeline access combined arrangement is selected, which can efficiently process the frequent access of a large amount of data, reduce access conflicts, and improve data throughput; when the data bandwidth is not limited, the synchronous access combined arrangement is selected, which can perform parallel access to data and reduce the number of conflicts.
[0105] In some embodiments, in the pipeline access combined arrangement, first search for the addressing request in the first storage unit. If there is a possible conflict, then transfer the conflicting request to the next storage unit for searching; the arrangement combination mode of the storage units includes the combination of sequential arrangement between memory banks and interleaved storage arrangement, the combination of sequential arrangement between memory banks and sequential arrangement in the memory, and the combination of sequential arrangement in the memory and interleaved storage arrangement.
[0106] In the synchronous access combined arrangement, first divide the vertices into two parts, and the two storage units respectively synchronously access the two parts of vertices; the arrangement combination mode of the storage units includes the combination of sequential arrangement between memory banks and interleaved storage arrangement, the combination of sequential arrangement between memory banks and sequential arrangement in the memory, the combination of sequential arrangement in the memory and interleaved storage arrangement, and the combination of sequential arrangement in the memory and sequential arrangement in the memory.
[0107] The pipeline access combined layout reduces access conflicts and improves the efficiency of data access by hierarchically searching for addressing requests, first searching in the first storage unit and transferring to the next storage unit if there is a conflict. At the same time, the storage units adopt various permutation and combination methods to adapt to the data access requirements in different scenarios, further optimizing the storage performance. The synchronous access combined layout divides the vertices into two parts, and two storage units access synchronously respectively. Through various permutation and combination methods of storage units, efficient parallel processing can be achieved, improving the data processing speed and the overall system performance, making data access and processing more flexible and efficient.
[0108] In some embodiments, the host side includes: a central processing unit, a graphics processing unit, and a controller;
[0109] The controller is communicatively connected to the central processing unit, the graphics processing unit, and the input / output interface respectively; the central processing unit is configured to execute instructions and detect and handle abnormal situations and interrupt requests; the controller is configured to control reading or writing data from / to the dynamic random access memory; the graphics processing unit is configured to process images according to the instructions.
[0110] It should be understood that the central processing unit (CPU) is the core component of a computer system, responsible for executing instructions, processing data, and controlling operations. It consists of an arithmetic logic unit, a control unit, registers, etc., and is connected to other components through a bus. The CPU executes instructions, obtains data from memory or the I / O interface in sequence, analyzes and operates on it, and then sends it back, ensuring the efficient and stable operation of the computer. At the same time, it detects and handles abnormal situations and interrupt requests to achieve efficient data processing and operations. The graphics processing unit (GPU) is a processor designed specifically for image and graphics operation tasks. The GPU processes image information according to instructions and achieves high-quality rendering tasks and display effects. The controller controls reading or writing data from / to the dynamic random access memory to ensure the accurate transmission and interaction of data between components, improving the overall operation efficiency.
[0111] The host side executes instructions through the central processing unit, and at the same time detects and handles abnormal situations and interrupt requests, achieving efficient data processing and operations. The controller controls reading or writing data from / to the dynamic random access memory to ensure the smoothness and stability of data transmission between components, improving the overall operation efficiency. The graphics processing unit processes images according to instructions, achieving high-quality image rendering and display effects. Each module works together, enabling the host side to exhibit powerful performance when processing complex tasks, enhancing the overall effectiveness of the system and the user experience.
[0112] In some embodiments, the multi-layer perceptron unit includes: a computing module and an on-chip cache module; the computing module and the on-chip cache module are communicatively connected; the computing module is configured to compute matrix multiplication operations in a neural network; the on-chip cache module is configured to store intermediate results during the operation of the multi-layer perceptron unit.
[0113] The multi-layer perceptron unit computes matrix multiplication operations in a neural network through the computing module, achieving efficient data processing and enabling fast computation of the forward and backward propagation processes of the MLP. The on-chip cache module is communicatively connected to the computing module and stores intermediate computation results during the operation, ensuring fast data reading and writing, reducing data transmission latency, and improving the overall computation efficiency. The two work together, making the multi-layer perceptron unit more fluent and efficient in processing complex data, enhancing the speed and accuracy of data processing, and providing more reliable technical support for related applications.
[0114] In some embodiments, the voxel processing unit further includes a 3D coordinate cache unit, a hash function computing unit, and an interpolation cache unit; the 3D coordinate cache unit, the hash function computing unit, and the interpolation cache unit are connected in sequence; the 3D coordinate cache unit is configured to temporarily store 3D coordinate data to be processed; the hash function computing unit is configured to receive the 3D coordinate data and perform a hash function computation thereon to generate a hash value; the interpolation cache unit is configured to store the addresses of the corresponding eight vertices in the embedded grid for the 3D coordinates to be accessed.
[0115] The voxel processing unit temporarily stores 3D coordinate data to be processed through the 3D coordinate cache unit, effectively improving the fluency and stability of data processing. The hash function computing unit receives the 3D coordinate data and performs a hash function computation to generate a hash value, achieving fast data retrieval and matching and improving the efficiency and accuracy of data processing. The interpolation cache unit stores the interpolation data addresses required for the 3D coordinates, facilitating fast call and reuse and further enhancing the overall processing speed. The three are connected in sequence and work together to optimize the data processing flow, reduce the consumption of computing resources, and bring higher performance and accuracy to voxel processing.
[0116] In some embodiments, the voxel processing unit includes: an interpolation computing unit and a gradient computing unit; the interpolation computing unit and the gradient computing unit are respectively connected to the hash table storage unit; the interpolation computing unit is configured to perform interpolation computation according to the data rearranged and stored in the hash table storage unit; the gradient computing unit is configured to compute gradient information according to the data rearranged and stored in the hash table storage unit.
[0117] Through the collaborative work of the interpolation calculation unit and the gradient calculation unit, the voxel processing unit achieves efficient processing of the rearranged data. The interpolation calculation unit performs precise interpolation calculations based on the data in the hash table storage unit, effectively improving data integrity and accuracy. The gradient calculation unit calculates gradient information based on the data also stored in the hash table storage unit, helping to more comprehensively analyze the data change trend. The two are respectively connected to the hash table storage unit to ensure fast and accurate data transmission, and overall improve the efficiency and quality of voxel processing.
[0118] As can be seen from the above technical solutions, the embodiment of the present application provides a hardware acceleration device for high-concurrency hash addressing, including: a dynamic random access memory, a host, an input / output interface, a multi-layer perceptron unit, and a voxel processing unit; the host is respectively communicatively connected to the dynamic random access memory and the input / output interface, and the input / output interface is respectively communicatively connected to the multi-layer perceptron unit and the voxel processing unit; the voxel processing unit includes a memory access mode judgment unit, a data rearrangement unit, and a hash table storage unit; the memory access mode judgment unit, the data rearrangement unit, and the hash table storage unit are communicatively connected in sequence; the memory access mode judgment unit is used to select a memory access mode according to the data bandwidth and memory access requirements; the memory access mode includes pipelined access combined layout and synchronous access combined layout; the data rearrangement unit is used to rearrange the data in the hash table storage unit according to the mapping rule of the hash function, the data layout method, the number of memory banks of the SRAM, and the memory access mode to reduce conflicts; the data layout method includes sequential layout between memory banks, sequential layout within a memory bank, and interleaved storage layout to solve the problems of frequent memory access and memory bank conflicts in high-concurrency hash addressing.
[0119] For the similarity parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation to the protection scope of the present application. For those skilled in the art, any other implementation manners extended based on the solution of the present application without creative efforts belong to the protection scope of the present application.
Claims
1. A hardware acceleration device for high-concurrency hash addressing, characterized in that Including: Dynamic random access memory, host side, input / output interface, multi-layer perceptron unit, and voxel processing unit; The host side is respectively communicatively connected to the dynamic random access memory and the input / output interface, and the input / output interface is respectively communicatively connected to the multi-layer perceptron unit and the voxel processing unit; The voxel processing unit includes an access mode judgment unit, a data rearrangement unit, and a hash table storage unit; the access mode judgment unit, the data rearrangement unit, and the hash table storage unit are communicatively connected in sequence; The access mode judgment unit is used to select an access mode according to the data bandwidth and the access requirement; the access mode includes pipeline access combined arrangement and synchronous access combined arrangement; The data rearrangement unit is used to rearrange the data in the hash table storage unit according to the mapping rule of the hash function, the data arrangement method, the number of memory banks of the SRAM, and the access mode to reduce conflicts; the data arrangement method includes sequential arrangement between memory banks, sequential arrangement within a memory bank, and interleaved storage arrangement.
2. The hardware acceleration device for high-concurrency hash addressing according to claim 1, characterized in that The calculation formula of the hash function is: h i = H(x i , y i , z i ) = (x i × π1) ⊕ (y i × π2) ⊕ (z i × π2) mod T; where x i , y i , z i are input parameters; π1, π2, π3 are constants, and are selected as 1, 2654435761, 805459861 respectively; ⊕ represents bitwise exclusive OR operation; T is a prime number.
3. The hardware acceleration device for high-concurrency hash addressing according to claim 2, wherein The input parameter of the hash function is the sampling point coordinates obtained by radiometric sampling from 2D image sets from different perspectives; the output of the hash function is the index value of the 1D hash table.
4. The hardware acceleration device for high-concurrency hash addressing according to claim 3, characterized in that, The data rearrangement unit is further used for: Obtaining the coordinates of the vertices around the sampling point and providing the data storage method in the SRAM; Grouping the coordinates of the surrounding vertices so that the y and z axis coordinates of each group are the same; When the SRAM size is 8 banks, the data arrangement method selects sequential arrangement within a memory bank; When the SRAM size is 16 banks, the data arrangement method selects interleaved storage arrangement; When the SRAM size is 32 banks, the data arrangement method selects sequential arrangement between memory banks.
5. The hardware acceleration device for high-concurrency hash addressing according to claim 4, characterized in that, The access mode judgment unit is further used for: If the data bandwidth is limited and the access requirement exceeds a preset amount, the access mode selects pipeline access combined arrangement; If the data bandwidth is not limited and the access requirement does not exceed a preset amount, the access mode selects synchronous access combined arrangement; Wherein, the access requirement includes the amount of data accessed, the access frequency, and whether there are potential conflicts.
6. The hardware acceleration device for high-concurrency hash addressing according to claim 5, characterized in that, In the pipeline access combined arrangement, first search for the addressing request in the first storage unit. If there may be a conflict, then transfer the conflicting request to the next storage unit for searching; The arrangement combination mode of the storage units includes the combination of sequential arrangement between memory banks and interleaved storage arrangement, the combination of sequential arrangement between memory banks and sequential arrangement within a memory bank, and the combination of sequential arrangement within a memory bank and interleaved storage arrangement; In the synchronous access combined arrangement, first divide the vertices into two parts, and the two storage units respectively synchronously access the two parts of vertices; the arrangement combination mode of the storage units includes the combination of sequential arrangement between memory banks and interleaved storage arrangement, the combination of sequential arrangement between memory banks and sequential arrangement within a memory bank, the combination of sequential arrangement within a memory bank and interleaved storage arrangement, and the combination of sequential arrangement within a memory bank and sequential arrangement within a memory bank.
7. The hardware acceleration device for high-concurrency hash addressing according to claim 1, characterized in that, The host side includes: a central processing unit, a graphics processing unit, and a controller; The controller is communicatively connected to the central processing unit, the graphics processing unit, and the input / output interface respectively; the central processing unit is configured to execute instructions and detect and handle exception situations and interrupt requests; the controller is configured to control the reading or writing of data from / to the dynamic random access memory; the graphics processing unit is configured to process images according to the instructions.
8. The hardware acceleration device for high-concurrency hash addressing according to claim 1, wherein The multi-layer perceptron unit includes: a computing module and an on-chip cache module; The computing module and the on-chip cache module are communicatively connected; the computing module is configured to calculate matrix multiplication operations in the neural network; The on-chip cache module is configured to store intermediate results during the operation of the multi-layer perceptron unit.
9. The hardware acceleration device for high-concurrency hash addressing according to claim 1, wherein The voxel processing unit further includes a 3D coordinate cache unit, a hash function calculation unit, and an interpolation cache unit; The 3D coordinate cache unit, the hash function calculation unit, and the interpolation cache unit are connected in sequence; The 3D coordinate cache unit is configured to temporarily store 3D coordinate data to be processed; The hash function calculation unit is configured to receive the 3D coordinate data and perform hash function calculation on it to generate a hash value; The interpolation cache unit is configured to store the addresses of the corresponding eight vertices in the embedded grid for the 3D coordinates to be accessed.
10. The hardware acceleration device for high-concurrency hash addressing according to claim 1, characterized in that, The voxel processing unit includes: an interpolation calculation unit and a gradient calculation unit; The interpolation calculation unit and the gradient calculation unit are respectively connected to the hash table storage unit; The interpolation calculation unit is configured to perform interpolation calculation according to the data rearranged and stored in the hash table storage unit; The gradient calculation unit is configured to calculate gradient information according to the data rearranged and stored in the hash table storage unit.