Hardware filter, graph neural network accelerator and off-chip memory access screening method thereof
By introducing hardware filters into the graph neural network accelerator, grouping and filtering read and memory access requests in the aggregation stage, the problem of low memory access efficiency in existing accelerators is solved, and more efficient memory access operations and performance improvements are achieved.
Patent Information
- Application Number
- CN202510069404.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
When processing large-scale graph data, existing graph neural network accelerators face the problems of large-scale memory access and low memory access efficiency, and fail to effectively utilize the robustness of graph neural networks to improve performance.
A hardware filter is designed to group and filter read and memory access requests in the aggregation stage, and schedule according to the minimum access unit burst and row sizes of DRAM, reducing the actual total amount of memory access and improving memory access efficiency.
Through grouping and filtering of hardware filters, the access stock is reduced, localization is improved, and acceleration effect is improved, while not affecting the accuracy of the model.
Smart Images

Figure CN119988246A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graph neural network accelerators, and specifically relates to a hardware filter, a graph neural network accelerator, and an off-chip memory access screening method thereof. Background Art
[0002] As graph neural networks (GNNs) are increasingly used in various fields, their memory access efficiency has become one of the key factors restricting performance improvement. Existing graph neural network accelerators face the problems of large total memory access and low memory access efficiency when processing large-scale graph data.
[0003] Graph neural networks have the interweaving behavior of coarse-grained random memory access in the aggregation phase and high-density regular computing in the combination phase. Due to the extreme sparsity and irregularity of the graph structure and the coarse-grained nature of the graph data, coarse-grained random memory access in the aggregation phase becomes the main execution bottleneck of graph neural networks. However, the robustness of graph neural networks allows a certain degree of data loss, so from the practical perspective of model accuracy, not all memory access behaviors are necessary. In summary, existing graph neural network accelerators have failed to seize the opportunity brought by the robustness of graph neural networks themselves to filter the memory access in the aggregation phase to obtain performance improvements. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention improves the existing graph neural network accelerator, adds a hardware filter, and provides a hardware filter, a graph neural network accelerator and an off-chip memory access screening method thereof, which aims to group the read memory access requests in the aggregation stage according to the two granularities of the minimum access unit burst and row of DRAM for screening, thereby reducing the actual total amount of memory access and improving the memory access efficiency without affecting the accuracy of the model, thereby achieving an acceleration effect.
[0005] In order to achieve the above object, the present invention provides a hardware filter, including:
[0006] An input interface for receiving sparse memory access requests from the graph neural network accelerator and grouping the sparse memory access requests into several burst requests according to the minimum unit burst of DRAM access;
[0007] at least one filter, configured to perform at least one round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in each round;
[0008] The output interface is used to output the burst requests to be retained in the final round and the burst requests to be filtered out in all rounds.
[0009] In an embodiment of the present invention, the at least one filter is further used to: in the current round of screening, screen the burst requests to be retained obtained in the previous round.
[0010] In one embodiment of the present invention, the hardware filter further comprises:
[0011] The first storage module is used to store the burst requests to be retained obtained in each round of screening in the form of a grouping table.
[0012] In one embodiment of the present invention, the burst requests to be retained obtained from each round of screening are stored in the form of a grouping table, including:
[0013] Indexing the rows of the grouping table by the unique identifier of the DRAM row;
[0014] In combination with the organization form and address mapping mode of DRAM, the rows of the burst requests to be retained obtained in each round of screening in the DRAM are obtained, and a row identifier corresponding to each burst request is generated;
[0015] If the row identifier is found in the grouping table, the burst request corresponding to the row identifier is stored in the grouping table in a row indexed by the row identifier;
[0016] If the row identifier is not found in the grouping table, an idle row is searched in the grouping table, the idle row index is set as the row identifier, and the burst request corresponding to the row identifier is stored in the grouping table in the row indexed as the row identifier.
[0017] In one embodiment of the present invention, the first storage module uses a content addressable memory combined with a FIFO queue for storage, wherein the content addressable memory is used to index a row identifier, and the FIFO queue stores the burst request of the row.
[0018] In one embodiment of the present invention, the at least one filter comprises:
[0019] A first filter, used for performing a first round of screening on the input burst requests, and identifying the burst requests to be retained and the burst requests to be screened out in the first round;
[0020] The second filter, during the current round of screening, is used to select at least part of the rows of the burst requests from the grouping table corresponding to the previous round of the first storage module to perform screening. During the screening process, retention is determined based on the granularity of the entire row, and the burst requests to be retained and the burst requests to be screened out in the current round are identified.
[0021] In one embodiment of the present invention, the second filter uses a total order of row sizes to perform filtering;
[0022] Calculate the balance of the second filter, where the value of the balance is the number of all burst requests filtered by the second filter*(1-the screening ratio of the second filter);
[0023] If the balance is greater than or equal to zero, find the largest row in the grouping table, and identify all burst requests stored in the row as the burst requests to be retained in the current round;
[0024] If the balance is less than zero, the smallest row in the grouping table is found, and all burst requests stored in the row are identified as the burst requests to be screened out in the current round.
[0025] In one embodiment of the present invention, the hardware filter further comprises:
[0026] A trigger is connected to the first storage module and the second filter, and is used to receive the external state change information generated by the grouping table, determine whether to trigger an output action, and when the trigger triggers the output action, select the burst request of at least part of the rows in the grouping table to perform filtering.
[0027] In one embodiment of the present invention, the external state change information includes at least one of the number of valid burst requests accommodated in the grouping table, the number of valid rows, the size of the rows inserted in the current round of screening, the number of burst requests inserted within a given time or the number of rows involved, the current timestamp, and the timestamp of the previous round of output actions, or a combination of more.
[0028] In one embodiment of the present invention, the hardware filter further comprises:
[0029] The second storage module is used to store the burst requests to be screened out obtained in each round of screening.
[0030] Another aspect of the present invention further provides a graph neural network accelerator, comprising:
[0031] A graph neural network acceleration module to generate sparse memory access requests;
[0032] A hardware filter that contains at least:
[0033] An input interface, connected to the graph neural network acceleration module, for receiving the sparse memory access request from the graph neural network accelerator module, and grouping the sparse memory access request into a plurality of burst requests according to the minimum unit burst of DRAM access;
[0034] at least one filter, configured to perform at least one round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in each round;
[0035] An output interface, used to output the burst requests to be retained in the final round and the burst requests to be screened out in all rounds;
[0036] A memory controller connected to the output interface, configured to receive the burst request to be retained in the final round and return a correct memory access result;
[0037] a false zero value generating module, connected to the output interface, for receiving the burst requests to be screened out of all rounds to generate a false zero value result;
[0038] A hybrid module, wherein the input end of the hybrid module is connected to the memory controller and the false zero value generation module, and the output end is connected to the graph neural network acceleration module, so as to obtain the correct memory access result and the false zero value result to generate a sparse memory access result which is fed back to the graph neural network acceleration module.
[0039] In another aspect, the present invention further provides an off-chip memory access screening method for applying a graph neural network accelerator, the method comprising:
[0040] Receive sparse memory access requests from the graph neural network accelerator, and group the sparse memory access requests into several burst requests according to the minimum unit burst of DRAM access;
[0041] Perform at least one round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in each round;
[0042] The memory controller receives the burst request to be retained in the final round and returns a correct memory access result;
[0043] Receiving the burst requests to be screened out from all rounds to generate a false zero value result;
[0044] The correct memory access result and the false zero value result are obtained to generate a sparse memory access result and feed it back to the graph neural network accelerator.
[0045] In one embodiment of the present invention, when performing at least one round of screening on the input burst requests, in the current round of screening, the burst requests to be retained obtained in the previous round are screened.
[0046] In one embodiment of the present invention, the method further comprises:
[0047] The burst requests to be retained obtained in each round of screening are stored in the form of a grouping table, including:
[0048] Indexing the rows of the grouping table by the unique identifier of the DRAM row;
[0049] In combination with the organization form and address mapping mode of DRAM, the rows of the burst requests to be retained obtained in each round of screening in the DRAM are obtained, and a row identifier corresponding to each burst request is generated;
[0050] If the row identifier is found in the grouping table, the burst request corresponding to the row identifier is stored in the grouping table in a row indexed by the row identifier;
[0051] If the row identifier is not found in the grouping table, an idle row is searched in the grouping table, the idle row index is set as the row identifier, and the burst request corresponding to the row identifier is stored in the grouping table in the row indexed as the row identifier.
[0052] In one embodiment of the present invention, when performing at least one round of screening on the input burst requests, at least part of the rows of the burst requests are selected from the grouping table corresponding to the previous round to perform screening. During the screening process, retention is determined based on the granularity of the entire row, and the burst requests to be retained and the burst requests to be screened out in the current round are identified.
[0053] In one embodiment of the present invention, during the screening process, the retention or removal is determined according to the granularity of the entire row, and the burst request to be retained and the burst request to be screened out in the current round are identified, including:
[0054] The filter is filtered by the total order of row size;
[0055] Calculate the balance of the filter, the value of which is the number of all burst requests filtered by the filter*(1-filter elimination ratio);
[0056] If the balance is greater than or equal to zero, find the largest row in the grouping table, and identify all burst requests stored in the row as the burst requests to be retained in the current round;
[0057] If the balance is less than zero, the smallest row in the grouping table is found, and all burst requests stored in the row are identified as the burst requests to be screened out in the current round.
[0058] It can be seen from the above scheme that the advantages of the present invention are:
[0059] Compared with existing graph neural network accelerators, the present invention utilizes the robustness of the graph neural network itself. Without affecting the accuracy of the model, it takes DRAM as the center and groups the read memory access requests in the aggregation stage according to the two granularities of DRAM's minimum access unit burst and row, and performs screening and scheduling for execution. This naturally ensures the effectiveness of DRAM access, eliminates inefficient memory access, reduces the amount of memory access, and improves locality, thereby improving the acceleration effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A schematic structural diagram of a hardware filter provided by an embodiment of the present invention is shown;
[0061] Figure 2 A schematic structural diagram of a hardware filter provided by another embodiment of the present invention is shown;
[0062] Figure 3 A schematic structural diagram of a graph neural network accelerator provided by an embodiment of the present invention is shown;
[0063] Figure 4 A schematic flow chart of an off-chip memory access screening method provided in accordance with an embodiment of the present invention is shown.
[0064] Wherein, the accompanying drawings are marked as follows:
[0065] 100: Graph neural network accelerator;
[0066] 10: Hardware filter;
[0067] 11: Input interface;
[0068] 12: filter;
[0069] 121: first filter;
[0070] 122_1...122_n: second filter;
[0071] 13: output interface;
[0072] 14: a first storage module;
[0073] 141: Grouping table;
[0074] 15: second storage module;
[0075] 16: Trigger;
[0076] 20: Graph neural network acceleration module;
[0077] 30: memory controller;
[0078] 40: False zero value generation module;
[0079] 50: Hybrid module. DETAILED DESCRIPTION
[0080] In order to make the above features and effects of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings.
[0081] The present invention designs a dedicated hardware filter for memory access requests of reading DRAM in the aggregation stage of a graph neural network accelerator. The hardware filter performs configurable grouping, screening, and scheduling to replace dropout at the pure algorithm level, thereby improving the locality of memory access and enhancing the performance of the system without affecting the accuracy of the model.
[0082] See also Figure 1 , Figure 1 FIG. 4 shows an overall schematic structural diagram of a hardware filter provided by an embodiment of the present invention. Figure 2 A specific schematic structural diagram of a hardware filter provided by an embodiment of the present invention is shown.
[0083] refer to Figure 1 As shown in, the present invention improves the existing graph neural network accelerator and adds a hardware filter 10, which specifically includes: an input interface 11, at least one filter 12, and an output interface 13, wherein: the input interface 11 is used to receive sparse memory access requests from the inside of the graph neural network accelerator, and group the sparse memory access requests into several burst requests according to the minimum unit burst of DRAM access. In one embodiment, the sparse memory access request can be a read memory access request or a write memory access request. At least one filter 12 is used to perform at least one round of screening on the input burst requests to identify the burst requests to be retained and the burst requests to be screened out in each round. The output interface 13 is used to output the burst requests to be retained in the final round and the burst requests to be screened out in all rounds.
[0084] In one embodiment, when the at least one filter performs each screening on the input burst requests, in the current round of screening, the burst requests to be retained obtained in the previous round are further screened, thereby further identifying the burst requests to be retained in the current round and the burst requests to be screened out from the burst requests to be retained obtained in the previous round.
[0085] In addition, in one embodiment, reference Figure 2As shown in , the hardware filter 10 further includes: a first storage module 14, which is used to store the burst requests to be retained obtained in each round of screening in the form of a group table 141. Specifically, in one embodiment, the standard for classification and storage of the group table is generally the same DRAM row, but it can be changed to other units that adapt to memory access delay or energy consumption as needed. The rows of the group table are indexed by the unique identifier of the DRAM row. In the storage process, in combination with the organizational form and address mapping method of DRAM, the rows of the burst requests to be retained obtained in each round of screening in the DRAM are obtained, and the row identifier corresponding to each burst request is generated. At this time, according to the row identifier corresponding to each burst request, each burst request is stored. If the row identifier is queried in the group table, the burst request corresponding to the row identifier is stored in the row indexed as the row identifier in the group table; if the row identifier is not queried in the group table, an idle row is searched in the group table, and the idle row index is set to the row identifier, and the burst request corresponding to the row identifier is stored in the row indexed as the row identifier in the group table.
[0086] In one embodiment, the grouping table is stored in a content addressable memory (CAM) combined with a first-in-first-out queue (FIFO), wherein the content addressable memory is used to index a row identifier, and the first-in-first-out queue stores the burst request of the row. In addition, other components that can implement index and sequence storage can also be used to implement storage of the grouping table, and the present invention is not limited thereto.
[0087] In addition, in one embodiment, the hardware filter 10 further comprises: a second storage module 15 for storing the burst requests to be filtered out obtained in each round of screening.
[0088] In one embodiment, reference Figure 2As shown in , at least one filter 12 specifically includes a first filter 121 and a second filter 122_1...122_n, wherein the output ends of the first filter 121 and the second filter 122_1...122_n are connected to the first storage module 14 and the second storage module 15, and the input ends of the second filter 122_1...122_n are simultaneously connected to the first storage module 14. The first filter 121 is used to perform a first round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in the first round. The second filter 122_1...122_n can be set to one or multiple in series, and each second filter can perform a round of screening. The burst requests to be retained obtained by each round of screening output by the first filter 121 and each second filter 122 are stored in the first storage module 14 for storage, and the burst requests to be screened out are stored in the second storage module 14 for storage. During the current round of screening, the second filter is used to select at least part of the rows of the burst requests from the grouping table 141 corresponding to the previous round of the first storage module 14 to perform screening. During the screening process, retention is determined based on the granularity of the entire row, and the burst requests to be retained and the burst requests to be screened out in the current round are identified.
[0089] Further references Figure 2 As shown in , in one embodiment, the hardware filter 10 is further configured with a trigger 16, which is connected to the first storage module 14 and each second filter 122_1...122_n. The trigger 16 is used to receive the external state change information generated by the grouping table, and decide whether to trigger the output action, so as to control the first storage module 14 and each second filter 122_1...122_n to select the burst request of at least part of the rows from the grouping table to perform filtering when the trigger triggers the output action. The external state change information includes but is not limited to: the number of valid burst requests accommodated by the grouping table, the number of valid rows, the size of the rows inserted in the current round of filtering, the number of burst requests inserted within a given time or the number of rows involved, the current timestamp, the timestamp of the previous round of output action, etc.
[0090] When the trigger 16 decides to trigger the output action, each second filter 122_1...122_n selects at least part of the burst requests in the grouping table 141 to perform filtering until the grouping table is empty or the predefined conditions are met. The second filter can specifically perform filtering through the set row comparison logic, which can be a full order method or content-independent, and can also add specific conditions. For example: the full order method is set to sort by the size of the row (including the number of bursts), and the large rows are retained first and the small rows are filtered out first; the additional specific condition is to sort independently by channel (or rank, bank, etc.) to achieve the purpose of load balancing. If a specific full order is adopted, the comparison logic can use a binary tree form to reduce delay, or a lower-overhead implementation method can be adopted, because the frequency of the output action is lower than the burst input, and the single delay can be higher. The predefined stop condition can be: reaching a certain number of outputs, each channel (or rank, bank, etc.) has output, the number of rows remaining in the grouping table / total size is lower than a certain proportion, etc. For example, in one embodiment, the second filter uses a full order method of row size for filtering. By calculating the surplus δ of the second filter, the value of the surplus δ is the number of all burst requests filtered by the second filter*(1-the screening ratio p of the second filter) r ). If the balance is greater than or equal to zero, find the largest row in the grouping table, and identify all burst requests stored in the row as the burst requests to be retained in the current round; if the balance is less than zero, find the smallest row in the grouping table, and identify all burst requests stored in the row as the burst requests to be filtered out in the current round, and update the balance δ accordingly.
[0091] In addition, in one embodiment, the screening logic of the first filter can be set to be independent of the content or related to the content. Applying the first filter can improve the memory access efficiency and reduce the actual total memory access at a lower cost, but still achieve the regularization effect of software dropout and ultimately do not affect the accuracy of the model.
[0092] In this embodiment, a hardware filter receives sparse memory access requests from inside the graph neural network accelerator, groups the sparse memory access requests into several burst requests according to the minimum unit burst of DRAM access, and groups the memory accesses according to the granularity of DRAM rows, and filters and schedules the execution in groups. The hardware filter outputs the burst requests to be retained in the final round and the burst requests to be screened out in all rounds. The burst requests to be retained are then sent to the memory controller for execution as normal, and the correct memory access results are returned after execution, while the requests that are screened out are assigned false zero value results and returned directly. At this point, the hardware filter effectively improves the effective ratio of memory access, improves locality, reduces the amount of memory access, and achieves an acceleration effect.
[0093] refer to Figure 3 As shown in , another embodiment of the present invention further discloses a graph neural network accelerator 100, comprising: a graph neural network acceleration module 20, a hardware filter 10, a memory controller 30, a false zero value generation module 40, and a mixing module 50, wherein the graph neural network acceleration module 20 is used to generate sparse memory access requests. The hardware filter 10 at least comprises: an input interface 11, at least one filter 12, and an output interface 13, wherein the input interface 11 is connected to the graph neural network acceleration module 20, and is used to receive the sparse memory access request from the inside of the graph neural network accelerator module, and group the sparse memory access request into a plurality of burst requests according to the minimum unit burst of DRAM access. At least one filter 12 is used to perform at least one round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in each round. The output interface 13 is used to output the burst requests to be retained in the final round and the burst requests to be screened out in all rounds. In addition, for the specific structure of the hardware filter 10, reference may be made to the corresponding solutions in the aforementioned embodiments, which will not be described in detail here for the convenience and brevity of description.
[0094] The memory controller 30 is connected to the output interface 13, and is used to receive the burst request to be retained in the final round and return the correct memory access result. The false zero value generation module 40 is connected to the output interface 13, and is used to receive the burst request to be screened out in all rounds and generate a false zero value result. The mixing module 50, the input end of the mixing module is connected to the memory controller 30 and the false zero value generation module 40, and the output end is connected to the graph neural network acceleration module 20, so as to obtain the correct memory access result and the false zero value result to generate a sparse memory access result and feed it back to the graph neural network acceleration module 20.
[0095] In this embodiment, the existing graph neural network accelerator is improved by adding a hardware filter, a false zero value generation module, a mixing module, etc. The hardware filter groups the sparse memory access requests received from the graph neural network accelerator into several burst requests, and groups the memory accesses according to the granularity of DRAM rows, and filters and schedules the execution in groups. The hardware filter outputs the burst requests to be retained in the final round and feeds them back to the memory controller for execution, returning the correct memory access results; at the same time, the hardware filter outputs the burst requests to be screened out in all rounds, and the false zero value generation module generates a false zero value result. Furthermore, the mixing module mixes the correct memory access result with the false zero value result to generate a sparse memory access result and feeds it back to the graph neural network acceleration module. At this point, the improved graph neural network accelerator groups the read memory access requests in the aggregation stage according to the two granularities of DRAM's minimum access unit burst and row, and performs screening and scheduling execution, which naturally ensures the effectiveness of DRAM access, eliminates inefficient memory access, and reduces expensive row activation operations, reducing energy consumption. In general, it reduces the amount of memory access and improves locality, thereby achieving an acceleration effect. In addition, any existing graph neural network accelerator can be superimposed with the hardware filter of the above embodiment to achieve performance and energy efficiency improvements.
[0096] In addition, based on the above Figure 3 The graph neural network accelerator structure shown in FIG. 1 is a graph neural network accelerator. In another embodiment of the present invention, an off-chip memory access screening method using a graph neural network accelerator is also disclosed. Figure 4 As shown, Figure 4 A detailed schematic diagram of the method is shown.
[0097] A method for selecting off-chip memory access using a graph neural network accelerator, the method comprising:
[0098] Step S1: Receive a sparse memory access request from inside the graph neural network accelerator, and group the sparse memory access request into several burst requests according to the minimum unit burst of DRAM access.
[0099] Step S2: The hardware filter performs at least one round of screening on the input burst requests, and identifies the burst requests to be retained and the burst requests to be screened out in each round.
[0100] Step S3: The memory controller receives the burst request to be retained in the final round and returns a correct memory access result;
[0101] Step S4: The false zero value generating module receives the burst requests to be screened out in all rounds and generates a false zero value result.
[0102] Step S5: Obtain the correct memory access result and the false zero value result to generate a sparse memory access result and feed it back to the graph neural network accelerator.
[0103] In one embodiment, in step S2, when performing at least one round of screening on the input burst requests, in the current round of screening, the burst requests to be retained obtained in the previous round are screened.
[0104] In one embodiment, in step S2, the burst requests to be retained obtained in each round of screening are stored in the form of a group table. Specifically, the rows of the group table are indexed by the unique identifier of the DRAM row; in combination with the organizational form and address mapping method of the DRAM, the rows of the burst requests to be retained obtained in each round of screening in the DRAM are obtained, and the row identifier corresponding to each burst request is generated. If the row identifier is found in the group table, the burst request corresponding to the row identifier is stored in the row of the group table indexed as the row identifier; if the row identifier is not found in the group table, a free row is searched in the group table, the free row index is set to the row identifier, and the burst request corresponding to the row identifier is stored in the row of the group table indexed as the row identifier.
[0105] In one embodiment, when performing at least one round of screening on the input burst requests, the burst requests of at least part of the rows are selected from the grouping table corresponding to the previous round for screening. During the screening process, the retention is determined according to the granularity of the entire row, and the burst requests to be retained and the burst requests to be screened out in the current round are identified. Specifically, the second screener uses the total order of the row size to perform screening. The balance of the second screener is calculated, and the value of the balance is the number of all burst requests screened by the second filter*(1-screening ratio of the filter). If the balance is greater than or equal to zero, the largest row in the grouping table is found, and all burst requests stored in the row are identified as the burst requests to be retained in the current round; if the balance is less than zero, the smallest row in the grouping table is found, and all burst requests stored in the row are identified as the burst requests to be screened out in the current round, and the balance is updated at the same time.
[0106] In addition, in order to distinguish the filterable sparse memory access requests from other normal requests on the input, in one embodiment, the QoS signal interface in the AXI protocol can be reused, and a specific value can be used as a "filterable" mark to coexist with levels such as "critical", "normal", and "best-effort". In one embodiment, a special signal interface is added, which, together with the memory access information, directly indicates whether the memory access can be filtered out. In one embodiment, several groups of configurable registers are added to record the address space that can be filtered out, and the read requests falling into these spaces can be filtered out by default.
[0107] Usually, in terms of output, there is no need to distinguish whether the returned result is a false value, but some scenarios require specific mask results, such as the back-propagation process of a neural network. In this case, the address information of the mask can be additionally configured, and a write request can be generated and executed synchronously while filtering. Usually, graph neural networks traverse in order of edges, so these write requests are usually continuous and of small granularity, so the additional overhead is relatively limited.
[0108] In addition, it should be noted that the off-chip memory access screening method disclosed in the present invention is not limited to being executed in the order of the above steps S1-S5, and the steps can be reordered, added or deleted. For example, the steps recorded in the present invention can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of the present invention can be achieved, and the present invention is not limited here.
[0109] According to an embodiment of the present invention, the present invention also provides a machine-readable medium, which may be a tangible medium, which may contain or store a program for use by an instruction execution system, apparatus or device or for use in conjunction with an instruction execution system, apparatus or device, and the computer program implements the steps of the above-mentioned off-chip memory access screening method when executed by a processor.
[0110] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored in a readable storage medium. When the computer program is executed by a processor, the computer can execute the off-chip memory access screening method provided by the above methods.
[0111] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A hardware filter, characterized in that: include: An input interface for receiving sparse memory access requests from the graph neural network accelerator and grouping the sparse memory access requests into several burst requests according to the minimum unit burst of DRAM access; at least one filter, configured to perform at least one round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in each round; The output interface is used to output the burst requests to be retained in the final round and the burst requests to be filtered out in all rounds.
2. The hardware filter according to claim 1, characterized in that: The at least one filter is further used to: in the current round of screening process, screen the burst requests to be retained obtained in the previous round.
3. The hardware filter according to claim 2, characterized in that: Also includes: The first storage module is used to store the burst requests to be retained obtained in each round of screening in the form of a grouping table.
4. The hardware filter according to claim 3, characterized in that: The burst requests to be retained obtained from each round of screening are stored in the form of a grouping table, including: Indexing the rows of the grouping table by the unique identifier of the DRAM row; In combination with the organization form and address mapping mode of DRAM, the rows of the burst requests to be retained obtained in each round of screening in the DRAM are obtained, and a row identifier corresponding to each burst request is generated; If the row identifier is found in the grouping table, the burst request corresponding to the row identifier is stored in the grouping table in a row indexed by the row identifier; If the row identifier is not found in the grouping table, an idle row is searched in the grouping table, the idle row index is set as the row identifier, and the burst request corresponding to the row identifier is stored in the grouping table in the row indexed as the row identifier.
5. The hardware filter according to claim 4, characterized in that: The first storage module adopts a content addressable memory combined with a FIFO first-in first-out queue for storage, wherein the content addressable memory is used to index a row identifier, and the first-in first-out queue stores the burst request of the row.
6. The hardware filter according to claim 4, characterized in that: The at least one filter comprises: A first filter, used for performing a first round of screening on the input burst requests, and identifying the burst requests to be retained and the burst requests to be screened out in the first round; The second filter, during the current round of screening, is used to select at least part of the rows of the burst requests from the grouping table corresponding to the previous round of the first storage module to perform screening. During the screening process, retention is determined based on the granularity of the entire row, and the burst requests to be retained and the burst requests to be screened out in the current round are identified.
7. The hardware filter according to claim 6, characterized in that: The second filter uses a full order of row sizes to perform filtering; Calculate the balance of the second filter, where the value of the balance is the number of all burst requests filtered by the second filter*(1-the screening ratio of the second filter); If the balance is greater than or equal to zero, find the largest row in the grouping table, and identify all burst requests stored in the row as the burst requests to be retained in the current round; If the balance is less than zero, the smallest row in the grouping table is found, and all burst requests stored in the row are identified as the burst requests to be screened out in the current round.
8. The hardware filter according to claim 6, characterized in that: Also includes: A trigger is connected to the first storage module and the second filter, and is used to receive the external state change information generated by the grouping table, determine whether to trigger an output action, and when the trigger triggers the output action, select the burst request of at least part of the rows in the grouping table to perform filtering.
9. The hardware filter according to claim 8, characterized in that: The external state change information includes at least one of the number of valid burst requests accommodated in the grouping table, the number of valid rows, the size of the rows inserted in the current round of screening, the number of burst requests inserted within a given time or the number of rows involved, the current timestamp, and the timestamp of the previous round of output actions, or a combination of more.
10. The hardware filter according to claim 1, characterized in that: Also includes: The second storage module is used to store the burst requests to be screened out obtained in each round of screening.
11. A graph neural network accelerator, characterized in that: Include: A graph neural network acceleration module to generate sparse memory access requests; The hardware filter is the hardware filter according to any one of claims 1 to 10, and the hardware filter at least comprises: An input interface, connected to the graph neural network acceleration module, for receiving the sparse memory access request from the graph neural network accelerator module, and grouping the sparse memory access request into a plurality of burst requests according to the minimum unit burst of DRAM access; at least one filter, configured to perform at least one round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in each round; An output interface, used to output the burst requests to be retained in the final round and the burst requests to be screened out in all rounds; A memory controller connected to the output interface, configured to receive the burst request to be retained in the final round and return a correct memory access result; a false zero value generating module, connected to the output interface, for receiving the burst requests to be screened out of all rounds to generate a false zero value result; A hybrid module, wherein the input end of the hybrid module is connected to the memory controller and the false zero value generation module, and the output end is connected to the graph neural network acceleration module, so as to obtain the correct memory access result and the false zero value result to generate a sparse memory access result which is fed back to the graph neural network acceleration module.
12. A method for selecting off-chip memory access using a graph neural network accelerator, characterized in that: The method includes: Receive sparse memory access requests from the graph neural network accelerator, and group the sparse memory access requests into several burst requests according to the minimum unit burst of DRAM access; Perform at least one round of screening on the input burst requests, and identify the burst requests to be retained and the burst requests to be screened out in each round; The memory controller receives the burst request to be retained in the final round and returns a correct memory access result; Receiving the burst requests to be screened out from all rounds to generate a false zero value result; The correct memory access result and the false zero value result are obtained to generate a sparse memory access result and feed it back to the graph neural network accelerator.
13. The method according to claim 12, characterized in that When performing at least one round of screening on the input burst requests, in the current round of screening, the burst requests to be retained obtained in the previous round are screened.
14. The method according to claim 12, characterized in that Also includes: The burst requests to be retained obtained in each round of screening are stored in the form of a grouping table, including: Indexing the rows of the grouping table by the unique identifier of the DRAM row; In combination with the organization form and address mapping mode of DRAM, the rows of the burst requests to be retained obtained in each round of screening in the DRAM are obtained, and a row identifier corresponding to each burst request is generated; If the row identifier is found in the grouping table, the burst request corresponding to the row identifier is stored in the grouping table in a row indexed by the row identifier; If the row identifier is not found in the grouping table, an idle row is searched in the grouping table, the idle row index is set as the row identifier, and the burst request corresponding to the row identifier is stored in the grouping table in the row indexed as the row identifier.
15. The method according to claim 14, characterized in that When performing at least one round of screening on the input burst requests, at least part of the rows of the burst requests are selected from the grouping table corresponding to the previous round to perform screening. During the screening process, retention is determined based on the granularity of the entire row, and the burst requests to be retained and the burst requests to be screened out in the current round are identified.
16. The method according to claim 15, characterized in that During the screening process, the retention or removal is determined based on the granularity of the entire row, and the burst requests to be retained and the burst requests to be screened out in the current round are identified, including: The filter is filtered by the total order of row size; Calculate the balance of the filter, the value of which is the number of all burst requests filtered by the filter*(1-filter elimination ratio); If the balance is greater than or equal to zero, find the largest row in the grouping table, and identify all burst requests stored in the row as the burst requests to be retained in the current round; If the balance is less than zero, the smallest row in the grouping table is found, and all burst requests stored in the row are identified as the burst requests to be screened out in the current round.