Data prefetcher for graph calculation, processor and equipment
By designing a graph-oriented data prefetcher in the processor, the problem of untimely and inaccurate data prefetching in graph calculations is solved, and more efficient data access and processor performance improvement is achieved.
Patent Information
- Application Number
- CN202510527371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The prior art is difficult to accurately and timely data prefetching in graph calculations, resulting in memory access bottlenecks and affecting the processor's graph calculation performance.
A data prefetcher for graph computing is designed, including a prefetch unit located between each processor core and its secondary cache L2. The prefetch unit includes a configuration register group, a status register group, a control state machine, a prefetch pipeline and a data cache. The graph data access prefetch logic is executed through a pipeline to improve the timeliness and accuracy of prefetching.
It effectively alleviates the memory access bottleneck in graph calculation, improves the timeliness and accuracy of prefetching, thereby improving the efficiency of graph calculation and the graph computing performance of processors.
Smart Images

Figure CN120066987A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of processors, and particularly relates to a data prefetcher, a processor and a device for graph computing. Background Art
[0002] Data prefetching is a technology that, before a user requests data, predicts the address of memory access in advance, fetches the potentially needed data in advance, and stores it in a cache or a temporary memory, thereby hiding the memory access latency and improving the working speed of the processor. Traditional data prefetching technologies predict future memory access addresses based on the memory access patterns of application programs. They search for data from memory or a lower-level cache according to the predicted address so that when a memory access request arrives, the corresponding data has been stored at the corresponding cache level, thereby reducing the memory access latency. Data prefetching can be divided into hardware prefetching and software prefetching according to the implementation method.
[0003] The most common prefetcher of the hardware prefetcher is the stride prefetcher design, which adopts a fixed and repeatable address difference in a sequence. However, traditional data prefetcher such as GHB and AMPM usually perform prefetching by learning the stride history in the address stream, so they can only recognize the streaming memory access patterns with deterministic rules. To further capture the indirect access data stream, Yu et al. proposed a hardware mechanism IMP that supports indirect data stream prefetching. First, it identifies the sequential data stream and, based on this, identifies the indirect access pattern. However, it is difficult for these prefetcher to fully capture the four-level indirect access pattern D[C[B[A[i]]+j]] in graph computing, where i is in [0,n-1] and j is in [B[A[i]],B[A[i]+1]], and because a large amount of memory is required to record the access history, the implementation overhead is very high. Ainsworth proposed a graph structure-aware data prefetcher. This prefetcher needs to listen to data requests to prefetch the data of subsequent nodes, but it needs to be specially optimized for the prefetching timing. At the same time, because the prefetched data is stored in the L1 cache, special consideration needs to be given to the resource waste caused by replacement.
[0004] Software prefetching is implemented by adding some prefetch instructions to the program by the programmer or the compiler. For example, Ainsworth proposed a compiler-based indirect mode prefetcher. They can usually capture the memory access pattern more accurately. Software prefetching has inherent limitations: it may not be portable across microarchitectures (for example, with different cache hierarchies), resulting in significant instruction overhead and increased power consumption, thus masking the advantage of improving the cache hit rate.
[0005] Ideal prefetching techniques introduce little to no additional overhead to memory accesses. However, in practical scenarios, prefetching is not always timely or accurate, and lagging or incorrect prefetching can affect performance. Therefore, prefetching techniques must consider three fundamental issues: how to determine the prefetched data, i.e., the address of the prefetched data; when to prefetch the data; and where to place the prefetched data. This design targets the data access pattern of graph computing, can accurately prefetch all relevant data, and stores all prefetched data in a dedicated prefetch cache, thus not contaminating the cache.
[0006] Graphs in real-world scenarios are highly sparse, with most nodes having no connected edges. A relatively small number of nodes in the graph structure are connected to a large number of nodes. Therefore, to improve the storage efficiency of graph data, the Compressed Sparse Row (CSR) format is usually used to store graph data, which has high time and space utilization. As Figure 1 shown, the compressed sparse row format uses four arrays to store graph data, where: the offset array stores the offset value of the first out-edge of each node in the edge array, the edge array stores the numbers (ids) of all out-edge nodes in node order, the weight array stores the weights (weights) of the edges corresponding to the edge array, and the property data of the nodes related to graph computing (such as V 0 ~V 4 ) are stored in the property array.
[0007] Each update operation in the graph computing process requires multiple memory access operations. When updating an out-edge vertex, multiple data including the out-edge vertex id, out-edge weight weight, and property need to be accessed, while the computation usually only performs reduction operations with relatively small computational amounts. Therefore, when graph computing applications are executed on a general-purpose processor platform, they have the characteristics of memory access-intensive and present three challenges: 1) Weak locality: Due to the complex and variable graph data structure, the memory access behavior of the program has the characteristic of low locality. Although a large number of redundant structures have been found in graph data for real-world application scenarios, current graph computing applications still have difficulty accurately predicting the location of redundant structures at runtime, thus limiting the optimization of data access locality. 2) High concurrency: The combination of a large number of data accesses generated by traversing vertex or edge data and simple reduction computations causes the application to generate a large number of memory access requests in a short period of time, resulting in the processor being unable to obtain the required data in a timely manner under high-concurrency conditions. 3) Fine-grained access: The computational operations in graph computing applications usually only involve a certain property of a vertex or an edge, and the space occupied by these properties is small, making the memory access behavior of graph computing applications exhibit the characteristic of fine-grained access.
[0008] The traditional von Neumann general-purpose processor architecture uses a unified off-chip memory resource to store instructions and data. However, for memory-intensive applications such as computing applications, due to the existence of large-scale and highly concurrent memory access requests, general-purpose processors based on the von Neumann architecture can only obtain the data required for the operation instructions after consuming a large number of idle waiting cycles, resulting in the instruction pipeline being unable to be filled with a sufficient number of computing instructions, and thus unable to fully utilize instruction-level parallelism.
[0009] Pipeline blocking is the direct cause of processor performance degradation. The main conflicts that cause out-of-order superscalar processor blocking are resource-related, data-related, and control-related. By analyzing the proportion of pipeline blocking cycles caused by data-related, we can analyze the graph data access bottleneck. Figure 2 The figure shows the memory access instruction blocking cycle ratio obtained by executing the Single-Source Shortest Path (SSSP) algorithm on different data sets. Figure 2 It can be seen that the average blocking cycle of load instructions accounts for 38.02%, while that of store instructions accounts for only 19.77%. This is mainly because load instructions are on the critical execution path, and subsequent calculation instructions depend on the results of load instructions. Therefore, optimizing load instructions has a greater impact on overall performance. The key to optimizing load instructions lies in data prefetching. As mentioned above, ideal prefetching technology will hardly cause any additional overhead to memory access. However, in actual situations, prefetching is not always timely or accurate. Therefore, for data prefetchers for graph computing, how to alleviate the memory access bottleneck of graph computing, improve the timeliness and accuracy of prefetching, and improve the efficiency of graph computing has become a key technical issue that needs to be solved urgently. Summary of the invention
[0010] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a data prefetcher, processor and device for graph computing are provided. The present invention aims to improve the timeliness and accuracy of prefetching for graph computing, so as to improve the efficiency of graph computing, alleviate the memory access bottleneck of graph computing, and improve the graph computing performance of the processor.
[0011] In order to solve the above technical problems, the technical solution adopted by the present invention is: A data prefetcher for graph computing includes a prefetch unit located between each processor core and its L2 cache L2, wherein the prefetch unit includes: Configuration register group, used to store data related to the graph computing data structure configured by the software, including array addresses related to the work queue and compressed sparse row format; A status register set, which is used to maintain intermediate index variables of the prefetch unit during the prefetch process and provide them to the control state machine to generate control signals. Specifically, it includes a work queue index, an out-edge offset index, and a flag indicating that the data cache is full. A control state machine, which is used to control the execution of the prefetch pipeline by reading the contents of the configuration register set and the status register set, including pausing and resuming the prefetch pipeline. A prefetch pipeline, which is used to execute each stage of the graph data access prefetch logic in a pipeline manner. The operations of each stage include calculating the prefetch address, sending a prefetch request, and receiving prefetch data. A data cache, which is used to store the data prefetched by the prefetch pipeline. The data cache maintains the prefetched data in terms of edges. Each entry contains all the information related to the edge, and the edge information is saved to the cache in the last stage of the pipeline.
[0012] Optionally, the execution of each stage of the graph data access prefetch logic in a pipeline manner includes: Stage 1, accessing the work queue listArray to prefetch the node n to be processed; Stage 2, accessing the attribute array dataArray and the offset array offsetArray to prefetch the attributes and offsets of the source node; Stage 3, accessing the edge array edgeArray and the weight array weightArray to prefetch the out-edge nodes and weights; Stage 4, accessing the attribute array dataArray to prefetch the weight of the destination node.
[0013] Optionally, the array addresses related to the work queue and the compressed sparse row format saved in the configuration register set include: the start address of the work queue, the number of elements in the work queue, the start address of the offset data, the start address of the edge array, the start address of the edge weight array, and the start address of the attribute array.
[0014] Optionally, the intermediate index variables in the status register set include a column index idx, a current column index current_idx, a start offset of the current row start, an end offset of the current row end, and a flag full indicating that the data cache is full. The flag full indicating that the data cache is full is used as a credential for the control state machine to pause and resume the prefetch pipeline. If the flag full indicating that the data cache is full is true, the prefetch pipeline is paused; otherwise, the prefetch pipeline is resumed.
[0015] Optionally, the fields of each piece of data in the data cache include a source node ID, a destination node ID, source node attributes, destination node attributes, and an out-edge weight.
[0016] Optionally, the data prefetcher provides configuration interface functions config(addr, data), fetch data interface function fetch_edge(), and buffer empty status judgment interface function empty() for software use, where addr is the address of the data prefetcher, data is the configuration data, and the ways for software to use the data prefetcher include: S1. Configure the data prefetcher by software calling the configuration interface function config(addr, data). The configuration information for configuration includes the starting addresses and sizes of the offset array offsetArray, edge array edgeArray, weight array weightArray, and attribute array dataArray. The offset array offsetArray is used to record the starting and ending offsets of out-edges, the edge array edgeArray is used to record the ids of out-edge nodes, the weight array weightArray is used to record out-edge weights, and the attribute array dataArray is used to record status data; S2. Configure the starting address and the number of elements len of the work queue listArray by software, so that the data prefetcher starts to prefetch the graph data related to all nodes to be processed in the work queue listArray. The graph data includes source node ID, destination node ID, source node attributes, destination node attributes, and out-edge weights, and stores the relevant graph data in the data cache; S3. Obtain the data in the data cache by software calling the fetch data interface function fetch_edge(); S4. Execute the corresponding graph calculation using the prefetched data by software and update the attributes of out-edge nodes; S5. Judge whether there is still data to be processed in the data cache of the data prefetcher by software calling the buffer empty status judgment interface function empty(). If there is still data to be processed, continue to execute step S3; otherwise, end and exit.
[0017] Optionally, the data prefetcher starts to prefetch the graph data related to all nodes to be processed in the work queue listArray in step S2, including: S2.1. Obtain the starting address and the number of elements len of the input work queue listArray; initialize the column index idx to 0; S2.2. Judge whether the column index idx is less than the number of elements len of the work queue. If the column index idx is not less than the number of elements len of the work queue, end; otherwise, jump to step S2.3; S2.3, Take out the index number i of the node to be processed corresponding to the column index idx from the work queue listArray; increment the column index idx by 1, and jump to step S2.3 until the numbers of all nodes to be processed are generated; S2.4, Prefetch the node attribute srcData of the node to be processed according to the starting address of the attribute array dataArray and the index number i of the node to be processed. Take the offset corresponding to the index number i of the node to be processed from the offset array offsetArray as the starting offset start of the current row, and the offset corresponding to the index number i + 1 of the next node as the ending offset end of the current row; S2.5, Assign the starting offset start of the current row to the current column index current_idx; S2.6, Determine whether it holds that the current column index current_idx is less than the ending offset end of the current row. If it does not hold, jump to step S2.2; otherwise, jump to the next step; S2.7, Determine whether the data cache is full. If it is full, jump back to step S2.7; otherwise, jump to the next step; S2.8, Take out the element corresponding to the current column index current_idx from the edge array edgeArray as the out-edge node node corresponding to the node to be processed i, and take out the element corresponding to the current column index current_idx from the weight array weightArray as the edge weight weight; S2.9, Read the attribute data of the out-edge node node corresponding to the node to be processed i from the attribute array dataArray according to the node node; increment the current column index current_idx by 1, and jump to step S2.6.
[0018] In addition, the present invention further provides a processor, including a processor body with multiple processor cores, and the data prefetcher for graph computing is provided in the processor body.
[0019] In addition, the present invention further provides a computer device, including a processor and a memory connected to each other, and the data prefetcher for graph computing according to any one of the above is provided in the processor.
[0020] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: To alleviate the memory access bottleneck of graph computing, the data prefetcher of the present invention includes a prefetch unit located between each processor core and its level-2 cache L2. The prefetch unit includes: a configuration register group for storing data related to the graph computing data structure configured by software; a status register group for maintaining intermediate index variables during the prefetch process of the prefetch unit; a control state machine for controlling the execution of the prefetch pipeline by reading the contents of the configuration register group and the status register group; a prefetch pipeline for executing each stage of the graph data access prefetch logic in a pipelined manner; and a data cache for storing the data prefetched by the prefetch pipeline. The present invention can improve the timeliness and accuracy of prefetching for graph computing, so as to improve the efficiency of graph computing, alleviate the memory access bottleneck of graph computing, and improve the graph computing performance of the processor. Description of the Drawings
[0021] Figure 1 It is a schematic diagram of the Compressed Sparse Row (CSR) format in the prior art.
[0022] Figure 2 It is the proportion of memory access instruction blocking cycles of the single-source shortest path algorithm in the prior art.
[0023] Figure 3 It is a flowchart of graph data access in the embodiment of the present invention.
[0024] Figure 4 It is a schematic diagram of the structure of the data prefetcher in the embodiment of the present invention.
[0025] Figure 5 It is a schematic diagram of the structure of the prefetch unit in the data prefetcher in the embodiment of the present invention.
[0026] Figure 6 It is an overall execution flowchart of the cooperation mechanism in the embodiment of the present invention.
[0027] Figure 7 It is a schematic diagram of the working process of the data prefetcher in the embodiment of the present invention.
[0028] Figure 8 It is the normalized execution time of different algorithms in the embodiment of the present invention.
[0029] Figure 9 It is the proportion of memory access instruction blocking cycles of different algorithms normalized in the embodiment of the present invention. Detailed Embodiments
[0030] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0031] As shown Figure 3 in the figure, the steps of the data prefetcher for graph data access in this embodiment include: S101, initialize the column index idx to 0, and the number of elements len in the work queue is the size listArray.size() of the work queue listArray; S102, determine whether the column index idx is less than the number of elements len in the work queue. If not, end; otherwise, jump to the next step; S103, take out the node listArray[idx] corresponding to the column index idx from the work queue listArray as the node i to be processed, and increment the column index idx by 1 (idx += 1); S104, take out the offset offsetArray[i] corresponding to the node i to be processed from the offset array offsetArray as the starting offset start of the current row, and the offset offsetArray[i + 1] corresponding to the next node i + 1 as the ending offset end of the current row; S105, assign the starting offset start of the current row to the current column index current_idx; S106, determine whether the current column index current_idx is less than the ending offset end of the current row. If not, jump to step S102; otherwise, jump to the next step; S107, take out the element edgeArray[current_idx] corresponding to the current column index current_idx from the edge array edgeArray as the out-edge node node corresponding to the node i to be processed, and take out the element weightArray[current_idx] corresponding to the current column index current_idx from the weight array weightArray as the edge weight weight; S108, read the attribute data dataArray[node] of the out-edge node node corresponding to the node i to be processed from the attribute array dataArray according to the node node; S109, increment the current column index current_idx by 1 (current_idx += 1); jump to step S106.
[0032] To alleviate the memory access bottleneck of graph computing, as Figure 4 shown in the figure, this embodiment provides a data prefetcher for graph computing, which includes a prefetch unit located between each processor core and its secondary cache L2. The prefetch unit is placed beside the L1 cache to reduce the latency of the processor accessing the prefetch data, as Figure 5 shown in the figure, the prefetch unit includes: A configuration register group for storing data related to the graph computing data structure configured by software, including the array addresses related to the work queue and the compressed sparse row format; A status register group is used to maintain intermediate index variables during the prefetching process of the prefetch unit and provide them to the control state machine to generate control signals, specifically including a work queue index, an out-edge offset index, and a flag indicating that the data cache is full; A control state machine is used to control the execution of the prefetch pipeline by reading the contents of the configuration register group and the status register group, including pausing and resuming the prefetch pipeline; A prefetch pipeline is used to execute each stage of the graph data access prefetch logic in a pipelined manner. The operations in each stage include calculating the prefetch address, sending a prefetch request, and receiving prefetch data; A data cache is used to store the data prefetched by the prefetch pipeline. The data cache maintains the prefetched data in terms of edges. Each entry contains all the information related to an edge, and the edge information is saved to the cache in the last stage of the pipeline.
[0033] As Figure 5 shown, in this embodiment, the stages of executing the graph data access prefetch logic in a pipelined manner include: Stage 1, accessing the work queue listArray to prefetch the node n to be processed; Stage 2, accessing the attribute array dataArray and the offset array offsetArray to prefetch the attributes and offsets of the source node; Stage 3, accessing the edge array edgeArray and the weight array weightArray to prefetch the out-edge nodes and weights; Stage 4, accessing the attribute array dataArray to prefetch the weight of the destination node.
[0034] As Figure 5 shown, the array addresses related to the work queue and the compressed sparse row format saved in the configuration register group in this embodiment include: the starting address of the work queue, the number of elements in the work queue, the starting address of the offset data, the starting address of the edge array, the starting address of the edge weight array, and the starting address of the attribute array.
[0035] As Figure 5 shown, the intermediate index variables in the status register group in this embodiment include a column index idx, a current column index current_idx, a starting offset start of the current row, an ending offset end of the current row, and a flag full indicating that the data cache is full. The flag full indicating that the data cache is full is used as a basis for the control state machine to pause and resume the prefetch pipeline. If the flag full indicating that the data cache is full is true, the prefetch pipeline is paused; otherwise, the prefetch pipeline is resumed.
[0036] As Figure 5 shown, the fields of each piece of data in the data cache in this embodiment include the source node ID, the destination node ID, the source node attributes, the destination node attributes, and the out-edge weight.
[0037] As Figure 6As shown in the figure, the data prefetcher in this embodiment provides a configuration interface function config(addr, data), a data fetching interface function fetch_edge(), and a buffer empty status judgment interface function empty() for software use. Here, addr is the address of the data prefetcher, data is the configuration data, and the ways for software to use the data prefetcher include: S1. Configure the data prefetcher by software calling the configuration interface function config(addr, data). The configuration information for configuration includes the start addresses and sizes of the offset array offsetArray, edge array edgeArray, weight array weightArray, and attribute array dataArray. The offset array offsetArray is used to record the start and end offsets of the outgoing edges. The edge array edgeArray is used to record the IDs of the outgoing edge nodes. The weight array weightArray is used to record the outgoing edge weights. The attribute array dataArray is used to record the status data. S2. Configure the start address of the work queue listArray and the number of work queue elements len by software, so that the data prefetcher starts to prefetch the graph data related to all the nodes to be processed in the work queue listArray. The graph data includes the source node ID, destination node ID, source node attributes, destination node attributes, and outgoing edge weights, and stores the relevant graph data in the data cache. S3. Obtain the data in the data cache by software calling the data fetching interface function fetch_edge(). S4. Execute the corresponding graph calculation using the prefetched data by software to update the attributes of the outgoing edge nodes. S5. Judge whether there is still data to be processed in the data cache of the data prefetcher by software calling the buffer empty status judgment interface function empty(). If there is still data to be processed, continue to execute step S3; otherwise, end and exit.
[0038] As Figure 7 shown, the graph data related to all the nodes to be processed in the work queue listArray that the data prefetcher starts to prefetch in step S2 of this embodiment includes: S2.1. Obtain the start address of the input work queue listArray and the number of work queue elements len; initialize the column index idx to 0. S2.2. Judge whether the column index idx is less than the number of work queue elements len. If the column index idx is not less than the number of work queue elements len, end; otherwise, jump to step S2.3. S2.3. Retrieve the index number i of the node to be processed corresponding to the column index idx from the work queue listArray; increment the column index idx by 1, and jump to step S2.3 until the numbers of all nodes to be processed are generated; S2.4. Prefetch the node attribute srcData of the node to be processed based on the starting address of the attribute array dataArray and the index number i of the node to be processed. Retrieve the offset corresponding to the index number i of the node to be processed from the offset array offsetArray as the starting offset start of the current row, and the offset corresponding to the index number i + 1 of the next node as the ending offset end of the current row; S2.5. Assign the starting offset start of the current row to the current column index current_idx; S2.6. Determine whether the current column index current_idx is less than the ending offset end of the current row. If not, jump to step S2.2; otherwise, jump to the next step; S2.7. Determine whether the data cache is full. If it is full, jump back to step S2.7; otherwise, jump to the next step; S2.8. Retrieve the element corresponding to the current column index current_idx from the edge array edgeArray as the out-edge node node of the node to be processed i obtained, and retrieve the element corresponding to the current column index current_idx from the weight array weightArray as the edge weight weight; S2.9. Read the attribute data of the out-edge node node of the node to be processed i from the attribute array dataArray according to the node node; increment the current column index current_idx by 1, and jump to step S2.6.
[0039] To verify the performance of the data prefetcher for graph computing in this embodiment, in this embodiment, the prefetcher is evaluated using a 4-core skylake processor on the ZSim simulator, and the specific configuration is shown in Table 1. At the same time, 6 graph data sets are used for testing, as shown in Table 2.
[0040] Table 1 Simulator Parameter Configuration
[0041] Table 2 Graph Data Sets
[0042] In this embodiment, three graph datasets in different fields are selected from the snap website of the Stanford University large-scale network dataset: Email-Eu-core, Wiki-Vote, and Soc-Epinions graph datasets. Each dataset has typical characteristics and complexities within its corresponding field, covering a variety of application scenarios and usage conditions. In addition, three automatically generated random graph datasets of different scales: graph1 - graph3 are designed to simulate various unknown or future possible graph structures, further enhancing the universality and robustness of the experiment.
[0043] To evaluate the performance of the data prefetcher on different graph algorithms, three graph algorithms are used in this embodiment, namely the BFS (Bread First Search) algorithm, the SSSP (Single-Source Shortest Path) algorithm, and the SSWP (Single-Source Widest Path) algorithm. The BFS algorithm traverses the graph to obtain the depth of all nodes in the graph from the origin, and can be used for friend relationships and community discovery problems in social network analysis, and for state space search and network link analysis in artificial intelligence and other fields. The SSSP algorithm is an algorithm for searching the single-source shortest path in a graph. In network routing, by analyzing the topological structure, the shortest path from one node to other nodes can be obtained, thus achieving efficient network routing and data transmission; in social network analysis, the shortest path from one node to other nodes can be obtained through this algorithm to achieve relationship analysis and influence evaluation of the social network. The SSWP algorithm is an algorithm for solving the single-source widest path. In network communication, the widest path refers to the maximum bandwidth path in the network, and can be used to search for the maximum bandwidth path between two nodes. In traffic planning, the widest path refers to the maximum traffic capacity of the road, and can be used to search for the maximum traffic capacity path between two places.
[0044] Figure 8 The normalized execution times of the three algorithms on different datasets are shown. Normalized to the benchmark system without an integrated data prefetcher, the smaller the time, the higher the performance. It can be seen that the data prefetcher in this embodiment can average improve the performance of the BFS algorithm, SSSP algorithm, and SSWP algorithm by 31%, 41%, and 36% respectively. At the same time, Figure 9The following shows the normalized proportion of memory access instruction blocking cycles of three algorithms on different data sets, normalized to the baseline system without an integrated prefetcher. The smaller the value, the higher the performance. The proportion of memory access instruction blocking cycles of all three algorithms has decreased significantly. Among them, the average degree of memory access instruction blocking of the BFS algorithm, SSSP algorithm, and SSWP algorithm has decreased by 95%, 66%, and 43% respectively. It can be seen that in graph computing, graph data access seriously blocks the pipeline execution, affecting the overall performance. Moreover, the data access pattern of graph computing is relatively fixed, and a dedicated data prefetcher can be designed according to this access pattern. The data prefetcher for graph computing in this embodiment, compared with the previous general data prefetching technology, can achieve more ideal prefetching performance due to its design for graph data access and customized prefetching scheme. The data prefetcher for graph computing in this embodiment addresses the deficiencies of the existing data prefetching mechanism. Starting from three key issues of data prefetching, it designs and customizes in a way that combines software configuration with hardware prefetching, and can be flexibly integrated into the existing graph computing framework to achieve user transparency. Existing graph computing usually uses existing frameworks for development. The method of adding software configuration in this embodiment can improve the efficiency of hardware prefetching without changing the original usage method. Through experiments on multiple algorithms and data sets, the data prefetcher for graph computing in this embodiment can improve the performance of the BFS algorithm, SSSP algorithm, and SSWP algorithm by 31%, 41%, and 36% on average.
[0045] In addition, this embodiment also provides a processor, including a processor body with multiple processor cores, and the data prefetcher for graph computing is provided in the processor body. This processor can be either a general-purpose processor or an accelerator processor for graph computing.
[0046] In addition, this embodiment also provides a computer device, including a processor and a memory connected to each other, and the data prefetcher for graph computing is provided in the processor. This computer device can be either a general-purpose computer device or a dedicated computer device for graph computing.
[0047] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A data prefetcher for graph computing, characterized in that: It includes a pre-fetch unit located between each processor core and its L2 cache, and the pre-fetch unit includes: Configuration register group, used to store data related to the graph computing data structure configured by the software, including array addresses related to the work queue and compressed sparse row format; The status register group is used to maintain the intermediate index variables of the prefetch unit during the prefetch process and provide them to the control state machine to generate control signals, including the work queue index, the outbound offset index, and the data cache full flag; A control state machine, for controlling the execution of the prefetch pipeline by reading the contents of the configuration register group and the status register group, including pausing and resuming the prefetch pipeline; The prefetch pipeline is used to execute various stages of graph data access prefetch logic in a pipeline manner. The operations of each stage include the calculation of prefetch addresses, the sending of prefetch requests, and the receiving of prefetch data. The data cache is used to store the data pre-fetched by the pre-fetch pipeline. The data cache maintains the pre-fetched data at the edge granularity. Each entry contains all the information related to the edge. The edge information is saved in the cache at the last stage of the pipeline.
2. The data prefetcher for graph computing according to claim 1, characterized in that: The various stages of executing graph data access prefetching logic in a pipeline manner include: stage 1, accessing the work queue listArray to prefetch the node n to be processed; stage 2, accessing the attribute array dataArray and the offset array offsetArray to prefetch the attributes and offsets of the source node; stage 3, accessing the edge array edgeArray and the weight array weightArray to prefetch the edge nodes and weights; stage 4, accessing the attribute array dataArray to prefetch the weight of the destination node.
3. The data prefetcher for graph computing according to claim 2, characterized in that: The array addresses related to the work queue and compressed sparse row format stored in the configuration register group include: the work queue starting address, the number of work queue elements, the offset data starting address, the edge array starting address, the edge weight array starting address and the attribute array starting address.
4. The data prefetcher for graph computing according to claim 3, characterized in that: The intermediate index variables in the status register group include the column index idx, the current column index current_idx, the current row start offset start, the current row end offset end and the data cache full flag full. The data cache full flag full is used as a credential for controlling the state machine to pause and resume the prefetch pipeline. If the data cache full flag full is true, the prefetch pipeline is paused, otherwise the prefetch pipeline is resumed.
5. The data prefetcher for graph computing according to claim 4, characterized in that: The fields of each data item in the data cache include a source node ID, a destination node ID, a source node attribute, a destination node attribute and an outbound edge weight.
6. The data prefetcher for graph computing according to claim 5, characterized in that: The data prefetcher provides a configuration interface function config(addr, data), a data fetching interface function fetch_edge(), and a buffer empty state judgment interface function empty() for software use, where addr is the address of the data prefetcher, data is the configuration data, and the software uses the data prefetcher in the following ways: S1, configure the data prefetcher by calling the configuration interface function config(addr, data) through software. The configuration information includes the starting address and size of the offset array offsetArray, the edge array edgeArray, the weight array weightArray, and the attribute array dataArray. The offset array offsetArray is used to record the starting offset and the ending offset of the outgoing edge, the edge array edgeArray is used to record the id of the outgoing edge node, the weight array weightArray is used to record the outgoing edge weight, and the attribute array dataArray is used to record the status data; S2, configuring the starting address of the work queue listArray and the number of work queue elements len through software, so that the data prefetcher starts to prefetch the graph data related to all the nodes to be processed in the work queue listArray, the graph data including the source node ID, the destination node ID, the source node attribute, the destination node attribute and the outgoing edge weight, and stores the related graph data in the data cache; S3, obtain the data in the data cache by calling the data fetching interface function fetch_edge() through software; S4, using the pre-fetched data to perform the corresponding graph calculations through software and update the outgoing edge node attributes; S5, calling the buffer empty status judgment interface function empty() by software to judge whether there is any data to be processed in the data cache of the data prefetcher, if there is any data to be processed, continue to execute step S3; otherwise, end and exit.
7. The data prefetcher for graph computing according to claim 6, characterized in that: In step S2, the data prefetcher starts to prefetch the graph data related to all the nodes to be processed in the work queue listArray, including: S2.1, get the starting address of the input work queue listArray and the number of work queue elements len; initialize the column index idx to 0; S2.2, determine whether the column index idx is less than the number of work queue elements len. If the column index idx is less than the number of work queue elements len, the process ends. Otherwise, the process jumps to step S2.
3. S2.3, take out the index number i of the node to be processed corresponding to the column index idx from the work queue listArray; increment the column index idx by 1, and jump to step S2.3 until the numbers of all the nodes to be processed are generated; S2.4, pre-acquire the node attribute srcData of the node to be processed according to the starting address of the attribute array dataArray and the index number i of the node to be processed, take out the offset corresponding to the index number i of the node to be processed from the offset array offsetArray as the starting offset start of the current row, and take the offset corresponding to the index number i+1 of the next node as the ending offset end of the current row; S2.5, assign the current row offset start to the current column index current_idx; S2.6, determine whether the current column index current_idx is less than the current row end offset end, if not, jump to step S2.2; otherwise, jump to the next step; S2.7, determine whether the data cache is full, if it is full, jump to step S2.7 again; otherwise jump to the next step; S2.8, taking out the element corresponding to the current column index current_idx from the edge array edgeArray as the outgoing edge node node corresponding to the node to be processed i, and taking out the element corresponding to the current column index current_idx from the weight array weightArray as the obtained edge weight weight; S2.9, read the attribute data of the outgoing node node corresponding to the node i to be processed from the attribute array dataArray according to the node node; add 1 to the current column index current_idx, and jump to step S2.
6.
8. A processor, comprising a processor body having a plurality of processor cores, characterized in that: A data prefetcher for graph-oriented computing as described in any one of claims 1 to 7 is provided in the processor body.
9. A computer device comprising a processor and a memory connected to each other, characterized in that: The processor is provided with a data prefetcher for graph computing as described in any one of claims 1 to 7.
Citation Information
Patent Citations
A data prefetching method and apparatus for a data structure oriented graphics processor
CN109461113A
Subgraph segmentation optimization method based on inter-core storage access and application
CN114756483A
Data prefetcher for coping with irregular memory access
CN118132464A
System and method for improving index performance through prefetching
US20030126116A1
Pipelined Prefetcher for Parallel Advancement Of Multiple Data Streams
US20170286304A1