A high-energy-efficiency collaborative graph calculation method and device
By converting the core dependency path into a direct dependency relationship on a multi-core processor, using the dependency path prefetching unit and the direct dependency management unit, the problems of low parallelism and dynamic graph processing of the iterative graph algorithm are solved, and efficient graph calculation is achieved.
Patent Information
- Application Number
- CN202210525819.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-12
AI Technical Summary
The prior art has low parallelism of iterative graph algorithms on multi-core processors, and the graph vertex state propagates slowly, making it impossible to efficiently process static and dynamic graphs, resulting in slow convergence speed of iterative graph algorithms and loss of timeliness of processing results.
The dependency path prefetching unit and the direct dependency management unit are adopted. By converting the dependency between the head and tail vertices on the core dependency path into direct dependencies, the multi-core processor is used for efficient parallel processing, and the dependency index is updated in dynamic graph processing to adapt to graph structure changes.
It improves the effective parallelism of multi-core processors, accelerates the propagation of graph vertex states, improves the convergence speed of iterative graph algorithms, and ensures the real-time and accuracy of graph processing results.
Smart Images

Figure CN114817648B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph computing, and particularly to an energy-efficient collaborative graph computing method and device. Background Art
[0002] With the advent of the big data era, graphs, as a data structure that can well represent data correlations, have been widely used in many fields such as Internet applications, data mining, and scientific computing. Many existing important graph applications use iterative graph algorithms to iteratively process graph data until convergence, such as path analysis, product recommendation, social network analysis, etc.
[0003] In iterative graph algorithms, the state update of graph vertices depends on the state values of their adjacent graph vertices. This dependence of the graph structure often results in long dependence chains between graph vertices. The new state of each vertex needs to be propagated along the dependence path for multiple rounds to reach its indirect neighbors. When propagating vertex states between multiple processor cores along the dependence path, high synchronization overhead will be generated. Many vertices are in an inactive state before the new states of their neighbors arrive. In addition, the outdated states of graph vertices may be read by their neighbors, resulting in unnecessary vertex state updates. Therefore, multi-core processors can often only execute effective updates of iterative graph algorithms on graph data with a low degree of parallelism, which seriously affects the graph processing efficiency.
[0004] To provide real-time results for graph applications, many existing research efforts have proposed various software and hardware solutions to accelerate the graph processing speed on multi-core processors. However, due to the neglect of the dependence relationships between vertices, existing solutions still suffer from the problem of insufficient utilization of multi-core processors and cannot efficiently propagate vertex states in the graph topology, resulting in a slow convergence rate of iterative graph algorithms. In addition, graphs in real-world applications are often dynamically changing, such as changes in social relationships and flight information. The rapid changes in the graph topology will cause the graph processing results to quickly become outdated. Therefore, graph applications have higher requirements for the real-time nature of dynamic graph processing results. However, existing solutions often cannot handle graph structure changes well and can only be used to accelerate static graph processing, and are not applicable to dynamic graph processing scenarios.
[0005] For example, Chinese Patent CN109919826A discloses a graph data compression method and a graph computing accelerator for a graph computing accelerator. The method includes: S1. A preprocessing circuit of the graph computing accelerator converts graph data represented by an adjacency sparse matrix to be processed into graph data in a compressed sparse column independent (CSCI) format. Each column of the compressed graph data includes a column identifier data pair and a non-zero element data pair. Each data pair includes an index and a value. The highest two bits of the index indicate the meaning of the remaining bits of the index and the value. S2. The preprocessing circuit of the graph computing accelerator stores the converted graph data in the CSCI format in the memory of the graph computing accelerator. The compression method of the invention can improve the parallelism and energy efficiency of the graph computing accelerator, and also does not consider improving the accelerator from the aspect of graph vertices.
[0006] In addition, on the one hand, there are differences in the understanding of those skilled in the art; on the other hand, when the applicant made this invention, a large number of documents and patents were studied, but due to space limitations, all details and contents were not listed in detail. However, this does not mean that the invention does not have the features of these prior arts. On the contrary, the invention already has all the features of the prior arts, and the applicant reserves the right to add relevant prior arts in the background art. Summary of the Invention
[0007] Since the technical solutions in the prior art ignore the dependency relationships between vertices, there is a problem of insufficient utilization of multi-core processors, and it is impossible to efficiently propagate vertex states in the graph topology, resulting in a slow convergence rate of iterative graph algorithms. In addition, graphs in real-world applications are often dynamically changing, such as changes in social relationships and flight information. The rapid change of the graph topology will cause the graph processing results to quickly lose timeliness. Therefore, graph applications have higher requirements for the real-time nature of dynamic graph processing results, while existing solutions often cannot handle graph structure changes well, so they can only be used to accelerate static graph processing and cannot be applied to dynamic graph processing scenarios.
[0008] To solve the defects of the prior art, the present invention proposes an energy-efficient collaborative graph computing method and device, aiming to solve the problems of low effective parallelism, slow propagation of graph vertex states, and low graph processing efficiency in iterative graph processing, and to efficiently process static graphs and dynamic graphs.
[0009] The present invention provides an energy-efficient collaborative graph computing device, which is characterized by at least including: a dependency path prefetching unit: configured to receive active vertex information and prefetch edges of a graph partition along the dependency path starting from the active vertices in the circular queue; and
[0010] a direct dependency management unit: configured to convert the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency.
[0011] Preferably, the direct dependency management unit is further configured to: after the state of the head vertex of the path is updated, provide direct dependency formula parameters for the processor core, and the processor core calculates the impact on the tail vertex of the path according to the direct dependency formula and updates the state of the tail vertex.
[0012] Preferably, the direct dependency management unit is further configured to: during the dynamic graph processing, obtain invalidated dependency indexes based on changes in the graph structure, and delete the invalidated dependency indexes to update the dependency indexes.
[0013] Preferably, the dependency path prefetching unit prefetches the edges of the graph partition along the dependency path starting from the active vertex in at least the following ways: in the case of accelerator initialization, complete the edge prefetching in the form of a four-stage pipeline, and output the obtained edges and the states of a pair of vertices corresponding to the edges to the FIFO edge buffer for access and processing by the processor core.
[0014] Preferably, the way that the dependency path prefetching unit completes the edge prefetching in the form of a four-stage pipeline at least includes:
[0015] If the stack is empty, obtain an active vertex from the circular queue and push it onto the stack;
[0016] Obtain the start / end offset of the out-edge of the top vertex of the stack from the offset array;
[0017] Obtain the ID of the unvisited neighbor vertex according to the unvisited edges of the top vertex of the stack, and push one of the neighbor vertices onto the stack;
[0018] Obtain the states of the relevant vertices from the vertex state array, and output the states of the edges and a pair of vertices corresponding to the edges to the FIFO edge buffer; if the top vertex of the stack belongs to the vertex set H m , then pop the top vertex of the stack and insert it into the circular queue as a new active vertex; if no unvisited vertex in the graph partition G m can be obtained from the neighbors of the top vertex of the stack, then pop the top vertex of the stack.
[0019] Preferably, the formula for the direct dependency management unit to convert the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency relationship is at least expressed as:
[0020]
[0021] where s j , s i represent the state values of vertices j and i, and μ and ξ represent constant parameters.
[0022] Preferably, the manner in which the direct dependency management unit converts the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency includes at least:
[0023] When the processing of the core dependency path l is completed for the first time, the numbers j and i of its head and tail vertices and the first state value s j , s i are saved to the direct dependency index array, and the index flag flag is set to I; wherein, the core dependency path l is a path whose head and tail vertices both belong to the vertex set H m of the path;
[0024] When the processing of the path l is completed for the second time, the second state values s' j , s' i are obtained, and the second state values s' j , s' i and the first state value s j , s i are substituted into the formula of the direct dependency relationship to calculate the values of the constant parameters μ and ξ.
[0025] The values of the constant parameters μ and ξ are saved to the direct dependency index array, and the index flag flag is set to A.
[0026] Preferably, the direct dependency management unit constructs a mapping relationship between the vertex ID and the direct dependency index address through a memory hash table. When performing dependency relationship conversion, the direct dependency management unit inserts or updates the memory hash table according to the generated direct dependency index. Among them,
[0027] When the vertex ID corresponding to the direct dependency index is not inserted into the memory hash table, the direct dependency management unit inserts the table entry <ID, start_offset, end_offset, weight> into the memory hash table, where the weight weight is set to the hash collision count N + 1;
[0028] When the vertex ID corresponding to the direct dependency index has been inserted into the memory hash table, the direct dependency management unit updates the start offset start_offset, the end offset end_offset, and the weight weight of the table entry, where the weight weight is updated to weight + 1.
[0029] Preferably, the device further includes an on-chip cache unit.
[0030] The on-chip cache unit establishes a data connection relationship with the direct dependency management unit.
[0031] The direct dependency management unit establishes a cache hash table in the on-chip cache unit. Among them,
[0032] The direct dependency management unit caches the frequently accessed entries and the conflicting entries in the memory hash table into the on-chip cache unit according to the customized insertion strategy and / or replacement strategy.
[0033] When the path start vertex is prefetched, the direct dependency management unit retrieves the corresponding dependency index through the vertex ID. The dependency index retrieval process at least includes:
[0034] First, obtain the storage address of the target dependency index from the on-chip cache unit. If the acquisition fails, obtain the storage address of the target dependency index from the memory hash table;
[0035] Obtain the direct dependency index information corresponding to the vertex from the direct dependency index array according to the storage address of the target dependency index.
[0036] Preferably, during the dynamic graph processing, the dependency index update process of the direct dependency management unit at least includes: traversing the graph structure update information to obtain the deleted edge <s, d>;
[0037] Starting from the destination vertex d of the deleted edge, perform a depth-first search traversal in the core subgraph and set the maximum traversal depth,
[0038] Add the core vertices visited during the traversal to the vertex set H d , and after the traversal ends, pass the vertex set H d to the direct dependency management unit for index update;
[0039] The direct dependency management unit traverses to obtain the direct dependency index whose tail vertex number belongs to the vertex set H d . If the start vertex of the dependency index does not belong to the vertex set H d , delete the dependency index. If the start vertex of the dependency index belongs to the vertex set H d , retain the dependency index;
[0040] Synchronously update the core subgraph, where the corresponding edge in the core subgraph is deleted, and the source vertex and destination vertex of the deleted edge are added to the core vertex set H m .
[0041] The present invention also provides an energy-efficient collaborative graph computing method implemented by the energy-efficient collaborative graph computing device of the present invention. The method at least includes:
[0042] Receive the active vertex information, and prefetch the edges of the graph partition along the dependency path starting from the active vertex in the circular queue;
[0043] Convert the dependency relationship between the start vertex and the end vertex on the core dependency path into a direct dependency; and / or
[0044] During the dynamic graph processing, the dependency index is updated according to the dynamic changes of the graph structure. Brief Description of the Drawings
[0045] Figure 1 is a schematic diagram of the hardware architecture of an accelerator according to a preferred embodiment provided by the present invention;
[0046] Figure 2 is a flowchart of a graph calculation method according to a preferred embodiment provided by the present invention;
[0047] Figure 3 is a flowchart of the dependency index update in graph calculation according to a preferred embodiment provided by the present invention;
[0048] Figure 4 is a flowchart of a preprocessing stage according to a preferred embodiment provided by the present invention;
[0049] Figure 5 is a flowchart of a graph calculation stage according to a preferred embodiment provided by the present invention.
[0050] List of Reference Numerals
[0051] 1: Processor core; 2: First cache unit; 3: Dependency path prefetch unit; 4: On-chip cache unit; 5: Direct dependency management unit; 6: Second cache unit; 8: Third cache unit; 9: Graph data. Detailed Embodiment
[0052] The following is a detailed description with reference to the drawings.
[0053] The present invention provides an energy-efficient collaborative graph calculation method and apparatus. The present invention can also provide a system for graph calculation. The present invention can also provide a processor capable of running the graph calculation method of the present invention. The present invention can also provide a storage medium storing the graph calculation running code of the present invention.
[0054] The cache unit in the present invention refers to a memory capable of performing efficient data exchange. The cache unit is, for example, RAM (Random-Access Memory), ROM (Read-Only Memory), memory-mapped registers, and the like.
[0055] As Figure 1 shown, the energy-efficient collaborative graph calculation apparatus of the present invention can also be referred to as a graph calculation accelerator.
[0056] The graph calculation accelerator establishes a communication connection through at least one second cache unit 6 and at least one third cache unit 8 for transmitting data information.
[0057] As shown Figure 1 in the figure, the graph computing accelerator at least includes a dependency path prefetching unit 3 and a direct dependency management unit 5. The dependency path prefetching unit 3 and the direct dependency management unit 5 establish a data transmission relationship.
[0058] For example, the dependency path prefetching unit 3 includes a sub-processor, an application-specific integrated circuit, a server, etc. with the dependency path prefetching function. For example, the sub-processor can run the encoded program of the dependency path prefetching method.
[0059] The direct dependency management unit 5 includes a sub-processor, an application-specific integrated circuit, a server, etc. with the direct dependency management function. For example, the sub-processor can run the encoded program of the direct dependency management method.
[0060] The graph computing accelerator can be a processor, an application-specific integrated circuit, a server, etc. integrated with the dependency path prefetching function, the direct dependency management function, and / or the on-chip cache function. Preferably, the graph computing accelerator is integrated in a multi-core processor.
[0061] Preferably, the graph computing accelerator can also be composed of at least two connected sub-processors. For example, the dependency path prefetching unit 3 and the direct dependency management unit 5 establish a connection for data transmission and data processing. Each accelerator is coupled to a core of the multi-core processor and accesses memory through a secondary cache.
[0062] Preferably, the graph computing accelerator is coupled to the processor core 1, as Figure 1 shown in the figure. The first cache unit 2 respectively establishes a data transmission connection with the processor core 1 and the second cache unit 6. The graph computing accelerator respectively establishes a data transmission connection with the processor core 1 and the second cache unit 6. The second cache unit 6 establishes a communication connection with the third cache unit 8. The first cache unit 2 is arranged in parallel with the graph computing accelerator.
[0063] The dependency path prefetching unit 3 is configured to: prefetch the edges to be processed along the dependency path starting from the active vertices.
[0064] The direct dependency management unit 5 is configured to: convert the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency and perform cache management on it.
[0065] Due to the power-law characteristics of graphs in the real world, a small number of graph vertices connect most of the edges of the graph, and the state propagation between the vast majority of graph vertices needs to be carried out through the core dependency path. The key design of the present invention is to convert the indirect dependency between the head and tail vertices of the core dependency path into a direct dependency, so as to parallelize the asynchronous vertex state propagation on the dependency path and accelerate the convergence of the iterative graph algorithm.
[0066] The direct dependency management unit 5 is also configured to update the dependency index according to the changes in the graph structure during the dynamic graph processing.
[0067] For the high - energy - efficient collaborative graph calculation method and device proposed by the present invention, the main idea of the graph calculation is as follows: In each iteration, the graph calculation accelerator coupled with the processor core 1 prefetches the graph data 9 along the dependency path for the processor core 1 to access and process, so that the graph vertex state can be efficiently propagated along the dependency path. At the same time, the graph calculation accelerator also maintains a set of direct dependency relationships between the head and tail vertices of the core dependency paths, thereby further accelerating the propagation of the graph vertex state and maximizing the effective parallelism of the multi - core processor.
[0068] In the present invention, the core sub - graph is divided into paths with the intersecting vertices being the head and tail vertices of the path, which are called core paths. At the same time, the vertices where two core paths intersect are obtained, which are called core vertices.
[0069] As Figure 2 shown, the graph calculation method of the graph calculation accelerator of the present invention at least includes steps S1 to S3. Preferably, the graph algorithm of the present invention at least satisfies two properties:
[0070] The first property: The graph algorithm can be represented by the Gather - Apply - Scatter (GAS) model.
[0071] The second property: The edge processing function of the graph algorithm is a linear expression, usually represented as multiplication or addition.
[0072] Most iterative graph algorithms satisfy these two properties, such as pagerank, adsorption, SSSP, WCC, k - core, etc. The present invention takes the SSSP algorithm as an example to implement the following calculation process.
[0073] S1: Pre - processing stage.
[0074] Traverse the graph vertices, and take the vertices with degrees greater than the degree threshold T as the central vertices. Secondly, based on the central vertices, traverse the graph data to obtain the central paths, that is, the paths with both the head and tail vertices being central vertices, so as to obtain the core sub - graph composed of the union of the central paths. Then traverse the core sub - graph, divide the core sub - graph into core paths with the intersecting vertices being the head and tail vertices of the path, and at the same time obtain the vertices where two core paths intersect, that is, the core vertices. After the processor core completes the pre - processing of the graph data, by calling the configuration interface of the accelerator, the graph data information is passed to the memory - mapped register accessible by the accelerator to initialize the accelerator. The memory - mapped register here is part of the accelerator.
[0075] S2: Graph calculation stage.
[0076] In each round of graph processing, the dependency path prefetching unit 3 of the accelerator starts from the active vertices in the local circular queue, and dynamically prefetches the edges of the corresponding graph partition for its corresponding processor core through depth-first search. While the dependency path prefetching unit 3 performs edge prefetching, the direct dependency management unit 5 converts the indirect dependencies of the head and tail vertices of the core dependency path into direct dependencies and performs cache management on them. In the SSSP algorithm, the direct dependency relationship between two vertices can be expressed as the formula where s j , s i are the state values of vertices j and i, and μ and ξ are constant parameters. (In the SSSP algorithm, the parameter μ is always 1).
[0077] When processing the core dependency path l for the first time, save the first state values of the first group of vertices at the head and tail of the path (s j , s i ). When processing the path l for the second time, save the second state values of the second group of vertices (s' j , s' i ), and then substitute them into the formula of the direct dependency relationship to calculate the values of the parameters μ and ξ.
[0078] In subsequent processing, after updating the head vertex of the path l, the parameters of the direct dependency formula can be obtained through the direct dependency index. Calculate the impact of the update of the head vertex on the tail vertex according to the direct dependency formula and update the state of the tail vertex, without waiting for the state of the head vertex of the path to be propagated to the tail vertex through multiple rounds of iteration. Thus, multiple paths can be processed concurrently on multiple processor cores, accelerating the propagation of graph vertex states and improving the convergence speed of graph computing.
[0079] In addition, when the accelerator processes a dynamic graph, the direct dependency management unit 5 updates the dependency index according to the dynamic changes of the graph structure to ensure the accuracy of the graph processing results.
[0080] Preferably, the local circular queue is located in the memory. The local circular queue stores the active vertices in the graph partition corresponding to the processor core. Each processor core will be assigned a graph partition for processing, so each processor core will have a corresponding local circular queue in the memory.
[0081] S3: Output stage.
[0082] The processing steps of the preprocessing stage at least include:
[0083] S11: Obtain the central vertex and central path of the given graph; including steps S50 - S55.
[0084] S12: Divide the core subgraph, core path and core vertex; including steps S56 - S60.
[0085] The specific steps in the preprocessing stage are as Figure 4 described.
[0086] S0: Start.
[0087] S50: Traverse the graph vertices.
[0088] S51: Determine whether the degree of the graph vertex is greater than the threshold T. If so, execute step S52; if not, execute step S53.
[0089] S52: Add to the central vertex set and execute step S54.
[0090] Specifically, take the vertices with degrees greater than the degree threshold Τ as central vertices and add them to the central vertex set. Among them, the calculation method of the degree threshold Τ is:
[0091] Calculate the number of central vertices λ·n (n is the number of all vertices) according to the ratio λ of the central vertices specified by the user, and then sort all vertices in descending order according to the vertex degrees, and take the degree of the λ·n-th vertex as the degree threshold Τ.
[0092] Preferably, since the cost of sorting all vertices is too high, a sampling method can also be used to quickly determine the degree threshold Τ. Take the sampling vertices with a proportion of β and sort them in descending order, and take the degree of the λ·β·n-th vertex as the degree threshold Τ.
[0093] S53: Determine whether the traversal of the graph vertices is completed. If so, execute step S54; if not, execute step S50.
[0094] S54: Obtain the unvisited central vertices.
[0095] Specifically, obtain the central vertices from the divided central vertex set H.
[0096] S55: Use depth-first search traversal to obtain the central path.
[0097] Specifically, perform depth-first search traversal with the central vertex as the root vertex. During the traversal, preferentially visit the vertices with higher degrees and specify the traversal depth (the default depth is 16). Let the first vertex of the traversal path l be v root , and the last vertex be v curr . If v curr belongs to the central vertex set H, then the path l is the central path. Mark the vertex v curr as visited and add the path l to the set G s . If all the vertices in the central vertex set H are marked as visited, directly end the traversal. S56: Determine whether all the central vertex sets are visited. If not, execute step S54; if so, execute step S57.
[0098] S57: Construct the core subgraph.
[0099] Specifically, after the current traversal is completed, if the vertices in the central vertex set H have not all been visited, then select the next unvisited vertex in the central vertex set H as the root vertex and continue to perform the above traversal until all the vertices in the central vertex set H have been visited, and finally obtain the set G of all central paths l s which is the core subgraph.
[0100] S58: Obtain the vertices in the core subgraph that have unvisited edges.
[0101] Specifically, starting from the core subgraph G s obtain the vertices with unvisited edges. Among them, the central vertices are preferentially selected.
[0102] S59: Perform a depth-first search traversal to obtain a core path and the corresponding core vertices.
[0103] Starting from the vertex with unvisited edges, obtain a path along the order of depth-first search traversal. The maximum length of the path is defaulted to 16. Mark all the edges on the path as visited, and add the head and tail vertices of the path to the core vertex set H m .
[0104] S60: Determine whether all the edges in the core subgraph have been visited. If not, execute step S58; if so, execute step S100.
[0105] Repeat the above steps until all the edges in the core subgraph have been visited.
[0106] Step S100: End.
[0107] Graph calculation stage: Through the dependency path prefetching unit 3 of the accelerator, start from the active vertices and prefetch the edges along the dependency paths for processing. At the same time, through the direct dependency management unit 5, convert the dependency relationship between the head and tail vertices of the core dependency path into a direct dependency relationship and perform cache management on it. After the state of the path head vertex is updated, quickly calculate the impact on the path tail vertex according to the formula of the direct dependency relationship and update the state of the tail vertex.
[0108] The specific steps of the graph calculation stage of the present invention are as Figure 5 shown.
[0109] S0: Start.
[0110] S61: Initialize the accelerator.
[0111] Specifically, by calling the accelerator configuration interface, the graph data information in graph data 9 is passed to the memory-mapped register accessible by the accelerator to initialize the accelerator. During the accelerator initialization process, the graph data information (such as the starting address of the CSR array, etc.) rather than the graph data itself is passed to the accelerator. The transmission route of the graph data information is: memory - the third cache unit 8 - the second cache unit 6 - the memory-mapped register of the accelerator.
[0112] The graph data information at least includes:
[0113] (a) The sizes and starting addresses of the offset array, edge array, and vertex status array included in the CSR-format graph data;
[0114] (b) The starting vertex ID and ending vertex ID of the graph partition G assigned to the corresponding processor core m ;
[0115] (c) The size and starting address of the core vertex set H in the graph partition G m ; m ;
[0116] (d) The size and starting address of the local circular queue of the corresponding processor core, and this circular queue is used to store the active vertices to be processed in the graph partition G m .
[0117] S62: Obtain the active vertices from the local circular queue.
[0118] S63: Obtain the graph data along the dependency path of the active vertices.
[0119] Specifically, the dependency path prefetch unit 3 dynamically prefetches the edges of the graph partition G for its corresponding processor core 1 through depth-first search. m ;
[0120] The dependency path prefetch unit 3 uses a stack with a fixed depth to record the prefetch information, and each entry in the stack contains the following information:
[0121] (a) The ID of the vertex visited during traversal;
[0122] (b) The current offset and ending offset of the unvisited edges of this vertex;
[0123] (c) The ID of the unvisited neighbor vertices of this vertex.
[0124] Specifically, the dependency path prefetch unit 3 completes the edge prefetch in the form of a four-stage pipeline. Each time, the obtained edges and the states of a pair of vertices corresponding to these edges are output to the FIFO edge buffer for the processor core 1 to access and process.
[0125] S63.1: If the stack is empty, obtain an active vertex from the local circular queue and push it onto the stack.
[0126] S63.2: Obtain the start / end offsets of the outgoing edges of the top vertex of the stack from the offset array.
[0127] S63.3: Obtain the IDs of the unvisited neighbor vertices according to the unvisited edges of this vertex, and push one of the neighbor vertices onto the stack.
[0128] S63.4: Obtain the status of the relevant vertex from the vertex status array, and output the status of this edge and the pair of vertices corresponding to this edge to the FIFO edge buffer. If the top vertex of the stack belongs to the vertex set H m , then pop the top vertex of the stack, and insert it into the local circular queue as a new active vertex, and go to step S63.1. If no unvisited vertex in the graph partition G m can be obtained from the neighbors of the top vertex of the stack, then pop the top vertex of the stack and go to step S63.1.
[0129] S64: Process the graph data.
[0130] For example, process the graph data according to the graph algorithm, and the graph algorithm is, for example, the SSSP algorithm.
[0131] S65: Determine whether there is a direct dependency index. If so, execute step S66; if not, execute step S75.
[0132] Specifically, while the edge prefetching is performed in the dependency path prefetching unit 3, the direct dependency management unit 5 converts the indirect dependencies of the head and tail vertices of the core dependency path into direct dependencies.
[0133] When the dependency relationship between vertices is a linear relationship, the direct dependency relationship between two vertices can be expressed by the formula: where s j , s i are the status values of vertices j and i, and μ and ξ are constant parameters. In the SSSP algorithm, the parameter μ is always 1.
[0134] The direct dependency management unit 5 uses the direct dependency index array to save the direct dependency indexes between the head and tail vertices of the path. As Figure 1 shown, each index in the array includes the head vertex number j, the tail vertex number i, the path identifier l, the parameter μ, the parameter ξ, and the index identifier flag. Among them, the index identifier flag represents the current status of this index, which is divided into three cases in total:
[0135] (a) If the index identifier is N, then this index is invalid;
[0136] (b) If the index identifier is I, the current values of the parameter μ and the parameter ξ are a set of state values s of vertices j and i j 、s i ;
[0137] (c) If the index identifier is A, the index is valid, and the values of the parameter μ and the parameter ξ are the parameter values directly dependent on the formula.
[0138] S66: Determine whether the directly dependent index status is A. If so, execute step S68; if not, execute step S67.
[0139] S67: Determine whether the directly dependent index status is I. If not, execute step S69; if so, execute step S72.
[0140] S68: Utilize the influence of the head vertex of the directly dependent calculation path on the tail vertex, and continue to execute step S75.
[0141] S69: Process the core dependent path.
[0142] S70: Save the state values of the head and tail vertices of the path to the index.
[0143] S71: Set the directly dependent index status to I, and continue to execute step S75.
[0144] S72: Process the core dependent path.
[0145] S73: Calculate the constant parameters of the directly dependent formula.
[0146] S74: Set the directly dependent index status to A.
[0147] S75: Determine whether the prefetch of the current dependent path is completed. If so, execute step S76; if not, execute step S63.
[0148] S76: Determine whether the local circular queue is empty. If so, execute step S77; if not, execute step S62.
[0149] S77: Output the result.
[0150] S100: End.
[0151] The conversion process of the dependency includes steps S69 to S74.
[0152] Specifically, the examples of steps S69 to S71 are as follows:
[0153] The index identifier flag of the directly dependent index is initially set to N. During the graph processing, when the processing of the core dependent path l (the path whose head and tail vertices both belong to the vertex set H m is completed for the first time), the numbers j and i of its head and tail vertices and the first state value sj , s i Save it to the direct dependency index array and set the index flag `flag` to I.
[0154] Specifically, the examples of steps S72 - S74 are as follows: When the processing of path l is completed for the second time, a set of second state values s' of the head and tail vertices is obtained j , s' i , combined with the first state values s stored at the indices μ and ξ j , s i , substitute them into the formula of the direct dependency relationship and solve the equations Calculate the values of the constant parameters μ and ξ, save the values of μ and ξ at the indices μ and ξ, and set the index flag `flag` to A.
[0155] When the dependency path prefetch unit 3 prefetches the head vertex of a path, the direct dependency management unit 5 retrieves the corresponding dependency index through the vertex ID, then obtains the dependency index information and provides it to the processor core. The processor core calculates the influence of the head vertex of the path on the tail vertex based on the dependency index information (parameters μ, parameter ξ) and updates the state of the tail vertex, and then inserts the tail vertex of the path into the local loop queue of the processor core.
[0156] Embodiment 2
[0157] This embodiment is a further improvement of Embodiment 1, and the repeated content will not be elaborated here.
[0158] Preferably, as Figure 1 shown, the graph computing accelerator further includes an on-chip cache unit 4. The on-chip cache functional unit 4 includes a sub-processor with on-chip cache function, an application-specific integrated circuit, a server, etc. For example, the sub-processor can run the encoding program of the on-chip cache method. The on-chip cache unit 4 establishes a data transmission relationship with the direct dependency management unit 5.
[0159] The on-chip cache unit 4 and the direct dependency management unit 5 are connected to perform data transmission and data storage. The on-chip cache unit 4 is used to store the entries of the in-memory hash table. The in-memory hash table is used by the direct dependency management unit 5 to quickly obtain the dependency index. Therefore, the on-chip cache unit 4 has a data transmission relationship with the direct dependency management unit.
[0160] Preferably, when the on-chip cache unit 4 is provided in the graph computing accelerator, to accelerate the retrieval speed of the dependency index, the direct dependency management unit 5 uses the in-memory hash table to quickly obtain the storage address of the target dependency index, and at the same time uses the on-chip cache unit 4 to cache the frequently accessed entries and the conflicting entries in the hash table.
[0161] The details are as follows.
[0162] Build the mapping relationship between vertex IDs and direct dependency index addresses through an in-memory hash table.
[0163] Each entry in the in-memory hash table can be represented as <ID, start_offset, end_offset, weight>, where start_offset and end_offset respectively represent the start offset and end offset of the dependency index corresponding to the vertex ID in the direct dependency index array. Weight represents the weight of this table entry. The weight is set to |M + N|, where M is the number of dependency indexes corresponding to the vertex ID, and N is the number of hash collisions generated when inserting this hash table entry. The number of entries in the in-memory hash table is set to |H| / d, where |H| is the number of core vertices. Preferably, d is set to 0.75. The conflict handling method uses linear probing.
[0164] While the direct dependency management unit 5 performs dependency relationship conversion, insert or update the in-memory hash table according to the generated direct dependency indexes.
[0165] The situations of inserting or updating the in-memory hash table according to the generated direct dependency indexes at least include the following.
[0166] The first situation: When the vertex ID corresponding to the direct dependency index is not inserted into the in-memory hash table, the direct dependency management unit 5 inserts the table entry <ID, start_offset, end_offset, weight> into the in-memory hash table, where the weight weight is set to the number of hash collisions N + 1.
[0167] The second situation: When the vertex ID corresponding to the direct dependency index has been inserted into the in-memory hash table, the direct dependency management unit 5 updates the start offset start_offset, end offset end_offset, and weight weight of the table entry, where the weight weight is updated to weight + 1.
[0168] When the direct dependency management unit 5 performs dependency index retrieval, first obtain the corresponding start offset and end offset (start_offset and end_offset) from the hash table through the vertex ID, and then obtain the direct dependency index information corresponding to the vertex from the direct dependency index array according to the offset.
[0169] The direct dependency management unit 5 caches the frequently accessed table entries and the conflict-generated table entries in the in-memory hash table to the on-chip cache unit 4, thereby further accelerating the retrieval speed of the dependency indexes.
[0170] Specifically, the direct dependency management unit 5 establishes a cache hash table in the on-chip cache unit 4, and caches hash table entries into the cache hash table using customized insertion and replacement policies.
[0171] Insertion policy: When the space in the on-chip cache unit is not full and the accessed hash table entry is not cached, insert the hash table entry into the on-chip cache unit.
[0172] Replacement policy: When the space in the on-chip cache unit is full and the accessed hash table entry is not cached, the hash table entry with the smallest weight in the on-chip cache unit will be replaced out of the cache space.
[0173] When the direct dependency management unit 5 obtains a hash table entry according to the vertex ID, it first obtains it from the on-chip cache unit. If the acquisition fails, it obtains it from the memory hash table, and caches the obtained hash table entry using a customized cache policy.
[0174] Preferably, when obtaining the direct dependency index, the direct dependency management unit 5 caches the direct dependency index into the multi-core processor cache using a customized cache policy. Details are as follows.
[0175] Reusability partitioning of dependency indices: First, arrange all dependency indices in descending order of the degree of the dependency source vertex. A region of LLC size at the beginning of the arrangement is divided into a high-reusability region, a region of LLC size after the high-reusability region is divided into a medium-reusability region, and the remaining region is divided into a low-reusability region. The dependency indices in each region have corresponding levels of reusability;
[0176] Insertion policy: When the accessed dependency index is not cached, insert the index into the cache, and set different cache priorities according to the reusability region where the index is located. Otherwise, do not insert the index. The indices in the high-reusability region are set to high priority, the indices in the medium-reusability region are set to medium priority, and the indices in the low-reusability region and the graph data are set to low priority.
[0177] Hit promotion policy: When a dependency index is hit, upgrade the cache priority of the index. The indices in the high-reusability region are immediately upgraded to the highest priority, and the indices in the medium-reusability region and the low-reusability region are only upgraded one level upward.
[0178] Eviction policy: When the cache space is full, the dependency index or graph data with the lowest cache priority will be preferentially replaced out of the cache, and the dependency indices that have not been hit for a long time will be gradually downgraded.
[0179] Embodiment 3
[0180] This embodiment is a further improvement of Embodiment 1 and Embodiment 2, and the repeated content will not be elaborated.
[0181] The direct dependency management unit 5 is further configured to update the dependency index according to the changes in the graph structure during the dynamic graph processing.
[0182] Obtain the invalidated dependency index based on the changes in the graph structure, and delete it through the direct dependency management unit.
[0183] Specifically, it includes the following steps, as Figure 3 shown.
[0184] S41: Traverse the graph structure update information to obtain the deleted edge <s, d>.
[0185] S42: Determine whether the deleted edge belongs to the core subgraph. If so, execute step S43; if not, execute step S48.
[0186] S43: If the deleted edge belongs to the core subgraph, perform a depth-first search traversal in the core subgraph starting from the destination vertex d of the deleted edge, and set the maximum traversal depth (the same as the traversal depth in the graph data preprocessing stage). Add the core vertices visited during the traversal to the vertex set H d . If the destination vertex d belongs to the core vertices, also add it to this vertex set. After the traversal, transfer the vertex set H d to the direct dependency management unit for index update.
[0187] S44: The direct dependency management unit 5 traverses to obtain the direct dependency index whose tail vertex number belongs to the vertex set H d .
[0188] S45: Determine whether the head vertex of the dependency index belongs to the vertex set H d . If so, execute step S46; if not, execute step S47.
[0189] S46: If the head vertex of this dependency index belongs to the vertex set H d , then retain this dependency index.
[0190] S47: If the head vertex of this dependency index does not belong to the vertex set H d , then delete this dependency index.
[0191] S48: Synchronously update the core subgraph, that is, delete the corresponding edge in the core subgraph, and add the source vertex and destination vertex of the deleted edge to the core vertex set H m . Determine whether the traversal of the graph structure update information is completed. If so, execute step S100; if not, return to step S41.
[0192] S100: If the traversal of the graph structure update information ends, then the current dependency index update phase is completed, and end.
[0193] It should be noted that the above specific embodiments are exemplary. Those skilled in the art can come up with various solutions inspired by the disclosed content of the present invention, and these solutions also fall within the scope of the disclosure of the present invention and within the protection scope of the present invention. Those skilled in the art should understand that the description and drawings of the present invention are illustrative and do not constitute a limitation on the claims. The protection scope of the present invention is defined by the claims and their equivalents. The description of the present invention contains multiple inventive concepts. Expressions such as "preferably", "according to a preferred embodiment", or "optionally" all indicate that the corresponding paragraphs disclose an independent concept. The applicant reserves the right to file divisional applications based on each inventive concept.
Claims
1. An energy-efficient collaborative graph computing device, characterized in that, Also known as a graph computing accelerator, the graph computing accelerator is a processor, an application-specific integrated circuit, and a server integrated with a dependency path prefetching function, a direct dependency management function, and / or an on-chip caching function; the graph computing accelerator is coupled to a processor core (1); the graph computing accelerator includes at least a dependency path prefetching unit (3) and a direct dependency management unit (5); The graph computing accelerator prefetches graph data (9) along a dependency path for access and processing by the processor core (1), and the processor core (1) obtains the central vertex and the central path of the given graph data, and divides the core subgraph, the core path, and the core vertex; After the processor core (1) completes the preprocessing of the graph data, by calling the configuration interface of the graph computing accelerator, the processor core (1) transfers the graph data information to a memory-mapped register accessible by the graph computing accelerator to initialize the graph computing accelerator; The dependency path prefetching unit (3) is configured to receive active vertex information and prefetch the edges of a graph partition along the dependency path starting from the active vertices in the local circular queue; The direct dependency management unit (5) is configured to convert the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency; wherein, the direct dependency management unit (5) uses a direct dependency index array to save the direct dependency index between the head and tail vertices of the path; During the dynamic graph processing, the direct dependency management unit (5) updates the dependency index according to the dynamic changes of the graph structure.
2. The high-energy efficiency collaborative graph computing device according to claim 1, wherein The direct dependency management unit (5) is further configured to: During the dynamic graph processing, obtain the invalidated dependency index based on the changes of the graph structure, and delete the invalidated dependency index to update the dependency index.
3. The high-energy efficiency collaborative graph computing device according to claim 2, wherein The manner in which the dependency path prefetching unit (3) prefetches the edges of a graph partition along the dependency path starting from the active vertices includes at least: In the case of initializing the graph computing accelerator, the prefetching of edges is completed in the form of a four-stage pipeline, and the obtained edges and the states of a pair of vertices corresponding to the edges are output to the edge buffer for the processor core to access and process.
4. The high energy efficiency collaborative graph computing device according to claim 3, wherein The manner in which the dependency path prefetching unit (3) completes the edge prefetching in the form of a four-stage pipeline includes at least: If the stack is empty, obtain an active vertex from the local circular queue and push it onto the stack; Obtain the start / end offset of the outgoing edges of the top vertex of the stack from the offset array; Obtain the unvisited neighbor vertices based on the unvisited edges of the top vertex of the stack , and push one of the neighbor vertices onto the stack; Obtain the status of relevant vertices from the vertex status array, and output the status of the edges and the pair of vertices corresponding to the edges to the edge buffer; if the vertex at the top of the stack belongs to the vertex set , then pop the vertex at the top of the stack and insert it into the local circular queue as a new active vertex; if an unvisited vertex in the graph partition cannot be obtained from the neighbors of the vertex at the top of the stack , then pop the vertex at the top of the stack.
5. The high-energy efficiency collaborative graph computing device according to claim 2, wherein The formula for the direct dependency management unit (5) to convert the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency relationship is at least expressed as: , Among them , represent the vertex , The state values of , represent constant parameters 6. The high energy efficiency collaborative graph computing device according to claim 2, characterized in that The manner in which the direct dependency management unit (5) converts the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency includes at least: When the processing of the core dependency path is completed for the first time, the numbers of its start and end vertices , as well as the first state value , are saved to the direct dependency index array, and the index identifier is set to ; wherein, the core dependency path is a path whose start and end vertices both belong to the vertex set . When the processing of the path is completed for the second time, the second state values of the start and end vertices are obtained . The second state value and the first state value are substituted into the formula of the direct dependency relationship to calculate the constant parameters , values. Save the values of the constant parameter and to the direct dependency index array, and set the index identifier to .
7. The high-energy efficiency collaborative graph computing device according to any one of claims 1 to 4, characterized in that, The device further includes an on-chip caching unit (4), The on-chip caching unit (4) establishes a data connection relationship with the direct dependency management unit (5), The direct dependency management unit (5) constructs a mapping relationship between the vertex ID and the direct dependency index address through a memory hash table, and establishes a cache hash table in the on-chip caching unit (4), wherein, The direct dependency management unit (5) caches the frequently accessed table entries and the conflicting table entries in the memory hash table into the on-chip caching unit (4) according to a customized insertion strategy and / or replacement strategy.
8. The high-energy efficiency collaborative graph calculation device according to claim 4, characterized in that When the path start vertex is prefetched, the direct dependency management unit (5) retrieves the corresponding dependency index through the vertex and the on-chip cache unit (4) is used to store the entries of the memory hash table. The dependency index retrieval process includes at least: Obtain the storage address of the target dependency index from the on-chip caching unit (4), and if the acquisition fails, obtain the storage address of the target dependency index from the memory hash table; Obtain the direct dependency index information corresponding to the vertex from the direct dependency index array according to the storage address of the target dependency index.
9. The high energy efficiency collaborative graph computing device according to claim 4, characterized in that, During the dynamic graph processing, the dependency index update process of the direct dependency management unit (5) at least includes: Traverse the graph structure to update information and obtain the deleted edges ; Starting from the destination vertex of the deleted edge perform a depth-first search traversal in the core subgraph and set the maximum traversal depth Add the core vertices visited during traversal to the vertex set , and after the traversal ends, transfer the vertex set to the direct dependency management unit (5) for index update; The direct dependency management unit (5) traverses and obtains the direct dependency index whose tail vertex number belongs to the vertex set If the head vertex of the dependency index does not belong to the vertex set , then the dependency index is deleted. If the head vertex of the dependency index belongs to the vertex set , then the dependency index is retained; Synchronously update the core subgraph, where the corresponding edges in the core subgraph are deleted, and the source vertices and destination vertices of the deleted edges are added to the core vertex set 。 10. An energy-efficient collaborative graph calculation method implemented by the energy-efficient collaborative graph calculation device according to any one of claims 1 to 9, characterized in that, The method at least includes: The graph computing accelerator prefetches graph data (9) along the dependency path for the processor core (1) to access and process. The processor core (1) obtains the central vertex and the central path of the given graph data, and divides the core subgraph, the core path, and the core vertex. After the processor core (1) completes the preprocessing of the graph data, by calling the configuration interface of the graph computing accelerator, the processor core (1) transfers the graph data information to the memory-mapped register accessible by the graph computing accelerator to initialize the graph computing accelerator. The dependency path prefetching unit (3) receives the active vertex information; and prefetches the edges of the graph partition along the dependency path starting from the active vertices in the local circular queue. The direct dependency management unit (5) converts the dependency relationship between the head and tail vertices on the core dependency path into a direct dependency. Among them, the direct dependency management unit (5) uses the direct dependency index array to save the direct dependency index between the head and tail vertices of the path. During the dynamic graph processing, the dependency index is updated according to the dynamic changes of the graph structure.
Citation Information
Patent Citations
A graph data compression method for a graph calculation accelerator and the graph calculation accelerator
CN109919826A
High-level comprehensive method and system for graph calculation
CN110750265A
Database partition pruning using dependency graph
US20200311067A1