Dynamic maximal clique enumeration device and method based on FPGA having hbm
By designing a dynamic huge cluster enumeration device with HBM on the FPGA, and using the collaborative design of matrix, sorting and update computing units, the problem of inefficient incremental huge cluster calculation in the prior art is solved, and efficient processing and pipelined computing of large-scale dynamic graph data are realized.
Patent Information
- Application Number
- PCT/CN2023/134667
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-07
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-15
AI Technical Summary
The existing incremental enumeration method cannot support pipelined incremental enumeration calculations, and in large-scale dynamic graph data processing, the existing FPGA graph computing design lacks collaborative algorithm design for incremental enumeration problems, resulting in inefficient computing.
A dynamic maximum cluster enumeration device based on FPGA with HBM is designed, including a matrix calculation unit, a sorting calculation unit and an update calculation unit. Through the coordinated use of FIFO and BRAM, rapid processing of dynamic edge flow and parallel update of candidate clusters are realized, and pipelined incremental maximum cluster calculation is supported.
It realizes fast response and efficient calculation of large-scale dynamic graph data, reduces the calculation delay of incremental huge cluster enumeration tasks, and improves the overall computing speed through parallel computing.
Smart Images

Figure CN2023134667_15052025_PF_FP_ABST
Abstract
Description
Dynamic maximum clique enumeration device and method based on FPGA with HBM Technical Field
[0001] The present invention belongs to the technical field of graph computing for big data processing, and in particular relates to a dynamic maximal clique enumeration device and method based on FPGA with HBM. Background Art
[0002] With the advent of the big data era, graph data has become an important data model for describing and mining connected data. Nodes represent data units, and edges represent relationships between data units. Maximal clique enumeration is a graph computing problem with broad applications and significant significance. For example, enumerating maximal cliques in social network graphs can reveal closely related groups or groups with similar interests. Enumerating maximal cliques in protein interaction networks can also help discover protein complexes and functional modules.
[0003] In the real world, most large graphs change dynamically over time, randomly generating new edges or disconnecting existing ones. These changes are typically small relative to the overall size of the graph. Therefore, many applications require incremental computation to update the corresponding maximal clique results in real time based on graph changes, rather than recalculating the entire maximal clique based on the changed graph structure.
[0004] The problem of incremental maximal clique enumeration based on dynamic graphs requires the algorithm to have high real-time performance and be able to quickly respond to pipeline computing tasks caused by random data changes. The existing incremental maximal clique enumeration method "Honour thy neighbor—clique maintenance in dynamic graphs" can handle the addition / deletion of a group of edges at the same time. The algorithm proposed in the document "Incremental maintenance of maximal cliques in a dynamic graph" can accurately calculate the changes in maximal cliques by enumerating newly added maximal cliques and disappeared maximal cliques. When the addition and deletion of edges are randomly mixed, the algorithm adopts a pseudo-mixing method to process the added edges and deleted edges separately and then summarize the final structure. Patent document CN114357264A discloses a dynamic maximal clique enumeration method based on SOMEi data structure fallback reconstruction, which handles the addition and deletion of edges simultaneously within a unified framework, realizing the enumeration and update of maximal cliques with real edge mixing changes. None of the above existing methods can support pipeline incremental maximal clique calculations.
[0005] With the development of new hardware such as FPGAs, incremental computing on large graphs faces the challenge of further speeding up and increasing efficiency. This requires algorithms with good parallelism to fully leverage the performance advantages brought by increased hardware resources. Existing FPGA-based graph computing research primarily focuses on general-purpose computing frameworks that adapt to a variety of classic graph operators. Currently, there is no work on the customized design of FPGA-based hardware-software collaborative algorithms for the incremental maximal clique enumeration problem.
[0006] HBM (High Bandwidth Memory) is a new type of CPU / GPU memory chip. By stacking multiple DDR chips together and packaging them together with the GPU, a large-capacity, high-bit-width DDR combination array is achieved. HBM is integrated with FPGA to form an FPGA with HBM.
[0007] Summary of the Invention
[0008] In view of the above, the purpose of the present invention is to provide a dynamic maximum clique enumeration device and method based on FPGA with HBM, which supports pipelined incremental maximum clique calculation and improves the overall computing efficiency of the task.
[0009] To achieve the above-mentioned purpose, the embodiment provides a dynamic maximal clique enumeration device based on an FPGA with HBM, comprising an HBM, a matrix calculation unit, a sorting calculation unit, and an update calculation unit constructed by functional isolation and algorithm functional division of the FPGA's built-in FIFO;
[0010] The HBM is used to store the dynamic edge flow, full-graph adjacency matrix, and candidate clusters transmitted from the external PC host for updating the graph structure;
[0011] The matrix calculation unit is used to update the full-graph adjacency matrix based on the dynamic edge flow and send the updated full-graph adjacency matrix to the HBM storage, and at the same time determine the head node to be updated of the candidate cluster that needs to be updated;
[0012] The sorting calculation unit is used to construct a sorted set of candidate cluster reconstructions by sorting data blocks according to the updated full-graph adjacency matrix and each head node to be updated;
[0013] The update calculation unit is used to execute the update task of the candidate cluster corresponding to each head node to be updated in parallel based on the sorted set reconstructed by the candidate cluster, and send the updated candidate cluster to the HBM storage, and send the updated candidate cluster to the PC host to perform filtering operations to extract the maximum cluster.
[0014] Preferably, the matrix calculation unit includes a first FIFO, and updates the full-graph adjacency matrix based on the dynamic edge flow, including:
[0015] The dynamic edge flow obtained from HBM is cached in the first FIFO, and the node to be updated and the neighbor node set of the node to be updated are determined based on the dynamic edge flow. The old adjacency list of all nodes to be updated is obtained from the full-graph adjacency matrix of HBM, and the old adjacency list is updated based on the neighbor node set. The updated adjacency list is written to HBM storage to implement the update of the full-graph adjacency matrix, where the neighbor node set includes the small neighbor node set and the large neighbor node set of the node to be updated.
[0016] Preferably, in the matrix calculation unit, determining the head node to be updated of the candidate cluster that needs to be updated includes:
[0017] The updated head node of the candidate cluster generated by the current batch of dynamic edge flows that needs to be updated is determined according to the small neighbor node set of the node to be updated, and the node corresponding to the position where the candidate cluster of each head node to be updated needs to be rolled back is recorded using the index.
[0018] Preferably, in the sorting calculation unit, constructing a sorted set for candidate clique reconstruction by data block sorting based on the updated full-graph adjacency matrix and each head node to be updated includes:
[0019] Obtain the updated adjacency list corresponding to each head node to be updated in the index record from the updated full-graph adjacency matrix in the HBM, and calculate the large neighbor node set required for updating the candidate clique of each head node to be updated based on the index record information and the updated adjacency list;
[0020] Obtain the small neighbor node set corresponding to each large neighbor node from the HBM, intersect each small neighbor node set with the large neighbor node set, and determine the common neighbor node set of each large neighbor node;
[0021] A second FIFO is constructed for each head node to be updated, and the large neighboring nodes of each head node to be updated and their corresponding public neighboring node sets are sorted according to the node sequence number and stored in the second FIFO to form a sorted set of candidate cluster reconstruction.
[0022] Preferably, in the update calculation unit, based on the sorted set reconstructed from the candidate clusters, the update task of the candidate clusters corresponding to the head nodes to be updated is executed in parallel, including:
[0023] For each head node to be updated in the index record, a subtask and three FIFO queues and BRAM blocks corresponding to the subtask are established. One FIFO queue is used to store the large neighbor nodes and their common neighbor node sets obtained from the sorted set. The other two FIFO queues alternately serve as temporary queues and active queues. The temporary queue stores the candidate team column corresponding to the node to be updated, and the active queue stores the candidate team column being updated. The BRAM block includes a node set record block and three length record blocks. The node set record block is used to store the intersection of the current common neighbor node set and the current candidate group. The three length record blocks respectively record the length of the current common neighbor node set, the length of the current candidate group, and the length of the intersection.
[0024] Based on the node sorting corresponding to each head node to be updated obtained from the sorted set, all subtasks corresponding to the head nodes to be updated use three FIFO queues and BRAM blocks to execute the candidate clique update task in parallel.
[0025] Preferably, each subtask performs a candidate clique update process, including:
[0026] Obtain the old candidate cluster corresponding to each head node to be updated from the HBM and store it in the active queue. Obtain the large neighbor nodes and their common neighbor node sets corresponding to each head node to be updated from the sorted set and store them in the first FIFO queue.
[0027] Sequentially access each large neighbor node and its corresponding public neighbor node set in the FIFO queue, and use each large neighbor node and its corresponding public neighbor node set to update the old candidate clique in the active queue. The updated subsequent maximal clique is transferred to the temporary queue, and the temporary queue and active queue marks are exchanged for the next round of update operations.
[0028] Preferably, the updating operation process includes:
[0029] If the candidate team list in the active queue is empty, or the public neighbor node set is empty, the large neighbor node corresponding to the current public neighbor node set is directly added to the temporary queue as a candidate team, ending the update operation of this round;
[0030] If the candidate cluster maximal queue in the active queue is not empty, each candidate cluster is traversed in turn, and a single comparison operation between the common neighbor node set and the candidate cluster is performed to update the candidate cluster until the candidate cluster list traversal is completed or the traversal stop condition is triggered. During the traversal process, two sets FSet and PSet are maintained to temporarily store candidate clusters that need further screening.
[0031] Preferably, a single comparison operation includes:
[0032] Store the common neighbor node set in BRAM, use the node as the BRAM address, count by judging whether the value of BRAM is 1, and store the length count of the common neighbor node set in the first length record block cnt A , take the current candidate group as the read address of BRAM, count by judging whether the value of BRAM is 1, and store its length count in the second length record block cnt B In the process, the intersection of the common neighbor node set and the candidate cluster is stored in the node set record block IList, and the length of the node set record block is stored in the third length record block cnt C In the example, the values in the three length record blocks are compared and the following four operations are performed:
[0033] When cnt A =cnt B =cnt C , that is, if the common neighbor node set and the candidate cluster set are equal, then the large neighbor node is added to the candidate cluster set and written into the temporary queue, and the other candidate clusters in the active queue are read out and written into the temporary queue, triggering the traversal stop condition;
[0034] When cnt C =cnt A <cnt B , that is, the public neighbor node set is a true subset of the candidate cluster set, then the candidate cluster set is written into the temporary queue, and after adding the large neighbor node to the public neighbor node set, it is written into the temporary queue, and the other candidate clusters in the active queue are read out and written into the temporary queue, triggering the traversal stop condition;
[0035] When cnt C =cnt B <cnt A , that is, the candidate cluster set is a true subset of the common neighbor node set, then after adding the large neighbor node to the candidate cluster set, it is written into the temporary queue and the IList is stored in the set FSet;
[0036] When cnt B ≠0 and cnt B <cnt C and cnt B <cnt A , that is, the candidate cluster set and the common neighbor node set do not contain each other, and the intersection is not empty, then the candidate cluster set is written into the temporary queue and the IList is stored in the set PSet.
[0037] Preferably, the update calculation unit further includes:
[0038] When the candidate group list traversal is completed and the traversal stop condition is not triggered, the candidate groups in the sets FSet and PSet are screened. The specific operations are as follows: traverse each candidate group in PSet. If the traversed candidate group is not a true subset of any other candidate group in PSet, nor a true subset of any candidate group in FSet, the traversed candidate group is written into the temporary queue after joining the large neighbor node.
[0039] To achieve the above-mentioned object of the invention, an embodiment further provides a dynamic maximum clique enumeration method based on an FPGA with HBM, the method adopting the dynamic maximum clique enumeration device of the claim, and the method comprising the following steps:
[0040] Obtain the dynamic edge stream from the external PC host for updating the graph structure and store it in the HBM. The HBM also stores the full graph adjacency matrix and candidate clusters.
[0041] Use the matrix calculation unit to update the full-graph adjacency matrix based on the dynamic edge flow and send the updated full-graph adjacency matrix to HBM storage, while determining the head node to be updated for the candidate cluster;
[0042] The sorting calculation unit is used to construct a sorted set of candidate cluster reconstructions based on the updated full-graph adjacency matrix and each head node to be updated by data block sorting;
[0043] The update calculation unit is used to execute the update task of the candidate cluster corresponding to each head node to be updated in parallel based on the sorted set reconstructed by the candidate cluster, and the updated candidate cluster is sent to the HBM storage;
[0044] Send the updated candidate cluster result data to the PC host, and the PC host extracts the maximum cluster through the candidate cluster filtering operation.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] To address the problem of incremental maximum clique enumeration of large-scale dynamic graph data, we designed a serial matrix calculation unit, sorting calculation unit, and update calculation unit using the FPGA pipeline parallel structure computing model. This unit can quickly respond to batches of dynamic edge streams sent by the user at different times. Multi-batch pipeline processing significantly reduces the computational latency of the incremental maximum clique enumeration task.
[0047] Furthermore, in the sorting calculation unit and the update calculation unit, the update tasks are divided and isolated according to the head node to be updated, supporting parallel calculation of multiple subtasks, with extremely high concurrency, which improves the overall calculation speed of the task. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] FIG1 is a schematic structural diagram of a dynamic maximum clique enumeration device based on an FPGA with HBM according to an embodiment;
[0050] FIG2 is a flow chart of a dynamic maximum clique enumeration method based on an FPGA with HBM according to an embodiment;
[0051] Figure 3 is an example of a dynamic maximal clique enumeration parallel computing task, where the data structure of the dynamic graph G changes from time t0 to time t1.
[0052] FIG4 is an example diagram showing changes in the adjacency list of the dynamic graph in the example.
[0053] FIG5 is an example diagram of the changes in the candidate cluster in the example.
[0054] FIG6 is an example diagram of the head node range and rollback position corresponding to the candidate cluster to be updated in the example.
[0055] FIG7 is a diagram showing an example of the data structure of the parallel computing task in the example.
[0056] FIG8 shows the calculation process of a subtask in an example of parallel computing. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0058] As shown in Figure 1, the dynamic maximum clique enumeration device based on FPGA with HBM provided in the embodiment includes HBM, a matrix calculation unit, a sorting calculation unit and an update calculation unit, wherein the matrix calculation unit, the sorting calculation unit and the update calculation unit are determined by functional partitioning in combination with the inherent hardware functions of FPGA, such as FIFO, and the designed algorithm for implementing dynamic maximum clique enumeration. In each unit, the function of FIFO is isolated. These three units cooperate with HBM to support pipelined incremental maximum clique calculation, thereby improving the overall computing efficiency of the task.
[0059] Among them, HBM is used to store the dynamic edge flow, full-graph adjacency matrix, and candidate clusters transmitted from the external PC host for updating the graph structure; the matrix calculation unit is used to update the full-graph adjacency matrix of the global graph data based on the dynamic edge flow and send the updated full-graph adjacency matrix to the HBM storage, and at the same time determine the head node to be updated for the candidate cluster; the sorting calculation unit is used to construct a sorted set for reconstructing the candidate cluster by data block sorting based on the updated full-graph adjacency matrix and each head node to be updated; the update calculation unit is used to execute the update task of the candidate cluster corresponding to each head node to be updated in parallel based on the sorted set reconstructed by the candidate cluster, and send the updated candidate cluster to the HBM storage.
[0060] Based on the above-mentioned dynamic maximum clique enumeration device, the embodiment further provides a dynamic maximum clique enumeration method, as shown in FIG2 , comprising the following steps:
[0061] Step 1: Obtain the dynamic edge stream for updating the graph structure from the external PC host and store it in the HBM. The HBM also stores the full graph adjacency matrix and candidate clusters.
[0062] Step 2: Use the matrix calculation unit to update the full-graph adjacency matrix based on the dynamic edge flow and send the updated full-graph adjacency matrix to HBM storage. At the same time, determine the head node to be updated for the candidate cluster;
[0063] Step 3: Use the sorting calculation unit to construct a sorted set of candidate cluster reconstructions based on the updated full-graph adjacency matrix and each head node to be updated by sorting the data blocks;
[0064] Step 4: Use the update calculation unit to execute the update task of the candidate cluster corresponding to each head node to be updated in parallel based on the sorted set reconstructed by the candidate cluster, and send the updated candidate cluster to HBM storage;
[0065] Step 5: Send the updated candidate cluster result data to the PC host, and the PC host extracts the maximum cluster through the candidate cluster filtering operation.
[0066] The following is a detailed description of each unit in the dynamic maximal clique enumeration device and the dynamic maximal clique enumeration method.
[0067] During the dynamic maximum clique enumeration process, a batch of dynamic edge streams are received from the PC host through the FPGA's PCLe main line and stored in the HBM. The number of dynamic edges is N.
[0068] In the embodiment, the matrix calculation unit updates the full graph adjacency matrix in the HBM according to the received dynamic edge flow. Specifically, the matrix calculation unit includes a first FIFO, in which the dynamic edge flow obtained from the HBM is cached, and the node to be updated and the neighboring node set of the node to be updated are determined according to the dynamic edge flow. Each dynamic edge is e + / - (vi , v j ), where v i and v j are both nodes to be updated, and e + (v i , v j ) means adding an edge between v i and v j , and e - (v i , v j ) means removing the edge between v i and v j . The set of adjacent nodes of the nodes to be updated includes the set of small adjacent nodes and the set of large adjacent nodes of the nodes to be updated. Assuming i < j, record v j into the set of large adjacent nodes to be updated of node v i , and record v i into the set of small adjacent nodes to be updated of node v j . Obtain the old adjacent list of all nodes to be updated from the full - graph adjacency matrix of HBM, and update the old adjacent list according to the set of adjacent nodes. The updated adjacent list is written into HBM storage to implement the update of the full - graph adjacency matrix.
[0069] When updating the graph adjacency matrix, the matrix calculation unit also needs to determine the head nodes to be updated of the candidate cliques that need to be updated. The process is as follows: Determine the head nodes to be updated of the candidate cliques that need to be updated generated by the current dynamic edge flow according to the set of small adjacent nodes of the nodes to be updated, and use the index H1 to record the node v<00000y ,...,v max The nodes in} are arranged in order according to the size of the subscript numbers.
[0072] Get each large neighbor node v from HBM y ,...,v max Corresponding small neighbor node sets, take the intersection of each small neighbor node set with the large neighbor node set, and determine each large neighbor node v y ,...,v max and the head node v to be updated x The set of common neighbor nodes of
[0073] For each head node v to be updated x Build a second FIFO and put each large neighbor node v of the head node to be updated y ,...,v max The corresponding common neighbor node sets are sorted according to the node sequence numbers and stored in the second FIFO to form a sorted set of candidate cluster reconstruction.
[0074] In the embodiment, the update calculation unit executes the update task of the candidate cluster corresponding to each head node to be updated in parallel based on the sorted set reconstructed by the candidate cluster. The specific process includes:
[0075] For each head node to be updated in the index record, a subtask and three FIFO queues and BRAM blocks corresponding to the subtask are established, where one FIFO queue FIFO A Used to store the large neighbor nodes and their common neighbor node sets obtained from the sorted set, and the other two FIFO queues FIFO B and FIFO C The temporary queue and the active queue are alternately used. The temporary queue stores the candidate team column corresponding to the node to be updated, and the active queue stores the candidate team column being updated. If FIFO B The updated candidate team column is stored in FIFO B For active queues, FIFO C for temporary queues and vice versa.
[0076] The BRAM block includes a node set record block IList and three length record blocks cnt A 、cnt B and cnt C , where the node set record block IList is used to store the intersection of the current common neighbor node set and the current candidate cluster, cnt A 、cnt B and cnt C Record the length of the current common neighbor node set, the length of the current candidate cluster, and the length of the intersection respectively;
[0077] Based on the node sorting corresponding to each head node to be updated obtained from the sorted set, all subtasks corresponding to the head nodes to be updated use three FIFO queues and BRAM blocks to execute the candidate clique update task in parallel.
[0078] Specifically, each subtask performs the candidate cluster update process, including:
[0079] Obtain the old candidate cluster corresponding to each head node to be updated from the HEM and store it in the active queue. Obtain the large neighbor nodes and their common neighbor node sets corresponding to each head node to be updated from the sorted set and store them in the first FIFO queue. A middle;
[0080] Sequential access FIFO queue FIFO A Each large neighbor node and its corresponding public neighbor node set in the active queue are updated using each large neighbor node u and its corresponding public neighbor node set. The updated subsequent maximum clique is transferred to the temporary queue, and the temporary queue and active queue marks are exchanged for the next round of update operations.
[0081] More specifically, the update operation process during the update process includes:
[0082] If the candidate team list in the active queue is empty, or the public neighbor node set is empty, then the large neighbor node u corresponding to the current public neighbor node set is directly added to the temporary queue as a candidate team, ending the update operation of this round;
[0083] If the candidate cluster maximal queue in the active queue is not empty, each candidate cluster is traversed in turn, and a single comparison operation between the common neighbor node set and the candidate cluster is performed to update the candidate cluster until the candidate cluster list traversal is completed or the traversal stop condition is triggered. During the traversal process, two sets FSet and PSet are maintained to temporarily store candidate clusters that need further screening.
[0084] More specifically, a single comparison operation during an update operation includes:
[0085] Store the common neighbor node set in BRAM, use the node as the BRAM address, count by judging whether the value of BRAM is 1, and store the length count of the common neighbor node set in the first length record block cnt A , take the current candidate group as the read address of BRAM, count by judging whether the value of BRAM is 1, and store its length count in the second length record block cnt B In the process, the intersection of the common neighbor node set and the candidate cluster is stored in the node set record block IList, and the length of the node set record block is stored in the third length record block cnt C In the comparison, cnt A、cnt B 、cnt C The value in , and operate according to the following 4 cases:
[0086] When cnt A =cnt B =cnt C , that is, if the common neighbor node set and the candidate cluster set are equal, then the large neighbor node u is added to the candidate cluster set and written into the temporary queue, and the other candidate clusters in the active queue are read out and written into the temporary queue, triggering the traversal stop condition;
[0087] When cnt C =cnt A <cnt B , that is, the public neighbor node set is a true subset of the candidate cluster set, then the candidate cluster set is written into the temporary queue, and after adding the large neighbor node u to the public neighbor node set, the candidate cluster is written into the temporary queue. The other candidate clusters in the active queue are read out and written into the temporary queue, triggering the traversal stop condition;
[0088] When cnt C =cnt B <cnt A , that is, the candidate cluster set is a true subset of the common neighbor node set, then add the large neighbor node u to the candidate cluster set and write it into the temporary queue, and store IList in the set FSet;
[0089] When cnt B ≠0 and cnt B <cnt C and cnt B <cnt A , that is, the candidate cluster set and the common neighbor node set do not contain each other, and the intersection is not empty, then the candidate cluster set is written into the temporary queue and the IList is stored in the set PSet.
[0090] During the update operation, when the candidate group list traversal is completed and the traversal stop condition is not triggered, the candidate groups in the sets FSet and PSet are screened. The specific operations are as follows: traverse each candidate group in PSet. If the traversed candidate group is not a true subset of any other candidate group in PSet, nor a true subset of any candidate group in FSet, the traversed candidate group is written into the temporary queue after joining the large neighbor node.
[0091] The following takes the dynamic maximal clique enumeration calculation task given in Figure 3 as an example to specifically illustrate the calculation process of the present invention. Figure 3 shows a dynamically changing undirected graph G = (V, E) whose changes from time t0 to time t1 are: an edge is added between v3 and v5, and an edge is reduced between v2 and v3, that is, the dynamic edge flow received from G0 to G1 is {e +(v3,v5),e - (v2, v3)}. Figures 4 and 5 show the changes in the adjacency list and candidate clusters from G0 to G1, respectively. The specific process includes:
[0092] 1. Receive 2 dynamic edge streams from the PC host {e + (v3,v5),e - (v2, v3)} and stored in HBM.
[0093] 2. The matrix calculation unit takes out the old adjacency lists corresponding to v2, v3, and v5 from HBM, that is, the old large / small neighbor node sets, and updates the large / small neighbor node sets after calculation. and Write back to HBM. As shown in (1) in Figure 6, the head node to be updated is determined to be {v2} from the changes in the small neighbor nodes of v3, and the head node to be updated is determined to be {v1, v2, v3} from the changes in the small neighbor nodes of v5. After merging, the head node to be updated is {v1, v2, v3}. Correspondingly, the rollback position of the candidate cluster of head node v1 is v5, the rollback position of the candidate cluster of head node v2 is v3, and the rollback position of the candidate cluster of head node v3 is v5. The rollback position is recorded in the index H1, that is, H1 = {v1:v5,v2:v3,v3:v5}.
[0094] 3. Based on the H1 index record, we determine that there are three head nodes to be updated: v1, v2, and v3. To update the candidate clique corresponding to v1, we need the large neighbor v5. We retrieve the small neighbor set of v5 and calculate the common neighbor set of v1 and v5 as {v2}. We then construct a FIFO for v1 and write {v2}-v5 into it.
[0095] To update the candidate cluster corresponding to v2, we need large neighbor nodes v4, v5, v6, and v7. We take out the small neighbor node sets of v4, v5, v6, and v7, calculate the common neighbor node set of v2 and v4 as an empty set, calculate the common neighbor node set of v2 and v5 as {v4}, calculate the common neighbor node set of v2 and v6 as {v4, v5}, and calculate the common neighbor node set of v2 and v7 as {v6}. We construct a FIFO for v2 and {v4}-v5, {v4,v5}-v6, {v6}-v7 are written in sequence.
[0096] To update the candidate clique corresponding to v3, we need the large neighbor nodes v5 and v6. We then extract the small neighbor node sets of v5 and v6, calculate the common neighbor node set of v3 and v5 as {v4}, and the common neighbor node set of v3 and v6 as {v4, v5}. We then construct a FIFO for v3 and write {v4}-v5 and {v4, v5}-v6 into it.
[0097] 4. Create three subtasks corresponding to the head nodes v1, v2, and v3 to be updated. Each subtask constructs three FIFO queues and allocates a corresponding BRAM block. Extract candidate clusters from the HBM and roll back the corresponding candidate clusters based on the H1 index record. Figure 6 (2) shows the result after the candidate clusters corresponding to v1, v2, and v3 are rolled back. Figure 7 shows the status of the FIFO and BRAM before the three subtasks begin execution.
[0098] Take subtask 2 as an example to illustrate the calculation process of candidate group update. As shown in Figure 8 (1), it is the initial state, traversing the FIFO A The public neighbor node set in . Take out the public neighbor node set of the first large neighbor node v4 FIFO B If the candidate cluster set is empty, {v4} is directly written into the temporary queue as the candidate cluster. The active queue and the temporary queue are swapped.
[0099] Take out the public neighbor node set {v4} of the second large neighbor node v5, FIFO c The candidate cluster set is {v4}, and IList = {v4}, cnt A =cnt B =cnt C =1, {v4,v5} is written into the temporary queue as a candidate clique, triggering the traversal stop condition.
[0100] Take out the public neighbor node set {v4, v5} of the third large neighbor node v6, FIFO B The candidate cluster set is {v4, v5}, and IList = {v4, v5}, cnt A =cnt B =cnt C =2, {v4,v5,v6} is written into the temporary queue as a candidate clique, triggering the traversal stop condition.
[0101] Take out the public neighbor node set {v6} of the fourth largest neighbor node v7, FIFO c The candidate cluster set is {v4, v5, v6}, and IList = {v6}, cnt C =cnt A <cnt B , {v4,v5,v6} and {v6,v7} are written into the temporary queue as candidate groups, triggering the traversal stop condition. At this time, FIFO A The traversal is completed and subtask 2 calculation is completed.
[0102] 5. Write the updated candidate cluster results calculated by the three subtasks back to the HBM.
[0103] 6. Send the updated candidate cluster result data to the PC host, and extract the maximum cluster through the candidate cluster filtering operation on the PC host.
[0104] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A dynamic maximum clique enumeration device based on FPGA with HBM, characterized in that: Including HBM, matrix calculation unit, sorting calculation unit and update calculation unit built by functional isolation and algorithm function division of FPGA's own FIFO; The HBM is used to store dynamic edge flows, full-graph adjacency matrices, and candidate clusters transmitted from an external PC host for updating the graph structure; The matrix calculation unit is used to update the full-graph adjacency matrix based on the dynamic edge flow and send the updated full-graph adjacency matrix to the HBM storage, and at the same time determine the head node to be updated of the candidate cluster that needs to be updated; The sorting calculation unit is used to construct a sorted set of candidate cluster reconstructions by sorting data blocks according to the updated full-graph adjacency matrix and each head node to be updated; The update calculation unit is used to execute the update task of the candidate clusters corresponding to each head node to be updated in parallel based on the sorted set reconstructed by the candidate clusters, and send the updated candidate clusters to the HBM storage, and send the updated candidate clusters to the PC host to take filtering operations to extract the maximum clusters.
2. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 1, characterized in that: The matrix calculation unit includes a first FIFO, and updates the full-graph adjacency matrix based on the dynamic edge flow, including: The dynamic edge flow obtained from HBM is cached in the first FIFO, and the node to be updated and the neighbor node set of the node to be updated are determined according to the dynamic edge flow. The old adjacency list of all nodes to be updated is obtained from the full-graph adjacency matrix of HBM, and the old adjacency list is updated according to the neighbor node set. The updated adjacency list is written to HBM storage to implement the update of the full-graph adjacency matrix, wherein the neighbor node set includes the small neighbor node set and the large neighbor node set of the node to be updated.
3. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 2, characterized in that: In the matrix calculation unit, determining the head node to be updated of the candidate cluster that needs to be updated includes: The head nodes to be updated of the candidate clusters generated by the current batch of dynamic edge flows are determined according to the small neighbor node set of the node to be updated, and the nodes corresponding to the positions where the candidate clusters of each head node to be updated need to be rolled back are recorded using indexes.
4. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 3, characterized in that: In the sorting calculation unit, a sorting set of candidate cluster reconstruction is constructed by data block sorting according to the updated full-graph adjacency matrix and each head node to be updated, including: Obtain the updated adjacency list corresponding to each head node to be updated in the index record from the updated full-graph adjacency matrix in the HBM, and calculate the large neighbor node set required for updating the candidate clique of each head node to be updated based on the index record information and the updated adjacency list; Obtain the small neighbor node set corresponding to each large neighbor node from the HBM, take the intersection of each small neighbor node set with the large neighbor node set, and determine the common neighbor node set of each large neighbor node; A second FIFO is constructed for each head node to be updated, and the large neighboring nodes of each head node to be updated and their corresponding public neighboring node sets are sorted according to the node sequence numbers and stored in the second FIFO to form a sorted set of candidate cluster reconstruction.
5. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 3, characterized in that: In the update calculation unit, based on the sorted set reconstructed by the candidate cluster, the update task of the candidate cluster corresponding to each head node to be updated is executed in parallel, including: For each head node to be updated in the index record, a subtask and three FIFO queues and BRAM blocks corresponding to the subtask are established. One FIFO queue is used to store the large neighbor nodes and their common neighbor node sets obtained from the sorted set. The other two FIFO queues are alternately temporary queues and active queues. The temporary queue stores the candidate team column corresponding to the node to be updated, and the active queue stores the candidate team column being updated. The BRAM block includes a node set record block and three length record blocks. The node set record block is used to store the intersection of the current common neighbor node set and the current candidate group. The three The length record block records the length of the current common neighbor node set, the length of the current candidate cluster, and the length of the intersection respectively; Based on the node ordering corresponding to each head node to be updated obtained from the sorted set, the subtasks corresponding to all head nodes to be updated use three FIFO queues and BRAM blocks to execute the update tasks of the candidate clique in parallel.
6. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 5, characterized in that: Each subtask performs the candidate group update process, including: Obtain the old candidate cluster corresponding to each head node to be updated from the HBM and store it in the active queue; obtain the large neighbor nodes and their common neighbor node sets corresponding to each head node to be updated from the sorted set and store them in the first FIFO queue; Sequentially access each large neighbor node and its corresponding public neighbor node set in the FIFO queue, and use each large neighbor node and its corresponding public neighbor node set to update the old candidate cluster in the active queue. The updated subsequent maximum cluster is transferred to the temporary queue, and the temporary queue and active queue marks are exchanged for the next round of update operations.
7. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 6, characterized in that: Update operation process, including: If the candidate team list in the active queue is empty, or the public neighbor node set is empty, the large neighbor node corresponding to the current public neighbor node set is directly added to the temporary queue as a candidate team, ending the update operation of this round; If the candidate cluster maximal queue in the active queue is not empty, each candidate cluster is traversed in turn, and a single comparison operation between the common neighbor node set and the candidate cluster is performed to update the candidate cluster until the candidate cluster list traversal is completed or the traversal stop condition is triggered. During the traversal process, two sets FSet and PSet are maintained to temporarily store the candidate clusters that need to be further screened.
8. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 7, characterized in that: A single comparison operation, including: Store the common neighbor node set in BRAM, use the node as the BRAM address, count by judging whether the value of BRAM is 1, and store the length count of the common neighbor node set in the first length record block cnt A , take the current candidate group as the read address of BRAM, count by judging whether the value of BRAM is 1, and store its length count in the second length record block cnt B In the process, the intersection of the common neighbor node set and the candidate cluster is stored in the node set record block IList, and the length of the node set record block is stored in the third length record block cnt C In the example, the values in the three length record blocks are compared and the following four operations are performed: When cnt A =cnt B =cnt C , that is, if the common neighbor node set and the candidate cluster set are equal, then the large neighbor node is added to the candidate cluster set and written into the temporary queue, and the other candidate clusters in the active queue are read out and written into the temporary queue, triggering the traversal stop condition; When cnt C =cnt A <cnt B , that is, the common neighbor node set is a true subset of the candidate cluster set, then the candidate cluster set is written into the temporary queue, and after adding the large neighbor node to the common neighbor node set, it is written into the temporary queue, and the other candidate clusters in the active queue are read out and written into the temporary queue, triggering the traversal stop condition; When cnt C =cnt B <cnt A , that is, the candidate cluster set is a true subset of the common neighbor node set, then the large neighbor node is added to the candidate cluster set and written into the temporary queue, and the IList is stored in the set FSet; When cnt B ≠0 and cnt B <cnt C and cnt B <cnt A , that is, the candidate cluster set and the common neighbor node set do not contain each other, and the intersection is not empty, then the candidate cluster set is written into the temporary queue and IList is stored in the set PSet.
9. The dynamic maximum clique enumeration device based on FPGA with HBM according to claim 8, characterized in that: Also includes: When the candidate team list traversal is completed and the traversal stop condition is not triggered, the candidate groups in the set FSet and PSet are screened. The specific operation is as follows: traverse each candidate group in PSet. If the traversed candidate group is not a true subset of any other candidate group in PSet, Any proper subset of a candidate cluster in FSet is written into the temporary queue after the traversed candidate cluster is added to the large neighbor node.
10. A dynamic maximum clique enumeration method based on FPGA with HBM, characterized in that: The method adopts the dynamic maximal group enumeration device according to any one of claims 1 to 9, and the method comprises the following steps: Obtain the dynamic edge flow from the external PC host for updating the graph structure and store it in HBM. HBM also stores the full graph adjacency matrix and candidate clusters. Use the matrix calculation unit to update the full-graph adjacency matrix based on the dynamic edge flow and send the updated full-graph adjacency matrix to HBM storage, and determine the head node to be updated of the candidate cluster that needs to be updated; Using the sorting calculation unit, a sorting set of candidate cluster reconstruction is constructed by sorting the data blocks according to the updated full-graph adjacency matrix and each head node to be updated; The update calculation unit is used to execute the update task of the candidate cluster corresponding to each head node to be updated in parallel based on the sorted set reconstructed by the candidate cluster, and the updated candidate cluster is sent to the HBM storage; Send updated candidate cluster result data to the PC host, and the PC host extracts the maximum cluster through candidate cluster filtering operation.
Citation Information
Patent Citations
Dynamic graph processing method based on FPGA
CN111104224A
Maximum group enumeration method oriented to social network flow data and based on MPICH parallel computing
CN115935080A
Graph data mining method and device and electronic equipment
CN116842073A
Dynamic tile sequencing in graphic processing
US20230095535A1