A Multi-threaded Loop Enumeration Algorithm Based on NumPy

By using a multi-threaded loopback enumeration algorithm based on NumPy in graph data processing, the problem of large memory consumption and long calculation time in the prior art is solved, and efficient loop calculation is realized.

CN113934976BActive Publication Date: 2025-06-13NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111060114.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-10
Publication Date
2025-06-13
Estimated Expiration
2041-09-10

AI Technical Summary

Technical Problem

The existing loopback enumeration algorithm has problems such as high memory consumption and long calculation time in graph data processing, especially when enumeration of finite length loops is required.

Method used

The multi-threaded loopback enumeration algorithm based on NumPy is adopted to improve the efficiency of loop computing through multi-threading technology and NumPy library. Specific steps include data preprocessing, the main thread calculates the path index and sends it to the auxiliary thread, the auxiliary thread expands the path index and returns to the loop path.

Benefits of technology

It significantly improves the efficiency of loopback enumeration algorithm, reduces memory consumption and calculation time, and is suitable for loop calculation in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113934976B_ABST
    Figure CN113934976B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-threaded loop enumeration algorithm based on NumPy. First, data preprocessing is performed to process multi-source heterogeneous data into graph data of point sets and edge sets that meet the requirements, and a NumPy adjacency matrix is further constructed. In terms of the main thread, the path indices of all loops are calculated respectively. Each iteration calculates and generates the path index of the current iteration relative to the previous iteration, and sends it to the auxiliary thread for expansion. After the iteration is completed, the main thread receives the paths of all loops sent by the auxiliary thread. The auxiliary thread receives the path index passed by the main thread, combines it with the stored path data between nodes to obtain the specific paths of the loops, and when the iteration ends, sends all loop results back to the main thread. The multi-threaded loop enumeration algorithm provided by the present invention can further improve the execution efficiency of the algorithm by using the efficient NumPy library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph data processing, and mainly relates to a multi-threaded loop enumeration algorithm based on NumPy. Background Art

[0002] The problem of loop enumeration for graph data is a common problem in the field of graph analysis, and is often used in graph analysis scenarios such as economic investigation and criminal investigation. In most real-world scenarios, the length of the loops to be enumerated is limited, and specifying the upper limit of the length of the enumerated loops can greatly reduce the computational cost. Existing loop enumeration algorithms are based on structured graph data for calculation, and have the disadvantages of large memory consumption and long calculation time. Summary of the Invention

[0003] Object of the Invention: Aiming at the problems existing in the above background art, the present invention provides a multi-threaded loop enumeration algorithm based on NumPy, which improves the efficiency of loop calculation through multi-threaded technology and the NumPy library, and provides an efficient solution for loop calculation in different scenarios.

[0004] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is as follows:

[0005] A multi-threaded loop enumeration algorithm based on NumPy, comprising the following steps:

[0006] Step S1, data preprocessing; processing multi-source heterogeneous data into graph data of a point set and an edge set that meet the requirements through an ETL process, and further constructing an adjacency matrix based on the graph data;

[0007] Step S2, calculating path indexes by the main thread; the main thread performs n iterations, calculates the path indexes of all loops with lengths from 1 to n respectively, calculates and generates the path indexes of the current iteration relative to the previous iteration each time, and sends them to the auxiliary threads for expansion; after the iteration is completed, the main thread receives the paths of all loops sent by the auxiliary threads;

[0008] Step S3, the auxiliary threads receive the path indexes sent by the main thread, and obtain the paths of the loops by combining the stored paths between nodes and the path indexes; when the iteration ends, the auxiliary threads send the paths of all loops to the main thread.

[0009] Further, constructing the adjacency matrix in step S1 specifically includes:

[0010] According to the input graph data and the maximum length of the loop to be calculated, obtain the adjacency matrix of the nodes from the graph data and construct it into a boolean NumPy matrix. Use True to represent that there is a connection between nodes and False to represent that there is no connection between nodes; construct a path dictionary with a 1-hop distance between nodes from the graph data, where the key is the id of the starting node and the ending node of the edge, and the value is the id of the edge; use the maximum length of the loop to be calculated as the total number of iterations N; start the auxiliary thread with the generated path dictionary as the initial parameter.

[0011] Further, the specific process of the main thread calculating the path index in step S2 includes:

[0012] Step S2.1: Calculate the dot product of the adjacency matrix representing i - 1 hops in the previous iteration round and the initial adjacency matrix based on the dot operator of NumPy, and use the dot product as the adjacency matrix of the i-th hop.

[0013] Step S2.2: Use the nonzero method of NumPy to obtain the positions of non-zero elements in the rows and columns of the i-th hop adjacency matrix; represent the positions in the form of row vectors and column vectors, obtain the row data of the adjacency matrix of i - 1 hops with the row vector as the coordinate, and obtain the column data of the initial adjacency matrix with the column vector as the column coordinate.

[0014] Step S2.3: Use the AND operation and where operator of NumPy to find the positions where the row data and column data are equal; obtain the path index of the current round based on the positions where the row data and column data are equal, and transmit the current iteration round i and the path index to the auxiliary thread for expansion; reset the diagonal elements in the result of the dot product in step S2.1 to 0.

[0015] Further, the specific iterative process executed by the auxiliary thread in step S3 includes:

[0016] Use the edges in the graph data as the initial value of the paths between nodes stored by the auxiliary thread. The auxiliary thread obtains the path index obtained in each iteration round and the total number of iterations to be performed from the main thread, obtains the path value between nodes of the current round based on the paths between nodes and the index stored by the auxiliary thread, selects all the loop paths obtained in this iteration from all the path values between nodes, and updates the node paths stored by the auxiliary thread to the path values between nodes obtained in the current round; after the iteration ends, send all the loop paths to the main thread and end the execution of the auxiliary thread.

[0017] Further, the auxiliary thread obtains the message queue for transmitting the path index and the message queue for transmitting the loop path result from the main thread;

[0018] First, the auxiliary thread obtains the initial path index from the message queue of the path index as the result of the expansion of the path index in the 0th iteration step; during the specific iteration process, the auxiliary thread obtains the current iteration round and the path index from the message queue of the path index; the path index refers to the path index in the previous iteration.

[0019] Expand the path index of this iteration according to the path index of the previous iteration. The key of the path index represents the start and end nodes of the path. Take the value of the path index with the same start and end nodes as the path of the loop and add it to the list representing all loop paths; update the variable representing the path index of the previous step according to the result of the path expansion.

[0020] When the iteration id of this time is greater than or equal to the maximum number of iterations, pass the list representing all loop paths to the main thread through the message queue for passing loop path results and end the loop.

[0021] Furthermore, in step S1, ETL refers to obtaining data that meets the requirements from multi-source heterogeneous data through extraction, transformation, and loading.

[0022] Beneficial effects:

[0023] This paper proposes a multi-threaded loop enumeration algorithm based on NumPy, decouples the discovery and expansion of loops with the help of multi-threaded technology, and further improves the execution efficiency of the algorithm with the help of the efficient operation of the NumPy library. In order to cope with the data differences in different scenarios, this paper also designs a link for preprocessing multi-source data to further improve the generality of the algorithm. Description of the drawings

[0024] Figure 1 is the flowchart of the multi-threaded loop enumeration algorithm based on NumPy of the present invention;

[0025] Figure 2 is the execution flowchart of the main thread and the auxiliary thread in the embodiment of the present invention. Detailed implementation manners

[0026] The following further describes the present invention with reference to the drawings.

[0027] The multi-threaded loop enumeration algorithm based on NumPy provided by the present invention is as Figure 1 shown, and specifically includes the following steps:

[0028] Step S1, data preprocessing; process multi-source heterogeneous data into graph data of point sets and edge sets that meet the requirements through the ETL process, and further construct an adjacency matrix based on the graph data. The ETL process here refers to obtaining data that meets the requirements from multi-source heterogeneous data through various methods such as extraction, transformation, and loading. The finally constructed adjacency matrix is a NumPy matrix.

[0029] According to the input graph data and the maximum length of the cycle to be calculated, obtain the adjacency matrix of the nodes from the graph data and construct it into a boolean NumPy matrix. Use True to represent that there is a connection between nodes and False to represent that there is no connection between nodes; construct a path dictionary with a distance of 1 hop between nodes from the graph data, where the key is the id of the start node and the end node of the edge, and the value is the id of the edge; use the maximum length of the cycle to be calculated as the total number of iterations N; start the auxiliary thread with the generated path dictionary as the initial parameter.

[0030] Step S2: Calculate the path index by the main thread; the main thread performs n iterations, calculates the path indexes of all cycles with lengths from 1 to n respectively, calculates the path index of the current iteration relative to the previous iteration each time, and sends it to the auxiliary thread for expansion; after the iteration is completed, the main thread receives all the cycle paths sent by the auxiliary thread. Specifically,

[0031] Step S2.1: Calculate the dot product of the adjacency matrix representing (i - 1) hops in the previous iteration and the initial adjacency matrix based on the dot operator of NumPy, and use the dot product as the adjacency matrix of the i-th hop.

[0032] Calculate the i-th power of the adjacency matrix of the graph. The calculation method is to take the dot product of the result of the (i - 1)-th power in the previous iteration and the adjacency matrix of the graph. All matrices are stored as boolean NumPy matrices, and the result is denoted as matrix_i.

[0033] Step S2.2: Use the nonzero method of NumPy to obtain the positions of the non-zero elements in matrix_i; record their positions with the row vector rows and the column vector columns. Extract the rows in matrix_i according to the row vector rows to generate a new row vector new_rows, and extract the columns of the original graph adjacency matrix according to the column vector to generate a new column vector new_columns; use the unique, cumsum, and split operators of NumPy to split the row vector new_rows and perform a groupby operation on the column vector new_columns. The split row vector is new_rows_unique, and the grouped column vector is new_columns_groupby;

[0034] Step S2.3: Generate a new path index based on the new row vector and column vector. The specific method is as follows: Take the elements of new_rows_unique and new_columns_groupby and pack them into tuples (row_index, columns_list) one by one, and generate a new path index with (rows[row_indx], columns[row_index]) as the key and (rows[row_index], column_list, columns[row_index]) as the value. Among them, (rows[row_index], column_list) represents (row number, column number list) of the path dictionary in the previous step, and (column_list, columns[row_index]) represents (row number list, column number) of the initial path dictionary; Send the current round and the path index generated in the previous step to the auxiliary thread; Set all diagonal elements of matrix_i to 0, and update the dot product of the adjacency matrix matrix_(i - 1) of the previous round with matrix_i.

[0035] Step S3: The auxiliary thread receives the path index sent by the main thread, and combines the stored paths between nodes and the path index to obtain the paths of the loops; When the iteration ends, the auxiliary thread sends all the paths of the loops to the main thread.

[0036] Use the edges in the graph data as the initial value of the paths between nodes stored by the auxiliary thread. The auxiliary thread obtains the path index obtained in each round of iteration and the total number of iterations required from the main thread, obtains the path value between nodes in the current round according to the paths between nodes and the index stored by the auxiliary thread, selects all the loop paths obtained in this iteration from all the path values between nodes, and updates the node paths stored by the auxiliary thread to the path values between nodes obtained in the current round; After the iteration ends, send all the paths of the loops to the main thread to end the execution of the auxiliary thread. Specifically,

[0037] The auxiliary thread accepts the following several parameters: the total number of iterations N, the queue path_message_queue for transmitting the path index, and the queue cycles_queue for transmitting all the loops. The auxiliary thread obtains the initial path dictionary from path_message_queue as the path index for the 0th step of iteration. During each iteration, the auxiliary thread obtains the updated path index from the main thread in the queue, expands the current path index in combination with the path dictionary of the previous iteration round to obtain the path dictionary of the current round, extracts all the loops from the path dictionary, and adds the loops to the list representing all the loop paths. After the iteration ends, send all the loop results back to the main thread.

[0038] Further, obtain the ID of the current iteration and the path index of the current iteration from the path_message_queue; the path index is a variable of dictionary type, and its value represents the index of the path dictionary in the previous iteration. Based on this, expand its value. The expansion method is to obtain the path list from the previous path dictionary according to (rows[row_index], column_list), obtain the path list from the initial path dictionary according to (column_list, columns[row_index]), perform the Cartesian product on the two lists, and store the result as the new value in the path index of the current iteration round; assign the variable representing the path dictionary of the previous round to the updated path index, and obtain all the loop paths from it, and store them in the list representing all the loop paths; determine whether the obtained ID of the current round is greater than the total number of iterations N. If so, send the list representing all the loop paths back to the main thread.

[0039] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A multi-threaded loop enumeration algorithm based on NumPy, characterized in that, it includes the following steps: Step S1, data preprocessing; process multi-source heterogeneous data into graph data of a point set and an edge set that meet the requirements through the ETL process, and further construct an adjacency matrix based on the graph data; Step S2, calculate path indexes by the main thread; the main thread performs n iterations, calculates the path indexes of all loops with lengths from 1 to n respectively, calculates and generates the path indexes of the current iteration relative to the previous iteration each time, and sends them to the auxiliary thread for expansion; after the iteration is completed, the main thread receives the paths of all loops sent by the auxiliary thread; Step S3, the auxiliary thread receives the path indexes sent by the main thread, and obtains the paths of the loops by combining the stored paths and path indexes between nodes; when the iteration ends, the auxiliary thread sends the paths of all loops to the main thread; The specific construction of the adjacency matrix in the step S1 includes: According to the input graph data and the maximum length of the loops to be calculated, obtain the adjacency matrix of the nodes from the graph data, and construct it into a boolean-type NumPy matrix, using True to represent that there is a connection between nodes, and using False to represent that there is no connection between nodes; construct a path dictionary with a node distance of 1 hop from the graph data, where the key is the id of the start node and the end node of the edge, and the value is the id of the edge; use the maximum length of the loops to be calculated as the total number of iterations N; start the auxiliary thread with the generated path dictionary as the initial parameter; The specific calculation of the path indexes by the main thread in the step S2 includes: Step S2.1, calculate the dot product of the adjacency matrix representing i-1 hops in the previous iteration round and the initial adjacency matrix based on the dot operator of NumPy, and use the dot product as the adjacency matrix of the i-th hop; Step S2.2, use the nonzero method of NumPy to obtain the positions of the non-zero elements in the rows and columns of the i-th hop adjacency matrix; represent the positions in the form of row vectors and column vectors, obtain the row data of the adjacency matrix of i-1 hops with the row vector as the coordinate, and obtain the column data of the initial adjacency matrix with the column vector as the column coordinate; Step S2.3, find the positions where the row data and the column data are equal through the AND operation and the where operator of NumPy; obtain the path indexes of the i-th iteration according to the positions where the row data and the column data are equal, and transmit the current iteration round i and the path indexes to the auxiliary thread for expansion; reset the diagonal elements in the result of the dot product in step S2.1 to 0; The specific iteration process executed by the auxiliary thread in the step S3 includes: Use the edges in the graph data as the initial value of the paths between nodes stored by the auxiliary thread. The auxiliary thread obtains the path indexes obtained in each iteration and the total number of iterations to be performed from the main thread, obtains the path values between nodes of the current round according to the paths and indexes between nodes stored by the auxiliary thread, selects all the loop paths obtained in the current iteration from all the path values between nodes, and updates the node paths stored by the auxiliary thread to the path values between nodes obtained in the current round; after the iteration ends, send the paths of all loops to the main thread and end the execution of the auxiliary thread; The auxiliary thread obtains the message queue for passing the path index and the message queue for passing the loop path result from the main thread; First, the auxiliary thread obtains the initial path index from the message queue of the path index as the result of the expansion of the path index in the 0th iteration step; in the specific iteration process, the auxiliary thread obtains the current iteration round and the path index from the message queue of the path index; the path index refers to the path index in the previous iteration; Expand the path index of the current iteration according to the path index of the previous iteration. The key of the path index represents the start and end nodes of the path. Take the value of the path index with the same start and end nodes as the loop path and add it to the list representing all loop paths; update the variable representing the path index of the previous step according to the result of the path expansion; When the current iteration id is greater than or equal to the maximum number of iterations, pass the list representing all loop paths to the main thread through the message queue for passing the loop path result and end the loop.

2. A multi-threaded loop enumeration algorithm based on NumPy according to claim 1, characterized in that, in step S1, ETL refers to obtaining the required data from multi-source heterogeneous data by means of extraction, transformation, and loading.

Citation Information

Patent Citations

  • Fully parallel in-place construction of 3D acceleration structures in a graphics processing unit

    CN103440238A

  • Arm architecture-based NumPy operation acceleration optimization method

    CN112783503A