Parallel-based datalog system equivalent data processing method
Patent Information
- Application Number
- CN202211341678.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-10-31
AI Technical Summary
[0005]本发明针对并行Datalog系统求解技术中对含有等价关系规则的计算存在冗余,造成计算资源浪费这一问题,提供一种基于并行的Datalog系统等价数据处理方法,在充分利用多核CPU的计算资源的同时,过滤掉求解中不必要的计算,以提高Datalog系统对此类常见规则的计算效率,并使得生成的新元组具有双倍信息量
[0027]本发明的有益效果是:本发明提出了用于Datalog规则的并行等价关系的求解技术,采用多线程技术,将计算的输入按照负载均衡策略分配到不同的线程中,多个线程同时进行计算并在计算过程中过滤掉了冗余的计算,减少了整体的计算量,使得硬件资源的得到了充分的利用,也提高了此类规则的计算效率。同时,本发明提出一种哈希位向量表来存储生成的元组,能辅助并行的实现,采用的数据结构中,每个存储元素包含了双倍的信息,减少了空间的开销。比起传统的方法,本发明能够减少求解中的计算量,同时减少空间的消耗,能提高Datalog系统的性能。
Smart Images

Figure CN115543581B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of knowledge bases, specifically relating to an equivalent data processing method based on a parallel Datalog system. Background Technology
[0002] Datalog is a declarative logic programming language, syntactically a subset of Prolog, and is commonly used as a query language for deductive databases. Similar to SQL, but with added recursive semantics, Datalog possesses powerful expressive capabilities, making it easy to represent tasks requiring recursive computation. In recent years, Datalog systems have been widely applied in data integration, information extraction, program optimization, security, and cloud computing. For some large-scale applications, the input of a Datalog program may contain massive amounts of tuples, posing a challenge to program efficiency. With hardware advancements and the widespread adoption of multi-core CPUs, many technologies have been developed for parallel Datalog system computation, aiming to improve the operational efficiency of Datalog programs.
[0003] Currently, many parallel Datalog solving techniques have been proposed, which can be broadly categorized into top-down and bottom-up methods, with the majority employing bottom-up evaluation. Equivalence relations exist in many common Datalog applications. An equivalence relation R is a binary relation that is reflexive, symmetric, and transitive. Examples include blockchain user group analysis, program alias analysis, strongly connected component analysis of graphs, and optimal network routing analysis. However, typical parallel Datalog solving techniques do not specialize in addressing the characteristics of equivalence relations; they are still explicitly represented and computed. This leads to the explicit generation of different tuples with the same semantics during computation, resulting in significant redundant calculations and high computational time overhead.
[0004] Existing methods propose modular frameworks for solving Datalog systems, integrating specialized methods for handling specific rules, but they do not introduce new data structures. Some methods use variants of the Union-find structure to handle equivalence rules, but their solutions do not reduce computational cost. Equivalence relations are common in Datalog applications, and adding specialized computation modules for equivalence relations in parallel Datalog systems can improve system computational efficiency and reduce computational time overhead. Summary of the Invention
[0005] This invention addresses the problem of redundant calculations in parallel Datalog system solving techniques that waste computational resources when dealing with rules containing equivalence relations. It provides a parallel equivalent data processing method for Datalog systems that fully utilizes the computing resources of multi-core CPUs while filtering out unnecessary calculations in the solving process. This improves the computational efficiency of Datalog systems for such common rules and ensures that the generated new tuples have twice the information content.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A method for equivalent data processing based on a parallel Datalog system, characterized by the following steps:
[0008] S1: Allocate memory space and initialize the hash bit vector table, read EDB tuple data from the file and store it in memory, build a B+ tree to index the EDB tuple data, calculate the entry rule, and store the result in the hash bit vector table.
[0009] S2: Request m child threads from the operating system and put them into the thread pool. Start the child threads to read EDB tuple data as input data.
[0010] S3: Each sub-thread performs calculations on the input data and returns the number of new tuples as a return value to the CPU's general-purpose register;
[0011] S4: After all tuples have been calculated, the main thread retrieves the return values of all child threads from the general-purpose register and accumulates them to obtain the number of newly generated tuples. If the number is not 0, the incremental tuples obtained after deduplication are used as input data and the process returns to step S3; otherwise, the calculation ends.
[0012] To optimize the above technical solution, the specific measures also include:
[0013] Furthermore, the specific steps of step S1 are as follows:
[0014] S11: Request a contiguous space in memory with a size equal to the preset size of the hash array; store a pointer in the element with index i in the hash array and point to a bit array of size ni bits, where n represents the number of vertices in the EDB tuple data; set all bits to zero; each bit in the hash bit vector table has a dual semantic meaning, representing both tuple (x, y) and tuple (y, x).
[0015] S12: Read EDB tuple data from the file and store it in memory. Build a B+ tree to index the EDB tuple data and select a column from the EDB tuple data as the key.
[0016] S13: The user defines a comparison function Compare(x, y) to set the size relationship in the EDB tuple data. If x is greater than y, then swap x and y, perform a hash calculation on x, store the tuple in the bit vector pointed to by the pointer at index hash(x) in the hash array, set the bit at offset yx-1 in the bit vector to 1, and mark it as the increment.
[0017] Furthermore, in step S2, the specific steps for the sub-thread to read the EDB tuple data are as follows:
[0018] S21: The main thread partitions the data based on the leaf nodes of the B+ tree index, grouping tuples with the same key column into a single data block; the main thread encapsulates the keys in the B+ tree index into tasks and puts them into a task queue; idle child threads read one task from the task queue and execute it each time while the task queue is not empty by polling.
[0019] S22: The sub-thread queries the corresponding leaf node in the B+ tree index based on the keyword read from the task, and reads the corresponding tuple from the EDB tuple data based on the offset recorded in the data block pointed to by the pointer in the leaf node, and uses it as the input data for calculation.
[0020] S23: After the child thread finishes its current calculation, it reads tasks from the task queue by polling until the task queue is empty.
[0021] Furthermore, in step S3, the specific steps for each sub-thread to perform calculations on the input data are as follows:
[0022] S31: Sub-thread i, i∈[1,m] applies recursive rules to the read data, puts the incremental tuple into the buffer according to the incremental flag bit, performs a join operation on the incremental tuple in the buffer according to the rules, and filters out tuples that do not need to be calculated before performing the join operation;
[0023] S32: Sub-thread i performs a deduplication operation on the new tuple generated by the connection, calculates the hash value of the column corresponding to the B+ tree key in the new tuple, indexes the bit vector at the corresponding position in the hash table according to the hash value, and uses the hash value of another column in the tuple as the offset to locate the specific position in the bit vector and reads the data at that position.
[0024] S33: Thread i reads the data at the position described in step S32 and determines whether the tuple should be inserted. If the bit at that position is 1, it means that the tuple already exists and is discarded. If the bit at that position is 0, it means that the tuple does not exist. Thread i sets the data at that position to 1, which means that the tuple should be inserted, and marks the tuple as an increment. In the hash bit vector table, the bit at position (x, y) being 1 means that both the (x, y) tuple and the (y, x) tuple are inserted into the table. Thread i uses the count variable to record the number of new tuples inserted into the hash bit vector table in this round of calculation.
[0025] Further, in step S31, filtering out tuples that do not need to be calculated before performing the join operation is as follows: for the current tuple (x, y), if the return value of the Compare function is less than 0 and x > y, then this tuple is filtered out and no calculation is performed on this tuple.
[0026] Furthermore, the present invention also proposes a computer-readable storage medium storing a computer program, characterized in that the computer program causes a computer to execute the parallel-based Datalog system equivalent data processing method as described above.
[0027] The beneficial effects of this invention are as follows: This invention proposes a parallel equivalence relation solving technique for Datalog rules. It employs multi-threading technology, distributing the computational input to different threads according to a load balancing strategy. Multiple threads perform computation simultaneously, filtering out redundant calculations during the process, reducing the overall computational load, fully utilizing hardware resources, and improving the computational efficiency of such rules. Simultaneously, this invention proposes a hash bit vector table to store the generated tuples, which aids in parallel implementation. In this data structure, each storage element contains double the information, reducing space overhead. Compared to traditional methods, this invention reduces the computational load and space consumption during the solution process, thereby improving the performance of the Datalog system. Attached Figure Description
[0028] Figure 1 This is a flowchart of an equivalent data processing method based on a parallel Datalog system.
[0029] Figure 2 This is a diagram of the hash bit vector table structure.
[0030] Figure 3 This is a load balancing framework diagram. Detailed Implementation
[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0032] Combination Figure 1 In one embodiment, the present invention proposes an equivalent data processing method based on a parallel Datalog system, the overall implementation steps of which are as follows:
[0033] Step 1: Allocate memory space and initialize the hash bit vector table, read EDB tuple data from the file and store it in memory, build a B+ tree to index the EDB tuple data, calculate the entry rule, and store the result in the hash bit vector table.
[0034] The specific steps for step 1 are as follows:
[0035] Step 11: Allocate space for the hash bit vector table: First, allocate a contiguous block of memory, the size of which is the preset size of the hash array; store a pointer in the element at index i of the hash array, allocate a contiguous block of space (bit array) of size n-i bits for this pointer, where n represents the number of vertices in the EDB tuple data, and set all bits to zero. Each bit in the hash bit vector table has a dual semantic meaning, that is, it represents both the tuple (x, y) and the tuple (y, x);
[0036] Step 12: Read the EDB array (i.e., EDB tuple data) from the file and store it in memory. Build a B+ tree to index the EDB tuple data and select a column in the EDB tuple data as the key.
[0037] Step 13: Calculate the entry rule and store the calculated tuple (x, y) into the hash bit vector table: The user defines a comparison function Compare(x, y) to manually set the size relationship in the EDB tuple data. If x is greater than y, then swap x and y, perform a hash calculation on x, and store the tuple in the bit vector (bit array) pointed to by the pointer at index hash(x) in the hash array. That is, set the bit at offset yx-1 in the bit vector to 1 and mark it as the increment.
[0038] Step 2: Request m child threads from the operating system and add them to the thread pool. Start the child threads to read data blocks from the EDB tuple data as input data.
[0039] In step 2, the specific steps for the child thread to read the input data are as follows:
[0040] Step 21: The main thread partitions the data based on the leaf nodes of the B+ tree index, grouping tuples with the same key column into a single data block. The main thread encapsulates the keys in the B+ tree index into tasks and places them into a task queue. Idle child threads then poll the task queue to read and execute one task at a time while it is not empty.
[0041] Step 22: The sub-thread queries the corresponding leaf node in the B+ tree index based on the keyword read from the task, and reads the corresponding tuple from the EDB array based on the offset recorded in the data block pointed to by the pointer in the leaf node, and uses it as the input data for calculation.
[0042] Step 23: After the child thread finishes its current calculation, it reads tasks from the task queue by polling until the task queue is empty.
[0043] Step 3: Each thread performs calculations on the input data and returns the number of new tuples as the method's return value to the CPU's general-purpose register;
[0044] In step 3, each sub-thread performs the following calculations on the input data:
[0045] Step 31: Sub-thread i (i∈[1,m]) applies recursive rules to the read data: First, it puts the incremental tuple into the buffer according to the incremental flag bit, and then performs a join operation on the incremental tuple in the buffer according to the rules. Before performing the join operation, it filters out tuples that do not need to be calculated: For the current tuple (x, y), if the return value of the Compare function is less than 0, i.e., x>y, then this tuple is filtered out and no calculation is performed on this tuple;
[0046] Step 32: Sub-thread i performs a deduplication operation on the new tuple generated by the join. It calculates the hash value of the column corresponding to the B+ tree key in the new tuple, indexes the corresponding bit vector in the hash table based on the hash value, and uses the hash value of another column in the tuple as an offset to locate the specific position in the bit vector, then reads the data at that position.
[0047] Step 33: Thread i reads the data at the position described in step S32 and determines whether the tuple should be inserted: if the bit at that position is 1, it means the tuple already exists and is discarded; if the bit at that position is 0, it means the tuple does not exist, thread i sets the data at that position to 1, representing the insertion of the tuple, and marks the tuple as an increment. In the hash bit vector table, a bit of 1 at position (x, y) indicates that both the (x, y) tuple and the (y, x) tuple have been inserted into the table. Thread i uses the `count` variable to record the number of new tuples inserted into the hash bit vector table in this round of calculation.
[0048] Step 4: After all tuples have been calculated, the main thread retrieves the return values of all child threads from the register and accumulates them, which is the number of new tuples generated. If the result is not 0, the incremental tuple is used as input data and the process returns to step S3; otherwise, the calculation ends.
[0049] Combination Figure 2 The implementation steps for deduplication and insertion of newly generated tuples in this embodiment are as follows:
[0050] 1) Calculate the position of the newly generated tuple (x, y) in the hash bit vector table HashBits. Let the total number of vertices be n, and the hash function used be Hash(). First, calculate the index value pos_x of x in the hash array:
[0051] pos_x = Hash(x)
[0052] 2) Locate the bit vector pointed to by the pointer at index pos_x in the hash array, and determine the specific position pos_y (index starts from 0) by calculating the y value in the tuple (x, y):
[0053] pos_y = yx - 1
[0054] 3) Based on the calculated position, determine whether the tuple has already been inserted:
[0055]
[0056] 4) For a tuple that needs to be inserted, set the bit at that position to 1 and mark the tuple as the increment, so that it can be used as the input data for the next round of the loop.
[0057] Combination Figure 3 The implementation steps of the multi-threaded load balancing method in this embodiment are as follows:
[0058] 1) The main thread requests m threads and puts them into the thread pool. At this time, all threads are in an idle state.
[0059] 2) The main thread encapsulates the keys in the B+ tree index (i.e., the vertices appearing in the EDB table) into tasks and puts them into the task queue. Idle child threads read one task from the task queue each time through a polling method.
[0060] 3) The sub-thread queries the corresponding EDB tuple data block in the B+ tree index based on the keyword read from the task, calculates the number of incremental tuples (count), and puts it into the CPU's general-purpose register. Then, it retrieves new tasks from the task queue by polling until the task queue is empty.
[0061] 4) The main thread reads the count of increment tuples returned by the child threads from the CPU's general-purpose registers and increments all counts returned by all threads. The main thread waits until all tasks in the task queue have been completed and returned, then checks if the total count of increment tuples is 0. If it is 0, the process ends; otherwise, it returns to step 2.
[0062] In another embodiment, the present invention provides a computer-readable storage medium storing a computer program that causes a computer to perform the equivalent data processing method for a parallel Datalog system as described in the first embodiment.
[0063] In the embodiments disclosed in this application, a computer storage medium may be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of computer storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0064] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A method for equivalent data processing based on a parallel Datalog system, characterized in that, Includes the following steps: S1: Allocate memory space and initialize the hash bit vector table, read EDB tuple data from the file and store it in memory, build a B+ tree to index the EDB tuple data, calculate the entry rule, and store the result in the hash bit vector table; the specific steps of step S1 are as follows: S11: Request a contiguous space in memory with a size equal to the preset size of the hash array; store a pointer in the element at index i of the hash array and point to a bit array of size ni bits, where n represents the number of vertices in the EDB tuple data; set all bits to zero; each bit in the hash bit vector table has a dual semantic meaning, representing both tuple (x, y) and tuple (y, x); S12: Read EDB tuple data from the file and store it in memory. Build a B+ tree to index the EDB tuple data and select a column from the EDB tuple data as the key. S13: The user defines a comparison function Compare(x, y), which sets the size relationship in the EDB tuple data. If x is greater than y, then x and y are swapped, a hash calculation is performed on x, and the tuple is stored in the bit vector pointed to by the pointer at index hash(x) in the hash array. The bit at offset yx-1 in the bit vector is set to 1 and marked as the increment. S2: Request m child threads from the operating system and put them into the thread pool. Start the child threads to read EDB tuple data as input data. S3: Each sub-thread performs calculations on the input data and returns the number of new tuples as a return value to the CPU's general-purpose register; S4: After all tuples have been calculated, the main thread retrieves the return values of all child threads from the general-purpose register and accumulates them to obtain the number of newly generated tuples. If the number is not 0, the incremental tuples obtained after deduplication are used as input data and the process returns to step S3. Otherwise, the calculation ends.
2. The equivalent data processing method for a parallel Datalog system as described in claim 1, characterized in that: In step S2, the specific steps for the sub-thread to read the EDB tuple data are as follows: S21: The main thread partitions the data based on the leaf nodes of the B+ tree index, grouping tuples with the same key column into a single data block; the main thread encapsulates the keys in the B+ tree index into tasks and puts them into a task queue; idle child threads read one task from the task queue and execute it each time while the task queue is not empty by polling. S22: The sub-thread queries the corresponding leaf node in the B+ tree index based on the keyword read from the task, and reads the corresponding tuple from the EDB tuple data based on the offset recorded in the data block pointed to by the pointer in the leaf node, and uses it as the input data for calculation. S23: After the child thread finishes its current calculation, it reads tasks from the task queue by polling until the task queue is empty.
3. The equivalent data processing method for a parallel Datalog system as described in claim 1, characterized in that: In step S3, the specific steps for each sub-thread to perform calculations on the input data are as follows: S31: Sub-thread i, i∈[1, m] applies recursive rules to the read data, puts the incremental tuple into the buffer according to the incremental flag, performs a join operation on the incremental tuple in the buffer according to the rules, and filters out tuples that do not need to be calculated before performing the join operation; S32: Sub-thread i performs a deduplication operation on the new tuple generated by the connection, calculates the hash value of the column corresponding to the B+ tree key in the new tuple, indexes the bit vector at the corresponding position in the hash table according to the hash value, and uses the hash value of another column in the tuple as the offset to locate the specific position in the bit vector and reads the data at that position. S33: Thread i reads the data at the position described in step S32 and determines whether the tuple should be inserted. If the bit at that position is 1, it means that the tuple already exists and is discarded. If the bit at that position is 0, it means that the tuple does not exist. Thread i sets the data at that position to 1, which means that the tuple should be inserted, and marks the tuple as an increment. In the hash bit vector table, the bit at position (x, y) being 1 means that both the (x, y) tuple and the (y, x) tuple have been inserted into the table. Thread i uses the count variable to record the number of new tuples inserted into the hash bit vector table in this round of calculation.
4. The equivalent data processing method for a parallel Datalog system as described in claim 3, characterized in that: In step S31, filtering out tuples that do not need to be calculated before performing the join operation is as follows: For the current tuple (x, y), if the return value of the Compare function is less than 0 and x > y, then this tuple is filtered out and no calculation is performed on this tuple.
5. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute the equivalent data processing method for a parallel Datalog system as described in any one of claims 1-4.
Citation Information
Patent Citations
Optimization method for effectively improving B+ tree retrieval efficiency on GPU
CN111966678A
Streaming data processing method, streaming data processing device and memory medium
WO2016035189A1