A Parallel Relational Table Joining Method and Model

Through the parallelized relation table connection method (POJ method), column storage technology and buffer multiplexing technology, the problems of large storage consumption and large number of threads in parallel processing during multi-field connection of multiple relationship tables are solved, and efficient processing efficiency and video memory utilization are achieved.

CN115374112BActive Publication Date: 2025-05-27TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210930722.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-05-27
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

In the prior art, when multiple relationship tables and multiple fields are connected simultaneously, storage consumption is large in parallel processing, the number of threads is large and the resource is large, making it difficult to improve processing efficiency.

Method used

A parallelized relation table joining method (POJ method) is proposed, which continuously stores possible results of field intersections, so that multiple threads can find computational tasks independently. This method uses column storage technology and buffer multiplexing technology to reduce storage consumption and thread resource usage.

Benefits of technology

The POJ method is implemented on the GPU and can achieve an acceleration effect of more than 22 times. The memory consumption is only 13.8% of the traditional connection method, which significantly improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115374112B_ABST
    Figure CN115374112B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of database technology, and in particular to a parallelized relational table connection method and model, which is used to solve the problems of large storage consumption, large number of threads and high resource consumption in the prior art parallel processing, and difficulty in improving processing efficiency. The present invention improves the efficiency of relational table connection through multi-threading technology. By continuously storing the possible results of the intersection of fields, it is convenient for multiple threads to find computing tasks independently. Compared with the computing speed of the traditional connection method on the GPU, the acceleration effect of the POJ method can reach more than 22 times. The present invention uses column storage technology and buffer reuse technology to try to avoid allocating additional space when the POJ method is running. The space consumption of the POJ method is at least only 13.8% of that of the traditional connection method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of databases, and particularly to a method and model for parallelizing relational table joins. Background Art

[0002] A join is an operation in the Structured Query Language (SQL) that combines multiple relational tables. The join operation can flexibly express various association information between relational tables, and thus is widely used in fields such as business data query and business intelligence. A relational table is composed of tuples with the same fields, and a tuple represents a data record. For example, a relational table representing student information in a class has three fields: student ID, name, and age. Then each tuple in the relational table represents the information of a student, and these information records the student ID, name, and age of the student.

[0003] In the prior art, the traditional join method has the problem of low efficiency, and there is no efficient parallel computing solution for the join method. The biggest challenge in the parallel computing join method lies in the task allocation of threads. In particular, when the number of relational tables to be joined is large, after recursively calculating the fields multiple times, the number of threads required will increase rapidly, and the storage consumed at this time will rise linearly. Even if the thread pool technology is used to reduce the total number of threads required for computing, since the thread pool needs to linearly allocate computing tasks to each thread, this results in the thread pool becoming the performance bottleneck of the method, restricting the join efficiency.

[0004] In summary, for the simultaneous join of multiple fields in multiple relational tables, in the prior art parallel processing, the storage consumption is large, the number of threads is large and the resource occupation is large, and it is difficult to improve the processing efficiency. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method and model for parallelizing relational table joins, which are used to solve the problems in the prior art that when multiple relational tables are joined with multiple fields simultaneously, the storage consumption is large, the number of threads is large and the resource occupation is large in parallel processing, and it is difficult to improve the processing efficiency, and at the same time solve the problem that it is difficult to perform parallel computing for the join method.

[0006] A method for parallelizing relational table joins, the method comprising:

[0007] Input a set of sets of relational tables to be joined;

[0008] Select the relational table R with the fewest number of tuples for each set of relational tables in the set E , and allocate a thread to each tuple in the relational table R E , connect the fields to be joined in the set of relational tables in each thread, obtain a set of sets of joined relational tables, and use the joined set as the input for joining the next field;

[0009] Repeat the above steps until all field connections are completed, and output the set of connected relationship tables.

[0010] Further, input the set of sets of relationship tables to be connected, including: the set of relationship tables to be connected is , where is the first relationship table in the set of relationship tables E 0 ; is the nth relationship table in the set of relationship tables E 0 , where n is a positive integer;

[0011] The order of the fields participating in the connection in the set of relationship tables E 0 is , where A m is the mth field in the set of relationship tables E 0 , where m is a positive integer;

[0012] For each relationship table in the set of relationship tables E 0 , the tuples are sorted in ascending order according to any combination order of the m fields in V and the tuple values.

[0013] Further, for each relationship table in the set of relationship tables E 0 , the tuples are sorted in ascending order according to any combination order of the m fields in V and the tuple values, including:

[0014] The tuples in the relationship table are sorted in the order of the fields in V, and when the fields are the same, they are sorted in ascending order of the tuple values; the order of the fields in V is not specified, and any order can be used;

[0015] When specifically sorting, first sort in ascending order according to the value of field A 1 ; if there are two tuples with the same value of A 1 , sort in ascending order according to the value of A 2 of the two tuples; at this time, if the tuple in the relationship table does not contain A 2 , then sort in ascending order according to the value of A 3 of the two tuples; and so on, to complete the sorting of the tuples in the relationship table.

[0016] Further, connect the fields to be connected in the set of relationship tables in each thread, including: first use the set of relationship tables to be connected {E 0} to connect field A 1 to obtain multiple sets of relationship tables , where represents the second set of relationship tables connected for the first time;

[0017] Use the set of relationship tables to connect field A 2 to obtain multiple sets of relationship tables , and then use the set of relationship tables to connect field A 3 , …, connect each field in turn until all fields are calculated.

[0018] Furthermore, a set of connected relationship table sets is obtained, including: after calculating the connection of all fields, , arbitrarily select one of the relationship table sets , where is the x-th relationship table set of the m-th connection, m is a positive integer, and x is a positive integer;

[0019] In all tuples of, field A 1 has only one value, field A 2 has only one value, …, field A m has only one value.

[0020] Furthermore, assign a thread to each tuple in the relationship table, including: when connecting the k-th field A k , k is a positive integer, assign a thread to each tuple in the relationship table R E , for the i-th tuple in the relationship table R E , i is a positive integer, the value of the field A k of this tuple is a i ; the thread corresponding to the i-th tuple needs to judge: Condition 1: i is equal to 1; Condition 2: a i is not equal to a i-1 , where a i-1 is the value of the field A E of the (i - 1)-th tuple in the relationship table R k , where i is a positive integer; if neither of these two conditions is satisfied, the thread returns None and is recycled.

[0021] Furthermore, the specific process of obtaining the set of connected relationship table sets includes: when connecting the k-th field A k , k is a positive integer, each thread uses the binary search method in to judge whether the relationship table exists a tuple with field A k equal to a i , where E k-1 is a set of relationship tables, is the i-th relationship table in the set of relationship tables E k-1 , a i is the value of field A k in the relationship table ;

[0022] If Each relation table in contains A k Equal to a i tuples, then a relational table set is returned , where if the relationship table Does not contain field A k , where x is a positive integer, then the relation table ,otherwise equal Field A k Equal to a i A relational table consisting of tuples of The xth relational table in the relational table set output by the kth field connection, where x is a positive integer.

[0023] Further, field connection includes: for the xth set of relational tables of the k-1 fields before the connection is completed , k is a positive integer greater than 1, x is a positive integer, Contains field A k The relational table with the least number of tuples is , The number of tuples is recorded as ;

[0024] Remember the xth set of relational tables The prefix sum is ; The prefix sum of the last relational table set is recorded as PrefixSumTotal, and PrefixSumTotal is a positive integer;

[0025] Allocate a sequence of length PrefixSumTotal and PrefixSumTotal threads, with the threads numbered from 1 to PrefixSumTotal;

[0026] For the i-th numbered thread, i is a positive integer, the thread finds the set of relation tables through binary search , so that ; The thread then The number x, and Middle Field A of the tuple k The value of a k,i , written into the i-th element of the sequence;

[0027] The i-th thread needs to determine the corresponding tuple (x, a k,i ) whether: 1) i is equal to ; 2) a k,i Not equal to a k,i-1 , where a k,i Indicates field A in the i-th tuple when connecting the k-th fieldk The value of, where both k and i are positive integers; if neither of these two conditions is satisfied, the thread returns None; if at least one condition is satisfied and a k,i in the relational table both appear, then return the set of relational tables , otherwise return None; among them, if the relational table does not contain the field A k , then the relational table , otherwise the relational table is equal to the relational table composed of the tuples in which the field A k is equal to a k,i .

[0028] A parallel relational table joining model, the model includes:

[0029] An input unit for inputting a set of relational tables to be joined;

[0030] A joining unit for joining the fields of the tuples of the relational tables, selecting the relational table R with the smallest number of tuples in the set of relational tables E , allocating a thread for each tuple in the relational table R E , joining the first field in the set of relational tables in each thread to obtain a joined set of relational tables, and using the joined set of relational tables as the input for joining the next field, recycling the allocated threads, and repeating the above steps until all fields are joined;

[0031] An output unit for outputting the joined set of relational tables.

[0032] Furthermore, the input unit is specifically used for: the set of relational tables to be joined is , where is the first relational table in the set of relational tables E 0 , is the nth relational table in the set of relational tables E 0 , and n is a positive integer;

[0033] The fields participating in the join in the set of relational tables E 0 are sorted as , where A m is the mth field in the set of relational tables E 0 , and m is a positive integer;

[0034] The tuples in each relational table in the set of relational tables E 0 are sorted in ascending order of the tuple values according to any combination order of the m fields in V.

[0035] Furthermore, the set of relational tables E of the input unit 0The tuples in each relational table are sorted according to any combination order of the m fields in V and the tuple value increment sorting subunit, specifically used for: the tuples in the relational table are sorted according to the field order in V, and when the fields are the same, they are sorted in ascending order of the tuple value; the order of the fields in V is not specified, and any order can be used;

[0036] When specifically sorting, first sort in ascending order according to the value of field A 1 ;

[0037] If there are two tuples with the same value of A 1 , sort in ascending order according to the value of A of the two tuples 2 ; at this time, if the tuples in the relational table do not contain A 2 , then sort in ascending order according to the value of A of the two tuples 3 ;

[0038] And so on, complete the sorting of the tuples in the relational table.

[0039] Furthermore, the order subunit for connecting the relational table set of the connection unit is specifically used for: first use the relational table set {E 0} to connect field A 1 to obtain multiple relational table sets , where represents the second relational table set connected for the first time;

[0040] Next, use the relational table set to connect field A 2 to obtain multiple relational table sets , then use the relational table set to connect field A 3 , …, connect each field in turn until all fields are calculated.

[0041] Furthermore, the result subunit for connecting the relational table set of the connection unit is specifically used for: after calculating the connection of all fields, obtain , arbitrarily select one of the relational table sets , where is the xth relational table set connected for the mth time, m is a positive integer, and x is a positive integer;

[0042] Among all the tuples in , field A 1 has only one value, field A 2 has only one value, …, field A m has only one value.

[0043] Furthermore, the subunit for allocating a thread to each tuple in the relational table of the connection unit is specifically used for: when connecting the kth field A kWhen k is a positive integer, it is the relation table R E Each tuple in the relation table R is assigned a thread. E The i-th tuple in the tuple, i is a positive integer, and the field A of the tuple k The value is a i ; The thread corresponding to the i-th tuple needs to judge: Condition 1: i is equal to 1; Condition 2: a i Not equal to a i-1 , where a i-1 For the relation table R E The i-1th tuple field A k The value of , where i is a positive integer; if both conditions are not met, the thread returns None and is recycled.

[0044] Furthermore, the process subunit of the connection unit's relational table set connection is specifically used to: connect the kth field A k When k is a positive integer, each thread Find the judgment relationship table by binary search Does field A exist? k Equal to a i A tuple of E k-1 is a set of relational tables, is the relational table set E k-1 The i-th relationship table in a i For field A k In the relationship table The value of ;

[0045] if Each relation table in contains A k Equal to a i tuples, then a relational table set is returned , where if the relationship table Does not contain field A k , where x is a positive integer, then the relation table ,otherwise equal Field A k Equal to a i A relational table consisting of tuples of The xth relational table in the relational table set output by the kth field connection, where x is a positive integer.

[0046] Further, the connection unit repeats the steps until all field connection sub-units are completed, specifically for: for the xth set of relational tables of the first k-1 fields after the connection is completed , k is a positive integer greater than 1, x is a positive integer, Contains field A k The relational table with the least number of tuples is , The number of tuples of is denoted as

[0047] Denote the prefix sum of the x-th relational table set as ; Denote the prefix sum of the last relational table set as PrefixSumTotal, and PrefixSumTotal is a positive integer;

[0048] Allocate a sequence of length PrefixSumTotal and PrefixSumTotal threads, and the thread numbers are 1 to PrefixSumTotal in sequence;

[0049] For the thread with the i-th number, where i is a positive integer, this thread uses binary search to find the relational table set such that ; Then the thread writes the number x of , and the value a of field A of the -th tuple in k into the i-th element of the sequence; k,i The i-th thread needs to determine whether the corresponding binary tuple (x, a

[0050] k,i k,i k,i-1 k,i k k,i k k k,i k,i k k,i k,i in the relational table satisfies: 1) i is equal to ; 2) a is not equal to a k k,i-1 is equal to k k,i k,i k k,i

[0051] Most of the computational cost of existing connection methods is spent on calculating field intersections and scheduling tasks for these intersections. The parallelized optimal join method (abbreviated as POJ method) proposed by the present invention stores the possible results of field intersections continuously, facilitating multi-threaded independent discovery of computational tasks. Compared with the traditional join method in terms of computational speed on the GPU, the acceleration effect of the POJ method can reach more than 22 times.

[0052] The present invention uses column storage technology and buffer reuse technology to avoid allocating additional space during the operation of the POJ method as much as possible. The minimum space consumption of the POJ method is only 13.8% of that of the traditional join method.

[0053] Other features and advantages of the present invention will be described in the following specification, and, in part, will become apparent from the specification or will be understood by practicing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 It is a flowchart for the POJ method to join m fields.

[0056] Figure 2 It is for the relational table in which each value of field A 1 generates a join result, where None indicates that the join result is empty.

[0057] Figure 3 It is a schematic diagram for removing the rows with the result of None generated by removing field A 1

[0058] Figure 4 It is for each set of relational tables to find within a relational table containing field A k and having the minimum number of tuples schematic diagram.

[0059] Figure 5 It is for each set of relational tables to generate multiple binary tuples, where the binary tuples are composed of the serial number of a set of relational tables and a value of field A k

[0060] ​​Figure 6 Schematic diagram for finding the first prefix sum greater than the thread number in the prefix sum sequence for each thread.

[0061] Figure 7 For the serial number of the set of relationship tables and field A k Calculate field A by taking values k Schematic diagram of the connection.

[0062] Figure 8 For removing field A k Schematic diagram of generating rows with the result of None. Specific implementation manners

[0063] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0064] The time complexity of traditional connection methods is usually several orders of magnitude higher than the space complexity of the connection result. Therefore, it often cannot meet the business requirements when dealing with large-scale data. When multiple relationship tables are connected simultaneously, in the prior art, when multiple fields of multiple relationship tables are connected simultaneously, the storage consumption is large in parallel processing, the number of threads is large and the resource occupation is large, and it is difficult to improve the processing efficiency.

[0065] For this reason, the present invention proposes a parallelized relationship table connection method (abbreviated as POJ method) and model, including a parallelized relationship table connection method and a parallelized relationship table connection model, aiming to solve the problem that it is difficult to perform parallel calculation for relationship table connection, and provides a parallelized relationship table connection method that can perform highly parallel calculation. The speedup ratio of the present invention can reach linearity. Reaching linearity of the speedup ratio means that when providing n threads, the time complexity required by the POJ method can reach 1 / n of the existing connection method; on the GPU, the POJ method can reach up to 22 times the speedup compared with the traditional connection method, and the video memory consumption is only 13.8% of the traditional connection method.

[0066] In addition, the POJ method also provides the following optimizations on the system: 1) Column storage technology, where the relationship table stores the data of the same field continuously; 2) Buffer reuse technology, where the POJ method allocates a buffer once to store field data, and this buffer can be used multiple times by the POJ method.

[0067] In a first aspect, the present invention provides a parallelized relationship table connection method, and the method includes:

[0068] A set of sets of relational tables to be joined;

[0069] Select the relational table R with the fewest number of tuples for each set of relational tables in the set E , for the relational table R E Allocate a thread to each tuple in it, connect the first field in the set of relational tables in each thread, obtain a set of joined relational tables, and use the joined set as the input for connecting the next field, and recycle the allocated threads;

[0070] Repeat the above steps until all fields are connected, and output the set of joined relational tables.

[0071] In a specific embodiment, selecting the relational table R with the fewest number of tuples in the set of relational tables E can reduce the number of allocated threads, save resources, and improve the running efficiency.

[0072] In this embodiment, the set of relational tables to be joined is , where is the first relational table in the set of relational tables E 0 , is the nth relational table in the set of relational tables E 0 , and n is a positive integer;

[0073] The fields participating in the join in the set of relational tables E 0 are sorted as , where A m is the mth field in the set of relational tables E 0 , and m is a positive integer;

[0074] The tuples in each relational table in the set of relational tables E 0 are sorted in the order of the fields in V.

[0075] In a specific embodiment, the unified input format is conducive to iterative loop and multi-threaded parallel operation, improving the processing efficiency.

[0076] In this embodiment, the tuples in the relational table are sorted in ascending order of the value of field A 1 ;

[0077] If there are two tuples with the same value of A 1 , they are sorted in ascending order of the value of A of the two tuples 2 ; at this time, if the tuples in the relational table do not contain A 2 , they are sorted in ascending order of the value of A of the two tuples 3 ;

[0078] And so on, to complete the sorting of the tuples in the relational table.

[0079] In specific embodiments, sorting enables more efficient subsequent binary searches, improving the running efficiency. At the same time, being ordered also means that data read into the video memory can be predicted and read each time, saving video memory space and improving the utilization rate of video memory space.

[0080] In the present invention, by storing the possible results of the field intersections continuously, it is convenient for multiple threads to independently find computing tasks. Compared with the traditional join method in terms of computing speed on the GPU, the acceleration effect of the POJ method can reach more than 22 times.

[0081] In this embodiment, first use the set of relationship tables {E 0} to join field A 1 to obtain multiple sets of relationship tables , where represents the second set of relationship tables for the first join;

[0082] Next, use the set of relationship tables to join field A 2 to obtain multiple sets of relationship tables , then use the set of relationship tables to join field A 3 , …, join each field in turn until all fields are calculated.

[0083] In this embodiment, after calculating all field joins, is obtained. Arbitrarily select one of the sets of relationship tables , where is the x-th set of relationship tables for the m-th join, m is a positive integer, and x is a positive integer;

[0084] In for all tuples, field A 1 has only one value, field A 2 has only one value, …, field A m has only one value.

[0085] In specific embodiments, the multi-field joins of multiple relationship tables form a loop similar to field joins, and each node field in the loop has only one value.

[0086] In this embodiment, when joining the k-th field A k , a thread is assigned to each tuple in the relationship table R E . For the i-th tuple in the relationship table R E , i is a positive integer, and the value of the field A k of this tuple is a i ; the thread corresponding to the i-th tuple needs to judge: Condition 1: i is equal to 1; Condition 2: a i is not equal to a i-1; where a i-1 is the (i - 1)-th tuple field A E in the relationship table R k where i is a positive integer; if neither of these two conditions is met, the thread returns None and is recycled.

[0087] In a specific embodiment, to adapt to the multi-threaded operation of the GPU, threads need to be pre-allocated. To improve the thread operation efficiency, duplicate threads are recycled to avoid duplication.

[0088] In this embodiment, when connecting the k-th field A k where k is a positive integer, each thread uses binary search to determine whether the relationship table contains a tuple with field A k equal to a i where E k-1 is a set of relationship tables, is the i-th relationship table in the set of relationship tables E k-1 and a i is the value of field A k in the relationship table ;

[0089] If each relationship table in k contains a tuple with A i equal to a , a set of relationship tables is returned, where if the relationship table k does not contain field A , where x is a positive integer, then the relationship table is equal to the relationship table composed of tuples with field A k equal to a i in , where

[0090] is the x-th relationship table in the set of relationship tables output by connecting the k-th field, and x is a positive integer.

[0091] In this embodiment, for the x-th set of relationship tables after connecting the first k - 1 fields , where k is a positive integer greater than 1 and x is a positive integer, the relationship table in k that contains field A and has the fewest number of tuples is , and the number of tuples of

[0092] is denoted as The prefix sum of ; Denote the prefix sum of the last set of relational tables as PrefixSumTotal, and PrefixSumTotal is a positive integer;

[0093] Allocate a sequence of length PrefixSumTotal and PrefixSumTotal threads, and the thread numbers are 1 to PrefixSumTotal in sequence;

[0094] For the thread with the i-th number, where i is a positive integer, this thread uses binary search to find the set of relational tables such that ; Then the thread writes the number x of , and the value a of field A in the -th tuple in k into the i-th element of the sequence; k,i

[0095] The i-th thread needs to determine whether the corresponding binary tuple (x, a k,i ) satisfies: 1) i is equal to ; 2) a k,i is not equal to a k,i-1 , where a k,i represents the value of field A in the i-th binary tuple when connecting the k-th field, and both k and i are positive integers; if neither of these two conditions is satisfied, this thread returns None; if at least one condition is satisfied and a k appears in all the relational tables k,i , then return the set of relational tables , otherwise return None; among them, if the relational table does not contain field A , then the relational table k , otherwise the relational table is equal to the relational table composed of the tuples in which field A is equal to a k ; k,i

[0096] In a second aspect, the present invention provides a parallel relational table connection model, and the model includes:

[0097] An input unit for inputting a set of relational tables to be connected;

[0098] A connection unit for connecting the fields of the tuples of the relational tables, and selecting the relational table R E with the fewest number of tuples in the set of relational tables, and for the relational table R EAllocate a thread to each tuple, and in each thread, concatenate the first field in the set of relational tables to obtain a set of concatenated relational tables, and use the set of concatenated relational tables as the input for concatenating the next field. Recycle the allocated threads, and repeat the above steps until all fields are concatenated;

[0099] An output unit for outputting the set of concatenated relational tables.

[0100] In this embodiment, the set of relational tables to be concatenated is , where is the first relational table in the set of relational tables E 0 , is the nth relational table in the set of relational tables E 0 , and n is a positive integer;

[0101] The set of relational tables E 0 The fields participating in the concatenation are sorted as , where A m is the mth field in the set of relational tables E 0 , and m is a positive integer;

[0102] The tuples in each relational table in the set of relational tables E 0 are sorted in the order of the fields in V.

[0103] In this embodiment, the tuples in the relational table are sorted in ascending order of the values of field A 1 ;

[0104] If there are two tuples with the same value of A 1 , they are sorted in ascending order of the values of A of the two tuples 2 ; at this time, if the tuples in the relational table do not contain A 2 , they are sorted in ascending order of the values of A of the two tuples 3 ;

[0105] And so on, to complete the sorting of the tuples in the relational table.

[0106] In this embodiment, first use the set of relational tables to be concatenated {E 0} to concatenate field A 1 to obtain multiple sets of relational tables , where represents the second set of relational tables after the first concatenation;

[0107] Next, use the set of relational tables to concatenate field A 2 to obtain multiple sets of relational tables , and then use the set of relational tables to concatenate field A 3, …, connect each field in turn until all fields are calculated.

[0108] In this embodiment, after calculating all field connections, we get , arbitrarily select one of the relationship table sets ,in is the xth set of relational tables connected for the mth time, where m is a positive integer and x is a positive integer;

[0109] exist In all tuples, field A 1 There is only one value, field A 2 There is only one value, ..., field A m There is only one value.

[0110] In this embodiment, when connecting the kth field A k When k is a positive integer, it is the relation table R E Each tuple in the relation table R is assigned a thread. E The i-th tuple in the tuple, i is a positive integer, and the field A of the tuple k The value is a i ; The thread corresponding to the i-th tuple needs to judge: Condition 1: i is equal to 1; Condition 2: a i Not equal to a i-1 , where a i-1 For the relation table R E The i-1th tuple field A k The value of , where i is a positive integer; if both conditions are not met, the thread returns None and is recycled.

[0111] In this embodiment, when connecting the kth field A k When k is a positive integer, each thread Determine the relationship table through binary search Does field A exist? k Equal to a i A tuple of E k-1 is a set of relational tables, is the relational table set E k-1 The i-th relationship table in A k For the relationship table The fields in a i For field A k In the relationship table The value of ;

[0112] if Each relation table in contains A k Equal to a i tuples, then a relational table set is returned , where if the relationship table Does not contain field A k , where x is a positive integer, then the relational table , otherwise equals field A in k equals a i The relational table composed of tuples, where is the x-th relational table in the set of relational tables output by connecting the k-th field, and x is a positive integer.

[0113] In this embodiment, for the x-th set of relational tables after connecting the first k - 1 fields , k is a positive integer greater than 1, and x is a positive integer, contains field A k and the relational table with the fewest number of tuples is , The number of tuples of is denoted as ;

[0114] Denote the prefix sum of the x-th set of relational tables as ; Denote the prefix sum of the last set of relational tables as PrefixSumTotal, and PrefixSumTotal is a positive integer;

[0115] Allocate a sequence of length PrefixSumTotal and PrefixSumTotal threads, and the thread numbers are 1 to PrefixSumTotal in sequence;

[0116] For the thread with the i-th number, where i is a positive integer, this thread uses binary search to find the set of relational tables such that ; The thread then writes the number x of, and the field A of the -th tuple in k value a k,i into the i-th element of the sequence;

[0117] The i-th thread needs to determine whether the corresponding binary tuple (x, a k,i ) satisfies: 1) i equals ; 2) a k,i is not equal to a k,i-1 , where a k,i represents the value of field A in the i-th binary tuple when connecting the k-th field k , and both k and i are positive integers; if neither of these two conditions is satisfied, then this thread returns None; if at least one condition is satisfied, and a k,i is in the relational table If it has appeared in all, then the relationship table set is returned , otherwise returns None; if the relationship table Does not contain field A k , then the relationship table , otherwise the relationship table equal Field A k Equal to a k,i A relational table consisting of tuples.

[0118] In a specific embodiment, the kth field is connected in an iterative loop until the connection of all fields is completed.

[0119] In the thread allocation process, the prefix sum is used to calculate the number of threads to be allocated, and two conditions are used to remove duplicate threads, thereby reducing thread resource waste and improving resource utilization.

[0120] The present invention uses column storage technology and buffer multiplexing technology to avoid allocating extra space when the POJ method is running. The space consumption of the POJ method is only 13.8% of the traditional connection method.

[0121] In order to enable those skilled in the art to better understand the present invention, the principle of the present invention is described as follows in conjunction with the accompanying drawings:

[0122] The present invention proposes a parallelized relational table connection method, namely the POJ method. The main calculation process of the POJ method is as follows: Figure 1 As shown, connect field A 1 Then connect field A 2 , until the connection field A m Then output the result.

[0123] For a query Q containing a connection, the set of relational tables it connects is recorded as , the ordered set of connected fields is The POJ algorithm first sorts the tuples in each relational table in ascending order according to the order of the field set V. The higher the tuple is, the smaller the value of its field is. , relationship table Contains fields ,So The tuple in the first press A 1 The values ​​are sorted in ascending order. If there are two tuples of A 1 If the values ​​are the same, continue to compare the two tuples A 3 The value size (because The tuple in does not contain field A 2 , so A is not considered 2). The order of fields in V is not specified, and any order can be used. The POJ algorithm splits Q into multiple "field joins" and calculates these joins in sequence, as Figure 1 shown. Each "field join" connects only one field, which takes a set of sets of relational tables as input and calculates to obtain another set of sets of relational tables. For example Figure 1 the join of field A 1 will take the set {E 0} as input and generate another set through the parallel join method. The parallel join method is explained in the specific implementation part. Then, field A 2 takes as input and calculates to obtain . The calculations for other fields are carried out in the same way. Any set of relational tables contains n relational tables , where contains the same types of fields as . If the set of relational tables is obtained after calculating the k-th field A k , then the POJ algorithm ensures that for all tuples in this , their field A 1 (if there is field A 1 , the same below) has only one value, field A 2 has only one value,..., field A k has only one value. The brief working method of the POJ algorithm includes the following steps:

[0124] S1. The POJ algorithm first uses {E 0} to calculate the join of field A 1 and obtains . Next, it uses to calculate the join of field A 2 and obtains . Then, it uses to calculate field A 3 ,..., until all fields are calculated.

[0125] S2. After calculating all field joins, is obtained. Arbitrarily select one set of relational tables . Among all the tuples in , field A 1 has only one value, field A 2 has only one value,..., field A mThere is only one value, and the only combination of values of these fields is the tuple that finally meets the join condition. Therefore, each set of relational tables corresponds to a tuple that meets the join condition, and all these tuples form the calculation result of join Q.

[0126] The present invention optimizes parallel computing and solves the problem of low efficiency of parallel computing joins in large-scale data scenarios. The POJ method serializes the join of each field and parallelizes the join of a single field, thus avoiding the problem of complex thread task scheduling in a multi-threaded environment for the optimal join method.

[0127] More specifically, to implement the POJ algorithm on a GPU, the working method of the system includes the following steps:

[0128] S1. Determine that the fields participating in the join are sorted as , and all relational tables are , and each relational table sorts its tuples in the order of the fields in V.

[0129] S2. The main calculation process is as follows: The POJ algorithm joins each field in turn, and the join of these fields takes a set of sets of relational tables as input and outputs another set of sets of relational tables, which is used as the input for joining the next field. The join of a single field can be parallelized. The process of parallelizing the join of a single field is as follows:

[0130] S2.1. Prepare to calculate the first field A 1 , that is, at this time, the POJ algorithm has not calculated other fields. Field A 1 takes {E 0} as input, and the following explains how to obtain . First, the POJ algorithm finds the relational table with the fewest number of tuples in the set . The POJ algorithm assigns a thread to each tuple in , and a total of threads are assigned. Taking the i-th tuple in as an example, the value of A of this tuple is a 1 . The thread corresponding to the i-th tuple needs to judge: 1) i equals 1; 2) a i is not equal to a i . If neither of these two conditions is met, the thread returns None. If one of the conditions is met, the thread uses binary search in i-1 (because the relational tables are sorted by V, and A is the first field, so the values of A 1 in all relational tables containing A 1 are monotonically increasing) to judge whether there exists a field A 1 equal to a 1 equal to ai tuples. If each relational table in 1 equals a i tuple, then return a set of relational tables , where if does not contain the field A 1 , then , otherwise equals the relational table composed of tuples where the field A 1 equals a i in. If there exists a relational table containing A 1 in, and there is no tuple in this relational table where A 1 equals a i , then this thread returns None. For example, the first thread returns , the second thread fails to meet the above conditions and returns None, the third thread returns , the fourth thread returns , …, as Figure 2 shown. Figure 2 In, the left column data is all the values of the field A in the relational table 1 , so it satisfies , Figure 2 the right side data is obtained according to the above process of S2.1. Next, we will remove the results that return None and store the set of relational tables in a continuous sequence, as Figure 3 shown. Figure 3 The leftmost column data is the result obtained from Figure 2 , and the rightmost column data is used to participate in the join of the field A 2 . First, allocate an array Eum of length . If the previous i-th thread returns None, then Eum[i] = 0, otherwise Eum[i] = 1. Calculate the prefix sum of Eum. The prefix sum of an array equals , for example, as Figure 3 shown, the second column data is filled with 0 or 1 according to whether the first column data is None, and the third column data is obtained by calculating the prefix sum of the second column data. The prefix sum can be calculated in parallel using the Scan algorithm. Next, if the i-th thread finds that Eum[i] = 1, then is stored in a new sequence according to the i-th prefix sum, for example, as Figure 3 shown in the fourth column. The first row data in the leftmost data corresponds to Eum[1] = 1, and the corresponding prefix sum is 1, so Put it in the first position of the new sequence. For the data in the second row, Eum[2]=0, so the result of the second row is not involved in subsequent calculations. For the data in the third row corresponds to Eum[1]=1, and the corresponding prefix sum is 2. Therefore, put it in the second position of the new sequence. By analogy for the data in other rows, finally re-number the relationship table sets in this sequence to obtain , for example Figure 3 as shown on the far right

[0131] S2.2. Assume that the calculations for the first k - 1 fields (k≥2) have been completed and the result has been obtained. Next, connect field A k . Each set of relationship tables has n relationship tables , and contains the same fields as . The tuples in are still sorted in the field order of V, and each field in has only one value in all tuples (that is, A 1 has only one value, A 2 has only one value, and so on). Use to represent: the relationship table that contains field A and has the fewest tuples k .

[0132] S2.3. For any set of relationship tables , the relationship table can be found. Therefore, a thread can be assigned to each to find the corresponding for . As Figure 4 shown, each set of relationship tables can find a relationship table within that contains field A k and has the fewest tuples . Put all the tuples' field A values and the numbers of k in into a binary structure and store them continuously in a sequence. For example Figure 5 shown, in the binary tuple sequence on the right side of Figure 5 , the left column data indicates which set of relationship tables the binary tuple data comes from (i.e., the relationship table the arrow comes from), and the right data is the value of field A k . For example, the first binary tuple (1, a k,1 ) in the right list means that this binary tuple is from the set of relationship tables Generated, a k,1 Is The first tuple field A k Takes the value, the second tuple in the right - hand list (1, a k,2 ) indicates that this tuple is from the set of relational tables Generated, a k,2 Is The second tuple field A k Takes the value, the third tuple in the right - hand list (2, a k,3 ) indicates that this tuple is from the set of relational tables Generated, a k,3 Is The first tuple field A k Takes the value, and so on. The steps to determine the values of each tuple in this sequence are as follows:

[0133] S.2.3.1. For the x - th set of relational tables , the relational table with the least number of tuples that contains A k and is corresponding to it is , and the number of its tuples is denoted as . Next, the prefix sum of these tuple numbers needs to be calculated. Then the prefix sum of the x - th set of relational tables is . In particular, denote the prefix sum of the last set of relational tables as PrefixSumTotal.

[0134] S.2.3.2. Allocate a sequence of length PrefixSumTotal, allocate PrefixSumTotal threads, and the numbers of these threads are 1~PrefixSumTotal in sequence. Each thread corresponds to an element with the same number in the sequence. Suppose the number of a certain thread is i, then this thread finds through binary search such that , as shown in Figure 6 . The thread then writes the number x of , and the value a of A in the - th tuple in k into the i - th element in the sequence, that is, the result as shown in k,i is obtained. Figure 5

[0135] S2.4. The POJ algorithm allocates a thread for each tuple in the sequence, and the number of the thread is equal to the serial number of the tuple. The i - th thread needs to determine whether the corresponding tuple (x, a k,i ) satisfies: 1) i is equal to ; 2) a k,i is not equal to a​k,i-1 (a k,i represents the value of A stored in the i-th pair in the sequence of pairs k value). If neither of these two conditions is satisfied, the thread returns None, i.e., stops the calculation. If one of the conditions is satisfied, and a k,i appears in both the E k relationship tables, a new set of relationship tables is returned , where, if does not contain the field A k , then , otherwise is equal to the relationship table composed of the tuples in which the field A k is equal to a k,i . Figure 7 shows the corresponding relationship for generating the set of relationship tables E k from the sequence of pairs, where None indicates that the corresponding pair does not produce a join result.

[0136] S2.5, as Figure 7 shown, the sets of relationship tables generated by the pairs in S2.4 are not stored continuously. Next, we will store these sets of relationship tables continuously. First, allocate an array Eum of length PrefixSumTotal. If the result returned by the i-th pair is not None, then Eum[i] is equal to 1, otherwise equal to 0, (for example Figure 7 the corresponding array is Eum = [1, 0, 1, 1, 0, 1,...]), and then use the Scan algorithm to calculate the prefix sum EumSum of this array in parallel ( Figure 7 the corresponding prefix sum is [1, 1, 2, 3, 3, 4,...]). Then allocate PrefixSumTotal threads. If the i-th thread finds that Eum[i] is equal to 1, it places the set of relationship tables generated by the i-th pair at the position indicated by EumSum[i]. At this time, the sets of relationship tables returned by S2.4 are stored continuously in the new sequence, and these sets are re-numbered as . The process is as Figure 8 shown. At this time, any meets the assumption in S2.1, so it can continue to be used for the join of the field A k+1 . Next, verify that E k meets the assumption in S2.1. Obviously, the relationship tables k in E contain the same fields as , so also contain the same fields as . And the tuples in all come from the A in ktuples with the same value, so it can ensure that the fields in each have only one value in all tuples. Therefore, E k also meets the assumptions of S2.1.

[0137] S3. After calculating the join of m fields, a set of sets of relational tables is obtained . For any set of relational tables among them, in all the tuples it contains, field A 1 has only one value, field A 2 has only one value,..., field A m has only one value. Therefore, the unique value combinations of these fields are the tuples that finally meet the join conditions.

[0138] Finally, there is one more optimization. Notice that in S2.5, if the tuples in E k-1 have only one value in these fields, and the relational table is sorted in ascending order by V, then if contains field A k , then all the tuples in it are monotonically increasing in the value of field A k . Then the tuples with the value of A k equal to a k,i in must also be continuously arranged, and these tuples are all the tuples in . Therefore, can use the starting number and ending number of the tuples in which A k is equal to a k,i to represent, thus avoiding copying the tuples in and reducing the calculation amount.

[0139] The present invention designs an embodiment for verifying the effect of the POJ method, and explains the principle and actual use effect of the present invention according to the embodiment.

[0140] In the embodiment, the operating system version used by the machine is Ubuntu 20.04.3 LTS, the CPU model is Intel Xeson Silver 4210 CPU @ 2.20GHz, the memory model is Micron 94G 2666MT / S, the disk model is HGST HUS726T4TAL W41G, the GPU model is Nvidia RTX 2080ti, the GPU video memory is 11G, and the interface model is a PCIe interface.

[0141] In the embodiments, the performance of three methods, namely POJ, LFTJ-GPU, and PG-Strom, is compared through experiments. PG-Strom is a GPU extension plugin running on PostgreSQL, which supports performing native joins on the GPU and is a mainstream GPU computing solution for PostgreSQL. The LFTJ-GPU method is an attempt at parallelized computing for the optimal join method, but its speedup does not reach linearity. To avoid excessive disk I / O overhead caused by multiple disk accesses, PostgreSQL sets up a buffer to store tuples. To reduce the number of disk accesses, 30G of shared memory is set for PostgreSQL in the experiment, and the calculation is preheated multiple times so that as many tuples as possible can reside in memory. The cost estimation of PG-Strom is relatively conservative. Therefore, when PostgreSQL selects an execution plan, it may choose an execution plan that runs on the CPU. It is necessary to confirm that the selected execution plan runs on the GPU before testing. Before the LFTJ-GPU calculation, the relational table needs to be converted into a tree heap and stored in the video memory, and these preprocessing steps are all included in the running time of LFTJ-GPU.

[0142] The test cases used in the embodiments are triangle joins , and quadrilateral joins , where R 1 -R 4 is a relational table, and A 1 -A 4 is a field in the relational table. These two test cases are commonly used test cases in the field of relational algebra. The data values in each relational table are uniformly distributed positive integers, with the minimum value being 1 and the maximum value being such that the number of join results is the largest.

[0143] Table 1: Time taken (in seconds) for the POJ method and other join algorithms to calculate the join query Q 3 Time taken (unit: seconds)

[0144]

[0145] Table 2: Time taken (in seconds) for the POJ method and other join algorithms to calculate the join query Q 4 Time taken (unit: seconds)

[0146]

[0147] S1. First, compare the computational performance of the PG-Strom, LFTJ-GPU, and POJ methods. Table 1 shows the time consumption of the PG-Strom, LFTJ-GPU, and POJ methods in computing datasets of different lengths for triangle joins, and Table 2 shows the time consumption of the PG-Strom, LFTJ-GPU, and POJ methods in computing datasets of different lengths for quadrilateral joins. In Q 3 terms of time consumption. The running efficiency of the POJ method has always been much higher than that of PG-Strom. Especially when the number of tuples exceeds 300K, the running efficiency of the POJ method is at least 55% higher than that of PG-Strom. When the number of tuples is 30M, PG-Strom runs for more than 10 hours without obtaining a result and is judged as timed out. Although the time complexity of the LFTJ-GPU method does not reach the optimal, the gap with the POJ method is not large, and the time consumption of the two is basically the same. When the number of tuples reaches 30M, the time consumed by the LFTJ-GPU method is 9.424 seconds less than that of the POJ method. This is because the LFTJ-GPU method stores all data in the video memory, reducing the data exchange overhead between the memory and the video memory. This gap also exists in Q 4 . When the data scale does not exceed 3M, the POJ method has no obvious performance advantage in Q 4 . However, when the number of tuples exceeds 3M, the running efficiency of the POJ method quickly overtakes PG-Strom, and the calculation speed increases by about 65% in the case of 30M tuples.

[0148] Table 3: Video memory consumption (unit: megabyte) of the POJ method and other join algorithms for computing the join query Q 3 Video memory consumption (unit: megabyte)

[0149]

[0150] Table 4: Video memory consumption (unit: megabyte) of the POJ method and other join algorithms for computing the join query Q 4 Video memory consumption (unit: megabyte)

[0151]

[0152] S2. Table 3 shows the video memory consumption of the PG-Strom, LFTJ-GPU method, and POJ method in calculating the triangle join for datasets of different lengths. Table 4 shows the video memory consumption of the PG-Strom, LFTJ-GPU method, and POJ method in calculating the quadrilateral join for datasets of different lengths. As the number of tuples grows, the video memory required for the PG-Strom method calculation increases from 153 MB to 6417 MB. Under normal circumstances, the resident memory of PG-Strom in the GPU is 151 MB, so the actual memory consumption ranges from 2 MB to 6266 MB, a 3133-fold increase. Judging from the growth rate of the video memory consumed by PG-Strom, PG-Strom lacks the ability to further improve the computing throughput. The video memory consumption of LFTJ-GPU is between that of PG-Strom and the POJ method. This is because the LFTJ-GPU method stores all the data of the relational table in the video memory. This results in LFTJ-GPU consuming 124% more video memory than the POJ method in the 30M tuple test. In contrast, the space complexity of the POJ method is much better. In the experiment, the POJ method requires at most 887 MB of video memory. Compared with the minimum required 196 MB of video memory, the growth rate of video memory consumption does not exceed 4 times. Overall, Q 4 relatively Q 3 will consume more memory. PG-Strom consumes at most more than 65% additional video memory. Due to the need to store the entire relational table, the video memory consumption of LFTJ-GPU also increases by more than 50%. The increase in the video memory consumed by the POJ method is only 9.9%, which indicates that the POJ method has a stronger ability to process large-scale data and can join more tuples simultaneously.

[0153] Next, combined with the complexity analysis of the POJ algorithm, the improvement of the efficiency of the present invention will be explained.

[0154] Assume that the set of relational tables involved in the join Q to be calculated is E, with a total of n relational tables, and the set of fields is V, with a total of m fields. In our subsequent proof, we use E i to represent all relational tables containing the first i fields, to represent all relational tables containing the field v i . According to the AGM theorem, the upper bound of the result quantity of Q is:

[0155]

[0156] where x i is called the fractional edge cover coefficient. Each relational table corresponds to a fractional edge cover coefficient and satisfies the condition:

[0157] By finding the minimum value on the right side of formula (1), we can obtain the upper limit of the result number of Q.

[0158] The time complexity of the optimal join algorithm is:

[0159] where, The symbol indicates that the possible log-term coefficients are ignored (because the search method may be binary search or hash search). Assume that the maximum number of threads used by the algorithm is p.

[0160] (1) When joining the first field A 1 at this time

[0161] When joining the first field, there is only one original range tuple S in the range table T 0 The shortest relationship table of S is denoted as R s . Then the complexity of calculating the join of the first field is:

[0162]

[0163] Because for each tuple in R s we can separately allocate a thread to find whether the field value exists in all relationship tables. Therefore, under the condition of parallel computing, the time complexity of joining the first field is:

[0164]

[0165] Compared with the optimal join algorithm using binary search, the speedup ratio of the POJ algorithm reaches linearity.

[0166] (2) When joining the first k fields

[0167] Let's assume that the range table generated by joining the first k - 1 fields is T k-1 , and the time complexity of joining the first k - 1 fields is:

[0168]

[0169] When joining the kth field, the number of field values to be judged is:

[0170]

[0171] where, S represents the range tuple in T k-1 , and R s represents the tuple in the shortest range in S. Then the time complexity of joining the kth field is:

[0172]

[0173] Also because R sEach tuple in may generate a new range tuple, so The maximum value of , which is the upper limit of the number of results calculated by concatenating the first k fields. Therefore, the time complexity of calculating the first k fields is:

[0174]

[0175] After all m fields are concatenated, that is, when k = m, the algorithm complexity of POJ can be obtained as:

[0176]

[0177] Compared with the time complexity of the OJ algorithm, the POJ algorithm only requires 1 / p of the OJ algorithm's time, and the speedup ratio reaches a linear effect, fully demonstrating the advantages of parallel computing.

[0178] Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for parallelizing relational table joins, characterized in that, the method includes: inputting a set of relational tables to be joined; Select the relation table R with the fewest number of tuples for each set of relation tables in the set E , for the relation table R E Allocate a thread to each tuple in, and in each thread, connect the fields to be joined in the set of relation tables to obtain a set of sets of joined relation tables, and use the set of sets of joined relation tables as the input for joining the next field; repeating the above steps until all fields are joined, and outputting a set of joined relational tables; Field connection, including: for the x-th set of relational tables that connect the first k - 1 fields before the connection is completed k is a positive integer greater than 1, and x is a positive integer contains field A k and the relational table with the fewest number of tuples is the number of tuples of is denoted as Denote the x-th set of relationship tables The prefix sum of Denote the prefix sum of the last set of relationship tables as PrefixSumTotal, and PrefixSumTotal is a positive integer; allocating a sequence of length PrefixSumTotal and PrefixSumTotal threads, with thread numbers sequentially from 1 to PrefixSumTotal; For the thread numbered i, where i is a positive integer, the thread finds the set of relationship tables through binary search such that The thread then the number x of and the value a of field A of the th tuple in k are written into the i-th element of the sequence; k,i ​ The i-th thread needs to determine the corresponding tuple (x, a k,i ) whether: 1) i is equal to 2)a k,i Not equal to a k,i-1 , where a k,i Indicates field A in the i-th tuple when connecting the k-th field k The value of k and i are both positive integers; if both conditions are not met, the thread returns None; if at least one condition is met, and a k,i In the relationship table If it has appeared in all, then the relationship table set is returned Otherwise, None is returned; if the relationship table Does not contain field A k , then the relationship table Otherwise the relationship table equal Field A k Equal to a k,i A relational table consisting of tuples.

2. The method for parallelizing relational table joins according to claim 1, characterized in that, A set of sets of relational tables to be joined, including: the set of relational tables to be joined is where is the first relational table in the set of relational tables E 0 in is the nth relational table in the set of relational tables E 0 where n is a positive integer; The set of relationship tables E 0 The sorting of the fields participating in the connection in 1 is V = {A 2 , A m}, where A m is the m-th field in the set of relationship tables E 0 , and m is a positive integer; Set E of relationship tables 0 In each relationship table in it, the tuples are sorted in the field order of V for the tuples.

3. The method for parallelizing relational table joins according to claim 2, characterized in that, Set E of relationship tables 0 The tuples in each relationship table in are sorted in the order of the fields in V, including: The tuples in the said relationship table are sorted in ascending order according to the value of field A 1 ; If the values of field A of two tuples 1 are the same, sort them in ascending order according to the values of field A of the two tuples 2 ; at this time, if the tuples in the relational table do not contain field A 2 , then sort them in ascending order according to the values of field A of the two tuples 3 ; and so on, to complete the sorting of tuples in the relational table.

4. The method for parallelizing relational table joins according to claim 2, characterized in that, Connect the fields to be joined in the set of connection relation tables in each thread, including: First, use the set of relation tables to be joined {E 0} to connect field A 1 to obtain multiple sets of relation tables where represents the second set of relation tables for the first connection; Using a set of relationship tables Connect field A 2 Obtain multiple sets of relationship tables Then, use the set of relationship tables again Connect field A 3 , …, connect each field in turn until all fields are calculated.

5. The method for parallelizing relational table joins according to claim 4, characterized in that, Obtain a set of connected relationship table sets, including: after calculating all field connections, obtain Arbitrarily select one of the relationship table sets Among them The xth relationship table set of the mth connection, where m is a positive integer and x is a positive integer; Among all tuples of 1 field A has only one value, field A 2 has only one value,..., field A m has only one value.

6. The method for parallelizing relational table joins according to claim 1, characterized in that, Assign a thread to each tuple in the relational table, including: k When k is a positive integer, it is the relation table R E Each tuple in the relation table R is assigned a thread. E The i-th tuple in the tuple, i is a positive integer, and the field A of the tuple k The value is a i ; The thread corresponding to the i-th tuple needs to judge: Condition 1: i is equal to 1; Condition 2: a i Not equal to a i-1 , where a i-1 For the relation table R E The i-1th tuple field A k The value of , where i is a positive integer; if both conditions are not met, the thread returns None and is recycled.

7. The method for parallelizing relational table joins according to claim 1, characterized in that, The specific process of obtaining the set of sets of connected relational tables includes: when connecting the k-th field A k , where k is a positive integer, each thread uses binary search to determine whether the relational table contains a tuple with field A k equal to a i , where E k-1 is the set of relational tables, is the i-th relational table in the set of relational tables E k-1 , a i is the value of field A k in the relational table ; If each relational table in k contains A i equal to a then return a set of relational tables where if the relational table k does not contain the field A otherwise is equal to the relational table composed of tuples where the field A k is equal to a i where is the x-th relational table in the set of relational tables output by concatenating the k-th field, and x is a positive integer.

8. A device for parallelizing relational table joins, characterized in that, it includes: an input unit for inputting a set of relational tables to be joined; A connection unit for connecting fields of tuples in a relational table, and selecting the relational table R with the smallest number of tuples in the set of relational tables E , for the relational table R E allocating a thread to each tuple in it, connecting the first field in the set of relational tables in each thread to obtain a set of connected relational table sets, and using the set of connected relational table sets as the input for connecting the next field, recycling the allocated threads, and repeating the above steps until all field connections are completed; an output unit for outputting a set of joined relational tables; Field connection, including: for the x-th set of relational tables that connect the first k - 1 fields before connection is completed k is a positive integer greater than 1, and x is a positive integer contains field A k and the relational table with the fewest number of tuples is The number of tuples is denoted as Denote the set of the x-th relationship tables The prefix sum of Denote the prefix sum of the last set of relationship tables as PrefixSumTotal, and PrefixSumTotal is a positive integer; allocating a sequence of length PrefixSumTotal and PrefixSumTotal threads, with thread numbers sequentially from 1 to PrefixSumTotal; For the thread numbered i, where i is a positive integer, the thread finds the set of relationship tables through binary search such that The thread then the number x of and the value a of field A of the k th tuple in k,i are written into the i-th element of the sequence; The i-th thread needs to determine whether the corresponding binary tuple (x, a k,i ) satisfies: 1) i is equal to 2) a k,i is not equal to a k,i-1 , where a k,i represents the value of field A in the i-th binary tuple when connecting the k-th field, and both k and i are positive integers; if neither of these two conditions is satisfied, the thread returns None; if at least one condition is satisfied and a k appears in the relation table k,i , then return the relation table set otherwise return None; among them, if the relation table does not contain field A , then the relation table k otherwise the relation table is equal to the relation table composed of the tuples where field A in k is equal to a k,i .

9. The device for parallelizing relational table joins according to claim 8, characterized in that, An input unit, specifically used for: the set of relationship tables to be connected is where is the first relationship table in the set of relationship tables E 0 and is the nth relationship table in the set of relationship tables E 0 where n is a positive integer; Set E of relationship tables 0 The fields participating in the connection in are sorted as V = {A 1 , A 2 , …, A m}, where A m is the m-th field in the set E of relationship tables 0 , and m is a positive integer; Set E of relationship tables 0 The tuples in each relationship table in it are sorted in the order of the fields in V.

10. The device for parallelizing relational table joins according to claim 9, characterized in that, Set E of relationship tables of the input unit 0 In each relationship table, the tuples in the tuple sub-unit are sorted in the field order of V, specifically for: the tuples in the relationship table are sorted in ascending order according to the value of field A 1 ; If there are two tuples with the same value in field A 1 , sort them in ascending order according to the value of field A of the two tuples 2 . At this time, if the tuples in the relation table do not contain field A 2 , then sort them in ascending order according to the value of field A of the two tuples 3 . and so on, to complete the sorting of tuples in the relational table.

11. The device for parallelizing relational table joins according to claim 9, characterized in that, The relational table set of the connection unit connects the sequential sub-units, specifically used for: first using the relational table set {E to be connected 0} to connect field A 1 to obtain multiple relational table sets where represents the second relational table set of the first connection; Next, use the relationship table set again to connect field A 2 to obtain multiple relationship table sets Then, use the relationship table set again to connect field A 3 , …, connect each field in turn until all fields are calculated.

12. The device for parallelizing relational table joins according to claim 11, characterized in that, The result subunit connected by the relationship table set of the connection unit is specifically used for: after calculating all field connections, obtain Arbitrarily select one of the relationship table sets Among them The xth relationship table set of the mth connection, where m is a positive integer and x is a positive integer; Among all tuples of 1 field A has only one value, field A 2 has only one value,..., field A m has only one value.

13. The device for parallelizing relational table joins according to claim 8, characterized in that, The connection unit allocates a thread subunit to each tuple in the relational table, which is specifically used to: connect the kth field A k When k is a positive integer, it is the relation table R E Each tuple in the relation table R is assigned a thread. E The i-th tuple in the tuple, i is a positive integer, and the field A of the tuple k The value is a i ; The thread corresponding to the i-th tuple needs to judge: Condition 1: i is equal to 1; Condition 2: a i Not equal to a i-1 , where a i-1 For the relation table R E The i-1th tuple field A k The value of , where i is a positive integer; if both conditions are not met, the thread returns None and is recycled.

14. The device for parallelizing relational table joins according to claim 8, characterized in that, The relational table set of the connection unit connects the process sub-units, specifically used for: when connecting the k-th field A k where k is a positive integer, each thread uses binary search to determine whether the relational table contains a tuple where the field A k is equal to a i where E k-1 is the relational table set, is the i-th relational table in the relational table set E k-1 and a i is the value of the field A k in the relational table ; If each relational table in k contains a tuple where i is equal to a then return a set of relational tables where if the relational table k does not contain the field A otherwise is equal to the relational table composed of tuples where the field A k is equal to a i where is the x-th relational table in the set of relational tables output by concatenating the k-th field, and x is a positive integer.

15. The device for parallelizing relational table joins according to any one of claims 8 - 14, characterized in that, Repeat the steps of the connection unit until all field connection subunits are completed, specifically for: for the xth set of relational tables that have completed the connection of the first k - 1 fields where k is a positive integer greater than 1 and x is a positive integer containing field A k and the relational table with the fewest number of tuples is The number of tuples is denoted as Denote the x-th set of relationship tables The prefix sum of Denote the prefix sum of the last set of relationship tables as PrefixSumTotal, and PrefixSumTotal is a positive integer; allocating a sequence of length PrefixSumTotal and PrefixSumTotal threads, with thread numbers sequentially from 1 to PrefixSumTotal; For the thread numbered i, where i is a positive integer, the thread finds the set of relationship tables through binary search such that The thread then the number x, and in the th tuple's field A k with value a k,i , are written into the i-th element of the sequence; The i-th thread needs to determine the corresponding tuple (x, a k,i ) whether: 1) i is equal to 2)a k,i Not equal to a k,i-1 , where a k,i Indicates field A in the i-th tuple when connecting the k-th field k The value of k and i are both positive integers; if both conditions are not met, the thread returns None; if at least one condition is met, and a k,i In the relationship table If it has appeared in all, then the relationship table set is returned Otherwise, None is returned; if the relationship table Does not contain field A k , then the relationship table Otherwise the relationship table equal Field A k Equal to a k,i A relational table consisting of tuples.

Citation Information

Patent Citations

  • Developing collective operations for a parallel computer

    CN103246508A

  • Method and device for generating SQL statement

    CN104050264A