Associated data set-oriented data table combination query method with maximized difference degree

By establishing a table connection index in the graph database and using feature-data column index and connection score calculation function, a data table combination query method is proposed for the associated data set, which solves the problem of failing to effectively consider the complexity of data lakes in the existing technology, and realizes efficient and scalable data table combination query.

CN120045592AActive Publication Date: 2025-05-27NANJING UNIV OF POSTS & TELECOMM
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510202473.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing data table query method for maximizing the degree of difference has failed to effectively consider the complexity of the tabular data in the real data lake, including proportional foreign key constraints and namespace inconsistencies, making it difficult to find a combination of data tables that meet budget constraints in the associated dataset.

Method used

A method of querying data table combinations for the associated data set is proposed. By establishing a table connection index in the graph database, using feature-data column index and connection score calculation function, a candidate data table combination set is constructed, and a greedy selection algorithm is used to select a data table combination that meets the budget conditions.

Benefits of technology

It effectively reduces the number of connected edges in the table connection graph, reduces the amount of calculation, shortens the query time, and improves the scalability of the table connection graph, while ensuring query efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045592A_ABST
    Figure CN120045592A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of data retrieval, and discloses a difference maximization data table combination query method for an associated data set, which comprises the following steps of: in a data processing stage, firstly performing data processing on a given table data set, establishing a feature-data column index, discovering a connectable table in the table data set according to the index, and performing data retrieval on the connectable table; meanwhile, constructing a data table connection graph index, and pre-calculating connection information between the tables; in the data query stage, according to a given sample query table and a given connection column set, a candidate connection column set is searched in the feature-data column index, a candidate data table set is obtained, and according to a given budget, a data table set which can be connected with the sample query table and enables the difference degree to be maximum is selected. The method for searching the connectable data table combination in the associated data set is proposed for the first time, the connectable data tables are filtered by using the feature indexes, the data table connection diagram is established to discover the connection path between the data tables, and the data table set which maximizes the difference degree under the budget constraint is returned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data retrieval, and specifically relates to a method for querying a combination of data tables with maximized difference degree for an associated data set. Background Art

[0002] Currently, many data engines are widely used in data query, and table data query has received extensive attention. As a kind of structured data, table data widely exists in various fields such as enterprises, scientific research, and finance. Its query and management are of great significance for improving work efficiency and supporting decision-making. There is an urgent practical need to efficiently obtain connectable data tables required.

[0003] In order to query table data more efficiently, some advanced technologies and tools have been proposed and applied. For example, table recognition technology based on deep learning can automatically detect tables, recognize table structures and contents through semantic segmentation algorithms, object detection algorithms, text sequence generation algorithms, etc., thereby improving the query and processing efficiency of table data.

[0004] In the process of connectable query of table data, some challenges are also faced, such as diversity evaluation, value evaluation, multi-attribute join query, etc. To address these challenges, researchers and developers have been continuously exploring new methods and technologies. For example, by constructing a document summary index structure and using multi-modal LLMs, table data can be efficiently retrieved and summarized, thereby improving the accuracy and efficiency of table data query. However, due to the variable row and column orders of different tables, the complexity is relatively high when calculating the overlapping area of table data, and various methods have been proposed in the existing literature. For example, Sloth generates a seed list by detecting the attribute pairs of shared data element values in two tables and gradually combines these seeds to calculate the maximum overlapping area between the two tables; Mate uses a Bloom filter to pre-screen rows that cannot overlap, and then obtains the overlapping area of the two tables. However, in the proposed methods for querying data tables with maximized difference degree, the complexity of table data in a real data lake, such as foreign key constraints, namespace inconsistencies, etc., has not been considered. Therefore, if a new method can be proposed to find a combination of data tables that meets the budget limit in an associated data set, this problem can be effectively solved. Summary of the Invention

[0005] To solve the deficiencies of the prior art, the present invention provides a method for querying a combination of data tables with maximized difference degree for an associated data set. This query method is oriented to a table data lake with multi-level connection relationships, establishes a table connection index based on a graph database, and proposes a method for querying a combination of data tables with maximized difference in an associated data set.

[0006] To achieve the above object, the present invention is implemented through the following technical solutions:

[0007] The present invention is a method for querying a combination of data tables with maximized difference degree for an associated data set, which refers to querying in a table data set C for a data table combination that satisfies a multi-attribute connection condition on a specified connection column set Q and maximizes the difference degree, including a data processing stage and a data query stage. Let the table data set composed of table data be denoted as C = {T Q , T 1 ,... T 2 ,... T n}, where T q is the query example data table, and the specified connection column set Q = {q 1 , q 2 ,... q m}, specifically:

[0008] The first stage: the data processing stage, specifically including the following steps:

[0009] Step 1.1: For each data table T i in the table data set C, according to the data type of each column, extract the feature set F i,j of each column. For the table data set C, for each feature type f l , construct a feature-data column index I l , and then construct a feature-data column index set I = {I 1 , I 2 ,... I L};

[0010] Step 1.2: For each data table T i in the table data set C, based on the feature-data column index set find all data columns in the table data set C that satisfy the similarity threshold θ according to each data column c i of the data table T j , and then form a candidate connection set R i of the data table T i = {<T k , c x , FK(c x ))>|c x ∈T i , FK(c x )∈T k , e(c x , FK(c x ))>θ}, where FK(c x ) represents the column c x in the candidate table T kFor the connectable columns in, where e is a column matching score calculation function, for each candidate table T in the candidate connection set R k , select a set of optimal connections <T k , c x , FK(c x ))>, and delete the candidate connection set R i in the data table T i and other connections of the candidate table T k ;

[0011] Step 1.3. For the table data set C, according to the candidate connection set R i of each data table T i in the table data set C, use the graph database to construct a graph index G representing the connection relationships between the tables in the table data set C;

[0012] Second stage: data query stage, which specifically includes the following steps:

[0013] Step 2.1. According to the given query example data table T q and the connection column set Q, retrieve all data columns in the feature-data column index set I in the connection column set Q that meet the similarity threshold θ, and then construct a sample candidate connection set R q of the query example data table T;

[0014] Step 2.2. According to the sample candidate connection set R q , obtain all data table combinations p q that can be connected to the query example data table T k on all data columns of the connection column set Q, have a connection path in the graph index G, and meet k , and then construct a candidate data table combination set where PC k is the set composed of the connection paths that can connect all the data tables in the data table combination p q ;

[0015] Step 2.3. According to each data table combination p k in the candidate data table combination set PT k and the corresponding connection path set PC k of p k , obtain the set of differential row tuples Diff(T q , p q ) between the data table combination p k and the query example data table T q ;

[0016] Step 2.4: According to the given budget B, search for a subset RP of the candidate data table combination set PT q such that the subset RP satisfies the following conditions:

[0017]

[0018] Furthermore, construct the result set R = {(p k , PC k ) | p ∈ RP}, and return the result set R.

[0019] A further improvement of the present invention lies in: In step 1.1, for each data table T in the table data set C i , construct a feature-data column index set Specifically, it includes the following steps;

[0020] Step 1.1.1: For the table data set C, for the j-th column of the data table T i , use a preset feature extraction algorithm to extract the feature set i of the j-th column of the data table T where L is the number of feature types;

[0021] Step 1.1.2: For the table data set C, for each feature type f l , construct a feature-data column index I l based on locality-sensitive hashing;

[0022] Step 1.1.3: For each column data feature , construct an index entry containing the data column information (T i , j) and the corresponding feature value, and add it to the corresponding feature-data column index I l , and then construct the feature-data column index set I = {I 1 , I 2 ,..., I L}.

[0023] A further improvement of the present invention lies in: In step 1.2, for each data table T in the table data set C i and the candidate connection set R i , and then for each candidate table T in the candidate connection set R i , select a set of optimal connections <T k , c k , FK(c x )>, specifically including the following steps: x Step 1.2.1: For each candidate table T in the candidate connection set R

[0024] in the candidate connection set R i ​k , filter out all connections in the candidate connection set R i that are pointed to the data columns in the data table T i and point to the data columns in the candidate table T k . The connection R ik = {<T k , c x , FK(c x )>|c x ∈ T i , FK(c x ) ∈ T k};

[0025] Step 1.2.2. For each group of connections <c ik , FK(c x )> in the connection R x , calculate its connection score:

[0026] FK-SCORE(c x , FK(c x )) = (e(c x , FK(c x )) + j(c x , FK(c x )) * max{u(c x ), u(FK(c x ))}

[0027] where e(c x , FK(c x ) is the column matching score calculation function, and j(c x , FK(c x )) is the matching score of each data element of the two columns, and the calculation is as follows:

[0028]

[0029] where u(c x ), u(FK(c x )) are the uniqueness scores of the two columns respectively, and the calculation is as follows:

[0030]

[0031] where unique(c x ) represents the set composed of unique values in the c x column, and u(FK(c x )) represents the set composed of unique values in the FK(c x ) column;

[0032] Step 1.2.3. Select FK_SCORE(c x , FK(cx )) The largest connected pair <c x , FK(c x )> is used as the best join for the data table T i to the candidate table T k .

[0033] A further improvement of the present invention lies in: in step 2.3, for each data table combination p q in the set PT of candidate data table combinations k and the corresponding set PC of connection paths k for p k , the set Diff(T k , p q ) of different row tuples between the data table combination p q and the query example data table T k is obtained, which specifically includes the following steps:

[0034] Step 2.3.1: For each connection path path k in the set PC of connection paths x , the edges <T x , T i > included in the connection path path j are sequentially selected and the data tables T i , T j are joined. The join conditions and join methods are as follows:

[0035]

[0036] Furthermore, the set PT(path x ) of row tuples on the connection column set Q after all data table combinations on the connection path path x are obtained:

[0037]

[0038] where |V(path x )| represents the number of vertices on this path;

[0039] Step 2.3.2: According to the set PT(path k ) of row tuples formed on each path in the set PC of connection paths x , the set of different row tuples between the data table combination p k and the query example data table T q is calculated:

[0040]

[0041] A further improvement of the present invention lies in: in step 2.4, according to the given budget B, a set PT of candidate data table combinations is constructed q A subset RP that meets the defined conditions of

[0042] Step 2.4.1: Initialize the remaining budget b = B, the set RS of selected row tuples = ∏ Q (T q ), the set TS of selected data tables, the candidate data table combination list PS = {<p k , g(p k )|p k ∈PT q}, and the set RP of selected data table combinations, where g(p k ) is the marginal benefit of this data table combination, and the calculation method is as follows:

[0043]

[0044] Step 2.4.2: Sequentially select from the candidate data table combinations the data table combination p k with the highest marginal benefit and a cost lower than the remaining budget b and add it to the set RP of data table combinations, and sequentially update the set TS of selected data tables = TS ∪ p k , the remaining budget b = B - ∑ T∈TS price(T), the set RS of selected row tuples = RS ∪ Diff(RS, p k ), and then update the marginal benefit of each remaining candidate table combination in the candidate data table combination list PS.

[0045] Step 2.4.3: When there is no data table combination that meets the conditions in the candidate data table combination list PS, return the set RP of data table combinations.

[0046] The beneficial effects of the present invention are:

[0047] The present invention can preliminarily filter the connectable data tables through feature indexing, effectively reducing the number of connection edges in the table connection graph, reducing the calculation amount, and thus shortening the query time.

[0048] The present invention judges the connectability between data tables through a specific connection score calculation function without the need to pre-provide the connection information in the data lake, improving the scalability of the table connection graph.

[0049] The present invention selects the data table combination that meets the constraint conditions by means of greedy selection, ensuring the query efficiency without significantly reducing the accuracy. Description of the Drawings

[0050] Figure 1It is the flowchart of the method for querying the combination of data tables with the maximum difference degree of the present invention.

[0051] Figure 2 It is the schematic diagram of the index of the table connection diagram of the present invention.

[0052] Figure 3 It is the schematic diagram of the data table query engine of the present invention. Detailed implementation manners

[0053] The following will disclose the implementation manners of the present invention. For the sake of clear description, many practical details will be described together in the following narration. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some implementation manners of the present invention, these practical details are not necessary.

[0054] For the convenience of description, the following definitions are made for relevant symbols: The table data set C = {T 1 , T 2 , … T n}, which contains n data tables, the query sample data table T q , and the set Q of specified connection columns = {q 1 , q 2 , …, q m}.

[0055] As Figures 1-3 shown, the present invention is a method for querying the combination of data tables with the maximum difference degree for associated data sets. This method refers to querying in the table data set C for the combination of data tables that satisfy the multi-attribute connection condition on the specified connection column set Q with T q and maximize the difference degree, including a data processing stage and a data query stage. Let the table data set composed of table data be denoted as C = {T 1 , T 2 , … T n}, T q is the query sample data table, and the specified connection column set Q = {q 1 , q 2 ,..., q m}, specifically:

[0056] The first stage: The data processing stage, specifically including the following steps:

[0057] Step 1.1: For each data table T i in the table data set C, according to the data type of each column, extract the feature set F i,j of each column. For the table data set C, for each feature type f l , construct the feature-data column index I l , and then construct the feature-data column index set I = {I1 , I 2 ,..., I L}}。

[0058] Among them, constructing the feature-data column index set specifically includes the following steps;

[0059] Step 1.1.1: For the table data set C, for the data table T i of the j-th column, use the preset feature extraction algorithm to extract the feature set i of the j-th column of the data table T where L is the number of feature types.

[0060] The feature extraction algorithm in this application can be flexibly selected according to factors such as data type and application scenario. For example, for numerical data, statistical features such as mean and variance can be extracted, and for text data, text features such as word frequency, TF-IDF, and word embedding can be extracted.

[0061] Step 1.1.2: For the table data set C, for each feature type f l , construct the feature-data column index I l ;

[0062] Step 1.1.3: For each column data feature construct an index item containing the data column information (T i , j) and the corresponding feature value, and add it to the corresponding feature-data column index I l , and then construct the feature-data column index set I = {I 1 , I 2 ,..., I L}}。

[0063] Step 1.2: For each data table T in the table data set C i , based on the feature-data column index set According to each data column c i of the data table T j in the table data set C that satisfies the similarity threshold θ, and then form the candidate connection set R i of the data table T i = {<T k , c x , FK(c x ))>|c x ∈T i , FK(c x )∈T k , e(c x , FK(c x ))>θ}, where FK(cx ) represents column c x In the candidate table T k Among the connectable columns, e is the column matching score calculation function. For each candidate table T in the candidate connection set Ri k , select a set of optimal connections <T k , c x , FK(c x ))>, and delete the other connections between the data table T i in the candidate connection set R i and the candidate table T k . The specific steps are as follows:

[0064] Step 1.2.1: For each candidate table T i in the candidate connection set R k , screen out all connections R i in the candidate connection set R i where the data columns in the data table T k point to the data columns in the candidate table T ik ={ <T k , c x , FK(c x )>| c x ∈T i , FK(c x )∈T k};

[0065] Step 1.2.2: For each group of connections <c ik , FK(c x )> in the connection R x , calculate its connection score:

[0066] FK_SCORE(c x , FK(c x ))=(e(c x , FK(c x ))+j(c x , FK(c x ))*max{u(c x ), u(FK(c x ))}

[0067] where e(c x , FK(c x ) is the column matching score calculation function, j(c x , FK(c x )) is the matching score of each data element of the two columns, and the calculation is as follows:

[0068]

[0069] where u(c x ), u(FK(c x )) are the uniqueness scores of the two columns, calculated as follows:

[0070]

[0071] where unique(c x ) represents the set composed of unique values in column c x , and u(FK(c x )) represents the set composed of unique values in column FK(c x );

[0072] Step 1.2.3. Select the connection pair <c x , FK(c x )> with the largest FK_SCORE(c x , FK(c x )> as the best connection of the data table T i to the candidate table T k .

[0073] Step 1.3. For the tabular data set C, according to the candidate connection set R i of each data table T i in the tabular data set C, use a graph database such as Neo4j to construct a graph index G representing the connection relationships between the tables in the tabular data set C.

[0074] Second stage: Data query stage, specifically including the following steps:

[0075] Step 2.1. According to the given query example data table T q and the connection column set Q, retrieve all data columns in each feature-data column index set in the connection column set Q that meet the similarity threshold θ, and then construct a sample candidate connection set R q of the query example data table Tq;

[0076] Step 2.2. According to the sample candidate connection set R q , obtain all data table combinations p q that can be connected to the query example data table T on all data columns of the connection column set Q, and there is a connection path in the graph index G and meet k , and then construct a candidate data table combination set where PC k is the set composed of all connection paths that can connect all data tables in the data table combination p k ;

[0077] Step 2.3. Combine data table combination set PT according to the candidate data table combination q For each data table combination p k in k the corresponding connection path set PC k , obtain the difference row tuple set Diff(T k , p q ) between the data table combination p q and the query example data table T k . The specific steps are as follows:

[0078] Step 2.3.1. For each connection path path k in the connection path set PC x , sequentially select the edge <T x , T i > contained in the connection path path j and connect the data tables T i , T j . The connection conditions and connection methods are as follows:

[0079]

[0080] Furthermore, obtain the row tuple set PT(path x ) on the connection column set Q after combining all the data table combinations on the connection path path x :

[0081]

[0082] where |V(path x )| represents the number of vertices on this path;

[0083] Step 2.3.2. Calculate the difference row tuple set between the data table combination p k and the query example data table T x according to the row tuple set PT(path k ) formed on each path in the connection path set PC q :

[0084]

[0085] Step 2.4. According to the given budget B, find a subset RP of the candidate data table combination set PT q such that the subset RP satisfies the following conditions:

[0086]

[0087] Furthermore, construct the result set R = {(p k , PCk ) | p ∈ RP}, and return the result set R.

[0088] Among them, construct the candidate data table combination set PT q A subset RP that meets the limited conditions, specifically including the following construction process:

[0089] Step 2.4.1: Initialize the remaining budget b = B, the set of selected row tuples RS = ∏ Q (T q ), the set of selected data tables TS, the candidate data table combination list PS = {<p k , g(p k ) | p k ∈ PT q}, and the set of selected data table combinations RP, where g(p k ) is the marginal benefit of this data table combination, and the calculation method is as follows:

[0090]

[0091] Step 2.4.2: Sequentially select the data table combination p with the highest marginal benefit and a cost lower than the remaining budget b in the candidate data table combinations k and add it to the data table combination set RP, and sequentially update the set of selected data tables TS = TS ∪ p k , the remaining budget b = B - ∑ T∈TS price(T), the set of selected row tuples RS = RS ∪ Diff(RS, p k ), and then update the marginal benefit of each remaining candidate table combination in the candidate data table combination list PS.

[0092] Step 2.4.3: When there is no data table combination that meets the conditions in the candidate data table combination list PS, return the data table combination set RP.

[0093] The above is only the implementation manner of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A data table combination query method for maximizing the difference of associated data sets, characterized by: This method refers to querying the table dataset C with T q The data table combination that satisfies the multi-attribute connection condition on the specified connection column set Q and maximizes the difference includes the data processing stage and the data query stage. Let the table data set composed of the table data be recorded as C = {T1, T2, ...T n }, T q To query the sample data table, specify the join column set Q = {q1, q2, ..., q m }, specifically: The first stage: data processing stage, which includes the following steps: Step 1.1: For each data table T in the table data set C i , according to the data type of each column, extract the feature set F of each column i,j , for the table dataset C, for each feature type f l , construct feature-data column index I l , and then construct a feature-data column index set I = {I1, I2, ..., I L }; Step 1.2: For each data table T in the table dataset C i , based on the feature-data column index set I, according to the data table T i Each data column c j All data columns that meet the similarity threshold θ in the table data set C constitute the data table T i The candidate connection set R i ={ <T k , c x , FK(c x ))>|c x ∈T i , FK(c x )∈T k ,e(c x , FK(c x ))>θ}, where FK(c x ) represents column c x In the candidate table T k The connectable columns in the candidate connection set R are as follows: k , select a set of best connections <T k , c x , FK(c x ))>, and delete the candidate connection set R i Medium Data Table T i With the candidate table T k Other connections; Step 1.3: For the table data set C, according to each data table T in the table data set C i The candidate connection set R i , use the graph database to build a graph index G that represents the connection relationship between the tables in the table dataset C; The second stage: data query stage, which includes the following steps: Step 2.1: Based on the given query sample data table T q And the connection column set Q, for each feature-data column index set I in the connection column set Q, retrieve all data columns that meet the similarity threshold θ, and then construct the sample candidate connection set R of the query sample data table T q ; Step 2.2: Based on the sample candidate connection set R q , get all the data columns that can be queried on the sample data table T on the connection column set Q q For connectable columns, there is a connection path in the graph index G and satisfies Data table combination p k , and then construct a candidate data table combination set PC k All combinations of tables that can be connected k The set of connection paths of all data tables in; Step 2.3: Combine the set PT according to the candidate data table q Each data table combination p k and p k The corresponding connection path set PC k , get the data table combination p k And query sample data table T q The difference row tuple set Diff(T q , p k ); Step 2.4: Find the candidate data table combination set PT based on the given budget B q A subset RP of , so that the subset RP satisfies the following conditions: Then construct the result set R = {(p k , PC k )|p∈RP}, and returns the result set R.

2. The method for maximizing the difference between related data sets according to claim 1, characterized in that: In step 1.1, for each data table T in the table data set C i , build a feature-data column index set , specifically including the following steps; Step 1.1.1: For the table data set C, for the data table T i The jth column of the data table T is extracted using the preset feature extraction algorithm i The feature set of the jth column Where L is the number of feature types; Step 1.1.2: For each feature type f l , construct feature-data column index I based on local sensitive hashing l ; Step 1.1.3: For each column data feature Construct a table containing data column information (T i , j) and the index item of the corresponding eigenvalue, and add it to the corresponding feature-data column index I l Then construct a feature-data column index set 3. The method for maximizing the difference between related data sets according to claim 1, characterized in that: In step 1.2, for each data table T in the table data set C, i and the candidate connection set R i , and then the candidate connection set R i Each candidate table T in k , select a set of best connections <T k , c x , FK(c x )>, which includes the following steps: Step 1.2.1: For the candidate connection set R i Each candidate table T in k , in the candidate connection set R i Filter out all the data from table T i The data columns in the candidate table T k Concatenation of data columns in R ik ={ <T k , c x , FK(c x )>|c x ∈T i , FK(c x )∈T k }; Step 1.2.2: For connecting R ik Each set of connections <c x , FK(c x )>, calculate its connection score: FK_SCORE(c x ,FK(c x ))=(e(c x ,FK(c x ))+j(c x ,FK(c x ))*max{u(c x ),u(FK(c x ))} where e(c x , FK(c x ) is the column matching score calculation function, j(c x , FK(c x )) is the matching score of each data element in the two columns, calculated as follows: Where u(c x ), u(FK(c x )) are the uniqueness scores of the two columns, calculated as follows: Among them unique(c x ) means c x The set of unique values ​​in the column, u(FK(c x )) represents FK(c x ) column; Step 1.2.3, select FK_SCORE(c x , FK(c x ))The largest connection pair <c x , FK(c x )> as data table T i To the candidate table T k The best connection.

4. The method for combined query of data tables for maximizing the difference of associated data sets according to claim 1, characterized in that: In step 2.3, the set PT is combined according to the candidate data table q Each data table combination p k and p k The corresponding connection path set PC k , get the data table combination p k And query sample data table T q The difference row tuple set Diff(T q , p k ), specifically including the following steps: Step 2.3.1: For the connection path set PC k Each connection path in path x , select the connection path path in turn x The edges included in <T i , T j > and connect to data table T i , T j , the connection conditions and connection methods are as follows: Then get the connection path path x The row tuple set PT (path x ): Where |V(path x )| represents the number of vertices on the path; Step 2.3.2: Collect PCs based on connection paths k The row tuple set PT (path x ), calculate the data table combination p k And query sample data table T q A collection of tuples of difference rows:

5. The method for combined query of data tables for maximizing the difference of associated data sets according to claim 1, characterized in that: Step 2.4 Based on the given budget B, construct the candidate data table combination set PT q A subset RP that meets the restricted conditions includes the following construction process: Step 2.4.

1. Initialize the remaining budget b = B and the selected row tuple set RS = ∏ Q (T q ), selected data table set TS, candidate data table combination list PS = { <p k , g(p k )>|p k ∈PT q } and the selected data table combination set RP, where g(p k ) is the marginal benefit of the data table combination, which is calculated as follows: Step 2.4.2: Select the data table combination p with the highest marginal benefit and a cost lower than the remaining budget b from the candidate data table combinations. k Add to the data table combination set RP, and update the selected data table set TS = TS∪p in turn k 、Remaining budget b=B-∑ T∈TS price(T), the set of selected row tuples RS = RS∪Diff(RS, p k ), and then update the marginal benefit of each remaining candidate table combination in the candidate data table combination list PS. Step 2.4.3: When there is no data table combination that meets the conditions in the candidate data table combination list PS, return the data table combination set RP.

Citation Information

Patent Citations

  • Optimal data representation and auxiliary structures for in-memory database query processing

    CN104737165A

  • Electricity big data querying method and system based on HBase secondary indexing

    CN106503243A

  • System for analysing data relationships to support query execution

    CN110168515A

  • Non-equivalent connectable data table direct query method based on LSH

    CN115374142A

  • Adaptive database index optimization and abnormal query detection method and system based on deep reinforcement learning

    CN119046502A