Data query service transaction processing method and device based on tuple information gain
Patent Information
- Application Number
- CN202410535931.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-04-30
AI Technical Summary
[0004]本申请实施例的目的是提供一种基于元组信息增益的数据查询服务交易处理方法及装置,以解决相关技术中存在的计算效率低、可解释性差的问题
[0023]Based on the above technical solutions, this invention provides data sellers with an efficient and arbitrage-free data query transaction processing method. A support set is constructed for each relation table in the database, with the size specified by the data seller. For a single-table query input by the data consumer, an auxiliary query is constructed. The information gain of all tuples in the corresponding table is calculated based on the original query results and the auxiliary query results. The query price is then published based on the overall information gain and a price function. For multi-table queries input by the data consumer, the original query is rewritten, and an auxiliary query is constructed for each table. The original query results and auxiliary query results are extracted and deduplicated. The information gain of all tuples in each table is then calculated. Finally, the query price is set based on the overall information gain and a price function. This method solves the problems of low efficiency and poor interpretability in data query transaction processing, supporting the practical application of data query transactions.
Smart Images

Figure CN118467810B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data processing technology, and in particular to a data query service transaction processing method and apparatus based on tuple information gain. Background Technology
[0002] With the popularization and development of IoT devices, 5G communication technology, and internet technology, data and its applications have generated enormous value. The demand for interaction, integration, and exchange of big data is increasing, giving rise to big data trading. Different organizations and individuals have different data analysis and trading needs. How to meet diverse data needs while achieving efficient and effective data trading is a major challenge facing the implementation of current data trading platforms. Query-based data trading models have wide application scenarios, allowing data consumers with limited budgets to express their data needs through queries and purchase the required data, avoiding the high cost of purchasing entire datasets. However, due to the diverse and complex forms of queries, simple query-based pricing and trading processes can lead to arbitrage problems.
[0003] Arbitrage refers to the practice of data consumers purchasing multiple inexpensive queries Q1, Q2, ..., Q... l The goal is to infer the result of a high-priced query Q. For example, query Q might be to select the age and gender data of users older than 20, i.e., Q = "select age,gender from User where age>20". The result of query Q can be inferred through queries Q1 = "select age from User where age>20" and Q2 = "select gender from User where age>20". If the price of query Q is greater than the sum of the prices of Q1 and Q2, data consumers can obtain the result of query Q by purchasing Q1 and Q2 at a lower price. If arbitrage exists in the query transaction processing, speculative data consumers will continuously try to obtain the required data at the lowest price, reducing transaction profits. The existence of arbitrage will also make ordinary data consumers feel unfairly treated, reducing their willingness to trade. Therefore, the data query price function needs to satisfy the arbitrage-free property while ensuring the efficiency of query transaction processing. Existing methods suffer from low computational efficiency and poor interpretability in data query transactions. Summary of the Invention
[0004] The purpose of this application is to provide a data query service transaction processing method and apparatus based on tuple information gain, so as to solve the problems of low computational efficiency and poor interpretability in related technologies.
[0005] According to a first aspect of the embodiments of this application, a data query service transaction processing method based on tuple information gain is provided, including:
[0006] Based on the support set size |S| specified by the data seller, for each relation table R in database D i Construct a support set S i Each support set S i Includes the corresponding relationship table R i The possible values of a tuple, and each support set S i Stored in the database server where D resides;
[0007] For a single table R input by a data consumer i The query Q will retrieve the table name R from the query Q. i Replace with S i , thus obtaining the support set S i The auxiliary query Q' is executed on the database server, and queries Q and Q' are executed to obtain query results O and O';
[0008] Calculate R based on the query results O and O'. i The information gain (|S) of each tuple t under Q i |-|E t |), cumulative R i The information gain of all tuples in the dataset is used to set the price of query Q based on the information gain-price function chosen by the data seller for the transaction, where E... t Let S represent the set of possible values for each tuple t, (|S i |-|E t |) represents the uncertainty that t eliminates under Q, i.e., the information gain of tuple t;
[0009] For data consumers inputting multiple tables R1, R2, ..., R... k Given a query Q, rewrite the query Q as Q', and construct a support set S1, S2, ..., S' based on Q'. k Auxiliary queries Q1, Q2, ..., Q on k Execute a query Q', Q1, Q2, ..., Q in the database server. k The query results W, W1, W2, ..., W are obtained. k For each relation table R i Extract query results W and W i In R i The data W' and W' on i The results are then deduplicated to obtain O. i and O' i , i = 1, ..., k;
[0010] Based on the results of each group O i and O' i Calculate R iThe information gain (|S) of each tuple t under Q i |-|E t |), cumulative R i The information gain of all tuples in R is used to obtain query Q based on the information gain-price function chosen by the data seller. i The prices on the table are summed up from all relation tables R. i The price on the query results is used to obtain the price of Q for the transaction.
[0011] According to a second aspect of the embodiments of this application, a data query service transaction processing apparatus based on tuple information gain is provided, comprising:
[0012] The support set construction module is used to construct support sets for each relation table R in database D, based on the support set size |S| specified by the data seller. i Construct a support set S i Each support set S i Includes the corresponding relationship table R i The possible values of a tuple, and each support set S i Stored in the database server where D resides;
[0013] The single-table query processing module is used to process the table name R in the single-table query Q input by the data consumer. i Replace with S i , thus obtaining the support set S i The auxiliary query Q' is executed on the database server, and queries Q and Q' are executed to obtain query results O and O';
[0014] The single-table query transaction module is used to calculate R based on the query results O and O'. i The information gain (|S) of each tuple t under Q i |-|E t |), cumulative R i The information gain of all tuples in the dataset is used to set the price of query Q based on the information gain-price function chosen by the data seller for the transaction, where E... t Let S represent the set of possible values for each tuple t, (|S i |-|E t |) represents the uncertainty that t eliminates under Q, i.e., the information gain of tuple t;
[0015] The multi-table query processing module is used to process multiple tables R1, R2, ..., R2 input by data consumers. k Rewrite the query Q as Q', and construct the support set S1, S2, ..., S' based on Q'. k Auxiliary queries Q1, Q2, ..., Q on k Execute a query Q', Q1, Q2, ..., Q in the database server. kThe query results W, W1, W2, ..., W are obtained. k For each relation table R i Extract query results W and W i In R i The data W' and W' on i The results are then deduplicated to obtain O. i and O' i , i = 1, ..., k;
[0016] The multi-table query transaction module is used to query each set of results O i and O' i Calculate R i The information gain (|S) of each tuple t under Q i |-|E t |), cumulative R i The information gain of all tuples in R is used to obtain query Q based on the information gain-price function chosen by the data seller. i The prices on the table are summed up from all relation tables R. i The price on the query results is used to obtain the price of Q for the transaction.
[0017] According to a third aspect of the embodiments of this application, an electronic device is provided, comprising:
[0018] One or more processors;
[0019] Memory, used to store one or more programs;
[0020] When the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in the first aspect.
[0021] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method as described in the first aspect.
[0022] Beneficial effects:
[0023] Based on the above technical solutions, this invention provides data sellers with an efficient and arbitrage-free data query transaction processing method. A support set is constructed for each relation table in the database, with the size specified by the data seller. For a single-table query input by the data consumer, an auxiliary query is constructed. The information gain of all tuples in the corresponding table is calculated based on the original query results and the auxiliary query results. The query price is then published based on the overall information gain and a price function. For multi-table queries input by the data consumer, the original query is rewritten, and an auxiliary query is constructed for each table. The original query results and auxiliary query results are extracted and deduplicated. The information gain of all tuples in each table is then calculated. Finally, the query price is set based on the overall information gain and a price function. This method solves the problems of low efficiency and poor interpretability in data query transaction processing, supporting the practical application of data query transactions. Attached Figure Description
[0024] Figure 1 A flowchart of a data query service transaction processing method based on tuple information gain provided in an embodiment of the present invention.
[0025] Figure 2 This is a database example diagram provided for an embodiment of the present invention.
[0026] Figure 3 An example diagram of the support set provided for an embodiment of the present invention.
[0027] Figure 4 This is an example diagram of a single-table query result provided in an embodiment of the present invention.
[0028] Figure 5 This is an example diagram of multi-table query results provided in an embodiment of the present invention.
[0029] Figure 6 An example diagram of query processing results provided in an embodiment of the present invention.
[0030] Figure 7 A comparison chart of transaction price and running time in the transaction processing method provided in this embodiment of the invention on the MoveLens movie rating dataset.
[0031] Figure 8 This is a block diagram of a data query transaction processing device based on tuple information gain, provided as an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the specific content of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention. Contents not described in detail in the embodiments of the present invention belong to the prior art known to those skilled in the art.
[0033] This invention provides a data query service transaction processing method based on tuple information gain, which can be used in online transaction scenarios where queries input by data consumers are priced. In this online data trading market, each data consumer queries data according to their individual needs; for example, a data analyst might query a movie rating dataset (such as...). Figure 2 The data consumer is interested in movies produced after 1900 (i.e., tuples) shown in the diagram, and selects these movies for querying and analysis. The query Q input by the data consumer can be a single-table query, such as query Q = "select * from movie where year >= 1990" on the movie table; or a multi-table query, such as query Q = "select title, name, rating from movie, user, rating where rating >= 4 and rating.userID = user.userID and rating.movieID = movie.movieID", used to query the movie name, user, and specific rating of a movie with a rating of four or higher. Data query transaction processing requires online pricing for such queries; the following describes the method of this invention in detail with reference to this scenario.
[0034] refer to Figure 1 The method may include the following steps:
[0035] S1: Based on the support set size |S| specified by the data seller, for each relation table R in database D. i Construct a support set S i Each support set S i Includes the corresponding relationship table R i The possible values of a tuple, and each support set S i Stored in the database server where D resides.
[0036] Specifically, firstly, based on the total size of the support set specified by the data seller, |S| = 12, and the size of each relation table R in the database... i Size | R i(i.e., |R1|=|R2|=|R3|=3) Calculate each support set S i The size, that is Where R1, R2, ..., R3 are all the movie tables, user tables and movie rating tables in the movie rating database D;
[0037] Then, for each relation table R i (For i = 1, 2, 3), let T i For R i The number of unique tuples in the set is T. i =3, because |S i |=4 is greater than T i , will R i Add all unique tuples to S i In, and according to the relation table R i Given the constraints, generate a unique tuple and add it to S. i In the end, each relation table R is obtained. i On the support set S i (like Figure 3 As shown in the figure, it is stored in the database server where D is located.
[0038] S2: For a single table R input by the data consumer. i The query Q = "select * from movie where year >= 1990" will retrieve the table name R from query Q. i Replace with S i Replacing "movie" with "movie_support" yields the support set S. i The auxiliary query Q' = "select * from movie_support where year >= 1990" is executed on the database server, yielding query results O and O', as follows: Figure 4 As shown.
[0039] S3: Calculate R based on the query results O and O'. i The information gain (|S) of each tuple t under Q i |-|E t |), cumulative R i The information gain of all tuples in the dataset is used to set the price of query Q based on the information gain-price function chosen by the data seller for the transaction, where E... t Let S represent the set of possible values for each tuple t, (|S i |-|E t |) represents the uncertainty eliminated by t under Q, i.e., the information gain of tuple t.
[0040] The possible value set E for each tuple t in the above process t The specific definitions include:
[0041] If the result of tuple t under query Q is v t And v t The set of possible values E of t is non-empty. t Includes S i The query result is v t The set of elements, namely E t ={v|v∈S i andQ(v)=v t}, v can be counted by traversing O'. t The frequency of occurrence is obtained from |E t |and information gain(|S i |-|E t |);
[0042] If the result of query Q for tuple t is empty, then the set of possible values E for t is... t Includes S i The set of elements that do not meet the query conditions, i.e. Then |E t |=|S i |-|O| and information gain|S i |-|E t |=|O|;
[0043] According to the above definition, if certain single-table queries are Q1, Q2, ..., Q... l If we can deduce the query Q, then for R... i Each tuple t in the set satisfies in Let Q be the set of possible values for tuple t. Is a tuple query Q j The set of possible values for (j = 1, 2, ..., l) guarantees that pricing can be directly based on information gain, and the price of each tuple t under Q is necessarily less than or equal to its price under Q1, Q2, ..., Q1. l The total price under this definition ensures the absence of arbitrage; at the same time, under this definition, the price of each tuple in a query corresponds to its information gain, providing an explanatory basis for pricing data query transactions.
[0044] Specifically, iterate through O' and count the frequency h of each element o. o This results in R having a frequency of 1 for each element. i The overall information gain of all tuples t in the set is:
[0045] Furthermore, the data seller can choose an information gain-price function to transform the information gain into a query price. This function needs to satisfy monotonically increasing and subadditivity within the range of positive integers to ensure the arbitrage-free nature of the query price. Substituting this total information gain into the information gain-price function chosen by the data seller yields the price of query Q, and the transaction is then conducted at this price.
[0046] Specifically, if the seller chooses the information gain-price function f(x) = log(x+1), then the price of query Q is log10.
[0047] S4: For multiple tables R1, R2, ..., R4 input by data consumers... k Given a query Q, rewrite the query Q as Q', and construct a support set S1, S2, ..., S' based on Q'. k Auxiliary queries Q1, Q2, ..., Q on k Execute a query Q', Q1, Q2, ..., Q in the database server. k The query results W, W1, W2, ..., W are obtained. k For each relation table R i (i = 1, ..., k), extract the query results W and W'. i In R i The data W' and W' on i The results are then deduplicated to obtain O. i and O' i .
[0048] For query Q, all tables R1, R2, ..., R... k Add all tables R1, R2, ..., R to the Selection clause of query Q. k The primary key attribute is used to rewrite query Q as query Q'; based on Q', for each table R involved in query Q... i (i = 1, 2, ..., k), change the table name R in Q'. i Replace with S i , obtain auxiliary query Q i (i = 1, 2, ..., k); Execute the query Q', Q1, Q2, ..., Q' in the database server. k The query results W, W1, W2, ..., W are obtained. k ;
[0049] Specifically, if query Q is Q = "select title,name,rating from movie,user,ratingwhere rating>=4and rating.userID=user.userID and rating.movieID=movie.movieID", by adding the primary key attributes (i.e., movieID and userID) from all tables to the Selection clause of query Q, query Q is rewritten as query Q' = "select movie.movieID,user.userID,rating.movieID,rating.userID,title,name,rating from movie,user,rating whererating>=4and rating.userID=user.userID and rating.movieID=movie.movieID". Further, by replacing table names, three auxiliary queries are obtained: Q1 = "select movie_support.movieID,user.userID,rating.movieID,rating.userID,title,name,rating from movie_support,user,rating where rating>=4and rating.userID=user.userID". andrating.movieID=movie_support.movieID”,Q2="select movie.movieID,user_support.userID,rating.movieID,rating.userID,title,name,rating from movie,user_support,rating where rating>=4and rating.userID=user_support.userIDand rating.movieID=movie.movieID",Q3="select movie.movieID,user.userID,rating_support.movieID,rating_support.userID,title,name,rating from movie,user,rating_support where rating>=4and rating_support.userID=user.userIDandrating_support.movieID = movie.movieID; the query results for Q', Q1, Q2, and Q3 are W, W1, W2, and W3, respectively. Figure 5 As shown.
[0050] Due to the query Q' (or Q) i The query filter criteria and the columns to be output are specified, and the query result W (or W...) i This includes multi-row, multi-column data; check W (or W...) sequentially. i For each column in ), if that column belongs to relation table R i If the condition is met, keep the column; otherwise, remove the column to obtain W (or W). i In R i The result W' (or W' i ).
[0051] Due to the query Q' (or Q) i This is a multi-table query, and the extracted query result is W' (or W'). i It may contain duplicate data; check W' (or W') in turn. i For each row in the query, remove duplicate query results and delete the primary key column added during the query rewriting process, resulting in O. i (or O') i ).
[0052] Specifically, for i = 1, 2, 3, each query result W and W i The processed result O after filtering, deduplication, and deletion of the primary key column is obtained. i and O' i like Figure 6 As shown. For example, for i=1 (i.e., the movie table), the movie.movieID and title columns on W and W1 are retained, and the remaining columns are deleted; then duplicate elements are deleted, and the primary key column movie.movieID is deleted, resulting in Figure 6 O1 and O'1 in the text.
[0053] S5: Based on the results of each group O i and O' i Calculate R i The information gain (|S) of each tuple t under Q i |-|E t |), cumulative R i The information gain of all tuples in R is used to obtain query Q based on the information gain-price function chosen by the data seller. i The prices on the table are summed up from all relation tables R. i The price on (i=1,…,k) is used to obtain the price of query Q for trading.
[0054] Traversing O' i Calculate the frequency h of each element 'o'. o R i The overall information gain of all tuples t in the set is: R i Substituting the overall information gain into the information gain-price function chosen by the data seller, we obtain the query Q in R. i The prices on the table are summed from all relation tables R1, R2, ..., R. k The price on the chart is used to obtain the final price of Q and complete the transaction.
[0055] Specifically, by traversing O'1, O'2, and O'3, the information gains of the three tables R1, R2, and R3 are respectively... and Based on the information gain-price function f(x) = log(x+1), we can obtain that the price of Q on the three tables is log 9, and the total price is 3·log 9.
[0056] Example
[0057] The transaction processing method of this invention was implemented on an Ubuntu 18.04 system on a server with 192GB of memory and an Intel core 2.80GHz. The performance of the transaction processing method of this invention on the MovieLens dataset was tested under different selectivity (i.e., the ratio of the size of the query result set to the size of the data table).
[0058] The performance of the data query transaction processing method (ARIA) proposed in this invention was tested and analyzed through simulation experiments. The results are as follows: Figure 7 As shown, ARIA is more efficient than existing database information gain-based processing methods (i.e., the QIRANA method). The ARIA method sets a higher query price because it considers more granular information gain and is more comprehensive.
[0059] Corresponding to the aforementioned embodiments of the data query service transaction processing method based on tuple information gain, this application also provides embodiments of the data query service transaction processing apparatus based on tuple information gain.
[0060] Reference Figure 8 The data query service transaction processing apparatus based on tuple information gain includes:
[0061] Support set construction module 1 is used to construct support sets for each relation table R in database D according to the support set size |S| specified by the data seller. i Construct a support set S i Each support set S iIncludes the corresponding relationship table R i The possible values of a tuple, and each support set S i Stored in the database server where D resides;
[0062] Single-table query processing module 2 is used to process the table name R in the single-table query Q input by the data consumer. i Replace with S i , thus obtaining the support set S i The auxiliary query Q' is executed on the database server, and queries Q and Q' are executed to obtain query results O and O';
[0063] Single-table query transaction module 3 is used to calculate R based on the query results O and O'. i The information gain (|S) of each tuple t under Q i |-|E t |), cumulative R i The information gain of all tuples in the dataset is used to set the price of query Q based on the information gain-price function chosen by the data seller for the transaction, where E... t Let S represent the set of possible values for each tuple t, (|S i |-|E t |) represents the uncertainty that t eliminates under Q, i.e., the information gain of tuple t;
[0064] Multi-table query processing module 4 is used to process the multi-table R1, R2, ..., R input by the data consumer. k Rewrite the query Q as Q', and construct the support set S1, S2, ..., S' based on Q'. k Auxiliary queries Q1, Q2, ..., Q on k Execute a query Q', Q1, Q2, ..., Q in the database server. k The query results W, W1, W2, ..., W are obtained. k For each relation table R i (i = 1, ..., k), extract the query results W and W'. i In R i The data W' and W' on i The results are then deduplicated to obtain O. i and O' i ;
[0065] Multi-table query transaction module 5 is used to query each set of results O i and O' i Calculate R i The information gain (|S) of each tuple t under Q i |-|E t |), cumulative R iThe information gain of all tuples in R is used to obtain query Q based on the information gain-price function chosen by the data seller. i The prices on the table are summed up from all relation tables R. i The price on (i=1,…,k) is used to obtain the price of query Q for trading.
[0066] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0067] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0068] Accordingly, this application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the data query transaction processing method based on tuple information gain as described above.
[0069] Accordingly, this application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the data query transaction processing method based on tuple information gain as described above.
[0070] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A data query service transaction processing method based on tuple information gain, characterized in that, include: Based on the support set size specified by the data seller | S |, for database D Each relation table R i Construct a support set S i Each support set S i Includes a correspondence table R i The possible value set of the tuple, and each support set S i Store to D In the database server where it is located; For a single table input by a data consumer R i Online query Q , will query Q Table name in R i Replace with S i , obtain support set S i Auxiliary queries on Q’ Execute a query in the database server Q and Q’ Get the query results O and O’ ; Based on the query results O and O’ calculate R i Each tuple t exist Q Information gain below (| S i |- | E t |), cumulative R i The information gain of all tuples in the query is based on the information gain-price function settings selected by the data seller. Q The price is used for trading, among which... E t Represent each tuple t The set of possible values, (| S i |- | E t | indicates t exist Q The uncertainty of elimination, i.e., tuples t Information gain; For data consumer input multiple tables R 1, R 2, …, R k Online query Q Rewrite query Q for Q’ and according to Q’ Building support sets S 1, S 2, …, S k Auxiliary queries on Q 1, Q 2, …, Q k Execute a query in the database server Q’ , Q 1, Q 2, …, Q k Get the query results W , W 1, W 2, …, W k For each relation table R i Extract query results W and W i exist R i Data on W’ and And remove duplicates from the results to obtain O i and , i =1, …, k ; Based on the results of each group O i and calculate R i Each tuple t exist Q Information gain below (| S i |- | E t |), cumulative R i The information gain of all tuples in the query is obtained based on the information gain-price function chosen by the data seller. Q exist R i The prices on the table are summed up across all relational tables. R i The price can be found on the website. Q Trade at the price; Wherein, each of the relation tables R i Support set on S i The construction method is as follows: First, based on the total support set size |S| specified by the data seller and each relation table in the database... R i Size | R i |Calculate each support set S i The size, that is ,in R 1, R 2, …, R m It is a database D All relation tables in the table; Then, for each relation table R i Corresponding support set S i The construction method is as follows: Let T i for R i The number of unique tuples in a given set, if | S i | Less than or equal to T i ,from R i Random selection | S i | Add unique tuples S i In the middle; if | S i | Greater than T i ,Will R i Add all unique tuples to S i In, and according to the relationship table R i Constraints, generate (| S i | - T i Add ) unique tuples. S i middle; Among them, according to the query results O and O’ calculate R i Each tuple t Information gain (| S i |- | E t |) and query Q The price at which transactions are conducted includes: If tuple t In the query Q The result is v t and v t Not empty, t The set of possible values E t Include S i The query results are v t The set of elements, i.e. It can be done by traversing O’ statistics v t The frequency of occurrence is obtained | E t |and information gain(| S i |- | E t |); If tuple t In the query Q The result is empty. t The set of possible values E t Include S i The set of elements that do not meet the query conditions, i.e. Then there is and information gain ; Traversal O’ Count each element o frequency of occurrence h o ,but R i All tuples t The overall information gain can be calculated as follows: Substituting this total information gain into the information gain-price function chosen by the data seller, we obtain the query... Q The price is set, and transactions are conducted at that price.
2. The method according to claim 1, characterized in that, The multi-table R 1, R 2, …, R k The query rewriting and auxiliary query generation methods are as follows: For query Q All tables involved R 1, R 2, …, R k In the query Q Add all tables to the Selection clause R 1, R 2, …, R k The primary key attribute on the query will then be used to determine the query. Q Rewrite as a query Q’ ; exist Q’ Based on this, for query Q Each table involved R i ,Will Q’ Table name in R i Replace with S i , obtain auxiliary query Q i , i = 1, 2, …, k .
3. The method according to claim 1, characterized in that, The query result extraction process is as follows: Due to query Q’ or Q i The query filter criteria and the columns to be output are specified, and the query results are displayed. W or W i Including multi-row, multi-column data, check them one by one. W or W i For each column in the table, if that column belongs to the relation table R i If the condition is met, keep the column; otherwise, remove the column to obtain the desired result. W or W i exist R i The results W’ or .
4. The method according to claim 1, characterized in that, The process of deduplicating query results includes: Due to query Q’ or Q i It's a multi-table query, and the extracted query results are... W’ or It may contain duplicate data; check sequentially. W’ or For each row in the query, remove duplicate query results and delete the primary key columns added during the query rewriting process, resulting in... O i or .
5. The method according to claim 1, characterized in that, Also includes: Traversal Count each element o frequency of occurrence h o , R i All tuples t The overall information gain is ,Will R i Substituting the overall information gain into the information gain-price function selected by the data seller, we obtain the query. Q exist R i The prices on the table are summed up across all relational tables. R 1, R 2, …, R k The price on the website can be found through inquiry. Q The final price and transaction.
6. A data query service transaction processing device based on tuple information gain, characterized in that, The apparatus for performing data query service transaction processing based on tuple information gain as described in claim 1 includes: The support set building module is used to determine the support set size specified by the data seller. S |, for database D Each relation table R i Construct a support set S i Each support set S i Includes a correspondence table R i The possible value set of the tuple, and each support set S i Store to D In the database server where it is located; The single-table query processing module is used to process single-table queries input by data consumers. Q Table name in R i Replace with S i , obtain support set S i Auxiliary queries on Q’ Execute a query in the database server Q and Q’ Get the query results O and O’ ; The single-table query transaction module is used to query the transaction results. O and O’ calculate R i Each tuple t exist Q Information gain below (| S i |- | E t |), cumulative R i The information gain of all tuples in the query is based on the information gain-price function settings selected by the data seller. Q The price is used for trading, among which... E t Represent each tuple t The set of possible values, (| S i |- | E t | indicates t exist Q The uncertainty of elimination, i.e., tuples t Information gain; The multi-table query processing module is used to process multi-table queries input by data consumers. R 1, R 2, …, R k Online query Q Rewritten as Q’ and according to Q’ Building support sets S 1, S 2, …, S k Auxiliary queries on Q 1, Q 2, …, Q k Execute a query in the database server Q’ , Q 1, Q 2, …, Q k Get the query results W , W 1, W 2, …, W k For each relation table R i Extract query results W and W i exist R i Data on W’ and And remove duplicates from the results to obtain O i and , i =1, …, k ; The multi-table query transaction module is used to query each set of results. O i and calculate R i Each tuple t exist Q Information gain below (| S i |- | E t |), cumulative R i The information gain of all tuples in the query is obtained based on the information gain-price function chosen by the data seller. Q exist R i The prices on the table are summed up across all relational tables. R i ( i =1, …, k The price can be found on the website. Q They trade at the price.
7. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Alias remittance malicious query method and device
CN110414993A
Decision tree training using a database system
US20200193332A1