Data query method and apparatus for distributed database

By extracting the characteristics of query keywords and generating density features, combined with the cardinality estimation model, the problem of inefficiency in data query in distributed databases is solved, and more efficient and accurate data query is achieved.

WO2025130336A1PCT designated stage expired Publication Date: 2025-06-26BEIJING OCEANBASE TECHNOLOGY CO LTD

Patent Information

Application Number
PCT/CN2024/127304
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-10-25
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing distributed databases are difficult to efficiently process query statements when querying data, resulting in inefficient responses.

Method used

By obtaining the query keywords in the query statement, extracting their characteristics, and determining the data distribution of keywords in the historical data distribution of the distributed database, the keyword density characteristics are generated. Enter these features into the cardinality estimation model for data cardinality estimation to obtain the cardinality estimation value of the target data.

Benefits of technology

It improves the efficiency and accuracy of data queries, reduces the direct query burden on distributed databases, and improves the response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127304_26062025_PF_FP_ABST
    Figure CN2024127304_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present description provide a data query method and apparatus for a distributed database. The data query method for the distributed database comprises: acquiring a query statement for performing data query in the distributed database; performing feature extraction on a query keyword comprised in the query statement to obtain a query keyword feature, determining, in historical data distribution of the distributed database, keyword data distribution corresponding to the query keyword, and generating a keyword density feature on the basis of the keyword data distribution obtained by query; and inputting the query keyword feature and the keyword density feature into a cardinality estimation model for data cardinality estimation to obtain a cardinality estimate of target data corresponding to the query keyword as a query result of the query statement.
Need to check novelty before this filing date? Find Prior Art

Description

Data query method and device for distributed database Technical Field

[0001] This document relates to the field of distributed technology, and in particular to a data query method and device for a distributed database. Background Art

[0002] With the continuous development and promotion of the Internet, in order to cope with the rapidly growing data storage and management needs on the Internet, distributed processing technology has been widely used. In particular, distributed databases have been widely used as an important means of distributed processing technology in the field of data storage. Distributed databases refer to the use of high-speed computer networks to connect multiple physically dispersed data storage nodes to form a logically unified database cluster. Based on the established distributed database, efficient data storage and data access can be achieved.

[0003] Summary of the Invention

[0004] One or more embodiments of this specification provide a data query method for a distributed database, comprising: obtaining a query statement for performing a data query in the distributed database; performing feature extraction on query keywords contained in the query statement to obtain query keyword features; determining a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating a keyword density feature based on the keyword data distribution; inputting the query keyword features and the keyword density features into a cardinality estimation model to perform data cardinality estimation, and obtaining a cardinality estimation value of target data corresponding to the query keyword.

[0005] One or more embodiments of the present specification provide a data query device for a distributed database, comprising: a query statement acquisition module, configured to acquire a query statement for performing data query in a distributed database. A feature extraction module, configured to perform feature extraction on the query keyword contained in the query statement to obtain query keyword features. A density feature generation module, configured to determine the keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generate keyword density features based on the keyword data distribution. A data cardinality estimation module, configured to input the query keyword features and the keyword density features into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

[0006] One or more embodiments of the present specification provide a data query device for a distributed database, comprising: a processor; and a memory configured to store computer-executable instructions, wherein when executed, the computer-executable instructions cause the processor to: obtain a query statement for performing a data query in a distributed database. Perform feature extraction on the query keywords contained in the query statement to obtain query keyword features. Determine the keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generate keyword density features based on the keyword data distribution. Input the query keyword features and the keyword density features into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

[0007] One or more embodiments of the present specification provide a storage medium for storing computer-executable instructions, which implement the following process when executed by a processor: obtaining a query statement for performing a data query in a distributed database. Performing feature extraction on the query keywords contained in the query statement to obtain query keyword features. Determining the keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating keyword density features based on the keyword data distribution. Inputting the query keyword features and the keyword density features into a cardinality estimation model to perform data cardinality estimation, and obtaining a cardinality estimation value of the target data corresponding to the query keyword. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or related technologies, the following briefly introduces the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings described below are only some embodiments described in this specification. Those skilled in the art can also derive other drawings based on these drawings without inventive efforts.

[0009] FIG1 is a schematic diagram of an implementation environment of a distributed database data query method provided by one or more embodiments of this specification;

[0010] FIG2 is a flowchart of a data query method for a distributed database provided by one or more embodiments of this specification;

[0011] FIG3 is a schematic diagram of the architecture of a cardinality estimation model provided by one or more embodiments of this specification;

[0012] FIG4 is a flow chart of a data query method for a distributed database applied to an actual data query scenario provided by one or more embodiments of this specification;

[0013] FIG5 is a schematic diagram of an embodiment of a data query device for a distributed database provided by one or more embodiments of this specification;

[0014] FIG6 is a schematic diagram of the structure of a data query device for a distributed database provided by one or more embodiments of this specification. DETAILED DESCRIPTION

[0015] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0016] The distributed database data query method provided in one or more embodiments of this specification is applicable to the implementation environment of a distributed database. Referring to FIG1 , the implementation environment at least includes:

[0017] A distributed database 101, a query system 102 for performing data query in the distributed database 101, the query system 102 comprising: a feature extraction module 102-1 for extracting features from query keywords, a density feature generation module 102-2 for generating keyword density features, and a cardinality estimation model 102-3 for estimating data cardinality;

[0018] The distributed database 101 is deployed on a storage device, and the query system 102 runs on a processing device. The storage device and the processing device can be one or more servers, a server cluster consisting of several servers, or a cloud server of a cloud computing platform;

[0019] In this implementation environment, during the data query process performed by the query system 102, for the query statement for data query in the distributed database 101, the query keyword features of the query keywords contained in the query statement are extracted by the feature extraction module 102-1, and the keyword density features corresponding to the query keywords are generated by the density feature generation module 102-2. The query keyword features and the keyword density features are transmitted to the cardinality estimation model 102-3 by the feature extraction module 102-1 and the density feature generation module 102-2. The query keyword features and the keyword density features are input into the cardinality estimation model 102-3 for data cardinality estimation, and the cardinality estimation value of the target data corresponding to the query keyword output by the cardinality estimation model 102-3 is used as the query result.

[0020] One or more embodiments of a distributed database data query method provided in this specification are as follows:

[0021] 2 , the data query method for a distributed database provided in this embodiment specifically includes steps S202 to S208 .

[0022] Step S202: Obtain a query statement for performing data query in a distributed database.

[0023] In the process of querying data in a distributed database, data query is performed by sending a query statement, such as an SQL query statement submitted through the query system of the distributed database. In this embodiment, the data query performed in the distributed database refers to querying the number of target data in the distributed data. The number of target data to be queried is called the cardinality.

[0024] During specific implementation, during the process of data query in a distributed database, on the premise that the distributed database maintains historical query records of its own historical data queries, the query statement of the current data query can be matched and verified. If the matching and verification determine that the data query in the historical query record of the distributed database is relatively close to the current query statement, the query result can be directly returned based on the relatively close data query in the historical query record, thereby improving the response efficiency of the data query.

[0025] In an optional implementation provided by this embodiment, after obtaining the query statement currently performing data query in the distributed database, the query statement currently performing data query is matched and verified in the following manner:

[0026] Performing a matching check on the query keywords included in the query statement based on the historical query record table of the distributed database;

[0027] If the matching verification result contains the target historical query record, the target historical query record is output as the query result of the query statement;

[0028] If the matching verification result is empty, the following step S204 is executed to extract features of the query keywords contained in the query statement to obtain query keyword features.

[0029] The historical query record table records the query input and query output of the distributed database's historical data queries. Each query input and query output of a historical data query constitutes a historical query record in the historical query record table.

[0030] During the specific execution process, when matching and verifying the query keywords contained in the query statement of the current data query with the historical query record table of the distributed database, matching and verification can be performed from the data range and time dimension contained in the query keywords. If there are historical query records in the historical query record table that match the current query statement in the data range and time dimension, it indicates that the data query of the matching historical query record is relatively close to the data query of the current query statement.

[0031] In an optional implementation manner provided by this embodiment, performing a matching check on the query keyword based on the historical query record table of the distributed database includes:

[0032] Verify whether there is a historical query record in the historical query record table whose overlap with the data range included in the query keyword meets the overlap condition;

[0033] If it exists, check whether the query interval duration corresponding to the query time of the historical query record is less than the preset duration threshold; if it is less, use the historical query record as the target historical query record; if it is greater than or equal to, determine that the matching verification result is empty;

[0034] If it does not exist, the matching verification result is determined to be empty.

[0035] In this embodiment, the data range included in the query keyword refers to the numerical range of the corresponding data to be queried included in the query keyword. For example, in a distributed database storing resident information of a certain area, when querying the number of residents aged between 18 and 45, the data range included in the query keyword of the SQL query statement is [18, 45]. This data range is also called a continuous data range.

[0036] In addition, the data range can also refer to the data set range composed of the data set of the corresponding data to be queried contained in the query keyword. For example, in a distributed database that stores resident information of a certain area, when querying the number of residents living in area A, area B, and area C, the data range contained in the query keyword of the SQL query statement is [A area, B area, C area]. This data range is also called a discrete data range.

[0037] In the above process of verifying whether there are historical query records in the historical query record table whose overlap with the data range included in the query keyword satisfies the overlap condition, if the data range included in the query keyword is a numerical range, the numerical overlap of the numerical range included in the query keyword and the numerical range of the field record corresponding to the numerical range in the historical query record table can be calculated. If the calculated numerical overlap is greater than a preset overlap threshold, it is determined that the overlap condition is met. Conversely, if the calculated numerical overlap is less than or equal to the preset overlap threshold, it is determined that the overlap condition is not met.

[0038] Similarly, when the data range contained in the query keyword is a data set range, the data overlap can be calculated between the data set range contained in the query keyword and the data set range recorded in the field corresponding to the data set range in the historical query record table. If the calculated data overlap is greater than a preset overlap threshold, it is determined that the overlap condition is met. Conversely, if the calculated data overlap is less than or equal to the preset overlap threshold, it is determined that the overlap condition is not met.

[0039] Step S204: extracting features of the query keywords contained in the query statement to obtain query keyword features.

[0040] In this embodiment, in the process of extracting features from the query keywords contained in the query statement, feature extraction can be performed on the query keywords from at least one keyword dimension to obtain query keyword features. Specifically, feature extraction can be performed on the query keywords from the data range dimension to determine the feature expression of the numerical range of the corresponding data to be queried contained in the query keywords, or feature extraction can be performed on the query keywords from the data table dimension to determine the data table to be queried in the distributed database based on the query keywords, or feature extraction can be performed on the query keywords from the table connection relationship dimension to determine the table connection relationship between the data tables to be queried in the distributed database. In addition, feature extraction can be performed on two or three of the three keyword dimensions: the data range dimension, the data table dimension, and the table connection relationship dimension to obtain corresponding query keyword features.

[0041] In an optional implementation provided by this embodiment, feature extraction is performed on the query keyword contained in the query statement to obtain query keyword features, including:

[0042] Performing query data range extraction on the query keyword to obtain data range features;

[0043] Extracting data table keywords from the query keywords to obtain data table features;

[0044] Performing table connection relationship extraction on the query keyword to obtain table connection features;

[0045] The data range feature, the data table feature, and the table connection feature are subjected to feature concatenation to obtain the query keyword feature.

[0046] During the specific execution process, extracting data table keywords from query keywords refers to extracting features from the table keywords of the data tables being queried that are contained in the query keywords. For example, a distributed database storing information on residents in a certain area has several data tables. If a corresponding data query is to be performed in one or more of the several data tables, the query keywords of the SQL query statement contain the table keywords of the one or more data tables to be queried. Here, the table keywords contained in the query keywords are extracted as data table features. By extracting data table features, data queries against distributed databases can be constrained to specific data tables, which helps improve the efficiency of data queries.

[0047] Extracting table connection relationships for query keywords refers to extracting features from the table connection relationships between the data tables for data query contained in the query keywords. For example, a distributed database storing information on residents in a certain area may contain several data tables. If a corresponding data query is to be performed on multiple data tables with table connection relationships among the multiple data tables, the query keywords of the SQL query statement contain the table connection relationship keywords of the multiple data tables for which the corresponding data query is to be performed. The table connection relationship keywords contained in the query keywords are extracted here as table connection features. By extracting table connection features, the amount of computation required for performing data queries on multiple data tables at the same time can be reduced. At the same time, the extraction of table connection features can also distinguish the differences between data queries at a finer granularity, thereby helping to improve the accuracy of data queries, that is, the accuracy of the cardinality estimation of the target data to be queried.

[0048] For example, when querying data in a distributed database storing information about residents in a certain area, the query keyword of the SQL query statement includes a data range of [18, 45], and the table keyword of the included data table is the table name or table code of the identity information data table and the table name or table code of the residence data table in the distributed database, and also includes the table connection relationship between the identity information data table and the residence data table;

[0049] In the process of extracting features from a query keyword in an SQL query statement, extracting a data range [18, 45] contained in the query keyword, and constructing a data range feature vector based on the extracted data range [18, 45], extracting table names or table codes of an identity information data table and a residence data table contained in the query keyword, and constructing a data table feature vector based on the extracted table names or table codes, and extracting a table connection keyword of a table connection relationship between the identity information data table and the residence data table contained in the query keyword, and constructing a table connection feature vector based on the extracted table connection keyword;

[0050] After obtaining the data range feature vector, the data table feature vector and the table connection feature vector, vector splicing is performed on the data range feature vector, the data table feature vector and the table connection feature vector, and the spliced ​​vector obtained after the splicing is obtained is used as the keyword feature vector of the entire query keyword in the SQL query statement.

[0051] As described above, the data range contained in the query keyword includes two types: a continuous data range, a numerical range, and a discrete data range, a data set range. In the case that the data range contained in the query keyword is a discrete data range, a data set range, the discrete data range can be converted into a continuous data range, and data range feature extraction can be performed when the data range is converted into a continuous data range. Specifically, in an optional implementation manner provided by this embodiment, query data range extraction is performed on the query keyword to obtain a data range feature, including: if it is detected that the query data range contained in the query keyword is a discrete data range, data range encoding is performed on the query data range, and based on the encoding result, the continuous data range of the query data range is determined as the data range feature.

[0052] It should be noted that the above-mentioned query data range extraction process for query keywords, data table keyword extraction process for query keywords, and table connection relationship extraction process for query keywords, the three processing processes are not limited in the specific execution process. Any one processing process can be executed first and then the remaining two processing processes. In addition, in order to improve processing efficiency and improve data query response efficiency, parallel processing can also be adopted to execute the three processing processes in parallel in different processing threads.

[0053] Step S206 , determining the keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating a keyword density feature based on the keyword data distribution.

[0054] During specific implementation, based on the query statement for data query in the distributed database, starting from the query keywords contained in the query statement, the keyword data distribution corresponding to the query keywords is determined in the historical data distribution of the distributed database, and further based on the determined keyword data distribution, the keyword density feature is generated. The keyword density feature is used to reflect the density distribution of the query keywords in the distributed database, and the density distribution of the query keywords is combined to reflect the distribution of the target data to be queried in the distributed database.

[0055] In an optional implementation manner provided by this embodiment, determining the keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database includes:

[0056] Determining a keyword type corresponding to the query keyword;

[0057] A keyword density histogram corresponding to the keyword type is searched in the historical data distribution as the keyword data distribution.

[0058] Specifically, in the process of determining the keyword data distribution corresponding to the query keywords contained in the query statement, in addition to the above-mentioned implementation method of querying the keyword density histogram corresponding to the keyword type in the historical data distribution of the distributed database, other forms of data distribution data can also be queried in the historical data distribution as keyword data distribution.

[0059] Taking an SQL query statement for data query in a distributed database storing resident information of a certain area as an example, the keyword type of the query keyword in the SQL query statement is an age keyword type. Then, an age density histogram representing the age distribution is queried in the historical density histogram of the distributed database. The horizontal axis of the age density histogram is the age value, and the vertical axis is the number of residents corresponding to each age value.

[0060] Based on the above-mentioned determination of the keyword density histogram corresponding to the keyword type in the historical data distribution as the keyword data distribution, in the process of generating keyword density features based on the determined keyword data distribution, the keyword density features are generated through the following optional implementation method: performing feature conversion on the keyword density histogram, and using the density histogram features obtained by the conversion as the keyword density features.

[0061] It should be noted that the above-mentioned processing process of extracting features from the query keywords contained in the query statement to obtain query keyword features and the above-mentioned processing process of determining the keyword data distribution corresponding to the query keywords in the historical data distribution of the distributed database and generating keyword density features are not limited in the specific execution order. The processing process of extracting features from the query keywords contained in the query statement to obtain query keyword features can be performed first, and then the processing process of determining the keyword data distribution corresponding to the query keywords in the historical data distribution of the distributed database and generating keyword density features can be performed. Alternatively, the processing process of determining the keyword data distribution corresponding to the query keywords in the historical data distribution of the distributed database and generating keyword density features can be performed first, and then the processing process of extracting features from the query keywords contained in the query statement to obtain query keyword features can be performed. Alternatively, a parallel processing method can be adopted to perform the two processing processes of extracting features from the query keywords contained in the query statement to obtain query keyword features and determining the keyword data distribution corresponding to the query keywords in the historical data distribution of the distributed database and generating keyword density features in parallel on different processing threads.

[0062] Step S208: Input the query keyword feature and the keyword density feature into a cardinality estimation model to perform data cardinality estimation to obtain a cardinality estimation value of the target data corresponding to the query keyword.

[0063] The above improves the accuracy and efficiency of data query by extracting query keyword features from query keywords contained in the query statement, and improves the accuracy of data query by determining the keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database and generating keyword density features. On the basis of the obtained query keyword features and keyword density features, the query keyword features and keyword density features are input into the cardinality estimation model to perform data cardinality estimation, which can perform data cardinality estimation more comprehensively and accurately, thereby obtaining a more accurate cardinality estimation value of the target data corresponding to the query keyword.

[0064] The data query performed in the distributed database in this embodiment refers to querying the number of target data in the distributed data. Here, the number of target data to be queried is called the cardinality; the cardinality estimation value of the target data corresponding to the query keyword output by the cardinality estimation model refers to the estimated value of the number of target data output by the cardinality estimation model after estimating the target data to be queried.

[0065] During the specific execution process, when the cardinality estimation model estimates the data cardinality based on the input query keyword features and keyword density features, the query keyword features and keyword density features are different types of features. At the same time, the input query keyword features may be one-dimensional features, while the keyword density features may be two-dimensional features. The query keyword features and keyword density features can be unified through feature normalization, and then the data cardinality can be estimated based on the two-part features obtained by feature normalization.

[0066] Optionally, the cardinality estimation model includes: a first perceptron module, a convolutional neural network, and a second perceptron module;

[0067] Based on the cardinality estimation model including the first perceptron module, the convolutional neural network and the second perceptron module, the query keyword features can be feature standardized through the first perceptron module, and the keyword density features can be feature standardized through the convolutional neural network. Then, the two parts of features obtained by feature standardization are input into the second perceptron module for data cardinality estimation.

[0068] Specifically, in an optional implementation provided by this embodiment, data cardinality estimation is implemented in the following manner:

[0069] Normalizing the query keyword features by the first perceptron module to obtain standard keyword features;

[0070] Normalizing the keyword density feature by using the convolutional neural network to obtain a standard density feature;

[0071] The standard keyword feature and the standard density feature are input into the second perceptron module to perform cardinality estimation of the target data corresponding to the query keyword to obtain the cardinality estimation value.

[0072] For example, in the cardinality estimation model shown in Figure 3, the keyword feature vector of the entire query keyword in the SQL query statement is input into the multilayer perceptron (MLP) for feature standardization to obtain the standardized keyword feature vector, and the keyword density feature vector obtained by performing feature conversion on the keyword density histogram is input into the convolutional neural network (CNN) for feature standardization to obtain the standardized keyword density feature vector, and the standardized keyword feature vector and the standardized keyword density feature vector are input into the enhanced multilayer perceptron to estimate the number of residents to be queried, and output the estimated value of the number of residents.

[0073] In this embodiment, when the distributed database maintains a historical query record table that records its own historical data queries, the cardinality estimation model can also be trained based on the historical query records of data queries in the distributed database recorded in the historical query record table. Specifically, in an optional implementation provided by this embodiment, the cardinality estimation model is trained in the following manner: training samples and corresponding sample labels are constructed based on the historical query records in the historical query record table of the distributed database; the model to be trained is trained based on the training samples and the sample labels, and the cardinality estimation model is obtained after the training is completed.

[0074] To sum up, the present embodiment provides one or more distributed database data query methods. In the process of performing data query in a distributed database, in order to improve the data query efficiency of the distributed database, starting from the query keywords contained in the query statement for data query, the cardinality of the target data to be queried is estimated through a cardinality estimation model, rather than consuming a large amount of computing resources to perform data query in the data table of the distributed database. Specifically, during the query process, the accuracy and efficiency of the data query are improved by extracting features of the query keywords contained in the query statement to obtain query keyword features, and the accuracy of the data query is improved by determining the keyword data distribution corresponding to the query keywords in the historical data distribution of the distributed database and generating keyword density features. On the basis of the obtained query keyword features and keyword density features, the query keyword features and keyword density features are input into the cardinality estimation model to perform data cardinality estimation, which can perform data cardinality estimation more comprehensively and accurately, thereby obtaining a more accurate cardinality estimation value of the target data corresponding to the query keyword.

[0075] The following takes the application of a distributed database data query method provided by this embodiment in an actual data query scenario as an example, and combines Figure 4 to further illustrate the distributed database data query method provided by this embodiment. Referring to Figure 4, the distributed database data query method applied to the actual data query scenario specifically includes the following steps.

[0076] Step S402: Obtain an SQL query statement for performing data query in a distributed database.

[0077] Step S404: extract the query data range of the query keyword contained in the SQL query statement to obtain a data range feature vector.

[0078] Step S406: extracting data table keywords from the query keywords contained in the SQL query statement to obtain a data table feature vector.

[0079] Step S408: extracting table connection relationships from the query keywords included in the SQL query statement to obtain table connection feature vectors.

[0080] Step S410 , performing feature concatenation on the data range feature vector, the data table feature vector, and the table connection feature vector to obtain a query keyword feature vector.

[0081] Step S412: query the keyword density histogram corresponding to the keyword type of the query keyword in the historical data distribution, and perform feature conversion on the keyword density histogram to obtain a keyword density feature vector.

[0082] Step S414 , performing feature normalization on the query keyword feature vector through the first perceptron module of the cardinality estimation model to obtain a standard keyword feature vector.

[0083] Step S416: normalize the keyword density feature vector using the convolutional neural network of the cardinality estimation model to obtain a standard density feature vector.

[0084] Step S418: input the standard keyword feature vector and the standard density feature vector into the second perceptron module of the cardinality estimation model to perform cardinality estimation of the target data corresponding to the query keyword, and output a cardinality estimation value.

[0085] An embodiment of a data query device for a distributed database provided in this specification is as follows:

[0086] In the above embodiment, a data query method for a distributed database is provided. Correspondingly, a data query device for a distributed database running on a server is also provided, which will be described below with reference to the accompanying drawings.

[0087] 5 , which shows a schematic diagram of an embodiment of a data query device for a distributed database provided by this embodiment.

[0088] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiment described below is only illustrative.

[0089] This embodiment provides a data query device for a distributed database, which runs on a server and includes:

[0090] A query statement acquisition module 502 is configured to acquire a query statement for performing data query in a distributed database;

[0091] The feature extraction module 504 is configured to extract features of the query keywords contained in the query statement to obtain query keyword features;

[0092] A density feature generating module 506 is configured to determine a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generate a keyword density feature based on the keyword data distribution;

[0093] The data cardinality estimation module 508 is configured to input the query keyword feature and the keyword density feature into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

[0094] An embodiment of a distributed database data query device provided in this specification is as follows:

[0095] Corresponding to the data query method of a distributed database described above, based on the same technical concept, one or more embodiments of this specification also provide a data query device for a distributed database, which is used to execute the data query method of a distributed database provided above. Figure 6 is a structural schematic diagram of a data query device for a distributed database provided by one or more embodiments of this specification.

[0096] This embodiment provides a distributed database data query device, including:

[0097] As shown in Figure 6, the data query device of a distributed database may have relatively large differences due to different configurations or performances, and may include one or more processors 601 and memory 602, and the memory 602 may store one or more storage applications or data. Among them, the memory 602 can be a temporary storage or a persistent storage. The application stored in the memory 602 may include one or more modules (not shown in the figure), each module may include a series of computer executable instructions in the data query device of the distributed database. Furthermore, the processor 601 can be configured to communicate with the memory 602 and execute a series of computer executable instructions in the memory 602 on the data query device of the distributed database. The data query device of the distributed database may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input / output interfaces 605, one or more keyboards 606, etc.

[0098] In a specific embodiment, a data query device for a distributed database includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for querying the data in the distributed database, and the one or more programs are configured to be executed by one or more processors, including computer-executable instructions for performing the following:

[0099] Get query statements for data query in distributed databases;

[0100] Extracting features of the query keywords contained in the query statement to obtain query keyword features;

[0101] Determining a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating a keyword density feature based on the keyword data distribution;

[0102] The query keyword feature and the keyword density feature are input into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

[0103] An embodiment of a storage medium provided in this specification is as follows:

[0104] Corresponding to the data query method of a distributed database described above, based on the same technical concept, one or more embodiments of this specification also provide a storage medium.

[0105] The storage medium provided in this embodiment is used to store computer-executable instructions. When the computer-executable instructions are executed by a processor, the following process is implemented:

[0106] Get query statements for data query in distributed databases;

[0107] Extracting features of the query keywords contained in the query statement to obtain query keyword features;

[0108] Determining a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating a keyword density feature based on the keyword data distribution;

[0109] The query keyword feature and the keyword density feature are input into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

[0110] It should be noted that the embodiment of a storage medium in this specification and the embodiment of a data query method for a distributed database in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned corresponding method, and the repeated parts will not be repeated.

[0111] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. For example, the device embodiment, equipment embodiment and storage medium embodiment are similar to the method embodiment, so the description is relatively simple. To read the relevant content in the device embodiment, equipment embodiment and storage medium embodiment, please refer to the partial description of the method embodiment.

[0112] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0113] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0114] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0115] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0116] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0117] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0119] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0121] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0122] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0123] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0124] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising at least one ..." does not exclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.

[0125] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0126] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.

Claims

1. A data query method for a distributed database, comprising: Get the query statement for data query in the distributed database; Extracting features of the query keywords contained in the query statement to obtain query keyword features; Determining a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating a keyword density feature based on the keyword data distribution; The query keyword feature and the keyword density feature are input into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

2. According to the distributed database data query method of claim 1, the step of extracting features of the query keywords contained in the query statement to obtain query keyword features comprises: Extracting the query data range of the query keyword to obtain data range features; Extracting data table keywords from the query keywords to obtain data table features; Extracting table connection relationships of the query keywords to obtain table connection features; The data range feature, the data table feature and the table connection feature are concatenated to obtain the query keyword feature.

3. The distributed database data query method according to claim 2, wherein the step of extracting the query data range of the query keyword to obtain the data range feature comprises: If it is detected that the query data range contained in the query keyword is a discrete data range, data range encoding is performed on the query data range, and a continuous data range of the query data range is determined as the data range feature based on the encoding result.

4. The data query method of a distributed database according to claim 1, wherein determining the keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database comprises: Determine the keyword type corresponding to the query keyword; A keyword density histogram corresponding to the keyword type is searched in the historical data distribution as the keyword data distribution.

5. The data query method for a distributed database according to claim 4, wherein generating a keyword density feature based on the keyword data distribution comprises: Perform feature conversion on the keyword density histogram, and use the density histogram features obtained by the conversion as the keyword density features.

6. The data query method for a distributed database according to claim 1, wherein the cardinality estimation model comprises: A first perceptron module, a convolutional neural network, and a second perceptron module; The data cardinality estimation is implemented in the following way: Standardizing the query keyword features by using the first sensor module to obtain standard keyword features; Standardizing the keyword density feature by using the convolutional neural network to obtain a standard density feature; The standard keyword feature and the standard density feature are input into the second perceptron module to perform cardinality estimation of the target data corresponding to the query keyword to obtain the cardinality estimation value.

7. According to the data query method of a distributed database according to claim 1, the cardinality estimation model is trained in the following manner: Constructing training samples and corresponding sample labels according to the historical query records in the historical query record table of the distributed database; The model to be trained is trained based on the training samples and the sample labels, and the cardinality estimation model is obtained after the training is completed.

8. The method for querying data in a distributed database according to claim 1, after the step of obtaining a query statement for querying data in a distributed database is executed, and before the step of extracting features of query keywords contained in the query statement and obtaining query keyword features is executed, further comprising: Performing a match check on the query keyword based on the historical query record table of the distributed database; If the matching verification result includes the target historical query record, the target historical query record is output as the query result of the query statement; If the matching verification result is empty, the step of extracting features of the query keywords contained in the query statement to obtain query keyword features is performed.

9. The data query method of a distributed database according to claim 8, wherein the matching verification of the query keyword based on the historical query record table of the distributed database comprises: Checking whether there is a historical query record in the historical query record table whose overlap with the data range included in the query keyword satisfies the overlap condition; If so, check whether the query interval duration corresponding to the query time of the historical query record is less than a preset duration threshold; if so, use the historical query record as the target historical query record; If it does not exist, the match verification result is determined to be empty.

10. A data query device for a distributed database, comprising: A query statement acquisition module is configured to acquire a query statement for performing data query in a distributed database; A feature extraction module is configured to extract features of the query keywords contained in the query statement to obtain query keyword features; a density feature generating module configured to determine a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and to generate a keyword density feature based on the keyword data distribution; The data cardinality estimation module is configured to input the query keyword feature and the keyword density feature into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

11. A data query device for a distributed database, comprising: processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: Get the query statement for data query in the distributed database; Extracting features of the query keywords contained in the query statement to obtain query keyword features; Determining a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating a keyword density feature based on the keyword data distribution; The query keyword feature and the keyword density feature are input into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

12. A storage medium for storing computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the following process: Get the query statement for data query in the distributed database; Extracting features of the query keywords contained in the query statement to obtain query keyword features; Determining a keyword data distribution corresponding to the query keyword in the historical data distribution of the distributed database, and generating a keyword density feature based on the keyword data distribution; The query keyword feature and the keyword density feature are input into a cardinality estimation model to perform data cardinality estimation, and obtain a cardinality estimation value of the target data corresponding to the query keyword.

Citation Information

Patent Citations

  • Cardinality estimation method and device for database query optimization

    CN115587111A

  • Automatic query predicate selective prediction using machine learning models

    CN116057518A

  • Rodinary number estimation processing method, device and equipment for relational database

    CN116361326A

  • Data query method and device for distributed database

    CN117743381A

Cited By

  • Cardinality estimation method and apparatus

    US12670132B2

  • Cardinality estimation method and apparatus

    US20250021531A1